OPEN INFRASTRUCTURE FOR ETHIOPIAN LANGUAGE AI
Dataset.ET is an open dataset initiative building speech, text, and translation datasets for 80+ Ethiopian languages to enable AI research and language technology.
LANGUAGES
SENTENCES
OUR MISSION
Every language carries a universe of knowledge. No language should be left behind as the world moves towards AI.
Every Ethiopian language carries centuries of knowledge, poetry, and identity. We create high- quality datasets that ensure no language faces digital extinction.
80+ languages at risk
From Amharic Ge'ez script to Oromo grammar, our corpus enables AI to understand, speak, and reason in Ethiopian language with true fluency.
Zero languages left behind
Open infrastructure built by researchers, developers, and language enthusiasts across Ethiopia. Your data, your rules, your future.
Open source
LANGUAGES
Building high-quality datasets to train the next generation of Ethiopian AI models, expanding to more every quarter.
01
አማርኛ
Semitic
Ge'ez script
02
Afaan Oromoo
Cushitic
Latin script
03
Af Soomaali
Cushitic
Latin script
04
ትግርኛ
Semitic
Ge'ez script
05
Sidaamu Afoo
Cushitic
Latin script
06
Wolaytta
Omotic
Latin script
+74 more languages on our roadmap
HOW IT WORKS
Every contribution, no matter how small, helps bridge the gap between Ethiopian languages and artificial intelligence.
1
Submit sentences, paragraphs, or documents in your native Ethiopian language through our portal.
2
Community reviewers verify every contribution for quality, accuracy, and linguistic structure.
3
Validated data is cleaned, annotated, and stored in our open-access corpus with full attribution.
4
Researchers and developers use the corpus to train models that understand Ethiopian languages
COMMUNITY
Volunteers, researchers, linguists, and developers who believe in the power of language technology for Ethiopia.
Join our Telegram GroupDonate your voice, translate text, and help expand our datasets across multiple Ethiopian languages.
Ensure quality by reviewing and verifying submitted audio and text contributions.
Build innovative AI tools and models using our open-source datasets and APIs.
Provide expert guidance on grammar, phonetics, and cultural nuances of Ethiopian languages.
GET INVOLVED
Every sentence you contribute builds the open infrastructure that will power Ethiopian AI for generations. No technical skill required.
Start Contributing NowBy contributing, you agree to our Privacy Policy