About Naija Speaker Audio
Building consented, high-quality Nigerian-language speech datasets for AI training.
Naija Speaker Audio is a contributor platform operated as part of the NaijaSpeaker ecosystem. We collect structured speech data from Nigerians across languages, dialects, and speaking styles to power speech recognition, text-to-speech, translation, and language identification systems.
Languages in scope
We collect Nigerian Pidgin, Yoruba, Hausa, Igbo, and Nigerian English, alongside Esan, Efik, Ibibio, Urhobo, Tiv, Nupe, and Igala — expanding coverage for regional voices often missing from speech datasets.
Our mission
Ensure Nigerian languages are represented accurately in AI systems through community-led data collection, transparent consent, and rigorous quality review.
Data principles
- Informed, versioned consent for every contributor
- Human review alongside automated quality checks
- Traceability from exported samples to source consent
- Fair compensation where projects are paid
- Community advisory input on sensitive content