About Naija Speaker Audio

Building consented, high-quality Nigerian-language speech datasets for AI training.

Naija Speaker Audio is a contributor platform operated as part of the NaijaSpeaker ecosystem. We collect structured speech data from Nigerians across languages, dialects, and speaking styles to power speech recognition, text-to-speech, translation, and language identification systems.

Languages in scope

We collect Nigerian Pidgin, Yoruba, Hausa, Igbo, and Nigerian English, alongside Esan, Efik, Ibibio, Urhobo, Tiv, Nupe, and Igala — expanding coverage for regional voices often missing from speech datasets.

Our mission

Ensure Nigerian languages are represented accurately in AI systems through community-led data collection, transparent consent, and rigorous quality review.

Data principles

  • Informed, versioned consent for every contributor
  • Human review alongside automated quality checks
  • Traceability from exported samples to source consent
  • Fair compensation where projects are paid
  • Community advisory input on sensitive content