Commissioned speech datasets · Nigeria & West Africa

NDPC REGISTERED · NDPC/DCP/13596

Registered with the Nigeria Data Protection Commission. Certificate available on request. See our Compliance page and Privacy Policy.

African speech data for AI training, built to your specification.

Commission human-sourced speech data for ASR, speech-to-text, voice assistants, TTS, call-centre intelligence, evaluation, and multilingual AI. BSG DataWorks plans the collection, recruits consenting contributors, runs quality checks, and delivers the documentation your technical and legal teams need.

Collection scope

Specify the speech data your product actually needs.

Most projects are commissioned. You define the target use case, voices, languages, environments, recordings, metadata and delivery format; we turn that brief into an operational collection plan.

01 · USE CASE

Model-relevant recordings

Plan scripted prompts, conversational-style speech, dialogue, voice-over, customer-service scenarios, or domain-specific collection around the task your model must perform.

02 · SPEAKER MIX

Brief the right voices

Set languages, regions, age bands, gender balance, first-language profile, recording environment, device and demographic constraints appropriate to the project.

03 · DELIVERY

Choose usable outputs

Receive agreed audio specifications with structured transcripts, time alignment, metadata and a delivery format designed around your pipeline.

Language coverage

Start with Nigerian English and Yorùbá. Commission the next language to your brief.

BSG DataWorks supports African-language speech work through native-speaker networks and commissioned collection. Availability depends on the language, scope and collection plan.

Available now

Nigerian-accented English

Speech samples and commissioned collection for voice, transcription and conversational AI work that needs Nigerian English rather than a generic English proxy.

Available now

Yorùbá

Speech and time-aligned transcription support for projects that require proper tonal orthography and native-speaker language oversight.

In pipeline

Igbo, Hausa & Nigerian Pidgin

Collection is being developed for these language needs. Early project conversations can shape the brief and priority requirements.

On commission

Kanuri, Efik & minority languages

For languages missing from major open datasets, commission a bespoke scope through an appropriate native-speaker collection plan.

How commissioning works

A speech-data project with clear operational stages.

We use an explicit process so your team can evaluate scope, evidence and delivery before committing to a full collection.

Brief

Share the intended model use, languages, target volume, speaker requirements and delivery deadline.

Scope

We return initial questions, a likely plan, quality criteria and a commercial path for the work.

Collect & QA

We recruit and consent contributors, collect to the agreed specification, and run the required language and technical checks.

Deliver

You receive the agreed files, metadata, documentation and licensing materials through secure transfer.

What to ask a speech-data supplier

Evaluate more than the audio.

A useful speech-data conversation should include the collection plan and the records behind it, not only a sample file.

  • How will speakers, languages and scenarios be selected against the brief?
  • Which transcription, time-alignment and metadata fields will be delivered?
  • What native-speaker and technical quality checks will be applied?
  • How will contributor consent, commercial rights and delivery records be documented?

Frequently asked questions

Questions teams ask before commissioning African speech data.

Can we commission a dataset for a language not shown on the homepage?

Yes. The languages shown reflect current availability and pipeline status. For a language not listed, send a brief with the intended use, target speakers, volume and timeline so we can assess a bespoke collection plan.

Can you provide transcripts and time alignment?

Yes, transcription and time-alignment requirements can be included in the scope. State the format, orthography, annotation and metadata requirements in your initial brief.

Do you work with commercial AI training and licensing requirements?

Yes. BSG DataWorks is set up for commissioned collection and commercial licensing discussions. Rights, documentation and the permitted use should be addressed in the project scope and agreement.

What should we send in the first email?

Tell us the model use, languages or speaker profile, target volume, delivery timeline, governance requirements and an approximate pilot budget or licensing range if known. If you are still scoping any item, say so; we can address it in a call.

Start a speech-data brief

Tell us what your model needs.

Share a few project details and we will reply within 48 hours with initial scoping questions, a likely timeline and the next step.

What helps us get started

  • Intended model use and speech-data scenario
  • Languages, markets or participant profile
  • Target volume, format and transcript requirements
  • Desired delivery timeline
  • Rights, governance or compliance requirements
  • Approximate pilot budget or licensing range, if known

If you do not have all of this yet, send what you have — including if you are still scoping the budget. We will fill in the rest on a call.

Prefer to write directly? dataservices@bsgdataworks.com

WhatsApp only: +234 806 790 4903

Start a speech-data brief by email