Annotation services · existing or commissioned data

NDPC REGISTERED · NDPC/DCP/13596

Registered with the Nigeria Data Protection Commission. Certificate available on request. See our Compliance page and Privacy Policy.

Human-led annotation and labelling for AI training data.

Make speech, text, image and video data usable for your model workflow. BSG DataWorks applies trained native-speaker and project-specific annotation support to your existing materials or to data collected through our contributor networks.

Service scope

Bring your data, your schema, or a collection brief.

We support teams that already hold data and need it made usable, as well as teams that want data collection and annotation planned together. The task definitions, quality criteria and delivery format are agreed before work begins.

01 · EXISTING DATA

Annotate what you already hold.

Apply a defined annotation task to your speech, text, image or video materials, with the workflow scoped to your model objective and label requirements.

02 · COLLECTION + LABELS

Plan data and annotation together.

Where BSG DataWorks collects the source material, annotation, transcription, metadata and documentation can be scoped as part of the same delivery.

03 · SPECIFICATION

Set a usable annotation brief.

Define the taxonomy, examples, format, review requirements, handoff criteria and timeline before the annotation work is released.

Annotation tasks

Label the data your AI system actually needs.

Choose a task that matches the model pipeline rather than a one-size-fits-all service. BSG DataWorks scopes the relevant annotation type, language treatment and delivery expectation with you.

Speech and audio

Transcription, time alignment, speaker diarisation and related speech-data annotations for voice, ASR, conversational and evaluation workflows.

Text and language

Classification, categorisation, entity and sentiment labelling, with native-speaker language support where the task requires it.

Image and video

Image and video tagging, classification and segmentation tasks aligned to the agreed schema and target model use.

Combined datasets

Use an agreed workflow for data that needs multiple task types, such as speech with transcripts and metadata or video with structured tags.

Quality and delivery

Agree the evidence behind the labels before work starts.

Annotation is documented around the task definition, expected outputs, quality criteria and delivery route. The final package is determined by the agreed project scope.

Schema and outputs

  • Classification, categorisation, segmentation or diarisation to brief
  • Transcription and time-alignment options for speech tasks
  • JSON, TXT or CSV delivery formats where agreed
  • Metadata fields and label structure aligned to the project brief

Quality and documentation

  • Trained native-speaker support for relevant language tasks
  • Native-speaker linguists review transcriptions
  • Inter-annotator agreement of at least 80%, reported per batch where applicable
  • Quality criteria, reporting and redelivery thresholds agreed for the project

From brief to delivery

A practical route from raw files to usable labels.

Share enough detail for a suitable workflow to be scoped, then agree the annotation instructions and quality criteria before production begins.

Brief

Share the model use, source data type, volume, target labels, language requirements and intended delivery date.

Scope

Agree the annotation task, label schema, delivery format, quality criteria and commercial route.

Annotate

Run the agreed workflow with trained annotators and native-speaker support where the task requires it.

Deliver

Receive the agreed labels, metadata, quality reporting and documentation through a client-specified secure transfer or approved destination.

Frequently asked questions

Questions teams ask before starting an annotation project.

Can you annotate data we already own?

Yes. BSG DataWorks provides annotation support for existing client data as well as for data collected through our networks. Start with the data type, model use, task definition and expected output.

Which annotation tasks can you support?

Current service scope includes transcription and time alignment, classification and categorisation, segmentation and diarisation, entity and sentiment labelling, and image or video tagging to an agreed schema.

Can you work to our label taxonomy?

Yes. Share your existing schema, examples, definitions and delivery requirements in the initial brief. The annotation workflow and quality criteria will be agreed before production.

How is annotation quality handled?

Quality criteria and reporting are agreed for the project. For relevant language work, BSG DataWorks uses native-speaker support; transcriptions are reviewed by native-speaker linguists, and inter-annotator agreement is reported per batch where applicable.

Start an annotation brief

Tell us what your model needs labelled.

Share a few project details and we will reply within 48 hours with initial scoping questions, a likely timeline and the next step.

What helps us get started

  • Model use and annotation task
  • Existing data type, language or modality
  • Target volume and current data format
  • Label schema, examples or quality requirements
  • Required delivery format and desired timeline
  • Approximate pilot budget or service range, if known

If you do not have all of this yet, send what you have — including if you are still scoping the budget. We will fill in the rest on a call.

Prefer to write directly? dataservices@bsgdataworks.com

WhatsApp only: +234 806 790 4903

Start an annotation brief by email