The outcome you want is a dataset that a model team can use without relabeling the same work twice. That starts before annotation begins.
A strong data annotation vendor does more than assign labels.
The vendor helps test whether the label guide, examples, file format, and review sample match the dataset. This matters more when the data is multilingual. Meaning, script, region, and spoken context can change the label decision.
The 5 checks before vendor selection
Use these 5 checks before choosing a vendor:
- Data type: text, image, video, audio, or a mixed dataset.
- Language coverage: source language, target language, dialect note, and script.
- Label guide: definitions, examples, edge cases, and disallowed labels.
- Feedback contact: sample size, feedback role, feedback format, and acceptance rule.
- Output format: CSV, JSON, XML, platform export, or internal schema.
If any one of those 5 is missing, the first batch may become a guessing exercise.
What “multilingual” changes in annotation
A monolingual English annotation task and a multilingual one look similar in a quote and very different in production. Three differences drive most of the risk:
Label meaning can shift by language. A sentiment label that is stable in English may split by dialect or register in another language. The label guide needs per-language examples, not just translated definitions.
The worker profile changes. Some tasks need native-language annotators; some need bilingual reviewers; some need subject specialists who also speak the language. These are different pools at different rates, and a vendor who quotes one blended rate for all three has not planned the staffing.
Quality evidence must be comparable across languages. Agreement scores, error categories, and review notes should be reported per language so a weak pair is visible instead of averaged away inside a portfolio-wide number.
| Annotation area | Extra question for multilingual work |
|---|---|
| Text classification | Do label examples exist in each language, or only translated from English? |
| Image annotation | Do region-specific scenes change label boundaries? |
| Audio and speech | Which dialects and accents are in scope per language? |
| Video | Do on-screen text and signage need transcription in the source script? |
| Model-output evaluation | Is the evaluator rating fluency, accuracy, or both, per language? |
Why a pilot should come first
A pilot does not need to be large. It needs to be representative. A small sample can expose unclear labels, missing examples, file issues, and cases where language expertise is required.
For multilingual annotation, a pilot is also where teams discover whether the task needs translators, native-language reviewers, subject reviewers, or general annotators. Those are different profiles.
A useful pilot plan names:
- Record count per language, not just a total
- The exact label-guide version used
- Who reviews the pilot output and in what format
- The acceptance rule: what agreement or accuracy level passes
- What happens to pilot learnings (guide revision, example additions, retraining)
What to ask in the request response
Ask for the vendor’s proposed label workflow, sample review plan, output format, and exception handling. Also ask which parts of the work need language review. A clean request response should separate mechanical tagging from language-dependent decisions.
For the annotation scope itself, the AI data annotation services page shows how DD structures label guides, pilots, and per-batch reporting. For a vendor-neutral worksheet you can send to any shortlisted vendor, use the data annotation vendor worksheet.
A 6-point vendor comparison checklist
Before a vendor is selected, compare each response against the same 6 fields:
- Sample design: the vendor names the number of records, files, minutes, or images in the pilot.
- Label rules: the response shows how edge cases, rejected labels, and unclear items will be handled.
- Language fit: the vendor separates script, dialect, region, and subject review instead of grouping all language work together.
- Review sample: the response names the percentage or count of items checked before the first full batch moves.
- Output test: the vendor confirms one delivery file can be opened by the client’s platform before production.
- Rework rule: the request states what counts as a defect, who reviews it, and how corrected labels are returned.
If 2 vendors quote the same dataset but only 1 names those 6 fields, the clearer request is usually safer than the lower line item. Price matters, but unlabeled rework is where annotation budgets drift.
Reading a per-batch quality report
Once production starts, the quality report is the only visibility you have. A useful per-batch report contains:
- Items delivered and items rejected, per language
- Agreement between annotators on double-labeled samples, per language
- Top error categories with example records
- Guideline questions raised by annotators and the answers given
- Any label-guide version change that applies from this batch forward
If a report shows only a total accuracy figure with no per-language split, ask for the split. A portfolio average can hide a failing pair for months.
Frequently asked questions
How big should a pilot be?
Large enough to include the hard cases: every label class, every language, and a deliberate sample of edge cases. For text, that is often a few hundred records per language; for audio, a few hours per language. The point is coverage of the decision space, not volume.
Should annotators see each other’s labels?
For a measured agreement sample, no — double-label a subset blind and compare. For production efficiency, annotators should see resolved edge-case rulings so the same question is not relitigated record by record.
Who owns the label guide?
You do. The vendor should propose revisions when annotators hit ambiguity, but every accepted change should be versioned and visible to your team. A guide that drifts silently produces datasets that cannot be compared across batches.
What output format should we request?
Whatever your training or evaluation pipeline reads directly, plus one lossless export (JSON or CSV) you control. Ask the vendor to open-test one delivery file against your platform during the pilot, not after the first full batch.
How do we compare vendors on languages we cannot read?
Ask for the per-language quality evidence described above, and run a small gold set — records your team already knows the right labels for — through each vendor’s pilot. Gold-set performance per language is comparable even when your team does not speak the language.
What should we do when the pilot fails?
Treat it as information, not disaster. A pilot that exposes a weak label guide, an underpowered reviewer bench, or an unclear acceptance threshold has done its job before real budget moved. Ask the vendor for a written root-cause read: which classes failed, whether the failure was guide ambiguity or annotator skill, and what changes would fix it. A vendor who can answer precisely is worth a second pilot; a vendor who blames the data without evidence is telling you how production escalations will go.
Dynamic Dialects plans multilingual annotation across text, image, video, and audio datasets. Requests can include 250+ language coverage, label guide review, pilot planning, and output files prepared for the client’s system. Start with the AI data services overview or send dataset details through the contact form.