A collaborative multilingual annotation platform.

Multilingual data,
annotated together.

ELMAnnotationStudio is a collaborative annotation platform for African languages — where communities build the datasets their languages have been missing, many people and many languages at once, with every correction reviewed by someone who speaks it.

Free to start · 1,000 prompts in a solo project · any language to any language

AcoliACH AfarAAR AfrikaansAFR AkanAKA BambaraBAM BasaaBAS BembaBEM DinkaDIN DioulaDYU EfikEFI EnglishENG EsanESN EweEWE FonFON FulfideFUL GaGAA GagnoabeteGNB GandaLUG HausaHAU IgboIBO IsokoISO JulaJOD KabuverdianuKEA KalangaKLQ KambaKAM KaondeKQN KikuyuKIK KimbunduKMB KinyarwandaKIN KongoKON KrioKRI LambaLAM LingalaLIN LoziLOZ LubakatangaLUB LubaluluaLUA LundaLUN LuoLUO LuvaleLUE AcoliACH AfarAAR AfrikaansAFR AkanAKA BambaraBAM BasaaBAS BembaBEM DinkaDIN DioulaDYU EfikEFI EnglishENG EsanESN EweEWE FonFON FulfideFUL GaGAA GagnoabeteGNB GandaLUG HausaHAU IgboIBO IsokoISO JulaJOD KabuverdianuKEA KalangaKLQ KambaKAM KaondeKQN KikuyuKIK KimbunduKMB KinyarwandaKIN KongoKON KrioKRI LambaLAM LingalaLIN LoziLOZ LubakatangaLUB LubaluluaLUA LundaLUN LuoLUO LuvaleLUE
MakhumaMAK MalagasyMLG MendeMEN NamaNAQ NandeNNE NdebeleNDE NdongaNDO NorthernsothoNSO NuerNUS NyanjaNYA NyankoleNYN PidginPCM PularFUF RundiRUN SangoSAG ShonaSNA SidamoSID SomaliSOM SouthernsothoSOT SouthndebeleNBL SukumaSUK SwahiliSWA SwatiSSW TigrinyaTIR TivTIV TongaTOI TsongaTSO TswanaTSN TumbukaTUM TwiTWI UmbunduUMB UrhoboURH VendaVEN WolayttaWAL WolofWOL XhosaXHO YorubaYOR ZuluZUL MakhumaMAK MalagasyMLG MendeMEN NamaNAQ NandeNNE NdebeleNDE NdongaNDO NorthernsothoNSO NuerNUS NyanjaNYA NyankoleNYN PidginPCM PularFUF RundiRUN SangoSAG ShonaSNA SidamoSID SomaliSOM SouthernsothoSOT SouthndebeleNBL SukumaSUK SwahiliSWA SwatiSSW TigrinyaTIR TivTIV TongaTOI TsongaTSO TswanaTSN TumbukaTUM TwiTWI UmbunduUMB UrhoboURH VendaVEN WolayttaWAL WolofWOL XhosaXHO YorubaYOR ZuluZUL
Collaboration, not a queue

One prompt. Many hands. Many languages.

Most annotation tools hand one item to one person and lock everyone else out. Here a prompt has a lane per language, and people work them at the same time.

Prompt 042 English

Take one tablet after meals, twice a day.

Four people, four languages, the same prompt — at the same time. A prompt is not finished until every lane has something in it, so a gap surfaces while there is still time to fill it.

Why this exists

Frontier models reason fluently in English. Elsewhere, the data runs out.

For Hausa, Bemba, Yoruba, Zulu and dozens of languages spoken by hundreds of millions, the depth of data English enjoys barely exists — and what does exist is mostly scraped, unreviewed, and written by nobody in particular.

What's abundant

  • Web text and news articles
  • Basic parallel translation pairs
  • Generic, unreviewed scraped data

What's missing

  • Corrections made by speakers, not crowds
  • Preference judgments between real alternatives
  • A record of who reviewed what, and whether they agreed
  • Domain-specific, culturally grounded material

ELMAnnotationStudio produces exactly that missing layer — one annotated example, one human review, one language at a time.

The real loop

Annotate, review, release.

This is the studio's own interface, working. Prefer a translation, score it, highlight a phrase you are unsure of, then send it through review.

Prompt 001 drafted English

Take one tablet after meals, twice a day.

English Bemba · either direction works · up to 6 translations per language, lettered within the lane

Preferences, scores and notes all travel with the row into the dataset.

Translation · Bemba

Kwata panuma ya , imiku ibili pa bushiku.

Select any phrase to check it against another language.

Saved as a segment against this exact span — reviewers see it, and it travels into the release.

Source

Take one tablet after meals, twice a day.

Bemba

Nwena akatembo kamo ilyo wapwisha ukulya, imiku ibili pa bushiku.

Sending work back needs a reason.

Cleared. This row can now go into a VAL release. Sent back to the annotator, with your reason attached to the pair. Dropped. It stays on the record but is kept out of clean releases.

Every decision records who made it and why. Nothing reaches a dataset on one person's say-so.

Dataset code name

Language

Modality

Review

Sequence

001

A set is only VAL when every row in it was cleared. Mixed rows come out PAR. The name is derived from the rows, not chosen — so it cannot overstate the work.

Every panel above is live — click through it.

What a project collects

One workspace, whatever the task is.

A project declares its source language, its targets, an objective and a modality when you open it. The interface is the same either way.

Objectives

Translation Summarization Question answering Named entity recognition Response evaluation RAG evaluation Instruction tuning

Modalities

TXTText SPHSpeech IMGImage VIDVideo MMMultimodal

Where the target text comes from

Generated from prompts

Members write a prompt; the model produces the target.

Uploaded source–target pairs

Both sides arrive in the file; members judge and correct.

One source, many targets

A single source rendered into several languages side by side.

Upload .csv, .tsv, .jsonl or .txt.

Built for communities

The language belongs to the people annotating it.

Work alone in a solo project, or open a community one and bring collaborators in with roles, assignments, a shared forum, and review that stays inside the community.

Roles that stack

Each carries everything below it, so you assign one seat rather than a checklist.

  • 01AnnotatorWrites and corrects translations, rates outputs, raises questions.
  • 02ReviewerEverything an annotator does, plus clearing or returning work.
  • 03AuditorEverything a reviewer does, plus a second pass and building releases.
  • 04StewardEverything an auditor does, plus seats, invites and assignments.
  • 05Community leadEverything a steward does, plus closing the project.

Language lanes

Several annotators hold the same prompt at once, each in their own language — nobody waits for a queue.

A forum that keeps its answers

Ambiguous phrase? Raise it. Each thread has its own page, takes an image, and stays searchable — including inside the replies.

Access on the record

Community projects run inside a grant or subscription window. When it ends the work stays and the project goes read-only — nothing is deleted.

Conduct with a paper trail

Members report violations to administrators, not to the lead. Leads remove people with a reason on the record.

For developers

The same models, behind an API key.

Any source, any target. Translate a string, or send one prompt into several languages at once. 100 requests a day on the free tier.

POST
curl https://www.everylanguagematters.com/api/studio/v1/translate \
  -H "Authorization: Bearer $ELM_KEY" \
  -H "Content-Type: application/json" \
  -d '{"text":"Take one tablet after meals.",
       "source":"en","target":"bem"}'
curl https://www.everylanguagematters.com/api/studio/v1/generate \
  -H "Authorization: Bearer $ELM_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"Take one tablet after meals.",
       "source":"en","targets":["bem","kab","yo"]}'

Keys with scope

A key declares whether it may translate, generate, or both — and which language pairs it may reach.

Usage in every response

Each reply carries how much of the day is left, so you never poll a second endpoint.

Refusals are free

Hitting the cap returns 429 with the reset time. Refused calls are never counted against you.

77 African languages

Sixty ways to say it.

Any language to any language. Each carries its ISO 639-3 code — the same code that goes into the dataset name, so a file can be cited without ambiguity.

Any of these can be the source, any can be the target — you pick the pair when you open a project.

Showing of 77 languages

UG Acoli ACH ET Afar AAR ZA Afrikaans AFR GH Akan AKA ML Bambara BAM CM Basaa BAS ZM Bemba BEM SS Dinka DIN CI Dioula DYU NG Efik EFI -- English ENG NG Esan ESN GH Ewe EWE BJ Fon FON WA Fulfide FUL GH Ga GAA CI Gagnoabete GNB UG Ganda LUG WA Hausa HAU NG Igbo IBO NG Isoko ISO CI Jula JOD CV Kabuverdianu KEA ZW Kalanga KLQ KE Kamba KAM ZM Kaonde KQN KE Kikuyu KIK AO Kimbundu KMB RW Kinyarwanda KIN CD Kongo KON SL Krio KRI ZM Lamba LAM CD Lingala LIN ZM Lozi LOZ CD Lubakatanga LUB CD Lubalulua LUA ZM Lunda LUN KE Luo LUO ZM Luvale LUE ZM Makhuma MAK MG Malagasy MLG SL Mende MEN NA Nama NAQ CD Nande NNE ZW Ndebele NDE NA Ndonga NDO ZA Northernsotho NSO SS Nuer NUS ZM Nyanja NYA UG Nyankole NYN NG Pidgin PCM GN Pular FUF BI Rundi RUN CF Sango SAG ZW Shona SNA ET Sidamo SID SO Somali SOM ZA Southernsotho SOT ZA Southndebele NBL TZ Sukuma SUK EA Swahili SWA ZA Swati SSW ET Tigrinya TIR NG Tiv TIV ZM Tonga TOI ZA Tsonga TSO ZA Tswana TSN ZM Tumbuka TUM GH Twi TWI AO Umbundu UMB NG Urhobo URH ZA Venda VEN ET Wolaytta WAL SN Wolof WOL ZA Xhosa XHO NG Yoruba YOR ZA Zulu ZUL

No matching language yet — tell us what to add next.

Open now

Bring your community.
Bring your language.

A solo project holds 1,000 prompts — enough to fine-tune on. If it earns its place, a community project brings your collaborators in.

ELMAnnotationStudio is part of the EveryLanguageMatters platform — building foundational multilingual AI infrastructure for every community, in every region and language.