By Every Language Matters
ELM Voice Studio
The Complete Workspace for Multilingual Speech Data Collection. Built for creating high-quality ASR and TTS datasets in the languages the world actually speaks.
Professional speech annotation, built for annotators — not engineers.
The tool
ELM Voice Studio is a desktop annotation application designed specifically for building high-quality speech datasets used in Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems. It gives annotators a clean, focused workspace to record, review, segment, and label audio — without needing a technical background to get started.
Who it's for
Built for native speakers, community contributors, and professional annotation teams working across Every Language Matters projects and beyond. Whether you are annotating your first audio file or managing hundreds of hours of speech data, ELM Voice Studio scales with your workflow.
Why it matters
The quality of a voice AI system is only as good as the data it was trained on. ELM Voice Studio puts the tools for producing that data directly in the hands of the communities whose languages matter most — removing the technical barriers that have historically kept underrepresented languages out of the AI pipeline.
Everything an annotator needs, nothing they don't.
Audio recording & playback
Record straight into the app, or open audio you already have. Play it back, slow it down, loop a section, and move through the waveform a fraction of a second at a time.
Filters & noise detection
The app points out background noise, clipping and long silences, and gives you filters to clean them up — so weak audio does not end up in your dataset.
Prompted recording
Load a list of sentences and record them one at a time. The app shows each prompt, saves the take, and moves on to the next — the quickest way to build a dataset from a script.
Export to training format
Export as WAV at the sample rate your training needs, with transcripts and labels saved alongside in a plain, readable format. Ready to hand straight to a training script.
Works offline, anywhere
Everything runs on your own machine. No internet needed to record, label or export, and your recordings never leave your computer unless you send them somewhere yourself.
Built for both ASR and TTS annotation workflows.
Automatic Speech Recognition
Type out what was said and line the text up with the audio. That pairing — the sound and the words — is what a speech recognition model learns from.
- Listen to the audio
- Type what you hear
- Review the transcript
- Validate the recording
Text-to-Speech
Read from a script and record clean, steady speech, then add the small details that make a synthetic voice sound natural — including tone languages and clicks.
- Read the prompt
- Record your voice
- Review your takes
- Validate recordings
From your first recording to a production-ready dataset.
Set up your project
Start a project, pick ASR or TTS, choose the language, and set the labels you want to use. It takes a couple of minutes and the app walks you through it.
Record or import audio
Record with your microphone, or open WAV files you already have. WAV is what we recommend and what the app exports, so your audio keeps its full quality from the first take to the finished dataset.
Annotate with precision
Run the filters to clear background noise, hum and clipping, cut the dead air at the ends, and flag anything that still sounds wrong. Cleaning as you go beats fixing a whole set at the end.
Export & deliver
Export the project and you get the WAV files, the transcripts and labels, and a short summary of what is in the set — ready to train on, or to hand to whoever is doing the training.
9 languages. One tool for the world.
ELM Voice Studio's interface is fully available in the following languages, so annotators can work comfortably in their own.
English
English
Français
French
Español
Spanish
Português
Portuguese
中文
Chinese
Kiswahili
Swahili
हिन्दी
Hindi
Русский
Russian
العربية
Arabic
More coming soon
Tell us what to build next.
ELM Voice Studio is shaped by the people annotating in it. If something is slowing you down, a language needs support we have not added, or you can think of a feature that would make the work easier — send it to us. Every submission is read.
Share feedback or request a featureTakes about 5-7 minutes. No account needed.
Start annotating today.
ELM Voice Studio is free to download. Available for Windows, macOS, and Linux.