Open Source

Built in the Open. Shared with the World.
Our Philosophy

Why open source is central to everything we do.

Giving back

The AI community has always moved fastest when knowledge is shared freely. Every Language Matters was built on the shoulders of open source tools, open research, and communities who chose to share rather than hoard. We believe the only credible response to that generosity is to give back — openly, consistently, and without reservation. Every dataset we produce, every tool we build, and every model we train that can be shared, will be shared.

Democratising AI

Proprietary data is one of the most significant barriers to inclusive AI development. When language datasets are locked behind commercial licences, only well-funded organisations can access them — which means the AI built on those datasets continues to serve the already-privileged. By releasing our work under open licences, we ensure that a researcher in Lusaka, a startup in Nairobi, or a university lab in Manila has the same access to high-quality multilingual data as any Silicon Valley company.

Transparency

Open source is also a commitment to accountability. When we publish our datasets, annotation guidelines, and model cards publicly, anyone can inspect our methodology, challenge our decisions, and improve on our work. We do not claim to be perfect — but we do commit to being transparent. Every release includes full documentation of how data was collected, who annotated it, what quality checks were applied, and what limitations exist.

Impact So Far

Our open source footprint.

0+

public datasets released

0+

languages represented in open data

0K

downloads across HuggingFace repos

0+

GitHub stars across our repositories

What We Share

What we publish and release openly.

Everything we release is documented, versioned, and free to use for research and non-commercial applications. Commercial licencing enquiries are welcome.

Annotated datasets

Text classification, NER, sentiment analysis, translation pairs, and transcription datasets across African, South Asian, and Southeast Asian languages. All datasets include full annotation guidelines and quality metadata.

CC BY 4.0

Evaluation benchmarks

Standardised evaluation sets for measuring LLM performance on low-resource languages. Designed to enable fair, reproducible comparison of multilingual models on tasks relevant to underrepresented communities.

Apache 2.0

Annotation tooling

Open source scripts, schemas, and workflow templates for setting up multilingual annotation pipelines. Includes our quality assurance frameworks and inter-annotator agreement calculators.

MIT

Fine-tuned models

Domain-adapted and language-specific fine-tuned models trained on our annotated corpora, published with full model cards, training details, and evaluation results. Ready to use or fine-tune further.

CC BY-NC 4.0

Research & papers

Our research findings, annotation methodology papers, and language documentation reports are published as preprints and open access wherever possible, with accompanying code and data repositories.

Open Access

Data schemas & standards

Our general-purpose annotation schema is designed to support translation, summarisation, NER, QA, RAG evaluation, chatbot scoring, and LLM fine-tuning without changing the underlying structure.

MIT
Find Our Work

Where to find everything we have published.

HuggingFace

huggingface.co/everylanguagematters

Our primary platform for datasets, models, benchmarks, and evaluation resources. Everything is versioned and ready to use via the HuggingFace ecosystem.

0 datasets 0 models 0K downloads
View HuggingFace

GitHub

github.com/everylanguagematters

All tooling, annotation systems, pipelines, and research code are open source. Contributions and discussions are welcome.

0 repositories 0+ stars 0 forks
View GitHub
Get Involved

How to contribute to our open source work.

You do not need to be a professional engineer to contribute. There are meaningful ways to help at every skill level.

Submit a pull request

Browse open issues on our GitHub repositories and submit fixes, improvements, or new features. All pull requests are reviewed and contributors are credited in release notes.

Review & validate datasets

Native speakers can review published datasets for accuracy, cultural appropriateness, and linguistic quality. Open a GitHub issue or contact us directly to join a validation review.

Improve documentation

Good documentation is what makes open source actually usable. Help us write clearer READMEs, better dataset cards, and more useful tutorials — especially in languages other than English.

Train and share models

If you use our datasets to train models, we encourage you to publish them back to the community on HuggingFace with a model card linking to our data. Tag us so we can amplify your work.

Report issues & suggest languages

Found an error in a dataset? Know a language we should cover next? Open a GitHub issue or email us. Community-driven language prioritisation is how we decide what to work on next.

Share & cite our work

If you use our datasets or tools in your research, please cite us. Visibility in the research community helps us attract contributors, funding, and partnerships that let us do more.

Open Source

The future of AI is built in the open.

Join researchers, engineers, and linguists from around the world who are building the data infrastructure for an AI that speaks every language. Star our repos, use our datasets, and help us go further — together.

0+

open datasets

0+

languages

0K

downloads

0+

GitHub stars