DPGs for AI Collection
All DPGs included in the DPGs for AI Collection must meet the requirements outlined in the DPG4AI collection criteria.
The information below reflects the self-reported responses provided as part of the assessment process.
Solution Type: Open Data
DPG Compliance & Profile Page: https://www.digitalpublicgoods.net/r/mozilla-common-voice-dataset
Description: A multilingual open voice dataset created by Mozilla Common Voice, an initiative to help teach machines how real people speak
Assessment Status: Under Review
Review Date: 2026-06-14
Category Fit:
Open Data (CC0-1.0). Mozilla Common Voice is a multilingual speech dataset explicitly created to train and evaluate speech technology (ASR/STT, language ID). Lifecycle roles: Data Acquisition & Preparation and Model Building & Training. https://commonvoice.mozilla.org/
AI Lifecycle Utility:
Documented Relevance/ Impact:
Designed-for-AI by mandate ('teach machines how real people speak'). One of the largest public multilingual voice corpora (130+ languages across releases) with demographic metadata (age, sex, accent). Standard training data for Whisper-class and community ASR models. Versioned corpora (…Corpus 22.0) on the site and HF (mozilla-foundation/common_voice_\*). Datasets: https://commonvoice.mozilla.org/en/datasets
Adoption Readiness Level:
L5 Productized / Plug-and-Play
Adoption Readiness Evidence:
Regular numbered releases (Corpus 2.0 → 22.0) with versioning/changelog metadata (https://github.com/common-voice/cv-dataset), per-language splits with TSV metadata, and plug-and-play HF dataset loaders. https://huggingface.co/datasets/mozilla-foundation/common_voice_17\_0
Interoperability Level:
L5 Croissant + API — Croissant metadata standard + programmatic API access
Interoperability Evidence:
MP3 audio + TSV metadata; distributed via Hugging Face datasets (mozilla-foundation/common_voice_\*) which exposes auto-generated Croissant metadata and the datasets API/streaming loader. ISO 639 language codes. https://huggingface.co/datasets/mozilla-foundation/common_voice_17\_0
Responsible Practices Level:
Equity & Inclusion:
Responsible Practices:
Inclusion & Autonomy:
Standards & Vocabularies:
L3 Full Disclosure — bias + representation + de-identification method + provenance documented (Croissant-RAI, Datasheet, or Data Statement)
Responsible Practices Evidence:
Demographic representation metadata (age/sex/accent) ships with the corpus, and known gender/language imbalances are documented. Collection context, consent, age limits, right-to-be-forgotten and de-identification (no PII in released audio/metadata) are documented in terms/privacy. Original data statement: Ardila et al., 'Common Voice: A Massively-Multilingual Speech Corpus' (LREC 2020). Terms: https://commonvoice.mozilla.org/en/terms