DPGs for AI Collection
All DPGs included in the DPGs for AI Collection must meet the requirements outlined in the DPG4AI collection criteria.
The information below reflects the self-reported responses provided as part of the assessment process.
Solution Type: Open Data
DPG Compliance & Profile Page: https://www.digitalpublicgoods.net/r/wikipedia
Description: Wikipedia is a free online encyclopedia, created and edited by volunteers around the world and hosted by the Wikimedia Foundation.
Assessment Status: Under Review
Review Date: 2026-06-14
Category Fit:
Open Data / Open Content (CC-BY-SA-4.0). Wikipedia is one of the most widely used text corpora for pretraining and evaluating LLMs and NLP systems; it is explicitly available for bulk reuse via database dumps and APIs and is demonstrably used across the AI/ML field. Lifecycle roles: Data Acquisition & Preparation and Model Building & Training (pretraining corpus). https://www.wikipedia.org/
AI Lifecycle Utility:
Documented Relevance/ Impact:
Foundational AI training data: Wikipedia is a core component of major pretraining corpora (e.g. C4, The Pile, ROOTS) and a standard NLP benchmark source. 300+ language editions. Official structured / ML-ready distribution via Wikimedia Enterprise and the Hugging Face dataset `wikimedia/wikipedia`. Export docs: https://en.wikipedia.org/wiki/Help:Export
Adoption Readiness Level:
L4 Orchestrated / Optimized
Adoption Readiness Evidence:
Regularly versioned XML dumps (dated releases act as a de-facto changelog), MediaWiki REST/Action APIs, and the curated `wikimedia/wikipedia` HF dataset with loaders and parquet. Field/structure documented. Dumps: https://dumps.wikimedia.org ; HF: https://huggingface.co/datasets/wikimedia/wikipedia
Interoperability Level:
L5 Croissant + API — Croissant metadata standard + programmatic API access
Interoperability Evidence:
Native: XML dumps + HTML + MediaWiki API (RDF via DBpedia). ML-ready: the official `wikimedia/wikipedia` dataset on Hugging Face exposes auto-generated Croissant metadata plus the datasets API/loader and parquet files. https://huggingface.co/datasets/wikimedia/wikipedia
Responsible Practices Level:
Equity & Inclusion:
Responsible Practices:
Standards & Vocabularies:
L2 Minimum Disclosure — known limitations and collection context stated (narrative form)
Responsible Practices Evidence:
Collection context is fully transparent (volunteer-edited, policy-governed). Known biases are extensively documented by the community — systemic bias, gender gap, geographic/language coverage gaps (e.g. https://en.wikipedia.org/wiki/Wikipedia:Systemic\_bias ). No single formal datasheet / Croissant-RAI; documentation is narrative/essay form → L2. A consolidated data statement would reach L3.