DPGs for AI Collection
All DPGs included in the DPGs for AI Collection must meet the requirements outlined in the DPG4AI collection criteria.
The information below reflects the self-reported responses provided as part of the assessment process.
Solution Type: Open Source Software
DPG Compliance & Profile Page: https://www.digitalpublicgoods.net/r/simpleaudit
Description: Lightweight AI safety auditing framework for red-teaming AI systems through adversarial probing. Supports multilingual testing across safety, healthcare, and RAG scenarios. Works with cloud APIs or fully local models.
Assessment Status: Under Review
Review Date: 2026-06-14
Category Fit:
Open Source Software (Python library, MIT). SimpleAudit provides specific functionality in the Verification & Validation stage — it red-teams and audits LLM-based systems through adversarial probing and LLM-as-judge scoring — and supports Operation & Monitoring via repeatable, versioned audit runs. Verified DPG. Repo: https://github.com/kelkalot/simpleaudit
AI Lifecycle Utility:
Documented Relevance/ Impact:
Purpose-built for AI: a multilingual safety-auditing / red-teaming framework. Methodology validated in a 2026 paper (arXiv:2605.06652) and applied to a Norwegian public-sector procurement case (Borealis vs Gemma 3). Built by Simula/SimulaMet with the Norwegian Directorate of Health. 13 built-in scenario packs (safety, RAG, health, epistemic). Live demo: https://simulamet-simpleauditvisualization.hf.space
Adoption Readiness Level:
L5 Productized / Plug-and-Play
Adoption Readiness Evidence:
Published on PyPI (`pip install simpleaudit`), semantic-versioned, CI test workflow, runnable via uvx, with a packaged visualization server. Governance docs present: DPG.md, CODE_OF_CONDUCT.md, SECURITY.md. Active maintenance (recent commits, 2026 paper). https://pypi.org/project/simpleaudit/
Interoperability Level:
L4 Integrated — versioned API/SDK, plug-in capable, automated synchronization
Interoperability Evidence:
Targets OpenAI-compatible chat-completions endpoints; provider-agnostic via any-llm-sdk (Ollama, vLLM, HF, OpenAI, Anthropic, etc.). Results export to standard JSON; charts to PNG. Pluggable scenario packs and judge configs (single-parameter swap). README target-API + results schema: https://github.com/kelkalot/simpleaudit#understanding-results
Responsible Practices Level:
Equity & Inclusion:
Responsible AI Tooling:
Responsible Practices:
Inclusion & Autonomy:
L3 Fully Documented — comprehensive limitations, biases, and failure modes statement
Responsible Practices Evidence:
The methodology paper (arXiv:2605.06652) documents explicit limitations: judge/auditor non-determinism; the requirement that auditor capability match the target range (too-strong auditors floor safe-target scores and erase the deltas the instrument reports); reproducibility bounds (~1 pt on the 0–100 scale by n=10); and a 'report the bundle, not a leaderboard' caveat. README 'Why SimpleAudit?' and Methodology sections reiterate these.