
AI Verify is Singapore's voluntary AI governance testing framework and open-source toolkit, developed by IMDA. It tests AI systems against 11 internationally aligned governance principles using technical tests plus process checks. It is self-assessment, not certification. For LLM applications, the companion Project Moonshot toolkit adds red-teaming and benchmarking.
What does AI Verify test?
AI systems, against 11 internationally aligned governance principles, using two kinds of evidence: technical tests run against the system itself and process checks on the governance around it. The principles are the ones that recur across international AI governance documents. Technical tests are quantitative; they run against the system and return measurements. Process checks are documentary; they ask whether a governance practice exists and whether records show it operating.
- Transparency
- Explainability
- Repeatability and reproducibility
- Safety
- Security
- Robustness
- Fairness
- Data governance
- Accountability
- Human agency and oversight
- Inclusive growth and societal well-being
The pairing of test and process is the design insight. A fairness metric without a documented process is a number without a promise. A process document without a test is a promise without a number. AI Verify asks for both, per principle, which is why the output reads as evidence rather than marketing. It is also what makes the reports hard to game. Numbers can be cherry-picked and prose can be polished, but producing both, aligned, across 11 principles is roughly as much work as doing the governance.
Who runs it and what's the Foundation?
You run it. IMDA developed AI Verify, but the testing is self-assessment: your team runs the toolkit against your systems and produces the report. No accredited body signs it, and no government agency reviews it. What you get is a structured account of how your system performs against the principles, produced by you, for whoever you choose to show it to. Self-assessment sounds weaker than certification, and for some audiences it is. It still beats silence by a distance. A buyer comparing two vendors, one holding a structured test report and one holding adjectives, learns something real from the difference, and the report's standard structure makes vendors comparable in a way bespoke claims never are.
Stewardship of the toolkit sits with the AI Verify Foundation, launched in June 2023 to develop it as an open-source project. The open-source structure has a practical consequence for adopters. The toolkit's development happens in public, so you can inspect exactly what a test measures before staking a customer conversation on its result. Closed testing products ask for trust in the tester; this one lets you read the code. It also means nothing about using the toolkit requires approval or membership. The code is public; you can start this week.
What is Project Moonshot?
The companion toolkit for LLM applications, and like AI Verify it is open-source. Moonshot adds two instruments the base framework's tests were not built around. Red-teaming, meaning structured adversarial prompting to find the inputs that make your application misbehave before a user or an attacker does. Benchmarking, meaning running the application against defined test sets so its performance can be compared and tracked over time. Benchmark results age quickly as the underlying models change, which argues for re-running them on a schedule rather than once.
The reason it exists as a companion rather than a feature is that LLM applications fail differently. A loan-scoring model gives wrong answers within a narrow output space. An LLM application can be talked into behaviour nobody designed, in fluent prose. Testing for that requires an adversary rather than a metric alone, and red-teaming is the adversary made systematic.
If your AI exposure is a chatbot or anything else built on a language model, Moonshot is the half of the ecosystem that applies to you most directly.
Should you use it?
The decision rule is what you ship. If you build or deploy AI models, yes: the toolkit gives you evidence about your own systems that no policy document can produce, and a self-assessment report is a strong artefact to hand a procurement team that asked what your governance amounts to in practice. Start with one system rather than a programme. Running the toolkit against your highest-stakes model teaches you more about your governance gaps in a fortnight than a quarter of policy drafting.
If you merely use SaaS AI tools that vendors run for you, the testing toolkit solves a problem you do not have. You cannot meaningfully run technical tests against a system you neither host nor control. Your governance effort belongs in acceptable-use rules and vendor checks instead, with AI Verify appearing only as a question you put to vendors: have they tested, and will they show you the results. Asked at renewal, that question is cheap and surprisingly clarifying.
The middle case is common in Singapore. A company deploys a vendor's model but tunes it on its own data, or wraps it in an application whose behaviour is now its responsibility. That company is a builder in the sense that matters here. The vendor tested their model; nobody has tested your application. Voluntary today is a low price for finding that out privately. The test is responsibility, not authorship. If your name is on the outcome, you are the one who needs the evidence.
Zavior files AI Verify outputs alongside the rest of your governance evidence, so a test report answers a buyer's question instead of sitting in a repo.
Frequently asked questions
Is AI Verify a certification?
No. It is self-assessment: you run the tests and produce the report yourself, and no accredited body attests the result. If a vendor waves an "AI Verify certificate" at you, ask sharper questions about who exactly signed it.
Is it free?
The toolkit is open-source, so there is no licence fee to run it. The real cost is engineering time: integrating your systems with the tests, and then acting on what the tests find, which is the part organisations underbudget.
Does it apply to generative AI?
Yes, through Project Moonshot, the companion toolkit built for LLM applications. It adds red-teaming and benchmarking designed for generative outputs, alongside the process checks that apply to any AI system regardless of type.
Zavior · AI Governance
A test report is worth more inside a governance programme than beside one. Zavior finds the AI tools your team already runs, wraps them in usage rules and training people will actually follow, and tracks what those tools do, so an AI Verify result reads as one line of a managed system rather than a one-off. Not sure you need it yet? The free AI Readiness Scan scores where you stand first.
Book a free 30-minute business assessment →This is general information, not legal advice.