What is AI assurance?
Most organisations have AI principles and a committee. Far fewer can produce the dated evidence that shows a specific AI system was assessed, reviewed by someone independent, and approved by a named person. That difference is AI assurance.
- Assurance is a process with outputs, not a statement of intent: classified risk, matched controls, dated evidence, independent review and a recorded decision.
- Self-attestation fails under scrutiny because it records an opinion rather than the basis for it.
- Control obligations scale with the risk tier of the use case, not with the size of the model.
- Australian obligations arrive from several directions at once, so the same evidence has to answer to several frameworks.
- The people who build an AI system should not be the people who judge whether its evidence is sufficient.
A working definition
AI assurance is the disciplined process of establishing, evidencing and independently reviewing whether a specific AI system is fit to operate in a specific context — and recording that position so it can be defended later.
The emphasis on specific matters. Assurance is not a property of a model. The same large language model can be entirely appropriate for drafting internal meeting notes and entirely inappropriate for triaging clinical referrals. Assurance attaches to the use case: what the system does, who it affects, how much autonomy it holds, and what happens when it is wrong.
The emphasis on recorded matters too. An assurance position that lives in someone's memory, or in a slide from eight months ago, is not assurance. The test is whether an informed outsider — a board member, an auditor, a regulator, a customer's procurement team — can follow the reasoning from the conclusion back to the evidence that supports it.
Why principles and checklists are not enough
Three artefacts are commonly mistaken for assurance. Each is useful. None of them, on its own, survives scrutiny.
Published AI principles
Principles establish intent and set a standard to be held to. They say nothing about whether a particular system met that standard, and they carry no evidence. An organisation can hold excellent principles and still deploy an unassessed system.
A completed questionnaire
A team confirming that it has tested for bias records an opinion. Assurance records the test: the dataset, the method, the measured result, the date, the person accountable and the threshold the result was judged against.
A vendor's trust page
Supplier documentation describes the supplier's controls in the supplier's words, at an unknown point in time. It is an input to assurance, not a substitute for it — and it changes without notice.
| Self-attestation | Verifiable assurance | |
|---|---|---|
| What is recorded | A judgement | The judgement and its basis |
| Who decides | The team that built the system | An independent reviewer, separate from the builder |
| Evidence | Optional, often undated | Required, dated, owned, with a refresh point |
| Residual risk | Implied | Accepted in writing by a named executive, with an expiry |
| When the system changes | Unnoticed | Triggers re-review against the recorded position |
| Under later questioning | Relies on recollection | Traceable to the record |
The four pillars of AI assurance
Assurance in practice rests on four disciplines. Weakness in any one of them undermines the other three.
1. Contextual risk classification
Classify by consequence, not by technology. The factors that drive risk are the affected population, the sensitivity of the data, the degree of autonomy, the severity of a plausible failure, and whether a person can realistically intervene in time. This classification is what determines how much assurance the system actually needs.
2. Matched control obligations
Once a system is classified, it inherits a defined set of controls. Higher tiers inherit more. This is where most organisations first see the real size of the task: a high-risk customer-facing system can carry well over a hundred applicable control obligations, each needing an owner and evidence.
3. Dated evidence with ownership and expiry
Every control is supported by something an outsider can inspect: an evaluation report, a model card, a penetration test, a data-flow description, a signed supplier assessment. Each carries a date, an accountable owner and a point at which it must be refreshed. Evidence without an expiry quietly becomes a historical document.
4. Role separation and governance gates
Builders do not judge their own evidence. A reviewer independent of delivery forms the assurance opinion; a named executive records the readiness decision and accepts any residual risk in writing. Two-person accountability on evidence judgement, finding closure and readiness decisions is what makes the record credible rather than merely complete.
Where evaluations fit
Assurance is the governance frame. Evaluations are how the frame gets filled with fact. Without measurement, an assurance record is a well-organised set of assertions.
Evaluations supply the empirical evidence behind the controls that matter most: accuracy and grounding, safety and guardrail resilience, fairness across affected groups, the quality of human oversight, and behaviour after deployment. Generic model benchmarks do not do this job, because they test the model rather than the organisation's application of it.
The Australian context
Australian organisations are answering to several overlapping expectations at once. Assurance is what lets one body of evidence serve all of them.
This guide describes assurance practice. It is not legal advice, and no assurance process can certify compliance on an organisation's behalf.
Voluntary AI Safety Standard
Ten guardrails covering accountability, risk management, data governance, testing, human oversight, transparency, contestability, supply chain, records and stakeholder engagement. Useful as a structure, but each guardrail has to become a control with evidence behind it.
Privacy Act and the Australian Privacy Principles
Applies wherever personal information reaches an AI system. The practical demands are data lineage, minimisation, disclosure to third parties including model providers, and preventing inadvertent retention or leakage.
ISO/IEC 42001 and the NIST AI RMF
ISO/IEC 42001 provides an auditable AI management system; the NIST AI RMF provides the Govern, Map, Measure and Manage structure. Mapping day-to-day evidence to both means one assurance record answers certification and risk-framework questions alike.
Sector and board expectations
Financial services, health, government and higher education each add their own expectations, and directors carry duties that do not pause for new technology. Boards increasingly ask for the basis of an AI decision, not a summary of it.
The assurance lifecycle in five steps
Register and classify
Record what the system does, who owns it, what data it touches and who it affects. Produce an indicative risk tier from that context.
Match control obligations
Derive the applicable controls from the tier and the use case, and give every one of them a named owner.
Collect and date evidence
Gather evaluation results, technical documentation, supplier assessments and operational procedures, each with a date and a refresh point.
Independent review
A reviewer outside the delivery team judges sufficiency, raises findings, and states residual risk.
Recorded decision and monitoring
A named executive approves, approves with conditions, or declines — then monitoring watches for incidents, material change, evidence expiry and supplier developments that would reopen the question.
A realistic first step
Do not begin with the whole portfolio. Take the single AI system with the most exposure — usually the one closest to customers, staff decisions or sensitive data — and take it through the full lifecycle once. One system carried end to end teaches an organisation more than a partial inventory of twenty.
The two things worth measuring at the outset are the number of control obligations the system actually carries, and the proportion of those obligations with no evidence behind them today. Those two figures turn AI governance from an ambition into a scope.
See the indicative risk tier and control obligations for one of your AI systems.
The free AI Defensibility Check takes a few minutes and produces an executive summary you can share.
