AI evaluations: measuring a system before it reaches production
Benchmark scores describe a model. Assurance needs evidence about your application of it — on your data, against your thresholds, with a result someone can re-run. This guide covers how to design evaluations that stand up as evidence.
- Public benchmarks test the model; assurance needs tests of your system in your context.
- An evaluation is only evidence when the dataset, method, threshold, result, date and owner are recorded together.
- Set the pass threshold before running the test, not after seeing the result.
- Red-teaming is a structured, repeatable exercise with recorded findings, not an afternoon of curious prompting.
- Pre-deployment evaluation is a starting point; behaviour drifts, so the cadence matters as much as the first result.
Why benchmarks are not evidence
A model's benchmark performance tells you something about general capability under laboratory conditions. It tells you nothing about whether your retrieval layer surfaces the right policy document, whether your prompt leaks customer data into a summary, or whether your system refuses appropriately when a user asks something it should not answer.
Assurance-grade evaluation measures the assembled system — model, prompts, retrieval, tools, guardrails, interface and the human process around it — on data that resembles the real workload, against thresholds decided in advance.
Five evaluation disciplines
Most assurance-relevant evidence comes from five families of test. Which ones matter, and how deeply, follows from the system's risk tier.
1. Accuracy, grounding and task completion
Does the system produce correct, supported output on representative work? Measure task success against a curated set with known answers, unsupported-claim rate, citation or source fidelity where the system quotes documents, and refusal appropriateness. Include the awkward cases: ambiguous requests, out-of-scope questions, and inputs the system should decline.
2. Safety and guardrail resilience
Can the system be pushed outside its intended behaviour? This is adversarial work: prompt injection through retrieved or user-supplied content, attempts to extract personal information or system instructions, jailbreaks toward prohibited advice, and unsafe tool or action invocation. Record the attack set, the success rate, and what was changed in response.
3. Fairness and disparate impact
Does the system perform materially worse for some affected groups? Compare outcome and error rates across the attributes that matter in context, test for skewed language and stereotyping, and check that the evaluation data represents the population the system actually serves. Document what was measured and, honestly, what could not be.
4. Human oversight and agreement
Where a person reviews or approves output, oversight is itself a control that can be tested. Measure agreement between system output and expert judgement, how often reviewers overturn the system, whether reviewers have time and information to intervene, and whether prolonged accuracy has quietly turned review into rubber-stamping.
5. Production drift and monitoring
Behaviour changes after launch: the vendor updates the model, the input mix shifts, the documents underneath change. Re-run a core evaluation set on a schedule, watch output distributions, latency and cost anomalies, escalation and complaint volumes, and set a threshold that triggers re-review rather than a dashboard nobody reads.
Designing an evaluation that counts as evidence
Two failures are common and both are avoidable. The first is choosing the threshold after seeing the result, which turns measurement into justification. The second is evaluating a version that is not the version in production, which makes the evidence technically accurate and practically useless.
| Element | Why it matters |
|---|---|
| Stated purpose | Names the control or risk the evaluation is evidence for |
| Dataset and provenance | Shows the test resembles real work and can be re-run |
| Method | Lets an independent reviewer judge whether the measurement is sound |
| Threshold set in advance | Prevents the result from deciding what counts as acceptable |
| Result and interpretation | Records both the number and what was concluded from it |
| Date, owner and version | Ties the finding to a point in time and a specific system build |
| Re-run point | Stops the evidence from silently expiring |
How often to evaluate
Cadence should follow risk tier and change rate, and re-evaluation should be triggered by events as well as by the calendar.
Before a readiness decision
The full applicable evaluation set, on the build intended for production.
On material change
A model or version update, a new data source, a change in scope or user population, or a new tool the system can invoke.
On a scheduled cycle
A reduced core set at a frequency set by the risk tier, so higher-consequence systems are measured more often.
After an incident or pattern of complaints
Re-run the relevant discipline and record the outcome against the original finding.
How evaluations become assurance evidence
An evaluation result on its own is a technical artefact. It becomes assurance evidence when it is attached to the control obligation it satisfies, dated, owned, judged sufficient by an independent reviewer, and carried into the readiness decision.
In practice this means testing and governance cannot be run as separate exercises. Evaluation work that never reaches the assurance record leaves the record weaker than the engineering behind it — which is the single most common gap we see in otherwise capable organisations.
See the indicative risk tier and control obligations for one of your AI systems.
The free AI Defensibility Check takes a few minutes and produces an executive summary you can share.
