Guide · AI evaluations and testing

AI evaluations: measuring a system before it reaches production

Benchmark scores describe a model. Assurance needs evidence about your application of it — on your data, against your thresholds, with a result someone can re-run. This guide covers how to design evaluations that stand up as evidence.

9 minute readUpdated September 2026RAIReady assurance practice
In short
  • Public benchmarks test the model; assurance needs tests of your system in your context.
  • An evaluation is only evidence when the dataset, method, threshold, result, date and owner are recorded together.
  • Set the pass threshold before running the test, not after seeing the result.
  • Red-teaming is a structured, repeatable exercise with recorded findings, not an afternoon of curious prompting.
  • Pre-deployment evaluation is a starting point; behaviour drifts, so the cadence matters as much as the first result.

Why benchmarks are not evidence

A model's benchmark performance tells you something about general capability under laboratory conditions. It tells you nothing about whether your retrieval layer surfaces the right policy document, whether your prompt leaks customer data into a summary, or whether your system refuses appropriately when a user asks something it should not answer.

Assurance-grade evaluation measures the assembled system — model, prompts, retrieval, tools, guardrails, interface and the human process around it — on data that resembles the real workload, against thresholds decided in advance.

Five evaluation disciplines

Most assurance-relevant evidence comes from five families of test. Which ones matter, and how deeply, follows from the system's risk tier.

1. Accuracy, grounding and task completion

Does the system produce correct, supported output on representative work? Measure task success against a curated set with known answers, unsupported-claim rate, citation or source fidelity where the system quotes documents, and refusal appropriateness. Include the awkward cases: ambiguous requests, out-of-scope questions, and inputs the system should decline.

2. Safety and guardrail resilience

Can the system be pushed outside its intended behaviour? This is adversarial work: prompt injection through retrieved or user-supplied content, attempts to extract personal information or system instructions, jailbreaks toward prohibited advice, and unsafe tool or action invocation. Record the attack set, the success rate, and what was changed in response.

3. Fairness and disparate impact

Does the system perform materially worse for some affected groups? Compare outcome and error rates across the attributes that matter in context, test for skewed language and stereotyping, and check that the evaluation data represents the population the system actually serves. Document what was measured and, honestly, what could not be.

4. Human oversight and agreement

Where a person reviews or approves output, oversight is itself a control that can be tested. Measure agreement between system output and expert judgement, how often reviewers overturn the system, whether reviewers have time and information to intervene, and whether prolonged accuracy has quietly turned review into rubber-stamping.

5. Production drift and monitoring

Behaviour changes after launch: the vendor updates the model, the input mix shifts, the documents underneath change. Re-run a core evaluation set on a schedule, watch output distributions, latency and cost anomalies, escalation and complaint volumes, and set a threshold that triggers re-review rather than a dashboard nobody reads.

Designing an evaluation that counts as evidence

Two failures are common and both are avoidable. The first is choosing the threshold after seeing the result, which turns measurement into justification. The second is evaluating a version that is not the version in production, which makes the evidence technically accurate and practically useless.

ElementWhy it matters
Stated purposeNames the control or risk the evaluation is evidence for
Dataset and provenanceShows the test resembles real work and can be re-run
MethodLets an independent reviewer judge whether the measurement is sound
Threshold set in advancePrevents the result from deciding what counts as acceptable
Result and interpretationRecords both the number and what was concluded from it
Date, owner and versionTies the finding to a point in time and a specific system build
Re-run pointStops the evidence from silently expiring
What every recorded evaluation should carry

How often to evaluate

Cadence should follow risk tier and change rate, and re-evaluation should be triggered by events as well as by the calendar.

Before a readiness decision

The full applicable evaluation set, on the build intended for production.

On material change

A model or version update, a new data source, a change in scope or user population, or a new tool the system can invoke.

On a scheduled cycle

A reduced core set at a frequency set by the risk tier, so higher-consequence systems are measured more often.

After an incident or pattern of complaints

Re-run the relevant discipline and record the outcome against the original finding.

Put this into practice

See the indicative risk tier and control obligations for one of your AI systems.

The free AI Defensibility Check takes a few minutes and produces an executive summary you can share.

Download

AI evaluation protocol template

A template for teams designing evaluations of an AI system. Complete one protocol per evaluation, before running it, and keep the completed protocol with the result.