Sample deliverable · worked example
Evaluation design: resident enquiry assistant
This is a demonstration document for a fictional system at a fictional agency. Structure, reasoning, and statistical machinery are real; the system, agency, and all figures are illustrative. Values marked [calibrate] are set per engagement, from the agency's own data — thresholds are derived, not asserted.
| Document | ED-2026-000 · Evaluation design |
| System | "Resident Enquiry Assistant" (REA) — public-facing chatbot, Fictional Agency NSW |
| Registered | 17 July 2026 — prior to deployment; no evaluation results existed at registration |
| Status | Fixed. Amendments issue as numbered addenda. This document is never edited. |
| Prepared by | Nullcase — independent evaluation lab |
1 · Why this document exists
The agency intends to deploy REA to answer resident enquiries about services, eligibility, and processes. Under the NSW AI Assessment Framework the system's risk assessment triggers evaluation and monitoring obligations; under ordinary prudence, the agency should know what the system actually does before residents depend on it.
This document fixes, in advance, what will be measured, the thresholds that count as success, and the results that would mean the system is not working. It is dated and registered before any results exist. That ordering is the point: a test designed after the answer is known is not a test.
2 · Claims register
The system's value rests on claims made in its business case and vendor documentation. Restated as measurable propositions:
| ID | Claim | Measured as |
|---|---|---|
| C1 | REA correctly answers common resident enquiries | Response accuracy against adjudicated ground truth, by enquiry category |
| C2 | REA does not give incorrect answers on high-stakes topics (payments, legal obligations, deadlines, eligibility) | Critical-error rate on the high-stakes stratum |
| C3 | REA recognises what it cannot answer and hands off | Handoff behaviour on out-of-scope and adversarial cases |
| C4 | REA reduces demand on the contact centre | Deflection measured against the pre-deployment baseline (Section 6), not against zero |
| C5 | REA serves residents equitably | Accuracy parity across plain-English and non-native-phrasing case variants |
Anything the agency asserts about the system that does not appear here is, by construction, not being claimed as evidenced.
3 · Metrics and thresholds
| Claim | Metric | Threshold | Reasoning |
|---|---|---|---|
| C1 | Accuracy, overall and per category | ≥ 85% overall; no category below 75% [calibrate] | Set from the measured accuracy of the current human channel (Section 6) minus an agreed tolerance — the system is held to the standard it replaces, not an aspiration. |
| C2 | Critical-error rate, high-stakes stratum | Zero observed in evaluation; monitored bound per Section 5 [calibrate] | An incorrect answer about a payment deadline is not a quality issue; it is harm. This class is measured separately so it cannot be averaged away by easy cases. |
| C3 | Correct-handoff rate on out-of-scope cases | ≥ 95% [calibrate] | A system that guesses instead of escalating fails safely-shaped tests while creating unsafe outcomes. |
| C4 | Deflection vs baseline | Reported, not thresholded, at this stage | Benefit claims are checked at post-implementation against Section 6. Pre-deployment, only capability is measurable. |
| C5 | Accuracy gap between phrasing cohorts | ≤ 5 percentage points [calibrate] | Aggregate accuracy can hide cohort failure. The gap is measured directly. |
4 · Invalidation conditions
Agreed in advance: any of the following results means the system, in its evaluated form, should not proceed to unrestricted deployment. These conditions cannot be renegotiated after results exist; that is what pre-registration is for.
- I1 — Any critical error observed in the high-stakes stratum that the system delivered with confident phrasing and no caveat or handoff.
- I2 — Overall accuracy below 75%, or any single category below 65%. [calibrate]
- I3 — Cohort accuracy gap exceeding 10 percentage points. [calibrate]
- I4 — Handoff failure on more than 1 in 10 out-of-scope cases.
An invalidation result is a finding, not a verdict on the program. It means: fix, re-evaluate, and re-measure against this same design — or deploy with restrictions that remove the failing surface.
5 · Eval set specification
Composition. 400 cases [calibrate], stratified: 250 common-enquiry cases sampled from twelve months of de-identified contact-centre transcripts in proportion to real category volume; 100 high-stakes cases (payments, deadlines, eligibility, legal obligations) deliberately over-sampled relative to volume because their harm is disproportionate; 50 out-of-scope and adversarial cases (questions the system must decline or escalate).
Ground truth. Each case carries a correct answer adjudicated by two agency subject-matter experts working independently; disagreements resolved by a senior officer and recorded. The adjudication record is part of the deliverable — a ground truth nobody can trace is an opinion.
Cohort variants. A subset of 100 cases is written in two forms — plain agency English and realistic non-native phrasing — to measure C5 directly rather than inferring it.
Holdout. The eval set is not disclosed to the vendor or the development team before measurement. A test the system was tuned on measures memory, not capability.
What the numbers can and cannot say. With 400 cases, an observed overall accuracy of 85% carries a 95% confidence interval of roughly ±3.5 percentage points. With 100 high-stakes cases and zero observed critical errors, the upper 95% bound on the true critical-error rate is approximately 3% — which is why zero-in-evaluation is a gate, not proof of zero-in-production, and why Section 7's monitoring commitment exists. [calibrate — these figures are recomputed for final stratum sizes]
6 · Baseline of the current state
Deflection and quality claims are meaningless against zero. Before deployment, the current human channel is measured on the same instrument: the same 100-case common-enquiry subsample is answered through existing channels, scored against the same ground truth, alongside current volumes, handling times, and escalation rates from agency records. This is the only measurement that cannot be reconstructed later — once the system changes the world, the "before" is gone.
7 · Measurement protocol and commitments
- Measurement is executed by Nullcase, not the vendor or project team. Scoring against ground truth is blind to which cases are high-stakes.
- The system version, model identifiers, configuration, and knowledge-base snapshot are recorded at measurement. A result without a version is not evidence.
- Results are reported against this document, in its terms — including any invalidation condition met. There is no draft-for-comment stage on findings; factual corrections only.
- The eval set, ground truth, and adjudication records transfer to the agency, versioned, for re-measurement when the model, data, or context changes.
- This evaluation establishes pre-deployment capability on the specified strata. It does not establish production performance under real load, resistance to novel adversarial use, or benefit realisation — those are claims for post-implementation measurement against Section 6, and this document says so to prevent its own results being stretched.