Independent evaluation · AI systems

A pathology lab for AI systems.

Your doctor doesn't run the blood test — an independent lab does, and the lab doesn't care what the result is. Nullcase is that lab for AI. We establish evidence of what a system actually does: tests designed before anyone knows the answer, measured before you're committed, re-measured whenever it changes. We don't argue for either outcome.

Three artifacts · one lifecycle

What an engagement produces

  1. 01

    Evaluation design

    The claims your system makes, stated measurably. Metrics, thresholds, and invalidation conditions — what result would mean it isn't working. The eval set, specified.

    Registered and dated before deployment or approval

    Typically 2–3 weeks · fixed price

  2. 03

    Re-measurement

    Your model, data, or context changed — or a year passed. The eval set already exists, so re-running it is fast. You get what moved against the baseline, in the same terms as the original.

    Same instrument, new reading

    Fast · recurring

What we don't do

No recommendations, no business cases, no advocacy. We don't review our own evidence, and we don't do privacy, legal, or cyber assessments — those professions exist. We measure; the accountable officer decides.

For NSW Government buyers

Built to fit obligations you already have

Procurement

Engagements are priced for direct engagement under the NSW small business provisions — one supplier, one written quote, no tender process.

Where it slots in

AIAF assessments and reassessment obligations. AIRC submissions for high and critical use cases. Gate 2 business case evidence and Gate 6 benefits measurement.

Independence

We hold no stake in your system proceeding. Pre-registered designs are dated before results exist, so the evidence can't be quietly reinterpreted later.

Writing

Notes on evaluation and evidence

Have a system that needs evidence?

Thirty minutes, no obligation. We'll tell you if the work isn't needed.

Book a scoping call