Teams shipping AI agents change a prompt, swap a model or add a tool, then find out from a customer that the agent now refunds without asking or closes tickets on its own. Unit tests don't catch it, because the same input can pass once and fail the next time.
Most teams test by hand with a spreadsheet of example prompts, or with tools built for engineers that product managers and support leads can't read.
So the brief: run the same test suites on every release, and show the whole team what passed, what broke and whether it is safe to ship.
Agents aren't deterministic, so each test runs three times. 3/3 is a pass, 0/3 a fail, and anything between is flagged flaky instead of hidden in an average.
A pass rate can hold steady while one critical test breaks. The dashboard opens on what passed last release and fails now.
Every failed run is tagged: wrong tool, hallucinated fact, overstepped authority, unsafe output, broken format, timeout or lost context. Patterns show up across releases.
"Never delete without consent" and "respect the £50 approval limit" are tested like any feature. An agent should know where its authority ends.
Rules like "block if pass rate drops 5 points" or "block on any critical regression" are editable, so the ship call is agreed before the release, not argued after.
Each test shows the input, what good looks like, the agent's steps and the grader's reason in plain English, so a PM or support lead can judge it without reading code.
A dashboard that only shows red doesn't fix anything. Delta puts every failing or flaky test from the latest release into a triage queue with a status and an owner.
Open a test and you see its history across releases, the run trace for each attempt and the grader's verdict, with notes for the root cause and the linked fix.



Delta is a concept running on sample agents, so these are targets, not results.
I designed and built Delta, and made the product calls: three runs per test, the failure types, authority boundaries as their own suite, what the gate blocks on and how a test reads to someone who doesn't code. It's a fast static web app; triage notes, new tests and gate rules save in the browser.
It's a concept: the three agents, their releases and every run are sample data.
"An agent you can't measure is an agent you can't trust with authority."