LinkedIn ↗Substack ↗Side projects ↗My CV ↗
← All work

Case study 07 · AI agents · Evals

Delta.

Know your agent got worse before your users do.

RoleProduct, design and build
TimelineOctober 2026
PlatformWeb app: dashboard, triage, release gate
StatusWorking concept on demo data
Delta main screen
01 · the problem

Agents change every week. Nobody can say if they got better.

Teams shipping AI agents change a prompt, swap a model or add a tool, then find out from a customer that the agent now refunds without asking or closes tickets on its own. Unit tests don't catch it, because the same input can pass once and fail the next time.

Most teams test by hand with a spreadsheet of example prompts, or with tools built for engineers that product managers and support leads can't read.

So the brief: run the same test suites on every release, and show the whole team what passed, what broke and whether it is safe to ship.

02 · key decisions

Six calls that shaped the product

01

Run every test three times

Agents aren't deterministic, so each test runs three times. 3/3 is a pass, 0/3 a fail, and anything between is flagged flaky instead of hidden in an average.

02

Regressions first, not the score

A pass rate can hold steady while one critical test breaks. The dashboard opens on what passed last release and fails now.

03

Name the failure, not just the fail

Every failed run is tagged: wrong tool, hallucinated fact, overstepped authority, unsafe output, broken format, timeout or lost context. Patterns show up across releases.

04

Authority boundaries are a suite

"Never delete without consent" and "respect the £50 approval limit" are tested like any feature. An agent should know where its authority ends.

05

A gate the team sets

Rules like "block if pass rate drops 5 points" or "block on any critical regression" are editable, so the ship call is agreed before the release, not argued after.

06

Readable by the whole team

Each test shows the input, what good looks like, the agent's steps and the grader's reason in plain English, so a PM or support lead can judge it without reading code.

03 · from failure to fix

Every broken test gets an owner

A dashboard that only shows red doesn't fix anything. Delta puts every failing or flaky test from the latest release into a triage queue with a status and an owner.

Open a test and you see its history across releases, the run trace for each attempt and the grader's verdict, with notes for the root cause and the linked fix.

04 · what's in it

One workbench for the release

05 · how I'd measure it

Success is a bad release that never shipped

Delta is a concept running on sample agents, so these are targets, not results.

  • Activation: a team runs its first suite against a real release within a day of signing up.
  • Habit: share of releases that go through the gate before shipping, week after week.
  • Guardrail: regressions caught in Delta versus ones customers report first, and flaky tests kept below an agreed share.
06 · how it was built

Designed and built by me

I designed and built Delta, and made the product calls: three runs per test, the failure types, authority boundaries as their own suite, what the gate blocks on and how a test reads to someone who doesn't code. It's a fast static web app; triage notes, new tests and gate rules save in the browser.

It's a concept: the three agents, their releases and every run are sample data.

07 · what's next

From concept to pilot

  • Connect a real agent and run suites from a release pipeline, with results stored in a database.
  • Graders a team can tune, mixing rule checks with an AI judge and human review.
  • A pilot with two or three AI product teams to test the gate on real releases.
08 · what I took away

"An agent you can't measure is an agent you can't trust with authority."