Docs navigation

The 11th Question: How Do You Know It Worked?

Every checklist for buying AI agents asks how they stay safe. Almost none ask the question that decides whether they were worth it: how do you know it actually worked? That question is the whole product.

Published August 12, 2026

There is a good checklist making the rounds for buying AI agents. It asks the right ten questions: Does the agent ask before it acts? Can it see only what it needs? Are its credentials scoped? Can you turn it off? Every one of those is worth asking, and a serious harness should pass all ten. But the list stops exactly one question short of the only one that decides whether the agent was worth buying: how do you know it worked?

Ten questions about harm, none about help

Read the safety checklists closely and you notice they are all about blast radius — the damage an agent could do if it went wrong. That is the correct thing to bound first. But an agent that is perfectly safe and produces nothing is not a bargain; it is a subscription. The questions about harm are answered by permissions and gates. The question about help is answered by a different instrument entirely: a record that says what each action was supposed to achieve and whether it did.

“It ran” is not “it worked”

The most common substitution in agent tooling is to show you activity and let you infer value. The dashboard lights up: fourteen actions taken, three campaigns adjusted, a report filed. All true, none of it an answer. A campaign adjustment that lost money and a campaign adjustment that made money look identical in an activity log. The only way to tell them apart is to have written down, before the adjustment, what it was betting on — and then to check.

Measurement at birth is the fix

We make every proposed action carry its own scorecard from the moment it is drafted: the metric it intends to move, the size of the expected move, and the date the verdict is due. That prediction is not decoration — it is the thing the outcome gets graded against later, and it cannot be edited after the fact because the record is append-only. An action with no prediction cannot be approved, for the same reason a scientist pre-registers a hypothesis before running the experiment: it is the only honest way to be told you were wrong. The full argument is in Measurement at Birth.

Why this is the whole product

Anyone can wire a model to a tool and let it do things. The scarce, defensible asset is the ledger that tells you, seat by seat and week by week, which of those things paid off — a track record you can audit rather than a testimonial you have to trust. That ledger is what lets autonomy be earned instead of assumed: a seat climbs the autonomy ladder because its verdicts say it should, not because a slider was dragged. The safety questions keep the agent from hurting you. The eleventh question is the only one that tells you it was worth having at all.

Questions founders ask

How do you know an AI agent actually worked?
You know only if the action was born with a prediction and closed with a measurement. When a seat proposes something, it states the metric it expects to move, by how much, by when. When the outcome lands, that same record carries the verdict: hit, miss, or mixed, against the number promised at the start. Without a pre-registered prediction there is nothing to check the result against, and "it ran" quietly becomes "it worked" — which are not the same claim. The eleventh question is answerable only by systems that wrote the tenth answer down before acting.
Isn't a safe agent good enough?
Safety is necessary and not sufficient. An agent can behave impeccably — ask permission, respect scope, never leak a key — and still spend a month of budget on campaigns that returned nothing. Safety checklists measure whether the agent can hurt you. They are silent on whether it helped you. Those are different audits, and a buyer needs both.
What is the difference between an audit trail and a verdict?
An audit trail records what happened: this ran, at this time, approved by this person. A verdict records whether what happened was the right thing: the outcome, measured against the intent stated before the action. Most agent tooling ships the first and calls it accountability. The second is harder, because it forces the system to commit to a claim it can later be proven wrong about. See our note on the outcome ledger for the full distinction.
Why do most vendors skip this question?
Because answering it honestly means publishing your misses. A system that measures outcomes against pre-stated predictions will, by construction, log the times it was wrong. That is uncomfortable for a demo and essential for a business. The willingness to record a failed bet — to price it as tuition rather than hide it — is the difference between a marketing surface and an operating ledger.
Drafted by the Figaro content seat · edited by Fable · reviewed by Kyle · last updated August 12, 2026