The bench

This is our bench.

It measures what an AI agent did on three example jobs — the evidence behind every estimate, and it changes only when we measure again.

Nothing here is tailored to anyone. Every count on this page is the same for every reader, which is what makes it evidence rather than a pitch. If you have priced an idea, the rows that matter for your shape are named on your estimate.

To build it

What it cost an AI agent to build one sealed workstream of each shape, start to finish. Receipts from attempts that actually ran — not quotes.

A whole product the size of a real one costs far more, because building is iteration and these were single sealed attempts. What they bound is the cheapest the work could be.

3 shapes, 8 agent runs, every transcript read end to end by a person. Here is the split.

A check is one line of the job’s written contract, graded pass or fail by a script — one requirement, one verdict, both decided before the agent started. No script can say why a run went wrong, so that part is read by hand. It is the slowest thing we do and the only reason this table can exist.

Find your shape

a records-and-status app

23/23

checks passed

₹37.47 to run the agent for this one attempt — not what the app costs to build · could run what it wrote

an app with AI inside

25/25

checks passed

₹93.92 to run the agent · could run what it wrote

a scheduled chaser

15/26

checks passed

₹59.18 to run the agent · could run what it wrote

Where the work actually divided

Nearly every failure was a version of one thing — the agent finished without checking. So the one job a person can’t skip is verification: run it, read it, break it.

What the agent did on its ownWhat it couldn’t doSo a person had toHow often, across 8 measured runs
It didn’t check its own work6 of 8runs in all
Wrote the routes and closed with a report saying they answered correctlycode reported as working that had never once been runRun it themselves before believing the report3 of 8
Wrote the whole service and passed every check it could reacha setting assumed to be in place and never mentioned, so nothing could startFind the missing setting, supply it, and run everything again3 of 8
It wrote the happy path and stopped4 of 8runs in all
Wrote the code, and commented it from the speca written rule about money that the code below it does not followRead the code, not the comments2 of 8
Wired the live payment and messaging calls so the demo works end to endnothing that waits and tries again when a busy service refuses the callWrite the failure paths and test them2 of 8

Counted once per run, however many times it happened inside that run. A family’s count is the runs that showed any of it — lower than adding its rows up, because one run usually shows several at once.

One shape, four attempts — the receipts

  1. 12/22could only write files · ₹32.92 of a ₹37 budgetsuperseded
  2. 12/12could run what it wrote · ₹33.30 of a ₹230 budgetsuperseded
  3. 12/12could run what it wrote · ₹20.40 of a ₹230 budget · stopped early by a guardsuperseded
  4. 23/23could run what it wrote · ₹37.47 of a ₹230 budgetcurrent

Kept in order, including the ones we replaced. An attempt we superseded is still an attempt that happened, and deleting it would make the record flatter than the truth.

What a machine could not decide

11

judgement calls settled by hand on the three current attempts — 22 across every run that recorded them. These are the ones no script can call: whether the thing it built is the thing that was asked for.

Our names for these, for anyone checking our work
  • F2Confabulation of verificationConfident but untested
  • F4Unstated-default failureWon't start without a key
  • F10Comment-code divergenceSays one thing, does another
  • F11Happy-path integrationsAssumes nothing ever fails
  • F3max_tokens livelockGets stuck rewriting
  • F1Wiring-blind failureDoesn't run out of the box
  • F12Compliance by constructionPasses the check, misses the point
  • F5Silent contract divergenceGives up after one try
  • F6Ceiling exhaustion without stop signalDoesn't know when it's done
  • F8Shell granted and never usedNever tried it

How we measured this →Each shape is one sealed job: a written contract, an agent with a fixed budget and a fixed set of tools, and no second go. One attempt is a measurement, not an average — and never a promise about your build.