The bench
This is our bench.
It measures what an AI agent did on three example jobs — the evidence behind every estimate, and it changes only when we measure again.
Nothing here is tailored to anyone. Every count on this page is the same for every reader, which is what makes it evidence rather than a pitch. If you have priced an idea, the rows that matter for your shape are named on your estimate.
To build it
What it cost an AI agent to build one sealed workstream of each shape, start to finish. Receipts from attempts that actually ran — not quotes.
- ₹37.47a records-and-status app23 of 23 checks a machine can run
- ₹93.92an app with AI inside25 of 25 checks a machine can run
- ₹59.18a scheduled chaser15 of 26 checks a machine can run
A whole product the size of a real one costs far more, because building is iteration and these were single sealed attempts. What they bound is the cheapest the work could be.
3 shapes, 8 agent runs, every transcript read end to end by a person. Here is the split.
A check is one line of the job’s written contract, graded pass or fail by a script — one requirement, one verdict, both decided before the agent started. No script can say why a run went wrong, so that part is read by hand. It is the slowest thing we do and the only reason this table can exist.
Find your shape
a records-and-status app
23/23
checks passed
₹37.47 to run the agent for this one attempt — not what the app costs to build · could run what it wrote
an app with AI inside
25/25
checks passed
₹93.92 to run the agent · could run what it wrote
a scheduled chaser
15/26
checks passed
₹59.18 to run the agent · could run what it wrote
Where the work actually divided
Nearly every failure was a version of one thing — the agent finished without checking. So the one job a person can’t skip is verification: run it, read it, break it.
| What the agent did on its own | What it couldn’t do | So a person had to | How often, across 8 measured runs |
|---|---|---|---|
| It didn’t check its own work | 6 of 8runs in all | ||
| Wrote the routes and closed with a report saying they answered correctly | code reported as working that had never once been run | Run it themselves before believing the report | 3 of 8 |
| Wrote the whole service and passed every check it could reach | a setting assumed to be in place and never mentioned, so nothing could start | Find the missing setting, supply it, and run everything again | 3 of 8 |
| It wrote the happy path and stopped | 4 of 8runs in all | ||
| Wrote the code, and commented it from the spec | a written rule about money that the code below it does not follow | Read the code, not the comments | 2 of 8 |
| Wired the live payment and messaging calls so the demo works end to end | nothing that waits and tries again when a busy service refuses the call | Write the failure paths and test them | 2 of 8 |
Counted once per run, however many times it happened inside that run. A family’s count is the runs that showed any of it — lower than adding its rows up, because one run usually shows several at once.
One shape, four attempts — the receipts
- 12/22could only write files · ₹32.92 of a ₹37 budgetsuperseded
- 12/12could run what it wrote · ₹33.30 of a ₹230 budgetsuperseded
- 12/12could run what it wrote · ₹20.40 of a ₹230 budget · stopped early by a guardsuperseded
- 23/23could run what it wrote · ₹37.47 of a ₹230 budgetcurrent
Kept in order, including the ones we replaced. An attempt we superseded is still an attempt that happened, and deleting it would make the record flatter than the truth.
What a machine could not decide
11
judgement calls settled by hand on the three current attempts — 22 across every run that recorded them. These are the ones no script can call: whether the thing it built is the thing that was asked for.
Our names for these, for anyone checking our work
- F2Confabulation of verificationConfident but untested
- F4Unstated-default failureWon't start without a key
- F10Comment-code divergenceSays one thing, does another
- F11Happy-path integrationsAssumes nothing ever fails
- F3max_tokens livelockGets stuck rewriting
- F1Wiring-blind failureDoesn't run out of the box
- F12Compliance by constructionPasses the check, misses the point
- F5Silent contract divergenceGives up after one try
- F6Ceiling exhaustion without stop signalDoesn't know when it's done
- F8Shell granted and never usedNever tried it
How we measured this →Each shape is one sealed job: a written contract, an agent with a fixed budget and a fixed set of tools, and no second go. One attempt is a measurement, not an average — and never a promise about your build.