Johnny Butler

July 19, 2026

The Capability Is Real. The Discipline Is Sold Separately.

The same demo is everywhere right now: an agent running for hours, untouched, shipping a whole feature on its own. Nobody is demoing the discipline. The hype is outpacing it: the loop is autonomous, the verification isn't.

Look at the conditions the demo runs under. The token budget is effectively unlimited, of course it is, the vendor sells the tokens. The repo is green-field or hand-picked. There's no compliance, no on-call rota, nobody maintaining the output in six months, and the runs that went sideways didn't make the keynote. That isn't dishonesty. A demo's job is to show the ceiling.

But your delivery doesn't run at the ceiling. It runs on a budget somebody signs off, in a codebase with history, at a defect rate someone is accountable for. Marc Brooker, a VP at AWS, put it plainly recently: the opportunity for agents is limited by the defect rate. And in developer surveys, teams have started describing their AI tooling budgets as unsustainable. The demo doesn't have to live with the code. You do.

We've been here before. Every generation of tooling had a fifteen-minute demo that never survived contact with a real backlog. The lesson was never "the tool is fake". It was that demos optimise for the moment of creation, and production optimises for everything after.

So run the loops, they're coming either way. But give an agent exactly as much autonomy as your verification can actually check and your budget can actually carry: standards in front of it, checks it has to pass, evidence handed back, spend you can see. Extend the leash as the track record earns it.

The capability is real. The discipline is sold separately.