Green Is Not Evidence: Reviewing Code Your Agent Wrote
An agent can write a pull request faster than I can read one. Almost everything I have built around my coding agents exists to answer a single question: how do I know this change is right, when every report I am handed says it is?
The most expensive rework I see has nothing to do with bad code. It starts with a ticket that says one thing. I read it and understand a second thing. The agent I hand it to implements a third. Every step is competent, the tests pass, the pull request is tidy, and the result matches none of the three.
I spend most of my working day with coding agents now, on a production codebase that real customers depend on. The agents write nearly all of the code. That sentence still sounds like a boast to some people and a confession to others. In practice it is neither. It is a change in where the work sits. Producing a change has become cheap. Knowing whether it is right has not, and every report the agents hand me is written by the same process that produced the change.
So almost everything I have built around my agents is verification. None of it is clever. Each stage exists because a cheaper version of it failed, usually in a way that looked like success. This is what that machinery looks like, and the failures that earned each part of it.
The plan is the first test
Back to the three readings. The tempting diagnosis is that the agent misunderstood, and the tempting fix is a better prompt. Neither is right. An agent that meets an ambiguity does not stop and ask. It resolves it, silently, in the most plausible direction, and then builds three hundred lines on top of that choice. By the time you see the result, the misunderstanding has been compiled.
So every task starts with discovery before anyone writes code. Subagents read the ticket, the code it touches, my own notes on that part of the system, and the linked specification where there is one. The plan that comes out of it is written by the strongest model I have, and it has to contain two things beyond the implementation steps.
- The test scenarios. What a tester will do, in the running application, to show the change works. Written before the code exists, from the ticket rather than from the implementation. If we cannot write the scenario, we do not understand the task yet.
- The open questions. Every ambiguity the discovery turned up, listed rather than resolved. I answer the ones I can. The rest go to whoever actually owns the answer — the tech lead, the product owner, design. The model is not allowed to guess on their behalf.
That second list is the whole fix for the three readings. The plan is the one artefact where the ticket, my understanding and the agent's understanding sit side by side in writing, before any of them has cost anything. Rework found at this stage costs a conversation. Found in review, it costs the implementation.
For a feature that spans several tasks there is one more layer: a long-running coordinating session that owns the big picture. It tracks every sub-task, receives status from each implementation session, and watches for scope drift — the slow accumulation of "while I was in there" changes that individually make sense and collectively turn a feature into a rewrite. Its most useful job turns out to be re-planning. Sub-task three almost always discovers something that changes sub-task five, and someone has to notice. For a one-ticket change it is pure overhead, and I do not use it.
Separate the hands from the judgment
Once the plan is agreed, implementation is split into pieces and handed to several implementers running in parallel, on a cheaper, faster model. When they finish, the stronger model that wrote the plan reviews every diff against it and sends work back. Plan with the expensive model, type with the cheap one, review with the expensive one again.
The rule that makes this work is about what happens when the plan turns out to be wrong, which it sometimes does. An implementer that finds the plan does not hold stops and says so. It does not improvise. An improvised fix from a model that never saw the discovery is exactly how the third reading of the ticket gets into the code, now with the authority of having been "in the plan". The broken plan goes back to planning.
The lead review also catches a class of damage that is invisible in the change itself. My favourite example: a three-line edit to a test file that landed as a hundred and twenty-six changed lines. The agent's scripted edit had read a file with Windows line endings and written it back with Unix ones, so every line in the file changed. Nothing broke. The diff was simply unreviewable, and a real change was buried in it. The cheapest check in the whole pipeline is the first one: look at the size of the diff before you read it. If a three-line fix touches the whole file, something other than the fix happened.
Green is not evidence
This is the section the article is named for, because these are the failures that cost me the most before I learned to look for them. They have one thing in common: the tests were green, the run exited cleanly, and the result was wrong.
- The guard nobody could break. A new authorisation check, with tests for it. Every test ran as a super administrator, who passes every check before the new one is reached. Delete the guard and the suite stays green. The tests documented the feature; they did not test it.
- The fixture that already knew the answer. Production code looked up records by a key it derived from a URL. The test fixture created those records with the same derivation. Locally, every lookup matched. Against real data, most of them matched nothing, because the real records had been written over the years by several generations of code, each with its own idea of the key. The test proved the rule agreed with itself.
- The counter that could only say zero. A dry run reported "missed: 0", and I quoted it as the critical check passing. The dry-run branch returned before the code that increments the counter. Zero was not a measurement; it was the only value that path could produce.
- The stub on the wrong object. A test replaced a global, but the code under test had imported its own reference to the real module. Every assertion ran against something the code never touched.
- The command that half-ran. A seeding script piped into
headto keep the output short.headclosed the pipe after four lines, the script died after the first of thirty-two batches, and the pipeline exited zero. Only a row count showed it.
None of these is an agent problem in particular. Humans write every one of them. What the agent changes is the volume: it produces tests faster than anyone reads them, and a test that passes for the wrong reason looks exactly like a test that passes for the right one.
The defence is old and it is called mutation testing, though I mostly do it by hand. For every gate that matters, ask one question: if I invert this, does anything go red? Delete the guard, run the suite, expect a failure, restore it. Swap the stub's return value. Feed the fixture a key from real data instead of from the function under test. And before quoting any number as evidence, find the line that produces it and check that the path you ran can actually reach it. A zero from code that cannot produce anything else is not a result.
Test against the plan, not against the code
Unit tests written by the agent that wrote the code share its understanding of the task, including its misunderstanding. If it read the ticket wrong, its tests read the ticket wrong in the same way, and they agree with each other perfectly.
So the next stage does not look at the tests. QA agents take the scenarios that were written into the plan before any code existed, start the application locally, and drive it end to end in a real browser where the change allows it. They report back to the lead model, not to the implementers, and the lead decides what goes back for rework. Their independence is the point: they are testing what was asked for, not what was built.
Reviewers hallucinate too
When the pull request opens, two things start in parallel. A scheduled job watches CI and fixes failures until the build is green. And a separate review session reads the whole change for architectural and technical problems, with several reviewers on different models looking at it independently.
Their findings do not go straight to the implementer. They go to a judge — one more model whose only job is to check each finding against the code, one by one, and decide which are real. This stage exists because reviewers are wrong in the same confident register as implementers. A reviewer will report a missing null check on a value that cannot be null. Worse, it works in the other direction too. I once had an implementing agent decline a reviewer's suggestion on the grounds that "no shared test builder exists" — and I passed that reason on in the pull request description. The builder existed. The human reviewer found it and quoted my sentence back to me. A claim that something does not exist is still a claim, and it gets checked like any other.
The review repeats until it comes back clean, with a hard cap of five rounds. That cap is the same argument I made about agent loops: termination is a feature. Two models disagreeing about style will happily ping-pong forever, and hitting the cap tells me something useful — usually that the finding is a judgement call that belongs to a human.
"Fixed" is a claim, too
This one took me longest to see, because it looked like a cost problem rather than a correctness problem. A review would come back with findings. The agent would fix them, everything would go green, I would ask for a re-check — and the re-check would report the same findings again. Not similar ones. The same ones.
What was happening is that an agent handed a list of review comments tends to rush them. It makes a change near each finding, sees the tests pass, and reports the list as done. Some of those changes fixed the problem; some fixed a neighbouring line; some satisfied the letter of the comment and missed its point. And each wasted round costs an entire review, which on a large change is not cheap.
The fix is a ledger. Every finding is logged on its own row: what the finding says, whether we accept it, what was changed, and — written before the fix — how we will check that the fix worked. That last column is the one that matters. "Add the missing permission check" is a task. "A user without the grant gets a 403 from this endpoint, and the test proving it fails when the check is removed" is a verification plan. A finding is closed when its check has run, not when the agent says so.
Summaries are optimistic, sources are not
If I had to compress everything above into one habit, it would be this one, and I learned it the slow way. An agent's report said a project's CI tested three Python versions. It tested one, three times. My own verification once reported a hundred and eleven passing tests for a command that CI does not actually run. Neither report was a lie. Both were summaries, and summaries are optimistic in a way sources are not.
So before acting on anything that matters, go to the artefact. Read the CI configuration, not the description of it. Run the suite the way CI runs it. Check the effect — the row count, the file list, the HTTP response — rather than the exit code. It costs minutes, and in my experience it changes the conclusion more often than anyone would like.
Every mistake becomes a rule; the important rules become hooks
Every example in this article is now written down, because the system is only as good as its memory of how it failed. When something goes wrong — a correction from me, a failing build, a reviewer's comment — the lesson gets captured in the same session, as a short rule with the incident that produced it. Not later. A session that ends without writing the lesson down loses it permanently. It is the same discipline I argued for in evaluating LLM systems: every production incident becomes an eval case.
But rules written in prose have a known weakness: they get followed most of the time. As the context fills up, older instructions lose their grip, and "most of the time" is the worst reliability profile there is. So the rules that must hold every single time are not prose at all. They are hooks — code that runs before a tool call and can refuse it. An agent once merged a pull request into the main branch because I had floated the idea and then picked an option from a menu without realising what it would do. That rule is no longer a sentence in an instruction file. It is a hook that denies the command, whoever asks.
If that sounds familiar, it should. It is the argument from the agent loop article about approvals being the system's job rather than the model's, applied to the development process itself. If a prompt instruction reads like an access-control rule, it is in the wrong layer.
What is left for me
At the end of all of this I get a pull request that has been planned, built, reviewed against its plan, tested end to end, reviewed again by a panel and a judge, and taken to green. Then I review it myself, and I test it locally myself.
That is not ceremony. What has changed is what I read for. I no longer review syntax, naming or missing null checks; the pipeline is better at those than I am at the end of a long day. I review intent. Does this do what the ticket meant, as opposed to what it said? Is this the change I would have wanted in this part of the system in a year? Did anything get in that nobody asked for? Those are questions about judgement, and the pipeline exists to make sure my judgement is spent on them rather than on things a machine can check.
What it costs
All of this costs tokens and wall-clock time, and it would be dishonest to pretend otherwise. A full run on a mid-sized feature is a lot of model calls, most of them reviewing other model calls. For a one-line fix it is absurd, and I do not run it.
It is the same ladder I described in Before You Build an Agent, Try a Cron Job: climb only as far as the problem needs. A small, well-understood change gets a plan, an implementation and a review. A feature that spans several tasks and three people's understanding of a ticket gets everything. The question is the same one as before — what does a mistake here cost, and when will I find out about it?
None of this is new, either. Test plans written before code, separation of duties, independent QA, review by someone other than the author, mutation testing, checklists for findings — software engineering has had all of it for decades, mostly because humans have the same failure modes as agents, only slower. What has changed is the ratio. When producing a change takes minutes, the rest of the process is the work.
The agents write the code. My job is making sure that green means something.
I'm Tihomir Tomašević, a software architect with 17+ years in enterprise systems, currently leading development of an agentic AI platform. This piece is about building software with agents; the rest of my writing is about building them — when a problem actually needs one, how to structure the loop, how to get retrieval right, and how to debug them in production. Through T2 Software I take on selected consulting work on exactly these problems, including helping teams set up AI-assisted development they can trust. Get in touch or find me on LinkedIn.