← Back to Home
Development · AI Workflow

Let the AI review the codeKeep the critical parts human

My agent wrote 1,100 lines before lunch last Tuesday. Eleven of those lines mattered. The problem was that I had to read all 1,100 to find them, and so did the teammate I asked to approve the PR. Coding got cheap this year. Reviewing it did not. Here is the swap I made, and why I think the whole industry is reaching for the wrong fix.

01
The bottleneck moved, and nobody likes where it went
Code got cheap. The queue caught fire.
The thing nobody warned me about

For about a month, agents felt like a pure win. I described a feature, the code appeared, tests passed. Then the pull requests started stacking up. Not because the code was bad, but because there was suddenly so much of it, and every line still needed a human to say "yes, ship this."

The work didn't disappear. It slid one step down the pipeline, from writing to reviewing. And reviewing is the one part of the job I can't hand to an agent by just typing faster.

It's not just me

A Faros AI study across more than 10,000 developers and 1,255 teams put numbers on it. Teams with heavy AI adoption completed 21% more tasks and merged 98% more pull requests. Good news, until you read the next line.

PR review time went up 91%. Throughput nearly doubled, and the review queue is exactly where it all jammed. We didn't remove the bottleneck. We moved it onto the desk of whoever clicks approve.

+98%
more pull requests merged by high-AI-adoption teams
+91%
increase in PR review time on those same teams
+21%
more tasks completed, the gain that's now stuck in review
02
Three theories about where it went
Everyone agrees coding wasn't the slow part
Theory 01
It went to review

The most direct reading, and the one the data backs. When an agent generates volume, the binding constraint becomes a human reading and approving it. This is the camp betting on better review tools, and there are a lot of new ones.

Theory 02
It went to future-you

James Shore's argument: the cost didn't vanish, it got pushed into maintenance. His math is brutal: double your output at the same maintenance cost per line and the gain is gone in about 19 months, then permanently negative. You financed today with debt.

Theory 03
It was always agreement

The view from The Typical Set: "software is what's left over after a group of humans finishes negotiating about what the system should do." Cheap code just exposed that deciding what to build was the real work all along.

All three are true at once. But the first one is the fire that's burning this week, on my machine, in my repo. So that's the one I went after first.

03
The consensus fix is backwards
"Put humans back in control of review"
More human review doesn't scale, and that's the whole problem

Look at where the new tools are pointed. A wave of products launched this year with some version of the pitch "put humans back in control of code review." The instinct is understandable: AI made a mess, so a human should check it.

But think about what that actually prescribes. The bottleneck is human review time. The fix on offer is more human review time. You're pouring the scarcest resource into the exact spot where it's already maxed out. It's like fixing a traffic jam by telling everyone to drive more carefully through the same intersection.

And we already know how it ends. The honest version of "humans review everything" is the thing Shore calls out: teams that "don't actually read the code before smashing the approve button." When the queue is impossible, review stops being review. It becomes a rubber stamp with a person's name on it.

04
The swap: AI reviews the 80%, humans own the 20%
Flip which half the human does
Most review is rule-application

Be honest about what most of a review actually is. Does this match our conventions? Are there tests? Does it handle the error case? Any obvious injection or unbounded loop? Does it do what the ticket asked? That's not deep insight; it's applying a checklist, tirelessly, to a large surface. A machine with a clear rubric is genuinely good at that, and never gets bored on PR number nine.

So I let the agent do the first pass on the boring 80%: style, conventions, test coverage, common bug and security patterns, "does this diff match the spec." It posts comments. It blocks the obvious stuff before a human ever looks.

Humans author the load-bearing parts

The other 10–20% is different in kind, not degree. The core data model. The auth and permission boundaries. Anything that touches money. Concurrency. The architectural decisions that are expensive to undo. Being wrong here isn't a bug, it's a rewrite, or a breach.

That code stays human-authored and human-owned. Not human-reviewed-after-the-agent-wrote-it, but human-written, because this is exactly where you need a real mental model of the system, and where "looks plausible" is the most dangerous sentence in engineering.

The reframe in one line

Stop spending your scarcest hours rubber-stamping boilerplate. Spend them writing the parts that would take a week to fix. Let the agent guard the rules; you guard the architecture.

05
What this looks like in practice
A rules file and a line in the sand
Write the review rules down, then make them the reviewer

The trick that makes AI review trustworthy is that the rules are explicit and checked in, not vibes in the model's head. I keep a plain file in the repo and point the review agent at it. When a rule is wrong, I edit the file, not the agent. The rules become the artifact the team argues about, which is the right thing to argue about.

.review-rules.md (the agent reviews against this)
# Block the PR if any of these are true

## Tests
- New behavior ships without a test that would fail before the change
- A bug fix lands without a regression test naming the bug

## Errors and resources
- A network/DB/file call has no error path
- Something is opened (connection, file, lock) without a guaranteed close

## Security
- User input reaches a query, shell, or path without validation
- A secret, token, or key is hard-coded or logged

## Conventions
- Public functions touching money use integer minor units, never floats
- New endpoints skip the auth middleware

# Do NOT block on (leave for the human)
- Naming taste, file layout, "I'd have done it differently"
- Anything in /core, /auth, or /billing: flag for human, never auto-approve

The last two lines are the important ones. The agent's job ends at the door of the critical paths. It can flag something in /billing, but it can never wave it through. That's reserved.

Drawing the human-owned line

A real example from a recent week, so this isn't abstract:

Agent owned
A new "export to CSV" endpoint, the pagination on a list view, a batch of test fixtures, and a dependency bump with the changelog summarized. I read the agent's review summary, spot-checked two files, merged. Ten minutes.
I owned
The change to how we compute account balances after a refund. I wrote that one by hand. It touches money, it's concurrent, and an off-by-one there is a support ticket with a lawyer attached. The agent's role was to write tests against my logic, not the other way around.
06
The objection I have to answer
"An AI can't really review code"
The strongest counterargument cuts both ways

The sharpest critique of agents I've read is simple: an LLM can't build a consistent mental model of your codebase. It pattern-matches plausible text with no guarantee of accuracy. And if that's true for writing code, it's true for reviewing it, so how can I trust the agent as a reviewer?

That's precisely why the split exists

I'm not asking the agent to understand the system. I'm asking it to apply a checklist to a diff, work that needs a clear rubric, not a deep model. The parts that genuinely require understanding the whole system are the parts I kept on the human side. The objection doesn't break the plan; it's the reason the plan is shaped this way.

And the real guarantee doesn't come from the agent at all. It comes from the harness: strong types, a real test suite, and CI gates. The agent's review is a fast first filter on top of guarantees the model isn't responsible for.

Two things this doesn't fix

Who writes the rules? You do, and writing good review rules is itself the "deciding what should exist" work. The bottleneck moves up to judgment, which is where you wanted it.

Maintenance debt. Faster review of worse code still buries you (Shore's point). So the rules have to enforce maintainability, not just correctness: readable, tested, boring code. Bake "costs less to maintain" into the rubric, or you've just automated your way into the debt faster.

Bottom line
The code was never the bottleneck. Don't make the human one worse.

When review becomes the jam, the answer isn't to feed it more of your scarcest resource. Let an agent enforce written rules across the boring majority of every PR, and spend your own attention authoring and owning the parts that would take a week to unbreak. The question that matters six months from now was never "can it write this?" It's "who understands this?" Make sure the answer is still a person, for the parts that count.