The hardest thing on our roadmap was document ingestion. Buyers upload tender packs in whatever format the authority issued them: a Word .docx, a scanned PDF, a spreadsheet of line items in Excel. Each has to be parsed, split into chunks, embedded, and made searchable, all behind a job queue with real retries because parsing is slow and fails in a dozen ways. This was not a demo. It was the piece everything else in the product leaned on.
I described the design to an agent over a couple of days: a queue, a parser per format, an embedding step, idempotent retries. By the end of the week it worked end to end. Genuinely hard code, not boilerplate, in less time than I'd scoped. My first reaction was relief. My second, once I sat with it, was the useful one: this was one of the more load-bearing things we'd shipped, and I hadn't written a line of it myself.
Here's the trap. The cost of producing that code dropped sharply. The cost of owning it (reading it, trusting it, debugging it a month from now, onboarding a teammate onto it) barely moved.
A demo ends at "it works." Production begins there and runs for years. Building a demo with AI and building software you'll maintain with AI are two different jobs, and the speed makes them look identical until the second bill comes due.
For most of my career the constraining question was can we build this in time? Estimation, scope-cutting, and sprint planning all exist to answer it. Agents take that question off the table for a lot of everyday work. When a feature costs an afternoon instead of a fortnight, "can we?" stops being the interesting part.
What replaces it is harder and less fun: should this exist at all? Every line we generate is a line someone has to understand, secure, migrate, and eventually delete. Cheap production doesn't lower that ongoing tax; it just lets you sign up for it faster. The scarce resource is no longer typing speed. It's the judgment to decide what deserves to be in the codebase in the first place.
A month after the ingestion pipeline shipped, I got cocky. A ticket asked for a small thing: a way to re-trigger a parse job that had failed. Instead of the one button that would have done it, I let an agent build a whole "operations dashboard" on top of the queue: bulk re-runs, filters, an audit log, the works. It looked useful. It was also a couple thousand lines of surface area nobody had asked to own.
Because it was so cheap to make, I skipped the question I'd have asked if it cost two weeks: do we actually want to maintain this forever? The answer was no. We ripped most of it out. Cheap to write, expensive to carry.
Left unsupervised, agents don't produce dramatic disasters. They produce noise: pull requests that touch forty files to fix one bug, a helper that already existed under a different name, a config option nobody will ever set. Individually each is defensible. In aggregate they bury the signal.
The most dangerous artifact is the plausible one: code and docs that read as correct, pass a skim, and are subtly wrong in exactly the place you didn't look.
While building the pipeline, an agent updated our README to describe its retry and backoff policy. The prose was confident and well-formatted. It also described a backoff curve the code didn't implement. Six weeks later someone tuned a queue timeout based on that README and got burned. Generated documentation fails in a uniquely nasty way: it's fluent enough to be trusted and detailed enough to be wrong.
Read enough agent-authored pull requests and a pattern shows up. The description tells you what changed in detail: "added a .docx parser, wired it into the queue, updated the embedding step." What it rarely tells you is why: which lost bid prompted this, which older parsing decision it quietly reverses, what we tried before that didn't work.
That "why" isn't in the repo. It's in your head, in a support thread, in a hallway conversation from two years ago. An agent can't reconstruct context it was never given, and the context is the expensive part.
Here's the honest part. I didn't hand-write that pipeline, and I don't think I should have. The agent's version was at least as good as what I'd have typed against a deadline. What changed wasn't who wrote it, it was how hard I looked. Because so much downstream depends on a tender being parsed correctly and never silently dropped, I reviewed that code far more thoroughly than I review a feature that can't hurt anyone. I traced the retry paths and each parser's failure modes: what a corrupt PDF does, whether an Excel with merged cells loses rows, whether a re-queued job embeds the same document twice.
One rule shows why the depth mattered. A scanned PDF with no extractable text has to be flagged for a human, never indexed as empty, because a single missed requirement in a tender can sink an entire bid. To the agent that check looked redundant. To us it was the whole point. You only catch that when you review in proportion to what the code can break, not to how long it took to write.
Agents accelerate execution. Humans supply direction, context, the judgment of how hard to look, and the call on whether the output is good enough to ship. The work didn't shrink; it moved up a level, from writing lines to setting intent, weighing tradeoffs, and reviewing in proportion to what a change can break.
I went looking for the magic configuration that makes a codebase good for agents. There isn't one. The things that make a repo pleasant for an agent are the exact things that make it pleasant for a new hire: clear structure, one obvious way to run each task, tests you can trust, docs that match reality, consistent patterns. The difference is that a human tolerates a messy repo and an agent amplifies it: point agents at spaghetti and you get more spaghetti, faster.
AGENTS.md. The house rules, written down once, so you're not re-explaining them every session.# How to work in this repo ## Commands - Install: pnpm install - Test: pnpm test (must pass before any PR) - Verify: pnpm verify:e2e (drives the real flow, not just units) - Lint: pnpm lint --fix ## Conventions - Money is stored in integer minor units. Never a float. - Every new endpoint goes through the auth middleware. No exceptions. - New behavior ships with a test that would fail before the change. ## Explain your reasoning - In the PR body, write WHY, not just what. Link the ticket. - If you touched the ingestion pipeline, embeddings, or auth: stop and flag for a human. ## Do not - Invent config options nobody asked for. - Update the docs to describe behavior the code does not have.
None of this is exotic. It's the checklist a good team already half-follows. Agents just remove the option of half-following it. The gaps that a patient human papered over become the potholes the agent drives straight into.
My first instinct was to bolt agents onto the workflow I already had: same planning, same tickets, same review dance, just with a faster author in the middle. That gets you a modest bump and a review queue that backs up. The whole cycle was designed around code being the expensive step. Remove that assumption and every other step is in the wrong shape.
The real gains show up only when you rethink the loop end to end: how you plan, how you generate, how you test, how you review, how you decide something is done.
We changed how we plan. Instead of a backlog of a hundred small issues, we now keep a short list of the user experiences that have to be true ("every uploaded tender is parsed and searchable within a minute," "no document is ever silently dropped from the index") and score how close we are to each.
When building is cheap, a pile of tickets is just a to-do list for a machine. The valuable artifact is the shared judgment about which outcomes matter and how we'll know we got there.
The teammates most resistant to this shift were my best engineers, and it took me a while to understand why. They've spent a decade building workflows that genuinely work. Telling them the hard thing is "easy now" lands as an insult, not a gift. And for a while they really are slower, because decomposing a task for an agent, feeding it context, and validating its output is a new skill on top of the ones they already mastered.
That discomfort is real and temporary. The move isn't to shame anyone into adopting the tools; it's to let people trade up gradually and to respect that "progress feels like regression" is an honest description of the first few weeks, not stubbornness.
The tools got much better at producing code. They did little for the obligation to understand, explain, and stand behind what ships, and that obligation is where the cost, and the value, actually lives. Treat generation as cheap and ownership as the scarce thing it is: decide what deserves to exist, keep the "why" human, make the repo trustworthy enough for an agent to work in, and redesign the loop instead of just flooring the old one. The question six months from now was never "could it write this?" It's "who understands this?" Keep the answer a person.