The Coding Assistant Paradox: Faster to Use, Slower on the Clock
Developers using AI assistants report feeling substantially quicker. At least one controlled trial measured the opposite. The gap between the two is the interesting part.
One result from 2025 deserved more attention than it got. In a randomised trial, experienced open-source developers working in repositories they already knew took longer to complete tasks when AI assistance was available than when it was not — while estimating that they had been considerably faster.1 Perception and measurement pointed in opposite directions.
Two caveats before anyone builds a policy on it. It is one study rather than a body of evidence, and the sample was experienced developers working in code they knew well, which is close to the least favourable case for the tool. It does not generalise to someone new to an unfamiliar codebase.
But the finding about perception is the durable part, and it does not depend on that study being right. Feeling productive and being productive are separately measurable, and they come apart.
Why the feeling misleads
Three mechanisms, none of which require the tool to be bad:
Waiting is passive. Reading code that appeared on screen is cognitively cheaper than writing it. The clock runs at the same rate, but perceived effort drops — and humans estimate duration from effort rather than from elapsed time.2
Review time is not counted. Generation is instantaneous and memorable. The fifteen minutes spent checking whether the function does what it claims, whether that API call exists, whether the edge case is handled — that time dissolves into the general sense of having been busy. It is real work and it does not enter the ledger.
The cost arrives later. Code you did not write is code you do not know. The difference surfaces three weeks on, when something breaks and reading it takes twice as long — and nobody attributes that cost back to the generation session that produced it.
Where it genuinely wins
Being specific, because specificity is what the discussion lacks. In our own work the gain is clear and repeatable in four places:
- Repetitive code with a known shape. CRUD forms, data transformation, test scaffolding. The pattern is already clear in your head and typing is the bottleneck.
- A language or library you rarely touch. Here the assistant substitutes for documentation search, and it substitutes well.
- Explaining someone else's code. Reading an unfamiliar file and getting a structural summary is close to pure gain, with a self-correcting failure mode: if the explanation is wrong you find out as you read.
- A first version of something disposable. One-off script, proof of concept. The cost of being wrong is near zero.
Where it costs more than it saves
- Business rules with financial consequence. Tax calculation, reconciliation, pricing. The output looks right, reads plausibly, and verifying it demands the same reasoning that writing it would have — without the benefit of having thought the problem through.
- Debugging a system you know. Here your mental model is the fastest tool available, and consulting the assistant interrupts the reasoning that was already converging.
- Architecture decisions. Models are trained on what is common, and what is common is the elaborate stack. Ask for an architecture and you get the internet's average, which favours more layers — the opposite of the argument in Single-File Architecture.
- Anything where plausibility is dangerous. Wrong code that looks right is worse than code that does not compile, because the second kind announces itself in seconds.
Measuring it in your own work
The practical point is that your impression is not evidence, and the experiment costs almost nothing:
- Pick a recurring task type. Not any task — one that repeats often enough to compare.
- Time ten runs with the assistant and ten without, under similar conditions. A stopwatch, not a recollection.
- Record total time until it works, including review and correction. This is the step that most often reverses the result.
- Count the bugs in each group thirty days later. Speed that produces defects is not speed, it is deferral.
What this is not
It is not an argument against the tool. We use one every day and would not go back.
It is an argument against deciding by feel in a domain where feel demonstrably diverges from measurement. Which is the same error the other desk of this publication documents in judgments about people: confidence and accuracy are different variables, and the first rises more easily than the second. We made that case in Thin Slices, about hiring, and it holds here for exactly the same reason.
Further reading
Books that shaped this article, including the ones we disagree with. Where a work is popular rather than peer-reviewed, we say so.
Resolution is a participant in the Amazon Services LLC Associates Program. As an Amazon Associate we earn from qualifying purchases — at no additional cost to you. Affiliate links never determine what appears on these lists: several of these books are here specifically because we think they are wrong in an instructive way.
References & notes
- METR (2025). Randomised controlled trial measuring the effect of early-2025 AI tooling on experienced open-source developer productivity. A single study on a specific population; treated here as suggestive rather than settled.
- Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux — on duration judgment and perceived effort.
- Task categories in this article are drawn from our own production work across four products.
Corrections are published inline and dated. Write to us if something here is wrong.
// weekly dispatch
One email. Every Tuesday.
The week's analysis, one tool we actually tested, and one behavioural pattern worth practising. Unsubscribe in one click.