Articles
The AI Coding Trap: Why Faster Tasks Don’t Mean Faster Delivery

AI can accelerate code creation. The real question is whether the engineering system can absorb, verify, and release the change.
Why AI Coding Tools Aren't Improving Enterprise Productivity
An AI agent can produce a pull request before a human has finished reading the ticket. That looks like a breakthrough until the pull request sits in review, fails due to an undocumented data rule, triggers security rework, and misses the release window.
The code was generated faster. The change was not delivered faster.
That is the AI coding trap.
Coding assistants and engineering agents are creating real value. They can search repositories, explain unfamiliar code, generate implementation, write tests, run tools, and prepare pull requests. Assuming that faster code creation automatically means faster software delivery is a mistake.
Software delivery is a system. A change must carry business intent, architecture, application knowledge, security, quality, approvals, and operational evidence all the way into production. Accelerating one station in that system does not increase total throughput when the stations around it remain constrained.
The task accelerates. The bottleneck migrates.
Key Takeaways
- AI coding tools can accelerate individual engineering tasks without reducing end-to-end delivery lead time.
- When code enters the system faster than teams can review, test, secure, and release it, queues grow downstream.
- Four hidden taxes absorb much of the apparent gain: context, verification, coordination, and rework.
- Brownfield systems magnify the problem because critical knowledge lives beyond the repository.
- Engineering leaders should measure accepted production change, not generated code, prompts, licenses or activity.
- The way out is not another point tool. It is a connected engineering system that carries context, quality, accountability, and evidence with the work.
The Productivity Inversion
Most productivity conversations begin with the developer: How much faster can a person complete a coding task? That is useful, but incomplete.
An engineering system has multiple service points: requirements, architecture, implementation, review, testing, security, release, and operations. If AI increases the arrival rate of code while the service rate of review and release remains unchanged, work in progress accumulates. Queue time rises. The organization produces more output without necessarily delivering more value.
This is a predictable systems effect. It is not a failure of the model.
DORA describes one part of the phenomenon as the verification tax: time saved during initial creation is often re-spent auditing, correcting and hardening the result before it can be trusted.[1] The stronger the generator becomes, the more important the receiving system becomes.
The gap between perceived and measured productivity can also be significant. In a narrow 2025 randomized study, METR observed 16 experienced open-source developers completing 246 tasks in repositories they knew well. The developers expected AI to make them faster and still believed it had done so after the study. In that specific setting, task completion time increased by 19% when AI tools were allowed.[2]
That result should not be generalized to every engineer, tool or workflow. It is a point-in-time study of a particular population and task set. But its lesson is important: felt productivity, generated output and measured delivery performance are different things.
A developer can finish a task faster while the project still delivers later.

Four Hidden Taxes that Absorb the Gain
The time saved in code creation does not simply vanish. It often reappears somewhere else in the delivery system.
Each tax originates at a different stage in the delivery system, but the underlying mechanism remains the same: tasks that AI can perform without human input still require human oversight to catch errors, coordinate, or redo work.
The breakdown provided below helps identify each type of tax in your workflow before it impacts lead time.

Brownfield Systems Magnify the Trap
Greenfield demonstrations are useful because the assumptions can be made explicit. Most enterprise engineering does not happen on a blank canvas.
It happens inside applications shaped by years of business rules, integrations, data relationships, exceptions, and operational workarounds. Some knowledge is visible in the code. Some live in tickets, diagrams, database schemas, and runbooks. Some exists only in the experience of the people who keep the system running.
That is why plausible code is not the same as safe change.
A model can reproduce the structural patterns of a repository and still be wrong about the database schema in which a record actually lives; a protected field that must come from a vault rather than a table; a soft-delete convention that is not documented; one of several required access roles; a runtime backfill or fallback behavior; or the true testing depth needed for a cross-repository change.
Field Evidence: Where the Economics Changed
In a controlled Xebia field experiment, one real brownfield enhancement spanning three repositories was taken to production under a fixed quality standard. A well-engineered standalone five-agent coding pipeline was faster than the measured human baseline. But it required 11 rework iterations and consumed 348.6 million tokens.
When deterministic application intelligence and the Xebia Ace engineering system were added, the same production change required three rework iterations and 2.3 million tokens. Compared with the standalone agent pipeline, the integrated configuration was 8.5 times faster, approximately 10 times cheaper and used roughly 150 times fewer tokens.[3]
This is one measured case, not a universal benchmark. Results will vary with codebase, team and starting point.
The largest gains did not come from typing code faster. They came from collapsing specification effort and preventing review debt.

Accuracy decided the economics. The value was created by preventing rework, not by generating speed.
This is the brownfield lesson in one sentence: AI cannot safely change what the engineering system does not understand.
Measure Accepted Change, Not Generated Output
Organizations often measure what is easiest to count: licenses activated, prompts submitted, lines of code generated, pull requests opened, self-reported time saved or token consumption.
Those signals help explain adoption and activity. They do not prove that software delivery improved.
The more useful unit is accepted production change: work that satisfies the requirement, passes the quality bar, reaches production safely and produces the intended outcome.
You can run a quick self-check using the measures outlined below:

Measure work that safely reaches production, not work that was generated.
How Engineering Organizations Escape the Trap
The answer is to redesign the system around the new rate of execution, not slowing AI down.
Carry Context with the Work
Requirements, architecture decisions, application dependencies, data rules, standards, and accepted operational behavior should travel with the artifact. Teams should not have to reconstruct the same context at every stage.
Shift Verification Upstream
Typed schemas, programmatic checks, static analysis, automated tests, security controls, and policy validation should run before output enters expensive human review. Humans should judge trade-offs and exceptions, not repeatedly catch malformed or incomplete output.
Keep Changes Small and Reviewable
AI makes it easy to generate large change sets. Large pull requests increase review time and hide risk. Decompose work into bounded, evidence-backed increments with explicit acceptance criteria, and a clear rollback path.
Put Humans at Consequential Gates
The future is not human or agent. It is bounded autonomy. Agents can research, generate, test, and prepare artifacts. Humans should remain accountable for business intent, architecture trade-offs, security exceptions, risk acceptance, and release decisions.
Preserve Lineage and Close the Loop
Every requirement, decision, code change, test, approval, and release should remain connected. Production signals and human corrections should improve the context, skills, tests, and controls used in the next change.
From local acceleration to trusted production change
A connected engineering system turns generated output into reviewable, traceable, and production-ready change.

Where Xebia Ace Fits
Xebia Ace does not attempt to replace the coding agents and developer tools teams already use.
Xebia Ace provides the engineering system around them.
Xebia Ace carries enterprise and application context across requirements, architecture, development, testing, and release. Specialized agents and reusable skills execute work, while structured outputs, deterministic workflows, automated evaluations, and a 12-pillar Quality Gate provide repeatable controls around probabilistic model behavior. Human approvals remain at the decisions that require judgment and accountability. Traceability connects intent to the resulting artifacts and production evidence.
The same system supports three engineering missions:
- Greenfield development, where intent must remain connected to architecture, code, tests, and release;
- Brownfield enhancement, where existing behavior and dependencies must be understood before change; and
- Legacy modernization, where application knowledge is reconstructed before technology is upgraded, refactored or decomposed.
The goal is to reduce the context reconstruction, verification, and rework that prevent local AI productivity from becoming delivery performance, not maximize generated output.
Point tools accelerate execution. Xebia Ace connects execution to enterprise delivery.
The Leadership Shift
Code generation will continue to improve. Every major engineering platform will gain more capable agents. That is good news. But it also means raw code production will become less differentiating.
The scarce capabilities will be the ones surrounding the model: trusted enterprise and application context; executable architecture and engineering standards; automated quality and security controls; human accountability at the right decisions; traceability from intent to production; and evidence that the outcome was worth the cost.
The next software engineering advantage will not come from producing more code than competitors. It will come from converting abundant generation into trusted change.
Faster tasks are useful. Faster delivery requires a better system.
Curious whether your own delivery pipeline is already paying these taxes? Talk to a Xebia Ace engineer.
Sources:
- DORA, “Balancing AI tensions: Moving from AI adoption to effective SDLC use,” 10 March 2026.
- Becker, Rush, Barnes and Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” 2025.
- Xebia controlled field experiment, July 2026: production-released brownfield enhancement across three repositories, comparing a standalone agent pipeline with deterministic application intelligence and Xebia Ace. Public use of client-identifying details remains subject to approval.
This is the second post in an ongoing series on AI-Native digital engineering. You can read our previous blog here: Moving Beyond AI Experiments to AI-Native Digital Engineering.
Stay tuned for our next blog: Context Is the New Engineering Infrastructure: Why AI Cannot Safely Change What It Does Not Understand.
Frequently Asked Questions
Our Ideas
Explore More Articles

Frontier Company: When AI Becomes an Enterprise Capability, Not an Experiment
Microsoft’s launch of Frontier Company confirms what we at Xebia have been seeing for years: the real challenge in enterprise AI is no longer about...
Marieke van der Werf
Contact



