The Constraint Didn't Disappear. It Moved.

Matthias Heim · 2026-07-10

Teams write 228% more code with AI and ship 10% more product. This interactive pipeline shows why, and what the teams that broke the plateau did differently.

Ask around your company and you'll hear it: "We already use AI." The developers have a copilot. The PMs draft with ChatGPT. Everyone feels faster. By every subjective measure, adoption is going great.

Now the uncomfortable data point. In a 2025 randomized trial by METR, experienced developers working on mature codebases were 19% slower with AI tools. The whole time, they believed they were 20% faster. The feeling of speed is real. The speed itself, at the level that shows up in your results, often isn't.

This post is about that gap: why genuine individual gains stall before they reach the P&L, and what the teams that broke through did differently. There's an interactive pipeline below: the whole argument fits inside it.

The Plateau Is Real, and Quantified

Across every major study of enterprise AI adoption, the same picture emerges: usage is nearly universal, results are not.

95%

78% → 6%

21%

Look at that third number again. Barely one company in five has changed how work actually flows. Everyone else has inserted new tools into old processes. That, not model quality and not skills, turns out to be the plateau.

The Conversion Gap: More Code, Same Output

The cleanest evidence comes from a 2026 NBER study by Demirer, Musolff and Yang, aptly titled "Writing Code vs. Shipping Code". Developers using AI autocomplete produced 228% more raw code. Shipped releases went up 10.2%.

With more autonomous coding agents the gap gets wider, not narrower: code volume rose 741%, and releases rose just 20.3%. Whatever limits organizational output, it is not typing speed.

more code written with AI autocomplete

more releases actually shipped

The reason is mechanical, and you can feel it yourself in the simulator below. A product pipeline ships at the rate of its slowest stage. Speed one stage up (even massively) and the constraint simply moves next door. Drag the Build slider and watch it happen.

The Value Pipeline: Where Does Your Speed Go?

Five stages take an idea from opportunity to customer. Add AI speed to any stage and watch what happens to actual shipped output. The day values are illustrative; the mechanics are not.

Insert AI into the old process

The 2024–2025 default: hand the tool to whoever types. Only Build is in reach.

Redesign the process around AI

Change how work flows between stages. Every slider unlocks.

Replay the study (+228% code)

Reset

Learn

talk to users, find the opportunity

Prioritize

decide what matters most

Build

write the code

Validate

review, test, judge whether it's good

Ship

release, documentation, go-live

days

The locked stages are the authority gap in miniature: the person at the keyboard can speed up their own step, but has no mandate to change the process around it. Switch to "Redesign the process" to unlock them.

A pipeline ships at the rate of its slowest stage, not at the sum of its speedups.

constraint

more code produced

more features shipped

time to ship one feature

The constraint now sits at:

Drag the Build slider up: that's the stage almost every company accelerated first.

Build got dramatically faster, but shipped output barely moved. The human-limited stage next door now caps the whole pipeline. This is the +228% / +10.2% gap in action. To go further, you need the toggle above.

Now the rest of the chain is in reach. Speed up whichever stage glows red: each constraint you clear reveals the next one.

This is what workflow redesign unlocks: gains of 30–50% and beyond, versus 10–15% for tool insertion (BCG's Deploy vs. Reshape gap). Notice that you had to touch stages no individual contributor controls.

Study values: AI autocomplete produced +228% more code but only +10.2% more shipped releases; more autonomous agents +741% code, +20.3% releases (Demirer, Musolff & Yang, NBER Working Paper 35275, 2026).

Why Teams Get Stuck (It's Not a Skill Problem)

The instinctive diagnosis is a training gap: send people to a prompt workshop and the plateau will resolve. The evidence says otherwise. Organizational research shows that a team with a fixed structure always settles at a local maximum: a "good enough" that more effort cannot escape (Siggelkow & Levinthal, 2005). Copilot-style AI clears the good-enough bar so cheaply that the search for a better way of working stops right there. More tool exposure doesn't dislodge it. Only changing the structure itself does.

And there is a second lock, the most corroborated finding in the 2025–2026 research: the authority gap. Organizations hand AI tools to the people doing the work, but keep the decision rights over how work is structured several levels up. The developer who can see that code review is now the constraint has no mandate to change the review process. The breakthrough sits above the pay grade of the person at the keyboard. It's exactly like the locked sliders you just met.

A mandate without infrastructure is pressure on employees. The same mandate on top of infrastructure is pressure on the status quo.

The lesson from Shopify's 2025 AI memo: the memo alone changed little. The infrastructure and the changed performance expectations underneath it did the work.

So the plateau is a local maximum nobody can escape by trying harder, held in place by an authority structure that separates the people who see the constraint from the people who could move it. The way out starts with naming what actually changed.

The Constraint Moved from Feasibility to Judgment

For the entire history of software, the scarce resource was building. Every gate in your process (the PRD sign-off, the roadmap fight, the quarterly prioritization) exists because engineering time was too expensive to waste. Scarcity did your quality filtering for free: only ideas that survived the gauntlet got built. Cheap code doesn't remove that constraint. It relocates it.

"Can we afford to build this?"

"Should we build this, and is it any good?"

Andrew Ng put numbers on it in 2026: with AI tools his engineers work roughly 10× faster, and product management hasn't sped up at all. His conclusion: the constraint has moved to the judgment layer, and teams are now discussing two product managers per engineer: an exact inversion of the traditional ratio.

This is why the plateau feels so disorienting. Companies removed the feasibility filter and discovered they had never built a real mechanism for judging whether an idea is good: the old gates were doing that job silently, by rationing. Taste, validation and decision speed are now the product bottleneck.

What Breakthrough Actually Looks Like

Only 21% of organizations have redesigned any workflow around AI. Yet in McKinsey's analysis, workflow redesign is the single highest-weighted lever, out of 25 tested, for getting bottom-line impact from AI. The gap between those two facts is the opportunity. BCG's maturity research puts bands on it:

Deploy

+10–15%

Insert AI tools into existing processes. Real gains, quickly capped. This is where 79% of companies still sit.

Reshape

+30–50%

Redesign how work flows across the whole chain: review processes, release cadence, decision rights. This is where the compounding starts.

You just experienced the difference between those two numbers as a toggle. In a real organization the toggle is harder to flip: it means changing review processes, release trains and decision rights, not dragging sliders. Which is why the practical question isn't "which tool?" but "how do we redesign without betting the company on a theory?"

Start With One Slice, Not a Transformation Program

You don't flip a 15-year-old organization to a new operating model by memo. You also don't need to. The pattern that works is a carve-out: demonstrate the redesigned workflow on one safe slice, with the authority-holder in the room, and let the velocity make the argument. The shape:

One new slice

Pick a zero-to-one piece of work: a new feature, a prototype, an internal tool. Never the legacy core, never anything safety-critical. New ground has no ceremony to defend.

3–5 volunteers

Volunteers, not conscripts, and crucially, including someone who owns the process, not just the people who use the tools. That closes the authority gap for one week.

One week

Long enough to go from idea to something a real user can try. Short enough that nobody has to bet a roadmap on it.

A one-page permission slip

One page from leadership: for this slice, for this week, the team may drop the standard ceremony (status meetings, tickets, spec documents) and work AI-first. The permission is the intervention: it changes the structure, not the people.

Quality rules doubled, not dropped

Rigor moves from documents into feedback loops: human review first, feature flags, test coverage, guardrails in code. The redesign removes ceremony, never safety.

That last point carries the weight of history. In the 1990s reengineering wave, 50–70% of initiatives failed, and the pattern separating winners from losers was brutal: winners redesigned the process before cutting anything; losers cut first and validated later. Killing your gates before you've proven that the replacement judges quality better repeats the losers' error. Redesign first. Then, and only then, retire the old process.

Sources

Demirer, Musolff & Yang, "Writing Code vs. Shipping Code", NBER Working Paper 35275 (2026): +228% code / +10.2% releases with AI autocomplete; +741% / +20.3% with more autonomous agents.

METR, randomized controlled trial with experienced developers (2025): 19% slower with AI tools while believing they were 20% faster.

MIT NANDA, "The GenAI Divide" (2025): 95% of GenAI pilots show zero measurable P&L impact, across $30–40 billion invested.

McKinsey, The State of AI (2025): 78% of organizations use AI somewhere; ~6% are high performers; 21% have redesigned any workflow; workflow redesign is the highest-weighted of 25 levers for EBIT impact.

BCG, AI maturity research (2024–2025): Deploy-phase gains of 10–15% vs. Reshape gains of 30–50%; 79% of companies have not moved past Deploy.

Andrew Ng (2026) on the shifted bottleneck: engineers roughly 10× faster with AI while PM capacity is unchanged, with teams proposing two PMs per engineer.

Siggelkow & Levinthal (2005) on local maxima: a fixed organizational structure always halts at a local optimum; only changing the structure itself dislodges it.

How we work · Work · Insights