---
title: "AI Engineering Productivity is Anything But Normal"
description: "Per-engineer productivity from AI coding runs from -19% to 8x. The distribution has three tiers, \u0026 the gap between them is the operating layer, not the model."
categories: ["AI","SaaS"]
keywords: ["ai coding productivity","engineering productivity","ai agents","software factory","replit","devin","cognition","nvidia cursor","anthropic claude code"]
ai_summary: "AI engineering productivity outcomes cluster in three tiers: a mean of 20-46% from distributing an IDE, a frontier of 2.5-3x from companies building the operating layer around agents (Replit, NVIDIA, Amplitude, Anthropic), \u0026 a factory tier at 8x+ where agents operate as first-class organizational units (Nubank with Devin). The gap is not the model but the operating discipline built around it."
date: 2026-07-21
lastmod: 2026-07-21
canonical_url: https://www.tomtunguz.com/ai-engineering-productivity-anything-but-normal/
author: "Tomasz Tunguz"
---


We are now in an era where we should expect 3x more from each other.

Over the last six months, one data point has followed another :

- NVIDIA reported a 3x increase in committed code across 30,000 developers with bug rates flat.[^nvidia]
- Amplitude tripled weekly production commits, with an AI agent now a top-three contributor to the codebase.[^amplitude]
- Anthropic measured a 2.5x increase in code written per engineer since adopting Claude Code internally, quality stable.[^anthropic]
- Replit doubled its team & tripled per-engineer output over the same period, with review times, reversions, & incidents all flat.[^replit]

{{< email_image src="hd4mzhiuhqfmdrzvinch" alt="ai-eng-productivity-distribution" width="540" height="342" >}}

[^chart]: The distribution above is illustrative, not statistical. Each point is a reported multiplier from a published study, RCT, or company disclosure. It is not drawn from a sampled population, & the curve is a right-skewed log-normal fit to the pattern of reported outcomes, not to raw data. Treat it as a shape argument, not an estimator.

[^nvidia]: [Cursor, "How NVIDIA uses Cursor," February 2026](https://cursor.com/blog/nvidia).

[^amplitude]: [Cursor, "Amplitude and Cursor cloud agents," April 2026](https://cursor.com/blog/amplitude).

[^anthropic]: [Boris Cherny, head of Claude Code, on the Big Technology podcast, July 2026](https://www.bigtechnology.com/p/boris-cherny-claude-code).

[^replit]: [Amjad Masad, "The Self-Driving Company," July 16, 2026](https://blog.replit.com/self-driving-company).

[^augment]: [Augment Code on X, 2026](https://x.com/augmentcode/status/2070243305385066973).

[^faros]: [Faros, "AI Engineering Report 2026"](https://www.faros.ai/blog/are-ai-coding-assistants-really-saving).

[^google]: Google internal randomized controlled trial, ~100 engineers, 2024. Referenced in DORA reports; roundup at [Value Add VC](https://valueaddvc.com/blog/ai-coding-productivity-study-data-what-metr-mckinsey-and-github-actually-found-in-2026).

[^copilot]: GitHub, Microsoft, and Accenture study with a large fintech, ~450 developers, 2024.

The chart above sorts the ecosystem into three unequal tranches, each defined by how much of the model's power the company captures.[^chart]

The first tranche is what most companies experience today. Distribute an AI IDE, change nothing else, & the outcome is modest.

> "Engineering leaders went into AI expecting 2-3x productivity gains but are landing closer to 30%."
>
> — Augment Code[^augment]

Faros's telemetry across 22,000 developers confirms this: engineers completed epics 66% faster, but bugs per developer increased by 54%.[^faros] The Google randomized controlled trial put the number at 21%, close to GitHub's 24%.[^google][^copilot] This is the default outcome.

The frontier tranche follows. Companies here have built harnesses around the model, orchestrating agents sharing context across GitHub, Linear, & Slack; escalating to engineers for their judgment.

> "Every employee gets a manager agent that spawns worker agents in loops. Our internal agent outperformed a seven-figure SaaS tool in security testing and incident triage at one-tenth the cost."
>
> — Amjad Masad, Replit, "The Self-Driving Company"[^replit]

Human PR review time dropped 30%. Complex support handling time dropped 60%. Total code contribution rose 5.8x. This is where the 3x number lives.

The third tranche are the software factories, & here the name is an apt descriptor. They are AI machines that produce software mechanistically. Cognition's Devin refactors monolithic codebases end-to-end. Factory.ai is deploying software factories at NVIDIA, Adobe, Blackstone, & EY.[^factory]

> "Nubank achieved an 8x improvement in engineering efficiency & a 20x cost reduction using Devin for large-scale refactoring."
>
> — Contrary Research, January 2026[^nubank]

Goldman Sachs is piloting Devin alongside 12,000 human developers & publicly estimates agentic AI could deliver 3-4x the rate of prior tools.[^goldman]

[^factory]: [Factory.ai, "Factory 2.0: From coding agents to software factories"](https://factory.ai/news/software-factory).

[^nubank]: [Contrary Research, "Cognition"](https://research.contrary.com/company/cognition), January 2026.

[^goldman]: [CNBC, "Goldman Sachs is piloting its first autonomous coder in major AI milestone for Wall Street," July 2025](https://www.cnbc.com/2025/07/11/goldman-sachs-autonomous-coder-pilot-marks-major-ai-milestone.html).

AI engineering productivity gains are here. The initial data shows what to expect: most teams should migrate from 20% productivity gains to a 3x productivity gain & they aren't normal.
