JevonsMaxxing

AI is making firms more prolific before it makes them more productive.

The Wrong Question

The wrong question in the AI productivity fight is whether AI reduces work. That is the labor-substitution question, and it is too tidy for the phase we are in.

The live debate runs through The Information’s Uber budget shock showing local utility outrunning the budget model; Anthropic, the interested protagonist behind Claude Code, selling both the accelerant and the relief machinery; McKinsey, Menlo, Gartner, Deloitte, and Stanford HAI mapping adoption, scaling lag, and the productivity paradox; Faros, Latent Space, and software-delivery researchers instrumenting the engineering floor; Zscaler ThreatLabz, SentinelOne, Sonatype, and AI-code security researchers warning about the vulnerability multiplier; and ROI skeptics such as Aswath Damodaran and Dan Rasmussen pushing back against treating usage as value.

The argument splits six ways: prolific versus productive, demand validation versus uncontrolled waste, verification versus value capture, productivity multiplier versus vulnerability multiplier, Jevons expansion versus labor substitution, and software-specific story versus organizational pattern. The shared question is whether more candidate work becomes accepted value, or merely more attempted work for the firm to process.

What JevonsMaxxing Names

“JevonsMaxxing is an economic diffusion pattern, not internet slang. Generation gets cheaper, usage expands, and scarcity moves downstream.”

Jevons Paradox says efficiency improvements can increase total consumption by lowering the effective cost of use. JevonsMaxxing translates that into AI lowering the cost of candidate work so attempted work expands before verification, security, budgeting, and value capture catch up.

JevonsMaxxing is an economic diffusion pattern, not internet slang. Generation gets cheaper, usage expands, and scarcity moves downstream. That does not mean the firm has become productive. It means the firm has become ingestive. Candidate output starts arriving faster than the organization can review, trust, secure, budget, procure, attribute, govern, own, and assimilate it.

The Budget Model Breaks

“Budget blowouts are demand signals before they are value signals.”

Uber is the cleanest hook because neither camp gets to declare victory. It is not the triumph story, and it is not the failure story.

In December 2025, Uber gave about 5,000 engineers access to Claude Code. Between December 2025 and February 2026, usage nearly doubled. By April 2026, CTO Praveen Neppalli Naga said the expected annual AI budget had already been blown away.

That sequence is the point. Uber was not simply foolish, and Anthropic had not proved the revolution. Local utility outran the budget model. Budget blowouts are demand signals before they are value signals.

The Firm Becomes Hungry Before It Becomes Healthy

ORGANIZATIONAL METABOLISM

The Firm's New Bottleneck

Cheap generation widens intake first. Productivity depends on what survives verification and value capture.

Intake
What gets cheaper
What expands
Drafts, pull requests, analyses, tickets, experiments, synthetic research, vendor workflows, and agent runs.
What the signal means
Usage can reveal local demand before the organization knows how to account for value.
Failure mode
More candidate work arrives than the firm can sort.
FIRST
More attempted work
Verification Gate
What becomes scarce
What the signal means
More generated output also means more exposure to validate.
Failure mode
Backlog, review drag, trust drag, and compounding security assumptions.
What must happen
Review, trust, security, quality, permissions, dependencies, configs, secrets, and acceptance have to keep up.
GATE 1
More inspection
Value-Capture Gate
What converts motion into value
What the signal means
A workflow can be locally useful and institutionally illegible.
Failure mode
Budget stress, unowned motion, and activity that cannot be recognized as return.
What must happen
Finance, procurement, governance, workflow ownership, attribution, and management adoption attach to accepted output.
GATE 2
Accepted throughput
Conceptual synthesis from the article's organizational metabolism sections; not a measured conversion funnel.

The broader enterprise picture has the same shape. Menlo’s 2025 work puts enterprise AI spending at roughly $37 billion, while only 16% of enterprise deployments qualify as true agents. Most deployments are still fixed workflows or routing systems, not autonomous systems that plan, act, and adapt.

McKinsey’s 2025 data lands nearby: 23% of organizations are scaling agentic AI in at least one function, 39% are experimenting, and no more than 10% have scaled agents enterprise-wide or across most functions. S&P Global sees rapid GenAI adoption with mixed results, dragged by budget constraints, data quality, privacy, and security.

This is the metabolism model. A firm has intake, digestion, waste, assimilation, and accepted throughput. Cheap generation increases intake. The overloaded loading dock fills with drafts, pull requests, analyses, tickets, experiment ideas, synthetic research, vendor workflows, and agent runs. Some of that is real work that previously would not have existed. Some is duplicated, low-quality, or institutionally homeless. The question is not how much arrives. The question is what the organization can metabolize.

The Verification Gate

“Local speedup becomes organizational drag when review, trust, quality, permissions, security, and acceptance become the first scarce resources.”

VERIFICATION PRESSURE

Software Shows The Loading Dock

The measured software case shows output expanding while review, delegation, and acceptance remain constrained.

+98%+98%
Merged PRs
AI coding tools produced more merged pull requests in software-delivery evidence cited by the article.
+91%+91%
Review time
Review times rose in the same PR-flow evidence; output did not eliminate verification.
27%27%
Net-new work
Anthropic reports this share of Claude-assisted software engineering work would not otherwise have been done.
~60%~60%
AI use in work
Anthropic reports engineers use AI in roughly this share of their work.
0-20%0-20%
Full delegation
Anthropic reports full delegation remains bounded rather than becoming the default.
Software-delivery PR-flow evidence and Anthropic's 2026 agentic coding evidence are source-separated; figures are not one common sample.

Software is the evidence-rich case because, for once, the loading dock has instrumentation.

Faros, Latent Space, and software-delivery researchers show AI coding tools producing 98% more merged PRs while increasing review times by 91%. Anthropic’s own numbers need the same careful reading: 27% of Claude-assisted software engineering work would not otherwise have been done; merged PRs per engineer per day rose about 67%; developers use AI in about 60% of work; full delegation remains only 0-20%. The work is not disappearing. It is being generated, shaped, reviewed, and accepted through a human organization with more in flight.

The tail thickens too. Claude Code’s 99.9th percentile turn duration nearly doubled, from under 25 minutes to over 45 minutes. Local speedup becomes organizational drag when review, trust, quality, permissions, security, and acceptance become the first scarce resources.

The pull request is only the visible carton on the dock. Inside it may be a changed authentication flow, a new dependency, a permissive config file, an exposed secret, an agent handoff, or a tool permission broader than the reviewer is consciously evaluating.

That is why security is not a side concern in JevonsMaxxing. Gartner’s April 2026 forecast expects GenAI application security incidents to rise materially. Current vulnerability evidence finds AI-generated code with higher exposure than human code across OWASP Top 10 issues, privilege-escalation paths, design flaws, and secrets exposure. AI coding tools also introduce distinct attack-surface classes: config-based injection, prompt injection through trusted tool or API access, command-execution vulnerabilities, and supply-chain amplification through hallucinated dependencies or autonomous package installation. Multi-agent detection can decay across chain hops, creating compounding review gaps rather than one isolated bad step.

The Value-Capture Gate

“A workflow can be locally useful and institutionally illegible: sensible to users while the firm has no clean way to assign cost, recognize output, or hold anyone accountable for return.”

VALUE-CAPTURE TESTS

Spend Clears Before Absorption

Enterprise AI demand is visible, but agency, scaling, and team-level value capture clear on slower evidence.

The firm becomes hungry before it becomes healthy: spend and experimentation show intake, while true agency and accountable return remain harder tests.

Cleared first

Spend

Demand evidence
What the evidence proves

Menlo's 2025 work puts enterprise AI spending at roughly $37B.

What remains unsettled

Spend reveals demand and budget pressure; it does not prove accepted throughput.

Usage is a signal before it is value.

Partial proof

True agency

Maturity lag
What the evidence proves

Menlo reports only 16% of enterprise deployments qualify as true agents; most remain fixed workflows or routing systems.

What remains unsettled

Autonomous planning, acting, adapting, and ownership remain uneven.

Partial proof

Scaling

Partial scaling
What the evidence proves

McKinsey reports 23% scaling agentic AI in at least one function, 39% experimenting, and no more than 10% scaled enterprise-wide or across most functions.

What remains unsettled

Scaling is visible, but broad institutional absorption remains limited.

Slower proof

Team value

Value-capture drag
What the evidence proves

Individual GenAI savings of 4.11 hours per week fall to 1.5 hours at team level in the article's Gartner finding.

What remains unsettled

Finance, procurement, governance, attribution, workflow ownership, and management adoption must still convert local usefulness into accountable return.

Menlo, McKinsey, Gartner, and the article's savings evidence use different bases, so this is a layer-separated reading rather than one arithmetic funnel.

Menlo 2025, McKinsey 2025, Gartner savings evidence, and the article's value-capture section.

The second scarcity is the value-capture gate.

Finance, procurement, governance, workflow ownership, attribution, and management adoption lag behind the new intake. A workflow can be locally useful and institutionally illegible: sensible to users while the firm has no clean way to assign cost, recognize output, or hold anyone accountable for return.

Gartner’s savings finding has this contour: individual GenAI savings of 4.11 hours per week fall to 1.5 hours at the team level. One non-coding workflow example consumed 334 million tokens per day across 3,264 sessions at $778 per day. That is not automatically cheap, and not automatically expensive. It is cheap if the workflow creates recognized value. It is expensive if it merely creates more unowned motion.

This is where the missing institutional owner matters. Often that owner has to be the CFO, procurement, or AI FinOps function, not because any already has a finished doctrine, but because somebody has to turn tokenized activity into workflow-specific budgets, accepted outputs, and accountable returns.

Demand Before Value

“The synthesis is JevonsMaxxing: cheap generation expands the surface area of work before the organization expands the surface area of verification and value capture.”

FOUR MEASURES, ONE SHAPE

What Survives the Trip from the Desk to the Institution

Four independent measures of the same gap — what the individual gets, and what the organisation keeps. In each, the institution keeps between a sixth and under half.

Time saved
4.11 h1.5 hper personper team
36%36%
of the individual's weekly saving survives at team level.
Gartner
Delegation
~60%0–20%AI in the workfully delegated
≤33%≤33%
at most, of the work AI touches is fully handed to it.
Anthropic, 2026
Scaling
23%≤10%in one functionenterprise-wide
≤43%≤43%
at most, of the organisations scaling agents in one function have scaled them across the enterprise.
McKinsey, 2025
Agency
All16%deploymentstrue agents
16%16%
of enterprise AI deployments qualify as true agents rather than fixed workflows.
Menlo, 2025
Gartner (time saved), Anthropic 2026 (delegation), McKinsey 2025 (scaling), Menlo 2025 (agency). Each share is taken within one source; the four are not one funnel. Where a source gives a range, its upper bound is used, so the share shown is at most that.

The reconciliation is less satisfying than any camp wants. Optimists are right that heavy usage can reveal local demand. Operators are right that the load is real, and that more AI-generated work can mean more review, security, platform, and governance work. Skeptics are right that neither spend nor output proves value.

Anthropic’s usage story is real but not neutral. Damodaran and Rasmussen are right to pressure the investor leap from spend to return. The synthesis is JevonsMaxxing: cheap generation expands the surface area of work before the organization expands the surface area of verification and value capture.

What Healthy Metabolism Looks Like

The healthy firm is not the one with the most model access or the steepest adoption curve. That is intake.

The signal is workflow sorting. Permissions get scoped. Work gets routed. Review checkpoints move closer to generation. Budgets attach to workflows, not vibes. Security assumptions become explicit. High-ROI use cases concentrate instead of dissolving into universal entitlement.

That is what digestion looks like. The loading dock is not productivity simply because more trucks arrive. AI may still become a clean productivity multiplier. But the visible phase is organizational metabolization under stress: more candidate output, more verification work, more security surface, more budget ambiguity, and more fights over accepted throughput. JevonsMaxxing is what happens before the accounting catches up.