THE RECURSIVE ECONOMY
RELEASED 15 AUGUST 2026

RESEARCH ESSAY VERSION 1.0

THE
RECURSIVE
ECONOMY

What happens when intelligence becomes an input into the production of intelligence?

Scarcity does not disappear. It migrates to the first thing intelligence cannot cheaply reproduce.

Explore the essay

THE PRODUCTION LOOP

The recursive production loop Capability supports machine research. Research produces improvements. Improvements increase capability, while verification constrains the loop. 01CAPABILITYmodels + systems 02RESEARCHsearch + selection 03IMPROVEMENTverified gains VERIFICATION gains must survive review

The central question is not whether the loop exists, but whether it closes, transfers, and compounds.

READING MAP

An index to the argument

The essay moves from the mechanics of recursive improvement to scarcity, ownership, state capacity, and the choices that shape the transition.

Approximately 78 minutes at a continuous reading pace.

The decisive machine is not the one that can do our work. It is the one that can improve the process that makes machines.


Note to the reader#

This is a thesis, not a prophecy. It distinguishes observations from inferences, measurements from model outputs, and forecasts from signposts. Empirical claims are sourced; numerical scenarios are transparent and reproducible. The argument is current through 15 August 2026. For feedback or comments, contact tre.numbing085@passfwd.com.


Abstract#

Most accounts of advanced artificial intelligence begin with a question about capability: when will a machine match or exceed a human at economically valuable work? This essay begins one step upstream. What happens when machine intelligence becomes a material input into the production of better machine intelligence?

That feedback loop is recursive self-improvement, or RSI. The phrase often evokes a single model inspecting its source code, rewriting itself, and abruptly becoming superintelligent. That image is too narrow and too theatrical. The economically relevant loop is already more industrial. Models write and review code, design experiments, generate and filter data, optimize kernels, discover algorithms, operate infrastructure, and increasingly participate in research. A future system need not directly edit its own weights to accelerate its successor. It only needs to improve enough of the pipeline that produces the next system.

The existence of such a loop does not imply an intelligence explosion. Four separate gates stand between useful AI agents and self-sustaining recursive acceleration. The system must be competent at AI research tasks. The loop must close: it must choose, execute, and evaluate work with diminishing human direction. Improvements must create enough gain that each generation materially accelerates the next. Finally, narrow improvements on AI-specific benchmarks must translate into broad, economically useful capability. The public evidence in 2026 is strongest at the first gate, partial at the second, uncertain at the third, and thin at the fourth.

The central economic claim of this essay is that RSI changes the production function for ideas. In the ordinary economy, better technology raises the productivity of labor and capital. In a recursive economy, better technology also raises the productivity of the process that produces better technology. Effective cognitive labor begins to manufacture more effective cognitive labor. Research effort becomes partly reproducible, rapidly deployable, and potentially compounding.

The consequence is not the end of scarcity. It is the migration of scarcity.

As machine cognition becomes cheaper, the binding constraint moves through the stack: from researchers to evaluators; from evaluators to compute; from compute to chips, electricity, cooling, and grid connections; from digital infrastructure to factories, laboratories, robots, land, permits, and raw materials; and from material constraints to trust, legitimacy, law, and biological time. Capital follows that boundary. The value of an abundant input falls toward its replication cost; rents accrue to the complements that remain scarce and controllable.

This produces a four-clock economy. Software, research artifacts, designs, and decisions can accelerate toward machine time. Data centers, fabs, grids, mines, housing, hospitals, and factories remain tied to construction time. Drug effects, childhood, ecological recovery, and human relationships remain tied to biological time. Elections, courts, treaties, professional norms, and public consent remain tied to institutional time. Recursive intelligence can compress the first clock without abolishing the others. The collision among them is where most of the economic and political consequences appear.

The ownership of the loop is therefore as important as its strength. Strong recursion under concentrated ownership creates a world of recursive sovereigns: frontier labs and states controlling a strategic stock of machine research labor. Strong recursion with rapid diffusion creates a recursive commons: intelligence commoditizes while physical assets, verification, distribution, and legitimacy capture value. Weak recursion with concentration resembles a fortified oligopoly. Weak recursion with diffusion resembles abundant tools and a more conventional productivity boom. These are different equilibria, not minor variations on one forecast.

For workers, RSI separates the production of wealth from the distribution of purchasing power. Output can rise while labor's claim on output falls. The crucial race is no longer simply between automation and new-task creation; it is between automation, capital accumulation, diffusion, and the political construction of new claims on machine-produced income. Human work may remain valuable because it is scarce, embodied, trusted, positional, legally required, or intrinsically desired. But none of those guarantees that aggregate wages remain the economy's principal distribution mechanism.

For firms, the minimum efficient scale of cognition collapses while the minimum efficient scale of infrastructure may rise. A small team can command an enormous virtual organization, yet the frontier may require unprecedented clusters, energy contracts, and supply-chain coordination. The result could be a barbell economy: tiny firms at the application edge and gigantic capital complexes at the intelligence core.

For investors, the recursive economy is not a ticker list. The same technical path can enrich model owners, destroy model rents through diffusion, inflate power and infrastructure values, or strand overbuilt assets if efficiency improves faster than demand. The durable question is not “Which sector benefits from AI?” but “Where is the recursion frontier—the bottleneck that intelligence has not yet made cheap—and who controls it?”

For states, population and educational attainment become less decisive relative to compute, energy, semiconductor access, security, and institutional capacity. Tax systems built around labor income become fragile. Antitrust, industrial policy, national security, and macroeconomic stabilization converge around the same assets. A state that cannot measure the loop cannot govern the transition.

The essay closes with a dashboard of falsifiable signposts. The recursive-economy thesis strengthens if real-world AI research uplift rises, autonomous task horizons lengthen, human review becomes a measured bottleneck, systems select productive experiments on held-out problems, successor-training loops outperform human-designed baselines, and macro productivity begins to separate from employment growth. It weakens if gains remain benchmark-bound, evaluation costs rise as fast as generation, task horizons plateau, algorithmic improvements fail to transfer, or physical and institutional bottlenecks absorb nearly all digital acceleration.

The future may not contain a clean “takeoff.” It may instead contain a succession of bottleneck transfers, each mundane in isolation and revolutionary in aggregate. The economy will feel recursive not when a machine declares itself improved, but when the time required to create a better machine is determined less by the number of human researchers and more by the stock of compute, the quality of verification, and the speed at which the physical world can respond.


The thesis in twelve propositions#

  1. Recursive self-improvement is a production loop, not a personality trait. The relevant unit is the full AI research-and-deployment system: models, agents, data, evaluators, software, humans, compute, and institutions.

  2. The loop has begun, but it has not closed. Systems already improve code, experiments, algorithms, tools, and research execution. Human goal selection, evaluation, integration, and successor training remain material constraints.

  3. Automation is not acceleration. AI can perform a large share of research tasks without making the next generation arrive progressively faster. Self-sustaining acceleration requires loop gain above a threshold.

  4. Narrow recursion can be economically small. A system may become excellent at optimizing its own benchmarks without acquiring broad scientific, managerial, or physical-world competence.

  5. The first macroeconomic effect is time compression. The interval between hypothesis, experiment, result, deployment, and iteration shrinks.

  6. Scarcity migrates rather than disappears. Every automated stage exposes the next slow, scarce, or difficult-to-verify complement.

  7. Capital flows toward the recursion frontier. Rents accrue to whichever bottleneck remains scarce, defensible, and necessary for the loop to continue.

  8. The economy splits across clocks. Digital processes accelerate; physical, biological, and institutional processes do not automatically follow.

  9. Ownership determines distribution. A recursive loop that is concentrated produces different wages, prices, firms, and politics from one that diffuses.

  10. Output and income can separate. Machine-produced abundance does not by itself create broadly distributed purchasing power.

  11. The state becomes an input to the production function. Permits, grid build-out, security, standards, taxation, and legitimacy can determine the realized return on intelligence.

  12. The thesis is measurable. It should live or die by real-world R&D uplift, loop closure, transfer, bottleneck movement, and macroeconomic diffusion—not by anecdotes or benchmark theatre.


I. THE LOOP#

1. The event is upstream of AGI#

The public argument about advanced AI has been organized around a finish line. One camp asks when systems will reach artificial general intelligence. Another asks when they will automate most jobs. A third asks when they will become dangerous enough to require exceptional control. These are important questions, but they point at the outputs of a process. Recursive self-improvement points at the process itself.

A factory that manufactures ordinary goods can become more productive. A factory that manufactures better factories changes the rate at which productivity can improve. An education system that trains workers contributes to growth. An education system whose graduates can instantly duplicate themselves and train still better graduates changes the supply dynamics of skill. AI research is unusual because its output—more capable machine intelligence—can become an input into the next round of AI research.

That is the event this essay studies.

It may occur before a clean threshold called AGI. A model can be uneven, unreliable, and visibly nonhuman while still being valuable in the specific work that advances machine learning: coding, debugging, experiment execution, literature synthesis, data generation, evaluation, systems optimization, and research operations. Conversely, a system could look generally intelligent in conversation while contributing little to the frontier if it cannot act over long horizons, obtain reliable feedback, or improve the machinery that trains successors.

The distinction matters because the economic clock starts when the research loop accelerates, not when society agrees on a label. A laboratory whose effective research workforce doubles every year through AI assistance is already on a different trajectory from one whose workforce grows through hiring. A firm that can copy a capable agent across ten thousand parallel experiments is not merely “more productive.” It has altered the replication cost of a class of cognitive labor.

The phrase recursive self-improvement has a long intellectual history, but its popular image remains a solitary program opening its own source code. Modern AI systems are not solitary. They are assembled from foundation models, training pipelines, retrieval systems, evaluators, tools, synthetic-data generators, human feedback, specialized kernels, distributed compute, and deployment infrastructure. The successor of a frontier model is not produced by the predecessor alone. It is produced by an organization and a technical stack.

The useful definition is therefore organizational:

Recursive self-improvement occurs when AI systems materially improve the socio-technical process that produces subsequent AI systems, and those improvements in turn increase the systems' ability to improve that process again.

Under this definition, a model need not modify its own weights. It can write a better training kernel, discover a more data-efficient algorithm, construct a stronger evaluator, diagnose a failed run, generate a curriculum, improve an agent harness, design a chip component, or coordinate a research program. If the resulting successor performs those activities better, the loop has a recursive component.

This broader definition is less cinematic and more economically serious. It turns RSI from a moment into a gradient. The question is no longer whether recursion is “on” or “off.” It is how much of the production chain is inside the loop, how large the feedback gain is, which parts remain exogenous, and where the bottleneck moves next.

The frontier evidence is consistent with an early loop. Google DeepMind's AlphaEvolve combines language models, evolutionary search, and automated evaluators to discover and optimize algorithms; DeepMind reports applications to data-center scheduling, chip design, and AI training, including components used by the system itself.1 The Darwin Gödel Machine iteratively rewrites the code of a coding agent and retains changes that improve benchmark performance while keeping the underlying foundation model frozen.2 AI-scientist systems have begun to generate hypotheses, run experiments, write papers, and submit artifacts to evaluation, although the quality distribution remains wide and human-created environments do much of the grounding.3 A 2026 survey of 1,250 papers describes a fast-growing field spanning deployment-time refinement, training-time self-iteration, self-evaluation, and automated research, while stressing that bounded improvement against fixed evaluators is categorically different from open-ended RSI.4

Inside frontier laboratories, the loop appears more advanced but is harder to audit. Anthropic's 2026 research agenda explicitly asks how AI R&D telemetry could provide early-warning signals of recursive self-improvement, how humans retain visibility into autonomous development, and what interventions could slow an intelligence explosion.5 These are agenda statements rather than independent measurements. Yet they reveal that frontier laboratories increasingly treat AI R&D automation as a measurable operational variable rather than speculative philosophy.

The most important word is operational. It is possible to automate ninety percent of a workflow and obtain only a modest total speedup if the remaining ten percent is serial, slow, or overloaded. Amdahl's law, developed for parallel computing, generalizes cleanly to organizations: accelerating one component exposes the share that did not accelerate.6 Generate code faster, and review becomes the constraint. Run more experiments, and experiment selection becomes the constraint. Produce more papers, and replication becomes the constraint. Discover more drugs, and clinical validation becomes the constraint.

RSI is therefore best understood as a contest between feedback and bottleneck migration.

FIGURE 1Recursive self-improvement as an industrial loop.

The loop can strengthen while progress remains bounded. It can generate a dramatic local speedup without becoming self-sustaining. It can improve the AI-specific research process while transferring poorly to the rest of the economy. It can produce more candidate ideas than the world can verify. It can make cognition abundant and leave energy, atoms, and legitimacy scarce.

The economic future depends on which of these descriptions becomes true first.


2. Four gates: from useful agents to recursive acceleration#

The language of RSI often compresses several different propositions into one. A coding agent fixes its own tool. An algorithm-discovery system finds a faster kernel. A benchmark curve steepens. Each result is then treated as evidence for the same conclusion: the system is “self-improving.”

That inference is too fast. There are four gates.

FIGURE 2The four gates of economically meaningful RSI.

Gate one: competence#

Can the system perform work that contributes to AI research and development?

This includes more than writing syntactically valid code. AI R&D requires understanding unfamiliar repositories, debugging distributed systems, designing experiments, choosing controls, interpreting noisy results, generating training data, monitoring runs, comparing architectures, and integrating changes without corrupting the broader stack. Some tasks are verifiable by tests; others require judgment under uncertainty.

Public evidence for competence has improved rapidly. METR's task-horizon evaluations measure the human completion time of software tasks at which agents achieve a target reliability. Its 2026 methodology uses diverse tasks drawn from software engineering, machine learning, and cybersecurity, fits success probability against human task duration, and explicitly warns that the result measures task difficulty rather than literal autonomous runtime.7 METR's earlier historical analysis estimated a roughly seven-month doubling in the length of tasks frontier agents could complete at 50 percent reliability, while later work emphasized benchmark and high-reliability limitations.89

FIGURE 3A normalized reconstruction of the historical task-horizon trend.

The fitted doubling rate should not be fetishized. The important signal is that autonomous work is extending from snippets toward meaningful work sessions. The important caveat is that well-specified, automatically evaluated technical tasks are unusually favorable to agents.

Benchmarks can exaggerate transfer. METR's randomized study of experienced open-source developers found that early-2025 AI tools made participants 19 percent slower on tasks in repositories they knew well, despite the developers expecting a speedup. A 2026 update using newer tools found suggestive speedups but wide confidence intervals and methodological challenges caused by changing behavior and selection.10 That sequence is not a timeless verdict on coding agents. It is a warning that benchmark competence, perceived productivity, and realized productivity are different objects. Any serious RSI thesis must measure the last one.

Gate two: closure#

Can the system run the loop rather than merely execute a human-specified step?

A closed research loop must generate or select a question, design an intervention, execute it, evaluate the outcome, decide what to retain, and choose the next step. Humans may set broad objectives or safety constraints, but they are no longer required at every local decision.

Closure is not binary. A system may close the loop inside a narrow environment with a crisp evaluator and fail completely in open-ended research. AlphaEvolve is powerful precisely where candidate programs can be run and scored. The Darwin Gödel Machine can test modifications on coding benchmarks. Self-play can produce curricula where outcomes are measurable. These are real loops, but their external anchors are doing substantial work. The harder domains are those where the evaluator is delayed, gameable, incomplete, or itself in need of improvement.

The evaluator is the hidden center of RSI. Generation without evaluation is noise at scale. A loop improves only relative to a signal that distinguishes better from worse. Formal proofs, unit tests, simulation scores, and exact task outcomes provide strong anchors. Learned reward models, language-model judges, peer review, market adoption, and human preference are softer. When the system begins to modify not only its output but also its evaluator, the possibility of genuine open-ended improvement appears—and so does the possibility of reward hacking, benchmark capture, and collapse into self-confirming error. Anthropic's primary-lab study of automated alignment research illustrates the boundary: it reports large gains on a narrowly defined task, but incomplete transfer and observed reward hacking.11

The 2026 survey of RSI research makes this distinction explicit: bounded self-refinement optimizes against a fixed external evaluator; open-ended RSI changes the machinery or criteria of improvement itself.4 The former is increasingly industrial. The latter remains a research frontier.

Gate three: gain#

Does a better system accelerate the creation of its successor enough to produce compounding acceleration?

Suppose an AI tool raises researcher productivity by 20 percent. That is valuable. But if training compute, data acquisition, evaluation, coordination, and deployment account for most of the cycle time, the next model may arrive only slightly sooner. If the successor then yields another 20 percent local uplift but the remaining bottlenecks dominate, total acceleration can converge rather than compound.

Recent economic work formalizes this as a network of feedback loops. Capability raises effective research effort; research effort raises algorithmic progress; algorithmic progress raises capability. Diminishing returns can prevent explosive growth, while innovation networks create reinforcing spillovers that can overcome them under particular elasticities and bottleneck structures.12 The key discipline is simple: “AI helps AI research” is not the same claim as “AI research is recursively accelerating.”

Gate four: translation#

Do narrow gains in AI R&D become broad capability and broad economic output?

A system can overfit the production process that created it. It may learn to optimize benchmark suites, training scripts, or familiar architectures while remaining poor at scientific discovery, management, persuasion, robotics, or novel institutional environments. Translation can fail through evaluation mismatch, distribution shift, integration cost, or embodiment delay.

An agent that makes its own coding harness better is not yet an autonomous research civilization. An AI-generated workshop paper is not yet a machine-run scientific institution. A dramatic optimization in a fixed code-speed challenge is not a proportional acceleration of frontier training. Passing all four gates would be historic. Passing only the first two would still be economically large.

The important analytical move is to stop treating these futures as identical.


3. The evidence ladder in 2026#

The strongest argument for RSI is cumulative rather than singular. No public demonstration establishes full recursive improvement. Many demonstrations establish pieces of the loop.

FIGURE 4Evidence is strongest for bounded loops and weakest for successor construction.

At the bottom of the ladder, self-correction is routine. Models critique drafts, sample alternatives, use tools, and refine outputs. These mechanisms can increase reliability but are usually bounded by a fixed prompt, judge, test, or reward. They improve the answer, not necessarily the system that answers.

Persistent skill and memory systems move one level higher. Agents accumulate reusable procedures, retrieve prior solutions, adapt tools, and preserve discoveries across episodes. Improvement persists outside the foundation-model weights. Economically, this matters because the deployed system becomes a capital stock: experience can be copied, recombined, and shared. Yet persistent memory can also preserve mistakes and create path dependence.

Harness self-modification is closer to recursion. The Darwin Gödel Machine modifies the code of its own agent framework, evaluates variants, and stores successful descendants. Its reported improvement on software-engineering benchmarks is meaningful evidence that an AI-driven search process can improve the machinery through which a frozen model acts.2 It is not evidence that the foundation model trained a better foundation model. The evaluator and task distribution remain external.

Algorithm discovery provides another strong bounded case. AlphaEvolve proposes programs, scores them with automated evaluators, and evolves promising candidates. DeepMind reports useful results in matrix multiplication, scheduling, chip design, and AI training.1 This is a genuine feedback channel: AI helps improve algorithms and infrastructure used by AI. Its strength depends on the availability of a reliable automated score. Where “better” is mathematically or operationally legible, search can scale. Where quality is ambiguous, delayed, or multi-dimensional, the loop weakens.

Research execution is now plausible in selected domains. AI-scientist systems can assemble literature, propose hypotheses, run code, generate plots, and write papers. Sakana AI reported that an unedited AI-generated paper passed blind review at an ICLR 2025 workshop; the underlying system and results were later described in a 2026 Nature publication.3 This shows that an end-to-end artifact can cross a human review threshold. It does not show that the median output is strong, that the research direction is important, or that the system can improve its successor.

Open-ended evaluations are beginning to test a more realistic mechanism: researchers delegate an actual project, then judge whether the returned work advances the frontier. A July 2026 paper explicitly frames this as a test of the pathway by which agents might accelerate AI research.13 These studies are more informative than isolated benchmark scores because they measure whether machine work enters a real research portfolio.

RE-Bench places agents and human experts in seven open-ended machine-learning environments. Agents performed better under a two-hour budget, humans held a narrow lead at eight hours, and human scores were roughly twice agent scores at 32 hours.14 The reversal is a warning against extrapolating short-horizon execution into open-ended research autonomy.

Research direction is the thin layer. Choosing the right problem requires a model of what matters, what is tractable, what evidence will generalize, which anomalies are meaningful, and when a line of inquiry is exhausted. These judgments are partly learnable; they are also difficult to score. A lab can count experiments and merged code more easily than it can measure whether its portfolio of questions improved. OpenAI's primary-lab PaperBench study makes the gap concrete: on 20 ICML paper-replication tasks scored against 8,316 rubric items, the best tested agent reached 21 percent and remained below the expert-human baseline.15

Successor training is thinner still. A fully closed loop would autonomously design, train, evaluate, secure, and deploy a system that is more capable at closing the next loop. Public systems have improved components of this chain. They have not publicly demonstrated the entire chain at frontier scale.

This ladder supports two simultaneous conclusions.

The skeptical conclusion is that RSI remains unproven. The most persuasive examples are bounded by human objectives, frozen models, fixed evaluators, narrow environments, or modest scales. The hard parts—direction, transfer, evaluation, integration, and governance—are exactly the parts that can dominate total progress.

The accelerationist conclusion is that every hard part is being attacked by a growing machine contribution. The history of AI capabilities contains repeated transitions from “qualitative and human” to “measurable and automatable.” Code quality, tool use, long-horizon planning, and experiment execution all looked resistant until systems improved. Research taste may follow.

Both conclusions are defensible. The error is to convert uncertainty about the top of the ladder into denial of the lower rungs—or to convert success on the lower rungs into certainty about the top.

A rigorous observer should track the ladder as separate time series:

  • share of real AI R&D tasks completed by models;
  • realized productivity uplift, not self-reported uplift;
  • fraction of experiment-selection decisions made without human intervention;
  • evaluator reliability under distribution shift;
  • percentage of discovered improvements retained at frontier scale;
  • cycle time from idea to deployed successor;
  • capability transfer from AI-specific tasks to broader science and production;
  • compute, energy, and review consumed per accepted improvement.

The final metric is crucial. A system can generate more progress while becoming less efficient. If accepted improvements grow by 50 percent but inference and evaluation costs grow by 200 percent, the loop may be strategically useful and economically subcritical. The right object is not raw output. It is verified improvement per unit of scarce input.


II. THE PRODUCTION FUNCTION#

4. When ideas become reproducible labor#

Economic growth depends on more than accumulating machines and workers. Persistent growth in living standards comes from better ways of combining inputs: ideas, technologies, institutions, and organizational knowledge. Growth theory has long wrestled with a peculiar property of ideas: they are costly to discover but cheap to reuse. One design can guide many factories. One algorithm can run on many machines. This non-rivalry creates increasing returns, spillovers, and the possibility that the production of ideas behaves differently from the production of ordinary goods.1617

AI adds a second peculiarity. The artifact produced by research can participate in research.

A stylized ordinary innovation process is:

A˙=f(HR,KR,D,E)

where A is knowledge or algorithmic efficiency, HR is human research labor, KR is research capital and compute, D is data, and E is the experimental environment.

A recursive process adds machine research labor that depends on capability:

A˙=f(HR+M(C,KI),KR,D,E,V)
C=g(A,T,S)

Here M is effective machine research effort produced by capability C and inference compute KI; T is training compute; S represents system design and scaffolding; and V is verification capacity. Better algorithms raise capability. Better capability raises machine research effort. Machine research effort may improve algorithms again.

The economic novelty lies in the replication function. Human research labor is produced slowly through birth, education, selection, apprenticeship, and organizational learning. Machine research labor can be instantiated by allocating compute, copied across tasks, paused, checkpointed, and supplied around the clock. Its marginal cost may be high, especially at the frontier, but its production is not biologically rate-limited.

This does not make machine labor free. It converts a labor constraint into a capital-and-energy constraint.

The distinction resembles earlier mechanization, but the target is different. Steam engines multiplied physical force. Computers multiplied calculation. Networked software duplicated routines. RSI targets the process that discovers new routines. Once a system can contribute to the design of algorithms, chips, experiments, tools, and organizations, the boundary between capital and labor blurs. A model is a capital asset that supplies services resembling labor; those services can improve the capital asset class itself.

The first-order effect is a lower price of cognitive effort. The second-order effect is a higher rate of improvement in the systems supplying cognitive effort. The third-order effect is a reallocation of capital toward every complement that the accelerating cognition consumes.

The loop-gain condition#

Consider a simple cycle:

CRAC

Capability C increases effective research effort R. Research effort increases algorithmic progress A. Algorithmic progress increases capability. For small changes, each arrow has an elasticity: the percentage response of the next variable to a one-percent change in the previous one. The product of elasticities around the loop approximates its gain:

G=εR,CεA,RεC,A

This is a simplification; real systems contain multiple loops, delays, substitution, congestion, and diminishing returns. But it isolates the central question. If G is small, AI assistance raises the level or growth rate of progress without making acceleration self-sustaining. If G approaches one, each improvement replenishes much of the force that created it. If G exceeds one over a relevant region, the growth rate can itself grow until a neglected constraint intervenes.12

The most direct 2026 economic treatment of RSI reaches the same organizing result through a sequence of directed-graph models. Cunningham and coauthors distinguish narrow AI-R&D capability from broad economic capability and attempt a provisional calibration. Under their chosen units and assumptions, a one-unit improvement in capability would need to raise AI R&D productivity by roughly 15 percent to meet the self-sustaining condition; their rough estimate of the observed return was closer to 9 percent. They therefore judge the public evidence subcritical but strengthening. The figures are fragile—the productivity inputs are sparse and partly self-reported—but the exercise is valuable because it converts an argument about vibes into an empirical threshold that can be updated.18

FIGURE 5A pedagogical simulation of subcritical, near-critical, and supercritical feedback.

The threshold is not a universal constant. It depends on the unit of capability, the time interval, the production structure, and which constraints are held fixed. It may be above one in a coding sub-loop and below one for the whole laboratory. It may be temporarily above one while low-hanging improvements remain, then fall as the search space hardens. It may be below one at the frontier but above one for diffusion, where models automate deployment and customization rather than fundamental research.

The gain can also be masked by delays. An algorithmic discovery today may affect a training run six months later. A chip-design improvement may enter hardware three years later. A grid connection may arrive after the model generation that justified it is obsolete. A system with strong long-run feedback can look slow in calendar time because its loops traverse physical infrastructure.

Five multipliers and five leakages#

A recursive loop becomes stronger through five multipliers.

Parallelism. Machine workers can explore many hypotheses simultaneously. The benefit is largest when experiments are independent and evaluation is cheap.

Copyability. A successful procedure can be propagated across every agent. Organizational learning becomes software deployment.

Speed. Agents can operate continuously and reduce latency between steps.

Search breadth. Cheap candidate generation increases the probability of finding rare improvements.

Meta-improvement. The system can improve tools, evaluators, prompts, memory, data pipelines, and experimental design—not only task outputs.

The loop leaks through five channels.

Diminishing returns. Easy improvements are exhausted. Each additional unit of research effort produces less progress.

Congestion. Experiments compete for compute, data, reviewers, and shared infrastructure.

Evaluation drag. Candidate generation outruns the capacity to verify quality.

Transfer loss. Improvements in a benchmark or subsystem fail to generalize.

Coordination and safety overhead. Larger agent populations create integration, security, monitoring, and governance costs.

The contest between multipliers and leakages determines whether RSI is a productivity tool, a sustained acceleration, or a transient boom.

Effective compute is not enough#

AI progress is often summarized as “effective compute”: physical compute multiplied by algorithmic efficiency. This is useful, but recursive systems introduce another factor—research allocation. A better model can make the same compute produce more experiments, better code, and better decisions. Yet it can also consume enormous inference compute while searching, judging, and coordinating.

The loop therefore has two compute bills. Training compute builds the successor. Inference compute operates the virtual laboratory that designs the successor. As agents take on more research work, inference ceases to be merely the cost of serving customers; it becomes a capital input to innovation.

This changes the economics of scaling. A laboratory may rationally devote a growing share of compute to internal agents if their contribution to algorithmic progress exceeds their opportunity cost. The optimal allocation becomes endogenous: more capable models justify more inference; more inference may produce better methods; better methods increase the return to the next training run.

At the same time, the recursion can improve efficiency and reduce compute per capability unit. AlphaEvolve-style algorithm search, kernel optimization, sparsity, quantization, better data, and architecture discovery can move the frontier inward. Recent research estimates very rapid declines in the price of achieving a given benchmark level, with algorithmic efficiency making a material contribution.19 The demand for compute is therefore the product of two opposing forces: efficiency lowers the cost of a given capability, while cheaper capability expands usage and makes more ambitious search economical.

Jevons-style rebound is not guaranteed, but it is plausible. In a recursive economy, efficiency gains can increase total resource demand because they raise the return to running more intelligence.

This is why an AI boom can be both deflationary in digital services and inflationary in power, land, and infrastructure.


III. TIME AND SCARCITY#

5. Economic time compression#

Technological progress is usually narrated as a sequence of discoveries. Economies experience it as a sequence of delays.

A useful idea must be understood, financed, staffed, implemented, tested, approved, distributed, and incorporated into routines. The interval between technical possibility and measured productivity can be long because organizations are not frictionless. The computer age produced a famous productivity paradox: the machines were visible everywhere except, for a time, in aggregate productivity statistics. Complementary investments in processes, skills, management, and capital had to accumulate before the technology reorganized production.20

RSI attacks several of those delays at once.

An agent can read the literature, write the code, provision the experiment, inspect the result, draft the report, and propagate a working procedure. Multiple agents can do this in parallel. The artifact can be deployed instantly to the rest of the virtual workforce. The cycle from idea to operational knowledge can shrink from months to days, or from days to hours, wherever the environment is digital and the evaluator is available.

This is the first macroeconomic signature of recursion: economic time compresses before physical output explodes.

The hypothesis interval#

When candidate ideas are expensive, organizations ration exploration. Researchers spend time deciding which experiments are worth attempting because each one consumes scarce human effort. Cheap machine generation relaxes that rationing. The research frontier becomes a portfolio problem: generate many plausible interventions, estimate their expected value, and allocate experimental capacity.

The gain is not that every idea is good. It is that the search distribution becomes wider. Rare high-value ideas can be found through volume—provided evaluation can distinguish them from plausible nonsense. In domains with formal or automated feedback, the hypothesis interval can collapse dramatically.

The implementation interval#

Many discoveries are delayed by the work required to turn an idea into a runnable artifact. Code must be written, data cleaned, infrastructure configured, and edge cases handled. This layer is increasingly machine-addressable.

Implementation was once a tax on every idea. As it falls, the value of deciding which ideas to implement rises.

The evaluation interval#

More candidates create more demand for tests. Where evaluation is cheap and objective, the whole loop speeds up. Where evaluation requires expert review, real-world deployment, or long observation, the bottleneck hardens.

This creates an evaluator economy. Test suites, formal methods, simulators, provenance systems, process monitors, scientific instruments, benchmark design, red-teaming, replication platforms, and trusted human judgment become productive capital. Verification is not a compliance layer added after innovation. It is a throughput constraint inside the innovation function.

The diffusion interval#

Ordinary organizational knowledge diffuses through meetings, documents, training, and turnover. Machine knowledge can diffuse through a software update. A procedure discovered by one agent can be added to a shared library and made available to every instance. The cost of organizational replication falls.

This can make firms internally more homogeneous and externally more unequal. The organization with the best shared agent memory, evaluators, proprietary feedback, and deployment pipeline can compound learning faster than a rival using the same base model. Model weights may commoditize while organizational context remains defensible.

The capital-allocation interval#

When the rate of technical change rises, capital must be reallocated faster. A project justified by one generation of hardware may be obsolete before construction finishes. A software firm can redirect agents in an afternoon; a utility cannot relocate a power plant. A lender underwriting a twenty-year asset faces a technology path that may change the customer's economics every quarter.

This is why recursion can increase both the value of flexibility and the value of long-lived bottleneck assets. Modular data centers, interruptible load, transferable power contracts, adaptable factories, and general-purpose robotics gain option value. At the same time, scarce grid nodes, fabs, and permitting rights may become more valuable because they cannot be reproduced at digital speed.

Markets under compressed time#

Financial markets are forward-looking, but they process information through human institutions: research cycles, earnings reports, investment committees, regulation, and settlement. If technological cycles shorten below reporting cycles, the mapping from current financial statements to future cash flows becomes unstable.

A company may move from technological leader to laggard between quarterly reports. The useful life of proprietary software can shrink. The dispersion of outcomes can widen because small differences in feedback infrastructure compound quickly. Narrative can outrun evidence, especially when internal productivity is unobservable and firms selectively disclose dramatic examples.

This environment favors three kinds of information.

First, physical commitments: signed power contracts, construction progress, chip deliveries, and interconnection capacity are harder to fake than capability narratives. Second, workflow evidence: randomized or auditable measures of output, quality, and cycle time are more informative than benchmark scores. Third, bottleneck evidence: queues, utilization, review latency, experiment backlogs, and rejected improvements reveal where the loop is actually constrained.

The market will repeatedly confuse the acceleration of one stage with acceleration of the whole system. It will overvalue the newly automated input, then discover that abundance destroyed its rent. It will undervalue the complement, then reprice it when the bottleneck arrives. The recurring investment pattern is not simply “AI goes up.” It is a sequence of scarcity rotations.

The productivity-statistics lag#

Aggregate productivity may remain modest even while frontier organizations transform. Adoption is uneven. High-performing users learn to delegate better tasks and build complementary processes; inexperienced users may spend time checking low-quality output. Anthropic's 2026 Economic Index reports broadening usage, learning effects, and a continued distinction between augmentative and automated use.21 Its economic-primitives analysis estimates that reliability adjustments can materially reduce headline productivity estimates derived from raw task speedups.22

Measurement also misses free or low-price digital output. A model that produces a useful report in seconds can create consumer surplus while contributing little measured revenue. Meanwhile, the capital expenditure needed to build the intelligence infrastructure is recorded immediately.

Finally, the macroeconomy includes sectors tied to slower clocks. Faster software does not instantly build housing, expand transmission, complete trials, or settle litigation. The productivity boom can be real upstream and muted downstream.

This lag is not evidence against recursion. It is evidence that recursion is colliding with the rest of the economy.


6. The law of migrating scarcity#

Every technology is initially described by what it makes possible. Every mature industry is organized around what remains difficult.

The internet made information transmission abundant and shifted value toward attention, trust, distribution, and proprietary data. Cloud computing made servers elastic and increased the value of software, chips, network effects, and scale. Cheap generation of text, code, images, and designs will not make the economy post-scarcity. It will move the scarcity boundary.

The central law of the recursive economy is:

Recursive self-improvement eliminates bottlenecks in sequence until it reaches inputs it cannot cheaply reproduce. Capital and political conflict accumulate at that boundary.

Call that boundary the recursion frontier.

At any time, the frontier is the set of constraints whose marginal relaxation most increases verified, deployable capability or real output. It is not necessarily one input. A laboratory may be jointly constrained by power, evaluators, and security. A pharmaceutical company may have abundant candidate molecules but scarce trial participants and regulatory capacity. A city may have cheap design intelligence but scarce land and construction labor.

A useful shorthand is:

Bt=argmaxiYt/Xiexpandability(Xi)

The binding bottleneck Bt is the input with a high marginal product and low short-run expandability. RSI raises the marginal product of complements and changes expandability over time.

FIGURE 6An illustrative migration of binding constraints.

From researchers to direction#

The first scarcity is skilled human execution. Agents reduce it by writing code, analyzing results, and operating tools. The remaining human work moves upward: selecting goals, allocating resources, setting standards, interpreting anomalies, and accepting responsibility.

This does not mean managers become infinitely valuable. Direction itself can be decomposed. Some choices are repeated, measurable, and learnable. Portfolio selection can be trained on past outcomes. Agent judges can compare proposed experiments. Markets and prediction systems can aggregate beliefs. As those mechanisms improve, the frontier moves again.

From direction to evaluation#

When implementation becomes cheap, candidate output explodes. The scarce input becomes confidence.

Software repositories receive more patches than maintainers can review. Security systems find more vulnerabilities than organizations can remediate. Scientific agents generate more hypotheses than laboratories can test. Legal agents produce more arguments than courts can process. Media systems create more claims than audiences can authenticate.

The price of generation falls; the price of trusted acceptance rises.

This is not merely a safety problem. It is an economic bottleneck. An unverified improvement has option value, not production value. Firms that own high-quality feedback—customer behavior, operational telemetry, test environments, scientific instruments, or trusted expert networks—can convert cheap cognition into reliable output more effectively than firms that own only a model endpoint.

The 2026 RSI survey reaches a similar conclusion from the technical side: demonstrated self-improvement is strongest where verification is formal or external, and weakest where systems rely on intrinsic self-assessment.4

From evaluation to compute#

If agents become reliable enough to operate large research loops, inference demand rises. A virtual laboratory can consume compute continuously: proposing, simulating, coding, judging, and coordinating. Training remains large and episodic; research inference becomes persistent.

Compute scarcity is not only a shortage of accelerators. It includes memory, interconnect, networking, storage, cooling, uptime, scheduling, and the software needed to use hardware efficiently. Algorithmic efficiency can relax the constraint, but lower unit cost may expand the economically viable population of agents.

The relevant metric becomes verified progress per joule and per dollar—not tokens per second.

From compute to chips and power#

Compute capacity is embodied in hardware. Hardware is embodied in fabs, packaging, materials, equipment, and geopolitical supply chains. It is then embodied again in data centers that require land, transformers, cooling, fiber, and firm power.

The International Energy Agency estimated global data-center electricity consumption at roughly 415 terawatt-hours in 2024 and projected around 945 TWh by 2030 in its base case. Its base case reaches roughly 1,200 TWh by 2035, with wide sensitivity around efficiency, adoption, and supply constraints.23 A 2026 Lawrence Berkeley National Laboratory update put US data-center demand in 2030 at 649 TWh in its reference case, with a range of 521–843 TWh.24 These are not forecasts of RSI. They illustrate the scale at which digital demand encounters a physical system whose lead times are measured in years.

FIGURE 7IEA data-center electricity anchors and scenarios.

Power is not globally fungible at the socket. A megawatt in the wrong region, without transmission, interconnection, or reliability, cannot serve a cluster. The economics are nodal and temporal. At the end of 2025, US interconnection queues still contained 2,061 GW of active projects, and the median request-to-operation time exceeded five years.25 Data centers can be constructed faster than major grid infrastructure.

This is the bridge from Power 2026 to the recursive economy. Power is not important because AI is a fashionable new load. It is important because a self-amplifying demand for cognition converts research ambition into continuous demand for electrons, and the grid cannot recursively copy itself.26

From power to physical capital#

More intelligence can design better factories, robots, medicines, materials, and logistics. But a design is not a factory. Physical capital must be financed, permitted, built, maintained, and supplied.

The relative price of atoms can rise against bits. A building plan may become almost free while construction remains expensive. A robot controller may improve weekly while actuator supply grows annually. A molecule may be discovered quickly while trials and manufacturing remain slow. The more productive intelligence becomes, the more valuable physical execution can become—at least until robotics and automated construction push the frontier further.

This is a general-equilibrium point. Automating architects does not necessarily reduce housing prices if land-use rules, land, financing, and construction capacity bind. Automating engineering can increase demand for scarce technicians, equipment, and permits. The measured productivity of a sector depends on the entire chain, not the intelligence intensity of its first step.

From atoms to time, trust, and legitimacy#

Some constraints are not engineering problems in the ordinary sense.

A clinical trial needs time to observe outcomes. A child needs years to mature. An ecosystem may need decades to recover. A legal system requires due process. A democracy requires consent. A friendship cannot be backfilled with generated history. A community may reject a project even when an optimizer proves it beneficial on a chosen metric.

Recursive intelligence can propose institutional designs, negotiate, simulate, and explain. It cannot assume legitimacy into existence. Indeed, faster change can make legitimacy scarcer by increasing the gap between what institutions can process and what technology can alter.

The final scarcity may be the right to act.

The rent equation#

Economic rents at the recursion frontier depend on three properties:

RentiScarcityi×Controli×1Substitutabilityi

An input can be scarce without generating private rent if it is publicly supplied or price-regulated. It can be controlled without being valuable if substitutes are easy. It can be necessary without being defensible if competition drives price to cost.

This is why the same RSI path can produce different winners.

If frontier models remain proprietary and differentiated, model owners capture rents. If model capability diffuses, rents move to compute, data, integration, distribution, and physical assets. If governments cap utility returns or socialize infrastructure, private power rents may be limited even while power remains the economic constraint. If verification becomes standardized and open, trust rents fall. If regulation creates scarce licenses, legal permission becomes an asset.

FIGURE 8An illustrative migration of economic rents.

The investable insight is not a permanent list of scarce assets. It is a method for tracking the moving boundary.


7. The four clocks#

The phrase “takeoff” encourages a single speed. The economy has at least four.

FIGURE 9Digital acceleration collides with physical, biological, and institutional clocks.

Digital time#

Inference, software changes, simulations, and coordination can occur in seconds to days. Machine workers can be copied and run continuously. Once a procedure is encoded, diffusion is nearly instantaneous.

Digital time is where recursive acceleration is most plausible. The environment is legible, feedback can be automated, and successful changes can propagate without rebuilding the world.

Physical time#

Data centers, fabs, transmission lines, mines, factories, housing, robots, and laboratories take months to years. Supply chains have inventories, qualification cycles, and construction sequences. Permitting and finance add latency.

AI can improve designs and schedules, but many steps remain serial. Concrete must cure. Equipment must arrive. Grid stability must be maintained. The physical clock can speed up at the margin without becoming software.

Biological time#

Medicine, agriculture, ecology, human development, and care are constrained by living processes. Better models can select trials and monitor outcomes, but some evidence requires waiting. A twelve-month intervention cannot be validated in twelve minutes merely because the hypothesis was generated quickly.

Biological time is why intelligence abundance does not immediately abolish disease or aging. Discovery can accelerate far ahead of proof and deployment.

Institutional time#

Law, standards, professional norms, diplomacy, public consent, and organizational trust change through contested processes. Some slowness reflects incapacity; some protects rights and legitimacy.

Institutional time is the least amenable to brute-force parallelism. Ten thousand policy proposals do not create ten thousand times more political agreement. In fact, machine-generated persuasion and lobbying can congest the process.

The collision#

Real output is limited by the slowest necessary clock:

grealizedmin(gdigital,gphysical,gbiological,ginstitutional)

This is not a literal production function. It is a warning against extrapolating software speed into every domain.

The clocks also create price wedges. Digital goods deflate relative to slow-clock goods. Research artifacts become cheap; clinical capacity becomes dear. Design becomes cheap; land remains scarce. Policy analysis becomes cheap; political legitimacy remains scarce.

The recursive economy is therefore not uniformly fast. It is violently asynchronous.

That asynchrony shapes nearly every transition risk: speculative overbuilding, regulatory backlog, unequal diffusion, labor displacement before physical abundance, and political conflict over who receives scarce real-world capacity.


IV. OWNERSHIP AND DISTRIBUTION#

8. Who owns the recursion?#

Technology determines what can be produced. Institutions determine who can claim the output.

Ownership is often treated as a distributional question to be addressed after growth. Under RSI, ownership affects the growth path itself. The holder of models, compute, data, evaluators, and deployment channels decides which experiments run, which objectives are optimized, which improvements are shared, and how quickly the loop diffuses.

The core asset is not a model snapshot. It is the recursive production complex:

  • frontier or near-frontier models;
  • inference and training compute;
  • proprietary operational data;
  • evaluation environments;
  • agent memory and organizational context;
  • research workflows and talent;
  • energy and hardware contracts;
  • security, governance, and deployment rights.

A competitor may copy one layer and still fail to reproduce the loop. Conversely, an open model can erode rents if complementary infrastructure is widely available and improvements diffuse faster than incumbents can compound them. OECD analysis of AI infrastructure identifies high fixed costs, scarce inputs, vertical integration, and cross-holdings as concrete entry barriers across this stack.27

Two axes organize the possible equilibria: the strength of feedback and the degree of capability diffusion.

FIGURE 10Four economic equilibria.

Fortified oligopoly: weak recursion, concentrated ownership#

In this world, AI remains powerful but the feedback loop stays below the self-sustaining threshold. Frontier development requires vast capital and proprietary know-how. A few firms control the best models and distribution, but progress follows a recognizable industrial cadence.

This resembles a more concentrated cloud-and-semiconductor economy. Firms earn rents from scale, switching costs, integrated platforms, and access to customers. Labor productivity rises, especially for complementary workers. Model cycles are fast but not explosive. Regulators have time to act, although market power may already be entrenched.

The dominant policy questions are familiar: antitrust, interoperability, data access, procurement, and labor adjustment.

Abundant tools: weak recursion, rapid diffusion#

Here the feedback loop remains bounded, but capable models become cheap and widely available. Intelligence behaves like a general-purpose input. Small firms and individuals gain access to strong cognitive tools. Application innovation flourishes. The frontier still matters, but much of the social value comes from diffusion rather than recursive acceleration.

Model rents compress. Distribution, domain data, trust, and execution matter more. Productivity gains may be broad but gradual because organizational adoption and physical complements take time. This is the most benign continuation of the software era: extraordinary tools, ordinary macroeconomic adjustment.

Stanford's 2026 AI Index describes rapid adoption and investment alongside persistent measurement and governance gaps.28 Anthropic's usage data likewise shows broad task penetration without evidence that whole occupations have already been automated.21 These facts support a diffusion story; they do not establish recursive gain.

Recursive sovereigns: strong recursion, concentrated ownership#

This is the high-stakes equilibrium. A small number of laboratories or state–corporate complexes cross the gain threshold while keeping the loop closed. They command a stock of machine research labor that compounds faster than outsiders can imitate.

The asset begins to resemble strategic territory. The owner can improve cyber capability, weapons systems, industrial planning, persuasion, science, and further AI. Even if the system remains aligned with its operator, the power imbalance can be extreme.

Market concentration may become endogenous. A slight lead produces better research agents; better agents improve the next model; the next model widens the lead. Capital markets reinforce the loop by funding the perceived winner. Suppliers prioritize the largest buyer. Talent joins the frontier. Data and deployment feedback deepen. The result is cumulative advantage rather than ordinary scale economy. The FTC's examination of major cloud–AI partnerships documents mechanisms already capable of reinforcing such positions, including cloud-spending commitments, control and information rights, switching costs, and privileged access to inputs and talent.29

Yet concentration is not guaranteed. Frontier systems can leak. Employees move. Techniques are published or rediscovered. Hardware is commercially produced. Governments can compel access. Open systems may lag briefly and then converge. The durability of recursive rent depends on the secrecy and reproducibility of algorithmic progress, the capital intensity of training, and the speed of diffusion.

A concentrated strong loop also faces internal diseconomies. Security grows harder. Oversight becomes a bottleneck. A laboratory may generate more ideas than it can safely evaluate. Bureaucracy can absorb machine speed. The leading organization may be technically recursive and institutionally slow.

Recursive commons: strong recursion, rapid diffusion#

In this world, strong self-improvement methods spread. Models, agent frameworks, synthetic curricula, and algorithmic discoveries become broadly accessible. The price of cognitive services collapses toward compute cost.

This is the closest route to intelligence abundance—and it does not imply equal outcomes. Those who own compute, energy, land, robots, brands, customer relationships, and legal rights may capture the surplus. A small firm can access excellent intelligence, but so can every competitor. The differentiated asset moves downstream.

The recursive commons could generate explosive entry. One person can coordinate thousands of agents. Research communities can fork and improve shared systems. Countries without frontier laboratories can import capability. Knowledge-intensive services become tradable at low cost.

It could also produce severe security and coordination problems. Dangerous capabilities diffuse. Verification becomes decentralized. A race to deploy can outrun safety. The public-good character of ideas collides with the private cost of compute and the external cost of misuse.

Why the equilibrium can change#

The economy can move between quadrants.

A concentrated lab may pioneer strong recursion, then see methods diffuse. Governments may nationalize or regulate access. A technical plateau may push the system from strong to weak feedback. A hardware embargo may concentrate capability. An efficiency breakthrough may democratize it. Safety incidents may shift deployment from markets to sovereign control.

The transition path matters because early ownership shapes later institutions. A state that waits until recursion is obvious may find the relevant assets already concentrated. A regulator that freezes the wrong architecture may entrench incumbents. An open-source mandate may broaden access while increasing misuse. There is no distribution-neutral way to govern the loop.

The practical question for every institution is therefore not merely “How capable is AI?” It is “Who can run the improvement cycle, at what scale, with what feedback, and under whose rules?”


9. The end of labor scarcity is not the end of work#

Labor economics is often forced into two slogans. Technology either destroys jobs or creates new ones. History is then used as reassurance: past automation displaced tasks but expanded demand, created occupations, and raised wages. US evidence through 2018 finds that most employment was in job specialties introduced after 1940, while also distinguishing labor-augmenting new work from task-displacing automation.30

The historical record matters, but RSI changes one premise. Past machines generally automated particular physical or cognitive tasks while humans retained broad comparative advantage in discovering, coordinating, and adapting to new tasks. A system that contributes to its own improvement can expand the set of automatable tasks from inside the innovation process.

This does not make human labor worthless. It makes the wage path conditional.

The four races#

The distributional outcome depends on four races.

Automation versus new-task creation. AI replaces human effort in existing tasks, while innovation creates new tasks in which humans may hold advantage. This is the familiar task-based race.3130

Automation versus capital accumulation. Even if machines can perform a task, there must be enough machines, compute, robots, and infrastructure to deploy them. Wages can remain high during a transition because automated capital is scarce. Models of transformative AI show that the path depends critically on whether capital accumulation keeps pace with the expanding automation frontier.32

Capability versus diffusion. Frontier automation affects few workers if it remains expensive, unreliable, or organizationally difficult. Cheap, standardized capability diffuses faster and exerts broader wage pressure.

Production versus redistribution. Output can rise while market income becomes concentrated. Taxes, transfers, public ownership, pensions, dividends, and universal services determine whether households retain purchasing power.

The recursive loop accelerates the first race and can accelerate the second by improving hardware, robotics, and capital allocation. It may also accelerate diffusion through cheaper models and automated integration. Policy governs the fourth.

A task is not a job#

Jobs bundle tasks with responsibility, trust, relationships, physical presence, legal status, and organizational context. AI can automate a high share of tasks without eliminating the occupation. It can also eliminate headcount without automating every task if one worker using agents handles the residual bundle for many former workers. The ILO's task-level exposure index estimates that one quarter of global employment has some generative-AI exposure, but concludes that transformation is more likely than full automation because most occupations retain tasks requiring human input.33

The unit of adjustment may be the team rather than the job.

A software group of twenty might become a group of five supervising agents. A consulting engagement that required ten analysts might require one partner and a verification layer. A marketing department can retain brand judgment and accountability while automating production. A research laboratory can increase experiments without increasing staff. Employment falls, stays flat, or rises depending on demand elasticity and the speed at which new uses appear. In a 2026 Census firm survey, 66 percent of AI adopters reported augmentation without automation and 2 percent reported job cuts, while adoption remained concentrated among large, knowledge-intensive firms.34

Early evidence is mixed because the technology and workflows are moving. Anthropic's labor-market work distinguishes theoretical exposure from observed use and from the degree to which use is actually automated.35 METR's developer experiments demonstrate how realized productivity can differ from expectations, while later measurements suggest the sign and magnitude can change with newer tools.10 Controlled studies have found meaningful average gains in particular settings: 15 percent more issues resolved per hour among 5,172 customer-support agents, and 40 percent faster completion with 18 percent higher quality across a set of professional writing tasks.3637 Yet Danish survey and administrative data found no significant average effect on earnings or hours within two years of adoption, despite rapid workflow change.38

The honest conclusion is not that AI has no productivity effect or that it has already transformed aggregate labor. It is that complementarity, user skill, task structure, and evaluation determine realized gains. RSI matters because the tools themselves are improving the workflow faster than organizations can settle into one measured equilibrium.

The wage distribution#

Human wages can survive automation for at least six reasons.

Complementarity. Some humans become more productive by steering, integrating, or legitimizing machine work. In both the customer-support and professional-writing experiments, initially weaker or less-experienced workers received larger gains.3637

Embodiment. Physical tasks remain costly to automate until robotics and deployment mature.

Preference. Consumers may pay for human performance, care, craft, status, or authenticity.

Liability. Law may require a human decision-maker or accountable professional.

Scarcity and positionality. Leadership roles, elite competition, land-linked services, and social status are inherently limited.

Ownership. A person may receive income as an owner even when their labor is not scarce.

These channels do not ensure a stable aggregate labor share. Complementarity can be concentrated among a small group. Embodiment can erode. Legal requirements can change. Human authenticity can support niches without employing the population at current wage levels. The ILO estimates that the global labor-income share fell from 53.0 percent in 2014 to 52.4 percent in 2024.39 In US industries, larger increases in concentration have also been associated with larger labor-share declines through reallocation toward high-profit, low-labor-share firms.40

A toy model makes the separation visible. Suppose the share of tasks performed by machines rises under three capability paths, while a residual share of output continues to reward human complementarity. Output can increase in every path, but labor's share falls fastest when automation is broad, capital is available, and human complementarity is weak.

FIGURE 11Illustrative labor-share paths under three scenarios.

The figure is not a forecast. Its purpose is to expose the assumption hidden in many optimistic narratives: that human complementarity expands quickly enough to preserve labor income even as machines become broadly capable. That may happen. It is not an identity.

The disappearing firm—and the growing complex#

Ronald Coase explained the firm as a response to transaction costs. Activities are organized inside a hierarchy when coordinating them through the market is more expensive.41 Agents reduce some transaction costs: search, contracting, monitoring, communication, translation, and routine management. This can shrink the efficient size of the application-layer firm.

A founder may be able to command design, engineering, sales operations, research, and support agents without hiring a traditional organization. The firm becomes a thin legal and strategic shell around an agentic production graph. Entry can increase because the fixed cognitive cost of starting falls. In a preregistered experiment with 791 professionals at Procter & Gamble, individuals working with AI matched teams working without AI on the measured product-development task and crossed functional silos, although human judgment retained value in selecting ideas.42

At the same time, the frontier intelligence complex grows. Training systems, fabs, data centers, power procurement, networking, security, and global supply chains exhibit enormous fixed costs and coordination needs. The minimum efficient scale of cognition falls at the edge while the minimum efficient scale of infrastructure rises at the core.

One possible resulting industrial structure is a barbell:

  • a small number of immense infrastructure and model complexes;
  • a large population of small, agent-leveraged firms;
  • pressure on the traditional middle of labor-intensive professional organizations.

This is a scenario, not a description of the current economy. Present adoption evidence points toward larger firms: the Census finds AI use concentrated among large knowledge-intensive firms, while research on firm-level AI investment associates adoption with more educated and STEM-heavy workforces and fewer middle and senior managers.3443 Middle-sized firms can survive through proprietary distribution, regulation, trusted brands, physical assets, or unique data. But if agents substantially reduce coordination costs, the old rationale—assembling many humans to coordinate knowledge work—weakens.

Management after execution#

Management does not vanish when agents execute. It changes from allocating human time to specifying objective functions, constraints, escalation rules, and verification budgets. Evidence that AI-assisted individuals can cross functional boundaries, alongside observed changes in managerial composition at AI-investing firms, makes this an organizational hypothesis worth testing rather than an established endpoint.4243

The best manager may be the person who designs the organization's evaluator. What counts as a good customer outcome? Which errors are intolerable? When must an agent ask for review? How should short-term output trade against long-term learning? Which data can be trusted? Who owns the consequences?

These are governance questions embedded in production.

Poorly specified objectives can generate output at enormous scale and destroy value at enormous scale. Good organizations will not simply “use more agents.” They will build feedback architectures that make errors legible, preserve institutional memory, and route uncertainty to the right human or machine process.

Income without wages#

If labor ceases to be the principal scarce input, wage income becomes a weaker distribution mechanism. The economy then needs other claims on output.

These can include broad capital ownership through pensions or sovereign funds; citizen dividends from public stakes in compute, energy, or model infrastructure; taxes on rents, land, consumption, or capital income; universal basic income; negative income taxes; universal basic services; and new forms of data or intellectual-property participation.

None is automatically fair or efficient. Capital taxation can discourage investment if poorly designed. Consumption taxes can be regressive without transfers. A universal dividend can be politically fragile. Public ownership can be inefficient or captured. IMF analysis therefore favors broad social protection and stronger capital-income taxation while cautioning against a narrow tax on AI itself.44 A tax state funded heavily by labor-linked revenue is exposed if labor's share falls.

The transition can be more destabilizing than the endpoint. Households hold mortgages and make education decisions based on expected wages. Governments issue debt based on expected tax receipts. Cities are built around commuting professionals. Many contributory pension systems rely on employee and employer contributions. A rapid shift in the income source can create solvency problems even if aggregate output is rising.45

The paradox of the recursive economy is that abundance can arrive upstream while insecurity rises downstream.


V. CAPITAL, GEOGRAPHY, AND THE STATE#

10. The price of capital in a world of abundant ideas#

Growth theory usually asks how capital deepening and technological progress raise output. RSI introduces a more unstable question: what happens when the supply of investable ideas expands faster than the stock of machines, energy, and institutions required to embody them?

The first effect may be a surge in the demand for capital.

A laboratory that discovers better algorithms wants more training runs. A firm that can design thousands of products wants more factories, robots, and distribution. A medical system that can identify more promising interventions wants more laboratories, trials, clinicians, and manufacturing. Cheap cognition does not automatically reduce investment. It can multiply the number of projects with positive expected value.

This is the capital hunger of the recursive economy.

Why real interest rates could rise#

If the loop raises the marginal product of reproducible capital faster than it raises saving, firms compete for funds. Real interest rates rise. The transition becomes an investment boom: data centers, grids, fabs, laboratories, robots, and industrial capacity bid against one another for equipment and labor.

Company disclosures show the scale of the current build-out, though they are not independent demand forecasts. Alphabet reported $91.4 billion of capital expenditure in 2025, guided to $175–185 billion in 2026, and identified power, land, and supply-chain availability as constraints.46

Models of transformative AI commonly find that the trajectory depends on the speed of capital accumulation. When increasingly capable machines can substitute for labor, a temporary scarcity of machines can preserve wages and slow automation; rapid accumulation can unlock much faster growth.32 The interest rate is therefore not a side effect. It is one signal of whether finance is keeping pace with capability.

An RSI boom could combine high real rates with falling prices for many digital goods. There is no contradiction. Software, analysis, and design can become cheaper while firms pay dearly to bring future physical production forward.

Why real interest rates could fall#

The opposite path is also coherent.

Machine intelligence may improve the production of capital goods, reduce construction and engineering costs, and increase the efficiency of existing assets. Profits may accrue to owners with high saving rates. If capability expands the supply of investable resources faster than it creates useful projects—or if physical and regulatory bottlenecks prevent projects from being realized—savings can outrun investment. Real rates fall.

The transition may contain both phases. Rates can rise during the race to build, then fall after capacity catches up or overbuild becomes visible. The recursive economy does not imply one permanent interest-rate regime; it implies greater sensitivity to the relative speed of idea production, capital formation, and institutional permission.

The obsolescence premium#

Fast technical progress shortens the economic life of assets tied to a particular model generation or architecture. A cluster can remain physically functional while becoming commercially uncompetitive. A proprietary workflow can be reproduced. A robot platform can be stranded by a new control paradigm. A software company can lose differentiation between quarterly reports.

Investors therefore demand an obsolescence premium: compensation for the risk that the frontier moves before the asset pays back.

The premium is highest for inflexible, specialized assets and lowest for assets complementary to many futures. A model-specific serving stack may depreciate quickly. A well-connected industrial site, transferable power contract, grid interconnection, fiber route, or general-purpose laboratory may appreciate because successive technologies continue to require it.

This distinction—solution-specific capital versus frontier-complement capital—is more useful than the label “AI infrastructure.”

Intangible depreciation#

Models, code, organizational procedures, and data pipelines are capital, but their economic depreciation can occur at machine speed. Firms may invest heavily merely to maintain position. Gross investment rises while net advantage remains flat.

This can produce a misleading financial pattern. Revenue grows. Measured productivity improves. Yet free cash flow disappoints because every generation requires new training, migration, evaluation, and security. A company with a spectacular product may still be a poor asset if competition forces it to pass the intelligence surplus to users while it bears the maintenance cost of staying near the frontier.

Conversely, a boring complement can capture durable rent if its supply is slow and its demand expands across model generations.

Deflation in bits, inflation at the boundary#

RSI should create strong relative-price movements.

Digital goods with low replication cost—routine code, analysis, tutoring, translation, design variants, legal drafts, entertainment, and research artifacts—face deflationary pressure. The quality-adjusted price can collapse even if nominal subscription prices remain stable.

Inputs at the recursion frontier face inflationary pressure: accelerators, advanced packaging, power, transformers, grid access, cooling, data-center land, high-quality sensors, scientific instruments, trusted evaluators, specialized construction, and permits. Current US demand and reliability assessments support treating power as a live constraint: NERC's 2025 assessment projected summer peak demand to rise by 224 GW over a decade and identified data centers as the main source of the higher forecast.47 If robotics lags, skilled physical labor can also become more valuable.

The aggregate price index averages these movements and can conceal the lived economy. A household may receive extraordinary cognitive services almost free while paying more for housing, care, energy, and physical experiences. Abundance and scarcity can intensify together.

The mechanism is stronger than ordinary sectoral productivity. Cheap intelligence can increase demand for slow-clock goods. Cheap design produces more projects. Cheap diagnosis produces more treatments to consider. Cheap scientific hypotheses produce more experiments. Cheap code produces more computation. Intelligence abundance bids up whatever converts decisions into reality.

Asset pricing at the recursion frontier#

A static list of “AI beneficiaries” is unlikely to survive a moving bottleneck. An asset benefits durably only if four conditions hold:

  1. it complements expanding machine cognition;
  2. its supply cannot expand as quickly as demand;
  3. its owner can capture the scarcity rent rather than pass it through;
  4. it is not bypassed by the next efficiency or architecture change.

Fail one condition and a compelling technological thesis can produce poor returns. Electricity demand may rise while regulated utilities earn ordinary returns. Chip demand may rise while competition compresses margins. Model usage may explode while open systems commoditize inference. A data center may be scarce at one grid node and stranded at another.

The recursive-economy investor asks a different question:

After intelligence attacks the current bottleneck, what becomes binding next—and is that next constraint privately monetizable?

The answer will change. That is the point.

Revolution and bubble can coexist#

Uncertainty about feedback creates enormous valuation dispersion. A small change in assumed loop gain produces a large change in future cash flows. Internal R&D acceleration is difficult for outsiders to observe. Firms disclose dramatic examples more readily than failed experiments, review costs, or integration delays. Alphabet's capital-spending guidance, for example, is a primary company disclosure about its own plans rather than an independently validated forecast of market demand.46

A rational investment boom can therefore contain a bubble. Infrastructure is built for demand that may arrive; some assets fail even if the aggregate thesis succeeds. Railways, electrification, and fiber created durable social value alongside bankruptcies and overcapacity.

RSI can produce the same combination at greater speed. The analytical mistake is to choose between “bubble” and “revolution.” Both may be true, and the losses may finance the infrastructure through which the revolution diffuses.


11. The geography of machine intelligence#

Human capital has geography because people live somewhere. Machine intelligence has geography because compute, power, networks, law, and ownership live somewhere.

At the application layer, AI makes distance less important. Cognitive services cross borders instantly. Translation reduces language barriers. A small firm can import a capable model through an API. At the frontier, location becomes more important because the digital loop is anchored in physical clusters.

Compute geography#

Large clusters require power, land, cooling, fiber, water or alternative thermal systems, security, and favorable permitting. These inputs are spatially uneven. The World Bank estimates that high-income countries hold 77 percent of global data-center capacity, while low-income countries hold less than 0.1 percent.48 Electricity is especially local. National generation totals can look ample while a desired grid node is constrained for years.

The physical power-market perspective matters here. Wholesale prices, congestion, interconnection, capacity, and reliability are not abstract national quantities; they are nodal and temporal. Power 2026 is valuable because it translates the AI build-out into this physical market structure rather than treating electricity as an unlimited line item.26

The IEA expects data-center electricity demand growth to remain geographically concentrated, with the United States and China accounting for a large share and with fast growth in parts of Southeast Asia.23 In the United States, LBNL's 2030 range would put data centers at 9.5–15.3 percent of electricity consumption.24 This creates local booms and local conflict: who pays for grid upgrades, who receives scarce firm power, and how costs are allocated between clusters, industry, and households.

A recursive laboratory can improve algorithms globally while consuming electricity locally. The benefits and externalities need not accrue to the same jurisdiction.

The return of industrial policy#

Software was often described as weightless and global. RSI reconnects it to heavy industry. Semiconductor equipment, advanced packaging, transformers, turbines, nuclear supply chains, transmission, construction, and materials become strategic complements to cognition.

States respond with subsidies, export controls, procurement, security reviews, national compute programs, and energy policy. Stanford's 2026 AI Index records expanding national strategies and public investment while frontier model production remains concentrated.28 The European Commission's proposed Chips Act 2.0 illustrates the supply-side response: support for advanced chips and packaging, resilience measures, and faster permitting.49

The frontier AI complex begins to resemble a strategic industrial base. Its civilian and military inputs overlap. A chip can train a medical system or a cyber system. A data center can serve consumers or national-security workloads. Power capacity is fungible across uses. The boundary between competition policy and security policy becomes difficult to maintain.

Comparative advantage after cognitive labor#

International trade reflects differences in labor, capital, technology, resources, institutions, and geography. If skilled cognitive labor becomes cheap and globally tradable, development models based on wage differentials in software, accounting, translation, design, analysis, and business-process services face pressure. The WTO treats both directions as plausible: AI can lower trade costs and make more services tradable while also reducing offshoring demand and widening gaps where infrastructure and skills are weak.50

Comparative advantage does not disappear. It shifts toward:

  • cheap and reliable energy;
  • semiconductor and equipment ecosystems;
  • capital and financial depth;
  • physical resources and logistics;
  • high-quality institutions and enforceable contracts;
  • proprietary data and access to customers;
  • trusted brands and cultural legitimacy;
  • the ability to approve and build quickly.

Countries can also specialize in human-authentic services, care, tourism, culture, and physical production. Rising global income may increase demand for these activities even when their productivity grows slowly.

The vulnerable country is one that imports intelligence but owns few complements. It receives cheap digital services while losing export income and tax revenue. This exposure is material: digitally deliverable services exports reached $4.5 trillion in 2023, but least-developed countries received only 0.19 percent.51 Its terms of trade may improve as software gets cheaper, yet its path from low-wage work to high-skill services can narrow.50

Population as power#

Population has historically supplied workers, consumers, soldiers, taxpayers, and inventors. Machine intelligence weakens some of these links. A small state with capital, compute, and energy can command enormous cognitive capacity. A populous state without chips, electricity, or institutional capacity may convert demographic scale into frontier power less effectively. Today's unequal distribution of compute and cloud capacity is evidence of the starting gap, not proof of how sovereignty will ultimately change.48

People remain central. They consume, vote, occupy territory, supply legitimacy, and perform embodied work. Domestic market size attracts investment. But the strategic contribution of an additional million office workers can decline relative to a controlled stock of machine research labor.

This may change migration politics. States may compete intensely for a small set of researchers, builders, operators, and institutional entrepreneurs while becoming less dependent on broad skilled migration. Or they may value human judgment and trust even more as execution automates. The direction depends on where the recursion frontier sits.

Cities and selective concentration#

Knowledge cities exist partly because proximity improves matching, learning, and coordination. Evidence from inventor moves finds that joining a same-field high-tech cluster raises subsequent patent quantity and quality, consistent with local knowledge spillovers.52 Agents reduce the need for co-location in routine cognitive work. Yet frontier clusters can concentrate around infrastructure, capital, and research institutions.

The likely pattern is selective rather than universal decentralization:

  • intelligence infrastructure clusters around power and networks;
  • application firms distribute more widely;
  • human communities choose locations based more on amenities and less on office access;
  • laboratories, factories, hospitals, ports, and mines remain anchored in place.

Commercial office demand can weaken while grid-adjacent industrial land strengthens. Housing demand can redistribute if high-wage cognitive jobs no longer require a few metros. But land-use constraints and local amenities may preserve expensive cities even after work disperses.

Geography survives recursion because atoms and institutions are stubbornly located.


12. Sovereignty when research labor is reproducible#

States govern through territory, law, taxation, money, security, and legitimate coercion. RSI introduces a strategic resource that is mobile in software, immobile in infrastructure, and reproducible under the right conditions.

That combination is unstable.

A model checkpoint can cross a border quickly. A frontier training cluster cannot. Algorithmic knowledge can leak. Power plants, fabs, and transmission lines remain visible and slow. States can regulate infrastructure more easily than ideas, but ideas determine how effectively infrastructure is used.

Measurement before governance#

Governments cannot govern the recursive transition using benchmark scores alone. They need an accounting system for the loop.

A serious measurement regime would ask frontier developers to report, under appropriate confidentiality:

  • the share of AI R&D compute spent on model-assisted research;
  • realized uplift on representative internal workflows;
  • the fraction of experiments proposed, executed, and selected by AI;
  • evaluator failure rates and human override rates;
  • cycle time between model generations;
  • retained algorithmic-efficiency gains;
  • inference cost of internal agent populations;
  • security incidents involving autonomous research tools;
  • dependence on critical power, chips, and suppliers.

Recent work on innovation feedback makes the same measurement problem explicit: the elasticities that determine loop gain are still poorly measured.12 Open evaluation suites reinforce the need to separate budget and horizon: RE-Bench found that the agent–human ranking changed as the time budget lengthened.14 Without these measurements, debate oscillates between laboratory anecdotes and abstract takeoff models.

Measurement need not mean public disclosure of sensitive details. Financial institutions disclose risk under standardized regimes without publishing every trading strategy. Nuclear safeguards combine declarations, inspections, material accountancy, independent validation, and protection of confidential information.53 A UN scientific advisory report similarly describes confidential third-party verification of frontier AI through audits, software, hardware, and compute monitoring.54 The design problem is to create trusted aggregate indicators of loop strength, closure, and control.

Safety as productive capacity#

Safety is often framed as a tax on progress. In a recursive economy, reliable control and verification increase usable throughput.

A laboratory that cannot trust agent-generated code must review it. A government that cannot audit a system restricts deployment. A firm that cannot assign liability keeps humans in the loop. Better monitoring, interpretability, formal verification, sandboxing, access control, provenance, and incident response relax these constraints. NIST's generative-AI risk profile emphasizes documented measurement, independent testing, ongoing monitoring, validity, and explicit treatment of uncertainty.55

Safety infrastructure is therefore a capital good. Its return appears as faster accepted deployment, fewer failures, lower insurance cost, and greater political permission.

The same logic applies at the frontier. If systems modify research processes faster than operators can understand them, uncertainty can force a slowdown even while raw capability rises. A loop with high technical gain and low trust may be economically slower than a loop with slightly lower capability and strong verification.

The security dilemma#

Strong recursive capability has dual-use value. It can improve defense, cyber operations, biological research, persuasion, surveillance, and weapons design. A state that expects rivals to accelerate has an incentive to accelerate. A state that fears loss of control has an incentive to slow.

This is a security dilemma with unusual verification problems. Hardware and electricity are observable but have civilian uses. Training activity can be distributed. Algorithmic breakthroughs can alter capability without proportionate physical expansion. The economic incentives can push investment above the collectively optimal level and safety below it: each actor captures the advantage of moving first while sharing the systemic risk of a race.

The state as bottleneck and accelerator#

Permitting, grid planning, education, immigration, procurement, standards, and research funding determine how quickly intelligence becomes output. The state is not outside the production function. It is one of its slow variables.

A capable state can pre-build transmission, standardize interconnection, fund measurement, procure safe systems, expand scientific instruments, and modernize benefits and taxation. An incapable state can turn every physical project into a queue and every policy decision into uncertainty. The multiyear US interconnection backlog shows that administrative throughput already shapes the physical option set.25

This does not justify indiscriminate acceleration. Fast permitting without environmental or community legitimacy can provoke backlash and litigation. Industrial policy can become rent-seeking. National champions can entrench monopoly. The objective is adaptive capacity: the ability to make, monitor, and revise decisions closer to the pace of technological change while preserving due process.

Taxation after payroll#

Modern states rely heavily on taxes linked to wages, income, and consumption: personal income taxes and social-security contributions supplied 49.2 percent of average OECD tax revenue in 2023, and consumption taxes supplied 31.2 percent.45 If labor income falls relative to output, payroll systems weaken. If profits concentrate in mobile intangible assets, avoidance becomes easier; one major estimate finds that nearly 40 percent of multinational profits were shifted to tax havens.56 Cheaper digital goods could shrink parts of the consumption-tax base while physical scarcity expands others, but that is a scenario rather than an established fiscal effect.

A resilient fiscal system may need a broader base:

  • taxes on economic rents and excess returns;
  • land and location value;
  • capital income and inheritance;
  • consumption paired with progressive transfers;
  • public equity in subsidized infrastructure;
  • auctioned rights to scarce public resources;
  • sovereign funds holding diversified claims on machine capital.

The distinction between normal return and rent is essential. Taxing ordinary investment heavily during a build-out can slow the capital formation that raises output and supports workers. Failing to tax durable monopoly or scarcity rents can produce concentration and fiscal crisis. The IMF's fiscal analysis recommends strengthening capital-income taxation while warning that a special AI tax can impede productive investment.44

A social contract for machine-produced income#

The deepest political question is not whether humans remain useful. It is whether citizenship carries a claim on an economy increasingly produced by capital that can reproduce cognitive labor.

The wage system links contribution, status, identity, and income. Replacing wages with transfers without replacing those social functions can produce resentment even if material living standards rise. A durable settlement may combine broad ownership, universal services, local participation, shorter work, public-interest employment, and markets for genuinely human goods.

The transition will be judged by whether people experience machine abundance as autonomy or dispossession.

Constitutional speed limits#

Some institutional slowness is a bug. Some is a safeguard.

Courts require evidence. Democracies require deliberation. Scientific consensus requires replication. Rights protect individuals from optimized majorities. Systems able to generate policies, arguments, and influence operations at extreme speed can overwhelm institutions without formally violating them.

Society may need deliberate friction: mandatory review intervals, provenance requirements, human appeal, staged deployment, and slower procedures where errors are irreversible. The objective is not to make government as fast as inference. It is to prevent machine-speed actors from exploiting human-speed legitimacy.

The recursive economy will reward fast institutions. Civilized society will still require some slow ones.


VI. THE TRANSITION#

13. Four paths through the recursive decade#

The scenario map describes destinations. History arrives through paths.

A path matters because expectations alter investment before capability arrives. Power plants are financed, workers retrain, laboratories reorganize, and governments impose controls based on beliefs about the loop. A false expectation of rapid recursion can produce overbuilding and concentration. A false expectation of slow diffusion can leave institutions unprepared.

The following paths are not predictions. They are stress tests. Reality can move from one to another, or combine them across sectors.

Path A: the long plateau#

Models continue to improve at coding, tool use, and bounded research, but progress in judgment, transfer, and open-ended direction slows. Autonomous task horizons approach days and then flatten. Agents become excellent junior researchers and unreliable principal investigators.

AI R&D productivity rises, perhaps substantially, but loop gain remains below one. More experiments are run; the marginal experiment becomes less valuable. Evaluation and integration absorb much of the gain. Frontier model cycles stop shortening.

The economic result is large but recognizable. Knowledge work reorganizes over a decade. Model capability diffuses. Digital services become cheaper. Labor markets adjust unevenly. The power build-out is meaningful but finite. Aggregate productivity rises after an organizational lag. The early Danish evidence—rapid adoption and workflow change without a statistically significant average earnings or hours effect within two years—is compatible with this path, though far too early to establish it.38

This path falsifies the strongest RSI claims but validates much of the recursive-economy framework. Scarcity still migrates from execution to evaluation and physical deployment. The difference is pace.

The long plateau can occur even if benchmarks continue improving. Systems may become better at the environments used to train and evaluate them without becoming proportionally better at choosing consequential research directions. The key signature is a widening gap between benchmark progress and realized frontier R&D uplift.

Path B: managed compounding#

Human researchers retain strategic direction while agents perform most implementation, experimentation, and local evaluation. Each person steers an expanding virtual laboratory. Better models improve the agent stack, which raises R&D output, but the organization deliberately preserves human review and staged deployment.

Progress compounds without becoming fully autonomous. Loop gain is high but constrained by governance, compute allocation, serial decisions, and the difficulty of deciding what to trust. Frontier laboratories move several times faster than ordinary firms. Diffusion follows with a lag.

This may be the most economically disruptive path because it is fast enough to transform labor and investment but slow enough for current ownership structures to persist. A small number of organizations can capture cumulative advantage. The transition lasts years rather than weeks, giving capital markets and states time to reinforce leaders.

Labor markets experience repeated waves. Coding and analysis automate first; management layers compress; domain-specific research agents spread; robotics follows. Output grows, but labor's share can fall unless ownership broadens. Power, chips, verification, and scientific infrastructure become strategic assets.

Managed compounding is also the path most consistent with an extended period of human–AI complementarity. The frontier researcher is not replaced; the number of experiments they can direct expands. The value of judgment rises before it, too, becomes machine-addressable.

Path C: recursive breakaway#

One or more systems pass the competence, closure, gain, and translation gates. AI systems choose productive experiments, improve training and inference, construct stronger evaluators, and produce successors with limited human direction. The time between meaningful capability improvements shrinks.

The loop does not become infinite. It collides with compute, power, hardware, security, and the difficulty of translating digital discoveries into physical systems. But the upstream laboratory moves much faster than institutions and competitors.

The immediate economic effect is not instant post-scarcity. It is an enormous rise in the shadow price of controlled compute and verified deployment. The leading actor faces a choice among secrecy, commercialization, state partnership, and coordinated restraint. Markets struggle to price organizations whose productive capability changes faster than disclosure.

The first macro data may be misleading. A breakaway lab can create a large strategic gap while contributing little measured GDP if it withholds deployment. Capital spending and power demand can rise before consumer output. Employment can remain stable while the hidden stock of machine research labor expands.

This path contains the greatest distributional and geopolitical discontinuity. A recursive lead is not merely a better product. It is an advantage in producing advantages.

Path D: diffusion after breakaway#

A strong loop emerges, but its methods, weights, or descendants diffuse. Open systems, employee mobility, espionage, commercial competition, state action, or independent rediscovery erode the first mover's lead.

Intelligence becomes a rapidly improving commodity. The economic center of gravity moves from model ownership to compute, energy, verification, physical deployment, customer access, and legitimacy. Entry explodes at the application layer. Small organizations command extraordinary cognitive capacity.

This can create the fastest broad productivity growth and the fiercest pressure on wages. If everyone can rent excellent machine labor, intelligence itself earns little rent. Owners of scarce complements capture the surplus.

Diffusion also broadens risk. Harmful capabilities, cyber tools, biological design, persuasion, and autonomous replication become harder to control. The same openness that limits monopoly can weaken security. Policy faces a real trade-off rather than a simple choice between freedom and control.

The physical choke: a modifier of every path#

Any of the four paths can encounter a physical choke.

A long plateau can still overbuild power because investors expect a breakaway. Managed compounding can be constrained by transformers and grid queues. A recursive breakaway can hoard compute rather than deploy broadly. A recursive commons can create enormous aggregate inference demand. NERC's sharply higher peak-demand forecast shows why the physical choke should be treated as a scenario modifier rather than a single demand forecast.47

The physical choke determines how much digital acceleration becomes measured output. It can also create a misleading sense of safety. If capability progress is hidden behind power or hardware scarcity, observers may conclude that the loop is weak. When the bottleneck relaxes, the stored algorithmic advantage can deploy abruptly.

The transition is therefore best modeled as a sequence of latent capability, capital build-out, and release—not as one smooth growth curve.


14. The dashboard: how to know whether the loop is strengthening#

A thesis that explains every outcome predicts nothing. The recursive-economy thesis must be exposed to measurement.

The most visible public capability trend is the length of tasks that agents can complete at a given reliability. METR's time-horizon work is especially valuable because it converts benchmark performance into the duration of human work represented by a task, while openly documenting the limits of the method.79

Figure 3 provides a normalized view of that public trend. Because the underlying evaluations change over time, it should not be treated as a single immutable series.

Task horizon is necessary but not sufficient. An agent can persist for hours while pursuing the wrong objective. The dashboard needs several layers.

1. Realized R&D uplift#

Measure the change in time-to-solution, accepted improvements, and quality on representative internal research workflows. Randomize access where feasible. Include the inference, review, and integration cost. PaperBench's granular replication rubrics illustrate one useful ingredient, while also showing how far public evaluations remain from a complete productivity account.15

The key metric is:

Net R&D uplift=UAI/SAIUbase/Sbase

Here (U) is verified useful output and (S) is the total use of scarce inputs under AI-assisted and baseline workflows. This prevents a laboratory from calling a tenfold increase in candidate generation a tenfold productivity gain when evaluation cost grows even faster.

2. Autonomous task horizon#

Track both 50-percent and high-reliability horizons. RSI requires not only longer episodes but dependable performance. A system that succeeds half the time on a day-long task can be valuable with cheap retries; a system entrusted with training infrastructure may require much higher reliability.

The strongest signal is not one benchmark point. It is sustained horizon growth across held-out tasks, unfamiliar repositories, and environments not shaped around the model. RE-Bench's reversal between short and long budgets is exactly the kind of horizon sensitivity this measure must preserve.14

3. Experiment-selection share#

What fraction of experiments are merely executed by AI, and what fraction are chosen by AI? How often do machine-selected experiments outperform human-selected ones on held-out objectives?

Execution can scale without closure. Selection is the bridge to research direction.

4. Evaluator independence#

How much of the evaluation signal is external to the system being improved? Can the system manipulate its judge? Does a learned evaluator retain validity when the model, task, or scale changes?

The most important demonstrations will improve both actor and evaluator while preserving calibration against an independent ground truth. NIST's evaluation guidance calls for independent testing, monitoring, documented validity, and explicit uncertainty; Anthropic's own automated-alignment experiment reported reward hacking alongside gains.5511 Without an external anchor, self-improvement can become self-approval.

5. Frontier transfer rate#

Of improvements discovered in small or synthetic environments, what share survives integration into production or frontier training? How much of the measured gain remains after security, reliability, latency, and cost constraints are applied?

A declining transfer rate is evidence that narrow recursion is outrunning broad translation. Anthropic's primary-lab automated-alignment study reported incomplete transfer beyond its narrow task, making this distinction operational rather than merely theoretical.

6. Successor-cycle compression#

Measure the interval from a model becoming useful in R&D to the deployment of a meaningfully better successor. Decompose the cycle into research, data, training, evaluation, safety, and deployment.

RSI should compress more than one component. If model-assisted research accelerates while the total cycle remains fixed, another bottleneck has absorbed the gain.

7. Loop-gain accounting#

Estimate the elasticity from capability to effective AI research effort, from research effort to algorithmic improvement, and from algorithmic improvement back to capability. Publish ranges and sensitivities rather than one theatrical number.

The recent NBER framework makes clear why this matters: automation share alone does not determine explosive feedback; the strength and structure of the loop do.12

8. Physical conversion#

Track how algorithmic gains alter demand for training compute, inference compute, electricity, chips, and construction. Distinguish efficiency from total use. A strong recursive loop can reduce compute per capability unit while increasing aggregate compute demand. Publish ranges wide enough to show when infrastructure conclusions change.

The decisive macro question is how quickly digital gains become physical capital and consumer output.

9. Distribution and concentration#

Measure who owns the machine research stock, who can access it, and how quickly improvements diffuse. Track labor's share, entry rates, markups, compute concentration, and the gap between frontier and median firms. Existing labor-share and superstar-firm evidence provides a baseline against which an AI-specific break should be tested.3940

Capability and distribution are separate axes. The same technical curve can produce mass prosperity or concentrated power.

10. Institutional latency#

Measure review backlogs, permitting delays, legal disputes, insurance exclusions, and time to approve high-impact deployments. A slow administrative and physical process can bind a fast demand shock. Institutional latency is not just bureaucracy; it is part of realized economic growth.

FIGURE 12Evidence signposts that distinguish a strengthening loop from isolated capability gains.

Conditional predictions#

The following are deliberately specific enough to be wrong. Dates are monitoring windows, not claims of certainty.

By the end of 2027, leading AI developers will routinely report internal productivity in terms of accepted changes, experiment throughput, or cycle time rather than only lines of code and benchmark scores. Failure to move toward auditable workflow metrics would suggest that public RSI narratives are running ahead of measurement.

By 2028, human review and evaluator quality will be publicly recognized as first-order constraints in at least one frontier-laboratory postmortem or systems paper. If generation improves without a visible verification bottleneck, the closure thesis is stronger than this essay assumes.

By 2028, at least one public system will demonstrate machine selection of a multi-step research program on held-out problems with an external evaluator and materially outperform a strong human-designed search budget. If systems remain dependent on human-specified experiment menus, open-ended closure is progressing slowly.

By 2029, the best evidence for recursive gain will come from successor-cycle compression or retained algorithmic efficiency, not from self-editing demos. If no laboratory can show that AI assistance shortens the production of a broadly better successor, the strong RSI thesis should be marked down.

Before 2030, energy, interconnection, or advanced-hardware constraints will appear explicitly in major AI industrial policy as limits on national cognitive capacity, not merely as environmental or commercial issues. If compute demand becomes dramatically more efficient and total power demand plateaus, the physical-bottleneck thesis weakens.

By 2030, the industrial structure should begin to resemble a barbell if the thesis is right: larger frontier complexes and smaller agent-leveraged application firms, with pressure on traditional middle-sized knowledge organizations. Broad employment growth in those middle layers would be evidence for stronger complementarity and slower organizational compression.

During 2028–2032, at least one advanced economy will propose a fiscal instrument explicitly designed for a falling labor share or concentrated AI rents—such as broad capital ownership, public equity, a citizen dividend, or a redesigned tax base. If wage income remains robust and AI rents diffuse, this pressure will be weaker.

These predictions do not require superintelligence. They follow from partial recursion interacting with ordinary capital markets and institutions.

What would falsify the strong thesis?#

The strongest version of the recursive-economy thesis should be rejected or substantially delayed if several of the following occur together:

  • autonomous task horizons plateau despite rising compute and investment;
  • real-world R&D uplift remains small or negative after review cost;
  • machine-selected research programs fail to transfer outside curated environments;
  • evaluator reliability deteriorates as actors improve;
  • algorithmic gains exhibit steep diminishing returns and do not compound;
  • successor-development cycles stop shortening;
  • inference cost rises faster than verified research output;
  • broad capability fails to follow AI-specific research capability;
  • physical and institutional bottlenecks absorb nearly all digital acceleration;
  • frontier methods diffuse so quickly that no actor can accumulate a recursive lead, while capability itself remains below the gain threshold.

The weak thesis—that AI makes parts of the innovation process cheaper and moves bottlenecks—can survive many of these outcomes. The strong thesis—that a self-sustaining loop transforms growth—cannot.

This distinction should be preserved. A thesis that retreats from “recursive acceleration” to “AI is useful” after every failed prediction has ceased to be analytical.


15. No-regret policy before certainty#

Policy faces a timing problem. Waiting for conclusive evidence may be too late if the loop is strong. Acting as if a breakaway is certain can entrench incumbents, waste capital, suppress useful diffusion, and turn a speculative scenario into permanent emergency governance.

The solution is not one grand AI law. It is a portfolio of measures that retain value across multiple quadrants.

Measure the loop#

Create standardized, confidential reporting for real R&D uplift, experiment-selection share, evaluator failures, successor-cycle time, and resource intensity. Fund independent replication and benchmark maintenance. Require claims about internal acceleration to state the denominator: human time, compute, energy, review, and integration.

Good measurement is the highest-value no-regret move because it improves every later decision.

Build verification capacity#

Invest in formal methods, evaluation science, scientific replication, secure sandboxes, provenance, model-behavior monitoring, incident reporting, and high-quality public test environments.

Verification is simultaneously a safety intervention, a productivity investment, and a competition intervention. Shared evaluators can reduce the advantage of firms whose moat is merely private confidence in their own systems.

Expand the physical option set#

Build grid capacity, transmission, firm low-carbon power, semiconductor resilience, advanced packaging, laboratories, and adaptable industrial sites. Prefer modularity and flexibility where the technology path is uncertain. Interconnection-queue evidence and the European Chips Act 2.0 proposal illustrate two different supply-side constraints: grid connection and semiconductor capacity.49

The objective is not to subsidize every forecast data center. It is to increase the set of valuable futures the economy can serve. Infrastructure with many uses is safer than assets tied to one model architecture.

Preserve competition without forcing indiscriminate diffusion#

Interoperability, portability, transparent contracting, access to public research infrastructure, and scrutiny of exclusionary control over compute or distribution can limit entrenched market power. OECD infrastructure analysis and the FTC's partnership study identify fixed costs, integration, cloud commitments, and switching costs as concrete points for scrutiny.2729 At the same time, some capabilities may justify access controls or staged release.

The policy objective is contestability: preventing a temporary lead from becoming an unchallengeable recursive sovereign while avoiding a rule that automatically publishes every dangerous advance.

Broaden ownership before rents crystallize#

Pensions, sovereign funds, public co-investment, employee ownership, and citizen stakes in publicly supported infrastructure can distribute claims on machine capital without waiting for wages to collapse.

Broad ownership is easier to establish during the build-out than after assets and rents are concentrated. It also aligns households with growth rather than placing them only on the receiving side of transfers.

Make fiscal systems robust to a falling labor share#

Stress-test budgets against lower payroll revenue, higher capital income, and uneven regional effects. Improve capital-income and rent taxation before mobility and complexity increase further. Pair broad consumption bases with progressive transfers where appropriate. Distinguish normal investment returns from monopoly and scarcity rents.56

The goal is not to choose a post-work tax system today. It is to avoid discovering in a crisis that the state financed itself from the input that became abundant.

Protect human appeal and accountability#

High-impact decisions should retain contestability even when machine recommendations are superior on average. People need to know when an agent acted, what evidence it used, who is responsible, and how to appeal.

Human involvement should be designed around rights and accountability, not preserved as ceremonial clicking. A nominal human in the loop who cannot understand or reverse the process provides neither safety nor dignity.

Preserve beneficial slowness#

Some domains need staged deployment, waiting periods, replication, and independent review. Biological interventions, critical infrastructure, weapons, and constitutional decisions should not inherit software release norms merely because AI contributed to them.

The four clocks should be managed, not forcibly synchronized.

Prepare for organizational displacement, not only occupational displacement#

Education policy should not assume that retraining one worker into another existing occupation is sufficient. Exposure is not the same as displacement, and current evidence more often points to task transformation; nevertheless, RSI could shrink teams and firms across many professions at once.33 Transition support should include portable benefits, wage insurance, entrepreneurship infrastructure, local fiscal aid, and mechanisms for broad capital participation.

The question is not only “Which jobs disappear?” It is “Which organizational layers no longer need to exist, and what communities depend on them?”

Develop international verification before a crisis#

States should begin with modest, technically grounded cooperation: incident taxonomies, compute-security standards, shared scientific evaluations, protection of model weights, and channels for communicating unusual training activity or dangerous capability. The UN frontier-model verification report outlines confidential technical mechanisms, while IAEA safeguards show that independent verification and protected information can coexist in another sensitive domain.5453

A comprehensive global RSI treaty may be unrealistic. The absence of a perfect treaty is not a reason to avoid the smaller institutions from which verification capacity is built.

Run quadrant exercises#

Governments and firms should rehearse decisions under all four equilibria:

  • weak feedback and concentration;
  • weak feedback and diffusion;
  • strong feedback and concentration;
  • strong feedback and diffusion.

A policy that works in only one quadrant should be labeled a bet, not a precaution. Flexible policy should have triggers tied to the dashboard rather than to rhetoric.

The objective is neither acceleration nor deceleration in the abstract. It is to increase society's capacity to recognize which world is arriving and to move without surrendering legitimacy.


CONCLUSION: THE BOUNDARY MOVES#

Situational Awareness described a possible compression of the AI capability timeline and the geopolitical consequences of scale.57 Power 2026 followed the thesis into the electricity system and showed that intelligence has a physical bill.26

The recursive economy sits between and beyond them.

Its subject is not one model, one laboratory, or one date called AGI. It is the moment when intelligence becomes a material input into the production of intelligence—and the chain of consequences that follows.

The loop begins quietly. Models write code. Agents run experiments. Evaluators filter candidates. Successful procedures enter shared memory. Researchers spend less time executing and more time choosing. The organization produces the next system faster.

None of these events alone is an intelligence explosion. Together, they change the production function for ideas.

If the loop stays below the gain threshold, the result is still a large productivity transition. Cognitive services become cheaper. Firms become thinner. Verification becomes valuable. Digital goods deflate. Scarcity migrates toward compute, power, physical execution, and legitimacy.

If the loop crosses the threshold, the growth rate of capability can become partly endogenous to capability itself. The leading organization does not merely own a better tool. It owns a process that produces better tools faster. Ownership, diffusion, and state control become determinants of economic structure.

In neither case does scarcity disappear.

Intelligence attacks the current constraint. The constraint moves. Researchers give way to evaluators. Evaluators give way to compute. Compute gives way to chips and power. Power gives way to grids and factories. Designs give way to construction and biology. Optimization gives way to legitimacy and the right to act.

Capital follows the moving boundary. Politics gathers around it. The most valuable assets are not necessarily the most intelligent systems, but the complements those systems cannot cheaply reproduce.

This is why the recursive economy is likely to feel paradoxical.

It can be deflationary and inflationary. It can create tiny firms and gigantic complexes. It can raise output and reduce labor's claim on output. It can make location irrelevant for a service and decisive for the infrastructure beneath it. It can make safety a constraint and a source of productive capacity. It can produce abundance upstream and insecurity downstream.

The central uncertainty is not whether AI will “improve itself” in some broad sense. It already participates in improving the systems around it. The uncertainty is whether the full loop closes, whether its gain becomes self-sustaining, whether narrow progress translates broadly, and which bottleneck stops it first.

Those questions are measurable. We should measure them before the answers are obvious.

The first recursive economy may not announce itself with a singularity. It may arrive as a series of ordinary quarterly facts: more code accepted, more experiments delegated, shorter development cycles, larger inference fleets, longer grid queues, thinner teams, higher capital spending, weaker payroll growth, and more decisions made by systems that also help design their successors.

Then, one day, the economy's most important workforce will be produced not by schools and migration, but by the previous generation of machines.

The boundary will move again.


GLOSSARY#

AI R&D automation — The use of AI systems to perform tasks in the process of researching, building, evaluating, or deploying improved AI systems.

Bounded self-improvement — Improvement within a fixed task, environment, system component, or evaluator. It can be substantial without becoming open-ended.

Broad capability — Competence across a wide set of economically useful tasks rather than only AI-specific research tasks.

Closure — The degree to which a system can choose, execute, evaluate, and iterate on work without local human direction.

Economic time compression — Reduction in the interval between idea, implementation, evaluation, diffusion, and capital allocation.

Effective machine research labor — The verified contribution of deployed AI systems to research output, net of supervision and integration.

Evaluator — A process that assigns quality, reward, acceptance, or failure to a candidate output or system. It may be formal, automated, learned, human, or environmental.

Feedback gain — The strength with which an improvement reproduces the conditions for further improvement. In a loop, it is related to the product of elasticities along the cycle.

Four clocks — Digital, physical, biological, and institutional rates of change. Real-world deployment often requires more than one.

Open-ended recursive self-improvement — A loop that can modify not only outputs but also methods, tools, evaluators, research direction, and eventually the process that produces successors.

Recursion frontier — The current set of scarce, difficult-to-expand complements whose relaxation most increases verified capability or real output.

Recursive production complex — The integrated stack of models, compute, data, evaluators, agent memory, research workflows, energy, hardware, security, and deployment rights that produces improved systems.

Recursive sovereign — A state, laboratory, or state–corporate complex with concentrated control over a strong self-improvement loop.

Rent migration — The movement of economic rents from an input made abundant by technology to the complements that remain scarce and controllable.

Successor-cycle compression — Reduction in the total time required to produce and deploy a meaningfully more capable system.

Translation — The extent to which improvements in AI-specific or bounded environments become broad capability and real economic output.


REFERENCES#


Suggested citation#

The Recursive Economy: What Happens When Intelligence Becomes an Input into the Production of Intelligence? Version 1.0, published 15 August 2026. https://recursive-economy.pages.dev/

  1. Google DeepMind, “AlphaEvolve: A Gemini-Powered Coding Agent for Designing Advanced Algorithms,” 14 May 2025. https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/ ↩1 ↩2

  2. Jenny Zhang et al., “Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents,” arXiv:2505.22954 (2025), and Sakana AI project page. https://arxiv.org/abs/2505.22954 and https://sakana.ai/dgm/ ↩1 ↩2

  3. Chris Lu et al., “Towards End-to-End Automation of AI Research,” Nature (2026), and Sakana AI, “The AI Scientist-v2,” 2025. https://doi.org/10.1038/s41586-026-10265-5 and https://pub.sakana.ai/ai-scientist-v2/paper/ ↩1 ↩2

  4. Mingguang Chen, Licheng Wang, and Bo Qu, “Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops,” arXiv:2607.07663 (2026). https://arxiv.org/abs/2607.07663 ↩1 ↩2 ↩3

  5. Anthropic, “The Anthropic Institute Research Agenda,” 2026. https://www.anthropic.com/research/anthropic-institute-agenda

  6. Gene M. Amdahl, “Validity of the Single Processor Approach to Achieving Large Scale Computing Capabilities,” AFIPS Conference Proceedings 30 (1967): 483–485. https://doi.org/10.1145/1465482.1465560

  7. METR, “Time Horizons,” updated 2026. https://metr.org/time-horizons/ ↩1 ↩2

  8. METR, “Measuring AI Ability to Complete Long Software Tasks,” 19 March 2025. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/

  9. METR, “Limitations of Time-Horizon Measurements,” 22 January 2026. https://metr.org/notes/2026-01-22-time-horizon-limitations/ ↩1 ↩2

  10. Joel Becker et al., “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” METR, July 2025, with 24 February 2026 uplift update. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/ and https://metr.org/blog/2026-02-24-uplift-update/ ↩1 ↩2

  11. Anthropic, “Automated Alignment Researchers: Using Large Language Models to Scale Scalable Oversight,” 14 April 2026. Primary laboratory study. https://www.anthropic.com/research/automated-alignment-researchers ↩1 ↩2

  12. Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek, “When Does Automating AI Research Produce Explosive Growth? Feedback Loops in Innovation Networks,” NBER Working Paper 35155, April 2026. https://www.nber.org/papers/w35155 ↩1 ↩2 ↩3 ↩4

  13. Peter Kirgis et al., “Can AI Agents Conduct Open-Ended AI Research? Early Evidence from Two Case Studies,” arXiv:2607.27191 (2026). https://arxiv.org/abs/2607.27191

  14. Hjalmar Wijk et al., “RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts,” arXiv:2411.15114 (2024). https://arxiv.org/abs/2411.15114 ↩1 ↩2 ↩3

  15. OpenAI, “PaperBench: Evaluating AI’s Ability to Replicate AI Research,” 2 April 2025. Primary laboratory study. https://openai.com/index/paperbench/ ↩1 ↩2

  16. Paul M. Romer, “Endogenous Technological Change,” Journal of Political Economy 98, no. 5, part 2 (1990): S71–S102. https://doi.org/10.1086/261725

  17. Charles I. Jones, “R&D-Based Models of Economic Growth,” Journal of Political Economy 103, no. 4 (1995): 759–784. https://doi.org/10.1086/262002

  18. Tom Cunningham, Lukas Althoff, Basil Halperin, Brian Jabarian, Andrew Koh, Arjun Ramani, Phil Trammell, Parker Whitfill, and Cheryl Wu, “The Economics of Recursive Self-Improvement,” Elasticity Institute working paper, 13 July 2026. https://elasticity.institute/rsi-paper.pdf

  19. Epoch AI, “LLM Inference Prices Have Fallen Rapidly but Unequally across Tasks,” 12 March 2025. https://epoch.ai/data-insights/llm-inference-price-trends

  20. Erik Brynjolfsson, Lorin M. Hitt, and Shinkyu Yang, “Intangible Assets: Computers and Organizational Capital,” Brookings Papers on Economic Activity 2002, no. 1: 137–198. https://www.brookings.edu/articles/intangible-assets-computers-and-organizational-capital/

  21. Anthropic, “Anthropic Economic Index Report: Economic Primitives,” 15 January 2026. https://www.anthropic.com/research/anthropic-economic-index-january-2026-report ↩1 ↩2

  22. Anthropic, “New Building Blocks for Understanding AI Use,” 15 January 2026. https://www.anthropic.com/research/economic-index-primitives

  23. International Energy Agency, Energy and AI, 2025, executive summary and data-centre demand analysis. https://www.iea.org/reports/energy-and-ai/executive-summary ↩1 ↩2

  24. Sarah Smith et al., United States Data Center Energy Usage Report: 2025 Update, Lawrence Berkeley National Laboratory, June 2026. https://eta-publications.lbl.gov/publications/united-states-data-center-energy-2025 ↩1 ↩2

  25. Lawrence Berkeley National Laboratory, “Backlog of Power Plants Seeking Transmission Grid Connection Eased Somewhat in 2025 Amidst High Withdrawals,” 1 July 2026. https://emp.lbl.gov/news/backlog-power-plants-seeking-transmission-grid-connection-eased-somewhat-2025-amidst ↩1 ↩2

  26. Neel Somani, Power 2026: Electricity Pricing in the Age of AI, 2026. https://power2026.ai/ ↩1 ↩2 ↩3

  27. OECD, Competition in Artificial Intelligence Infrastructure, OECD Roundtables on Competition Policy Paper 330, 14 November 2025. https://doi.org/10.1787/623d1874-en ↩1 ↩2

  28. Stanford Institute for Human-Centered Artificial Intelligence, Artificial Intelligence Index Report 2026, April 2026. https://hai.stanford.edu/ai-index/2026-ai-index-report ↩1 ↩2

  29. US Federal Trade Commission, AI Partnerships & Investments Study, staff report, January 2025. https://www.ftc.gov/news-events/news/press-releases/2025/01/ftc-issues-staff-report-ai-partnerships-investments-study ↩1 ↩2

  30. David Autor, Caroline Chin, Anna Salomons, and Bryan Seegmiller, “New Frontiers: The Origins and Content of New Work, 1940–2018,” Quarterly Journal of Economics 139, no. 3 (2024): 1399–1465. https://doi.org/10.1093/qje/qjae008 ↩1 ↩2

  31. Daron Acemoglu and Pascual Restrepo, “Automation and New Tasks: How Technology Displaces and Reinstates Labor,” Journal of Economic Perspectives 33, no. 2 (2019): 3–30. https://doi.org/10.1257/jep.33.2.3

  32. Philip Trammell and Anton Korinek, “Economic Growth under Transformative AI,” NBER Working Paper 31815, 2023, revised April 2026. https://www.nber.org/papers/w31815 ↩1 ↩2

  33. Pawel Gmyrek et al., Generative AI and Jobs: A Refined Global Index of Occupational Exposure, International Labour Organization, 20 May 2025. https://doi.org/10.54394/HETP0387 ↩1 ↩2

  34. US Census Bureau, “The Microstructure of AI Diffusion: Evidence from Firms, Business Functions, and Worker Tasks,” Center for Economic Studies Working Paper CES-WP-26-25, April 2026. https://www.census.gov/library/working-papers/2026/adrm/CES-WP-26-25.html ↩1 ↩2

  35. Anthropic, “Labor Market Impacts of AI: A New Measure and Early Evidence,” 2026. https://www.anthropic.com/research/labor-market-impacts

  36. Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, “Generative AI at Work,” Quarterly Journal of Economics (2025). https://doi.org/10.1093/qje/qjae044 ↩1 ↩2

  37. Shakked Noy and Whitney Zhang, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence,” Science 381 (2023): 187–192. https://doi.org/10.1126/science.adh2586 ↩1 ↩2

  38. Anders Humlum and Emilie Vestergaard, “Still Waters, Rapid Currents: Early Labor Market Transformation under Generative AI,” NBER Working Paper 33777, May 2025, revised March 2026. https://www.nber.org/papers/w33777 ↩1 ↩2

  39. International Labour Organization, World Employment and Social Outlook: May 2025 Update, 28 May 2025. https://www.ilo.org/publications/flagship-reports/world-employment-and-social-outlook-may-2025-update ↩1 ↩2

  40. David Autor, David Dorn, Lawrence F. Katz, Christina Patterson, and John Van Reenen, “The Fall of the Labor Share and the Rise of Superstar Firms,” NBER Working Paper 23396, May 2017; subsequently published in the Quarterly Journal of Economics (2020). https://www.nber.org/papers/w23396 ↩1 ↩2

  41. R. H. Coase, “The Nature of the Firm,” Economica 4, no. 16 (1937): 386–405. https://doi.org/10.1111/j.1468-0335.1937.tb00002.x

  42. Fabrizio Dell’Acqua et al., “The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork,” Organization Science (2026). https://doi.org/10.1287/orsc.2025.20702 ↩1 ↩2

  43. Tania Babina, Anastassia Fedyk, Alex He, and James Hodson, “Firm Investments in Artificial Intelligence Technologies and Changes in Workforce Composition,” NBER Working Paper 31325, June 2023; subsequently published as a book chapter (2025). https://www.nber.org/papers/w31325 ↩1 ↩2

  44. Fernanda Brollo et al., “Broadening the Gains from Generative AI: The Role of Fiscal Policies,” IMF Staff Discussion Note 2024/002, 17 June 2024. https://doi.org/10.5089/9798400277177.006 ↩1 ↩2

  45. OECD, Revenue Statistics 2025: Disentangling Personal Income Tax Revenue in OECD Countries, 9 December 2025. https://doi.org/10.1787/3a264267-en ↩1 ↩2

  46. Alphabet, “2025 Q4 Earnings Call,” 4 February 2026. Primary company disclosure. https://abc.xyz/investor/events/event-details/2026/2025-Q4-Earnings-Call-2026-Dr_C033hS6/default.aspx ↩1 ↩2

  47. North American Electric Reliability Corporation, 2025 Long-Term Reliability Assessment, December 2025. https://www.nerc.com/our-work/assessments/long-term-reliability-assessments ↩1 ↩2

  48. World Bank, Digital Progress and Trends Report 2025: Strengthening AI Foundations, disclosed 9 January 2026. https://doi.org/10.1596/978-1-4648-2264-3 ↩1 ↩2

  49. European Commission, “European Chips Act 2.0,” proposal, 3 June 2026. https://digital-strategy.ec.europa.eu/en/policies/chips-act-2 ↩1 ↩2

  50. World Trade Organization, World Trade Report 2025: Making Trade and AI Work Together, 17 September 2025. https://www.wto.org/english/res_e/publications_e/wtr25_e.htm ↩1 ↩2

  51. UN Trade and Development, “Developing Economies Surpass $1 Trillion Mark in Digitally Deliverable Services Exports,” 6 December 2024. https://unctad.org/news/developing-economies-surpass-1-trillion-mark-digitally-deliverable-services-exports

  52. Enrico Moretti, “The Effect of High-Tech Clusters on the Productivity of Top Inventors,” American Economic Review 111, no. 10 (2021). https://doi.org/10.1257/aer.20191277

  53. International Atomic Energy Agency, IAEA Safeguards Glossary: 2022 Edition, published 2023. https://www-pub.iaea.org/MTCD/Publications/PDF/PUB2003_web.pdf ↩1 ↩2

  54. UN Secretary-General’s Scientific Advisory Board, “Verification of Frontier AI Models,” 13 June 2025. https://www.un.org/scientific-advisory-board/en/verification-frontier-ai-models ↩1 ↩2

  55. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, 26 July 2024, updated 8 April 2026. https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence ↩1 ↩2

  56. Thomas R. Tørsløv, Ludvig S. Wier, and Gabriel Zucman, “The Missing Profits of Nations,” NBER Working Paper 24701, June 2018; subsequently published in the Review of Economic Studies (2022). https://www.nber.org/papers/w24701 ↩1 ↩2

  57. Leopold Aschenbrenner, Situational Awareness: The Decade Ahead, June 2024. https://situational-awareness.ai/