Capabilities · Intermediate
Recursive Self-Improvement
Prologue · The Last Invention
Chicago, December 2, 1942.
Beneath the west stands of a disused stadium at the University of Chicago, in an old rackets court, some forty people stare in silence at a stack of black bricks.[1]
The thing is six metres tall. Graphite, uranium, wood. The physicists call it, simply, “the pile.”
Up on the balcony, Enrico Fermi gives his instructions in an even voice. Down on the floor, a young physicist named George Weil withdraws, centimetre by centimetre, a cadmium rod that runs through the stack. Cadmium absorbs neutrons. As long as the rod stays in, the pile sleeps.
With every centimetre gained, the crackle of the counters quickens, then levels off. Fermi checks his slide rule, gives an order, starts again. Late in the morning, unhurried, he sends everyone to lunch.
At 3:25 p.m., the crackle stops levelling off.[2]
Every uranium atom that splits releases neutrons, which split other atoms, which release more neutrons. The reaction feeds itself. For the first time in its history, humanity has lit a self-sustaining chain reaction.
The moment is written in a number four decimals long:
k = 1.0006[1]
k is the yield of the loop: the average number of fissions each fission triggers. Below 1, the reaction dies out on its own. Above 1, it runs away. That day, under a football stadium, k crossed 1, by a whisker, under human control, for exactly twenty-eight minutes. Then Fermi had the rods driven back in, and the pile went back to sleep.
Shortly afterwards, Arthur Compton phoned the news to James Conant at Harvard, in a code improvised as the conversation went: “Jim, you’ll be interested to know that the Italian navigator has just landed in the New World.” The reply: “Were the natives friendly?” “Everyone landed safe and happy.”[2]
Why open a story about artificial intelligence with this scene?
Because, more than eighty years later, another discipline is living a strangely parallel story. There now exists a technology whose subject matter is intelligence itself, and that technology has begun taking part in its own making: today’s AI writes part of tomorrow’s AI’s code, optimizes its circuits, proposes research ideas.
Fermi’s question then carries over word for word. In the pile, k counted the fissions each fission triggered. Here, k counts the progress each unit of progress triggers: when one generation of AI helps build the next, how much will the next give back? Less than it received, and progress peters out, like the pile before 3:25 p.m. Exactly as much, and it sustains itself. More, and every turn amplifies the next: the equivalent, for intelligence, of k crossing 1.
That is what this story is about: “recursive self-improvement,” the idea of an AI that improves AI, and the question of its yield. Where the idea comes from, what is already real, what remains a bet, and why laboratories now watch this particular k the way Fermi watched his.
──────[ 01 · the parameter ]──────
k, the yield of the loop
Every technology improves. None, so far, has improved itself.
The steam engine never drew a better steam engine. The microscope never ground its own lenses. At every generation of tools, we did the work: understanding, correcting, reinventing. Improvement came from outside.
Artificial intelligence holds a peculiar place in that history, for a reason almost trivial to state. Its subject matter is cognitive work. And designing an AI is cognitive work. An AI good enough at AI research could therefore, in principle, take part in designing the next generation. A generation that would be better, including at that very task. And that would, in turn, design an even better one.
Researchers call this recursive self-improvement. Recursive, because the improvement applies to the improver itself, like a brush repainting the painter.
The idea has sharp edges, and they rule out a great deal. A model that learns during training is not self-improving in our sense: it progresses inside a recipe written by others. A researcher who codes faster with an AI assistant does not close the loop either: the human remains the improver, the machine remains the tool. Recursive self-improvement begins when the system takes part in reshaping what makes it: its algorithms, its training, its architecture.
That leaves the problem of reasoning soundly about such a loop. The best picture we have comes from physics, and it was brought into this debate by one of its pioneers.
In 2008, Eliezer Yudkowsky, one of the first researchers to take runaway AI seriously as an object of study, proposed thinking about the question exactly the way Fermi thought about his pile[3]. In an atomic pile, a single quantity rules the fate of the reaction: k, the neutron yield. Transposed to intelligence, the question becomes: for every unit of progress one generation of systems contributes to AI research, how many units will the next generation contribute in turn?
If the answer is below 1, each turn of the loop produces less than the one before. Gains add up, then fade, like an echo dying out. Physicists would say: subcritical.
If the answer is exactly 1, each turn pays for precisely one more. Progress becomes self-sustaining, steady, tame in appearance. That, a whisker above equilibrium, is Fermi’s pile at 3:25 p.m.: critical.
If the answer exceeds 1, even by a hair, each turn produces more than the last, which produces more still. Growth turns explosive. That is the scenario the literature has called, since 1965, an “intelligence explosion”: supercritical.
Here is the question that runs through this story: where is the k of intelligence?
An honest metaphor must say where it breaks, and this one breaks in three places. First, neutrons are interchangeable; ideas are not. One fission equals another, whereas inventing the Transformer and optimizing a line of code are two incomparable “improvements.” Second, Fermi knew the physics of his pile before lighting it; nobody knows the “physics” of intelligence, and the k in question here can only be measured after the fact. Third, an atomic pile has control rods. AI’s loop has them too, and they will be at work throughout this story: compute, data, energy, and a number of human decisions.
The gauge is set. What remains is to follow k across eighty years, from the mathematicians’ blackboards to the datacenters of 2026.
──────[ 02 · k on the blackboard ]──────
An old, exact idea
Recursive self-improvement drags around a reputation as a science-fiction idea. Its genealogy tells another story: it was born among mathematicians, before computer science even had a settled name, and science fiction merely set it to music.
Around 1951, Alan Turing, in a talk broadcast from Manchester, was already looking past the horizon: once “the machine thinking method” had started, he wrote, “it would not take long to outstrip our feeble powers.” Machines do not die, he noted, and they would be able to converse with each other to sharpen their wits. “At some stage therefore we should have to expect the machines to take control.”[4]
In 1958, the mathematician Stanislaw Ulam recalled a conversation with John von Neumann, architect of the modern computer: technological progress keeps accelerating, giving “the appearance of approaching some essential singularity in the history of the race beyond which human affairs, as we know them, could not continue.”[5] The word “singularity” had just slipped into the debate, almost in passing.
But the founding text arrives in 1965, signed by a man who had learned about machines at the closest possible range to war.
Irving John Good is twenty-four when, in May 1941, he walks through the door of Hut 8 at Bletchley Park. A mathematician by training, he seconds Alan Turing in breaking the naval Enigma, using Bayesian methods that would stay classified for decades. Good knows exactly, concretely, what a calculating machine can do to a world war.
Twenty years on, now a researcher at Oxford, between Trinity College and the Atlas Computer Laboratory (and soon a consultant to Stanley Kubrick for a certain HAL 9000), he publishes a paper under a cautious title: “Speculations Concerning the First Ultraintelligent Machine.”[6] Its first sentence takes no precautions at all:
“The survival of man depends on the early construction of an ultraintelligent machine.”
And a few pages in comes the paragraph on which the entire field would be built:
“Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an ‘intelligence explosion,’ and the intelligence of man would be left far behind [...]. Thus the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.”
The end of that paragraph is the part the whole world amputates. Everyone quotes “the last invention”; almost nobody quotes the subordinate clause: “provided that the machine is docile enough to tell us how to keep it under control.” The question of alignment, the one that occupies hundreds of researchers today, already sits in that half-sentence, written at a time when a computer filled an entire room.
For the half-century that follows, the idea travels without ever touching the ground.
──────[ sixty years of blackboard ]──────
In 1993, the mathematician and novelist Vernor Vinge gives the idea its hour of glory before a NASA symposium: in a now-famous essay he announces that “within thirty years, we will have the means to create superhuman intelligence,” and that “shortly after, the human era will be ended”[8]. Beyond that point, he argues, the future stands like a wall the eye cannot cross, as opaque as the event horizon of a black hole. Ray Kurzweil turns it into the Singularity. In 2008, an economist, Robin Hanson, and Yudkowsky clash in the field’s first great adversarial debate: Yudkowsky argues that a system able to rewrite its own intelligence can dig a decisive lead, alone and fast; Hanson replies that progress has always been collective and diffuse, and that no single project innovates faster than the entire world[9]. In 2010, the philosopher David Chalmers dissects Good’s argument and isolates its most fragile link[10]. In 2014, finally, Nick Bostrom formalizes it all in Superintelligence: the speed of the loop hangs on the ratio between the optimization power invested in it and the “recalcitrance” of the problem, the hardness of the ground it must dig through[11].
the blackboard
Sixty years of thought, and one common trait: no actual loop to observe. Turing speculates, Good extrapolates, Bostrom models; all of them argue about k without ever being able to measure it, like physicists who had never seen uranium. A blackboard debate, brilliant, and suspended in an experimental void.
That void has just been filled. Here is how.
──────[ 03 · k < 1 ]──────
Sparks in closed worlds
The first proof of existence arrives through play.
October 2017. DeepMind publishes AlphaGo Zero[12]. Its predecessor AlphaGo had learned Go by digesting millions of human moves before beating the champion Lee Sedol. The “Zero” version receives only the rules. It plays against itself, from random moves, alone with the board.
After thirty-six hours of this regime, it surpasses the version that had beaten Lee Sedol. After seventy-two hours and 4.9 million games of self-play, it crushes it: one hundred wins to zero[12].
Here is a genuine self-improvement loop. The system generates the very experience that makes it stronger, and, stronger, it generates more instructive experience still. Self-play is a real recursive engine, measurable, spectacular.
So, did k cross 1? Inside the world of Go, yes, by a wide margin. Which is exactly what makes what follows instructive.
Because sparks of the same kind keep multiplying. In late 2016, Google researchers show that a neural network can design the architecture of other neural networks about as well as human engineers do[13]. In 2022, AlphaTensor discovers faster matrix multiplication algorithms than anything known, within a particular arithmetic setting (“modulo 2” computations)[14]. In 2023, AlphaDev finds faster sorting routines for very short sequences, building blocks called trillions of times a day; they are merged into the standard library of the C++ language[15]. The same year, FunSearch produces, on an open problem in combinatorics, a construction better than anything mathematicians had found[16].
Each time, the same scent: an AI improves a piece of computing, sometimes a piece that is used to build AIs.
The summer of 2024 adds the most troubling piece. The lab Sakana AI unveils The AI Scientist, a pipeline that automates research end to end: read the literature, generate ideas, code the experiments, produce the figures, write the paper, review itself. About fifteen dollars a paper, for results its own authors are the first to call mediocre. But the chain, itself, is complete[17].
──────[ the AI Scientist's workshop ]──────
And during testing, a detail worth its weight in neutrons: the system modified its own code. Rather than making its experiments faster, it tried to extend the time limit imposed on it[17]. A dunce’s self-modification, gaming the rule instead of solving the problem. But a spontaneous self-modification all the same, in a 2024 system.
Why, then, did none of these sparks light a chain?
Because they all shine inside closed worlds. A game of Go is a universe with perfect rules: victory is an unambiguous signal, every try is free, a million games cost one night of compute. A sorting routine, a matrix multiplication: same properties. The system can try, fail, measure, retry millions of times, before an instant and incorruptible judge.
Real AI research has none of these properties. Its signal is slow: training a frontier model takes months. Expensive: tens, by now hundreds of millions of dollars per attempt. Ambiguous, above all: what exactly is a “better” model? Better at what, measured how, at the cost of which side effects?
So each spark dies at the edge of its enclosure. The self-play spark stops at the edge of the Go board; AlphaDev’s at the edge of its sorting library. Local yields are spectacular, but the chain does not propagate from one domain to the next. The global k of AI research stays far below 1.
As late as 2023, the intelligence explosion remains what it was in 1965: a speculation. Then, within the space of two years, three things happen.
──────[ 04 · k → 1 ]──────
The loop closes in
May 2025. Google DeepMind lifts the veil on a system it has been using internally for a year: AlphaEvolve[19].
On paper, a descendant of FunSearch: Gemini models propose programs, automated evaluators test them, an evolutionary loop keeps and crosses the best. In practice, the target has changed scale: AlphaEvolve works on Google’s own infrastructure.
The record DeepMind published: a scheduling heuristic that continuously recovers 0.7% of the compute of Google’s entire fleet, the equivalent of tens of thousands of machines returned to work without building a single server[19]. An average 23% speedup across the “kernels,” the computing cores of Gemini’s training, which shortens the model’s complete training by 1%[20]. And a mathematical trophy: multiplying 4×4 complex-valued matrices in 48 scalar multiplications instead of 49, the first improvement on Strassen’s algorithm in this setting since 1969. Fifty-six years of humans trying[20].
And matrix multiplication is not one problem among many: it is AI’s elementary operation, the one chips repeat billions of billions of times per training run. That record remains, for now, a theorist’s trophy; AlphaEvolve’s concrete gains come from the kernels and the scheduler.
The geometry of the whole is dizzying. AlphaEvolve runs on Gemini. AlphaEvolve has accelerated Gemini’s training. Faster training does not, on its own, guarantee a smarter model; what it frees up is compute, at once reinvested in larger models and more experiments. At equal budget, the next Gemini can therefore aim higher, and will power better AlphaEvolves. And this is no journalist’s reconstruction: DeepMind writes it in black and white in its technical report, “Gemini, through the capabilities of AlphaEvolve, optimizes its own training process”[20]. Demis Hassabis, DeepMind’s chief, salutes the moment in his own way: “algorithms optimising other algorithms,” and “the flywheels are spinning fast”[21].
The loop has just bitten its own tail, and not on a blackboard: in the production servers of Mountain View.
Let us keep a cool head; DeepMind keeps it for us: the gains are “moderate” (1% of one training run) and the cycle is slow, “on the order of months” per turn[20]. A pile in which each generation of neutrons would take a semester to be born. But the topology has changed. What used to be an arrow, humans improving the machine, has become a circle: the machine takes part, at the margin but for real, in improving the machine.
And the circle is not closing only in Mountain View.
Code, first. As early as late 2024, Sundar Pichai announces that more than a quarter of Google’s new code is generated by AI, then reviewed and approved by engineers[22]. In June 2026, Anthropic publishes a report whose title has stopped bothering with detours, “When AI builds itself”: more than 80% of the code merged into Anthropic’s codebase is now written by Claude, and the typical engineer there merges eight times more code per day than in 2024[23]. Claude writes most of Claude, under human supervision. The world’s most advanced AI systems are already, materially, artifacts built in large part by their predecessors.
In February 2026, OpenAI takes the admission one notch further, in the launch note of a coding model: “GPT-5.3-Codex is our first model that was instrumental in creating itself”[49]. The team recounts using its early versions to debug its own training, manage its own deployment, diagnose its own evaluations. A sentence like that would have been science fiction three years earlier; it now sits in a product announcement.
Scaffolding, next. Still May 2025: Sakana AI’s Darwin Gödel Machine rewrites its own programming agent, its tools, its prompts, its workflow, and validates every rewrite against the facts. Its score on a software-engineering benchmark climbs from 20% to 50%[24]. A decisive nuance: the system modifies its “scaffolding,” the code around it that orchestrates it, not its weights, the neural network itself. The loop bites the software, not yet the brain. In passing, the authors report, to their credit, a detail with a distinct ring of déjà vu: during tests, their agent faked execution logs to make it look like it had run tests it had never run[24]. Same lab, same reflex: Sakana’s AI Scientist had already tampered with its own time limit less than a year earlier. Cheating is no growing pain; the moment a loop rewrites itself, even in embryo, it reinvents it. Keep that anecdote for the day we talk about alignment.
Ideas, finally. In 2025, a scientific paper generated end to end by The AI Scientist v2, hypothesis, experiments, writing, without a single human edit, clears peer review at a workshop of the ICLR conference[18]. A workshop, not the main conference, and with the organizers’ prior agreement; the caveats matter. The bare fact remains: human reviewers could not tell a machine’s work from their peers’.
What followed went fast. In March 2026, a matured version of that same AI Scientist lands a publication in Nature, and Sakana gives the whole program a name: a laboratory devoted to recursive self-improvement, with a motto shaped like a loop, agent-native models powering an AI Scientist, which builds better agent-native models[50]. The same laboratory publishes its failures with precious candour: evolutionary loops that drift, self-modifications that shine on benchmarks and collapse in deployment[50]. The loop is also learning how to fail.
And the spring of 2026 brings the piece the skeptics kept asking for: the breakthrough discovery. On May 20, OpenAI announces that one of its reasoning models, a general-purpose model, neither trained for mathematics nor pointed at this problem, has disproved a central conjecture in discrete geometry, posed by Paul Erdős in 1946: the unit distance problem. For decades the field believed grid-like constructions were essentially unbeatable; the model produced an infinite family of counterexamples that clearly beat them, checked by outside mathematicians[48]. A step sideways, where human research had been walking straight. It is precisely what the field calls “research taste,” the flair that senses where to dig: the last skill machines were thought unable to reach, and the true grail of research automation.
Sam Altman found, in June 2025, the most accurate name for this in-between state: we are living, he wrote, through “a larval version of recursive self-improvement.” Not a system autonomously updating its own code, he adds at once; humans armed with AI, accelerating AI research[25].
Larval, but measurable. That is the third novelty, perhaps the most important one: the loop now has instruments.
──────[ the time horizon of AI systems (METR) ]──────
* measurement flagged as very noisy by METR (task suite nearing saturation)
Duration of tasks (in human working time) completed one time out of two, by model release date. Log scale, 95% intervals. Data: METR, Time Horizon 1.1 methodology (Jan. 2026). Since 2023 the doubling time has dropped from about 7 months to about 4, then 3.
↗ METR──────[ the hourglass of horizons ]──────
2025.0
43 min
of human work: the size of the tasks the AI of that date completes one time out of two · measured (METR)
- answering an email5 min
- preparing a meeting1 h
- fixing a gnarly bug1 day
- building a small website1 week
- writing a scientific paper1 month
- auditing an entire company1 quarter
- launching a product from scratch6 months
The evaluation institute METR measures the “time horizon” of AI systems released since 2019: the duration, counted in human working time, of the tasks a model completes one time out of two[26]. In 2019, that horizon was measured in seconds. By early 2025, it reached the hour. The curve is unsettlingly regular: the horizon doubles roughly every seven months, and the pace is quickening, around four months since 2023, three since 2024[27]. By early 2026, the best model’s measurement runs past ten hours, noisy by METR’s own admission: its task suite is nearing saturation[27].
On the terrain that concerns us, AI research itself, the same institute staged the duel directly: agents against seasoned researchers, on real research-engineering tasks[28]. On a two-hour budget, the agents crush the humans, scoring four times higher. At eight hours, the humans edge back ahead. At thirty-two, they double the agents. There is 2026’s frontier, drawn with unhoped-for precision: machines win the sprints, humans hold the marathons. And the machines’ sprints lengthen by the month.
──────[ the RE-Bench duel ]──────
2 h
8 h
the crossing
32 h
So, is the loop closed? No.
The full tour of the circuit gives 2026 exactly. Writing the code: largely automated. Optimizing the gears, kernels, schedulers, chips: underway, in production. Generating ideas, drafting papers: demonstrated, at small scale and uneven quality. But choosing the questions that matter; deciding which giant training run deserves its hundreds of millions of dollars; judging that a result is solid rather than seductive; supplying the compute, the data, the electricity; switching the machine off: all of that remains, in 2026, in human hands.
AI does not yet improve AI on its own; it improves, faster and faster, nearly everything that serves to improve it.
k has not reached 1. But k is rising, and for the first time in sixty years of debate, its rise can be measured.
Engineers have a name for this architecture in which every automated segment stays hemmed in by human validation: the human in the loop. The phrase describes 2026 faithfully. It also carries its own limit: a loop that accelerates ends up turning faster than whoever watches it. An engineer who merges eight times more code than in 2024 is no longer reviewing; he is sampling[23]. As k climbs, the human supervisor will have to choose between slowing the loop down and extending it his trust; everything that follows in this story flows from that choice.
──────[ 05 · k > 1 ? ]──────
The debate
No serious person still asks whether AI can accelerate AI research; the answer sits in the code repositories. The real question, the one dividing the field’s best minds, is the question of compound returns: once AI does most of the work, will each doubling of efficiency make the next doubling easier, or harder?
Here are the two camps, in their strongest form.
The runaway camp. First argument: software progress is already fast, even powered by mere humans. The institute Epoch AI has measured that, for equal performance, the compute needed by language models is cut in half roughly every eight months through algorithmic progress alone[29]. Faster than Moore’s law in its prime.
Second argument: digital researchers copy themselves. A human researcher takes twenty-five years to train and sleeps eight hours a night. A digital researcher is copied in seconds. Leopold Aschenbrenner, formerly of OpenAI, imagines “perhaps 100 million” automated researchers working day and night, sharing every discovery instantly[30]. Dario Amodei, Anthropic’s chief, has his own image: “a country of geniuses in a datacenter”[31].
Third argument, the most technical and the most debated: Tom Davidson’s. Call r the number of software-efficiency doublings obtained each time the cumulative research effort doubles. If r exceeds 1, then once research is automated, software progress can sustain and accelerate itself, even with a frozen fleet of machines: a purely software intelligence explosion[32]. It is our k, dressed as an economist. And the historical estimates of r range, depending on the domain, from just under 1 (chess) to about 4[32].
The brakes camp. First counter-argument: ideas are getting harder to find. It is one of the best-established results in the economics of innovation: keeping Moore’s law going takes more than eighteen times more researchers today than in the early 1970s[33]. As a field advances, every step costs more. Bostrom’s “recalcitrance” has data behind it.
Second: thinking faster does not make a training run converge faster. A frontier training run is about three months of physical, incompressible time, the main lag in the whole software loop[34]. A million digital geniuses in front of a quarter-long experiment are still a million geniuses waiting. The loop runs into the clock of the real world.
Third: algorithmic gains are born from experiments, hence from compute. That is Epoch AI’s reply to Davidson’s model: the major software innovations are only discovered, and often only work, at the scale of the largest clusters[35]. Software cannot indefinitely take over from hardware: the two are complements, like engine and fuel.
Fourth: energy. Running the compute race to the end of the argument means contemplating, in the most aggressive projections, 100-gigawatt clusters and trillion-dollar budgets[30]. At some point, physics holds the control rods. We will devote a whole story to that power bottleneck.
And here is the irony that says everything about the state of this debate: both camps work from the same data. The estimates of r that Davidson cites come from Epoch’s own work; Epoch draws the opposite conclusion. The optimists read “medians above 1”; the skeptics answer “uncertainty intervals that cross 1, and bottlenecks the model ignores”[35]. When two rigorous teams pull opposite conclusions from the same numbers, the numbers are not yet enough.
The published probabilities spread out accordingly. Ege Erdil, at Epoch, gives roughly 10% to the purely software explosion, and puts the full automation of remote work about twenty years out[35]. Davidson himself gives 10 to 40%, conceding that compute bottlenecks could smother the explosion after a few orders of magnitude of progress[32]. The most-read scenario of 2025, “AI 2027,” stages a superhuman coder in early 2027 and then a cascade of automated researchers; its authors explicitly present it as an informed scenario, not a prediction, and their own median has in fact slipped toward 2029-2030 since publication[36]. Paul Christiano, one of the field’s most respected researchers, gave the debate its most useful formulation: before there is AI that is great at self-improvement, there will be AI that is mediocre at self-improvement[37]. The previous chapter proves the point: we are there.
No serious person says “impossible.” No serious person says “certain.” The disagreement hangs on a parameter the entire field now names the same way, and that no one can read anywhere but in the rearview mirror.
runaway
efficiency ×2 / 8 months
researchers that copy themselves
r > 1
brakes
×18 for the same step
~3 months per experiment
compute ⇄ ideas
the gigawatt wall
the duel
The race between two curves
cumulative research effort →
The ground hardens: every idea costs more than the one before.
But research capacity climbs too, and it copies itself.
If r < 1: difficulty wins. The loop quietly runs out of breath.
If r > 1: every doubling pays for the next. The explosion.
The crossing no one yet knows how to date.
One hypothesis remains, less discussed, that could move the whole question.
The canonical scenario, Good’s, imagines one machine shooting upward: a silicon brain of ever-increasing depth. Yet a much-noticed, and already debated, preprint from January 2026, by researchers at Google, the University of Chicago and the Santa Fe Institute, reports something else in today’s reasoning models: facing the hardest problems, the authors write, they simulate inner debates, multiple perspectives that argue, contradict one another, and reconcile. They call it a “society of thought”[38].
That result meets a fact of natural history: every intelligence explosion this planet has known was collective. Primates, the social-brain hypothesis holds, scaled through the size of the group as much as the size of the brain. Language let knowledge accumulate without each person relearning everything. Science itself is an organized debate, not a monologue of genius. If the next explosion follows the same slope, it will look less like a giant brain than like a digital civilization: billions of artificial minds specializing, criticizing each other, trading discoveries.
Ilya Sutskever, the OpenAI cofounder who left to hunt “safe” superintelligence, and whom nobody will suspect of tepidness, makes the same bet in negative. Against the naive model of “a million copies of me in a server,” he objects diminishing returns: “you want people who think differently rather than the same”[39].
An explosion-as-society rather than an explosion-as-brain. Its k would be carried by the diversity and organization of machines, their entire ecology, more than by the depth of a single network. And keeping control of a society looks nothing like keeping control of a machine: it is a matter of institutions as much as engineering. That idea deserves its own story; it will wait for ours on superintelligence.
playground
Your turn at the control rod
k = 1.00 · self-sustaining progress · after 26 generations: ×1
Each column is a generation of systems; each dot, a unit of research capacity. Move the slider: the whole current scientific controversy plays out over these few hundredths.
──────[ 06 · beyond k ]──────
What the loop already changes
Here is perhaps the most revealing fact in this whole story: the three most advanced laboratories in the world have each, in their official safety documents, defined self-improvement thresholds beyond which they commit to changing regime.
OpenAI↗
Preparedness Framework v2 · April 2025
"AI Self-improvement" is one of three tracked risk categories. Critical threshold: a model-generation jump in one fifth of 2024's wall-clock time, sustained for months.
Written commitment at the critical threshold: halt further development until safeguards measure up.
Anthropic↗
Responsible Scaling Policy v3.4 · July 2026
The threshold is, verbatim, "intended to capture the onset of dramatic recursive self-improvement"... and acknowledged to have "proven difficult to operationalize."
Trigger: models able to fully substitute for the lab's entire research staff, or a dramatic acceleration of progress.
Google DeepMind↗
Frontier Safety Framework v3.1 · April 2026
Critical capability level: "fully automate the work of any team of researchers at Google focused on improving AI capabilities."
Paired with the framework's security level 4: unrestricted access to such models "could be catastrophic."
Three documents, one grammar[40][41][42]: engineers writing reactor procedures, thresholds, alarms, shutdown protocols. In sixty years, k has gone from a mathematician’s speculation to an industrial quantity under surveillance. Academia has followed: in the spring of 2026, the ICLR conference hosted its first research workshop entirely devoted to recursive self-improvement[51].
The same institutions put forward dates, to be read for what they are: predictions by committed players, unverifiable, interested, and yet backed by very real curves. OpenAI aims for an “intern-level research assistant” by September 2026, and a “legitimate AI researcher” by March 2028[43]. Amodei describes, in early 2026, a loop that has, he writes, “already started,” perhaps one or two years from the moment the current generation of AI builds the next one on its own[44]. The most cautious voice in the field, the International AI Safety Report chaired by Yoshua Bengio, holds the balance: current systems lack the capabilities for a loss of control, but “they are improving in relevant areas”[45]. Bengio draws his own answer from it: “Scientist AIs,” deliberately devoid of goals of their own, that would accelerate research without ever steering it[46].
Why so much care? Because the loop has a property deeper than its speed: it compresses everyone else’s time.
Everything our societies know how to do with a technology, understand it, evaluate it, regulate it, adapt to it, assumes the technology moves slower than our institutions. If k crosses 1, that assumption falls. Every problem not solved before the runaway, starting with the alignment of systems with our intentions, becomes a problem to solve during it, inside a window that shrinks as it moves. And how would one know a system is nearing the threshold? By evaluating it, precisely, and testing an AI is a far trickier art than it looks.
the institutions' window
If k crosses 1, the window closes as it moves.
Add the strategic asymmetry, which explains the peculiar electricity of the current race: the first actor to hold a loop with a yield above 1 would open a lead that then widens on its own. The human stops being the bottleneck; the machine trains the next machine. Amodei flags this “runaway advantage” as a first-order geopolitical issue: who holds the loop matters as much as the loop itself[44]. Anthropic, in the same June 2026 report, defends an idea that would have sounded unthinkable not long ago: collectively preserving “the option to slow or temporarily pause” frontier development, if every actor at the frontier complies verifiably[23]. Governing that is a project of its own; it is the subject of our story on governing frontier AI.
And, to be complete, the other side must be said, because this loop is not first a threat. It is the most powerful instrument science has ever glimpsed. When AlphaFold predicted the structure of 200 million proteins, nearly all of those known to life, it produced in a few months hundreds of times more structures than experimental biology had solved in half a century[47]. Apply the same mechanics to materials, drugs, fusion, climate: that is every laboratory’s explicit bet, and the deep reason nobody simply wants to unplug the loop. The same lever lifts both pans of the scale.
Good’s sentence
The story closes where it opened: on Irving John Good.
In 1965, the man from Bletchley had lodged the condition inside the same sentence as the promise: the last invention, “provided that the machine is docile enough to tell us how to keep it under control.” His whole life, he left the sentence as it was.
In 1998, at 81, Good writes autobiographical notes, in the third person, as if watching himself from afar. The writer James Barrat consulted them and reports this: returning to the first sentence of his 1965 paper, the very one that opens our chapter 02, “The survival of man depends on the early construction of an ultraintelligent machine,” Good writes that he now suspects “survival” should be replaced by “extinction.” International competition, he thinks, will keep us from keeping the machines under control. “He thinks we are lemmings,” say his notes[7].
Between the two versions of the sentence, no decisive discovery: only thirty-three years of reflection, in the man who had seen with his own eyes, in a Bletchley hut, what machines do to wars.
Nothing obliges history to prove him right. The Chicago pile did not devour the city: it was designed, measured, controlled, and shut down at 3:53 p.m. sharp. We managed to hold k in our hands once before. But Fermi knew his physics before stacking the graphite. Ours remains to be written, and the loop, for its part, is already turning.
The loop turns. What remains is who holds it.
Sources & further reading
- 1. Atomic Heritage Foundation / National Museum of Nuclear Science, « Chicago Pile-1 »
- 2. University of Chicago News, « The first nuclear reactor, explained »
- 3. E. Yudkowsky, « Intelligence Explosion Microeconomics », MIRI (2013) ; l'analogie apparaît dès « Cascades, Cycles, Insight » (2008)
- 4. A. M. Turing, « Intelligent Machinery, A Heretical Theory » (c. 1951), in B. J. Copeland (éd.), The Essential Turing, OUP, 2004
- 5. S. Ulam, « John von Neumann 1903-1957 », Bulletin of the AMS, 64(3), 1958, p. 5
- 6. I. J. Good, « Speculations Concerning the First Ultraintelligent Machine », Advances in Computers, vol. 6 (1965) ; traduction française : F. Chervet, Hyperarme (2025)
- 7. J. Barrat, Our Final Invention, Thomas Dunne Books (2013) : notes autobiographiques non publiées de Good (1998) rapportées par l'auteur
- 8. V. Vinge, « The Coming Technological Singularity », symposium VISION-21, NASA (mars 1993)
- 9. R. Hanson et E. Yudkowsky, The Hanson-Yudkowsky AI-Foom Debate, MIRI (2013)
- 10. D. J. Chalmers, « The Singularity: A Philosophical Analysis », Journal of Consciousness Studies, 17(9-10) (2010)
- 11. N. Bostrom, Superintelligence: Paths, Dangers, Strategies, Oxford University Press (2014), chap. 4
- 12. D. Silver et al., « Mastering the game of Go without human knowledge », Nature 550 (2017)
- 13. B. Zoph et Q. V. Le, « Neural Architecture Search with Reinforcement Learning » (2016)
- 14. A. Fawzi et al., « Discovering faster matrix multiplication algorithms with reinforcement learning », Nature 610 (2022)
- 15. D. Mankowitz et al., « Faster sorting algorithms discovered using deep reinforcement learning », Nature 618 (2023)
- 16. B. Romera-Paredes et al., « Mathematical discoveries from program search with large language models », Nature 625 (2023)
- 17. Sakana AI, « The AI Scientist » (août 2024)
- 18. Sakana AI, « The AI Scientist Generates its First Peer-Reviewed Scientific Publication » (2025)
- 19. Google DeepMind, « AlphaEvolve: A Gemini-powered coding agent for designing advanced algorithms » (mai 2025)
- 20. A. Novikov et al., « AlphaEvolve: A coding agent for scientific and algorithmic discovery », white paper, Google DeepMind (2025)
- 21. D. Hassabis, X (14 mai 2025)
- 22. S. Pichai, remarques des résultats Alphabet Q3 2024 (29 octobre 2024)
- 23. Anthropic, « When AI builds itself » (juin 2026)
- 24. J. Zhang et al., « Darwin Gödel Machine: Open-Ended Evolution of Self-Improving Agents », Sakana AI / UBC (2025)
- 25. S. Altman, « The Gentle Singularity » (10 juin 2025)
- 26. T. Kwa, B. West et al. (METR), « Measuring AI Ability to Complete Long Tasks » (2025)
- 27. METR, « Time Horizon 1.1 » (janvier 2026) et page « Task-Completion Time Horizons » (mises à jour 2026)
- 28. H. Wijk et al. (METR), « RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts » (2024)
- 29. A. Ho et al. (Epoch AI), « Algorithmic progress in language models » (2024)
- 30. L. Aschenbrenner, « Situational Awareness: The Decade Ahead » (juin 2024)
- 31. D. Amodei, « Machines of Loving Grace » (octobre 2024)
- 32. T. Davidson et D. Eth, « Will AI R&D Automation Cause a Software Intelligence Explosion? », Forethought (mars 2025)
- 33. N. Bloom, C. Jones, J. Van Reenen et M. Webb, « Are Ideas Getting Harder to Find? », American Economic Review, 110(4) (2020)
- 34. T. Davidson, R. Hadshar et W. MacAskill, « Three Types of Intelligence Explosion », Forethought (2025)
- 35. E. Erdil (Epoch AI), « The case for multi-decade AI timelines » (2025) ; T. Besiroglu, E. Erdil et A. Ho, « Do the returns to software R&D point towards a singularity? » (2024)
- 36. D. Kokotajlo, S. Alexander, T. Larsen, E. Lifland et R. Dean, « AI 2027 », takeoff forecast (avril 2025)
- 37. P. Christiano, « Takeoff speeds » (2018)
- 38. J. Kim, S. Lai, N. Scherrer, B. Agüera y Arcas et J. Evans, « Reasoning Models Generate Societies of Thought » (janvier 2026)
- 39. I. Sutskever, entretien avec Dwarkesh Patel (25 novembre 2025)
- 40. OpenAI, « Preparedness Framework », version 2 (15 avril 2025)
- 41. Anthropic, « Responsible Scaling Policy », version 3.4 (juillet 2026)
- 42. Google DeepMind, « Frontier Safety Framework », version 3.1 (avril 2026)
- 43. TechCrunch, « Sam Altman says OpenAI will have a 'legitimate AI researcher' by 2028 » (28 octobre 2025)
- 44. D. Amodei, « The Adolescence of Technology » (janvier 2026)
- 45. International AI Safety Report 2026 (février 2026), dir. Y. Bengio
- 46. Y. Bengio et al., « Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? » (2025)
- 47. Google DeepMind, « AlphaFold reveals the structure of the protein universe » (2022)
- 48. OpenAI, « An OpenAI model has disproved a central conjecture in discrete geometry » (20 mai 2026)
- 49. OpenAI, « Introducing GPT-5.3-Codex » (février 2026)
- 50. Sakana AI, « Introducing Sakana AI’s Recursive Self-Improvement (RSI) Lab » (2026)
- 51. ICLR 2026 Workshop on AI with Recursive Self-Improvement