Read · · 22 min read
What the Google “RSI” rumor is, what recursive self-improvement looks like inside the labs today, and where it breaks
A Lab Read on the week’s loudest acronym. Diary framing throughout: this is how we are reading the story and what we are watching, not what anyone should do with it.
The week the acronym went viral
On 9 September an account known for AI leaks posted five words: “huge congRatulationS Indeed! @GoogleDeepMind.” The capitals spell RSI, short for recursive self-improvement. Within hours a screenshot went round that showed an internal model list with an entry called rsi-model-liverl-le.
That is the whole rumor. Four words, a screenshot nobody has authenticated, and no comment from Google or DeepMind. By the weekend it had a Reddit thread with a thousand comments, a crypto-trading feed treating it as a market event, and a trending topic on X.
On its own it would have lasted an afternoon. What made it catch fire is the four weeks around it.

On 12 August Reuters reported that a Google co-founder had pushed DeepMind to prioritise recursive self-improvement in order to speed up Gemini, alongside a leadership reshuffle that moved the DeepMind chief executive into a chair role and put Google’s AI chief in day-to-day charge of the model. On 3 September Fortune counted four Gemini Flash releases in 106 days and no flagship. A DeepMind executive had already described Alphabet’s 2026 capital spending of $195–205 billion as a precursor to recursive self-improvement, because the loop needs compute to run experiments.
On 6 and 7 September OpenAI said it had reached its “automated research intern” milestone on schedule, and put a number on it: 3.1 agent-workdays for every human workday inside its research organisation. On 8 September a senior pretraining researcher resigned from Anthropic with a public warning that the labs are racing toward self-improving systems. On 10 and 11 September more than a dozen OpenAI and Anthropic researchers called for a slowdown, and the White House dismissed the concern.
On 12 September Anthropic’s chief executive published “We Must Pace the Frontier”, proposing embedded third-party evaluators, common standards, and a speed limit on recursive self-improvement. OpenAI’s chief executive committed to evaluator access the same day. xAI’s replied in three words: “Dario is right.” Three competing labs agreeing to slow something is rarer than any leak.
So the rumor landed on prepared ground. Add the unconfirmed talk that Gemini 4 finished pre-training early, and a story assembled itself: Google has closed the loop.
Our read, up front: nothing verifiable says it has. Plenty says every lab is running the open version of the loop today, and that the people running it are worried about the closed one. The rest of this piece is about the difference.
Inside this study
The full piece — free with the C account — works through:
- What recursive self-improvement means: three rungs — bounded self-refinement, the weak form with a human in the loop, and the closed loop the rumor claims.
- The weak form, as it runs today — what Anthropic, OpenAI and Google have said on the record about AI building AI, with the numbers.
- The strong forms and the methods that reach for them — verified self-training, evolutionary code search, automated research agents, self-adapting weights, test-time search. Where each one works and where it stops.
- Why Google, and what Google is working on — chips, data, a self-play tradition, AlphaEvolve inside the training stack, and the reasons for scepticism.
- Why the loop is still open — compute for experiments, the data wall, task horizon, and the problem of grading your own homework.
- The dangers, including drift — model collapse, model and concept drift, reward hacking, evaluator drift, silent failure, and the loss of oversight.
- How the labs try to catch it — thought monitoring, activation probes, deliberative alignment, inoculation, and the hole in each.
- The race that keeps the pedal down — closed labs, open weights, and why a speed limit is sincere and fragile at the same time.
- What it means for the tape — compute, verification, ground-truth data, and the software layer, read through the Rubin build-out and the Agentic Winners.
- How we hold it — what would count as evidence, and what we watch.
Headline read: every frontier lab is on rung two of a three-rung ladder. The rumor is about rung three. The difference is whether a human still decides what the next model is, and whether anyone can check the answer the machine grades itself with.