Minbook
KO
AI 2040 Plan A: Why 'Just Slow Down' Is Already Too Late, and Where I Choose to Stand

AI 2040 Plan A: Why 'Just Slow Down' Is Already Too Late, and Where I Choose to Stand

M. · · 12 min read

A “best plan” its own authors believe in only 3-15% of the time left me with something other than a policy debate. AI 2040 Plan A recommends a verified slowdown to push superintelligence out to 2040. But the real weight of the piece is why “just slow down” is already too late, and where a single person should stand once the window for verification has closed.

AI 2040: Plan A is a scenario from the AI Futures Project. It is the same team, Daniel Kokotajlo among them, that became known last year for “AI 2027.” The plan-a-root route is what unfolds in the interactive branching scenario when you reach the 2029 “Choose a Path” fork and pick A out of five options (A/B/C/D/S).

I spent a long conversation with this piece. It started as an autopsy of a policy scenario and, by the end, had turned into a conversation about myself. I am keeping that arc intact.

Five branches, and a best plan the authors themselves don’t believe

The first distinction to get right: “AI 2027” was a prediction of what seemed likely to happen, and its endings were either extinction from loss of control or an irreversible concentration of power. Plan A is the opposite kind of document. It is not “this is what will happen” but a recommendation for what should be done.

The authors nail this down themselves. The implementation part is a recommendation, not something they expect to actually occur; only the subsequent effects are a prediction. Daniel personally thinks reality will move faster than this scenario. The point is to put a policy on top of a concrete scenario so its own holes become visible. Instead of only tearing apart other people’s policy, you put your own on the test bench first.

Plan A tries to block two failure modes at once.

  • Uncontrolled superintelligence: if the race continues, even the “winner” cannot open much of a lead and will not voluntarily slow down for safety, so control is lost.
  • Concentration of power: even if alignment (making the AI follow human intent) succeeds, one person or a small group commanding a superintelligent workforce for a few months is itself a takeover of the world.

Even a scenario where alignment succeeds can end badly. Where most safety discourse stops at “how do we make the AI obey,” this team pushes to “who holds the perfectly obedient AI.” That is where the 2040 in the title comes from.

YearEvent
2029The US and China agree to avoid a reckless superintelligence race
2030Without the deal, AI research fully automates and reaches superintelligence within the year. The deal averts it
2030-2035Scaling slowly within the range of human experts
2035A deliberate pause at top-human-expert level to keep human control
2040Unpause, and scaling toward superintelligence

The backbone of the deal is a verified slowdown plus total research transparency. Instead of a secret race, dozens of companies across several countries scale slowly together. Four mechanisms hold it up: total research transparency, verification of whether the other side keeps the deal, a deterrent called MACD (Mutually Assured Compute Destruction) that lets each side destroy the other’s compute if the deal breaks, and a scaling strategy that raises capability only within the human range. MACD is the nuclear MAD (Mutually Assured Destruction) moved onto compute.

The authors even scored all five plans on the same yardstick. These are Eli Lifland’s median estimates.

MetricPlan APlan B-kineticPlan C+Plan CPlan D
Takeoff length (AC to takeover-capable)6 yr3 yr1.5 yr1.13 yr1.02 yr
Training-compute safety tax (OOMs)2.810.460.130.02
Safety-researcher-years1M2,0001,500500200
Power distribution (1-5)51221
p(alignment success)72%50%45%40%25%
p(great future)42%25%25%20%10%

AC in the table means the stage of fully automated AI research, and takeoff length is the time from there until AI reaches a takeover-capable level. Plan B leads by sabotaging China (B-kinetic goes as far as physical strikes) and then spends the lead on safety; Plan C slows via regulation and export controls, or a leading firm voluntarily spends part of its lead on safety; Plan D is essentially a full-speed race; Plan S is an indefinite halt.

On the numbers alone, A wins by a wide margin. Then one table overturns all of it: the probability each author assigns to each plan actually happening.

AuthorPlan APlan D
Daniel15%30%
Eli5%25%
Romeo8%30%
Ryan4%28%
Thomas3%50%

Every author rates the odds of their own best plan at 3-15%, and treats the race (Plan D) as the most likely outcome. “This is the best, but it probably won’t happen.” The separation of recommendation from prediction shows up right here. It is intellectual honesty and, at the same time, a weakness as a policy proposal, because “a 5%-likely optimum” is a hard sell to any policymaker.

The real trap in “just slow down”

At this point my own cynicism kicked in. In the end, how is this different from “let’s all agree not to build nukes”?

The answer that came back: it is harder than that, not the same as that. The goal of nuclear nonproliferation is “don’t build” (halt), so the object of verification is clear: enrichment, yes or no. Plan A’s goal, by contrast, is “build, but slowly.” Translated back to nukes, it is closer to “keep adding warheads, but publish each other’s blueprints and only synchronize the production speed.” That is not disarmament; it is joint arms management, and it needs far more precise mutual trust. On top of that, the first country to get a nuke cannot end the world (deterrence balance), whereas with AI, if recursive self-improvement is real, a few months of lead hardens into permanent dominance. That nonlinearity makes the incentive to break a slowdown deal much larger than for nukes.

That much was logic. But from my own hands-on sense, something snags before the logic even starts.

Model release cycles have sped up partly because capability is rising, but I read it mostly as the time to verify before shipping having nearly vanished. This is not to say nobody cares about guardrails. It is that the speed at which things now ship is a speed verification simply cannot keep up with. If so, the “verified” in “verified slowdown” is already half fiction. The distance between “build slowly” and “don’t build” shrinks toward zero, at least in political feasibility.

Dig into it and the implication is worse than it first looks. “Slowly” mixes two different axes.

  • Deployment cadence: the interval of shipping outward. The verify-then-release window has been compressed. This is what I observed.
  • Internal R&D loop: the speed at which capability itself climbs. Recursive self-improvement happens here. This is what Plan A actually wants to stop.

And the verification window has not literally gone to zero; competitive pressure has quietly redefined verification downward. The primary sources of the last year or two point exactly this way.

  • Anthropic Responsible Scaling Policy (RSP) v3.0 (effective 2026-02-24): it split the AI R&D capability threshold into two levels and moved part of the future AI Safety Level (ASL) safeguard plans into non-binding descriptions. It explicitly aims for “realistic unilateral commitments that are achievable in the current environment.”
  • OpenAI Preparedness Framework v2 (2025-04-15): it simplified the gating thresholds down to two, High and Critical.
  • Even the government’s pre-deployment evaluation body dropped “Safety” from its name. In June 2025 the US Commerce Department renamed the US AI Safety Institute to CAISI (Center for AI Standards and Innovation) and shifted its focus to “demonstrable risks” such as cyber, bio, and chemical, along with national security and competitiveness.

So the actual mechanism behind my feeling that “there is no time to verify” is not the disappearance of time but a narrowing of scope, and a shift from pre-deployment gating to post-deployment monitoring. The regime changed from “pass it before release” to “ship it, watch, and patch.”

Here is the problem. Post-hoc monitoring works on reversible harm. If a chatbot says something strange, you patch it. That is why the current shipping cadence is survivable: most harm can be undone. But the one thing Plan A cares about is the irreversible kind, namely loss of control. Loss of control cannot be patched after the fact.

---
config:
  look: handDrawn
  theme: neutral
---
flowchart TB
    subgraph EXT["External deployment (what I observed)"]
        E1["User, press, reputation pressure present"]
        E2["Some verification pressure remains"]
        E3["Takeover risk low (mostly reversible)"]
    end
    subgraph INT["Internal R&D loop (Plan A's target)"]
        I1["No external scrutiny"]
        I2["Pure speed race"]
        I3["Takeover risk originates here (irreversible)"]
    end
    EXT -->|"Where the verify window already collapsed"| INT
    INT -->|"And here it is even more hopeless"| RESULT["A domain where ship-then-patch is impossible in principle"]

And this collapse is worst exactly where Plan A needs it most. If the verify window has broken even in external deployment, the one place with some public scrutiny, then verification in the internal R&D loop that nobody watches is structurally more hopeless still. My observation does not refute Plan A; it sharpens the pessimism about the very part Plan A depends on.

So is “slowly” possible? Not impossible, logically. There is exactly one condition: only when an external forcing function is strong enough to overpower the competitive gradient. And here “verified slowdown” is not about lifting your foot off the pedal; it is a demand to reattach a brake that has already been torn off. It asks the industry to reverse the flow of the past two years, in which pre-deployment gating was swapped for post-hoc monitoring and thresholds were softened into non-binding targets. That is a far heavier order than slowing down. Plan A being at 3-15% is not only about a lack of will; it is also this structural cost of reversal.

Plan C, and why the gradient beats individual will

Of the five plans, I actually put a little more hope in Plan C. If it becomes winner-take-all, then let the winner be a sensible winner. For that, many people have to keep raising the issue and demanding, and the public has to keep studying so it does not grow ignorant.

I did not stake this hope on individual goodwill. I paired it with an external lever: the public’s sustained pressure. That is the right target. The real content of Plan C is not “let us pray for a good winner” but laying down, in advance, the pressure and the informed public that make the sensible path the one of least resistance when the winner reaches the moment of choice. That version is not naive.

But here an old memory snags: a documentary I saw a while ago, “The Social Dilemma.” Until I watched it, I had never once thought about the “like” button. That one feature, which survived the A/B tests by keeping people just a little longer, made so many social problems and shifted so many standards of value. Its effect on the generation whose values had not yet formed was something I never even considered. A decision by a handful of people in Silicon Valley went that far.

That very button shows the trap. The people who built it were neither stupid nor malicious. They were all sensible individuals. The problem was not individual sense but that the optimization target (time on site, surviving the A/B test) quietly overwrote the stated values. Nobody decided “let us change society this way.” The gradient decided.

So if “let us be sensible this time” means individual resolve, that is precisely the method that already failed. The like-button team had resolve too. The AI race is the same shape. However sensible the winner is personally, the gradient pulling them (competition, “raise capability”) is stronger than individual will. The retreat we saw earlier, from pre-deployment gating to post-hoc monitoring, was not done by bad people; it was done by the gradient. So the strong version of my hope reads like this: not the winner’s character, but an external forcing function built in advance, strong enough to beat that gradient. The public education and sustained pressure I named are that thing, and they bear more weight than character.

The real flaw is timing. The feedback loop that partly corrected social media was slow: from the like button (2009) through “The Social Dilemma” (2020) to a few belated regulations. And that was only possible because the harm was legible and reversible. It took years of visible damage before the public caught up. But what Plan C tries to block is fast and irreversible. There is simply no time for that correction loop to run. The point where the public “understands” an AI takeover the way it understood the like button is already after the fact, and after the fact does not work on the irreversible.

So the ground for hope and the ground for pessimism come from the same observation. The formula that worked for social media, from late awakening to late correction, structurally fails for AI. The only version that works is to front-load the awakening: to make the norms and pressure exist before the irreversible moment arrives. Doing beforehand what took more than a decade last time is a much heavier order, and unlike the like button, there is no retake.

There is one asymmetry in my favor, though. When social media arrived, nobody knew it was a civilizational issue; it was just a fun app. AI, at the same stage, is already treated openly as a big deal, with evaluation bodies, government pre-deployment agreements, and people paying attention attached to it. Far more attention is already front-loaded than social media had at the equivalent stage. The very fact that I am writing this now is a condition the like button never had.

So where do I choose to stand

By this point the conversation had come down to myself. I had been talking about the gradient beating individual will at the scale of organizations and nations, and I ended up applying it to me.

Honestly, since coding agents arrived in earnest early this year, my life has genuinely changed. I do not think I can go back to before. The way I think about many things has shifted, and the way I live has shifted. I do not know what change is coming, but because I know how big it is, I want to respond, and yet the people around me who run the way I do are fewer than I expected, so there is little shared ground, it is unclear how to prepare, and there is a gap with the actual work I do. The gap between the capability rising up and what I can absorb. In the end my own ability will be the limit, and how to raise it within that limit has always been the worry.

But this one frame is the thing I had to flip, and the conversation is where I learned it. “My ability is the limit, so how do I raise it” is a race I cannot win by construction. The premise is that capability rises faster than me, so trying to keep pace with raw ability is chasing an asymptote that the agent catches up to again next month. Coding speed is a resetting asset. However much you raise it, it resets.

So it is better to change the question. There is no way off the treadmill. But you can choose where to stand so that each step compounds rather than resets. What the agent eats is implementation speed; what it cannot eat yet is the judgment and the taste for choosing what to build and why. My anxiety right now comes from clutching the layer that is becoming a commodity (coding speed) while sitting on top of the layer whose value is rising (judgment) without counting that asset. Move the allocation of learning from the resetting side to the compounding side, and the same anxiety becomes less draining.

One more layer of honesty. That judgment layer is no permanent safe harbor either. If the water keeps rising, taste too eventually gets modeled. There is no permanent safe harbor. The treadmill sensation I feel is accurate and cannot be removed. Only continually moving where I stand toward the compounding side is the sustainable version.

And the point from this conversation that stayed with me the longest was elsewhere.

Having worked in performance marketing and come to understand recommendation logic, I try hard to stay neutral. When I lean one way, I deliberately turn the other, taking various actions that influence the algorithm. I thought this was a refined habit. But I had not seen the trap in it. Understanding the mechanism creates the illusion of immunity. The moment I believe I am not swayed because I know the algorithm, I become more defenseless in the areas where I am not spending vigilance, because I have spent the whole vigilance budget on the domain I understand. On top of that, “deliberately turning the other way” assumes I know which direction I am being pushed, so the overcorrection becomes another bias in itself. In the end it is a structure where I grade my own bias.

So the conclusion inverts. What I should value is not my capacity for neutrality but the external input that disagrees with me. External counter-input breaks the loop of grading myself. Running fast and having few peers ties in here too. Lose peer calibration and you lose the way to tell “conviction” from “being right.” The faster you run, the less you should trust your own conviction and the more you should seek counter-input, and yet people usually go the other way: lonely, they lean harder on conviction.

Closing

Plan A is a best case its own authors believe in only 3-15% of the time. Half of that low number is a problem of will, but the other half is a problem of structure: the time to verify has already vanished. “Just slow down” is not lifting your foot off the pedal but reattaching a brake that was torn off, and even the hope of Plan C runs into a wall of timing, because the danger moves faster than the correction loop can turn. So the only remaining version is to front-load the awakening, and this time there is no retake.

What this argument left me was not an answer in policy but a stance. Do not clutch the layer becoming a commodity; keep moving to stand on the layer that compounds. And do not claim immunity just because I understand the mechanism. If the gradient beats individual will, the one thing an individual can do is trust their own conviction less and seek counter-input more. That is what “not growing ignorant,” a phrase at the scale of a nation, is called at the scale of a person.


Sources: AI 2040: Plan A (AI Futures Project) · Anthropic Responsible Scaling Policy v3.0 · OpenAI Preparedness Framework v2 · US AI Safety Institute renamed CAISI

Share

Related Posts