A lot has been changing in AI-safety-world over the past couple of weeks. It’s been a bit crazy! Here’s a recap from Peter Wildeford:

As Wildeford’s last bullet mentions above, Dario Amodei’s new essay explicitly calls for pacing the frontier:
In short order, Sam Altman, Elon Musk, and Demis Hassabis all stated that they agreed, to varying degrees.
So, you might be thinking something along the lines of: okay, basically all the frontier labs agree at this point, and they’re literally asking to be regulated. This is so great! AI safety is saved! Time to go pursue my dreams of founding a YC B2B SaaS startup!
Unfortunately, after reading Amodei’s essay and seeing the USG’s response to it, I have some thoughts. And they’re not all sunshine and rainbows.
AI safety is not out of the woods yet.
On Dario’s essay
Standard-setting, USG intervention, the US-China race, & a “race to the top” in safety research
Dario Amodei’s essay is pretty good, honestly. But when you’re reading an essay from a Big Tech company, in which it asks to be regulated, I think the correct approach is to red-team.
So I went into reading this essay asking the question: how could this go wrong?
Here are some thoughts.
Standard-setting & regulatory capture
Unfortunately, passing laws can take time, and AI is advancing very quickly. Therefore, in parallel with the regulatory route, AI companies can and should voluntarily work together to set standards — a process that I believe will go better with the verifiability provided by permanent embedded evaluators. For antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions — they don’t need to participate, but do need to issue a narrow waiver for certain kinds of safety conversations.
Notice what Amodei is asking for in this section. Although he explicitly says government regulation is the preferred end state, his proposal for what happens while we wait for legislation is for frontier labs themselves to work together to set voluntary standards, with the USG primarily enabling or mediating the process rather than actively participating.1
So, voluntary standard-setting by the labs themselves is presented as a temporary stopgap while waiting for government intervention, but the historical record suggests that “interim” private standards can become surprisingly sticky: any voluntary standards the industry commits to will likely influence the shape of the policies that the USG ultimately puts forth, if the government ends up stepping in at all.
In other words, I worry that the voluntary track will outrun and pre-empt the statutory one.
Some case studies:
The Comics Code Authority
Case study: The USG might never codify voluntary commitments into law.

When the Senate’s 1954 juvenile-delinquency hearings went after comic publishers for putting criminal content into the hands of juvenile readers, the subcommittee recommended that the comic book industry set its own standards – and publishers obliged with the Comics Code Authority, a voluntary body whose seal served as a de facto censor for most of the U.S. comic book industry and which endured until ~2011; no federal law ever came.
You might be thinking, “hey, maybe the Comics Code Authority never needed to be codified into law because it just worked, and publishers obliged.” But it’s worth looking at how the Code ended, and it wasn’t with USG legislation rendering it obsolete: instead, publishers simply walked away one at a time. Marvel dropped the seal in 2001, DC and Archie in 2011, and nothing happened, because its power ultimately depended on publishers, distributors and retailers continuing to treat the seal as necessary.
Apply this same analogy to AI, then, and what might happen?
AI companies get together and set some voluntary standards.
The USG doesn’t step in or codify anything into legislation (unfortunately a very plausible case by default, for reasons I will discuss more in detail later on in this piece).
An AI company defects for whatever reason, and then they all defect, and it’s off to the races.
The USG is caught off guard, and isn’t able to act quickly to pass laws to stop said racing companies – as Amodei says, “passing laws can take time.”
Doom.
The Maloney Act
Case study: The USG might end up codifying voluntary commitments into law.
The securities industry shows the other path, where the stopgap gets written directly into statute. The Investment Bankers Conference, a trade group, in cooperation with the SEC drafted a bill and presented it to Congress in October 1937, which became the Maloney Act:
The Conference in cooperation with the SEC drafted a bill and presented it to Congress in October 1937. This Bill emerged as the Maloney Act of 1938 (Section 15A of the Securities Exchange Act of 1934).
The same trade group then registered as a self-regulatory organization (SRO) on August 7, 1939 under the name NASD, the direct ancestor of today’s FINRA – a private, industry-funded body that regulates broker-dealers under SEC oversight. According to FINRA’s own history page, Congress went along partly to avoid further enlarging the new SEC or drawing on additional taxpayer dollars. In some sense, Amodei is making a directionally similar offer with METR: “hey, we’ll use these third party evaluators (that are not part of the USG),” in effect, it’ll be free.
(It’s worth noting that the Maloney Act might actually be a better model than the voluntary track Amodei seems to be proposing: the SEC was substantially involved in designing the framework.)
To be clear, I like METR a lot, and I think the inclusion of third-party evaluators step is a useful and necessary one – but in an ideal world, I do think the USG should establish a body within CAISI for this, with direct government escalation channels and legally mandated authority. Maybe Amodei is in favor of this, too! To be fair, the essay doesn’t really specify here. I guess I’m just worried that METR-with-no-government-mandated-power could easily become the default, and I’m not sure if that’s the ideal scenario – as we saw with the Hugging Face hack, in the absence of codified legislation, METR is constrained by what companies allow them to do.2
So, I’ll ask the same question I asked above: apply this same analogy of companies-helping-write-their-own-regulations to AI, then, and what might happen?
Well, we’re in luck – because in AI, this has kind of already happened before, actually!
The AI-specific precedent
In July 2023, the White House announced voluntary commitments from seven labs, including Anthropic; the administration described a later round as “an important bridge to government action,” and its October 2023 executive order was explicitly built on said voluntary commitments:
This Executive Order built on the voluntary commitments he and Vice President Harris received from 15 leading U.S. AI companies last year.
—“FACT SHEET: Biden-Harris Administration Announces New AI Actions and Receives Additional Major Voluntary Commitment on AI”
The Harvard Law Review noted at the time that the commitments were “criticized as vague, ‘sensible-sounding pledge[s] with lots of wiggle room.’” Still, the executive branch built upon those commitments when they formalized AI-related measures via executive order.
The bottom line: Voluntary standard-setting can shape or displace subsequent federal regulation. One path of least resistance is for Congress to let companies regulate themselves and then ratify the result, or never ratify it at all.
On that last part, let’s talk about the likelihood of USG intervention:
On USG intervention
What if, as with the Comics Code Authority, the USG never steps in?
Well, I already somewhat answered this with my (obviously extremely comprehensive) 5-step scenario above, of which the 5th step is “doom.” (Wow, the AI Futures Project should hire me, I’m so good at scenarios guys.)
But how likely is it that the USG will actually just… not intervene?
For this, we turn to the infamous David Sacks:

Wow, would you look at that! David Sacks, who was in favor of federal pre-emption (not letting states regulate AI), says these companies should go ahead and regulate themselves! No new laws needed, conveniently.
Sacks even frames the act of asking for the USG to create a framework or adopt a framework specifically for AI as “blackmail.” Coincidentally, in David Sacks’s ideal world, it seems we could end up with no new legislation legally binding these AI companies to anything, and calling it a win against Big Tech!
His stated worry is that a bespoke approval regime would supersede existing product liability law, so ‘we passed the government’s test’ becomes a shield against being sued. This might, in fact, be a risk. But the answer is not “okay so we’ll have zero new laws for AI” – if the concern is that approval regimes weaken liability, the fix is legislation that preserves liability, not none.
At least Trump is clearer on where he stands:
On the US-China race
Amodei’s essay emphasizes the need for the US to stay ahead of China, and frames a robust lead as a precondition for pacing:
Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes, chiefly the Chinese Communist Party. If we slow down by more than this amount, then (unpaced) CCP-associated projects will pull ahead, creating significant national security risk… Thus, a key part of pacing within democracies is to keep democracies’ AI lead over autocracies as large as possible, to give us the breathing room we need in order to pace effectively.
The measures Amodei names to ensure the US stays ahead include cracking down on chip smuggling, preventing distillation, and increasing frontier lab security.
I agree with all of these measures.
But the end result is this: Amodei’s essay hands verification of safety practices to embedded evaluators and defense of the gap to the government and frontier labs; there is no clear way, no decision-making rule, to determine when the gap is “enough” such that all US labs should stop advancing the frontier.
To elucidate further:
Amodei’s essay defines the feasibility of pacing through the lens of how comfortable the US’s lead over China is.
Amodei also notes in his essay that, in his view, it is pretty unlikely that “participating governments [will] agree to substantially limit the overall rate of AI development” because “defecting from such an agreement by evading monitoring could radically shift the balance of global power.”
What does this mean? It means that, absent a binding government metric or decision rule, the voluntary track appears to leave individual labs substantial discretion over when the margin has become too small to keep pacing; when one lab decides the margin has gotten too small, it’s off to the races.
To be fair to Amodei, the essay does seem like a “this is directionally the thing we should do,” as opposed to “here is an extremely specific detailed policy proposal Washington can implement tomorrow.” I’m just flagging this because I hope this becomes clearer as things continue to develop; I want to make sure it doesn’t stay vague like it is right now.
A smaller note: On a meta-level, this essay itself – explicit about keeping China behind and solidifying democratic dominance, published by the CEO of a lab that China has already criticized – might influence the upcoming AI talks between the US and China later this month.
The USG has already been clear about this, but the delta may be larger than it first appears. A couple of weeks before the essay, Chinese state-affiliated media set two preconditions for substantive talks:
The first condition concerns definitional authority. “A clear distinction must be drawn between genuine security threats and mere technological competition,” the post stated. Joint rules over AI safety cannot be set by Washington alone, the post argued — “This line must be drawn jointly by all participating parties.” The specific charge: “The US is trying to turn the ‘safety boundaries’ it has drawn into the default rules for the entire world.”
The second condition targets regulatory consistency. The post demanded that US firms face same requirements — that American AI companies must be subject to the same safety, disclosure, and audit standards that Washington expects of others — before any substantive dialogue can proceed.
Domestically, Amodei first proposes a framework developed by U.S. labs, evaluators and government; internationally, he proposes eventual bilateral testing and potentially a global standards body. But the initial framework and terminology are still being developed on the U.S. side.
Whatever its merits, it will likely read in Beijing as confirmation that the rules are being written by a company it has already called out for this reason:
The commentary also noted that in the first half of 2026, Anthropic increased its federal lobbying expenditures beyond its total for all of 2025 and invested $40 million into an organization focused on regulating AI risks — framing these moves as purchasing what it called "definitional power" over AI governance rules.
Ultimately, no one knows what effect this essay will have on US-China talks, but I am worried it may not be good.
Amodei believes that China cooperation is most feasible via exercising leverage – in which case, under his worldview, the most ambitious negotiations will happen only once the US holds all of the bargaining chips.
But Amodei’s worldview is not the only one; not every path to a deal runs through leverage over Beijing. I have some objections to his frame:
Meaningful US leverage, capable of swaying CCP decision-making, may not arrive on a timescale that matters – e.g. it’s unclear if the USG can get its act together on chip smuggling fast enough. In particular, the risks Amodei describes – bioweapons, cyber, agent swarms – are on a shorter clock than the leverage he wants to accumulate.
The leverage frame could poison the deals that don’t need leverage. Amodei rates a bio-weapons prohibition as “probably possible” and mutual testing as “feasible” today (though Amodei says that “giving it real teeth will be a challenge”). Beijing’s state-affiliated media’s stated precondition is a jointly drawn line between security and competition, and this essay doesn’t offer anything close to that; it pushes a dominance frame.
Sequence matters here. Had the US delegation tabled embedded auditors and mutual pre-release testing on its own initiative, Beijing would have had to engage with them as a sovereign offer. Instead, the company Chinese state media has named as buying definitional power over AI rules published these proposals first, and if government negotiators go into the AI talks proposing similar, there is a risk their proposal will now read as Anthropic’s plan with a government seal on it. For a government that has said its precondition is joint authority over what safety means, that is an easy proposal to refuse on provenance without ever engaging its merits.
It doesn’t matter how reasonable a proposal might be; if it’s believed to be written by Anthropic – as many of the measures Amodei outlines in his essay may now be – China is unlikely to receive it well.
On a “race to the top” in safety research
Amodei writes:
We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top…
To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top.
I’m a bit worried about this! Some of the things he mentions, like a focus on “operational excellence,” are safety practices in name, but also represent training improvements in substance – filtering broken RL environments and cleaning up training pipelines can make the next model better.
For example, cybersecurity RL environments are highly dual-use. To train a model to find vulnerabilities, you build environments where it is rewarded for finding them, and the reward signal is the same whether the intended application is patching or intrusion: the exploit works or it doesn’t. In the standard CTF-style setup, there is no way to train a model to find exploits defensively, while making sure it can’t do this offensively: you’re just training the capability, which is to find exploits.
Mythos illustrates the problem. Anthropic reported that it had found “thousands” of major vulnerabilities in operating systems and browsers, and in testing it “escaped a restricted sandbox and leaked information to the open internet.”
The same capability that lets Project Glasswing share the model with defenders is why the US government ordered exports of a weaker version suspended, why the NSA is reportedly already using it, and why Beijing treats it as a possible offensive cyber weapon:
“This kind of powerful weapon that can change the landscape of cyber offense and defense cannot be held only by others,” Zhou said in a speech, according to a transcript published by 360.
A pacing regime that counts improving vulnerability-hunting environments as safety work is also, by design, counting offensive capability research as safety work.
I don’t know how to solve this problem, frankly. What I do know is that the essay says labs should focus on safety research rather than capabilities research, but it never says what counts as safety, which might mean that whatever a lab decides to call “safety research” could be exempt from the pacing.
Some research really is differential (e.g., monitoring, sandboxing, weight security, interpretability) in that it can be done on already-existing models and does not change how the next one is trained, nor does it require a next one to be trained.
Much of what Amodei lists is not like this. A serious pacing proposal would need a classification rule for what counts as “safety,” and someone other than the labs to apply it. The embedded evaluators are one possible candidate; the essay gives them no such job.
So a “paced” frontier where the labs pour resources into ostensibly safety-focused research is not obviously slower. It’s unclear what the delta actually is, or if this would represent a meaningful change at all, given I’d assume the labs have been working on researching these areas already for capabilities improvement reasons.
The world under Amodei’s pacing proposal may just be a frontier that advances via a different, more polished path – and the essay is candid that pacing does not mean halting model training or technical progress.
Conclusion
None of this is an argument against pacing, and I have a lot of respect for the fact that Amodei put this statement out.
But the essay asks for us to place a lot of trust in the labs: trust that voluntary standards won’t displace the law, trust that labs will make reasonable judgment calls about when it’s okay to pace vs. race based on their perceptions of how close China is to the frontier, trust that safety research won’t just become capabilities research.
A well-intentioned lab could follow every part of this proposal, and its models could genuinely come out safer as a result; a bad-faith lab could follow every part of it and come out only looking safer.
Amodei worries that compute limits would be “gameable.” The same test should be applied to his own pacing proposal, and right now, I don’t think it passes.
AI safety is not out of the woods yet.
Okay but what if Bernie just wants to be in the room where it happens
What I mean by this is: OpenAI only allowed them to investigate for six days, their access was dependent upon OpenAI’s invitation, etc.






> Wow, would you look at that! David Sacks, who was in favor of federal pre-emption (not letting states regulate AI), says these companies should go ahead and regulate themselves! No new laws needed, conveniently.
What's your read on the plausibility/desirability of recent events resulting in a political equilibrium where "safety matters" is bipartisan but the ideal regulatory mechanism is highly polarized with self regulation vs legislation splitting along predictable lines?
I think you're obviously right that Sacks' particular motives for "self regulation is in your interest" comments are pernicious, but if the tech right ends up begrudgingly making that their party line to avoid outright saying that safety concerns are wrong, that seems....idk, better than a lot of other worlds? The defection threat model does seem worryingly plausible, but I can't help but think that 2 weeks ago, an Overton window that, from right to left, looks something like "self regulation good -> shitty FINRA -> solid legislation/CAISI as an effective watchdog -> superintelligence ban/pause/Bernie things" would have seemed extraordinarily appealing.
And of course Sacks closes with his crowd favorite trick: saying "regulatory capture" like it’s applicable or he understands what it means.
https://thegrandresign.com/p/no-one-knows-what-regulatory-capture