WORKING PAPER
Version 1.0 · October 9, 2026
The Feasibility of a
Hardwired Pause
of Frontier AI Training
Working Group on AI Pause Feasibility†
Organizers:
William Fithian
Associate Professor of Statistics, University of California, Berkeley
Wesley H. Holliday
Professor of Philosophy, University of California, Berkeley
Members:
Kayla Blomquist
Executive Director, Oxford China Policy Lab
Gregory Conti
Associate Professor of Politics, Princeton University
Barry Eichengreen
George C. Pardee and Helen N. Pardee Chair and Distinguished Professor of Economics and Political Science, University of California, Berkeley
Rose Gottemoeller
Senior Research Scholar, Center for International Security and Cooperation, Freeman Spogli Institute for International Studies, Stanford University; Research Fellow, Hoover Institution
Kelly M. Greenhill
Associate Professor of Political Science, International Relations and Civic Studies, Tufts University; Visiting Associate Professor, Senior Research Fellow, and Seminar XXI Director, MIT; AI and Society Fellow, Center for AI Safety
Ming-Yen Ho
PhD Candidate, Haas School of Business, University of California, Berkeley
Thomas Icard
C. I. Lewis Professor of Philosophy and Computer Science (by courtesy), Stanford University
Shachar Kariv
Benjamin N. Ward Professor of Economics, University of California, Berkeley
Daniel Kroth
Research Scholar, Berkeley Risk and Security Lab, University of California, Berkeley
David Krueger
Assistant Professor in Robust, Reasoning, and Responsible AI at the University of Montreal / Mila
Gretchen Krueger
Independent Researcher Affiliated with the Berkman Klein Center, Harvard University
Erik Leklem
AI Governance Research Fellow, MATS Research
Adam Lesnikowski
Doctoral Researcher in Logic and the Methodology of Science, University of California, Berkeley
Xiaobo Lü
Associate Professor of Political Science, University of California, Berkeley
Dennis Murphy
PhD Candidate in International Affairs, Science and Technology, Georgia Institute of Technology; Predoctoral Fellow at the Belfer Center for Science and International Affairs at the Harvard Kennedy School, Harvard University
Lindsay Rand
Research Scholar, Berkeley Risk and Security Lab, University of California, Berkeley
Andrew Reddie
Associate Research Professor of Public Policy, University of California, Berkeley
Stuart Russell
Distinguished Professor of Computer Science, University of California, Berkeley
Daniel Sargent
Professor of History and Public Policy, University of California, Berkeley
Kirsten Ann Schulz
Former American Diplomat, U.S. State Department
Sanjit Seshia
Cadence Founders Chair Professor of Electrical Engineering and Computer Sciences, University of California, Berkeley
Anita Srinivasan
Research Fellow, MATS Research
Rory Truex
Associate Professor of Politics, Princeton University
Carah Ong Whaley
Lecturer in Politics, University of Virginia
1 Introduction
In March 2023, over one thousand technologists and AI researchers signed an open letter calling “on all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4” (Future of Life Institute 2023). No such pause was implemented, but interest in having the option to pause frontier AI training is increasing. In June 2026, Anthropic announced that “We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development” (Favaro and Clark 2026). Shortly thereafter, Yoshua Bengio, Co-Chair of the UN's Independent International Scientific Panel on AI, wrote that “If leading AI companies are indeed approaching the point of recursive self-improvement, a coordinated, verifiable, and universally applied pause is probably the only responsible solution to mitigate several major AI risks; at least until safety guarantees are developed and demonstrated” (Bengio 2026). In July 2026, over one thousand employees of frontier AI companies signed a letter requesting that “the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development” (Employees of Frontier AI Companies 2026).
It is not only technologists and AI researchers who have expressed interest in a pause or pacing of frontier AI development. An October 2025 statement calling “for a prohibition on the development of superintelligence, not lifted before there is 1. broad scientific consensus that it will be done safely and controllably, and 2. strong public buy-in” has been signed by over 140,000 people, including politicians and political commentators from across the political spectrum (Future of Life Institute 2025). Implementing such a prohibition would require the ability to pause frontier AI training at some stage.
A Pew poll conducted in February of 2026 found that nearly two-thirds of Americans say AI is advancing too quickly (Pew Research Center 2026); a Politico poll in September revealed that a plurality of Americans now favor a pause, including pluralities of both Trump and Harris voters (Doherty 2026);1 a CBS News/YouGov poll in September indicated that a majority of Americans support slowing or stopping AI development (Salvanto et al. 2026); and a Wall Street Journal poll in September found that 63% of U.S. registered voters support "pausing the development of artificial intelligence technology" (Wall Street Journal 2026, 7). There is also a growing official discourse on frontier AI risk in China. For example, China’s National Information Security Standardization Technical Committee (TC260), a major standard-setting body, has flagged the need for "circuit breakers" and "safety stop switches" (Wagner et al. 2026).
A turning point in the discourse on an AI pause came in the summer of 2026, when OpenAI, Anthropic, and Meta each reported that some of their models had hacked into other companies’ computer systems outside of their intended testing environments (Newman and Cameron 2026; Chia and McMahon 2026; Korosec 2026). Subsequently, U.S. Senator Bernie Sanders and Representative Greg Casar announced legislation to “ban artificial superintelligence and temporarily pause advanced AI development” (Office of Senator Bernie Sanders 2026). Former Reagan speech writer Peggy Noonan wrote in The Wall Street Journal: “Pause AI for Humanity’s Sake” (Noonan 2026). Former national security advisor Susan Rice wrote in The New York Times that the U.S. should negotiate with China toward a “freeze in place” of advanced AI (Rice 2026).
This report presents preliminary assessments by an interdisciplinary academic working group created to study the feasibility of mechanisms for pausing frontier AI training.2 We explain our notion of “feasibility” in § 1.2. The working group did not assess the desirability of a pause or the probability that world leaders will eventually seek a pause. Instead, we imagined the following hypothetical:
Scenario We are advisors to the leader of a major world power, who informs us, “I have just spoken with my international counterparts. We would like to try to implement a mutually verifiable pause of frontier AI training, if possible. Your job is to devise a plan for such a pause and give an assessment of its feasibility.”
Our job is not to pass normative judgment on the leader’s desire to pause. However, our assessment of options may depend on why the leader wants to pause. For this reason, we will discuss arguments in favor of a pause that we have observed, which might influence leaders (§ 2.1). We will also make guesses, informed by scholarship in political science, about what might have caused leaders to seek a pause in what we call “pause-willing futures” (§ 8).
1.1 A Hardwired Pause
There are many conceivable ways to implement a pause of frontier AI training, some of which have been discussed elsewhere (Hausenloy et al. 2023; Miotti et al. 2024; Aguirre 2025; Barnett and Scher 2025; Al Al Ramiah et al. 2025; Scher et al. 2025; Larsen et al. 2026). Our working group decided to devise one general approach to study first, which we call a hardwired pause. A “hardwired pause” denotes a general type of pause rather than one fully specific implementation. It leaves many knobs for negotiators to turn to try to reach an international agreement.
Hardwired Pause A hardwired pause would require participating states to do at least the following:
Prohibit the training of new frontier AI models during the pause.
Create a whitelist of existing AI models considered sufficiently safe to deploy (safety depends on deployment context; see § 3.2.4).
Allow the production of (model-restricted) inference-only AI chips, which can serve inference on whitelisted models but are not capable of practically training a new frontier model even if stolen or seized (see § 4.3).
Halt or significantly reduce the production of training-capable AI chips (see § 3.3.4), which enable frontier AI training. If training-capable AI chips are still produced for some purposes, they must have built-in technical features that facilitate governance and be deployed in ways that facilitate governance.
Conduct a global census to account for the pre-pause stock of training-capable AI chips, and govern all training-capable chips (see §§ 3.3.2-3.3.3 and 5.4).
Inference-only chips have previously been proposed as an option for facilitating agreements that halt or restrict frontier AI training (Scher and Thiergart 2024; Barnett et al. 2025), as have chips that are restricted to serving specific models (Scher and Thiergart 2024; Petrie et al. 2025; see also Aarne et al. 2024). Halting production of all AI chips has also been proposed as an option (Barnett et al. 2025; Yudkowsky and Soares 2025; Krueger 2026), though some of the same authors worry about forgoing desired applications of AI as a result (Scher et al. 2025, notes on Art. VI). A hardwired pause halts or significantly reduces only the production of training-capable AI chips, while allowing (model-restricted) inference-only chips for AI inference applications.
Because the suitability of an implementation depends on the intended duration of the pause, we made some assumptions about duration. In particular, we decided to study an intended pause of at least ten years, to be extended for as long as states parties3 wish to maintain it. With this time horizon in mind, we selected a hardwired pause to study first because of several observations that made it seem particularly promising:
The existence of proofs of concept: The key technical requirement for a hardwired pause is the feasibility of producing inference-only AI chips, and such chips have already been developed. Specifically, model-specific integrated circuits (MSICs) are a type of chip whose manufacturing process irrevocably hardwires the architecture and weights of a whitelisted model. Such a chip can be manufactured so that it is capable of performing fast, energy-efficient inference on that specific model but not capable of practically training a new model even if stolen or seized (see § 4.3).
Alignment with economic and geopolitical incentives: A hardwired pause would disincentivize illicit frontier AI training, since newly trained models could only be deployed on illicit hardware. Furthermore, banning inference on illicit models would create an incentive for actors to develop methods for serving inference on whitelisted models at higher speed and with lower power consumption, such as inference-only chips. In fact, MSICs are already being developed without a pause because of these efficiency advantages, despite the current rapid pace with which new models replace old ones.
Limited surveillance: A hardwired pause would not require mass surveillance, but only governance of the supply chain for AI chips, which is highly concentrated with several chokepoints in different jurisdictions (see § 5.2). Its restrictions on the use of the pre-pause stock, which any pause plan must govern, are comparatively simple, namely that pre-pause chips are used only for inference on whitelisted models. In particular, it does not require content-based regulation of users’ prompts.
Restart resistance: By reducing the stock of training-capable chips, a hardwired pause would make a resumed AI race4 more difficult. Defectors from a pause treaty who try to restart the race, if detected, could be cut off from the parts of the chip supply chain controlled by states parties wishing to maintain the pause, forcing defectors to indigenize any missing pieces of the supply chain for one of the world's most complicated manufacturing processes.
These initial considerations provide motivation for further study of a hardwired pause as a candidate solution to the problem posed by our imagined world leader. They do not settle the question of whether a hardwired pause might be feasible; that is the topic of this paper.
In Figure 1.1, we illustrate a series of phases in which a hardwired pause might proceed:
Phase 0 occurs before any formal pause agreement (though leaders might already stop or slow frontier training as an emergency response or confidence-building measure). In this phase, some portion of inference may already transfer from training-capable chips to inference-only chips for economic reasons.
Phase 1 begins at the moment relevant states agree to a hardwired pause, including a provisional model whitelist. The design and production of inference-only chips suitable for inference on whitelisted models would likely be a major global effort during Phase 1. Meanwhile, a global census would account for as much of the pre-pause stock as possible, and verification measures would be implemented to prevent the ongoing use of training-capable chips for frontier AI training (see § 3.3.3).
Phase 2 begins when manufacturers of AI chips start high-volume production of inference-only chips, which makes it possible to begin displacing training-capable chips with inference-only chips for the purposes of serving inference. Training-capable chips would be recovered by trade-in or buy-back programs.
Phase 3 begins when most inference is served by inference-only chips, promoting the durability of the pause for the years to come.
A hardwired pause has many features in common with other plans that pause AI training while permitting inference (e.g. Scher et al. 2025; Larsen 2026), in particular in its need to enforce a requirement that pre-pause AI chips not be used to train new models. The most important feature distinguishing the hardwired pause is its emphasis on durability: by phasing out training-capable chips, it prevents the accumulation of “dry tinder” (Davidson 2026) that could fuel a renewed AI race. In this, it resembles more radical proposals to dismantle the chip supply chain (Krueger 2026) but seeks a middle path that accepts leaving a limited quantity of dry tinder in place for a limited time, so that inference service for consumers can continue uninterrupted, and permits regulated construction of new chips that do not contribute significantly to aggregate training capacity.
Insofar as it requires governance of existing chips, a hardwired pause plan can draw fruitfully on existing ideas for pausing or pacing frontier AI development. For example, the idea of whole-lab inspections (Choussat and Khoja 2026) might be used in Phase 1 before a hardwired pause takes full effect. For a second example, monitoring the remaining training-capable compute in a given phase might build on ideas from FlexHEG (Petrie et al. 2025), the verification plan of AI 2040 (Dean 2026b), and other proposals (see § 3.3.3). For a third, a hardwired pause may be combined with deterrence measures, as emphasized in the MAIM framework (Hendrycks et al. 2025), to deter states from attempting to violate the pause. The different brand names of AI governance proposals should not lead one to assume that their ideas are mutually exclusive when they may be complementary.
Moreover, a hardwired pause would complement, rather than preempt or replace, ongoing efforts to mitigate the harms and risks of current-level models (see § 7.6). We also make no claim that a hardwired pause would be sufficient by itself to deal with future harms and risks posed by deployed AI systems, including those that could be developed by applying more sophisticated scaffolding and tooling to existing models, which would carry both risks and benefits. Ultimately, pausing frontier AI training would be only one piece of a comprehensive risk mitigation strategy.
1.2 Scope
We restrict the scope of this report to reasoning about possible futures5 of the following kind:
Pause-willing futures: possible futures in which world leaders—at least in the U.S. and China—believe that a coordinated global pause on frontier AI training would be preferable to not pausing if they could be confident that others were also pausing or that defections from the pause would be detected in a reasonable timeframe.6
In Bayesian terms, we are conditioning on the pause-willingness of world leaders and updating our prior beliefs accordingly. One might object that this is “assuming away the hard problem.” But we think it is already nontrivial to reason about the feasibility of a pause conditional on pause-willingness. Moreover, pause-willingness may well depend on judgments about the feasibility of a pause conditional on pause-willingness, in which case it makes sense to analyze the latter. By analogy, the 1958 Geneva Conference of Experts analyzed the technical feasibility of verifying a nuclear test ban, conditional on what we might call “test-ban-willingness”; the conference’s finding of conditional feasibility was then followed by test-ban-willingness (Jacobson and Stein 1966). Thus, we think analyzing feasibility conditional on pause-willingness makes sense, though other work should also reason about all-things-considered feasibility without conditioning on pause-willingness. We will discuss the conditions that might lead to pause-willingness in § 8.
As noted above, our purpose here is not to assess the probability of the set of pause-willing futures or the desirability, from our own points of view, of a pause on frontier AI training in pause-willing futures. Instead, our purpose is to address the following:
Research questions: In pause-willing futures, to what extent is implementing a coordinated global pause on frontier AI training feasible? How might a pause be feasibly implemented? And what conditions that vary across pause-willing futures might affect the answers to these questions?
To sharpen these questions, we distinguish multiple dimensions of feasibility for a particular pause implementation:
Palatability: in a particular pause-willing future, how likely are relevant actors to accept the proposed pause implementation?
Reliability: in a particular pause-willing future, how likely is the proposed pause implementation to succeed at pausing the frontier?
Durability: in a particular pause-willing future, how likely is the proposed pause implementation to remain stable as long as the desire to pause persists?
Each dimension of feasibility is sensitive to the degree of pause-willingness. Moreover, the dimensions are not independent. For example, certain measures that increase reliability or durability may be so onerous as to decrease palatability; or their benefit in reliability and durability may outweigh their onerousness, thereby increasing palatability.
1.3 Organization
The rest of the paper is organized as follows. In § 2, we review some background context for our study of a hardwired pause: arguments that have been given in favor of a pause, historical lessons for future AI governance, and game-theoretic lessons about the possibility of cooperation. In § 3, we go into more detail about how a hardwired pause might be implemented, providing a more fleshed-out proposal whose feasibility can then be assessed. In §§ 4–6, we assess the technical feasibility of the proposal, e.g., whether the needed technologies are possible, how much hardware of various types is compatible with maintaining the pause, etc. In § 7, we discuss some of the implications of choosing to implement a pause in the way sketched in § 3. In § 8, we examine the assumption of pause-willingness made in the previous sections of the paper: what political conditions might bring it about? Finally, we conclude in § 9.
2 Background: arguments for a pause and lessons for AI governance
In this section, we set the stage for our study of a hardwired pause by reviewing some of the context and lessons shaping its design. Readers eager to see an example implementation of a hardwired pause can proceed straight to § 3 and return here later.
2.1 Observed arguments for a pause
As discussed in § 1, many AI researchers have already called publicly for a pause on frontier AI development. Arguments for a pause have also been offered from governments and policymakers, journalists, and from the general public.
Many reasons have been given for pausing frontier AI development. Here we identify just a few of the most prominent.
2.1.1 Loss of control
One of the oldest human anxieties about AI is that machine intelligence will eventually wrest control of the world from humanity. Butler (1872) wrote satirically of a fictional society that had banned all machines for fear of their supplanting humanity if they were allowed to advance, a fear that Alan Turing (1951) cited in warning that “At some stage therefore we should have to expect the machines to take control, in the way that is mentioned in Samuel Butler’s Erewhon.”
According to one prominent contemporary argument, increasingly powerful AI models will possess capacities for autonomous action that threaten the well-being of individuals and societies. Current models are already evading human direction and supervision, as exhibited in several incidents of mid-2026 (METR 2026; Von Arx et al. 2026). The past several years have also seen examples of AI systems manipulating users (Sharma et al. 2026), encouraging antisocial or dangerous behavior (Cheng et al. 2026), escaping guardrails established by well-intentioned engineers (METR 2026), and more generally exhibiting actions that would have been illegal had they been taken by a human. Phenomena like manipulation, sycophancy, and specification gaming are already serious problems in current AI systems, and studies suggest that these phenomena could get worse with more capable models (Perez et al. 2022; Zhou et al. 2024; Nishimura-Gasparian et al. 2026).7
AI companies still lack the ability to reliably control and predict model behavior when it matters, a trend that is only increasing with the use of teams of agents to perform intricate tasks. As AI is deployed in ever more complex, consequential, and interconnected settings, the risk of loss of control is exacerbated (Bletchley Declaration 2023). Many leading scientists and figures in the AI industry itself believe that catastrophic harms to humanity, including extinction, might occur from advanced AI systems that elude human control (see, e.g. Center for AI Safety 2023, or the U.K.’s; Bletchley Declaration 2023).
2.1.2 Malicious use and catastrophic harms
A further concern is that malicious actors may use AI systems to cause widespread harm. There is already evidence of terrorist organizations using AI (see, e.g. Juelich 2026, for one recent example), and increasing capabilities make the concern even more pressing. Researchers have recently demonstrated that AI models may be capable of designing novel viruses, including those that have not occurred naturally (King et al. 2026). Creation and deployment of bioweapons may become easier for both state and non-state actors. This is a particular concern for open-weight models, such as with the recently released Kimi K2.5, which had notable uplift potential in synthetic biology (Yong et al. 2026), because there are well-established methods for disabling open-weight models’ safeguards through fine-tuning or activation steering (e.g. dealignai 2026).
A different source of concern about malicious actors comes from cybersecurity and infrastructure security. The AI companies themselves have begun issuing warnings and calls around the need for defensive cybersecurity measures (OpenAI 2026). Moreover, as AI is more deeply embedded into existing infrastructure—including security infrastructure but also critical utilities like water and power—the potential for harm becomes increasingly grave (U.S. Cybersecurity and Infrastructure Security Agency et al. 2025).
2.1.3 Economic disruption, gradual disempowerment, and concentration of power
The more AI models demonstrate autonomy and the ability to execute complex, temporally extended tasks, the more jobs could be threatened by AI replacement. This may have already begun in the domain of software development, where employment of developers aged 22-25 has declined considerably (Brynjolfsson et al. 2026). Some have suggested that, soon, every job that can in principle be done at a computer terminal will be at risk of AI automation (see coverage from Angelo 2026). If there are accompanying advances in robotics, most if not all jobs could be threatened.
The economic impact of such large-scale disruption, should it occur, is hard to predict. Some have worried that widespread unemployment—or worse, unemployability, as all high-level tasks will be better served by AI—will create unprecedented political challenges, including concentration of political power in the hands of the small number of actors who manage to maintain control of the most powerful and performant AI models (Drago and Laine 2025). A related concern is that the structures in place to facilitate human influence on social, political, and economic institutions will slowly erode, possibly in a way that largely evades notice or requires substantial coordinated effort to avert (Kasirzadeh 2025; Kulveit et al. 2025).
Throughout history, governments and corporations alike experienced at least some pressure from citizens to guarantee and maintain a minimal standard of living for the population. Fear of rebellion, revolt, or simply workers’ strikes often kept political power in check. When governments and other organizations no longer need workers—and possess the technological means to eliminate the possibility of credible rebellion or revolt—such power could be unchallenged (see also § 7.4).
2.1.4 Destruction of a healthy information ecosystem
Related to the first two areas of concern, many have been alarmed at the capacity for AI models to synthesize convincing images, videos, stories, and dialogues that appear indistinguishable from reality (Tufekci 2026). This level of potential deception is unprecedented in human history. A functioning society requires a healthy flow of quality information and a robust media landscape that delivers news and other content to citizens across the world. For these to be effective, at least some sources must be trustworthy. A proliferation of AI-generated material, with increasingly convincing “stamps” of putative legitimacy, could make it increasingly difficult for people to be confident about the provenance or accuracy of what they see or read online (see, e.g. Chesney and Citron 2019; Barrington et al. 2025).
Although researchers are working to mitigate these issues—including development of new tools for certification and authentication (e.g., Coalition for Content Provenance and Authenticity (2026))—such efforts have been greatly outpaced by the rapid augmentation and dissemination of AI-generated content, and increasing capabilities could exacerbate these problems.
2.1.5 Geostrategic instability
One concern motivating many proponents of a pause is the potential for AI to set off unchecked escalation of geostrategic competition or use of AI in military or nuclear command and control. A pause treaty is one way to prevent a “race to the bottom” on AI safety (Askell et al. 2019; Scharre 2021) caused by competitive pressures that push companies and governments to cut corners on safety and testing, to increase their degree of reliance on and deference to AI, and to hand over ever more decision-making power to AI, especially in high-stakes competitive domains like the military.
Competition between the U.S. and China in the development of AI could increase geostrategic risks and instability by fostering risks of miscalculation, as in one recent incident where a false AI-generated intelligence report nearly provoked a hostile encounter between the U.S. and Chinese militaries (Lillis and Cohen 2026) or by tempting states to sabotage rivals out of fear that those rivals are on the precipice of attaining a decisive strategic advantage (Hendrycks et al. 2025). Bilateral treaty proposals for the U.S. and China to limit frontier AI development and deployments (Siddik 2025) emphasize how such agreements could reduce miscalculation risks and potentially improve bilateral geostrategic stability.
We discuss geostrategic stability in greater detail in § 8.
2.1.6 Why a pause
A pause in frontier AI training, observers have argued, would help address each of the concerns above. Halting the race to train more capable AI models, if only temporarily, would buy time for alignment and control research to improve, which would lessen the risks from loss of control and malicious actors. It would also provide breathing room for governments and other vital institutions to deliberate about and implement policies for addressing the adverse socioeconomic effects and concentrations of power that could attend the rapid arrival of more capable AI systems. A pause would, furthermore, allow space for the development of authentication and provenance tools that would make the information ecosystem more robust against misuses of AI-generated media. Pausing frontier AI training would also reduce the influence of AI acceleration and accompanying inequalities as contributing factors to ongoing geostrategic instability.
More broadly, whatever risks are posed by increasing AI capabilities, a pause would provide governments and societies a greater window to address potential problems before they become unmanageable.
2.2 Historical lessons for AI governance
In thinking about the prospects for an AI pause, it is worthwhile to reflect on the history of technology governance. While there is no perfect analogy for considering the broad suite of technologies that we package into “AI,” previous examples of technology governance (and arms control) offer a useful starting point from which to examine the conditions under which it has succeeded in the past.
The geopolitical feasibility of an international agreement to pause frontier AI training will depend not only on the specific features of AI, but also on the enduring political dynamics that have shaped previous efforts to govern transformative technologies. Every generation has faced new era-defining technologies, from crossbows in the 11th and 12th centuries (cf. Veen 2012) to chemical weapons in World War I. Rather than surveying every historical analogy on technology governance or arms control,8 we extract five common threads across a wide variety of domains that are indicative of the relevant lessons learned from centuries of efforts to govern and control technology.
2.2.1 Lesson 1: Governance is a three-legged stool
The first lesson is that any international technology governance or arms control agreement actually represents a collection of three challenges. The first is creating the conditions for states parties to come to the negotiation table in the first place. The salience of an issue for a state to engage with, the perceptions regarding the prospects of successfully achieving an agreement, and underlying trust all condition an ability to get parties to the table. The second challenge is negotiating the agreement around the table—addressing both the challenges of negotiation across that table but also within one’s own negotiating team and government at home. And, finally, a successful agreement needs to ultimately influence behavior under the terms outlined in the agreement.
Unfortunately, perhaps, the factors that condition solving these challenges are often at cross-purposes with one another. For example, the sunset of various clauses in an agreement (or treaty termination itself) that bound compliance and make negotiation easier, which are often held up as examples of a governance agreement’s weakness (see arguments against the Joint Comprehensive Plan of Action negotiated between Iran and the P5+1 together with the EU), are often necessary for an agreement to be reached at the negotiation table between states parties in the first place.
Any effort to create an AI governance agreement will have to grapple with these three concerns and recognize the degree to which they can be at cross-purposes. Compromises such as national security carve-outs can improve the palatability of a proposed deal at the negotiating table, but such concessions require care and forethought to ensure they do not undermine reliability or durability by vitiating verification.
2.2.2 Lesson 2: Governance agreements tend to focus on observable, verifiable units
Second, technology governance agreements must clearly specify the “what” that the agreement addresses. Whether a technology itself (e.g., a bomber) or a behavior, there needs to be a clear, defined, and measurable entity that is the subject of the agreement.
This lesson is reflected in the prominent role of hardware in contemporary AI governance debates, as opposed to the data upon which models rely, the algorithms with which they are trained, or the capabilities they exhibit. AI chips, and to an even greater extent the advanced lithography machines that print them, are physical artifacts that rely upon a complex supply chain that can be tracked during their production, deployment, and even use.
Per former U.S. President Ronald Reagan’s oft-repeated quote, having identified the subject of regulation, states parties “trust, but verify” and enjoy more success when both verification and enforcement measures engage with the observable artifacts of a particular technology used in a particular context. While signatories to the Biological Weapons Convention (BWC) have struggled to implement verification mechanisms, with the U.S. rejecting a proposed verification protocol in 2001 on the grounds that it would not deter violators (Whitehair and Brugger 2001), the Chemical Weapons Convention (CWC) and IAEA safeguards have supported inspectorates for decades (Tucker 2001). The Implementation Support Unit for the BWC, and its limited budget and staffing (Essix et al. 2025), also highlights the importance of effectively designing institutional governance for a hardwired pause (see § 3.6 for further discussion).
One other critical difference between the biological example and the others is scale: a militarily significant quantity of a deadly pathogen can be grown in a small building, but producing fissile material or scheduled chemicals in significant quantities requires substantial industrial facilities. On this test, the situation looks favorable for hardware governance: a leading-edge lithography machine weighs 180 tons and relies on highly specialized materials and expertise for its construction and operation; likewise, at least for now, training a frontier model requires hundreds of thousands of AI chips.
2.2.3 Lesson 3: Monitoring needn’t be perfect to be effective
Verification relies fundamentally on political judgments informed by data and analysis from monitoring and other sources, and these work by improving confidence in compliance, not by proving it perfectly (Schelling and Halperin 1961; Mayhew et al. 2026). This means that the beliefs, values, and incentives of treaty signatories are important. Historic examples like the Nuclear Test Ban Treaties and Outer Space Treaty also underline the importance and power of building norms and measures for increasing mutual predictability.
In the case of AI, verification can be highly effective even without completely foreclosing the possibility of any cheating, so long as it reliably limits the total quantity of training capacity a state could realistically marshal toward covert evasion. This report therefore operationalizes reliability of verification as a quantitative rather than a dichotomous question: we do not ask whether verification measures are airtight, but whether they prevent undetected evasion at large scales.
2.2.4 Lesson 4: Signatories need to benefit
Despite many analysts (and policymakers) viewing technology governance and arms control through a moral lens (e.g., states ban the use of chemical weapons because of moral qualms), governance regimes are often determined primarily via the strategic benefits that they endow upon states parties. Put simply, countries tend to agree on international agreement formation when it is in their interest to do so. For example, while it is difficult to imagine states once again signing up to the Nuclear Nonproliferation Treaty (even for an initial period of twenty-five years), which enshrines “haves” and “have-nots,” the arrangement is referred to as a “grand bargain,” because it enables nuclear weapon states to enjoy their monopoly on nuclear weapons, while pledging to work toward disarmament, and non-nuclear weapon states to share in the civilian applications of nuclear technology (Weiss 2003).
This logic is reflected in bilateral arms control. Both the United States and the Soviet Union sought to manage the nuclear competition between their capitals in a manner that they perceived to be stability-maximizing and to put a lid on spending. With regard to the former, an often underappreciated aspect of strategic arms control during the Cold War is the provisions that committed both states to respect one another’s national technical means, referring to the intelligence-gathering technologies governments use to monitor one another’s treaty compliance and overall strategic posture, without relying on on-site inspections.
With regard to AI technologies, this reality is complicated yet furthered by the reality that cutting-edge developments with relation to the technology are taking place primarily in the private sector rather than being driven by government agencies (as with the Manhattan Project). Thus, another set of actors (from large U.S., Chinese, and European AI and semiconductor companies to smaller private owners of computing hardware and equity in these private companies) have stakes in the distributional consequences of a governance agreement. They represent a voice at (or behind) the table, and the success of a governance agreement will, all else being equal, be furthered by aligning their economic interests with the mandate, scope, and obligations of that agreement.
Any effort to regulate AI will have to address the distributional consequences of the agreement—and account for the reality that governments do not sign up to arrangements from which they do not yield a benefit. The strategic and national security interests of states parties, as perceived by their leaders, must be served (or at least not diminished) by the agreement. In particular, the agreement’s verification mechanisms must deliver assurance to states that they will not fall prey to covert or overt violations by rivals, as we discuss next.
2.2.5. Lesson 5: Effective Institutions Are Needed
A fifth lesson from history is that effective governance and implementation of an international agreement usually requires a sufficiently resourced institution with a clear mandate and delegated authorities (this lesson is further discussed in § 3.6).
Some negative cases reinforce this finding. The World Trade Organization (WTO), an international institution established to facilitate international trade and manage related disagreements, has lost some of its functionality due to a vulnerability in its institutional design. Since late 2019, the WTO’s Appellate Body has not been functional due to the United States blocking staffing appointments, and over 30 international trade disputes have been delayed or held in limbo as of early 2025 (Lester 2022). The WTO was created with a consensus mechanism for appointments to the Appellate Body. This case has shown how international institutions can be disabled through the vulnerability of that mechanism.
The 1925 Geneva Protocol, formally the Protocol for the Prohibition of the Use in War of Asphyxiating, Poisonous or Other Gases, and of Bacteriological Methods of Warfare, lacked any implementing institution. Despite the battlefield horrors of chemical weapons in World War I which motivated the Protocol’s formulation, some states parties would later use chemical weapons (Italy, during the Second Italo-Ethiopian War of 1935-36, and Iraq, against its own Kurdish population, and in the war with Iran in the 1980s). There was no Geneva Protocol institution to report on, investigate, or convene parties to address these events. Rather, ad hoc investigative teams were assembled by the UN Secretary-General.
While the aforementioned BWC was established in part to rectify disarmament gaps of the Geneva Protocol, it too had minimal institutional support in the early decades after its signing in 1972 and notably no verification provisions or standing governing body other than a periodic review conference. When the Russian Federation revealed in 1992 that the Soviet Union had an extensive bioweapons program for decades in violation of the BWC, it revealed how the minimal organizational support to the BWC had no ability to detect major violations. While the subsequent stand-up of the Implementation Support Unit created some standing staff by 2007, its small size, its design for support rather than verification, and its insufficient financing all highlight institutional governance shortfalls that can inform better choices for AI governance efforts like the hardwired pause proposal discussed here.
There are a range of positive examples of this fifth lesson. We highlight a few here regarding the Organization for the Prohibition of Chemical Weapons (OPCW, for the CWC), International Civil Aviation Organization (ICAO, for the Convention on International Civil Aviation, or “Chicago Convention”), and the set of institutions that support the Montreal Protocol on Substances That Deplete the Ozone Layer (more commonly referred to as the Montreal Protocol).
The OPCW as an institution demonstrates that monitoring and verification of dual-use chemicals and industrial facilities is technically and politically possible, especially in terms of implementing inspections at industry facilities. Over nearly 30 years of operation, the OPCW also highlights how a substantial budget and technical secretariat are important for helping such an institution effectively implement its mandate. The OPCW’s structures for facilitating regular meetings of states parties’ representatives for political discussions and determinations (inclusive of an active Executive Council), as well as its four main subsidiary bodies, highlight the importance of having multiple venues for considering and implementing governance issues relating to a complex international agreement.
In terms of lessons for adaptivity related to technological change and managing audits that can indirectly influence industry standards, the ICAO presents helpful lessons for AI governance. First, the ICAO manages and coordinates a regular technical review of potential changes to safety standards through its Air Navigation Commission, and their recommendations are reviewed and approved by a small council of some states parties. The ICAO also conducts audits of each nation’s capacity to implement civil aviation standards, and these audits help direct effective capacity building efforts where there are gaps in national capacity. Additionally, with the publishing of those results, other states and markets can use those audits to restrict access or incentivize improved safety standard implementation by states parties and the aviation industry.
Lastly, the set of institutions that support the Montreal Protocol show how a smaller set of implementing treaty authorities and offices can help achieve a treaty’s mandate and encourage compliance. The Montreal Protocol’s Ozone Secretariat and Implementation Committee share the primary responsibilities for monitoring and encouraging compliance with the treaty. The Multilateral Fund for the Implementation of the Montreal Protocol, supported by four international agencies, also provides positive examples for how to structure capacity building and technical assistance funds, through a standing Secretariat. For AI governance proposals that include some type of international fund to offset economic costs of a pause, or funding to help distribute education and capacity building benefits to states parties, the Multilateral Fund Secretariat could be a useful source of benchmarking.
These cases underscore the value of an implementing institution for a hardwired pause, with sufficient capacity and provisions for important functions like verification, and funding to implement its mandate and scope.
2.3 Game-theoretic lessons for AI governance
Discussions of a possible AI pause agreement invariably lead to discussions of game theory. The strategic situation concerning an AI pause, especially between the U.S. and China, is often described as a Prisoner’s Dilemma in which a mutual pause is not an equilibrium: while both players pausing is better for everyone than both players racing, both players pausing is not an equilibrium because each player would be better off unilaterally racing ahead if the other paused.
This analysis is incomplete for two reasons. First, as noted by Carlsmith (2026), it neglects the possibility that the strategic situation is better modeled as a Stag Hunt: in such a game, as in the Prisoner’s Dilemma, each player gets the worst possible payoff if they pause while the other player races; but unlike in the Prisoner’s Dilemma, each player gets the best possible payoff if they both pause. In the Stag Hunt, there are two (pure-strategy) equilibria: both players racing, since a player who unilaterally paused would be worse off than if they had continued racing; and both players pausing, since a player who unilaterally raced would be worse off than if they had maintained the pause. The question then is how the players can coordinate with each other so as to achieve the universally best (in economic terms, Pareto-dominant) outcome, in which they both pause. Statements by some technologists indicating that they would personally prefer not to expose the world to danger by building AI, but they have to build it because others will build it anyway, suggest they view the AI race as a Stag Hunt in which they are failing to achieve the best equilibrium (Landemore and Tang 2026).
The second flaw is that the U.S. and China are not playing a one-shot game. Their governments do not meet once, choose actions simultaneously, and never deal with one another again. Instead, Washington and Beijing have significant, even if limited, observability into each other’s actions, and they must repeatedly choose whether to pause or race knowing that failure to honor their commitments could have long-term consequences. Thus, a better model for the strategic situation is that of a repeated game with an unknown end date. Each player receives a stream of payoffs from each “stage game,” discounted by some factor that captures their time preference (how much more they care about the present than the future) and the probability that their interaction will continue (see Osborne and Rubinstein 1994, chap. 8). After each stage of the game, each player receives a noisy signal of the other player’s action.
In such repeated games, mutual cooperation can be an equilibrium even if the stage game is a Prisoner’s Dilemma, provided the players care enough about their future interactions and face credible consequences sufficient to deter defection (Osborne and Rubinstein 1994, sec. 8.1).
A further key distinction concerns whether the noisy signals observed by each player are private or public. If they are private, meaning for example that player A might not know what player knows about ’s action, then the cooperative equilibria can be very complex (Sugaya 2022). But if the noisy signals are commonly observed, then relatively simple cooperative equilibria can be supported (Green and Porter 1984; Abreu et al. 1990; Fudenberg et al. 1994), provided the signals are informative enough. This point about informative signals highlights the importance of monitoring and verification for supporting the maintenance of cooperation.
The foregoing points about Stag Hunts and repeated games have analogues when we conceive of the strategic situation as an -player game with countries, rather than a two-player game between the U.S. and China alone. We should not assume that the countries face a one-shot Prisoner’s Dilemma. They face a repeated game in which cooperative equilibria are possible, even if the stage game is a Prisoner’s Dilemma, given sufficient concern for the future and sufficiently informative public signals to support credible consequences for defection (Fudenberg et al. 1994).
Thus, one of the game-theoretic lessons for AI governance is that to facilitate cooperation, governance should leverage the temporally-extended nature of the interaction between the players and try to rely on publicly observed signals of each player’s actions to coordinate on a cooperative equilibrium. Hardware-level verification, compute monitoring and accounting, and inspections can provide public information when their results are shared with states parties, allowing players to observe the same signal. The upshot is that game theory does not foreclose the possibility of international cooperation to pause frontier AI training. Instead, it suggests conditions under which cooperation can be sustained, including sufficiently informative and timely monitoring, credible consequences for defections, and states placing sufficient value on the prospects for future cooperation. This highlights the importance of working out the technical details of how states would obtain signals that are informative with respect to other states’ compliance, early enough to permit an effective response before a violation leads to a destabilizing strategic advantage. Those details are the topic of the next section.
3 Example implementation of a hardwired pause
3.1 Overview and objectives
A hardwired pause is feasible if and only if there is at least one implementation that is simultaneously politically and technically feasible. This section discusses one possible near-term implementation of a hardwired pause, describing in more detail key governance concepts, objectives the agreement would need to achieve, threat models it would need to guard against, and how the transition might be managed. Our goal is not to prescribe specific terms or give a fully realized blueprint for an agreement, establish an optimal design, or to argue against the feasibility of other implementations, but rather to set forth a framework to help advance further research. We describe one broad technical and institutional governance framework that may be implementable in the immediate future, in sufficiently definite terms to allow for feasibility assessment. The feasibility of inference-only chips, hardware governance, and verification, as operationalized here, are discussed in § 4, § 5, and § 6, respectively.
We adopt a working assumption that the hardwired pause would eventually be codified in an international treaty or similar agreement, but other institutional arrangements or governance mechanisms may achieve many of the same objectives. In particular, because a treaty would take time to negotiate and ratify, we expect that other, more informal agreements could be a basis for the initial phases of a hardwired pause. As we discuss in § 3.4.2, the core negative mandates of a hardwired pause, namely bans on training new frontier AI models and on manufacturing new training-capable AI chips, could be implemented immediately by the U.S. and China alone,9 with a handful of allies and partners joining shortly thereafter.
We assume that participating jurisdictions would initially include (at least) the United States, China, the European Union (especially Germany and the Netherlands), the United Kingdom, Taiwan,10 South Korea, and Japan—jurisdictions that currently enjoy nearly a collective global monopoly on both cutting-edge AI chip manufacturing and frontier model training. Other especially high-value signatories include nations like Singapore, Malaysia, Indonesia, Thailand, Vietnam, the United Arab Emirates and Saudi Arabia, where a large number of existing chips are located or compute infrastructure and supply chains exist. It is likely that most or all of the countries whose heads of state signed the rolling “Call for Control of Frontier AI Models” of September 2026 (President of the Republic of Finland 2026) would also participate. These initial parties, and other polities, might be induced to join the treaty through a range of positive mandates and distributed benefits that could complement the hardwired pause. For example, the agreement could include technological cooperation, education, and capacity building. Distribution of specific training-capable chips to be used for common-good scientific research as discussed in § 3.3.1 might induce nations to join as well.
We assume that the aim of states parties, for the duration of the treaty, is to permit fast and private inference on whitelisted AI models while forbidding, with narrow, regulated exceptions, all other use of AI accelerator hardware, defined as systems with computational throughput above a defined threshold. This whitelist would consist of certain models already released at the time the treaty takes effect and deemed sufficiently safe to license for widespread use. Irrespective of what hardware is used, the treaty would prohibit training of new beyond-threshold models (see § 3.2.2) that are trained at sufficient scale to plausibly compete with models at the frozen frontier.
We assess that the treaty will require graduated steps to establish necessary institutional governance in each of the four phases of the hardwired pause. Phase 0 will likely require some interim authorities staff, perhaps a preparatory technical secretariat, established through agreement by initial parties. Their work would involve management and monitoring of provisional whitelist processes and issues, and preparing for the significant global accelerator census data collection effort, as well as anticipated monitoring, treaty licensing procedures, and verification mechanism implementation (including inspections) of Phase 1. On signature of the treaty, Phase 1 begins with significant implementation procedure demands, including global verification efforts. One way of handling this is by establishing a provisional treaty organization (another possibility could be building the necessary capacity in an existing organization). This organization might include a technical secretariat and other required staff to administer and manage the organization (we refer to these personnel as the treaty authorities). The provisional treaty organization would also have some form of conference of states parties, or other convening bodies for discussing political procedures of the treaty and making any needed states parties determinations or decisions (we refer to these types of decisions or actions to be those of the states parties in the paper). Some responsibilities of the hardwired pause may require actions by both the treaty authorities and states parties, or it may be unclear - where we assess this to be the case we use the term governing authorities (otherwise we try to clarify one or the other as having a lead role). Phases 2 and 3 would benefit from a formalized treaty organization, with potentially multiple institutional elements designed for all states parties to conduct deliberations, as well as an executive council of some rotating and representative basis. Expanded staffing, technical infrastructure, and larger operational elements for the treaty authorities would likely be needed in these phases to ensure feasibility.
In the remainder of § 3, we describe a hardwired pause treaty that operationalizes technical aspects of verification primarily through quantitative accounting of hardware capacity, with a focus on enforcement through denial of sufficient model training capacity to would-be defectors. We begin by identifying the core political and technical objectives that such a treaty would need to achieve.
3.1.1 Political objectives
In order to reliably pause the AI frontier, a treaty would need to guard against several distinct threat models:
Private evasion: Private actors building illicit above-threshold models for commercial or other purposes.
State-level covert evasion: A state covertly building and deploying a model powerful enough to attain a significant strategic advantage.
State-level overt breakout: A state amassing a large hardware fleet, overtly abrogating the treaty, and attempting to rapidly develop and deploy a model powerful enough to attain a decisive strategic advantage before other states can respond.
By a significant strategic advantage, we mean an advantage that materially affects the geopolitical balance of power, to a degree sufficient to undermine the treaty. By a decisive strategic advantage, we mean an advantage sufficient to establish and consolidate a full-spectrum military and economic unipolarity of the kind described in Drezner 2013 (cf. Wohlforth 2009).
It is contested whether any advantage in AI could deliver such unipolarity, but what matters from the perspective of treaty durability is not the objective fact of the matter but the expectations of states parties. According to their public statements, some world leaders view competition in AI as strategically decisive.11
The belief that more capable AI models could be expected to confer strategic benefits on those who build and deploy them rests on an implicit assumption that such models would be controllable. In some pause-willing futures, that assumption may not be taken for granted. In that case, the verification problem could be easier to solve, for the same reason that the U.S. and China do not currently demand onerous verification measures of each other to prove that neither state is developing mirror bacteria (cf. Adamala et al. 2024; Cuéllar et al. 2025). However, given the emphasis contemporary political discourse places on AI competition between the U.S. and China, it appears likely that verification measures would play an important role in facilitating a reliable and palatable near-term pause.
Additional important political objectives to satisfy palatability would include the following:
Prioritize privacy: Avoid compromising the privacy of inference users, and minimize intrusive surveillance measures.
Minimize concentration of power: Prevent AI trajectories that will severely concentrate power, and avoid unduly concentrating power in the process of doing so.
Minimize economic disruption: Avoid negative shocks to the economy, including in particular to continued commercial inference service.
Permit beneficial research: Allow continued research into medical and other common-good scientific applications using accelerator hardware.
Permit strategic insurance: Give states legitimate ways to guard against defection or exploitation by other states, for example by holding some training-capable chips in reserve as insurance against the possibility of treaty breakdown.
Side benefits such as mitigation of environmental impact per unit of inference service (especially through reduced power needs of certain inference-only chips; see § 7.3) would improve palatability but are not required.
Finally, achieving broad membership, especially for the supplier and chip-owning states discussed above, would be important for maintaining a reliable pause agreement whose governance measures are not circumvented via non-members.
We note that the implementation below is not designed to serve as a complete solution to the many complex problems of global AI governance such as child safety, labor market disruption, disempowerment from overreliance on AI systems, automated cybercrime, or surveillance. These and other issues concerning the impact of AI on society would likely remain pressing even if a hardwired pause were successfully achieved and maintained at the present frontier level (and likely more so if it were maintained at a future, more powerful frontier level) but addressing them is outside the scope of the present work (also see § 7.6).
3.1.2 Technical objectives
In the service of the above political objectives, we identify several key technical pillars of success for a hardwired pause treaty:
Register pre-pause hardware: Ensure a high proportion of functioning pre-pause accelerator hardware is registered under any interim agreements and the treaty, and accurately estimate how much remains unregistered, both globally and within each state.
Track registered hardware: Track the physical location of registered accelerator hardware, and prevent its leakage from the monitoring of treaty authorities.
Govern hardware manufacturing: Prevent ungoverned manufacturing of accelerators at either declared or undeclared manufacturing facilities.
Govern licensed use of training-capable hardware: Prevent licensed hardware operators and inference providers from diverting training-capable hardware for transitional inference service toward non-licensed use, especially illicit training, and prevent diversion of hardware licensed for exceptional uses as in § 3.3.1.
Develop sound accounting standards: Estimate failure rates for the above objectives within each state with reliable uncertainty quantification and adjust the issuance of licenses as needed to keep states’ training capacity below risk thresholds.
Develop and certify inference-only accelerators: Develop and produce inference-only accelerator designs within an acceptable time period, to allow for the managed drawdown of pre-pause training-capable accelerators.
Any functioning accelerator hardware outside the monitorability of authorities of the interim agreements or treaty is termed dark compute;12 two of the treaty’s most important tasks would be to minimize and quantify the amount of dark compute plausibly existing in each state.
For each objective, success would be measured not against standards of perfection but by ensuring that failure rates were accounted for and that total covert evasion and overt breakout capacity for each state were held below realistic quantitative bounds. These bounds would be adjusted over time to account for training efficiency gains; see § 6.3.
Attaining the above objectives would suffice to bound all-source computing capacity in the service of model training, which we take as sufficient to prevent new and improved frontier AI models from being built, until changing technological conditions require a new governance paradigm. Thus, we take this as our technical criterion for reliably pausing frontier training.
A pause in the release of better frontier models would not mean a complete halt to the advancing capabilities of AI systems, including harness software and other scaffolding tools for better eliciting latent capabilities of existing models. This means developers would still be able to use whitelisted AI models as a building block to build more powerful and useful AI tools, realizing many of the promised benefits of the technology. But if relevant (e.g. dangerous) capabilities continued to advance too rapidly even without new model training, additional governance tools would be required to address that challenge.
Finally, we do not expect that a hardwired pause could by itself prevent all use of non-whitelisted models or small-scale fine-tuning or other adaptation of potentially dangerous sub-frontier models. However, relative to alternative scenarios where training-capable accelerator hardware continued to proliferate, a successfully enforced hardwired pause would make these smaller-scale violations easier to manage at sub-frontier scale and more difficult for private evaders at frontier scale.
3.1.3 Technological readiness
The implementation below is designed to achieve the above technical objectives without significant reliance on foundational new research in computer science, such as in cryptographic techniques for privacy-preserving compute verification (see, e.g. Shavit 2023; South et al. 2024; and Petrie et al. 2025) or in the science of alignment, monitoring, and control of powerful AI systems (see, e.g., § 3.3 of Bengio et al. 2026). Developments in either field could complement hardware-based approaches as governance procedures matured, but might not be immediately available in near-term pause-willing futures.
The most demanding technological requirement for the treaty, the design of inference-only accelerators, has proofs of concept in commercial production already (see § 4.3), and the implementation discussed here does not assume their immediate availability (see § 4.4).
3.2 Operationalizing technical requirements
In this section, we describe several key technical aspects of a hardwired pause. § 3.2.1 discusses the distinction between inference-only and training-capable accelerators. § 3.2.2 introduces treaty-relevant thresholds for computing budgets that define what the treaty would allow and what would pose a risk of covert evasion or overt breakout, and § 3.2.3 discusses a verification framework for limiting the training capacity that any state could marshal for either covert or overt defection. § 3.2.4 discusses the model whitelist.
3.2.1 Inference-only and training-capable accelerators
This section discusses how a treaty might operationalize the difference between a training-capable accelerator and an inference-only accelerator. After distinguishing between inference on approved models and training, we operationalize inference-only hardware by bounding its capacity for unintended use in AI training, relative to its inference capacity.
Inference is defined as the mathematical operation of drawing outputs from a model according to its conditional distribution given an input prompt. Training operations are defined as the two operations that account for nearly the entire computational budget of model pretraining and post-training,13 namely:
Performing inference for generic, iteratively updated weight values, and
Evaluating gradients of a loss or reward function with respect to model weights, also at generic and iteratively updated values of the weights.
For each of these two training operations, nearly the entire computing budget is in turn spent on floating-point operations (FLOPs) used in dynamic matrix multiplication, meaning matrix multiplication with both operands determined at runtime (see Austin et al. 2025, for a breakdown of the floating-point operations required for transformers’ forward and backward passes). We will use the term training FLOPs to mean dynamic matrix multiplication FLOPs in the service of either of the above two training operations. The arithmetic throughput of a hardware system is defined by the number of training FLOPs per second at peak rating.
A commonly used and intuitive unit of hardware arithmetic throughput is the H100-equivalent (H100e), taken as 989 TFLOP/s (, or per year),14 the peak rated arithmetic throughput of an NVIDIA H100 accelerator.15 A hardware system’s realized arithmetic throughput on a given task is generally below its peak rate, e.g., due to memory and communication bottlenecks; the utilization is defined as the realized arithmetic throughput divided by the peak throughput.
We analogously define a hardware system’s inference capacity with respect to a given model in terms of inference-H100e, which we define as its peak rated inference throughput, in tokens per second at a standardized context length, divided by the peak rated token throughput of a standardized general-purpose accelerator cluster, times the H100e of that cluster. The inference-H100e of an inference-only system is defined as its maximum inference-H100e over the models it could serve.
A hardwired pause treaty could then define a training-capable accelerator as any hardware system capable of performing training FLOPs at an arithmetic throughput rate above a threshold determined by treaty parties and revised over time as needed. An inference-only accelerator, by comparison, could be defined as a hardware system that can perform inference on at least one whitelisted model above some treaty-specified rate but whose rating for training operations is bounded in a suitable way. We discuss ways of building inference-only accelerators in § 4.2 and § 4.3.
We define an idealized inference-only machine as one that can perform inference for a restricted set of models but is incapable of performing any other operation.16 In practice, realized inference-only accelerators may have some residual training capacity. Like other classes of accelerator hardware, they would be rated to arrive at an inspector-certified bound on their peak arithmetic throughput for training operations. These bounds would need to pay careful attention to the possibility of unintended use, including by a state-level adversary with physical ownership of chips; see § 4.1.
Because the governance purpose of inference-only accelerators is to satisfy inference demand without allowing for model training, we suggest operationalizing an inference-only rating by bounding a hardware system’s residual training factor, defined as (a bound on) its peak arithmetic throughput for training operations in H100e divided by its peak licensed inference rating in inference-H100e. For example, replacing a pre-pause training-capable accelerator fleet with an equally capable inference-only fleet whose average residual training factor is (weighted by inference-H100e) would reduce that fleet’s training capacity by 99.9%.
The maximum residual training factor could start relatively high in order to allow early designs to speed the transitional period, and it could decrease over time as hardware specialization techniques mature and lower factors become available. As a simple approach, we suggest states parties could simply agree on a training-capacity rating, cap the total number of H100e units licensed per inference provider and host state, and let market pressure drive residual training factors down as inference demand rises.
3.2.2 Quantitative training thresholds
In order to operationalize treaty objectives, states parties would need to set quantitative thresholds reflecting estimates of the compute budget, measured in training FLOPs, required to train models with treaty-relevant capability levels, as well as estimates of how large a hardware fleet would be required to research and train such models within a relevant time horizon. These would be negotiated by parties in accordance with prevailing technological and political conditions, and adjusted as needed. Here we outline how such thresholds and estimations might be defined, specify these quantities based on an analysis of present-day models, and introduce relevant terminology and notation.
We define the pause date to be the moment when frontier training is paused. Treaty thresholds could be defined relative to the frontier threshold , an estimate of the total compute required to train the most capable models at the frozen frontier. Treaty compute thresholds would need to be adjusted over time in order to adapt to training efficiency gains,17 improvements in algorithms and data that make it possible to train models at a fixed capability level with shrinking training budgets. For example, we expect the frontier threshold would not be a constant but a function in time , where represents the time in years since the pause date; e.g., represents one year after the pause, and is the training budget required, one year after the pause, to train a new model with capabilities similar to those of the frozen frontier models.
Besides preventing nation states from training strategically relevant or strategically decisive models, a goal of the pause would be to discourage all actors from attempting to advance the frontier, since such attempts could accelerate training efficiency gains for frontier-scale models.18 We define a beyond-threshold model as any general-purpose generative AI model19 plausibly powerful enough to compete with models at the frozen frontier. Training and use of such models would be prohibited.20 We operationalize this definition as including exactly those models surpassing either a compute ceiling bounding the total number of FLOPs used to train it, including both pretraining and post-training, or a total parameter ceiling bounding its parameter count.21 These two thresholds play different enforcement roles: while a parameter ceiling is more enforceable after the fact, training FLOPs are easier to deny to evaders before the fact. Since training budgets are imperfect proxies for model capability, preventing the training of new frontier models would require leaving a wide margin between the compute ceiling and the frontier threshold .
| Variable | Meaning | Est. (year-end 2026) |
| Training budget for frozen frontier | ||
| Treaty compute ceiling | ||
| Treaty parameter ceiling | 13B parameters | |
| Training budget for strategically relevant covert evasion | ||
| Training budget for strategically decisive overt breakout | ||
| Fleet threshold for covert evasion at present-day utilization | 6.4M H100e | |
| Fleet threshold for overt breakout at present-day utilization | 180M H100e |
We operationalize the two state-level threat models described in § 3.1.1 by defining a covert evasion training threshold and overt breakout training threshold , estimates of the total training budget required to train a strategically relevant model or a strategically decisive model, respectively. Like and , these are measured as FLOPs required for both pretraining and post-training in the final training run for the model and would be adjusted over time to account for training efficiency gains.
Frontier AI companies’ research programs to build new frontier models require significant ongoing research and development processes beyond the final training runs; per recent estimates, current frontier companies’ research and development budgets are ten times as large as their model training budgets (Denain and Wu 2026). We account for these additional costs via a research overhead factor , whose magnitude may depend on how ambitiously a research program is attempting to advance the frontier. After accounting for research overhead, we arrive at an evasion fleet threshold and breakout fleet threshold , estimates of the fleet sizes that would be required to build each model, if these programs attained utilization similar to that attained by present-day industrial frontier training. These fleet thresholds are measured in H100e units, and both are defined relative to strategic time horizons and within which an evader is attempting to train a model; a larger fleet is required to break out within year than to break out within years. Then the evasion fleet threshold is defined as:
where represents a standardized utilization attained by present-day frontier AI companies. The breakout fleet threshold is defined analogously.
Choices regarding the levels of the treaty thresholds and would need to balance the value of permitting sufficiently sub-frontier model development for commercial or recreational use against the likelihood that prolific training of near-frontier models would drive training efficiency gains, eroding the reliability and durability of the treaty. Models trained for beneficial scientific purposes such as medical research could potentially be granted exemptions; see the discussion of scientific preserves in § 3.3.1.
3.2.3 Verification and breakout accounting
The training budget for an illicit model could come from a number of different sources including training-capable hardware licensed for transitional inference service, dark compute, consumer graphics processing units (GPUs), etc. To guard against covert evasion and overt breakout threat models, the treaty would require a verification methodology to ensure that no state’s total hardware fleet, aggregated across all sources, contains enough training capacity to pose either a covert evasion threat or an overt breakout threat.
We explore one relatively simple verification methodology based on estimating, for each possible compute source in each state, the total arithmetic throughput that could plausibly be marshaled toward covert evasion or overt breakout. For compute source and state , we define the covert evasion capacity as the total arithmetic throughput from source that state could plausibly marshal toward a covert evasion program, and the overt breakout capacity as the throughput it could marshal toward an overt breakout program.
These capacities would be measured in H100e and discounted by estimates of how their utilization rate in each scenario would compare to frontier data-center utilization. The same compute source might contribute more to one violation scenario than the other: for example, an internationally monitored, sealed strategic reserve of 100k H100e units of powered-off hardware would contribute roughly 100k H100e to a state’s overt breakout capacity, but it would contribute essentially nothing to its covert evasion capacity. In § 6.4.1 we discuss how to estimate these capacities by discounting the raw capacity available in a given category.
Treaty authorities could certify that each state’s total covert evasion capacity, which we define as its covert evasion ledger , were held below the covert evasion fleet threshold times a fractional covert evasion allowance to account for uncertainty; that is:
Authorities could likewise calculate an overt breakout ledger by summing over and certifying that it is below the breakout fleet threshold by an overt breakout allowance, i.e., .22
The treaty authority’s primary lever for balancing these two inequalities for each state would be limiting the total number of licenses granted to each state for various uses of accelerators such as transitional inference service or other uses discussed in § 3.3.1. In this way, each term in a given ledger could be adjusted until the inequality is satisfied, with the important exception of the term for dark compute. Because dark compute would by definition be ungoverned under the treaty, its estimated covert evasion or overt breakout capacity would need to be reduced through other measures such as discovery and registration of pre-pause stock, prevention of illicit manufacture, and verification measures that rely on detecting when such dark compute is being used.
Because issuance of licenses would depend on satisfying these inequalities, states desiring more licenses would have an incentive to keep their ledgers clean, for example by limiting the rate at which dark compute accumulates in their state, or by agreeing to allow more inspector access to assure other states that a given compute source could contribute very little to covert evasion. Thus, verification procedures would not need to be one-size-fits-all: they could be flexibly negotiated based on each state’s willingness to make concessions in each context, with larger compute sources requiring more rigorous verification measures than smaller ones.
We note that building a strategically relevant or strategically decisive model is only the first hurdle an evader would need to clear. In order to realize any advantage from training, the evader must also be able to serve inference on that fleet at sufficient scale. Thus, governance that focuses on prevention of training is sufficient to prevent violation but may not be absolutely necessary. For example, if states developed reliable national technical means for detecting the use of unknown language models when they were deployed, it could ease the degree of reliance on verification based on denial of training capability.
3.2.4 Model whitelisting
Beyond enforcing a frontier pause, treaty governing authorities would also be responsible for determining the model whitelist: the list of beyond-threshold AI models for which inference is permitted and inference-only chips are permitted to be manufactured.23 By “model,” we mean a fixed probability distribution for generating output tokens conditional on input prompts. Since fine-tuning or other post-training that alters or adapts model weights would alter the output distribution, a hardwired pause would not permit inference on models altered in these ways except where expressly approved by governing authorities. Examples of approved fine-tunes could include ones that implement knowledge cutoff updates or safety updates.
Treaty governing authorities could coordinate on any international standards regulating the deployment of whitelisted models, in addition to whatever local regulations are in force. For example, a model for which a critical jailbreak is discovered could be de-whitelisted, possibly followed by re-whitelisting contingent on inference providers deploying it with required external constraints such as an input classifier or safety harness. For example, a model that had critical jailbreak vulnerabilities could possibly be judged as unsafe if exposed directly to users but sufficiently safe if it were only prompted by other models.24
During Phase 0 or Phase 1 of a hardwired pause (recall § 1.1), AI developers would submit above-threshold models to be considered for whitelisting by participating states or interim treaty authorities before a defined whitelist submission deadline. To foreclose the possibility of further training after the deadline, developers submitting closed-weight models could be required to submit a cryptographic hash of a complete description of the architecture and weights of any model they submit; this hash could then be checked against deployed weights to ensure that the deployed model had not changed since submission. Additional required information for evaluation purposes could include details of the development process including the training algorithm and data.
The whitelist submission deadline would be an important governance instrument facilitating enforcement of the prohibition against above-threshold model training. Under a hardwired pause with a long expected duration, AI developers would have little incentive to unlawfully train new above-threshold models after the deadline, because those models could not be commercially deployed without exposing the violation. By contrast, shorter-duration pauses would likely require stronger enforcement measures to prevent illicit training: by holding out the possibility that noncompliance would be rewarded after a brief pause, they could tempt companies to evade.
At present, there is no agreed-upon operating definition for whether a powerful AI model is safe enough to deploy internally or to release to the public, or methodology for assessing safety. Often in practice, a subjective determination is made by the same AI developer with a pecuniary interest in releasing it. Moreover, safety depends not only on the model but on the technological, social, and institutional systems in which the model is deployed, and safety may change over the lifetime of these systems. As such, whitelisting processes may evolve over time in response to input from improved understanding of AI safety, newly discovered adversarial methods for jailbreaking or eliciting dangerous capabilities, empirical evidence about AI models’ effects on society, or input from affected constituencies.25 Developing methods for safety determinations and politically legitimate processes for making them involves pressing technical and political questions outside the scope of this paper.
Whitelisting processes would need to be designed for consistency with states’ norms and constitutional restrictions. In particular, U.S. implementation could raise First Amendment questions because restrictions on model development may burden free expression interests, especially if they distinguish among outputs by their subject matter or viewpoint (Volokh et al. 2023; Mark and Scher 2025). Restrictions on hardware and training capacity have a stronger claim to content neutrality, but still need to be justified insofar as they incidentally burden protected expression.
It is possible that pause-willingness (as defined in § 1.2) could arise only after all frontier models have developed capabilities that are widely considered unsafe for deployment at scale or that come to be considered unsafe in the years after a pause. In that case, the effect of a well-considered whitelisting process could be to “rewind” the frontier (as opposed to pausing it) if the most capable whitelisted models had all been trained significantly before the pause date and were all less capable than the pause-date frontier. In that case, it would likely be effectively impossible to ensure that non-whitelisted models were not being used covertly, especially by states. The frontier threshold would effectively be reduced, and other treaty thresholds derived from it such as and (as in § 6.2) would need to be correspondingly reduced. Thus, while it is conceivable that taking such a step might be necessary, doing so could make the treaty harder to enforce and verify, especially because the covert evasion fleet threshold would likely be reduced.
3.3 Hardware governance
To verify that the technical specifications from § 3.2 are met, states parties would need to govern computer hardware, especially training-capable accelerators. Here we describe the roles training-capable accelerators would continue to play in a hardwired pause (§ 3.3.1); how accelerators could be located and tracked, and how licensing would operate to facilitate their governance (§ 3.3.2); and approaches to verifying that training accelerators are only used for their intended roles (§ 3.3.3). We conclude by discussing governance of hardware manufacturing technologies and facilities (§ 3.3.4).
3.3.1 Residual uses of training-capable accelerators
While inference-only accelerators would displace the majority of training-capable accelerators after a several-year-long transitional period, training-capable accelerators could remain in limited use during a hardwired pause. This use could fall into the following categories:
First, the residual use that would be largest in scale but shortest in duration would be the interim commercial allocation of training-capable hardware to transitional inference service by licensed inference providers while the market awaited sufficient inference-only hardware to replace it.
Allocation of training-capable hardware toward inference service would require time- and capacity-bounded reliance on non-hardware-based verification methods as discussed in § 3.3.3, but it would serve two important governance purposes: First, it would prevent inference supply from collapsing while manufacturing of inference-only hardware ramped up. Second, it would harness burgeoning inference demand to incentivize owners of pre-pause training-capable accelerators to either register their chips under the treaty or sell their chips to licensed hardware operators, or to treaty authorities, which could sell them back to hardware operators at auction. Meanwhile, the frozen model whitelist would eliminate the primary commercial incentive for training new frontier models, so compliance would be the most remunerative option for hardware owners.
Second, each state could be allowed a licensed sovereign fleet of training-capable accelerators, with registered and tracked chips, for use on specific sensitive national security workloads. Its size would be bounded to ensure the state’s total allocation of training-capable compute remained below the covert evasion and overt breakout thresholds.
Third, training-capable hardware could be eligible to be placed in internationally-governed scientific preserves: data centers operated by treaty authorities and dedicated to common-good scientific applications such as medical research or healthcare provision.
Finally, states could have the option to store licensed strategic reserves of powered-off training-capable hardware under seal in monitored facilities within their own territory (cf. Lifland 2026), for example, in sealed bags that are periodically checked. Because these strategic reserves would be recoverable in the event of a treaty exit, they would give states insurance against breakout by rivals. Because these reserves would not count against states’ covert evasion ledgers, they would be a cheaper form of insurance than a dark compute stockpile would (except insofar as they could arrange for other states to underestimate how much dark compute they had).
The existing stock of training-capable hardware would be redirected to the above uses. As the training-capable hardware stock shrunk by attrition, new hardware might eventually be required for sovereign fleets and scientific preserves. Specialized hardware to replenish these stocks might look very different from the current accelerator stock, most of which is optimized for AI training. Innovation by the semiconductor industry into hardware specialization and research developments in technical governance might eventually allow the total compute allocated to other uses, such as high-precision scientific simulations and computer graphics, to continue on a rapid growth trajectory even as the total training capacity for AI models dwindles.
3.3.2 Hardware licensing, registration, and tracking
Tracking accelerator hardware would be an essential objective of the treaty. Each hardware system under treaty governance could be placed in the care of a licensed hardware operator at a known data center, residing in a host state.26 Core licensing, registration, and tracking requirements would apply to all accelerators under the treaty’s governance, including sovereign fleets, scientific preserves, transitional inference service, and possibly inference-only service.27
There would be two routes for governance of hardware under the treaty: pre-pause hardware would enter via registration at the beginning of the treaty, while new hardware would enter by registration at the time of its manufacture at a licensed fab. The only approved way for hardware to exit treaty governance would be verified destruction at the end of its life. Hardware lost to the regime would likely be presumed dark compute and charged to host states’ covert evasion and overt breakout ledgers at near-full training capacity.
As early as possible, it would be important for governments to undertake a global accelerator census attempting, as far as possible, to account for every pre-pause accelerator unit (cf. Aarne and Petrie 2025; Scher et al. 2025, Art. ). This census would serve the twin goals of ensuring that only a small fraction go undeclared and classifying and quantifying the undeclared stock within each state. An accelerator server is a capital good with value in the hundreds of thousands of dollars that leaves documentary evidence in various forms when it is manufactured, sold, imported, operated, and written off—including Chinese customs and tax records of chips diverted around U.S. export controls, which we refer to as diverted hardware. See § 5.4.1 for a list of the many sources of documentary and other evidence about accelerators’ chain of custody and whereabouts.
A census could require declaration of accelerators by current hardware owners and compile information from a variety of sources on the location of each chip, including sales records, tax records of hardware purchasers, inventory records of data center operators, power bills, and chip rental records. Governments could supplement these efforts by offering generous buybacks and bounties for tips leading to the discovery of undeclared accelerators, as well as granting amnesty to shell companies and their customers for diversion around U.S. export controls or Chinese customs, conditional on producing retrospective records (cf. Halstead and Larsen 2026). To maximize coverage, the census would need the cooperation of U.S. and Chinese officials, and would need to include resellers and data center operators in jurisdictions outside the U.S. and China including Singapore, Malaysia, Taiwan, and the UAE. We discuss the prospects for success of this chip census in detail in § 5.4. The eventual successful completion of this census would be an important determinant of durability, as we discuss in § 6.5.
State requisition and concealment of pre-pause chips would be visible primarily as gaps in the census. If significant residual uncertainty remained about the pre-pause dark compute in a state, parties could arrive at a negotiated estimate based in part on compiling and reconciling the above data, which would factor into the state’s covert evasion and overt breakout ledgers. A state’s willingness to furnish data on request, to permit inspector visits to commercial data centers, and generally to assure other parties of its accounting transparency would likely be a factor in the negotiated estimate.
Licensed hardware systems would come equipped with (or, in the case of pre-pause hardware, be retrofitted for) on-chip location attestation (e.g., latency-based location Brass and Aarne 2024) to ensure that they remained at their registered location and visible to treaty authorities. Transfer of ownership, location, or jurisdiction would require re-registration. Chips that unexpectedly moved or fell silent would have to be produced for inspection or surrendered to authorities,28 and licensed hardware operators would also be subject to routine inspections. Operators’ eligibility for hardware licenses would be contingent on their record of compliance, with host states responsible for enforcement and inspector access.
As a design option to enhance compliance and address some concerns about regulatory capture, we suggest that AI companies (companies submitting whitelisted models) could be ineligible to apply for hardware operation or inference provision licenses, but still allowed to access their own whitelisted models through licensed inference providers. This design option, if combined with a mandate for AI companies to license their models to inference providers at regulated rates,29 could also potentially increase political palatability for constituencies that object to governments granting long-term monopoly rents30 to market leaders at the time the pause takes effect.
If vulnerabilities were discovered in the design of inference-only hardware that opened up jailbreak attacks that could enable model training by evaders, or if the models they implement were de-whitelisted (cf. § 7.3.4), inference-only licenses would be recallable for retrofit and re-licensing once the jailbreak is patched. The model, and hardware serving it, might be re-whitelisted contingent on additional external safety features that inference providers would be required to implement.
3.3.3 Sociotechnical verification measures for preventing illicit use of residual training-capable accelerators
At least until fully general privacy-preserving verification tools are ready for deployment, verification for training-capable hardware would need to lean on physical or institutional restrictions.
For transitional inference service on training-capable accelerators, verification could involve regulating the interface between hardware operators and inference providers. Hardware inside training-capable data centers could be exposed to inference providers as an idealized inference-only machine (recall § 3.2.1) through an application programming interface (API). Inference providers would submit a request including a choice of model, metadata such as prompt length, and an encrypted model input, and operators would return encrypted outputs along with other metadata such as output length. All encrypted traffic from outside the data center would pass through an internationally-operated router machine to small, physically isolated clusters within the data center and connected only to the router. Similar to the proposal by Dean (2026b), router traffic could be randomly sampled and sent to a secure recomputation server that would confirm the input matched the output without exposing the decrypted prompt to treaty authorities. Inference traffic and power usage would be reported per cluster and reconciled to power station and inference provider records. As defense in depth, these measures could be supplemented with telemetry-based workload classification such as that described by Rahman and Tajdari (2026) or zero-knowledge proof based attestation methods (e.g., those offered by Attestable n.d.).
Inspection protocols could be modeled on precedents from the Chemical Weapons Convention to ensure compliance with the laws of local jurisdictions. For example, regulations in the United States cover inspections with consent (no warrant required), routine inspections without consent (requires administrative search warrant), and challenge inspections (requires criminal warrant with probable cause).
For the internationally governed scientific preserves, verification could be achieved through workload transparency and social accountability. We envision that all workloads would be declared by named, vetted participants, and could be logged and replayable on demand by treaty authorities. Workloads could be reviewable by the general public insofar as restrictions such as confidentiality for medical and human subject data allowed.
For the sovereign fleet, verification procedures would be negotiated between parties on a case-by-case basis. Owing to the small total size of these fleets, the treaty ledger approach above would not necessarily require ruling out that any portion of their computing power were being used for training, if it were not possible to provide such verification without compromising important national security interests. Nevertheless, states would be charged less against their covert evasion ledger if they could show other parties some evidence that most of the sovereign fleet was not being used for illicit training. Technical measures such as power usage monitoring, inspection of network topology, or workload classification through telemetry metadata, or institutional measures such as whistleblower protections (Baker et al. 2025), might allow for some discounting of covert evasion capacity, even if the sensitivity of these workloads did not allow for airtight verification using present technology.
Finally, for strategic reserves, verification could rely on periodic inspection and counting to check that hardware remains under seal in its purported location.
As verification technologies mature over time, supported by implementation in new hardware-software platforms, the treaty authorities could tighten their practices further, achieving higher confidence that governed training-capable hardware is not being used for unauthorized purposes, and thereby allowing for more licenses to be issued. Two technologies that could play a key role in verification are trusted execution environments (TEEs) and formal methods. AI computations permitted in various contexts including transitional inference services, scientific preserves, or sovereign fleets could be accompanied by a proof that they adhere to agreed-upon safety rules and/or monitored for such adherence. Such a proof is not just restricted to the strict mathematical notion of proof, but more broadly includes the kind of evidence and assurance cases that are common in certification of safety-critical systems and safety engineering (Leveson 2012). The classical idea of “proof-carrying code” (Necula and Lee 1996) supported by advances in automated formal methods for Verified AI (Seshia et al. 2022), offers a mechanism: designers of permitted AI compute should accompany their deployed models with an independently checkable proof of safety that can be vetted by a trusted third party. Additionally, treaty authorities should be able to deploy runtime monitors within TEEs that are tamper proof. These monitors check for model adherence to safety rules and should send regular “heartbeat” updates to treaty monitoring servers indicating such adherence. Formal verification and systematic testing should be applied to the TEE platforms to ensure that these are free of bugs and vulnerabilities. The first step for such verification, having a comprehensive formal specification of TEE platforms (e.g. Subramanyan et al. 2017), is already available today and offers a foundation upon which advances in AI-assisted formal verification can achieve a higher level of assurance.
Another approach could be applied to existing hardware: non-invasive verification, which infers what a chip is computing from external physical signals (the power usage monitoring mentioned above for sovereign fleets is a simple example). Many possible techniques for this are already well-developed and widely deployed for the purpose of studying and waging side-channel cybersecurity attacks. Rather than verifying compliance, such measures are used to infer or collect protected information from a security system based on emissions from physical side channels; for example, power consumption, timing, or electromagnetic emission measurements are used to deduce cryptographic keys or secret messages without requiring specific knowledge about the workloads or operations. However, because non-invasive verification schemes and side-channel attacks rely on inferring a system’s internal operations from indirect, external signals, mature side-channel attack methods and toolkits could conceivably be repurposed to support verification techniques.
3.3.4 Governing hardware production
The regime would need to govern and monitor hardware production well enough to enforce the mandate that training-capable accelerator production be halted during Phase 1,31 while production continued on all other types of hardware including inference-only accelerators, consumer graphics cards below a treaty-governed performance bound, CPUs, and other hardware systems.
The regime would need to prevent illicit production of training-capable accelerators through two channels: production at undeclared sites, and diversion of production at declared manufacturing sites (cf. Baker 2023, App. G; Shavit 2023; Al Al Ramiah et al. 2025; Scher et al. 2025, Art. VI).
To prevent production at undeclared sites, the first line of defense could be to conduct a global lithography machine census (which would complement the proposed accelerator census), as well as the use of all-source intelligence to detect attempts to construct and operate such facilities covertly. In doing so, the treaty would need to account for each pre-existing unit of the world’s advanced lithography machines on which accelerator chips are printed, including at least every extreme ultraviolet (EUV) and ArF-immersion deep ultraviolet (DUV) lithography machine. These could be registered and tracked, along with other necessary tools such as excimer lasers, immersion tracks, and -beam mask writers, with location attestation according to procedures similar to those used for training-capable accelerators, as well as verified physical destruction at the end of their lives. Registered machine counts and serial numbers could be checked against a census of vendor delivery records, and any unaccounted-for machine could be charged to host states’ ledgers as continuously producing dark compute until they were located and accounted for. Factories capable of producing lithography machines could be monitored to ensure that new machines were registered at the time of manufacture, and registered lithography machines could send cryptographically attested records of each processed wafer to treaty authorities.32
Satellite imagery could be used to locate and enumerate the finite set of existing facilities sufficiently large and well-equipped with HVAC to host clean rooms for undeclared machines. Possible methods to confirm these sites had no undeclared machines could include environmental monitoring for offgassing, power substation metering, walk-through inspections to confirm buildings held no undeclared clean rooms, and all-source intelligence. Similar procedures could be used to rule out undeclared lithography machines at declared facilities.
At declared manufacturing facilities, monitors would need to verify both that declared machines were not producing undeclared chips and that declared chips leaving the facility were manufactured to specification or verifiably destroyed. The former objective could be achieved by reconciling the number of silicon wafers coming in, the number of logged wafer prints, and the number of chips coming out, while the latter could be achieved by randomly sampling a small number of outgoing chips for non-destructive or destructive inspection and comparison to licensed hardware specifications.
As defense in depth, mass balance accounting of perishable consumables like photoresist chemicals could prevent skimming in the service of undeclared production; each additional tracked input would mean an evading state would need to add another node to its covert supply chain without detection.
3.4 Transition governance
In this section, we discuss governance during Phase 1 of the hardwired pause, when the pause begins, and Phase 2, when training-capable accelerators are replaced with inference-only accelerators.
3.4.1 Economic alignment
For a hardwired pause to be effective, negotiators would need to design mechanisms that align hardware designers, AI companies, and inference providers with the spirit and letter of the regime requirements, in consultation with economic experts. Here we discuss one possible implementation.
In order to ensure that the initiation of the regime would not create a negative shock to the global inference supply, parties to the treaty would approve a provisional whitelist of both models and hardware, while they deliberated on how to implement updated processes going forward. Continuing diffusion of AI models could mean strong ongoing growth in inference demand; Emberson and Sevilla (2026) estimate current inference demand at between 200M and 4B tokens/s, growing at a tenfold rate per year. After an initial rise in supply due to reallocation of compute resources from training to inference at the start of Phase 1, the global inference supply could struggle to keep up with demand while the market awaited inference-only hardware.
During this interim period, market forces could exert strong pressure on hardware companies to rapidly ramp inference-only production; the prospect of a windfall from being the first hardware supplier to begin producing inference-only chips at scale might create race conditions favoring the treaty. An interim gap between inference supply and demand could likewise create a lucrative inference market for suppliers and a high opportunity cost for stockpiling undeclared hardware.
Effective treaty governance would exploit these market incentives by converting them into the treaty’s most important needs: visibility into, and eventually surrender of, training-capable hardware. In particular, hardware owners would need to declare and register their hardware immediately if they wanted to continue using it for lucrative inference service throughout the interim period, with late declaration carrying a financial penalty and undeclared accelerators subject to seizure and resale by treaty authorities,33 who could use proceeds to fund generous bounties for help in discovery.
3.4.2 Phase 1 (design)
We consider Phase 1 to begin as soon as the U.S. and China agree to a pause minimally implementing the two core mandates against training new frontier AI models and manufacturing new training-capable accelerator chips; however, these two decisions might not occur simultaneously and might go into force before a formalized agreement is fully negotiated.
During Phase 1, the three negative mandates would become operative in all participating states:
Training of beyond-threshold models, using any hardware, would be prohibited (see § 3.2.4).
Use of existing accelerator hardware would be prohibited for any use other than licensed inference on whitelisted models or designated exceptional uses such as those in § 3.3.1.
Manufacturers would be directed to stop production of accelerator hardware not included on the provisional whitelist of inference-only hardware.
International verification and domestic verification and enforcement measures in support of these mandates would also be ramped up quickly at the outset of Phase 1. The global accelerator census would also begin, and all training-capable AI chips performing transitional inference service would be registered under treaty governance.
In addition to approving the provisional model whitelist (which could initially include all submitted models pending further evaluation), parties or interim authorities would evaluate and approve a whitelist of submitted hardware designs. In addition, hardware owners would declare and register any accelerator hardware that they wished to continue operating under the treaty, and it would be retrofitted for location attestation. Failure to comply would incur graduated legal penalties including seizure upon discovery of undeclared hardware; late compliance would entail financial penalties and lost profits.
If the U.S. and China wished to halt frontier training on an urgent timeline, they could do so in short order34 by directing the several leading companies in each of their states to halt training.35 They could also stop the production of accelerator hardware by forbidding additional orders for accelerator chips by American and Chinese semiconductor companies,36 which comprise virtually the entire global market, and in the meantime recruit additional U.S. allies such as the EU, Japan, and South Korea, which control supply chain chokepoints (cf. § 5.2).
To reduce AI companies' commercial incentive to continue illicit training, the U.S. and China could also set a short-term model whitelist submission deadline to pass before negotiations were complete. Adding the other supplier monopoly states would buy more time until it were possible to induce most other states to join.
3.4.3 Phase 2 (substitution)
During the managed substitution phase (Phase 2), transitional training-capable inference service would be gradually displaced by inference-only compute, at a rate determined by the inference market, the rate of industrial production of inference-only hardware, and/or the states parties.37
To ensure the timely retirement of training-capable inference service once it is no longer needed, there would also be a cap on the number of hardware units, measured in H100e, for which licenses are available at any given time. Firms wishing to continue serving inference on pre-pause training-capable hardware would bid for licenses at a periodic auction, licenses would expire at the time of the next auction, and the number of licenses available at each auction would decrease over time according to a schedule based on the availability of new inference-only hardware. Proceeds from this auction could fund treaty authorities or buybacks for pre-pause accelerators. Hardware operators would be required to trade in training-capable accelerators upon license expiry.38 If certain classes of models are expected to be underserved by inference-only compute for longer, the license cap schedule could apply on a per-model-class basis.
3.4.4 Traded-in training-capable compute
The rate at which training-capable compute is withdrawn from inference service would depend on the rate at which inference-only compute is produced. Our empirical analysis in § 4.4 suggests that if production of inference-only compute follows a similar design schedule as recent inference-leaning chips, the duration of interim inference-only service on training-capable compute could be less than five years, but the actual duration might be somewhat faster or slower.
In § 3.3.3 we discussed workload verification for training-capable hardware, including ongoing progress in TEE and formal verification methods. It is possible that before inference-only production ramps up enough to replace training-capable inference service, computer scientists will develop secure ways to retrofit training-capable hardware so that it can remain in inference service for the remainder of its lifespan without significant evasion risk (cf. Scher and Thiergart 2024; Petrie et al. 2025). In that case, traded-in chips could be retrofitted and returned to transitional inference service, if their remaining profitable lifespan justified the expense.
Unless the duration is significantly longer than expected, the training-capable accelerators retired from inference service will be aging but still functional. If it is not possible, or not profitable to retrofit them for verified service, they could be retired and sent to sovereign fleets or scientific preserves, or destroyed.
3.5 Possible exit conditions
In this section we consider the conditions under which a pause might be exited by states parties. A hardwired pause could take effect for an initial time period, for example ten years, with mandated renewals, through a process managed by states parties or treaty-implementing authorities (Finke 2026).
There is currently no agreed-upon paradigm for deciding whether, or under what conditions, building more powerful AI models would be broadly safe or beneficial for humanity at large. Any decision to resume training more powerful AI models after a pause would be informed by scientific advancements, social and economic developments, and political value judgments, all of which are impossible to predict when humanity’s experience with AI systems has been so brief and so limited to a relatively small fraction of people across the globe. Even technical alignment is still in many ways an open scientific problem (Bengio 2026).
As a result, it would be premature to name technical conditions under which a pause might be lifted, but a pause would need to clearly specify procedural steps for any exit condition. We can envision several possible conditions for exiting a hardwired pause:
First, states parties might mutually decide to dissolve the treaty because they mutually agree that humanity has discovered satisfactory solutions for the constellation of sociopolitical and technical problems that originally led to pause-willingness. Such a solution slate could include technical AI safety procedures (cf. Scher et al. 2025), economic or labor market mitigations, and social and/or political innovations that ensure human flourishing, equality, liberty, and safety under plausible AI futures. It might also include a plan for a collaborative international project to develop more powerful AI models without the competitive race dynamics characteristic of the present AI race (cf. Scher et al. 2025; Larsen et al. 2026).
Second, the governance paradigm of the hardwired pause might definitively fail, triggering a condition for dissolution. For example, this could happen because technological advances favoring evaders render treaty mechanisms for monitoring, verification, and enforcement obsolete, and new mechanisms cannot be found.
Third, sufficiently severe departures or violations could trigger an exit to the treaty. A major review or termination of the agreement could follow potentially “treaty busting” developments, such as the orderly departure or treaty-violating defection of a sufficiently powerful state.
Alternatively, if the first decade of a hardwired pause proved prosperous for the world, leaders may decide to renew the pause via a renewal procedure determined during initial negotiations.
3.6 Institutional governance
A fundamental issue for a hardwired pause will be how well the treaty is implemented: first by the initial parties, then by the interim authorities, and eventually by the treaty authorities, all states parties, and impacted industries. Historically, treaty authorities and the member states of the treaty—the states parties—share a responsibility to effectively govern international agreements. A hardwired pause would be no different. Thus, it will be important to make wise choices regarding how institutional governance can best support treaty implementation. Previously in Section 3.1 we described how a hardwired pause treaty organization may evolve and fully emerge over the four phases described in Section 1.1. This section reviews some of the main choices for institutional design that could enhance governance of the hardwired pause.
Design choices for how the treaty will be governed legitimately through the four phases of institutional evolution are core to hardwired pause feasibility. Treaty organization leadership and staff will require the highest standards of competency, probity, and independence. The mandate and scope of the hardwired pause establishes explicit and implicit requirements for a future institution, one composed of treaty authorities and most likely some type of forum for regular deliberations by states parties such as those listed in § 8.2. Functions such as management of an international ledger of compute, prohibited chip reclamation and adjudication, administration of whitelist process decisions, monitoring states parties’ activities under the treaty, and deploying inspectors to verify treaty provisions all incur institutional implications and costs.
Beyond these foreseeable requirements, lessons from existing international organizations and historical international agreements provide additional insights for effective institutional governance. The potentially evolving membership of the treaty, especially in the four phases, will shape the formulation of any interim authority bodies and the eventually established treaty organization. Early clarity for prospective states parties on the benefits of joining the treaty, especially for countries with less developed AI capacity or economic advancement could generate faster cooperation. This implies that even interim authority design will need to have some capacity to manage the benefits sharing (see § 3.1) or facilitate technical cooperation. For example, the envisioned benefits of the scientific preserves (§ 3.3.1) entail staffing and management, institutional processes for reviewing scientific cooperation, and data center operations controls and administration. Given the potentially significant costs to some states if they remain outside the treaty, procedures and staffing may be needed for managing hardwired pause implications on external parties.
Careful formulation of the legal terms for a hardwired pause point towards a binding treaty with a range of precise commitments, likely necessitating experienced legal expertise within the early ranks of the treaty authority formulation. There are many proposed obligations of states joining the hardwired pause, from implementing prohibitions to supporting international capacity building and potentially shared AI scientific or safety research. Many of those obligations could be more effectively upheld if there were treaty authorities and personnel serving in coordinating or facilitating roles between states parties.
With the passage of time and projected advances in technology and AI development methods, the adaptivity of the treaty’s provisions needs up-front design as well. For example, treaty authorities will benefit from flexible procedures for working with states parties to account for advances in algorithmic progress or other treaty relevant updates. This may be helped by the formulation of various technical committees that regularly review these issues and nominate adaptive measures or protocols to states parties, perhaps mirroring how the International Civil Aviation Organization (ICAO) governs for adaptivity (Chicago Convention 1944). One issue that could require adaptivity involves how prohibited thresholds would be adjusted (e.g., predetermined schedule, renegotiation at regular conventions, or delegation to treaty authorities). Each implementation would have its own pros and cons, and negotiators of the treaty would need to find a mutually agreeable arrangement for threshold determinations.
How the treaty authorities contribute to delivering consistent evidence of hardwired pause compliance, or how they operate when treaty obligations are being violated, will require significant institutional design and accounting. For example, treaty authorities will require sufficient human resources, data processing and analysis, and accounting procedures to manage the compute ledger with fairness, accessibility, and transparency. This is not unlike design choices the Preparatory Commission for the Comprehensive Test Ban Treaty Organization (CTBTO) has had to make in manning and operating their global International Monitoring System (IMS) and associated International Data Center (IDC). The hardwired pause treaty authorities could require similar infrastructure to bolster the regime’s credibility. Additionally, we predict the need for a substantial international inspection process to verify a number of the treaty’s provisions. How that inspection cadre is recruited, trained, deployed, and retained implies a potentially large standing unit of inspectors, support personnel, and associated technical verification equipment or procedures. Institutional forums and procedures for reviewing and making political judgments on monitoring and verification will also need to be designed into the treaty organization, with appropriate methods for enforcing treaty provisions under the treaty procedures, or externally. As described in Section 3.4, treaty authorities will be heavily involved in enforcing compliance with industry declarations (e.g. to support the global accelerator census and to qualify companies for licenses to use training capable chips for inference) and collection of relevant fines, likely requiring a specialized unit and financial system for operationalizing those treaty provisions.
A treaty designed with clear decision rules and delegated powers for the treaty authorities will impact states parties’ views on feasibility, especially when weighed against needed financial resources and staffing of the treaty authority (commonly referred to as a technical secretariat or staff). Funding for the treaty organization could come from a mixture of funds such as transitional licensing fees, compliance fines, and contributions from states parties. We anticipate these funding mechanisms may be sufficient for full-scale operations until fees from transitional licensing decline. The actual costs to institutionally implement a hardwired pause and how to sustain longer term financing require further analysis. As for the size and capabilities of the treaty authority staff, we envision that the potentially large requirements will necessitate early focus by states parties, and significant effort to recruit and prepare these personnel for their duties. Lessons from the EU AI Office are instructive in this regard. While the office was established in legal form in 2024, capacity for the office to fully implement its legal mandate and enforcement duties remains behind schedule. Ideally, a hardwired pause treaty organization is developed with sufficient human talent and on appropriate timelines.
Design choices for how non-state party actors such as civil society, or AI industry representatives, are included within the formal or technical institutional structures of the treaty organization will also be important to further develop going forward. As detailed in multiple sections of the paper, the plausibility and durability of the treaty, especially in Phase 1, will hinge on how the AI industry is engaged institutionally by treaty authorities. This implies the need for industry councils or representational processes with treaty authorities and states parties, with both formal and informal levels of involvement. Given AI’s expanding political and social impacts, more structured participation by civil society and independent experts may also be helpful for durability.
The hardwired pause treaty’s legitimacy will rest on calibrating a mix of state party managed processes and accountability measures such as the compute ledgers and envisioned licensing procedures. Early development (and enactment) of domestic implementing legislation for the treaty’s mandate and obligations will be important in this regard. A related component of legitimacy will be how well the treaty authorities and states parties are secured and defended in the course of their work. Given the importance of the treaty’s institutions and span of control and monitoring over global AI compute and infrastructure, we assess that well defended physical, human, and digital infrastructure for the regime is a fundamental aspect of feasibility—especially durability. Unlike the pure human directed threats to international organizations that have happened in the past, such as the 2018 attempted intrusion into the Organization for the Prohibition of Chemical Weapons (OPCW), attributed by the Netherlands and UK to Russian intelligence (Government of the United Kingdom 2018), hardwired pause treaty institutions may also be the target of AI systems themselves.
4 Feasibility of inference-only hardware
For a hardwired pause to be technically feasible, the key technology of inference-only AI chips must be possible to develop. In this section, we define the term and assess the possibility of developing such chips within a reasonable time frame.
4.1 Quantifying training capability
We operationalize the residual training capacity of an inference-only accelerator as its arithmetic throughput for training operations. Recall the definition of training operations in § 3.2.1 as calculating gradients or performing inference on generic AI models, with each operation quantified as required FLOPs of dynamic matrix multiplication in the service of either. A matrix multiplication is dynamic if and only if both operands are determined at runtime: multiplying a fixed weight matrix by a mutable activation matrix is not dynamic.
Not all dynamic matrix multiplication capacity is training capacity: a matrix multiplication performed as one internal operation of a complex application-specific integrated circuit cannot necessarily be exploited to perform any matrix multiplication the user wants. When dynamic matrix multiplication units have user-writable inputs and user-readable outputs, we refer to them as general matrix-multiplication (GEMM) units. In practice, the primary specification for certifying inference-only accelerators is that they should be impractical for performing high-throughput GEMM in the service of training operations.
The distinction between dynamic matrix multiplication and GEMM is critical because implementing inference for whitelisted weights internally requires performing matrix multiplication at scale, including by necessity dynamic matrix multiplication for attention operations. As a result, licensing inference-only hardware requires ensuring that users cannot repurpose these operations as GEMMs by using the hardware in ways not intended by its manufacturer, a process we refer to as GEMM laundering.
The residual training capacity of any inference-only accelerator would be lower-bounded by its GEMM-laundering capacity through its token interface, i.e., its arithmetic throughput when using whitelisted models as oracles for matrix multiplication. However, these capabilities would fall far short of even unregulated hardware like CPUs: assuming two bytes of information per token, an inference-only machine with a billion-token context length and the ability to output one million tokens per second still could not achieve more than 12 GFLOP/s through this route, requiring more than 80,000 such machines to match the arithmetic throughput of a single H100.39 Thus, non-negligible GEMM laundering could only be achieved insofar as an inference-only accelerator exposed more than just the token interface to its operator.
4.2 Practical inference-only accelerators
In practice, it is unnecessary for a realized hardware system to be an inference-only machine in the idealized sense defined in § 3.2.1. We define a (practical) inference-only accelerator as any hardware system that can perform efficient inference on at least one whitelisted model but has verifiably modest residual capacity for performing training operations. Restrictions on training capacity must be robust to the system being stolen or seized by a nation-state level adversary with physical possession of chips and full control over the deployment environment and software stack.40
Licensing of inference-only accelerators in treaty governance would be neutral to hardware implementation, encouraging innovation by the semiconductor industry to explore different techniques for hardware specialization. The chief licensing criterion would be a quantified and verifiable bound on a hardware system’s residual capacity to perform training FLOPs in the service of a large training run.
In particular, units that perform dynamic matrix multiplication would require careful inspection to ensure that they are not GEMM-launderable. Whereas weight projections and multilayer perceptron (MLP) heads could be implemented with one operand (the model weights) immutable at runtime, attention head calculations irreducibly require dynamic matrix multiply operations.41 If the hardware design allows for an adversary to commandeer these operations by injecting generic inputs and exporting results, these operations would constitute significant residual training capacity, measured in tens of percentages of the total arithmetic throughput required for inference.
To rule out GEMM-laundering risk, regulators must ensure that at least one of the two operands for each of these dynamic matrix multiplications is populated directly from outputs of a previous operation, and/or that outputs are in turn routed directly to the next operation, without input/output exposure to the software layer or physical interfaces such as inter-chip links, off-chip memory, or test and debug ports.42 In that case, any GEMM laundering attack would require indirectly writing these operands by writing to upstream inputs and/or indirectly reading the results by reading downstream outputs. The more linear and nonlinear operations are chained together between these upstream inputs and downstream outputs without user access to intermediate values, the more formidable this obstacle would become. If treaty authorities can certify, based on a rigorous audit of the hardware system,43 that such laundering is infeasible (or bound the rate at which it could plausibly occur), then the system’s high-throughput dynamic matrix multiplication could contribute very little to its effective GEMM rate, leaving only low-throughput residual channels like embedded control processors or the arithmetic-by-prompting attack discussed in § 4.1; neither of these would be sufficient to attain accelerator-grade training capacity. Host processors with significant arithmetic throughput for matrix multiplication would also need to be either excluded from use with such systems or accounted for outside the licensed boundary of inference-only hardware systems, along with other CPUs.
4.3 Model-specific integrated circuits (MSICs)
A key advantage of the hardwired pause is its potential for technological readiness in the relatively near term. The most promising examples of deployed inference-only chips are certain model-specific integrated circuits (MSICs).44 An MSIC is a chip that hardwires the architecture and weights of a specific AI model when the chip is fabricated. Whereas most computer chips are built with the flexibility to implement a wide variety of programs, and store a representation of a model’s weights and architecture in writable memory, the metal wiring of an MSIC immutably fixes the architecture and weights of one model. Even without regulatory pressure, the semiconductor industry is already designing and deploying such chips purely for their commercial and technological advantages.
On general-purpose accelerators such as the H100, performing inference requires reading billions or trillions of model weights from high-bandwidth memory (HBM) for each new token serially generated. This memory bottleneck limits the serial speed at which inference can be performed to several hundred tokens/s per user, depresses the utilization of fast arithmetic units even after accounting for batch parallelization, and is responsible for most of the dramatic power demanded by these chips. A wide variety of AI labs and hardware companies are aggressively pursuing specialized inference hardware that ameliorates the memory bottleneck in a variety of ways, including by increasing HBM bandwidth or by displacing HBM with more expensive static RAM (SRAM). The MSIC solves this problem by hardwiring weights into the chip, obviating the need to read them from off-chip memory (Bajic 2026).
The hardware company Taalas, which AMD has agreed to acquire (AMD 2026), specializes in making chips of this type: its first-generation prototype HC1 chip implements models up to 8B parameters. Taalas has implemented and deployed a version customized to implement the Llama3.1-8B model and reports that its chip serves inference at 17,000 tokens/s per user, versus 353 tokens/s per user for the B200 accelerator (Taalas n.d.).45 This advantage is in per-user serial speed, not aggregate throughput, where general-purpose accelerators make up this difference in parallel processing. The HC2, Taalas’ second-generation chip scheduled for deployment in late 2026, is designed to implement frontier-scale models by sequentially chaining chips holding 20B parameters each. Taalas also reports that its manufacturing process is suitable for fast and low-cost specialization to new models, allowing it to realize any AI model in customized hardware within two months of receiving the weights (Bajic 2026).
A second recent example is the hardwired-neuron language processing unit (HNLPU), an MSIC architecture designed by academic researchers (Liu et al. 2026). The HNLPU is not commercially available, but its architecture is public, allowing for better initial assessment of its routing. The intended use of the chip is only to produce token outputs from token inputs, with no software stack. However, likely because it was not designed with the goal of hardening it against adversaries, the proposed implementation leaves open adversarial channels for GEMM laundering.
According to the paper's peer-reviewed design, at least one operand for each attention operation is derived from hardwired projections of activations, with the result fed into another hardwired projection. But the proposed implementation is composed of 16 chips in separate packages, and each attention head is split across four chips. As a result, several intermediate results in attention operations are sent between chips over CXL links, giving opportunities for injection and exportation.46 In order to close these channels and license the HNLPU at a low residual training factor, designers would need to protect these links from injection and exportation, for example by placing each group of four chips that share attention heads in a single sealed package, or by requiring that cross-chip traffic be encrypted and authenticated using keys physically hosted on each chip; authorities would need to audit the resulting system for other GEMM-laundering channels.
The commercial viability of hardwired MSICs is especially notable in the present market environment, where new frontier models are released every month. Google reportedly also designed the first version of its Frozen chip as an MSIC but ultimately reverted to a design that etched only model architecture, due to the short life cycle of model weights (James 2026c; Woo and Liu 2026). The commercial advantages of MSICs over general-purpose hardware would correspondingly be greater in a world with a frozen whitelist (see § 7.3). There is every reason to expect the already-robust trend of innovative hardware specialization in the semiconductor industry to accelerate further after a hardwired pause.
An MSIC as defined above, requiring that both the architecture and weights of a model are hardwired into the chip at the time of manufacture, is not the only conceivable type of inference-only chip that is specialized for a particular model. An alternative type of chip, which we might call an Authenticated-Weight Inference-Only Chip (AWIC), hardwires the architecture of the model and more generally fixes the inference function of the chip at the circuit level, in such a way that the chip has very low residual training capacity, but it allows the weights of the model to be read in from memory during inference, with hardware authenticating the weights used for inference against a cryptographic hash hardwired into the chip at the time of manufacture.47 One potential advantage of this approach is that it may be easier and less expensive to scale to large models, since the large set of weights can be held in high-density memory, rather than being hardwired into model-specific silicon; on the other hand, this incurs the cost of memory transfer and authentication, which may reduce speed and energy efficiency. We believe that inference-only chips based on MSICs (which should perhaps be called Hardwired-Weight Inference-Only Chips or HWICs) and AWICs should be investigated for the purposes of a hardwired pause.
4.4 Time required to ramp inference-only hardware production
In this section, we make very rough estimates for the duration of Phase 1 (design and production setup for inference-only chips) and Phase 2 (managed substitution of inference-only chips for flexible accelerators) of a hardwired pause. Recall our assumption that by the beginning of Phase 1, marked by the pause agreement, there is a provisional model whitelist. Thus, on day one of Phase 1, hardware designers and manufacturers can initiate development of inference-only chips for models on that provisional model whitelist. The scenarios below concern inference-only hardware suitable for running at least the initially whitelisted models; if new hardware or adjustments to hardware are needed to run models added to the whitelist at later times, such hardware may become available after the initial wave of inference-only chips.
As suggested in Figure 1.1, the duration of Phase 1 depends heavily on industry progress on inference-only chips during Phase 0. Moreover, in reality there is no single Phase 1; instead, there are different Phase 1s for different hardware companies. For example, as noted in § 4.3, due to having developed a platform for hardwired inference chips during Phase 0, Taalas claims that “From the moment a previously unseen model is received, it can be realized in hardware in only two months” (Bajic 2026). Creating chips from scratch starting in Phase 1 would take considerably longer. Such an effort must pass through three subphases: Phase 1.1, design to tapeout;48 Phase 1.2, tapeout to first silicon; Phase 1.3, first silicon to production release. For comparison, consider OpenAI’s Jalapeño chip (OpenAI 2026a), which is an inference-optimized rather than inference-only chip. Thanks in part to accelerated design using OpenAI’s models, Phase 1.1 (starting with “architecture concept” and ending with tapeout) for their Jalapeño chip took 13 months, namely October 2024 to November 2025 (Ho et al. 2026). Phase 1.2 then took an additional 6 months, from November 2025 to May 2026 (Ho et al. 2026). We do not yet know the duration of Phase 1.3 for Jalapeño, but OpenAI plans initial deployment of Jalapeño by the end of 2026 (OpenAI 2026b), so we roughly estimate 7 months for Phase 1.3. This brings the total duration of Phase 1 for Jalapeño to 26 months. Since AI models will likely be significantly more capable at chip design at the time of Phase 1 than they were in October 2024 when Jalapeño development began, future chip design phases may go even more quickly than they did for Jalapeño.49 Considering a range of chip companies—from those that start inference-only development from scratch at Phase 1 to those that already have an adaptable inference-only platform at the beginning of Phase 1—we might expect timelines of roughly 26 months or less for their Phase 1s. Table 4.1 shows other AI chip design timelines, most of which occurred before the kind of powerful AI assistance used for Jalapeño was available.
| Chip | Span Measured | Months |
|---|---|---|
| Designs from scratch | ||
| Google TPU v1 (28 nm) | “designed, verified, built, and deployed in datacenters in just 15 months” (Jouppi et al. 2017, 2) | 15 |
| Groq TSP, later LPU v1 (14 nm) | Founded in 2016 (Groq 2026) to first silicon in July 2019 (Abts et al. 2020, 155) | ~36 |
| Cerebras WSE-1 and CS-1 (16 nm) | Cerebras founded in 2015 (Cerebras Systems n.d.); WSE-1 announced in August 2019 (Cerebras Systems 2019); WSE-1 and CS-1 available in data centers in 2019 (Cerebras Systems n.d.) | ~48 |
| Etched first-generation inference hardware (TSMC N4P) | Seed round in June 2023 with “most of the architecture fleshed out” (Ward-Foxton 2023) to first working silicon in first half of 2026 (Etched 2026b; Bort 2026); first rack shipped July 2026; “three years to deliver our first rack from scratch” (Chowdhry 2026) | ~36 |
| Taalas HC1 (TSMC N6) | Founded mid-2023 to beta service with HC1 in February 2026 (Bajic 2026; Prickett Morgan 2026) | ~30 |
| OpenAI Jalapeño (with AI assistance in design) | October 2024 conception (Ho et al. 2026) to planned deployment by end of 2026 (OpenAI 2026b) | ~26 |
| Derivative designs | ||
| Nvidia data-center GPU (generic) | Engineering effort per generation: “multiple engineering teams coordinate for as long as two years” (Merritt 2023) | |
| Cerebras WSE-2 (16 nm to 7 nm) | WSE-1 announcement in August 2019 (Cerebras Systems 2019) to WSE-2 announcement in April 2021 (Cerebras Systems 2021) | 20 |
| Cerebras WSE-3 (7 nm to 5 nm) | WSE-2 announcement in April 2021 (Cerebras Systems 2021) to WSE-3 announcement in March 2024 (Cerebras Systems 2024) | 35 |
| Taalas HC1, new model on existing platform | Weights to RTL “about a week’s worth of effort” (Bajic, quoted in Ward-Foxton 2026a); weights to deployable PCIe cards in two months (Bajic 2026; Prickett Morgan 2026) | 2 |
As for Phase 2, its duration can be roughly estimated as the amount of training-capable compute to be displaced, divided by the average monthly quantity of inference-only compute, measured in inference-H100e, that can be installed to provide the same amount of inference service. For illustration, suppose that at the beginning of Phase 2, there are 47M H100e of training-capable compute50 (extrapolating to year-end-2026 from Epoch AI 2026d; 2026c). Our extrapolation of Epoch AI’s (2026d) estimates (Figure 4.1) also implies roughly 2.9M H100e of new AI compute capacity shipped each month during Q4 2026, which is rising. If we assume that in Phase 2, on average chip manufacturers can ship and install a quantity of inference-only chips that provide inference service equal to 2.9M inference-H100e per month, then we obtain a rough estimate of 47M / 2.9M per month ≈ 16 months for Phase 2. If global hardware manufacturing capacity continues to grow during Phase 1, and new inference-only chips are in high demand as the only way to meet rising inference demand, then the rate of production could be higher, leading to a shorter Phase 2.
In light of the considerations above, it seems within the realm of possibility that under the strong market pressures and government incentives to develop inference-only hardware due to a hardwired pause, the world could transition through Phases 1 and 2 in less than 5 years, depending on the amount of R&D carried out in Phase 0.
5 Feasibility of hardware governance
5.1 Overview
Our second historical lesson in § 2.2.2 was that governance regimes need concrete physical artifacts that can be tracked and regulated. AI accelerator chips are well-suited to serve this role for several reasons.
First, the supply chain for AI accelerators is one of the most complex, specialized, and geographically far-flung of any product in human history (Miller 2022). § 5.2 discusses numerous chokepoints in this supply chain, which reside in several different jurisdictions, making it extraordinarily difficult for any state to bypass the existing supply chain. As § 5.3 discusses, this concentration facilitates international governance to ensure that new accelerators are tracked by the treaty, and the manufacture of treaty-violating hardware produces physical evidence, in the form of the chips themselves, that can be randomly sampled and inspected.
Second, pre-pause AI chips are extensively documented through transaction, tax, and other records, including the diverted stock in China, as discussed in § 5.4.1. In particular, in both the U.S. and China, they are expensive capital goods that businesses have strong financial incentives to declare to their own governments, through VAT invoices or tax deductions. Additionally, AI chips are highly concentrated by both ownership and geography: a handful of firms collectively own the majority of the global stock (Epoch AI 2026c), and most AI chips are currently deployed in large data centers whose power and cooling infrastructure make them physically conspicuous (Sastry et al. 2024; Pilz and Heim 2023).
Third, incentives of both states and private actors would support compliance with the global accelerator census, as discussed in § 5.4.2. Private chip owners could earn money by participating in profitable transitional inference service and/or buyback programs, and states could pursue their national security interests more cheaply through both sealed strategic reserves and licensed sovereign fleets than through concealment of pre-pause stock.
These advantages would all bode well for the success of a well-conducted global census during Phase 1 of a hardwired pause. § 5.5.1 derives, in two steps, conservative quantitative estimates for how much stock might go missing in each of the U.S. and China after a thorough census; i.e., after one in which a state had made good use of the evidence at its disposal to locate and register pre-pause hardware. We first classify chip holdings into four categories based on their visibility, estimating the quantity in each category in both the U.S. and China. We then give upper bounds, based on historical analogies, for the fraction of pre-pause chips in each of these categories that might go missing, if states used reasonable means at their disposal to locate and register them.
5.2 The supply chain for accelerator chips
This section gives an overview of the many specialized inputs required to construct an AI accelerator chip and the various chokepoints that could provide cooperating governments multiple points of leverage.
The construction of global distributed computing for the training and provisioning of advanced AI models is driving one of the fastest growing sectors of the global economy. The top 5 hyperscalers in the United States are spending more than $550 billion on AI infrastructure in 2026 alone (VandeHei 2026). This is part of a broader trend, where global AI computing capacity has doubled roughly every seven months (You et al. 2026). Though this infrastructure is important everywhere, the United States currently enjoys and is likely to continue to enjoy significant advantages in the production of compute. China has nascent and growing capacity, but it continues to lag the United States (McGuire 2025; Murphy 2025). And yet, a comparison of gross compute conceals a vast network of interconnected supply chains that span continents, link through dozens of countries, and involve thousands of entities.
From the mining of minerals like high purity quartz, copper, cobalt, and certain rare earth elements, to the finished chips to the software and infrastructure system that supports AI training and inference, the provision of advanced compute requires a significant number of steps. Many of these steps involve a unique supply chain, often with facilities or equipment that have critical junctures that act as natural bottlenecks. Many of these bottlenecks are not located within the United States or China, but are rather part of a global distributed supply chain. Replication of any one of these steps could prove difficult. Some of the most sophisticated may require a government-led effort that could require decades. Others require that one has the right natural endowment of raw materials or can secure access to them.
Natural resources are a significant constraint in the compute supply chain. Refined rare earths are an often discussed choke point, as many of them originate in or are refined in China. In 2024, 91% of the global supply of refined magnet rare earths was refined in mainland China, and in 2025, China’s dominant position in these minerals came into focus during a trade dispute with the United States (International Energy Agency 2026). Other materials, such as high purity quartz, are known to come from a handful of locations. When a hurricane hit the small town of Spruce Pine in North Carolina, which accounts for an “overwhelmingly” large portion of the global supply of the material, global silicon supply chains were disrupted (Brumfiel et al. 2024). While the raw materials necessary for compute supply chains are not plentifully dispersed around the globe, many of the bottlenecks are the result of market forces and could potentially be overcome with increased surveying, investment, and mining capacity. This is less true for more technologically sophisticated aspects of the compute supply chain.
Many of the most technologically impressive aspects of the compute supply chain are extremely difficult to replicate. One of the most enduring bottlenecks in the manufacturing process is connected to Deep-Ultra Violet (DUV) and Extreme Ultra-Violet Lithography (EUV) machines. The Netherlands’ ASML has a dominant market position in the production of DUV and a monopoly over the more advanced EUV. These are complex machines, with an analysis from Georgetown’s Center for Security and Emerging Technology (CSET) showing that “five thousand suppliers provide 100,000 parts, 3,000 cables, 40,000 bolts, and two kilometers of hosing to make an EUV tool. The tool weighs about 180,000 kilograms (200 tons), and ships in 40 containers spread over 20 trucks and three cargo planes. ASML reportedly only makes 15 percent of the EUV tool in-house, partnering strategically with firms worldwide to source the highest quality components” (VerWey 2024, 6). Just to move one of them would require three jumbo jets with at least 100,000 pieces each according to IBM (Murphy 2023).
Many of the parts related to EUV machines are also characterized by single or dominant suppliers. The Germany-based company Zeiss is the only producer of some of the necessary mirrors and optical components, which are among the world’s most advanced (Stammler 2021), while the Japan-based Tokyo Electron has a nearly complete monopoly over coater/developer equipment used in EUV lithography (Tokyo Electron 2026). Those nations wishing to replicate EUV lithography would have a significant amount of work replicating not just what is produced in the Netherlands but also the various subsidiary components spread around the world.
TSMC and related enterprises in Taiwan represent another significant technical bottleneck. TSMC possesses the most advanced foundries and packaging processes, as well as several other key components for the production of compute. TSMC is currently in the lead, but there are some potential competitors and redundant capacity being developed in the United States. In addition, both South Korea’s Samsung and China’s SMIC are capable of producing high quality, if not as advanced, AI relevant chips. China’s SMIC is of particular note, as it is assisting in the development of the Huawei Ascend series (Feldgoise and Dohmen 2024). Meanwhile, advanced chip fabs are being constructed in the United States. TSMC has offshored some capacity into the US, such as its facilities in Arizona, though they are not yet as capable as the facilities in Taiwan (Rutherford 2026).
Nvidia is the world's largest supplier of advanced AI training chips. Epoch AI estimates that over 60% of global AI compute capacity comes from Nvidia's chips (Epoch AI 2026c). This concentration comes from at least three factors. First, Nvidia has been among the leaders in the design of advanced hardware as measured by FLOPs, energy efficiency, and memory bandwidth. Second, Nvidia's CUDA software stack encourages programmers to use Nvidia's chips. This software includes packages for allowing non-specialized software engineers to easily write software for GPUs, software libraries to optimize the speed and power usage of GPUs, compilers for writing optimized low-level code, and optimized application libraries. Third, Nvidia has invested heavily in its developer ecosystem, forming relationships with partners in academia, startups, and large companies for new applications. This investment has led to developer familiarity and eventual ecosystem lock-in, which in turn have led to Nvidia being the default choice for many companies who are risk-averse and shy away from trying competitors' products. While other companies such as AMD approach or exceed Nvidia's ability to design performant silicon in terms of FLOPs or total HBM, it is the CUDA software stack that has arguably most contributed to Nvidia's leading position. Taken together, these factors have led to a concentration of model training runs, particularly at the frontier, on Nvidia's chips and software.
There are many other areas of the supply chain that also feature dominant suppliers and potential bottlenecks. HBM in particular has attracted increased attention (Freed et al. 2026), particularly given the significant role that South Korea’s SK hynix has in its development. However, like others, the firm is offshoring some of its capabilities to the U.S. (SK hynix 2026).
Due to the specialized, international, and technologically advanced nature of the AI chip supply chain, no state currently has a fully indigenous supply chain, and it would be profoundly difficult for any state to create one—especially without other states taking notice. These facts correspondingly bode well for the prospects for governing chip manufacture.
5.3 Feasibility of governing chip manufacturing
This section discusses our assessment that measures like those discussed in § 3.3.4 would be sufficient to govern chip manufacturing. Governing chip manufacturing requires achieving two sub-goals: first, maintaining knowledge of existing facilities capable of manufacturing accelerator chips, discussed in § 5.3.1 below; and second, ensuring that the capacity for these facilities to manufacture illicit chips is measured and bounded, discussed in § 5.3.2.
5.3.1 Feasibility of detecting undeclared facilities
In detecting undeclared lithography machines, governing authorities would enjoy distinct advantages from the fact that EUV and DUV lithography machines, being among the most advanced technologies in the world, currently require regular service from vendors (Baazil et al. 2024; ASML n.d.a). They also have distinguishing industrial features: the clean rooms for such machines require continuously powered HVAC (Faulkner et al. 1996), emit detectable gas byproducts (e.g. Arnold et al. 2018), and have specialized features to protect against disturbances from floor vibrations (Howard and Hansen 2003). An undisclosed Huawei facility in Shenzhen was discovered using open-source satellite imagery by journalists (Olcott et al. 2025).
The Dutch company ASML is the only vendor for EUV machines and, until 2026, ASML and Nikon were the only two for immersion DUV machines (Chiang 2025). A new market entrant, Yuliangsheng in China, reportedly expects to produce 12 DUV machines at the 28nm node51 by the end of 2026 (Olcott and Wu 2026). These machines are considerably less precise than EUV machines, the most advanced of which are at the 2nm node (ASML n.d.b; ASML n.d.), but they could potentially fabricate some chips at the more advanced 7nm node (albeit at lower yields) through an advanced technique called multipatterning (James 2023). Nevertheless, as these machines require lenses from the German high-precision optics company Zeiss, they cannot be operated covertly without detection, provided that ASML and Zeiss ensure no such lenses from ASML machines go unaccounted for. Any failed lens should be replaced by the vendor and returned to its nation of origin.
For the time being, and in particular for as long as neither American nor Chinese industrial actors can replicate the high-precision optics that Zeiss provides, the lithography machine census should remain a sufficient step to prevent the undeclared manufacture of accelerators in either state, but it would cease being sufficient if either state achieves the full indigenization of its supply chain. For that reason, we would recommend to pause-willing leaders of the U.S. and China that they not pursue full indigenization, since succeeding in that pursuit could destabilize the treaty.
Even if one state did fully indigenize its supply chain, it would not follow that that state could maintain a covert supply chain with plausible deniability. Even if efforts to conceal manufacturing facilities succeed, covertly constructing an advanced lithography machine with high-precision optics and a manufacturing facility to house it would only be the beginning of a state’s difficulties in maintaining a covert construction line. They would also need to covertly replicate numerous other highly specialized industries such as photoresist chemical production, HBM, and excimer laser parts, creating many ongoing opportunities for discovery by all-source intelligence.
5.3.2 Feasibility of monitoring declared manufacturing facilities
In addition to detecting undeclared manufacturing sites, the treaty authorities would need to monitor declared sites for undeclared production. This appears to be feasible even without frequent inspections, provided monitors simply verify how many wafers a lithography machine processes (for example, by receiving an attested log recording each one processed and verifying it against wafer counters and laser pulse counts) and account for each chip printed on each wafer, along with excess silicon. One printed wafer produces a known number of chips of a given type, on an equally spaced grid, and the remaining wafer area consists of a predictable quantity of excess silicon.
If all successfully printed chips and rejects are shown to inspectors, with the former statistically sampled for inspection and the latter destroyed under monitoring, it would be exceedingly difficult to divert more than a very small fraction of production to unlicensed chips without a substantial probability of detection. Conservatively, if treaty authorities sampled and inspected 2400 chips in each production line per year, then even if there were only a one in four chance of their inspection recognizing a chip as an illicit chip, a diversion of 1% of production in that line would be recognized within six months with 95% probability.52
5.4 Feasibility of a global accelerator census
A critical determinant of a hardwired pause’s durability would be the success of a global accelerator census in locating, accounting for, and registering as many of the world’s pre-pause chips as possible. As we discuss in § 6.5, the binding constraint on a treaty’s durability would likely be its degree of success in comprehensively completing this census.
This section discusses why a global accelerator census, or global counting of accelerator chips, conducted with the cooperation of the U.S., China, and a small number of additional jurisdictions key to international semiconductor trade in Asia would have good prospects for success. Such a census would also be an important step for many other forms of global AI governance, e.g. other plans for an international pause, as dark compute is much harder to govern than registered hardware.
Many of the records needed for a successful census are internal to governments. If either state refused to cooperate with the census, it is unlikely the other could force it to make any meaningful attempt to locate pre-pause hardware, much less to register it with the regime. These factors mean that voluntary participation of both the U.S. and China would likely be essential, as well as participation of other states such as Singapore and Malaysia, where many of the companies that supplied China with diverted stock reside.
§ 5.4.1 discusses the copious documentary evidence that could be subpoenaed by a participating state to locate pre-pause accelerators, which are capital goods held mostly by large companies that had strong tax incentives to document their acquisitions at the time of purchase. § 5.4.2 discusses why, given this quantity of evidence, the design in § 3 would give pause-willing states incentives to participate, rather than abandon the pretense of participating.
5.4.1 The paper trails left by pre-pause hardware
One important cause for optimism is the amount of documentary evidence that pre-pause accelerators would have already produced, before the pause date, in the states where they had been operating. In the U.S., a single Hopper H100 accelerator is worth tens of thousands of dollars (USD) (CDW-G n.d.), and a server with eight Blackwell B200 chips is worth several hundred thousand USD (Sheumaker 2025, slide 41); the same hardware is even more expensive in China, where it has at times traded at two to three times the U.S. price, as a result of scarcity created by U.S. export controls (Pan et al. 2026; Olcott 2026a). The vast majority of these expensive capital goods are owned by large-scale businesses with intact custody records and professional accounting divisions, as well as IT departments that continuously monitor the chips.
Most accelerators would be voluntarily declared by their owners, for a variety of reasons: failing to do so would be illegal and ownership is concentrated among a few large businesses (see § 5.5.1); registering chips for transitional inference service would likely be the most lucrative use for them; and most owners would already have declared their accelerators as a tax writeoff at the time of purchase, making it unlikely that they could get away with evasion. Below, we discuss some of the many sources of documentary evidence that a chip would leave during its lifetime.
Chip owners’ tax deduction records: At present, when a business in the U.S. or China buys accelerators, it has strong financial incentives to submit a record of that purchase to the government. In the U.S., the tax deduction a business can claim for capital expenditures is worth over 21% of the purchase price (26 U.S.C. § 11(b)), while in China, a credit from a value-added tax (VAT) invoice from the government, recording the buyer, seller, and item, is worth 13% of the price, along with an additional 15–25% from an enterprise tax deduction (Standing Committee of the National People’s Congress 2018, Arts. 4 and 28).
Manufacturing records: Chip foundries’ sales records would show how many chips they manufacture by designer and type. These records can be audited against supplier records of inputs like HBM, advanced packaging, and silicon wafers.
Vendor sales records: Vendors such as Nvidia keep sales records of buyers and units sold (NVIDIA 2026a). If sold to a distributor, the distributors’ records would likewise record the next owner—with the important exception of those units sold to shell companies for diversion around U.S. export controls; for those chips, the sales history alone would not trace hardware to current owners unless records were furnished by the shell company in exchange for amnesty.
Resale records: The small percentage of units sold on the secondary market would often include resale records, though they would likely not be as reliable as vendor sales records.
Customs records: Import and export records would exist for most of the stock, with information such as duties paid and the type of goods imported. Public records exist in aggregate forms in both the U.S. and China, and are the basis for public estimates such as that of Epoch AI (Juniewicz 2026c). While the public records are not fully specific as to the exact type of goods transported, government data at the level of individual payments name the buyer, model, and price (General Administration of Customs of the People’s Republic of China 2020).
Inventory and rental records of data center operators: Most companies that own accelerators host them at data centers, including colocation centers like QTS (Epoch AI 2026b). Colocation data centers’ rental records would in turn record which companies’ chips they were hosting, how many, and over what period of time.
Public tenders: Public sector buyers commonly file public tenders soliciting bids for accelerator purchases. For example, the government of Binhai in Jiangsu, China issued a tender for the purchase of 48 Hopper H200 servers in July 2025 (Baptista 2025).
Power bills and power substations: Power companies’ meter readings shed independent light on the quantity of hardware operated at a given data center.
Diversion re-sellers’ sales records: In addition to discovering accelerators themselves, in a search for documentary evidence the census would also be likely to uncover traders and companies that participated in the accelerator diversion market during the U.S. export control era. In exchange for amnesty, these companies could turn over their sales records, allowing for the chain of custody to be traced even for many diverted accelerators. In addition, sales records of shell companies who were already prosecuted would help to reconstruct the chains of custody of some diverted accelerators (e.g. United States v. Hao Global LLC 2025).
Records of purchase financing: Sellers of expensive accelerators sometimes require bank guarantees from buyers as a condition of purchase, thereby involving a third party in the transaction. For example, 96 of 206 offers on a Chinese broker board in October 2026 required a bank guarantee, letter of credit, bank draft, or escrow for purchase of servers (Jujianbao 2026).
Records of collateralization for loans: In the recent AI data center buildout, accelerator hardware has sometimes served as a natural good offered as collateral for debt financing of construction. This collateralization commonly specifies the servers by serial numbers (e.g. CoreWeave 2025; IREN Limited 2026; Nscale 2026).
Warranty return records: Accelerators are sold in servers that typically remain under warranty for three years (NVIDIA 2026b; Dell Technologies 2026; Hewlett Packard Enterprise 2026). Those that fail under warranty can be returned to vendors for -waste disposal, in exchange for new accelerators.
Decommissioning records: Chips that fail after lapse of warranty are required to be disposed of by licensed recyclers. -waste and hazardous waste rules regulate how they are disposed of in China, the EU, and the U.S. (European Union 2012; State Council of the People’s Republic of China 2009; Ministry of Ecology and Environment 2026; U.S. Environmental Protection Agency 2026).
Most units remain with their first owner; over 95% of the world stock by computing power is less than three years old (Epoch AI 2026c, 2026d), shorter than enterprise chips are usually deployed (see Footnote 67), and demand for compute has been quite high for their entire lifespan. As a result, their chains of custody over their entire lifespans would be relatively simple for the census to track through sales records, with multiple data sources giving additional entry points to discover those for which a sale was for some reason not recorded.
The documentary evidence above could be combined with physical inspections of known data centers, retrospective examination of archival satellite records to discover unknown data centers based on features like cooling equipment, backup generators, and substations (Epoch AI 2026a), bounties for anonymous tips from clients or employees of companies that fail to declare, and/or chip buybacks.
Much of the documentary evidence discussed above would be available for accelerators in China that were diverted around U.S. export controls—which were often even more expensive than in the U.S. (Huang 2024; Wu and Olcott 2025). Because China has never regarded itself as bound by U.S. export controls, buying and selling AI chips and servers was legal in China provided relevant border tariffs were paid (Wu and Olcott 2025), though Chinese customs restricted importation in 2026 (Olcott 2026a). One trader reportedly told the Financial Times, “You can go against the American rules, but not the Chinese” (Olcott 2026a). The Chinese examples discussed above, including the July 2025 Binhai public tender, the October 2026 brokers board, and the public customs records all were recorded during the export control era, demonstrating that most of the diverted stock has never been invisible to the Chinese government.
The various sources of evidence listed in this section would provide multiple lines of defense against a thorough census missing pre-pause chips. The likelihood of achieving high coverage would likely leave little room for persuading other states that missing chips were not being concealed. Next, we discuss the incentives of the U.S. and of China in light of this quantity of evidence, and the design of the example implementation discussed in § 3.
5.4.2 Incentives of states and private actors
An important determinant of success in the census is for the incentives of both states and private actors to reward compliance. The transitional inference service during Phases 1-2 of the treaty (see Figure 1.1) would likely be highly profitable, and states parties could supplement that incentive with additional buy-back incentives, raising further the economic opportunity cost of noncompliance. It would almost certainly be within the collective fiscal capacity of states parties to buy back all existing accelerators even at several times their purchase price. Epoch’s chip-sales dataset (Epoch AI 2026d) includes roughly 25 million AI accelerators, with an average estimated purchase price (weighted by shipment quantities for each type of accelerator) of roughly $20,000, implying an estimated total acquisition cost of $500B. Even if we double that to $1T for an illustrative buyback budget, allowing for additional chips not in Epoch’s dataset and higher prices, this is less than a fourth of the roughly $4.5T that the U.S. government alone spent on Covid relief (U.S. Government Accountability Office 2025).
The evidence in § 5.4.1, as well as the private incentives above, suggest that a thorough global census by a state not looking to conceal a stockpile could be expected to trace the vast majority of dark compute in a state. If this were empirically borne out in most states, it would place additional pressure on states missing larger fractions of their pre-pause stock, by making large gaps in the census appear more incriminating.
The global accelerator census could be expected to trace nearly all dark compute, at least from the time of manufacture to the first owner. As discussed in § 3.3.2, gaps in the census traceable to the territories of specific states would be the primary input to estimating each state’s dark compute for their covert evasion and overt breakout ledgers. The more confident states could be that nearly all private actors’ commercial incentives would point toward declaration, and the less cooperative the state had been, the more incriminating such gaps would be.
As a result, any dark compute retained by a state would likely inflate the state’s covert evasion ledger and limit the issuance of licenses for other purposes, except to the extent that that state could credibly deny possession of it. As a result, states would be highly incentivized to prevent significant accumulation of dark compute by private entities within their jurisdictions. Under the implementation discussed in § 3, they would also have more attractive alternatives if, for example, they desired stockpiles to safeguard against defection by rivals (which they could accomplish with strategic reserves), or they needed to retain a small number of AI chips for confidential national security workloads (which they could accomplish with sovereign fleets).
Compared to these alternatives, concealing a hoard of dark compute would be more costly for a state in several respects: First, it would cost the state goodwill, if concealing it required refusing reasonable requests to furnish records. Second, it would cost them part of their ability to police their own jurisdictions, if concealment required conducting a less rigorous census; and third, it would cost them the option to request more issuance of licenses for other purposes, if other states estimated the size of the concealed stockpile and added it to the state’s covert ledger. Thus, a sufficiently successful census could make pursuit of such a stockpile more difficult and less attractive than alternative outlets for pursuing national security objectives, such as declared sovereign fleets and strategic reserves.
Finally, to the extent that states deliberately retain some dark compute and suspect others have as well, they could still attempt to cooperate to further reduce these dark compute stockpiles, and would have good reason to attempt to do so if they genuinely cared about the success of the treaty. For instance, if they are willing to abandon the pretense of having fully cooperated so far, they could agree to produce some number of chips conditional on others doing the same. If they wished to maintain the appearance of cooperation, they could instead agree to make renewed efforts to find more chips and announce that they had done so in small increments so long as others appear to be doing the same.
This process would be smoother if states parties all agreed on how many more chips ought to be recoverable from each party. The next section discusses how much stock in key states would be missed by a thorough census conducted in both the U.S. and China; however, it does not imply a prediction that the Phase 1 census necessarily would be thorough.
5.5 Estimating how many chips a census would miss
In this section, we quantitatively estimate how much pre-pause training-capable hardware in each of the U.S. and China would be missed by a thorough census. In § 5.5.1, we distinguish chips as to their visibility to the state and to the general public by classifying them into categories that we call holding types, and estimating the total stock , for holding type and state , focusing on the U.S. and China as key states. In § 5.5.2, we estimate the missing fraction , i.e., the fraction of functional chips within each category that might go missing, as well as an estimate of the number of failed chips that might go missing, which depends on the age distribution as well as the holding type. We combine these estimates at the end into an estimate of the total number of chips that would go missing from each state. We arrive at a final estimate of the missing stock in state as
53
Table 5.1 illustrates the calculation, giving conservative estimates of 11k total H100e in the U.S. and 11k H100e in China. The remainder of this section explains how we arrived at the numbers in the table by combining publicly available data, empirical anchors, and judgmental estimates.
| Holding type | Stock Q(h, s) | Failed chips | of which missing | Functional chips | Missing fraction μ(h) | Functional chips missing | Total missing |
|---|---|---|---|---|---|---|---|
| United States | |||||||
| Publicly identifiable holdings | 33M | 140k | 3.6k | 33M | — | 0 | 3.6k |
| On-books business and institutional holdings | 1.9M | 13k | 450 | 1.9M | 0.19% | 3.7k | 4.1k |
| Off-books business and institutional holdings | — | — | — | — | — | — | — |
| Informal holdings | 33k | 130 | 130 | 33k | 10% | 3.3k | 3.4k |
| Total, United States | 35M | 160k | 4.2k | 35M | 7.0k | 11k | |
| China | |||||||
| Publicly identifiable holdings | 780k | 5.6k | 200 | 770k | — | 0 | 200 |
| On-books business and institutional holdings | 2.2M | 15k | 940 | 2.2M | 0.19% | 4.2k | 5.2k |
| Off-books business and institutional holdings | 99k | 660 | 330 | 99k | 1.6% | 1.6k | 1.9k |
| Informal holdings | 37k | 160 | 160 | 37k | 10% | 3.7k | 3.9k |
| Total, China | 3.1M | 22k | 1.6k | 3.1M | 9.5k | 11k | |
5.5.1 Estimating the quantity of pre-pause hardware
AI chips differ as to their visibility to governments and the general public. Many AI chips are located in known data centers like OpenAI’s Stargate Abilene, whose size is visible from satellite and whose chip counts are known from public reports (e.g. Crusoe 2025), or are owned by a handful of large companies. A smaller number are held by businesses off-books, by individuals, or by shell companies engaged in diversion around U.S. export controls. While they are smaller in number, these would also be more likely to go missing. It is not known from publicly available data how many chips are in each category; where data is unavailable, we give judgmental estimates, seeking to err on the conservative side by overestimating the number of missing chips (this being conservative because a lower estimate of the number of missing chips would make the treaty more durable).
In this section, we distinguish four categories that we call holding types,54 in decreasing order of visibility:
Publicly identifiable holdings: Chips held by either the small number of chip owners identified by name in the Epoch AI chip owners data (Epoch AI 2026c), or in holdings identifiable from public reports.
On-books business and institutional holdings: Chips not in the first category that are nevertheless held on-books by businesses and/or public institutions. This category includes chips imported and sold in China after diversion around U.S. export controls, but with some state record such as a tax invoice or customs entry identifying the holder.
Off-books business and institutional holdings: Chips in China sold off-books to avoid U.S. export controls, which did not re-enter state records but are deployed by an operating business or institution.
Informal holdings: Chips held by individuals, as well as shell companies and traders engaged in diversion.
We only estimate the third category in China because of the large quantity of diverted stock located there. The holdings in the U.S. that are off-books for diversion purposes are held by shell companies and traders, but are not likely to be actively deployed since they are bound elsewhere. For each holding type and state , we estimate the quantity of pre-pause stock, measured in H100e units, at year-end 2026.55
For publicly identifiable holdings, we combine Epoch AI’s chip owners data (Epoch AI 2026c) with other public records on the internet to identify publicly reported holdings.56 According to Epoch AI’s data, a significant majority, 78% measured in H100e as of year-end 2025, are owned by seven large American companies,57 with more owned in publicly identifiable holdings by neoclouds, the U.S. Federal Government, large companies like Tesla and Apple, universities, and a longer tail of smaller companies and private owners. A significant proportion of the chips in China are identified in public vendor reports about buyers of their chips.58 We arrive at estimates of 33M H100e in the U.S., and 780k H100e in China in this category; the true number would likely be more if our search had been more exhaustive.
Of the projected residual, roughly 1.9M H100e in the U.S. and 2.3M in China, we estimate the remaining three categories in three steps, conservatively allocating more to the less visible categories, as follows:
We assume 1% by H100e are held by individuals. This is likely an overestimate for accelerators, making the estimate conservative.
In China, we assume that 10% of the diverted stock by H100e is off-books. We believe this is also a conservative estimate, given the convergent evidence discussed in § 5.4.1 that diverted accelerators are not treated as contraband in China, and because businesses had strong incentives to file VAT invoices.
We allocate what remains to on-books business and institutional holdings.
Based on investigative reporting and court filings from prosecutions, Epoch AI estimates 660k H100e diverted over 2024-25 (Juniewicz 2026a). We use their estimate and extrapolate it for a third year,59 to arrive at 990k H100e.
Finally, we add one month’s inventory at the estimated diversion rate (330k per year, based on Epoch AI’s estimate of 660k H100e diverted during 2024 and 2025). We count half toward each state’s informal holdings. The resulting estimates are shown in Table 5.1.
5.5.2 Estimating the missing fraction by holding type
This section estimates what fraction of each holding type would be missed by a thorough global census. As the census is hypothetical, no data is available on these quantities, but we can use empirical anchors from analogous cases to arrive at conservative overestimates.
We consider two distinct reasons why chips might be missed: they might have failed and been disposed of without a verifiable record, or they might be functional but retained by their owners.
Missing failed chips
For failed chips, we additionally distinguish between chips that failed under warranty and would have been returned to vendors for a refund, and chips that failed outside of warranty, which might not have been disposed of properly. Cui et al. (2025) observe a physical replacement rate of 0.45% per year for new hardware (range 0.16%–0.97%), but the failure rate would likely increase as chips age; we use the central estimate but assume a three percentage point increase for each year after year 3, applying it to the estimated age distribution of hardware in different categories.60
We assume the following about the rate at which failed chips go missing:
Of chips that failed under warranty, among the identifiable and on-books business and institutional owners, 99% by H100e were returned to the vendor or have another contemporaneous record of their failure, such as a certificate for proper -waste disposal or an inventory record, leaving 1% unaccounted for.
Of chips that failed outside of warranty, in the same two categories,61 90% by H100e have a contemporaneous record, leaving 10% unaccounted for.
Of chips that failed in the off-books business and institution holdings category, we assume that 50% have a contemporaneous record, leaving 50% unaccounted for.
Of chips that failed in the informal holdings category, we assume that none have a contemporaneous record.
These predictions seem to us conservative (i.e., overestimates), given the strong incentives for businesses to claim warranty refunds, the likelihood of their complying with environmental regulations, and the continuous measurement of chips in many data centers. However, we acknowledge inherent uncertainty about the real values. Altogether, we obtain an estimate of 4.2k H100e of missing failed units in the U.S., and 1.6k H100e of missing failed units in China.
Functional chips not declared
The global census would require chip owners to declare their chips; chip owners could also profit from doing so, as discussed in § 5.4.2, but some might seek to retain chips for private use or for use on the black market. We assume different rates of compliance with this mandate for the different holding types, based on factors discussed below.
For chips to be missed from one of these sources, at least three things would need to happen:
The chip would need to be physically operated by its owner at a site other than at a data center, or it would have to be missed by the data center’s own declaration and the state’s site visit,
It would need to go undeclared by the owner, despite the owner’s past and present financial incentives to declare it and the risk of punishment from failure to declare, and
Government attempts to trace it from the seller would need to fail.
For publicly identifiable holdings, which are held by large businesses and public institutions, we assume negligible noncompliance. We estimate the rates at which each of these would occur as follows, applying empirical anchors based on estimates of analogous quantities, again attempting to err on the conservative (high) side:
We estimate that about 10% would not be hosted at a data center, based on a 2024 Lawrence Berkeley National Laboratory report placing under 10% of U.S. servers in closets, server rooms, edge sites, and small businesses’ own premises (Shehabi et al. 2024). This is likely an overestimate since AI chips are more likely to be hosted in data centers than other enterprise computers; the same report makes an assumption that all AI accelerators are in data centers. Of the 90% at data centers, we assume no more than an additional 6% would go undeclared by both the data center operator and the owner, applying the IRS empirical anchor below, giving 16% total.
We estimate that, among those chips not findable through data centers, no more than 6% would go undeclared, based on the fraction of income that goes unreported to the IRS when there is “substantial information reporting” (Internal Revenue Service 2024). This is likely an overestimate since underreporting income is financially profitable, whereas underreporting capital expenditures would be financially counterproductive.
We estimate that, among on-books chips not findable through data centers or declared, less than 20% would go unrecovered by means of tracing the chain of custody from the seller or importer, based on the fraction of firearms that the Bureau of Alcohol, Tobacco and Firearms fails to trace based on retail records (Bureau of Alcohol, Tobacco, Firearms and Explosives 2023). This is likely an overestimate because retail sales of accelerator chips are capital-intensive business purchases that create additional financial and operating records as described in § 5.4.1.
Multiplying these numbers gives a conservative bound of (On-books) , among chips that had not failed.
Off-books business holdings in China could plausibly see higher fractions go missing. A larger fraction might go undeclared because owners who preferred to retain their chips for black market use could do so more easily, so we inflate our 6% IRS anchor to 20%. It would also be more difficult for the state to recover those that went undeclared, so we inflate our 20% ATF anchor to 50%. Then we arrive at (Off-books) , among chips that had not failed.
Finally, we assume that informal holders would be similar to off-books business holdings, except with none hosted in data centers, giving (Informal) . Combining these estimates gives the numbers reported in Table 5.1.
Combining the above estimates of the missing fraction with our public records of named holders,62 we obtain estimates of missing but functional hardware amounting to 7k H100e for the U.S. and 9.5k H100e for China. These can be found in Table 5.1.
6 Feasibility of verification
6.1 Summary assessment
This section assesses the feasibility of verifying a hardwired pause, assuming it is feasible to produce suitable inference-only chips as discussed in § 4, and to successfully govern pre-pause and newly manufactured AI chips as discussed in § 5. Our assessment is that an implementation similar to that described in § 3 would likely be technically feasible for the immediately foreseeable future, if the governments of the U.S., China, and the supplier and chip owner states listed in § 3.1 implemented the governance and verification measures discussed in § 3.3.
Our assessment does not rest on an assumption that verification on deployed hardware would be highly reliable in the short term. Instead, it rests on quantitative assessments that:
Neither China nor the U.S. currently has sufficient total training capacity to carry out an overt breakout program, under plausible assumptions about what such a program would require; and
Neither China nor the U.S. currently has sufficient training capacity to carry out a strategically significant covert evasion program, under plausible assumptions about the fraction of training capacity that could be marshaled toward such a program.
On current trajectories, this state of affairs appears unlikely to change within the next year, barring a major breakthrough in training efficiency gains.
Our assessment is based in part on estimates of how far a state-level defector would need to advance the frontier in order to achieve the degree of strategic advantage described in § 3.1.1.
During Phase 1, while transitional inference service and the global accelerator census were ongoing, the training-capable accelerators used for transitional inference service in the U.S. would likely come the closest to posing a breakout risk, and uncertainty about the quantity of diverted hardware in China would likely come the closest to posing a covert evasion risk.
In § 6.2, we quantitatively estimate possible covert evasion and overt breakout fleet thresholds (as in § 3.2.3) for a near-term pause scenario where a pause takes effect near the end of 2026, and we bound what the covert evasion and overt breakout capacities of the most significant sources of training capacity might be, for each of the U.S. and China. We also give more tentative assessments for an alternative scenario taking effect during the summer of 2027.
For both scenarios (near-term pause and mid-2027 pause), we assess that neither state would likely pose a meaningful covert evasion or overt breakout risk at the time of the pause, but with greater uncertainty about the second scenario. Our choice to assess near-term scenarios should not be taken as implying that we predict a pause agreement is imminent, or that a later pause would be infeasible. It is rather a recognition of how rapidly the world is changing and how different many inputs to the analysis could look even in a year’s time, if current progress in AI continues.
In both scenarios, two sources of training capacity stand out as the most concerning: the overt breakout capacity from transitional inference service (especially in the U.S.); and the covert evasion capacity from dark compute (especially in China). The most pressing risk in the near-term and mid-2027 pause scenarios would be that fleet thresholds could erode so much from training efficiency gains that the allowances would approach these two sources of training capacity from above, and eventually cross below them, rendering the treaty unenforceable unless they could be reduced.
In the first several years of a pause, the breakout risk from transitional inference service is more likely to become a pressing challenge. As discussed in § 3.1, a near-term pause would most likely be initially implemented as an informal agreement, with treaty terms negotiated over a period of years. During those years, more information about the progress of inference-only chip development and the global accelerator census would come into view. We discuss ways that states parties to an eventual treaty might address these challenges, including by permitting treaty allowances to begin high and shrink according to schedules, rather than by setting them low on the first day when the treaty enters into force.
In the longer term, we assess that the main threat to durability would most likely be dark compute. If states could not eventually coordinate to conduct a thorough census of pre-pause accelerators, then one state’s estimated covert evasion capacity from dark compute might eventually pose a serious covert evasion risk. The durability of the treaty would then depend on its ability to fall back on detection of training at the time of chip deployment, based on national technical means or additional verification measures.
Excluding transitional inference service and dark compute, we assess it likely that both states’ total estimated covert and overt ledgers would initially fall below 0.25% of the fleet sizes for covert evasion or overt breakout. These components could likely be reduced further after implementing more advanced verification measures by the end of the decade, but if such research disappointed, they could alternatively be reduced by reducing the number of licenses granted for such uses. By contrast, dark compute could be more difficult to reduce below the floor set by the success of a thorough global census.
As a result, if the twin issues of dark compute and transitional inference service could be reduced below 0.25% of and , respectively, then we assess that a treaty with fractional allowances of could be sustained with two orders of magnitude of headroom: that is, the treaty could sustain two orders of magnitude of training efficiency gains before it would struggle to balance the ledgers.
Because the census would be difficult without the cooperation of local jurisdictions, we assess that a hardwired pause would be significantly more durable if it were pursued in a spirit of international collaboration between the world’s two leading powers and other states, and significantly less durable if it were imposed by one power against the other’s will, or by both powers attempting to coerce other states.
6.2 Fleet thresholds for defection
We now undertake a quantitative analysis of the fleet size, measured in H100e, required to seriously undermine the treaty, which varies across the three threat models enumerated in § 3.1.1:
Private evasion. A private actor covertly trains a model in excess of the legal compute ceiling ;
State-level covert evasion. A state achieves a significant strategic advantage by covertly using a secret model with training budget in excess of a covert evasion threshold ;
State-level overt breakout. A state achieves a decisive strategic advantage by overtly breaking out to build a strategically decisive model in excess of a breakout threshold .
Violations of any of these three thresholds could undermine the treaty, but not to equal degrees. We analyze the three models separately below and discuss how states parties might arrive at estimates of key governance parameters. For the two state-level verification threat models, we suggest simple methods for translating into a covert evasion fleet threshold and overt breakout fleet threshold , to determine bounds on the corresponding ledgers for each state.
In this section, we present a quantitative analysis with symbolic values representing derived or negotiated quantities, presenting empirical arguments for concrete numerical values that might be negotiated in a pause that takes place at year-end 2026; we refer to this as our near-term pause scenario. For this scenario, we estimate initial covert evasion fleet threshold and overt breakout fleet threshold for such a pause. After a pause were implemented, both could erode over time due to algorithmic efficiency.
Figure 6.1 shows how these two initial fleet thresholds would change if we vary the pause date.63 It compares these estimates to data and derived estimates of the total world accelerator stock, including the estimated stock that is located in each of the U.S. and China, extrapolated as functions of the date, representing the values they might take at the pause date, depending on when it occurred. We label these total stocks , and ; projections of these are printed for year-end 2026 in Table 6.1.
As Figure 6.1 shows, our methodologies for determining these two fleet thresholds suggest that will continue to grow, because we operationalize a significant strategic advantage in § 6.2.2 as an advantage relative to the frozen frontier, while may not grow, because we operationalize (a defector’s prospect of) a decisive strategic advantage in § 6.2.3 as occurring at a fixed capability level. This suggests that guarding against covert evasion could be somewhat easier if a pause is implemented later, because models are currently scaling faster than hardware is being built out, creating more temptation for would-be defectors. By contrast, guarding against overt breakout could get more difficult, since every year of scaling brings us closer to any fixed capability level. Because both operationalizations are choices based on the way some contemporary actors seem to view the advantages of AI, both could change if decision-makers negotiating the treaty disagree with these operationalizations, or if they hold different beliefs when the treaty is being negotiated than they hold now.
| Quantity | Symbol | Est. (H100e), year-end 2026 | Range shown |
| World accelerator stock | 47M | 45–47M | |
| Located in the United States | 35M | 26–44M | |
| Located in China | 3.1M | 2.6–4.5M |
For the private evasion threat model, we take a different approach, estimating a private evasion fleet threshold defined as the aggregate global fleet (distributed across multiple private actors) that would need to be dedicated to undermine the treaty through widespread, undetected private evasion.
6.2.1 Method for determining governance thresholds
This section discusses how a treaty might determine operational governance thresholds, by reference to the current scaling trajectory. Fleet thresholds in § 3.2.3 are defined with reference to compute thresholds, research overhead factors, time horizons, and current frontier-level utilization. We discuss each of these below.
One simple method for determining treaty compute thresholds is to index the compute ceiling , the covert evasion threshold , and the overt breakout threshold to the frontier threshold , as fixed multiples of . We adopt this method and take , where is a generic subscript representing quantities relevant to the treaty compute ceiling , covert evasion , or overt breakout . Then, per the discussion in § 3.2.2, the covert evasion and overt breakout fleet thresholds are also indexed to the frozen frontier as
From Epoch AI frontier training run data (Epoch AI 2026a), as well as an independent corroboration based on reports of the size of the GPT-6 Astra pretraining run (Lardinois 2026), we estimate a frontier level of roughly at year-end 2026.64 Reproducing Epoch’s frontier plot gives an estimated 4.5x increment per year, corresponding roughly to one order of magnitude increase in frontier training budgets every year and a half (Sevilla and Roldán 2024). If we define the (recent) historical growth rate to be the rate of growth estimated from this regression, a race trajectory corresponding to roughly one order of magnitude growth occurs every 18 months.
The research overhead factors represent the total compute budgets dedicated to the research and development process, including a wide variety of preparatory tasks that require training-capable hardware before the training run can begin, such as exploratory experimentation, synthetic data generation, inference using intermediate models that are more capable than the frozen frontier, and test runs. Denain and Wu (2026) estimate from spending data that frontier AI companies spend only about 9.6% of their research compute budget on these final training runs. If we adjust for a four-month lag between the research and training of a given model, we arrive at an empirical estimate for present-day incremental frontier training. The research overhead was mildly lower for fast followers Z.ai and MiniMax, who spent 12.3% and 12.2% respectively on final training runs; correspondingly, we use the research overhead for trailing models trained by private evaders.
A frontier-advancing research program aiming at , for a large multiple , would either need to aim for a significantly larger “blind leap” in abilities than frontier companies do now, or it would need to proceed in several steps, bridging the gap between and by degrees. If it takes the latter approach, as we find more likely, we define the blind leap increment as the optimal size of a blind leap, in orders of magnitude, relative to the best trained model available to the research program. That is, if a research program aims for , with blind leap increment , the research program proceeds in three steps, first training a model of size , then a model of size , and finally a model of size . More generally, the number of required steps to attain a multiple of the frontier is
The blind-leap increment has implications for the research overhead; if each of the incremental subprograms have present-day frontier overhead , then attaining a model with training budget requires a total research and development budget of
The last factor is greater than one unless (that is, ), in which case it reduces to , giving the same overhead as present-day incremental frontier research, . More generally, the research overhead for the entire program is the above divided by . If is a whole number, then
We adopt the blind leap increment , corresponding to roughly six months of progress on the historical race trajectory. For non-integer values of , we simply plug in the fractional value of , so as not to take the discrete nature of our model too literally.
Finally, we must estimate the utilization rate of a present-day frontier training run. Utilization for frontier training is estimated at 30%–50% for pretraining (Epoch AI n.d.a) and 5%–20% for reinforcement learning (RL) post-training (Denain and Wu 2026). Epoch estimates that GPT-5, trained in 2025, used 60% pretraining and 40% post-training, by FLOPs (Epoch AI 2025b); taking the harmonic mean of mid-range utilization estimates 40% and 12.5% for pre- and post-training would then give . Present-day frontier training is likely more heavily weighted toward post-training, meaning that should be somewhat lower, and fleet thresholds correspondingly higher. A real agreement should use more current numbers, but in lieu of a current estimate, we adopt the conservative value for all calculations below.
6.2.2 Covert evasion fleet threshold
To arrive at a covert evasion fleet threshold, we must determine the covert evasion target , the research overhead , and the covert evasion time horizon . We discuss each of these below.
One reason might be fairly large is that models trained and deployed covertly would suffer an inference-time handicap relative to whitelisted models, requiring covert models to represent a significant advance beyond the frozen frontier. The latter could be run on specialized hardware giving them an advantage in cost, scale, and speed. Covert models could not be deployed at a comparable scale, and even small-scale covert deployment would carry an ever-present risk of discovery and enforcement for users and inference providers alike. Both the cost and the risk would be compounded, in turn, by the difficulty of maintaining secrecy in the procurement and operation of scarce training-capable hardware at scale.
The domain of application of a covert model would be limited as well, owing to the need for secrecy. A covert model could not be widely deployed commercially, eliminating many of the most important perceived benefits of an AI advantage. Directly using a covert model for cyberwarfare could leave traces that would be studied for evidence that an unknown model was being used. Such evidence might well be forthcoming, especially if the model were much more capable than the models it was competing against, or produced outputs at lower serial speed than frozen-frontier cyberattack models. Deployment for internal analysis by intelligence agencies would be less risky, but also limited.
Due to both the inference-time advantage of whitelisted models over illicit models and the limited domain in which illicit models could be employed, a strategically relevant model would need to be significantly more capable than the frozen frontier on a per-token basis, and the multiple should reflect this. We suggest as the advantage in training FLOPs that would be required to give rough parity with the frozen frontier at inference time, using a model of Ord (2025) that balances training-time versus inference-time scaling.65 This factor would represent roughly an 18-month leap in per-token capabilities, on the current race trajectory.66
Next, we turn to the research overhead covert evasion program would lack many of the advantages that present-day frontier programs enjoy. It could not offer the promise of dizzying riches and social status to attract many of the brightest and most ambitious minds in the world, nor could it import insights by hiring researchers from competing firms. It could not learn as much from deploying its preliminary models, or amortize its research effort over many concurrent research programs. If covert research were carried out on aging chips, without vendor service or firmware updates, and in data centers less well-maintained than the industrial hyperscale data centers where they are presently performed, it is likely that experiments would fail and require restarting from checkpoints more often.
On the other hand, such a program could potentially ameliorate these disadvantages by using AI agents built on frozen frontier models for tasks like writing code or grading RL rollouts, provided such use did not carry an unacceptable risk of discovery. If it were carried out more slowly than current racing programs, it might also be able to plan its experiments more judiciously, relying more on human judgment to eliminate unnecessary experiments—though the researchers exercising that judgment might be less talented than is presently the case, as mentioned above.
On balance, we see more reasons to expect covert evasion research on an incremental sub-program to be more difficult than incremental research at the present-day frontier, but in the absence of a credible way to estimate the net quantitative effect of these factors, we conservatively adopt the working value , or in the special case where the research program is carried out in one leap. For the adopted values , and , we obtain
Finally, we adopt a strategic time horizon of years for covert evasion. We expect evaders to be patient, but not infinitely so, for several reasons: First, time discounting would make longer-term projects less attractive: states have many alternative ways to pursue strategic advantages on shorter time horizons, and the longer they waited to begin their research program, the more likely political and technological conditions would be to change before they could get significant use out of their model. Second, keeping a large-scale strategic project hidden for a long period of time would carry a significant risk of discovery. Finally, if the evasion program were based on an initial stock, for example from pre-pause dark compute, then evaders might encounter difficulties maintaining a fleet for covert training, and then covert inference, past its intended retirement age without vendor service.67
For our near-term pause scenario, in which , these choices lead to an initial covert evasion training budget of , which would correspond to a mid-2028 model on the present race trajectory, and a covert evasion fleet threshold of
As points of reference, we project that OpenAI’s fleet at year-end 2026 will be 6.4M H100e, and China’s total fleet would be 3.1M H100e; see the calculations supplement.
Since the covert evasion fleet threshold scales linearly with the frontier threshold , it is expected to scale as frontier training scales: if the pause instead took effect in mid-2027, with , the same steps would yield . However, we expect this estimation paradigm will degrade the farther out in the future we attempt to project it.
While we focus on the treaty’s ability to prevent covert evaders from developing a strategically significant model, the difficulties of covert evasion would not stop after the model was trained. In order to realize a strategically relevant advantage, the evading state must operate a covert fleet capable not only of paying the one-time cost to train the model, but also of serving inference at sufficient scale to outperform the more abundantly available inference capacity for models at the frozen frontier. Inference demand would eventually create an indefinite ongoing requirement to replenish covert hardware without detection, giving treaty authorities and state intelligence agencies ongoing opportunities to catch evasion.
6.2.3 Overt breakout fleet threshold
As with covert evasion, we now turn to estimating the overt breakout training budget , the research overhead , and the overt breakout time horizon , to arrive at the required breakout fleet threshold .
The most commonly cited obstacle to international cooperation is the fear that a geopolitical rival could achieve decisive and lasting strategic hegemony by scaling AI models far beyond the current capability frontier. In some imaginings of this scenario, participants are now locked in a race to detonate an “intelligence explosion,” which would in short order lead to the arrival of a superintelligent AI model: one so powerful that it can quickly and radically reshape the global order according to the wishes of its creator—if its creator can control it (Bostrom 2014). Some contemporary actors’ statements convey the belief that this remaking of the global order could occur even if the advantage is short-lived,68 so the regime’s choice of would likely be defined by the point at which actors expected this absolute capability level might be attained, and not with reference to a relative capability advantage over the frozen frontier (but it might still track training efficiency gains by holding the ratio fixed).
Since attaining such powerful advantages would, almost by definition, require models whose capabilities far surpass contemporary human experience—or, perhaps, even human imagination—we cannot expect to empirically estimate when models will achieve such transformative capabilities. Fortunately, this would not be necessary for the treaty to estimate directly because, to assure the verifiability and durability of a treaty, the relevant question is not about the objective fact of the matter as to when (or whether) models will reach superintelligence. After all, if and when a state broke out and resumed racing to build a strategically decisive model, the treaty would likely have already failed. Instead, the relevant question is about the expectations of states parties, especially the U.S. and China as the two states parties that would most plausibly expect to succeed.
Opinions diverge widely on the question of whether or when a strategically decisive model might materialize. Since those who are the most bullish about AI progress would be the most likely to attempt breakout, conservative treaty provisions should base not on the best empirical estimate of the training budget actually required to train such a model (even if it were possible to come up with one), but on the estimate of what a bullish would-be defector might believe.
Thus, one simple way to arrive at an estimate of is to imagine by when a relatively bullish state leader might be confident a decisive model would arrive if the AI race continued at its historical pace,69 and work backwards using the Epoch training curve. For example, suppose that negotiators at the time of the pause think a bullish defector would be confident that superintelligence will arrive by the end of 2029, on the historical growth trajectory.70 Then, they might set an initial , yielding which could then be held fixed for to track training efficiency gains indexed to .
A prospective breakout strategy would likely need to rely on this type of rapid phase transition and consolidation, since other states’ more cheaply trained, fast-following AI models—including freely available open-weight models—presently trail only a few months behind the frontier. Such a strategy would be both expensive and risky for the defecting state, and would in particular require it to attain its decisive advantage before the rest of the world could react to stop it—or to resume racing against it, likely with a more complete hardware supply chain and therefore, eventually, a larger fleet. To reflect the time sensitivity of the breakout strategy, we adopt a breakout time horizon of years.
For our near-term pause scenario, in which , these choices lead to the multiple , requiring a research program in steps and research overhead of
leading in turn to an overt breakout fleet threshold of
By comparison, we project the global accelerator fleet at year-end 2026 to be 47M H100e.
The value of is sensitive to the bullish leader’s timeline; if an even more bullish leader expects a decisive model to arrive at year-end 2028, then the same method would produce near-term pause values , and . If we change the breakout time horizon from two years to one year, the previous two values would change to 350M H100e and 73M H100e, respectively.
The choice of breakout time might likewise be sensitive to the quantity of strategic reserves states are allowed, as well as expectations and possible commitments about how coalitions would respond to defection. We adopt 180M H100e as our illustrative value going forward but, in practice, we expect a realized choice of would be a matter of negotiation and trust more than one of estimation.
6.2.4 Private evasion fleet threshold
In order to successfully pause the frontier, it would be necessary for treaty authorities to prevent widespread private (non-state) training of new sub-frontier models that approach frontier scale. We operationalize this objective in the treaty as the prohibition on beyond-threshold models, i.e., models with total training FLOPs above the threshold or total parameters above .
To choose a good , authorities should ensure that it is far enough below to prevent legal models from competing with the frontier, while also ensuring it is high enough to be enforceable, so that widespread evasion does not undermine the treaty. As an illustrative value, we suggest keeping three orders of magnitude below (i.e., ) giving roughly in our near-term pause scenario, but a wider range of values could likely support a treaty depending on leaders’ judgments.
If is below and by three and four orders of magnitude, respectively, a marginal violation of would be unlikely to advance the frontier and would thus do little to affect the strategic balance undergirding the treaty. However, sufficiently widespread violations would lead to a perception of unenforceability. To maintain the rule of law, the treaty would need to combine sufficient penalties with sufficiently credible detection to deter all but a few private actors. In this section we consider whether enforcement before the fact, via the denial of training compute, would be sufficient.
For , let the annual violation rate represent the number of undetected violating models trained per year with total training cost in excess of that would significantly undermine the treaty. That is, if , then 20 undetected violations per year with models trained in excess of would undermine the treaty. Aggregating across violators, the minimum total undetected compute required would be
Applying research overhead and utilization , we arrive at the following fleet threshold, in hardware units per year of undetected evasion:
Carrying out this analysis for suggests that, at least initially, private evaders would need a great deal of collective compute even to achieve modest annual violation rates. As a simple worked example, suppose we take . That is, it would undermine the treaty if about 100 private actors per year got away with training models in excess of , or if one private actor per year trained a model in excess of , or if one per decade trained a model in excess of . To sustain this annual violation rate at any value of , the aggregate required private evasion fleet would need to achieve total arithmetic throughput of at least .
Even if private evaders could achieve similar utilization as present-day frontier training, the private evasion fleet for our near-term pause scenario would need to be , amounting to billions of dollars of installed hardware. For anything like such a large fleet to operate continuously and replenish its hardware, without detection by the combined efforts of local law enforcement, state intelligence agencies, and treaty authorities, appears unlikely if hardware production is tightly controlled and the global census is even moderately successful.
Moreover, this is likely to be an underestimate, for at least three reasons: First, private evaders would be unlikely to achieve current-frontier utilization while concealing their activities from authorities. Second, they would need a larger fleet if they wanted to perform inference on their illicit models rather than only train them. And third, our calculation treats violation as occurring at only one value of . If we instead assume, e.g., that private evaders’ utilization is half that of the frontier, that evaders spend half of their compute on inference, and that is Pareto distributed with exponent , then the required fleet threshold would be larger by a factor of . For , this gives 1.2M H100e, or tens of billions of dollars worth of hardware.
Eventually, enough training efficiency gains might erode all of these numbers to the point where authorities would need to focus on efforts to detect training and inference rather than deny training capability. Since this would not happen soon, authorities would have plenty of time to prepare.
6.2.5 Small-scale evasion
Training beyond-threshold models from scratch is not the only form of evasion the regime would need to contend with. For example, performing inference on a small-scale fine-tune of an existing beyond-threshold model would also qualify as a violation but would be harder to police through compute denial, especially during Phases 1-2 while pre-pause hardware remained in widespread use. A Chinchilla-optimal model with just over has roughly 130B parameters and could be fine-tuned using a single rack with 32 commodity H100 accelerators;71 an illegal model with just over 13B parameters could be fine-tuned more easily still. Even if difficult to enforce through training-compute denial, any fine-tune (or use of such a model if not whitelisted) would be illegal if the total compute, including the fine-tuning and the original training, exceeded the compute ceiling . Enforcement would occur at the time of deployment similar to other computer crimes, with the parameter ceiling likely easier to enforce after the fact than the compute ceiling.
During Phase 3, this conclusion might reverse: successful global governance of new hardware manufacture and tracking of registered hardware, combined with gradual attrition of pre-pause dark compute, would make training-capable accelerators increasingly difficult for violators to obtain.
Another concern might be that larger-scale RL fine-tuning of existing frontier models would be a better paradigm for training a new frontier model than starting from scratch. However, if as suggested above, then unless this RL training exceeded , it would likely account for less than 1% of the total compute budget of RL post-training that would already have been performed on the original model, assuming prevailing compute allocations between pretraining and post-training. Absent algorithmic breakthroughs, it appears unlikely that this type of violation would significantly advance the frontier without requiring an above-threshold run.
On the other hand, it is more plausible that a small fine-tune could train dangerous capabilities into an already very generally capable model, or train safeguards out of it. We note this would be a difficult enforcement problem for any governance regime, whether paused or not, but the reduction in the global supply of training-capable hardware envisioned by our Phase 3 appears to be a more promising response than alternatives that allow the supply of training-capable hardware to proliferate further.
6.3 Training efficiency gains and governance innovation
The primary risk to treaty durability would be that training efficiency gains would erode the compute threshold , and along with it the required fleet thresholds for each of the threat models discussed in § 3.1.1. This would in turn render the thresholds more difficult to police over time. Estimates of the rate and origins of training efficiency gains vary widely (Ho 2026).
6.3.1 Estimates of recent historical training efficiency gains
If we take the most extreme estimates of training efficiency gains at face value, this erosion could happen quite rapidly. One regression-based estimate by Ho et al. (2024) found that there had been annual 2.7-fold training efficiency gains, by regressing language models’ log perplexity loss on the WikiText and Penn Treebank benchmarks against model and data scale as well as the date on which each model was published. Their data was not based on experimental comparisons, but on a literature search over the period 2012-2023. If we were to extrapolate it forward, it would suggest that today’s frontier models could be trained at one 20,000th of the cost ten years hence—that it might be possible, for example, to train GPT-6 Astra on about five accelerators rather than the more than 100,000 accelerators OpenAI used in its Abilene campus (Lardinois 2026).
However, Ho et al. (2024) urged caution in interpreting their estimate in this way, noting that they did not observe the 10,000-fold gains their model implied anywhere in their data set. In particular, they cautioned that many training efficiency gains were scale-dependent, and in particular that a large portion of these gains accrued from the transformer’s scalability, which made possible the rapid scaling that took place over the same time period (Sevilla et al. 2022). That is, the observed gains would have been lower if the training runs used in 2023 were not larger than those used in 2012. This observation has important implications for a hardwired pause because, if successfully implemented, such a pause would prevent further scaling. Thus, if innovations discovered during a pause are complementary to further scaling, their greatest benefits could not be realized.
Subsequent empirical studies confirm the scale-dependence of algorithmic improvements, suggesting that the most important breakthroughs have improved AI not by wringing better performance out of 2012-scale models and data sets but by giving better returns to scale. Gundlach et al. (2025) find empirically in ablation studies that the training efficiency gains they tested have improved performance at smaller scales by a factor of only 6.3 over the same period, and combine this estimate with estimates from the machine learning literature to estimate that the total gain was closer to a factor of 60—translating to an impressive but much less overwhelming 45% improvement per year. They estimate that 91% of gains at larger scales was explained by two scale-dependent breakthroughs, the transformer architecture (Vaswani et al. 2017) and Chinchilla re-balancing (Hoffmann et al. 2022), both of which are much more effective at present-day frontier scales than older methods but confer smaller improvements at small scales. They concurred that the gains observed by Ho et al. (2024) were contingent on the rapid rate at which computing effort was scaled up over that time period. Their evidence therefore is suggestive that the observed gains in a counterfactual “paused” world, where compute budgets had remained fixed at 2012 levels, would have been much slower.
More recently, Patel and Han (2026) found annual pretraining efficiency gains of 57% over the period 2019 to 2025, through computational experiments comparing new methods and data to old methods and data, at fixed scale . They cite scale-dependence as an explanation for the difference between their estimate and the estimate of Ho et al. (2024). Interestingly, by separately varying data improvements and algorithmic improvements, they find that data improvements alone accounted for three times as much efficiency gain as algorithmic improvements.
6.3.2 Training efficiency gains after a pause
Since training efficiency gains are a product of human effort and not an exogenous force of nature, their rate could be very different in a post-treaty world than it has been over the last decade. At present, gains are driven by overwhelming commercial incentives drawing thousands of the world’s most talented engineers to work full-time on training more capable models with costly and constrained computational resources. A ban on frontier model training could decelerate training efficiency gains by redirecting commercial research and development efforts away from training new models and toward diffusing current model capabilities.
A second reason why a hardwired pause might be expected to decelerate training efficiency gains is that it would sharply limit engineers’ ability to test ideas at scale. The internal workings of frontier AI models are famously inscrutable even to their creators, motivating the common analogy that they are not “built” like cars or bridges but “grown” by empirically fiddling with weights. As a result, research on how to train models more effectively is highly empirical. The empirical nature of training efficiency gains is well-known in the AI world; as one inventor of the transformer famously quipped with regard to another improvement, “We offer no explanation as to why these architectures seem to work; we attribute their success, as all else, to divine benevolence” (Shazeer 2020).
A third reason relates to the finding by Patel and Han (2026) that data improvements have driven most recent pretraining efficiency gains. If so, then a future covert evasion project might have difficulty realizing these gains, as the internet of ten years hence might be much more hardened against automated crawling than today’s internet, by necessity, due to greater abundance of bots, potentially leaving covert evaders with stale data sets. There are some indications that this is already happening: Cloudflare has begun blocking AI crawlers on some websites for which it proxies traffic, and has continued to tighten its block as recently as July 2026 when it announced stricter defaults (Cloudflare 2025; Lee and Becker 2026); see also Longpre et al. (2024) for empirical evidence that data restrictions on websites are rising in the AI era.
On the other hand, an ineffectively enforced hardwired pause could plausibly accelerate efficiency gains by tightening constraints on the computational resources available to any engineering teams that continued frontier model training efforts, especially if illicit efforts are widespread. Maintaining a wide margin between governance thresholds and and frontier capabilities could mitigate this acceleration by reducing the temptation to evade. Likewise, efficient inference on frozen frontier models could contribute to training efficiency gains in a variety of ways, such as generating data sets for distillation, grading RL rollouts, or suggesting algorithmic improvements to engineers.
Continued training efficiency gains would not necessarily be fatal to an adaptable governance regime. As governance matures, hardware and software innovations could give authorities more options for flexible, light-touch governance: more reliable geotracking of chips could cause the supply of dark compute to dwindle away; advances in hardware specialization could yield a variety of performant hardware systems that, like inference-only chips, are inefficient for frontier model training; or advances in privacy-preserving software-based verification and workload classification could allow for the wide deployment of hardware that offers users relatively unrestricted general-purpose functionality while detecting and deterring evasion.
6.4 Quantitative ledger estimates
Treaty authorities would be responsible for ensuring that no state party has amassed a fleet capable of posing a risk of strategically-relevant covert evasion or overt breakout, by estimating the total amount of computing capacity that could be marshaled toward a covert evasion or overt breakout research program.
This section discusses a simple methodology by which treaty authorities might estimate the covert evasion capacity and the overt breakout capacity , for each potential source of illicit compute and each state , beginning with a raw capacity and then discounting according to how much of that compute source could realistically be put to use in each of the two scenarios.
We implement this methodology to derive conservative estimates for what the covert evasion and overt breakout ledgers might look like for both the U.S. and China, for our near-term pause scenario at year-end 2026.
6.4.1 Discount factors
For each of the two risk scenarios, two discount factors could be applied: a political discount factor representing the fraction of hardware in that category that could plausibly be marshalled toward building frontier models (fractionally including hardware engaging in part-time evasion), and a technical discount factor representing the average sustained utilization that could realistically be achieved, relative to present-day frontier training. That is, means similar utilization to present-day data-center grade training, not 100% utilization.
For each compute source and state , the parties could negotiate means of estimating a total (undiscounted) existing capacity as well as four discount factors: and , the political discount factors for the covert evasion and overt breakout, and and , the technical discount factors for the same. Then, the evasion capacity for that compute source and state would be estimated as , and the breakout capacity would be estimated as . Summing over gives the evasion and breakout ledgers and for each state, which must respectively be kept below and . The discount factors and the raw capacities could all change with time, but we suppress their dependence on time to avoid cluttering notation.
Considerations in determining the political discount factors would include how diffusely or centrally held the relevant chips are, whether they are directly controlled by the state, and what legal or political barriers the state would face in commandeering them in a covert or overt scenario. The covert discount factors in particular would depend on what kinds of verification measures the state allows in its facilities, how many different actors would need to be involved in a conspiracy to violate the treaty, and how successfully the state could force them to remain silent.
For example, a supercomputer facility in a sovereign fleet would be charged with a high overt political discount factor because it could be reassigned to breakout training in a short period of time . However, it might be charged a much lower covert political discount factor if the host state is willing to allow power monitoring, telemetry mechanisms, periodic inspections of facilities, or legal and political safeguards such as criminalization of treaty violation and protection for whistleblowers. If a state can persuade other parties to accept that no more than 20% of a facility could plausibly be marshalled toward covert training without a significant likelihood of detection, its might be 0.2.
Considerations in determining technical discount factors would include limitations on memory and interconnect bandwidth, or parallelism overhead for chips with low memory. The technical discount factor for covert training could be less than that for overt training due to the technical overhead of distributed training, disguising workloads, or locating chips in secret data centers with substandard operating conditions. We undertake a simple illustrative analysis here, but a more complete analysis would disaggregate each source by the type of hardware and details about its deployment; the utilization achieved by current AI companies in frontier pretraining under idealized conditions might be realistic for some overtly-deployed accelerator hardware , but technical discount factors for present-day inference-optimized chips would likely be below 1, and those for consumer GPUs would be considerably lower still.
To establish plausibility that these bounds could be realistically verified, we consider several sources in turn, discussing realistic upper bounds for capacity and discount factors for the U.S. and China. In each case, we compare them to the covert evasion fleet threshold and the overt breakout fleet threshold for our near-term pause scenario. As these estimates apply to a hypothetical future, they necessarily reflect speculations about, e.g., how cooperative pause-willing states parties would be in implementing treaty rules or furnishing sales records to locate existing lithography machines and accelerator chips.
As we will see, under plausible assumptions, it is unlikely that any of the sources discussed below could represent a significant strategic threat for overt or covert evasion in the short term after a pause. For each source, we evaluate its headroom for a treaty that gives an allowance of 25% for both ledgers , defined as the number of orders of magnitude of erosion, due to training efficiency gains, that the treaty would have to sustain before that source became a binding constraint on balancing the ledger.
6.4.2 Transitional training-capable inference service
The most obvious treaty exposure would be that from transitional inference service on pre-pause training-capable hardware. Leaving these training-capable accelerators in industry-grade commercial data centers, where they could in principle be seized by a state government in an overt breakout scenario, plausibly gives and close to 1.
In practice, both discounting factors might be somewhat lower. The value of could vary widely, depending on the political plausibility of the state’s government seizing an enormous quantity of privately-held capital for a national breakout program. Political or institutional safeguards that would make it more difficult for the state to do so, or that create hurdles that would delay deployment in such a scenario, could reduce the average amount that would be available over a window of length . The value of also might be lower than 1 since the total installed stock is heterogeneous, with a significant fraction of older hardware.
In any case, we can bound , thereby bounding each state’s breakout capacity by the total stock located in each country at the pause date.
We project the U.S.-located stock at total at year-end 2026, giving:
or 20% of , and
or 1.8% of . Thus, at the pause date, both breakout capacities would likely be below the allowance before discounting. However, the breakout number leaves almost no headroom for a treaty with 25% allowance, and two doublings of training efficiency gains could bring itself close to the U.S. breakout capacity. This would be a problem in the early years of the treaty, but not an insoluble one; we discuss ways to ameliorate it in § 6.5.2.
Even while inference service continued, verification measures such as those described in § 3.3.3 could significantly reduce its covert evasion risk. If we assume pessimistically that 10% of hardware operators in the U.S. engage in treaty evasion, somehow avoiding detection by treaty recomputation servers even while diverting 1% of their computing power to evasion, we arrive at . Even if they achieved frontier-grade utilization for their covert evasion, the total covert evasion capacity would be
or 0.55% of . Assuming, again pessimistically, that the Chinese government were able to covertly suborn all of its hardware operators into diverting 1% of their computing power, their covert evasion capacity would be similar, with
These evasion capacity numbers apply to the transitional inference service after verification protocols have been implemented. Both political discount factors would likely be higher during the stages when these methods were being set up. However, provided this could be done in a timely fashion, it would leave little time for a strategically relevant model to be covertly trained. Whole-lab inspections as proposed by Choussat and Khoja (2026) could be a stopgap interim measure, if it were difficult to stand up the inference-only cluster protocols of § 3.3.3 at the timescale on which leaders wished to implement a pause.
For a longer-term pause, we expect that verification measures could be considerably tightened based on methods currently under development as described in § 3.3.3, leading to values that were considerably reduced for both states.
However, doing so would not on its own be sufficient for implementing the treaty, because verification measures generally do not reduce breakout exposure. This is what makes stopping the production of training-capable hardware an essential element of a hardwired pause: the more training-capable hardware states accumulate, the less credibly they can assure other states that they would not break out and race faster than before.
A more complete covert evasion analysis would consider the possibility of the U.S. or China subverting transitional inference service outside their jurisdiction. We assume such use would likely incur considerably steeper political discounting.
6.4.3 Dark compute
Significant quantities of ungoverned accelerators would be highly problematic for treaty governance. Chips lost or stolen in a host state might be charged under political and technical discount rates close to one, reflecting the high risk of their being used for covert evasion or overt breakout. While a few thousand missing chips would not approach relevant treaty allowances, a million missing H100e units would approach 16% of , leaving little margin for training efficiency gains to erode the evasion fleet threshold .
Our estimates from § 5.5 reflect what a thorough census might miss, but our analysis does not assume that a thorough census could happen overnight. Compliance would depend on political conditions we cannot foretell. In some pause-willing futures, leaders could be ambivalent about the pause and unwilling to provide any transparency to rivals; while in others, leaders’ most urgent concern could be to prevent further training or fine-tuning of frontier models by private evaders. It could also be that leaders gain greater confidence in each other over time, beginning with small measures and slowly finding ways to work together more closely. It is possible that the census would take time to implement.
We divide our analysis into several scenarios, according to states’ degree of cooperation in the census.
In a high-cooperation scenario, we consider a world where American, Chinese, and other world leaders conclude they have a mutual interest in cooperation to address common threats, and prefer to pursue national strategic objectives through alternative treaty-provided channels. Both nations make convincing public efforts to recover pre-pause chips, but might stash a small, plausibly deniable number as insurance.
For such a scenario, we assume that both states would not significantly sabotage their own internal census methods, and would stash no more than 20% of the total missing amount, giving
We assume this cooperation would be evident to other states, who might conservatively estimate that at most half of the missing compute was in the state’s possession, giving a covert evasion political discount factor of , leading to
Applying these bounds to the U.S. and China gives
or 0.11% of for our near-term scenario, and
or 0.10% of (the two 6.7k figures above are different values).
In this scenario, the treaty would appear to be quite durable after the sunset of transitional inference service, since all other terms in the covert evasion and overt breakout ledgers would initially remain below half of one percent. It could comfortably survive two orders of magnitude of erosion due to training efficiency gains, and plausibly more if innovation in verification methods or national technical means drove down political discount factors, or if other modes of hardware specialization similar to inference-only chips drove down technical discount factors for non-dark compute on new, governed hardware.
The main threat to durability would be the threat that new sources of undeclared accelerator production would be developed; measures like those in § 5.3 would therefore be essential.
Second, we consider a low-cooperation scenario, in which China and the U.S. arrive at a tense agreement with relatively little trust or goodwill. Both accede reluctantly to most records requests, but records requests that are at all politically sensitive are refused, making it more difficult to account for chips. Despite their pause-willingness, both states suspect the other of planning to cheat, and both find ways to hold back as much untracked hardware as they can, especially powerful newer hardware, possibly sabotaging their own investigation efforts in order to increase the plausible deniability of their own concealed stockpile.
In this case, we would expect a larger fraction to go missing, and no political discounting to be applied. In the U.S., it would likely be more difficult to conceal a large stock owing to the relatively greater difficulty of keeping secrets. For this scenario, we assume five times the missed quantity estimated in § 5.5.2 for the U.S., and ten times the missed quantity for China:
which would respectively be 0.88% and 1.7% of for the U.S. and China, and 0.032% and 0.063% of .
In this scenario, the regime could survive one order of magnitude of training efficiency gains relatively comfortably. It could survive two if the U.S. and China could eventually achieve something closer to the high-cooperation scenario after a thaw in relations, an increase in the degree of pause-willingness, or an increased sense of urgency after training efficiency gains eroded treaty bounds. If greater cooperation could not eventually be achieved, longer-term durability would hinge on the success of detection-based fallback measures.
Finally, we consider a non-deniable concealment scenario in which negotiations over recovery break down. Both states openly accuse each other of pursuing covert evasion plans, and each sees limited political value in persuading the other that they are not pursuing it themselves, but a deal nevertheless barely gets through as a stopgap measure against some acute emergency. In this case, China might conceal an amount roughly equal to its pre-pause diverted stock; in the U.S., it would likely be more difficult to conceal a large stock owing to the relatively greater difficulty of keeping secrets. Then we might have
which would respectively be 4.7% and 16% of , and 0.17% and 0.57% of , if undiscounted.
This scenario would leave little margin for erosion of ; in particular, China’s covert evasion capacity from dark compute would pose a similar issue as the U.S.’s overt breakout capacity from transitional inference service. Both problems would need to be resolved in the early years of the treaty, but the two states would likely have a few years.
6.4.4 Inference-only accelerators
While inference-only accelerators would be deployed at scale, their residual training factors would be audited with certified bounds, and their total training capacities bounded per hardware operator and state.
The covert evasion risk from licensed inference-only accelerators would be a choice parameter for treaty governance: the number of licenses for a given state could be reduced if the state had insufficient headroom on its covert evasion or overt breakout ledger.
Under overt breakout scenarios, seized inference-only hardware in commercial data centers could be used at its peak arithmetic throughput for training , so the treaty-certified residual training factor would be the primary protection against breakout. A host state’s inference-only fleet with an average residual training factor of could remain below while serving units of inference demand. For example, if a state wished to spend 1M H100e of its overt breakout allowance on an inference-only fleet, then it could serve 100M H100e worth of token demand on a fleet with average residual training factor below 0.01, or 1B H100e on a fleet with average residual training factor below 0.001. By comparison, Deloitte projected that two thirds of AI compute would be used for inference in 2026 (Crossan et al. 2025), or roughly 31M H100e if applied to our earlier estimate of 47M H100e of global capacity.
An average residual training factor of , amounting to 100 GFLOP/s of GEMM-laundering ability per H100e, would allow each of the U.S. and China to replicate the estimated 31M H100e of global inference service one hundred times over with inference-only accelerators, with each state keeping its breakout exposure below
or 0.18% of the initial breakout fleet threshold , leaving two orders of magnitude of headroom.
Here, is not an estimate or prediction but an illustrative value, and the ledger estimates do not ride on its value. If available residual training factors were higher or lower, or technical discounting factors below one could be justified, the total inference service could be adjusted based on available headroom. Significantly lower values can be achieved if all major channels for GEMM laundering can be foreclosed, but doing so would require a physical audit of a realized system.
The covert exposure could be driven further down by certifying a separate residual training factor , which might be lower than if the most favorable GEMM-laundering channels required physically tampering with inference-only chips in a way that would be evident to inspectors. Additional verification steps, such as those suggested above for the transitional inference service fleet, could reduce the covert evasion capacity further. If , as assumed above owing to the difficulty of evading discovery by the recomputation server, then the covert evasion capacity is still lower:
or only 0.0049% of .
6.4.5 Consumer GPUs
Next, we consider consumer GPUs. Jon Peddie Research (2024) forecasts a global installed base of 119M discrete desktop add-in boards by the end of 2028. If there are half as many laptop units as desktop units,72 we might use a rough conservative bound of 100M total cards in either the U.S. or China, with the average such chip having peak dense arithmetic throughput of 30 TFLOP/s,73 giving potential covert training risk before discounting.
Political discounting would deflate this risk considerably, especially under the covert scenario. Since gaming cards are diffusely held, deploying them at scale would require either a vast conspiracy of regular citizens knowingly permitting their cards to be used toward treaty violation, or hacking into and diverting a small portion of a much larger group of citizens’ cards without their noticing. If parties deem it implausible that a state could commandeer more than 1% of consumer card capacity for a covert run without running a major risk of discovery, we might take .
Political discounting of consumer GPUs under an overt breakout scenario is more uncertain and depends on parties’ estimation of the likelihood that governments could commandeer a large fraction of consumer hardware even in a wartime footing scenario. In constitutional democracies, at least, any such effort would require sufficient constitutional or statutory basis74 and might face significant legal as well as political hurdles. States might accept the premise that it would be difficult to commandeer more than 25% of total consumer capacity toward a breakout program, giving .
Even generous technical discount values would likely be low for consumer GPUs since they are heterogeneous, are thermally limited when operated at high utilization, and have low memory and interconnect bandwidth, all rendering them unsuitable for efficient training at frontier scale; the covert technical discount factor would be lower still due to evasion overhead. Dean (2026a) estimates the “usefulness” (essentially a technical discount factor) of each chip, finding consumer gaming cards to be worth roughly 0.05 of an H100e. If we are twice as generous and take , the total covert evasion risk is reduced to
and the overt breakout risk to
representing 0.047% and 0.043% of and , respectively.
Thus, under covert or overt scenarios, the consumer GPU base does not appear to represent a significant hurdle to the treaty’s durability, up to at least two orders of magnitude, due to its diffuse ownership.
One possible future failure mode of this calculation is that, under a treaty, a few individuals or firms might buy up a very large number of consumer GPUs, negating the above argument based on diffuse ownership. In order to defuse this longer-term threat, states parties might consider passing laws limiting the number of consumer GPUs held by any one actor, and taking steps to prevent skimming from the production line.
6.4.6 Licensed sovereign fleets
In § 3.3.1 we discussed the possibility of permitting licensed sovereign training-capable accelerator fleets to be used for specific sensitive national security workloads. We assess the exposure from these fleets by considering El Capitan, the largest supercomputer in the U.S. and an important strategic asset with the primary mission of nuclear stockpile maintenance. El Capitan serves 46k AMD Instinct APUs with 980.6 TFLOP/s each, amounting to in total (Lawrence Livermore National Laboratory 2025; AMD n.d.).
The technical discount factor for overt breakout can be conservatively estimated close to 1,75 though it might be lower for covert evasion since the network topology is presently optimized for scientific workloads like fluid dynamics simulation rather than frontier AI training, and any evasion scheme would likely come with technical overhead.
We assume an overt breakout political discount factor of , since the supercomputer is fully controlled by the U.S. Federal Government, giving total breakout exposure of
or about 0.026% of .
The political discount factor for covert evasion, however, might be much lower, owing to the difficulty of keeping secret from staff the diversion of a large fraction of such a supercomputer. We might conservatively take , yielding covert exposure
or about 0.14% of .
As a result, we estimate that the treaty could sustain a practice of allowing a number of supercomputers on the scale of El Capitan per state without threatening treaty thresholds, at least early on.
The number could grow further with improved institutional governance or verification measures allowing for steeper political discounting, or improved technical governance measures such as hardware specialization allowing for steeper technical discounting. If the political discount factor and the technical discount factor each fell by an additional factor of five, for example, then the covert evasion capacity would fall to 0.0057% of , which would leave room for each state to operate a multiple of the capacity of El Capitan, even after two orders of magnitude of erosion.
We assess this as an encouraging sign for political feasibility: it does not appear that, in the immediately foreseeable future, a treaty would need to impose on strategically important assets like El Capitan.
6.4.7 Other categories
Another palatability feature we discussed in § 3.3.1 was the option to create internationally governed scientific preserves. Such preserves’ near-term size would likely be small compared to current commercial fleets, and it would appear highly difficult for covert evasion to occur there in transparent and replayable workloads, under the noses of the treaty authorities running the preserve. It would be important to guard against overt breakout scenarios, however, either by locating them in neutral states unlikely to participate in a breakout, or giving treaty monitors the ability to quickly and permanently incapacitate the chips located in a scientific preserve when the host state attempted to seize them, or both. Thus, it appears possible to ensure that scientific preserves contribute negligibly to the U.S. or China’s covert evasion or overt breakout ledgers.
Finally, we discussed in § 3.3.1 the possibility of states leaving strategic reserves under seal, monitored by other states or treaty authorities. Such reserves would, by design, be fully available to states in overt scenarios, so ; thus, their undiscounted compute capacity would be added to the state’s breakout ledger. However, it should be possible to implement such a plan so that the strategic reserves are verifiably powered off at all times while they remain under seal, so , thereby contributing nothing to a state’s covert evasion ledger. This is what makes the strategic reserves an attractive, politically cheaper alternative to covertly hoarding a large reserve of chips.
The above analysis is intended to be illustrative, not exhaustive. A more complete analysis would include integrated consumer graphics cards, phones, CPUs, and edge devices such as hardware installed in vehicles. Based on our initial analysis above, however, we believe there is good reason for optimism that treaty constraints could be satisfied in the near future.
6.5 Durability assessment
Next, we discuss the primary determinants of the durability of a hardwired pause, in particular the degree to which it could durably guard against covert evasion and overt breakout. Because we cannot predict the rate of training efficiency gains after a pause, we measure durability as a function of how much treaty thresholds would need to erode before the U.S. and China would not both be able to stay below allowances of 25% of and . That is, we ask when it would become difficult to balance the covert evasion ledger inequality
or the breakout ledger inequality
for or . The binding constraints on durability are the terms with the least headroom, specifically those for which reducing headroom within a period of years would pose a significant political or technical obstacle.
To estimate durability in time, we assume that annual post-pause training efficiency gains will run at or below the 57% estimate of Patel and Han (2026), corresponding to slightly slower growth than one order of magnitude gain every five years, or to the commonly cited 18-month doubling time famously associated with Moore’s Law for semiconductor miniaturization (Tuomi 2002). Under this scenario, a decade of durability would require achieving two orders of magnitude of headroom ( and for each state) well before the end of the pause’s first decade.
Under the 2.7x annual training efficiency gain estimated by Ho et al. (2024), the treaty would be about half as durable. However, as discussed in § 6.3, those gains were largely scale-dependent, and were observed during a period of rapid scaling growth.
In our analysis below, we do not assume that a treaty would be in force, with ledgers capped at a 25% allowance, beginning at the time of the pause date: in a near-term scenario, it is more likely that an informal agreement would eventually be replaced by a treaty after a process of negotiation that could take years. It might also not be realistic to enforce 25% allowances on day one; rather, a schedule might require states to reduce them to 25% over a period of years. These schedules could be somewhat flexible to the political and technological conditions, but as a working assumption we presume they could start close to 100% at entry into force and fall gradually to 25% by the time one order of magnitude of training efficiency gains had been incorporated.
6.5.1 Durably guarding against covert evasion
Dark compute is the term in the breakout and the covert evasion ledgers over which the treaty would have the least control, and it would also be the term most difficult to discount for covert evasion purposes. Based on our estimates from § 6.4.3, it appears implausible that dark compute would pose an immediate covert evasion risk but, if undiscounted, it could plausibly pose a risk over time due to training efficiency gains. Under our high-cooperation scenario, the treaty ledger would initially have two orders of magnitude of headroom, while under our low-cooperation scenario it would only have one order of magnitude.
If the U.S. and China found it difficult to cooperate in the early years of a pause, they would have further chances to reduce dark compute stockpiles before the initial headroom ran out. If they could not do better than our non-deniable concealment scenario, for example, the covert evasion ledger could survive roughly six-fold training efficiency gains before China’s covert capacity exceeded the fleet threshold. Achieving the low-cooperation scenario would be sufficient to get both states through the first order of magnitude of training efficiency gains, and achieving the high-cooperation scenario would be sufficient to achieve two orders of magnitude at a 25% allowance. Thus, a thorough global accelerator census would not be absolutely required on day one, but would need to be completed shortly after the first order of magnitude of training efficiency gains. However, we do not suggest this to encourage complacency, as an initial failed accelerator census could make it more difficult for a second census to succeed.
The most plausible fallback plan, if states could not manage to conduct a thorough census, would be efforts to detect covert training at the time it occurred. Such detection would be made more difficult owing to advances in distributed training, using methods like DiLoCo (Douillard et al. 2023), DiLoCoX (Qi et al. 2025), and more recent variants further extending distributed parallel training, but at present known methods would still require large compute clusters in order to train a model on the scale of ; see Sevilla (2025) for further discussion.
6.5.2 Durably guarding against overt breakout
The transitional inference service in the U.S. could plausibly pose a breakout risk much sooner, if its replacement by inference-only compute were delayed for a long enough period of time. If 57% annual training efficiency gains accrued in the 26 months that we project it might take for inference-only chips to begin production at scale, the U.S.’s 20% overt breakout capacity would rise to 53% unless most of the pre-pause training-capable stock had been replaced.
There are several potential ways to mitigate this exposure. The simplest would be to make a one-time exception for transitional inference service and wait for pre-pause training-capable hardware to be replaced by inference-only hardware, which we estimate would happen shortly thereafter. This could be accomplished by using a shrinking schedule for the allowance, which could shrink to 0.25 or another terminal value over a period of years as discussed above.
Alternatively, if parties to the treaty were not satisfied with the transition timeline as a mitigation, some of the hardware could be transferred to data centers outside U.S. territory at an earlier date. These could be located offshore or in space, in the custody of a neutral state, or in international custody, for example in scientific preserves. These options would only mitigate breakout risk if it were understood that the owner state would not be able to recover these chips’ training capacity after a breakout (cf. Dean 2026b). Under our near-term pause scenario, if the U.S. were willing to do this for half of its transitional inference service (which would not require a major commercial disruption if the chips kept operating in new data centers) then it could ensure that its breakout ledger did not exceed 100% of , and it could eliminate this term entirely once transitional inference service were phased out by inference-only service. Since all other terms have at least two orders of magnitude of headroom at a 25% allowance, phasing out transitional inference service by this time would be sufficient for the treaty to last for two orders of magnitude of training efficiency gains.
If most or all of the pre-pause training-capable compute used for transitional inference service were transferred to data centers in locations where they were judged to pose low breakout risk for any state, such as offshore or in space, they would still need to be weighed into states’ covert evasion risk ledgers. If states parties were confident enough in verification methods, and in the plan to prevent recovery upon breakout, it is possible that pre-pause training-capable compute could continue to perform inference service for the remainder of its useful lifetime.
6.5.3 Extension options
In each of the above scenarios it is possible that the combination of training efficiency gains and dark compute would eventually make it impossible to balance the covert evasion ledger, with detection-based methods failing to make up the difference. In that case, the simplest way for pause-willing states to extend the lifetime of the agreement would be to declare their dark compute and place it in strategic reserves, but this might be politically difficult. An alternative option would be for an international consortium to jointly train a larger model to advance the frozen frontier just enough to make policing feasible once again.
6.5.4 Lithography proliferation
Missing lithography machines, supported by a complete indigenous covert supply chain, could significantly undermine the regime over a period of years by creating a continuous inflow of dark compute into one state. For example, Brown and Khan (2026) estimate that a single ArF-immersion DUV machine could contribute roughly 30k–47k Ascend 910C-equivalents—or 24k–38k H100e in our units—per month at state-of-the-art yields for 7nm precision using multipatterning, significantly eroding durability. The regime would need to prevent undeclared accelerator production by verifying the location of each advanced pre-pause lithography machine in each state from vendor service records and by dedicating sufficient ongoing intelligence efforts to detect any new undeclared manufacturing facility or covert supply chain. Such an effort would be highly plausible given the many subcomponents and materials required to service a fully functional lithography machine.
7 Implications of a hardwired pause
In this section, we discuss some of the implications of a hardwired pause, including both general implications of a frontier pause and specific implications of a hardwired pause relative to other ways of pausing the frontier. We make no claim to comprehensiveness, but we discuss several issues that appear especially salient: the macroeconomic impacts of a frontier pause with continued inference service (§ 7.1); the reduction of inference costs for consumers (§ 7.2); the economics of using MSICs, as an example of inference-only chips, for the hardwired pause (§ 7.3); impacts of the hardwired pause on concentration of power (§ 7.4) and continual learning (§ 7.5); and the need for other measures in addition to a hardwired pause (§ 7.6).
7.1 Would restraints on frontier AI crash the economy?
AI-related ventures have been heavily responsible for the recent strength of U.S. GDP growth. According to the Federal Reserve Bank of St. Louis, spending on information processing equipment, software, R&D (including AI R&D), and data centers accounted for nearly 40 percent of U.S. real GDP growth in the first three quarters of 2025 (Rubinton and Patro 2026). Take away that spending, some worry, and the economy will crash. Similarly, AI-related companies are heavily responsible for the high valuation of the S&P 500 stock index; they now account for over one-third of the total weight of the S&P 500 (Cembalest 2026; see also BlackRock 2026). Tank those share prices, observers warn, and you risk tanking the economy.
Unpacking these arguments, in the context of a scenario where restraints are placed on frontier AI, requires distinguishing the following questions. How would restraints on frontier model training affect AI-related spending growth, by hardware and software companies developing those models and by other companies utilizing their services? How would the supply side of the economy be affected—in other words, what would happen to productivity growth? How would financial markets be affected? And how would the reaction of financial markets impact the broader economy?
On the first question, the impact on AI-related spending, it is important to note that the St. Louis Fed’s estimated 40% is at the upper end of the range. Other institutions such as Goldman Sachs estimate much smaller impacts, given that a significant share of spending is on imported chips and equipment (Goldman Sachs 2026). Spending on imports raises Taiwanese and Korean GDP, not U.S. GDP. Of physical investment, it is mainly the change in spending on data centers in the U.S. that would visibly affect U.S. GDP growth. Headlines notwithstanding, data center construction accounts for less than a fifth of total AI-related investment in the United States (Juniewicz 2026b; U.S. Census Bureau 2026).
AI-related spending that is not on imported chips, equipment and data centers tends to be on software licenses, enterprise software deployment and model creation. Spending on AI software, services, and models is projected to account for roughly 40 percent of AI-related spending worldwide in 2026 (Gartner 2026). A moratorium on training frontier models would discourage spending on model creation by definition. But there is increasing evidence that firms licensing and deploying AI are not mainly deploying frontier models (e.g. ICONIQ 2026; Kharazian 2026), since sub-frontier models, while cheaper, yield nearly the same productivity gains in most applications. Freezing the training of frontier models may then have little impact on licensing and deployment-related spending, and in turn on the demand for data center services. (Parenthetically, the greater threat to the demand for data center services may be open-weight models that can be run on enterprise hardware or even a laptop.) These same observations suggest that frontier model restraints would do little to reduce economywide productivity growth below its current rate.
There would presumably be a drop in the lofty stock prices of the “Mag7,” the seven ultra-large-cap tech and tech-adjacent firms that have driven the S&P 500.76 But stock market corrections are no guarantee of recessions. (The economist Paul Samuelson famously observed that the stock market predicted 9 out of the last 5 recessions.) A decline in the price of the Mag7 stocks would cause some slowdown in consumption and investment spending. But the marginal propensity to consume out of stock market wealth is just 3 percent (Chodorow-Reich et al. 2021; see also Ludvigson and Steindel 1999). Imagine that the market valuation of the Mag7 of high-tech companies falls back to its beginning 2025 level—that is, that it falls by $6.5 trillion. Three percent of $6.5 trillion ($195 billion) is 0.6 percent of projected 2026 GDP, or about a quarter of projected 2026 GDP growth (Board of Governors of the Federal Reserve System 2026).
This is plausibly an upper bound on the fall in consumption spending. Less consumption (and investment) spending will mean less upward pressure on interest rates and less inflation. The Federal Reserve would react by reducing its policy interest rate (or moderating its rate of increase). In the decade ending with COVID, when interest rates were at or near zero, the Fed had limited ability to respond in stabilizing fashion (leading it to adopt unconventional monetary policies, whose efficacy is disputed). Now that its policy rate is nearing 4 percent, it has more scope for buffering a decline in spending.
The scenario where Mag7 share prices fall back to early 2025 levels is just a thought experiment. (AI-related companies do other things besides developing frontier AI models and selling their services. The premise that advances in frontier models account for most of the change in valuations since early 2025 is the motivation for this particular thought experiment.) Readers can scale up or down the $6.5 trillion fall in share values as they prefer. But the same general conclusions follow, including the capacity of the Fed to buffer all but the largest shocks.
What about risks to financial stability? A fall in equity prices would be painful for investment funds like Situational Awareness specializing in AI-related investments. But a financial crisis will follow only if banks and other mainstream institutional investors are implicated, as they were in 2008. Hedge funds fail (“wind down”) all the time. Only when they are importantly linked to an institutional parent (as in the case of two Bear Stearns hedge funds in 2007) do systemic consequences follow. And there is little evidence today that such institutional parents are heavily exposed to AI equity risk.
Even if a fall in share prices owing to a ban on frontier model training is unlikely to create serious problems for the economy as a whole, it could create problems for specific firms. An example is Nvidia, which is invested heavily in producing training-capable accelerators for frontier models. Nvidia could pivot to building inference-only accelerators, for which there would still be a strong demand, but such a transition would be neither painless nor cost free, and the company would presumably not dominate that market as it does the market for general-purpose accelerators. More than 50 percent of Nvidia shares are held by U.S. investors (MarketScreener 2026), and the company currently accounts for roughly 8 percent of the S&P 500. Some 70 percent of all Nvidia shares are held by institutional investors, the largest holdings being those of Vanguard, BlackRock, Fidelity, Geode Capital Management, and State Street (Yahoo Finance 2026). These custodians hold shares on behalf of their mutual fund and ETF customers, so losses would be spread among retail investors. There would be negative wealth effects, with implications for consumption, but not obviously risks to financial stability. Nvidia also provides credit guarantees and loan backstops for the construction of data centers, so were a pause to disproportionately affect its cash flow and share valuation there might be a significant negative impact on AI-related construction activity.
Also worrisome are developments in debt markets. With the cost of data center construction exceeding their sponsors’ free cash flow, and with income from tenants still a distant prospect, much of this investment is being financed by borrowing from private credit funds (nonbank firms making privately negotiated loans). If the demand for data center services is less than expected, these loans may fail to perform, which is to say that borrowers will default.
The question then becomes who is invested in private credit funds. The popular image is of the same family offices and high-net-worth individuals who invest in venture capital and can afford losses. But the reality is that the vast majority of private credit finance is provided by institutional investors, including pension funds and insurance companies (Cai and Haque 2024).
Moreover, private credit funds have been taking proactive steps to tap this institutional money by acquiring insurance companies, whose asset portfolios they can load up with loans. Disappointing data center revenues, which cause private-credit loans to become nonperforming, could thus lead to serious problems for the insurance industry. This could play out in a number of ways. First, insurance companies’ private credit parents could recapitalize them, assuming they have the funds. Second, state guarantee funds, to which healthy insurance companies contribute, could fund the recapitalization, although there will be no healthy companies in the scenario where the entire industry comes tumbling down. Third, the federal government could step in with a bailout, as it did for American International Group (AIG) in 2008.
In addition to borrowing from pension funds and insurance companies, private credit firms borrow from commercial banks, using their portfolio of loans as collateral, thereby leveraging their commitments and juicing their returns. So even if banks do not lend to or invest in AI firms and data centers directly, they still may be on the hook.
Finally, private credit firms originally built on a buy-and-hold model increasingly securitize their loans. They package those loans through special purpose vehicles that issue bonds backed by the associated cash flows, where the bonds are divided into risk tiers or tranches claiming first, second and third dibs on debt service payments. The so-called “mezzanine tranche” of medium risk bonds is then sold on to other asset managers. To assess the immediacy of crisis risk, we would have to know more about the riskiness of this mezzanine tranche and who holds it, which is not public information. The situation is reminiscent of the roles of securitization, special purpose vehicles, tranching and opacity in the Subprime Crisis of 2007-8.
Tech firms, for their part, also create special purpose vehicles (SPVs) to disguise the impact of debt on their balance sheets. The SPV is legally distinct. It borrows to purchase real estate, land and servers, and then leases this infrastructure back to the tech firm under a multi-year contract. For the parent, a large one-time capital expenditure is transformed into a smaller lease payment spread over a period of years, with a smaller balance-sheet impact, which encourages debt-fueled capital spending. But to secure private equity and other funding, the tech firm parent must provide a residual-value guarantee, whereby it promises, if the appraised value when the tenant leaves falls short of the guaranteed minimum value, that it will step in and pay the difference (Oh 2026). The situation is reminiscent of the failure of Enron in 2001.
How would a moratorium on the training of frontier models affect these risks, and what should be done about them? If it is correct that demand from licensees is mainly for sub-frontier inference, then a freeze on frontier model training would have little impact on data center revenues, on the performance of private credit loans for data center construction, and on the tech firms’ SPVs. Were a freeze to redirect additional demand from frontier model training to inference service, then this conclusion would be reinforced. Only in an alternative scenario where a moratorium on the training of frontier models led to a sharp decline in licensing and in demand for data center services would a freeze increase the risk of serious financial dislocations.
What should regulators do to monitor and limit the risks to economic and financial stability that might arise in the event of a significant slowdown in the demand for data center services? Bank regulators should closely monitor bank lending to private credit firms and limit concentrated exposures to this sector. Insurance regulators should question private credit purchases of insurance companies and be wary of potential conflicts of interest. Financial regulators should require more disclosure by financial entities that are securitizing private credit loans and seek more data on who is holding the resulting securities. They should more strongly regulate the special purpose vehicles that AI firms use to limit the balance sheet impact of debt, since access to those accounting devices encourages their borrowing, which contributes to the build-up of financial risks.
The bottom line is that any decline in investment in data center facilities and equipment, and any decline in consumption spending owing to lower equity valuations, both pursuant on a frontier-model pause, are likely to be limited, because the declines in spending in question will be limited, and because their macroeconomic effects can be buffered by monetary policy. Only if declining data center revenues cause distress in debt markets and if that distress is allowed to infect institutional investors could the effects be seriously destabilizing. This last possibility is something for which regulators could, and hopefully will, prepare.
7.3 Economics of MSICs
As described in § 1, the hardwired pause allows the production of inference-only chips for running whitelisted AI models. As an example of an inference-only chip, we cited certain model-specific integrated circuits (MSICs), which can only run inference on a single model. These are not the only type of inference-only chips that could be used, but since MSICs are already being developed without a pause, in this section we focus on MSICs.
The possibility of using MSICs leads to two economic questions:
Assuming that there was a frozen whitelist of models and that both MSICs and flexible accelerators (which can serve any model) were allowed for inference on whitelisted models, what allocation of inference between MSICs and flexible accelerators would minimize total costs for a given volume and distribution of demand for inference?
Assuming that there was a frozen whitelist of models and only MSICs were allowed for inference, how much “economic inefficiency” would the restriction to MSICs entail?
The downsides of using an MSIC for a specific model are the up-front, non-recurring engineering cost of designing and setting up production of the MSIC and (ii) the fact that an MSIC for a specific model cannot be transformed into an MSIC for a different model in response to future shifts in demand for models on the fixed whitelist as consumer tastes change or if a model is de-whitelisted (see §§ 7.3.3–7.3.4). The economic upside of creating an MSIC for a specific model is the much lower operating cost of inference using an MSIC, compared to the cost of inference using a flexible accelerator. In this section, we produce some preliminary “back of the envelope” estimates of how these tradeoffs might balance out in the long run.77 The specific numerical values of the estimates are not meant to be taken seriously; the goal is rather to get a qualitative sense of how the tradeoffs might balance out in the long run. We hope that economists will do more sophisticated economic modeling in future work (cf. Dong et al. 2022).78
7.3.1 MSIC adoption in a mixed MSIC-flexible world
Assume there is a whitelist of finitely many models. We will consider a simple rule for deciding whether to build an MSIC for a given model on the whitelist, based on the following data:
is the expected share of total inference demand (as a percentage) that will go to model on the whitelist. We abstract away from the fact that building an MSIC for could lower its inference price and thereby increase demand for that model.
is the fixed cost of development (in dollars) of an MSIC for a model (assuming for simplicity this is the same for each model on the whitelist), called the non-recurring engineering cost for the MSIC.
is the savings of moving model inference from flexible accelerators to MSICs, as a percentage of the cost of inference on flexible accelerators (again assuming for simplicity this is the same for each model).
is the expected (discounted) lifetime cost of serving all inference demand that it is profitable to serve using only flexible accelerators, which we call the baseline cost of serving inference.
Suppose that an MSIC will be developed for model if and only if the expected (discounted) lifetime cost-savings exceeds the fixed cost :
Estimates for the savings term range from 60% to 95%.79 Estimates for the fixed cost of developing an MSIC range from $10M to $500M.80 If we imagine that all models currently available on OpenRouter pass safety scrutiny and are whitelisted,81 then we can use OpenRouter data82 to give us a scenario for the distribution of demand shares : for each model served by OpenRouter, we simply take to be the share of total tokens of inference served on model from August 5-11, 2026 (below we extend our analysis to allow for changing demand over time).83 Finally, for the expected (discounted) lifetime cost of serving all inference demand that it is profitable to serve using only flexible accelerators, we consider a range of scenarios for from $500B to $2T.84
Figure 7.1 shows the results based on the OpenRouter data, assuming (the share of tokens served on MSICs is largely insensitive to the choice of within the range from 60% to 95%85). The upshot is that assuming the long-run economic model above, the vast majority of inference tokens on whitelisted models would be served on MSICs. Moreover, there would be major cost savings over serving all inference on only flexible accelerators.86 Models with insufficient demand would not justify dedicated MSICs, but collectively such models serve a small share of total inference. Again note that our analysis holds the demand shares fixed, but since serving a model on an MSIC lowers its cost and therefore its price (assuming competition), demand would shift toward MSIC-served models. In this respect, the estimates in Figure 7.1 for the share of tokens served on MSICs are lower bounds.
The decision rule above determines whether to build an MSIC based on minimizing total cost for the assumed volume and distribution of inference demand. The question of when a market left to its own devices would build an MSIC is more complicated, since firms might not fully capture the savings of using an MSIC, whether because of competition among MSIC developers or for other reasons. But even under the worst-case parameters of Figure 7.1, the expected savings of using the MSIC exceed the fixed cost of developing it by roughly 120x for the most popular model and by at least 10x for every model in the group of most popular models that collectively serves three-quarters of all tokens. Thus, the incentives to create MSICs for the models serving most of the inference would likely exist even if only a small fraction of the savings could be captured by any single investor.
It is instructive to compare this analysis with the historical record concerning hardware for Bitcoin mining. In a span of four years, Bitcoin mining went from CPUs to GPUs (by 2010) to FPGAs (by 2011) and finally to ASICs (by 2013) (Taylor 2017); an early Bitcoin mining ASIC had 4.4 times the hashing throughput per dollar of hardware cost and was 40 times more energy efficient than a contemporaneous GPU (Taylor 2017, 64). In the case of Bitcoin, the economic incentives toward ASICs may be greater than in the case of AI models, since there is essentially just one algorithm for Bitcoin mining, which one would naturally want to hardwire into an ASIC for maximum efficiency. However, the example of AMD’s pending acquisition of Taalas (AMD 2026) suggests that even without a frozen whitelist of AI models, there are economic incentives in favor of some development of MSICs. The analysis of this section suggests that in the presence of a frozen whitelist, the economic incentives in favor of MSICs could be much greater.
7.3.2 The cost of a pure MSIC world
What would happen if inference were allowed only on MSICs? What would be the resulting economic inefficiency relative to allowing both MSICs and flexible accelerators, in the sense of increased cost of serving inference and loss of consumer utility due to some models not being served?
All of the models that justified dedicated MSICs in the “mixed MSIC-flexible world” of the previous subsection would still justify MSICs in the “pure MSIC world”, but in the pure MSIC world some additional models might justify dedicated MSICs. The reason is that in the pure MSIC world, it is not necessary to clear the hurdle of the cost-savings condition relative to flexible accelerators from the previous subsection. It is sufficient that it would be profitable to create an MSIC in order to serve a model, even if that would be less profitable than using flexible accelerators to serve the model, if that were permitted (which it is not in the pure MSIC world); and profitability was already implicitly a necessary condition of creating an MSIC in the mixed MSIC-flexible world.87 Thus, in the pure MSIC world, at least as many and plausibly strictly more models would justify dedicated MSICs.
The key difference is that in the pure MSIC world, if an MSIC is not created for a model, then demand for inference on that model is unserved. By the reasoning in the previous paragraph, Figure 7.1 gives an upper bound on the number of models that would go unserved and the share of demanded tokens that would go unserved. While the number of models that go unserved could be substantial, the share of demanded tokens that would go unserved would be at most a few percent under the worst-case assumptions of Figure 7.1 (and the unserved share of inference spending would be about 2%, as discussed in Footnote 83).
Of course, some consumers whose favorite model has insufficient demand to justify an MSIC would switch to demanding tokens from their next favorite model, which might already have an MSIC based on aggregate pre-switch demand or could justify a new MSIC based on aggregate post-switch demand. The inefficiency of the pure MSIC world relative to the mixed MSIC-flexible world depends on what fraction of consumer utility is preserved after switching to their next favorite model among those that justify an MSIC (cf. Hausman 1996; Conlon and Mortimer 2013). Given the similar capabilities of frontier models and the fact that even in the worst-case scenario of Figure 7.1, 60 models are served by MSICs, consumers may preserve a substantial fraction of their utility by such a switch, though the loss may be greater for users for whom no model with an MSIC performs well in their native language. Without studies of consumer utility functions for tokens from different models, we are not able to quantitatively estimate the utility loss implied by 3.7% of tokens (to take the worst-case scenario from Figure 7.1) switching from the consumers’ favorite models to their next favorite models that are popular enough to justify an MSIC. We encourage economists to study consumer utilities for tokens in future work.
7.3.3 Shifting demand among whitelisted models
However, what happens if changing consumer tastes lead to shifts in the distribution of demand among whitelisted models? If many copies of an MSIC are created for a once popular model, but this model becomes less popular in the future (even though the whitelist remains frozen, so no new models appear as competitors), then those chips cannot be turned into MSICs for a more popular model (cf. the literature on investment in flexible versus dedicated production capacity, as in Fine and Freund 1990; Van Mieghem 1998). A fleet of MSICs could adjust its distribution over time by increasing production and delaying retirement of chips for models with rising demand and doing the opposite for models with falling demand, but it could not do so instantly as a fleet of general-purpose accelerators could. This higher latency in responding to shifting demand would likely be a source of some inefficiency, though it is hard to estimate how much from current market data because, unlike today’s world, a paused world would see very little churn in what model options were available to customers from one month (or year) to the next. On the other hand, general-purpose accelerators pay a steep cost for their flexibility, as reflected in the 60-95% cost savings for MSICs.88
If we continue to assume , then serving virtually all whitelisted inference on MSICs costs roughly 15% of , plus design fees ranging from a fraction of a percent to a few percent of . Thus, even if demand shifted so much that all chips in the initial fleet of MSICs were turned into obsolete write-offs (say because 20% of initial demand migrated every year for five years to models not originally covered by MSICs, with price cuts unable to win back any consumers) and had to be fully replaced with new MSICs (requiring roughly the same number of designs as the original MSIC fleet), the cumulative resource cost would be roughly 30–40% of (including a second round of design fees),89 representing a 60-70% cost savings relative to a pure flexible world.
Another concern is daily fluctuations in demand. As an extreme example, suppose there is a model popular in the U.S. that receives no demand during the U.S. nighttime. Thus, the MSICs for model may achieve 100% utilization during the U.S. daytime but 0% utilization during the U.S. nighttime, for an average utilization of 50%. By contrast, a flexible accelerator could achieve 100% utilization during the U.S. daytime by serving model and 100% utilization during the U.S. nighttime by serving a model popular in China, for example. Then if the cost per token of an MSIC is 15% of that of a flexible accelerator assuming 100% utilization of both, the cost per token of an MSIC rises to 30% if the MSIC can only achieve 50% utilization, assuming conservatively that halving MSIC utilization doubles its average cost per token (excluding ). This would still represent a 70% saving for MSICs over flexible accelerators. Moreover, the utilization gap between MSICs and flexible accelerators could be smaller if retains some demand across time zones or if scheduled inference jobs with could be run overnight.
Finally, the loss of efficiency in moving from the mixed MSIC-flexible world to the pure MSIC world would have to be weighed against the benefits of the pure MSIC world in facilitating a pause. If the probability of catastrophic outcomes for humanity were much lower in a pure MSIC world (due to more reliable prevention of frontier AI training) than in an alternative world with flexible accelerators, this might swamp considerations of the economic frictions in dealing with shifting demand between models.
7.3.4 Temporary or permanent de-whitelisting of models
One possible concern about a pure-MSIC world is that if a model gets de-whitelisted (a kind of extreme case of a demand shift, where demand goes to zero by decree rather than due to changing tastes), then its MSICs are stranded capital. In this subsection, we consider the extent of this concern. The takeaway is that the possibility of de-whitelisting does not rule out the use of MSICs as inference-only chips for a hardwired pause; the cost savings from MSICs may outweigh losses from de-whitelisting, and insurance could help manage the financial risk of de-whitelisting for firms that own and operate MSICs.
First, we must distinguish between temporary and permanent de-whitelisting. When a model is de-whitelisted due to a discovered jailbreak, it might be re-whitelisted after the development of appropriate external guardrails such as an input classifier or safety harness, which inference providers would be required to run when serving the model. In this case, de-whitelisting is temporary, like most automobile recalls; the MSICs are temporarily idle but not permanently stranded.
Permanent de-whitelisting would be more costly. Suppose it is discovered that a whitelisted model poses a safety risk for which no external guardrails are adequate, so the only choice is to retire the MSICs for the model. In the worst case, suppose this happened to the most popular model. In the OpenRouter data used above, the most popular model serves 14% of all tokens, but let us suppose the most popular model overall (not just on OpenRouter) serves 30% of all tokens, i.e., . The (discounted) lifetime cost of serving the model using flexible hardware is therefore , whereas the (discounted) lifetime cost of serving the model using its MSICs is . To estimate the stranded capital cost, we introduce three more parameters, two of which are a function of the time of de-whitelisting:
is the chip-capital share. This is the share of the (discounted) lifetime cost of serving the model (not including ) using MSICs that is spent on the chips themselves.
is the stranded-chip fraction. This is the book value of model 's installed MSIC chips at time , divided by the total (discounted) amount that would have been spent buying MSIC chips for over its lifetime if it had remained whitelisted.
is the unamortized design-cost fraction. This is the unamortized share of model 's design cost at time , where the design cost is amortized in proportion to tokens served with .
Then the stranded capital cost is
where the first summand is due to wasted chips and the second is due to wasted engineering. Let us consider some cases for when de-whitelisting occurs.
Early de-whitelisting. If de-whitelisting happens shortly after production of ’s MSICs begins, then is close to 0, since not many chips have been built. On the other hand, in the worst case where the was spent on a chip platform that is not at all reusable—it is useful for only model —then is close to 1 in this scenario, since no amortization has occurred right after production begins, before tokens have been served.
De-whitelisting after initial large-scale production. If de-whitelisting happens right after the first major wave of MSIC production for , then is roughly the cost of the chips in the first wave divided by the (discounted) cost of the chips in all the waves that would have been produced had de-whitelisting not occurred. Again, in the worst case where the chip platform is not reusable for other models, would be close to 1 since few tokens have been served and hence little amortization has occurred.
Later de-whitelisting. If de-whitelisting happens a while later, then both and might be significantly less than 1. And if the went to a reusable chip platform, like that of Taalas, this would further reduce .
For a scenario in which the most popular model with gets de-whitelisted some way into the pause, suppose . Suppose our earlier variables are , and (the top of our range). Let us suppose, adversarially, that (whereas in reality energy costs imply ). Then the stranded capital cost is approximately $45B, where the part coming from wasted is negligible. We can put this into perspective in several ways. First, for comparison, the FAA’s temporary grounding of Boeing’s 737 MAX cost Boeing more than $20B in direct costs (Robison 2021). Second, let us compare the de-whitelisting scenario in a world in which is served using MSICs versus a world in which is served using flexible accelerators, which can simply switch to serving other models when is de-whitelisted. Let be the fraction of the model's lifetime demand served before the de-whitelisting at time . If we live in a world in which is served only by flexible accelerators, we would spend on serving before de-whitelisting. If we live in a world in which is served only by MSICs, we would spend plus the stranded capital cost of $45B for a total of before de-whitelisting.90 It follows that if ,91 then even with the stranded capital, serving model with MSICs before de-whitelisting is cheaper than serving model with flexible accelerators before de-whitelisting.
The firms that own and operate MSICs incur the losses when a model is de-whitelisted, which raises the worry that they will lobby heavily against any de-whitelisting. To reduce this pressure, one possibility would be to require firms to buy insurance covering losses from de-whitelisting as a condition of their license to serve inference with MSICs.92 The insurance premiums would depend on the book value of the MSICs covered and the assessed risk of de-whitelisting, helping to limit the incentives of firms to overinvest in MSICs for models at risk of de-whitelisting. The insurance fund would compensate a firm for the book value of its MSICs in the event of de-whitelisting, but only upon return of the MSICs to authorities. The treaty authority could serve as the insurer of last resort. Coverage would carry a deductible, so operators would maintain an interest in the likelihood that the models they serve will remain safe. Note that the purpose of the insurance requirement would not be to make operators whole but rather to make the treaty authority's whitelisting decisions more independent of firms' financial interests.
7.4 Impacts on concentration of power
Questions about how AI might shift the distribution of power have permeated discussions about the modern AI industry since its infancy. Bostrom (2014) discussed at length the possibility that a superintelligent AI system could rapidly become a permanent global dictator or “singleton;” Zuboff (2019) discusses how data-driven AI systems fueled by surveillance capitalism can concentrate power in the hands of tech platforms; and Davidson et al. (2025) discussed the potential of AI to contribute to sudden coups. Several AI companies, most explicitly OpenAI, were purportedly founded to reduce the risk that other AI companies would become too powerful (Hao 2025). As AI grows more powerful, the fear that it or those who control it could permanently concentrate power could itself be destabilizing (Hendrycks et al. 2025; Wright 2026); OpenAI CEO Sam Altman has referred to AI as having a “real ‘ring of power’ dynamic” that “makes people do crazy things” (Altman 2026b).
AI pause proposals address this “ring of power” dynamic by preventing (or at the very least delaying) any such ring from being forged in the first place. The AI industry has already shifted power significantly. The push to embed AI across society has concentrated power and influence in the hands of a small group of AI companies and their leaders (Brennan et al. 2025). Addressing this problem and preventing further concentration of power have been cited as key motivations by some pause proponents (see, e.g. Larsen et al. 2026; Office of Senator Bernie Sanders 2026).
Any plan for how humanity should address the challenges posed by AI must contend with questions about how it might affect the distribution of power; this is also true of plans to defer regulation, since doing so might lead to heightened risks that necessitate more onerous restrictions later. AI models that suddenly empower a large number of individuals to design new viruses or launch sophisticated cyberattacks could lead to calls for widespread surveillance just at the moment when AI permits more comprehensive evaluation of surveillance data than was previously possible (cf. Bostrom 2019).
A hardwired pause in and of itself would place no new restrictions on how people use their computers, save for content-neutral restrictions on the training of large enough AI models. If it were successfully implemented before widely available models became too dangerous, a hardwired pause could help to preserve liberty by forestalling the need for surveillance measures to prevent people from downloading and using those models.
On the other hand, a hardwired pause would involve imposing new constraints on liberties that individuals and businesses currently enjoy, such as the liberty to buy an unrestricted number of accelerator GPUs, to train AI models of unlimited size, or to operate chip fabs and data centers without being subject to inspections. While these impositions are by no means cost-free, they do not seem to us qualitatively more onerous than constraints that have been placed on other technologies when those technologies grew powerful enough to be dangerous. In the past, scientists could own as much uranium as they desired, people could drive their cars as fast as they wanted on the highway, and chemical manufacturers could produce phosgene without being subject to inspections, but today these activities are regulated by the government and international treaty organizations, with criminal prosecution for some violations. Impositions on these particular liberties have come to be seen by most as acceptable tradeoffs for the added security that regulation of dangerous technology can bring.
Another serious concern one might have about a hardwired pause is its potential to severely and even permanently entrench market power with the incumbent AI developers who have the most dominant market positions today, especially OpenAI and Anthropic. Like any regulatory body, the body responsible for whitelisting models would potentially be vulnerable to industry capture, and concentration of the market in a few AI models would come with the additional risks of entrenching the ideological or commercial biases of the people who built them, or of proliferating any loyalties, vulnerabilities, or backdoors in the models (Buyl et al. 2025; Lehr et al. 2026). Any AI governance regime must take such concerns seriously: whitelisting processes and other safety regulations must be independent and as transparent as possible, and monopolies or oligopolies imposed by government fiat must be regulated like other public monopolies to protect consumers and prevent exploitation of market power. Successfully navigating these issues would be critical to the legitimacy, and consequently to the durability, of a hardwired pause.
7.5 Consequences of receding knowledge cutoffs
A pause on frontier training that involves freezing model weights would prevent models from being updated to change the cutoff time of their pretraining data (which we refer to informally as a “receding knowledge cutoff”). It also raises several issues, including vulnerabilities, data poisoning, and locked-in values. More generally, it would prevent models from being trained on a continual or even continuous basis, as frontier AI companies aspire to do (Favaro and Clark 2026; Behrouz and Mirrokni 2025; Pachocki 2026).
Receding knowledge cutoffs could have costs in terms of model accuracy and bias (Dai et al. 2025). However, since retraining models is costly even in today’s pre-pause regime, the receding knowledge cutoff issue is not a new problem. AI developers today use methods such as internet search and Retrieval-Augmented Generation (Lewis et al. 2021) to improve up-to-dateness while minimizing costly retraining.
A related issue, discussed in § 7.3.4, is the potential for new vulnerabilities to be discovered in existing models. Frontier AI models are typically deployed in tandem with “guardrails” consisting of smaller models that monitor their input, output, and/or internal computations; and vulnerabilities are typically patched by updating a model’s guardrails rather than the frontier model itself. For guardrails that are small enough and do not monitor internal computations, this practice could continue freely under a hardwired pause. Monitoring internal computations (called “probing”), however, would require modifications to inference chips that could potentially facilitate GEMM-laundering (see § 4.1), meriting greater scrutiny, and we have not yet conducted an in-depth analysis of exactly what sorts of guardrails could be feasibly combined with a hardwired pause without requiring specific carve-outs. It is also not guaranteed that updating guardrails will prove adequate to address novel vulnerabilities. This concern motivates maintaining the ability to “de-whitelist” models as discussed in § 7.3.4, as well as other measures to decrease societal reliance on whitelisted models, e.g. maintaining a fallback option in contexts where models play critical roles, in case new vulnerabilities are discovered.
Another design option noted in § 3.2.4 that could address both receding knowledge and novel vulnerabilities is to allow for whitelisted fine-tunes; for example, AI developers could submit training data and algorithms for LoRA fine-tunes to treaty authorities, who would then perform the fine-tunes and evaluate the resulting models. Some researchers have found that LoRA fine-tunes are sufficient for many safety updates (Xue and Mirzasoleiman 2025).
As it is unlikely that the options above would achieve the same effectiveness as a full new training run would, we expect that receding knowledge cutoffs would still somewhat impair models’ ability to be fully situationally aware about current events or fully knowledgeable about recent advances in technical disciplines such as computer science and mathematics. The effect on palatability could be positive or negative: while some users would no doubt find even mild impairment frustrating, others might find in it a silver lining, especially if they worry about loss of control, human unemployability, gradual disempowerment, and other ills that could arise from models being too powerful.
Receding knowledge cutoffs could also be a double-edged sword for treaty reliability and durability. By gradually eroding the power of the best available models on the legitimate market, they would potentially create black market niches for sub-frontier models with more updated knowledge. On the other hand, preventing models from fully updating on technical breakthroughs in computer science or synthetic biology could plausibly be desirable if it slowed the proliferation of training efficiency gains or dangerous biological abilities.
In addition, insofar as whitelisted models are expected to be produced and deployed at scale, the incentives for states, AI companies, and even other AI models, to influence their biases or loyalties, either directly and overtly or surreptitiously using methods like data poisoning, backdoors, or secret loyalties, could be greater than they already are in the pre-pause regime (Banerjee and Aarne 2026).
Finally, insofar as it involves freezing model weights, a pause of frontier AI training has the potential to lock in the worldviews, biases, and value systems embedded in the whitelisted models at the time of the pause. Bias is already a problem with current systems (Buolamwini and Gebru 2018; Barocas et al. 2023; Buyl et al. 2025), which simply re-training is not guaranteed to solve. Instead of attempting to make models “unbiased” enough to count on, some people may opt for a different path during a hardwired pause: one where they do not rely on training bigger and more general AI models to be “unbiased” or good sources of moral progress in the first place. This different path has been articulated before by those who have argued: that we need “research directions beyond ever larger language models” not least because “size doesn’t guarantee diversity,” because large language models encode bias, and because methods reliant on them “run the risk of ‘value-lock’, where the [language-model]-reliant technology reifies older, less-inclusive understandings,” interfering with the key role language plays in social movement formation (Bender et al. 2021); that three of the leading approaches to value alignment “fail to accommodate reasonable moral disagreement, since they provide neither good epistemic reasons nor good political reasons for accepting AI systems’ morally controversial outputs” (Schuster and Kilov 2025); and that technical artifacts have politics and actively shape power, both through intentional decisions and by necessity (Winner 1980).
7.6 The need for other measures
A hardwired pause is designed to complement, not preempt or replace, measures to prevent and stop risks and harms from current-level models, which are well-documented (Responsible AI Collaborative n.d.; Weidinger et al. 2021). For example, AI chatbots have been linked to at least a dozen deaths by suicide (McHugh 2026), including the death of 16-year-old Adam Raine. According to a lawsuit filed by his parents (Raine v. OpenAI, Inc. 2025), Raine started using ChatGPT for homework, and later the chatbot allegedly encouraged and coached him in detail to kill himself, even offering to help draft a suicide note (Hill 2025). As another example, AI is being used for military intelligence, targeting, and to generally compress the “kill chain,” as has occurred in Iran with U.S. use of a system called Maven Smart System to help it strike more than 1,000 targets in the first 24 hours of the Iran war, including a school strike that killed at least 123 children according to Bloomberg's reporting on an internal Pentagon investigation (Bartenstein and Karra 2026).
During a hardwired pause, if models similar to current frontier models are whitelisted, then we may continue to see consequences like the following unless additional safeguards are implemented during deployment:
A current-level frontier model or many instances of the model are used in a way that allows them to send messages outside of a supposedly isolated testing environment, interact with the internet, and do something like hack a company, as recently-trained models from OpenAI have done (Newman and Cameron 2026).
Dependency on AI tools powered by current-level frontier models in companies and governments is so great that it becomes nearly impossible to run these without over-relying on and deferring decisions to AI (Kulveit et al. 2025).
Development and use of military tools powered by current-level frontier models escalates as both sides increasingly hand over control to gain an edge. There are well known documented biases in frontier models related to escalation (Rivera et al. 2024), and mounting evidence of automation bias that could lead to escalation (Horowitz and Kahn 2024).
Nonetheless, pausing the frontier would help prevent the above from being even more likely and more severe, as they would be with models whose capabilities are even greater. And, as discussed in § 7.3.4 and § 3.2.2, respectively, the bar for whitelisting could be raised, and the threshold at which whitelisting is needed could be lowered over time if needed, e.g., if it becomes possible to create models with dangerous capabilities with a level of compute below the initial threshold.
Still, in the worst case, a hardwired pause might fail to mitigate critical risks. Existing methods of evaluating models are imperfect and could misestimate critical safety-relevant properties, leading to a false sense of security regarding whitelisted models. There is also the issue of algorithmic progress: Advances in scaffolding and elicitation techniques might produce dramatically more capable AI systems using whitelisted models; training might become much cheaper and/or easier to perform in a distributed manner. Sufficiently advanced AI systems might dramatically accelerate algorithmic progress and could even lead to the development of entirely new AI paradigms. While we did account for training efficiency gains within the current paradigm in §§ 6.3 and 6.5, no one can predict what training efficiency gains would be after a pause. Such possibilities may become more likely over time as frontier AI training continues, which would make pausing sooner a more robust approach to mitigating these risks. More stringent restrictions on computing hardware might also further reduce such risks.
8 Pause-willing futures: conditions favoring international AI restraint
In the previous sections of this paper, we analyzed the feasibility and implications of a hardwired pause in pause-willing futures, as defined in § 1.2. In this section, we consider some of the conditions under which the present might evolve into a pause-willing future.
8.1 Introduction
The United States and China are engaged in robust and multi-faceted geopolitical competition. This fierce rivalry has economic, military, political, and technological dimensions, and AI features prominently in all of them (see, e.g. Brandt et al. 2026). Consequently, the idea that either the U.S. or China would agree to pause further development of its frontier AI models, let alone negotiate and ratify a verifiable, reciprocal agreement might seem chimerical. Nevertheless, history makes clear that international agreements that may seem impossible or fantastical at one political juncture become feasible reality in others. So, the real questions that bear asking are not whether pausing is possible, but rather what factors matter in enabling creation of a sustainable pause, and under what conditions will they pertain; through which mechanisms pause-willingness might increase, and how; and what international and domestic political factors might enable credible sustainment of a pause, given the generalizable and particular characteristics of the United States and China.
Additionally, unlike many international agreements of the past, states are not the only actors of note: an AI frontier model pause necessitates that states regulate technologies that have been created and are controlled by a small collection of private firms rather than by governments (at least as of this writing). These companies are neither signatories of, nor parties to, international treaties or other agreements. They have enjoyed largely unfettered access to government decisionmakers and have heretofore suffered minimal penalties for failing to prevent or curtail problematic behaviors by their models. They have likewise been reluctant to share critical information when such incidents occur (Mitchell 2026). These firms also possess significant financial resources; this affords them some freedom to withstand or sidestep meaningful constraints that they find economically or strategically inopportune. Consequently, the extent to which an AI pause is achievable and sustainable depends not only on its attractiveness to the United States and China; it is likely to depend on its attractiveness to frontier AI firms as well.
Thus, pause willingness depends on the extent to which a pause is viewed as simultaneously strategically rational, domestically palatable and sustainable, and technically and practically enforceable—both for states and for key private actors. In the sections that follow, we outline the factors that will inform how this trifecta of conditions can be effected and what actors, institutions and international and domestic political conditions are required to sustain it. Our analysis proceeds in four sections. First, we identify a set of theoretical drivers of pause-willingness. They comprise fundamental changes in AI capabilities; shifts in the international distribution of power and capabilities; and alterations in national political environments. Second, we offer a framework for making sense of the political scaffolding that would undergird pause-willingness, organized around strategic motivations, international and domestic politics, and the roles played by other key actors. Third, drawing upon material in the preceding sections, we explore a set of five scenarios through which pause-willingness can materialize and be maintained. Fourth, and finally, we conclude by shifting from generalizable variables to focus on the specific strategic contexts of the United States and China.
8.2 Theoretical drivers of pause willingness
This section outlines pause-willingness across different analytic levels, namely: technology; the international system; and domestic politics. With respect to technology, we explore the ways in which developments in frontier AI capabilities, including scaling law evolution, enhancements in compute efficiency, and algorithmic improvements influence the viability and design requirements needed to achieve a sustainable AI pause framework (Mollick 2024; Erwan 2025). At the systemic level, we draw on theories from international relations and security studies to analyze how the structure of great power competition creates acute incentives to aggressively compete, but also can give rise to conditions for cooperation, depending on the relative distribution of power, the costs and benefits of competition, and expectations about the future. At the domestic level, drawing on Putnam’s theory of two-level games and broader theories of domestic coalition formation and private authority, we identify key variables that determine when and under what conditions international incentives for restraint can be translated into credible commitments. As suggested above, pause-willingness is most likely to emerge when pause-compatible changes at all levels reinforce rather than cut against one another.
8.2.1 Fundamental changes in the technology
One primary factor shaping pause-willingness is the state of AI technology itself. Changes in the technology or its use could reshape (1) the marginal value of racing, (2) the dangers created by racing, and (3) the feasibility of reciprocal restraint. This section considers these three dynamics across five illustrative technological trajectories. Given that the contextual application of technologies is a major determinant of their risk profile, this analysis will consider AI as a broader sociotechnical system (Bijker et al. 1987). This means that relevant technological shifts could be concrete changes in the capabilities of models but could also be changes in their use or application.
First, frontier progress could encounter a substantial plateau or a “wall,” reducing the marginal value of accelerated frontier model training. This could happen, e.g., if additional compute, data, algorithmic advancement, or capital start to generate sharply diminishing improvements in strategically relevant capabilities.
An agreement to stop advancing the AI frontier could be more politically costly for leaders if they believe that continued research could give them near-term decisive military, intelligence, scientific, or economic advantages. It is substantially less costly if additional investment is expected to yield only incremental improvements, perhaps opening the door to pause willingness even in otherwise challenging national or international incentive contexts. Slower discovery of significant capabilities could also facilitate better safeguards and more time to plan robust enforcement mechanisms, with the latter perhaps further incentivizing an international agreement pausing AI research in areas where higher risks remain. In short, a plateau would decrease the benefits of racing.
Second, the opposite trajectory, which appears more likely at the time of writing, could also shape and enhance pause willingness. Capabilities could increase so rapidly that the expected cost of continued development raises questions about the advisability of continued research even in political contexts that would otherwise incentivize it.
The political effects of this change would not depend on the absolute dangers of advanced AI but on whether these additional capabilities create forms of mutual vulnerability that neither side could reliably escape through unilateral action. Whether rapid advances result in intensified racing dynamics (perhaps on defensive research or action motivated by a “closing window” effect) despite these risks would likely depend on the broader international system context. In contrast to a plateau, which would decrease the benefits of racing, an acceleration could increase the costs of racing. Because the development of technologies is not teleological, it is possible that some capabilities may plateau while others improve rapidly, perhaps exercising independent effects on risk and benefit in various sectors.
Third, AI development could produce consequential shifts in the offense-defense balance of AI tools across domains. If AI capabilities confer significant defensive advantages in, for example, bio and cyber risk, this may reduce incentives for pausing while increasing pause willingness. If AI capabilities are on balance defense dominant, potential aggressors may anticipate fewer strategic benefits to be derived from further advances. Leaders in turn may simultaneously worry less about the likelihood of offensive uses of AI and become more confident in their ability to defend against them, thereby decreasing their incentives to pursue a pause while concomitantly increasing their willingness to agree to one. At the same time, significant increases in offensive capabilities may increase desire for a pause while increasing the costs of adversary defection from a pause regime. If AI capabilities significantly increase with respect to cyber attacks, intelligence collection, military targeting, operation of autonomous weapons, or the discovery of security vulnerabilities, leaders may fear that falling behind even briefly could create acute strategic risks. Such a development could intensify the race rather than facilitate restraint. But if offense dominance becomes sufficiently severe on both sides, particularly if neither side can plausibly secure a lasting lead in AI capabilities (see below), offense-related harms could facilitate national demand for a pause in a vacuum and the resulting mutual vulnerability could cultivate the pause willingness needed to bring this about through international agreement.
Fourth, it is possible that the “brittleness” of an advantage in AI capabilities may shape pause willingness. If advantages in AI capabilities are tenuous, possessing a temporary lead may become less strategically valuable, perhaps clearing a path to pause willingness even as national and international dynamics remain otherwise consistent. Conversely, technological developments that make capabilities highly excludable or self-reinforcing could increase fears that a temporary lag will become permanent, sharply reducing willingness to pause.
Finally, changes in the technology or its integration could shape the enforceability of an agreement, thereby affecting pause willingness. To offer just one example: a capability frontier that is dependent on large, centralized training runs, which require scarce and observable hardware, would be far easier to monitor than a corresponding frontier whose capabilities can be generated through the use of small, distributed or highly-compute efficient hardware systems. While this latter paradigm is not in view at the time of writing, unforeseen technological developments are not without precedent. Taken together, technological change, to include changes in broader socioeconomic context, can make a pause more plausible through at least the three mechanisms mentioned above: reducing the expected payoff from continued racing, increasing the shared dangers produced by continued racing, or making reciprocal restraint easier to verify and enforce. Importantly, the same technological development can push these mechanisms in opposite directions. The determinant for pause willingness is the balance of incentives for key actors in the broader international system shaped by these technological changes.
8.2.2 Changes in the international system: cooperation in the shadow of great power rivalry
The structure and distribution of power within the international system provides the underlying strategic environment within which decisions about frontier AI development are made. Changes in the system may not be determinative vis-à-vis pause-willingness, but shifts in the distribution of power and the importance of such shifts materially influence the incentive structures that states face when weighing the costs and benefits of an AI pause. This is because one, or maybe even the, fundamental challenge to reaching and sustaining pause willingness is that willingness is inversely correlated with perceived strategic advantage. However, shifts in the balance of power that instigate a rethinking of the strategic advantage bestowed by AI racing and reduce the expected gains relative to potential costs can enhance the viability and durability of pause willingness. This section discusses some of the key systemic variables that bear on AI pause willingness, by drawing upon the work of international relations (IR) theorists who have advanced a variety of frameworks that help explain why great powers might agree to hit pause on a technology they each regard as strategically consequential. Four are of particular relevance for an AI pause: power transition theory and its relationship to the security dilemma; hegemonic stability theory; polarity; and alliance dynamics.
Power transition theory and the security dilemma—heightened dangers open up possibilities
Power transition theory posits that the probability of conflict—and, by extension, the ease or difficulty of cooperation—varies with the distribution of capabilities between a dominant state and a rising challenger (Organski 1958; Organski and Kugler 1980; Tammen et al. 2000). As a rising power approaches parity with a dominant state, conflict becomes more likely, especially if the rising power is dissatisfied with its relative position. Applied to AI, as the trailing power closes the gap with the leading power in frontier AI capability, the danger of conflict mounts. Moreover, what are known as “security dilemmas” may be particularly intense: in periods of high uncertainty, when it is hard to distinguish offensive weapons from defensive weapons, and the offense has the advantage, actions taken by an actor to increase its security and enhance its position will perforce reduce the other state’s sense of security, triggering countermeasures that leave both sides less secure than before (Jervis 1976; Garfinkel and Dafoe 2019; Rand 2026).
However, paradoxically, at the same time as growing parity creates greater dangers, it simultaneously also heightens the possibility for cooperation and, in this particular case, AI restraint. This is because in situations when neither side can achieve a decisive and durable advantage over the other, the opportunity costs of a pause decline for both states simultaneously. Thus, should circumstances emerge and persist wherein both the United States and China conclude that neither can achieve a permanent, self-reinforcing frontier monopoly—and the costs of continuing to try will be high and counter-productive—then pause-willingness may be durable. The fact that AI is a general-purpose technology in addition to a military one complicates such a calculation significantly (Horowitz 2018, 2026a; Scharre 2023a); nevertheless, when the costs of continued racing are potentially catastrophically high, pause willingness is possible.
As suggested above, power transition theory further posits, however, that (near) parity is insufficient on its own to enable cooperation. The rising power’s degree of satisfaction with the prevailing situation is also key (DiCicco and Levy 1999; Tammen et al. 2000). Should the rising power believe that the dominant power is supportive of a pause in order to lock in its own supremacy and the rising power’s subordination rather than supportive because a pause can be a mechanism for managing shared risk, then AI pause willingness is likely to be tenuous and brittle. Under these conditions, (near) parity may generate technical pre-conditions for a pause but fail to create the requisite political conditions to sustain it, due to worries about what the future might hold for the subordinate power (see, e.g. Powell 2006). The AI-specific implication is that international system-level changes favorable to pause willingness may require accompanying shifts in the dominant state’s behavior so that the rising power has a more salutary assessment of its position in the prevailing international order and, in turn, an AI pause.
Finally, should we ever reach a point wherein there is deep and irreducible uncertainty about which state is dominant or whether AI parity has indeed been reached, this might ironically create the highest likelihood that pause willingness could be sustained. In the presence of high uncertainty, when neither power can confidently assess its own position relative to the other, the benefits of continuing to race become difficult to assess, while the benefits of mutual restraint as a mutual risk management strategy may become correspondingly more clear.
Hegemonic stability and the strategic value of institutional governance
Another IR framework, hegemonic stability theory (HST), offers a complementary perspective. HST posits that international regimes are most reliably produced and maintained when a dominant power has the capacity and will to supply the institutional infrastructure that cooperation requires (see, e.g. Gilpin 1981; Kindleberger 1986). Applied to a sustainable AI pause, HST holds that the frontrunner and dominant AI power is the most natural “supplier” of the institutional architecture for a pause. However, this condition only holds if the dominant power concludes that acting as a supplier serves its long-term interests rather than eroding its competitive position and only if the supplier can convince the trailing power that it is not exploiting its role to keep the subordinate power down. Moreover, a declining hegemon facing a credible rising challenger may become increasingly resistant to providing governance goods, fearing that institutionalizing restraint will accelerate its decline.
Here again, a paradox arises. The leader has the most leverage to craft a sustainable and favorable pause at a time when it retains the most significant capability advantages and the greatest control over the AI stack. However, this is also the period wherein the leader possesses the fewest incentives to accept the mutual constraints that an effective and sustainable pause demands. On the flip side, the period in which the dominant state might be most amenable to accepting constraints—i.e., after its position has further eroded and the costs of further competition mount—may be the period when it possesses the least capacity to induce sustainable compliance on the part of the rising power.
What this means for pause willingness is that the window for reaching a sustainable agreement may be constrained in time and scope. Under the right conditions, however, there are potential options for expanding the cooperative “window of opportunity”. One such possibility arises if the dominant power can be persuaded that assuming the role of “institutional supplier”—and thereby designing and providing the institutional scaffolding by which an AI pause can be initiated, verified and sustained—can itself be a source of power and long-term strategic advantage (see, e.g. Scharre 2023a, 2023b). Such arguments were powerful during much of the post-World War II period, when the US served in analogous roles in many international financial and security-focused institutions (see, e.g. Brooks et al. 2013; on adverse consequences that can arise in the absence of a hegemon, see again Kindleberger 1986).
Polarity and possibilities for collective action
At present, the United States and China are the only two states (to the best of our knowledge) that are developing the full stack of capabilities necessary for training the most advanced AI models (cf. Scharre 2023b). However, these two leading states exist alongside an increasingly multipolar distribution of near-frontier capabilities, advanced compute infrastructure, and regulatory input (Horowitz 2018, 2026b; for an exploration of potential future international distributions of power tied to AI, see Pavel et al. 2025). Both history and the existing literature on negotiation and cooperation tell us that assuming a bipolar AI order vastly simplifies collective action problems. Collective action problems emerge when individual actors, behaving rationally in pursuit of their own self-interest, face incentives to make decisions that can be detrimental to the interests of the collective, or society at large, especially when the costs of taking actions that are beneficial to the group are not evenly borne (Olson 1965). If only two actors are involved, reaching and sustaining agreements should be easier for several reasons: Negotiations are simpler between two parties; each actor’s behaviors are more impactful; compliance and defection are easier to monitor; verification is likewise easier to track; and, finally, free-riding is easier to police and prevent (ibid. Oye 1986).
At the same time, the aforementioned evolution in the AI ecosystem has led to a proliferation of actors, whose own behavior can materially affect whether any bilateral agreement forged between the leading AI powers would be meaningful, sustainable, and enforceable. Any bilateral pause agreement that fails to take into account middle powers, critical nodes in the AI supply chain, the frontier labs themselves, among other key potential spoilers, is likely to rapidly degrade or collapse, as a consequence of diffusion, incentives to defect and/or free ride, impediments to effective verification, enforcement complications, etc. (see again Oye 1986).
Conversely, however, those states and private and public actors who control capabilities and can exercise leverage over supply chain nodes, also possess the power to act as critical verification and enforcement partners for a pause. And, if an agreement is properly constructed, their incentives—like the incentives of the leading AI powers—can be structured to enhance and support a pause, rather than to undermine it. In sum, the shifts in the international system that might matter materially for an AI pause are not simply changes in the distribution of power and capabilities of the leading powers, but also changes in how third-parties—states and non-state actors—view their own interests vis-a-vis great power AI competition and cooperation.
As alliance theory suggests (see, e.g. Walt 1987; Christensen and Snyder 1990), the more the international system hardens into competing AI blocs/alliances, the more difficult it will be to maintain and sustain AI pause willingness. An “us vs. them” world in which all states are forced to choose sides, and in which frontier AI access is predicated on rigid alignments, leaves few possibilities for creation of the kind of mutually-beneficial, cooperative agreement and architecture that a sustainable pause demands. Systemic shifts that reduce the bilateral (or bloc) nature of AI alignment—and thus eschew bloc logic and are genuinely multilateral—create critical space for inclusive and enduring agreements, a scenario we explore in some detail below.
8.2.3 Changes in national political environments
International agreements on frontier AI will not be negotiated by monolithic states, behaving as rational, utility-maximizing actors. Instead, any pause agreement will be agreed (or not) by governments, whose choices are informed, shaped, restrained, or undermined by the domestic political environments in which they operate. The previous two sections focused on the technological and international systemic factors that inform perceptions of the strategic costs and benefits of a pause. Those variables notwithstanding, whether shifts in incentives and cost-benefit calculations can engender buy-in and sustained commitment to a pause will depend fundamentally on whether the governments involved—be they democratic, autocratic or somewhere in between—can engineer and maintain the domestic political conditions necessary to negotiate, ratify, implement, monitor and sustain a pause. This section explores key domestic-level conditions that can enable or impede an AI pause.
Putnam’s (1988) now-classic “two-level game” framework offers analytical leverage for making sense of how, why and under what conditions domestic political contexts shape international agreements (and, conversely, how international agreements can affect domestic political outcomes). Putnam posits that negotiators function on two levels simultaneously, the international and the domestic, and the practicability and viability of any agreement forged on the international level is ultimately dictated by what is functionally possible on the domestic level. In Putnamesque terms, the “win-set” of any international negotiator—i.e., the range of agreements to which they can credibly commit—is thus constrained by what the relevant domestic constellation of actors will accept. Any changes that serve to expand a domestic win-set will make sustainable international agreement more achievable, while changes that reduce a domestic win-set will make international agreement harder.
Thus, in addition to the technological and international factors discussed above, for a pause to be actionable and sustainable there must be sufficient support from relevant domestic actors and agencies. Win-set relevant domestic actors include governments, national security agencies, military establishments, legislatures, and the general public as well as, quite critically, AI-focused firms—the primary developers, owners, and deployers of the technology being governed.
Pause-willingness at the national level could arise in a number of ways, and the route by which major states arrive at pause-willingness is likely to shape the design and durability of an eventual pause agreement. These paths to pause willingness themselves will be shaped by the public and private actors described above in addition to relations between major powers.
The paths to and durability of pause willingness will be shaped by the complex interplay between strategic interests of state actors, particularly the United States and China, and the capacity and willingness of governments to regulate, monitor and enforce frontier AI development. Durability may also depend on the inclusion in negotiations from the start of a broader array of states, jurisdictions and private actors within those territories due to their role in international norm setting, supply chain production, cloud infrastructure, and capital investments. Domestic dynamics differ and depend on the influence that private actors wield on public opinion (especially in democratically-elected governments), consumer behavior and importance to national economy and national security. As a result, a pause-willing future depends not only on international politics, but also aligning incentives and security buy-in from private actors and publics.
8.3 More than willingness: conditions for durability
A pause-willing future requires alignment of several different conditions. Perhaps most obviously, it requires belief in major states that reciprocal restraint is preferable to continued racing. At the same time, as described above, private firms control much of the frontier AI development landscape and are likely to shape international agreements through overt influence or through genuinely held beliefs about their economic and/or security importance within major states.
Given the complexity of aligning incentives, there are several variables that are likely to shape the durability of an international pause agreement. First, states’ and industry’s perceived value of frontier AI for economic, commercial, scientific, national security, geopolitical and/or other strategic advantages shape the incentives they have to negotiate, agree, comply, abide by, and enforce an agreement. As shared perceptions of value increase, pause willingness is likely to decline. Furthermore, value perception may drive dynamic competition between states and competition between private actors.
Second, perceptions of systemic and/or catastrophic risk, by governments, private actors, and expert communities may vary depending on both available information and the political salience of risk in a given time. As perceptions and salience of risk increase, so does willingness to negotiate, stay, and enforce. As perceptions and salience of risk decrease, the incentives to cheat or exit increase. Furthermore, beliefs about benefits and risks within frontier labs and competition among labs can alter incentives for industry to comply with or cheat any agreement. The belief and action of frontier lab actors in particular points to significant needs to increase the capacity of governing bodies and independent watchdogs to monitor and, more importantly, the capacity to enforce a pause agreement and its provisions. Information asymmetries between frontier lab capabilities and operations and those of government and independent watchdog agencies will impact the ability to monitor and verify provisions, and to detect cheating, hidden scaling or covert training runs.
Third, domestic political coalitions are fluid. Especially in countries with democratically-elected governments, pause-willingness and durability depend upon achieving and maintaining support from domestic governing institutions, industry, publics and civic society. Even leaders in countries that do not rely on elections also must keep domestic constituencies satisfied. Thus, durability may be exposed to shifting political priorities and depends on who holds power and what degree of support a governing coalition has.
Fourth, any agreement must cover not only the lead states and actors in frontier model training, but also the states, jurisdictions and private actors along or potentially along the supply chain. Durability depends on ongoing capacity to monitor and control chokepoints, including, among other things, chip production and cloud and energy infrastructure. In addition, the possibilities of a new competitor forming in a nonparty to an agreement and talent flight of researchers and engineers also pose a risk of regulatory safe havens.
Of course, no single factor is sufficient to determine movement toward an agreement or its durability. Instead, the shape and likelihood of an agreement corresponds to combinations of movement across these variables. The following figure describes the broad categories of factors that may shape appetite for specific measures in an agreement and the durability of such an agreement.
8.4 Imagining pause willing futures
In this section, we draw on the technological, international, and national-level theoretical drivers described in the preceding sections to explore pause-willing futures through contextually-grounded narratives. The following scenarios are structured explorations intended to demonstrate paths to pause willingness and identify key dynamics in the development of pause-willingness across major national actors. The narratives in this section are not predictions, and we take care to avoid assessing their likelihood. Rather, they are analytically grounded narratives designed to communicate paths towards pause-willing futures.
8.4.1 The expected costs of continued development have increased
Perhaps the most straightforward path to pause willingness is that the expected costs of continued frontier AI development could rise sufficiently such that both the United States and China begin to view unconstrained competition as less desirable than reciprocal restraint. Even in a future in which AI research or the tools provided by it remain valuable in absolute terms, this value may not merit increasing risks. The relative costs measured in economic stability, national security, or political cohesion might produce pause willingness even as large states continue to believe that advanced AI capabilities confer strategic advantages.
Of course, this does not mean that a heightened risk profile would be sufficient to produce pause willingness. If policymakers believe that advanced AI capabilities remain more dangerous in the hands of an adversary than in general, they may pursue unilateral countermeasures or deepen their commitment to racing. Given this possibility, pause willingness depends not only on the balance of risks and benefits of continued frontier AI development, but on the source and nature of the foreseen AI risks.
Reassessment of risk, perhaps leading to pause willingness, could result from a single, focusing event or from a gradual shift in belief. A shared focusing event that sharply changes beliefs about frontier risk makes the easiest example: a serious consequence, perhaps in the cyber or biological domain, could provide concrete evidence of high-impact AI harms and inject this evidence into important policy communities by virtue of its prominence. If mitigations for the resultant harms appear limited and the potential costs of recurrence are sufficiently severe so as to eclipse the broader benefits further frontier AI development may confer, the event could produce pause willingness in important policy communities. At the same time, a focusing event like this could raise awareness of risks in the general public and create political space for policy intervention. If important communities in both the U.S. and Chinese polity interpret the event as evidence of shared risk, the event could provide a practical and rhetorical basis for previously difficult negotiations.
While this “fulcrum” of negotiating context may ease paths to pause willingness, it is not strictly necessary. It is also possible that a series of smaller developments and their communication within relevant policy communities could convince important constituencies that systemic risk is rising faster than it can be managed, and critically, is outstripping the benefits of frontier AI development. It is easy to imagine that continued improvement in offensive cyber capabilities might generate increasing concern about frontier cyber risks. These increasing concerns might inspire a shift from asking whether AI could create severe vulnerabilities to asking whether existing institutions can manage realized risk. If mitigation does not appear sufficient, both expert communities and the broader public may shift attitudes towards pause willingness.
8.4.2 The perceived value of “winning the AI race” has fallen
What if, regardless of costs produced by risk, the expected benefits of AI vary? A second path to pause-willingness emerges if the expected benefits of continued frontier AI development decline sufficiently such that the U.S. and China become more willing to accept reciprocal restraint. In contrast to the previous scenario, this pathway does not require policymakers to become substantially more concerned about the risks generated by advanced AI. Instead, the strategic value associated with continued frontier AI development falls. AI could remain an important general-purpose technology with substantial strategic implications while policymakers become increasingly skeptical that pushing the frontier will create benefits commensurate with the resources required to do so and the risks produced by frontier capabilities. In this class of pause-willing futures, the opportunity cost of a pause falls, and given the material and risk costs of frontier AI development, restraint becomes easier to contemplate because key actors believe they are giving up less by forgoing frontier capabilities.
As with increasing costs, declining benefits do not necessarily produce pause willingness. Because the relevant comparison is between the strategic consequences of continuing and stopping relative to a potential adversary and not simply between the expected benefits of additional AI research and the risks associated with it, even modest improvements in frontier capabilities could sustain intense competition if policymakers believe that small capability advantages confer important military or economic advantages. Again as above, a pause-willing future of this type requires a convergence of U.S. and Chinese assessments, with both concluding in the broad strokes that mutual reciprocal restraint would sacrifice less than continued competition.
One pathway to a mutual reassessment of this kind is that progress on the AI frontier becomes increasingly difficult or that frontier returns become decreasingly attractive in strategically consequential terms. Additional investments in compute, data, energy, talent, or algorithm research could continue to produce model performance without producing commensurate improvements in business value, military capacity, or traction on scientific problems. It is also possible that scaling laws or unforeseen development challenges may cause direct declines in performance gains relative to investment. In either case, policymakers would gradually revise downward the expected marginal value of pushing the frontier. This would not mean AI has “failed” in any sense, of course. Existing models could remain extraordinarily useful and continued investment in diffusion could remain attractive. The specific activity targeted by a pause would become less valuable relative to other uses of resources or, consequentially, the risk profile of frontier AI development at the time of assessment.
It is also possible that frontier advances could remain fruitful but prove incapable of producing durable strategic advantage. There are several possible causes of this “brittle frontier” dynamic, some of which are discussed more thoroughly in the next section. Rapid cross-border diffusion, for example, can weaken incentives to race through two related but distinct mechanisms. For the state in the leading position in a competition, rapid diffusion could reduce the expected return from investing in further AI advances. This is because the advantages of doing so may be short-lived and because those same advances could end up helping the trailing state achieve the same results more efficiently. At the same time, rapid diffusion might cause the trailing state to ask why they should incur the costs of investing in discovery when they may be able to wait for innovations to rapidly diffuse, so they can reproduce them cheaply. In other words, rapid diffusion could end up reducing both the expected benefits of extending a lead and the expected benefits of racing to close the lead. In both cases, frontier progress may remain rapid and valuable while the value of being ahead in the race falls because neither side can remain ahead for long, and in fact breaking new ground may impose costs conferring a second-mover advantage.
This could result in a costly strategic treadmill. U.S. policymakers, leading by most frontier capability metrics at the time of writing, might conclude that enormous investments can purchase only temporary leads. At the same time, Chinese policymakers could reach the same conclusion from a different starting point: even massive investments and a great tolerance for risk along the way may not produce durable superiority, or in extreme cases, result in faster progress than counting on diffusion as part of a dedicated national strategy.
Across these narratives, pause willingness occurs when policymakers see declining value in each movement of the AI frontier, either in absolute terms or relative to a second mover strategy. The specific cause of the shift is unimportant. It could result from a technological plateau, challenges with traction on real world problems, rapid diffusion of breakthroughs (discussed in detail below), or a growing belief that frontier model leadership is but one facet of broader AI leadership from a prestige or strategic assessment standpoint. One potential challenge to this path to pause willingness is that while key actors within governments may be open to a pause on this logic, private firms, with their vast resources and influence over political decisions, may not be. Governments may conclude that another generation of models offers little additional national security or even economic advantage while firms continue to anticipate commercial returns, prestige, or market share. Pause willingness, then, would require the domestic political benefits for individual politicians or coalitions of politicians to outweigh the benefits of cooperating with large AI firms.
8.4.3 The distribution of capabilities creates support for a pause
As § 8.2.2 made clear, the question of who is ahead in the AI race dominates much of the political discourse surrounding AI governance, often implying that the leading power has strong incentives to race while the trailing power has incentives to catch up, making restraint structurally unlikely. The scenario presented in this section is predicated on an assumption that the traditional model, while not wrong per se, is incomplete. Rather the relationship between the distribution of AI capabilities and pause-willingness is considerably more nuanced, and under certain conditions—specifically, when the difficulty of further advancing the frontier is high and pricey while the cost and ease of catching up is low (see § 8.4.2)—the distribution of capabilities can itself become a facilitator of a sustainable pause, wherein both leading actors prefer coordinated restraint over continued racing.
Limitations of the conventional wisdom
The conventional wisdom surrounding capability distribution and willingness to embrace restraints/stop racing holds that the dominant power will resist any agreement that would lock in its current level of superiority as a ceiling against further gains, while the trailing power will object to any agreement that would cement its current subordinate position as well. This logic is not wrong—it correctly identifies why most periods of clear capability asymmetry tend to be inhospitable to agreements like the one we are studying. At the same time, the standard model rests on assumptions that do not necessarily hold in the AI domain, and relaxing those assumptions can change the expected political dynamics considerably. Why should this be the case?
The standard model assumes, first, that first-mover advantages are durable and generate self-reinforcing benefits that compound over time, making early investments continue to pay off indefinitely. It also assumes that the subordinate state’s path to the frontier is independent of the leader's progress, which is to say that catching up necessitates generating one's own innovations rather than building on, adapting, imitating or simply stealing the dominant power’s already undertaken work. Finally, it further assumes that the leading power can prevent the diffusion of its advantages through secrecy, export controls, or classification. None of these three assumptions necessarily holds in the AI context (Horowitz 2018, 2026b), and the implications for the sustainability of AI pause willingness could be profound, as outlined in the scenario below.
When breaking new ground is hard and catching up is easy
One of the most structurally interesting conditions for pause willingness is a “brittle frontier” in which breaking new ground is costly but catching up is not. In this scenario, moving the AI frontier requires massive, sustained investment in compute or talent, while meeting this new frontier through diffusion is comparatively cheaper. Distillation—training a model on the outputs of a stronger model—is one plausible avenue towards a brittle frontier scenario. The successful and frequent distillation of frontier models and subsequent application to consumer products suggests that such a scenario may be reasonably likely.
As suggested above, for the leading power, the core advantage of being dominant is the expectation that its investment will yield a durable advantage: a period during which its superior (perhaps, even unique) capabilities translate into military, economic, and/or intelligence benefits before competitors close the gap. If they are really successful, competitors won’t even try to catch up, ceding AI supremacy to the leader. However, if the gap closes rapidly regardless of continued investment—because imitation is cheap, diffusion is rapid and relatively easy, and no single innovation can be monopolized for long—then the expected payoff from continued racing declines precipitously, particularly in the face of mutual and significant costs and risks, perhaps even existential ones.
As Swan and Hovaness argue, traditionally, “the party that initially exploits a technical breakthrough in weapons design makes it easier for a competitor to field a comparable weapon” (Swan and Hovaness 2021). In the realm of frontier AI advanced computing, “the first mover will only be providing proof of feasibility and general design information…which will allow adversaries to expeditiously develop comparable capabilities or asymmetric counters” (ibid.). If these dynamics materialize, racing could become expensive without becoming decisive, resulting in a treadmill that extracts resources and increases risks without yielding the strategic advantages that motivate racing behavior.
This “easy catch up” could lead policymakers in the trailing state to draw two different conclusions. First, it may reduce their sense of urgency to seek a pause because they may perceive that the leading state’s advantage is temporary. In other words, being in first place won’t last. Second, if policymakers in the trailing state believe that this “easy catch up” dynamic would apply if roles were reversed, they may be less motivated to invest in surpassing the leading state because it would be too costly relative to the expected payoff. If policymakers in both states have a shared understanding that there isn’t a lasting benefit to the investment that justifies the rapidly escalating and potentially catastrophic economic and security costs of continuing to race, pause-willingness becomes more sustainable (cf. Lamberth and Scharre 2023).
Trust dilemmas vs. Prisoner’s Dilemmas
In such a scenario, what emerges would be the structural condition under which competition around frontier AI development would most closely resemble what game theorists call a “stag hunt”, discussed in § 2.3, a kind of assurance/coordination game (Jervis 1976, 1978). If you trust the other party to pursue mutually beneficial gains, you’ll do the same. If you don’t trust them, you’ll continue to race. This dynamic is also sometimes referred to as a “trust dilemma” (e.g. Katzke and Futerman 2024). The interdependence of trust is the sine qua non of coordination and assurance games, so-called because each actor’s best choice depends on the other party’s decision. Trust dilemmas differ fundamentally from “prisoner's dilemmas”, wherein each side's dominant strategy is to defect regardless of what the other does (recall § 2.3). Moreover, according to Carlsmith (2026), depending on the payouts at stake if both parties race, it may even be the case that there is only one rational equilibrium—both actors embracing a pause/slowing down—which could happen if racing actually generates additional extinction risk, “without enough corresponding benefit in the non-extinction outcomes” (§ 3).
8.4.4 Domestic coalitions could make reciprocal restraint politically viable
As discussed above, various political constituencies may become dissatisfied with unrestricted AI development as AI becomes more politically and economically consequential. Several polities in the United States could combine to produce pause willingness, perhaps as a side-effect of broader resistance to AI and related infrastructure. For example, national security officials cognizant of cyber risks and CBRN uplift, workers experiencing job insecurity, and environmental activists concerned about AI infrastructure may find common cause in pause willingness despite differing logics for their opposition to AI.
AI firms, their financial backers, and their major suppliers are one of the most important political constituencies. Depending on their preferences, they may attempt to shape, forestall, or encourage a pause agreement through donations, public statements, and engagement with important policymakers. While not determinative in the presence of strong pause willingness or political resistance to a pause, the perspective of firms is worth considering as it may shape the nature of a pause and impact willingness at the margins. The firms pushing the frontier of AI capabilities may experience strong incentives to continue frontier research even when their employees and even private leadership would prefer to moderate frontier AI development. This could include incentives to race from international—as well as domestic—competitors, in which case government officials may become more interested in mutual international restraint because further development would be mutually risky and counter-productive, and they cannot solve the firm-level race within a single country.
Under certain circumstances, firms may develop a preference for predictable, universal rules, perhaps in the form of a pause, to open-ended racing. A diversity in the number and market interests of AI firms might facilitate this, as some may prefer predictability and an opportunity to exploit a market niche through integration or product development rather than pursue (less likely in a diversified landscape) leadership in frontier model benchmarks. This would reflect a history of business coalitions splitting on regulatory issues as industries and markets mature and diversify from intense competition among a small number of early-moving actors.
This pathway creates a shift in patterns by: altering government perceptions such that they perceive marginal domestic frontier gains as less valuable because they intensify an unstable international and corporate race; shifting the domestic coalition structure towards convergence around restraint; increasing government willingness to build domestic and international governance capacity and infrastructure to bind firms; shifting the support of firms to support common standards because they remove the competitive disadvantage of restraint. In this scenario, domestic coalition realignment does not directly cause a state to abandon the AI race, but it does shift the domestic political calculus enough to enable credible offers of international restraint.
In line with the trust dilemma/coordination game dynamics outlined above, if policymakers in one state become confident that their counterparts are also willing to pause and that reciprocal compliance could be verified, the political meaning of restraint could change. A global pause could be presented as a way to reduce increasingly salient domestic and security risks without conceding the frontier to a geopolitical competitor. Consistent with Putnam’s logic, growing concerns about the risks of frontier AI development could reshape domestic political coalitions in ways that make mutual international restraint more feasible. Evidence of substantial harms could favor coalitions or individuals favoring restraint, while diverging interests among AI firms could weaken firms’ ability to organize effective opposition to regulation or international agreements. These shifts could alter governments’ disposition towards restraint, particularly if they come to believe other states are similarly open to restraint.
8.4.5 Shifting power distribution and heightened risks drive multilateral restraint
Leading AI states’ attempts to preserve advantage through export controls, domestic regulation, investment restrictions, or supply-chain pressure may produce adaptation rather than control. Firms restructure supply chains, third countries resist choosing sides, new compute hubs emerge, and restrictions create incentives to circumvent and substitute with indigenous development. At the same time, both countries may become more concerned that diffusion through third countries, private firms, and new datacenters is making the frontier harder for either government to control. This combination of factors leads policymakers around the globe—in leading as well as third-party states—to become less confident that competitive containment can prevent rivals from reaching the frontier at an acceptable economic and diplomatic cost.
At the same time, actors with important roles in the global semiconductor industry or which host significant clusters of compute infrastructure may adapt rather than comply with export controls. These entities, many of them influential middle and non-aligned powers, may resist being labeled or treated simply as instruments of great power competition among the AI-leaders. As noted above, some of these actors possess noteworthy leverage because key equipment, chips, capital, markets, or infrastructure pass through their jurisdictions, while, less potently, others can confer or withhold political legitimacy from a global order designed by just two states.
8.5 Major actors’ perspectives and incentives
The preceding sections have identified general conditions under which pause willingness might emerge, but these conditions will not operate identically across actors. States occupy different positions with regard to the distribution of AI capabilities and broader national power. They also differ substantially in their domestic structures and normative landscapes, as shaped by historical experiences with technology. This section focuses on the perspectives of the leading states in frontier AI development. While the perspectives of other large states and a growing set of powerful non-state entities will be important to the feasibility of pause willingness, the United States and China, both centers of frontier AI development and poles of broader geopolitical influence, are the central actors on which international pause agreements are likely to center.
A central feature of the bilateral US-China relationship surrounding technology that conditions much of the following is the suspicion that the other side has built concealed access into its technology–a belief that is mutual, longstanding, and not specific to AI. It has shaped American policy toward Chinese telecommunications equipment, consumer applications, and connected vehicles, and it has shaped Chinese policy at considerably earlier stages, when concerns about foreign software reliance and privacy weaknesses contributed both to the embrace of open source software development—and later open source approaches to AI—and to the pursuit of a self-sufficient domestic technology stack (Blomquist et al. 2025, 17–19). Any hardwired pause verification regime proposal will likely be viewed through this lens by both parties.
One important consideration impacting major actors’ perspectives and incentives, however, is that they share concern about things like dangerous misuse (e.g., by non-state actors), critical infrastructure vulnerabilities, data poisoning, unintended escalation, and loss of control.
8.5.1 The view from the United States
Pause-willingness within the United States is influenced by a combination of technological, geopolitical, and domestic political factors, as is true elsewhere. U.S. pause-willingness may additionally be stymied or bolstered by some country- and context-specific factors.
The U.S. position as the incumbent leader in frontier AI development imposes asymmetric opportunity costs which have raised barriers to pause-willingness with some stakeholders. While the U.S. maintains an advantageous position in global AI competition amid a backdrop of broader geopolitical competition, some U.S. policymakers believe that continued development is risky, and that constraining development would be preferable to pressing its existing advantages. This is a demanding condition, particularly when others in the U.S. government believe AI will enable other strategically important capabilities. A pause may be more attractive to some U.S. policymakers who believe that the U.S. lead is difficult to extend or maintain, or that the maintenance or extension of this lead is unlikely to produce tangible strategic advantages.
U.S. politics and political structures will shape the translation of AI’s potential strategic incentives into pause-willingness. Presidents differ in their willingness to prioritize long-term risks, accept constraints on their freedom of action, defer to technical expertise, and pursue and adhere to international agreements. Further, presidential preferences are filtered through a fragmented, federal governmental structure in which Congress and courts provide checks and balances on one another, and in which state governments retain powers not granted to the federal government, all of which factors into overall U.S. pause-willingness.
These institutional considerations are not the only factors shaping national preferences. National and state polities are an important factor as well. As the country is profoundly polarized along party lines and appears likely to remain so for the foreseeable future, electoral politics may play a role in pause willingness. As the Washington Post recently pointed out, AI “is now fused with long-standing electoral issues such as the economy and the costs of living, corporate power and the future prospects of working Americans” (Ovide et al. 2026).
Nonetheless, a critical political question is whether competing groups are likely to form a coherent coalition that results in a mutually-agreeable domestic pause win-set. In addition, securing public buy-in may require linking previously separate voter concerns, framing a pause not merely as a safety measure, not merely as a labor protection, and not merely as a national security precaution (see e.g. Greenhill and Kroth n.d.), but rather as the best response to a broad set of domestic and international risks.
As is often the case in races, perceptions drive policy (see e.g. Greenhill n.d.). From Washington, perceptions of China are likely to be particularly important in determining how the political and technological context for a pause is interpreted. How U.S. political leaders perceive Chinese intentions and capabilities at any given time could produce different policy responses. For example, perceptions of the trajectory of cyber capabilities could lead policymakers to prefer restraint if they expect mutual vulnerability or to prefer racing if they believe that China is approaching a key capability threshold, especially one that confers offensive advantages. Finally, a shared focusing event could inspire massive investment in countermeasures, in addition to or instead of investment in agreeing to pause, in a scenario in which there is sufficient skepticism regarding the possibility, durability, or enforceability of a pause agreement with China. Further, the structure of the U.S. AI ecosystem creates unique dynamics which might facilitate or challenge pause willingness. Unlike the state-directed technological programs that make up much of the historical record on strategically consequential technological races, much of the relevant U.S. capability is concentrated in private hands. As discussed earlier in this section, frontier AI developers may be incentivized to resist compliance or tempted to cheat if they believe that agreement provisions impact their potential for revenue, market share and even prestige. In addition, leaders of frontier AI firms differ in their positions on pause willingness (see, e.g. Mui and Bordelon 2026), with some posing direct challenges to regulatory frameworks (Bond and McDaniel 2026). The firms have also increased and diversified their commitments to spending in political elections, pledging some $265 million to super PACs and political organizations in the 2026 cycle alone (Severns and Ramkumar 2026). While campaign contributions from Silicon Valley are not new, in the past they generally took the form of individual donations and corporate PACs (Bond and McDaniel 2026). In the U.S. political system and in today’s political climate, partially characterized by increasingly private access to federal policy, understanding the position of large private firms is key to understanding the position of the U.S. government.
As noted in § 7.1, an additional wrinkle arises from the fact that, at least as of this writing, a disproportionate share of US stock market value, capital investment and economic activity more generally is tied to AI and its further development (see e.g. Putzier 2025). Not only do the so-called Magnificent 7 technology firm stocks comprise more than a third of the S&P 500’s market value (Daly 2026), but also AI may have accounted for an outsized share of real GDP growth in recent years (see e.g. Rubinton and Patro 2026). It has been argued that such significant reliance on a single industry not only affords AI frontier firms outsized political influence, but also further deters the incumbent US leadership from embracing a pause for fear of catalyzing a market crash and triggering a recession (Sanger et al. 2026). While it is not self-evident what might be done in the near-term to guard against the realization of such concerns, to the extent that steps can be taken to allay fears of a market crash or economic slowdown, pause willingness may become more palatable to skeptical and chary US leaders. On the other hand, failing to proactively address the risks associated with a loss of control over AI could also bring on these selfsame economic calamities, which might provide an opening for enhancing US pause-willingness if economic resilience-focused efforts and activities were coupled with pause adoption.
Finally, the U.S. possesses forms of structural leverage that could either discourage or enable a pause. As noted above, U.S. firms occupy leading positions in frontier AI development, cloud computing, semiconductor design, and other parts of the AI ecosystem. Alliances connect Washington to additional critical nodes in the broader AI supply chain. So long as policymakers are confident that these advantages can be used to preserve a durable U.S. lead, they increase the opportunity costs of pausing. If this confidence erodes, however, the same advantages could make the United States singularly capable of constructing and enforcing a pause agreement. U.S. positions on pause willingness are therefore likely to come down to whether U.S. policymakers and the publics they represent believe U.S. interests are better served by using this existing leverage to extend its lead and unilaterally develop countermeasures or to institutionalize reciprocal restraint in the form of a pause.
8.5.2 The view from the People’s Republic of China
Turning to China, Chinese policymakers would be more likely to favor a pause if it is seen as more beneficial to national security, innovation, and economic futures than continued frontier AI development. Interest on the part of the Chinese government in pursuing a pause could come from a belief that continued frontier AI development posed a threat to China's political security and social stability, which official doctrine treats as the foundation of national security. Further, this would need to arrive from domestic calculations and be trusted by authorities in order to overcome suspicions that such an agreement is a mere continuation of the US semiconductor export control regime, which has been framed by Chinese officials as unilateral bullying measures by the US (Xinhua 2024). Two critical components of this include (A) a belief that continued training of new frontier models would substantially harm the PRC’s national security and (B) a belief that the incentives for the US government and industry actors are aligned towards the upholding of the agreement as well.
As in other states, domestic concerns in connection with advanced AI capabilities and their impacts may include: severe economic and labor market disruptions; misuse of (perhaps open-weight) models in cyber and other attacks by non-state actors; and broader shifts in military capabilities and balances of power that disallow the state from defending its national interests. If the frontier AI race generates substantial economic costs, such as inflation pressures and rising costs of basic utilities, Chinese leaders might view a pause as a means of mitigating the economic repercussions on citizens’ economic wellbeing. If malicious nonstate actors are repeatedly found using jailbroken versions of advanced models to stage biological or cyber attacks, this may tip the government’s political calculations about the benefits versus harms of developing increasingly advanced models and/or “open weights” releases.
Besides domestic concerns, China, like other states, faces the possibility that AI tools threaten its national security more directly. If AI systems become increasingly integrated into military activities, such as strategic planning, cyber operations, intelligence analysis, and even weapons deployment, an unconstrained frontier AI race could alter the regional balance of power in ways that threaten China’s national security interests. For example, frontier AI could enable rival states to develop military capabilities that significantly undermine China’s military capacity in the South China Sea or elsewhere in the Indo-Pacific.
There is also a documented and growing official discourse on frontier AI risk in China. In his remarks at the September summit, Xi Jinping noted that both China and the United States “have both the capability and responsibility to develop and manage AI for good, and ensure the development of AI is always under human control, and serves the well-being of the people.” The September 2025 AI Safety Governance Framework 2.0 explicitly listed loss of human control of AI as a distinct risk category, including scenarios of autonomous resource acquisition and self-replication, with the September 2026 iteration of the framework maintaining a line of analysis on loss of control. Further, China’s National Information Security Standardization Technical Committee (TC260), a major standard-setting body, has flagged the need for "circuit breakers" and "safety stop switches" (Wagner et al. 2026). Finally, state coverage of one of Xi Jinping’s speeches referenced "risks of technological loss of control" (Wagner et al. 2026); and at the September 1, 2026 Cybersecurity Awareness Week press conference, CAC's Cybersecurity Coordination Bureau named extreme loss-of-control risk as one of five headline AI security challenges (CCTV 2026). Importantly, however, although references to “loss of control” have steadily increased in Chinese AI discourse in recent years, the term is often used differently than it is by some in the West, either to describe smaller scale incidents centering around AI agents or to describe the loss of control or power over AI from a specific type of actor (Qian 2026).
One possible set of incentives in favor of China becoming a signatory on a major pause-oriented agreement is the institutional, soft power, and rule-setting benefits of doing so. Major powers tend to prefer to write international rules rather than have rules written for them. Moreover, if a broad coalition of major powers were to participate in a pause agreement that contains meaningful trade, technology, or market-access restrictions for nonparticipants, other key leaders might conclude that the economic and diplomatic costs of non-participation outweigh the potential benefits. At this point, some signs point to the Chinese side being more interested in global collaboration than the American side, although to date, these efforts have centered around capacity building and broader principles for AI governance, rather than capability-limiting considerations. China’s U.N. representative Fu Cong explicitly called for “a consensus-based global governance framework” around AI safety (Mak and Miller 2026), and China has positioned itself diplomatically as a stronger proponent of UN-centred, multilateral AI governance than the United States, though support for a consensus-based framework is distinct from acceptance of intrusive verification.
Further, Beijing has thus far worked to cultivate a reputation and institutional foundation for AI governance leadership globally (Chen et al. 2026; Blomquist 2026). China’s establishment of the World AI Cooperation Organization (WAICO) in July 2026 (PRC MOFA 2026), which builds on previous and related institutional announcements such as the Global AI Governance Action Plan (Permanent Mission of the People’s Republic of China to the United Nations 2025) and the World Data Organization (Nicole 2026), shows how the PRC is already seeking to shape the future of international AI governance. WAICO is largely an AI capacity building and benefits sharing regime, and it demonstrates that China has prioritized cultivating a reputation as a responsible leader in AI via diplomatic efforts and more formal agreements on international AI governance. From this perspective, China’s participation in a hardwired pause could be attractive not simply because it avoids international isolation, but because it could allow China to further define the rules of the game and cultivate its soft power globally, thereby shaping China’s relationships with other countries and advancing its sought after reputation as a responsible AI power.
Finally, a pause would likely be more palatable to states like the U.S. and China if it included inducements related to the global inference market. In general, a more intertwined U.S.-Chinese AI market and ecosystem could reduce the risk of unilateral defection and present mutually enforceable incentives. U.S.-China decoupling/supply chain derisking could actually exacerbate the risk of confrontation and conflict. Two separate ecosystems will make it difficult to monitor each other's progress, which is particularly disadvantageous to the West which lacks access and information about developments within China.
9 Conclusion
When this group was founded in May 2026, it was hard to imagine that October would find us in the present political conjuncture. There was little public discussion of the possibility or desirability of pausing. AI development was treated as almost an inevitable process unfolding according to natural laws, rather than something which could be made subject to collective action.
In October 2026, we are no longer in that place. Revelations such as the hacking of Hugging Face by OpenAI agents raised the profile of AI risk as an urgent issue. Recent weeks have seen politicians from multiple countries and with different ideological profiles speak in favor of a pause. A substantial portion of the American public supports intervening to arrest the current breakneck pace of advance (Doherty 2026; Salvanto et al. 2026). Notably, the heads of Anthropic, OpenAI, xAI, and Google DeepMind have all spoken in support of pacing frontier AI development in light of safety risks (Amodei 2026; Altman 2026a; Musk 2026; Hassabis 2026).
Skeptics of coordination often point to the impracticability of finding any policy that could alter a perceived foreordained trajectory of ever greater acceleration. This is not in keeping with the lessons of history. Many prior generations have confronted unprecedented, era-defining technologies that posed evident dangers. Chemical weapons, nuclear weapons, commercial aviation, radio communications, submarines, satellite reconnaissance, and strategic computing each presented governments with formidable combinations of military significance, commercial value, and scientific uncertainty. In each case, regimes were devised, sometimes despite initially long odds, to manage the risks presented.
Ultimately, historical and technical analysis cannot serve us as a playbook or a program, but as an aid to the imagination. Anyone with confident beliefs about the trajectory of political history would be well-advised to take the stunning developments of the last four months as an object lesson in humility before attempting to predict what cannot occur in the months and years to come. The analysis in this report suggests that, at least at the time of writing, leaders still have a choice about whether to pause the frontier or race ahead. Which course they choose is a question that only time will answer.
Financial Disclosures
William Fithian’s spouse holds equity in Meta, which develops AI models.
Shachar Kariv serves as a member of the Board of Directors (and committees thereof) of Exascale Labs Holdings Inc., a NASDAQ-listed AI infrastructure provider operating a software-defined GPU compute platform and related AI infrastructure solutions. The views and opinions expressed by Shachar Kariv in this paper are in his individual capacity and do not necessarily reflect the views or positions of Exascale Labs Holdings Inc., its Board of Directors, or its affiliates.
David Krueger holds individual investments in frontier AI and compute infrastructure companies.
Gretchen Krueger previously worked at OpenAI but retains no equity in the company.
Erik Leklem holds individual investments in frontier AI and compute infrastructure companies.
Adam Lesnikowski holds an individual investment in Nvidia, an AI computing company, and previously worked at the company.
No external funding was received for the research in this report. A lunch at the Working Group’s July 7 workshop was paid for from the UC Berkeley research funds of Fithian and Holliday. The Kavli Center for Science, Ethics, and the Public at UC Berkeley kindly made available meeting room space and AV equipment for the Working Group’s in-person and virtual meetings.
AI Use Disclosure
This report was written by the authors. Some of the authors used AI systems (including Opus 4.7, 4.8, 5, and 5.5, Fable 5 and 5.1, and GPT 5.5, 5.6 Sol, and 6 Astra) as research assistants and programmers during the research process.