# The frontier game

*Why two AI labs might want to slow down—and why neither might do it.*

A frontier lab faces two questions that sound similar but can have different answers: **Would we earn more if everyone paced development? Would we earn more if we paced while our rival accelerated?** The first asks about a shared outcome. The second asks about an individual incentive. The space between them is where a prisoner's dilemma can appear.

This simulator tests that possibility. OpenAI and Anthropic are the two named players. Each can **cooperate by pacing** within an imagined capability constraint, or **defect by accelerating** beyond it. Those are game labels, not claims about an actual agreement or a company's conduct. Pacing need not mean stopping product releases, efficiency work, or safety research. A real proposal would have to specify what activity is constrained and how compliance is measured.

The simulator does not assume that pacing is profitable, stable, or socially desirable. It calculates the incentives under six sets of assumptions. Sometimes the familiar dilemma emerges. Sometimes both labs have a commercial reason to race. Sometimes pacing needs no bargain at all.

## What the machine is doing

One turn represents a twelve-month cycle. Acceleration makes a lab's offering more attractive but requires extra investment. Customers divide their usage among the two labs and an outside option. That outside option represents all other suppliers together, including other closed providers and open-weight models. It is not a third strategic player making its own decisions.

Outside models improve through two separate channels. **Independent progress** is already present throughout the cycle, whether or not either lab accelerates. **Release-induced improvement** arrives after a chosen delay. This second channel represents a possible spillover from a frontier release, including imitation or diffusion. With the default three-month delay, customers spend one quarter of the year facing the early competitive landscape and three quarters facing improved outside models. The model uses the strongest accelerated release; it does not double the spillover when both labs accelerate equally.

This separation matters. An observed open/closed capability gap cannot tell us how much progress came from copying. [Epoch's gap measurement](https://epoch.ai/data-insights/open-closed-eci-gap) is useful context, not a stopwatch for distillation. The simulator's imitation settings are assumptions, not estimates extracted from that research.

The machine then calculates each lab's revenue, subtracts serving costs, extra investment, fixed commitments, and any hypothetical acceleration penalty, and displays the remaining **payoff**. All financial values are normalized units. A payoff of 15.55 is not $15.55 billion, a margin, or an estimate of company profit. At the default prices, 100 units of total token demand produce 100 revenue units across the three suppliers.

Token share and spending share are shown separately. A low-price supplier can win many tokens while collecting a smaller fraction of spending; [Vercel's gateway sample](https://vercel.com/blog/ai-gateway-production-index-september-2026) illustrates why that distinction matters. Neither the default share nor that sample is a global market estimate.

## How to read the four outcomes

The matrix asks what happens if both pace, either lab accelerates alone, or both accelerate. Every cell lists OpenAI's payoff first and Anthropic's second. Compare a lab's payoff while holding its rival's choice fixed: can changing only its own action improve the result?

A **Nash equilibrium** is a cell where neither lab benefits from switching alone. It is a statement about incentives, not fairness, total value, or what will inevitably happen. A better shared outcome can exist outside equilibrium. The simulator reports equilibria in definite actions; it does not solve strategies that randomly mix pacing and acceleration.

The presets are symmetric so the logic is easy to see. OpenAI and Anthropic can exchange places without changing the result. That is a simplifying assumption, not a claim that their businesses are identical. Separate investment and commitment controls let you explore asymmetry.

## 1. An illustrative dilemma

**The trap: each wants the advantage of accelerating, but both prefer the payoff from mutual pacing.**

With the default assumptions, joint pacing pays **15.55 each**. If one lab accelerates alone, its payoff rises to **17.29** while its rival falls to **7.71**. If both accelerate, each receives **12.63**.

Imagine OpenAI believes Anthropic will pace. OpenAI can gain 1.74 by accelerating. Now imagine Anthropic will accelerate. OpenAI can gain 4.91 by accelerating too, avoiding the worst outcome of being left behind. Anthropic faces the same comparisons. Acceleration is therefore the best individual choice against either rival action.

Yet both would earn 2.93 more from pacing together than from racing together. That makes this a strict prisoner's dilemma: the only Nash equilibrium is mutual acceleration, while mutual pacing is better for both. Release advantages, investment costs, and outside spillovers combine to create the trap; the label was not built into the payoffs.

**Try this:** reduce the imitation lag from three months to zero. Immediate outside improvement reduces the temptation to accelerate first, and the game becomes a coordination problem. This is an experiment about an assumed mechanism, not evidence that instantaneous imitation is possible.

## 2. Outside pressure, weak imitation

**The race can be a rational defense of a shrinking franchise.**

Here outside suppliers are much stronger independently, release-induced imitation is weaker, and acceleration costs fall from 6 to 5. Joint pacing pays only **5.03 each**. A lone accelerator earns **8.61**, leaving the pacing rival with **3.36**. Mutual acceleration pays **6.49 each**.

Both labs still prefer acceleration against either rival choice, so mutual acceleration remains the only Nash equilibrium. But it now pays more than mutual pacing. The defining conflict of the prisoner's dilemma has disappeared. Pacing together sacrifices demand to outsiders that are improving anyway; acceleration protects enough demand to cover its extra cost.

This scenario changes three assumptions together. It does not establish that stronger outsiders always make labs race. If the remaining opportunity becomes too small, acceleration may stop paying for itself.

**Try this:** raise independent outside progress from 1.5 to 2.5 while leaving the other settings alone. Pacing becomes individually optimal, even though all four outcomes now give the labs negative payoffs. The model can identify the less costly choice without claiming that choice restores profitability.

## 3. Expensive acceleration

**Sometimes the commercial answer is simply to stop spending so much.**

This preset changes only the extra investment required to accelerate, raising it from 6 to **20 per lab**. Customers still respond to releases exactly as they did in the default case. The added demand is now too expensive to win.

Joint pacing still pays **15.55 each**. Accelerating alone pays **3.29**, while the pacing rival receives **7.71**. Mutual acceleration produces **−1.37 each**. Against a pacing rival, a lab is much better off pacing. Against an accelerating rival, pacing also pays more than joining the race.

Mutual pacing is therefore the only Nash equilibrium. No promise of future retaliation is needed to support it in this one-cycle game. This is different from a dilemma in which restraint is desirable but individually unstable. A negative racing payoff is also not a prediction of bankruptcy: the simulator has no balance sheet, financing process, or default mechanism.

**Try this:** return both extra-investment controls to 6. The original dilemma returns, showing how avoidable acceleration cost can change the strategic category while the demand assumptions stay fixed.

## 4. A coordination problem

**Both prefer restraint, but nobody wants to be the only lab exercising it.**

Set extra investment to **9 each**, between the default and expensive-acceleration cases. Joint pacing pays **15.55 each**; joint acceleration pays **9.63 each**. A lone accelerator earns **14.29**, while its pacing rival receives **7.71**.

Against a pacing rival, acceleration would reduce a lab's payoff from 15.55 to 14.29. Against an accelerating rival, however, acceleration improves its payoff from 7.71 to 9.63. Each lab's best response is to match the other's action.

There are consequently **two Nash equilibria**: both pace or both accelerate. Pacing is better for both, but that fact does not automatically select it. If each expects the other to race, neither wants to pace alone. Expectations, observability, and a credible description of the restraint could matter in a richer model. This simulator identifies the two stable outcomes; it does not model a negotiation that chooses between them.

**Try this:** reduce both investment controls from 9 to 6. Acceleration becomes tempting even against a pacing rival, converting the coordination problem into the default prisoner's dilemma.

## 5. Releases expand demand

**Better models can create a larger market instead of merely reallocating an existing one.**

This preset adds one assumption: if both labs accelerate, total token demand rises **80%**, from 100 to 180 units. One accelerator creates half that expansion, taking demand to 140. All other default settings remain unchanged.

Joint pacing pays **15.55 each**. A lone accelerator earns **28.61** while its rival earns **12.80**. Mutual acceleration pays **31.53 each**, the best payoff for each lab among the four outcomes. Acceleration is individually attractive and mutually preferred to pacing, so mutual acceleration is the sole Nash equilibrium without a strict dilemma.

The mechanism is commercially plausible: new capabilities can make previously uneconomic applications useful. [Research on model releases](https://arxiv.org/abs/2504.15440) describes both substitution and market expansion. It does not supply this simulator's 80% assumption. Extra token volume also need not represent equally valuable work, and the model holds token prices and serving costs fixed while volume expands.

**Try this:** lower demand growth from 80% to zero. The default dilemma returns. The result turns on how much additional business acceleration creates, not just whose existing customers it attracts.

## 6. One accelerator is enough

**Each lab wants to lead, but neither wants an expensive head-to-head race.**

Here release-induced imitation is zero and extra investment is **13 each**. Independent outside progress remains. Joint pacing pays **15.55 each**. Accelerating alone pays **16.79**, leaving the pacing rival with **10.63**. Mutual acceleration pays **10.07 each**.

Against a pacing rival, acceleration wins an advantage worth its cost. Against an accelerating rival, joining the race costs more than the additional demand is worth: pacing earns 10.63 instead of 10.07. Each wants to choose the opposite action from its rival.

This is chicken, or anti-coordination. It has two Nash equilibria: OpenAI accelerates alone, or Anthropic accelerates alone. Neither mutual pacing nor mutual acceleration is stable against a unilateral switch. Both labs prefer the equilibrium in which they are the accelerator, so identifying an equilibrium does not settle who gets the favorable role. “One is enough” describes the modeled incentives, not a social recommendation to appoint a permanent leader.

**Try this:** restore imitation strength to 0.8. At these investment costs, pacing becomes the best response to either rival action.

## What repetition and financial pressure change

In the default dilemma, today's temptation can be outweighed by tomorrow's lost cooperation. The repeated-game panel assumes both labs pace until either accelerates, then both accelerate forever. Under that particular strategy, cooperation can be sustained when the weight placed on the next annual cycle is at least **0.373**, or about **37.3%**.

That number is a discount factor, not a probability that a company survives or cooperates. The calculation assumes indefinitely repeated, identical payoffs, perfectly observed actions, and no commonly known final cycle. Passing the threshold makes cooperation supportable under those assumptions; it does not predict actual coordination. Rapidly changing capabilities and markets are a major limitation of this stationary exercise.

Fixed commitments have a different role. Increasing a lab's commitments lowers every one of its payoffs by the same amount. It therefore changes neither its best response nor this repeated-game threshold. Real obligations can affect financing and runway, but those mechanisms would need to be added explicitly. A large infrastructure headline alone cannot establish an incentive to defect.

## What this says about “pace the frontier”

The simulator makes a commercial motive for pacing intelligible without treating it as proven. It also shows when that motive fails: outside competition or new demand can make racing attractive to both labs. Public proposals need to be assessed against their actual rules, participation, verification, and enforcement—not a generic promise to slow down.

The trading app and squirrel tracker expose another boundary. Two customers can buy comparable model usage and obtain very different value from it. This model charges for tokens; it does not measure the customer’s downstream earnings or let a lab charge a share of them. Testing whether applications, specialized products, or licensing capture more of that value would require another layer of the model. The present game cannot establish motives for watermarking or particular licensing policies.

Private incentives and public welfare remain separate questions. The payoffs omit consumer benefits, scientific gains, employment effects, and potential AI harms. A policy that increases both labs' payoffs might harm users; a policy that reduces profits might improve safety. Commercial self-interest and sincere safety concerns can coexist.

Use the game to make the disagreement precise: which assumption changes the outcome, what evidence could measure it, and whose benefits or costs are missing? The result is a map of conditional incentives, not a verdict on either company's motives.
