# How the frontier game works

**This simulator is a structural thought experiment. It is not calibrated to OpenAI or Anthropic, does not estimate their actual finances, and does not establish that either company is cooperating, defecting, or coordinating with the other.** The names make the proposed strategic story concrete; the model tests whether particular assumptions produce a prisoner's dilemma.

One cycle lasts 12 months. Each lab chooses **C: pace** or **D: accelerate**. Acceleration adds a demand advantage and an extra investment cost. Outside suppliers can improve independently and can receive an additional improvement after a release, representing an assumed imitation or diffusion channel. Six presets show how these mechanisms can produce different games.

## From assumptions to payoffs

Write `aᵢ = 0` for pacing and `aᵢ = 1` for acceleration. The two labs begin with equal appeal. If their combined starting token share is `F`, the outside option's baseline appeal is:

`z = ln[2(1 − F)/F]`

This construction reproduces the chosen starting share when prices are equal and no outside progress or acceleration has occurred. Appeal is a dimensionless index, not a benchmark score. Adding one point multiplies a supplier's demand weight by approximately 2.72 before shares are normalized.

In the early phase, the lab and outside appeal indices are:

`vᵢ = g aᵢ − β ln(p)`

`vₒ = z + o − β ln(pₒ)`

Here `g` is the release advantage, `o` independent outside progress, `β` price sensitivity, `p` the labs' common token price, and `pₒ` outside price. Independent outside progress is a **level shift present throughout this cycle**, even when both labs pace.

Demand uses a softmax, or logit, rule:

`sᵢ = exp(vᵢ) / [exp(v_A) + exp(v_B) + exp(vₒ)]`

After imitation arrives, outside appeal increases by `m g max(a_A, a_B)`, where `m` is imitation strength. The labs retain their own release advantages. The `max` assumption lets outsiders learn from the strongest release without treating two equally strong releases as twice the capability improvement. It is a modeling choice, not a measured fact about distillation.

For imitation lag `L` months, the cycle's token share is:

`s̄ᵢ = (L/12) sᵢ,early + (1 − L/12) sᵢ,late`

This is a continuous two-phase weighting, with uniform token volume over time. A three-month lag assigns 25% of the cycle to early shares and 75% to post-imitation shares. Fractional-month values work the same way. A zero-month lag uses only post-imitation shares; a 12-month lag places imitation beyond the modeled cycle. With imitation strength zero, changing the lag has no effect.

Acceleration can expand the market as well as redistribute it. If `e` is demand expansion:

`Q = Q₀ [1 + e(a_A + a_B)/2]`

One accelerator creates half the selected expansion; two create all of it. This response is assumed, not estimated. The model's payoff for lab `i` is:

`πᵢ = Q s̄ᵢ (p − c) − aᵢ kᵢ − fᵢ − aᵢ h`

`c` is serving cost per token, `kᵢ` extra acceleration investment, `fᵢ` fixed commitments, and `h` the expected acceleration penalty. The penalty is a hypothetical, probability-weighted policy cost; zero assumes no such penalty. It is not an allegation about any current legal obligation.

Token shares are converted separately into **spend shares** using each supplier's price. Cheap outside supply can win more tokens than spending. The simulator does not calculate outside profits because their costs are unspecified. More tokens need not mean more revenue when prices fall.

## Inputs and defaults

Financial quantities are normalized, not dollars. At the default price of 1, 100 token units generate 100 revenue units across suppliers before any market expansion. The table lists every engine parameter; the interface groups common assumptions and places others in advanced controls.

| Input | Default | Meaning and unit |
|---|---:|---|
| Starting frontier share `F` | 60% | Combined token share at equal reference prices, before outside progress |
| Release advantage `g` | 0.80 | Appeal-index increase for an accelerator |
| Extra investment `k_A`, `k_B` | 6 each | Revenue units per accelerated cycle; can differ by lab |
| Independent outside progress `o` | 0.35 | Outside appeal-index increase throughout the cycle |
| Imitation strength `m` | 0.80 | Multiplier on the strongest release advantage |
| Imitation lag `L` | 3 months | Delay inside the 12-month cycle |
| Demand expansion `e` | 0% | Extra token volume when both accelerate |
| Expected penalty `h` | 0 | Revenue units per accelerator per cycle |
| Baseline token volume `Q₀` | 100 | Normalized token units per cycle |
| Lab token price `p` | 1 | Revenue units per token unit |
| Lab serving cost `c` | 0.20 | Cost units per token unit |
| Outside token price `pₒ` | 1 | Revenue units per token unit |
| Price sensitivity `β` | 1 | Coefficient on minus log price; not an estimated elasticity |
| Fixed commitments `f_A`, `f_B` | 5 each | Choice-independent cost units per cycle |
| Future-cycle weight `δ` | 0.80 | Discount factor used only for the repeated-game illustration |

## What counts as a prisoner's dilemma

The simulator derives all four payoff cells from those inputs. Values are in **OpenAI, Anthropic** order. The defaults produce:

| OpenAI / Anthropic | Anthropic paces | Anthropic accelerates |
|---|---:|---:|
| **OpenAI paces** | 15.55, 15.55 | 7.71, 17.29 |
| **OpenAI accelerates** | 17.29, 7.71 | 12.63, 12.63 |

Both firms would earn more from joint pacing than joint acceleration, yet acceleration is each firm's best response to either rival choice. For each lab, let `T` be unilateral acceleration, `R` joint pacing, `P` joint acceleration, and `S` unilateral pacing. A **strict prisoner's dilemma** requires:

`T > R > P > S`

These inequalities must hold for **both** labs, including when costs differ. The code tests preferences directly; it does not assign the familiar ranks 4, 3, 2, 1 to predetermine the result.

A cell is a **pure Nash equilibrium** if neither lab can improve by changing its action alone. Payoff differences within `10⁻⁸` are treated as ties. Ties retain both best responses and are not labeled a strict dilemma. The other categories are pacing-dominant, coordination, chicken/anti-coordination, acceleration-dominant without a strict dilemma, and asymmetric or boundary cases. Only pure equilibria are reported; mixed equilibria are not solved.

Fixed commitments cancel out of every unilateral comparison: subtracting the same `fᵢ` from two alternatives cannot change which is larger. They also cancel from the repeated-game threshold below. They can make the displayed cash payoffs negative, but the model has no funding constraint or bankruptcy mechanism. Consequently, it cannot turn obligations alone into an incentive to accelerate.

## Six verified scenarios

| Preset | Verified result | Pure Nash choices |
|---|---|---|
| **An illustrative dilemma** | Joint pacing pays 15.55 each; joint acceleration 12.63; unilateral acceleration 17.29 | Both accelerate |
| **Outside pressure, weak imitation** | Joint pacing pays 5.03 each; joint acceleration 6.49. Acceleration dominates without a strict dilemma | Both accelerate |
| **Expensive acceleration** | Extra investment rises to 20 each; pacing is best against either rival action | Both pace |
| **A coordination problem** | Extra investment is 9 each; pacing is best against pacing, acceleration against acceleration | Both pace **or** both accelerate |
| **Releases expand demand** | Full acceleration expands volume by 80%; joint acceleration pays 31.53 each versus 15.55 from pacing | Both accelerate |
| **One accelerator is enough** | No imitation spillover, extra investment 13; each prefers to accelerate only when the other paces | Either lab accelerates alone |

**Stronger outsiders do not automatically force a race.** Raising only independent outside progress from the default 0.35 to 1 to 2 changes this model from a prisoner's dilemma to coordination to pacing-dominant behavior. The remaining franchise eventually becomes too small to justify the extra investment.

The outside-pressure preset changes three assumptions together: outside progress rises to 1.5, imitation strength falls to 0.1, and extra investment falls to 5. Its result should not be read as the isolated effect of stronger outsiders. By contrast, the expanding-demand preset changes one assumption and shows why releasing better models need not be a zero-sum contest for an unchanged pool of spending.

## Repeated interaction

Only for a strict dilemma, the simulator compares perpetual pacing with a one-cycle deviation followed by perpetual acceleration. Under **grim trigger**, both pace until either accelerates; after that, both accelerate forever.

`Rᵢ/(1 − δ) ≥ Tᵢ + δ Pᵢ/(1 − δ)`

Rearranging gives the minimum future-cycle weight for each lab:

`δᵢ* = (Tᵢ − Rᵢ)/(Tᵢ − Pᵢ)`

Both firms' conditions hold when `δ ≥ max(δ_A*, δ_B*)`. The default threshold is approximately **0.373**. This is the standard grim-trigger argument, generalized from the numerical example in [MIT's lecture on cooperation in repeated games](https://live.ocw.mit.edu/courses/14-15-networks-spring-2022/mit14_15s22_lec19.pdf).

The calculation requires stationary payoffs, indefinite repetition with no commonly known final cycle, perfectly observed actions, and the specified punishment strategy. Here the cycles are annual. The discount factor is a weight on future payoffs, **not a firm-survival probability**. Passing the threshold establishes that this strategy can sustain cooperation under those assumptions; it does not predict cooperation or establish actual coordination. Persistent changes in markets or technology invalidate the stationary calculation.

## Boundaries and checks

The model does not include financing, default, survival, fundraising, consumer welfare, or the benefits and harms of AI safety. It does not model learning across cycles, accumulated capability, persistent outside-technology stocks, research-driven serving efficiency, endogenous prices, compute limits, enterprise contracts, heterogeneous users, uncertain imitation, or strategic outside suppliers. The demand rule and imitation channel are assumptions, not empirically fitted laws. A higher modeled corporate payoff is not necessarily a better social outcome.

The implementation passes **20 checks**, including a deterministic sweep of 1,000 parameter combinations. These verify the six preset classifications, payoff accounting, share totals, asymmetric costs, tie handling, lag endpoints, the difference between token and spend shares, fixed-commitment invariance, and repeated-game thresholds against discounted present values. Every input is bounded and validated; invalid or nonfinite restored state is rejected.
