Scope & evidence charter
The Global Simulator — scope of application
A compact operating charter for interpretation, evidence and responsible use.
The Global Simulator is a counterfactual exploration tool — structured "what-if" investigation, policy debugging, scenario stress-testing — and not a predictor of real-world events. Every run requires a human in the loop.
What the system does NOT claim
- "We predict the dates or outcomes of real conflicts and elections" — forecast-style claims are permitted only once Track C validation has passed, with an open Brier score.
- "The agents think like real leaders" — they are LLM personas assembled from public sources; plausible, but not identical to the people they're modeled on.
- "Fit for autonomous decision-making" — the system's output always requires human interpretation and accountability.
- "Our arbiter is objective and unbiased" — the underlying models carry bias; validation measures consistency, not "objectivity".
Ethics — prohibited uses
- Optimizing protest-suppression tactics against real groups of people.
- Targeting disinformation campaigns at real ethnic, religious, or other minorities.
- Processing real citizens' personal data to simulate their individual behavior.
Reliability horizon
A meaningful session runs 20–30 ticks: beyond that, documented cognitive degradation sets in for the LLM agents, and trajectories become illustrative rather than analytical.
Transparency and reproducibility
- Every actor action carries a mandatory rationale; every arbiter decision keeps a full trace (arguments, the 2d6 roll, effects applied).
- Numbers are computed by fixed rule and effect tables, not by the LLM.
- Every run gets a replication manifest (models, temperatures, seed), so its configuration can be inspected and seeded numeric rolls can be recomputed. A newly created run has a new identity, and changing model output means end-to-end replay is best-effort rather than a byte-for-byte guarantee.
- Scenarios touching acute real-world conflicts can be marked for a
closed research mode through
sim_runs.acute_conflict. The run pages and run-bound data then require operator+; adjacent institution/chat channels are being brought under the same access contract. These scenarios are not published as forecasts.
Evidence status taxonomy
Every run and every eval track carries one of five statuses (MEMO-2026-07.md §5.2; full list — docs/EVAL.md):
- Synthetic — Tests the mechanism on a made-up trajectory. No claim about the real world.
- Calibrated — Parameters tuned on a historical window (in-sample). Direction/sign/lag, not point prediction.
- Retrodictive — Reproduced outside the training window; the outcome was not used to fit parameters.
- Forecast — Probability was registered before the outcome was known.
- Scenario — Conditional counterfactual: 'if these assumptions hold, the model gives...'. Not a prediction.
All analytical runs default to Scenario status. Most mechanism eval tracks are Synthetic, connectivity/population checks are Calibrated, and registered forecasting is Forecast. The deferred backtest track is Retrodictive; after a resolved registered check, a specific run can move from Scenario to Retrodictive or Forecast under the backtest gates. Unchecked runs remain Scenario.
Bias declaration
PILOT (n=5, not representative) — dev smoke test from 2026-07-07 (Track A — Arbiter Reliability, Track B — Committee Baselines; methodology — docs/EVAL.md). Published as-is, including the arbiter's own target it failed to meet:
- Inter-Model Consistency (IMC): 0.2 (target ≥ 0.85 — NOT met)
- Prompt Sensitivity (PSI mean): 0.0 (target ≤ 0.15)
- Committee dissent entropy (staged): 0.8 (target ≥ 0.5)