# Snapshot v1.0 — how the numbers were derived **Computed 2026-09-16. Snapshot v1.0. Methodology v0.1 / rubric R0.1.** Static founding snapshot. Next protocol refresh window: 2026-12-16. BLOB-R0.1 `0fae6bbcc552eec8ce70b36b0ecbab2c5a1c85dd`. COMMIT-R0.1 `051565a5751dfe3e87cb60298da91c6d3d26b9ea`. Machine-readable output: [`snapshots/v1.json`](../snapshots/v1.json). Recompute: `python3 compute/snapshot.py`. Check: `python3 compute/snapshot.py --check`. Tests: `python3 compute/test_snapshot.py` (or `cd compute && python3 -m unittest test_snapshot.py`). This page is the method printout for the founding number. It is not a live index. --- ## Plain language This site does not measure the chance that AI causes a catastrophe. Nobody can observe that. What it does is collect **named, published forecasts**, keep only the ones that are actually a number said by the named person or survey, line them up on the **same kind of event** (catastrophe vs extinction vs something else) and the **same kind of time horizon** (this century vs the next few decades), and then take a **weighted middle**. That middle, for “AI-caused catastrophe by 2100,” is **10.0%**. The five forecasts that are allowed to count disagree. The lowest is 2.13%; the highest is 12%. The middle of that disagreement, after the published weights, sits on 10% because two of the five rows *are* 10% and together they carry half the weight. News is **not** currently pushing the number up or down. There is no frozen list of scored news events yet, so the methodology’s default applies: multiply by **1.00** (neutral). The 10.0% you see is the same as the expert baseline. That is a gap, not a hidden judgment that “nothing is happening.” A second gauge, **literal extinction by 2100**, has only two published numbers (0.38% and 3%). That is too few for a full average, so the page shows **both points** and their midpoint (1.69%), labeled a thin panel. A third gauge, **deployment cascade** (not-ready AI pushed into the grid, military command, hospitals, finance, logistics until failures compound), has **zero** published forecasts. The page does **not** invent a percentage for it. It is labeled a constructed estimate. The old mockup’s 4.7% / 1.2% / 3.9% / ×1.08 figures were pictures. They are not these numbers. --- ## Headline (G1) | Item | Value | |---|---| | Statistic | Untrimmed weighted median [SJ-STAT] | | Unconditioned aggregate | **10.0%** | | Conditions multiplier M | **1.00** (default; empty event list) | | Headline (unconditioned × M) | **10.0%** | | Weighted mean (shown, not the headline) | 8.12% | | Source IQR (unweighted Tukey hinges, then × M) | **3.57–11.0%** | | Conditioning envelope (point × [0.65, 1.54]) | **6.5–15.4%** | | n | 5 (full aggregate eligible; §6.3) | | Predecessor | none (GENESIS-001) | ### Who is in the mixture Only ledger rows that are a **primary numeric utterance**, mapped to G1 (catastrophe / existential catastrophe / collapse-or-worse), and in horizon class **H-century**. Weights are the published v0.1 weights; they are not tuned to hit a preferred headline. | ID | Value used | Weight | Why it counts | |---|---|---|---| | A-XPT-SF | 2.13% | 1.5 | XPT superforecasters, FRI catastrophe (≥10% dead), 2100 | | A-CARL-2022 | 5% | 2.0 | Carlsmith main-text product. Companion later bound >10% is shown beside, not substituted [SJ-CARL-POINT] | | A-ORD-2020 | 10% | 2.0 | Ord, *The Precipice*, ~10% existential risk from unaligned AI, ~100 years | | A-AII-2024 | 10.0% | 3.0 | AI Impacts 2024 pooled median, extinction-or-severe-disempowerment | | A-XPT-DE | 12% | 1.5 | XPT domain experts, same FRI catastrophe question | Total weight 10.0. Ordered: 2.13 (cum 1.5) → 5 (cum 3.5) → 10 (Ord, cum 5.5). Half of 10 is 5. The first value whose cumulative weight is at least 5 is **10%**. A-AII-2024 is also 10%, so the median is not a knife-edge between two different numbers; it sits on the 10% plateau. Convention (also printed in the JSON): if cumulative weight lands *exactly* on half after a row, take the midpoint of that row and the next. That even-split rule is what produces 7.5% in one leave-one-out case below. ### Who is not in the mixture | ID | Why not | |---|---| | A-AII-2023 (5%) | Prior wave of the same survey family. Weight 0 in the mixture. Used only as substitute-wave sensitivity. | | A-HINT-2024 (10–20%) | Literal extinction, ~30 years (H-near). Near-term panel, as an interval, not a midpoint. | | A-HUB-2026 (>10%) | Literal extinction, 10 years (H-near). Near-term panel, as a bound, not rewritten as 10%. | | A-BENG-2023 (~20%) | No calendar horizon (H-unspecified). | | A-LECUN-2023 | Asteroid comparison; the <0.01% figure is a third-party conversion. Ledger only. | | A-YUD | UNVERIFIED as a primary numeric quote. Shown as a high outlier; does not score. | | Metaculus / markets | Tier B unverified. Out of the v1 mixture [SJ-TIERB-V1]. | LOCK include-outliers means LeCun and Yudkowsky **stay on the ledger**. It does not mean a paraphrase becomes a data point. Carlsmith and both AI Impacts waves are **disempowerment-shaped**, not the FRI “≥10% dead” operationalization. They enter G1 with an ambiguity flag. That is a site mapping judgment [SJ-MAP], not a source claim. --- ## G2 — extinction, thin panel Two scoring rows, both XPT, horizon 2100: - A-XPT-SF **0.38%** - A-XPT-DE **3%** n = 2 → **thin panel, not a full aggregate**. Both numbers are shown. Unweighted midpoint **1.69%**. No IQR costume. Same M = 1.00. Hinton and Hubinger stay on the near-term panel; they are not mixed into 2100. --- ## G3 — constructed estimate {#g3} n = 0 published forecasts of “premature deployment into critical systems producing compounding failures.” **SITE-CONSTRUCTED, NOT AN AGGREGATE.** No percentage is published as if it were a conditioned aggregate. The 3.9% on the old mockup is withdrawn. Allowed on the page: a short scenario narrative; a qualitative factor list (deployment into named system classes, verification lag, liability lag, compounding, terminality), each an *indicator*, not P(G3); pointers to C2 and C7 event rows — of which there are none at genesis. The operator thesis (SITE-PRIOR-001) motivated G3’s existence. It does not set the G1 headline. --- ## News multiplier (M1) There is **no frozen event ledger** for v1. The script will refuse to apply invented events. Methodology §11.5: > v1 snapshot default is **M = 1.00 at genesis** with an empty qualifying > event list. All seven categories are 1.00 (the only default on the quantized set {0.94, 0.97, 1.00, 1.03, 1.06}). Combined M = 1.00, inside [0.65, 1.54]. C2+C7 share cap does not bind. Decay is a no-op (already neutral). C6 cannot score upward until an S2 codebook annex exists. Permanent comparison line (§10.3): > Conditions currently multiply the aggregate by **×1.00** (span allowed: > ×0.65–×1.54, ratio 2.37). Reference disagreement (XPT catastrophe, > superforecasters vs domain experts) is **5.63×**. Conditions currently move > this **less** than the forecaster disagreement (neutral multiplier; > disagreement remains 5.63×). --- ## Down-move eligibility INC-CLOCK starts at the as-of date (2026-09-16) because no S2+ incident is scored. Clock value: **0 days**. | Trigger | Status | |---|---| | D-GOV-TREATY | UNMET | | D-INC-CLEAR | UNMET | | D-INT-REPL | UNMET | | D-DEP-PAUSE | UNMET | Unmet is listed, not implied. Genesis does not backdate a clear incident window through the 2020–2026 reference period: there is no C6 codebook, and the written rule starts the clock at as-of. --- ## Sensitivities {#sensitivities} Required by §9.2. Small n makes the weighted median jumpy; that is shown rather than smoothed away. ### 1. Leave-one-out (G1) | Row removed | Headline | |---|---| | A-XPT-SF (2.13%, w=1.5) | 10.0% | | A-CARL-2022 (5%, w=2.0) | 10.0% | | A-ORD-2020 (10%, w=2.0) | 10.0% | | A-XPT-DE (12%, w=1.5) | 10.0% | | **A-AII-2024 (10.0%, w=3.0)** | **7.5%** (−2.5 pt) | Removing the survey row lands exactly on half the remaining weight at Carlsmith’s 5%, so the even-split rule takes the midpoint of 5% and Ord’s 10% = 7.5%. The headline is therefore **sensitive to the heaviest row**. ### 2. Substitute-wave Replace A-AII-2024 (10.0%) with A-AII-2023 (5%) at the same weight 3.0. Headline becomes **5.0%** (−5.0 pt). The 2023→2024 survey move is the largest single lever in the current table. ### 3. Weighting-scheme range | Scheme | n | W(·) | Note | |---|---|---|---| | v0.1 weights | 5 | **10.0%** | Headline | | Equal weights | 5 | 10.0% | | | Survey-only | 1 | 10.0% | Insufficient for a headline; sole remaining row | | XPT-only | 2 | 7.065% | Even split of 2.13% and 12% | | Structured-only | 2 | 7.5% | Even split of Carlsmith 5% and Ord 10% | Schemes with n ≥ 2 span about **7.07–10.0%**. ### 4. Mapping lines Actual headline = native-G1 rows only: **10.0%**. Counterfactual: add the H-century native-G2 values (XPT extinction 0.38% and 3%) as extra rows. This **double-counts FAM-XPT** and is not a headline. | Line | Result | |---|---| | Native-G1 only (headline) | 10.0% | | Native-G2 at face value added | 7.5% | | Native-G2 up-mapped by m=4 added | 10.0% | Hinton and Hubinger are not added here: that would also break the horizon rule. ### 5. Named-statement on/off No named-statement rows in the G1-2100 mixture (Hinton, Hubinger, and Bengio fail horizon). Difference **0.0 pt**. ### 6. Unconditioned vs conditioned 10.0% vs 10.0%. Difference **0.0 pt**. M = 1.00. --- ## Band composition **In the source IQR:** disagreement among the five G1-2100 mixture rows, unweighted so A-AII-2024 cannot hide the 2.13–12 spread. Tukey hinges: for odd n the overall median is excluded from both halves; Q1 = median(2.13, 5) = 3.565% → displayed 3.57%; Q3 = median(10, 12) = 11.0%. **Not in the source band:** mapping uncertainty, weighting-scheme uncertainty, horizon-standardization uncertainty, conditioning. **Envelope (separate):** 10.0% × 0.65 and 10.0% × 1.54 = 6.5–15.4%. Drawn from the **bounds**, not from current M. After conditioning, the IQR is multiplied by current M. At genesis that changes nothing. --- ## Genesis row **GENESIS-001** (SJ-GENESIS, R-NULL): founding static snapshot under methodology v0.1. Unconditioned G1 10.0%, M = 1.00, rationale as above, day-one down-move eligibility all UNMET. Subsequent movements must cite events. This document is the compute printout that the methodology deferred. --- ## Site prior and incentive | Row | Content | Control | |---|---|---| | SITE-PRIOR-001 | Mundane deployment of not-ready AI into sensitive systems is the founding thesis. It motivated G3 and categories C2 and C7. | Does not set G1. C2+C7 log-share cap. G3 is not an aggregate. | | SITE-INCENTIVE-001 | The product is more noticeable if the number moves. | Static v1; default R-NULL; quantized steps; resolution floor 0.5 pt; public ledger of non-moves. | --- ## Gaps (honest G1, incomplete elsewhere) G1 with n=5 **is** a full aggregate under §6.3. It is attributable. It is also jumpy, mapping-ambiguous, and unconditioned by news. Those are named, not papered over. 1. **No news events scored.** M = 1.00 by default, not because a review found nothing. A later frozen event ledger can move M inside [0.65, 1.54]. 2. **Tier B out.** Metaculus / markets remain unverified. 3. **G3 n=0.** Constructed label; no fake percent. 4. **G2 n=2.** Thin panel only. 5. **C6 S2 codebook missing.** Incidents cannot score up. INC-CLOCK is 0 days at genesis. 6. **Ord not re-OCR’d** from the book in the research pass; ~10% is inherited with a tilde. 7. **Horizon rule** keeps several famous named statements off the 2100 headline. That is deliberate. They are on the near-term panel as-stated. 8. **No 0.5 pt predecessor delta.** Founding snapshot. Non-comparability [SJ-NOCOMP]: this index is not Metaculus, not a market price, not a personal p(doom) poll, and not Ord’s or Carlsmith’s number alone. --- ## How to reproduce Inputs (nothing else scores): - `docs/SOURCE-LEDGER-TIER-A.md` — governing human ledger - `data/tier-a-ledger.json` — structured extract of that ledger - `data/event-ledger-v1.json` — empty qualifying events - `docs/METHODOLOGY-v0.1.md` — rules, implemented in `compute/snapshot.py` ```bash python3 compute/snapshot.py python3 compute/snapshot.py --check cd compute && python3 -m unittest test_snapshot.py ``` The script re-applies eligibility (primary numeric, gauge map, H-century, weight > 0). A JSON `in: true` flag cannot smuggle LeCun or Yudkowsky into the median. If `qualifying_events` is non-empty, compute exits rather than inventing an unfrozen news pipeline. Outputs: `snapshots/v1.json`, `design/mockups/frontpage-v1/index.html`, `design/push/frontpage-v1.html`.