Forecast accuracy
Forecast accuracy scorecard
World Monitor scores every forecast it publishes once the outcome is knowable, over a rolling 180-day window. This is the standing record: the scores, the calibration, the sample sizes, and the parts that are not measurable yet.
Record status
- Availability: Captured
- Read from the live scoring service on 2026-09-10 and frozen into this page.
- Freshness: Current
- The scoring service generated these numbers 12 hours before the capture read them. The threshold is 36 hours, matching the one the scoring service applies to itself.
- Coverage: Measurable
- The headline cohort has 180 scored forecasts, enough to publish a score.
Over the current 180-day window, World Monitor's headline cohort scores a Brier of 0.118 across 180 scored forecasts, against 0.25 for answering 0.5 to everything.
Lower Brier is better. A Brier score is the mean squared error of a probability forecast, so 0 is perfect and answering 0.5 to everything scores 0.25. Log score is harsher on confident mistakes, and lower is better there too.
What the headline number counts
The headline Brier and log score cover 180 scored forecasts: every scored entry except those from the origins the scorecard excludes, currently bet_engine, state_derived. That exclusion costs 310 scored entries. An entry that carries no origin at all is filed as unknown and is counted, so this cohort is defined by what it leaves out and not by any property of the entries it keeps. The all-scored figure beside it is the same window with every origin put back, which is why the two numbers differ.
Resolution ledger
| Ledger stage | Entries |
|---|---|
| Entries in the rolling window | 958 |
| Resolved | 772 |
| Scored | 490 |
| Voided | 36.5% of 772 resolved entries (282 entries) |
| Awaiting a judge | 96 |
| Still open, not yet resolvable | 90 |
| Scored share of the ledger | 51.1% of 958 entries |
Brier/log score over resolved YES/NO published forecast windows; VOID and pending entries are counted for coverage but excluded from accuracy math.
Calibration
| Predicted probability | Forecasts | Average predicted | Actually happened | Brier |
|---|---|---|---|---|
| 0% to 10% | 40 | 0.060 | 0.0% of 40 forecasts | 0.005 |
| 10% to 20% | 53 | 0.143 | 1.9% of 53 forecasts | 0.034 |
| 20% to 30% | 80 | 0.254 | 32.5% of 80 forecasts | 0.227 |
| 30% to 40% | 168 | 0.351 | 35.1% of 168 forecasts | 0.227 |
| 40% to 50% | 106 | 0.412 | 36.8% of 106 forecasts | 0.234 |
| 50% to 60% | 29 | 0.531 | 27.6% of 29 forecasts | 0.266 |
| 60% to 70% | 13 | 0.640 | 46.2% of 13 forecasts | 0.274 |
| 90% to 100% | 1 | 0.930 | 100.0% of 1 forecasts | 0.005 |
Accuracy by domain
| Domain | Resolved | Scored | Voided | Brier | Log score |
|---|---|---|---|---|---|
| conflict | 36 | 12 | 66.7% of 36 resolved | 0.273 | 0.730 |
| cyber | 146 | 144 | 1.4% of 146 resolved | 0.076 | 0.278 |
| energy | 24 | 24 | 0.0% of 24 resolved | 0.192 | 0.567 |
| geopolitical | 6 | 6 | 0.0% of 6 resolved | 0.242 | 0.645 |
| infrastructure | 89 | 11 | 87.6% of 89 resolved | 0.286 | 0.764 |
| macro | 5 | 5 | 0.0% of 5 resolved | 0.235 | 0.662 |
| market | 346 | 272 | 21.4% of 346 resolved | 0.242 | 0.679 |
| military | 20 | 6 | 70.0% of 20 resolved | 0.270 | 0.731 |
| political | 6 | 0 | 100.0% of 6 resolved | Insufficient sample | Insufficient sample |
| supply_chain | 94 | 10 | 89.4% of 94 resolved | 0.237 | 0.666 |
Accuracy by generation origin
| Generation origin | In the headline cohort | Resolved | Scored | Voided | Brier | Log score |
|---|---|---|---|---|---|---|
| bet_engine | No, excluded | 299 | 299 | 0.0% of 299 resolved | 0.236 | 0.664 |
| legacy_detector | Yes | 362 | 176 | 51.4% of 362 resolved | 0.114 | 0.366 |
| state_derived | No, excluded | 78 | 11 | 85.9% of 78 resolved | 0.241 | 0.676 |
| unknown | Yes | 33 | 4 | 87.9% of 33 resolved | 0.289 | 0.773 |
Against prediction markets
Measured over all scored entries that overlapped a liquid market, not over the narrower headline cohort. On 78 such resolved questions the forecast Brier was 0.155 and the market Brier was 0.073. The published delta is the market Brier minus the forecast Brier, and lower is better, so a negative delta means the market scored better. Here the delta is -0.081, so on this sample the market scored better.
What this page does not publish
- Calibration buckets that scored nothing are omitted from the table rather than drawn as a zero: 70-80, 80-90 are empty in this window.
- No confidence intervals. Each figure is published with the number of forecasts behind it instead, because stating a sample size without an interval is honest and inventing an interval is not. Tracking: issue #7072.
- No accuracy for the 24-hour, 7-day and 30-day projections shown in the product. Those horizons are not scored yet, so nothing here describes them. Tracking: issue #7075.
- No individual forecasts, resolution evidence, judge inputs or archive locations. This page publishes aggregates only.
Related reference
Download: scorecard.json. Source: docs/snapshots/crawlable-live-pulse-2026-09-09.json. Numbers generated 2026-09-10 06:02:24 UTC and read on 2026-09-10. Live results come from the credentialed forecast scorecard endpoint, which this page freezes so it can be read without one.