Forecast accuracy
Forecast accuracy scorecard
World Monitor scores every forecast it publishes once the outcome is knowable, over a rolling 180-day window. This is the standing record: the scores, the calibration, the sample sizes, and the parts that are not measurable yet.
Record status
- Availability: Captured
- Read from the live scoring service on 2026-09-28 and frozen into this page.
- Freshness: Current
- The scoring service generated these numbers 23 hours before the capture read them. The threshold is 36 hours, matching the one the scoring service applies to itself.
- Coverage: Measurable
- The headline cohort has 223 scored forecasts, enough to publish a score.
Over the current 180-day window, World Monitor's headline cohort scores a Brier of 0.125 across 223 scored forecasts, against 0.25 for answering 0.5 to everything.
Lower Brier is better. A Brier score is the mean squared error of a probability forecast, so 0 is perfect and answering 0.5 to everything scores 0.25. Log score is harsher on confident mistakes, and lower is better there too.
What the headline number counts
The headline Brier and log score cover 223 scored forecasts: every scored entry except those from the origins the scorecard excludes, currently bet_engine, state_derived. That exclusion costs 381 scored entries. An entry that carries no origin at all is filed as unknown and is counted, so this cohort is defined by what it leaves out and not by any property of the entries it keeps. The all-scored figure beside it is the same window with every origin put back, which is why the two numbers differ.
Resolution ledger
| Ledger stage | Entries |
|---|---|
| Entries in the rolling window | 1,153 |
| Resolved | 944 |
| Scored | 604 |
| Voided | 36.0% of 944 resolved entries (340 entries) |
| Awaiting a judge | 98 |
| Still open, not yet resolvable | 111 |
| Scored share of the ledger | 52.4% of 1,153 entries |
Brier/log score over resolved YES/NO published forecast windows; VOID and pending entries are counted for coverage but excluded from accuracy math.
Calibration
| Predicted probability | Forecasts | Average predicted | Actually happened | Brier |
|---|---|---|---|---|
| 0% to 10% | 41 | 0.060 | 0.0% of 41 forecasts | 0.004 |
| 10% to 20% | 80 | 0.142 | 5.0% of 80 forecasts | 0.058 |
| 20% to 30% | 101 | 0.253 | 31.7% of 101 forecasts | 0.223 |
| 30% to 40% | 203 | 0.352 | 35.0% of 203 forecasts | 0.227 |
| 40% to 50% | 126 | 0.412 | 38.9% of 126 forecasts | 0.237 |
| 50% to 60% | 34 | 0.529 | 23.5% of 34 forecasts | 0.267 |
| 60% to 70% | 13 | 0.640 | 46.2% of 13 forecasts | 0.274 |
| 80% to 90% | 5 | 0.846 | 0.0% of 5 forecasts | 0.715 |
| 90% to 100% | 1 | 0.930 | 100.0% of 1 forecasts | 0.005 |
Accuracy by domain
| Domain | Resolved | Scored | Voided | Brier | Log score |
|---|---|---|---|---|---|
| conflict | 50 | 18 | 64.0% of 50 resolved | 0.395 | 1.045 |
| cyber | 183 | 181 | 1.1% of 183 resolved | 0.075 | 0.280 |
| energy | 36 | 36 | 0.0% of 36 resolved | 0.243 | 0.696 |
| geopolitical | 6 | 6 | 0.0% of 6 resolved | 0.242 | 0.645 |
| infrastructure | 89 | 11 | 87.6% of 89 resolved | 0.286 | 0.764 |
| macro | 8 | 8 | 0.0% of 8 resolved | 0.315 | 0.833 |
| market | 424 | 328 | 22.6% of 424 resolved | 0.240 | 0.676 |
| military | 23 | 6 | 73.9% of 23 resolved | 0.270 | 0.731 |
| political | 7 | 0 | 100.0% of 7 resolved | Insufficient sample | Insufficient sample |
| supply_chain | 118 | 10 | 91.5% of 118 resolved | 0.237 | 0.666 |
Accuracy by generation origin
| Generation origin | In the headline cohort | Resolved | Scored | Voided | Brier | Log score |
|---|---|---|---|---|---|---|
| bet_engine | No, excluded | 370 | 370 | 0.0% of 370 resolved | 0.241 | 0.677 |
| legacy_detector | Yes | 430 | 214 | 50.2% of 430 resolved | 0.110 | 0.360 |
| state_derived | No, excluded | 106 | 11 | 89.6% of 106 resolved | 0.241 | 0.676 |
| unknown | Yes | 38 | 9 | 76.3% of 38 resolved | 0.473 | 1.248 |
Against prediction markets
Measured over all scored entries that overlapped a liquid market, not over the narrower headline cohort. On 97 such resolved questions the forecast Brier was 0.147 and the market Brier was 0.068. The published delta is the market Brier minus the forecast Brier, and lower is better, so a negative delta means the market scored better. Here the delta is -0.079, so on this sample the market scored better.
What this page does not publish
- Calibration buckets that scored nothing are omitted from the table rather than drawn as a zero: 70-80 are empty in this window.
- No confidence intervals. Each figure is published with the number of forecasts behind it instead, because stating a sample size without an interval is honest and inventing an interval is not. Tracking: issue #7072.
- No accuracy for the 24-hour, 7-day and 30-day projections shown in the product. Those horizons are not scored yet, so nothing here describes them. Tracking: issue #7075.
- No individual forecasts, resolution evidence, judge inputs or archive locations. This page publishes aggregates only.
Related reference
Download: scorecard.json. Source: World Monitor forecast scorecard snapshot. Numbers generated 2026-09-27 06:05:24 UTC and read on 2026-09-28. Live results come from the credentialed forecast scorecard endpoint, which this page freezes so it can be read without one.