Are the probabilities honest?
A forecaster who says "30%" should be right about 30% of the time. Below, every resolved range-break call from every board is grouped by the chance it gave, and compared with how often those ranges actually broke. Points on the dashed line are well calibrated; above it, ranges broke more often than forecast; below it, less often.
Dots: share that broke, sized by number of calls. Bars: 95% interval for that share (Wilson). This is a description, not a verdict: calls on the same pool on neighbouring days share most of their window, so the effective sample is far smaller than the count. The pass/fail tests are the sealed verdicts on the verification page.
By forecast band
| Forecast | Calls | Avg forecast | Broke | 95% interval |
|---|---|---|---|---|
| 0–10% | 27 | 4.1% | 0.0% | 0%–12% |
| 10–20% | 21 | 15.6% | 19.0% | 8%–40% |
| 20–30% | 17 | 25.6% | 23.5% | 10%–47% |
| 30–40% | 8 | 32.1% | 25.0% | 7%–59% |
| 40–50% | 7 | 46.7% | 14.3% | 3%–51% |
| 50–60% | 7 | 58.5% | 0.0% | 0%–35% |
| 60–70% | 4 | 63.8% | 25.0% | 5%–70% |
| 70–80% | 8 | 76.1% | 75.0% | 41%–93% |
| 80–90% | 28 | 85.7% | 82.1% | 64%–92% |
| 90–100% | 51 | 99.4% | 96.1% | 87%–99% |
By board
| Board | Resolved | Avg forecast | Broke | Brier | Skill vs flat rate |
|---|---|---|---|---|---|
| Market Weather (coins) | 102 | 60.8% | 53.9% | 0.114 | +0.54 |
| Market Weather (pools) | 76 | 52.5% | 46.1% | 0.121 | +0.51 |
Brier: average squared error, lower is better. Skill vs flat rate: 1 − Brier ÷ the Brier of always forecasting the board's own observed break rate (a rate known only with hindsight). Above 0 means the board told pools apart; at or below 0 means it did not.