Our own forecasts, dated when made and scored when they resolve. This is the only page here that can be wrong, which is the reason it exists.
Every other page on this site describes how forecasting should be judged. This one submits to it.
A date, a source of truth, and a binary outcome fixed in advance. "Will the incumbent win" resolves. "Will the economy improve" does not, and does not get entered.
Timestamped at entry. An entry is never edited after the fact. If reasoning changes, a second dated entry is added and both stand.
An entry stands whatever the outcome, because a ledger that can be quietly pruned carries no information. The whole value of this page rests on it being impossible to edit backwards.
Scores are reported against what simply predicting the historical frequency would have achieved. Beating nothing is not skill.
The ledger opens with the site, so the first entry made here will genuinely be the first. Scores begin the moment forecasts start resolving.
The calibration curve appears once twenty forecasts have resolved, because that is roughly where a curve starts carrying information rather than decoration. It is the same standard applied everywhere else here, turned on ourselves.
Illustrative only. No forecast below has been made.
Below is the instrument, running on adjustable inputs rather than on real entries. Drag the slider to see what different kinds of wrong look like. A perfectly calibrated forecaster sits on the diagonal. Everyone starts off it.
Reliability measures the gap between what was said and what happened. Resolution measures whether the forecasts discriminated between cases at all. A forecaster who says fifty per cent to everything scores perfectly on the first and zero on the second, and is useless.