The Currency Times

The Ledger

Our own forecasts, dated when made and scored when they resolve. This is the only page here that can be wrong, which is the reason it exists.

Every other page on this site describes how forecasting should be judged. This one submits to it.

The rules

1

The question must resolve without argument

A date, a source of truth, and a binary outcome fixed in advance. "Will the incumbent win" resolves. "Will the economy improve" does not, and does not get entered.

2

The probability is recorded before the outcome is known

Timestamped at entry. An entry is never edited after the fact. If reasoning changes, a second dated entry is added and both stand.

3

Every entry made is an entry scored

An entry stands whatever the outcome, because a ledger that can be quietly pruned carries no information. The whole value of this page rests on it being impossible to edit backwards.

4

The comparison is the base rate, not zero

Scores are reported against what simply predicting the historical frequency would have achieved. Beating nothing is not skill.

Current standing

0forecasts resolved
0open
Brier score
skill against base rate

The ledger opens with the site, so the first entry made here will genuinely be the first. Scores begin the moment forecasts start resolving.

The calibration curve appears once twenty forecasts have resolved, because that is roughly where a curve starts carrying information rather than decoration. It is the same standard applied everywhere else here, turned on ourselves.

What an entry looks like

Illustrative only. No forecast below has been made.

Made
Question
Stated
Outcome
Named storm makes landfall as a major hurricane
0.35
open
Treaty signed by all parties before the stated deadline
0.60
open
Reported figure revised downward at first revision
0.72
open

Calibration, and why it is drawn the way it is

Below is the instrument, running on adjustable inputs rather than on real entries. Drag the slider to see what different kinds of wrong look like. A perfectly calibrated forecaster sits on the diagonal. Everyone starts off it.

overconfident
stated probability observed frequency perfect calibration
0.00reliability, lower is better
0.00resolution, higher is better
0.00Brier score

Reliability measures the gap between what was said and what happened. Resolution measures whether the forecasts discriminated between cases at all. A forecaster who says fifty per cent to everything scores perfectly on the first and zero on the second, and is useless.

NextProblemswhat we still cannot answer