Questions this project cannot currently answer, published in the hope that somebody else can. Correspondence on any of them is welcome.
A list of things that are wrong, or unresolved, or where the honest answer is that nobody appears to know.
The correction for how many candidates were searched requires knowing how many were searched. Results produced before anyone counted cannot be corrected, only discarded. Almost every published finding in applied forecasting is in this position. Is there any principled recovery, or is discarding genuinely the only option?
Regime-switching models fit history beautifully and identify the switch retrospectively. Every practical use requires knowing the state now. Is the retrospective advantage irreducible, or an artefact of how these models are usually estimated?
Twenty-five thousand observations with a two-hundred-period window are not twenty-five thousand independent observations. Rules of thumb exist. A defensible general formula does not appear to.
Simple averaging beats most individuals, and stops working when the individuals read each other. Extremising helps empirically without a satisfying theory behind it. What is the right aggregation when the dependence structure is unknown?
Random matrix theory transfers cleanly because the mathematics does not care what the matrix describes. Critical phenomena transfer less cleanly. Agent-as-spin models may not transfer at all. Is there a criterion for which borrowings are legitimate, other than seeing whether they work?
Maximum entropy answers this when the constraints are known. Most interesting questions arrive with no reference class and no constraints anyone can state. Reference-class forecasting begs the question of which class.
The stated principle is to maximise sharpness subject to calibration. In practice both are estimated with error on small samples and the constraint is never exactly satisfied. What is the correct decision rule when neither can be measured precisely?
Every model has conditions under which it stops working. Almost none carry a signal that those conditions have arrived. Is a general early-warning statistic possible, or is this necessarily domain-specific?
The open questions are not scattered evenly. Four of the eight touch Validation, and three touch Forecast evaluation, which says something about where the next useful work sits.
A problem that touches two fields is usually a problem about the boundary between them, which is generally where the interesting difficulty lives.
If any of these has a settled answer, the useful correspondence is a reference rather than an argument.