Workstream B · Research dossier
Forecasts
and limits.
When does the available information justify predicting a later readout, despite noise and calibration uncertainty?
Model specified, evaluation not yet run
Definitions
The forecast rule
- :
- information available before the forecast is issued.
- :
- the declared nonempty set of calibration parameters being considered.
- :
- the model probability of later outcome for parameter and information .
- :
- a chosen error tolerance with .
Issue a prediction of only when
Otherwise, no prediction is issued.
Under the stated model, with the true calibration parameter inside the declared set, the rule bounds the probability of disagreement with the specified later readout.
This is a forecast rule. It does not assert that other candidates have become physically impossible or that a realised record already exists.
Limits
What the guarantee does not cover
- A numerical grid over calibration values is not automatically a certified lower bound over the whole set. A demonstration that uses a grid is numerical screening unless a valid bound is implemented.
- High accuracy must not be reported while concealing a very low prediction rate.
- A marginal calibration-coverage statement does not become an identical error guarantee conditional on a rare prediction being issued.
- Retrospective fitting is separate from prospective evaluation. A model adjusted after seeing outcomes must be evaluated again on fresh held-out data.
Evaluation
What any reported evaluation must show
- The information available at prediction time.
- The model and how the calibration set was constructed.
- The chosen error tolerance.
- The fraction of eligible cases receiving a prediction.
- The error rate among issued predictions.
- The no-prediction cases.
- Dataset size and uncertainty intervals.
- Calibration-set coverage or its stated assumptions.
- Performance on held-out data.
Dossier summary
Model, evidence and criteria
Proposed evaluation
Evaluate a fixed prediction rule prospectively on suitable held-out records, reporting both errors and occasions when no prediction is issued.
Inputs and assumptions
Information available before the forecast, a declared calibration set, a chosen tolerance, and suitable held-out records.
Proposed measurements
Prediction rate, error rate among issued predictions and no-prediction cases on held-out records.
Evidence currently available
The rule is defined. No prospective evaluation on held-out data has been reported.
Comparison or rejection criterion
Pre-registered rule and tolerance; reported prediction rate, error rate among issued predictions, no-prediction cases and uncertainty intervals on held-out data. Reject the rule for a dataset when the error rate among issued predictions is significantly above δ under the stated coverage assumptions.
Reproducibility
No dataset or code has been published. A pre-registered protocol would fix h, Θ, δ and the held-out split before evaluation.
Current foundational references
