Workstream B · Research dossier

Forecasts
and limits.

When does the available information justify predicting a later readout, despite noise and calibration uncertainty?

Model specified, evaluation not yet run

Definitions

The forecast rule

hh:
information available before the forecast is issued.
Θ\Theta:
the declared nonempty set of calibration parameters being considered.
pr(θ∣h)p_r(\theta \mid h):
the model probability of later outcome rr for parameter θ\theta and information hh.
δ\delta:
a chosen error tolerance with 0<δ<120 < \delta < \tfrac12.

Issue a prediction of rr only when

inf⁡θ∈Θpr(θ∣h)≥1−δ\inf_{\theta \in \Theta} p_r(\theta \mid h) \ge 1 - \delta

Otherwise, no prediction is issued.

Under the stated model, with the true calibration parameter inside the declared set, the rule bounds the probability of disagreement with the specified later readout.

This is a forecast rule. It does not assert that other candidates have become physically impossible or that a realised record already exists.

Limits

What the guarantee does not cover

  • A numerical grid over calibration values is not automatically a certified lower bound over the whole set. A demonstration that uses a grid is numerical screening unless a valid bound is implemented.
  • High accuracy must not be reported while concealing a very low prediction rate.
  • A marginal calibration-coverage statement does not become an identical error guarantee conditional on a rare prediction being issued.
  • Retrospective fitting is separate from prospective evaluation. A model adjusted after seeing outcomes must be evaluated again on fresh held-out data.

Evaluation

What any reported evaluation must show

  1. The information available at prediction time.
  2. The model and how the calibration set was constructed.
  3. The chosen error tolerance.
  4. The fraction of eligible cases receiving a prediction.
  5. The error rate among issued predictions.
  6. The no-prediction cases.
  7. Dataset size and uncertainty intervals.
  8. Calibration-set coverage or its stated assumptions.
  9. Performance on held-out data.

Dossier summary

Model, evidence and criteria

Proposed evaluation

Evaluate a fixed prediction rule prospectively on suitable held-out records, reporting both errors and occasions when no prediction is issued.

Inputs and assumptions

Information available before the forecast, a declared calibration set, a chosen tolerance, and suitable held-out records.

Proposed measurements

Prediction rate, error rate among issued predictions and no-prediction cases on held-out records.

Evidence currently available

The rule is defined. No prospective evaluation on held-out data has been reported.

Comparison or rejection criterion

Pre-registered rule and tolerance; reported prediction rate, error rate among issued predictions, no-prediction cases and uncertainty intervals on held-out data. Reject the rule for a dataset when the error rate among issued predictions is significantly above δ under the stated coverage assumptions.

Reproducibility

No dataset or code has been published. A pre-registered protocol would fix h, Θ, δ and the held-out split before evaluation.

Current foundational references