How Many Predictions Are Needed to Judge a Forecaster?

There is no single number of predictions that proves whether someone is a good forecaster.

A record of three correct predictions may be encouraging, but it provides little evidence about long-term ability. A record of 300 comparable predictions is more informative, although it still needs to be interpreted in light of the subject, difficulty, time horizon, and scoring rules.

The right question is not simply “How many predictions?” It is:

How much evidence does this record provide about repeatable performance on this particular forecasting task?

Why short records are unstable

With a small number of predictions, each result has a large effect on the overall percentage.

For example:

RecordHitsDecided predictionsHit rate
A2367%
B203067%
C20030067%

The hit rate is the same in all three rows, but the amount of evidence is not. One additional result changes Record A from 67% to either 50% or 75%, depending on the outcome. The same result changes Record C by less than one percentage point.

This is why a percentage should always be accompanied by the number of decided predictions.

A record is evidence, not proof

Even a large record cannot prove that a forecaster has permanent skill. Conditions may change, the forecaster may change their method, and future events may differ from the historical sample.

A record can provide evidence that performance was better, worse, or indistinguishable from a reasonable comparison during a defined period. It cannot guarantee that the same performance will continue.

This distinction is particularly important when a record is used to make decisions about future investments, policies, or bets. Historical results can inform judgment, but they do not remove uncertainty.

The amount of evidence depends on the task

Different forecasts require different sample sizes.

A sports forecaster who makes several hundred similar game predictions may build a useful record over one season. A geopolitical analyst who makes only a few carefully defined long-term forecasts each year may need many years before the record becomes large.

The following factors affect how much evidence is needed:

A low-frequency task cannot be evaluated fairly using the same sample-size expectations as a high-frequency task.

Comparable predictions matter more than raw volume

One hundred unrelated predictions may be less useful than 50 well-defined predictions about a consistent task.

For example, a record that combines election forecasts, cryptocurrency price targets, sports picks, and macroeconomic claims may contain a large number of cases but still be difficult to interpret. The forecaster may be stronger in one area and weaker in another.

Records are more informative when predictions are grouped by relevant characteristics such as:

The goal is not to discard broad records. It is to avoid treating a large mixed collection as if it measured one simple ability.

Confidence intervals help show uncertainty

An observed hit rate is an estimate, not an exact measurement of a forecaster’s underlying ability.

Suppose a forecaster gets 6 of 10 predictions right. The observed rate is 60%, but the true long-run rate could plausibly be considerably higher or lower. With 600 hits out of 1,000 predictions, the observed rate is still an estimate, but it is usually more precise.

Statistical confidence intervals can express this uncertainty. They provide a range of values that are consistent with the observed record under a chosen model and confidence level.

Confidence intervals are not magic guarantees. They depend on assumptions, and ordinary binary intervals may be inappropriate when predictions are correlated, selected over time, or drawn from changing conditions. Still, showing uncertainty is generally more informative than presenting a small-sample percentage as if it were exact.

Avoid arbitrary thresholds

It is tempting to create rules such as:

These labels may be useful for presentation, but they should not be mistaken for universal scientific boundaries. Ten easy and comparable predictions may tell us more than 50 vague or highly dependent predictions.

If categories are used, they should be described as practical reporting conventions. The underlying counts and limitations should remain visible.

Dependence can make a record look larger than it is

Predictions are not always independent.

Imagine a forecaster making ten predictions about the same event, changing the wording slightly as new information arrives. Those ten entries do not provide the same amount of evidence as ten forecasts about unrelated events.

Other examples of dependence include:

Dependent cases can still be useful to record, but they should not automatically be treated as independent observations when judging consistency.

Selection effects can inflate apparent performance

The sample must also be checked for how predictions entered the record.

Performance may look stronger if:

A transparent record should explain its inclusion rules. Otherwise, a large sample can create false confidence if the sample itself is selective.

Rolling performance is often more useful than one total

A single lifetime hit rate can hide changes over time.

Consider a forecaster with this sequence:

The overall rate may appear reasonable while concealing a period of deterioration or improvement. Rolling windows, annual summaries, or performance by forecast horizon can reveal whether results are stable.

Recent performance should not automatically replace the full record. It should be shown alongside it, with enough context to avoid overreacting to short-term variation.

What a reader should ask

When reviewing a forecaster’s record, ask:

  1. How many predictions were actually decided?
  2. How many remain pending or unverifiable?
  3. Are the predictions comparable?
  4. Were they recorded before the outcomes?
  5. Are multiple entries dependent on the same event?
  6. Were unsuccessful or vague predictions included consistently?
  7. Is the performance stable across time and subject?
  8. How much uncertainty surrounds the reported rate?

These questions often matter more than whether the headline number is above or below a particular threshold.

A practical reporting approach

A useful track record can report several levels of evidence:

This approach avoids pretending that one number can summarize every relevant feature of a forecasting record.

Conclusion

More predictions generally provide more evidence, but quantity alone does not establish forecasting skill. The predictions must be clearly defined, consistently recorded, relevant to the same task, and interpreted with uncertainty.

A short record can be a useful early signal. A long record can provide stronger evidence. Neither should be treated as a guarantee.

The most responsible evaluation combines sample size with transparency about what was predicted, how it was scored, how the cases were selected, and whether the results remain consistent over time.

This article is educational and is not investment, political, or sports-betting advice. Historical forecasting performance does not guarantee future results.