How Many Predictions Are Needed to Judge a Forecaster?
There is no single number of predictions that proves whether someone is a good forecaster.
A record of three correct predictions may be encouraging, but it provides little evidence about long-term ability. A record of 300 comparable predictions is more informative, although it still needs to be interpreted in light of the subject, difficulty, time horizon, and scoring rules.
The right question is not simply “How many predictions?” It is:
How much evidence does this record provide about repeatable performance on this particular forecasting task?
Why short records are unstable
With a small number of predictions, each result has a large effect on the overall percentage.
For example:
| Record | Hits | Decided predictions | Hit rate |
|---|---|---|---|
| A | 2 | 3 | 67% |
| B | 20 | 30 | 67% |
| C | 200 | 300 | 67% |
The hit rate is the same in all three rows, but the amount of evidence is not. One additional result changes Record A from 67% to either 50% or 75%, depending on the outcome. The same result changes Record C by less than one percentage point.
This is why a percentage should always be accompanied by the number of decided predictions.
A record is evidence, not proof
Even a large record cannot prove that a forecaster has permanent skill. Conditions may change, the forecaster may change their method, and future events may differ from the historical sample.
A record can provide evidence that performance was better, worse, or indistinguishable from a reasonable comparison during a defined period. It cannot guarantee that the same performance will continue.
This distinction is particularly important when a record is used to make decisions about future investments, policies, or bets. Historical results can inform judgment, but they do not remove uncertainty.
The amount of evidence depends on the task
Different forecasts require different sample sizes.
A sports forecaster who makes several hundred similar game predictions may build a useful record over one season. A geopolitical analyst who makes only a few carefully defined long-term forecasts each year may need many years before the record becomes large.
The following factors affect how much evidence is needed:
- the natural frequency of predictions;
- the length of the forecast horizon;
- the variability of outcomes;
- the difficulty of the subject;
- whether the predictions are independent;
- whether the scoring rules are binary or graded;
- how much performance variation is acceptable.
A low-frequency task cannot be evaluated fairly using the same sample-size expectations as a high-frequency task.
Comparable predictions matter more than raw volume
One hundred unrelated predictions may be less useful than 50 well-defined predictions about a consistent task.
For example, a record that combines election forecasts, cryptocurrency price targets, sports picks, and macroeconomic claims may contain a large number of cases but still be difficult to interpret. The forecaster may be stronger in one area and weaker in another.
Records are more informative when predictions are grouped by relevant characteristics such as:
- subject matter;
- forecast horizon;
- prediction type;
- confidence level;
- market or geographic region;
- scoring method.
The goal is not to discard broad records. It is to avoid treating a large mixed collection as if it measured one simple ability.
Confidence intervals help show uncertainty
An observed hit rate is an estimate, not an exact measurement of a forecaster’s underlying ability.
Suppose a forecaster gets 6 of 10 predictions right. The observed rate is 60%, but the true long-run rate could plausibly be considerably higher or lower. With 600 hits out of 1,000 predictions, the observed rate is still an estimate, but it is usually more precise.
Statistical confidence intervals can express this uncertainty. They provide a range of values that are consistent with the observed record under a chosen model and confidence level.
Confidence intervals are not magic guarantees. They depend on assumptions, and ordinary binary intervals may be inappropriate when predictions are correlated, selected over time, or drawn from changing conditions. Still, showing uncertainty is generally more informative than presenting a small-sample percentage as if it were exact.
Avoid arbitrary thresholds
It is tempting to create rules such as:
- fewer than 10 predictions: no record;
- 10 to 49 predictions: preliminary record;
- 50 or more predictions: reliable record.
These labels may be useful for presentation, but they should not be mistaken for universal scientific boundaries. Ten easy and comparable predictions may tell us more than 50 vague or highly dependent predictions.
If categories are used, they should be described as practical reporting conventions. The underlying counts and limitations should remain visible.
Dependence can make a record look larger than it is
Predictions are not always independent.
Imagine a forecaster making ten predictions about the same event, changing the wording slightly as new information arrives. Those ten entries do not provide the same amount of evidence as ten forecasts about unrelated events.
Other examples of dependence include:
- multiple predictions based on one common assumption;
- a series of price targets that all depend on the same market movement;
- several forecasts about one election or tournament;
- repeated predictions copied from the same source.
Dependent cases can still be useful to record, but they should not automatically be treated as independent observations when judging consistency.
Selection effects can inflate apparent performance
The sample must also be checked for how predictions entered the record.
Performance may look stronger if:
- only specific claims were recorded;
- unsuccessful predictions were deleted or ignored;
- the record begins after a period of poor performance;
- only predictions with clear outcomes were included;
- forecasts were selected because they were unusually confident;
- the forecaster changed style after early failures.
A transparent record should explain its inclusion rules. Otherwise, a large sample can create false confidence if the sample itself is selective.
Rolling performance is often more useful than one total
A single lifetime hit rate can hide changes over time.
Consider a forecaster with this sequence:
- Years 1–2: 70% hit rate;
- Years 3–4: 48% hit rate;
- Year 5: 65% hit rate.
The overall rate may appear reasonable while concealing a period of deterioration or improvement. Rolling windows, annual summaries, or performance by forecast horizon can reveal whether results are stable.
Recent performance should not automatically replace the full record. It should be shown alongside it, with enough context to avoid overreacting to short-term variation.
What a reader should ask
When reviewing a forecaster’s record, ask:
- How many predictions were actually decided?
- How many remain pending or unverifiable?
- Are the predictions comparable?
- Were they recorded before the outcomes?
- Are multiple entries dependent on the same event?
- Were unsuccessful or vague predictions included consistently?
- Is the performance stable across time and subject?
- How much uncertainty surrounds the reported rate?
These questions often matter more than whether the headline number is above or below a particular threshold.
A practical reporting approach
A useful track record can report several levels of evidence:
- the full count of recorded predictions;
- the count of decided predictions;
- hits, misses, partials, pending, and unverifiable cases;
- the hit rate and its formula;
- performance by period and category;
- confidence or calibration information when available;
- notes about selection and dependence.
This approach avoids pretending that one number can summarize every relevant feature of a forecasting record.
Conclusion
More predictions generally provide more evidence, but quantity alone does not establish forecasting skill. The predictions must be clearly defined, consistently recorded, relevant to the same task, and interpreted with uncertainty.
A short record can be a useful early signal. A long record can provide stronger evidence. Neither should be treated as a guarantee.
The most responsible evaluation combines sample size with transparency about what was predicted, how it was scored, how the cases were selected, and whether the results remain consistent over time.
This article is educational and is not investment, political, or sports-betting advice. Historical forecasting performance does not guarantee future results.