How to Score a Prediction Fairly
Scoring a prediction means comparing the original claim with an outcome using rules that were defined clearly enough to apply consistently.
The arithmetic is usually easy. The difficult part is deciding what the prediction actually said, when it should be evaluated, what evidence counts, and how to handle outcomes that are incomplete or ambiguous.
A fair scoring process should make those decisions transparent. It should not change the standard depending on whether the prediction succeeded or failed.
Start with the original claim
The first step is to preserve the prediction as it was originally made.
Record:
- the exact wording or a faithful quotation;
- the source URL or publication;
- the date and time of publication;
- the subject of the prediction;
- any stated conditions or assumptions;
- the relevant target, direction, or outcome.
The original version matters because later commentary may be clearer than the initial claim. A forecaster may explain what they “really meant” after the outcome is known, but the score should be based primarily on what a reasonable reader could have understood at the time.
That does not mean every prediction must be interpreted in the most literal or hostile way. It means the interpretation should be documented and grounded in the original context.
Define the prediction before judging it
Before looking at the result, write down what would count as success.
For example, consider the statement:
“I expect Company A to trade above $150 by the end of the year.”
A scoring rule might specify:
- the relevant security and exchange;
- whether an intraday price or closing price is required;
- whether adjusted or unadjusted prices are used;
- the exact evaluation date and time zone;
- whether the prediction succeeds if the price briefly crosses $150 or must close above it.
If these details are decided only after the outcome is known, the scorer may unintentionally favor one interpretation. Defining them in advance reduces that risk.
Use a clear evaluation deadline
Every time-bound prediction needs a deadline. The deadline may be explicit in the original statement or inferred from a conventional period such as “this season” or “by year-end.” If it must be inferred, the assumption should be recorded.
Useful deadline information includes:
- the calendar date;
- the time zone;
- the end of the relevant trading session, game, election, or reporting period;
- the treatment of events that occur after the deadline.
An outcome that happens after the deadline may be informative, but it should not automatically turn a missed prediction into a successful one. Timing is part of the claim.
Choose appropriate evidence
The evidence should measure the outcome described by the prediction as directly as possible.
Examples include:
- official election results for an election forecast;
- final scores and league records for sports predictions;
- exchange or market data for price forecasts;
- official economic releases for macroeconomic predictions;
- primary documents or reputable reporting for event-based claims.
Secondary commentary can help explain what happened, but it should not replace primary evidence when primary evidence is available. The source and retrieval date should be retained so that another reviewer can reproduce the evaluation.
Use outcome categories when binary scoring is too simple
Some predictions can be scored cleanly as hit or miss. Others cannot.
A useful classification may include:
| Category | Meaning |
|---|---|
| Hit | The defined outcome occurred within the stated conditions and deadline. |
| Miss | The defined outcome did not occur within the stated conditions and deadline. |
| Partial | Some material part of the prediction occurred, but the full claim was not satisfied. |
| Pending | The deadline or outcome has not yet arrived. |
| Unverifiable | The claim or outcome cannot be evaluated reliably from available evidence. |
These categories should not be confused with a probability score. A prediction can be a “miss” even if it was assigned a 30% probability and therefore had a reasonable chance of failing. The category describes the result; the probability describes the forecast’s uncertainty.
Be careful with partial outcomes
Partial outcomes are one of the most difficult cases because they require judgment.
Suppose someone predicts that a stock will rise above $150 within six months. It reaches $145, then falls. The direction may have been broadly correct, but the stated threshold was not reached. Whether this is a partial result depends on the published scoring rules.
Similarly, a political prediction may correctly identify a broad development but give the wrong timing or specific mechanism. A sports prediction may identify the winning team but miss a spread or total-points condition.
The important principles are:
- Define partial success before reviewing individual results when possible.
- Do not silently count partial outcomes as full hits.
- Explain which part of the claim was satisfied and which part was not.
- Apply the same standard to similar predictions.
If a numerical score is used for partial outcomes, the formula should be published and tested against examples.
Handle conditional predictions explicitly
Conditional statements often have the form:
“If interest rates remain high, housing prices could decline.”
This is not the same as saying:
“Housing prices will decline.”
The first statement contains a condition. To score it, the evaluation must determine whether the condition was satisfied and whether the predicted result followed within the relevant period.
Conditional claims that never activate should usually be labeled separately rather than treated as ordinary hits or misses. Otherwise, a collection of many possible scenarios can create the appearance of accurate forecasting without requiring a clear unconditional commitment.
Do not score vague opinions as precise predictions
Statements such as “volatility is likely,” “the market looks weak,” or “major changes may be coming” may be useful commentary. But they are often too broad to score objectively.
The right response is not to force them into a numerical result. Mark the claim as unclear or unverifiable unless the surrounding context supplies a reasonable, documented interpretation.
A scoring system should reward specificity by making specific claims easier to evaluate. It should not manufacture precision that the original speaker did not provide.
Use independent review for difficult cases
A second reviewer can improve consistency when a prediction is ambiguous or the result is disputed.
Independent review works best when:
- the reviewers see the original claim and scoring rules;
- each reviewer records their reasoning before discussing the case;
- disagreements are documented rather than hidden;
- the final decision can be revisited if new evidence appears.
The purpose is not to create the appearance of mathematical certainty. It is to reduce the influence of one person’s assumptions and make disagreements visible.
Publish the reasoning, not only the label
A result such as “Hit” or “Miss” is easier to trust when readers can see why it was assigned.
A useful evaluation note should identify:
- the relevant facts;
- the evidence source;
- the comparison with the original claim;
- the reason for the selected category;
- any unresolved limitation.
This is especially important for claims involving multiple conditions, long time horizons, or outcomes that cannot be represented by a single number.
Review the rules for consistency
Scoring rules should be tested against a sample of past predictions before they are used for rankings or comparisons. This can reveal unclear definitions, unexpected edge cases, or rules that systematically favor one type of claim.
After the rules are in use, changes should be documented. A methodology may improve over time, but changing the rules without identifying which historical records were affected makes comparisons difficult.
Historical scores should not be quietly rewritten simply because a later rule is more convenient. If a correction is necessary, the original result, revised result, reason for the change, and date of revision should be recorded.
A practical scoring checklist
Before assigning a result, ask:
- What exactly was claimed?
- Is the original source and timestamp available?
- What subject, direction, target, and conditions were specified?
- What is the evaluation deadline and time zone?
- What evidence directly measures the outcome?
- Does the outcome satisfy the stated conditions?
- Is the result a hit, miss, partial, pending, or unverifiable case?
- Would the same rule produce the same result for a similar prediction?
- Can another reader understand the reasoning from the record?
If these questions cannot be answered, the appropriate result may be “unverifiable” rather than a forced judgment.
Conclusion
Fair prediction scoring is less about producing a neat percentage than about preserving a consistent chain from claim to evidence to result.
The strongest systems define predictions before outcomes are known, use appropriate evidence, account for timing and conditions, separate partial and unresolved cases, and publish the reasoning behind difficult decisions.
These practices do not remove uncertainty from forecasting. They make uncertainty and disagreement easier to see—and that is what allows readers to evaluate prediction records responsibly.
This article is educational and is not investment, political, or sports-betting advice. Historical forecasting performance does not guarantee future results.