Most forecasting work is evaluated against a feeling. It looks reasonable. It tracks roughly. Nobody asks the one question that would establish whether it is worth anything.
What would we have got for free?
The naive baseline
A naive baseline is the simplest possible forecast:
- Last period repeated. Next month equals this month.
- Seasonal naive. This month next year equals this month last year.
- Trailing average. Next month equals the average of the last three.
These take minutes to construct and cost nothing. In stable businesses they are often startlingly accurate — frequently within a few percent.
That is the bar. A model that does not clear it has not added information; it has added confidence.
Why this is not a technicality
A forecast carries the authority of analysis. People plan against it. An organization with no forecast plans for a range of outcomes. An organization with a confident, inaccurate forecast plans for one specific wrong outcome — and the authority of the method is exactly what makes the error expensive.
So a model that performs worse than last-period-repeated is not neutral. It is worse than having nothing, because it displaces the appropriate uncertainty.
How to construct the standard properly
Three steps, all before any modelling begins:
1. Measure the baseline on held-back data. Separate a validation period — say the most recent six months — before you start. Run the naive method against it. Record the error in the units the business actually uses.
2. Ask the business what accuracy would change a decision. Not what would be impressive. If the forecast were within X%, would you act on it? That number comes from the person who owns the decision, not from the analyst.
3. State the standard as both. The model must beat the baseline by an agreed margin and land inside the band where the business would act. Either alone is insufficient — a model can beat the baseline and still be too imprecise to be useful.
Lock the validation set
Separate it before modelling and do not look at it. Every additional evaluation against the same held-back data is a form of fitting to it. A model tuned against a validation set you have inspected is not validated; it is rehearsed.
One evaluation run. If it fails, you may not re-tune and re-run against the same data and call it a pass.
The uncomfortable implication
This standard means forecasting work can fail — publicly, against a number agreed in advance.
That is the point. A forecasting method that cannot fail cannot be trusted when it succeeds, and the discipline of agreeing the bar beforehand is what separates a forecast from an opinion with a chart attached.
Practical check
Put the baseline beside every model result in the management readout. Report both errors in business units, not only a statistical score. That makes the incremental value visible and prevents complexity from being mistaken for accuracy.
Predictive Intelligence engagements agree the accuracy standard before any model is built, and report honestly when it is not met.