A demand forecast should be evaluated using only information that would have been available when the prediction was made. Joining historical sales to today's corrected tables can give a model knowledge that the real planning team never possessed. The resulting backtest may look impressive while offering little evidence about future performance.
Feature stores and historical retrieval tools make this problem easier to address, but they do not remove the need to define time correctly. The useful question for a forecasting project is not just “which model predicts best?” It is “which facts did the planner actually know at each decision?”
Two timestamps tell different stories
Event time describes when something happened: a sale, a delivery or a stock count. Availability time describes when the forecasting system could use the information. A transaction can occur on Monday and reach the warehouse on Wednesday. A correction entered on Friday can describe an event from the previous month.
If a Monday forecast uses that Friday correction during evaluation, the model has received information from its future. Sorting the final dataset by the original event date does not fix the problem.
A third timestamp can also matter: when a business assumption became approved. A promotional calendar may contain an event planned for next month, but the planner knew about it last week. Future event dates are not automatically leakage; the question is whether the information was legitimately available at the forecast cutoff.
Rebuild one planning meeting
Consider an illustrative wholesaler placing replenishment orders every Tuesday morning. It forecasts demand for the following weeks. Its inputs include sales, stock availability, supplier lead times and approved promotions.
For a historical Tuesday, reconstruct the information that existed before the ordering decision. Exclude late-arriving sales records until they became available. Use the promotion plan known at that time, even if marketing revised it later. Preserve the lead-time estimate the buyer could see, rather than replacing it with the eventual delivery duration.
This does not require recreating every screen of an old system. It requires a reproducible snapshot of the inputs and rules relevant to the prediction. The scenario is fictional and illustrates a proposed design, not a measured forecasting result.
A point-in-time join needs a precise meaning
Feast's point-in-time join documentation explains historical feature retrieval relative to entity timestamps. Its current guidance also distinguishes event-time matching from filtering by when a value was created or made available, with support depending on the offline store.
That distinction matters when selecting a feature platform. Ask what each configured timestamp means and how backfilled corrections are handled. A field named “created at” is useful only if it represents the availability boundary your evaluation needs.
For a small forecasting workload, versioned warehouse tables and carefully written historical queries may be sufficient. A feature store can become valuable when several models reuse features or when offline training and online serving need coordinated definitions. The objective is a trustworthy dataset, not ownership of a particular category of infrastructure.
Keep the test period ahead of the training period
The scikit-learn TimeSeriesSplit documentation explains why ordinary cross-validation can train on future data and evaluate on the past. It provides time-ordered splits and describes conditions needed for comparable folds.
For the wholesaler, use successive historical ordering dates. At each date, train or update using only eligible past information, generate the relevant forecast horizon and evaluate against later observed outcomes. Keep feature calculation inside this boundary.
A temporal split alone is insufficient if a feature was computed from the full dataset beforehand. For example, an average supplier delay calculated across the entire year can leak later information into an early-year fold. Version the feature transformation and record its input window alongside the model.
Examine the errors that change a purchasing decision
A single average metric can hide useful differences. Separate fast-moving products, intermittent products, new products and periods affected by promotions. Compare the model with a clearly defined baseline that follows the same information restrictions.
For the illustrative project, the review might inspect these cases:
- ▸A promotion was cancelled after the forecast cutoff.
- ▸Sales arrived late from one branch.
- ▸A product was unavailable, reducing observed sales.
- ▸A supplier lead-time correction was backdated.
- ▸A new item had no meaningful sales history.
Stockouts require particular care. Observed sales can understate demand when customers could not buy the item. Document whether the project forecasts recorded sales or estimates underlying demand, and how unavailable-stock periods are represented. Do not conceal that modelling choice inside a feature name.
Make the dataset part of the deliverable
The forecasting handover should include the cutoff definition, source timestamp meanings, correction policy and query that builds each evaluation snapshot. Save dataset versions and model outputs so an analyst can reproduce a disputed result.
A useful review starts with a small set of historical ordering decisions. Ask a planner to inspect what the system believes was known on each date. This can reveal a data problem before a larger training run makes it expensive to investigate.
Our MLOps production guide covers the wider model lifecycle. Historical correctness is one foundation for that lifecycle, especially when a production model is retrained repeatedly as source records change.
When to pause the model purchase
If the project cannot distinguish original records from later corrections, begin with data reconstruction and a modest baseline. An elaborate model comparison cannot repair an evaluation dataset that contains future knowledge.
That does not mean delaying all value. A first delivery can expose missing arrival timestamps, produce a reproducible rolling backtest and show planners where data arrives too late for the current decision schedule. Those findings can improve the workflow independently of a new algorithm.
For AI and data integration, describe your planning cutoff and source systems to TuniCyberLabs. We can help scope the data work needed before forecast accuracy becomes a meaningful buying criterion.
