The backtest is a mirror, not a window
A price-prediction model that looks brilliant on its own training data has done something closer to revision than prophecy. It has been shown the answers, in the order they happened, and asked to fit them. When the same model is then described as “predicting the market”, the quiet assumption is that the future will keep grading the model the way the past did. That assumption is the thing worth examining.
The honest unit of evidence is out-of-sample, walk-forward performance: train on a window, predict the next slice, roll the window forward, repeat. The number that comes out of that procedure is smaller than the in-sample number, sometimes dramatically so, and it is the only number that resembles the situation a reader would actually face.
What we look for instead
When we read an AI price-prediction claim, we ask three questions before we read any further:
- Was the test walk-forward? A single train/test split on a trending series can flatter almost any model.
- Did the test include costs? Spread, slippage and financing eat most paper alpha; a claim that ignores them is a claim about a market that does not exist.
- Did the test survive its own regime? A model trained on a decade of falling rates has not been tested against a rising-rate regime, and vice versa.
The limit, plainly
A model can be genuinely useful as a descriptive lens on past structure and still be a poor predictor of next week’s price. The useful claim is usually the smaller one: “this model detected a pattern that held, on average, across these windows, under these costs”. That is a long way from “this model predicts the market”, and the distance between the two is where most of the trouble lives.