Risk & Volatility
Model Risk And The Limits Of Backtests
A strategy tested against historical data will almost always look good, because the testing process itself selects for rules that happened to fit the past.

Testing a rule against historical data seems like the obvious way to evaluate it. The process contains several biases that make favourable results the expected outcome rather than an informative one.
Searching a dataset finds patterns
Any sufficiently large body of data contains apparent relationships that arose by chance. Trying many rules against it will produce some that fit well.
Because only the successful rule is presented, the number of variations tried before arriving at it is invisible to whoever reads the result.
The reported performance therefore describes the outcome of a search rather than the merit of the rule, and the two are easily confused.
Rules get adjusted to fit
Thresholds, holding periods and filters are typically tuned until results improve. Each adjustment fits the rule more closely to the particular history being tested.
A rule with several tunable settings can be made to fit almost any dataset, which tells us about the flexibility of the rule rather than about the market.
The more settings a strategy has, the more scepticism its historical results warrant, since fitting capacity grows with them.
Survivorship distorts the data itself
Databases of securities and funds often exclude those that ceased to exist, so testing against them measures performance among the survivors only.
Since failures are systematically absent, results are biased upwards by an amount that is difficult to estimate after the fact.
Whether a dataset is free of this bias is a question worth asking before any conclusion drawn from it is taken seriously.
Costs and execution are usually understated
Tests commonly assume trades occur at closing prices with modest costs, which understates spreads, market impact and the difficulty of dealing in size.
Strategies trading frequently or in less liquid securities are affected most, and the gap between tested and achievable results widens accordingly.
A rule whose advantage is smaller than realistic dealing costs has no advantage at all once implemented.
Conditions change beneath the rule
Market structure, participants and regulation all change over time, so a relationship that held in one period need not persist into another.
Widely adopted rules also change the conditions that produced them, since capital following a pattern affects the prices generating it.
None of this makes historical analysis useless, but its output is a hypothesis about how something behaved rather than a projection of what it will do.
Also by Nour Haddad
- Comparing two funds properlyFunds & ETFs
- What happens if a fund or platform failsFunds & ETFs
- Thematic and sector fundsFunds & ETFs
- Factor investing, explained honestlyFunds & ETFs





