Four thousand backtests. That's what a modest grid search produces — one trend strategy, lookbacks from 10 to 200 days, a handful of exit thresholds. In my second year of building systems I ran exactly that, and the top row of the results table showed a 61-day lookback returning 212%, close to triple anything near it.
The 55-day version barely broke even. At 70 days the strategy lost money. And I very nearly traded the 61.
An edge is a region, not a point
That gap between neighbors is the tell. A real edge doesn't live on one setting. If trend following works at 61 days but fails at 55 and 70, you haven't found an edge — you've found the one path through historical noise that happened to sidestep every drawdown.
The market doesn't know what a 61-day lookback is. Noise does.
Quants have a name for results like mine: parameter islands. An island is a combination that scores brilliantly while everything around it drowns. The opposite is a plateau — a wide, flat region where 50, 60, and 70 days all produce similar, unremarkable, positive numbers.
Robustness looks boring on a map. That's how you recognize it.
The degradation data backs this up. A standard rule of thumb in quant research is to expect a strategy's Sharpe ratio to drop by a third to a half the moment it trades data it wasn't fit on.
And that's the average case. When Suhonen, Lennkh and Perez examined 215 quantitative strategies marketed by investment banks, the median Sharpe deterioration between backtest and live performance was 73%. Professional desks, with real budgets and review committees, still shipped peaks.
Why everyone picks the peak anyway
The software is built for it. Every optimizer sorts its output by return or Sharpe, and the whole interface funnels your eyes to row one. Deliberately choosing the 40th-best result feels like ordering the 40th-best meal on the menu.
There's a psychological trap underneath. A 212% backtest feels like a discovery, like you dug something real out of the data. The honest read is usually the reverse — the more spectacular a single cell looks against its neighbors, the larger the share of its performance that is noise.
Excitement about a backtest is inversely correlated with its reliability. That might just be my experience talking, but I haven't found a counterexample yet.
And the standard verification makes things worse. Most people confirm an optimized strategy by running it again on the same data with a different metric, which confirms nothing. Robert Pardo published the fix in 1992 — walk-forward analysis: optimize on one window, test on the next, repeat. Thirty-four years later, most retail backtesting still skips it.
Select for the plateau
After any optimization, ignore the ranking column and look at neighborhoods instead. Take your candidate setting and move every parameter 10 to 20% in each direction. A 20-day lookback should still work at 18 and 22; a 2% stop should still work at 1.8 and 2.4. If performance collapses anywhere in that range, the setting is an island — discard it, whatever the headline number was.
Then pick the middle of the widest plateau you can find, even if that cell ranks 40th.
A 40th-best backtest that survives perturbation will beat the best one that doesn't, because live trading is one long perturbation. Fees drift. Volatility regimes shift. Fills arrive late. A setting that couldn't survive a 10% nudge in-sample has no chance against all of that at once.
This selection habit is most of what separates a system that survives its third year from a backtest that dies in its third month. Our BTC strategy's parameters sit on plateaus, not peaks — the full backtest data and methodology are public at v33systematic.com if you want to check that claim yourself.