Flip a coin ten times and count the heads. You might get seven. Flip it a thousand times and you'll land close to five hundred. Nobody judges a coin by ten flips.
Traders judge their systems by ten trades constantly. I did it for the better part of two years — three losing weeks and I was already rewriting the entry logic.
The math nobody checks before judging a strategy
Here's the number that changed how I think about this. Researchers who've run the math on strategy evaluation put the danger zone at roughly 10 to 30 trades. In that range, a strategy with a genuinely positive 45% win rate can show anywhere from 10% to 80% winners, just from randomness.
One documented example: a GBP/USD system posted a 33% win rate over its first 12 trades and looked broken. Its actual expectancy, once the sample grew, was +0.122R per trade — solidly positive. The strategy was never the problem. Twelve trades was.
Most sources agree that reliability starts showing up around 100 trades, gets convincing past 200, and only becomes strongly significant past 500. A calendar month of systematic trading, depending on your frequency, might hand you five to twenty trades. That's not a sample. That's a rumor.
I know this because I closed out a strategy in its fourth month, annoyed at a string of red weeks, and rebuilt something "better." Eight months later I ran the numbers on the version I'd killed. It was profitable across the full period. My replacement wasn't.
Why traders fall for it anyway
The reason this keeps happening isn't stupidity. It's a well-documented pattern in how people process small samples — the assumption that even a short run of results represents the true, long-run picture. Kahneman and Tversky named it back in 1971: the law of small numbers.
Applied to markets, it means we treat a bad month like it's telling us something about our edge. Usually it isn't. It's just what variance looks like on the way to a good year.
Here's why traders fall for it specifically, not just people in general. Months are the unit we're wired to track. Rent is monthly. Reviews are monthly.
Statements arrive monthly, too. So we grade our trading on the same calendar, even though a strategy doesn't know what a month is. It only knows trades, and thirty of them tell you almost nothing.
There's a second trap hiding inside the first. A losing month feels personal in a way a losing trade doesn't. One bad trade is a data point.
A bad month feels like a verdict on you — your judgment, your system, your whole approach. That emotional weight is exactly why people abandon working strategies at the worst possible time: right after the variance, right before it reverts.
Count trades, not weeks
I don't grade a system by the month anymore. I set a trade-count threshold before I'm even allowed to have an opinion — usually 100 closed trades, sometimes fewer for higher-frequency setups where that number arrives in weeks instead of months. Below that threshold, the only thing I'm allowed to check is whether a hard risk rule got breached: max drawdown, a broken assumption, an execution failure. Above it, I look at the actual distribution, not the story my last four weeks are telling me.
If you want a rule you can use tonight: write down the trade count where your strategy's stats start meaning something, and don't let yourself re-evaluate before you hit it. Track it the way you'd track a countdown, not a calendar. It's a small habit, but it's the difference between killing a strategy because of noise and killing one because it actually stopped working.
Most of what looks like "the market changed" is really "my sample was too small to know." Systematic trading doesn't remove that trap by itself — the system still runs on your account, and you still have to decide when to trust it. But it does give you a number to check instead of a feeling to manage, and that's worth something. If you want to see what that looks like in practice, we publish our full trade-by-trade data at v33systematic.com.