How many trades before a number means anything
This is the chapter that would prevent the most damage if it were the first thing anyone read. Almost every confident statement traders make about their own performance — this setup works, mornings are better for me, that instrument does not suit me — is made from a sample far too small to support it.
The intuition to fix is this: the noise in a win rate falls with the square root of the number of trades, not with the number of trades. Going from 25 trades to 100 does not quarter the uncertainty, it halves it. To halve it again you need 400. This is why the honest answer to "how many trades do I need" is usually a much larger number than people want to hear.
Some rough anchors. From 20 trades you can see gross incompetence and nothing else. From 50 you can begin to trust the sign of the expectancy if it is large. From 100 you can compare two setups if the difference between them is big. Below about 30 observations, a subgroup — one setup, one session, one weekday — should not be reported with a win rate at all, because the number will look precise and be meaningless.
The trap that follows is subgroup fishing. Slice a hundred trades by setup, session, weekday, instrument and hour and you have made dozens of comparisons; some of them will look impressive from chance alone. The discovery that "my Tuesday NY-session gold longs win 78 per cent" from nine trades is not a discovery. It is the arithmetic of looking at many small groups.
The defence is to decide what you are testing before you look, to require a minimum sample per group, and to treat anything found by slicing as a hypothesis rather than a finding — something to check on the next fifty trades rather than to act on today. Evidence enforces the minimum for you in several places and refuses to print a rate below it; that refusal is a feature, and it is the difference between a statistic and a decoration.
What to take away
- Uncertainty falls with the square root of the sample — 4× the trades to halve the error.
- Below ~30 observations a subgroup win rate should not be reported at all.
- Anything found by slicing is a hypothesis to test, not a finding to act on.
Where it goes wrong
- Declaring a setup dead after eight losing trades.
- Acting on a subgroup discovered by slicing the same data many ways.
- Reporting a win rate to one decimal place from twelve trades.
This chapter, measured against your own trades
In the app the same chapter ends in your figures rather than an example: how often you did the thing it describes, over your last ninety days. You pick one change to make, and Evidence checks afterwards whether it actually changed — from your journal, arithmetic, no opinion involved. Questions you get wrong come back a week later and again a month after that.
Open the free plan →Check that it stuck
Answers shown — in the app these are asked before you see them, and the ones you get wrong come back after a week.