Expectancy per setup, not per account
An account-level expectancy is an average over everything you did, and averages hide the structure that matters. Most traders with a flat curve are not running one mediocre method — they are running one good method and one bad one, and the bad one is eating the good one.
Splitting expectancy by setup is therefore the highest-value slice there is, and unlike most slicing it is not fishing: the groups were defined by you in advance, when you named the setup, rather than discovered by searching. That distinction is exactly the one the sample-size chapter draws, and it is what makes this analysis legitimate where weekday-and-hour slicing is not.
What you are looking for is not the best performer. It is the one with a clearly negative expectancy over a decent sample, because removing it is the single change with the largest and most reliable effect on the curve — and it costs nothing, requires no new skill, and needs no market to cooperate.
The same analysis is worth repeating one level down for the setups that survive: within a working setup, is there a condition that separates the good instances from the bad? Time of day and instrument are the two that most often carry real signal, and both should still meet the minimum sample before you believe them.
A warning about acting on this. Cutting a setup after a sample of fifteen is the same error as trusting one after fifteen, pointed the other way. The rule is symmetric: the sample that would let you add a setup is the sample you need to drop one. Anything smaller and you are trading on noise in both directions.
What to take away
- A flat curve is usually one good method plus one bad one, not one mediocre one.
- Setup groups are defined in advance, which is why slicing by them is legitimate.
- Removing a negative-expectancy setup is the cheapest improvement available.
Where it goes wrong
- Optimising the best setup instead of removing the worst.
- Dropping a setup on a sample too small to have added it.
- Slicing by setup and then by five more dimensions and believing the result.
This chapter, measured against your own trades
In the app the same chapter ends in your figures rather than an example: how often you did the thing it describes, over your last ninety days. You pick one change to make, and Evidence checks afterwards whether it actually changed — from your journal, arithmetic, no opinion involved. Questions you get wrong come back a week later and again a month after that.
Open the free plan →Check that it stuck
Answers shown — in the app these are asked before you see them, and the ones you get wrong come back after a week.