LEARN · EN · easyquanttrading.com

How to compare two trading strategies fairly

Comparing strategies is where most selection mistakes happen, and the mistakes are rarely about statistics. They are about measuring two things in two different ways and then treating the numbers as comparable. The fixes are simple, which makes it frustrating that they are so often skipped.

The four things that must match

Before any comparison is worth doing, four conditions have to hold. If any of them fails, the difference you are measuring is the difference in conditions, not the difference in strategies.

  • Same time period. A strategy tested over a bull market will beat one tested over a range, regardless of quality.
  • Same instrument and data source. Different brokers supply different bars; two strategies tested on different feeds are not comparable.
  • Same cost assumptions. Spread, commission and slippage must be applied identically.
  • Same position sizing. Fixed lot versus risk-based sizing produces different equity curves from identical signals.

Compare against a benchmark, not only against each other

Choosing the better of two bad strategies is a real failure mode, and it happens because the comparison is framed as a choice between them. Add a third column: what would happen if you simply held the instrument.

The benchmark is crude and it is also remarkably hard to beat on many markets after costs. A strategy that beats buy-and-hold on return but with three times the drawdown is not obviously better — it may simply be a leveraged version of the benchmark with extra steps.

Look at the worst case before the average

Averages are where bad comparisons hide. Two strategies with the same average return can have very different worst quarters, and it is the worst quarter that determines whether you keep going.

CompareRather thanBecause
Maximum drawdownAverage returnIt decides usability
Worst rolling 3-month returnTotal returnAverages hide the pain
Trade count in each periodTotal trade countReveals uneven activity
Behaviour in the worst periodBehaviour overallThat is when you will judge it
Largest single lossAverage lossOne trade can define the year

Five substitutions that make a comparison decision-relevant rather than flattering.

Account for the search that produced them

If one of the two strategies was selected from two hundred candidates and the other was chosen for a reason you can state, they are not on equal footing even if their measured metrics are identical. The first one's numbers are inflated by the selection process.

Practical version: count how many variants you tried for each. If the counts differ substantially, discount the one with the larger search, and note that this is exactly the correction the Deflated Sharpe Ratio formalises.

Prefer the one you can explain

When two candidates are genuinely close on every measurable dimension, the tie-break should be explanation. Which one can you describe in a sentence, including who is on the other side of the trade and why they would keep taking it?

This is not a soft criterion. A strategy you can explain is one where you will notice when its premise stops holding. A strategy you cannot explain is one you will keep running on faith, past the point where the edge has gone.

A comparison checklist

  • Have I tested both over exactly the same period on the same data?
  • Are the cost and fill assumptions identical, and are they realistic?
  • Is the position sizing rule the same, including rounding?
  • Have I included a buy-and-hold benchmark for the same period?
  • Am I comparing worst cases, or only averages?
  • Do I know how many variants of each I tried before choosing?
  • Can I state in one sentence why the winner should keep working?

FAQ

Is the strategy with the higher Sharpe always better?
No. The Sharpe ratio is blind to the sequence of losses, which is the part you actually experience. Two strategies with equal Sharpe can differ by a factor of four in maximum drawdown.
What if I only have one strategy to evaluate?
Compare it to a benchmark — holding the instrument, or holding cash. A strategy that cannot beat a simple benchmark after realistic costs is not adding value, however good its internal statistics look.
Does a longer backtest always make a better comparison?
Longer is generally better because it includes more conditions, but only if the data is reliable over the whole period and the market's structure has not changed fundamentally. A twenty-year test on an instrument that only became liquid five years ago is not twenty years of evidence.

More guides

Not investment advice. Historical results do not guarantee future performance. EasyQuant is a research factory — you execute on accounts you control.