Why Most Backtests Lie (And How to Build One That Doesn’t)

There’s a moment every discretionary trader has. You build a strategy, run a backtest, and the equity curve goes up and to the right. You feel a surge of confidence. You go live. And then nothing works.

This isn’t bad luck. It’s a measurement problem.

Most retail backtests are wrong by design — not because the trader is incompetent, but because the way we’re taught to backtest is fundamentally flawed. The result is a number that tells you what you want to hear, not what the market will actually deliver.

Here’s what’s actually going on.

The backtest is always in your favour

When you build a strategy and test it on the same data you used to build it, you’re not measuring performance. You’re measuring fit. The strategy has been — consciously or not — shaped by the very data it’s now being tested on.

This is called in-sample overfitting, and it’s the reason most backtests don’t survive contact with live markets.

Every parameter you tuned, every signal threshold you adjusted, every filter you added — each one was informed by the historical data in front of you. The backtest looks good because you’ve already seen the answer.

The split that changes everything

The fix is structural, not cosmetic. You divide your historical data into two parts:

  • In-sample (IS): Where you build and tune the strategy
  • Out-of-sample (OOS): Data the strategy has never seen — used only for final validation

The OOS result is the number that matters. Not your IS Sharpe ratio. Not your IS win rate. Only what the strategy does on data it wasn’t built on.

The exact ratio between the two matters less than the principle. What you choose should reflect how much historical data you have and how many walk-forward windows you need downstream. The split is always applied chronologically — never randomly. Random splits leak future information into the past, and that contaminates everything that follows.

Why a single OOS test isn’t enough either

One OOS window is better than none. But a single test period can still get lucky. If your OOS window happened to coincide with a market regime that suited your signal, you’ll look good by accident.

The solution is walk-forward validation: you test across multiple OOS windows, each representing a different market period. Trending markets, ranging markets, high volatility, low volatility. If your strategy holds up across all of them, that’s a result worth taking seriously.

If it only works in some windows, you need to understand why — before going anywhere near live capital.

The number most traders ignore

The Sharpe ratio is the standard performance metric, and it matters. But most traders focus on the average Sharpe across windows and miss the minimum.

A strategy that averages SR 2.0 across four windows sounds solid. But if one of those windows produced SR 0.4, you have a regime dependency problem. Your strategy might stop working during the next market phase you haven’t seen yet.

The minimum OOS Sharpe across all walk-forward windows is the number I look at first. That floor has to clear a hard threshold — not the average. The floor. If it doesn’t clear it, the strategy doesn’t get deployed. No exceptions.

What a dishonest backtest looks like

Here are the most common ways backtests mislead — usually unintentionally:

  • Lookahead bias: The signal uses data that wouldn’t have been available at the time of the trade
  • Survivorship bias: Testing only on assets that still exist, ignoring those that failed or were delisted
  • Overfitting: Too many parameters relative to the number of trades in the dataset
  • Fee neglect: Ignoring execution costs — slippage and commissions that quietly destroy live returns
  • Tiny trade count: A 90% win rate on 20 trades means almost nothing statistically

Any one of these can turn a mediocre strategy into a beautiful backtest.

What an honest backtest requires

It requires committing to rules before you look at results. The moment you start adjusting parameters because the equity curve looked wrong, you’ve contaminated the test.

Some principles I hold to:

  1. Define the signal logic completely before running any backtest
  2. Split data chronologically before touching the model
  3. Never look at OOS results until IS development is finished
  4. Include realistic fees and slippage — especially for futures
  5. Require a minimum trade count before treating any result as meaningful
  6. Test across multiple OOS windows, not just one

Systematic trading is only as trustworthy as the testing behind it. A beautiful backtest built on contaminated data is worse than no backtest at all — because it gives you false confidence to deploy real capital.

The goal isn’t to produce a good-looking chart. It’s to produce a reliable estimate of what to expect from a live strategy.

That’s a harder problem. And it’s exactly the kind of problem worth solving correctly.


AlgoBlueprint covers the methodology behind building and deploying systematic trading strategies — without the noise. Subscribe to the newsletter or follow on Twitter and YouTube.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.