← Back to the blog
Notes

Deflated Sharpe and why an overfitting test must be built into every optimisation

López de Prado's deflated Sharpe ratio, plateau detection, and forward out-of-sample windows — what a real overfitting test looks like in practice.

In the first two posts in this series we covered the problem (TradingView backtests lie) and one of the data causes (Dukascopy M1 isn't enough). This post is about the test.

The question we want an overfitting test to answer is simple:

Given that I just searched N parameter combinations on a train set, how much should I deflate the Sharpe ratio I found before I trust it as an out-of-sample estimate?

The math behind this was worked out properly in López de Prado's 2014 paper The Deflated Sharpe Ratio and How to Predict Strategy Performance. The intuition is older than that — every quant who's been burned by parameter search knows the rule of thumb.

Three independent signals

Deepwick's overfitting test is not one number. It's three independent signals, each catching a different failure mode, combined into a verdict.

Signal 1 — Deflated Sharpe Ratio

The deflated Sharpe ratio adjusts your in-sample Sharpe for the number of trials. If you ran 1,000 parameter sets and picked the highest Sharpe, the DSR tells you the Sharpe you should expect out-of-sample, accounting for that selection effect.

The formula is exact: it uses the expected maximum of N i.i.d. Sharpe ratios (where N is the number of trials), corrected for skewness and kurtosis of the return distribution. The output is a probability: the chance that your reported Sharpe is a false positive.

When this fires: any time you optimised over more than ~10 parameters. With 100+ combinations the deflation is severe. With 1,000+ it's brutal.

Signal 2 — Plateau detection

A robust strategy has a plateau in parameter space — the metric is similar across a wide range of nearby parameter values. An overfit strategy has a peak — one parameter combination is dramatically better than its neighbours.

Deepwick measures the plateau width by computing the metric across the full parameter grid and looking at the ratio of the peak value to the median value of the top quartile. A plateau ratio below ~0.3 (peak is at most 30% better than the median of the top region) suggests robustness. Above 1.0 strongly suggests overfitting.

When this fires: any time you have at least a 2D parameter grid (which is most real strategies). 1D grid searches are especially prone to missing this signal because the plateau can look wide by accident.

Signal 3 — Forward out-of-sample window

The third signal is the simplest and the most informative: we hold out the last 20% of the data and run the optimised strategy on it without any further tuning.

This isn't novel. It's the basic train/test split every ML practitioner learns in week one. But it's surprising how rarely it appears as a mandatory part of an optimisation workflow.

Deepwick runs the forward window by default. You can disable it, but the disable is logged and visible — both to you and to anyone reviewing the strategy later.

When this fires: always. If your forward window performance is dramatically worse than your in-sample performance, that's the smoking gun. If it's similar, that's the actual evidence you needed.

Why we don't ship just Monte Carlo

A lot of traders hear "overfitting test" and assume it means Monte Carlo simulation. Monte Carlo is great at estimating the expected distribution of returns under the null hypothesis. It is silent on the parameter search bias problem.

To be concrete: if I run 1,000 Monte Carlo paths on your in-sample equity curve, I'll tell you how likely your Sharpe ratio is under random trade ordering. That's useful. It tells me nothing about whether your parameters were cherry-picked. The two questions are different. Monte Carlo only answers one of them.

Deepwick's overfitting test runs all three signals and reports a verdict:

  • Robust — all three signals pass. Forward window within tolerance, plateau wide, deflated Sharpe above threshold.
  • Caution — one signal failed. Reasonable to proceed but worth inspecting.
  • Likely overfit — two or more signals failed. The strategy's in-sample performance is not a reliable estimate of forward performance.

The verdict is shown next to the strategy's reported metrics in the Strategy Tester. It's not a black-box score — each signal's value is reported, so you can disagree with the verdict if you have context.

The marketplace consequence

This matters most for the marketplace we'll launch alongside Deepwick.

If authors can publish strategies that look great in-sample but are overfit, the marketplace becomes a junk drawer. Buyers can't tell from the headline Sharpe ratio alone whether the strategy is real. The overfitting test gives them a single signal they can compare strategies on — same way shoppers compare products by review score, not by trusting the merchant's description.

Authors with genuinely robust strategies get to advertise that. The marketplace's reputation accrues to the authors whose strategies survive the test, not to those who happen to find a peak.

This is the structural reason Deepwick has the overfitting test built in. Not as a feature. As the basis of trust for everything else.

— Enrique

overfittingdeflated-sharpebacktestingmonte-carlo