Statistics reference
What each test measures, what it does not, and how to read a result that says no.
Every test below is implemented in quant-integrity,
AGPL-3.0, and can be read, audited and run without this server. What the
hosted server adds is the trial count these corrections are applied against.
Deflated Sharpe ratio#
stats_deflated_sharpe(strategy_id="...")
Bailey and López de Prado's correction. A Sharpe ratio selected as the best of N attempts is biased upward even when every underlying edge is zero — the maximum of N noisy draws is large by construction. The deflated Sharpe asks: given N trials, given the skew and kurtosis of these returns, what is the probability the true Sharpe exceeds zero?
There is no n_trials parameter. The server reads the ledger. See
The trial ledger.
Non-normality matters here, and crypto returns are not normal. The correction accounts for skew and excess kurtosis directly, which is why a fat tailed series with a nominally attractive Sharpe can deflate to nothing.
Probability of backtest overfitting (PBO / CSCV)#
stats_pbo(...)
Combinatorially symmetric cross-validation. The data is split into equal subsets; every combination is used as in-sample with the complement as out-of-sample; the configuration that ranked best in-sample is located in the out-of-sample ranking.
PBO is the frequency with which the in-sample winner lands below the out-of-sample median.
- PBO near 0 — selection is picking up something real.
- PBO near 0.5 — selection is uninformative.
- PBO above 0.5 — selection is anti-predictive. Choosing the in-sample best is worse than choosing at random, which usually means the thing being selected on is noise plus a cost artefact.
CSCV requires equal subsets. Unequal ones break the symmetric construction the statistic is defined on.
Minimum backtest length#
stats_min_backtest_length(...)
How many years of history N trials actually require before an observed Sharpe of a given size is distinguishable from the best of N random draws:
MinBTL ≈ 2 · ln(N) / E[max SR]²
It is denominated in years. This is the tool that most often ends a line of research, which is the point.
Resampling to a finer timeframe does not escape it. Finer bars are autocorrelated, not new information, and they multiply cost drag.
Stationary bootstrap confidence intervals#
stats_bootstrap_ci(...)
Politis–Romano stationary bootstrap on Sharpe and expectancy. Blocks of geometrically distributed length preserve the autocorrelation structure that an i.i.d. bootstrap would destroy.
On a marginal strategy the interval will straddle zero. That is the tool working. An interval that straddles zero is not a failure to be re-run with different settings; it is the answer.
Bootstrap p-values use (1 + count) / (B + 1), so a finite number of
replications can never report exactly zero.
Reality check and SPA#
stats_reality_check(...)
White's Reality Check and Hansen's Superior Predictive Ability test: is the best strategy in a set better than a benchmark, accounting for the fact that you looked at the whole set?
Romano–Wolf stepdown reports the full rejection count alongside the truncated display, so a truncated view cannot be mistaken for the complete result.
Effective number of tests#
stats_effective_tests(universe="...")
M_eff from the correlation eigenspectrum. See Effective breadth — the short version is that 22 liquid Coinbase pairs are about 3 independent tests by participation ratio and 7 by entropy, and which of those two you should use depends on whether you are raising a hurdle or widening an interval.
Triple-barrier labelling#
stats_triple_barrier(...)
Labels candidate entries by which barrier is hit first — profit target, stop, or time limit. This lets you label many candidate entries without increasing trading frequency, which matters because frequency is what costs money.
Reports a breakeven hit rate. When barriers are tighter than the round-trip fee this exceeds 1.0, meaning no achievable hit rate suffices.
Alternative bar sampling#
data_resample(kind="dollar"|"volume", ...)
Dollar and volume bars sample on activity rather than clock time, which usually improves the normality of returns. The tool reports a before/after normality comparison rather than assuming resampling helped, because sometimes it does not.
Purged combinatorial cross-validation#
CPCV improves the reliability of an estimate and enables PBO. It is not extra sample: it reuses the same data and never reduces N.
Purging covers both edges of each test block, and an embargo separates folds. See Splits and the holdout.
How to read a result that says no#
A confidence interval straddling zero, a PBO above 0.5, a deflated Sharpe with a p-value of 0.4, a MinBTL of 40 years — these are not malfunctions and they are not invitations to adjust parameters until they change.
They are the tool doing the only thing that makes the other numbers worth anything. A backtest suite that never says no is a backtest suite whose yes means nothing.
Read this page as Markdown:
/docs/statistics.md