Who this is for — Anyone who needs to interpret a strategy report without relying on a single number and wants to compare results using coherent units, costs and risks.
Backtest metrics summarise different aspects of a simulation: return, variability, tail loss, path, trade quality, implementation and uncertainty. A benchmark defines the term of comparison. No single metric establishes that a strategy has an edge or is suitable for an investor.
The contract comes before the number: return series, frequency, currency, capital, leverage, cash treatment, reinvestment, costs, missing data, period and benchmark. Changing one of these conventions may change the result without changing a single trade.
Families of measures
| Family | Examples | Question | Limitation |
|---|---|---|---|
| Return | cumulative return, CAGR, mean return, excess return | how much did capital grow under the stated convention? | does not describe path or risk |
| Variability | volatility, downside deviation | how much do returns fluctuate? | penalises or ignores movements according to the definition |
| Risk-adjusted | Sharpe, Sortino, information ratio | how much return is associated with the chosen risk measure? | sensitive to dependence, distribution and benchmark |
| Path | maximum drawdown, duration, recovery time, Calmar | what peak-to-trough loss appears in the sample? | one path does not identify the future distribution |
| Tail | quantiles, Expected Shortfall, worst period | what happens in the worst observations or scenarios? | few events and model risk make estimation uncertain |
| Trade-based | expectancy, profit factor, win rate, payoff ratio | how are results distributed per defined trade? | overlapping trades and the definition of the unit |
| Implementation | turnover, spread, slippage, impact, capacity | how much of gross performance can be captured? | costs depend on size and market state |
| Uncertainty | standard error, intervals, bootstrap, fold results | how precise and stable is the estimate? | depends on the method's assumptions |
A metric may have several definitions. “Annual return” can be an annualised arithmetic mean or a compound rate; “Sortino” can use different targets and denominators; “profit factor” depends on how trades, costs and positions are aggregated. Formula and convention must be visible.
Sharpe: formula and annualisation
For excess returns xₜ = rₜ − rᶠₜ, a simple sample estimate is:
Sharpe = mean(xₜ) / standard_deviation(xₜ)The frequency of the series is part of the measure. Multiplying a daily Sharpe ratio by the square root of the number of periods is exact only under restrictive conditions concerning dependence and temporal scale. Lo shows that with autocorrelation, smoothed returns or overlapping positions, the naive conversion can be materially misleading.
The risk-free rate or hurdle, its source, currency, day-count convention and handling of non-trading days must be stated. A gross Sharpe ratio is not comparable with a net one; an unlevered ratio is not equivalent to one with a different exposure profile.
Drawdown and trade-based metrics
Maximum drawdown is the largest decline from a previous peak in the equity curve within the sample. It depends on the path and the start and end dates. Two strategies with the same marginal distribution can have different drawdowns because returns occur in a different order. Duration and recovery time add information that magnitude alone does not show.
Win rate, average payoff, expectancy and profit factor describe trades under a chosen segmentation. If a position is scaled through several orders, the number of “trades” changes with the grouping rule. Many trades exposed to the same shock are not independent observations. Presenting a profit factor above a threshold as standalone proof of edge ignores sample, dependence, selection and tails.
Four types of benchmark
| Benchmark | Comparison | Caveat |
|---|---|---|
| Cash / risk-free | remunerates capital that is not put at risk | currency, duration and rate must be coherent |
| Market index or portfolio | represents an investable opportunity or mandate | price return and total return are not equivalent |
| Strategy baseline | measures incremental value against a simple rule | must use the same data, costs, universe and leverage |
| Execution benchmark | compares the price obtained with arrival price, VWAP or another reference | evaluates execution, not signal alpha |
The benchmark should match the question and preferably be selected before the result. A market-neutral strategy is not well explained by a long-only index alone; a timing strategy can be compared both with cash and passive exposure, while also showing beta and capital employed. Index and portfolio comparisons must align dividends, rebalancing, currency and costs.
IOSCO's Principles for Financial Benchmarks address financial benchmarks and their governance. They help explain integrity, methodology and conflicts, but they do not turn every internal baseline into a regulated benchmark.
Gross, net and capacity
A useful report builds a reconcilable bridge:
gross result − commissions − spread − slippage − impact − borrow/funding/roll = simulated net resultThese entries are not universal constants. Impact and fill probability depend on quantity, urgency, venue, depth and regime. Showing several levels of size and motivated assumptions helps distinguish theoretical return from economically attainable capacity.
Turnover without units is ambiguous. The report should state whether it uses purchases, sales, half of the absolute sum, weights or notional, and over which interval. The same applies to margin utilisation, gross and net exposure, and idle capital.
Selection and uncertainty
The Sharpe ratio of the best result among many attempts has a different distribution from that of a pre-specified strategy. Within its model, the Deflated Sharpe Ratio considers non-normality and selection among trials. The Reality Check and other methods address multiple comparisons. These tools require a credible ledger of candidates and do not correct leakage or execution errors.
Intervals and bootstrap procedures must respect dependence. The number of rows or trades does not automatically equal the number of independent observations. In addition to a point estimate, it is useful to show results by period, instrument, regime and fold, as well as sensitivity to costs and specifications.
Illustrative example
Two reports state a Sharpe ratio of 1.2. The first uses daily net returns, a rate coherent with the currency and a correction for autocorrelation; the second uses smoothed monthly gross P&L, annualised under an unstated convention. The two numbers are not comparable.
A proper factsheet reports formula, series, frequency, sample, costs, benchmark, exposure and uncertainty. The value 1.2 is purely illustrative and does not constitute a quality threshold.
Minimum reproducible factsheet
- Period, point-in-time universe, time zone and frequency.
- Definition of returns, currency, capital, leverage and reinvestment.
- Gross results, cost bridge and net results.
- Selected benchmark and baselines, including dividends and costs.
- Return, variability, drawdown/tail, trade and implementation measures.
- Turnover, capacity and concentration of profits.
- Dependence, number of attempts and uncertainty method.
- Results by segment and code/data versions.
The purpose is not to accumulate indicators. It is to prevent an elegant number from concealing the convention that produced it.
Sources
- William F. Sharpe, The Sharpe Ratio — definition, interpretation and relationships of the measure.
- Andrew W. Lo, The Statistics of Sharpe Ratios — distribution and limitations of annualisation in the presence of dependence.
- David H. Bailey and Marcos López de Prado, The Deflated Sharpe Ratio — adjustment for selection bias and non-normal returns within the proposed model.
- International Organization of Securities Commissions, Principles for Financial Benchmarks — principles for benchmark governance, quality and methodology within their scope.
- Joseph Simonian, CFA Institute Research Foundation, Investment Model Validation: A Guide for Practitioners — joint use of performance, benchmarks, stress and uncertainty in validation.