Skip to content
Learning path Silver Repeatable method

Backtest metrics and benchmarks

How to read return, risk, drawdown, costs, capacity and uncertainty by stating the series, frequency, annualisation and term of comparison.

Who this is for — Anyone who needs to interpret a strategy report without relying on a single number and wants to compare results using coherent units, costs and risks.

Backtest metrics summarise different aspects of a simulation: return, variability, tail loss, path, trade quality, implementation and uncertainty. A benchmark defines the term of comparison. No single metric establishes that a strategy has an edge or is suitable for an investor.

The contract comes before the number: return series, frequency, currency, capital, leverage, cash treatment, reinvestment, costs, missing data, period and benchmark. Changing one of these conventions may change the result without changing a single trade.

Backtest metrics and benchmarksReturn, path, implementation and uncertainty remain separate. Illustrative values omitted: every measure needs stated units, frequency, capital and costs.Backtest metrics and benchmarksReturn, path, implementation and uncertainty remain separateIllustrative values omitted: every measure needs stated units, frequency, capital and costs.ReturnTotal, annualised orexcess return needsconsistent flows andcompounding.ExpectancyMean per trade orrisk unit depends onobservationdefinition.Volatility anddownsideEstimator, frequencyand distributiondetermine meaning.Sharpe /SortinoNumerator, benchmark,annualisation andautocorrelation arestated.Drawdown andrecoveryPath-dependentmeasures tied to theobserved window, notfuture limits.Turnover andcostsConnect theoreticalsignal toeconomicallyimplementable result.Exposure andcapacityLeverage,concentration andscale explain partsof return and risk.Benchmark anduncertaintyRelevant comparison,intervals and trialcount accompany theestimate.Cyclepedia · conditional teaching diagram, not a forecast or promise
A factsheet combines families of metrics and preserves the conventions needed to interpret them.

Families of measures

Family Examples Question Limitation
Return cumulative return, CAGR, mean return, excess return how much did capital grow under the stated convention? does not describe path or risk
Variability volatility, downside deviation how much do returns fluctuate? penalises or ignores movements according to the definition
Risk-adjusted Sharpe, Sortino, information ratio how much return is associated with the chosen risk measure? sensitive to dependence, distribution and benchmark
Path maximum drawdown, duration, recovery time, Calmar what peak-to-trough loss appears in the sample? one path does not identify the future distribution
Tail quantiles, Expected Shortfall, worst period what happens in the worst observations or scenarios? few events and model risk make estimation uncertain
Trade-based expectancy, profit factor, win rate, payoff ratio how are results distributed per defined trade? overlapping trades and the definition of the unit
Implementation turnover, spread, slippage, impact, capacity how much of gross performance can be captured? costs depend on size and market state
Uncertainty standard error, intervals, bootstrap, fold results how precise and stable is the estimate? depends on the method's assumptions

A metric may have several definitions. “Annual return” can be an annualised arithmetic mean or a compound rate; “Sortino” can use different targets and denominators; “profit factor” depends on how trades, costs and positions are aggregated. Formula and convention must be visible.


Sharpe: formula and annualisation

For excess returns xₜ = rₜ − rᶠₜ, a simple sample estimate is:

Sharpe = mean(xₜ) / standard_deviation(xₜ)

The frequency of the series is part of the measure. Multiplying a daily Sharpe ratio by the square root of the number of periods is exact only under restrictive conditions concerning dependence and temporal scale. Lo shows that with autocorrelation, smoothed returns or overlapping positions, the naive conversion can be materially misleading.

The risk-free rate or hurdle, its source, currency, day-count convention and handling of non-trading days must be stated. A gross Sharpe ratio is not comparable with a net one; an unlevered ratio is not equivalent to one with a different exposure profile.


Drawdown and trade-based metrics

Maximum drawdown is the largest decline from a previous peak in the equity curve within the sample. It depends on the path and the start and end dates. Two strategies with the same marginal distribution can have different drawdowns because returns occur in a different order. Duration and recovery time add information that magnitude alone does not show.

Win rate, average payoff, expectancy and profit factor describe trades under a chosen segmentation. If a position is scaled through several orders, the number of “trades” changes with the grouping rule. Many trades exposed to the same shock are not independent observations. Presenting a profit factor above a threshold as standalone proof of edge ignores sample, dependence, selection and tails.


Four types of benchmark

Benchmark Comparison Caveat
Cash / risk-free remunerates capital that is not put at risk currency, duration and rate must be coherent
Market index or portfolio represents an investable opportunity or mandate price return and total return are not equivalent
Strategy baseline measures incremental value against a simple rule must use the same data, costs, universe and leverage
Execution benchmark compares the price obtained with arrival price, VWAP or another reference evaluates execution, not signal alpha

The benchmark should match the question and preferably be selected before the result. A market-neutral strategy is not well explained by a long-only index alone; a timing strategy can be compared both with cash and passive exposure, while also showing beta and capital employed. Index and portfolio comparisons must align dividends, rebalancing, currency and costs.

IOSCO's Principles for Financial Benchmarks address financial benchmarks and their governance. They help explain integrity, methodology and conflicts, but they do not turn every internal baseline into a regulated benchmark.


Gross, net and capacity

A useful report builds a reconcilable bridge:

gross result − commissions − spread − slippage − impact − borrow/funding/roll = simulated net result

These entries are not universal constants. Impact and fill probability depend on quantity, urgency, venue, depth and regime. Showing several levels of size and motivated assumptions helps distinguish theoretical return from economically attainable capacity.

Turnover without units is ambiguous. The report should state whether it uses purchases, sales, half of the absolute sum, weights or notional, and over which interval. The same applies to margin utilisation, gross and net exposure, and idle capital.


Selection and uncertainty

The Sharpe ratio of the best result among many attempts has a different distribution from that of a pre-specified strategy. Within its model, the Deflated Sharpe Ratio considers non-normality and selection among trials. The Reality Check and other methods address multiple comparisons. These tools require a credible ledger of candidates and do not correct leakage or execution errors.

Intervals and bootstrap procedures must respect dependence. The number of rows or trades does not automatically equal the number of independent observations. In addition to a point estimate, it is useful to show results by period, instrument, regime and fold, as well as sensitivity to costs and specifications.


Illustrative example

Two reports state a Sharpe ratio of 1.2. The first uses daily net returns, a rate coherent with the currency and a correction for autocorrelation; the second uses smoothed monthly gross P&L, annualised under an unstated convention. The two numbers are not comparable.

A proper factsheet reports formula, series, frequency, sample, costs, benchmark, exposure and uncertainty. The value 1.2 is purely illustrative and does not constitute a quality threshold.


Minimum reproducible factsheet

  1. Period, point-in-time universe, time zone and frequency.
  2. Definition of returns, currency, capital, leverage and reinvestment.
  3. Gross results, cost bridge and net results.
  4. Selected benchmark and baselines, including dividends and costs.
  5. Return, variability, drawdown/tail, trade and implementation measures.
  6. Turnover, capacity and concentration of profits.
  7. Dependence, number of attempts and uncertainty method.
  8. Results by segment and code/data versions.

The purpose is not to accumulate indicators. It is to prevent an elegant number from concealing the convention that produced it.


Sources