Common metrics in four categories, each with "how it's calculated / how to use it / where its blind spots are"
Sharpe ratio
[How it's calculated] Excess return ÷ total volatility. [How to use it] Measures return per unit of volatility — the most widely used risk-adjusted metric. [Blind spots] It penalises upside volatility too, so trend strategies look worse than they are, and it is highly sensitive to the sample period.
Sortino ratio
[How it's calculated] Excess return ÷ downside volatility. [How to use it] Penalises only losing volatility, which matches trader intuition better. [Blind spots] Downside deviation is unstable with small samples, and definitions vary across platforms — compare only within the same methodology.
Calmar ratio
[How it's calculated] Annualised return ÷ max drawdown. [How to use it] Directly answers "how many times the drawdown did we earn". [Blind spots] It only looks at the single worst drawdown, so one extreme event can dominate; strategies with small but long drawdowns get overstated.
Max drawdown (MDD)
[How it's calculated] The largest peak-to-trough decline in equity. [How to use it] The most intuitive evidence of "couldn't take it" — the core metric of the survival layer. [Blind spots] It has no time dimension: losing 30% over three months versus three years are entirely different experiences, so always read it alongside drawdown recovery time.
Drawdown recovery time
[How it's calculated] Time needed to return from the trough to the previous peak. [How to use it] Harder than drawdown depth — only recovering counts, and professional institutions require it. [Blind spots] In extreme cases recovery may never happen; pair it with "did it make a new high".
Profit factor (PF)
[How it's calculated] Gross profit ÷ gross loss. [How to use it] A single number containing both win rate and R:R; above 1.5 is good, above 2 is excellent. [Blind spots] Sensitive to one very large trade, and easily inflated by a single outlier when the sample is small.
Expectancy (ER / R-multiple)
[How it's calculated] Win rate × average win − loss rate × average loss, expressed in R-multiples. [How to use it] Once expressed in R, traders with different capital sizes become comparable. [Blind spots] Depends on sample size, and R calculated with different stop-loss conventions is not directly comparable.
Reward-to-risk (R:R)
[How it's calculated] Average win ÷ average loss. [How to use it] Must always be read together with win rate; either one alone misleads. [Blind spots] High R:R often comes with a low win rate and long losing streaks, which is a real test of discipline.
Trimmed test
[How it's calculated] Remove the most profitable trades (e.g. top 5%) and check whether the result is still profitable. [How to use it] Distinguishes "the system makes money" from "I was right once" — the hardest robustness check. [Blind spots] The trim ratio must use a fixed convention, otherwise it can be presented selectively.
Cost erosion rate
[How it's calculated] (Fees + funding + slippage) ÷ gross profit. [How to use it] Essential for high-frequency and derivatives trading — funding and slippage can consume the entire profit. [Blind spots] Live slippage is hard to reconstruct precisely, and backtests often assume costs optimistically.
Loss-chasing rate
[How it's calculated] The proportion of re-entries shortly after a loss where position size significantly exceeds the trader's average. [How to use it] Captures the substance of "revenge doubling down" without penalising normal re-entry. [Blind spots] Threshold choice affects results and must be calibrated to each trader's cycle.
Ulcer Index
[How it's calculated] A root-mean-square measure combining drawdown depth and duration. [How to use it] Merges "how deep" and "how long" into one number, closer to real pain than MDD. [Blind spots] More complex to compute and harder to explain intuitively, so usually a supporting rather than primary metric.