Which statistics actually matter when testing the Santa Rally?
A single average return is not enough. A useful seasonality study should show frequency, median outcome, outliers, sample size and the individual historical years behind the headline.
Core metrics for a seasonal backtest
| Metric | What it tells you | Common mistake |
|---|---|---|
| Average return | Mean result across all occurrences. | Letting a few extreme years dominate the conclusion. |
| Median return | The midpoint historical outcome. | Ignoring the difference between mean and median. |
| Win rate | Share of positive occurrences. | Assuming a high win rate means low risk. |
| Sample size | How many historical periods are included. | Drawing conclusions from too few years. |
| Worst year | Shows downside that an average can hide. | Reporting only the best historical cases. |
Compare multiple lookbacks
Run the same dates over 10, 20 and 25 years. If the result disappears when the sample changes, the pattern may be unstable.
Keep the dates fixed
Changing the start or end date after viewing the results can create an apparently impressive backtest that is largely the product of selection bias.
A practical research workflow
Define
Choose a precise recurring calendar window before looking at the results.
Measure
Record return, win rate, median and every annual observation.
Validate
Compare alternative lookbacks and related markets rather than trusting one sample.
What is a good Santa Rally win rate?
There is no universal threshold. A win rate should be evaluated together with sample size, average return, median return and downside years.
Should I use calendar days or trading days?
For a recurring market pattern, be explicit about the convention and apply it consistently across the historical sample.
See the historical years, not just the headline
Use the free dashboard to inspect the selected period year by year.