Quantmodel · Strategy research

LETF Baseline: UPRO / ZROZ / GLD

LETF Baseline: UPRO / ZROZ / GLD receives a research grade of B (composite 82 / 100) on Paper-to-Profit robustness tests run against published monthly ours versus SPY. The statistical edge and stress tests are strong enough to keep as a candidate; this is still a historical scorecard, not a forecast.

B score 82 BF 293.0 months 2002-03 → 2026-07

Factor grades

Statistical edge · 20% B 82
Bootstrap vs SPY · 18% A 86
Alpha decay · 16% A 87
Overfit / stability · 16% A 94
Cost robustness · 15% A 96
Regime / stress · 15% F 47

Published CAGR 15.18% / Sharpe 0.73 / max DD -66.6% over 293 months. Those headline stats are context only; the grade is the robustness tests above.

Assessment

Statistical edge

Monthly mean is 1.42%. t-test vs zero p=0.000; vs SPY p=0.003 (excess 0.52% per month). The path is statistically distinct from SPY on this window.

Bootstrap versus SPY

Block bootstrap of consecutive months (400 paths): P(beat SPY) over 6.0 months is 72%; over 24.0 months is 86%. This is the Paper-to-Profit 'realistic path' test, not a reshuffled iid sample.

Alpha decay

60% of 12-month OLS windows have a significantly positive slope and 11% are significantly negative. Terminal strategy/SPY wealth ratio is 3.04. A steadily rising log-relative-wealth line is the decay test from the tutorial; a late collapse is a warning.

Overfit / stability

70/30 OOS/IS Sharpe ratio is 1.24; second/first-half Sharpe ratio is 1.49. Walk-forward parameter search is not estimable on a frozen catalog snapshot; this split is the overfit proxy. Return-space noise did not halve Sharpe on the tested grid.

Transaction costs

Average daily turnover is 0.008%. Additional one-way costs through 24 bps did not cut Sharpe in half. That is the tutorial's cost-robust case.

Regimes and crises

The path beats SPY in 81% of classified months (3.0/6.0 regimes). In bear/strong-bear/vol-down months it beats SPY 0% of the time. Crisis windows: Dot-com bust: strategy -34.8% vs SPY -25.6%; Global Financial Crisis: strategy -61.1% vs SPY -46.0%; COVID crash: strategy -26.4% vs SPY -19.4%; 2022 inflation shock: strategy -44.2% vs SPY -23.9%.

Limitations

Tests use published monthly ours, already lagged one bar in the engine, versus SPY. t-tests assume roughly comparable monthly means; Wilcoxon is shown as a non-normal check. Bootstrap paths are overlapping consecutive months (not iid reshuffles). Noise is added to monthly returns, not to OHLC, because this producer does not rerun backtests. Walk-forward parameter optimization, SPA, and PBO are not estimable without a trial ledger. Fee drag is extra one-way bps on catalog average daily turnover, on top of the published costed path.

Edge tests

Test Statistic p-value Significant?
t-test vs zero 3.59 0.0004 yes
t-test vs SPY (paired) 3.03 0.0027 yes
Wilcoxon vs zero 27777.0 0.0000 p<0.05

Part 1: a strategy that only beats zero but not SPY has no statistically significant edge versus the market.

Monthly return distributions

Strategy versus SPY on the shared window. Input to the t-tests above.

Block bootstrap — 6 months

400 overlapping consecutive-month paths. P(beat SPY) = 72% . Fan is p10 / p50 / p90 of terminal wealth from $1.

Block bootstrap — 24 months

Same construction at a two-year horizon. P(beat SPY) = 86% .

Bootstrap excess CDF

Sorted 24-month strategy minus SPY terminal wealth. Fraction above zero is P(beat SPY).

Relative wealth versus SPY

Strategy equity divided by SPY equity. A rising line is persistent outperformance; a late roll-over is decay.

Rolling OLS slope of log relative wealth

12-month window on ln(strategy/SPY). Positive significant slopes are the tutorial's "edge is still there" test.

70/30 chronological split

Train through 2019-04. OOS/IS Sharpe ratio 1.24 . Walk-forward parameter search is not estimable here; this is the overfit proxy.

Transaction-cost sensitivity

Extra one-way bps applied to average daily turnover, on top of the published (~10 bps) path. Half-Sharpe at beyond 24 bps.

Return-space noise

Gaussian noise scaled by 12-month rolling vol. Not an OHLC re-backtest. Critical multiplier (Sharpe halves): not reached.

SPY regimes

Seven trend/vol states on a 12-month SPY window. Mean monthly return of strategy versus SPY in each state.

RegimeMonthsStrategy meanSPY meanBeats SPY?
Strong bull 114.0 2.14% 1.39% yes
Bull 74.0 1.50% 0.97% yes
Sideways 10.0 -3.74% -2.85% no
Bear 11.0 -2.48% -1.84% no
Strong bear 2.0 — — —
Volatility up 48.0 3.44% 2.24% yes
Volatility down 34.0 -0.48% -0.32% no

Named crisis windows

EpisodeStatusStrategySPY
Dot-com bust ok -34.8% -25.6%
Global Financial Crisis ok -61.1% -46.0%
COVID crash ok -26.4% -19.4%
2022 inflation shock ok -44.2% -23.9%

Growth of $1

Published path versus SPY, 60/40 when present, and a leave-one-out catalog mix. Context only — not a scoring factor.

Drawdown

Related tools