Objective
Reproduce and explain why leveraged component/meta-strategy variants have lower CAGR than unlevered versions; verify whether the result is mathematical, selector/optimization behavior, leverage implementation, data/backtest artifact, or bug.
Approach
Reproduced published payloads (docs/site-data/meta-strategies/-lev-.json, catalog-leverage-meta.json, catalog-leverage-weighted-meta.json, weighted-meta/.json). Independently recomputed CAGR via geometric prod^(12/n)-1, vol via sample stdsqrt12, Sharpe/Sortino/maxDD/worst/total; verified lev catalog_metrics diff 0.0 (no display bug). Traced order: individual strategy weights -> evaluate_month(L) -> substitute ETF weights -> instrument monthly Close returns (live/sim) -> weighted sum (already levered); meta: optimizer weights (unlevered expanding history) -> Jan signal filter feasible <=L -> renorm to 1 -> hold forward -> each month sum w_held r_path(sleeve,Lf). No post-aggregation L multiply. Quantified allocation differences via constituent_weights, selection-change counts, and breadth policy. Ran five families: F1 base, F2 lev cohort, F3 frozen selections mechanically scaled (Lportfolio), F4 frozen allocs substituted lev sleeve returns, F5 constant 1x/1.5x/2x.
Files changed
- research/experiments/qm-wtfm.json — experiment record for investigation
- docs/site-data/experiments/qm-wtfm.json, docs/experiments/qm-wtfm.html — published experiment
- docs/meta-strategies.html — strengthened lede/methodology disclosure (feasibility-filtered window, post-filter renorm not re-capped, substitute-ETF path returns not Lx, apples-to-apples window note)
- research/results/qm-wtfm.md — this note
Validation
- python3 /tmp/run_wtfm.py — base 680m 11.69%, base-on-lev 385m 10.64% vs lev2 10.32% (-0.32pp same window; -1.36pp full window), mechanical 2x 21.78%; F2 monthly log gap vs mechanical -3.168 (top10 months 33% of gap: 2003-02, 2005-05, 2001-10 etc.)
- python3 /tmp/run_weighted.py — weighted base 679m 9.97% vs lev2 54m 9.92% (disjoint windows)
- No implementation bug found (single L scaling at sleeve, CASH skip, renorm guard, window inclusive checks all correct); misleading-but-correct behavior documented.
Results
Causal classification: combination — (1) window artifact ~75% of full-window gap from earliest levered-ETF start (1991-01 / 2012-07) truncation, (2) intentional selection change from inclusive <=L filter+renorm (38->10 sleeves, 0.2pp drag same window, 31% months active set differs), (3) volatility drag/daily-reset path dependence vs naive Lx (-9.4% CAGR drag vs mechanical 2x), (4) missing-month skipping when levered series incomplete. No look-ahead, annualization bug, or hidden financing cost. Leveraged meta display is correct given disclosed cohort construction but was misleading when compared naively to full-window unlevered CAGR; clarified on Meta Strategies page. No code fix required; regression guard is geometric CAGR equality and F3>F2 assertion.
Links
- Task:
qm-wtfm - Detailed analysis:
research/experiments/qm-wtfm.json(5-family controlled comparisons F1–F5, leverage order trace, allocation/period attribution) - Code changes:
research/levered_path_returns.py(levered path construction),research/reports/catalog_leverage.py/catalog_leverage_meta.py(feasibility filtering + renorm),research/reports/meta_export.py(meta payloads) - Experiments / data:
docs/site-data/experiments/qm-wtfm.json+docs/experiments/qm-wtfm.html,docs/site-data/meta-strategies/<em>-lev-</em>.json,docs/site-data/catalog-leverage<em>.json,docs/site-data/weighted-meta/</em>.json, validationpython3 /tmp/run_wtfm.pyandpython3 /tmp/run_weighted.py - Affected website pages:
docs/meta-strategies.html(methodology disclosure strengthened),docs/meta-strategies/meta-max-sharpe-lev-2.htmletc.,docs/experiments/qm-wtfm.html,docs/results/qm-wtfm.html