Research campaign · Iteration 3 · unassessed
3x S&P ensemble frontier — walk-forward confirmation
Does the issue-1310 causal UPRO/cash feasible frontier survive an expanding annual January walk-forward that freezes parameters from training prefixes only?
- Expressions
- 3
- Logged trials
- Not recorded
- Independent events
- Not assessed
- Evidence
- unassessed
What the research found
Primary local neighborhood stitched OOS CAGR 12.23% with max drawdown -25.45%; fails the preregistered stitched-path -25% floor.
All 17 annual OOS segments of the primary family individually clear -25%; the breach appears only after stitching.
Safety-margin family (training DD >= -22%) stitches to 11.01% CAGR / -21.53% DD and clears -25%.
Focused-grid sensitivity stitches to 12.41% CAGR / -24.26% DD and clears -25%.
No full-sample parameter choice was used; issue-1313 non-causal oracles were excluded.
Results are exploratory reconstructed OOS because prior campaign iterations already inspected 1999-2026; not a promotion candidate.
Mechanism and falsifiers
Not recorded in this iteration.
Not recorded in this iteration.
Confidence and limitations
Primary local stitched OOS path fails the -25% floor by ~0.45pp even though every annual segment clears.
Reconstructed expanding-window evidence is not a researcher-blind holdout because issues 1310/1313 already inspected this history.
Pre-2009 3x returns remain modeled; breadth and macro remain proxies/lagged finals.
Focused-grid and safety-margin clears do not authorize catalog promotion without a separate reviewed fork.
Not recorded in this iteration.
Compare expressions
Download evidenceExploratory results. Check each period, proxy and cost assumption before comparing. — means not recorded.
| Expression / family | CAGR | Sharpe | Max drawdown | Test period | Assessment |
|---|---|---|---|---|---|
| l3x.wf.primary_local_stitchedh1-local-c5-expanding-selection | 12.2% | — | -25.5% | Not recorded |
exploratory causal |
| l3x.c5_vol_scaled_ensemble.37764b997305h3-safety-margin-local-selection | 11.0% | — | -21.5% | Not recorded |
exploratory causal |
| l3x.wf.focused_grid_stitchedh2-focused-grid-expanding-selection | 12.4% | — | -24.3% | Not recorded |
exploratory causal |
Open an expression to inspect its rules and request confirmation. The request must be submitted by a trusted repository collaborator.
l3x.wf.primary_local_stitched
- variant id
l3x.wf.primary_local_stitched
- family id
h1-local-c5-expanding-selection
- parameters
- protocol id
expanding-annual-jan-local-c5/v1
- unique selected
l3x.c5_vol_scaled_ensemble.547a9b4d36f5
l3x.c5_vol_scaled_ensemble.e5003d8780f5
- created by
walk_forward_confirmation
- reason tested
Primary preregistered local expanding-window path
- status
exploratory_causal
- tags
interesting_negative_result
parameter_sensitive
needs_more_research
l3x.c5_vol_scaled_ensemble.37764b997305
- variant id
l3x.c5_vol_scaled_ensemble.37764b997305
- family id
h3-safety-margin-local-selection
- parameters
- dd floor
-0.22
- created by
walk_forward_confirmation
- reason tested
Training drawdown buffer sensitivity
- status
exploratory_causal
- tags
candidate_for_confirmation
needs_more_research
l3x.wf.focused_grid_stitched
- variant id
l3x.wf.focused_grid_stitched
- family id
h2-focused-grid-expanding-selection
- parameters
- candidate count
161
- created by
walk_forward_confirmation
- reason tested
Broader focused-grid multiplicity sensitivity
- status
exploratory_causal
- tags
candidate_for_confirmation
needs_more_research
Interactive lab
Interactive history unavailableLoading available evidence…
Research record
Source claims, inferred rules, experiments and the evidence behind the assessment.
Source
- kind
user_hypothesis
- issue
1312
- parent issue
1310
- related issue
1313
- primary window
1999-03-05 to 2026-07-24
- objective
preregister expanding-window selection and test whether the feasible CAGR/drawdown frontier survives without full-history parameter choice
What the source claims
Full-window CAGR must be north of 30%, max drawdown must be no worse than -25%, and the only risky holding is levered 3x S&P (UPRO/cash).
Analyze the parent frontier's large drawdowns and identify indicators that would have avoided those episodes.
Aggressive in-sample and hindsight overfitting is explicitly in scope, but the instruction does not turn future-return leakage into a deployable rule.
Evaluate causal full-window fits separately from memorized-event and forward-return oracle ceilings.
Use trend, shock, volatility, drawdown, breadth, macro, vote, veto, and calendar/oracle exposure maps to span conceptually diverse defenses.
Realized UPRO history before 2009, exact point-in-time constituent breadth, and prospective performance of a hindsight-fit rule are unavailable.
Rules actually disclosed
Full-window CAGR must be north of 30%, max drawdown must be no worse than -25%, and the only risky holding is levered 3x S&P (UPRO/cash).
Analyze the parent frontier's large drawdowns and identify indicators that would have avoided those episodes.
Aggressive in-sample and hindsight overfitting is explicitly in scope, but the instruction does not turn future-return leakage into a deployable rule.
What had to be inferred
Evaluate causal full-window fits separately from memorized-event and forward-return oracle ceilings.
Use trend, shock, volatility, drawdown, breadth, macro, vote, veto, and calendar/oracle exposure maps to span conceptually diverse defenses.
Research questions
Does the issue-1310 causal UPRO/cash feasible frontier survive an expanding annual January walk-forward that freezes parameters from training prefixes only?
Data
- inputs
SPY
UPRO
TB3MS
UNRATE
__SYNTH_SPXBREADTH50__
- market sqlite sha256
1537834d27b92bbf1872d98708186870e59d24f5de7a95728a3e08c959d7cff3
- synthetic method
backtest.sim_history.synthetic_levered_daily/testfol_daily-v1
- synthetic parameters
- leverage
3.0
- swap exposure
1.1
- financing spread percent
0.4
- expense ratio percent
1.0
- annualizer
252
- causal execution
close-derived decision applies one trading bar later
- hindsight execution
memorized dates or same-session/forward return; explicitly non-causal and non-deployable
- public data policy
publish derived metrics and normalized series, not bulk provider rows or subscriber workbook data
- trial count
1873
- causal trial count
1816
- hindsight trial count
57
- acceptance hit count
16
- causal acceptance hit count
0
- winner variant id
l3x.hindsight.036f423a11de
- walk forward issue
1312
- walk forward protocol
- cost bps
5.0
- execution lag bars
1
- first cutoff
2010-01-01
- fold cadence
annual_january
- primary verdict
stitched OOS path max_drawdown >= -0.25
- protocol id
expanding-annual-jan-local-c5/v1
- schedule
expanding_window
- training objective
max CAGR subject to training max_drawdown >= floor
- full sample parameter choice
False
Baseline implementation
- variant id
l3x.wf.primary_local_stitched
- family id
h1-local-c5-expanding-selection
- parameters
- protocol id
expanding-annual-jan-local-c5/v1
- unique selected
l3x.c5_vol_scaled_ensemble.547a9b4d36f5
l3x.c5_vol_scaled_ensemble.e5003d8780f5
- created by
walk_forward_confirmation
- reason tested
Primary preregistered local expanding-window path
- status
exploratory_causal
- metrics
- cagr
0.12231073495697942
- max drawdown
-0.2545491086115719
- acceptance hit
False
- tags
interesting_negative_result
parameter_sensitive
needs_more_research
What to try interactively
- name
trend_sma_days
- type
integer
- default
150
- min
5
- max
252
- name
momentum_days
- type
integer
- default
189
- min
5
- max
252
- name
fast_vol_threshold
- type
float
- default
0.25
- min
0.1
- max
0.3
- name
breadth_threshold
- type
float
- default
0.5
- min
0.3
- max
0.7
- name
oracle_return_cut
- type
float
- default
0.0
- min
-0.01
- max
0.005
- name
execution_lag_bars
- type
integer
- default
1
- min
0
- max
2
Suggested next research
Should confirmation privilege the stitched path or per-fold clearance?
Would a longer first training window or biennial cadence remove the 0.45pp primary breach without retuning?
Should the cleared safety-margin rule be forked into a separate indicator campaign after owner review?
Trial ledger
- iteration number
1
- issue
1310
- objective
maximize causal exploratory CAGR under the -25% drawdown and UPRO-only constraints
- iteration number
2
- issue
1313
- objective
find >30% full-window hindsight frontier and attribute the parent baseline's large drawdowns
- iteration number
3
- issue
1312
- objective
walk-forward confirmation without full-sample parameter choice
What the signal looks like
Expanding annual January selection freezes a causal UPRO/cash rule from each training prefix and applies it one bar later on the next OOS year.
Primary menu is the 33-variant local neighborhood around the issue-1310 c5 winner; focused grid and -22% training buffer are sensitivities.
State-space exploration
17 January cutoffs from 2010-01-01; local 33 + focused 161 + safety-margin local candidates.
Selections never see OOS rows; issue-1313 oracles excluded.
What worked
Safety-margin and focused-grid stitched paths clear -25% (-21.53%, -24.26%).
Primary annual OOS segments all clear individually; OOS CAGR exceeds the full-history parent winner.
What did not work
Primary local stitched path max drawdown -25.45% fails the preregistered floor.
Full-history winner parameters are not stably reselected on every prefix.
Agent assessment
Mixed confirmation: the feasible frontier does not unambiguously survive the primary local protocol, but buffered/broader menus still clear -25% out of sample. Keep exploratory; do not promote.
Next questions
Should confirmation privilege the stitched path or per-fold clearance?
Would a longer first training window or biennial cadence remove the 0.45pp primary breach without retuning?
Should the cleared safety-margin rule be forked into a separate indicator campaign after owner review?
Review status
published. Research publication does not imply official admission.
Return to pending research