Prometheus Basic Trend: stats match, holdout edge does not

Casual takeaway: the long-run published numbers are reproducible; the frozen economic holdout is not strong enough for a public strategy catalog entry.

This notebook reads only frozen repository artifacts. It does not retune lookbacks, costs, or windows.

In [1]:
from pathlib import Path
import json
import pandas as pd
import matplotlib.pyplot as plt
from IPython.display import HTML, display

ROOT = Path.cwd()
evaluation = json.loads((ROOT/'research/findings/prometheus-basic-trend-evaluation.json').read_text())
equity = json.loads((ROOT/'research/findings/prometheus-basic-trend-equity.json').read_text())
evidence = json.loads((ROOT/'research/strategy_evidence/prometheus-basic-trend.json').read_text())
pd.Series({
  'G13 decision': evidence['gates']['G13']['decision'],
  'publisher fidelity all': evaluation['gates']['fidelity'] and all(evaluation['gates']['fidelity'].values()),
  'all promotion checks': evaluation['gates']['all_promotion_checks'],
  'untouched Sharpe improvement': evaluation['primary_windows']['untouched_confirmatory']['comparison']['sharpe_improvement'],
  'untouched MDD improvement': evaluation['primary_windows']['untouched_confirmatory']['comparison']['maximum_drawdown_improvement'],
}, name='frozen gates').to_frame()
Out[1]:
frozen gates
G13 decision fidelity_only
publisher fidelity all True
all promotion checks False
untouched Sharpe improvement -0.638614
untouched MDD improvement 0.033538

1. Documented vs inferred

Prometheus discloses the architecture. Exact instruments, estimators, lag, costs, and the 3x cap are QuantModel inferences frozen before returns.

In [2]:
pd.DataFrame([
 {'piece':'Universe','documented':'stocks / bonds / gold / bitcoin','inferred':'SPY, TLT←VUSTX, GLD←GC=F, Coin Metrics BTC'},
 {'piece':'Trend','documented':'1m & 6m full/half/flat','inferred':'21 / 126 NYSE sessions'},
 {'piece':'Risk','documented':'inverse-vol + breadth-scaled 15% target','inferred':'126-session vol, 21-session covariance'},
 {'piece':'Execution','documented':'daily histories shown','inferred':'next_close + 10 bps one-way'},
])
Out[2]:
piece documented inferred
0 Universe stocks / bonds / gold / bitcoin SPY, TLT←VUSTX, GLD←GC=F, Coin Metrics BTC
1 Trend 1m & 6m full/half/flat 21 / 126 NYSE sessions
2 Risk inverse-vol + breadth-scaled 15% target 126-session vol, 21-session covariance
3 Execution daily histories shown next_close + 10 bps one-way

2. Growth of $1

Solid line: Basic Trend replica. Dashed line: monthly static beta 60/15/15/10. Same 10 bps costs and next_close lag.

In [3]:
labels = pd.to_datetime(equity['labels'])
fig, ax = plt.subplots(figsize=(11, 5.5))
ax.plot(labels, equity['candidate'], lw=2.4, label='Prometheus Basic Trend replica')
ax.plot(labels, equity['beta'], '--', lw=1.8, label='Static beta 60/15/15/10')
ax.set(title='Annualized path summaries · base 10 bps · through 2026-07-24', xlabel='Date', ylabel='Growth of $1')
ax.set_yscale('log')
ax.grid(alpha=.25); ax.legend(frameon=False); plt.show()
No description has been provided for this image

3. Publisher fidelity vs untouched holdout

Fidelity asks: do rounded long-run published stats match? Holdout asks: does the replica beat its own beta after freeze?

In [4]:
gross = evaluation['publisher_fidelity']['gross_candidate']
targets = evaluation['publisher_fidelity']['targets']['prometheus']
fidelity = pd.DataFrame([
 {'metric':'excess return','observed':gross['annualized_arithmetic_excess_return'],'publisher':targets['annualized_arithmetic_excess_return']},
 {'metric':'volatility','observed':gross['volatility'],'publisher':targets['volatility']},
 {'metric':'max drawdown','observed':gross['maximum_drawdown'],'publisher':targets['maximum_drawdown']},
 {'metric':'Sharpe','observed':gross['sharpe'],'publisher':targets['sharpe']},
 {'metric':'corr to beta','observed':gross['correlation_to_beta'],'publisher':targets['correlation_to_beta']},
])
display(fidelity)
untouched = evaluation['primary_windows']['untouched_confirmatory']
pd.DataFrame([
 {'side':'candidate','Sharpe':untouched['candidate']['sharpe'],'MDD':untouched['candidate']['maximum_drawdown'],'excess':untouched['candidate']['annualized_arithmetic_excess_return']},
 {'side':'static beta','Sharpe':untouched['beta']['sharpe'],'MDD':untouched['beta']['maximum_drawdown'],'excess':untouched['beta']['annualized_arithmetic_excess_return']},
 {'side':'improvement','Sharpe':untouched['comparison']['sharpe_improvement'],'MDD':untouched['comparison']['maximum_drawdown_improvement'],'excess':untouched['comparison']['annualized_arithmetic_excess_return_difference']},
])
metric observed publisher
0 excess return 0.149000 0.158
1 volatility 0.103081 0.110
2 max drawdown -0.196791 -0.180
3 Sharpe 1.445463 1.440
4 corr to beta 0.463833 0.490
Out[4]:
side Sharpe MDD excess
0 candidate 0.940236 -0.098414 0.112949
1 static beta 1.578850 -0.131952 0.198356
2 improvement -0.638614 0.033538 -0.085407

4. Cost viewer (frozen 0 / 10 / 20 bps)

This control swaps among three already-computed untouched stresses. It is a sensitivity viewer, not an optimizer.

In [5]:
rows = {r['id']: r for r in evaluation['robustness_rows']}
payload = {
  'variants': {
    '0': {'label':'gross 0 bps','sharpe':rows['RET-GROSS']['candidate']['sharpe'],'improv':rows['RET-GROSS']['comparison']['sharpe_improvement'],'mdd':rows['RET-GROSS']['comparison']['maximum_drawdown_improvement']},
    '10': {'label':'base 10 bps','sharpe':rows['RET-PRIMARY']['candidate']['sharpe'],'improv':rows['RET-PRIMARY']['comparison']['sharpe_improvement'],'mdd':rows['RET-PRIMARY']['comparison']['maximum_drawdown_improvement']},
    '20': {'label':'stress 20 bps','sharpe':rows['RET-COST2X']['candidate']['sharpe'],'improv':rows['RET-COST2X']['comparison']['sharpe_improvement'],'mdd':rows['RET-COST2X']['comparison']['maximum_drawdown_improvement']},
  }
}
DATA = json.dumps(payload, separators=(',', ':'))
display(HTML('''<div id="pbt-cost" style="border:1px solid #ddd;padding:16px;border-radius:10px">
<label><b>One-way cost:</b> <span id="pbt-cost-value">10</span> bps</label>
<input id="pbt-cost-slider" type="range" min="0" max="20" value="10" step="10" style="width:100%">
<p id="pbt-cost-metrics"></p></div>
<script>(function(){const D=''' + DATA + ''';const s=document.getElementById('pbt-cost-slider'),v=document.getElementById('pbt-cost-value'),m=document.getElementById('pbt-cost-metrics');
function draw(){const b=s.value,x=D.variants[b];v.textContent=b;m.textContent=x.label+' · untouched Sharpe '+x.sharpe.toFixed(3)+' · Sharpe improvement '+x.improv.toFixed(3)+' · MDD improvement '+(100*x.mdd).toFixed(2)+' pp';}
s.addEventListener('input',draw);draw();})();</script>'''))

5. Decision

Fidelity only. Publisher tolerances pass; untouched economic admission floors fail. The implementation stays in the repository as research evidence. It does not enter Other / Strategies.

In [6]:
pd.DataFrame({'failure': evaluation['failures']})
Out[6]:
failure
0 untouched_sharpe_improvement_at_least_0_20
1 untouched_drawdown_improvement_at_least_0_08
2 adjacent_majority_three_of_five
3 extra_lag_nonnegative_improvements
4 double_cost_nonnegative_improvements