C10 Backtest Results — 2018 → 2025
Delta-hedged straddle portfolio · 3 strategies (S1C + S3 + S4) · $1M initial capital
NAV: $809,972 · P&L: $-190,028
Portfolio Equity Curve — % Return from $1M
── Portfolio
── S1C Contrarian
── S3 Dispersion
── S4 VRP
── S1 (ref)
C18 — A Strike-Rolling Bug Contaminated Every Prior Backtest Number
The straddle engine re-entered a position the same day it exited, so the position series never
returned to flat. A continuously-held position is one trade: the strike was fixed at first entry
and the tenor decayed to its 1-day floor — instead of rolling a fresh 30-day ATM straddle every
max-hold. So the book held a single stale-strike, perpetually-expiring
straddle — a directional |S−K| bet, not a vol trade. This contaminated all P&L from C10–C17.
Fixed: positions now roll correctly, and the correction inverts the
contrarian thesis. The "catastrophic" S1 (−$1.53M) that motivated flipping its direction into S1C was the
artifact — corrected, S1 is a regime-gated short-VRP harvester (+$463K, 76% win)
with a fat crisis tail, and the contrarian inversion S1C collapses to −$404K.
The (S1C + S3 + S4) portfolio is −19.0% cumulative, Sharpe −1.57, max DD −25.9%.
S3 dispersion timing (+$23K, 73% win) is the lone signal positive across every correction.
Still a documented research result, not a deployable strategy: mark-to-model, concentrated, tail-heavy.
── S1C Contrarian
── S4 VRP
The stale-strike bug, mechanically: max-hold exits that still met the entry
condition re-entered the same day, so
position never returned to 0. In the P&L engine a
never-zero position is a single trade whose 30-day straddle is struck once and then marked past expiry at the
1-day tenor floor. Fixing it (forcing a flat day, so straddles roll) re-prices every signal — and shows the
contrarian story was built on the artifact, not on a real edge.
S1 — IVR SHORT-VRP (CORRECTED)
Total P&L
$462,976
Sharpe
0.257
Win Rate
75.93%
N Trades
54
53 of 54 trades are short straddles selling rich implied vol in R1 — a regime-gated VRP harvester.
Its prior −$1.53M "catastrophe" was the stale-strike bug.
Corrected: +$463K, 76% win, positive in 2018/19/21/23/24.
But this is not a green light — the payoff is fat-tailed (a single −$333K COVID trade; 2022 −$99K)
and concentrated (top-5 trades = 73% of profit), and it is mark-to-model. Honest read: short-VRP
earns the premium in calm regimes and pays it back in spikes.
S1C — CONTRARIAN PDV (DEMOTED)
Total P&L
$-404,183
Sharpe
-0.622
Win Rate
45.45%
N Trades
88
S1C flipped S1's direction on the premise that S1 lost both ways, so the PDV spread had to be contrarian.
C18 verdict: that premise was the bug. Once S1 rolls its straddles
properly it is profitable, so inverting it is the wrong trade — S1C is
−$404K on 2018–2025. Demoted, like the ML classifier in C17:
a thesis that only held under a bug is not a signal.
S4 — VOLATILITY RISK PREMIUM (VRP)
Total P&L
$-188,519
Sharpe
-1.164
Win Rate
32.26%
N Trades
31
Entry requires z_vrp declining from peak (post-spike mean-reversion, not pre-event entry).
Under correct straddle rolling + accurate regime labels + margin caps:
−$189K on 2018–2025 —
consistent conclusion across every revision: VRP is not extractable at retail costs.
Key Performance Metrics
| Cumulative Return | -19.00% |
| Ann. Return | -2.89% |
| Sharpe (rf=T-bill) | -1.568 |
| Sortino | -3.004 |
| Max Drawdown | -25.92% |
| DD Duration | 1729 days |
| Calmar | -0.112 |
| Win Rate | 29.55% |
| Avg P&L / Trade | $-1,080 |
| Best Day | 4.13% |
| Worst Day | -0.91% |
| N Trades | 176 |
| N Days | 1808 |
Total P&L by Signal (in $000s)
Annual Return — 2018 → 2025
Walk-Forward Sharpe — 6-Month OOS Windows
10/13 negative
Per-Signal Performance Breakdown
| Signal | Sharpe | Ann. Return | Win Rate | N Trades | Total P&L | Status |
|---|---|---|---|---|---|---|
| S1C Contrarian MAIN | -0.622 | -6.97% | 45.45% | 88 | $-404,183 | NEGATIVE |
| S3 Dispersion MAIN | N/A* | 0.32% | 72.73% | 22 | $23,212 | POSITIVE |
| * Sharpe N/A — S3 has 22 trades over 7 years (mostly flat NAV → near-zero std → ratio diverges). Use win rate (72.73%) and total P&L ($23,212) instead. | ||||||
| S4 VRP MAIN | -1.164 | -2.88% | 32.26% | 31 | $-188,519 | NEGATIVE |
| S1 IVR (ref) REF | 0.257 | 5.45% | 75.93% | 54 | $462,976 | REFERENCE |
| S2 VIX TS (ref) REF | -0.789 | -1.79% | 35.14% | 37 | $-121,370 | REFERENCE |
MAIN signals (S1C, S3, S4) form the audited portfolio — all numbers under the fully corrected C17 pipeline. REF signals (S1, S2) are retained for historical comparison only.
Regime-Exit Variants — Historical S1/S2 Research
PRE-C15 REFERENCE
Before C15, S1 and S2 were the main signals. These variants show the impact of exiting positions when the regime transitions to R2 (rather than blocking new entries only). S1 was replaced by S1C Contrarian in the main portfolio at C15 — a decision C18 reverses in spirit: once straddles roll correctly, S1 (the short-VRP harvester) is the stronger leg and the contrarian inversion S1C is the weaker one.
| Signal | Original P&L | R2-Exit P&L | Improvement |
|---|---|---|---|
| S1 IVR REF | $462,976 | $+275,954 | $-187,022 |
| S2 VIX TS REF | $-121,370 | $-73,236 | $+48,135 |
| S3 Dispersion MAIN | $+23,212 | unchanged | — |
R2 = VOMMA_ACTIVE (deterministic rule: VVIX > 100). C17: regime labels are the lagged rule itself (persistence predictor) — the ML classifier is research-only. R2-exit: forced closure + state reset on R2 transition (S1C, S1X, S2X, S4). Backtest: 2018–2025, $1M notional.
C18 — The Contrarian Thesis Was a Strike-Rolling Artifact
S1 +$463K
The "S1 loses both ways" diagnosis that launched the contrarian idea was the bug:
positions never rolled, so S1 held one stale-strike straddle past expiry. Rolling
properly, S1 is a regime-gated short-VRP harvester
— profitable in calm regimes, with a fat crisis tail.
S1C −$404K
Inverting a signal that is actually profitable is the wrong trade. The contrarian
flip only looked good while S1 looked catastrophic — and that was the artifact.
S1C is demoted, exactly like the ML classifier in C17:
a thesis that holds only under a bug is not a signal.
PORTFOLIO
The (S1C + S3 + S4) construction, still anchored on the discredited contrarian leg:
−19.0% cumulative, max DD −25.9%, Sharpe −1.57
vs the actual T-bill path. Crisis windows stay positive (COVID, Fed, Tariffs all small +)
but cannot pay for the bleed between them.
Signal Breakdown — Only S3's Timing Survives
S3 +$23K
Dispersion proxy fires only at VIX/VVIX extremes — 22 trades in 7 years,
73% win rate, positive at every tested P&L
scale (0.5%–4%/z-unit) and in pseudo-OOS. Unchanged by the strike-rolling fix — the one
robust finding; dollar magnitude is proxy-scaled, not market-derived.
S1 +$463K
76% win, but not deployable: fat-tailed (one −$333K
COVID trade), concentrated (top-5 = 73% of profit), and mark-to-model. Short-VRP earns the
premium in calm regimes and pays it back in spikes.
S4 −$189K
VRP mean-reversion negative in-sample under every revision — consistent:
VRP is not extractable at retail cost levels.
C16–C18 Methodology Audit — Ten Flaws Found, Eight Fixed, Two Quantified
Self-audit of the backtest methodology (C16: 2026-06-11, C17 redteam: same week). All numbers on this page are produced under the fully corrected pipeline.
Each successive fix made the backtest worse — the original "edge" was a stack of artifacts.
| # | Flaw | Resolution | Evidence |
|---|---|---|---|
| 1 | PDV look-ahead bias | FIXED — PDVLinear retrained per backtest year on strictly pre-year returns (walk-forward), injected into the pdv_iv_spread feature. | S1/S1C now trade a true causal PDV forecast |
| 2 | Circular classifier feature | FIXED — vvix removed from training features (the R2 label is defined as VVIX>threshold, so the old 86.2% accuracy partly measured rule-recovery). Honest accuracy: 63.4%. | See flaw #6 — the honest number then lost to the baseline |
| 3 | Retrospective Kelly netting | FIXED — correlation/Sharpe multipliers now shifted 1 day and applied to position size in a second simulation pass (contracts & costs reflect real size). | Ex-ante sizing; no same-day information |
| 4 | S1C data snooping | QUANTIFIED — direction was chosen after observing S1 fail on 2018–2025; cannot be fixed in code. Tested on 2013–2017, a window never used in the discovery. C17 re-ran the test under corrected labels. | Pseudo-OOS under C17: S1C −$333K — the apparent edge was an artifact |
| 5 | S3 arbitrary P&L scale | QUANTIFIED — the 2%/z-unit proxy scale is assumed (no single-stock options data). Sensitivity run at 0.5%/1%/2%/4% per z-unit. | P&L +$5K…+$52K, win 15–17 of 17 at every scale |
| 6 | Classifier never benchmarked vs persistence | FIXED (C17) — the no-skill baseline "predict yesterday's regime" scores 90.0% on 2020+ (labels are same-day observables, so y(t−1) is known at t−1). The honest 63.4% classifier LOSES by 27 points. ML demoted from the trading loop; backtest uses lagged rule labels directly. | 63.4% vs 90.0% — the ML added negative value |
| 7 | Margin model documented but not enforced | FIXED (C17) — MARGIN_RATIO was defined but never used: short straddles had unlimited capacity. Now capped at nav/(0.20·K·100) contracts at entry. | Cost model now matches its documentation |
| 8 | S1C had no R2 exit (unbounded tail) | FIXED (C17) — an open S1C short straddle could ride a VVIX spike for up to 21 days with no stop. S1C now force-closes when the regime transitions into R2, same as S1X/S2X/S4. | Worst-case short-gamma exposure now bounded by the R2 rule |
| 9 | Dead "monthly Heston recalibration" + wrong rf | FIXED (C17) — the monthly-recalibration loop never produced a single calibration (no pre-2026 options data) and its output was never consumed; removed. Sharpe/Sortino now measured against the per-date ^IRX T-bill path instead of a flat 5%. | Backtest claims now match what the code does |
| 10 | Straddles never rolled (stale strike, decayed tenor) | FIXED (C18) — the state machine re-entered the same day it exited, so the position never went flat: the engine held one straddle struck once and marked past expiry at the 1-day tenor floor, instead of rolling a fresh 30-day ATM straddle each cycle. Now forces a flat day so positions roll. This contaminated every P&L number in C10–C17. | Inverted the thesis: S1 −$1.53M→+$463K, S1C −$97K→−$404K, portfolio −9.1%→−19.0%; S3 unchanged |
Honest read: after removing every identified artifact — including the strike-rolling bug that had been
distorting the P&L the whole time — the portfolio is −19.0% and the
only signal with cross-correction robustness is S3's timing. S1 is a real short-VRP harvester (+$463K in-sample)
but fat-tailed, concentrated and mark-to-model — not deployable. This dashboard documents a negative result with
full methodology: the infrastructure (calibration, Greeks, hedging, walk-forward harness) is the deliverable;
the trading signals are the case study in why most backtests flatter themselves.
Annual Returns Heatmap
2018
-3.7%
2019
-11.8%
2020
+3.1%
2021
-1.0%
2022
+3.4%
2023
-9.0%
2024
-3.9%
2025
+3.2%
System Improvements — Steps 1–8 Summary
| Step | Name | What Changed | Before | After |
|---|---|---|---|---|
| 1 | Disable VIX options leg | JOINT_W3 = 0.0 (was 0.2). Heston CIR density is structurally mis-specified for VIX options. | VIX opt RMSE 37.14 vp (corrupts calibration) | SPX-only objective. RMSE expected ↓ from 5.3 → ~2.5 vp. |
| 2 | Vega-weighted anchor selection | Select 9 highest-BS-vega options per expiry (was log-uniform grid). Moneyness lo 0.70 (was 0.75). | Log-uniform includes low-information deep OTM options | Near-ATM anchors only; more stable Heston gradient. |
| 3 | SVI smoothing (SSVI surface) | Fit Gatheral (2004) SVI per expiry; calibrate Heston to smooth surface not noisy quotes. | SPX RMSE ~5.3 vol pts (market microstructure noise) | Target SPX RMSE ~2.0 vol pts. Butterfly-free surface. |
| 4 | Per-date T-bill rate | ^IRX 3M T-bill from DB replaces fixed r=0.045 in backtest. 18 new tests. | Fixed r=0.045 throughout 2018–2025 | Per-date rate (2018: ~1.5%, 2023: ~5.3%, 2025: ~4.3%). |
| 5 | Adaptive VVIX threshold | Rolling 252-day 80th percentile of VVIX as R2 gate (was fixed 100). | R2 frequency ~53% (over-triggered in low-vol periods) | Expected R2 frequency ~25% (calibrated to market regime). |
| 6 | Isotonic calibration + step sizing | CalibratedClassifierCV(FrozenEstimator(XGB), isotonic). S1S step: full/half/flat at P(R2) 0.4/0.6. | Raw XGB probabilities (overconfident). Linear 1−p scaling. | Calibrated probabilities. Cleaner 3-level position sizing. |
| 7 | Portfolio Kelly netting | Rolling 60-day pairwise P&L corr; mult = sqrt(2/(1+r_ij)). C16: multipliers shifted 1 day, applied to position size in a second simulation pass (ex-ante; contracts & costs reflect real size). | No correlation netting; each signal sizes off full NAV ÷ 4 | Ex-ante diversification sizing; no same-day information in sizing. |
| 8 | Monthly Heston recalibration — REMOVED (C17) | The C13 monthly-recalibration loop was dead code: no historical options data exists before the 2026 snapshots, so every monthly calibration fell back to static defaults — and the resulting parameters were never consumed by any P&L path. Removed in C17. | ASSUMPTIONS claimed time-varying Heston params in the backtest | Honest: backtest P&L is BS on VIX-TS ATM vol; the calibration stack serves C6 Greeks / C7 hedge sim, not C10 P&L. |
Roadmap — Where This Goes With More Resources
Everything above was built on free data, and the single ceiling is
data, not method. The honest negative result proves the infrastructure and the
discipline; the paid-tier version is where you find out whether the edge is real.
Rigor first, then capital — that progression is the plan. Each item
below maps a limitation we documented to the concrete unlock that would lift it.
1 · Real historical option prices
Now: all P&L is Black-Scholes on VIX-ATM vol — the database holds a single options snapshot (2026-03-24), so every dollar is mark-to-model, not a traded fill.
Unlocks: real fillable prices across the smile turn the backtest from an estimate into evidence. The single highest-leverage change in the whole project.
≈ $50–250/mo for EOD option chains; OptionMetrics-grade is institutional.
Unlocks: real fillable prices across the smile turn the backtest from an estimate into evidence. The single highest-leverage change in the whole project.
≈ $50–250/mo for EOD option chains; OptionMetrics-grade is institutional.
2 · Intraday / high-frequency data
Now: jump parameters are unidentifiable on daily returns (Bates λ collapses) — daily data can't separate frequent small jumps from rare large ones.
Unlocks: the Barndorff-Nielsen-Shephard bipower-variation decomposition separates continuous vol from jumps, making the Bates / Merton leg estimable and the VIX-options fit tractable.
High-frequency bars; the missing ingredient for the jump models.
Unlocks: the Barndorff-Nielsen-Shephard bipower-variation decomposition separates continuous vol from jumps, making the Bates / Merton leg estimable and the VIX-options fit tractable.
High-frequency bars; the missing ingredient for the jump models.
3 · Breadth — a real book, not a few trades
Now: the only signal robust across every correction (S3 dispersion) fires 22 times in 7 years — by Grinold's Fundamental Law, that breadth is fatal.
Unlocks: running dispersion across S&P constituents and more underlyings turns a handful of bets into hundreds of near-independent ones — the regime where a thin edge becomes a usable information ratio.
Single-name vol data + universe construction.
Unlocks: running dispersion across S&P constituents and more underlyings turns a handful of bets into hundreds of near-independent ones — the regime where a thin edge becomes a usable information ratio.
Single-name vol data + universe construction.
4 · Realistic execution + more compute
Now: a flat cost assumption, and the two-factor Quintic OU calibration takes ~46 min on 8 cores while the VIX-options fit sits at 17.5 vp (target < 10).
Unlocks: fitting bid-ask and slippage curves to real quotes tests whether the thin vol-carry edge survives the round trip; more cores or a GPU plus a stronger global optimizer push the joint fit under 10 vp and let it run intraday.
Quote-level cost data + compute.
Unlocks: fitting bid-ask and slippage curves to real quotes tests whether the thin vol-carry edge survives the round trip; more cores or a GPU plus a stronger global optimizer push the joint fit under 10 vp and let it run intraday.
Quote-level cost data + compute.
None of these change the honest conclusion that the current edge is thin. They change
whether it can be measured and scaled — which is the actual research
program, and the part that needs resources rather than more code.