Joint SPX/VIX Volatility Research System  ·  14 components  ·  637 tests  ·  end-of-day data  ·  by Navnoor Bawa
Portfolio Equity Curve — % Return from $1M
── Portfolio ── S1C Contrarian ── S3 Dispersion ── S4 VRP ── S1 (ref)
C18 — A Strike-Rolling Bug Contaminated Every Prior Backtest Number
The straddle engine re-entered a position the same day it exited, so the position series never returned to flat. A continuously-held position is one trade: the strike was fixed at first entry and the tenor decayed to its 1-day floor — instead of rolling a fresh 30-day ATM straddle every max-hold. So the book held a single stale-strike, perpetually-expiring straddle — a directional |S−K| bet, not a vol trade. This contaminated all P&L from C10–C17. Fixed: positions now roll correctly, and the correction inverts the contrarian thesis. The "catastrophic" S1 (−$1.53M) that motivated flipping its direction into S1C was the artifact — corrected, S1 is a regime-gated short-VRP harvester (+$463K, 76% win) with a fat crisis tail, and the contrarian inversion S1C collapses to −$404K. The (S1C + S3 + S4) portfolio is −19.0% cumulative, Sharpe −1.57, max DD −25.9%. S3 dispersion timing (+$23K, 73% win) is the lone signal positive across every correction. Still a documented research result, not a deployable strategy: mark-to-model, concentrated, tail-heavy.
── S1C Contrarian ── S4 VRP
The stale-strike bug, mechanically: max-hold exits that still met the entry condition re-entered the same day, so position never returned to 0. In the P&L engine a never-zero position is a single trade whose 30-day straddle is struck once and then marked past expiry at the 1-day tenor floor. Fixing it (forcing a flat day, so straddles roll) re-prices every signal — and shows the contrarian story was built on the artifact, not on a real edge.
S1 — IVR SHORT-VRP (CORRECTED)
Total P&L $462,976 Sharpe 0.257 Win Rate 75.93% N Trades 54
53 of 54 trades are short straddles selling rich implied vol in R1 — a regime-gated VRP harvester. Its prior −$1.53M "catastrophe" was the stale-strike bug. Corrected: +$463K, 76% win, positive in 2018/19/21/23/24. But this is not a green light — the payoff is fat-tailed (a single −$333K COVID trade; 2022 −$99K) and concentrated (top-5 trades = 73% of profit), and it is mark-to-model. Honest read: short-VRP earns the premium in calm regimes and pays it back in spikes.
S1C — CONTRARIAN PDV (DEMOTED)
Total P&L $-404,183 Sharpe -0.622 Win Rate 45.45% N Trades 88
S1C flipped S1's direction on the premise that S1 lost both ways, so the PDV spread had to be contrarian. C18 verdict: that premise was the bug. Once S1 rolls its straddles properly it is profitable, so inverting it is the wrong trade — S1C is −$404K on 2018–2025. Demoted, like the ML classifier in C17: a thesis that only held under a bug is not a signal.
S4 — VOLATILITY RISK PREMIUM (VRP)
Total P&L $-188,519 Sharpe -1.164 Win Rate 32.26% N Trades 31
Entry requires z_vrp declining from peak (post-spike mean-reversion, not pre-event entry). Under correct straddle rolling + accurate regime labels + margin caps: −$189K on 2018–2025 — consistent conclusion across every revision: VRP is not extractable at retail costs.
Key Performance Metrics
Cumulative Return -19.00%
Ann. Return -2.89%
Sharpe (rf=T-bill) -1.568
Sortino -3.004
Max Drawdown -25.92%
DD Duration 1729 days
Calmar -0.112
Win Rate 29.55%
Avg P&L / Trade $-1,080
Best Day 4.13%
Worst Day -0.91%
N Trades 176
N Days 1808
Total P&L by Signal (in $000s)
Annual Return — 2018 → 2025
Walk-Forward Sharpe — 6-Month OOS Windows
10/13 negative
Per-Signal Performance Breakdown
Signal Sharpe Ann. Return Win Rate N Trades Total P&L Status
S1C Contrarian MAIN -0.622 -6.97% 45.45% 88 $-404,183 NEGATIVE
S3 Dispersion MAIN N/A* 0.32% 72.73% 22 $23,212 POSITIVE
* Sharpe N/A — S3 has 22 trades over 7 years (mostly flat NAV → near-zero std → ratio diverges). Use win rate (72.73%) and total P&L ($23,212) instead.
S4 VRP MAIN -1.164 -2.88% 32.26% 31 $-188,519 NEGATIVE
S1 IVR (ref) REF 0.257 5.45% 75.93% 54 $462,976 REFERENCE
S2 VIX TS (ref) REF -0.789 -1.79% 35.14% 37 $-121,370 REFERENCE
MAIN signals (S1C, S3, S4) form the audited portfolio — all numbers under the fully corrected C17 pipeline. REF signals (S1, S2) are retained for historical comparison only.
Regime-Exit Variants — Historical S1/S2 Research
PRE-C15 REFERENCE
Before C15, S1 and S2 were the main signals. These variants show the impact of exiting positions when the regime transitions to R2 (rather than blocking new entries only). S1 was replaced by S1C Contrarian in the main portfolio at C15 — a decision C18 reverses in spirit: once straddles roll correctly, S1 (the short-VRP harvester) is the stronger leg and the contrarian inversion S1C is the weaker one.
Signal Original P&L R2-Exit P&L Improvement
S1 IVR REF $462,976 $+275,954 $-187,022
S2 VIX TS REF $-121,370 $-73,236 $+48,135
S3 Dispersion MAIN $+23,212 unchanged
R2 = VOMMA_ACTIVE (deterministic rule: VVIX > 100). C17: regime labels are the lagged rule itself (persistence predictor) — the ML classifier is research-only. R2-exit: forced closure + state reset on R2 transition (S1C, S1X, S2X, S4). Backtest: 2018–2025, $1M notional.
C18 — The Contrarian Thesis Was a Strike-Rolling Artifact
S1 +$463K The "S1 loses both ways" diagnosis that launched the contrarian idea was the bug: positions never rolled, so S1 held one stale-strike straddle past expiry. Rolling properly, S1 is a regime-gated short-VRP harvester — profitable in calm regimes, with a fat crisis tail.
S1C −$404K Inverting a signal that is actually profitable is the wrong trade. The contrarian flip only looked good while S1 looked catastrophic — and that was the artifact. S1C is demoted, exactly like the ML classifier in C17: a thesis that holds only under a bug is not a signal.
PORTFOLIO The (S1C + S3 + S4) construction, still anchored on the discredited contrarian leg: −19.0% cumulative, max DD −25.9%, Sharpe −1.57 vs the actual T-bill path. Crisis windows stay positive (COVID, Fed, Tariffs all small +) but cannot pay for the bleed between them.
Signal Breakdown — Only S3's Timing Survives
S3 +$23K Dispersion proxy fires only at VIX/VVIX extremes — 22 trades in 7 years, 73% win rate, positive at every tested P&L scale (0.5%–4%/z-unit) and in pseudo-OOS. Unchanged by the strike-rolling fix — the one robust finding; dollar magnitude is proxy-scaled, not market-derived.
S1 +$463K 76% win, but not deployable: fat-tailed (one −$333K COVID trade), concentrated (top-5 = 73% of profit), and mark-to-model. Short-VRP earns the premium in calm regimes and pays it back in spikes.
S4 −$189K VRP mean-reversion negative in-sample under every revision — consistent: VRP is not extractable at retail cost levels.
C16–C18 Methodology Audit — Ten Flaws Found, Eight Fixed, Two Quantified
Self-audit of the backtest methodology (C16: 2026-06-11, C17 redteam: same week). All numbers on this page are produced under the fully corrected pipeline. Each successive fix made the backtest worse — the original "edge" was a stack of artifacts.
# Flaw Resolution Evidence
1 PDV look-ahead bias FIXED — PDVLinear retrained per backtest year on strictly pre-year returns (walk-forward), injected into the pdv_iv_spread feature. S1/S1C now trade a true causal PDV forecast
2 Circular classifier feature FIXED — vvix removed from training features (the R2 label is defined as VVIX>threshold, so the old 86.2% accuracy partly measured rule-recovery). Honest accuracy: 63.4%. See flaw #6 — the honest number then lost to the baseline
3 Retrospective Kelly netting FIXED — correlation/Sharpe multipliers now shifted 1 day and applied to position size in a second simulation pass (contracts & costs reflect real size). Ex-ante sizing; no same-day information
4 S1C data snooping QUANTIFIED — direction was chosen after observing S1 fail on 2018–2025; cannot be fixed in code. Tested on 2013–2017, a window never used in the discovery. C17 re-ran the test under corrected labels. Pseudo-OOS under C17: S1C −$333K — the apparent edge was an artifact
5 S3 arbitrary P&L scale QUANTIFIED — the 2%/z-unit proxy scale is assumed (no single-stock options data). Sensitivity run at 0.5%/1%/2%/4% per z-unit. P&L +$5K…+$52K, win 15–17 of 17 at every scale
6 Classifier never benchmarked vs persistence FIXED (C17) — the no-skill baseline "predict yesterday's regime" scores 90.0% on 2020+ (labels are same-day observables, so y(t−1) is known at t−1). The honest 63.4% classifier LOSES by 27 points. ML demoted from the trading loop; backtest uses lagged rule labels directly. 63.4% vs 90.0% — the ML added negative value
7 Margin model documented but not enforced FIXED (C17) — MARGIN_RATIO was defined but never used: short straddles had unlimited capacity. Now capped at nav/(0.20·K·100) contracts at entry. Cost model now matches its documentation
8 S1C had no R2 exit (unbounded tail) FIXED (C17) — an open S1C short straddle could ride a VVIX spike for up to 21 days with no stop. S1C now force-closes when the regime transitions into R2, same as S1X/S2X/S4. Worst-case short-gamma exposure now bounded by the R2 rule
9 Dead "monthly Heston recalibration" + wrong rf FIXED (C17) — the monthly-recalibration loop never produced a single calibration (no pre-2026 options data) and its output was never consumed; removed. Sharpe/Sortino now measured against the per-date ^IRX T-bill path instead of a flat 5%. Backtest claims now match what the code does
10 Straddles never rolled (stale strike, decayed tenor) FIXED (C18) — the state machine re-entered the same day it exited, so the position never went flat: the engine held one straddle struck once and marked past expiry at the 1-day tenor floor, instead of rolling a fresh 30-day ATM straddle each cycle. Now forces a flat day so positions roll. This contaminated every P&L number in C10–C17. Inverted the thesis: S1 −$1.53M→+$463K, S1C −$97K→−$404K, portfolio −9.1%→−19.0%; S3 unchanged
Honest read: after removing every identified artifact — including the strike-rolling bug that had been distorting the P&L the whole time — the portfolio is −19.0% and the only signal with cross-correction robustness is S3's timing. S1 is a real short-VRP harvester (+$463K in-sample) but fat-tailed, concentrated and mark-to-model — not deployable. This dashboard documents a negative result with full methodology: the infrastructure (calibration, Greeks, hedging, walk-forward harness) is the deliverable; the trading signals are the case study in why most backtests flatter themselves.
Annual Returns Heatmap
2018
-3.7%
2019
-11.8%
2020
+3.1%
2021
-1.0%
2022
+3.4%
2023
-9.0%
2024
-3.9%
2025
+3.2%
System Improvements — Steps 1–8 Summary
Step Name What Changed Before After
1 Disable VIX options leg JOINT_W3 = 0.0 (was 0.2). Heston CIR density is structurally mis-specified for VIX options. VIX opt RMSE 37.14 vp (corrupts calibration) SPX-only objective. RMSE expected ↓ from 5.3 → ~2.5 vp.
2 Vega-weighted anchor selection Select 9 highest-BS-vega options per expiry (was log-uniform grid). Moneyness lo 0.70 (was 0.75). Log-uniform includes low-information deep OTM options Near-ATM anchors only; more stable Heston gradient.
3 SVI smoothing (SSVI surface) Fit Gatheral (2004) SVI per expiry; calibrate Heston to smooth surface not noisy quotes. SPX RMSE ~5.3 vol pts (market microstructure noise) Target SPX RMSE ~2.0 vol pts. Butterfly-free surface.
4 Per-date T-bill rate ^IRX 3M T-bill from DB replaces fixed r=0.045 in backtest. 18 new tests. Fixed r=0.045 throughout 2018–2025 Per-date rate (2018: ~1.5%, 2023: ~5.3%, 2025: ~4.3%).
5 Adaptive VVIX threshold Rolling 252-day 80th percentile of VVIX as R2 gate (was fixed 100). R2 frequency ~53% (over-triggered in low-vol periods) Expected R2 frequency ~25% (calibrated to market regime).
6 Isotonic calibration + step sizing CalibratedClassifierCV(FrozenEstimator(XGB), isotonic). S1S step: full/half/flat at P(R2) 0.4/0.6. Raw XGB probabilities (overconfident). Linear 1−p scaling. Calibrated probabilities. Cleaner 3-level position sizing.
7 Portfolio Kelly netting Rolling 60-day pairwise P&L corr; mult = sqrt(2/(1+r_ij)). C16: multipliers shifted 1 day, applied to position size in a second simulation pass (ex-ante; contracts & costs reflect real size). No correlation netting; each signal sizes off full NAV ÷ 4 Ex-ante diversification sizing; no same-day information in sizing.
8 Monthly Heston recalibration — REMOVED (C17) The C13 monthly-recalibration loop was dead code: no historical options data exists before the 2026 snapshots, so every monthly calibration fell back to static defaults — and the resulting parameters were never consumed by any P&L path. Removed in C17. ASSUMPTIONS claimed time-varying Heston params in the backtest Honest: backtest P&L is BS on VIX-TS ATM vol; the calibration stack serves C6 Greeks / C7 hedge sim, not C10 P&L.
Roadmap — Where This Goes With More Resources
Everything above was built on free data, and the single ceiling is data, not method. The honest negative result proves the infrastructure and the discipline; the paid-tier version is where you find out whether the edge is real. Rigor first, then capital — that progression is the plan. Each item below maps a limitation we documented to the concrete unlock that would lift it.
1 · Real historical option prices
Now: all P&L is Black-Scholes on VIX-ATM vol — the database holds a single options snapshot (2026-03-24), so every dollar is mark-to-model, not a traded fill.
Unlocks: real fillable prices across the smile turn the backtest from an estimate into evidence. The single highest-leverage change in the whole project.
≈ $50–250/mo for EOD option chains; OptionMetrics-grade is institutional.
2 · Intraday / high-frequency data
Now: jump parameters are unidentifiable on daily returns (Bates λ collapses) — daily data can't separate frequent small jumps from rare large ones.
Unlocks: the Barndorff-Nielsen-Shephard bipower-variation decomposition separates continuous vol from jumps, making the Bates / Merton leg estimable and the VIX-options fit tractable.
High-frequency bars; the missing ingredient for the jump models.
3 · Breadth — a real book, not a few trades
Now: the only signal robust across every correction (S3 dispersion) fires 22 times in 7 years — by Grinold's Fundamental Law, that breadth is fatal.
Unlocks: running dispersion across S&P constituents and more underlyings turns a handful of bets into hundreds of near-independent ones — the regime where a thin edge becomes a usable information ratio.
Single-name vol data + universe construction.
4 · Realistic execution + more compute
Now: a flat cost assumption, and the two-factor Quintic OU calibration takes ~46 min on 8 cores while the VIX-options fit sits at 17.5 vp (target < 10).
Unlocks: fitting bid-ask and slippage curves to real quotes tests whether the thin vol-carry edge survives the round trip; more cores or a GPU plus a stronger global optimizer push the joint fit under 10 vp and let it run intraday.
Quote-level cost data + compute.
None of these change the honest conclusion that the current edge is thin. They change whether it can be measured and scaled — which is the actual research program, and the part that needs resources rather than more code.