Volatility Risk Forecasting Platform
The platform asks not only which model forecasts volatility best, but whether the resulting risk estimates behave as expected when market regimes change.
Central findingThe strongest model materially improves QLIKE forecast loss over rolling, EWMA and GARCH baselines while maintaining plausible VaR breach frequencies.
Evidence
The result in context
- QLIKE improvement vs rolling-21 baseline
- 40.41%
- Improvement vs EWMA(0.94)
- 83.15%
- Improvement vs GARCH(1,1)
- 91.50%
- Persisted analytical rows
- 2.87m
Question
Which forecasting methods improve on simple baselines, and do their risk estimates survive formal backtesting?
A reproducible forecasting and risk platform comparing statistical and machine-learning models through volatility and VaR backtests.
Forecasting
Simple baselines remain part of the model set
Rolling volatility, exponentially weighted volatility, HAR and GARCH sit beside machine-learning models and ensembles. QLIKE, a loss measure suited to volatility forecasts, provides a common comparison.
The best model improves QLIKE by 40.41% versus a rolling-21 baseline, 83.15% versus EWMA and 91.50% versus GARCH in the committed evaluation.
Model comparison
Forecast model explorer
Compare what each method assumes before comparing its error.
Rolling 21-day baseline
A simple recent-history estimate; transparent and difficult to justify omitting.
Interactive evidence
Rolling HAR leads the full 19-model QLIKE ranking
Average QLIKE across ten assets; lower is better.
Read: The rolling-update HAR records 0.00558 average QLIKE, modestly ahead of the next rolling GARCH-family models.
Data table · 19 verified rows
| Model | Asset Count | Avg Qlike | Avg Rank | Top Two Share Pct | Improvement Vs Rolling21 Pct | Improvement Vs Ewma Pct | Improvement Vs Garch Pct |
|---|---|---|---|---|---|---|---|
| Har Rolling Update | 10 | 0.005582 | 1.8 | 80 | 40.411555 | 83.153034 | 91.502378 |
| Garch Rolling Update | 10 | 0.005611 | 2.4 | 50 | 40.098847 | 83.072203 | 91.452533 |
| Gjr Rolling Update | 10 | 0.005633 | 2.5 | 60 | 39.787084 | 82.98345 | 91.378658 |
| Ewma Rolling Update | 10 | 0.005679 | 3.3 | 10 | 39.37768 | 82.863334 | 91.345209 |
| Validation Weighted Ensemble | 10 | 0.007927 | 5.6 | 0 | 16.534282 | 76.379379 | 88.148684 |
| Har Rv Market Huber | 10 | 0.00914 | 7.3 | 0 | 1.281171 | 72.042367 | 85.860129 |
| Previous Day Rv | 10 | 0.009263 | 7.9 | 0 | 0 | 71.72236 | 85.666932 |
| Rolling 21 | 10 | 0.009263 | 7.9 | 0 | 0 | 71.72236 | 85.666932 |
| Har Rv Log Ridge | 10 | 0.009417 | 7.9 | 0 | -1.263502 | 71.328464 | 85.618467 |
| Simple Average Ensemble | 10 | 0.009565 | 9.3 | 0 | -2.823233 | 70.923319 | 85.465859 |
| Har Rv Market Log Ridge | 10 | 0.009657 | 9.1 | 0 | -3.675557 | 70.65161 | 85.302816 |
| Random Forest | 10 | 0.023479 | 12.8 | 0 | -175.924532 | 21.393393 | 61.049185 |
| Hist Gradient Boosting | 10 | 0.026876 | 13.6 | 0 | -210.828291 | 11.374642 | 57.072549 |
| Ewma Tuned | 10 | 0.029268 | 13.8 | 0 | -215.362441 | 10.727555 | 55.012938 |
| Ewma 094 | 10 | 0.0331 | 14.4 | 0 | -256.779806 | 0 | 48.809942 |
| Garch 11 | 10 | 0.07411 | 16.5 | 0 | -668.658624 | -117.891564 | 0 |
| Egarch T | 10 | 0.076058 | 16.9 | 0 | -723.912752 | -132.668788 | -12.670258 |
| Rolling 63 | 10 | 0.099935 | 17.8 | 0 | -999.837601 | -207.68123 | -61.976987 |
| Gjr Garch | 10 | 0.158004 | 18.2 | 0 | -1,617.207943 | -366.905553 | -151.149655 |
Risk
Forecasts become risk estimates
Value at Risk, or VaR, estimates a loss threshold expected to be exceeded only at a stated frequency. The 95% VaR breach rate is 4.76%; the 99% rate is 0.81%.
Kupiec and Christoffersen tests examine whether breaches occur at the expected rate and whether they cluster through time.
Interactive evidence
Observed VaR breach rates remain close to their expected levels
Mean across ten assets and nineteen forecasting models.
Read: The realised rates are 4.76% at 95% confidence and 0.81% at 99%, compared with 5% and 1% expected.
Data table · 2 verified rows
| Confidence Level Pct | Expected Breach Rate Pct | Observed Breach Rate Pct | Kupiec Pass Share Pct | Christoffersen Pass Share Pct |
|---|---|---|---|---|
| 95 | 5 | 4.760787 | 92.105263 | 89.473684 |
| 99 | 1 | 0.813179 | 79.473684 | 86.315789 |
Platform
The analytical layer is designed for repeated investigation
DuckDB and SQL support 2.87 million analytical rows, while 88 automated checks protect rolling-window logic, forecasts and risk outputs. The dashboard’s measured p95 response is 16.14ms on the documented workload.
Interactive evidence
High-volatility regimes capture most top-decile volatility days
Top-decile capture and high-regime precision by asset.
Read: Capture is intentionally prioritised over precision: the state label catches severe volatility days while accepting false high-regime flags.
Data table · 10 verified rows
| Asset | Top Decile Capture Pct | Precision Pct | High Regime Share Pct | Median Regime Days |
|---|---|---|---|---|
| AAPL | 82.105263 | 44.571429 | 18.479409 | 8 |
| GLD | 80.701754 | 40.564374 | 19.957761 | 5 |
| IWM | 62.45614 | 35.528942 | 17.634636 | 5 |
| JPM | 70.877193 | 40.319361 | 17.634636 | 5 |
| MSFT | 71.22807 | 43.376068 | 16.473073 | 6 |
| NVDA | 96.842105 | 55.08982 | 17.634636 | 7 |
| QQQ | 78.596491 | 43.243243 | 18.233017 | 7.5 |
| SPY | 71.22807 | 41.260163 | 17.317846 | 9 |
| TLT | 66.666667 | 39.915966 | 16.754664 | 4 |
| USO | 73.684211 | 50.724638 | 14.572334 | 4 |
Limitations
What this evidence does not establish
- Forecast improvements are specific to the dataset, horizon and evaluation window used.
- A calibrated VaR breach rate does not capture every form of tail risk or liquidity stress.
Source and reproducibility
Trace the evidence
Source code, evaluation outputs and supporting material are available in the repository.
View repository- Model comparisonreports/model_ranking_results.csvCommit / evidence ID: ae7dd4781b6e039ad70b9e57d448f232ca0af3d6
- VaR backtestingreports/var_backtest_results.csvCommit / evidence ID: ae7dd4781b6e039ad70b9e57d448f232ca0af3d6
- Regime detectionreports/regime_detection_results.csvCommit / evidence ID: ae7dd4781b6e039ad70b9e57d448f232ca0af3d6