Econometric models beat machine learning in volatility forecasts
A Federal Reserve Board study shows targeted econometric models outperform machine learning when forecasting realized volatility. Analyzing the S&P 500 and 40 equities, the research finds regime-switching and long-memory specifications consistently deliver lower out-of-sample forecast errors.
Regimes and memory dictate horizons
Forecast accuracy follows a strict horizon-dependent structure across the 1996–2025 S&P 500 sample and 40 individual equities.
At the one-day horizon, Markov-switching heterogeneous autoregression (MSHAR) achieves the lowest mean squared forecast error of 0.051, compared with 0.093 for standard HAR and 0.102 for XGBoost.
MSHAR wins up to 100 percent of one-day cross-sectional comparisons across equities.
At the 22-day horizon, the fractionally integrated ARFIMA model dominates with an error of 0.104, outperforming neural networks that cluster above 0.260. The five-day horizon forms an intermediate transition where MSHAR and ARFIMA compete closely.
Machine-learning architectures occasionally beat basic linear HAR benchmarks, but they fail to displace the broader econometric frontier.
Information sets fail to rescue neural nets
Expanding the information set through Elastic Net screening of macroeconomic variables, VIX, and news sentiment does not alter model rankings.
While auxiliary predictors add marginal value in specific subperiods, they often introduce estimation noise rather than structural signal.
The findings remain robust under quarterly retraining, log-volatility transformations, and noise-robust realized-kernel measures.
Tail-risk evaluations using Value-at-Risk coverage and joint expected-shortfall loss further confirm that latent regime switching and fractional integration drive predictive gains.
Domain physics beats brute computation
Domain-specific econometric structure easily outperforms brute computational power in volatile financial series.
Generic machine learning fails because raw flexibility cannot replicate targeted regime shifts and long memory.
Risk managers should rely on disciplined parametric models rather than complex neural networks.