ForeComp package controls size distortion in forecast comparisons
BIS Paper

ForeComp package controls size distortion in forecast comparisons

Researchers from the Federal Reserve Bank of Philadelphia and Baltimore Orioles introduced ForeComp, an R package designed to correct small-sample size distortions in Diebold-Mariano forecast evaluation tests using fixed-smoothing asymptotics.

Tackling long-run variance noise

Standard Diebold-Mariano tests often over-reject the null hypothesis of equal predictive ability when evaluation samples are small.

This failure stems from estimation error in the long-run variance of loss differentials.

The ForeComp package addresses this by incorporating fixed-smoothing asymptotics, including fixed-b Bartlett, equal-weighted cosine (DM-EWC), weighted periodogram (DM-WPE), and Ibragimov-Müller block t-tests (DM-IM).

It also provides Plot Tradeoff, a diagnostic tool that visualizes size distortion and power loss across bandwidth parameters.

Applications using Survey of Professional Forecasters data show that standard normal tests reject equal accuracy, whereas fixed-smoothing methods maintain correct size.

Simulation evidence favors fixed smoothing

Monte Carlo simulations across 5,000 replications highlight the finite-sample advantages of fixed-smoothing procedures.

Under an unconditional-rolling data-generating process with an evaluation sample of P=75 and horizon h=12, the standard DM rectangular test registers an empirical size of 0.160 at a nominal 5 percent level.

Newey-West test size reaches 0.127 under the same conditions.

In contrast, fixed-smoothing methods such as DM-FB and DM-EWC achieve empirical rejection rates of 0.052 and 0.045, maintaining size control without sacrificing size-corrected power.

Essential tooling for empirical rigour

ForeComp consolidates fragmented econometric tests into a unified R toolkit for forecast comparison.

By proving that standard tests generate spurious rejections in small samples, the paper exposes key vulnerabilities in empirical literature.

Researchers should adopt fixed-smoothing diagnostics to ensure robust policy analysis.