Calibrated noise cuts disclosure risk in synthetic firm panel data
BDI Paper

Calibrated noise cuts disclosure risk in synthetic firm panel data

A Banca d'Italia study establishes a framework to generate synthetic unbalanced panel microdata from 13,837 Italian firms observed between 1993 and 2024. Sequential tree models paired with calibrated measurement error reproduce firm dynamics while safeguarding confidential records.

Trees outperform copulas across 100 specifications

Researchers Marco Langiulli, Alessandro Moro, Mattia Orzincolo, and Daniele Piras tested synthetic data generators on the INVIND panel spanning 1993 to 2024.

To account for heterogeneous histories, the dataset was partitioned at a ten-year threshold, placing 31 percent of firms and 65 percent of observations into the longitudinal component.

Across 20 quantitative and qualitative variables—including turnover, investment, and forward-looking expectations—sequential tree models based on CART outperformed Gaussian copulas and naive generators.

In 100 randomized regressions evaluating 3,500 coefficients, the tree-based method centered standardized errors at zero, with more than 50 percent of estimates falling within two standard errors of the original microdata.

Calibrating noise against nearest-neighbor matching

High distributional fidelity creates disclosure vulnerabilities, as synthetic records lie close to real observations.

Evaluated through the Nearest Neighbor Distance Ratio under Euclidean and weighted Gower metrics, tree generators initially exhibited elevated re-identification risks.

To resolve this tension, the authors introduced controlled measurement error with a tuning parameter of 0.20, raising series variance by roughly 2 percent.

This perturbation elevated distance ratios to secure benchmark levels while preserving underlying empirical structures.

A viable bridge for research labs

The paper delivers a viable blueprint for research data centers managing confidential firm panels.

Yet retaining fixed sector and regional codes leaves prominent enterprise outliers vulnerable to residual re-identification.

Controlled noise proves that privacy protections do not have to destroy empirical utility.

Report an error