An Open Benchmark for Evaluating Time Series Forecasting Methods across Financial Markets

Two-panel line chart comparing realized values with 12-step forecasts from ARIMA, Theta, NLinear, and N-BEATS. In panel (a), forecasts of the Treasury spot-futures basis diverge widely from one another and from the realized path; in panel (b), forecasts of bank cash liquidity cluster tightly around the realized series.

Abstract

Accurate time series forecasts underpin asset pricing, risk management, monetary and macroprudential policy, and other applications. The set of available forecasting methods is expanding rapidly, driven by new machine learning models. This raises a practical question: Do these methods deliver real forecasting power gains on financial data? Since no single method is best across all data generating processes, the question can be answered only by direct evaluation on domain-specific data, and a fair comparison requires holding the data fixed across methods. Yet, new methods are rarely evaluated on the financial data commonly used in academic finance; when they are, each is assessed on its own dataset and cleaning conventions. Apparent method rankings entangle method skill with cleaning choices that themselves require domain expertise. We address this by assembling, in one place, canonical cleaning procedures for many financial datasets to hold the data fixed across forecasting experiments. We introduce a standardized open-source dataset covering equities, corporate bonds, U.S. Treasuries, foreign exchange, commodities, credit default swaps, options, five basis spread datasets, and bank and intermediary indicators, each cleaned per the canonical paper for that asset class. Holding the data fixed and evaluating roughly a dozen univariate methods without exogenous regressors, we find that asset returns remain near unforecastable across every method family and that hybrid and machine learning methods exhibit additional forecasting power on basis spreads and bank indicators.

View related blog

Keywords: time series forecasting; machine learning; forecast evaluation; benchmark datasets; return predictability; basis spreads; financial intermediaries; financial stability; neural networks; open-source data

JEL Codes: C53, G17, C58, C45, C55, C81, E44