Large Language Models in Macroeconomics
Paper Session
Monday, Jan. 4, 2027 2:30 PM - 4:30 PM (EST)
- Chair: Tatevik Sekhposyan, Texas A&M University
Information Leakage in Large Language Models: A Cautionary Tale for Economic Surveillance
Abstract
In this study we use Federal Open Market Committee (FOMC) meeting minutes to assess the performance of large language models (LLMs) in classifying topics in economic text. We create a unique, expert-labeled dataset to use as our ground-truth benchmark for evaluating LLM performance. We find the models perform extremely well classifying topics of pre-2025 text, but when assigned truly novel data, they exhibit substantial performance degradation. Our results suggest that information leakage partly explains the impressive performance of LLMs, and thus caution should be used when relying on the predictions of these models using text-as-data in real time.Inflation Attitudes of Large Language Models
Abstract
This paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behavior of LLMs for use in the social sciences, to compare different models, or to assist in survey design.ChatMacro: Evaluating Inflation Forecasts of Generative AI
Abstract
Recent research suggests that generic large language models (LLMs) can match the accuracy of traditional methods when forecasting macroeconomic variables in pseudo out-of-sample settings generated via prompts. This paper assesses the out-of-sample forecasting accuracy of LLMs by eliciting real-time forecasts of U.S. inflation from ChatGPT. We find that out-of-sample predictions are largely inaccurate and stale, even though forecasts generated in pseudo out-of-sample environments are comparable to existing benchmarks. Our results underscore the importance of out-of-sample benchmarking for LLM predictions.Discussant(s)
Rubén Fernández-Fuertes
,
HEC Montreal
Pablo Guerron
,
Boston College
Olena Kostyshyna
,
Bank of Canada
Jane Ryngaert
,
University of Notre Dame
JEL Classifications
- E3 - Prices, Business Fluctuations, and Cycles
- C4 - Econometric and Statistical Methods: Special Topics