« Back to Results

Large Language Models in Macroeconomics

Paper Session

Monday, Jan. 4, 2027 2:30 PM - 4:30 PM (EST)

Marriott Marquis Washington DC
Hosted By: American Economic Association
  • Chair: Tatevik Sekhposyan, Texas A&M University

FOMC in Silico: A Multi-Agent System for Monetary Policy Decision Making

Sophia Kazinnik
,
Stanford University
Tara Sinclair
,
George Washington University

Abstract

We develop an LLM multi-agent simulation of the Federal Open Market Committee that models policy deliberation, disagreement, and consensus formation from member specific priors grounded in real time data. The framework supports counterfactual analysis of how policy outcomes change under alternative economic conditions, political pressure, and revisions to incoming information. Across the applications we study, political pressure increases dissent and dispersion, while weaker labor market data shifts outcomes modestly in a more dovish direction. We validate the framework against the historical record: in a conditional backtest over 215 FOMC meetings from 2000 to 2026, it predicts the policy rate within 25 basis points in 93% of meetings, with a mean absolute error of 9.1 basis points. In a longer horizon trajectory exercise, it tracks the broad evolution of the federal funds rate more closely than Taylor rule and other rule based benchmarks. Comparison to a fixed voting benchmark helps identify how deliberation and institutional frictions shape policy outcomes, creating an in silico framework for monetary policy experimentation.

Information Leakage in Large Language Models: A Cautionary Tale for Economic Surveillance

Wendy Dunn
,
Federal Reserve Board
Ellen Meade
,
Duke University
Nitish Sinha
,
Federal Reserve Board
Raakin Kabir
,
George Mason University

Abstract

In this study we use Federal Open Market Committee (FOMC) meeting minutes to assess the performance of large language models (LLMs) in classifying topics in economic text. We create a unique, expert-labeled dataset to use as our ground-truth benchmark for evaluating LLM performance. We find the models perform extremely well classifying topics of pre-2025 text, but when assigned truly novel data, they exhibit substantial performance degradation. Our results suggest that information leakage partly explains the impressive performance of LLMs, and thus caution should be used when relying on the predictions of these models using text-as-data in real time.

Inflation Attitudes of Large Language Models

Nikoleta Anesti
,
Bank of England
Edward Hill
,
Bank of England
Andreas Joseph
,
Bank of England

Abstract

This paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behavior of LLMs for use in the social sciences, to compare different models, or to assist in survey design.

ChatMacro: Evaluating Inflation Forecasts of Generative AI

Jahangir Alam
,
Wilfrid Laurier University
Shane Boyle
,
Federal Reserve Bank of San Francisco
Huiyu Li
,
Federal Reserve Bank of San Francisco
Tatevik Sekhposyan
,
Texas A&M University

Abstract

Recent research suggests that generic large language models (LLMs) can match the accuracy of traditional methods when forecasting macroeconomic variables in pseudo out-of-sample settings generated via prompts. This paper assesses the out-of-sample forecasting accuracy of LLMs by eliciting real-time forecasts of U.S. inflation from ChatGPT. We find that out-of-sample predictions are largely inaccurate and stale, even though forecasts generated in pseudo out-of-sample environments are comparable to existing benchmarks. Our results underscore the importance of out-of-sample benchmarking for LLM predictions.

Discussant(s)
Rubén Fernández-Fuertes
,
HEC Montreal
Pablo Guerron
,
Boston College
Olena Kostyshyna
,
Bank of Canada
Jane Ryngaert
,
University of Notre Dame
JEL Classifications
  • E3 - Prices, Business Fluctuations, and Cycles
  • C4 - Econometric and Statistical Methods: Special Topics