« Back to Results

New AI Tools for Economic Science

Paper Session

Sunday, Jan. 3, 2027 10:15 AM - 12:15 PM (EST)

Marriott Marquis Washington DC
Hosted By: American Economic Association
  • Chair: Alex Imas, University of Chicago

Replicating Experimental Economics with LLM Simulations

Jacob Snyder
,
NBER
Benjamin Manning
,
Massachusetts Institute of Technology
John Horton
,
Massachusetts Institute of Technology
Abhishek Nagaraj
,
University of California-Berkeley

Abstract

We develop an automated pipeline that reads experimental economics papers, extracts their design, and runs simulated replications using LLM agents as participants. The pipeline operates in five meta-stages: ingesting a paper’s text, generating a structured summary, converting the experiment into deterministic machine-executable artifacts, running the simulation with LLM-based participants, and producing a simulated dataset. Two post-pipeline verification stages then assess whether the simulation design is faithful to the original paper and whether key findings replicate in direction and magnitude. A core contribution is a schema that decomposes experiments into explicit, inspectable artifacts that allow the simulation to run deterministically while confining LLM non-determinism to participant responses. We apply the pipeline to a corpus of survey- and vignette-style experimental economics papers, using automated fidelity verification to confirm that simulated designs faithfully match original experiments before comparing outcomes. Among validated replications, we observe meaningful variation in how closely LLM-simulated results match original findings. The framework offers a scalable approach to replication, exploration of design-space sensitivity, and hypothesis generation in experimental economics.

Lean-ing into Formal Theory: Toward EconLib and Automated Proving in Economics

Ruize Chen
,
Axiom Math
Bo Cowgill
,
University of Toronto
Piotr Dworczak
,
Northwestern University
Ben Eltschig
,
Axiom Math
Nikhil Garg
,
Cornell University
Ziv Hellman
,
Bar-Ilan University
Carina Hong
,
Axiom Math
Scott Kominers
,
Harvard University
Moran Koren
,
Ben-Gurion University
Daniel Lyng
,
Princeton University
Ken Ono
,
Axiom Math
Jujian Zhang
,
Axiom Math

Abstract

Much of modern economic theory rests on intricate mathematical arguments, and as generative AI makes it increasingly easy to produce and assist with economic proofs, the need for rigorous, machine-checkable verification becomes more acute. We present a project that begins to address this verification gap by formalizing core results in auction theory and game theory within the Lean proof assistant, building on the infrastructure of Mathlib. We describe the formalization of key theorems—including Nash equilibrium existence, revenue equivalence, and incentive compatibility results—and discuss the design decisions that arise when translating economically natural objects such as type spaces, strategies, and mechanisms into dependent type theory. The broader aim is to seed a library of formally verified economic theory analogous to Mathlib’s role in mathematics. We outline our library's architecture and demonstrate its usefulness for verifying existing results and proving new ones. We then discuss what formalization reveals about the implicit assumptions in textbook economics arguments, where current Mathlib coverage creates bottlenecks, and how a shared formal library could change the way economic theorists communicate, verify, and build on one another’s work.
NOTE: this paper has six authors but the form only allows five, so we have combined to authors in the fifth slot

Human or Machine? Assessing AI’s Ability to Generate Game-Theory Questions

Benjamin Golub
,
Northwestern University
Annie Liang
,
Northwestern University
Marciano Siniscalchi
,
Northwestern University

Abstract

AI models now excel at solving difficult applied mathematics problems; we ask how well they can compose such problems, focusing on undergraduate game theory. Adapting the Turing test to problem generation, we collect problems from professors and GPT-5, standardizing presentation so evaluation focuses on content rather than style. Sixty-seven experts—undergraduate and graduate students who have taken game theory—classify problems as human- or LLM-generated. We find that AI output is indistinguishable to any single evaluator yet different in aggregate. Individually, evaluators perform at chance (mean accuracy 50.9%). However, pooling 2,680 classifications rejects the null that the two distributions are identical. The signal resides in solutions, not problem statements: restricting to evaluators who observe solutions and report medium or high confidence raises pooled accuracy, while without solutions we cannot reject the null. We train a classifier to distinguish the sources; the strongest objective feature separating human problems is the ratio of solution word count to problem word count. Human-authored problems tend to require more reasoning per unit of setup. We discuss implications for organizations that delegate knowledge work to AI.

GPT as a Measurement Tool

Hemanth Asirvatham
,
OpenAI
Elliott Mokski
,
Independent Researcher
Andrei Shleifer
,
Harvard University

Abstract

We present the GABRIEL software package, which uses GPT to quantify attributes in qualitative data, such as how “pro-innovation” a speech is. GPT is evaluated on classification and attribute-rating performance against more than 1,000 human-annotated tasks across a range of topics and data. We find that GPT as a measurement tool is accurate across domains and generally indistinguishable from human evaluators. Our evidence indicates that labeling results do not depend on the exact prompting strategy used, and that GPT is not relying on training data contamination or inferring attributes from other attributes. We showcase the possibilities of GABRIEL by quantifying novel and granular trends in congressional remarks, social media toxicity, and county-level school curricula. We then apply GABRIEL to study the history of technology adoption, assembling a novel dataset of 37,000 technologies. Our analysis documents a tenfold decline in lags from invention to adoption over the industrial age, from roughly 50 years to roughly 5 years today. We also quantify the increasing dominance of companies and the United States in innovation, alongside characteristics associated with faster or slower adoption.

Discussant(s)
Joshua Gans
,
University of Toronto
Alex Imas
,
University of Chicago
JEL Classifications
  • C0 - General
  • B4 - Economic Methodology