Back to 2026 Programme

Aligning Large Language Model Agents with Rational and Moral Preferences

Fontainebleau, France 13 July 2026 – 15 July 2026

Daniel Chen (TSE, IAST, CNRS); Wei Lu (CUNY); Amit Dhanda (Amazon); Chris Hansen (University of Chicago)

A1 AI, Markets, and Algorithmic Decisions
Chair: Sepehr Shahshahani
Amphi De Vitry
Economics / Other

Abstract

Understanding how large language model (LLM) agents behave in strategic interactions is essential as these systems increasingly participate autonomously in economically and morally consequential decisions in multi-agent systems. We first evaluate LLM agents’ strategic reasoning using canonical economic games, finding substantial deviations from human behavior. Models like GPT-4o show excessive cooperation and limited incentive sensitivity, while reasoning models, such as o3-mini, align more consistently with payoff-maximizing strategies. We propose a supervised fine-tuning pipeline that uses synthetic datasets derived from economic reasoning to align LLM agents with economic preferences, focusing on two stylized preference structures. In the first, utility depends only on individual payoffs (homo economicus), while utility also depends on a notion of Kantian universalizability in the second preference structure (homo moralis). We find that fine-tuning based on small datasets shifts LLM agent behavior toward the corresponding economic agent. We further assess the finetuned agents’ behavior in two applications: Moral dilemmas involving autonomous vehicles and algorithmic pricing in competitive markets. These examples illustrate how different normative objectives embedded via realizations from structured preference structures can influence market and moral outcomes. We further show that preference-aligned fine-tuning improves performance on standard AI safety benchmarks, such as bias, jailbreak robustness, and overrefusal, while leaving short-form factual accuracy largely unchanged, suggesting that theory-driven alignment can enhance safety properties without degrading core capabilities. This work contributes a replicable, cost-efficient, and interpretable method for shaping AI behavior in strategic, multi-stakeholder environments.

Download paper (PDF)