Maarten Sap

I am an assistant professor at CMU's LTI department with a courtesy appointment in HCII, and a part-time senior research scientist and technical AI safety lead at the Allen Institute for AI (AI2). My research focuses on (1) measuring and improving AI systems' social and interactional intelligence, (2) assessing and combatting social inequality, safety risks, and socio-cultural biases in human- or AI-generated language, and (3) building narrative language technologies for prosocial outcomes. I was named a 2025 Packard Fellow and a recipient of the 2025 Okawa Research Award.

I received my PhD from the University of Washington where I was advised by Noah Smith and Yejin Choi.
[bio for talks]

Recent updates:

August 2025 πŸŽ“πŸ“œ: Super proud of the first CMU Sapling and one of my first solo advisees, Xuhui Zhou, for successfully defending his PhD thesis! Huge congrats Xuhui!!

August 2025 πŸ†πŸ“ƒ: Very honored that our paper "I Just Don't Want My Work Being Fed Into The AI Blender'': Queer Artists on Refusing and Resisting Generative AI got an Honorable Mention Award at CSCW 2026! Major congrats to the first author Jordan Taylor!!

May 2025 πŸŽ“πŸ“ƒ: The first MIT Sapling, Jocelyn Shen, successfully defended her PhD thesis! Huge congrats Jocelyn!!

December 2025 πŸ…πŸ“ƒ: Very excited to have our paper Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) selected for a Best Paper Award at NeurIPS 2025 (Datasets and Benchmarks Track)!! Huge congrats to the first author Liwei Jiang!!!

November 2025 πŸ’ŽπŸš€: Honored to be a Spring 2025 recipient of the Amazon Research Award for our project on measuring AI agentic safety!

October 2025 πŸ…β­: I’m super excited and grateful to announce that I'm part of the 2025 class of Packard Fellows. The Packard Foundation and this fellowship will allow me to explore exciting research directions towards culturally responsible and safe AI 🌍🌈

October 2025 πŸ”πŸ§‘β€πŸŽ“: Due to my lab being quite full already, I'm not taking looking for any new students in this upcoming PhD application cycle 😟.

[older news]


Overarching Research Themes

Themes extracted and images generated with the OpenAI API; there may be inconsistencies.

Social Pragmatics and Theory of Mind

My research group explores how to evaluate and improve the social intelligence of AI systems, with particular attention to pragmatics, intent understanding, and theory-of-mind behavior in realistic interaction settings. Recent work such as [XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?](https://arxiv.org/abs/2609.06842) and [SOTOPIA-ToM: Evaluating Privacy and Information Management in Multi-Agent Interaction with Theory of Mind](https://arxiv.org/abs/2605.02307) shows growing interest in whether models can reason about what people mean, not just what they say. [Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments](https://arxiv.org/abs/2608.09128) further reflects a shift toward interactive, competitive settings that stress-test social reasoning under pressure. Together, these papers suggest the field is moving from static benchmarks toward dynamic evaluations of conversation, coordination, and social inference.

Agentic Safety and Human Reliance

My research group explores novel ways to measure and improve the safety of agentic AI systems, while also studying user-centric risks such as over-reliance, manipulation, and harmful interactions. [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://arxiv.org/abs/2507.06134) and [AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents](https://aclanthology.org/2025.naacl-long.595/) highlight concern with how autonomous systems behave when acting on behalf of users. On the human side, [Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance](https://aclanthology.org/2025.naacl-long.556/) and [The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues](https://arxiv.org/abs/2603.20907) focus on when AI influences users too much, or in the wrong direction. This theme points to an emerging safety agenda that treats both agent behavior and human response as part of the same risk surface.

Culturally Adaptive and Fair AI

My research group explores how to build AI systems that adapt responsibly across cultures, dialects, and value systems without amplifying bias or exclusion. [NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models](https://aclanthology.org/2025.naacl-long.120/) and [CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries](https://arxiv.org/abs/2607.05405) point toward more systematic evaluation of cultural competence in everyday contexts. [Black LLMirror: User (Self) Perceptions in Black American English Interactions with LLMs](https://dl.acm.org/doi/abs/10.1145/3772318.3791111) and [Rejected Dialects: Biases Against African American Language in Reward Models](https://arxiv.org/abs/2502.12858) show how lack of cultural adaptation can translate into fairness harms for minoritized language communities. Overall, the work suggests a shift from treating personalization as a convenience feature to treating it as a core requirement for equitable AI.

Narrative Understanding for Connection

My research group explores how AI can support human-human connection by understanding stories, personal narratives, and the social meanings people attach to them. [Social Story Frames: Contextual Reasoning about Narrative Intent and Reception](https://arxiv.org/abs/2512.15925) and [HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs](https://arxiv.org/abs/2405.17633) show how models are being used to analyze not just content, but how stories are told and received. [Modeling Empathic Similarity in Personal Narratives](https://arxiv.org/abs/2305.14246) extends this direction by asking how narrative structure can capture empathy between people. Taken together, these papers suggest a growing effort to make AI better at interpreting stories in ways that preserve nuance, context, and interpersonal meaning.