Maarten Sap

I am an assistant professor at CMU's LTI department with a courtesy appointment in HCII, and a part-time research scientist and AI safety lead at the Allen Institute for AI (AI2). My research focuses on (1) measuring and improving AI systems' social and interactional intelligence, (2) assessing and combatting social inequality, safety risks, and socio-cultural biases in human- or AI-generated language, and (3) building narrative language technologies for prosocial outcomes. I was named a 2025 Packard Fellow and a recipient of the 2025 Okawa Research Award.

I received my PhD from the University of Washington where I was advised by Noah Smith and Yejin Choi.
[bio for talks]

Recent updates:

August 2025 πŸŽ“πŸ“œ: Super proud of the first CMU Sapling and one of my first solo advisees, Xuhui Zhou, for successfully defending his PhD thesis! Huge congrats Xuhui!!

August 2025 πŸ†πŸ“ƒ: Very honored that our paper "I Just Don't Want My Work Being Fed Into The AI Blender'': Queer Artists on Refusing and Resisting Generative AI got an Honorable Mention Award at CSCW 2026! Major congrats to the first author Jordan Taylor!!

May 2025 πŸŽ“πŸ“ƒ: The first MIT Sapling, Jocelyn Shen, successfully defended her PhD thesis! Huge congrats Jocelyn!!

December 2025 πŸ…πŸ“ƒ: Very excited to have our paper Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) selected for a Best Paper Award at NeurIPS 2025 (Datasets and Benchmarks Track)!! Huge congrats to the first author Liwei Jiang!!!

November 2025 πŸ’ŽπŸš€: Honored to be a Spring 2025 recipient of the Amazon Research Award for our project on measuring AI agentic safety!

October 2025 πŸ…β­: I’m super excited and grateful to announce that I'm part of the 2025 class of Packard Fellows. The Packard Foundation and this fellowship will allow me to explore exciting research directions towards culturally responsible and safe AI 🌍🌈

October 2025 πŸ”πŸ§‘β€πŸŽ“: Due to my lab being quite full already, I'm not taking looking for any new students in this upcoming PhD application cycle 😟.

[older news]


Overarching Research Themes

Themes extracted and images generated with the OpenAI API; there may be inconsistencies.

Pragmatic Social Intelligence

My research group explores how to evaluate and improve AI systems’ social intelligence, especially in pragmatic language use, theory of mind, and interaction in multi-party settings. Recent work such as [XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?](https://arxiv.org/abs/2510.21903) and [Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations](https://arxiv.org/abs/2507.20409) shows growing interest in whether models can reason about what people mean, not just what they say. Benchmarks like [SOTOPIA-ToM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind](https://arxiv.org/abs/2605.02307) and [Social World Models](https://arxiv.org/abs/2509.00559) push toward richer evaluations of social reasoning, information sharing, and anticipation of others’ beliefs. Overall, the field is moving beyond surface-level social fluency toward grounded, interactive measures of whether AI can β€œread the room” in realistic social contexts.

Agentic Safety and Reliance

My research group explores novel measures of agentic AI safety and user-centric risk, spanning unsafe autonomous behavior, harmful truthfulness trade-offs, and how people rely on or are influenced by AI. Important new benchmarks such as [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://arxiv.org/abs/2507.06134) and [PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm](https://arxiv.org/abs/2601.08951) reflect a shift toward measuring risks in realistic deployments rather than narrow lab settings. On the human side, [Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance](https://aclanthology.org/2025.naacl-long.556/) and [The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues](https://arxiv.org/abs/2603.20907) examine overreliance, persuasion, and manipulation, while [AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents](https://aclanthology.org/2025.naacl-long.595/) highlights the tension between usefulness and honesty. Together, these papers show a broadening safety agenda: not only making agents less dangerous, but also understanding how AI systems shape human decisions, trust, and autonomy.

Culturally Adaptive Responsible AI

My research group explores responsible AI that can adapt across cultures while remaining fair, respectful, and robust to linguistic and social variation. Work such as [NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models](https://aclanthology.org/2025.naacl-long.120/) and [CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries](https://arxiv.org/abs/2607.05405) focuses on whether models can interpret culturally embedded norms rather than defaulting to a single assumed worldview. At the same time, [Rejected Dialects: Biases Against African American Language in Reward Models](https://arxiv.org/abs/2502.12858) and [Out of Style: RAG's Fragility to Linguistic Variation](https://arxiv.org/abs/2504.08231) highlight how personalization and retrieval systems can fail for minority dialects and nonstandard language. This line of research is increasingly concerned with both adaptation and equity: ensuring systems are locally appropriate without encoding new biases or erasing marginalized ways of speaking.

Stories, Empathy, and Connection

My research group explores AI for human-human connection through story understanding, narrative interpretation, and empathy modeling. Papers like [Social Story Frames: Contextual Reasoning about Narrative Intent and Reception](https://arxiv.org/abs/2512.15925) and [HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs](https://arxiv.org/abs/2405.17633) investigate how narratives convey intent, emotion, and social meaning to readers. Complementing this, [Modeling Empathic Similarity in Personal Narratives](https://arxiv.org/abs/2305.14246) studies what makes stories feel emotionally aligned across people, while [How Do People Challenge Racial Stereotypes Online? Counter-Story Detection Across Reddit Communities](https://arxiv.org/abs/2403.00179) points to the role of counter-stories in social understanding and solidarity. Together, these works suggest a growing research direction where story-aware AI can better support empathy, interpretation, and connection between people.