Overarching Research Themes
Themes extracted and images generated with the OpenAI API; there may be inconsistencies.
Pragmatic Social Reasoning

My research group explores how to evaluate and improve the social intelligence of AI systems in realistic interaction settings. A recurring focus is whether models can handle misunderstanding, non-literal intent, and theory-of-mind style reasoning rather than only producing fluent replies. Recent work such as [XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?](https://arxiv.org/abs/2502.14860) and [Cognitive Chain-of-Thought: Structured Multimodal Reasoning about Social Situations](https://arxiv.org/abs/2507.20409) pushes beyond static benchmarks toward richer social-pragmatic evaluation. We also see this theme in [SOTOPIA-ToM: Evaluating Information Management in Multi-Agent Interaction with Theory of Mind](https://arxiv.org/abs/2605.02307), which examines how models manage beliefs and information in interactive settings. Together, these papers suggest a shift from surface-level social fluency to grounded assessment of how well AI can reason about people, context, and conversational goals.
Safe Agentic and Human-AI Use

My research group explores new measures for agentic AI safety and user-centric safety risks, including reliance, manipulation, truthfulness, and the consequences of guardrails. A major thread is understanding how AI systems behave when they act as agents or shape user beliefs, rather than merely answering questions. For example, [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://arxiv.org/abs/2507.06134) broadens safety evaluation for deployed agents, while [Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance](https://aclanthology.org/2025.naacl-long.556/) studies when people depend on model outputs in ways that may be harmful. Relatedly, [The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues](https://arxiv.org/abs/2603.20907) and [AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents](https://aclanthology.org/2025.naacl-long.595/) highlight concerns about persuasion and truthfulness in agentic settings. Overall, this line of work is building practical ways to measure not just whether AI is capable, but whether it is safe to trust, safe to follow, and safe to deploy.
Culturally Adapted Responsible AI

My research group explores how to make AI systems culturally competent and adaptable while also avoiding harms caused by poor personalization or neglect of minority language practices. A central question is how models should respond differently across cultures, dialects, and socially meaningful norms without reinforcing bias or erasure. Important recent work includes [NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models](https://aclanthology.org/2025.naacl-long.120/), which formalizes cultural adaptability as an evaluative target, and [CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries](https://arxiv.org/abs/2607.05405), which tests whether models can infer culturally grounded expectations from subtle cues. The broader bias and language-variation dimension is also visible in [Rejected Dialects: Biases Against African American Language in Reward Models](https://arxiv.org/abs/2502.12858), showing how training and ranking systems can disadvantage minoritized varieties. Taken together, these papers point to responsible AI that is not only multilingual, but culturally aware, context-sensitive, and fair in how it treats diverse users and communities.
Storytelling and Social Connection

My research group explores how AI can support human-human connection by understanding stories, narratives, and the social meanings people attach to them. The key idea is that stories are not just text to summarize; they encode identity, empathy, conflict, and perspective-taking. Recent work such as [Social Story Frames: Contextual Reasoning about Narrative Intent and Reception](https://arxiv.org/abs/2512.15925) studies how narratives are interpreted socially, while [HEART-felt Narratives: Tracing Empathy and Narrative Style in Personal Stories with LLMs](https://arxiv.org/abs/2405.17633) investigates empathy and style in personal storytelling. Similarly, [Modeling Empathic Similarity in Personal Narratives](https://arxiv.org/abs/2305.14246) suggests that story understanding can be used to model interpersonal resonance, not just semantic similarity. This research direction is helping AI move toward better interpretation of lived experience, relational cues, and the connective power of stories in everyday communication.