TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
Details
Knowing when to speak, listen, or yield: evaluating turn-taking across six styles of natural conversation.
I’m Satya. I train AI models to be useful. I’m currently a researcher at Sesame.
Knowing when to speak, listen, or yield: evaluating turn-taking across six styles of natural conversation.
Evaluating factuality, retrieval, and reasoning together, with questions that require information from multiple sources.
Detecting the gap between apparently benign answers and deceptive reasoning in language models.
Understanding model predictions through a conversation: explanations, data analysis, and follow-up questions in natural language.
Turning model explanations into natural language rationales that help language models correct their mistakes.
Finding safety failures shared by a language model and its reward model, then repairing both.
Studying how removing refusals in one domain can weaken safety in other domains.
An automatically validated benchmark measuring the security of LLM-generated code across software weakness categories.
Recovering reward functions from aligned language models to investigate what their training objectives encode.
Measuring how explanation methods disagree, and how practitioners navigate conflicting explanations.
Studying when repeated prompting reduces truthfulness and calibration, and how prompting strategies can help.
Tools for comparing the faithfulness, stability, and fairness of model explanations.
A dataset and evaluation metrics for measuring social biases in open-ended language generation.
Exploring text transformations designed to protect privacy while retaining useful semantic information.
Knowing when to speak, listen, or yield: evaluating turn-taking across six styles of natural conversation.
Evaluating factuality, retrieval, and reasoning together, with questions that require information from multiple sources.
Knowing when to speak, listen, or yield: evaluating turn-taking across six styles of natural conversation.
Evaluating factuality, retrieval, and reasoning together, with questions that require information from multiple sources.
Detecting the gap between apparently benign answers and deceptive reasoning in language models.
Understanding model predictions through a conversation: explanations, data analysis, and follow-up questions in natural language.
Turning model explanations into natural language rationales that help language models correct their mistakes.
Finding safety failures shared by a language model and its reward model, then repairing both.
Studying how removing refusals in one domain can weaken safety in other domains.
An automatically validated benchmark measuring the security of LLM-generated code across software weakness categories.
Recovering reward functions from aligned language models to investigate what their training objectives encode.
Measuring how explanation methods disagree, and how practitioners navigate conflicting explanations.
Studying when repeated prompting reduces truthfulness and calibration, and how prompting strategies can help.
Tools for comparing the faithfulness, stability, and fairness of model explanations.
A dataset and evaluation metrics for measuring social biases in open-ended language generation.
Exploring text transformations designed to protect privacy while retaining useful semantic information.
I completed my PhD at Harvard’s School of Engineering and Applied Sciences, advised by Hima Lakkaraju and Finale Doshi-Velez. My research focused on understanding and improving the trustworthy aspects of generative models. I also collaborated with Sameer Singh at UC Irvine and Steven Wu at Carnegie Mellon.
Previously, I worked at Google DeepMind and Meta FAIR, and led research initiatives across AWS, Alexa, and Amazon Search. At Carnegie Mellon, I studied conversational agents and reinforcement learning.
School of Engineering and Applied Sciences
School of Computer Science
Doctoral thesis: From Understanding to Improving Artificial Intelligence: New Frontiers in Machine Learning Explanations.
View my CVI’ve joined Sesame, where I am building the best voice model.
Introducing TurnBench: a benchmark, dataset, and leaderboard for turn-taking in spoken dialogue.
Released D-REX, a benchmark for detecting deceptive reasoning in language models.
Harvard PhD thesis: From Understanding to Improving Artificial Intelligence: New Frontiers in Machine Learning Explanations.
Introduced FRAMES, a benchmark for factuality, retrieval, and reasoning. Published at NAACL 2025.
Teaching Fellow for COMPSCI 236R: Topics at the Interface between Computer Science and Economics.
Amazon Science highlighted my contributions to trustworthy machine learning at Alexa AI.
Presented The Disagreement Problem at the TRAIT Workshop at CHI. Two papers on algorithmic fairness accepted at ACL: paper one and paper two.
Released The Disagreement Problem, later published in TMLR. Covered by Fortune and Analytics India Magazine. Received the Nicole A. Chen and Karina A. Chen Graduate Student Research Fellowship.
Joined Harvard SEAS as a Computer Science PhD student.
Does Robustness Improve Fairness? accepted at Findings of ACL 2021.
VentureBeat covered BOLD, our benchmark for measuring bias in open-ended language generation.
ADePT, our work on private text transformation, accepted at EACL 2021.
Announced BOLD, a fairness benchmark for open-ended language generation, presented at FAccT 2021.