“What I cannot create, I do not understand.”— Richard Feynman

I’m Satya. I train AI models to be useful. I’m currently a researcher at Sesame.

Selected work

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

Preprint · 2026Dataset & leaderboard
Details

Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe

Knowing when to speak, listen, or yield: evaluating turn-taking across six styles of natural conversation.

All papers 14

TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue

Preprint · 2026Dataset & leaderboard
Details

Freeman Jiang, Ramon Sanabria, Soham Deshmukh, Bandhav Veluri, Simon Michael Vuch Williams, Elliott K. Suen, Garreth Lee, Kevin Yoonho Choi, Takuya Umeki, Riku Kubo, Sathvik Udupa, Chien-yu Huang, Shih-Yun Shan Kuan, Zhuoyan Tao, Satyapriya Krishna, Sefik Emre Eskimez, Yu Tsao, Hung-yi Lee, Shinji Watanabe

Knowing when to speak, listen, or yield: evaluating turn-taking across six styles of natural conversation.

AutoSUIT Bench - Automated Security UnIt Test Benchmark for LLM Coding

Findings of ACL 2026
Details

Samuel Osebe, Fan Yang, Junyi Li, Yue Gu, Yongxin Wang, Satyapriya Krishna, Kai-Wei Chang, Aram Galstyan, Rahul Gupta, Weitong Ruan

An automatically validated benchmark measuring the security of LLM-generated code across software weakness categories.

Background

I completed my PhD at Harvard’s School of Engineering and Applied Sciences, advised by Hima Lakkaraju and Finale Doshi-Velez. My research focused on understanding and improving the trustworthy aspects of generative models. I also collaborated with Sameer Singh at UC Irvine and Steven Wu at Carnegie Mellon.

Previously, I worked at Google DeepMind and Meta FAIR, and led research initiatives across AWS, Alexa, and Amazon Search. At Carnegie Mellon, I studied conversational agents and reinforcement learning.

PhD · Computer ScienceHarvard University

School of Engineering and Applied Sciences

Master’s degreeCarnegie Mellon University

School of Computer Science

Doctoral thesis: From Understanding to Improving Artificial Intelligence: New Frontiers in Machine Learning Explanations.

View my CV
Academic service & community
  • Co-President · Harvard SEAS Graduate Council, 2022
  • Program Committee · ICML Workshop on Human-Machine Collaboration and Teaming, 2022
  • Reviewer · NeurIPS Datasets and Benchmarks, 2021–2022; NeurIPS Main Conference, 2022; AAAI, 2021
  • Organizing Committee · Workshop on Trustworthy NLP at NAACL, 2021
  • Student Volunteer Award · ACL, 2022
  • Volunteer · ACM FAccT, 2022
News

I’ve joined Sesame, where I am building the best voice model.

Introducing TurnBench: a benchmark, dataset, and leaderboard for turn-taking in spoken dialogue.

Released D-REX, a benchmark for detecting deceptive reasoning in language models.

Harvard PhD thesis: From Understanding to Improving Artificial Intelligence: New Frontiers in Machine Learning Explanations.

Introduced FRAMES, a benchmark for factuality, retrieval, and reasoning. Published at NAACL 2025.

Earlier news

Teaching Fellow for COMPSCI 236R: Topics at the Interface between Computer Science and Economics.

Amazon Science highlighted my contributions to trustworthy machine learning at Alexa AI.

Presented The Disagreement Problem at the TRAIT Workshop at CHI. Two papers on algorithmic fairness accepted at ACL: paper one and paper two.

Released The Disagreement Problem, later published in TMLR. Covered by Fortune and Analytics India Magazine. Received the Nicole A. Chen and Karina A. Chen Graduate Student Research Fellowship.

Joined Harvard SEAS as a Computer Science PhD student.

Does Robustness Improve Fairness? accepted at Findings of ACL 2021.

VentureBeat covered BOLD, our benchmark for measuring bias in open-ended language generation.

ADePT, our work on private text transformation, accepted at EACL 2021.

Announced BOLD, a fairness benchmark for open-ended language generation, presented at FAccT 2021.