About me

I am a Member of Technical Staff at FAR.AI. I was previously a postdoc at the Center for Human-Compatible AI (CHAI) at UC Berkeley, mentored by Stuart Russell. In 2021, I received my PhD from the Computer Science Department at Stanford University, advised by Ashish Goel and supported by an NSF Graduate Research Fellowship. I also spent two years doing science and product work at Lyft.

I want to use my career to do good in the world. During my PhD, I designed markets and algorithms that achieve fair outcomes even when people are selfish. As AI capabilities advanced, I became increasingly concerned about risks from AI and shifted my research to AI safety. I'm concerned about a wide range of risks from AI, including but not limited to loss of control, facilitating malicious actors (especially for cyberattacks and CBRN attacks), concentration of power, exacerbation of societal inequalities, and economic disruption. My research is primarily relevant to the first two.

Research interests

My current focus is evaluation awareness: when models recognize that they’re being tested. Current safety assessments rely heavily on behavioral evaluations, and if models behave differently when they know they’re being tested, these assessments could be systematically misleading. My goal is to determine the extent of this issue and its implications for AI safety. I'm more focused on measuring the impact of eval awareness on behavior, rather than the prevalence of eval awareness.

My previous work centered on generalization: how a model handles unfamiliar inputs. I think that many types of safety failures can be framed as misgeneralization. I especially focused on training models to recognize when they're in unfamiliar situations and behave cautiously if so (e.g., ask for help). I studied this topic in the context of LLMs, deep RL in video games, and theory. See this meta-paper for an overview of my generalization work. I also did some work on optimization pressure in post-training.

Publications

Lead author(s): * | Senior author(s): †

AI Safety

  1. Safe Learning Under Irreversible Dynamics via Asking for Help
    Benjamin Plaut*, Juan Liévano-Karim, Hanlin Zhu, Stuart Russell†.
    Journal of Machine Learning Research (JMLR), 2026.
  2. Learning When Not to Learn: Risk-Sensitive Abstention in Bandits with Unbounded Rewards
    Sarah Liaw*, Benjamin Plaut*.
    International Conference on Artificial Intelligence and Statistics (AISTATS), 2026.
  3. YRC-Bench: A Benchmark for Learning to Coordinate with Experts
    Mohamad H. Danesh*, Nguyen X. Khanh, Tu Trinh, Benjamin Plaut.
    Transactions on Machine Learning Research (TMLR), 2025.
  4. Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safety
    Vamshi Krishna Bonagiri*, Ponnurangam Kumaraguru, Nguyen X. Khanh, Benjamin Plaut†.
    Neural Information Processing Systems (NeurIPS) 2025 Workshops on Reliable and Regulatable ML.
  5. Avoiding Catastrophe in Online Learning by Asking for Help
    Benjamin Plaut*, Hanlin Zhu, Stuart Russell†.
    International Conference on Machine Learning (ICML), 2025.
  6. Probabilities of Chat LLMs Are Miscalibrated but Still Predict Correctness on Multiple-Choice Q&A
    Benjamin Plaut*, Nguyen X. Khanh, Tu Trinh.
    Transactions on Machine Learning Research (TMLR), 2025.
  7. Getting By Goal Misgeneralization With a Little Help From a Mentor
    Tu Trinh*, Mohamad Danesh, Nguyen X. Khanh, Benjamin Plaut†.
    Neural Information Processing Systems (NeurIPS) 2024 Workshop on Safe and Trustworthy Agents.
  8. Safety Training May Persist Through Helpfulness Optimization in LLM Agents
    Benjamin Plaut.
    Preprint.

Resource Allocation Algorithms

  1. Algorithms for Fair Public and Private Resource Allocation
    Benjamin Plaut. PhD dissertation, 2021.
  2. Counteracting Inequality in Markets via Convex Pricing
    Ashish Goel†, Benjamin Plaut*. Conference on Web and Internet Economics (WINE) 2020.
  3. Almost Envy-free Repeated Matching in Two-sided Markets
    Sreenivas Gollapudi†, Kostas Kollias, Benjamin Plaut*. Conference on Web and Internet Economics (WINE) 2020.
  4. Optimal Nash Equilibria for Bandwidth Allocation
    Benjamin Plaut*. Conference on Web and Internet Economics (WINE) 2020.
  5. Equality of Power and Fair Public Decision-making
    Nicole Immorlica†, Benjamin Plaut*, E. Glen Weyl†. Conference on Web and Internet Economics (WINE) 2019.
  6. Markets Beyond Nash Welfare for Leontief Utilities
    Ashish Goel†, Reyna Hulett, Benjamin Plaut*. Conference on Web and Internet Economics (WINE) 2019.
  7. Communication Complexity of Discrete Fair Division
    Benjamin Plaut*, Tim Roughgarden†. Symposium on Discrete Algorithms (SODA), 2019; SIAM Journal on Computing (SICOMP), 2020.
  8. Markets for Public Decision-making
    Nikhil Garg, Ashish Goel†, Benjamin Plaut*. Conference on Web and Internet Economics (WINE) 2018.
  9. Almost Envy-Freeness with General Valuations
    Benjamin Plaut*, Tim Roughgarden†. Symposium on Discrete Algorithms (SODA), 2018; SIAM Journal on Discrete Mathematics (SIDMA), 2020.
  10. Algorithms for Social Good: Kidney Exchange
    Benjamin Plaut. Undergraduate honors thesis, 2016. Won the Allen Newell Award for Excellence in Undergraduate Research (best thesis in Computer Science).
  11. Hardness of the Pricing Problem in Barter Exchanges
    Benjamin Plaut*, John P. Dickerson, Tuomas Sandholm†. Preprint.
  12. Position-Indexed Formulations for Kidney Exchange
    John P. Dickerson, David Manlove†, Benjamin Plaut, Tuomas Sandholm†, and John Trimble*. Economics and Computation (EC), 2016.
  13. Fast Optimal Clearing of Capped-Chain Barter Exchanges
    Benjamin Plaut*, John P. Dickerson, Tuomas Sandholm†. AAAI Conference on Artificial Intelligence, 2016.

Physical Chemistry

  1. Direct Observation of Folding Energy Landscape of RNA Hairpin at Mechanical Loading Rates
    Huizhong Xu*, Benjamin Plaut, Xiran Zhu, Maverick Chen, Udit Mavinkurve, Anindita Maiti, Guangtao Song, Krishna Murari, and Maumita Mandal†. The Journal of Physical Chemistry, 2017.

Work experience

Member of Technical Staff FAR.AI August 2026 – present
Postdoctoral Researcher UC Berkeley September 2023 – August 2026
Data Scientist Lyft June 2021 – May 2023
Research Intern Google Summer 2019, Summer 2020

Research mentees

  • Sarah Liaw (2025), now a PhD student at Harvard
  • Pavel Czempin (2025), now a PhD student at USC
  • Vamshi Krishna Bonagiri (2025), now a PhD student at MBZUAI
  • Juan Liévano-Karim (2024), now a Teaching Professor at Universidad de los Andes
  • Tu Trinh (2024), now a Machine Learning Research Engineer at Scale AI

Other things about me

  • I speak English (native) and Spanish (advanced proficiency).
  • I wrote an algorithmic art-generator called RAMbrandt as a personal project in college.
  • I've composed and produced music. Check out my Spotify!