Research

Trustworthy AI in healthcare, NLP for the Sustainable Development Goals, and — now — how to evaluate agents.

Questions
  1. How do you attribute an agent's failure to the step that caused it, rather than grading only the final state?
  2. What does a simulated world need to contain for an eval result to transfer to production behavior?
  3. What makes a scenario cheap enough to write that engineers write one when they find a bug, and expensive enough to be right that the suite can be trusted?
  4. What has to be true of a preference signal collected from real usage for training on it to improve the product rather than the metric?
  5. What is the right unit of an eval scenario for systems that both build automations and repair them when the world changes?
  6. Which properties of a benchmark predict that improvements on it will show up for real users?

Updated Sep 2026. I revise this list twice a year.

Papers

Also on Google Scholar.

Research projects
Coverage
  1. Apr 2021
    SDG Data Catalog funded by Microsoft
    AI4Good / Microsoft
  2. Jun 2019
    MEng alumnus Andy Spezzatti speaks at SIGKDD 2019
    Berkeley Fung Institute