Andy Spezzatti

I work on how to know whether AI agents actually work.

At Zapier I build the evaluation infrastructure for agentic automation: simulated worlds that agents act in, end-to-end scenarios grounded in real outcomes, and the tooling that turns those measurements into release decisions.

I've been circling the same question for a while. At Genentech I built simulation engines so a portfolio of three hundred clinical trials could be planned before reality arrived. At Taltrics, which I co-founded, we shipped multi-agent systems to enterprises and post-trained them to stop hallucinating tool calls. The instrument keeps changing; the question — how do you trust a system before it meets the world? — hasn't. Along the way I published on trustworthy AI in healthcare.

Now — world simulation, eval-driven development, benchmarking agent performance. Sep 2026

Views here are my own.

Writing
  1. Aug 2026
    Which branch would you bet on
    A leaderboard reports scores. A decision needs a preference and an uncertainty, and frontier scenarios are exactly where you have neither by default.
  2. Jun 2026
    Eval-Driven Development
    Agents move decisions out of control flow and into the model. That part of the specification can only be recovered by watching the system act, which makes the eval the place where it gets written.
  3. May 2026
    Known-bad candidates
    An eval that has never been shown a wrong answer has not been shown to measure anything.
  4. Apr 2026
    Five of five is not a statistic
    What repeated eval runs can tell you, what they cannot, and why the two get confused.
  5. Dec 2022
    Priors for a portfolio
    A model's job is to be honest about what it doesn't know, in a form a decision can use. Three years of simulating clinical trials taught me what that means.

All writing →

Selected work
  1. 2025 –
    Eval infrastructure for agentic automation
    Zapier — simulated worlds and release gates for an agent platform used by 1.5M+ people.
  2. 2022 – 25
    Scoping and pricing a deal on the call
    Taltrics — agents that take an RFP to a priced proposal in minutes, post-trained so the number is right.
  3. 2019 – 22
    Simulating a clinical-trial portfolio
    Genentech / Roche — a discrete-event engine behind planning for 300+ trials in 70 countries.

More →

Selected papers
  1. 2023
    Sailing the Data Sea to Advance Research on the Sustainable Development Goals
    The Ethics of Artificial Intelligence for the Sustainable Development Goals, Springer
  2. 2022
    To explain or not to explain? — Artificial intelligence explainability in clinical decision support systems
    PLOS Digital Health
  3. 2021
    Co-design of a trustworthy AI system in healthcare: deep learning based skin lesion classifier
    Frontiers in Human Dynamics

All research →

Talks & media
  1. Feb 2026
    Build-Along Workshop on Agentic AI
    Zapier × Gamma
  2. Sep 2024
    AI and the Future of Global Sustainability
    Socially Conscious AI Podcast
  3. Jun 2022
    Leveraging AI to build a Data Catalog and support research on the SDGs
    ACM COMPASS
  4. Jul 2020
    Hidden in Plain Sight: Building a Global Sustainable Development Data Catalogue
    ICT4SD

More →