Research
Trustworthy AI in healthcare, NLP for the Sustainable Development Goals, and — now — how to evaluate agents.
Questions
- How do you attribute an agent's failure to the step that caused it, rather than grading only the final state?
- What does a simulated world need to contain for an eval result to transfer to production behavior?
- What makes a scenario cheap enough to write that engineers write one when they find a bug, and expensive enough to be right that the suite can be trusted?
- What has to be true of a preference signal collected from real usage for training on it to improve the product rather than the metric?
- What is the right unit of an eval scenario for systems that both build automations and repair them when the world changes?
- Which properties of a benchmark predict that improvements on it will show up for real users?
Papers
- 2023Sailing the Data Sea to Advance Research on the Sustainable Development GoalsArgues that progress on the SDGs is bottlenecked by data discoverability, and sets out the design and ethics of a global catalog to fix it.The Ethics of Artificial Intelligence for the Sustainable Development Goals, Springer · DOI
- 2022
- 2022Assessing trustworthy AI in times of COVID-19: deep learning for predicting a multiregional score conveying the degree of lung compromiseA post-hoc trustworthiness assessment of a deployed chest-X-ray severity model in Brescia, applying the EU HLEG framework under pandemic pressure.IEEE Transactions on Technology and Society · DOI
- 2022Note: Leveraging artificial intelligence to build a data catalog and support research on the sustainable development goalsPresents the SDG Data Catalog system: NLP over open-access literature to build a global database of datasets, metadata and research networks.ACM SIGCAS/SIGCHI COMPASS · DOI
- 2021
- 2021On assessing trustworthy AI in healthcare: Machine learning as a supportive tool to recognize cardiac arrest in emergency callsApplies the EU trustworthy-AI guidelines to a live emergency-call cardiac-arrest detector in Copenhagen and surfaces the trade-offs an interdisciplinary team actually faces.
- 2020Hidden in Plain Sight: Building a Global Sustainable Development Data CatalogueDescribes mining millions of open-access papers to build an open, extensible catalogue of SDG-relevant datasets.ICT Analysis and Applications, Springer · DOI
Research projects
- 2018 — 2021SDG Data CatalogLed the UN SDG Data Catalog at the AI for Good Foundation: an NLP extraction pipeline (BiLSTM-CRF, fine-tuned BERT) over millions of academic papers that maps datasets, research, and organizations across the Sustainable Development Goals, reaching over 92% recall on dataset discovery. Secured $300K+ in Microsoft funding and was presented at SIGKDD 2019.
- 2020 — 2023Z-InspectionCo-developed a process, grounded in applied ethics and the EU High-Level Expert Group definition of trustworthy AI, for assessing whether an AI system can be trusted in practice. Applied it to four production healthcare AI systems with clinicians, ethicists, and legal scholars; the assessments became four peer-reviewed publications.
- 2024MicroBio LLMFine-tuned Mistral and LLaMA on k-mer DNA sequences for metagenomic classification, reaching 82% family-level taxonomic accuracy by treating genomic sequences as a language.
- 2019
- 2020 — 2021
- 2019
Coverage
- Apr 2021SDG Data Catalog funded by MicrosoftAI4Good / Microsoft
- Jun 2019MEng alumnus Andy Spezzatti speaks at SIGKDD 2019Berkeley Fung Institute