Evaluation and Experimental Design

Strategies to evaluate and test GenAI tools in psychiatry and behavioral science.

Generative AI tools are entering psychiatry and behavioral science faster than evaluation practice can keep up. We design strategies to test these systems rigorously, covering validity, reliability, fairness across populations, and fitness for clinical or research use.

Our goal is evaluation methods that help labs and clinics decide when a tool is ready, where it fails, and how to measure improvement over time.

Projects in this area

  • Uncertainty in LLM psychiatric risk assessments

    LLM outputs vary depending on which clinical details are presented. We audit four models across four prompt framings and show that adding clinically irrelevant information significantly shifts predicted hospitalization risk and output variability, underscoring the need for governance around these models before clinical deployment.

    Shevya Panda, Shinjini Bose

  • Detecting drift in chatbot input

    The people talking to a mental-health chatbot do not all talk the same way: teenagers and older adults use different words and phrasing for the same concerns, and the input a deployed system sees shifts over time. We are working on detecting that shift and on mathematical models that quantify how much drift has occurred.

    Rosa Jahankhah

  • How mental health providers use chatbots

    A survey of mental health providers on whether and how they use chatbots in practice, and how they feel about them, to ground evaluation criteria in what clinicians actually need from these tools.

    Ariel Kim

Selected work

All publications →

← Back to research