Statistical Monitoring and Measurement

Developing new metrics, measurements, and methods to monitor for critical events related to psychiatry and mood disorders towards improved patient care.

Clinical psychiatry depends on timely detection of change: relapse risk, treatment response, and other events that matter for patient care. Our work builds measurement frameworks and monitoring methods that can surface these signals reliably from real-world data streams.

We focus on metrics that are statistically sound, clinically interpretable, and practical to deploy alongside existing workflows, so monitoring systems reduce noise rather than add alert fatigue.

Projects in this area

  • NAVIGATOR

    A public repository that surfaces mental-health benchmark examples where AI models show the highest uncertainty or disagreement with compliance rules. Behavioral health benchmarks carry erroneous labels and drift over time, so NAVIGATOR separates two sources of uncertainty: a multi-agent system constrained by a directed acyclic graph simulates an organization’s process map to surface context-dependent (aleatoric) uncertainty, while Monte Carlo simulations at each agent surface model gaps (epistemic uncertainty). Internally validated labels feed back into the compliance monitoring system, which re-evaluates the datasets to identify the next set of high-uncertainty examples. The system now covers more than 60 open-source datasets, improves labeling accuracy by 16%, and reduces human review by up to 85x.

    Minseo Choi, Rosa Jahankhah, Ronald Deng, with Michael Rudow

  • Patient heterogeneity

    Developing multimodal AI methods that integrate biological, electronic health record (EHR), and longitudinal patient data to provide adaptive clinical decision support under diagnostic uncertainty while improving our understanding of patient heterogeneity and psychiatric comorbidity.

    Rhea Makkuni, Ram Chitti, Minseo Choi

  • Sensitivity classification

    Building a classifier that flags sensitive content in mental-health text, closely related to the compliance monitoring work in NAVIGATOR.

    Ronald Deng

  • Reliable self-harm risk screening

    Multi-agent LLM pipelines are being used to assess self-harm risk, but common evaluation approaches do not indicate when a decision is reliable or how errors accumulate across agents. We give these pipelines a statistical footing with agent-level confidence bounds, bandit-based adaptive sampling, and regret guarantees, cutting the false positive rate by 40% against single-agent models without losing recall.

    Meghana Karnam

Selected work

All publications →

← Back to research