Hosted by the Center for Advancement and Dissemination of Intervention Optimization (CADIO)
There is a growing interest in adding generative-AI components to existing behavioral or digital health interventions. However, conversational AI introduces failure modes, misinformation, missed escalation, and the cultivation of emotional dependency. This webinar makes the case that these risks are best surfaced during the preparation phase, before a system is ever optimized or placed in front of participants. Drawing on work developing a generative-AI assistant for a substance-use recovery platform, Drs. William Nardi and Annika Schoene will introduce a two-part approach to safety testing: first, attacking a system with validated single- and multi-turn adversarial prompts to confirm it clears established hurdles; and second, developing domain-specific "stress tests", grounded in clinical expertise and patient lived experience, that probe how a clinical condition can show up in conversation and challenge a system in subtler ways over time. A central theme is a distinction that reframes how we think about safety evidence, where a good adversarial test is defined by its validity, not by whether it breaks the system. Breaking a system in the lab is a useful and informative exercise, not a failure and helps to fix problems before real-world exposure. The talk closes by connecting these ideas back to optimization, where continuous safety monitoring can become one more criterion informing component-selection decisions.