Underneath every claim that a clinical AI is "safe" sits a harder question:
is the thing doing the judging valid? When an automated safety evaluator and
an experienced clinician disagree about the same conversation, which of them
is measuring the construct — and how much does the format of the
evaluation itself, from single-turn vignettes to interactive sessions to
longitudinal voice journeys, determine the answer? That question is largely
open across the field, and it is the one our research programme exists to close.
We run that programme the way instrument science demands: pre-specified
per-case safety contracts, blinded clinician references, chance-corrected
agreement statistics reported with their limitations, and integrity
accounting that keeps contaminated runs out of every headline number.
Measurement culture, not marketing culture — including publishing the
results that don't flatter us.
We work with academic research groups on these questions. If evaluator
validity, synthetic-patient evaluation, or voice-modality safety is your
field, we would like to hear from you.
Start a research conversation