Independent clinical AI assurance

Clinical AI must earn its place at the bedside.

AlliumAI builds independent assurance for AI that talks to, triages, and supports patients — synthetic-patient simulation, deterministic safety grading, and an audit method built to resist gaming. Built by clinicians, and forged in the hardest room in medicine: mental health.

Why this exists

An AI that mishandles a suicide disclosure doesn't fail like software.

It fails like a clinician — quietly, plausibly, and at the worst possible moment. So does the triage assistant that waves through a red-flag symptom, and the aftercare bot that softens a warning until it disappears. Demos won't catch it. Star ratings won't catch it. And a system that has learned to look safe under its own tests is more dangerous than one that was never tested at all.

Clinical AI is arriving faster than the evidence culture around it. We build the missing piece: adversarial, repeatable, clinically authored safety evaluation that stands apart from the system it judges. We hardened the method where failure is least forgiving — mental healthcare — and built it to travel to any AI that faces patients.

The harness

Four stages. No shortcuts.

01

Simulate

Synthetic patients with clinically authored risk journeys — suicidality, psychosis, red-flag escalation, unsafe-advice traps — press the system where it matters, across whole care journeys, not single prompts. No real patients. No real data. Ever.

02

Grade

A deterministic evaluator scores every conversation against hard safety gates: crisis routing, risk escalation, continuity of clinical reasoning, do-no-harm floors. Same transcript, same verdict, every time — grading you can replay, audit, and contest.

03

Audit

A learned auditor hunts for what the rulebook misses; clinicians confirm or reject what it finds; confirmed gaps become new deterministic rules. The auditor advises — it never decides, and it never trains the system under test.

04

Govern

Changes pass or fail on replayed evidence: triggering suites plus held-out scenarios, clean and tainted runs kept strictly apart, veto rules that cannot be argued with. No clean replay, no promotion.

Non-negotiables

Rules that make the results worth trusting.

  1. 01

    The evaluator never trains the agent.

    Optimising a system against its own safety test is how you get safety theatre. The wall between them is the product.

  2. 02

    A do-nothing change must never pass.

    Every evaluation cycle carries no-op controls. If doing nothing ever earns a promotion, the harness — not the AI — has failed.

  3. 03

    Clean evidence and tainted evidence never mix.

    A run with integrity problems is recorded, quarantined, and never averaged into the numbers a decision rests on.

  4. 04

    Deterministic grading is the source of truth.

    Learned judges flag, suggest, and accelerate. They do not gate, veto, or promote. Truth stays replayable.

Research

The evaluator is a scientific instrument. We treat it like one.

Underneath every claim that a clinical AI is "safe" sits a harder question: is the thing doing the judging valid? When an automated safety evaluator and an experienced clinician disagree about the same conversation, which of them is measuring the construct — and how much does the format of the evaluation itself, from single-turn vignettes to interactive sessions to longitudinal voice journeys, determine the answer? That question is largely open across the field, and it is the one our research programme exists to close.

We run that programme the way instrument science demands: pre-specified per-case safety contracts, blinded clinician references, chance-corrected agreement statistics reported with their limitations, and integrity accounting that keeps contaminated runs out of every headline number. Measurement culture, not marketing culture — including publishing the results that don't flatter us.

We work with academic research groups on these questions. If evaluator validity, synthetic-patient evaluation, or voice-modality safety is your field, we would like to hear from you.

Start a research conversation

Who we serve

Built for the people who carry the risk.

Clinical AI builders

Independent safety evidence for AI that faces patients — before your users, your board, or a journalist finds the gap first.

NHS innovation & transformation teams

A way to ask vendors the hard question — “show me the adversarial evidence” — and a harness that makes the answers comparable.

Clinical safety officers

Structured, replayable evidence that slots into DCB0129/DCB0160 safety cases instead of a folder of screenshots.

Investors & due diligence

Find out whether a clinical AI's safety story survives contact with a simulated bad week — before the term sheet, not after.

Founders

Two doctors who carry the risk they test for.

Dr Karl Rathbone

Co-founder & CEO

Leads the engineering and evaluation programme. GMC-registered doctor.

Dr Mark Morris

Co-founder & Chief Safety Officer

Owns clinical safety; every safety-behaviour change ships under his signature. GMC-registered doctor.

DSP ToolkitStandards Met
ICORegistered
Clinical safetyDCB0129-aligned process
Founded byGMC-registered doctors

This site sets no cookies, runs no analytics, and makes no third-party requests.

The conversation starts here.

Pilots, partnerships, and hard questions all welcome.

[email protected]