Our research

Counted on 8 September 2026, from the files named beside each figure.

What we measure about the models we run, and what we have not measured yet

The Wellness runs models inside a clinic and publishes what that costs in risk. 156 places a model acts were read in source and written down, 71 hazards were scored against them, and 29 of those actions still reach a patient or the record with nobody accepting. Every figure here names the file it was counted from.

What we count

  • Places a model acts or proposes, inventoried156 actions
  • Hazards logged against those actions71 hazards
  • Actions that reach a patient or the record with nobody accepting29 actions
  • Hazards scored unacceptable before controls2 hazards
  • Hazards that stay at a residual 3 after every control we can think of5 hazards

What we have not measured

  • Median minutes to a first reply on a messageNot recorded
  • Median working days from referral to a first appointment offeredNot recorded
  • Share of drafts a clinician sends with nothing changedNot recorded

Median minutes to a first reply on a message. The board that measures this has not merged. A number here before then would be somebody remembering.

Median working days from referral to a first appointment offered. Counted by hand today and not yet instrumented, so it is not published.

Share of drafts a clinician sends with nothing changed. The table this reads lands with the cross cutting pass. Nought would read as the desk editing every draft, which is not what is true.

How it was counted

Every place a model acts or proposes anything in our system was read in source rather than inferred from a filename, and written down as a row. That inventory is the scope of everything below it.

Each of those actions was scored for how badly a patient could be harmed if it went wrong, and for how likely the harm pathway is to complete rather than how likely the model is to be wrong. A model that is wrong often and always caught scores low. A model that is wrong rarely and never checked does not.

Where a figure comes from published research rather than from us, the study is named beside it. Where a figure is ours, the file it was counted from is named beside it. Where we have not measured something, it says so.

  • docs/safety/AGENT-ACTIONS.md
  • docs/safety/HAZARD-LOG.md
  • The help desk, through the Traffic board
  • The accept telemetry

What the counting found

Two hazards score unacceptable and both are live
Under our own plan a risk of 5 does not ship and is turned off if it is already running. Two rows score 5 and both are live today. One is clinical text reaching a model call without redaction. The other is an approver who is anonymous by default. Both are fixable in a week and neither was fixed by writing them down.
Twenty nine actions reach a patient or the record with nobody accepting
Out of 156 inventoried, 29 take effect without a person pressing anything. That is the number the hazard log is built around, and it is the number that has to come down before any claim about supervision means much.
Five hazards cannot be designed out
Hallucination in a signed note, omission from one, the wrong patient, automation bias at the signature, and a voice agent speaking to a patient without review. Every control we can think of leaves these at a residual 3. They are managed rather than eliminated, and a safety case that claimed otherwise would be the thing to worry about.
Model drafted clinical summaries hallucinated in 42% of cases and omitted in 47%
From our own clinical study review rather than from a vendor. Omission is the more common failure and the harder one to see, because a reviewer reads what is on the page and cannot read what is missing from it. That asymmetry is why an unedited acceptance rate near 100% is a signal rather than a compliment.
Nothing on the public site was being measured until this month
The analytics script returned a 404 on the live domain, so every judgement about which pages worked was a guess. It is named here because a page about measurement that hid its own measurement gap would not be worth reading.

Who approved this

No Clinical Safety Officer is appointed, so no individual has approved the scores below. Every figure is a count of rows in a file named beside it and every score is a proposal awaiting that approval.

Written and reviewed by the clinical team at The Wellness. Last reviewed 8 September 2026.

The numbers on this page move when the files move, because a test recounts them.