eMERIT.ai / Clinical care decision support

Most patients who
deteriorate in hospital
are visible first.

Hours before harm, the signs are usually already in the record. The failure is rarely detection. It is the distance between something being visible and somebody acting on it. That distance is what this platform is built to close.

November 2005

Vanessa Anderson was sixteen.

She was struck on the head by a golf ball. She was taken to Royal North Shore Hospital in Sydney, where she was treated for a fracture of the skull — an injury that was not, in itself, expected to kill her.

She died in hospital, of respiratory arrest, caused by the depressant effect of the pain medication she had been given.

Nothing exotic killed her. What the coronial inquest found instead was a sequence of ordinary things that did not connect: an observation that was not escalated, a medicine that should not have been given, a handover that did not carry the risk, a record that never held the whole picture. Every piece of what was needed to save her existed somewhere in that hospital. None of it arrived in the same place at the same time.

“…not enough doctors, not enough nurses, inexperienced staff, poor communication, poor record keeping and poor management. These are systemic problems that have existed for a number of years and regrettably they all surface in the death of Vanessa Anderson.” NSW Deputy State Coroner · inquest findings · 24 January 2008

Her case helped precipitate a Special Commission of Inquiry into acute care in NSW public hospitals, and through it, the state-wide escalation system that every NSW hospital runs today. Two decades on, her inquest is still the clearest short description of how a modern hospital fails a patient it could have saved.

Four failures, one patient

No single instrument sees all of this.

Read closely, the inquest describes four different failures happening at once — and that is the difficulty. A warning score built from vital signs cannot see a medication error. A medication checker cannot see an observation nobody escalated. Each tool is blind in exactly the place the next one is looking.

Not escalated

A sign was visible in the record, and the concern it should have raised never reached anyone able to act on it.

The wrong medicine

A drug was given that, for this patient with this injury, carried a risk that was not brought forward at the moment of prescribing.

Lost at handover

The picture changed hands between shifts and teams, and the part that mattered did not survive the journey.

Alone at 3am

The least experienced person present was the one who had to recognise it, with no support standing behind the decision.

March 1951

The poem he could not show his father.

Dylan Thomas sent nineteen lines to an editor in Rome. He had written them for his father — a Swansea schoolmaster, ill for twenty years, losing his sight, and dying. They became some of the most quoted lines in the language: Do not go gentle into that good night. Rage, rage against the dying of the light.

In the postscript of the letter that carried the poem, he added one sentence.

“…the only person I can't show the little enclosed poem to is, of course, my father, who doesn't know he's dying.” Dylan Thomas to Marguerite Caetani · 28 March 1951

The man the poem was written for had never been told. He died the following December, his son holding his hand. The rage in those lines is what was left over when the conversation did not happen.

This is the second failure this platform is built around, and it is not Vanessa Anderson's. Hers was a death that should have been prevented. This is a death that should have been spoken about — a decline visible for years, in front of everyone who loved him, and silence where the conversation belonged.

A hospital gets both kinds wrong for the same underlying reason: something was knowable, and nobody acted on it in time. That a patient is dying is a prediction like any other, and recognising it early buys the one thing that cannot be bought late — time to talk, while there is still someone to talk to.

It is not a rare problem. About a third of rapid response calls reach a patient who is dying rather than one who can be rescued — a finding from our own trial data. An emergency service is arriving, at speed, where a conversation was what was needed. 2008

And that is only the half that gets noticed. Most people approaching the end of life never have a sudden deterioration at all. They decline slowly, across weeks and admissions, and never cross the threshold that summons anyone — so a system built to catch rapid deterioration never sees them, and the conversation comes late, or not at all.

So this platform is built to hold two questions at once, and to tell them apart. Is this a deterioration that can be reversed? Or is this someone near the end of a long decline, where the right answer is not a rescue but a conversation? At the bedside, at three in the morning, those can look alike — and they call for opposite things. Separating them is among the harder problems in this field, and meeting it is a deliberate part of the design, not a by-product of predicting deterioration well.

What twenty years of measurement showed

The teams worked. The calling did not.

Twenty years ago our group ran MERIT — still the only large cluster-randomised trial of the hospital emergency team, anywhere. It found no reduction in its combined outcome, and it is routinely quoted as the verdict on rapid response teams. It is not a verdict. It is inconclusive, and its own investigators published why: the intervention was delivered at a fraction of the dose mature systems run at, the comparison hospitals changed their behaviour too, and the trial itself estimated that more than a hundred hospitals would have been needed to see the difference it was looking for. And, as a modelling re-analysis later showed, with only 23 hospitals a conventional subgroup analysis could not see how differently those hospitals were already performing. 2005 2009

What it did establish is where the failure sits. When the team was called, it worked. Only about a third of patients who met the criteria for a call ever generated one — and later work showed the cost of the delay directly: when the call came late, more patients died. The gap was never the response. It was the recognition, and the escalation. 2009 2015

Then a whole jurisdiction did it at once

NSW put a standardised escalation system into every public hospital — the first time anywhere that a whole large health jurisdiction had done so. Evaluating something at that scale needs a different instrument than a trial: an interrupted time series across 9,799,081 admissions in all 232 hospitals, comparing before with after. 2016

46%
fewer cardiac arrests
54%
fewer arrest deaths
19%
lower hospital mortality
35%
less failure to rescue

Stated with its design attached, because that is the honest form. These reductions are associated with the programme; the design is an interrupted time series, supported by a dose–response relationship — observational evidence, not randomised certainty. Definitive randomised evidence is no longer obtainable, because escalation is now mandated standard of care and withholding it would not be ethical. What both designs agree on is the mechanism: the benefit tracks how early and how often the call is made. That is the quantity this platform exists to raise. 2016 2022

How we got here

Twenty-four years, as dates rather than assertions.

Every entry is a public, checkable event — a published trial, a coronial finding, a commission of inquiry, a state-wide programme — and each carries the year to look it up under in the sources at the foot of this page.

  1. 2002–03

    The MERIT trial runs across 23 Australian hospitals — still the only large cluster-randomised trial of the medical emergency team, anywhere.

  2. 2005

    MERIT reports no reduction in cardiac arrest, unexpected death or unplanned ICU admission. In November, Vanessa Anderson dies at Royal North Shore Hospital.

  3. 2008

    In January the Deputy State Coroner finds systemic failure in her care. In November, Peter Garling's Special Commission of Inquiry reports a prevalent problem in the care of the deteriorating patient. The same year, MERIT's own data show about a third of emergency team calls are made to patients who are dying, not deteriorating — rescue arriving where a conversation was needed.

  4. 2009

    Two analyses explain the null. Hospitals that made a higher proportion of early calls had significantly fewer arrests and unexpected deaths — a dose–response. And a modelling re-analysis shows that with only 23 hospitals, a conventional subgroup analysis cannot see how differently those hospitals were already performing.

  5. 2010

    NSW introduces Between the Flags — colour-coded observation charts and a mandatory escalation pathway. Every state hospital has it by 2012.

  6. 2014

    Two studies track what happens as rapid response systems expand: falling arrest and mortality trends across NSW, and a four-hospital comparison of arrests and deaths before and after implementation.

  7. 2015

    A multicentre study puts a cost on lateness: when the call is delayed, more patients die.

  8. 2016

    The jurisdiction-scale evaluation — 9,799,081 admissions, all 232 NSW public hospitals — finds cardiac arrests down 46%, arrest-related mortality down 54%, hospital mortality down 19% and failure-to-rescue down 35% after Between the Flags.

  9. 2022

    The same design, applied to 5,114,170 female admissions, finds new mortality reductions in age groups where deterioration had been under-recognised.

  10. 2026

    This platform: built against the four failures in Vanessa Anderson's case — escalation, communication, medication, and the junior left alone — and against the second failure, with a continuous end-of-life screen so the conversation happens in time, and the third of rapid response calls that reach a dying patient directed to the right pathway — one system across the whole inpatient journey: ED, ward and ICU.

What eMERIT.ai is

A place for everything, and everything in its place.

Not another risk score. A bedside system whose job is to make sure that what is already knowable about a patient reaches a person who can act, while it still matters.

  • It watches everyone, continuously

    Not only the patients somebody already thought to worry about. Every inpatient, all the time — which is precisely the population a busy ward cannot watch evenly.

  • It reads more than one kind of signal

    Observations, results, what clinicians wrote, what was imaged, what the monitors traced, what was said at handover. Vanessa Anderson's inquest is the argument for this: four failures at once cannot be caught by an instrument that sees one thing.

  • It turns a signal into a specific action

    Grounded in the guidelines and the standards the hospital already works to, and expressed in the hospital's own language — a handover in the format the ward uses, an escalation on the pathway the state mandates. The structure is the standard, not a generic summary wearing local labels.

  • A clinician decides. Always.

    Everything it produces is a draft put in front of a person, who can check it, question it, or overrule it. It is built to stand behind the least experienced person on the ward at three in the morning — never to replace their judgement, and never to act in their place.

  • It treats dying as its own question

    Almost all work in this field treats death only as the outcome to avoid. Here, anticipated death is its own pathway with its own recognition — including for the patients who never deteriorate sharply enough to be noticed.

How it has been built

What is standing today.

Principles are easy to write. This is what has actually been built and runs in the research and field-test environment — none of it deployed, and none of it yet used in the care of a patient.

A reasoning engine that can cite itself

Generative search over a knowledge graph built from the guidelines, standards and policies the hospital already works to — so an answer traces back to the clause it came from, rather than to something a model once absorbed.

Five clinical pathways, working together

Unexpected death, end of life, sepsis, acute kidney injury, and medication safety. Not five products in a box: they consult one another, and the end-of-life pathway can stop a rescue that was about to be proposed.

Three settings, each fitted on its own terms

Emergency department, ward and intensive care. Each keeps its own models and its own calibration, because a pattern learned in intensive care does not transfer to a ward unexamined — and pretending otherwise is how clinical AI has failed before.

A reasoning partner — and a teaching one

It shows its evidence and its uncertainty rather than a verdict, so the least experienced person on the ward at three in the morning has something to reason against. The same grounded cases become teaching material — which answers the fourth failure in Vanessa Anderson's inquest, and the only one no alert can close.

More than sixty trained prediction models

Across deterioration, anticipated death and outcomes after discharge — each reported with its calibration and its net benefit, never with accuracy alone.

Encoders for everything that is not text

Images, waveforms and the spoken handover, so the reasoning can draw on a film, a trace or a conversation and not only on what somebody found time to type.

Documents drafted for a signature

Shift handover, rapid-response brief and discharge summary — written to the local standard, and never filed until a clinician signs.

A clinical trial engine

Two jobs. It runs single and multi-centre prospective trials end to end, where the binding cost is coordination rather than inference — most acutely in rare disease. And it takes the question routine care can already answer: a pre-registered trial emulated across whole populations, then run forward inside ordinary practice, before anyone spends years and millions randomising. Built so that a clinician without a biostatistician is held to the standard rather than excused from it.

The discipline

What it refuses to do is the part that matters.

Any system can be asked to be careful. These are built in as structural limits rather than instructions, which is the difference between a promise and a property.

  • It does not

    diagnose. It raises possibilities with the evidence for and against them, and names what nothing explains. The conclusion is the clinician's.

  • It does not

    file anything by itself. Every document is a draft until a clinician signs it. Nothing reaches a patient record, a GP or another team unsigned.

  • It does not

    guess to fill a gap. Where the record is silent it says so, in the open. A claim it cannot ground is withheld, not shown anyway.

  • It does not

    change while nobody is looking. The version running at a bedside is fixed and identified. Which version was running, for which patient, at what time, has exactly one answer.

Where it actually is, today

Research and field test. Nothing more, and we will not say otherwise.

The console behind this page runs on de-identified historical hospital data from Australia and the United States — emergency department, wards and intensive care — as a research and field-test environment.

  • It is not deployed, not registered, and not in clinical use.
  • Everything runs in shadow — real logic, real logging, zero clinical effect. Nothing it produces reaches a patient.
  • No clinician has yet evaluated it in practice. Nothing here supports a claim that it improves care. That evidence has to be collected, and collecting it properly is the next piece of work.
The measure we ask to be judged against

A hospital in which Vanessa Anderson survives.

The death that should have been prevented.

A hospital in which a son can show his father the poem.

The death that should have been spoken about.

That is the whole of the ambition, and it has not been met. It is a standard, not a claim — and the difference between the two is the point of this page.

Get in touch

Who we would like to hear from.

The platform is built as separable parts, so a site can take one of them without taking all of them. If any of this is close to a problem you have, tell us which part and we will reply with what testing it would actually involve.

The form asks who you are, where you work, which setting you are in and which parts interest you. It takes about two minutes.

The form opens in a new tab and runs on UNSW's Microsoft 365. Responses come to Prof Jack Chen and are used only to reply to you. Please do not include any patient information.

Sources

Where the numbers come from.

Every quantitative claim on this page resolves to an entry here, keyed by the year it happened. A claim without one does not appear.

  1. 2005

    Hillman K, Chen J, Cretikos M, Bellomo R, Brown D, Doig G, Finfer S, Flabouris A; MERIT study investigators. Introduction of the medical emergency team (MET) system: a cluster-randomised controlled trial. Lancet. 2005;365(9477):2091–2097. — the trial, the null composite, and the >100-hospital estimate quoted above.

  2. 2008

    Chen J, Flabouris A, Bellomo R, Hillman K, Finfer S. The Medical Emergency Team System and Not-for-Resuscitation Orders: results from the MERIT study. Resuscitation. 2008;79(3):391–397. doi:10.1016/j.resuscitation.2008.07.021

  3. 2008

    Inquest into the death of Vanessa Anderson. NSW Deputy State Coroner, 24 January 2008. — Garling P. Final Report of the Special Commission of Inquiry into Acute Care Services in NSW Public Hospitals, 27 November 2008.

  4. 2009

    Chen J, Bellomo R, Flabouris A, Hillman K, Finfer S. The relationship between early emergency team calls and serious adverse events. Crit Care Med. 2009;37(1):148–153. — the dose–response.

  5. 2009

    Chen J, Flabouris A, Bellomo R, Hillman K, Finfer S; MERIT investigators for the Simpson Centre and the ANZICS Clinical Trials Group. Baseline hospital performance and the impact of medical emergency teams: modelling vs. conventional subgroup analysis. Trials. 2009;10:117. doi:10.1186/1745-6215-10-117

  6. 2009

    Jones D, Bellomo R, DeVita MA. Effectiveness of the Medical Emergency Team: the importance of dose. Crit Care. 2009;13(5):313. — the 25.8–56.4 calls per 1,000 admissions range.

  7. 2010

    NSW Health / Clinical Excellence Commission. Between the Flags: standard calling criteria and the Clinical Emergency Response System. Australian Commission on Safety and Quality in Health Care, NSQHS Standard 8: Recognising and Responding to Acute Deterioration (2nd ed., 2017).

  8. 2014

    Chen J, Ou L, Hillman KM, Flabouris A, Bellomo R, Hollis SJ, Assareh H. Cardiopulmonary arrest and mortality trends, and their association with rapid response system expansion. Med J Aust. 2014;201(3):167–170.

  9. 2014

    Chen J, Ou L, Hillman K, Flabouris A, Bellomo R, Hollis SJ, Assareh H. The impact of implementing a rapid response system: a comparison of cardiopulmonary arrests and mortality among four teaching hospitals in Australia. Resuscitation. 2014;85(9):1275–1281.

  10. 2015

    Chen J, Bellomo R, Flabouris A, Hillman K, Assareh H, Ou L. Delayed emergency team calls and associated hospital mortality: a multicenter study. Crit Care Med. 2015;43(10):2059–2065. — the cost of lateness.

  11. 2016

    Chen J, Ou L, Flabouris A, Hillman K, Bellomo R, Parr M. Impact of a standardized rapid response system on outcomes in a large healthcare jurisdiction. Resuscitation. 2016;107:47–56. doi:10.1016/j.resuscitation.2016.07.240 — all four effect sizes above.

  12. 2021

    Bhonagiri D, Lander H, Green M, Straney L, Jones D, Pilcher D. Reduction of in-hospital cardiac arrest rates in intensive care-equipped New South Wales hospitals in association with implementation of Between the Flags. Intern Med J. 2021;51(3):375–384. — independent corroboration.

  13. 2022

    Chen J, Ou L, Hillman K, Parr M, Flabouris A, Green M. Impact of a standardised rapid response system on clinical outcomes of female patients: an interrupted time series approach. BMJ Open Qual. 2022;11(3):e001614.