The Long Journey Home
Mortality, clinical trials and the race to the South Pole
The Black Flag
The first sign of failure was a dark speck on a white horizon.
At first, it could have been anything. Then it resolved into a flag.
It was black.
On 16 January 1912, Robert Falcon Scott and his four companions were less than a day’s march from the South Pole. Around the flag were the remains of a camp, ski tracks and the prints of many dogs. “This told us the whole story,” Scott wrote. Roald Amundsen’s party had been there before them.
Yet the greater catastrophe had not even begun.
The following day, Scott’s party reached what their observations placed as the Pole. “Great God! this is an awful place,” Scott wrote. His thoughts then turned northwards: “I wonder if we can do it” The next morning, after recalculating their position, they found Amundsen’s tent. Inside was a record dated 16 December 1911; the Norwegians had reached the Pole more than a month earlier1
They could not.
Amundsen and his four companions returned safely to their base. Scott’s five-man polar party did not. Nearly eight months after Scott’s final diary entry, the bodies of Scott, Wilson and Bowers were found in a tent just 11 miles from One Ton Depot.1,2
On one map, both expeditions reached the destination. On another, only one completed the journey.
Mortality is the second map. It asks whether the patient came home. That is its extraordinary strength. Its limitation is that survival and coming home are not always the same.
An expedition is a system. Its outcome depends on the route, the weather, the food, the fuel, the depots, the sledges, the animals, the clothing, the decisions and the people. If a clinical trial tests just one component of that system—a drug, a ventilator setting or a pair of boots—should the fate of the entire expedition be the only way it is judged?
The Destination Is Part of the Design
Before an expedition leaves, it must decide what will count as success. A clinical trial is no different. Its primary outcome is the destination declared before departure; secondary outcomes describe what happened along the route.
If an intervention can plausibly alter survival, mortality may be the right destination. If it acts on only one component of the journey, an outcome closer to that mechanism may provide the more informative test. Choosing mortality does not automatically make a trial more rigorous. It simply asks whether the intervention altered survival by a specified time.3,4
The Outcome That Cannot Be Argued Away
Mortality’s authority is earned.
It is directly important to patients, unambiguous and resistant to observer judgement. It does not depend on whether a clinician believes that an infiltrate has improved, whether a patient is considered ready for extubation or whether a lack of available beds prevents discharge from the ICU. Ascertainment is generally simple, follow-up can be highly complete and the meaning of the event requires little interpretation.3,4
Mortality also captures the net effect of treatment. An intervention may improve the variable it was designed to change while causing an unexpected complication elsewhere. Death integrates intended benefit, recognised toxicity and harms that investigators never thought to measure.
That capacity has repeatedly protected patients. Lower tidal volume ventilation, prolonged prone positioning in severe acute respiratory distress syndrome and dexamethasone in appropriately selected patients with severe COVID-19 produced survival benefits that transformed practice.5,6,7 In the opposite direction, CRASH, FEAST and OSCILLATE exposed lethal consequences that physiology, prevailing practice or intermediate outcomes had failed to predict.8,9,10
A biomarker can flatter an intervention. Mortality is harder to persuade.
The Expedition Is Not the Boot
Now change the question.
Suppose the expedition was being used to test a new pair of boots. The boots were lighter, warmer and kept the wearer’s feet dry. The wearer walked farther each day and developed less frostbite. But the expedition still failed after a fuel shortage, a missed depot and an overwhelming blizzard.
Did the boots fail?
Judged by expedition mortality, they did. Judged by the outcome they could reasonably influence, they may have performed exactly as intended.
The reverse is equally possible. A boot may produce reassuring measurements of warmth while retaining sweat, impairing movement or causing tissue injury. A narrowly physiological endpoint could declare success while the wider expedition reveals harm. The nearer outcome may therefore be more sensitive, but the farther outcome may be more complete.
Scott’s polar journey depended on motor sledges, ponies, dogs and ultimately man-hauling. It depended on the siting of depots, the quantity of food and fuel, the reliability of equipment, the condition of individual men, the route, the weather and a sequence of decisions made before and during the journey. More than a century later, the relative contribution of these factors remains debated. The expedition’s outcome is certain. The amount attributable to any single component is not.1,2
The same is true in critical illness. A ventilator strategy is only one part of a journey that may also involve antibiotics, fluids, vasopressors, sedation, surgery, renal replacement therapy, nutrition, complications, rescue treatments, decisions to limit care and rehabilitation. Mortality is a whole-expedition endpoint. Many treatments are boot-sized interventions.
The correct outcome depends on how much of the journey the intervention can plausibly own.
The Signal Fades Across the Ice
The farther the endpoint lies from the treatment, the more opportunities there are for its signal to disappear. Between randomisation and death stand physiological response, organ function, complications, co-interventions, cross-over, treatment withdrawal, rehabilitation, recurrent illness and chance.
A therapy directed at pulmonary oedema may improve lung function but have little influence on deaths caused by refractory sepsis, malignancy, neurological injury or later treatment-limitation decisions. All-cause mortality avoids the uncertainty and bias of adjudicating why a patient died, but it includes deaths that the intervention could never reasonably have prevented. Cause-specific mortality lies closer to the mechanism, yet introduces uncertainty about attribution.3,4
Mortality can also conceal opposing effects. An intervention may prevent deaths from respiratory failure while causing deaths from cardiovascular collapse, or shorten ventilation while increasing renal injury. At the chosen time point, the net mortality difference may be zero even though the treatment has produced both benefit and harm.
This is the reverse side of its completeness. Mortality reports the net survival balance, but can erase the mechanisms that produced it. Component outcomes, adverse events, causes of death and treatment trajectories are therefore essential to interpretation, although they cannot be used post hoc to rescue a neutral randomised comparison.
Randomisation protects the comparison from systematic imbalance. It does not make every subsequent death informative about the treatment. A mortality estimate can therefore be unbiased yet still be too imprecise, too diluted or too causally distant to answer the question well.
Mortality tells us the balance at journey’s end. It does not provide an inventory of what was gained and lost along the way.
The Arithmetic of the Impossible Journey
Causal distance has a statistical price.
Most modern critical care interventions are unlikely to reduce mortality by 10 or 15 absolute percentage points. Their effects, if real, are generally smaller.11,12 Detecting those effects requires large numbers of patients and dependable assumptions about baseline risk.
Those assumptions frequently fail. Baseline mortality may fall while a trial is being designed and recruited. The enrolled population may be less severely ill than expected. Treatment separation may be smaller than planned. Heterogeneous syndromes may include many patients whose illness cannot respond to the intervention.11,12
An analysis of adult critical care randomised trials found substantial heterogeneity in outcome selection and frequent limitations in statistical power.11 A subsequent analysis of 101 major mortality trials found that 77.3% had overestimated control-group mortality, 47.0% had been powered to detect an absolute mortality reduction of at least 10%, and 67.3% could not exclude clinically important benefit or harm.12
This is why a statistically non-significant mortality result is not necessarily evidence of no effect. A confidence interval that spans worthwhile benefit, no difference and important harm describes uncertainty—not equivalence.
Mortality is easy to count. It is extraordinarily expensive to estimate precisely.
The Last Depot
Death is unambiguous. Its address in time is not.
ICU mortality, hospital mortality, 28-day mortality, 60-day mortality, 90-day mortality and six-month mortality are not interchangeable endpoints. Each places the last depot at a different point on the journey.
Measure too early and the intervention may not have had time to produce its full benefit or harm. Measure too late and deaths unrelated to the original illness progressively dilute the treatment signal. In-hospital mortality introduces an additional problem because discharge practices differ between hospitals and health systems. A fixed time point avoids that particular distortion, but still requires a biological and clinical justification.4
The correct time horizon depends on the treatment. A rapidly acting resuscitation intervention may exert its principal effect within hours or days. A strategy that alters organ injury, recovery or late complications may require much longer observation.
“Mortality” is therefore not one endpoint. It is an event attached to a chosen coordinate in time.
The Scientific Expedition
A failed expedition can still produce knowledge.
Scott’s party did not return, but the wider Terra Nova expedition conducted an extensive scientific programme. That achievement does not undo the deaths or transform the polar journey into a success. It shows that one dominant outcome does not exhaust everything that may be learned from an expedition.1
A trial can similarly fail to demonstrate a mortality effect while identifying a clinically meaningful effect elsewhere. In FACTT, conservative fluid management did not significantly reduce 60-day mortality, but improved lung function and increased ventilator-free and ICU-free days without increasing non-pulmonary organ failure.13
That result is not equivalent to a survival benefit. Neither is it equivalent to no benefit.
The distinction depends on prespecification and coherence. A neutral primary endpoint cannot be rescued by searching through dozens of secondary outcomes until one crosses a significance threshold. But prespecified outcomes that are important to patients, closely connected to the treatment and supported by consistent findings remain informative.
The scientific specimens collected along the route matter. They are not the same as reaching the declared destination.
Survival Is Not the Same as Coming Home
Even survival may stop the story too soon.
A patient who is alive at day 30 may have returned to independent life. The same binary outcome is assigned to a patient with profound neurological disability, persistent organ dependence, severe weakness, cognitive impairment or an intolerable symptom burden.
PARAMEDIC2 exposed this tension. Adrenaline increased 30-day survival after out-of-hospital cardiac arrest, but did not significantly increase survival with a favourable neurological outcome. Severe neurological impairment was more frequent among survivors allocated to adrenaline.14
RESCUEicp presented an equally difficult distribution of outcomes. Decompressive craniectomy reduced mortality after severe traumatic brain injury, but increased survival in vegetative and severely disabled states; rates of moderate disability and good recovery were similar at six months.15
Neither trial makes mortality unimportant. Both show why mortality may be insufficient.
Survival is the prerequisite for recovery. It is not a description of the recovery achieved. Mortality tells us who remained alive; it does not tell us what life survived.
Reaching home is not merely crossing the threshold with a pulse.
The Danger of Nearer Landmarks
Once mortality’s limitations are visible, the opposite error becomes tempting: stop at the first measurable cairn.
Biomarkers, physiological variables, organ-support duration and hospital length of stay may be more responsive to treatment. They also lie closer to clinical practice and judgement. Extubation depends partly on sedation, staffing and local weaning practice. ICU discharge depends on bed availability and ward capacity. Length of stay can be prolonged by survival itself.
Composite outcomes increase event rates, but a common and less important component may dominate the result. A treatment may therefore appear beneficial because it changes laboratory-defined kidney injury while having no detectable effect on dialysis, survival or long-term health. Reporting the composite alone can conceal this imbalance.16
Outcomes measured only among survivors create another problem. If treatment changes who survives, the surviving groups may no longer be comparable. Better cognitive scores among survivors could reflect recovery, a different survivor population, or both.
Core outcome sets improve consistency by ensuring that trials measure outcomes important to patients and clinicians. They do not dictate that every outcome should be primary, or remove the need to match the primary endpoint to the intervention.17
A nearer landmark is useful only if it is important, valid and not mistaken for the destination. Importance and sensitivity are not the same. Mortality may be the most important outcome yet remain relatively insensitive to a modest, mechanistically narrow intervention; a physiological endpoint may respond dramatically while mattering little to patients. A trial of boots should assess what happened to the feet, while still recording frostbite, falls and deaths. The endpoint should fit the scale of the intervention.3,4
Declare the Mission Before Setting Out
The endpoint should be declared before the first patient is enrolled, but only after the causal route has been drawn.
Four questions should determine whether mortality belongs at the centre of a trial:
- Can the intervention plausibly alter survival? The treatment must act on a process that contributes materially to death in a meaningful proportion of the enrolled population. The farther the intervention lies from the causes of death, the weaker the rationale for mortality as the primary endpoint.
- Can the trial detect a realistic effect? Baseline mortality, the preventable fraction of deaths, treatment separation and the anticipated absolute effect must be credible. A design that requires an implausibly large survival benefit to remain feasible is not made rigorous by choosing mortality.
- Is the time horizon biologically appropriate? Follow-up must be long enough to capture the intervention’s important consequences, but not so long that unrelated events overwhelm the signal.
- Would survival alone answer the patient’s question? When treatment may exchange death for severe disability, dependence or prolonged suffering, function, quality of life and symptom burden are indispensable.
If these conditions are met, mortality may be the correct primary outcome. If they are not, it should usually remain an important safety or secondary outcome while the primary endpoint is placed closer to the effect the intervention is designed to produce.3,12
The Tyranny of the Furthest Endpoint
Mortality becomes dangerous when it stops being an outcome and becomes a badge of seriousness.
A culture in which only mortality-changing treatments are considered important risks dismissing meaningful reductions in ventilation, organ failure, disability, symptoms and time spent in hospital. It may also produce trials that are enormous yet underpowered because a modest intervention has been asked to move the most distant possible endpoint.
The opposite culture is no better. A treatment should not be declared successful because it improves a biomarker, pressure, ratio or score while leaving patients no better—or making them worse.
The solution is not to place every trial at the Pole. It is to choose the most patient-important endpoint that the intervention can plausibly and detectably influence.
The best endpoint is not necessarily the furthest one. It is the furthest endpoint the treatment can reasonably be expected to move.
The Expedition Log
An expedition log is more useful than a scoreboard.
This is the terrain charted by the collection Critical Care Trials Reporting a Mortality Effect. It brings together major randomised trials reporting Mortality Benefit, Mortality Harm and Signals Not Confirmed, while distinguishing primary mortality outcomes from secondary and subgroup findings.
The collection is not a trophy cabinet of positive trials. It is a record of routes attempted.
Some interventions produced robust survival benefits and changed practice. Others exposed unexpected harm. Some early mortality signals appeared persuasive but were not reproduced when the journey was repeated in a subsequent large multicentre randomised trial.
Non-confirmation does not automatically prove that the original finding was false. The later trial may have enrolled different patients, tested another dose, begun treatment at another time or operated within a changed landscape of usual care. But chance, multiplicity, optimistic effect estimates and selective attention to secondary or subgroup results also leave tracks that may disappear when the route is retraced.
The evidential weight of a mortality finding therefore depends on more than whether its P value crossed 0.05. It depends on whether mortality was primary, whether the effect was precise and plausible, whether the intervention achieved meaningful separation, whether the finding survived multiplicity and whether another expedition found the same path.
Read longitudinally, the collection shows both the power and the fragility of mortality evidence. Few individual trials end the story. Most become part of a longer journey.
Eleven Miles from Home
Scott’s final camp was 11 miles from One Ton Depot.1
On a map, the remaining distance was small. On the Ross Ice Shelf, after months of exhaustion, hunger, injury and cold, it was insurmountable.
Mortality is the furthest endpoint in critical care research. We should not abandon it because it is difficult. Neither should we demand it from every intervention simply because it matters most.
When a treatment can plausibly alter survival, when the population is appropriate and when the trial can detect a realistic effect, mortality may be the destination that matters above all others. When the causal path is long, the expected effect modest or the quality of survival central, the trial must also measure the nearer landmarks that reveal what happened along the way.
The expedition must be judged by whether its members returned. The boots must be judged by what happened to the feet. Confusing those questions does justice to neither.
Mortality tells us whether the patient came home. It does not always tell us which part of the expedition brought them there.
References
Related Materials
- This blog was written with the assistance of AI



