Two claims, two grades
This review carries the highest grade in the journal and it is important to be precise about what has earned it.
The prognostic claim is that a person's measured cardiorespiratory fitness tells you a great deal about their probability of dying in the following years, over and above what smoking status, blood pressure, cholesterol and diabetes status tell you. That claim is supported by an unusually consistent body of large cohort evidence and is graded A.
The interventional claim is that if a specific person raises their fitness, their own risk falls by a corresponding amount. That is a causal claim about a change over time within an individual, and it is not the same statement. It is plausible, it is supported by mechanism and by the graded nature of the association, and it has not been tested by randomised trials with survival endpoints because such trials are close to impossible to run. We grade it C.
Conflating the two is the most common error made with this literature, including by people who should know better. Our explainer on surrogate endpoints covers the general form of the mistake.
What the measurement actually is
VO2 max is the maximum rate at which a person can take in, transport and use oxygen during exercise of increasing intensity. It is limited principally by the heart's capacity to deliver oxygenated blood, and secondarily by peripheral factors including capillary density and mitochondrial content in working muscle.
That is why it works as a prognostic marker. It is not a narrow measure of athletic ability. It is an integrated readout of cardiac function, vascular health, pulmonary function, blood oxygen carrying capacity, muscle quality and, indirectly, of whatever chronic disease is quietly consuming reserve. A single number that summarises the whole oxygen cascade will inevitably correlate with survival.
The reference measurement is cardiopulmonary exercise testing with analysis of expired gas, conducted to volitional exhaustion. Everything else is an estimate. Treadmill or cycle protocols without gas analysis estimate fitness from workload achieved. Non-exercise equations estimate it from age, sex, body composition and reported activity. Consumer wearables estimate it from heart rate response during ordinary activity. These estimates are useful for tracking a trend in one person over time. They are considerably weaker than they appear when used to place someone in a population distribution.
One technical point matters for interpretation. Fitness is often expressed relative to body mass, which means the number falls when weight rises even if the heart and muscle are unchanged. That mixes a fitness signal with an adiposity signal, and part of the association with mortality in some datasets runs through that conflation.
What the human evidence shows
The prognostic literature is large, old and consistent. Cohorts recruited in different countries and different decades, following people for years to decades, have reported the same pattern: those with lower measured fitness die sooner, and the relationship is graded across the whole distribution rather than being a threshold effect at the bottom.
Four features of that literature are what earn the grade. It replicates across independent cohorts and populations. It shows an orderly dose-response rather than a single cut-point. It holds in both sexes and across age bands. And it survives adjustment for the conventional risk factors, meaning fitness carries prognostic information that those factors do not.
The largest gains in predicted risk appear at the bottom of the distribution. Moving from the least fit group towards merely below average is associated with a much larger difference than moving from good to excellent. This is the single most practically useful feature of the whole literature, and it is the opposite of how fitness is usually discussed.
Reverse causation is the standing objection and it is a serious one. Undiagnosed disease lowers exercise capacity before it becomes clinically apparent, so some of the association between low fitness and early death is disease causing low fitness rather than low fitness causing death. Cohorts address this by excluding early deaths and by adjusting for baseline disease, and the association persists, which weakens but does not eliminate the objection.
On the interventional side, what is well established is that structured aerobic training raises measured fitness in randomised trials, in young and old, and that trainability is retained into later life even though the achievable ceiling falls with age. What has not been established by randomised trial is that the resulting change in fitness produces the mortality difference seen in the cohorts. A trial of that kind would need to randomise people to years of training and count deaths, and adherence, contamination and cost make it close to impractical.
The limitations that constrain what this can tell you
| Limitation | What it affects |
|---|---|
| Reverse causation | Subclinical disease lowers fitness before diagnosis, inflating the apparent prognostic effect. |
| Estimated rather than measured fitness | Wearable and equation-based estimates carry error that is large relative to the differences being interpreted. |
| Scaling to body mass | Expressing fitness per kilogram blends a cardiovascular signal with an adiposity signal. |
| Healthy cohort selection | People able and willing to complete a maximal exercise test differ from those who are not. |
| No randomised survival trial | The step from prediction to personal causation has not been tested directly. |
| Genetic component of trainability | Baseline fitness and the response to training both vary between individuals for reasons outside behaviour. |
None of this argues against exercise. UK physical activity guidance already recommends regular aerobic activity for reasons that do not depend on this literature at all, and the evidence for exercise on intermediate outcomes such as blood pressure, glycaemic control and physical function is randomised and robust.[2] The narrow point is about what a fitness number can and cannot tell an individual about their own future.
What would change the grade
The prognostic grade is stable. It would only fall if large, well conducted cohorts began failing to replicate the association, or if the relationship proved to be an artefact of measurement in a way not previously appreciated. Neither looks likely.
The interventional grade would move to B if randomised trials of structured training in adults reported reductions in hard clinical endpoints, with the effect tracking the change in measured fitness. Trials with cardiovascular endpoints in specific patient groups are the most likely source of such evidence, and they are more feasible than a general population survival trial.
For a companion measure with a similar structure and a weaker evidence base, see our review of resistance training and mortality.