What a surrogate endpoint is
Clinical trials would ideally measure the outcomes that matter: whether people died, developed a disease, lost their independence, or felt better. Those outcomes are slow to appear and expensive to count. Measuring whether a drug lowers a laboratory value is fast and cheap.
So trials substitute. Blood pressure stands in for stroke. A cholesterol fraction stands in for heart attack. Bone density stands in for fracture. Tumour shrinkage stands in for survival. A biological age score stands in for ageing itself.
The substitution is not illegitimate. Without it, drug development would be impossibly slow, and many important treatments reached patients years earlier because a surrogate allowed a smaller, faster trial. The problem is what happens when the substitution is forgotten, and a result about the marker is reported as a result about the disease.
Why the substitution fails
For a surrogate to be valid, two conditions must hold. The marker must lie on the causal path to the outcome, and the treatment's entire effect on the outcome must run through the marker. The second condition is the one that fails, and it fails often.
A drug can lower a marker by a route that bypasses the disease process entirely. Two drugs can move the same marker equally and have opposite effects on patients. And a drug can improve the marker while causing harm through a mechanism nobody was measuring, so the net effect on the person is negative even though every measured number improved.
The most instructive cases in medical history follow exactly this pattern: a physiological abnormality is identified, a treatment corrects it, the correction is confirmed in trials, the treatment is adopted widely, and a later trial with clinical endpoints finds that patients did worse. The reasoning was sound at every step except the assumption that fixing the number fixed the problem.
Validation is possible in principle. It requires showing across multiple treatments and populations that the change in the surrogate reliably predicts the change in the outcome. Very few surrogates have been validated to that standard, and validation is specific to a treatment class rather than general to the marker.
Why this is the central problem in ageing research
Ageing research has a structural incentive towards surrogates that is stronger than in any other field. The outcome of interest is measured in decades. No funder will support a forty year trial, no company can wait for it, and no researcher's career survives it.
So the field runs on markers: blood metabolite concentrations, inflammatory panels, epigenetic clocks, telomere length, grip strength, gait speed, composite biological age scores. Some of these are respectable predictors in cohort studies. Almost none has been validated as a surrogate in the technical sense, which would require showing that changing it changes the outcome.
This is why our grading scheme treats surrogate-only evidence as capped at grade C regardless of how consistent it is. A hundred trials showing that a compound moves a marker do not add up to one trial showing that it helps a person. They add up to strong evidence about the marker.
The commercial consequence is direct. A supplement that shifts a biological age readout can be sold on that basis, and the reader has no way to know whether the shift means anything. Our review of NAD precursors is the clearest example in this journal of a literature that is entirely surrogate-based and is marketed as though it were not.
How to spot the substitution
Ask one question of any health claim: what was actually measured, and in whom?
If the answer is a laboratory value, a scan result, a score or a composite index, you are looking at a surrogate. That is not a reason to dismiss the finding. It is a reason to describe it accurately: the treatment changed the marker, and the effect on health is unknown.
Some specific signals are worth learning. Language that describes a mechanism rather than an outcome, phrasing such as supports, promotes, optimises or targets, usually indicates that no outcome was measured. Timescales that are implausibly short for the claimed benefit indicate the same. Claims framed around a score that the seller also supplies the test for should be treated with particular care.
Regulators face the same problem in a more consequential form. Approving a treatment on a surrogate endpoint gets it to patients faster, and sometimes the confirmatory trial with clinical endpoints is delayed, or never completed, or completed and does not confirm. Appraisal bodies weigh surrogate-based evidence differently from outcome-based evidence for precisely this reason.
None of this means markers are useless. They guide which hypotheses deserve an expensive trial, they detect harm early, and they are often the only feasible measurement. The discipline is simply to never let the sentence slide from the marker moved to the person benefited. That slide is where nearly all the damage happens. Our guide to reading a clinical trial covers where in a paper to check which was measured.