Longevity Magazine A review journal of healthspan, preventative medicine and ageing
Explainer 03 · Method

Biological age tests: what they measure

The output is a number of years. What produced it is a statistical model trained to predict something, and which something matters enormously.

Method and interpretationLast checked 31 July 2026Not medical advice

In short

A biological age test applies a statistical model to a biological sample and returns a number expressed in years. The model was trained to predict something, and what it was trained to predict determines what the result means. Clocks trained to predict chronological age mostly measure how well the model recovers your date of birth. Clocks trained to predict mortality or disease carry more information. Neither has been shown to respond usefully to intervention, test-retest reliability is a recognised weakness, and no biological age result currently supports a clinical decision.

What these tests actually do

The most common form measures DNA methylation: small chemical marks attached to DNA at specific positions, which change in patterned ways across the lifespan. A test measures the methylation state at a set of positions and feeds those values into a statistical model, which returns a single number described as biological age.

The key point is that the model is a prediction machine, and it was trained on something. It learned which combination of methylation values best predicts a target variable in a training dataset. Everything the output means derives from what that target was.

Other tests exist alongside methylation clocks. Some use blood biochemistry panels combined into a composite score. Some measure telomere length. Some use proteins or metabolites. All share the same structure: measure something, put it through a model, return an age. All share the same interpretive question.

What a biological age test measures, and where the inference stops.Sample takenModel appliedAge in years returnedHealth outcome changes
FigureWhat a biological age test measures, and where the inference stops.

What the model was trained to predict

Two broad generations exist and the distinction is more important than any difference between brands.

First generation clocks were trained to predict chronological age. The training target was the date on the birth certificate. Such a model is judged by how closely it recovers known ages, and a good one does so quite closely. The consequence is uncomfortable: the better the model gets at its training task, the less room is left for anything else, because the residual is precisely the part not explained by calendar age.

Second generation clocks were trained on health outcomes instead, typically mortality or a composite of clinical measures alongside age. These carry more information about health, because that is what they were built to detect, and they generally outperform the first generation at predicting outcomes in cohort studies.

A commercially sold test may use either, or a proprietary model that is not described. If the test does not say what its model was trained on, no interpretation of the result is possible, and that alone is grounds to disregard the output.

Why two tests disagree, and why the same sample can too

Send a sample to two providers and you may get results that differ by years. This is expected rather than scandalous. Different clocks use different positions, different models and different training targets, so they are not measuring the same construct. Agreement would be the surprise.

More troubling is within-test variability. Test-retest reliability, the extent to which the same sample or the same person measured again produces the same answer, is a recognised weakness of methylation clocks. Technical noise in the measurement, differences in the mix of cell types in a blood sample, sample handling and processing batch all shift the result.

Cell composition deserves particular attention. A blood sample contains a mixture of cell types, that mixture shifts with recent infection, stress, exercise and time of day, and methylation differs between cell types. Some clocks adjust for this and some do not. A result that moved because you had a cold last week is not a result about ageing.

The practical implication is direct. If the measurement error is of similar size to the changes people are trying to detect, then tracking your own score over time is largely tracking noise, and any intervention will appear to work roughly half the time.

What a result can and cannot support

These tests are genuinely valuable in research. At population scale, where individual noise averages out, second generation clocks predict outcomes and have become a useful tool for studying ageing biology. That is a real contribution.

What has not been established is that they are valid surrogate endpoints. Showing that a clock predicts mortality in a cohort is not the same as showing that changing the clock changes mortality. That second demonstration would require an intervention trial with clinical outcomes, and it has not been done for any clock. Our explainer on surrogate endpoints sets out exactly why the first does not imply the second.

For an individual, the position is weaker still. A single result carries measurement error large enough to matter, no clinical guideline attaches an action to it, and there is no evidence-based response to an unfavourable score other than the general advice that would apply anyway. A test whose result cannot change what you do is not a useful test, whatever it costs.

There is also a commercial structure worth naming. Where the same organisation sells the test and the intervention intended to improve the score, the incentive is to sell a measurement that responds to the product. That is not an allegation against any particular company. It is a structural conflict that a careful reader should account for. Our review of NAD precursors discusses a field where this pattern is common.

If a test result would worry you and there is no action attached to it, that is a reasonable ground for not taking it. The measurements with established clinical value in this space, blood pressure, glycaemic measures, lipids, and where indicated cardiorespiratory fitness, are available through ordinary clinical care and come attached to guidance about what to do with them.[1]

References
  1. NHS, on the health checks available through routine United Kingdom care and what their results are used for.
  2. National Institute for Health and Care Excellence, on the standards a predictive test must meet to enter clinical practice.
  3. PubMed, National Library of Medicine, for the primary literature on epigenetic clocks and their reliability.
Frequently asked

Are biological age tests accurate?

Accurate against what? They are reasonably good at predicting the thing they were trained to predict, in populations. As a measurement of an individual's ageing they have substantial technical noise, and test-retest reliability is a recognised limitation. Accuracy is not really the right frame; validity of the construct is.

Why did two tests give me different ages?

Because they are different models, built on different sets of methylation positions and trained on different targets. They are not two measurements of one quantity, so disagreement is the expected outcome rather than evidence that one is faulty.

Can I lower my biological age?

Scores can be made to move. Whether moving the score means anything for health has not been shown, because no intervention trial has demonstrated that changing a clock changes a clinical outcome. Given the measurement noise, a favourable change in one person's score over a few months is as likely to be noise as signal.

Should I buy one?

There is no clinical guideline that attaches an action to a biological age result, and no evidence-based response to an unfavourable score beyond general health advice that applies regardless. If a result would cause worry without changing what you do, that is a reasonable ground to decline. Blood pressure, glycaemic measures and lipids have established meaning and come with guidance attached.

Do doctors use these?

Not in routine practice, and not because of conservatism. There is no validated clinical use, no threshold that triggers an intervention, and no evidence that acting on a result improves outcomes. That is the standard any test has to meet before it enters clinical care.

Sources and further reading

We link to institution-level sources only. This journal names no individual study, author, journal or numerical result, for the reasons set out in the editorial policy.