A grade attaches to a claim, never to a substance
The most common failure in health writing is grading a thing rather than a statement about a thing. Exercise is not graded. The claim that measured cardiorespiratory fitness predicts mortality is graded, and it receives a different grade from the claim that raising your own fitness lowers your own risk.
So every review states the claim being graded at the top, in a single sentence, before any evidence is discussed. If a review makes several claims they are graded separately or the additional ones are discussed in the body with their own grade stated in words.
This has a consequence readers should expect. A compound can appear with a low grade while the underlying science is excellent, because the human question has not been asked. Grade D is not an accusation of bad science. It is a statement that the claim being made in public has not been tested in people.
Two kinds of claim, graded on different criteria
Not every question can be randomised. Nobody will randomise people to decades of low fitness, and it would be unethical to try. Grading prognostic claims on interventional criteria would mean permanently understating evidence that is as good as evidence about that question can be.
So we distinguish two claim types and state which applies on every review.
Intervention claims assert that doing something changes an outcome. These are graded principally on randomised evidence, because randomisation is the only design that balances the factors nobody measured.
Prognostic claims assert that a measurement predicts an outcome. These are graded on the size, number and independence of the cohorts, the consistency of the association, whether a dose-response relationship is present and orderly, whether the association survives adjustment, and how well the field has addressed reverse causation.
A prognostic grade never licenses an interventional conclusion. Where a review carries a high prognostic grade, it states the interventional grade separately and explicitly, because the gap between the two is where most public confusion lives.
The four bands, in full
| Grade | Intervention claims | Prognostic claims |
|---|---|---|
| A | Adequately powered randomised trials with clinical outcomes, reporting on a pre-registered primary endpoint, replicated by an independent group in a different population, with consistent direction and rough magnitude. | Large cohorts across multiple populations and eras reporting the same association, with an orderly dose-response, surviving adjustment, with reverse causation adequately addressed. |
| B | At least one adequately powered randomised trial with a clinical or robust functional primary endpoint, reporting benefit on the pre-registered analysis, not yet independently replicated. | Multiple large cohorts agreeing in direction, but with a material weakness such as self-reported exposure, inconsistent dose-response, or limited geographic range. |
| C | Human randomised evidence exists but is short, small, or confined to surrogate or intermediate outcomes. Or a large observational literature exists with no randomised test of the claim. | Consistent associations from observational data where reverse causation or confounding cannot be adequately excluded, or where measurement of the exposure is poor. |
| D | Evidence is preclinical, or human data are early phase, uncontrolled, or too conflicted to support the claim. Includes fields where target engagement cannot be demonstrated. | Associations reported but inconsistent, or derived from small or unrepresentative samples, or the measurement itself is not validated. |
Rules that cap a grade regardless of volume
Some features of an evidence base limit the grade no matter how many studies exist. Volume of publication is not a proxy for certainty, and a field can accumulate a hundred papers without answering its central question.
- Surrogate outcomes only. Capped at C. A biomarker moving is evidence about the biomarker. Our explainer on surrogate endpoints sets out why volume does not fix this.
- No demonstrable target engagement. Capped at D. If nobody can show the intervention reached and acted on its target, a null result is uninterpretable and a positive result cannot be attributed.
- Animal evidence only. Capped at D, however strong and however well replicated. Multi-site replication in mammals raises confidence that the animal finding is real, not that it transfers.
- Uncontrolled human reports. Contributes nothing to a grade. Self-selected users reporting their own outcomes carry no information about causation.
- Sponsor-controlled evidence base. Capped at B where all substantive trials are funded or conducted by parties selling the intervention, until independent replication exists.
- Trial duration far shorter than the claim. Capped at C where a claim about years or decades rests on trials of weeks.
What we weigh within a band
Within a band, several things move our confidence without changing the letter. We say so in the body of the review rather than inventing intermediate grades.
Pre-registration and whether the reported primary outcome matches the registered one. Allocation concealment and blinding, especially for subjective outcomes. Attrition and whether analysis kept participants in their assigned groups. Whether the comparator was a fair one. Whether the population resembles anyone the claim is aimed at. Whether the effect, if real, would be large enough to matter to a person. And whether the field has produced null results, because a literature with no null results is a literature with a publication problem.
Funding is recorded and weighed but never used alone to dismiss a study. Disclosed industry funding is a functioning system doing its job. The concern is a whole evidence base controlled by interested parties, which is why that appears as a cap rather than a judgement on any single trial.
What we do not do
We do not score substances out of ten, produce league tables, or rank interventions against one another. Different claims are supported by different kinds of evidence, and a single ranking would imply a comparability that does not exist.
We do not give doses, protocols or personal recommendations. Several interventions covered here are prescription only medicines, and prescribing is a clinical decision made by a doctor who knows the patient. Nothing on this site is medical advice, and this is stated on every page rather than buried in a disclaimer.
We do not name individual studies, authors, journals or numerical results in our reviews. This is a deliberate editorial constraint. Describing a literature qualitatively, by trial phase, population, duration, species, funding type and replication status, is more useful to a reader deciding how much confidence to place in a field, and it removes the temptation to lend false precision to a summary. Readers who want the primary sources are pointed to the institutional indexes where they can be found and read directly.
We do not accept payment, product, hospitality or advance sight of coverage from any party with an interest in a grade. The editorial policy sets out the full position.
Review cycle, and how grades change
Every graded review states, in its own words, what would raise the grade and what would lower it. That statement is written before the evidence arrives, which is the same discipline we ask of trialists. It makes the grade a testable position rather than an opinion.
Reviews are re-examined at least annually and sooner when a substantial trial reports. When a grade changes, the page records that it changed and why, rather than being quietly rewritten.
As of this revision, no claim reviewed in this journal holds grade A on an interventional basis. One prognostic claim does. That distribution is not editorial pessimism. It is what the field looks like when claims are graded against the evidence for the specific thing being asserted, and it is the most useful single fact this journal can offer a reader.