Endpoints that mean something
Surrogate versus clinical endpoints, why appearance scales are harder than they look, and how endpoint choice decides what a study can claim.

An endpoint is what a study measures to decide whether the treatment worked. A clinical endpoint is something the patient experiences directly. A surrogate endpoint is a measurement expected to correlate with it, such as a molecular marker or an instrument reading.
Surrogates are convenient and can mislead, because a treatment can move a marker without changing anything a person notices. In aesthetics the primary clinical endpoint is appearance, which is difficult to measure reliably, so studies often reach for surrogates or for scales that look objective and are not.
The endpoint decides the claim
A study can only support a claim about what it measured. This sounds trivial and is routinely violated. A study measuring gene expression in biopsies supports a claim about gene expression. A study measuring a hydration instrument reading supports a claim about that reading. Neither supports a claim about how skin looks after six months unless that was also measured.
Endpoints used in this field
Molecular and histological
Gene expression, protein levels, collagen density in biopsy. Objective in the sense of being instrument-measured, and distant from clinical outcome. They also require biopsy, which limits sample size and timepoints. Where they are reported, the important detail is which collagen types and what organisation, not total quantity, for the reasons in what regenerative actually means.
Instrument-measured skin properties
Hydration, elasticity, roughness, colour, thickness by ultrasound. These sit between surrogate and clinical. They are reproducible under controlled conditions and sensitive to those conditions: room humidity, time since washing, time of day and probe pressure all matter. A study using them should describe how conditions were standardised.
Photographic assessment
Standardised photographs assessed by blinded raters on a defined scale. This is closest to what a patient cares about and is highly dependent on standardisation. Lighting, position, expression, camera settings and post-processing must be fixed. Unstandardised before and after images are not an endpoint at all, which is why we do not publish them, as set out in our editorial standards.
Participant-reported outcomes
Satisfaction and self-assessed improvement. Genuinely relevant, since the treatment is elective and satisfaction is part of the point. Highly susceptible to expectation, which makes blinding essential rather than optional when they are used.
An increase in a laboratory marker after treatment demonstrates clinical benefit.
- Proposed mechanism
- The marker tracks the tissue change that produces the visible outcome.
- What has been shown
- Surrogate endpoints are validated by demonstrating that changes in the surrogate predict changes in the clinical outcome, which is a substantial evidential exercise. We are not aware of validated surrogate endpoints for aesthetic outcomes in this field, so markers are being used as surrogates without the validation that would make them one.
- Highest level reached
- Not shown
- Main confounders
- Markers can respond to the procedure rather than the product. Biopsy timing captures a moment in a dynamic process. Single timepoints cannot distinguish transient from durable change.
GradeNOT SUPPORTED
What would change thisValidation work demonstrating that a specified marker change predicts a specified clinical outcome in this population, or studies that simply measure the clinical outcome directly, which in aesthetics is usually feasible.
The timing problem
When an endpoint is measured determines what it can show. Assess too early and you are measuring procedural swelling and hydration. Assess only once and you cannot distinguish a transient change from a durable one.
In this field, where the delivery procedure produces its own effects over weeks, the informative timepoint is beyond the resolution of those effects. Studies assessing at short intervals are measuring the procedure as much as the product. Long follow up is rare, expensive, and the thing most likely to change what we know.
| Endpoint | Distance from what a patient experiences | Main vulnerability |
|---|---|---|
| Gene expression in biopsy | Far | Requires validation to mean anything clinically |
| Collagen density in biopsy | Far | Type and organisation matter more than quantity |
| Instrument-measured hydration | Middle | Highly sensitive to measurement conditions |
| Ultrasound skin thickness | Middle | Operator and probe pressure dependent |
| Blinded photographic assessment | Close | Requires strict standardisation of imaging |
| Participant-reported satisfaction | Closest | Highly susceptible to expectation, needs blinding |
Composite scores and scale drift
Some studies use composite scores combining several measures. Composites can obscure which component moved, and a composite improving because one insensitive component barely changed while another moved a lot is a different result from uniform improvement. Component-level reporting should accompany any composite.
Validated scales exist in dermatology for various indications and are preferable to ad hoc scales invented for a study, because a validated scale has known reliability between raters. A scale created for one study has none, and its numbers look identical to numbers from a validated one.
Durability is an endpoint in its own right
Almost every treatment in this family is sold as a course, repeated at intervals, which means the question a patient is really asking is how long the effect lasts. That is an endpoint, and it is measured by continuing to assess after the treatment stops.
Very few studies in this field do so. Follow up frequently ends at or shortly after the final session, which is the point at which procedural effects, hydration and swelling are all still contributing. A study designed that way cannot distinguish a durable tissue change from a temporary one, and it is the design most likely to produce a favourable-looking result, which is worth noticing when a supplier chooses it.
The remedy is neither expensive nor technically difficult: keep assessing. Follow up costs money and patience and nothing else. Its absence across most of this literature is a choice rather than a constraint, and it is the single feature we would most like to see change.
What we would want to see in this field
A study with a pre-registered primary endpoint that is blinded photographic assessment on a validated scale, at a timepoint beyond the resolution of procedural effects, with instrument measures as secondary endpoints and standardisation conditions described. Nothing about that is technically demanding. It is a matter of design discipline and follow up duration, both of which cost money and neither of which requires new methods.
The rarity of such studies is the reason so many claims in our panels sit at the earlier grades. It is not that the treatments have been tested and failed. It is that the studies which would settle the question have largely not been run.
Questions readers ask
What is an endpoint?
What a study measures to decide whether the treatment worked. A clinical endpoint is something the patient experiences; a surrogate is a measurement assumed to track it.
Why are surrogate endpoints risky?
Because a treatment can move a marker without changing what a person experiences. A surrogate is only reliable once it has been validated by showing that changes in it predict changes in the clinical outcome, which is a substantial exercise.
Are before and after photographs evidence?
Only if standardised for lighting, position, expression, camera settings and processing, and assessed by blinded raters on a defined scale. Unstandardised images are not an endpoint, which is why we do not publish them.
Why does the timing of assessment matter?
Because the delivery procedure produces effects of its own over weeks. Assessing early measures procedural swelling and hydration, and a single assessment cannot distinguish a transient change from a durable one.
What is pre-specification?
Naming the primary endpoint before the data are collected, usually through registration. It prevents the outcome that happened to move from becoming the outcome the study claims to have tested.