Sample size and what it buys you
Why small studies produce unreliable estimates, what statistical power means in practice, and why a small positive study should not raise your confidence much.

Sample size buys precision. A small study produces an estimate with wide uncertainty around it, which means the true effect could be much larger or much smaller than the number reported, or absent.
Two consequences follow that are less intuitive. A small study that finds nothing has usually not shown there is no effect, only that it lacked the ability to detect one. And a small study that finds a large effect is more likely, not less, to be overstating it, because in a small study only a large apparent effect reaches significance at all.
What a study is estimating
A clinical study measures an effect in a sample and uses it to estimate the effect in the wider population. The estimate carries uncertainty, and that uncertainty shrinks as the sample grows. This is the whole of the intuition, and everything else follows from it.
The uncertainty is usually expressed as a confidence interval: a range within which the true value is likely to sit. A wide interval means the study is compatible with many different truths. When only a point estimate is quoted, the uncertainty has been hidden, and the point estimate is the least informative part of the result.
Power, and what a null result means
Statistical power is the probability that a study will detect an effect of a given size if that effect is real. Power rises with sample size and with the size of the effect being sought. Underpowered studies are common in small clinical fields, and their null results are frequently misread.
The correct reading of a null result from an underpowered study is that the study did not detect an effect, which is different from showing there is none. The distinction is not pedantic. Absence of evidence and evidence of absence require different study designs, and only the second requires adequate power to rule an effect out.
A small study reporting a large improvement provides strong support for a treatment.
- Proposed mechanism
- A large observed effect indicates a large real effect.
- What has been shown
- In small samples, estimates vary widely by chance, and only large apparent effects reach conventional significance thresholds. The consequence, sometimes described as an inflation effect, is that statistically significant results from small studies systematically overstate effect size. This is a property of the statistics rather than a criticism of any particular study.
- Highest level reached
- Not shown
- Main confounders
- Selective reporting amplifies the problem, because small null studies are less likely to be published. Multiple endpoints without adjustment increase the chance of a spurious positive.
GradeNOT SUPPORTED
What would change thisReplication in an adequately powered independent study. If the effect is real, a larger study will find it, usually smaller than the first estimate. If it is not, the larger study is where that becomes apparent.
Why small positive studies mislead in a predictable direction
This is the most useful statistical idea for a lay reader of this literature, so it is worth stating carefully.
In a small study, chance variation is large. To reach statistical significance, an observed effect must be large relative to that variation. So among small studies, only those that happened to observe a large effect will report a significant result. The published small positive study is therefore drawn from the upper tail of possible outcomes, and its effect estimate is biased upward.
This predicts a pattern seen repeatedly across medicine: an early small study reports a striking effect, and later larger studies report a smaller one or none. That pattern is not evidence of misconduct. It is what the statistics predict.
| Study size | A positive result means | A null result means |
|---|---|---|
| Very small | Possibly a real effect, probably overstated | Very little, the study likely lacked power |
| Moderate | Reasonable support, effect size still uncertain | Some evidence against a large effect |
| Large and well designed | Good support with a usable effect estimate | Reasonable evidence against a clinically useful effect |
Multiple endpoints
A study measuring many outcomes has many chances to find something significant. Where a primary endpoint is specified in advance and the analysis is pre-registered, this is controlled. Where outcomes are reported without a pre-specified primary, a significant finding among many may be chance.
The practical check is to ask what the primary endpoint was and whether it was stated before the data were collected. Pre-registration makes this checkable, and its absence in much of the aesthetics literature is a real limitation, discussed further in endpoints that mean something.
Why we do not quote effect sizes
This publication describes the shape of evidence rather than reproducing figures. That policy exists because a number quoted without its study, its population, its interval and its design is worse than no number: it carries an air of precision it has not earned, and it propagates.
If you need a figure, go to the paper, find the confidence interval, note the sample size, and check whether the endpoint was pre-specified. That is more work than reading a number in an article, and it is the only version of the exercise that produces a reliable belief.
What adequate looks like in this field
We will not name a required participant number, because the answer depends on the effect being sought, the variability of the endpoint and the design. What we will say is that studies in aesthetic regenerative treatments are frequently in the range where the considerations above dominate, and that a reader encountering a small positive study should treat it as a reason to look for replication rather than as a reason to conclude.
That is the whole practical lesson: small positive studies generate hypotheses. Replication in larger, independent, adequately powered studies is what turns a hypothesis into a finding, and it is what moves a claim up our grading scale.
Questions readers ask
Why does sample size matter?
Because it determines the precision of the estimate. A small study produces a result with wide uncertainty, meaning the true effect could be much larger, much smaller or absent.
What is statistical power?
The probability that a study will detect an effect of a given size if that effect is real. Low power means a study can miss a real effect, so a null result from an underpowered study does not show there is no effect.
Why do small studies overstate effects?
Because chance variation is large in small samples, and only large apparent effects reach significance. Published small positive studies are therefore drawn from the upper tail of possible outcomes and their estimates are biased upward.
What is a confidence interval?
A range around an estimate indicating its uncertainty. A wide interval means the data are compatible with many different true values, often including no effect. A point estimate quoted without an interval hides this.
Why do you not quote effect sizes?
Because a figure quoted without its study, population, interval and design carries unearned precision and propagates. We describe the shape of the evidence and direct readers to the primary source for numbers.