Research

Why a twelve-person study cannot carry a population claim

A small study is not a small version of a big one. Below a certain size a positive result becomes less likely to be true and more likely to be exaggerated, which is the opposite of how most people read one.

By Nora Castellan, Standards Editor

The shape of the claim is the problem

A page says participants in a study saw an improvement. It does not say how many participants, and the answer is often a dozen.

A twelve-person study is a real piece of research. It can generate a hypothesis, show that something is feasible, or turn up a problem worth chasing.

What it cannot do is support a sentence about what happens to people in general. The claim being made is about a population, and the evidence is about a room.

What drug development uses each size for

The clearest way to see the mismatch is to look at what regulators expect each stage of a trial program to answer.

FDA describes a first-in-human phase as involving 20 to 100 healthy volunteers or people with the condition, running several months, with the purpose of studying safety and dosage.

The next phase involves up to several hundred people who have the disease or condition, runs several months to two years, and looks at effectiveness and side effects.

The third phase involves 300 to 3,000 volunteers with the condition and runs one to four years. FDA states its purpose directly: researchers design these studies to demonstrate whether or not a product offers a treatment benefit to a specific population.

Population claims are what the third phase is for. A study of a dozen people is smaller than the smallest stage of that sequence, and that stage is not asking whether the thing works.

Low power does something worse than miss

Most people know that a small study can fail to detect a real effect. The second consequence is less familiar and matters more when reading a marketing page.

A widely cited analysis of statistical power put it plainly. A study with low power has a reduced chance of detecting a true effect. Less well appreciated, the authors add, is that low power also reduces the likelihood that a statistically significant result reflects a true effect.

Read that twice. In a low-powered field, the positive results are the ones most likely to be wrong.

The same analysis lists the other consequences: overestimates of effect size, and low reproducibility of results.

That work measured average power in neuroscience specifically, and the measured figure belongs to that field. The statistical relationship it describes does not, and it applies wherever studies are small.

Why the number is usually too big as well as unreliable

The exaggeration point deserves its own paragraph, because it explains a pattern people notice without being able to name.

When a study is small, only an unusually large apparent effect will clear the bar for statistical significance. Ordinary-sized true effects do not make it through.

So the results that get published from small studies are systematically the ones where the estimate came out high. The effect size in the paper is larger than the effect size in the world.

This is why a striking early finding so often shrinks when a bigger study repeats it, and why shrinking is the normal outcome rather than a scandal.

The counting problem with rare harms

Size limits what a study can find on the safety side too, and the limit is a matter of arithmetic rather than diligence.

A harm that affects a small share of users will simply not appear in a room of twelve people. Nobody missed it. There was nobody for it to happen to.

Federal advertising guidance builds this in. Where product safety may be a concern, it says a study should be of sufficient size and duration to detect potential side effects.

The reading rule follows. No side effects were reported is a sentence whose meaning depends entirely on how many people were watched and for how long. Without those two numbers it says nothing.

A control group is not optional at any size

Size is not the only thing a small study usually lacks. Most are also uncontrolled, and that removes the comparison the claim depends on.

Federal regulation defines what an adequate and well-controlled study looks like. It uses a design that permits a valid comparison with a control, to provide a quantitative assessment of drug effect.

The regulation explains why the comparison exists at all. The purpose is to distinguish the effect of a drug from other influences, such as spontaneous change in the course of the condition, placebo effect, or biased observation.

A dozen people improving over eight weeks is consistent with the treatment working. It is equally consistent with all three of those. Nothing in the design separates them.

Decided in advance, or decided afterward

The same regulation asks something else of a study report, and it is the question that catches the weakest small studies.

A protocol and report should describe the design precisely, including whether the sample size was predetermined or based on some interim analysis.

It also requires a clear statement of the objectives and a summary of the methods of analysis, in the protocol and in the report. If the protocol did not describe the proposed analysis, the report should explain how the methods used were chosen.

A study that decided how many people to enroll after watching the results, or chose its analysis once the data were in, is a different object from one that committed first. Both can produce a number. Only one of them tested a hypothesis.

What happens when someone runs the bigger study

Federal advertising guidance includes a worked example that is worth reading as a general pattern.

A drink was promoted as proven to promote cardiovascular health. One small human trial had found a significant difference against a placebo drink. A later, larger study found no significant difference on that measure or others. A third large trial also found no difference, though a post hoc look at a subgroup suggested some benefit.

The guidance concludes that, given the totality of the evidence, the claim is unsubstantiated.

Two things in that example travel. The small positive study did not become false; it stayed on the record and was outweighed. And the subgroup found after the fact did not rescue the claim.

What a small study is genuinely good for

None of this makes small studies worthless, and treating them as worthless would be its own reading error.

A small study can establish that something is feasible, that a procedure can be carried out, or that a measurement behaves as expected. It can surface an unexpected signal that justifies a larger trial.

It can also do real damage on the safety side, since a serious harm appearing in a room of twelve is a strong signal precisely because it should have been unlikely.

The asymmetry is the point. A small study is far better at raising a question than at answering one.

Five things to ask before accepting a number

How many people, in each group. A total is not enough when the comparison is what matters.

Compared with what. No control means no way to separate the treatment from time, expectation and observer bias.

Chosen in advance or afterward. Whether the size and the analysis were set before the data arrived.

For how long. A number from eight weeks does not describe a year.

And what happened next. If the study is more than a few years old and nobody has repeated it at scale, that silence is part of the evidence.

Key takeaways

Frequently asked questions

How small is too small?

There is no universal cutoff, but drug development gives a useful scale. FDA describes first-in-human studies as 20 to 100 participants examining safety and dosage. Second-phase studies run to several hundred people with the condition and look at effectiveness and side effects. Third-phase studies involve 300 to 3,000 people over one to four years. FDA states that the third phase is where researchers set out to demonstrate whether a product offers a treatment benefit to a specific population.

Does a small study just risk missing a real effect?

That is the familiar half. The less familiar half is that low power also reduces the likelihood that a statistically significant result reflects a true effect. In other words, among small studies, the positive findings are the ones most likely to be wrong. The same analysis lists overestimated effect sizes and poor reproducibility as further consequences.

Why do impressive early results shrink later?

Because in a small study only an unusually large apparent effect clears the threshold for statistical significance. True effects of ordinary size do not make it through. The positive results that get published from small studies are therefore drawn from the high end of the range, so the published estimate is larger than the real one. Shrinking on replication is the expected outcome, not an anomaly.

What does "no side effects were reported" mean in a small study?

Very little on its own. A harm affecting a small share of users will not appear in a dozen people, because there is nobody for it to happen to. Federal advertising guidance says that where safety may be a concern, a study should be of sufficient size and duration to detect potential side effects. The sentence only carries meaning alongside how many people were observed and for how long.

Is an uncontrolled study of any use?

For raising questions, yes. For settling them, no. Federal regulation defines an adequate and well-controlled study as one whose design permits a valid comparison with a control. The stated point is to distinguish a drug's effect from spontaneous change in the condition, from placebo effect, and from biased observation. Without a comparison group, an improvement is consistent with all of those at once.

Does a later negative study cancel an earlier positive one?

It outweighs it rather than erasing it. Federal guidance works through a case where one small trial found a benefit, a larger study found none, and a third large trial also found none apart from a subgroup identified after the fact. The conclusion was that given the totality of the evidence the claim was unsubstantiated. A body of evidence is weighed, and an after-the-fact subgroup does not rescue a claim.

Sources

Each document below is named as it names itself, with the date printed on that document rather than the day it was read.

  1. Power failure: why small sample size undermines the reliability of neuroscienceNature Reviews Neuroscience, volume 14, pages 365-376 (PubMed identifier 23571845), May 2013
  2. Step 3: Clinical ResearchU.S. Food and Drug Administration, January 2018
  3. 21 CFR 314.126 — Adequate and well-controlled studiesOffice of the Federal Register, Electronic Code of Federal Regulations, August 2026
  4. Health Products Compliance GuidanceU.S. Federal Trade Commission, December 2022