Research
What statistically significant does and does not say
The phrase answers one narrow question, and it is not the question most readers think. It says nothing about how large an effect was, how useful it would be, or how many other things were measured first.
The phrase answers one narrow question
A significance test starts from an assumption that the treatment did nothing, and asks how surprising the observed result would be under that assumption.
FDA describes the calculation directly in its guidance on trials with more than one endpoint.
A result is called significant when the probability of observing a result at least as extreme as the study result, assuming there is no true difference, is sufficiently low.
That probability is the p-value. It is a statement about the data given an assumption, and it is not a statement about whether the treatment works.
The guidance is careful about how much the test settles. Rejecting the assumption supports a conclusion that there is a difference between treatment groups.
It does not, in the agency’s words, constitute absolute proof that the assumption is false.
The threshold, in plain numbers
The threshold is a convention, and it is a number somebody chose. The guidance prints the conventional values and what they buy.
For a two-sided test at the usual level, the probability of falsely concluding the drug differs from the control in either direction, when no difference exists, is no more than 5 percent.
FDA puts that in ordinary words as 1 chance in 20. For the one-sided version it is no more than 2.5 percent, or 1 chance in 40.
Then it adds a condition most summaries drop. These error rates are correct if the statistical test is appropriate.
If there are problems with the test, such as its underlying assumptions not holding, the error rate could be even larger.
It is not a statement about how large the effect was
This is the single most common misreading, and it is the one a marketing page depends on.
A significant result can be tiny. In a large enough study, a difference too small to notice will clear the threshold comfortably.
A non-significant result can be large, and simply measured in too few people to separate from noise.
The statistical guidance FDA issues for trial design asks for the size to be reported alongside, not instead.
It says estimates of treatment effects should be accompanied by confidence intervals whenever possible.
It repeats the point when discussing which methods to choose, noting the need for estimates of the size of treatment effects together with confidence intervals, in addition to significance tests.
The same guidance expects the study to have decided beforehand what size would matter. The difference to be detected may rest on a judgment about the minimal effect which has clinical relevance in the management of patients.
So a page reporting that a result was significant, without saying how big it was, has published the less interesting half.
Measure enough things and one of them clears the bar
The threshold governs one test. Nobody runs one test.
FDA works the arithmetic through in its own guidance, and the numbers are worth reading slowly.
With a single endpoint at the conventional two-sided level, the chance of falsely finding a favorable effect is about 2.5 percent.
With two independent endpoints, each tested at that level, the guidance says the overall error rate in favor of the drug nearly doubles.
For three independent endpoints it is about 7 percent. For ten independent endpoints it is about 22 percent.
The guidance then closes the obvious escape route. The problem is not limited to counting separate endpoints.
Even with a single outcome variable, analyzing multiple facets of it can inflate the same error rate.
It names the facets: multiple dose groups, multiple time points, or multiple subject subgroups based on demographic or other characteristics.
The consequence, in the agency’s words, is that conclusions about whether effectiveness has been demonstrated become unreliable.
Which is why deciding first is the whole game
If the arithmetic above is the disease, prespecification is the treatment, and it is a matter of order rather than effort.
FDA asks sponsors to specify all endpoints and all planned analyses in advance, then apply the adjustments that keep the overall error rate where it was meant to be.
It warns that changing the plan to add analyses can reintroduce the problem, and it draws a hard line about when.
The statistical analysis plan should not be changed after unmasking of treatment assignments and performing statistical analyses.
The design guidance says the same thing about analyses of particular groups within a trial.
In most cases subgroup or interaction analyses are exploratory and should be clearly identified as such.
And it states the consequence in one sentence. Any conclusion of treatment efficacy, or lack of it, or safety, based solely on exploratory subgroup analyses is unlikely to be accepted.
The results that get written up are not a fair sample of the results
There is direct evidence for what happens between a trial and the paper describing it, and it was obtained by reading protocols.
One study compared the protocols of randomized trials with the articles eventually published from them.
It covered 102 trials, 122 published articles and 3,736 outcomes, all approved by two Danish research ethics committees in 1994 and 1995.
Half of the efficacy outcomes and 65 percent of the harm outcomes per trial were incompletely reported.
Statistically significant outcomes had higher odds of being fully reported than non-significant ones, with a pooled odds ratio of 2.4 for efficacy data.
For harm data the pooled odds ratio was 4.7, on a confidence interval running from 1.8 to 12.0.
The comparison against protocols found that 62 percent of trials had at least one primary outcome that was changed, introduced, or omitted.
The authors surveyed the trialists as well. Eighty-six percent of those who responded denied the existence of unreported outcomes, despite the protocols showing otherwise.
That study is about a specific cohort of trials from a specific decade. It is not a finding about any company, and it is not a finding about peptides.
The registry has a field for this, and it is worth opening
When a study is covered by the federal results-reporting rule, the record it has to file separates planned analyses from ones done afterward.
The regulation requires an outcome measure type for each measure, and it gives four permitted values.
They are primary, secondary, other pre-specified, and post-hoc. The last of those is a label the sponsor has to apply to its own result.
The same paragraph sets what must accompany a statistical analysis. Either the p-value and the procedure used, or the estimation parameter, the estimated value and a confidence interval.
It also limits which analyses have to be filed at all, and the first category is the one to notice: analyses pre-specified in the protocol or the statistical analysis plan.
So where a covered trial has a registry record with results, a reader can often see whether the number being marketed was planned or found.
Reading a significance claim in about a minute
Find the size. If the page says a result was significant and never says how large it was, the useful half is missing.
Find the interval. A range around the estimate tells you how precisely it was pinned down, which a bare threshold cannot.
Count the outcomes. Ask how many things were measured, at how many time points, in how many groups.
Ask when it was chosen. A result named in the protocol and a result found afterward are different objects with the same arithmetic.
Then look for the word instead of the number. Significant means the result cleared a threshold, and important means somebody judged it worth having. Only one of those is a statistical term.
Key takeaways
- Significance describes how surprising a result would be if the treatment did nothing; it is not a probability that the treatment works.
- FDA states that rejecting the no-difference assumption does not constitute absolute proof that the assumption is false.
- The conventional two-sided threshold allows a false favorable conclusion about 1 time in 20, and only if the test used was appropriate.
- Significance says nothing about size, which is why guidance asks for the effect estimate and a confidence interval alongside it.
- FDA prints the multiplicity arithmetic: about 7 percent for three independent endpoints and about 22 percent for ten.
- The same inflation arises from analyzing one outcome across multiple doses, time points or subgroups.
- In one cohort of 102 trials, significant outcomes were more likely to be fully reported, and 62 percent of trials changed, added or dropped a primary outcome relative to their protocol.
- A covered trial’s results record labels each outcome primary, secondary, other pre-specified, or post-hoc, which is often the fastest way to see whether a number was planned.
Frequently asked questions
What does statistically significant actually mean?
That the result would be unlikely if the treatment did nothing. FDA describes the test as asking for the probability of observing a result at least as extreme as the one seen. The assumption behind that probability is that there is no true difference, and the result is called significant when the probability is sufficiently low. The agency also notes that rejecting the no-difference assumption does not constitute absolute proof that the assumption is false.
Does significant mean the effect was big?
No, and this is where most claims go wrong. Significance is about how confidently a difference can be told apart from chance, not about its size. In a large study a very small difference will clear the threshold. The statistical guidance FDA issues for trial design asks for estimates of the size of the treatment effect together with confidence intervals. It asks for those in addition to significance tests, precisely because the test alone does not answer the question.
Why does the number of things measured matter?
Because every additional test is another chance for a false positive. FDA works it through in its guidance: with two independent endpoints, the overall error rate in favor of the drug nearly doubles. With three it is about 7 percent, and with ten it is about 22 percent. The same inflation happens when one outcome is analyzed across multiple dose groups, time points or subgroups.
What is a post-hoc result?
An analysis chosen after the data were seen rather than planned in advance. The federal results-reporting rule treats it as a distinct category: each outcome measure in a covered trial’s results record is labeled primary, secondary, other pre-specified, or post-hoc. FDA’s design guidance adds that a conclusion of efficacy or safety based solely on exploratory subgroup analyses is unlikely to be accepted.
If a study is significant, why would the finding not hold up?
Several reasons stack. The threshold itself allows a false positive roughly 1 time in 20 for a two-sided test at the usual level. FDA adds that the stated error rate holds only if the statistical test is appropriate, and could be larger if its assumptions do not. And the more analyses that were run without adjustment, the further the real error rate has drifted from the advertised one.
What does a confidence interval add?
A range instead of a verdict. It shows how precisely the effect was pinned down, so a reader can see whether the plausible values include effects too small to be worth anything. The design guidance says estimates of treatment effects should be accompanied by confidence intervals whenever possible, and the federal results rule accepts an estimated value with an interval as an alternative to reporting a p-value.
Sources
Each document below is named as it names itself, with the date printed on that document rather than the day it was read.
- Multiple Endpoints in Clinical Trials — Guidance for Industry — U.S. Food and Drug Administration, CDER and CBER, October 2022
- E9 Statistical Principles for Clinical Trials — Guidance for Industry — U.S. Food and Drug Administration, CDER and CBER, September 1998
- 42 CFR 11.48 — What constitutes clinical trial results information? — Office of the Federal Register, eCFR, September 2026
- Empirical evidence for selective reporting of outcomes in randomized trials: comparison of protocols to published articles — JAMA, volume 291, pages 2457 to 2465 (PubMed identifier 15161896), May 2004