Research

Who was still in the study at the end

Enrollment and analysis are two different numbers, and the second one is smaller. The rules behind real trials say who is supposed to count, and researchers have measured what the missing people could do to an answer.

By Nora Castellan, Standards Editor

The number that gets quoted is rarely the number who enrolled

A study starts with a group of people and ends with a result calculated from a group of people. They are not the same group.

Some stop taking the treatment. Some leave the study. Some are still enrolled but were never measured at the end.

How those people are handled decides the answer, and it can decide it entirely. This is a choice made by whoever ran the analysis.

A page citing a study almost never mentions it. The documents behind real trials do, at length, because the choice is where the argument is.

Two different exits, often described as one

Stopping the treatment and leaving the study are separate events, and a 2021 addendum to the statistical guidance FDA issues for trial design draws the line explicitly.

It distinguishes discontinuation of randomized treatment from study withdrawal, and it says the two create different problems.

Discontinuation of the assigned treatment is treated as an event during the trial that the trial’s stated question has to account for.

Study withdrawal is different. It gives rise to missing data, to be addressed in the statistical analysis.

The distinction matters to a reader for one simple reason. Someone who stopped the treatment can still be followed and measured.

Someone who left cannot. A study that conflated the two has hidden which of its problems it actually had.

The default rule is that everyone assigned still counts

The design guidance sets out a principle with a name that sounds bureaucratic and is not.

The intention-to-treat principle implies that the primary analysis should include all randomized subjects.

The guidance calls the practical version of this the full analysis set, described as the set that is as complete as possible and as close as possible to that ideal.

It gives the reason in one line. Preservation of the initial randomization in analysis is important in preventing bias and in providing a secure foundation for statistical tests.

Exclusions are allowed, and the guidance keeps the list short: failure to satisfy major entry criteria, failure to take at least one dose, and the lack of any data after assignment.

It then adds four words that carry the weight. Such exclusions should always be justified.

One of its conditions on excluding people connects directly to how a study was run. All subjects must receive equal scrutiny for eligibility violations.

The guidance notes in passing that this may be difficult to ensure in an open-label study, which ties the analysis question back to the design question.

The other set, and why it flatters

There is a second, smaller group, and the guidance is unusually direct about what using it does.

The per protocol set is described as the subjects who are more compliant with the protocol, and it lists names the same group travels under.

Those names are the valid cases, the efficacy sample, and the evaluable subjects sample. All three sound like quality control.

Membership turns on completing a prespecified minimum exposure, having measurements of the primary variable, and the absence of major protocol violations.

The guidance states the consequence plainly. Using this set may maximize the opportunity for a new treatment to show additional efficacy in the analysis.

Then it names the mechanism, and this is the sentence to keep. The bias, which may be severe, arises from the fact that adherence to the study protocol may be related to treatment and outcome.

Read that twice. People who stuck with a treatment may have stuck with it because it was working, or because it was not making them ill.

Analyzing only them compares a self-selected group against everyone else, which is a different comparison from the one the study set up.

The honest way is to show both

The guidance does not ask a trial to pick one and hide the other.

In confirmatory trials, it says, it is usually appropriate to plan both analyses so that any differences between them can be the subject of explicit discussion.

It explains what agreement between them buys. When both lead to essentially the same conclusions, confidence in the trial results is increased.

And it attaches a caution to that reassurance. The need to exclude a substantial proportion of subjects from the per protocol analysis throws some doubt on the overall validity of the trial.

For a study trying to show one treatment is better than another, it names the default. The full analysis set is used in the primary analysis, apart from exceptional circumstances.

The reason is stated without decoration: it tends to avoid over-optimistic estimates of efficacy resulting from a per protocol analysis.

Somebody measured what the missing people could do to the answer

A systematic review set out to test how much loss to follow-up could change published conclusions, and it worked on trials that had already reported a positive result.

It searched five leading general medical journals for the years 2005 to 2007, and kept randomized trials reporting a significant primary outcome that mattered to patients.

Of the 235 eligible reports, 31 did not report whether loss to follow-up had occurred at all. That is 13 percent saying nothing about it.

Among those that did report it, the median share of participants lost was 6 percent, with an interquartile range of 2 to 14 percent.

The method used to handle the loss was unclear in 37 studies, which is 19 percent of the set.

The researchers then re-ran the results under different assumptions about what happened to the missing people.

Assuming none of them had the event of interest, 19 percent of the trials were no longer significant. Assuming all of them did, 17 percent were no longer significant.

Under what the authors call more plausible assumptions, between none and 33 percent of trials lost significance depending on the assumption applied.

Their own conclusion is stated carefully, and so is the scope. Plausible assumptions about patients lost to follow-up could change the interpretation of results of randomized trials published in top medical journals.

Every one of those percentages belongs to that set of trials in those journals in those years. It is not a claim about peptide research and not a claim about any seller.

Why nobody can just fix it afterward

A reasonable reaction is to ask why the analysis does not simply fill the gaps. The guidance answers that, and the answer is uncomfortable.

Missing values represent a potential source of bias in a clinical trial, and there will almost always be some.

The guidance says no universally applicable methods of handling missing values can be recommended.

What it asks for instead is a test of how much the answer depends on the choice made.

An investigation should be made into the sensitivity of the results to the method of handling missing values, especially if the number of missing values is substantial.

It also expects the trial to have planned for this before it started, sizing the study for the analysis that includes everybody.

The estimated effect may need to be reduced to allow for the dilution arising from data from patients who withdrew from treatment or whose compliance was poor.

The registry turns the flow into a required table

For studies covered by the federal results-reporting rule, none of this is optional and none of it is buried.

The regulation requires participant flow, defined as information for a table documenting the progress of human subjects through the trial, by arm.

It specifies what that table must contain, including the number of human subjects that started and completed the clinical trial, by arm.

It also requires a description of significant events that occur after enrollment and before people are assigned to a group.

Two further requirements do the real work for a reader, and they use nearly identical wording in two places.

Where the number of participants analyzed differs from the number assigned to that group, the record must carry a brief description of the reasons for the difference.

The same duty applies to the baseline table, where the number measured at the start differs from the number assigned.

So a covered trial with a results record has been made to write down who fell out and why. It is one of the few places that answer is published.

Five things to look for

Find both numbers. How many people were assigned, and how many appear in the result being quoted.

Find the gap’s explanation. A study that reports a difference without a reason has left the most interesting part out.

Ask which analysis the headline came from. Everyone assigned, or only the people who completed the course as written.

Look for both analyses. A study reporting the two and discussing the difference is doing what the guidance asks.

And watch the direction of the loss. If more people dropped out of one group than the other, the reason for that is part of the result, not a footnote to it.

Key takeaways

Frequently asked questions

What does intention-to-treat mean?

That people are counted in the group they were assigned to, whether or not they stuck with it. The statistical guidance FDA issues for trial design states the principle as the primary analysis including all randomized subjects, and explains that preserving the original assignment prevents bias and gives statistical tests a secure foundation. Its practical form is called the full analysis set, kept as close to that ideal as the data allow.

What is a per-protocol analysis, and why is it a warning sign?

It restricts the analysis to participants who were more compliant with the protocol, which the guidance also calls the valid cases, the efficacy sample or the evaluable subjects sample. The guidance says using it may maximize the opportunity for a treatment to show additional efficacy. It then names the mechanism, which is that the bias may be severe because adherence to the protocol may be related to both treatment and outcome. It is not forbidden, but it is not a neutral choice.

Is stopping the treatment the same as leaving the study?

No, and a 2021 addendum to that guidance separates them. Discontinuing the assigned treatment is an event the trial’s question has to account for, and someone who stops can still be followed and measured. Withdrawing from the study creates missing data, which has to be handled in the analysis. A report that blurs the two has concealed which problem it had.

How much can dropouts change a published result?

It has been measured on trials that already reported a positive finding. A systematic review of 235 reports in five leading general medical journals from 2005 to 2007 re-ran their results under different assumptions about the missing participants. Nineteen percent were no longer significant assuming none of the missing had the event, and 17 percent assuming all did. Under the authors’ more plausible assumptions the figure ranged from none to 33 percent.

Can the missing data just be filled in?

Only with an assumption, and the assumption has to be declared. The guidance says missing values are a potential source of bias, that there will almost always be some, and that no universally applicable method of handling them can be recommended. What it asks for is a check of how sensitive the result is to the method chosen, especially where the number of missing values is substantial.

Where would I find the dropout numbers for a study?

In the results record, if the study is covered by the federal reporting rule. That rule requires a participant flow table documenting progress through the trial by arm, including the number who started and completed. It also requires a brief description of the reasons whenever the number analyzed differs from the number assigned. That is a rare case of the question being answered in public by requirement rather than by choice.

Sources

Each document below is named as it names itself, with the date printed on that document rather than the day it was read.

  1. E9 Statistical Principles for Clinical Trials — Guidance for IndustryU.S. Food and Drug Administration, CDER and CBER, September 1998
  2. E9(R1) Statistical Principles for Clinical Trials: Addendum: Estimands and Sensitivity Analysis in Clinical Trials — Guidance for IndustryU.S. Food and Drug Administration, CDER and CBER, May 2021
  3. 42 CFR 11.48 — What constitutes clinical trial results information?Office of the Federal Register, eCFR, September 2026
  4. Potential impact on estimated treatment effects of information lost to follow-up in randomised controlled trials (LOST-IT): systematic reviewBMJ, volume 344, article e2809 (PubMed identifier 22611167), May 2012