By Jamal Hossain, PhD — Founder and Principal Statistician · Published
It is usually the first statistical question a research team asks, and often the one with the least satisfying answer: it depends. Not because the question is vague, but because sample size is a consequence of other decisions rather than a starting point. Change the primary outcome and the number changes. Change the design and it changes again.
What follows is how that number is actually arrived at, what information you need before it can be, and the mistakes that most often cause trouble at review.
What actually determines the number
For a study designed to detect a difference, four quantities drive the calculation.
The target difference. This is the difference the study is designed to detect, and it is the input that most affects the answer. It should be justified rather than assumed — in relation to what would be important to patients, clinicians or decision-makers, and to what is realistic for the intervention and outcome concerned. It is not necessarily the absolute minimum clinically important difference: a study may reasonably be designed around a difference that is both important and plausible given what the intervention could achieve. The DELTA2 guidance sets out the recognised approaches to specifying and justifying it, and expects the reasoning to be reported alongside the number.2
Variability. For a continuous outcome, the standard deviation. For a binary outcome, the expected event rate in the control group. For time-to-event data, the expected number of events rather than the number of participants.
The significance level (type I error). Conventionally 5%, two-sided.
Power. The probability of detecting the target difference if it is real. Commonly 80% or 90%, but this is a convention rather than a requirement, and the right choice depends on the consequences of missing a real effect.
Two of these are statistical conventions. The other two are substantive judgements, and they are where the real work lies. A sample size is only as defensible as the assumptions behind it, which is why the assumptions — not just the number — belong in the protocol. Reporting guidelines expect exactly that.1
Hypothesis testing or precision?
Not every study is testing a hypothesis, and not every sample size calculation should be based on power.
If the aim is to estimate something — a prevalence, a diagnostic accuracy, a mean score — the relevant question is not “what power do I have?” but “how precise will my estimate be?” You choose the sample size that makes the confidence interval narrow enough to be useful.
To illustrate: estimating a proportion near 50%, a 95% confidence interval of roughly ±5 percentage points needs about 385 participants; tightening that to ±4 points needs about 600. Nothing is being tested — the question is simply how much uncertainty you can live with.
Precision-based planning is often the more honest framing for observational and descriptive work, and it avoids inventing a target difference for a study that was never designed to detect one.4
Different designs, different calculations
The same research area can require quite different calculations depending on how the study is built:
- Comparing two means. Driven by the target difference and the standard deviation.
- Binary outcomes. Driven by the expected proportions in each group. Rare outcomes need substantially larger samples.
- Time-to-event outcomes. Driven by the expected number of events, not the number of people. A study with long follow-up and few events can be underpowered despite recruiting well.
- Cluster-randomised trials. Randomising practices, wards or schools rather than individuals means outcomes within a cluster are correlated, and the sample must be inflated.
- Longitudinal and repeated-measures studies. Repeated measurements add information, but correlated information; the correlation between measurements matters as much as their number.
- Diagnostic accuracy studies. Usually precision-based, and sized separately for sensitivity and specificity — which means the number of people with the condition often drives recruitment.
Observational studies are not exempt. Reporting guidance for observational research asks explicitly how the study size was arrived at.9 Design choice comes before the arithmetic; if you are still weighing options, that is a study design conversation rather than a calculation.
A worked example
This is an illustration, not a recommendation. The numbers are hypothetical and are here to show the logic, not to be reused.
Suppose a two-arm parallel-group trial with a continuous primary outcome measured once at follow-up.
| Assumption | Value |
|---|---|
| Target difference | 5 points |
| Standard deviation (both arms) | 10 points |
| Significance level | 5%, two-sided |
| Power | 90% |
| Allocation | 1:1 |
The familiar textbook formula uses the normal distribution:
n per group ≈ 2 × (1.96 + 1.2816)² × 10² / 5² = 84.1 → 85
That is useful for seeing what drives the answer, but it is not the calculation to rely on. Because the standard deviation is estimated from the data, the analysis uses a t distribution, and the sample size should too. Solving the t-based calculation gives 86 participants per group, 172 in total, with complete outcome data.3
The difference matters more than it looks. At 85 per group the exact power is 89.99% — just short of the 90% target — while at 86 per group it is 90.3%. Round up; never down.
Now allow for attrition. If 15% are expected not to provide primary outcome data, divide rather than multiply:
86 ÷ (1 − 0.15) = 101.18 → 102 per group
So the recruitment target is 102 per group, 204 in total.
Two things are worth noticing. First, the assumptions do the work: for fixed variability, significance level and power, the required sample size is approximately inversely proportional to the square of the target difference. That is why reducing the target difference from 5 points to 3 points would increase the required sample substantially. Second, dropping from 90% to 80% power would reduce the requirement to 64 per group — which is why power should be chosen deliberately rather than defaulted to.
Any real study should be calculated with appropriate software, using assumptions and a method that match the planned analysis — including any covariate adjustment, interim analyses or allocation ratio other than 1:1. The figures above are a teaching illustration of the logic, not a template.
Allowing for clustering
Where participants are grouped — patients within practices, pupils within schools — outcomes within a cluster tend to be more similar than outcomes across clusters, and the required sample increases. The standard inflation factor is:
design effect = 1 + (m − 1) × ICC
where m is the average cluster size and ICC the intracluster correlation.
With 20 participants per cluster and an ICC of 0.02, that is
1 + 19 × 0.02 = 1.38. Applied to the 172 participants above, that is
about 238 participants, which corresponds arithmetically to roughly 12
clusters of 20.
That arithmetic is only the starting point. It does not by itself establish an adequate design, because a cluster trial must also account for the number of clusters (few large clusters and many small ones are not equivalent), variation in cluster size, attrition of whole clusters, the degrees of freedom available when clusters are few, and the intended analysis. Cluster trials are one of the areas where an apparently small technical detail changes the answer substantially.8
Pilot and feasibility studies are different
A pilot or feasibility study conducted in preparation for a definitive trial has feasibility as its primary purpose. It should not generally be designed or interpreted as a definitive test of effectiveness. Its job is to answer questions such as: can we recruit at the expected rate, is the intervention deliverable, is the outcome measure acceptable, and do the study procedures work?
Effectiveness hypothesis testing is generally inappropriate in this setting, simply because a study sized for feasibility is not adequately powered for it. A non-significant result in a pilot tells you very little, and a significant one is unreliable.
Pilot data can help inform an estimate of variability for the definitive study’s calculation, and that is a legitimate use. But estimates of the standard deviation from small samples are themselves imprecise, and that imprecision should be acknowledged — for example by using a conservative value, or by examining how the required sample size changes across a plausible range. What a pilot does not give is a reliable estimate of the treatment effect: effect estimates from small samples are unstable, and using one as the target difference for the main trial is a common route to an underpowered study.
The CONSORT extension for pilot and feasibility trials sets out how these studies should be designed and reported,5 and Whitehead and colleagues give a principled basis for choosing their size.6
Why “30 participants” is not a general rule
The figure recurs, and it is worth being clear about where it comes from.
It is not a universal sample-size rule, and it has no connection to the research question in a definitive study. Part of its currency comes from a statistical folk memory about the central limit theorem, which concerns the behaviour of sampling distributions, not whether a study can answer a clinical question.
Around 30 per arm has, however, appeared as a recommendation in particular pilot-study contexts, including Lancaster and colleagues’ guidance on pilot design and analysis.7 That is a rule of thumb offered for a specific purpose, not a general standard — and modern practice is to justify a pilot’s size against its specific feasibility objective and the parameter being estimated, rather than reaching for a default figure.
The practical test is straightforward: a defensible sample-size justification should be tied to the study objective, primary outcome, design and the assumptions that drive the calculation — whether that means a target difference and power, or a required level of precision.
Allowing for dropout and missing data
Inflating for expected attrition — dividing by (1 − dropout proportion) — is standard and necessary. But it is an arithmetic adjustment, not a solution to missing data.
Recruiting extra participants preserves power. It does not protect against bias if the people with missing outcomes differ systematically from those without. That is a separate problem, addressed in the analysis plan through a pre-specified strategy for missing data, and White and colleagues set out how to think about it for randomised trials.10
Base the dropout estimate on something: comparable studies, the length of follow-up, the burden of the outcome measure. An unexplained 20% is as arbitrary as an unexplained sample size.
Common mistakes
- Working backwards from a feasible number. Deciding the study can recruit 60 and reverse-engineering a target difference that makes 60 sufficient. Reviewers recognise it.
- Taking the target difference from a small previous study. Small studies that reach significance tend to overstate effects, so this systematically underpowers the new study.
- Post-hoc power. Calculating power from the observed effect after the study has finished adds nothing; the confidence interval already conveys what the data support.
- Ignoring clustering. Analysing cluster-randomised data as if individually randomised overstates precision, in both planning and analysis.
- Multiple primary outcomes. More than one primary outcome raises the question of multiplicity, which affects the significance level and hence the sample size.
- Assuming bigger is always better. Beyond a point, extra participants buy little additional information while consuming resources and exposing more people to study procedures — a consideration with ethical weight, not just a budgetary one.
What to bring to a statistician
A sample size conversation moves quickly if you can supply:
- The primary research question, stated as precisely as you can
- The design — parallel groups, cluster, crossover, cohort, cross-sectional, diagnostic
- The primary outcome, how it is measured and when
- Expected values: control-group mean and standard deviation, or expected event rate — and where those figures come from
- The target difference, and the reasoning behind it
- The planned analysis, including any adjustment or covariates
- Expected attrition, and the basis for that expectation
- Recruitment constraints: available population, sites, timescale, budget
If some of these are unknown, that is normal, and identifying which assumptions the answer is most sensitive to is part of the work. A single number is rarely the most useful output — a small range showing how the requirement changes across plausible assumptions is usually more informative, and more persuasive to a funding committee.
References
- Schulz KF, Altman DG, Moher D. CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials. BMJ 2010;340:c332. doi.org/10.1136/bmj.c332
- Cook JA, Julious SA, Sones W, et al. DELTA2 guidance on choosing the target difference and undertaking and reporting the sample size calculation for a randomised controlled trial. BMJ 2018;363:k3750. doi.org/10.1136/bmj.k3750
- Julious SA. Sample sizes for clinical trials with Normal data. Statistics in Medicine 2004;23(12):1921–86. doi.org/10.1002/sim.1783
- Bland JM. The tyranny of power: is there a better way to calculate sample size? BMJ 2009;339:b3985. doi.org/10.1136/bmj.b3985
- Eldridge SM, Chan CL, Campbell MJ, et al. CONSORT 2010 statement: extension to randomised pilot and feasibility trials. BMJ 2016;355:i5239. doi.org/10.1136/bmj.i5239
- Whitehead AL, Julious SA, Cooper CL, Campbell MJ. Estimating the sample size for a pilot randomised trial to minimise the overall trial sample size for the external pilot and main trial for a continuous outcome variable. Statistical Methods in Medical Research 2016;25(3):1057–73. doi.org/10.1177/0962280215588241
- Lancaster GA, Dodd S, Williamson PR. Design and analysis of pilot studies: recommendations for good practice. Journal of Evaluation in Clinical Practice 2004;10(2):307–12. doi.org/10.1111/j..2002.384.doc.x
- Campbell MK, Piaggio G, Elbourne DR, Altman DG. CONSORT 2010 statement: extension to cluster randomised trials. BMJ 2012;345:e5661. doi.org/10.1136/bmj.e5661
- von Elm E, Altman DG, Egger M, et al. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: guidelines for reporting observational studies. Journal of Clinical Epidemiology 2008;61(4):344–9. doi.org/10.1016/j.jclinepi.2007.11.008
- White IR, Horton NJ, Carpenter J, Pocock SJ. Strategy for intention to treat analysis in randomised trials with missing outcome data. BMJ 2011;342:d40. doi.org/10.1136/bmj.d40