⚠️

For educational and reference use only.

Always verify all calculations independently before clinical application. Sample size calculations for regulatory submissions, IRB applications, or grant proposals must be verified by a qualified biostatistician. Results from this tool are approximations based on normal approximation methods.

Clinical Trial Sample Size Calculator Power Analysis for Research — RCT, Two Means, Chi-Square & Survival

Calculate sample sizes for four trial designs: two proportions (RCT), two means, chi-square, and survival analysis. Includes APA citation generator.

Two Proportions (RCT / Cohort)

Between 0 and 1 (e.g. 0.30 = 30%)

Two Means (Continuous Outcome)

Chi-Square (Categorical)

df = groups − 1 =

Survival Analysis (Log-Rank)

Results

n per group

Total N

APA Methods Paragraph

z-Value Reference

Parameter z
α=0.05, two-sided1.960
α=0.025, two-sided2.241
α=0.01, two-sided2.576
Power 80% (β=0.20)0.842
Power 85% (β=0.15)1.036
Power 90% (β=0.10)1.282

Cohen's Effect Size Reference

Size d h w
Small0.20.20.1
Medium0.50.50.3
Large0.80.80.5

Quick n Reference

n per arm, 80% power, α=0.05 two-sided, two means

Cohen's d n/arm
0.2 (small)394
0.3176
0.4100
0.5 (medium)64
0.645
0.734
0.8 (large)26

Frequently Asked Questions

Statistical power (1−β) is the probability that a study will detect a true effect when one actually exists — i.e., correctly reject the null hypothesis. A power of 80% means that if the treatment truly works, the study has an 80% chance of detecting it. Conversely, there is a 20% chance of a false negative (Type II error). Higher power requires a larger sample size. 80% is the conventional minimum; 90% is used when the consequences of missing a true effect are severe.
A Type I error (false positive, α) occurs when you conclude there is an effect when there is none — the null hypothesis is rejected incorrectly. The significance level α controls this: at α=0.05, you accept a 5% chance of a Type I error. A Type II error (false negative, β) occurs when you miss a real effect — the null hypothesis is not rejected when it should be. Power (1−β) controls the Type II error rate. The two error rates are inversely related: reducing α (being stricter) increases the risk of Type II errors unless sample size is increased.
The 80% convention was popularised by Jacob Cohen in his 1977 textbook Statistical Power Analysis for the Behavioral Sciences. Cohen proposed that a 4:1 ratio of Type II to Type I error risk (β/α = 0.20/0.05 = 4) was reasonable for most research contexts. It became entrenched in regulatory guidance and institutional practice. 80% power does not mean a study is underpowered — it is a deliberate balance between the cost of a larger sample and the risk of missing a real effect.
A two-sided test (two-tailed) is appropriate when you want to detect an effect in either direction — either the treatment is better or worse than control. This is the default in most clinical trials. A one-sided test is used only when: (1) you have strong a priori evidence that the effect can only go in one direction, and (2) you would take the same action regardless of which direction a significant result occurred. Regulatory agencies (FDA, EMA) generally require two-sided tests for confirmatory trials. One-sided tests are occasionally used in non-inferiority trials.
Effect sizes are standardised metrics of the magnitude of a difference. Cohen's conventions for d (two means): small = 0.2, medium = 0.5, large = 0.8. For h (two proportions): small = 0.2, medium = 0.5, large = 0.8. For w (chi-square): small = 0.1, medium = 0.3, large = 0.5. These are rough guidelines — the clinically meaningful effect size for your specific domain (minimal clinically important difference, MCID) is far more important than generic labels.
Intent-to-treat (ITT) analysis includes all randomised participants in the groups to which they were assigned, regardless of compliance or withdrawal. It preserves the benefit of randomisation and provides the most conservative, real-world estimate of efficacy. Per-protocol (PP) analysis includes only participants who completed the study as planned. PP analyses tend to show larger treatment effects but are subject to bias because non-compliers may differ systematically. Regulatory agencies require ITT as the primary analysis for confirmatory trials; PP serves as a supportive analysis.
If you expect a dropout rate of d% (expressed as a decimal), inflate the calculated n by dividing by (1−d). For example, if you need 100 participants per arm and expect 20% dropout: 100 ÷ 0.80 = 125 participants per arm need to be enrolled. This ensures that even with dropouts, you have 100 evaluable participants providing adequate power. Dropout assumptions should be based on pilot data or literature from similar populations and interventions.
An adaptive design allows pre-planned modifications to the trial based on interim data, without compromising its integrity. Common adaptations include: sample size re-estimation (if the interim data suggests a different effect size than assumed), response-adaptive randomisation (allocating more patients to the better-performing arm), seamless phase II/III designs, and stopping for efficacy or futility (group sequential designs). Adaptive designs can be more efficient than fixed designs but require complex statistical methodology, prospective planning in the protocol, and regulatory engagement before trial start.
Pilot studies are designed to assess feasibility, not to test the primary hypothesis. Common guidance: 12 participants per group (24 total for a two-arm study) or a total of 30–50 participants provides enough information to estimate standard deviation, recruitment rates, and protocol adherence — the inputs needed to power the main trial. The "rule of 12" and "rule of 30" are pragmatic guidelines rather than formula-based. Some methodologists use a precision-based approach: n = 50 participants provides a 95% CI for the proportion that has a width of approximately ±14%.
A complete IRB sample size justification should state: (1) the primary outcome and its expected values in each group; (2) the expected standard deviation (for continuous outcomes) or the proportion in each group (for binary outcomes); (3) the significance level (α) and its justification (usually 0.05, two-sided); (4) the statistical power (usually 80% or 90%) and its justification; (5) the sample size formula or software used (cite the reference); (6) the calculated sample size per group and total; (7) any inflation for dropout; and (8) any adjustment for multiple testing. Our APA citation generator produces a draft methods-section paragraph you can adapt.

Clinical Trial Sample Size: Power Analysis Theory and Practice

Determining the correct sample size is one of the most consequential decisions in clinical research design. An underpowered study wastes resources, exposes participants to interventions without a reasonable chance of producing actionable knowledge, and may provide misleadingly negative results. An overpowered study exposes more participants than necessary to experimental risks. Regulatory agencies — including the U.S. Food and Drug Administration (FDA) and the European Medicines Agency (EMA) — require prospective sample size justification in all pivotal trial protocols.

The Foundations of Power Analysis

A sample size calculation requires five inputs: the primary outcome measure and its expected distribution, the effect size (the difference you want to detect), the significance level (α), the desired power (1−β), and the allocation ratio between arms. Each of these is a deliberate research decision, not a statistical artifact.

Effect size is the most important and most frequently misspecified input. It should represent the minimum clinically important difference (MCID) — the smallest effect that would change clinical practice or patient management — not the largest effect that might plausibly occur. Inflating the assumed effect size to achieve a smaller sample size leads to a trial that is underpowered for the true treatment effect, even if statistically significant findings were observed in the pilot work.

Significance level controls the Type I error rate. The conventional α=0.05 (two-sided) means you accept a 5% probability of concluding an ineffective treatment works. In trials with multiple comparisons (multiple primary outcomes, multiple interim analyses, or multiple arms), the familywise error rate must be controlled — through Bonferroni correction, hierarchical testing, or group sequential boundaries — to maintain the overall Type I error at α.

Power controls the Type II error rate. The 80% convention (introduced by Jacob Cohen) remains standard, though 90% is increasingly expected for confirmatory Phase III trials by regulatory bodies, particularly for trials where the treatment addresses a serious condition and missing a real effect would be costly.

Two Proportions: The Most Common RCT Design

In a superiority randomised controlled trial with a binary primary endpoint — responder rate, mortality, complication rate — the sample size formula using the normal approximation is:

n = (z_α/2 + z_β)² × [p1(1−p1) + p2(1−p2)] / (p1−p2)²

Where p1 and p2 are the expected proportions in the control and treatment groups, z_α/2 is the critical value for the desired α (1.96 for α=0.05, two-sided), and z_β is the critical value for the desired power (0.842 for 80% power). The formula calculates n per arm for a 1:1 allocation; unequal allocation requires adjustment.

Cohen's h (the arcsine transformation effect size for proportions) provides a standardised measure: h = 2 × arcsin(√p1) − 2 × arcsin(√p2). This transformation stabilises the variance across different baseline proportions, making h more directly comparable across studies than the raw difference p1−p2.

Two Means: Continuous Primary Outcomes

For trials where the primary outcome is continuous (blood pressure, HbA1c, pain score, cognitive test score), the standard formula for equal-sized groups is:

n per group = 2(z_α/2 + z_β)² × σ² / (μ1−μ2)²

This simplifies to n = 2(z_α/2 + z_β)² / d², where d is Cohen's d (the standardised effect size: mean difference divided by the pooled standard deviation). The critical challenge in applying this formula is obtaining a reliable estimate of σ. Sources include: prior studies in the same population, systematic reviews and meta-analyses, pilot data from your own institution, or the literature's reported standard deviations for the same outcome measure. Using an underestimate of σ leads to an underpowered study; using an overestimate leads to a larger-than-necessary sample.

Survival Analysis: Number of Events, Not Participants

For time-to-event outcomes (overall survival, progression-free survival, time to readmission), sample size determination differs fundamentally from other designs: what matters is the number of events (deaths, progressions, etc.), not the number of participants enrolled. The required number of events for a log-rank test is:

E = 4(z_α/2 + z_β)² / [ln(HR)]²

Where HR is the target hazard ratio (the ratio of median survival times, approximately). Converting from events to participants requires assumptions about event rates during accrual and follow-up — this is why our survival mode asks for accrual and follow-up periods. Shorter follow-up relative to accrual means lower event rates during accrual, requiring more participants to achieve the target events.

Common Mistakes in Sample Size Calculation

Using observed pilot data as the assumed effect size. Pilot studies are inherently imprecise; observed differences in a 30-person pilot can vary widely around the true population difference. Using the pilot's observed difference as the target in a power calculation leads to optimistic (too-small) sample sizes and underpowered trials. Instead, use the MCID and treat the pilot's SD as the estimate of variability.

Failing to account for multiple comparisons. If a trial has three co-primary endpoints tested at α=0.05 each, the overall false-positive rate exceeds 5%. Each additional endpoint requires Bonferroni or other correction. The corrected α for each test then feeds into a higher required sample size.

Ignoring correlation in crossover designs. Crossover trials, where each participant receives both treatments, benefit from within-subject correlation — subjects serve as their own controls. The sample size formula for crossover trials includes a (1−ρ) term where ρ is the intra-subject correlation coefficient. Not accounting for this leads to gross overestimation of required participants.

Underestimating dropout. In oncology trials and long-duration trials, dropout rates of 20–30% are common. Failing to inflate the sample size for expected dropout leads to an evaluable population that is smaller than the powered estimate.

Regulatory Expectations: FDA and EMA

Both FDA and EMA guidelines for clinical trial design require sample size justification in the statistical analysis plan (SAP). FDA guidance (FDA E9, E10, and E9(R1) addendum) emphasises the importance of pre-specifying the estimand — the precise treatment effect being estimated, accounting for intercurrent events (treatment discontinuation, rescue medication use) — and linking the sample size to the power for detecting that estimand.

EMA's guideline on the investigation of subgroups in confirmatory clinical trials notes that subgroup analyses must be powered separately if they are to be confirmatory; otherwise, they are exploratory and hypothesis-generating only. This has direct implications for sample size when the label claim is intended for a specific population subset.

For adaptive trials, both agencies require that the adaptation rule be pre-specified in the protocol, that appropriate type I error control be demonstrated through simulation or analytical methods, and that an independent data monitoring committee (IDMC) oversee interim analyses to preserve the trial's integrity and the blind of the sponsor's analysis team.

Related Tools

APGAR Score Calculator
Interactive APGAR newborn assessment tool with built-in timer, auto-generated charting documentation, and clinical interpretation. Used by OB nurses and midwives.
Use Tool →
Nurse Medication Dosage Calculator
Weight-based dosing, IV drip rates, and reconstitution calculations for nurses. Drug library with 200+ medications, pediatric dosing, and renal adjustment.
Use Tool →

Full statistical analysis with IBM SPSS

Try SPSS →

User Reviews

Loading reviews…

Write a Review

Reviews are moderated and published within 24 hours.

Send Feedback

Found a bug? Wrong result? Have a suggestion? We read every message.