Why it matters most
Required sample size grows with the inverse square of the effect. Halve the effect and you need about four times as many participants. For a two-sample t test with a two-sided α of .05 and 80% power:
| Cohen’s d | Participants per group |
|---|---|
| 0.80 | 26 |
| 0.50 | 64 |
| 0.40 | 100 |
| 0.33 | 143 |
| 0.25 | 253 |
| 0.20 | 394 |
A wrong guess is costly. Plan for d = 0.5 when the true effect is d = 0.33, and 64 per group gives you 47% power: closer to a coin flip than to the 80% you planned for.
Every number on this page comes from the StatPower calculator and was checked against R’s pwr package.
1. The smallest effect that would matter
The strongest justification is the smallest effect size of interest: the smallest effect that would change a decision, a practice, or a theory. It doesn’t require guessing the true effect. It asks what size of effect would be worth finding, and plans to detect at least that.
Start in the outcome’s own units, then standardize:
d = smallest meaningful difference ÷ standard deviation of the outcome
Example. A tutoring program is worth adopting only if it raises test scores by at least 5 points, and the test’s standard deviation is 15 (from its published norms). Then d = 5 ÷ 15 = 0.33, and the study needs 143 students per group.
Where the “smallest meaningful” difference can come from:
- A published minimal important difference for your measure (common for health and quality-of-life scales).
- Cost and benefit: the effect a program would need to justify its cost.
- Stakeholders: the change practitioners, funders, or patients say they would notice.
- The precision of the measure: a difference smaller than the measure can reliably detect isn’t a sensible target.
Report where the number came from. “The smallest difference we consider meaningful” is only persuasive with a reason attached.
2. Prior evidence, discounted
Effects from earlier studies of the same intervention and outcome are the next-best basis. A meta-analysis beats a single study, and a single well-powered study beats a pilot. Two cautions apply.
Published effects run large
Studies with larger effects are more likely to reach significance and be published, so the literature overstates typical effects. When 100 psychology studies were replicated, the replication effects averaged about half the size of the originals (Open Science Collaboration, 2015).
One systematic fix is safeguard power (Perugini et al., 2014): plan for the lower end of a 60% confidence interval around the published estimate instead of the estimate itself. A published d = 0.5 from a study with 50 per group gives a safeguard d of about 0.33, and a plan of 148 per group instead of 64. That is more expensive, but it holds up if the original estimate was optimistic.
Pilot studies can’t estimate the effect
Pilots are too small to pin down an effect size (Leon et al., 2011). A pilot with 20 per group that observes d = 0.5 has a 95% confidence interval of −0.13 to 1.13, which is consistent with no effect and with a very large one. Depending on where in that range the true effect sits, the “right” sample size could be 14 per group or far more than 400.
There is also a selection problem: pilots that happen to look promising are the ones that get followed up, and those are disproportionately the overestimates (Albers & Lakens, 2018). Use a pilot for what it does well: recruitment, retention, procedures, and the outcome’s standard deviation. Choose the effect by option 1 or from other evidence.
3. Cohen’s benchmarks, as a last resort
Cohen (1988) proposed conventional values for when nothing better is available:
| Effect size | Small | Medium | Large |
|---|---|---|---|
| d (two means) | 0.20 | 0.50 | 0.80 |
| r (correlation) | 0.10 | 0.30 | 0.50 |
| h (two proportions) | 0.20 | 0.50 | 0.80 |
They are generic labels, not facts about your field. In education research, for example, effects that are small by Cohen’s standards are large compared with what most real-world interventions achieve (Kraft, 2020). Planning for a “medium” d = 0.5 there would badly underpower most studies. If you do use a benchmark, say so, and say why nothing more specific was available.
Getting the right metric
The calculator needs the effect in the form its test uses. The common translations:
Two independent groups: d
d = (mean₁ − mean₂) ÷ pooled standard deviation
Paired or pre–post designs: dz, not d
A paired test’s effect is the mean change divided by the standard deviation of the changes. That depends on how strongly the two measurements correlate (ρ). When the two time points have equal standard deviations:
dz = d ÷ √(2(1 − ρ))
For the same d = 0.5, the number of pairs needed changes a lot with ρ:
| Correlation ρ | dz | Pairs needed |
|---|---|---|
| 0.3 | 0.42 | 46 |
| 0.5 | 0.50 | 34 |
| 0.6 | 0.56 | 28 |
| 0.8 | 0.79 | 15 |
Entering a between-groups d as if it were dz is only right when ρ happens to be 0.5. Base ρ on earlier data from the same measure and time gap.
Correlations: r
Enter the correlation directly. For r = 0.3, α = .05 (two-sided) and 80% power, you need 85 participants.
Two proportions: baseline and change
StatPower asks for the baseline rate and the absolute change, and converts them to Cohen’s h for you. Going from 30% to 40% gives h = 0.21 and 356 per group. If earlier studies report an odds ratio instead, convert it to rates at your expected baseline. (For meta-analysis, an odds ratio can be converted to d as ln(OR) × √3 ÷ π (Chinn, 2000); an odds ratio of 2 corresponds to d ≈ 0.38.)
Mistakes reviewers catch
- Choosing the effect to fit the sample you can get. If the sample is fixed, say so and report the minimum detectable effect instead (the calculator’s MDE mode): with 50 per group and 80% power, a two-sample t test can detect d = 0.57. Then argue whether an effect that large is plausible and meaningful.
- Using an observed effect after the study. Power computed from the effect you just found adds nothing beyond the p-value. Report a confidence interval instead.
- Mixing effect-size types: a between-groups d in a paired design, or a benchmark from one metric applied to another.
- A pilot effect entered at face value (see above).
- Forgetting attrition. Divide the required sample by (1 − expected attrition rate): 143 per group with 15% attrition means recruiting 169 per group.
What to write in your proposal
A defensible justification names the effect, its source, the test, α, power, and the tool. Adapt this:
We powered the study to detect the smallest effect we consider meaningful: a 5-point difference on [measure] (SD = 15; [source]), d = 0.33. With a two-sided α of .05 and 80% power, a two-sample t test requires 143 participants per group (exact noncentral t; StatPower, statpower.org). Allowing for 15% attrition, we will recruit 169 per group.
The calculator writes a reporting sentence for your own inputs; add the “why this effect size” part yourself.
Not sure which effect size you can defend, or designing something the calculator doesn’t cover?
Get a design review ↗References
- Albers, C., & Lakens, D. (2018). When power analyses based on pilot data are biased: Inaccurate effect size estimators and follow-up bias. Journal of Experimental Social Psychology, 74, 187–195.
- Champely, S. (2020). pwr: Basic functions for power analysis (R package version 1.3-0).
- Chinn, S. (2000). A simple method for converting an odds ratio to effect size for use in meta-analysis. Statistics in Medicine, 19(22), 3127–3131.
- Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates.
- Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241–253.
- Lakens, D. (2022). Sample size justification. Collabra: Psychology, 8(1), 33267.
- Leon, A. C., Davis, L. L., & Kraemer, H. C. (2011). The role and interpretation of pilot studies in clinical research. Journal of Psychiatric Research, 45(5), 626–629.
- Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.
- Perugini, M., Gallucci, M., & Costantini, G. (2014). Safeguard power as a protection against imprecise power estimates. Perspectives on Psychological Science, 9(3), 319–332.