Set your study assumptions
Enter percentages as whole numbers. For example, enter 20 for 20%.
Formula used
The calculator applies a normal approximation for a primary binary predictor in logistic regression. It first converts the target odds ratio into the event probability for the exposed group.
p₁ = (OR × p₀) / [1 − p₀ + (OR × p₀)]
p̄ = q × p₁ + (1 − q) × p₀
N₀ = [zα√{p̄(1−p̄)(1/q + 1/(1−q))} + zβ√{p₁(1−p₁)/q + p₀(1−p₀)/(1−q)}]² / (p₁−p₀)²
The result is inflated for correlation with other predictors using N = N₀ / (1 − R²). It is then checked against the selected events-per-parameter target and inflated again for expected loss.
How to use this calculator
- Choose whether you need a recruitment target or achieved power.
- Enter the reference-group event probability from evidence or pilot data.
- Enter the smallest odds ratio that would be scientifically meaningful.
- Add expected predictor prevalence, covariate correlation, dropout, and model complexity.
- Review the recruitment target, expected events, and sensitivity to your assumptions.
Example planning data
| Input | Example | Why it matters |
|---|---|---|
| Baseline event probability | 20% | Defines expected outcome frequency in the reference group. |
| Target odds ratio | 1.50 | Represents the primary effect worth detecting. |
| Predictor prevalence | 40% | Controls the balance between exposure groups. |
| Predictor correlation | 0.10 | Inflates the sample for overlap with covariates. |
| Loss allowance | 10% | Converts evaluable participants into recruitment needs. |
Planning Logistic Regression Studies
Start with the outcome
Logistic regression links a binary outcome with one or more predictors. It suits events with two possible states. Examples include recovery, failure, purchase, survival, or diagnosis. Good planning starts before any participant is enrolled. Power analysis estimates whether the study can detect a meaningful association. It also exposes designs that are too small.
Choose a meaningful effect
The odds ratio is the central effect measure. An odds ratio above one indicates higher outcome odds. An odds ratio below one indicates lower outcome odds. The calculator converts this value into an exposed-group event probability. That conversion needs a baseline outcome probability. Small baseline probabilities can make events scarce. Scarce events often require larger samples.
Consider group balance
Predictor prevalence also changes the required sample. A binary predictor needs enough participants in both groups. A rare exposure creates an imbalanced design. Imbalance weakens precision and increases the necessary total sample. The same problem occurs when the exposure is extremely common. Balanced groups usually provide more information per enrolled participant.
Set the testing threshold
Statistical power is the chance of detecting the planned effect. Researchers often choose eighty or ninety percent power. Higher power requires more observations. Alpha controls the false-positive threshold. A two-sided alpha evaluates effects in both directions. A one-sided alpha uses fewer observations, but needs strong directional justification.
Protect the primary coefficient
Real models often contain adjustment variables. These variables may correlate with the primary predictor. Collinearity reduces the independent information available for the primary coefficient. The calculator uses an R-squared adjustment for this loss. Larger predictor correlation increases the recommended sample. This approximation is useful during early protocol development. Final analysis plans may require simulation.
Plan for sparse data
Expected dropout should be included before recruitment begins. The calculator inflates the evaluable sample for anticipated losses. It also estimates total expected events. Event counts matter because logistic models can become unstable with sparse outcomes. The events-per-variable target provides a practical warning. It is not an absolute rule. Complex models, interactions, nonlinear terms, and missing data can require more events.
Use the result responsibly
The displayed calculation uses a normal approximation for a binary predictor. It is suitable for planning a single primary logistic coefficient. It does not replace specialized simulation when assumptions are complex. Simulations are preferred for continuous predictors, clustered designs, repeated measures, rare events, or multiple testing. They can model the intended analysis more directly.
Check several scenarios
Use sensitivity checks instead of relying on one scenario. Test a smaller odds ratio. Test a lower event rate. Test higher dropout and stronger predictor correlation. Record the assumptions in the protocol. This creates a transparent recruitment target. It also helps reviewers assess feasibility. A well-planned sample protects time, funding, and participant effort.
Check input definitions carefully. Enter probabilities as percentages, not proportions. Keep the odds ratio clinically meaningful. Choose the number of model parameters, including planned interaction terms. Include known covariates in the correlation estimate. Do not count a predictor twice. Revisit calculations after pilot data clarify assumptions. Document revisions with a dated rationale.
Frequently asked questions
What does statistical power mean?
Power is the probability of detecting the planned association when that association truly exists. A higher target reduces the chance of missing an important effect, but usually requires more participants.
Why does the calculator need a baseline event probability?
The baseline probability converts the odds ratio into an expected exposed-group probability. It also determines how many outcome events the study is likely to observe.
Can I enter an odds ratio below one?
Yes. An odds ratio below one represents lower outcome odds in the exposed group. The calculator uses the absolute probability difference when estimating information.
Should I use a one-sided or two-sided test?
Two-sided testing is usually safer because it allows either direction. Use a one-sided test only when a directional hypothesis was specified before data collection and the opposite direction would not change the decision.
What does predictor correlation represent?
It is the approximate R-squared from predicting the main binary predictor with other model covariates. Higher values mean less independent information for the primary coefficient.
Why is there an events-per-parameter target?
Logistic models can be unreliable when too few outcome events support too many estimated parameters. This check helps flag sparse-data risk, although the best target depends on model complexity and validation methods.
Does this work for a continuous predictor?
Not directly. This approximation is designed for a primary binary predictor. Continuous predictors require assumptions about their distribution, scaling, and association with other covariates. Simulation is often more suitable.
Can I use this for rare outcomes?
Use it cautiously. Very rare outcomes can create sparse cells and unstable coefficients. Consider simulation, penalized estimation, or a different study design when events are uncommon.
Why is the recruitment target larger than the evaluable sample?
The tool inflates the required analysis sample for expected loss. This allows for withdrawal, exclusion, missing outcomes, and other anticipated reductions before the primary analysis.
Does this replace a statistical analysis plan?
No. It provides an early planning estimate. Your final plan should define outcomes, covariates, missing-data handling, interactions, multiplicity, model diagnostics, and any simulation work.
How should I report the calculation?
Report the baseline event probability, target odds ratio, alpha, power, predictor prevalence, covariate correlation, dropout assumption, parameter count, event target, and resulting recruitment total.