Study Planning Tool
Enter Study Assumptions
Percent values describe diagnostic performance. Alpha is entered as a decimal. Results use a two-sided normal approximation.
Example Data Table
This example matches the prefilled calculator assumptions.
| Input | Example Value | Planning Meaning |
|---|---|---|
| Minimum sensitivity | 80% | Lowest acceptable correct detection rate among cases. |
| Expected sensitivity | 90% | Performance anticipated for participants with the condition. |
| Minimum specificity | 80% | Lowest acceptable correct exclusion rate among controls. |
| Expected specificity | 92% | Performance anticipated for participants without the condition. |
| Prevalence or case share | 30% | About 90 cases and 210 controls in a sample of 300. |
| Alpha and target power | 0.05 and 80% | Decision threshold and design objective for sample planning. |
Formula Used
The calculator applies a two-sided normal approximation. Let P be prevalence, N be total sample size, p0 be the benchmark, and p1 be expected performance.
Expected cases: ND = N × P
Expected controls: NC = N × (1 − P)
Critical value: zα = Φ−1(1 − α / 2)
Endpoint shift: μ = (p1 − p0) / √[p0(1 − p0) / n]
Variance ratio: r = √[p1(1 − p1) / p0(1 − p0)]
Power: 1 − Φ[(zα − μ) / r] + Φ[(−zα − μ) / r]
Required endpoint sample: n = [(zα√p0(1 − p0) + zpower√p1(1 − p1)) / |p1 − p0|]2
Total requirement: max(nsensitivity / P, nspecificity / (1 − P))
How to Use This Calculator
- Set a minimum acceptable sensitivity and specificity.
- Enter the performance you expect from the tested method.
- Add the planned case share or condition prevalence.
- Enter total sample size, alpha, and desired target power.
- Select Calculate Power to view endpoint and overall estimates.
- Use the required sample result when revising recruitment targets.
Why Diagnostic Power Matters
Power measures the chance that a study detects a meaningful diagnostic difference. It protects a study from weak conclusions. Sensitivity describes correct identification of people with disease. Specificity describes correct identification of people without disease. Both values need enough supporting observations. Small samples produce unstable percentages. They can miss important improvements. This calculator estimates power for each measure separately. It also reports a conservative overall value. That value uses the lower estimate. It reveals the weakest part of a design.
Balance Cases and Controls
A diagnostic study contains two groups. Cases have the target condition. Controls do not have the condition. Prevalence determines their balance in a sampled population. Low prevalence creates fewer cases. Sensitivity then becomes harder to assess precisely. High prevalence creates fewer controls. Specificity then becomes harder to assess precisely. Total sample size alone does not tell the full story. The number within each group matters. This calculator estimates cases and controls before comparing power. It exposes designs lacking useful balance.
Set Performance Assumptions
The null value represents the acceptable performance. Expected value represents performance under the design. Their difference is the effect size. A larger difference is easier to detect. Smaller differences require more participants. The confidence level affects the decision threshold. A stricter threshold requires stronger evidence. It reduces power when conditions remain unchanged. Alpha represents the probability of a false positive result. Many use 0.05. The calculator uses a two-sided normal approximation. This supports planning with reasonably large counts.
Review Each Endpoint
Sensitivity power depends on the number of cases. Specificity power depends on the number of controls. The calculator evaluates both using separate proportions. It then reports conservative power. This prevents one strong measure from hiding another weakness. A screening test may show excellent sensitivity power. Yet low control counts may leave specificity uncertain. Study decisions should review both values. This matters when prevalence differs from recruitment mix. Deliberate case-control sampling changes the balance. Enter the planned case share honestly.
Plan the Sample Size
Required sample size supports early planning. It estimates required cases and controls. It converts needs into total sample size. The larger requirement controls recommendation. This avoids meeting one endpoint while failing another. Add allowance for exclusions and missing verification. Consider clustered recruitment when participants come from shared sites. Clustered designs need more observations. Also consider imperfect reference standards. Misclassification can reduce observed sensitivity and specificity. A sample size calculation guides planning. It does not replace protocol review or statistical oversight.
Apply Results Carefully
Use results with practical judgment. Confirm expected values using credible evidence. Select null values matching real requirements. Use prevalence matching the sample. Check whether a paired comparison or independent design is planned. Designs can require different methods. Report assumptions with every power estimate. Explain the chosen alpha and target power. Recalculate carefully when recruitment conditions change. Keep the final analysis method aligned with the planning model. Good power supports reliable decisions. Good design makes decisions useful, transparent, repeatable, and defensible for real applications.
Frequently Asked Questions
What does power mean in this calculator?
Power is the probability of detecting a true difference between expected performance and the selected benchmark. Higher power reduces the chance of missing a meaningful diagnostic improvement.
Why are sensitivity and specificity calculated separately?
Sensitivity uses participants with the condition. Specificity uses participants without it. Different group sizes create different precision and power, even within one total sample.
What is overall conservative power?
It is the lower of sensitivity power and specificity power. This prevents a stronger endpoint from concealing an underpowered endpoint.
How does prevalence affect the result?
Prevalence determines expected case and control counts. Lower prevalence usually reduces sensitivity power. Higher prevalence can reduce specificity power because controls become fewer.
Why do I enter benchmark and expected values?
The benchmark states minimum acceptable performance. The expected value states anticipated performance. Their difference determines the effect size used in the power calculation.
Which alpha value should I use?
Many studies use 0.05. Use a smaller alpha when false positive claims require stronger protection. Follow the standards of your protocol or field.
Is this an exact binomial power calculation?
No. It uses a normal approximation for planning. It is most useful when expected case and control counts are reasonably large. Use specialized methods for small samples.
What does the required total sample represent?
It is the estimated total sample needed to reach the selected target power for both endpoints. The larger sensitivity or specificity requirement determines the result.
Can I use this for a case-control design?
Yes. Enter the planned case share rather than population prevalence. The result then reflects the allocation you intend to recruit.
What happens if expected performance is lower than the benchmark?
The calculator can quantify power to detect a decline. Review the direction carefully. A lower expected value may signal an unacceptable test outcome.
When should I recalculate?
Recalculate whenever assumptions, prevalence, recruitment, or endpoints materially change.