AI Evaluation Sample Size

Estimate a binary-evaluation sample size for measuring AI task success rate.

How this calculator works

Enter baseline success rate (%), desired margin of error (%), z score. Select Calculate to apply the displayed formula and review the labeled results.

Formula / method

n = z² × p(1 − p) ÷ margin²

Worked example

An 80% baseline, 3% margin and 95% z score needs about 683 examples.

Assumptions and limitations

Examples are independent and the normal proportion approximation is appropriate.

This is not a power analysis for A/B comparisons and does not correct for stratification or multiple tests.