Tool
When is one AI agent really better than another?
A vendor's resolution rate, or a pilot on 50 tickets, comes with an uncertainty that is rarely shown. Type two rates and a sample size to see it.
Interval explorer
Two made-up products. Change the number of reviews and the shares to see when one product is really ahead.
A: 70.0% (56.2% to 80.9%). B: 64.0% (50.1% to 75.9%). The intervals overlap, so A and B would share a rank range.
| Reviews | Interval | Width |
|---|---|---|
| 10 | 39.7% to 89.2% | ±24.8 pts |
| 20 | 48.1% to 85.5% | ±18.7 pts |
| 50 | 56.2% to 80.9% | ±12.3 pts |
| 100 | 60.4% to 78.1% | ±8.8 pts |
| 300 | 64.6% to 74.9% | ±5.2 pts |
Wilson score interval, 95%. The shares here are examples you type in, not results.