The draft appeared in a minute. Then a manager spent ten minutes checking the terms, rewriting the answer and moving it into the working system. To understand whether AI helps, a business owner needs to count all of that work, including any effort shifted from the person doing the task to the person checking it.
This worksheet is for a routine you have already chosen, such as preparing replies to standard service enquiries. It supports a decision about whether to continue a trial. Every number below is invented to explain the calculation; Nadali has not conducted this study.
1. Define a fair comparison
Choose a unit of work: one finished, checked reply that is ready to use. Apply the same requirements for accuracy, completeness and tone to both methods.
Measure the existing process before changing it. Compare similar requests, with similar complexity and staff experience. Do not give AI only the easy cases. If you still need to choose a routine, start with our guide to one working task.
Write down the trial period, case count and continuation criteria in advance. Twenty cases, used below, makes the arithmetic easy; it is not a universal sample-size recommendation. Rare but costly failures might never appear in such a small group.
The UK government’s guidance on evaluating AI interventions recommends planning baseline measurement early. Our practical application is to agree how you will count before seeing the results.
2. Give every case a row
Use these fields in an ordinary worksheet:
- Case number, date, method and complexity: routine or exception
- Minutes preparing inputs, drafting and transferring the result
- Minutes for all checks, including repeat checks
- Minutes correcting, retrying prompts or rewriting manually
- Accepted at the first check: yes or no
- Final result meets requirements: yes or no; full manual fallback required: yes or no
- Error, consequence and reason for rejection, where relevant
Use identical fields for the manual process. Add everyone’s time: six minutes from an employee plus four from a manager equals ten staff minutes. Include colleagues’ clarification work. Do not count the same interval as both checking and correcting.
Exclude system waiting time when someone can do other work; include waiting that keeps them occupied. Track elapsed time from request to finished reply separately if customer response speed matters.
Keep setup, training and instruction-writing outside the case rows as separate costs. Add ongoing maintenance to the relevant period. Include failed attempts and later rework. Leave unfinished cases visible with the time already spent; comparisons remain provisional until their outcomes are known.
3. Count the whole process: an illustrative example
Imagine two comparable groups of 20 replies. All eventually pass the same final check. These are total staff minutes for each group.
| Stage | Existing process | AI-assisted process |
|---|---|---|
| Preparation, drafting and transfer | 240 | 70 |
| All checks | 48 | 70 |
| Corrections and manual rewrites | 12 | 40 |
| Total ongoing work | 300 | 180 |
| Additional initial setup and training | 0 | 150 |
| Total for the first group | 300 | 330 |
The ongoing difference is 300 − 180 = 120 minutes, or six minutes per reply. That is 40% of the 300-minute baseline. Including setup, however, the first group took 30 minutes longer: 330 − 300 = 30.
If the six-minute difference really continues, with no additional costs, the 150 minutes of setup would be recovered after 25 cases: 150 ÷ 6 = 25. A positive time balance starts with case 26. This is a conditional calculation to test, not a forecast.
4. Keep quality visible
Add quality to the same fictional example. Under the existing process, 18 of 20 replies pass the first check: 90%. With AI assistance, 14 of 20 pass: 70%. Four require corrections; two are rewritten entirely by hand. That work is already included in the 40 minutes of rework.
Eventually, 20 of 20 replies in both groups meet the requirements: 100%. The AI group’s full manual fallback rate is 2 ÷ 20 = 10%. These measures answer different questions: how often the initial result is acceptable, how often the complete process succeeds, and how often the new method has to be abandoned.
Keep every allocated case in its group’s denominator. If one reply never reaches the required standard, final acceptance is 19 ÷ 20 = 95%. Removing it would hide a failure.
Check outputs against sources and agreed conditions, as explained in our human-review guide. Examine critical errors separately, even when the overall acceptance rate looks good.
5. Separate capacity from cash
If a later comparable group of 20 replies also frees two hours, after setup time has been recovered, that time might help clear overdue enquiries or handle harder orders. With unchanged payroll and other spending, that is additional team capacity. Cash savings require an actual reduction in payments, such as overtime or outsourced processing. The Government Efficiency Framework makes this distinction between benefits that release cash and those that do not.
For your budget, list subscriptions, usage charges, support and launch costs separately. Never subtract currency from minutes. If valuing time, use the relevant role’s hourly cost: managerial review may cost more than initial preparation. Label the result as the estimated value of staff time; actual cash savings need evidence of reduced spending.
6. Agree when to stop
Before testing, set your own conditions:
- Pause immediately if data boundaries are breached or a critical error passes the final check
- Return to the existing method if quality falls below the agreed level
- Rework the process if checking and corrections consume the expected benefit
- Continue only when the pre-agreed minimum time saving and acceptable costs are achieved
These are Nadali’s practical rules to adapt to your task. The NIST AI Risk Management Framework recommends assessment in conditions resembling actual use and ongoing monitoring.
A small pilot informs your next decision. It does not establish that AI caused the difference: request mix, workload or staff experience may have changed. Avoid turning 20 cases into a promise of annual savings. Repeat measurement with another comparable group and test again after material changes to the tool or process.



