TL;DR: A 90-day automation pilot should end with a funding decision. Set success targets, spending limits, and stop rules before launch. Scale only when the evidence shows the workflow is useful, affordable, and safe to run.
What does a pilot project mean in business?
An automation pilot is a limited test of one real workflow before a wider rollout. It checks whether the work gets done better, what it costs, and whether people can handle failures. A working demo alone cannot answer those questions.
Your 90 day automation pilot business case needs a decision date and a clear reason to spend more. Start with the business process automation ROI framework for the value model. Then use this plan to decide what evidence must exist before you approve expansion.
A 2023 GSA OIG audit found that a claim of more than 240,000 hours saved annually lacked reliable evidence. The Inspector General's report said GSA had not checked actual hours saved with bot users or tracked bot costs. This public-sector finding shows why a savings claim needs an audit trail; it is not an SMB performance benchmark.
Choose one narrow job with enough cases to observe:
- E-commerce: flag orders with missing shipping details; keep refunds outside the pilot.
- Home services: create office tasks from completed job forms; keep invoice approval manual.
- B2B sales: route qualified forms into a customer relationship management system, or CRM; exclude existing-account disputes.
- Professional services: check signed onboarding packets for missing fields; leave contract interpretation to staff.
- Customer support: draft answers for one routine topic; require staff approval before sending.
How to write a business case for a 90-day pilot
The first step is to define one business problem and measure the current process. Name the person who owns the result, the cases you will include, and the evidence that could change the funding decision. Buy tools only after that agreement exists.
Create a shared document titled 90-Day Automation Pilot Plan: Success Criteria, Guardrails, and the Go/No-Go Review. The title can be long; the agreement should fit on one page. Fill these fields before development starts.
| Pilot charter field | What to write |
|---|---|
| Problem and scope | One workflow, team, case type, and finish event |
| Baseline | Current labor time, error rate, volume, and waiting time |
| Benefit hypothesis | The measurable change worth funding |
| Success and stop rules | Targets, denominators, review window, and unacceptable events |
| Spending authority | Total pilot cap and who can approve a change |
| Evidence owner | Who checks case logs, labor records, and bills |
| Decision owner | Sponsor who can approve, stop, or extend on day 90 |
| Fallback | Who resumes manual work and how pending items are found |
Use the two-week time study worksheet if the baseline is weak. Keep waiting time separate from hands-on labor. A shorter queue can help customers without cutting payroll.
How to pilot a project over 90 days
Use the first month to define and test the workflow, the second to run a limited live trial, and the third to collect stable results and decide. Each phase earns permission for the next. A calendar date never overrides a failed test.
- Days 1–14: record the baseline. Export case IDs from the CRM or order system into a controlled Excel workbook. Record manual work, fixes, unfinished cases, and exclusions. Freeze the rules for counting cases.
- Days 15–30: build and test in shadow mode. Connect a form or test inbox to Power Automate and a test list. Shadow mode means the system proposes an action while staff keep doing the real work. Test missing fields, duplicates, and timeouts.
- Days 31–45: open a small live group. Route only the approved case type through the workflow. Require approval before customer messages or financial changes. Check that failed runs reach a named person's queue.
- Days 46–60: compare matched work. Where practical, randomly assign eligible cases to pilot and manual groups. Use the same finish rule and follow-up window. Record changes in staffing, demand, or case mix that could explain a difference.
- Days 61–75: hold the version steady. Measure normal review and repair work after launch support tapers off. Test the fallback during staffed hours. If a material fix changes behavior, start a fresh evaluation window and record the delay.
- Days 76–90: close the evidence and review. Stop adding evaluation cases early enough to observe their outcomes. Reconcile pending items and costs. Give the sponsor the decision sheet below, including failures and missing evidence.
The 2023 NBER working paper studied 5,179 support agents and reported a 14% average productivity gain. This is evidence from that setting, not a pilot target. The November 2023 working-paper version also found gains varied across workers. Compare similar case types and staff groups before attributing a change to automation.
Set success criteria before the first live run
An automation pilot needs a primary benefit target plus quality, coverage, cost, and operating limits. Define each calculation before results arrive. Passing the average benefit target cannot cancel a serious failure.
The following scorecard uses illustrative targets for a low-risk office workflow. They are starting points for a sponsor discussion, not industry standards. Replace them with limits that fit the workflow and its downside.
| Criterion | Measurement rule | Example pass condition |
|---|---|---|
| Primary benefit | Total staff minutes, including review and fixes, per eligible case | At least 25% lower than the matched manual group |
| Quality | Cases with a material error found within seven days / all evaluated cases | At most 2%, with no critical error |
| Coverage | Eligible cases assigned to the pilot that actually enter it / all assigned eligible cases | At least 90%; skipped work stays visible |
| Completion | Every assigned case has a verified result or a named pending owner | No unowned or unaccounted-for case |
| Cash benefit | Verified avoided spending minus recurring cash costs | Positive, with a forecast inside the approved payback limit |
| Operation | Staff can find a failure, stop new runs, and resume manual work | Fallback drill passes; operating owner accepts the workload |
Count a case once even if the automation retries it several times. Include manual fallback and rework in its labor total. Do not quietly remove failed, delayed, or difficult cases from the pilot group.
For the final comparison, wait for the agreed outcome window to close. If cases remain open, report them and their work to date; do not claim a final time-saving rate from completed cases alone. If you cannot randomize assignment, describe the comparison as observational and list likely sources of bias.
Which guardrails stop the pilot?
Pause new automated actions when a critical error, missing record trail, failed fallback, or breached spending limit appears. Keep the affected work under a named person's control while the team investigates. Resume only after the fix and recovery have been checked.
For AI-enabled workflows, NIST's 2023 AI Risk Management Framework supports explicit go/no-go decisions and assigned responsibility for disabling systems that fail their intended use. The controls below are our proposed pilot rules, not a NIST checklist.
| Trigger | Immediate response | Evidence needed to resume |
|---|---|---|
| Wrong recipient, unauthorized payment, or sensitive data exposure | Disable new actions; alert the business and system owners | Affected cases identified, incident handled, fix retested |
| Duplicate task or write after a retry | Pause the affected path and reconcile destinations | Repeated-event test produces one intended result |
| Missing logs or mismatch between source and destination counts | Stop expansion; send new work to the manual queue | Every case reconciled with a result or pending owner |
| Review queue exceeds its staffed limit | Reduce intake or fall back to manual handling | Owner confirms capacity and clears overdue work |
| Approved spending cap is reached | Stop billable pilot activity and notify the sponsor | Explicit budget decision or closure |
Set the review queue limit in both case count and oldest-case age before launch. Assign stop authority to the operating owner, with a backup for absences. The human approval gate guide helps define who may approve a risky action and what they need to see.
Stopping a workflow does not undo messages already sent or changes already made. Keep an action log, identify the affected records, and reconcile pending work before replaying anything. Never replay the entire queue just because a connection works again.
What should the pilot program cost?
Budget for setup, staff time, software, monitoring, and recovery work. There is no useful universal pilot price: a simple task handoff and an unattended desktop bot have different costs. Approve a spending cap and record contractual commitments as well as monthly usage.
Microsoft’s June 2025 licensing guide lists Power Automate Premium at $15 per user/month and Process at $150 per bot/month, billed annually in USD. These are historical license references from the June 2025 guide, not complete project prices. Check the license rights, commitment, and current quote for your exact workflow.
| Budget item | USD amount or calculation | How to treat it |
|---|---|---|
| Power Automate Premium reference | $15/user/month; $180/year | Historical annual-billing reference; not a three-month contract |
| Power Automate Process reference | $150/bot/month; $1,800/year | Alternative licensing path; not automatically needed with Premium |
| Build and integration | 80 hours × $50 = $4,000 | Illustrative project labor budget |
| Baseline, training, and acceptance | 20 hours × $40 = $800 | Illustrative internal labor allocation |
| Data cleanup and recovery rehearsal | $1,200 | Illustrative scoped allowance |
| Ongoing operation | $300/month | Illustrative software and monitoring budget; avoid double-counting licenses |
| Total modeled pilot resources | $6,000 setup + $900 operation = $6,900 | Three-month model; extra commitments must be added |
Case study: a faster workflow still gets a no-go
This operator composite is a hypothetical worked example, not a public customer claim or measured That'sGonnaHelp engagement. A 12-person service firm handles 600 standard job packets per month. Each packet takes six staff minutes to check and turn into an office task, including routine corrections.
The team budgets the $6,000 setup and $300 monthly operating cost shown above. It proposes a Microsoft Forms-to-Power Automate workflow that creates tasks in a SharePoint list and alerts staff in Teams. Invoices and customer messages stay under manual control.
During testing, a repeated form submission creates two tasks for one job. The team pauses that path, adds a check against the stable job ID, and tests repeated submissions again. The final comparison uses the repaired version; earlier failures remain in the incident record.
In the final four-week evaluation, 300 eligible cases are randomly assigned to the pilot and 300 to the manual group. All finish and complete the seven-day quality window. The assumed results are three labor minutes per pilot case versus six per manual case, with three material errors in each group and no critical errors.
That is 15 hours saved across the 300 pilot cases, or a projected 30 hours per month at 600 cases. The projection assumes the full workload behaves like the tested groups. Only eight of those hours replace paid overtime at $50 per hour; the other 22 create capacity without cutting spending.
The cash model therefore counts $400 of avoided overtime and subtracts $300 in recurring costs: $100 net cash benefit per month at full volume. If the entire $6,000 setup were incremental cash spending, simple payback would be 60 months after full-volume operation starts. Internal salary allocations are resource costs; finance must identify the actual cash portion before using that payback figure.
The sponsor's rule required payback within 12 months, so the all-cash scenario receives a no-go for expansion despite the speed gain. A separate capacity-funded proposal could make sense, but it needs an explicit budget decision. Use the ROI calculator to compare assumptions, and the automation payback guide to account for rollout timing and uneven cash flows.
When is this pilot plan not a good fit?
A 90-day trial is a poor fit when outcomes take longer to observe, the process changes every week, or there is no safe manual fallback. Repair the scope or measurement design before committing to the calendar. High-consequence decisions also need domain review beyond this office-workflow plan.
A quarterly renewal workflow may produce too few independent cases. A long sales cycle may not reveal revenue outcomes by day 90. In both situations, use earlier operational evidence only for the limited decision it supports; do not present it as proven revenue lift.
Common mistakes
- Moving the target: changing the success threshold after seeing results makes the final decision harder to defend.
- Counting runs as outcomes: one case can generate retries and still never reach the destination.
- Hiding review time: faster software can leave staff with more checking and repair work.
- Turning capacity into cash: saved minutes are not payroll savings unless spending actually falls.
- Scaling the exception away: excluding every hard case can make a pilot look strong while leaving most of the business problem untouched.
Run the go/no-go review with a decision sheet
The final step in the automation business case process is a documented funding decision against the agreed criteria. Approve expansion, stop, or grant a bounded extension for a specific evidence gap. Give every decision an owner, spending limit, and next review date.
Apply the rules in this order: critical failures first, completeness of evidence second, then benefit and cost. A failed safety control blocks expansion even if the benefit is large. An incomplete sample warrants an extension only when the remaining evidence is worth its cost.
| Verdict | Required finding | Next action |
|---|---|---|
| Go | All required gates pass; owner accepts the workload and budget | Expand only the tested scope in a limited next stage |
| No-go | A required gate fails, economics miss the agreed bar, or a key risk cannot be controlled | Stop new automated work and complete the fallback handoff |
| Extend | No unresolved critical issue; one named evidence gap can be closed within a new cap | State the exact test, deadline, and maximum extra spend |
Copy this decision record into the charter. Attach the case log, cost ledger, error review, comparison results, and fallback drill rather than a slide full of totals.
Decision: Go / No-go / Extend
Scope and version reviewed: ___
Evaluation dates and eligible case counts: ___
Target versus result, including pending work: ___
Incidents and unresolved limits: ___
Cash benefit and capacity benefit, separately: ___
Approved next action, owner, budget, and review date: ___
FAQ
A pilot can support a limited next decision without proving every future benefit. These answers define the boundaries of the 90-day plan.
How long should a pilot program last?
Long enough to observe a normal work cycle, meaningful case volume, failures, and the outcomes you intend to measure. Ninety days is a planning window, not a success standard. Set the required observation window first, then check whether the schedule can provide it.
What is a pilot project versus a proof of concept?
A proof of concept checks whether an approach can work technically. A pilot tests it with a limited real operating group, including staff effort, exceptions, and costs. A successful proof of concept supports a pilot decision; it does not justify broad rollout by itself.
What comes after a pilot program?
A go decision leads to a limited expansion with continued monitoring, support ownership, and a new review date. A no-go leads to an orderly manual handoff and closure. An extension leads only to the agreed test, not open-ended rollout.
Can a pilot pass if it has not paid back its setup cost?
Yes, if the agreed approval criteria allow future payback and the measured run rate supports that forecast. Report actual pilot-period results separately from projected future returns. Do not describe the project as already paid back when cumulative cash remains negative.
What if 90 days produces too few cases?
Do not lower the target to fit the sample. Extend within a new spending cap, narrow the claim, or stop. Small samples and rare failures need careful analysis; NIST's confidence-interval guidance explains why simple approximations can be unreliable with few failures. Zero observed incidents does not prove zero risk.
Who should approve the result?
The sponsor owns the funding decision, while the operating owner accepts the work and fallback duties. Have someone other than the builder check the evidence. In a small team, that can be a finance or operations colleague who can inspect the records and challenge assumptions.
Answer clarity notes
- Dates: this article uses a November 21, 2025 publication context. The NBER figures refer to its November 2023 working-paper version; Microsoft prices refer to its June 2025 guide. Check current prices before buying.
- Evidence: linked public sources support the cited facts. GSA's hours figure was an unsupported claim challenged by its auditor, not a verified saving.
- Examples: the service-firm operator composite, budget, thresholds, schedule, and results are hypothetical planning inputs, not public customer claims or measured That'sGonnaHelp results.
- Economics: capacity, resource cost, and cash flow are different measures. Forecast savings and payback depend on stated assumptions and are not guarantees.
- Scope: this framework supports US SMB operating decisions. It does not replace legal, financial, tax, medical, or specialist risk advice for high-consequence workflows.
Sources
These primary sources support the public facts and measurement guidance. The pilot schedule, scorecard, and decision sheet are That'sGonnaHelp planning recommendations.
- NIST: AI Risk Management Framework 1.0, January 2023
- GSA OIG: RPA savings evidence audit, November 30, 2023
- NBER: Generative AI at Work, working paper revised November 2023
- Microsoft: Power Platform Licensing Guide, June 2025
- NIST: Confidence intervals for proportions
That'sGonnaHelp can help turn one workflow into a scoped pilot with an agreed scorecard and review date. Bring the current process, case volume, and spending limit to start the discussion.

