AI STRATEGY
Your AI pilot worked. Should you scale it, change it or stop it?
Decide whether an AI pilot is ready for everyday business using evidence about value, quality, adoption, risk, ownership and continuing cost.

An AI pilot can produce a polished answer in a meeting and still be unfit for everyday business.
The pilot may have used clean sample data. The project team may have corrected problems quietly. Employees may have tried it because senior management was watching. None of that is dishonest, but it means the demonstration has answered only one question: can this idea work under the pilot conditions?
Before scaling, leadership needs a harder answer:
Can this process keep producing a useful result, at an acceptable cost and risk, when ordinary employees use it on ordinary work?
That decision should lead to one of four outcomes: scale it, change it, test it again or stop it. “The demo was impressive” is not a fifth outcome.
The executive answer
- Return to the business problem and the baseline agreed before the pilot.
- Separate technical performance from business value.
- Measure checking, corrections, exceptions and adoption—not just speed.
- Test representative work, including awkward and incomplete cases.
- Confirm who will own the live process, information, controls and support.
- Recalculate the business case using production costs rather than pilot costs.
- Decide what evidence would justify expansion and what result would stop it.
- Scale in bounded stages. Do not move from 50 reviewed cases to unrestricted automation in one jump.
This is not excessive caution. PwC's 2026 survey of 4,454 chief executives found that 56% reported no significant revenue or cost benefit from AI, while 12% reported both. The same research found that companies with strong AI foundations were substantially more likely to report useful financial returns. Read PwC's 2026 Global CEO Survey.
The management problem is no longer whether AI can produce something. It is whether the organisation can turn that capability into dependable work.
Proof of concept, pilot and production are different decisions
These terms are often used interchangeably, which makes expectations difficult to manage.
| Stage | Main question | Typical evidence | What it does not prove |
|---|---|---|---|
| Proof of concept | Is the idea technically possible? | A small demonstration using a limited dataset | That employees will use it or the economics will work |
| Pilot | Does it improve a defined process under controlled real conditions? | Baseline comparison, user feedback, errors, exceptions and operating cost | That it is ready for every team, customer or case |
| Production | Can the organisation operate it safely and reliably over time? | Named ownership, monitoring, support, controls, incident handling and continuing measurement | That the process will remain valuable without review |
Australian government guidance makes a similar distinction: an AI proof of concept should have a clear problem, success criteria, suitable data and testing, with business, operational, legal and technical ownership involved. Although written for the public sector, those disciplines transfer well to commercial projects. Review the Australian Government's AI proof-of-concept guidance.
Return to the decision the pilot was supposed to support
A weak pilot objective sounds like:
Test an AI customer-service assistant.
A decision-ready objective sounds like:
Determine whether an assistant can prepare accurate first replies for three routine enquiry types, reducing median preparation time without increasing corrections or customer complaints.
The second version identifies the users, work, intended improvement and quality boundary. Management can disagree about whether the result is good enough, but at least there is something concrete to discuss.
If the pilot began without a clear objective, do not manufacture one after seeing the results. Write down what the business actually learned, what remains unknown and which decision can honestly be made now.
Use eight gates before scaling
Treat each gate as a management conversation supported by evidence. A high score in one area should not hide a dangerous failure in another.
1. Business-value gate
Did the process improve an outcome the business cares about?
Useful measures include response time, staff time, throughput, missed follow-ups, correction work, customer waiting, revenue leakage or the time required to prepare a decision.
Tool usage is not a business outcome. One hundred employees opening an assistant shows exposure, not value. A thousand generated summaries do not matter if nobody uses them to decide or act.
Compare the pilot result with a baseline captured before or during the test. Where possible, compare similar cases with and without the new workflow. Seasonal demand, a new employee or a quieter week can otherwise be mistaken for an AI improvement.
2. Quality gate
Did speed improve without creating unacceptable mistakes?
Pair every efficiency measure with a quality measure:
| Efficiency measure | Quality partner |
|---|---|
| Reply preparation time | Factual corrections and complaints |
| Documents processed per hour | Missing fields and incorrect classifications |
| Reports produced | Reconciliations, omissions and management corrections |
| Leads contacted | Relevant replies, opt-outs and qualified opportunities |
| Orders prepared | Quantity, price, delivery and customer-confirmation errors |
Do not report only an average accuracy percentage. Ask what the errors were, who was affected and how costly they would be in normal operation. Confusing two product colours is different from confirming a payment that never arrived.
3. Exception gate
What happened when the work was unclear, incomplete or outside the normal pattern?
Pilots often concentrate on the happy path: a clear invoice, a familiar customer question or a spreadsheet with consistent headings. Production work includes blurry documents, old policies, duplicate records, conflicting totals, unusual customers and requests nobody anticipated.
Create an exception set deliberately. Include cases with:
- missing information;
- conflicting source records;
- unusual wording or document layouts;
- requests outside policy;
- sensitive personal or commercial information;
- very high values or consequences; and
- cases where the correct response is to stop and ask a person.
A system that recognises uncertainty and routes an exception may be more useful than one that always produces a fluent answer.
4. Adoption gate
Did the intended users choose to use the process when the project team was not standing beside them?
Low adoption can have several causes: the workflow adds steps, the answer is hard to trust, the approved tool is slower than an unofficial one, employees were not trained, or the project solved a problem management noticed rather than one employees actually had.
Talk to users separately from the project sponsor. Ask:
- Which step became easier?
- Which step became slower?
- When did you avoid the tool?
- What did you still do outside the system?
- Which result did you feel compelled to check twice?
- What would make you stop using it next month?
Resistance is not automatically a training problem. Sometimes it is evidence that the process design is poor.
5. Ownership gate
Can you name the people responsible after the pilot team leaves?
Production needs owners for:
- the business outcome;
- the source information and its freshness;
- user access and permissions;
- output quality and exceptions;
- supplier or technical support;
- incidents and complaints; and
- the decision to change or retire the system.
“IT owns it” is usually incomplete. IT may operate the connection, but the customer-service lead still owns the approved answer, finance owns the payment rule and management owns the risk accepted by the process.
6. Economics gate
Does the business case still work using the full production cost?
Add the costs the pilot may have hidden:
- integration and data preparation;
- employee and manager training;
- human review and correction;
- monitoring, support and incident response;
- usage-based model or platform charges;
- privacy, security and legal work;
- maintaining approved information;
- supplier management; and
- replacing or leaving the system later.
Use conservative, expected and optimistic scenarios. If the project only makes sense when every saved minute becomes cash and support costs stay unusually low, the business case is fragile.
For a more complete method, use Garatropic's guide to measuring AI ROI without fooling yourself.
7. Control gate
Can the system's authority be explained and enforced?
Document what it may read, prepare, recommend, send or update. Important limits should be enforced by the application or business system, not left as polite wording in a prompt.
Before scaling, confirm:
- which records and fields are accessible;
- which actions require approval;
- who is authorised to approve them;
- how requests, evidence and actions are logged;
- what happens when confidence is low;
- how access is revoked; and
- how the process is stopped during an incident.
Grant Thornton's 2026 survey found a large governance-confidence gap between organisations still piloting AI and those reporting fully integrated use. That does not prove governance alone causes success, but it reinforces the need to build evidence and oversight alongside deployment. Review the Grant Thornton AI Impact Survey.
8. Operating-readiness gate
Will the process continue to work on an ordinary Tuesday?
Check the unglamorous but necessary parts:
- What happens when a source system is unavailable?
- How will users know that information is out of date?
- Who receives a failed-job alert?
- How quickly must a serious incident be handled?
- Can completed actions be reversed?
- How are changes tested before release?
- What records must be retained?
- How will performance be reviewed after three months?
An AI system is part of an operating process. It needs the same attention to continuity, support and change that other important systems receive.
Worked example: a customer-reply pilot
Consider an illustrative service business. Six employees answer routine questions about appointment availability, required documents and service areas.
The company pilots an assistant on 120 suitable enquiries. It uses approved service information to prepare a draft, and an employee checks every reply before sending it.
The results are:
| Measure | Before | Pilot result |
|---|---|---|
| Median preparation time | 7 minutes | 4 minutes |
| Replies requiring a factual correction | Not previously recorded | 11% |
| Suitable enquiries answered within one hour | 46% | 71% |
| Enquiries escalated because information was missing | Not previously recorded | 14% |
| Customer complaints attributed to the pilot | Not previously recorded | 0 recorded |
This is encouraging, but it does not justify unrestricted automation.
Management should ask what the 11% corrections involved. If most were harmless formatting changes, the quality picture differs from a situation where the assistant repeatedly gave the wrong service area. The 14% escalation rate may reveal a useful control—or a weak knowledge source. The company also needs to know whether review time is included in the four minutes and whether the current information can be maintained.
A reasonable next step might be to expand the reviewed pilot to a larger sample and a second team while fixing the information gaps. It would be premature to let the assistant send every answer automatically.
Choose the decision honestly
The next stage should remain a series of explicit learning and decision gates. The following 90-day path is illustrative; higher-risk work may need a longer period.
Earn the right to expand at each stage.
Problem · baseline · people · risks
Rules · testing · training · measures
Small group · review · feedback · incidents
Expand · revise · restrict · stop
Scale
Scale when the pilot produced a meaningful business improvement, quality stayed within an agreed boundary, users adopted the workflow, continuing costs remain defensible and live ownership is clear.
Expansion can still be staged: another enquiry type, one more branch, a larger volume or carefully increased permissions.
Change
Change the design when the underlying problem remains valuable but the workflow, information, controls or user experience caused avoidable failure.
Examples include improving the source material, narrowing the task, adding an approval, removing a redundant handoff or using ordinary automation for the predictable parts.
Test again
Repeat a bounded pilot when the evidence is too small, the sample was unrepresentative or an important operating condition changed. Define the unanswered question first. “Try it for another month” is not a test plan.
Stop
Stop when the benefit is too small, correction work absorbs the saving, the data cannot support the task, the risk is disproportionate or nobody will own the live process.
Stopping a weak pilot is not a failure. Continuing it to protect the original decision is.
A one-page scale decision
Ask the sponsor to complete this sentence:
We recommend [scale/change/test again/stop] because the pilot changed [business measure] from [baseline] to [result], while [quality and risk measures] remained [within/outside] the agreed boundary. The live process will be owned by [name or role], at an estimated continuing cost of [amount or range], with [approval, monitoring and exception controls].
If the sentence cannot be completed without words such as “probably,” “roughly” and “should,” the missing evidence is visible.
Scale the learning before the technology
The most valuable output of a pilot is not the prototype. It is the evidence that helps the business make a better decision.
Keep what worked. Fix what did not. Expand only the permissions, users and cases the evidence supports. A careful scale decision may feel slower than a company-wide announcement, but it is much faster than discovering at full volume that the pilot never tested the difficult parts of the work.
Research and helpful links
- PwC's 29th Global CEO Survey
- PwC's 2026 AI Performance Study
- Australian Government guidance for moving an AI proof of concept toward scale
- NIST AI Risk Management Framework resources
- Grant Thornton's 2026 AI Impact Survey
Research checked on 20 September 2026. The appropriate evidence, controls and professional advice depend on the process, industry and jurisdictions involved.
