AI pilot to production: why most trials never go live
AI pilot to production stalls when nobody owns cost. Accenture says 23% of C-suites get lasting impact — a $3k–$6k Slack agent you own skips the seat tax.
Anthropic published a CIO guide this week, written with Accenture, on moving an AI pilot to production. The number that matters is not a model score. Accenture’s Pulse of Change report (July 2026), cited in that guide, says only 23% of C-suite leaders report sustained, enterprise-wide impact from AI. The rest ran a demo that never became a job.
I deploy agents for shops in the 10–250 range. That 23% figure is the same failure I see without a CIO title: a ChatGPT workspace, a two-week Slack bot trial, then nothing wired to the system of record.
What does AI pilot to production actually require?
It requires a named job — user, task, output, and a quality bar you can fail — plus a cost owner before the trial starts, not after the demo looks good. Anthropic’s blueprint, again with Accenture, puts that four-part job definition and a lightweight total-cost model before the pilot. That is the part most small teams skip because a chatbot is cheap to stand up and expensive to own.
Accenture’s September 2026 Tokenomics research, quoted in the same guide, found 42% of organizations share AI cost and outcome accountability between IT and finance, with no single owner. If two departments “share” the meter, nobody kills a bad trial and nobody promotes a good one.
A production agent in my world looks like this:
- Trigger: a ticket, a lead, or a staff question hits Slack (or email, or the phone — pick one lane).
- AI action: classify, draft, or route against a written playbook.
- System of record: HubSpot, Jira, a shared sheet, or the inbox that already runs the company.
- Human escalation: a named person who gets the exception, not a channel everyone ignores.
If you cannot fill those four lines on a sticky note, you are still in pilot theater. That is the same first-lane rule I use on AI for small business work: one workflow, one record, one human.
Anthropic also sketches a four-tier review model — automated, sampled, reviewed, advisory — so human time matches the risk of the output. Invoice drafts get sampled. Refunds get reviewed. Strategy stays advisory.
Who should own the agent after the trial?
One operator, not a committee. The 42% “shared IT and finance” pattern is why pilots die: finance will not fund what IT will not measure, and IT will not measure what the business never defined.
On a 40-person agency, that owner is usually the COO or the founder who still reads the shared inbox. On a multi-location clinic, it is the office manager who already owns the front desk, not the IT vendor. AI agent maintenance is the same argument after go-live: if nobody is named on the calendar for a monthly check, the agent rots.
Do not promote a pilot to production because the champion liked the screenshots. Promote it when the job definition still holds under normal load — real tickets, real exceptions, real hours, not the hand-picked week.
What should a 10–250 person shop do with this?
Ignore the CIO theater. Steal the ownership rule. You do not need seven sequential leadership work-outs. You need one live lane, one cost owner, and a kill switch.
That is why I still sell a Slack AI Agent as a one-time $3,000–$6,000 deployment for companies that already live in Slack — you own the setup; I am not metering seats. Copilot-style per-user plans are a forever tax on a trial that never became a job. If the leak is the public phone instead of internal ops, that is a different product; this guide is about the internal one.
If your “pilot” is still a shared ChatGPT login with no system of record, you are in the 77% Accenture is describing. Name the job, name the owner, then go live — or stop spending.
For a shop that already runs on Slack, the change is simple: stop running unpaid internships for the model. Put one agent in production with a human on the exceptions, or admit the trial is done.