All resources
Playbook8 min read

Designing a 14-day AI pilot that actually proves something

Most AI pilots fail not because the technology underperformed, but because nobody agreed in advance what would count as success. A pilot that ends in a debate about interpretation cost the same as one that ended in a decision. This is how to design for the second outcome.

Published

Pick one journey, not one department

"Automate support" is not a pilot; it is a budget line. "Resolve order-status questions on WhatsApp in Arabic and English, end to end, including the lookup" is a pilot. It has a start, an end, a system of record, and an unambiguous definition of done.

The narrower the journey, the faster the disagreement surfaces — and surfacing it in week one is the entire point of piloting.

Pick the channel your volume is actually on

Run the pilot where the traffic already is. In most Saudi consumer businesses that is WhatsApp; in B2B it is the website; in real estate it is the phone. Piloting on a low-volume channel produces a clean result and no evidence.

You need enough conversations in fourteen days for the numbers to mean something. If the chosen channel cannot supply them, choose a different channel before you choose a different vendor.

Pick one primary metric and write down the threshold

One number, agreed before launch, with the bar written down: "40% of order-status conversations resolved without a human" or "30% more qualified meetings booked from the same inbound volume".

Secondary metrics are fine to collect and dangerous to decide on. The moment two metrics disagree, the pilot review becomes a negotiation — and the number chosen after the fact is always the flattering one.

The 14 days

Day 0–3 is co-design: the persona, the knowledge sources, the system connections, and the written success criteria. Nothing is built until the criteria are agreed, because that document is what the review is scored against.

Day 4–14 is deployment in your chosen posture — shared, VPC, self-hosted, or in-Kingdom — live on a real channel with real customers and real data. A pilot on synthetic traffic proves the demo worked.

Day 15 onward is measurement. Resist changing the prompt, the scope, or the metric during the measurement window; every mid-flight change costs you the comparison you were running the pilot to get.

Instrument before you launch, not after

Decide up front how each of these is captured, because retrofitting them at day 12 produces numbers nobody trusts:

  • Containment: conversations closed with no human touch, distinguished from conversations abandoned.
  • Escalation quality: what the human received, and whether they had to re-ask anything.
  • Write-back success: actions attempted versus actions committed to the system of record.
  • Latency: time to first response, and time to resolution end to end.
  • Language split: volume and outcome by Arabic and English, reported separately.
  • Customer signal: one post-conversation question, in the customer's own language.

Name the decision rule before you start

Write the three outcomes down and have the sponsor sign them: at or above the threshold we scale to journey two; within a defined margin we refine and rerun for two weeks; below it we stop.

"We stop" has to be a real option, stated out loud in week zero. A pilot that cannot fail is not a pilot — it is a procurement process with extra steps, and everyone in the room knows it.

Run the review against the document

At day 30, open the success criteria written on day 0 and read them first, before any dashboard. Compare, decide, and record the decision and its reason.

Then pick journey two. The second pilot is dramatically cheaper than the first — the connections, the persona, and the deployment posture are already done — and that compounding is the actual return on the fourteen days.

Book a pilot

Bring one of these to a real customer journey. We deploy a Hyper Human on it in 14 days.

Book a pilot

Keep reading