INSIGHTS / FLAGSHIP
When not to use AI: the checklist we score against
Every NO-GO we write has to fail at least one of six tests. Here they are, with a worked example each - so you can check our verdicts against them.
Most of the AI disappointment we see was predictable in advance. Not with hindsight, and not with unusual talent — with a checklist. This is the one we score every candidate use case against in an AI Reality Check. A use case that fails any test below gets a NO-GO in the report, in writing, with the failing test named.
We publish the list for two reasons. First, so you can check our verdicts against it — a scoring method you can't inspect is a vibe. Second, because the most useful thing a consultant can tell you is usually what not to spend on.
The context matters here. Australian adoption is already broad — the Reserve Bank calls it piecemeal1, and only 12% of organisations report transformation-grade value from it2. The missing skill isn't access to AI. It's discrimination: knowing which workflows it pays in, and which it doesn't.
A use case doesn't get built because someone senior likes it. It gets built because the numbers survive scrutiny.
Test 1 — The process isn't stable enough to automate
Automating chaos just makes faster chaos. If the process changes shape depending on who runs it, the model learns noise, and every exception becomes a support ticket.
Worked example. A services firm wants AI to draft quotes from inbound enquiries. Every estimator quotes differently — different inclusions, different margins, different templates. The verdict is NO-GO on automation and a two-week standardisation exercise instead. The standardised process later scores GO.
Test 2 — There's no measurement baseline
If nobody can say what the process costs today — hours, error rate, cycle time — nobody will ever know whether the automation worked. That isn't a technicality; it's the difference between an asset and a subscription.
Worked example. A firm wants AI email triage. Nobody knows the current response time, volume mix or misrouting rate. The roadmap's first item isn't a model — it is two weeks of measurement. The unlock condition is written into the report.
Test 3 — The volume doesn't justify the build
A checklist or a template beats a model at low volume, and it beats it forever: no drift, no retraining, no failure modes. Numerals matter here — we count.
Worked example. A board report assembled once a month, taking three hours. The proposed AI build would have cost more to maintain each quarter than it saved each year. NO-GO; a report template and one saved query does the job.
Test 4 — The data can't support it yet
Machine learning needs labels, volume and a measurement loop. Most organisations don't have them yet, and pretending otherwise is how pilots die. The honest move is fixing the data first — which is cheaper anyway.
Worked example. A multi-site operator wants demand forecasting. Fourteen months of revenue data, recategorised twice, across three systems. NOT YET, with the unlock condition named: one consistent category scheme, six months of it.
What "fixing the data first" usually means
It's smaller than it sounds. In most mid-market engagements it comes down to three moves:
- One scheme, kept. A single category or coding standard, with an owner, applied forward — not a retrofit of history.
- A baseline that survives a quarter. Hours, error rate or cycle time, measured the same way twice.
- A join that works. The two systems that matter agreeing on one identifier.
Each is boring. Each is also the unlock condition on a lane of the roadmap — which is why they get done.
Test 5 — The failure cost lands on a customer or a regulator
Accuracy ceilings are real. When the cost of a wrong answer lands on a client file or a compliance obligation, the ceiling has to clear the bar with margin — not on average, but on the worst day.
Worked example. A licensee wanted AI-drafted responses to client advice queries. The accuracy ceiling was high — but not high enough for a document a regulator might read. NO-GO for client-facing drafting; GO for the internal research summary behind it, with a person signing every response.
Test 6 — A person with better incentives beats the model
Some bottlenecks are organisational, not technical. If the delay exists because someone is paid to hold the sign-off, no model removes it — it just gives the delay a dashboard.
Worked example. An approvals queue averaging nine days. The analysis finds seven of them are one manager's holding pattern, and the incentive was real: errors were career-fatal, speed was invisible. The fix is a delegation rule and an error-tolerance agreement. No AI is harmed.
Most of the AI disappointment we see was predictable in advance.
What passing looks like
A use case that clears all six tests isn't automatically a GO — it still has to out-score its rivals on value, complexity and readiness. But a use case that fails one is a NO-GO no matter how good the demo looked. When our approach isn't likely to beat your current process, we put the probability of failure in the report.
NEXT STEP
Run this checklist against your own workflows
That's what the AI Reality Check is: 2-4 weeks, one fixed fee, and a verdict on every candidate — scored against the six tests above, in writing.
FOOTNOTES
1. Reserve Bank of Australia, Bulletin, November 2025 — adoption described as "piecemeal". Accessed July 2026. ↩
2. Deloitte, Enterprise AI 2026 — 12% of Australian organisations report transformation-grade value, vs 25% globally. Accessed July 2026. ↩
NEXT ARTICLE
Australia has adopted AI. It hasn't been paid yet PUBLISHING SOON