27 April 2026 · 7 min · By Jordan Foord

Agent washing: how to tell a real agentic system from a chatbot in a trench coat

When only a hundred-odd vendors out of thousands are the real thing, the buyer's job is mostly fraud detection. Here's the six-question test.

In mid-2025, Gartner put a number on something buyers had been sensing for a while: of the thousands of vendors describing their products as “agentic AI”, only around 130 were the real thing. The rest were doing what Gartner called agent washing: rebadging chatbots, RPA scripts and workflow tools with the season’s word.

The same Gartner research predicted that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Those two findings are connected. A lot of those doomed projects were never agentic to begin with. They were a chatbot in a trench coat, sold to a buyer who had no way to check.

This piece is the way to check. We run our own company on agents, so we’ll also answer our own test at the end, partly as proof the bar is passable, mostly because a test the examiner can’t pass isn’t a test.

What “agentic” actually means

Strip the marketing off and an agentic system has four properties:

Goal-directed. You give it an outcome (“close the books for May”, “triage today’s support queue”), not a script. It works out the steps.

Multi-step. It plans and executes a sequence (gather, decide, act, check) rather than answering one prompt with one response. If every action requires a fresh human prompt, it’s a tool you’re operating, not an agent that’s working.

Tool-using. It reads and writes to real systems (your ledger, your CRM, your ticketing queue) through actual integrations. An agent that can only talk about your invoices, rather than fetch and process them, is a chatbot with good manners.

Supervised. This is the one vendors skip, and it’s the most important. A real agentic system knows what it’s allowed to do alone, what needs human sign-off, and what to do when it’s uncertain. Unsupervised autonomy isn’t a more advanced version of this. It’s a liability with a roadmap slide.

A wrapped chatbot, by contrast, is a language model with a user interface and maybe access to your help docs. Useful, sometimes genuinely so. But it converses; it doesn’t execute. The trench coat is the pitch deck implying otherwise.

The six questions

Ask these of any vendor or consultant selling you “agents”. None require technical knowledge. All of them are difficult to bluff.

1. “Show me it running on your business.” Not a demo environment, not a customer video: their own operations. Anyone selling agentic transformation while running their own company on spreadsheets and elbow grease is selling you a theory. Watch for the deflection: “client confidentiality” doesn’t explain why their own finance process is off-limits.

2. “What happens when it’s wrong? Show me the failure log.” Every real agentic system fails sometimes. The honest vendor shows you the log: what broke, how it was caught, what changed afterwards. A vendor who claims theirs doesn’t fail is telling you either that it doesn’t do real work, or that nobody’s checking. Both are disqualifying.

3. “What does it cost to run monthly?” Agents have operating costs: model calls, integrations, monitoring, the human review time. A vendor who can’t quote run-costs has either never run one in production or doesn’t want you doing the ROI maths. (For calibration: one Melbourne boutique, Sandlabs, publishes typical agent run-costs of A$200–600 a month. Publishing the number is the tell that it’s real.)

4. “Who reviews its output, and when?” The answer should be specific: which role, at which checkpoints, with what authority to stop it. “It’s fully autonomous” is not a mark of sophistication. It’s the loudest agent-washing tell there is, because nobody who has operated agents in production talks that way.

5. “What did it do unsupervised last Tuesday?” Beautifully concrete and almost impossible to fake. A real operator can pull up the run history and tell you: it processed these invoices, drafted these responses, flagged these two anomalies for review. A washer gives you architecture metaphors, because there is no last Tuesday.

6. “Can my team read the runbook?” When the engagement ends, your people need to know how the thing works: what it does, when it escalates, how to adjust it, how to switch it off. If the answer is “it’s proprietary”, you’re not buying a capability. You’re renting a dependency with a logo on it.

Our own answers, briefly

Fair’s fair. Here’s how we’d answer our own test, compressed.

Running on our business? Yes: nollie, our AI CRM for hospitality, operates across four markets with an agent fleet handling finance close, customer onboarding, support triage and marketing production. That fleet is the day-to-day operating model, not a lab project. It’s also why we have opinions about supervision: we’re the ones reading the logs at 7am.

Failure log? Exists, and gets used. Recent honest entry: a marketing-production agent shipped a template with structural rules inverted. It was caught in review, fixed, and the fix written back into the agent’s instructions so the same mistake can’t recur quietly. The point of the log isn’t that failures are rare; it’s that every failure becomes a rule.

Monthly run-cost? Known per workflow and tracked, because we pay the bill. Model usage, integration overheads and (the line everyone forgets) human review time all sit in the number. We won’t pretend review time is zero; budgeting it at zero is how “AI savings” evaporate in audit.

Who reviews, when? A named human per workflow, at defined checkpoints. Anything customer-facing or money-touching gets sign-off before it leaves the building; internal, low-blast-radius work runs with after-the-fact review. The supervision tightens or loosens based on the failure log, not on vibes.

Last Tuesday? Pull the run history: research digests assembled, new customer accounts prepared and verified, support items triaged and routed, finance reconciliation items matched, with the exceptions queued for a human. Unglamorous, specific, real.

Readable runbook? Yes, and it’s the deliverable we care most about. Every workflow ships with plain-language documentation: triggers, permissions, escalation rules, the off switch. Our view is that a client who can’t run the system without us hasn’t been enabled. They’ve been acquired.

Why the washing works (and why it’s ending)

Agent washing succeeds for the same reason most washing does: the buyer can’t inspect the claim. “Agentic” has no trading-standards definition, demos are theatre by design, and in 2024–25 budgets were approved on the word alone.

That window is closing. Deloitte expects half of GenAI-using companies to be running agent pilots by 2027, which means a rapidly growing population of buyers who have seen a real one, and burned buyers are the hardest audience to wash. CFOs are already demanding proof of business value as the default posture. The six questions are simply that posture, written down and made portable.

The wider lesson sits in BCG’s 10-20-70 finding: successful AI deployment is roughly 10% algorithms, 20% technology and data, and 70% people and process. Agent washing inverts that: it sells you the 10% wearing the other 90% as a costume. The questions about supervision, runbooks and review aren’t compliance box-ticking; they’re how you check the 70% exists.

Do this next week

Take the six questions and put them in your procurement template, verbatim, before the next vendor meeting. They work best asked in order, in writing, with answers you can hold them to.

Then run the test the other way: if you already have an “AI agent” in the building, ask your own team question five. What did it do unsupervised last Tuesday? If nobody can answer from a log, you don’t have an agent problem or an agent asset. You have a chatbot in a trench coat on the payroll, and now you know.