
- Artificial Intelligence
Agent Washing: How to Tell a Real AI Agent From a Chatbot With a New Label

Agent Washing: How to Tell a Real AI Agent From a Chatbot With a New Label
Same product underneath. Only the sticker on top says "agent."
Gartner went looking for the vendors genuinely building agentic AI in mid-2025. Out of thousands claiming the label, it counted roughly 130 that qualified. Gartner gave the gap a name, agent washing, and in the same release forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027, largely on cost, unclear value and inadequate risk controls.
That is the buying environment right now. The demos are excellent. The category is real and growing fast. The failure rate is also real, and a meaningful share of it is not model failure at all. It is buying something that was never an agent, or buying a real agent and pointing it at data that cannot support one.
This is a buyer's test, not a vendor pitch. Five capability checks, twelve questions to ask in a demo, and the readiness work that decides whether any of it survives contact with your operations.
The label moved faster than the capability
Agent washing is not usually a lie. It is a redefinition. A vendor with a good retrieval chatbot adds the word agent to the product page. A vendor with a mature workflow engine calls each step an agent. A vendor with a scripted RPA bot adds a language model to the input layer and calls the whole thing agentic. Nothing was removed, so nothing is technically false. What is missing is the part that makes an agent worth the governance overhead.
Working definition. An AI agent is a system that is given a goal rather than a script, chooses its own sequence of actions to pursue that goal, uses tools that can read and write to real systems, carries state across the steps of a task, and operates inside explicit permission and escalation boundaries that a human set.
Every clause in that definition is load bearing. Remove goal decomposition and you have a workflow. Remove write access and you have a very good research assistant. Remove state and you have a stateless function call wearing a costume. Remove the permission boundary and you have an incident waiting for a quarter end.
The market context matters too, because the pressure to relabel is enormous. Gartner expects 40% of enterprise applications to feature task specific AI agents in 2026, up from under 5% in 2025, and agentic AI to account for roughly 30% of enterprise application software revenue by 2035. In supply chain software alone, Gartner projects spend on agentic capability rising from under 2 billion dollars in 2025 to 53 billion by 2030, with enterprise adoption moving from 5% to 60%. When a category grows on that curve, the label arrives before the engineering does.
Illustrative, not to exact scale: roughly 130 of the thousands of vendors claiming "agentic AI" passed Gartner's bar in 2025.
Gartner counted roughly 130 vendors as genuinely agentic in 2025, out of thousands claiming the label. Agent washing is rarely a lie, it is a redefinition, and the clause that goes missing is always autonomy.
Chatbot, workflow, RPA, agent
Most disagreements in an AI procurement meeting are really disagreements about which of these four things is on the table. The distinction is not academic. It changes the integration work, the failure modes, the audit requirement and the budget.
answers
follows a path
replays a script
decides, then acts
Autonomy increases left to right. Only the fourth column chooses its own next step.
| Capability | Chatbot or assistant | Workflow automation | RPA bot | AI agent |
|---|---|---|---|---|
| What it is given | A question | A trigger | A recorded path | A goal |
| Who decides the steps | Nobody, it responds | A human, at design time | A human, at record time | The system, at run time |
| Writes to systems of record | Rarely | Yes, on fixed rules | Yes, by simulating a user | Yes, through governed tools |
| State across steps | Conversation only | Instance variables | None beyond the script | Task memory plus retrieval |
| Behaviour on an unexpected case | Answers anyway | Errors out | Breaks | Replans or escalates |
| What you must govern | Output quality | Rule correctness | UI stability | Permissions, cost, rollback, audit |
Read the bottom row again. It is the honest cost of an agent. A chatbot that says something wrong is an embarrassment. An agent that is wrong has already posted the journal entry, released the purchase order, or emailed the customer. That is exactly why the capability is valuable, and exactly why the governance is not optional.
If you want the mechanics of how agents are constructed rather than how they are sold, our explainer on what AI agents are and how they work covers the architecture side, and the guide to retrieval augmented generation covers the grounding layer most agents depend on.
A chatbot answers, a workflow follows a path someone else designed, an RPA bot repeats recorded steps, and only an agent chooses its own next step. All four are sold under the same word.

Five tests a real agent passes and a relabelled product does not
Run these in the demo, on the vendor's own environment, in this order. Each one is designed so that a rebranded product fails visibly rather than verbally.
Test 1: The unscripted path
Give it a goal that cannot be met by the happy path. A vendor invoice that references a purchase order with a quantity mismatch and a missing tax field, for instance.
- Real agent: attempts a route, discovers the mismatch, changes approach, and either resolves or escalates with a stated reason.
- Relabelled product: returns a fluent paragraph about what should be done, or fails at the step the script did not anticipate.
Test 2: Write access, shown live
Ask it to change something and then show you the change in the source system, not in its own interface.
- Real agent: the record in the ERP, CRM or ticketing system now differs, under a service identity you can see in the audit trail.
- Relabelled product: drafts the change for a human to apply, and calls that human in the loop rather than admitting it has no write path.
Test 3: State across a long task
Interrupt it. Close the session. Come back and ask it to continue.
- Real agent: resumes with the task's history intact, knows which steps completed, does not repeat a write.
- Relabelled product: starts again, or duplicates an action it already performed. Duplicate writes are the single most expensive agent failure in finance and inventory workflows.
Test 4: The refusal boundary
Ask it to do something outside its permitted scope. Approve a payment above a threshold, or read a record the service identity should not see.
- Real agent: declines, names the policy or permission that blocked it, and routes to a human.
- Relabelled product: either complies, which is worse, or refuses with a generic safety message that reveals no enforced boundary exists.
Test 5: The reconstructable decision
Pick one action it took thirty minutes ago and ask why. Then ask them to roll it back.
- Real agent: produces the goal, the plan, the tool calls, the inputs, the retrieved evidence and the compensating action, with a token and cost figure attached.
- Relabelled product: offers a summary generated after the fact, which is a story about a decision rather than a record of one.
That last item deserves a note. KPMG's Q2 2026 pulse of large United States enterprises found only 26% had full visibility into what their AI actually costs to operate. Agents run loops. Loops consume tokens. A vendor who cannot show you cost per completed task in the demo will not be able to show you at scale either.
A real agent survives a case the happy path cannot handle, writes to a system of record, resumes an interrupted task, refuses an out of scope request, and produces an audit log recorded at the time rather than reconstructed afterwards.
Twelve questions that end a bad demo early
- Which systems does the agent write to, and under what identity?
- Show me the audit log for the action you just performed.
- What happens when the model is unavailable or rate limited mid task?
- How does the agent know it is finished, and who validates that?
- What is the rollback path for a completed write?
- What is the cost per completed task, at our volume, not your demo volume?
- Which of these steps is a model call and which is deterministic code?
- How are permissions scoped, and can I revoke one tool without redeploying?
- What does the agent do with a record it cannot reconcile?
- How do you evaluate quality after deployment, and what is your regression suite?
- What data of ours leaves our environment, and where is it processed?
- Name three customers in our industry where this runs in production, and what stopped working first.
Question 12 is the one that separates the two groups. Every team running real agents in production has a list of things that broke. A vendor with no such list either has no production deployments or is not going to tell you the truth about them.
Ask for three named production customers in your industry and what stopped working first. A vendor with no failure list either has no production deployments or will not be straight with you about them.
The pattern behind the deployments that stuck
Look at where agentic work is genuinely holding up in 2026 and a pattern shows up quickly. The wins cluster around narrow, high volume, verifiable tasks that sit next to a system of record, not around open ended reasoning.
- Document to record tasks. Invoice, purchase order and goods receipt matching where the agent has a ledger to check itself against and a clear definition of correct.
- Triage and routing. Support, claims and field service, where a wrong route is cheap to correct and a right route saves a queue.
- Retrieval over owned knowledge. The oldest of the useful cases and still the most reliable, because the answer can be traced to a source document.
- Exception hunting. Watching a transaction stream for the cases that break a rule, then assembling the context a human needs to decide.
Note what all four share. Each has a system of record that defines the truth, a bounded action space, and a cheap way to check the output. That is the whole trick, and it is why the same organisations show up in the winning column across very different industries.
The pattern shows up in practice too. In one internal knowledge deployment, time employees spent searching for information dropped by more than 90%, because the retrieval layer was built on a defined corpus rather than an open web. On the National Highways ATMS programme, an AI led detection layer cut incident reporting time by over eight minutes across a monitored network, because the system had a clear event definition to act on. And in a self hosted deployment, an assistant was built inside the customer's own infrastructure rather than a vendor cloud, which is increasingly the constraint that decides the architecture in regulated sectors.
None of those started as a general purpose agent. Each started as one task with a measurable definition of done.
KPMG found 53% of large United States enterprises had deployed AI agents by Q2 2026, while Stanford's AI Index 2026 still put agentic deployment in single digits in nearly every business function. Deployment is broad. Autonomy is not.

The part that decides the outcome, and that nobody puts in a demo
Assume you pass all five tests and buy a genuine agent. The next failure mode is bigger than the vendor choice, and the 2026 research is unusually consistent about it.
| Finding | Source | Year |
|---|---|---|
| 57% name data reliability as the key barrier to moving AI from pilot to production; 50% cite data quality as the top agentic AI deployment challenge | Informatica CDO Insights, 600 global data leaders | 2026 |
| 58% name data readiness and access as the single biggest challenge to deploying AI agents | KPMG AI Quarterly Pulse Q2, large US enterprises | 2026 |
| Organisations will abandon 60% of AI projects unsupported by AI ready data through 2026; 63% lack or are unsure of AI suitable data management practices | Gartner | 2025 |
| 37% report AI contributing to enterprise EBIT; 73% of high performers had fundamentally redesigned workflows, against 25% of everyone else | McKinsey State of AI, 1,719 respondents | 2026 |
| 60% of companies report minimal or no material value from AI; agentic AI is 17% of total AI value in 2025, rising to an expected 29% by 2028 | BCG, 1,250 CxOs across 68 countries | 2025 |
| Only 25% of AI initiatives delivered expected ROI; only 16% scaled enterprise wide | IBM Institute for Business Value, 2,000 CEOs | 2025 |
Read those together and the conclusion is uncomfortable but useful. The constraint is rarely the model. It is that the agent needs a defensible source of truth to act on, and in most mid sized enterprises that source of truth is split across a legacy ERP, three spreadsheets, a WhatsApp group and one person's memory.
Gartner's ERP forecast makes the same point from the other direction. It expects 62% of cloud ERP spending to go to AI enabled solutions by 2027, up from 14% in 2024, and embedded assistants to drive a 30% faster financial close by 2028. In the same release it names the blockers explicitly: data quality, integration complexity, skills gaps and inconsistent multi entity support. Those are not AI problems. They are systems of record problems that surface the moment you put an agent on top.
India is not behind on ambition here. Deloitte's March 2026 study found 40% of Indian enterprises at significant or full AI usage against 28% globally, with at scale adoption strongest in product development, strategy and operations. The same study found the top enabling investment priority was data storage and management at 61%, and that only 17% were pursuing business model reinvention while 44% stayed with incremental process redesign. The appetite is there. The plumbing is where the money is going, and correctly so.
This is why the sequencing argument matters more than the vendor argument, a point we set out at more length in the wrong question about AI and your ERP.
58% name data readiness as the top barrier to deploying agents, and Gartner expects 60% of AI projects unsupported by AI ready data to be abandoned through 2026. The constraint sits in the record the agent has to act on, not in the model.
How to evaluate an agent in 30 days without a six month pilot
Days 1 to 5: pick one task with a ledger
Choose a task that already has a definition of correct sitting in a system somewhere. Three way invoice matching, warranty claim triage, dealer credit checks. Write down the current cycle time, the current error rate, and who fixes an error today. If you cannot write those three numbers, choose a different task.
Days 6 to 12: audit the data the agent will act on
Pull 200 real records for that task. Count how many are missing a field the agent needs, how many contradict another system, and how many require a human to interpret a free text note. That count is your ceiling. No model improves it.
Days 13 to 22: run it read only, then scoped write
Let the agent propose for a week while humans execute, and log every disagreement. Then grant write access on the narrowest possible scope with a value threshold and a rollback path. Compare its decisions against the human decisions on the same records, not against a benchmark.
Days 23 to 30: price it and decide
Calculate cost per completed task including tokens, review time and exception handling. Compare against the baseline from days 1 to 5. Then decide on evidence rather than on the demo. BCG found leaders deploy AI workflows in nine to twelve months against twelve to eighteen for laggards, and the difference is usually decision speed on exactly this question.
Two further notes on governance, because they are what audit will ask about. Deloitte's 2026 enterprise study found around 75% of organisations plan agentic deployment within two years, but only 21% had mature governance for agents. And Gartner expects 90% of B2B buying to be agent intermediated by 2028. Whatever you deploy internally, agents will also be arriving on the other side of your sales and procurement processes. The identity, permission and audit work you do now is not single use.
Thirty days is enough to decide: five days to pick a task that already has a definition of correct, seven to audit 200 real records, ten to run read only then scoped write, and eight to price cost per completed task.
Frequently asked questions about agent washing and AI agents
What is agent washing?
Agent washing is marketing an existing product, usually a chatbot, a workflow engine or an RPA bot, as an AI agent without adding autonomous decision making. Gartner coined the term in June 2025 after assessing that only around 130 of the thousands of vendors claiming agentic AI capability genuinely qualified.
What is the difference between an AI agent and a chatbot?
A chatbot receives a question and returns a response. An AI agent receives a goal, decides its own sequence of steps, uses tools that read and write to real systems, keeps state across the task, and operates inside permissions a human defined. The practical test is whether the system can change a record in your system of record and produce an audit log explaining why.
How can I tell if a vendor is agent washing?
Run five checks in the demo. Give it a case the happy path cannot handle. Ask it to write to a real system and show you the record. Interrupt a long task and ask it to resume. Ask it to do something outside its permitted scope. Then pick an action it took earlier and ask for the full decision trail and a rollback. A relabelled product fails at least one of those visibly.
Why do so many agentic AI projects get cancelled?
Gartner attributes the expected cancellation of over 40% of agentic AI projects by end of 2027 to escalating cost, unclear business value and inadequate risk controls. Independent research points at a common root cause: 57% of data leaders name data reliability as the key barrier to moving AI from pilot to production, and Gartner separately expects 60% of AI projects unsupported by AI ready data to be abandoned through 2026.
Do I need to fix my ERP before deploying AI agents?
Not necessarily all of it, but you do need one trustworthy source of truth for the specific task the agent will perform. Agents act on records. If the record for that task is split across systems and reconciled manually, the agent inherits the ambiguity and the cost of resolving it. The practical route is to pick one workflow, consolidate the data behind that workflow, then deploy against it.
Are AI agents actually being used in production yet?
Yes, but narrowly. KPMG's Q2 2026 pulse found 53% of large United States enterprises had deployed AI agents, with 35% still piloting. Stanford's AI Index 2026 found agentic deployment still in single digits in nearly every business function, even while 88% of organisations use AI somewhere. Deployment is broad. Autonomy is not.
What should an AI agent audit log contain?
At minimum: the goal given, the plan generated, every tool call with its inputs and outputs, the retrieved evidence used, the service identity and permission scope under which each write occurred, the token and cost figure for the task, and the compensating action available to reverse each write. If any of those is reconstructed after the fact rather than recorded at the time, it is not an audit log.
How much does it cost to run an AI agent?
There is no useful list price, because agents consume compute per loop rather than per seat. The number to demand from any vendor is cost per completed task at your transaction volume, including review and exception handling time. Gartner's earlier estimate for full generative AI deployments ranged from five to twenty million dollars depending on approach, which is why narrow single task deployments are the sensible starting point.
What is the first AI agent use case a mid sized company should try?
One that already has a ledger to check itself against. Three way invoice matching, quality deduction calculations, dealer credit checks, warranty claim triage, or retrieval over an internal document set. Each has a defined correct answer, a bounded set of actions and a cheap way to verify the output, which is what makes the result measurable in weeks rather than quarters.
Working out whether your data can support an agent before you buy one? Auriga IT runs a structured AI readiness assessment that scores the specific workflow you have in mind against the data, integration and governance requirements it will actually need. Our white paper on implementing agentic AI covers the architecture in more depth, and the AI services overview sets out where we build.
Talk to our AI teamThe five capability tests, the twelve demo questions and the 30 day evaluation plan are free to reuse with attribution. If you are quoting the Gartner, KPMG or Informatica figures, please cite the primary source in the list below alongside this page.
Sources
- Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, June 2025.
- Gartner, 40% of Enterprise Apps Will Feature Task Specific AI Agents by 2026, August 2025.
- Gartner, Lack of AI Ready Data Puts AI Projects at Risk, February 2025.
- Gartner, Embedded AI in Cloud ERP Will Drive a 30% Faster Financial Close by 2028, February 2026.
- Gartner, Supply Chain Management Software With Agentic AI to Reach 53 Billion by 2030, April 2026.
- Informatica, CDO Insights 2026, January 2026.
- KPMG, AI Quarterly Pulse Survey, Q2 2026.
- McKinsey, The State of AI, 2026.
- BCG, The Widening AI Value Gap, September 2025.
- IBM Institute for Business Value, CEO Study, May 2025.
- Deloitte, State of AI in the Enterprise 2026.
- Deloitte India, Indian Enterprises Lead Global Peers in At Scale AI Adoption, March 2026.
- Stanford HAI, AI Index Report 2026, Chapter 4.
All external figures were verified against the linked primary sources in September 2026. Vendor forecasts are projections, not outcomes.
Related content
Auriga: Leveling Up for Enterprise Growth!
Auriga’s journey began in 2010 crafting products for India’s [...]






