Agentic CRM: What Actually Happens When An Agent Gets It Wrong

This is the question underneath most of the hesitation around agentic CRM, and most people asking it do not have an agent running yet.

That is actually good news, because the useful question is not whether an agent will eventually get something wrong. It will. Humans already do too.

Introduction

The useful question is what happens next: whether the mistake is caught before it reaches a customer, whether the system knows when to escalate, and whether the same mistake becomes less likely to happen again.

That changes how you should think about introducing agents. The goal is not to prove they are infallible before they are allowed anywhere near CRM. It is to build the guardrails, measurement and review process first, so autonomy increases only when the evidence says it should.

 

Step One: Write Your Guardrails Down, Before You Have Anything To Apply Them To

You do not need any automation in place to do this, and it is probably the cheapest, highest-leverage step on this entire list.

Write down, in plain sentences, the rules you would want enforced no matter who or what is running your campaigns.

No marketing message goes to somebody with an unresolved complaint. No discount above a certain size goes out without a specific person approving it. No customer who has opted out gets added back into a campaign because another system says they are eligible. No message referencing a sensitive account issue gets sent without review.

Then separate those rules into two types. 

The first are hard rules. These should never be broken and can usually be expressed deterministically: eligible or ineligible, above or below a threshold, consent present or absent.

The second are judgement calls. These require context. Is this tone appropriate after a complaint? Is this offer sensible for this particular customer? Does this message feel too aggressive?

That distinction matters enormously once agents arrive.

Wherever possible, hard rules should sit outside the agent. If somebody is ineligible for a campaign, the safest system is one where the agent cannot select them in the first place, rather than one where a prompt politely asks it not to.

Step Two: Measure How Often Your Current, Fully Manual Process Already Gets Things Wrong

Before comparing an agent to some imagined perfect standard, establish an honest baseline for what happens today. Look back over the last few months of campaigns and count the mistakes a customer could actually have noticed.

A wrong name. A broken personalisation field. An outdated offer. A customer included in a campaign they should have been excluded from. A message sent too soon after a complaint. A promotion for something they already own. A discount that should have required approval. Do not put everything into one generic “error” bucket either. Start classifying mistakes by consequence:

Low severity: formatting problems, awkward copy or minor personalisation errors that are unlikely to materially affect the customer.

Medium severity: irrelevant offers, incorrect customer information or badly timed messages that could damage the experience.

High severity: consent failures, sensitive-data exposure, inappropriate contact after a serious complaint, significant pricing errors or anything likely to create regulatory, financial or reputational consequences.

This gives you two baselines rather than one: how often things go wrong and how badly they go wrong. That distinction becomes important later. An agent that produces fewer errors overall but occasionally produces much more serious ones is not necessarily an improvement.

Most CRM teams have never formally measured their manual error rate because human review creates a feeling of control. But human review and error-free execution are not the same thing. Your existing process is the benchmark. Measure it before automation changes anything.

Book A Call

Expert help is only a call away. We are always happy to give advice, offer an impartial opinion and put you on the right track. Book a call with a member of our friendly team today.

Step Three: Automate The Simplest, Most Mechanical Checks First

Before AI writes a single line of customer-facing copy, automate the boring checks.

Validate that personalisation tokens are populated. Check that every recipient has the required consent status. Confirm that suppression lists have been applied. Flag impossible or unusual values. Prevent offers above an agreed threshold.Check that required fields are present before a campaign can move forward.

None of this is particularly agentic. That is exactly the point.

A deterministic rule is preferable when the answer genuinely is deterministic. You do not need an AI model to decide whether a required field is null or whether somebody appears on a suppression list.

These simple checks remove whole categories of mistakes cheaply and consistently. They also create an important separation of responsibilities: automation handles rules; agents handle judgement.

By the time AI enters the workflow, it should not be wasting intelligence on problems a basic validation rule could have prevented.

Step Four: Introduce Drafting With Full Human Review, Not Autonomy

When you start using AI to draft CRM messages, resist the temptation to measure success by how quickly you can remove the human. Start with every output being reviewed before it ships. At this stage, efficiency is not the primary objective. You are collecting evidence. For every draft, record whether the reviewer:

Approved it unchanged.

Made a minor edit.

Made a substantial edit.

Rejected it entirely.

And importantly, record why.

Was the problem factual accuracy? Brand voice? Offer selection? Customer context? Compliance? Personalisation? Timing? Something else? After enough volume, this becomes much more useful than a vague impression that “the AI seems pretty good”.

You might discover that it is excellent at routine reactivation messages but consistently weak when a customer has recently contacted support. Or that the copy is almost always usable but humans repeatedly change the same type of call to action. Those patterns tell you what to fix. They also tell you where autonomy may eventually be appropriate and where human judgement is still earning its place.

 
 

Step Five: Give Every Action A Risk Level

This is the step that makes gradual autonomy practical. Not every CRM action deserves the same level of oversight.

Sending a routine reminder to an established customer is not the same risk as issuing a large retention discount. Changing a subject line is not the same as deciding whether somebody with an unresolved complaint should receive an upsell. Give actions a simple risk classification.

Low risk: routine messages, low-value offers, established templates and segments with few sensitive edge cases.

Medium risk: meaningful discounts, behavioural personalisation, win-back campaigns or messages where incorrect context could noticeably damage the customer experience.

High risk: sensitive customer situations, complaints, vulnerable customers, unusual financial decisions, significant discounts or anything with material compliance implications.

Then make the review requirement follow the risk. Low-risk actions can eventually move towards sampling and automated QA. Medium-risk actions may require approval when particular conditions are met. High-risk actions stay human-reviewed. This is much more useful than asking whether “the agent” is autonomous, because autonomy is not one switch. A good agentic CRM system might be highly autonomous in one narrow part of the customer journey and deliberately unable to act without approval in another.

Step Six: Reduce Review Gradually, Using Your Own Numbers As The Signal

Once you have real volume through full human review, you can start deciding what happens next. 

Your override rate is one of the most useful signals. If humans are approving 500 routine messages and changing only a tiny proportion, you have evidence that reviewing every single one may no longer be the best use of their time.

But do not look at the headline override rate alone. Break it down by campaign type, customer segment, action and risk level. An agent could have an excellent overall approval rate because 90% of its work consists of easy messages while still performing poorly on the 10% where mistakes matter most.

This is why autonomy should be earned in slices. Move one low-risk use case from full review to sample review. Keep measuring it. If performance remains stable over meaningful volume, consider the next one. If the override rate increases, stop.

That is not a failed experiment. It is the system doing exactly what it was designed to do: showing you where autonomy has outrun the evidence.

Step Seven: Decide What Happens When Something Actually Goes Wrong

Eventually, something will. Plan for that before it happens.

For each risk level, decide what a failure should trigger. A minor copy issue might simply be logged and added to the next evaluation cycle. A repeated personalisation mistake might automatically move that campaign type back into full human review. A serious eligibility or compliance failure might pause the workflow entirely until somebody investigates it. The important thing is that autonomy should be reversible.

If an agent’s behaviour changes after a model update, prompt change, new data source or unusual customer situation, you should be able to move that workflow back to a safer review state quickly.

This is also why keeping an audit trail matters. For every meaningful automated decision, you should be able to reconstruct what happened: what customer context was available, what action was proposed, which rules were checked, whether a human intervened and what ultimately reached the customer.

Without that, investigating a mistake becomes guesswork.

The Data Worth Watching At Every Stage

Do not build one giant agent-performance score. Track a small number of measures that tell you different things.

Manual error baseline: How often the existing human process produces customer-visible errors, split by severity.

Approval rate: How often AI-generated work passes human review unchanged.

Override rate: How often a human changes the proposed action or message, and why.

Critical error rate: The frequency of high-severity mistakes. This deserves separate attention even when the overall error rate looks excellent.

Escalation rate: How often the agent asks for human intervention, broken down by reason and use case.

False escalation rate: How often humans review something that was actually straightforward. This helps you improve efficiency without simply suppressing escalation.

Post-send incident rate: Mistakes that made it through every control and reached the customer.

Performance by risk tier: Because an excellent aggregate score can hide poor performance on the cases that matter most.

And watch all of these over time. The average matters, but so does drift. A system that performed extremely well last quarter does not automatically deserve the same trust after its model, prompt, underlying data or available tools have changed.

Get In Touch

Our friendly team are always on hand to answer questions, troubleshoot problems and point you in the right direction.