Blog

Agentic AI Chatbots: What They Can Complete, Not Just Answer

Article hero image
In brief. Most chatbots answer questions. An agentic chatbot finishes the task: it checks availability, takes the reservation, drafts the quote, files the intake, and writes the result back to your system of record. That difference is what is worth paying for, and it is also where the risk lives. So this post covers what they can complete today, based on agents we run in production on Amazon Connect, and the four places they still fail.

Answering is search with better manners. Completing is revenue.

A customer types "do you have a 10×10 near Elmwood Park?" A conventional chatbot looks that up and tells them. An agentic chatbot looks it up, offers the two units that fit, holds one, sends a payment link, and texts a confirmation with the gate code, then updates the CRM so the store manager sees a rental, not a chat log.

The first is helpful. The second is a rental that would otherwise have waited until the office opened, which for anything a customer wants at 7:42 PM often means a rental lost to whoever answered first.

The technical difference is that the second bot has tools: a small set of functions it is allowed to call (search units, reserve a unit, fetch promotions, add a protection plan), each with a narrow contract and a system behind it. "Agentic" means the model decides which tool to call, in what order, with what arguments, based on the conversation. Everything useful and everything dangerous about these systems follows from that.

Three workflows agentic chatbots complete today

1. A storage rental, after hours, by voice or text

This is the one we have run longest, for a multi-site self-storage operator, so it is the one we can describe in the most detail.

The caller says what they need. The agent asks the two or three questions that narrow it (size, climate control, drive-up, move-in date) and calls the unit search. It reads back the matches with the current promotion applied, because promotions come from a tool call, not from a document that was accurate last month. The caller picks one. The agent reserves it, collects the name and phone number, and sends a payment link by SMS rather than taking a card over the phone. When the link is paid, the reservation is confirmed in the property management system and the caller gets a text with the details.

What the human never touches: the availability check, the hold, the promotion math, the CRM update, the confirmation. What the human still owns: the exceptions. A caller with a past-due account, a question about a lien, or a size the agent cannot resolve gets handed to a person with the transcript attached, during business hours, or gets a callback booked outside them.

On the phone this runs on Amazon Connect with a speech-to-speech model, so the agent listens and talks without a separate transcription step in the middle. By SMS the same tools sit behind a text agent with a knowledge base for the questions that are not transactions: access hours, what fits in a 10×10, whether you can store a car.

2. A quote that routes itself for approval

The pattern for quoting is the same shape with different tools. The agent asks the intake questions your sales team would ask (what, how many, where, when) and feeds the answers to your pricing rules. It drafts the quote. If the total is under the approval threshold, it sends it. If it is over, or if the customer asked for something the rules do not cover, it routes the draft to a person with the conversation attached and tells the customer when to expect it.

Two details decide whether this works. The pricing rules have to be a tool the agent calls, not a paragraph in its prompt; a prompt will negotiate, a tool will not. And every sent quote has to be logged against the account in the CRM as the sent version, because the first question a salesperson asks when the customer calls back is "what did we tell them?"

3. Intake that creates the record, not a ticket about the record

Claims, service requests, applications, warranty registrations: anything that arrives as a form and then waits for someone to re-key it. The agent captures the fields in conversation, validates them against your rules as it goes, so a policy number that does not match the pattern gets corrected in the moment rather than rejected three days later, and creates the record in the system of record directly. Attachments come in by a link the agent sends. Missing information generates a follow-up on a schedule rather than a dead intake.

The person reviews the exceptions: the intake that failed validation twice, the request that matched a fraud rule, the claim over a dollar threshold. Not the pile.

What makes an agent safe enough to let it write to your systems

Four things, and they are all boring.

  • Few tools, narrow contracts. A production storage agent has a handful of tools, each doing one thing with typed arguments. An agent with forty tools picks the wrong one more often, and picks it confidently.
  • A domain lock. The agent is told what it is for and instructed to decline everything else, politely. A storage agent that will also discuss the weather will eventually discuss your competitor's pricing.
  • Memory that lives outside the model. Session state in a database keyed by the conversation, so a customer who comes back after the payment link can pick up where they left off, and so a person picking up an escalation can see what happened.
  • An exit the agent is allowed to take. A defined handoff, to a queue during hours or a callback outside them, with the transcript attached. An agent that cannot escalate will improvise, and improvisation is where wrong data gets written.

Then you measure it. Amazon Connect now scores whether the agent achieved the goal and whether it called its tools correctly; we wrote up which metrics to gate a go-live on and why handoff rate is the wrong target.

Where they still fail

These are the four failure modes we spend the most time on in production, and none of them is solved by a better prompt.

Capturing names and email addresses by voice

Speech recognition is very good at sentences and poor at "S-Z-C-Z-E-P-A-N-S-K-I at gmail dot com." Names get normalized to the nearest common spelling; email addresses come back with the domain right and the local part wrong. Reading it back for confirmation helps less than you would hope, because callers say "yes" to a read-back that is close.

What works: do not collect email by voice at all. Collect the mobile number, which the telephony already has and the caller can confirm in four digits, and move the email capture into the payment link or a text reply, where the customer types it. Reserve voice for what voice is good at.

Latency on tool calls

A tool call that takes four seconds is a four-second silence on a phone line, and silence reads as a dropped call. Chat is more forgiving; voice is not. The fixes stack: keep tool chains short, so a single turn does not trigger three sequential calls; have the agent acknowledge before it calls ("let me check that for you") so the pause is expected; and where a step is unavoidably slow, run it asynchronously and let the agent carry on. Speech-to-speech models with response pacing mask some of this, but they mask it. They do not remove it. Measure invocation latency as a percentile, not a mean, or the tail will surprise you on the recordings.

Escalation timing

Agents fail in both directions. Too eager, and every mildly unusual request becomes a transfer, which makes the agent a more expensive IVR. Too reluctant, and the agent loops, asking the same clarifying question a third time while the customer's patience runs out, because nothing told it when to stop. The second is worse, and it is the one you create by targeting a low handoff rate.

What works: explicit exit conditions in the agent's instructions (two failed attempts at the same field, any mention of a dispute or a lien, any request outside the domain), and a transfer that carries the transcript so the customer does not repeat themselves. Then read handoff rate only next to goal success. Handoffs falling while goal success holds is progress. Handoffs falling while goal success falls is an agent that has learned to keep people on the line.

Answering from a stale knowledge base

The text agent answers "what's the promotion this month?" from a document. The document was right when it was uploaded. This is the quietest failure, because the answer sounds right and the customer only finds out at the counter. Anything that changes weekly (promotions, availability, prices, hours over a holiday) should be a tool call to the system that owns it, and the knowledge base should be reserved for things that are true for a year.

Should you deploy one?

An agentic chatbot pays for itself when three things are true: the workflow is repeatable and has rules a person could write down; there is a system of record it can read from and write to; and there is enough volume that the after-hours and overflow share is worth capturing, which usually means a few hundred transactions a month at minimum. If your questions are mostly informational, a conventional chatbot with a good knowledge base is cheaper and fails more gracefully. If you have the rental, the quote, or the intake, the agent is the one that changes the numbers.

Where ExecuteCX fits

We build and operate agentic voice and chat agents on Amazon Connect that complete the work: the storage rental above has been in production since late 2025. Start with our AI voice and chat agent services, see how the same agents run workflow automation behind the scenes, or request a free assessment: thirty days, your call data, and a ranked list of which workflows an agent could complete for you.

Find out if your contact center is ready for AI

Start with a free 30-minute call: an AI Readiness Check with a one-page summary. No pitch.

Book a Free 30-Minute Call →