"Just call ChatGPT" isn't the answer

A general-purpose language model does not reliably know your current hours, prices, or stock. Without authoritative source content, it can produce plausible but unsupported answers. A business chatbot needs more than access to a model.

The fix is to constrain the model to your actual business content and to put a thin rule layer on top that handles the things language models are unreliable at: deciding when to ask the visitor for their name, knowing when to stop pitching, recognizing a request for a live person, and routing that request to you quickly. None of this is exotic, but doing it well is the difference between a chatbot you tolerate and one you'd actually put on your site.

How grounded retrieval actually works

Grounded means the bot's reply has to come from content you gave it: your website text, an FAQ you maintain, policy documents, maybe an inventory feed. When a visitor asks a question, the system first looks up the relevant chunks of your content, then hands those to the language model with the instruction "answer only from this."

The retrieval step is where most of the quality lives. Two different techniques each catch things the other misses:

A well-built assistant runs both in parallel and fuses their rankings. The industry usually calls this reciprocal rank fusion, or RRF. That way a question like "I need something durable for daily commuting" gets help from the dense side (catches "durable" in a product description that uses "reinforced construction"), and a question like "ZT-500 price" gets help from the lexical side (pins the exact product to #1). Neither approach alone is enough for a business with both narrative answers and branded products.

Retrieval is also cost-aware. Every chunk of your content the system pulls into the prompt is a chunk the language model has to read, and prompt size is a real cost, both in money and in answer quality, since more context can dilute focus. A well-built assistant doesn't stuff every possibly-relevant FAQ entry and product chunk into every turn; it picks a number of chunks appropriate to the question shape. A two-word product lookup gets a tight inventory slice; a comparative recommendation gets a wider catalog; a hours-and-location question gets just the location FAQs. The system also skips retrieval entirely when the rule layer already knows the next step. For example, on a Fit Check qualifying-answer turn where the bot just needs to record "$8,000" as the budget answer, there's nothing for the FAQ to add. That's how a chatbot stays affordable at scale without quietly degrading the answers that need full context.

What you should not need to know as an operator: vector dimensions, embedding models, chunk sizes, rank fusion formulas. You should need to know: the bot answers from your content, it handles paraphrase, and it catches your product names exactly. If a vendor can't explain how the second of those works, that's a flag.

A rule-based layer on top of the LLM

Here's a thing that sounds boring but matters a lot: pure prompt engineering is unreliable. If you just tell a language model "always ask for the visitor's name before closing the lead," it will comply often enough to look fine in a demo, and not reliably enough to run a real lead-capture flow on. Missing the name on some leads is a real business cost, and the failures are hard to catch during testing because they look identical to the successes until you review a week of transcripts.

The fix is a small deterministic policy layer: a rule-based decision per turn about what shape the next reply should take. Should the bot just answer? Should it answer and then offer to connect the visitor with a real person? Is it time to ask for a name? A contact method? Is the visitor winding down, in which case the bot should stop pitching? Is the visitor frustrated, in which case the bot should acknowledge that before anything else?

Code applies the workflow rules, while language models may help interpret the visitor's intent and write the reply. That distinction matters: a rule can be explicit even when the language interpretation feeding it is uncertain. Both need testing.

Why this matters for you: it's how the bot stops asking for your email on the third turn in a row, how it refuses to offer to connect twice to the same visitor, how it knows to wind down the conversation instead of keeping the pitch going when you've clearly decided not to buy. None of that is a prompt trick. It's a policy layer.

The clearest example is qualifying before booking. Should the bot share your calendar link with every visitor, or only ones that fit your engagement criteria? Should it ask one criteria question at a time so the conversation doesn't feel like an interrogation? When a visitor doesn't fit, should they get a polite decline plus a helpful resource (a DIY guide, a partner referral) instead of a calendar slot? Those are policy decisions, not prompt decisions. A pure-prompt bot can be told "ask qualifying questions before sharing the booking link" and will mostly comply, but the moments where it doesn't are the moments your sales calendar fills with bad-fit meetings. The application should check the qualification result before offering booking, and provide a configured alternative when the visitor does not qualify. Test unclear answers and interrupted conversations as well as the straightforward path.

Session state and the "snapshot" lead

A chat is a conversation. The bot needs to remember things across turns. When a visitor says their name on turn 2 and their phone number on turn 5, the bot has to hold onto "name given" the whole time, even if the intervening turns talk about something else. That's session state.

The important constraint: state is per-session and time-bounded. It's not a running memory of "everything this visitor has ever told us." It's scoped to the conversation, it expires, and it doesn't train the underlying model. That matters for privacy and for predictability.

Lead capture is a related but separate event. SBB treats the initial lead as a snapshot of the information collected so far and uses deduplication to limit repeat notifications. Later actions, such as a booking or a new message for the owner, are tracked separately. Ask any vendor how it distinguishes those actions from retries, and how it surfaces delivery failures.

What a well-built chatbot deliberately doesn't do

Some of the most useful architectural choices are about what the bot won't do:

These are choices about source control. For business-specific answers, SBB uses content you configure rather than open-web search or automatic learning from chats. If another product uses broader sources or conversation learning, ask how those sources are reviewed and how errors are corrected.

"No open-internet search" doesn't mean the bot can't use anything beyond your FAQ. It means it can only use systems you've explicitly connected: an inventory feed (so it can check whether the ZT-500 is in stock), a Google Calendar (so a qualified visitor sees real available time slots in the chat instead of a "click here to book" link out to another page), a CRM (so the lead lands in the same pipeline your team already uses). These are bounded, auditable connections. Each one is opt-in, configured once, and revocable at any time. These connections define which external systems the bot can use. It can do useful things within the systems you've authorized, like presenting a list of bookable slots that you actually have free, instead of guessing at availability.

Per-industry tuning, briefly

Different industries need different routing rules. A dental chatbot may flag urgent language and show configured urgent-care guidance; it should not diagnose the visitor. An e-commerce chatbot may prioritize product identifiers and order questions. Test the language and escalation paths relevant to your business.

None of this is glamorous. It's the difference between a generic "AI for any business" that kind of works for everyone and a per-industry assistant that actually belongs on your site.

How Simple Business Bots handles each of these

Questions to ask any chat vendor

If you want a short checklist for evaluating any AI chat vendor, these five questions help you assess the system beyond its chat window:

  1. "Can your bot look things up on the open internet, or only from content and systems I've explicitly provided or connected?" The right answer is only content and systems I've connected (your FAQ, website, inventory feed, calendar, CRM, etc.), not the open internet.
  2. "How does the bot decide when to hand off to me?" The right answer is some version of a rule, not the model decides on its own.
  3. "What does the bot do when it can't answer?" The right answer is admits it, captures contact info, offers a handoff, not guesses and not replies with a dead end.
  4. "Do customer conversations train the underlying model?" The right answer is no.
  5. "Is there behavior specific to my industry, or is it the same bot for everyone?" Industry tuning isn't always essential, but it's a meaningful signal about how much care went into the product.