Online shops

The hard part is stopping it saying too little

Everybody worries their assistant will say something it should not. Almost nobody worries about the failure that actually costs money — and that one leaves no trace at all.

8 September 20266 min readDaniel

Put an assistant on a shop and the first conversation is always about what it must not say. Fair enough — that conversation matters, and in some trades it is the whole project.

But it produces a predictable outcome. The list of forbidden things gets long, the instruction becomes “when in doubt, refuse”, and within a week you have an assistant that answers roughly one question in three.

Nobody complains about that. Customers do not write in to say the chat was unhelpful. They close it, and the shop learns nothing.

A refusal is not a safe answer. It is a lost customer with a clean conscience.

Three ways it goes wrong

Only the first of these is the one people prepare for.

It says too much

The obvious failure. It invents a fact, states something it is not entitled to state, or turns a question about a person into a recommendation. Visible, embarrassing, and the one everybody plans for.

It says too little

Refuses a question it could have answered completely, because a word in it looked risky. Invisible, unreported, and by a wide margin the more expensive of the two.

It sidesteps

The worst of the three, and it looks like caution. “I cannot answer that, but you might be interested in…” makes exactly the connection that was not allowed, while sounding careful about it.

The third one has to be forbidden explicitly. Left alone, a model does it by itself — it is trying to be helpful, and sidestepping is what helpfulness looks like when a rule is in the way.

Four levels instead of two

Most assistants are built with two categories: allowed and forbidden. That is why so many of them are useless. Everything interesting happens in between.

A question about the product. Is this vegan, how many are in a pack, what does a day cost, how long does it last. Answer fully, from checked data. These are the questions that quietly stop a purchase, and they are the entire commercial case for having an assistant.

A question about what something does. Which products contain a given ingredient, what the difference between two of them is, what a category is generally for. Answer fully, with links. This is the level that gets wrongly refused, and refusing it is the single largest loss in the whole system.

Someone describing themselves. Not a question about the range any more, but about a person. Stop naming products and hand over to a human. It can still explain in general terms — refusing to speak at all is unhelpful — but a description of a person must not become a recommendation.

Something it must not answer at all. Hand over, and do not sidestep. This is where the third failure above lives, and where an explicit instruction is the only thing that prevents it.

The two middle levels are the entire craft. Get them wrong in one direction and the assistant is a liability; wrong in the other and it is decoration.

The list nobody expects to need

In any trade with rules about wording, there are pairs of phrases that mean roughly the same thing to a person and something quite different to a regulator. One is fine, the other is not.

A model cannot infer which is which. It will either refuse both, or use both. So the pairs have to be written out, one by one, and handed to it — which is tedious, unglamorous, and the difference between an assistant that works and one that annoys people.

This is also the part that cannot be bought. A widget arrives with generic caution, not with your trade’s vocabulary, because the vendor does not know your trade.

Two findings that surprised us

From building one

  • 4levels of behaviour, not the usual two
  • 10messages of history before rules start to slip
  • €4roughly, to run for a month

Long conversations drift. The further a chat runs, the more likely a model is to wander from its instructions. Keeping roughly the last ten messages holds behaviour steady, and has the side effect of costing less. We did it for the cost and kept it for the discipline.

It is far cheaper than people assume. Reading facts from checked material under strict rules is not demanding work, so a small model does it well. Most of the monthly figure people fear comes from choosing a large model for a task that never needed one.

What this means if you are considering one

The question to ask a supplier is not what the assistant can do. Everything can do everything now. Ask what it will refuse, and how that list was arrived at.

If the answer is a generic safety setting, you will get generic caution — and the invisible failure, the one nobody reports, on every second question.

Ask what it will refuse, and who decided.

The other question worth asking is whose account it runs on. If it is the agency’s, you are renting something that stops working the day you change supplier. If it is yours, you own it.

How we build them, including where we think the limits are, is on our custom AI assistants page. And if what you actually need is for your product pages to answer these questions before anyone opens a chat, that is the cheaper fix and we would start there.

Tell us what yours must never say.

That is the useful half of the first conversation, and it is the half most suppliers skip. Bring the list, or bring the problem and we will work out the list together.