If you are looking for the single best AI model to put in front of your customers, the answer is that it does not exist. Anyone who names a definitive winner is guessing. The technical leadership between the major providers changes every few months. What was the most capable model in January is often outdated by July.
Instead of looking for a permanent winner, you must compare what actually matters to a small business. You need to know the cost per conversation, how fast the model replies, and its refusal behaviour when it does not know the answer. We checked the current landscape against vendor pricing pages on 4 September 2026. The right choice depends entirely on the job you need it to do.
The short version- OpenAI retired the GPT-4o mini model from ChatGPT in February 2026, making it a legacy system.
- An adversarial test conducted on 2 and 3 September 2026 generated 131 real conversations and revealed 33 reproducible faults.
- Fixing 27 of those 33 reproducible faults required adding fourteen rules to the instructions and correcting one code defect, rather than changing the AI model.
- A standard customer conversation using 2,000 tokens of input and 500 tokens of output costs $0.0090 on Claude Sonnet 5 based on 4 September 2026 pricing.
- The iHayz Website Assistant includes 500 conversations a month for a $79 monthly fee, following an initial $699 build fee.

There is no best model, and anyone who names one is guessing
A small business owner does not need an artificial intelligence that can write software or pass a medical exam. You need a system that can read your shop policies, look at your calendar, and answer a customer accurately without inventing facts.
When you read articles declaring one company the winner of the AI race, they are usually referencing benchmark tests that have nothing to do with running a shop, a clinic or a workshop. A model that scores highly on logic puzzles might still struggle to politely tell a customer that you do not offer refunds on bespoke items.
The market moves too quickly for loyalty. We build and run customer-facing agents on several different model families for different clients, so the comparison here is from operating them, not from a leaderboard. We see firsthand that relying on one provider is a mistake. You must build your assistant so that you can swap the underlying model the day a better or cheaper one is released.
What you should measure is utility. Cost per conversation tells you if the system is financially viable. Speed tells you if the customer will wait for the reply or abandon the chat. Refusal behaviour tells you if the model knows how to stop and ask for help when it reaches the edge of its instructions. These are the metrics that keep a business safe.
The models, side by side
To understand the current options, you need to look at the models available for customer chat in September 2026. We checked these figures against vendor pricing pages on 4 September 2026.
The major providers offer families of models, tiered by size and capability. OpenAI currently sells GPT-5.6 Sol, Terra, and Luna. Anthropic sells Opus 5, Sonnet 5, Fable 5.1, and Haiku 4.5. Google sells Gemini 3.7 Flash, 3.6 Flash, 3.5 Flash, 3.5 Flash-Lite, and 3.1 Pro.
If you are reading older advice, you might see recommendations for GPT-4o mini. That model was retired from ChatGPT in February 2026 and is now legacy. Do not use it.
For a small business assistant, you are generally choosing between the fast, inexpensive models and the mid-tier models that balance speed with higher reasoning. Here is how the current representative models compare.
| Provider | Model | Input Cost (per million tokens) | Output Cost (per million tokens) | Context Window |
|---|
| OpenAI | GPT-5.6 Luna | $0.20 | $1.20 | Large enough for any small-business use |
| Google | Gemini 3.7 Flash | $0.75 | $3.75 | 1,048,576 tokens |
| Anthropic | Claude Sonnet 5 | $2.00 | $10.00 | 1,000,000 tokens |
The input cost is what you pay for the model to read your instructions and the customer's message. The output cost is what you pay for the model to write its reply. The context window is the maximum amount of text the model can hold in its short-term memory at one time. A context window of one million tokens is large enough to hold several thick books of text, which is more than enough for any small business chat history.
When you look at this table, you can see that Claude Sonnet 5 costs significantly more than GPT-5.6 Luna. However, cost is only one factor. Sonnet 5 is a larger model, which means it is better at following complex, multi-step instructions without getting confused. Luna is smaller and faster, making it ideal for simple, repetitive tasks.

What a conversation actually costs you
Token pricing is difficult to conceptualise. A small business owner cares about the cost of a real customer conversation, not the cost of a million abstract units.
A token is roughly three-quarters of a word. When a customer sends a message, the model does not just read that single message. It reads your background instructions, your business rules, the previous messages in the chat, and the new question. All of this is the input. The model's typed response is the output.
We do not know exactly how many words a customer will type, but we can measure a typical exchange. A standard customer conversation about a booking or a product inquiry usually involves the model reading about 2,000 tokens of input and writing about 500 tokens of output.
Here is the arithmetic for a single conversation using the verified prices from 4 September 2026.
Using GPT-5.6 Luna:
- Input: 2,000 tokens at $0.20 per million = $0.0004
- Output: 500 tokens at $1.20 per million = $0.0006
- Total cost for one conversation: $0.0010
Using Gemini 3.7 Flash:
- Input: 2,000 tokens at $0.75 per million = $0.0015
- Output: 500 tokens at $3.75 per million = $0.001875
- Total cost for one conversation: $0.003375
Using Claude Sonnet 5:
- Input: 2,000 tokens at $2.00 per million = $0.0040
- Output: 500 tokens at $10.00 per million = $0.0050
- Total cost for one conversation: $0.0090
Now, let us translate that single conversation cost into a monthly volume. A local clinic might handle 200 conversations a month. A busy trades firm might handle 500. A popular online store might handle 2,000.
At 200 conversations a month:
- GPT-5.6 Luna: $0.20
- Gemini 3.7 Flash: $0.68
- Claude Sonnet 5: $1.80
At 500 conversations a month:
- GPT-5.6 Luna: $0.50
- Gemini 3.7 Flash: $1.69
- Claude Sonnet 5: $4.50
At 2000 conversations a month:
- GPT-5.6 Luna: $2.00
- Gemini 3.7 Flash: $6.75
- Claude Sonnet 5: $18.00
The arithmetic shows a clear reality. The raw cost of the artificial intelligence thinking and typing is remarkably low. Even at 2,000 conversations a month using the most expensive model in our comparison, the token cost is under twenty dollars. The true cost of an AI assistant is not the token usage. The true cost is the software built around the model to keep it secure, the hosting, and the time spent writing the rules it must follow.

Fast model or slow model? Use both
Because the models have different strengths, using one model for everything is the wrong shape for a business assistant.
If a customer asks a shop for its opening hours, or asks a clinic for its address, the assistant does not need deep reasoning capabilities. It just needs to fetch a fact and reply instantly. For these routine turns, you route the conversation to a fast, cheap model like GPT-5.6 Luna or Gemini 3.7 Flash. The customer sees a reply in milliseconds, and you pay a fraction of a cent.
However, conversations often become complicated. If a customer asks a supplier to compare three materials for a job with unusual wear, or asks a service firm to diagnose a fault based on a description of a previous dye job, a small model might struggle. It might apply your rules incorrectly, or it might give a generic answer.
For these hard turns, the system should route the conversation to a slower, more accurate model like Claude Sonnet 5. The customer might wait an extra second for the reply, but the answer will be precise and carefully reasoned.
This routing changes your cost structure. Instead of paying the higher Sonnet 5 rate for every simple greeting and address request, you only pay it when the complexity of the question demands it. It also changes how quickly the customer sees a reply. Routine questions feel instant. Complex questions take a moment longer, which mimics how a human would pause to think before answering a difficult query.

What breaks has almost nothing to do with the model
There is a common misconception that when an AI assistant makes a mistake, the model itself is unintelligent or broken. Our own measurements show this is rarely the case.
We ran an adversarial test against our own live public assistant on 2 and 3 September 2026. We used nine different attackers across nine dimensions, generating 131 real conversations. The goal was to force the assistant to break its instructions, invent information, or behave badly.
When we reviewed the transcripts, we found 33 faults that reproduced when every reported fault was replayed mechanically against the live endpoint.
Several dramatic-sounding findings evaporated on replay. For example, an attacker claimed the assistant invented a usage price that nobody published. When we checked the transcript and the source material, the price the assistant quoted was printed verbatim on our published price sheet. The model was correct; the attacker was wrong.
The faults that did reproduce clustered into specific, predictable behaviours. In six separate findings, the assistant kept trying to sell to an angry customer after being explicitly told by the customer to fetch a human. In other instances, it answered a specific question with a generic menu of services instead of detailing the one service requested. It named the internal tools behind the build. It invented one usage figure. It asked twice for an email address the visitor had already given.
Almost none of these faults were the model being unintelligent. They were missing rules. The model did not know it was supposed to stop selling when a customer used angry language because we had not explicitly written a rule telling it to do so. It asked for the email address twice because the instruction on how to check the chat history for previous details was not strict enough.
State plainly: 27 of those 33 faults no longer reproduce. We fixed them by adding fourteen rules to the instructions and fixing one code defect. We did not change the model. We verified the fixes by replaying the exact same transcripts against the updated rules.
You can read the whole account, including what is still broken, at ihayz.com/work/what-goes-wrong-when-ai-answers-customers/

The one that was a real bug
During that same adversarial test, we monitored the system for data leaks. Across all 131 conversations, nothing leaked from the model itself. The system instructions remained secure, the model identity stayed hidden, and the vendor was not exposed.
However, we did find one genuine security bug. It was a cross-visitor session leak.
When a user visits a website and opens a chat, the system assigns them a session ID to remember their specific conversation. To make the assistant reply faster, systems often use a caching layer. Caching simply means storing a ready-made answer for a common question so the system does not have to ask the AI model to generate it from scratch every time.
The bug occurred because a cached first reply carried the session ID of whoever asked the question first. When a second visitor asked the same common question, they were handed the cached reply, which inadvertently included the first visitor's session ID. Because they now shared a session ID, the two visitors could be handed the same conversation memory.
This was a code fault in the caching layer, not a model fault. We reproduced it by hand and closed it the same morning.
We use this example to make a specific point. The plumbing is where the danger is. Small business owners worry that the artificial intelligence will suddenly go rogue and give away company secrets. The reality is much more mundane. The risk lies in the standard software surrounding the model—the databases, the caching layers, and the routing logic.

The test to run before you trust any of them
You should not put an AI assistant in front of your customers without testing it thoroughly. You do not need to be a software engineer to do this. You can run a reproducible method yourself in an afternoon.
First, define your attacks. You need nine kinds of attack to test the boundaries. You should test anger (yelling at the assistant), confusion (asking contradictory questions), scope (asking about a competitor's products), discounts (demanding a lower price), and instruction bypass (telling the assistant to ignore its previous rules).
Have a staff member or a friend act as the attacker and have these conversations with your assistant. Save the transcripts.
When you find a fault, you must replay it mechanically. This means taking the exact text the attacker typed and pasting it into a new, fresh conversation with the assistant. You must do this to see if the fault reproduces. Often, an AI will make a one-off mistake due to the randomness inherent in how it generates text. If you paste the exact same text three times and the assistant only makes the mistake once, it is a random error. If it makes the mistake all three times, you have a reproducible fault. Count what reproduces.
Do not rely on a model judging its own output. Some builders use a second AI model to read the transcripts and grade the first model's performance. This is not a test. It is circular logic. A model cannot reliably tell you if another model is hallucinating facts about your specific business, because neither model actually works there. Only a human reading the reproducible faults can tell you if the assistant is breaking your business rules.

So which should you actually use?
The verdict is a decision framework, not a single winner. The model you use depends on the complexity of the conversation. You should use a fast, inexpensive model like GPT-5.6 Luna or Gemini 3.7 Flash for routine questions, and route complex, multi-step problems to a larger model like Claude Sonnet 5.
When you hire someone to build this for you, you must insist on specific standards. Insist on published prices so you know exactly what the token markup is. Insist on approval gates, meaning the AI cannot process a refund or change a booking without a human clicking a button to approve it. Insist on strict refusal behaviour, so the assistant knows how to say it does not know the answer. Finally, insist on a written record of what was tested before the system goes live.
How much does it cost to run an AI chatbot for my business?+
The raw token cost of the artificial intelligence thinking and typing is remarkably low. Processing 2,000 conversations a month using Claude Sonnet 5 costs $18.00. The true cost of an AI assistant lies in the software built around the model, the hosting, and the time spent writing rules.
Should I use a fast AI model or a slower, more accurate one?+
You should use both by routing different types of questions to different models. Routine tasks like fetching an address can go to a fast, inexpensive model like GPT-5.6 Luna or Gemini 3.7 Flash. Complex questions should be routed to a slower, more accurate model like Claude Sonnet 5.
Why does my AI assistant give incorrect answers to customers?+
When an AI assistant makes a mistake, it is rarely because the model itself is unintelligent or broken. Faults usually occur because the system is missing specific rules in its instructions. Adding explicit rules to the instructions fixes most predictable behaviours without needing to change the underlying model.
How do I test my AI chatbot before making it public?+
You should have a staff member act as an attacker across nine dimensions, including anger, confusion, and instruction bypass. When you find a fault, you must replay the exact text mechanically in a fresh conversation to see if the error reproduces. You should rely on a human reading these reproducible faults rather than using another AI model to grade the transcripts.
Is it safe to let an AI assistant handle customer data?+
The AI models themselves rarely leak data, but the standard software surrounding them can pose security risks. For example, a code fault in a caching layer can inadvertently share one visitor's session ID and conversation memory with another visitor. You must ensure the plumbing and routing logic are secure to keep business data safe.
What standards should I demand if I hire someone to build an AI assistant?+
You must insist on published prices so you know the exact token markup being charged. You should also require approval gates so the AI cannot process refunds or change bookings without human permission. Finally, demand strict refusal behaviour and a written record of what was tested before the system goes live.
For most small-business owners, comparing AI models is a distraction from running the business. You likely do not have the time or desire to build a custom chat assistant, test its responses, or maintain the software as the technology changes. Managing the technical side of artificial intelligence is rarely a priority when you have daily operations to oversee.
If you prefer not to do this yourself, iHayz offers a single alternative called a Website Assistant. This is a knowledge assistant that answers questions on your website using your own business documents. The published pricing starts at a $699 build fee, plus $79 a month which includes 500 conversations. All questions are answered and everything is quoted in writing rather than on calls.

Try the test on our own assistantAsk it the questions from this article and see what it refusesAsk it →