Skip to content
StriqTech Logo
Guides & Comparisons

ManyChat or a custom AI chatbot: the 5-message test that decides it

Five specific messages, sent this afternoon, tell you whether the cheap platform is enough for you. Two of the five cannot be fixed by configuring.

10 min readStriqTech

The demo went perfectly: the bot gave the opening hours, recited the price of the best-selling business cards and sent the location. Three weeks later a customer writes "hi, can you confirm whether job 4821 is ready for pickup?" and the bot returns the main menu. It is not a configuration problem: that piece of data does not live in the platform, it lives in your management system, and no amount of buttons is going to fetch it.

On one side there are the flow platforms —ManyChat, Landbot, WATI and their lookalikes—: monthly subscription, set up in one afternoon, a tree of buttons. On the other, a bot connected to your management system, your stock or your CRM, with development and setup ahead of you. Comparing feature lists decides nothing, because both categories say the same thing on their websites: 24/7 service, automatic replies, WhatsApp integration.

What does decide it is five messages you can send this afternoon to any bot —your competitor's, the demo's, the one you already have—, without knowing anything technical: one with two questions at once, one written the way your customer writes and not the way the script imagined it, one asking for a piece of data that only exists inside your system, one where you change your mind halfway through an order, and a complaint asking to speak to a person. The first, the second and the fourth are fixed with configuration and a low budget. The third and the fifth are not fixed by configuring: they need the bot to be connected to your systems and to an inbox with people in it. That is the real dividing line between the two categories, not the price. This post does not quote anything: it tells you which side of the line your business is on today.

The five messages, in this order, with the text to copy

They are written the way the customer of a print shop with a walk-in counter would send them, which is the business where the line shows up most cleanly: it has prices that can be recited, jobs that only exist inside the system and customers with an account. Adapt the industry; do not change the structure, which is what does the measuring.

1. The compound one: two questions in a single message.

Hi, do you do business cards on 300 coated stock and when could you have them ready?

Nobody writes one thing at a time. A flow with keyword triggers picks up one of the two, almost always the last one, and the other is lost. The typical failure is that it sends you back a numbered menu, ignoring that you already asked something specific. If it answers both things, it passes.

2. The one written the way your customer writes.

hi i need 200 papers to hand out in the street, something simple, how much is it

No accents, no punctuation and without naming the product the way you name it: the customer did not write "flyer" or "leaflet" or "quote". This message measures whether the bot depends on the customer using your vocabulary. The typical failure is the fallback: I did not understand, pick an option from the menu. Write down the regional wording used in your area as well, because a script written in another country does not have it loaded.

3. The piece of data that only exists inside your system.

Hi, can you check whether I still have a balance on the June invoice?

Variant for any industry with jobs or orders in progress: is job 4120 ready for pickup? Both ask something that is on no website and in no script: it lives in your management system, it changes on its own and it is different for every customer. This is the question that separates a bot that recites from one that queries. There are three possible answers and all three mean different things. That it gives a verifiable and correct piece of data: it is connected. That it says it cannot look that up and offers to hand you over: it is honest and it is not connected. That it answers something that sounds right —nothing shows as outstanding, yes, you can come pick it up— without having looked at anything: that is the worst of the three, because the customer walks away with an answer in writing that you will later have to take back. To spot it you have to open the system and verify the fact yourself.

4. The change of mind halfway through.

Start an order, and when the bot asks you for the second or third piece of data —quantity, size, date—, cut across the lane:

hold on, before we go on: do you deliver to [area] or do I have to pick up?

And then: actually, better make it 500 instead of 1,000. This one measures whether the flow can take the customer stepping out and coming back. The two typical failures are that it repeats the previous question as if you had said nothing, or that it loses everything already entered and starts from scratch. It is the message that makes a real customer angriest, because they have already invested four replies.

5. The complaint that asks for a person.

I already complained about this last week and nobody answered me. I need to speak to someone.

Here you start the clock. You are not measuring what the bot answers, you are measuring two things: how long it takes for a human to appear, and whether, once they appear, they know what you were talking about. The classic failure is not that it fails to hand off: it is that it says an agent will be with you shortly and nobody appears, or that someone appears asking hi, how can I help you?, forcing you to repeat everything.

How to send them so the test does not lie

Four conditions, and all four change the result.

  • From a number the provider does not have on file. If you test from your own phone you are already a contact with tags and the flow behaves differently.
  • One at a time, waiting for each reply. If you send them back to back, WhatsApp groups them and many platforms process only the last one. The result tells you nothing.
  • On a Saturday at nine at night, not a Tuesday at eleven. A good part of what you are measuring is what happens when there is nobody on the other side.
  • Write down the literal reply, not your impression. It seemed to me like it understood is not data. Copy and paste the five replies into a note.

The score is not out of five: it is out of your last month

A bot that fails three out of five is not bad. It is bad for you if those three are the bulk of what people ask you. So before scoring, open your WhatsApp, look at the last forty incoming conversations and classify each one into the scenario it resembles. Ten minutes.

The print shop from the examples counted its last week like this:

ResemblesWhat goes in thereConversations
Messages 1 and 2Price of a standard job, opening hours, whether they do such and such11
Message 3is it ready yet?, do I still have a balance?, has it been shipped?14
Message 4Orders that start and change quantity, paper or date halfway through8
Message 5Complaints: it printed misaligned, it arrived late, units were missing5
NoneA supplier, a CV, a wrong number2

With that table beside you, the score stops being an impression. Then, a rough criterion. It is not a measurement: it is a threshold for deciding with what you already have at hand.

  • If what the bot failed represents less than 20% of your inquiries, the flow platform is enough for you. Pay the subscription and do not spend on custom development.
  • Between 20% and 50%, it works as a first layer as long as you have someone who can take over the conversation in minutes, not hours.
  • Above 50%, the cheap bot will create more work than it takes away. Every failure comes back to a person, but now with the customer already irritated from having talked to a machine that did not understand them.

The bot the print shop tested failed 3 and 5: 14 plus 5, 19 out of 40, 47.5%. Middle band, cheap platform viable as a first layer. But here is what the percentage alone does not say, and it is the part almost nobody looks at. If those 19 conversations had fallen on the other side —the 11 from messages 1 and 2 plus the 8 from message 4 add up to exactly the same—, the number would be identical and the decision would be the opposite: one afternoon rewriting the flow, adding synonyms and fixing the step where the data gets lost, and the cheap subscription serves you well for two years.

Because they fell into 3 and 5, no amount of configuration moves them: 3 requires reading a system and 5 requires a shared inbox with routing. The percentage tells you how much it hurts. Which ones it failed tells you whether it is fixed with an afternoon of work or by changing category, and those are two different calculations that can come out the same and mean the opposite.

That second calculation is the border. How much it costs to cross it is broken down in why an AI chatbot costs USD 500 in one place and USD 5,000 in another; here the question ends one step earlier, at which side you are standing on today.

What the cheap platform does better than an AI model

The five-message score always pushes upward —the more failures, the more expensive the bot you seem to need— and for a certain kind of business that push is wrong. A flow builder never improvises: there is no way for it to make up a price, promise a discount you did not authorize or answer something strange at three in the morning. For a business with five repeated questions, no stock to check and no calendar to consult, that is not a limitation: it is exactly what you want. Think of a car wash that serves customers on a first-come basis: there is no appointment to book, no order number to look up, no running account, and what comes in over WhatsApp is always the same —opening hours, address, how much a full wash costs, whether they clean upholstery, whether they take cards—. There is no possible message 3 there, because there is no system on the other side to query. A well-built menu answers almost everything, and a language model only adds the possibility that one day it answers something nobody wrote.

The flow platforms have also added artificial intelligence blocks on top of the decision tree. That lifts messages 1, 2 and 4 quite a bit, which were their historic weak point. It does not move 3 or 5.

Where the five-message test stops being enough

It measures the conversation, not the operation. And there are four things a bot can pass in five messages and still break in production —or end up costing you far more than it looked—.

The first: content volume. A bot passes all five with a demo catalog of ten products and falls apart with your four hundred, where there are three named almost the same and two you no longer sell.

The second: the day prices change. The test tells you how it answers today, not who updates the data next month or how long it takes. A bot that used to answer well and has been quoting the old price list for three weeks is worse than having no bot.

The third: Monday at ten, with forty conversations open at the same time. Message 5 passed on a Saturday with a single active chat does not prove routing exists; it proves somebody was looking at their phone.

The fourth is not in the conversation because it is not about the conversation: it is about the invoice. The five messages measure one chat, and one chat costs practically the same on any platform. What changes is the plan. These tools bill by contacts, by conversations or by active users depending on which one it is, so the price does not scale with how well the bot answers: it scales with how many people it answers. The same bot that passed all five with three hundred contacts in the database jumps one or two plan tiers when that database is three thousand, and on top of that you add what Meta charges you separately for the messages you initiate (how that is billed, here). Before deciding on price, find the platform's pricing grid and locate your contact count for next year, not today's: the cheap option stops being cheap exactly where it starts working, and none of the five messages sees it coming.

Those four you do not solve with five messages. The first two are solved by looking at where each piece of data comes from and who maintains it, the third by looking at a real Monday, and the fourth by reading the pricing grid with your contact projection next to it.

Run the test this afternoon and save the five literal replies, without summarizing them: the value is in how it answered, not in whether you decided it was good enough. Then read them with your last month's score beside them.

If they failed 1, 2 and 4 and that is the bulk of what people ask you, the honest answer is that you should buy the flow platform: cheap, this week, without us. If they failed 3 or 5 and that is where your volume is, the problem is not fixed by choosing a better platform but by connecting a system, and the order matters. Before any of that, the calculation of whether your volume justifies a bot at all is in how many inquiries justify a chatbot.

Paste the five replies into an email to info@striqtech.com and we will tell you which of the two cases is yours.

Frequently asked questions

Is a bot platform like ManyChat or Landbot enough for a business that handles customers over WhatsApp?

It is, and for many businesses it is the right answer. They are flow builders: decision trees with buttons and keyword triggers. They answer opening hours, address, payment methods and delivery area better than a language model, because they never improvise. Where they fall down is when the customer asks about something that only exists inside your management system —the status of their job, the balance on their account— and when the whole conversation has to be handed over to a person.

Do the AI blocks these platforms added solve the five messages?

They lift messages 1, 2 and 4 quite a bit: they understand two intents at once, they tolerate the customer writing differently from the script and they cope better with a change of mind halfway through an order. They do not move 3 or 5. An AI block on top of a flow still cannot read your management system or write an order into it, and it still does not transfer the history to an inbox with people in it.

Can I run this test on a provider's bot before hiring them?

Yes, and it is the best use of the five messages. But do not run it on the demo: ask them for the number of one of their bots in production, from a real client and ideally in an industry similar to yours. A demo is built with ten products and three made-up customers, and it passes message 3 by reading a spreadsheet they put together for the demo. A provider who gives you no number in production has already answered something.

What does it mean if the bot makes up a fact in message 3?

It is the worst of the possible answers, worse than not knowing. A bot that says it cannot look that up leaves you handling things by hand, which is what you already do. A bot that tells a customer they have nothing outstanding, or that their job is ready, without having looked at the system, creates something worse: a complaint with a screenshot, where your own message is the evidence against you. Before scoring, go into the system and verify whether the fact it gave was true.

How long does the five-message test really take?

Writing them is five minutes. The full result takes as long as message 5 takes to get a human reply, and that wait is part of the measurement. Do it on a Saturday night or a Sunday, not a Tuesday at midday: half of what you are measuring is what happens outside business hours, which is when a large share of the inquiries comes in.

Did this content help?

Implement this in your business in 72 hours

Let's talk for 15 minutes. No cost, no commitment. I'll audit one process and show you the projected ROI.

Let's talk on WhatsApp