Your AI passed the test. That doesn't mean it can do the job.
A test score tells you one narrow thing about an AI tool. Here are the five questions to ask before you sign anything.
You were shown a demo. Or a leaderboard position. Or an accuracy figure printed on a one-pager from a sales rep. It looked convincing, so you gave the tool a try.
A few weeks in, some of the answers are coming back slowly. Some of them are wrong — and not on obscure edge cases, but on the kinds of questions your team actually asks. And you’re wondering what happened between that impressive number and the thing sitting on your desk right now.
What happened is that the test and your business were measuring completely different things.
What a benchmark actually measures
A benchmark is a standardised test. Someone gives the AI a set of questions — drawn from chemistry papers, legal documents, coding problems, whatever the test covers — and grades the answers against known correct ones.
That works well as a rough guide to whether the tool can reason at all. Think of it like a driving theory test: it tells you the candidate knows the road rules. It does not tell you whether they can parallel-park a van in the rain outside your shop on a Saturday morning.
That everything else is your questions, your data, your customers, and the way people actually phrase things when they’re frustrated or in a hurry.
The trade-off nobody advertises
Any AI system you buy or build is balancing three things at once:
- Getting the answer right. Is it accurate enough that you’d trust it?
- Getting the answer fast. How long does a customer wait, or how long does your team sit waiting for a result?
- What it costs. Per month, per request, per busy season.
You can usually get two of those. The third one suffers.
A very accurate tool running fast enough to handle your whole team at once? That tends to be expensive. A cheap tool that replies quickly? The answers may not be reliable enough to act on. A cheap, reliable one? Responses may be slow enough that customers give up.
Think of it like staffing. You can hire someone brilliant who works round the clock — but you’ll pay for it. You can hire cheap and fast — but the work may need checking. You can hire cheap and careful — but there won’t be enough hours in the day. The benchmark score on the job advert tells you almost nothing about where this particular candidate lands on that triangle.
A score on a leaderboard does not tell you which corner of the triangle you’re actually buying.
Two very different questions
“Is this AI clever?” and “will this AI hold up on a busy Tuesday?” are not the same question.
The first question is what benchmarks answer. The second is what your business needs to know.
Passing the first tells you almost nothing about the second. A tool can ace every accuracy test and still fall over when your whole team is using it at 9am, or when the question is phrased slightly differently from how the test was written.
System performance — how fast it responds under real traffic, how many people it can handle at once, what happens when demand spikes — gets tested separately, and it is tested against your workload, not a generic one. A customer service chatbot and a document-reading tool handle completely different patterns of use. Benchmarked with the wrong pattern, you’ll see numbers that look fine and a live tool that doesn’t.
Why checking the final answer isn’t enough
If the AI is doing a single, simple task — answer this one question — checking whether the answer is right is straightforward.
Most real business tools don’t work that way. They’re a chain of steps. The AI reads what you asked, decides what that means, finds the right piece of information, pulls the right document, and writes a reply. Any one of those steps can be where things went wrong — and if you only check the final reply, you won’t know which one.
The IBM Technology video describes it as a pyramid of checks: you need to know whether the system can handle the load, whether it’s formatting responses correctly, whether it’s staying factually grounded, and whether it’s giving domain-specific answers that actually hold up — not just whether the last sentence looked plausible.
If you skip to checking only the top — only the final output — everything underneath can collapse without you noticing until a customer tells you.
Who checks the checker
Some AI suppliers use automated grading: another AI reads the first AI’s responses and scores them. That is genuinely useful for volume — you can evaluate thousands of answers that a human team couldn’t get through in a week.
But it has a limit. The grading AI is working from a rubric — a set of rules about what a good answer looks like. Someone who actually knows the job has to have written that rubric, and someone who actually knows the job has to periodically confirm that the scoring is still right.
In practice, that means people who know your work need to spot-check what the automated system is calling correct. The IBM framing calls this “humans in the loop” — domain experts who label what good and bad look like, so the automated system can scale that judgment rather than replace it.
If a supplier can’t tell you how they do this, that is a question worth pressing on.
What benchmarks are actually good for
None of this means tests are worthless. They do something useful: they eliminate the bottom of the market quickly. A tool that scores poorly on standard benchmarks is almost certainly not worth your time. One that scores well has cleared a basic bar — it can reason, it isn’t making things up constantly, it has some domain knowledge.
That is a shortlist, not a finish line. The video ends on exactly that point, and it is worth having in the presenter’s own words (where he says model, he means the AI software itself):
“a leaderboard score for a model’s performance in a certain area is just a starting point. It’s not the finish line, because real benchmarking happens when you test with your data, with your realistic traffic patterns, and your own definition of success.”
Your data. Your busiest hour. Your own definition of a good result.
You cannot test everything before you start. At some point you have to run the tool on real work and see what happens. The useful discipline is knowing what you’re testing and what you’re not — so when something goes wrong, you know where to look.
The questions to ask before you sign
When an AI supplier shows you a score, here are the questions that will tell you more than the number does:
- What did you test it on that actually looks like my work? Not chemistry papers. Not coding benchmarks. Your kind of questions, your kind of data.
- How fast is it when everyone’s using it at the same time? Not in a quiet demo environment. Under real load.
- What does a busy month cost? Not the cheapest tier — the tier you’d actually need.
- Who checks that it’s getting things right? Not just another AI. A person who knows the job.
- What happens when it gets one wrong? How do you find out? What’s the process?
A supplier who can answer those plainly is worth talking to further. One who redirects you back to the leaderboard number probably isn’t.
How this gets handled
We should declare an interest: this is what Operio does, so read this part as the interested party talking. What we’d say is that the five questions above are the right ones, and that a supplier should be able to answer them about your work rather than about a test. We build one workflow at a time around a business’s own documents, prices and approval rules, then run it and keep tuning it. That means it was tested on that business’s real enquiries rather than someone else’s; a person approves the work before it reaches a customer; and when something comes out wrong we look at where in the chain it broke instead of just rewriting the last reply.
On cost, in plain dollars: a front desk handling enquiries, bookings and follow-ups — one shift a day, in one language — runs about $700–$1,400 a month staffed by a person, and plans that include a specialist doing that work start at $400 a month. Whether that trade is worth making is a judgement about your own week, and it was never something a test score could tell you.
If you’d like someone to work through those questions with you — against your actual work, not a generic demo — book a free consultation with Operio. We’ll tell you what we think is worth the money and what isn’t.