This is not usually deliberate deception. It is that demos are built to show capability, and what you need to know is reliability, which is a different thing and harder to show.
Here is a checklist that surfaces the difference.
Test it on something you already know
The single most useful evaluation technique, and the most skipped.
Do not ask the tool a question you want answered. Ask it a question you already know the answer to, ideally something specific to your business or your jurisdiction where a wrong answer is obviously wrong to you.
You learn two things. Whether it is accurate. And more importantly, what it does when it is wrong. Good systems hedge or say they are unsure. Poor ones state incorrect things with complete confidence, which is far more dangerous, because in production you will not know the answer.
Run five or six of these. It takes twenty minutes and tells you more than any demo.
Ask what it does when it does not know
Follow up directly. Ask the vendor how the system behaves outside its competence.
The answer you want describes a mechanism: it declines, it flags uncertainty, it escalates to a human, it cites sources you can check. The answer that should worry you is a reassurance that it is very accurate. That is not an answer to the question.
Separate what is live from what is coming
Almost every AI product page mixes shipped features with roadmap items, usually with a small label that is easy to miss in a grid.
Ask for a written list of what is available today. Then check the changelog and the roadmap, which most vendors publish and which are more honest than the marketing pages because they are written for existing customers.
Pay particular attention when something on the pricing page is quantified but not built. A plan that allocates a monthly quota of a feature still marked as planned is a signal about how the whole page was written.
Look hard at the integrations
For business AI, integrations are where the value is. An assistant that cannot see your data is a general chatbot with a different logo.
Ask three questions about each integration you care about:
Is it live today, or planned? Does it read only, or can it write? What is the sync frequency?
The read versus write distinction matters for risk. An integration that reads your bank transactions is very different from one that can initiate payments. Most of the value sits in reading. Most of the risk sits in writing.
Read the data processing agreement
Under GDPR, any vendor processing personal data for you must have one. If they cannot produce it, that is your answer.
Go to the sub-processor annex at the back. It has to name every third party touching your data, what they do and where they are. It is the most honest page a vendor publishes, and it tells you the real architecture in a way the marketing site never will.
Check that the model providers are named, that non-EU entities have a transfer safeguard listed, and that there is a notification clause for changes.
Then ask one question that catches a lot of vendors out: is the commitment not to train on your data written in the contract, or only on the website? A great many homepages carry that promise and a great many contracts do not.
Check the claims that can be checked
If a product advertises a review score, click the link. It should go to the vendor's profile on the review platform, not the platform's homepage. A score with no verifiable profile behind it is not evidence, and it tells you something about how the rest of the page was written.
read more..