How we actually test AI software, step by step
This page opens up the process: how a tool is selected, how we set up an account, what devices and scenarios we test on, how we check AI output for accuracy and hallucinations, and how a score actually gets calculated. Editorial Standards covers governance, independence, and conflicts of interest. Review Methodology covers the exact scoring math. This page is the middle step: the hands-on mechanics that produce the numbers those other pages explain. Every scenario below is built around a realistic small-business task, not a synthetic benchmark chosen to make a tool look good.
Minimum active testing window
1–2 weeks per tool
Re-test cadence
Every 90 days
Account type used
Real, paid, small-business tier
Testing focus
Real business tasks, not demo benchmarks
Selection, accounts, environments, and devices
Before any hands-on testing starts, several things get locked in: which tool, which account tier, which environment, and which devices. Skipping any of these makes every later result harder to trust.
Product selection
A tool is shortlisted against real reader demand and relevance to small business workflows. The full criteria live on Editorial Standards.
Account creation
A real, paid account is created at the pricing tier a small business would actually choose, never a vendor-provided demo with unlocked enterprise features.
Testing environment
Testing happens under normal business conditions: typical home or office internet, no synthetic lab network optimized to flatter response times.
Testing devices
Desktop browser, mobile browser, and the dedicated app where one exists, since a business owner may set up on a laptop and manage day-to-day from a phone.
Browser and OS coverage
Core testing runs on Chrome and Safari at minimum, with Windows and macOS on desktop and iOS and Android on mobile, since a tool that works cleanly in one browser can still break in another.
Regional considerations
Where a tool’s pricing, language support, or feature availability changes by region, testing notes which region was used and flags anything a reader in a different one should expect to differ.
What earns a tool a spot on the shortlist. Reader demand comes first: search patterns and direct requests tell us what businesses are actually trying to solve. From there we weigh business relevance (does this fit a small business budget and skill level, not just an enterprise IT department) and market maturity (is this a stable product with a track record, or a beta feature likely to change before the review is even useful).
What keeps a tool off the shortlist. A tool gets excluded if it requires an enterprise sales call just to see pricing, if it’s clearly built for a different market segment than small business, or if it’s changed so recently that a review would go stale within weeks. None of this is about a vendor’s size or reputation. It’s about whether we can test it under the same standard as everything else on the site.
Prompt, workflow, and automation testing
Once the account is live, testing moves through three connected stages. Each one builds on the last rather than testing the tool in isolation.
Prompt testing
Individual prompts are tested in isolation first, from simple single-step requests to complex and ambiguous ones, before anything is chained together.
Workflow testing
Prompts are chained into a full realistic business scenario, such as a caller booking, rescheduling, then asking a follow-up question, to see whether context carries through correctly.
Automation testing
Any triggers, integrations, or scheduled actions the tool sets up on its own are checked for reliability over repeated runs, not just a single successful demo pass.
Inside prompt testing. We run simple prompts (a single direct question), complex prompts (multiple requirements in one message), edge cases (unusual or malformed input), and deliberately ambiguous instructions. We also send follow-up prompts to check contextual memory (does the tool remember what was said two messages ago), and repeat the same request several times to check whether the answer stays consistent or drifts.
Inside workflow testing. Scenarios are pulled from real business use, not isolated commands: a customer support exchange that starts with a complaint and ends with a resolution, an appointment that gets booked and then moved, a document generated from a template, a piece of marketing content drafted against a brand brief, a CRM record updated after a call, an automation sequence with several steps, or a question answered by pulling from internal knowledge the business already has on file. We test the complete scenario, not just the first step.
Inside automation testing. We check that triggers actually fire when they should, that integrations pass data correctly between systems, and that scheduled actions run on time. Reliability gets checked over repeated execution, not a single pass, and we specifically look at failure handling: what happens when a step fails, whether the tool recovers on its own, and whether repeated runs stay consistent with each other.
Accuracy, hallucination checks, response quality, and speed
These scorecards describe the rubric bands we score against, not a specific tool’s actual result, since every AI platform gets its own individual scorecard on its review page. Some of what’s below is checked objectively (right or wrong, present or absent); some requires a reviewer’s judgment call. We separate the two below rather than blending them into one unexplained number.
AI accuracy
RubricMeasured against real business content we already know the correct answer for (a menu, a service list, a pricing sheet), not a synthetic benchmark set. This checks factual accuracy, completeness of the answer, and whether the response stays relevant to what was actually asked.
Hallucination rate
RubricA hallucination is a confidently stated claim that’s verifiably wrong when checked against the input it was given or a primary source, not a vendor’s own marketing claims.
Response quality
RubricJudged on whether a real customer receiving the response would find it clear, professionally toned, well structured, and usable without editing. Grammar and formatting are checked directly; tone and overall usefulness are a reviewer’s judgment call.
Speed
RubricTimed under normal business-hours conditions, not off-peak, since that’s when a real customer is actually waiting on the other end. We also note consistency across repeated requests and any interruptions, retries, or downtime observed during the testing window.
Objective checks versus subjective judgment. Some criteria have a clear right answer: did the tool state a fact correctly, did it follow the instruction it was given, did the formatting come out clean, did it handle a second language correctly where that was tested. Others don’t: whether a response reads as clear and professional, whether the reasoning behind an answer actually holds up, whether the tone fits the business it’s written for. We score both, but we don’t pretend the second group is as exact as the first. Reviewers apply the same rubric bands shown above rather than an unstated personal impression.
How we test for hallucinations. A hallucination, in our testing, is a confidently stated factual claim that’s wrong. That’s a specific definition, and it matters: an AI tool saying “I’m not sure” about something it doesn’t know is not a hallucination, it’s appropriate uncertainty. The problem is a wrong answer delivered with the same confidence as a right one. We identify hallucinations by running the tool against content we already know the correct answer for, then checking the output line by line. When we find one, we note its severity (a wrong store hour is not the same as a fabricated legal claim), whether it happened once or repeated across several tries, and how confidently it was stated. Factual claims are checked against the real business content we gave the tool or against a primary source, never against a vendor’s own marketing.
Ease of use, feature validation, pricing, and support
Four practical categories, each with its own check method and pass bar.
| Category | What we check | Method | Pass bar |
|---|---|---|---|
| Ease of use | Core setup task completion | A non-expert completes onboarding unaided, timed | Completed without external help |
| Feature validation | Every advertised feature | Used directly, not read from a spec sheet | Works as described at the tested tier |
| Pricing verification | Every listed price | Checked on the vendor’s live pricing page | Matches exactly, or discrepancy is disclosed |
| Customer support | Response time and answer quality | A real support ticket is filed during testing | Resolved accurately within a reasonable window |
Usability, in detail. Beyond the timed onboarding task, we check navigation (can a feature be found without a search), the learning curve for anything past the basics, whether useful features are easy to discover on their own, and the state of the documentation. Setup complexity and the first-time experience get weighed together: a tool that takes 10 minutes to configure but works correctly afterward scores differently from one that takes 10 minutes and still needs support.
Feature validation, in detail. Every major advertised capability is used directly whenever our account tier allows it, not taken on faith from a spec sheet or a vendor’s marketing page. Where a feature sits behind an enterprise tier we didn’t purchase, the review says so explicitly rather than scoring it as untested or assuming it works.
Pricing verification, in detail. Every listed price is checked against the vendor’s own live pricing page, not a press release or a stale affiliate feed. We note free plans and their real limits, trial length and what happens after it ends, how enterprise or “contact sales” pricing is disclosed, and any hidden costs: usage caps, per-seat charges, or required add-ons that aren’t obvious from the main pricing table. Pricing is a snapshot from the date of testing, and it’s reviewed again on our 90-day cycle, since vendors change prices without much notice.
Customer support, in detail. We check the documentation and knowledge base first, since that’s what most users try before contacting anyone. Then we file a real support ticket during testing and evaluate the channel (live chat, email, or ticket queue), the actual response time, whether the answer resolved the issue correctly, and, where a first response didn’t fully solve it, how the escalation was handled.
Security review and privacy review
These are editorial reviews of what a vendor publicly documents, not a security audit we performed ourselves.
Security review
We check what a vendor’s own documentation discloses: encryption in transit and at rest, authentication options like two-factor login, account-level access controls, any published certifications, compliance documentation for regulated industries, and transparency reports where a vendor publishes one. We are not a penetration-testing firm, and we say so rather than implying a security audit we didn’t perform.
Privacy review
We read the actual privacy policy rather than summarizing a marketing page: what data is collected, how long it’s retained, whether customer data is used to train the vendor’s models, and how it’s disclosed if so. We check whether a business owner can export or delete their data on their own, what consent is required from end customers, and what account-level controls exist. Compliance-relevant context for regulated industries is flagged, not assumed.
How performance is weighted into a final score
Every category above feeds into one of five weighted groups. The exact per-criterion math lives on our Review Methodology page; this is the weighting at a glance.
Comparison methodology. When several AI tools are placed in a single comparison table, each one is scored independently against this same weighting before the table is assembled. The table reflects scores that already existed; it isn’t built first and scored backward to fit a preferred outcome.
Normalization and consistency. Raw observations (a response time in seconds, a count of hallucinations across a set of test prompts) get converted into the same 1-to-5 or percentage scale used across every review, so a “4” in ease of use means the same thing on two different reviews. Reviewers are calibrated against the same rubric bands shown in the scorecards above, and a second editor spot-checks scores during the fact-check pass to catch a rating that’s drifted from the standard. This is quality assurance for consistency, not a second full retest.
How often we retest, and what actually changes a score
Every review is fully retested at least every 90 days on a routine schedule. Below is an illustrative, generic example of the kind of change that triggers an update sooner than that, not a real tool’s actual history.
Before
After vendor change
Illustrative example only. Pricing changes and feature removals are the two most common triggers for an off-cycle re-test.
What triggers a re-test outside the routine cycle. A pricing update, a new feature launch, a vendor acquisition or merger, an update to the underlying AI model that changes output quality, a disclosed security incident, a major product redesign, a pattern of reader-reported errors, or a vendor reaching out with a correction all trigger an off-cycle look, on top of the standard 90-day schedule that applies regardless.
Testing limitations, transparency, and where we’re still improving
Testing reflects a snapshot in time, a specific pricing tier, and the specific real-world scenarios we chose to run. It cannot capture every configuration, every integration a reader might use, or guarantee an identical experience on every reader’s own setup.
Specific limits worth naming directly: our testing window is 1 to 2 weeks per tool, not months of use. We test the small-business pricing tier, not every enterprise configuration. Regional availability of features and pricing can differ from what we tested. We don’t test every third-party integration a tool supports, only the ones a small business would realistically connect. And AI models themselves change quickly, sometimes faster than a 90-day cycle, so a specific output quality result can shift between our visits. Where a limitation is specific to how we tested a particular tool, it’s stated in that review directly.
Every review states the pricing tier tested, the date of the most recent verification, and at least one genuine limitation of the tool. If a reader’s own experience differs from ours, we want to hear about it. See Contact.
As AI software itself evolves, with new AI agent capabilities, new generative AI models, and new automation features, our testing process is expected to evolve with it. Material changes to this methodology will be reflected on this page directly, not published elsewhere and left unlinked.
Reader feedback plays a direct role here: recurring questions or complaints about a testing gap tell us where the methodology itself needs work, not just where a single review needs an update. We also watch for industry-wide shifts, like a new category of AI tool that doesn’t fit our current evaluation matrix cleanly, and adjust the framework rather than forcing a square peg into an old rubric.
Quick answers
The questions we hear most about our testing process specifically.
We test the tier a small business would realistically choose. Higher enterprise tiers are noted in the review but are not the basis for scoring unless that tier is the only one with a usable feature set.
We run the tool against real business content we already know the correct answer for, then check the output against that known-correct baseline rather than relying on a vendor’s self-reported benchmark scores.
A confidently stated factual claim in the AI’s output that is verifiably wrong when checked against the real business content it was given or against a primary source.
At least once every 90 days, and immediately if we become aware of a pricing, feature, or policy change significant enough to affect the published score.
Testing reflects a snapshot in time, a specific pricing tier, and specific real-world scenarios we chose. It cannot capture every configuration, every industry use case, or guarantee identical results for every reader’s setup.
A tool is shortlisted based on real reader demand and relevance to small business workflows, screened for market maturity, then excluded if it requires an enterprise sales call just to see pricing or is clearly built for a different market segment.
At least 1 to 2 weeks of active hands-on testing per tool, covering prompt testing, full workflow scenarios, and automation reliability checks before a score is finalized.
Every listed price is checked directly against the vendor’s own live pricing page at the time of testing, including free plan limits, trial terms, and any hidden costs like usage caps or required add-ons. Pricing is reviewed again on our 90-day cycle.
No. Vendors get no advance access to our testing process or draft scores, and an affiliate relationship with a vendor has no bearing on how thoroughly a tool is tested or what score it receives.
A non-expert completes core onboarding unaided while timed, and we separately check navigation, discoverability of features, documentation quality, and the overall first-time user experience.
Want the short version, every Friday?
One practical AI idea for your business, tested by us first. No hype, no spam.