Inside the Lab

How we actually test AI software, step by step

This page opens up the process: how a tool is selected, how we set up an account, what devices and scenarios we test on, how we check AI output for accuracy and hallucinations, and how a score actually gets calculated. Editorial Standards covers governance, independence, and conflicts of interest. Review Methodology covers the exact scoring math. This page is the middle step: the hands-on mechanics that produce the numbers those other pages explain. Every scenario below is built around a realistic small-business task, not a synthetic benchmark chosen to make a tool look good.

Minimum active testing window

1–2 weeks per tool

Re-test cadence

Every 90 days

Account type used

Real, paid, small-business tier

Testing focus

Real business tasks, not demo benchmarks

Lab Setup

Selection, accounts, environments, and devices

Before any hands-on testing starts, several things get locked in: which tool, which account tier, which environment, and which devices. Skipping any of these makes every later result harder to trust.

Step 1

Product selection

A tool is shortlisted against real reader demand and relevance to small business workflows. The full criteria live on Editorial Standards.

Step 2

Account creation

A real, paid account is created at the pricing tier a small business would actually choose, never a vendor-provided demo with unlocked enterprise features.

Step 3

Testing environment

Testing happens under normal business conditions: typical home or office internet, no synthetic lab network optimized to flatter response times.

Step 4

Testing devices

Desktop browser, mobile browser, and the dedicated app where one exists, since a business owner may set up on a laptop and manage day-to-day from a phone.

Step 5

Browser and OS coverage

Core testing runs on Chrome and Safari at minimum, with Windows and macOS on desktop and iOS and Android on mobile, since a tool that works cleanly in one browser can still break in another.

Step 6

Regional considerations

Where a tool’s pricing, language support, or feature availability changes by region, testing notes which region was used and flags anything a reader in a different one should expect to differ.

What earns a tool a spot on the shortlist. Reader demand comes first: search patterns and direct requests tell us what businesses are actually trying to solve. From there we weigh business relevance (does this fit a small business budget and skill level, not just an enterprise IT department) and market maturity (is this a stable product with a track record, or a beta feature likely to change before the review is even useful).

What keeps a tool off the shortlist. A tool gets excluded if it requires an enterprise sales call just to see pricing, if it’s clearly built for a different market segment than small business, or if it’s changed so recently that a review would go stale within weeks. None of this is about a vendor’s size or reputation. It’s about whether we can test it under the same standard as everything else on the site.


Testing Pipeline

Prompt, workflow, and automation testing

Once the account is live, testing moves through three connected stages. Each one builds on the last rather than testing the tool in isolation.

Prompt testing

Individual prompts are tested in isolation first, from simple single-step requests to complex and ambiguous ones, before anything is chained together.

Workflow testing

Prompts are chained into a full realistic business scenario, such as a caller booking, rescheduling, then asking a follow-up question, to see whether context carries through correctly.

Automation testing

Any triggers, integrations, or scheduled actions the tool sets up on its own are checked for reliability over repeated runs, not just a single successful demo pass.

Inside prompt testing. We run simple prompts (a single direct question), complex prompts (multiple requirements in one message), edge cases (unusual or malformed input), and deliberately ambiguous instructions. We also send follow-up prompts to check contextual memory (does the tool remember what was said two messages ago), and repeat the same request several times to check whether the answer stays consistent or drifts.

Inside workflow testing. Scenarios are pulled from real business use, not isolated commands: a customer support exchange that starts with a complaint and ends with a resolution, an appointment that gets booked and then moved, a document generated from a template, a piece of marketing content drafted against a brand brief, a CRM record updated after a call, an automation sequence with several steps, or a question answered by pulling from internal knowledge the business already has on file. We test the complete scenario, not just the first step.

Inside automation testing. We check that triggers actually fire when they should, that integrations pass data correctly between systems, and that scheduled actions run on time. Reliability gets checked over repeated execution, not a single pass, and we specifically look at failure handling: what happens when a step fails, whether the tool recovers on its own, and whether repeated runs stay consistent with each other.


Benchmark Scorecards

Accuracy, hallucination checks, response quality, and speed

These scorecards describe the rubric bands we score against, not a specific tool’s actual result, since every AI platform gets its own individual scorecard on its review page. Some of what’s below is checked objectively (right or wrong, present or absent); some requires a reviewer’s judgment call. We separate the two below rather than blending them into one unexplained number.

AI accuracy

Rubric
Excellent
Acceptable
Needs work

Measured against real business content we already know the correct answer for (a menu, a service list, a pricing sheet), not a synthetic benchmark set. This checks factual accuracy, completeness of the answer, and whether the response stays relevant to what was actually asked.

Hallucination rate

Rubric
None observed
Occasional
Frequent

A hallucination is a confidently stated claim that’s verifiably wrong when checked against the input it was given or a primary source, not a vendor’s own marketing claims.

Response quality

Rubric
On-brand tone
Generic tone
Off-tone

Judged on whether a real customer receiving the response would find it clear, professionally toned, well structured, and usable without editing. Grammar and formatting are checked directly; tone and overall usefulness are a reviewer’s judgment call.

Speed

Rubric
Near-instant
Noticeable delay
Disruptive lag

Timed under normal business-hours conditions, not off-peak, since that’s when a real customer is actually waiting on the other end. We also note consistency across repeated requests and any interruptions, retries, or downtime observed during the testing window.

Objective checks versus subjective judgment. Some criteria have a clear right answer: did the tool state a fact correctly, did it follow the instruction it was given, did the formatting come out clean, did it handle a second language correctly where that was tested. Others don’t: whether a response reads as clear and professional, whether the reasoning behind an answer actually holds up, whether the tone fits the business it’s written for. We score both, but we don’t pretend the second group is as exact as the first. Reviewers apply the same rubric bands shown above rather than an unstated personal impression.

How we test for hallucinations. A hallucination, in our testing, is a confidently stated factual claim that’s wrong. That’s a specific definition, and it matters: an AI tool saying “I’m not sure” about something it doesn’t know is not a hallucination, it’s appropriate uncertainty. The problem is a wrong answer delivered with the same confidence as a right one. We identify hallucinations by running the tool against content we already know the correct answer for, then checking the output line by line. When we find one, we note its severity (a wrong store hour is not the same as a fabricated legal claim), whether it happened once or repeated across several tries, and how confidently it was stated. Factual claims are checked against the real business content we gave the tool or against a primary source, never against a vendor’s own marketing.


Evaluation Matrix

Ease of use, feature validation, pricing, and support

Four practical categories, each with its own check method and pass bar.

Evaluation matrix for ease of use, feature validation, pricing verification, and customer support testing
CategoryWhat we checkMethodPass bar
Ease of useCore setup task completionA non-expert completes onboarding unaided, timedCompleted without external help
Feature validationEvery advertised featureUsed directly, not read from a spec sheetWorks as described at the tested tier
Pricing verificationEvery listed priceChecked on the vendor’s live pricing pageMatches exactly, or discrepancy is disclosed
Customer supportResponse time and answer qualityA real support ticket is filed during testingResolved accurately within a reasonable window

Usability, in detail. Beyond the timed onboarding task, we check navigation (can a feature be found without a search), the learning curve for anything past the basics, whether useful features are easy to discover on their own, and the state of the documentation. Setup complexity and the first-time experience get weighed together: a tool that takes 10 minutes to configure but works correctly afterward scores differently from one that takes 10 minutes and still needs support.

Feature validation, in detail. Every major advertised capability is used directly whenever our account tier allows it, not taken on faith from a spec sheet or a vendor’s marketing page. Where a feature sits behind an enterprise tier we didn’t purchase, the review says so explicitly rather than scoring it as untested or assuming it works.

Pricing verification, in detail. Every listed price is checked against the vendor’s own live pricing page, not a press release or a stale affiliate feed. We note free plans and their real limits, trial length and what happens after it ends, how enterprise or “contact sales” pricing is disclosed, and any hidden costs: usage caps, per-seat charges, or required add-ons that aren’t obvious from the main pricing table. Pricing is a snapshot from the date of testing, and it’s reviewed again on our 90-day cycle, since vendors change prices without much notice.

Customer support, in detail. We check the documentation and knowledge base first, since that’s what most users try before contacting anyone. Then we file a real support ticket during testing and evaluate the channel (live chat, email, or ticket queue), the actual response time, whether the answer resolved the issue correctly, and, where a first response didn’t fully solve it, how the escalation was handled.


Callout Panels

Security review and privacy review

These are editorial reviews of what a vendor publicly documents, not a security audit we performed ourselves.

Security review

We check what a vendor’s own documentation discloses: encryption in transit and at rest, authentication options like two-factor login, account-level access controls, any published certifications, compliance documentation for regulated industries, and transparency reports where a vendor publishes one. We are not a penetration-testing firm, and we say so rather than implying a security audit we didn’t perform.

Privacy review

We read the actual privacy policy rather than summarizing a marketing page: what data is collected, how long it’s retained, whether customer data is used to train the vendor’s models, and how it’s disclosed if so. We check whether a business owner can export or delete their data on their own, what consent is required from end customers, and what account-level controls exist. Compliance-relevant context for regulated industries is flagged, not assumed.


Scoring Breakdown

How performance is weighted into a final score

Every category above feeds into one of five weighted groups. The exact per-criterion math lives on our Review Methodology page; this is the weighting at a glance.

AI output quality25%
Feature validation25%
Ease of use20%
Pricing value15%
Customer support15%

Comparison methodology. When several AI tools are placed in a single comparison table, each one is scored independently against this same weighting before the table is assembled. The table reflects scores that already existed; it isn’t built first and scored backward to fit a preferred outcome.

Normalization and consistency. Raw observations (a response time in seconds, a count of hallucinations across a set of test prompts) get converted into the same 1-to-5 or percentage scale used across every review, so a “4” in ease of use means the same thing on two different reviews. Reviewers are calibrated against the same rubric bands shown in the scorecards above, and a second editor spot-checks scores during the fact-check pass to catch a rating that’s drifted from the standard. This is quality assurance for consistency, not a second full retest.


Re-Testing

How often we retest, and what actually changes a score

Every review is fully retested at least every 90 days on a routine schedule. Below is an illustrative, generic example of the kind of change that triggers an update sooner than that, not a real tool’s actual history.

Before

CategoryPricing value
Feature statusIncluded in Starter plan
Score impactNo adjustment needed

After vendor change

CategoryPricing value
Feature statusMoved to a paid add-on
Score impactAdjusted same day

Illustrative example only. Pricing changes and feature removals are the two most common triggers for an off-cycle re-test.

What triggers a re-test outside the routine cycle. A pricing update, a new feature launch, a vendor acquisition or merger, an update to the underlying AI model that changes output quality, a disclosed security incident, a major product redesign, a pattern of reader-reported errors, or a vendor reaching out with a correction all trigger an off-cycle look, on top of the standard 90-day schedule that applies regardless.


Being Honest With You

Testing limitations, transparency, and where we’re still improving


Common Questions

Quick answers

The questions we hear most about our testing process specifically.

Want the short version, every Friday?

One practical AI idea for your business, tested by us first. No hype, no spam.