Review Methodology

Exactly how an AI software score gets calculated

Every review on AIBizMaster ends in a single number out of 10. This page shows the math behind it: the weighted categories, the bonus points and penalties, how ties get broken, what our badges actually require, and how confident we are in any given score. Editorial Standards covers governance and independence. How We Test covers the hands-on process that produces the raw testing notes. This page is where those notes turn into the number a reader actually sees, and every published score runs through the exact formula documented here. Adjust the calculator below to see it for yourself.

8.4out of 10

Sample dial for illustration, not a real product’s score.


Why Weighted Scoring

Why every category doesn’t count equally

Treating a tool’s customer support quality the same as its core AI output quality would let a mediocre product with a great support team outscore a genuinely excellent one. Categories are weighted by how much they actually affect whether a small business gets value from the tool.

Unweighted (equal categories)

A tool with flawless support but unreliable AI output can outrank a tool that does the core job well but has average support, a misleading result for the decision that actually matters.

Weighted (AIBizMaster)

AI output quality and feature validation carry the most weight because they determine whether the tool solves the actual business problem, with pricing and support weighted appropriately below that.

What a score is trying to measure. A high score means real, measurable usefulness to a small business, not polish, not marketing copy, and not a long feature list that goes mostly unused. A tool with a beautiful interface and a thin feature set scores lower than a plain-looking tool that reliably does the job it’s bought for. Scores are built from what testing actually found, not from what a vendor’s homepage claims.

A concrete example. Take an AI scheduling assistant. If it books appointments accurately but its interface looks dated, that’s a minor Ease of Use deduction on a category worth 20% of the total. If it has a polished interface but double-books appointments during testing, that’s a serious AI Output Quality and Feature Validation problem on categories worth 25% each. The weighting reflects which failure actually costs a business money and customers.

Why weighting stays fixed across reviews. The same five weights (25/25/20/15/15) apply to every AI software review under this methodology version, regardless of category or vendor. A weight doesn’t shift because a particular tool would score better under a different formula. If the weighting itself changes, that’s a new methodology version, logged with a date, not a silent adjustment applied to make one review’s math work out.

Normalization and fairness. Every category is scored on the same 0 to 10 scale using the same rubric bands before any weighting is applied, so a 7 in Pricing Value and a 7 in AI Output Quality represent the same relative performance before the weights pull them apart. That’s what keeps a “Very Good” tool in one comparison table comparable to a “Very Good” tool in a different one.


Scoring Categories & Scale

The 10-point scale and what each category measures

Every category is scored independently on the same 0–10 scale before weighting is applied.

1–2Poor
3–4.9Below Average
5–7.4Good
7.5–8.9Very Good
9–10Exceptional

Poor (1–2)

The tool fails at its core job in normal use: broken workflows, frequent hallucinations, or advertised features that don’t work. We rarely publish a full review at this level; more often the tool is dropped from the shortlist during initial testing.

Below Average (3–4.9)

The tool works, but with real friction: unreliable output on anything beyond simple requests, a confusing setup, or gaps between what’s advertised and what’s actually usable at the tested tier.

Good (5–7.4)

A solid, usable tool with clear limitations. This is where most competent AI software actually lands: it does the job, but a specific weakness (price, a missing feature, a rough onboarding) keeps it from a stronger recommendation.

Very Good (7.5–8.9)

Strong performance across nearly every weighted category, with at most one minor limitation. Most tools we’d actually recommend to a small business land here.

Exceptional (9–10)

Near-flawless execution across every category with no meaningful limitation found during testing. Very few tools should reach this band; if most of our published scores clustered here, that would be a sign our rubric had gone soft, not that the AI software industry is uniformly excellent.

What each category actually measures

AI Output Quality

Measures accuracy, tone appropriateness, and freedom from confident factual errors, checked against real business content we already know the correct answer for. It does not measure raw response speed (that’s its own category) or the length of a feature list. Strong performance looks like correct, on-brand answers across repeated tries; weak performance is a hallucination stated as confidently as a correct answer. Evidence comes from the prompt and workflow testing described on How We Test.

Feature Validation

Measures whether every advertised feature works as described at the tested pricing tier, verified through direct use rather than a spec sheet. It does not measure whether a feature is useful to every reader, only whether it works as claimed. A feature gated behind an untested enterprise tier is disclosed, not scored as broken.

Ease of Use

Measures whether a non-technical owner can complete core setup and daily tasks unaided, timed during testing. It does not measure visual design on its own; a plain interface that’s easy to navigate scores higher than a polished one that hides basic settings.

Pricing Value

Measures cost relative to what a small business actually gets, not simply the lowest sticker price. A $99 tool that replaces three separate subscriptions can score higher than a $29 tool with a thin feature set. It does not measure whether a business can afford the price, only whether the price is fair for what’s delivered.

Customer Support

Measures response time and answer quality on a real support ticket filed during testing, plus documentation quality. It does not measure enterprise-tier support (a dedicated account manager, for example) unless that’s the tier actually tested.


Try It Yourself

The scoring calculator

Move the sliders to see exactly how five category scores combine into one final number using our published weights.

8.0
7.5
9.0
7.0
8.5

Weighted total

7.9

Very Good

This is the same weighted formula used on every published review. No hidden adjustment happens after this number is calculated.

How the weighted average works. Each category score is multiplied by its weight, then the five results are added together: (AI Output Quality x 0.25) + (Feature Validation x 0.25) + (Ease of Use x 0.20) + (Pricing Value x 0.15) + (Customer Support x 0.15). Try setting every slider to the same number, say 8.0, and the total comes out to 8.0 too, since the weights sum to 1. That’s the check we use to confirm the math is right.

Why the total isn’t a simple average. A plain average would divide by 5 and treat every category as equally important, which is the exact problem the “Why Weighted Scoring” section above walks through. Two tools with identical category scores but different weighting priorities would never land on the same simple average, but they will land on the same weighted total under this formula, because the formula is what’s fixed, not the raw numbers.

Rounding. The weighted total is calculated to two decimal places internally, then rounded to one decimal place for display, using standard rounding (a result of 7.85 displays as 7.9, not 7.8). Bonuses and penalties, covered next, are applied to this rounded weighted total before the final published number is set.


Bonus Points & Penalties

Adjustments applied after the weighted total

A small number of adjustments apply on top of the weighted score, capped so they can never override a genuinely poor category result.

Bonus points

Genuinely novel feature not offered by comparable tools+0.2
Easy-to-find pricing with no hidden fees+0.1
Free tier genuinely usable for a real small business+0.2
Onboarding that gets a new user to a working result unaided+0.1

Penalties

Advertised feature not available at the tested tier−0.5
Confirmed hallucination in AI output during testing−0.4
Cancellation or data export made deliberately difficult−0.3
Pricing that’s misleading or hidden behind a sales call−0.3
Automation that fails inconsistently across repeated runs−0.3

What bonuses can’t do. Bonus points reward something genuinely above the baseline, not a marketing checkbox. They’re capped at +0.5 total per review, and they can’t lift a tool with a poor weighted score into a strong-looking published number. A tool with a 4.2 weighted average and every bonus available still publishes around 4.7, not a score that hides the underlying problem.

How penalties are applied. Penalties are evidence-based: each one requires something we observed directly during testing (a hallucination we can point to, a feature we tried to use and couldn’t, a cancellation flow we walked through ourselves), not a suspicion or a vendor’s competitor complaining. Where a security concern is based on publicly documented evidence, such as a vendor’s own disclosed data-handling practices rather than a claim we can’t verify, it’s treated the same way: cited to its source and applied consistently to any tool with the same issue.

How the final score is built. The published score is the weighted total, plus any bonuses, minus any penalties, capped at 10.0 and floored at 0.0. That combined number, not the pre-adjustment weighted total, is what displays as the review’s headline score and what’s used for badges, rankings, and comparison tables. It’s rounded to one decimal place for display, the same rounding rule used in the calculator above.


Tie-Breaking & Rankings

How ties get broken and rankings determined

Rankings in a comparison table are simply the weighted totals sorted highest to lowest. When two tools land on the exact same total, this order applies:

Are overall weighted scores tied?

Step 1

Higher AI Output Quality score wins

If still tied

Higher Ease of Use score wins

If still tied

Lower price at an equivalent tier wins

Still tied

Both tools are shown, marked as tied. No artificial winner is forced.

Why a tie can stay a tie. Two tools landing on the same number after all three tie-breakers is a real, honest outcome, not a formula failure to cover up. When it happens, a comparison table lists both tools together at that rank rather than assigning one an arbitrary edge it didn’t earn. Editorially, we’ll still describe what actually differs between them in the write-up (one might suit a solo operator better, the other a small team) even when the numbers alone can’t separate them.


Our Badges

What each badge actually requires

Most badges are calculated automatically from scores. One, Editors’ Choice, involves editorial judgment, and we say so directly rather than presenting it as purely formulaic.

Best Overall

Highest weighted total in its comparison, with no single category scoring below 6.

Score-driven

Best Value

Highest score-per-dollar ratio among compared tools, not simply the cheapest option.

Score-driven

Best for Beginners

Ease of Use score of 9 or higher, regardless of overall rank.

Score-driven

Editors’ Choice

Awarded for exceptional execution in one standout area, decided editorially rather than by formula alone.

Editorial judgment

Industry-specific recommendations. A tool can also be recommended for a single vertical, such as dental or health, when it scores Very Good or higher specifically on criteria relevant to that industry, like compliance-aware data handling, even if its general-purpose score is more middling. This recommendation is separate from the four badges above and is always labeled by industry so it isn’t confused with a general endorsement.

What “exceptional execution in one standout area” means for Editors’ Choice. This isn’t a vague catch-all. It requires a specific, named reason: a tool that solved a real testing problem no comparable product handled well, for example, or one whose AI output quality was consistently ahead of its category during testing. That reason is stated in the review itself, not left implicit.


Confidence Levels

How confident we are in any given score

Not every review carries the same weight of evidence behind it. Each one is labeled with a confidence level so readers know how much testing history stands behind the number.

High Confidence

Full hands-on testing completed and re-verified within the last 90-day cycle, with at least one prior re-test showing consistent results.

Medium Confidence

Fully tested, but approaching or slightly past the standard re-verification window, or tested only once so far.

Limited Confidence

Initial testing only completed; insufficient long-term data to fully confirm consistency yet.

Why a brand-new review starts at Limited or Medium confidence. A single testing pass tells us how a tool performed once. It doesn’t yet tell us whether that performance holds up after a vendor update, a pricing change, or a second look three months later. Confidence rises as a tool goes through repeat verification and the score holds steady, and it’s not a mark against a new review; it’s an honest signal that less history exists yet.

Data verification process. Every score is checked against the underlying testing notes before publication, cross-referencing that the written verdict actually matches the numeric result rather than the two drifting apart during editing.


Reviewer Consistency

Keeping scores consistent across reviewers and time

A score should mean the same thing whether it was assigned in January or in October, and whether one reviewer or another did the testing.

Independent scoring

Testing notes are scored against the published rubric before a written verdict is drafted, so the number isn’t reverse-engineered from the conclusion.

Cross-check pass

A second person reviews the score against the same testing notes independently, without seeing the first reviewer’s numbers in advance.

Conflict resolution

A category disagreement of more than one point triggers a third, independent pass rather than averaging it away.

Final sign-off

The reconciled score is locked before the review is scheduled to publish, and the math is checked one more time against the weighting formula.

Score quality control before publication. Beyond the reviewer cross-check above, every score goes through a separate quality pass: the underlying facts get verified (does the review cite what testing actually found), the arithmetic gets checked (does the weighted formula produce the displayed number), and a final editorial read confirms the written verdict matches the score rather than reading more positive or negative than the number suggests.

How scores stay comparable across different products. A score for a customer support chatbot and a score for a marketing content generator use the same five categories, the same weights, and the same 10-point rubric bands, calibrated through the reviewer cross-check process above. That’s what makes an 8.2 mean roughly the same thing in two completely different software categories: the categories being scored may look different in substance, but the scale, the weighting, and the review process behind them don’t change.


Updates & History

How updates, reader feedback, and revisions are tracked

How updates affect scores. A re-test can move a score up or down. The previous score isn’t preserved as a hidden “original”; the current published number is always the most accurate one we have.

What triggers a rescore outside the routine cycle. A major product update, a pricing change, an update to the underlying AI model that affects output quality, a vendor acquisition, a disclosed security incident, a removed feature, or a vendor reaching out with a correction all trigger a look before the standard 90-day cycle would otherwise catch it. Reader feedback that surfaces a verifiable factual error works the same way.

Historical score tracking. Where a score changes meaningfully between refresh cycles, the review notes what changed and when, so the shift is explained rather than silently overwritten.

Reader feedback adjustments. Feedback identifying a verifiable factual error is treated as a correction under our Corrections Policy. Feedback reflecting a different subjective experience doesn’t change a score without new testing evidence, but repeated similar feedback can move a tool up the re-test queue.

Methodology version history

This is the first published version of our Review Methodology. We’re not going to invent a fake history of past revisions to look more established than we are. From this point forward, any material change to how scores are calculated (a new category, a changed weight, a new badge) gets logged directly on this page with the date of the change. A small clarification to wording gets a minor version note; a change to weights, categories, or the point scale gets a full version number, and reviews scored under an earlier version stay attributed to that version rather than being silently recalculated under the new one unless we explicitly say a full rescore happened.


Common Questions

Quick answers

The questions we hear most about how scores work.

Want the short version, every Friday?

One practical AI idea for your business, tested by us first. No hype, no spam.