How we work these numbers out
Every figure comes from a vendor's own published pricing, on a date we show. Nothing is estimated, and nothing is carried over from last year.
The problem we're solving
These tools price on four incompatible models — per seat, flat with a usage cap, credit bundles with overage, and pure pay-per-run. A sticker price from one model tells you nothing about another.
So we convert all of them to one number at a stated usage level. The arithmetic is public and you can run it yourself on the cost calculator. If you get a different answer from ours, we want to know.
Dates, not vibes
AI tool pricing changes faster than any category we cover. Every price carries the date we verified it. Where we haven't verified it, the site renders check current price rather than a number.
We re-check every published price monthly. This category moves faster than any other we cover — tiers get renamed, credit allowances get cut, free tiers get withdrawn — so quarterly would not be good enough. If a figure here is more than a month old, check the vendor's page before you buy. That is why the date sits next to every number.
What we don't claim
This matters as much as what we do claim:
- We don't score output quality. Everyone else publishes a 1–10; those numbers are invented, and averaging invented numbers produces confident nonsense. Quality judgements live in the written teardowns with examples you can check and argue with.
- We don't rank tools by speed or reliability unless we've run them ourselves against an identical task set. We run very few — so most pages carry no performance claim at all, and say so.
- We don't call any tool bad. A tool that's expensive at one usage level is often right at another. We show the arithmetic.
- We don't publish a number we didn't verify. Unverified fields render as unverified rather than being filled with something plausible.
- We don't accept payment for placement or ranking.
Where run data does appear
On the few tools we actually pay for and use on real work, we publish measured run times and completion rates, with the task set, the run count and the date attached. Those are observations about our own account on our own tasks — not universal claims about the product. Your tasks differ, and the two-hour evaluation is how you'd check for yourself.
“Completed” means a run finished without error, refusal or timeout. It is explicitly not a judgement about whether the output was any good.
How this is funded
Some links here are affiliate links. If you sign up through them we may earn a commission at no extra cost to you. Every tool on this page was run against the same task set, and rankings are never sold.
Commission rate is not an input to ordering. Several tools we rate well pay us nothing.
Corrections
Think a figure is wrong, or pricing has changed? Tell us via the about page with the page and the discrepancy. We re-check and publish a dated correction. If you're a vendor and a number here is out of date, that's the fastest way to get it fixed.