Our whole method fits in one sentence: we send the exact same prompt to each model on our bench on the same day, and show you the raw outputs, dated and versioned. Everything below is how we keep that honest.
No scores. No winners. No sponsored rankings. We believe "best" depends on who's asking, so we show the actual outputs instead of hiding them behind a number. Every page states the test date, the model names, and which tier each answer came from — so you can check our work.
A real everyday task, phrased the way a person would actually type it. The exact prompt is shown in full on every page.
All five assistants receive the identical prompt on the same day, through the same API pipeline, with identical settings.
We publish each answer in full. Nothing is cut except for light formatting — the text is the model's, not ours.
Each answer gets a descriptive note — what it covers, how it's structured, anything surprising. We never rank or recommend.
Answers are snapshots: models change fast, and a free tier today may not be the same model next quarter. That's exactly why we date everything and refresh key pages quarterly.
About tiers: every test labels the tier each answer came from — free, low-cost, or paid — so you can compare fairly. Entry tiers are not flagship models, and a vendor's strongest model may not be the one on the bench for a given test. The tier is always shown, so you know exactly what you're reading.
We run every test ourselves through the same pipeline. No tool pays for placement, no ranking is for sale, and if anything about a test changes, we say so on the page.
How the site is funded: the site may carry affiliate links to third-party tools. If you sign up for one of those tools through our link, we may earn a commission at no extra cost to you.
What that does and doesn't affect: affiliate links are kept physically separate from every test and every answer. The models' outputs are raw snapshots taken through our own API pipeline — no sponsor, advertiser or affiliate partner has any influence on which models we test, how we run a test, or what appears in an answer. We never recommend or rank any tool in the tests, and we don't write our notes to favor any partner.
Every test is run exactly as described — one prompt, five models, full raw answers.