Our whole method fits in one sentence: we send the exact same prompt to each model on our bench on the same day, and show you the raw outputs, dated and versioned. Everything below is how we keep that honest.
No scores. No winners. No sponsored rankings. We believe "best" depends on who's asking, so we show the actual outputs instead of hiding them behind a number. Every page states the test date, the model names, and which tier each answer came from — so you can check our work.
A real everyday task, phrased the way a person would actually type it. The exact prompt is shown in full on every page.
All five assistants receive the identical prompt on the same day, through the same API pipeline, with identical settings.
We publish each answer in full. Nothing is cut except for light formatting — the text is the model's, not ours.
Each answer gets a descriptive note — what it covers, how it's structured, anything surprising. We never rank or recommend.
Answers are snapshots: models change fast, and a free tier today may not be the same model next quarter. That's exactly why we date everything and refresh key pages quarterly.
About tiers: every test labels the tier each answer came from — free, low-cost, or paid — so you can compare fairly. Entry tiers are not flagship models, and a vendor's strongest model may not be the one on the bench for a given test. The tier is always shown, so you know exactly what you're reading.
We run every test ourselves through the same pipeline. No tool pays for placement, no ranking is for sale, and if anything about a test changes, we say so on the page. Affiliate links may appear on the site — they never affect what a model outputs, because we don't touch the outputs.
Every test is run exactly as described — one prompt, five models, full raw answers.