Which AI Model Should You Use?
Five quick questions. A specific model recommendation — not just "it depends."
Recommended for you
Ranked from real, auto-synced pricing data. See exact monthly cost for your usage →
Rankings use price tier and general model-family positioning, not benchmark scores — for a purely capability-driven choice, cross-check independent leaderboards too.
How this recommendation works
There are 19+ actively maintained frontier models across six providers, and every one of them claims to be great at everything. Nobody has time to read six pricing pages and three benchmark leaderboards just to send an email or debug a function. This tool collapses the decision into two things that actually predict which model is right for you: what you're using it for, and what you're willing to pay — then ranks real, currently-priced models against those answers.
The ranking runs in two steps. First, your budget answer filters to a price tier — budget, balanced or flagship — using the same tier classification that powers the rest of this site, not a fuzzy guess. Second, within that tier, models tagged as strong for your stated use case (coding, writing, reasoning, automation, or general-purpose) are surfaced first; the rest are sorted cheapest-first as a tiebreak. If you said your workload can wait for a batch run, the ranking prefers models that actually offer batch pricing — about a third of the catalog doesn't, so this step can meaningfully change the result.
The use-case tags are hand-curated general positioning (e.g. "Anthropic's Claude family is a common default for coding and long-form writing") — well-established reputational framing, not claims about exact context windows or benchmark scores we haven't independently verified. If precise capability comparisons matter more to you than general positioning, treat this as a shortlist and cross-check it against a model-capability leaderboard before committing.
Why it also checks your subscription status
A flat-fee subscription like ChatGPT Plus or Claude Pro and a pay-as-you-go API key solve the same problem in opposite ways, and for a lot of people the honest answer to "which model should I use" is actually "you don't need the API at all yet." If you say you're already on a subscription, or that your usage is occasional, this tool points you at the Subscription vs API calculator instead of just picking a model — that tool answers the more fundamental question with your real numbers.
What this tool doesn't do
It doesn't run your prompts through every model and score the outputs — nobody's calculator does that credibly, and anyone claiming otherwise is guessing. It doesn't factor in context-window size, since publicly comparable context limits change too often to hand-maintain honestly here. And it isn't a substitute for actually trying the top 1-2 candidates on your real task for five minutes — pricing and positioning narrow the field, but the final call on quality is still yours to make.
FAQ
How is this different from just asking an AI which model to use?
An AI assistant answering that question is, at best, reciting roughly what's written here — and it doesn't have live pricing to check its answer against. This tool starts from the actual current price and tier of every model in the dataset, so the recommendation can't drift out of date the way a model's training-time knowledge can.
Why did the recommendation change when I changed "budget" but not "use case"?
Budget is the primary filter — it picks the price tier first. Use case only changes the order within that tier. If every model in a tier happens to share the same use-case tags, changing the use case alone won't visibly reorder the list; that's expected, not a bug.
What if my use case doesn't fit neatly into one category?
Pick whichever option is closest, or use "General everyday use" — it's a safe default that favors well-rounded models over narrow specialists. You can also just try two different answers and compare both shortlists; the tool is free to re-run as many times as you like.
Does this account for context window size or benchmark scores?
No — deliberately. Context limits and benchmark results change frequently enough, and vary enough by task, that hand-maintaining them here would go stale faster than the price data does. This tool optimizes for price-fit and general positioning; pair it with an independent capability leaderboard if raw benchmark performance is your deciding factor.