AI chat comparison · 2026-09-30
AI chat comparison: same prompt, every model, side by side
What a usable AI chat comparison looks like
Most “AI chat comparison” pages hand you a ranking built from marketing claims or cherry-picked demos. A comparison is only usable when three things hold: every model gets the same prompt, the answers are read side by side, and the conclusion is tied to a task you actually have. This page keeps both halves of that standard: the latest public test data we track (dated and sourced below), and the method to run the same comparison on your own prompts in one click.
The comparison table: latest public test results
Below are the current public results for the models ModelAny can reach, grouped by test. Each table lists only the models with a public result in that exact test—models without one are omitted, and scores from different tests cannot be added up. Exact model versions and retrieval dates are shown so newer generations are not judged against older ones.
SWE-bench Verified · real software-issue fixing
SWE-bench Verified measures how often an AI coding setup can fix real GitHub issues. A higher resolved percentage means more issues were fixed in that specific test setup.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| Claude | Sonar Foundation Agent + Claude 4.5 Opus | 1 | 79.2% | Resolved (%) |
| Doubao | TRAE + Doubao-Seed-Code | 3 | 78.8% | Resolved (%) |
| Gemini | live-SWE-agent + Gemini 3 Pro Preview (2025-11-18) | 4 | 77.4% | Resolved (%) |
| ChatGPT | JoyCode + Claude 4 Sonnet + GPT-4.1 | 17 | 74.6% | Resolved (%) |
| GLM | GLM 5 (high) | 25 | 72.8% | Resolved (%) |
| Kimi | Lingxi v1.5 x Kimi K2 | 35 | 71.2% | Resolved (%) |
| DeepSeek | DeepSeek V3.2 (high) | 46 | 70% | Resolved (%) |
| Qwen | Nebius AI Qwen 2.5 72B Generator + LLama 3.1 70B Critic | 137 | 40.6% | Resolved (%) |
LiveBench · agentic coding
LiveBench scores models on regularly refreshed objective tasks. Higher category scores mean better measured performance on that task type—not a universal ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| DeepSeek | deepseek-v4.1-flash-max | 1 | 77.27266666666667 | Category score (points) |
| Claude | claude-opus-5-5-max-effort | 2 | 71.717 | Category score (points) |
| Kimi | kimi-k3 | 11 | 62.17166666666666 | Category score (points) |
| GLM | glm-5.3 | 15 | 60.909 | Category score (points) |
| Gemini | gemini-3.7-flash-high | 19 | 58.282999999999994 | Category score (points) |
| ChatGPT | gpt-6-astra-max | 21 | 57.323 | Category score (points) |
LiveBench · coding tasks
LiveBench scores models on regularly refreshed objective tasks. Higher category scores mean better measured performance on that task type—not a universal ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| Claude | claude-sonnet-5-5-max-effort | 1 | 91.366 | Category score (points) |
| ChatGPT | gpt-5.6-sol-max | 7 | 83.941 | Category score (points) |
| Kimi | kimi-k3 | 15 | 81.4455 | Category score (points) |
| DeepSeek | deepseek-v4.1-flash-max | 20 | 80.037 | Category score (points) |
| GLM | glm-5.2 | 22 | 79.654 | Category score (points) |
| Gemini | gemini-3.7-flash-high | 27 | 78.8885 | Category score (points) |
LiveBench · data analysis
LiveBench scores models on regularly refreshed objective tasks. Higher category scores mean better measured performance on that task type—not a universal ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| ChatGPT | gpt-6-astra-max | 1 | 82.97333333333334 | Category score (points) |
| Claude | claude-fable-5-max-effort | 4 | 80.53766666666667 | Category score (points) |
| DeepSeek | deepseek-v4-flash-vision-exp | 11 | 79.47666666666667 | Category score (points) |
| Kimi | kimi-k3 | 19 | 78.73366666666668 | Category score (points) |
| Gemini | gemini-3.1-pro-preview-high | 21 | 78.54133333333334 | Category score (points) |
| GLM | glm-5.3-flash | 31 | 76.401 | Category score (points) |
LiveBench · instruction following
LiveBench scores models on regularly refreshed objective tasks. Higher category scores mean better measured performance on that task type—not a universal ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| Gemini | gemini-3.8-flash-high | 1 | 81.4125 | Category score (points) |
| Claude | claude-fable-5-max-effort | 6 | 75.77074999999999 | Category score (points) |
| ChatGPT | gpt-6-astra-max | 8 | 75.57900000000001 | Category score (points) |
| Kimi | kimi-k3 | 22 | 71.36275 | Category score (points) |
| DeepSeek | deepseek-v4-flash-vision-exp | 24 | 70.95824999999999 | Category score (points) |
| GLM | glm-5.3 | 31 | 69.304 | Category score (points) |
LiveBench · language tasks
LiveBench scores models on regularly refreshed objective tasks. Higher category scores mean better measured performance on that task type—not a universal ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| Claude | claude-fable-5-max-effort | 1 | 90.68400000000001 | Category score (points) |
| ChatGPT | gpt-6-astra-max | 3 | 89.43066666666668 | Category score (points) |
| Gemini | gemini-3.8-flash-high | 5 | 87.79266666666666 | Category score (points) |
| Kimi | kimi-k3 | 9 | 85.528 | Category score (points) |
| DeepSeek | deepseek-v4-pro-0813 | 24 | 82.074 | Category score (points) |
| GLM | glm-5.3 | 29 | 79.85633333333332 | Category score (points) |
LiveBench · mathematics
LiveBench scores models on regularly refreshed objective tasks. Higher category scores mean better measured performance on that task type—not a universal ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| Claude | claude-opus-5-5-max-effort | 1 | 97.07749999999999 | Category score (points) |
| ChatGPT | gpt-6-astra-max | 3 | 96.80624999999999 | Category score (points) |
| DeepSeek | deepseek-v4-pro-0813 | 14 | 95.08624999999999 | Category score (points) |
| Gemini | gemini-3.7-flash-high | 18 | 93.46799999999999 | Category score (points) |
| GLM | glm-5.2 | 33 | 89.78125 | Category score (points) |
| Kimi | kimi-k3 | 51 | 84.43675 | Category score (points) |
LiveBench · reasoning
LiveBench scores models on regularly refreshed objective tasks. Higher category scores mean better measured performance on that task type—not a universal ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| ChatGPT | gpt-6-astra-max | 1 | 92.65375 | Category score (points) |
| Claude | claude-opus-5-5-max-effort | 2 | 92.15375 | Category score (points) |
| Kimi | kimi-k3 | 7 | 90.673 | Category score (points) |
| Gemini | gemini-3.8-flash-high | 16 | 89.29325 | Category score (points) |
| DeepSeek | deepseek-v4.1-flash-max | 29 | 86.69225 | Category score (points) |
| GLM | glm-5.3 | 33 | 85.803 | Category score (points) |
Arena · coding preference
Arena asks people to pick the better answer without knowing which model wrote it. A higher Elo means more preference votes in that category—not an overall product ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| Claude | claude-opus-5.5-max | 1 | 1827 | Elo (Elo) |
| ChatGPT | gpt-6-astra-max | 2 | 1792 | Elo (Elo) |
| Kimi | kimi-k3-max | 9 | 1660 | Elo (Elo) |
| DeepSeek | deepseek-v4.1-flash-max | 16 | 1621 | Elo (Elo) |
| GLM | glm-5.3-max | 18 | 1619 | Elo (Elo) |
| Gemini | gemini-3.7-flash-high | 23 | 1593 | Elo (Elo) |
Arena · search-style preference
Arena asks people to pick the better answer without knowing which model wrote it. A higher Elo means more preference votes in that category—not an overall product ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| ChatGPT | gpt-5.6-sol-xhigh | 1 | 1257 | Elo (Elo) |
| Claude | claude-opus-4-6-search | 2 | 1253 | Elo (Elo) |
| Wenxin | ernie-5.1 | 6 | 1227 | Elo (Elo) |
| Gemini | gemini-3.1-pro-grounding | 9 | 1210 | Elo (Elo) |
Arena · general chat preference
Arena asks people to pick the better answer without knowing which model wrote it. A higher Elo means more preference votes in that category—not an overall product ranking.
| Product | Exact model version | Rank | Score | Metric |
|---|---|---|---|---|
| Claude | claude-opus-5.5-high | 1 | 1509 | Elo (Elo) |
| Gemini | gemini-3.8-flash-high | 10 | 1492 | Elo (Elo) |
| Kimi | kimi-k3-max | 16 | 1488 | Elo (Elo) |
| ChatGPT | gpt-5.6-sol-xhigh | 19 | 1483 | Elo (Elo) |
| GLM | glm-5.3-max | 24 | 1480 | Elo (Elo) |
| DeepSeek | deepseek-v4.1-flash-max | 29 | 1477 | Elo (Elo) |
Prefer one test per model pair with methodology notes? Browse all public benchmark snapshots by scenario, or open a pair such as ChatGPT vs Claude, ChatGPT vs DeepSeek or ChatGPT vs Gemini.
Run the comparison on your own prompts: the five-step method
- Fix the task first. Pick a real task—fix this failing test, draft this reply, summarize this contract clause—and write down what a good answer must contain before you read any model.
- Freeze one prompt. Word it once and send the identical text everywhere. Retyped prompts drift, and drifted prompts measure your typing, not the models.
- Send to every model at once. In ModelAny, tick the models—ChatGPT, Claude, Gemini, DeepSeek, Grok, Tencent Yuanbao, Wenxin, Qwen, Doubao, Kimi, GLM—and send. Each official site answers under your own account.
- Score on four axes. Factual accuracy (verifiable claims), completeness against your criteria, edit cost (what you had to fix), and speed. Note who cited sources, who hedged, who invented.
- Re-run on a second task before deciding. One prompt makes a data point; two or three different task types make a pattern. Models lead in different categories—the tables above say the same thing.
The full walkthrough with a scoring rubric is in the side-by-side comparison method; the three-way version is in ChatGPT vs Claude vs Gemini on the same prompt.
One click instead of eleven tabs
Manual comparison means opening each site, pasting the prompt, and juggling tabs until context falls apart. ModelAny is a free Chrome and Edge extension that sends one prompt to up to 11 official AI sites—ChatGPT, Claude, Gemini, DeepSeek, Grok, Kimi, Qwen, Doubao, GLM, Tencent Yuanbao and Wenxiaoyan—and lines the answers up side by side. It uses the sessions you already have: no API key, no subscription, and no ModelAny server in the middle. Drafts, settings, and history stay in your browser.
New here? Start with how to ask multiple AI at once, or the same-prompt model comparison workflow.
Frequently asked questions
Which AI chat is best right now?
There is no single answer: the honest ranking changes by task and by model version. The public test tables on this page show measured results with dates and sources, and the same-prompt method shows you how to settle it for your own task in minutes.
How do I compare AI chatbots fairly?
Freeze one prompt, send the identical text to every model you care about, and score the answers against success criteria you wrote before reading them—accuracy, completeness, edit cost. Comparing answers to differently-worded prompts tells you about prompts, not models.
Which models can I compare with ModelAny?
Eleven official sites: ChatGPT, Claude, Gemini, DeepSeek, Grok, Kimi, Qwen, Doubao, GLM (ChatGLM), Tencent Yuanbao and Wenxiaoyan (ERNIE)—side by side, from one prompt box.
Is there a free AI chat comparison tool?
ModelAny is free with no subscription or credits: it sends one prompt to the official sites under the accounts you already have. The benchmark tables on this page are also free to read, with sources linked.
What does the comparison table on this page measure?
Public, third-party test results—Arena preference votes, SWE-bench Verified issue fixing, LiveBench task scores—shown per category with the exact model version, retrieval date and a link to the original leaderboard. They are measurements under stated conditions, not a universal ranking.
Run this comparison on your own prompts
Install ModelAny free on Chrome or Edge, ask once, and read answers from 11 AI sites side by side. Chrome Web Store · Microsoft Edge Add-ons
Install extension