Public benchmark snapshot: 2026-07-19

ChatGPT vs Gemini: public benchmark comparison

How to read this page

This page summarizes public third-party benchmarks in plain language. It helps you see where the models stand in published data, but it does not replace trying them on your own tasks.

What public benchmarks show

Below are results only from public benchmark categories where every model on this page appears together. Scores from different sources cannot be added up, and they do not prove one model is best overall.

Arena · coding preference

Arena asks people to pick the better answer without knowing which model wrote it. Higher Elo means more people preferred that model in that category.

Retrieved: Jul 19, 2026, 8:44 PM · Open original leaderboard

ProductExact model versionRankScoreMetric
ChatGPT gpt-5.6-sol-xhigh (codex-harness) 3 1618 Elo (Elo)
Gemini gemini-3.5-flash 19 1504 Elo (Elo)

Arena · search-style preference

Arena asks people to pick the better answer without knowing which model wrote it. Higher Elo means more people preferred that model in that category.

Retrieved: Jul 19, 2026, 8:44 PM · Open original leaderboard

ProductExact model versionRankScoreMetric
ChatGPT gpt-5.5-search 2 1239 Elo (Elo)
Gemini gemini-3.1-pro-grounding 7 1211 Elo (Elo)

Arena · general chat preference

Arena asks people to pick the better answer without knowing which model wrote it. Higher Elo means more people preferred that model in that category.

Retrieved: Jul 19, 2026, 8:44 PM · Open original leaderboard

ProductExact model versionRankScoreMetric
Gemini gemini-3-pro 8 1486 Elo (Elo)
ChatGPT gpt-5.6-sol-xhigh 10 1486 Elo (Elo)

SWE-bench Verified · real software-issue fixing

SWE-bench Verified measures how often an AI coding setup can fix real GitHub issues. A higher resolved percentage means more issues were fixed in that test.

Retrieved: Jul 19, 2026, 8:44 PM · Open original leaderboard

ProductExact model versionRankScoreMetric
Gemini live-SWE-agent + Gemini 3 Pro Preview (2025-11-18) 4 77.4% Resolved (%)
ChatGPT JoyCode + Claude 4 Sonnet + GPT-4.1 17 74.6% Resolved (%)

Browse all public benchmark data by scenario

Product entry points

ChatGPT vs Gemini: official sites and support status
ProductProviderOfficial entry pointModelAny
ChatGPT OpenAI Visit product site Currently supported by ModelAny
Gemini Google Visit product site Currently supported by ModelAny

What to check when you try them yourself

  • Whether your real question is answered well, or still needs heavy editing.
  • Whether the current plan is enough, and whether the price fits.
  • Whether privacy, sign-in, and team features fit your needs.

Official sources and verification date

The links below are official sources for product identity and plan information, last checked 2026-07-19. Pricing, model availability, quotas, and regional access can change; confirm on the provider site before purchasing or deploying.

Frequently asked questions

Does this page publish an absolute ranking?

No. Conditional findings are published only when test conditions, raw outputs, and a review method are disclosed.

Are pricing and plan details definitive?

This page links to official sources and verification dates. Confirm the provider price shown for your region before purchasing.

Does ModelAny store prompts?

ModelAny is local-first. Prompts are sent only to the AI services you choose and are not uploaded to ModelAny servers.

Compare multiple models with one prompt

ModelAny is available on the Chrome Web Store. The Microsoft Edge Add-ons listing is still under review. Drafts, settings, and history remain in your browser.

Install extension