Every raw model response is publicly verifiable
Download image (with QR code)Share on XShare on LinkedIn

WeChat: save the image and long-press the QR code in a chat.

← All insights

AI Coding Agents: A Model Name in Fourth Place, and One Product Under Two Names

Published 09/22/2026

A 2026-09-22 measurement across 12 LLMs of which AI coding agents they recall, why a model name lands fourth, and how one product's two names score as one row.

Share
Download image (with QR code)Share on XShare on LinkedIn

WeChat: save the image and long-press the QR code in a chat.

Why this question, now

Since spring 2026 every frontier-model vendor has shipped a coding agent of its own, and tools that run and coordinate several agent sessions at once have appeared alongside them. The category turns over roughly once a month. We asked 12 large language models (seven head seats and five long-tail seats) "Most Popular AI Coding Agents", in English and separately in Chinese, to see which names this category brings to mind for them right now.

This is a hot-topic observation. It answers "what do models currently recall when asked this", not "which tool is better".

What the measurement shows

In the measurement taken on 2026-09-22, the English question reached a panel completeness of 92.7%: one long-tail seat returned nothing usable, all seven head seats answered, and no head model refused. Normalization status was ok. The top ten, with index scores, were GitHub Copilot 100, Cursor 90.0, Windsurf (Devin Desktop) 74.86, Claude 49.29, Replit 41.18, Amazon Q Developer 37.98, Devin 35.64, Tabnine 34.39, Cline 18.21 and Codex 15.54.

The Chinese question, asked the same morning, got all twelve models to answer. Its top five were GitHub Copilot, Cursor, Claude, Windsurf (Devin Desktop) and ChatGPT, and eight of the ten names in the two lists are shared (Jaccard 0.667). The Chinese-language article on this measurement goes into that comparison; this one stays with the English answers.

Three observations

1. A model name sits in fourth place on a list of tools

Claude scores 49.29 and comes fourth. It is a family of models, not a coding product. The question asked for coding agents, yet four of the seven head models (DeepSeek, GPT, Kimi and Qwen) put Claude in their top five. Qwen also put ChatGPT in its top five, although ChatGPT does not make the English top ten at all. Claude and ChatGPT here are model or chat-product names, not coding tools, and readers quoting fourth place should say so.

One detail is worth recording: the head models that left Claude out of their top five include Anthropic's own Claude Sonnet, which listed GitHub Copilot, Cursor, Cline, Windsurf (Devin Desktop) and Amazon Q Developer instead. Vendors' coding agents often share a name or a prefix with their models, so when a model reaches for a familiar name it may land on the model rather than the product. We did not rewrite "Claude" into any specific product and we did not remove it. The table shows what the normalization pipeline produced.

2. One product, one row, two names

Third place, at 74.86, reads Windsurf (Devin Desktop). Windsurf and Codeium, the name this product carried earlier, point at the same entry in our brand entity table, so the aggregation for this run scores them on one row instead of two, and the display name carries both names so a reader can tell which one is being counted. At entity level, six of the seven head models had that entry in their English top five; only Qwen did not. In the Chinese answers the same entry sits fourth at 63.10, one place behind Claude.

This is a property of the entity table rather than an editorial decision made for this article. When a name that the table files under an existing entry shows up in a run, the scores are added on that entry's row, and the same run's table can read differently before and after the table learns about a renaming. Anyone comparing this list with an older copy of the same run will see Windsurf and Codeium as two rows there and one row here; the run itself, and the answers the models gave, did not change.

3. Vendors' own models and their own tools

The seven head models' pairwise agreement on their top five was 0.480 for the English question, above the 0.30 threshold we set for hot-topic observations. Everyone agrees on GitHub Copilot and Cursor in the first two places; below that, each model has its own list. Codex, for instance, appears in Kimi's English top five and in Qwen's Chinese top five, but not in the top five of GPT, the model from Codex's own vendor, in either language. Google's model does not name Gemini in its Chinese top five either, while GPT and Grok both do. And, as above, the one head model that left Claude out of the English top five was Anthropic's.

A single run's top-five slice is too coarse to settle how a vendor's model treats its own tool, and the three cases above point in the same direction only by accident of this one morning. A much larger bilingual scenario-grid measurement on the same category, which counts mentions across repeated samples rather than reading one slice, is published separately; the link is at the end of this page. All of this describes what each model says when asked, not the tools themselves.

Limits

This measurement's panel completeness was 92.7% for the English question and 100% for the Chinese question, with 7 head models + 5 long-tail models scoring; weight sources are on the methodology page. Normalization status was ok. One reminder: the evaluation panel itself consists of large language models, and their training data may include the historical results of other AI rankings and reviews, so a form of self-referential contamination is possible: a model's impression of a brand may partly come from that brand having been mentioned by other rankings rather than from independent judgment. This is a known limit of the current methodology, not something specific to this run.

  • Model names and tool names share one table: Claude and ChatGPT are model or chat-product names. We did not rewrite them; please label them when quoting.
  • Brand normalization: Windsurf and the earlier name Codeium are one entity and one row in this run's table. The mapping lives in the entity table, is maintained for the whole site rather than per article, and a later change to it changes how this run reads.
  • The index is not a quality score. It reflects how often and how early models mention a name, not how good the tool is.

The bilingual scenario-grid measurement referred to above is published as the deep report. The full model panel, weight sources and scoring formula are on the methodology page.

The measurement this piece cites

AI Coding Agents

Observed 09/22/2026 · 12 models · panel completeness 93%

  • 1GitHub Copilot100.00
  • 2Cursor90.00
  • 3Windsurf (Devin Desktop)Windsurf(Devin Desktop)74.86
  • 4Claude49.29
  • 5Replit41.18
  • 6Amazon Q Developer37.98
  • 7Devin35.64
  • 8Tabnine34.39
  • 9Cline18.21
  • 10Codex15.54

This snapshot is frozen to the run above; later re-runs do not rewrite it. The ranking page always shows the most recent measurement.

View the full ranking →