From memory. For the record.

The certificate is for sale.
The rank is not.

The same question in two languages — once in Chinese, once in English — to the same panel of mainstream LLMs, on the same day, with both rankings side by side. Change the language and the list changes. Every raw response stays on file for anyone to check, including the ones that never answered.

Verified record

Pet Dryer Box / Dog Washer

PETKIT69.54/ 100

Ranked 1st of 56 candidate brands

2026-09-1312 models

Sample layout, built from the most recent evaluation — real brand, real score, real per-model ranks. A certificate number is issued at purchase; this brand has not bought one.

Behind every ranking
25
Rankings published
775
Raw model responses kept on file
12
Models queried per ranking

These aren't marketing figures — every one of those responses is stored verbatim and reprinted on the certificate page, so you can open any ranking and check the arithmetic yourself. How we compute it →

How a ranking is produced

Three steps. All three are logged, and you can check any of them.

  1. 01

    The same question, in two languages

    One question goes to every model on the panel in the same run, then the same question goes out again in the other language, on the same day. No follow-up prompts, no re-rolls, and no discarding an answer we don't like.

  2. 02

    Normalise across languages, then weight

    The answers come back in different languages and spellings. Brands are merged into one canonical name, then each model's ranking is weighted by its panel weight and aggregated into a single index.

  3. 03

    Keep every raw answer

    Each response is stored verbatim alongside its model id, weight and contribution — and reprinted in full on the verification page of every certificate.

Evaluation panel · 12 mainstream LLMs currently in use
openai/gpt-5.6-lunatencent/hy4-previewxiaomi/mimo-v2.5z-ai/glm-5.3-flash
Show 8 more models
tencent/hy3qwen/qwen3.8-maxgoogle/gemini-3.7-flashnvidia/nemotron-3-ultra-550b-a55b:freedeepseek/deepseek-v4-flashx-ai/grok-4.6anthropic/claude-sonnet-5moonshotai/kimi-k3

Hot Topics

The same method, pointed at a question people are asking right now.

Observed 2026-09-02

AI Glasses

1Rokid66.52
2XREAL64.40
3Meta59.39
4Ray-Ban52.33

2026-09-0212 models4 of 30 candidates shownon the board 2026-09-02

View full ranking →
Observed 2026-09-03

LLM API Plans

1OpenAI100.00
2Anthropic90.00
3Google46.96
4Mistral AI44.84

2026-09-0312 models4 of 29 candidates shownon the board 2026-09-03

View full ranking →

Weekly monitor

Re-measured every Monday: three questions, each asked once in Chinese and once in English, six cards in matched pairs — each card is one dated measurement.

Smart Pet Feeder / Water Dispenser (asked in Chinese)

1PETKIT89.97
2PETLIBRO71.86
3Xiaomi62.88
4PetSafe62.73

2026-09-1312 models4 of 35 candidates shown

View full ranking →
pet-hardware

Smart Pet Feeder / Water Dispenser

1PETLIBRO86.92
2PETKIT76.17
3PetSafe73.99
4WOPET64.45

2026-09-1312 models4 of 32 candidates shown

View full ranking →
pet-hardware

Smart Cat Litter Box (asked in Chinese)

1PETKIT90.71
2Litter-Robot86.95
3CATLINK77.13
4Neakasa37.80

2026-09-1312 models4 of 43 candidates shown

View full ranking →
pet-hardware

Smart Cat Litter Box

1Litter-Robot90.62
2PETKIT87.88
3CATLINK65.38
4PetSafe51.64

2026-09-1312 models4 of 44 candidates shown

View full ranking →
pet-hardware

Pet Dryer Box / Dog Washer (asked in Chinese)

1PETKIT76.81
2Homerunpet53.88
3Shernbao27.64
4SHELANDY21.14

2026-09-1312 models4 of 72 candidates shown

View full ranking →
pet-hardware

Pet Dryer Box / Dog Washer

1PETKIT69.54
2Flying Pig Grooming39.50
3Homerunpet33.17
4Bissell24.47

2026-09-1312 models4 of 55 candidates shown

View full ranking →

AI tooling

For readers who evaluate AI engineering. This group stays online and is re-measured with exactly the same method as every other ranking.

ai-tooling

Most Popular Embedding Model API Brands

1OpenAI59.50
2Cohere49.77
3Hugging Face32.41
4Voyage AI29.96

2026-08-3112 models4 of 36 candidates shown

View full ranking →
ai-tooling

Most Popular LLM Observability Platforms

1LangSmith87.93
2Weights & Biases68.32
3Arize AI64.65
4Helicone59.29

2026-08-3112 models4 of 30 candidates shown

View full ranking →
ai-tooling

Most Popular AI Model Gateway / Routing Platforms

1OpenRouter91.52
2LiteLLM85.17
3Portkey66.73
4Helicone43.59

2026-08-3112 models4 of 39 candidates shown

View full ranking →
ai-tooling

Most Popular Code Execution Sandboxes for AI Agents

1E2B86.98
2Docker75.25
3Modal69.41
4Replit51.00

2026-08-3112 models4 of 43 candidates shown

View full ranking →
ai-tooling

Most Popular Browser Automation Frameworks for AI Agents

1Playwright91.70
2Selenium79.22
3Puppeteer78.53
4Cypress41.91

2026-08-3112 models4 of 34 candidates shown

View full ranking →
ai-tooling

Most Popular Web Search APIs for AI Agents

1Google84.75
2Tavily80.79
3Bing77.03
4SerpApi76.26

2026-08-3112 models4 of 25 candidates shown

View full ranking →
ai-tooling

Most Popular LLM Application / Agent Development Frameworks

1LangChain100.00
2LlamaIndex83.76
3AutoGen61.64
4CrewAI56.02

2026-08-3112 models4 of 29 candidates shown

View full ranking →
ai-tooling

Most Popular Vector Database Brands

1Pinecone96.90
2Milvus92.75
3Weaviate76.57
4Chroma67.67

2026-08-3112 models4 of 17 candidates shown

View full ranking →