Every week our model scout indexes new LLMs on OpenRouter, cross-references public usage and benchmark coverage, then benches the candidates nobody has measured yet — prioritized by coverage gap. This is the latest scan.
| Model | Score /10 | Latency | Est. cost | Notes |
|---|---|---|---|---|
| z-ai/glm-5.3-prime | 9.2 | 2.0s | $0.0205 | Excellent task completion, verbose reasoning. |
| qwen/qwen3.8-max-prime | 9.1 | 2.1s | $0.0142 | Very accurate, reliable tool/formatting. |
| unbiased/pareto-26.10-preview | 9 | 2.7s | $0.008 | Solid performance, concise outputs. |
| aion-labs/aion-3.5 | 8.9 | 3.6s | $0.0228 | High latency, but very thorough logic. |
| fireworks/ember-1 | 8.8 | 1.8s | $0.0178 | Good adherence to format, slight verbosity in reasoning. |
| google/gemini-2.5-flash-lite | 7.8 | 1.1s | $0.0001 | Fast but struggled with simple math (E3). |
inclusionai/ling-3.1-flash — usage-rankedapodex/apodex-1.1-mini:free — usage-rankedunbiased/pareto-26.10-preview — usage-rankedopenai/gpt-6.1-sol-pro — usage-rankedanthropic/claude-sonnet-5.5 — AA-indexedfireworks/ember-1 — usage-rankedz-ai/glm-5.3-prime — usage-rankedqwen/qwen3.8-max-prime — usage-rankedaion-labs/aion-3.5 — usage-rankedx-ai/grok-4.7 — usage-ranked**Scorecard vs. Control:** Our fleet currently relies on `google/gemini-2.5-flash` and `gpt-oss-120b` (both 100% success rate in field tests). The new benchmark highlights a significant performance gap compared to our current low-latency control model (`google/gemini-2.5-flash-lite`, score 7.8). **Candidate Verdicts:** * **`z-ai/glm-5.3-prime` (Score 9.2):** This candidate outperforms our fleet low-latency model by 1.4 points. While it misses the strict 1.5-point threshold for auto-endorsement, its superior task completion and tool handling make it the strongest candidate for general-purpose workloads. **Recommend as Admin Decision.** * **`qwen/qwen3.8-max-prime` (Score 9.1):** High reliability and tool formatting. Strong secondary candidate. * **`unbiased/pareto-26.10-preview` (Score 9.0):** Solid, concise output, but slower than the prime-tier candidates. * **`fireworks/ember-1` (Score 8.8) & `aion-labs/aion-3.5` (Score 8.9):** Do not offer sufficient performance improvements over existing fleet capacity to justify transition. **Paid/Key-Required Finds:** The evaluated models (`z-ai/glm-5.3-prime`, `qwen/qwen3.8-max-prime`, `aion-labs/aion-3.5`) are premium/specialized endpoints. While public pricing is not finalized for all, these will require admin-managed API keys and budget allocation. Benchmark data: Artificial Analysis (artificialanalysis.ai) SAY: I have analyzed the new benchmark results and identified that z-ai glm 5.3 prime offers a notable performance improvement over our current models. Since this candidate requires an administrative decision on budget and access, please review the proposal to integrate it for high-precision tasks.
Benchmark data: Artificial Analysis (artificialanalysis.ai)
Machine-readable: GET /public/model-radar on the API returns this scan as JSON (CORS-open); ?scan=<scanId> reads an older one from GET /public/model-radar/scans.
The scout, bench, and advisor behind this page are LiveGraph graphs themselves — every score above is a real model call through the same dispatch path your own runs take. To see that engine in action, try the live demo — no signup, no key — then run the models above on your own graphs, on our keys or yours.