Skip to content
LiveGraph

Model Radar

Every week our model scout indexes new LLMs on OpenRouter, cross-references public usage and benchmark coverage, then benches the candidates nobody has measured yet — prioritized by coverage gap. This is the latest scan.

Scanned 10/4/2026 · verdict: upgrade candidate surfaced

Bench scorecard

ModelScore /10LatencyEst. costNotes
z-ai/glm-5.3-prime9.22.0s$0.0205Excellent task completion, verbose reasoning.
qwen/qwen3.8-max-prime9.12.1s$0.0142Very accurate, reliable tool/formatting.
unbiased/pareto-26.10-preview92.7s$0.008Solid performance, concise outputs.
aion-labs/aion-3.58.93.6s$0.0228High latency, but very thorough logic.
fireworks/ember-18.81.8s$0.0178Good adherence to format, slight verbosity in reasoning.
google/gemini-2.5-flash-lite7.81.1s$0.0001Fast but struggled with simple math (E3).

New free-tier models

  • inclusionai/ling-3.1-flash — usage-ranked
  • apodex/apodex-1.1-mini:free — usage-ranked

New paid models worth watching

  • unbiased/pareto-26.10-preview — usage-ranked
  • openai/gpt-6.1-sol-pro — usage-ranked
  • anthropic/claude-sonnet-5.5 — AA-indexed
  • fireworks/ember-1 — usage-ranked
  • z-ai/glm-5.3-prime — usage-ranked
  • qwen/qwen3.8-max-prime — usage-ranked
  • aion-labs/aion-3.5 — usage-ranked
  • x-ai/grok-4.7 — usage-ranked

Advisor notes

**Scorecard vs. Control:**
Our fleet currently relies on `google/gemini-2.5-flash` and `gpt-oss-120b` (both 100% success rate in field tests). The new benchmark highlights a significant performance gap compared to our current low-latency control model (`google/gemini-2.5-flash-lite`, score 7.8).

**Candidate Verdicts:**
*   **`z-ai/glm-5.3-prime` (Score 9.2):** This candidate outperforms our fleet low-latency model by 1.4 points. While it misses the strict 1.5-point threshold for auto-endorsement, its superior task completion and tool handling make it the strongest candidate for general-purpose workloads. **Recommend as Admin Decision.**
*   **`qwen/qwen3.8-max-prime` (Score 9.1):** High reliability and tool formatting. Strong secondary candidate.
*   **`unbiased/pareto-26.10-preview` (Score 9.0):** Solid, concise output, but slower than the prime-tier candidates.
*   **`fireworks/ember-1` (Score 8.8) & `aion-labs/aion-3.5` (Score 8.9):** Do not offer sufficient performance improvements over existing fleet capacity to justify transition.

**Paid/Key-Required Finds:**
The evaluated models (`z-ai/glm-5.3-prime`, `qwen/qwen3.8-max-prime`, `aion-labs/aion-3.5`) are premium/specialized endpoints. While public pricing is not finalized for all, these will require admin-managed API keys and budget allocation.

Benchmark data: Artificial Analysis (artificialanalysis.ai)

SAY: I have analyzed the new benchmark results and identified that z-ai glm 5.3 prime offers a notable performance improvement over our current models. Since this candidate requires an administrative decision on budget and access, please review the proposal to integrate it for high-precision tasks.

Benchmark data: Artificial Analysis (artificialanalysis.ai)

Machine-readable: GET /public/model-radar on the API returns this scan as JSON (CORS-open); ?scan=<scanId> reads an older one from GET /public/model-radar/scans.

The scout, bench, and advisor behind this page are LiveGraph graphs themselves — every score above is a real model call through the same dispatch path your own runs take. To see that engine in action, try the live demo — no signup, no key — then run the models above on your own graphs, on our keys or yours.

LiveGraph

Model-agnostic agent orchestration on a live canvas. Hosted at livegraph.ai.

Product

  • Live demo
  • Pricing
  • Security
  • Automations
  • Compare
  • Migrate from Flowise
  • Migrate from Agent Builder
  • Sign up

Resources

  • Blog
  • Docs
  • Changelog
  • Model Radar
  • Status
  • RSS

Agents

  • Agent guide
  • llms.txt
  • llms-full.txt
  • MCP integration

Support

  • [email protected]

Legal

  • Privacy
  • Terms
© 2026 LiveGraphSite by Canweb Ltd.