Z.ai's flagship for complex, long-horizon coding. Text only.
- 1M context
- 128K output
- Tools
- Reasoning
- Intelligence
- 45
- DeepSWE
- 69%
- Terminal-Bench
- 83.9%
- Input
- $1.40
- Output
- $4.40
- Cache read
- $0.26
USD per 1M tokens
Models
Every model we serve, side by side: list prices, context windows, and scores on the benchmarks we trust for agentic coding.
6 models in production
Ranked by the Artificial Analysis Intelligence Index, highest first.
Z.ai's flagship for complex, long-horizon coding. Text only.
USD per 1M tokens
Moonshot's 2.8T flagship for repository-scale agentic work, with native vision. The premium tier.
USD per 1M tokens
Z.ai's fast multimodal coder: 18B active parameters on a 1M window, at a flash price.
USD per 1M tokens
DeepSeek's latest flash model, built for fast agentic coding, with native image input.
USD per 1M tokens
The lowest price in the lineup for real agentic work. Text only.
USD per 1M tokens
Qwen3.6-35B-A3B, open weights by Qwen
A light, fast complement for the roles around your main coder: scouting, summaries, quick edits.
USD per 1M tokens
Experiments open for a short window, free with a Labs seat while they run. Seats are limited and rotate, and an experiment is not meant for production work.
MiMo-V2.6-Pro-RL, open weights by Xiaomi
Xiaomi's flagship open model, post-trained for agentic coding. Free to try while the experiment runs.
Free with a Labs seat while the experiment runs
Free to try while the experiment runs. For production work, use umans-deepseek-v4.1-flash.
Free with a Labs seat while the experiment runs
Scores as of October 3, 2026. Each model's page links every score to its source.
By Artificial Analysis
An independent composite of reasoning, knowledge, maths, coding and agentic evaluations, run by Artificial Analysis on every model the same way.
MethodologyBy Datacurve
Long-horizon software engineering: real repository tasks an agent has to finish end to end.
MethodologyBy tbench.ai
Hard, realistic tasks an agent completes in a terminal sandbox: building, debugging, configuring.
Methodology* Reported by the model's publisher on its model card. Publishers run their own harnesses, so these compare less cleanly than the independent measurements, which carry no mark. No number here is estimated: a model without one has not been measured on that benchmark yet.
Create a key once, then switch models by changing one string in your tool.
Looking for a retired model? Its history stays on the status page.