Models

Pick the right model.

Every model we serve, side by side: list prices, context windows, and scores on the benchmarks we trust for agentic coding.

6 models in production

Ranked by the Artificial Analysis Intelligence Index, highest first.

In Labs

Experiments open for a short window, free with a Labs seat while they run. Seats are limited and rotate, and an experiment is not meant for production work.

Get a seat

About the scores

Scores as of October 3, 2026. Each model's page links every score to its source.

Artificial Analysis Intelligence Index

By Artificial Analysis

An independent composite of reasoning, knowledge, maths, coding and agentic evaluations, run by Artificial Analysis on every model the same way.

Methodology

DeepSWE

By Datacurve

Long-horizon software engineering: real repository tasks an agent has to finish end to end.

Methodology

Terminal-Bench 2.1

By tbench.ai

Hard, realistic tasks an agent completes in a terminal sandbox: building, debugging, configuring.

Methodology

* Reported by the model's publisher on its model card. Publishers run their own harnesses, so these compare less cleanly than the independent measurements, which carry no mark. No number here is estimated: a model without one has not been measured on that benchmark yet.

One key, every model.

Create a key once, then switch models by changing one string in your tool.

Looking for a retired model? Its history stays on the status page.

Models · Umans AI