Every model included at $1.99/seat

Your inference.
Your choice of models.

Pick any open-weight model up to 70B from the catalog — text, vision, image, video, and voice. Same API, same flat price. You choose what thinks for you.

Model catalog

Browse by what you're building

Filter the catalog by workload. Every card shows parameters and context window — and every model is included in your seat.

deepseek-r1-70b
DeepSeek
REASONING

Reasoning-first. Emits explicit chain-of-thought for math, logic, and hard multi-step problem solving.

Params70.56B
Context32k
llama-3.3-70b-instruct
Meta Llama
FLAGSHIP

The open generalist — strong reasoning, chat, and instruction-following with a large context window for RAG.

Params70.55B
Context32k
qwen3-32b
Qwen
REASONING

Multilingual reasoning tuned for coding and agentic tool use, at a fraction of a 70B footprint.

Params32.76B
Context32k
qwq-32b
Qwen
REASONING

Reasoning-tuned mid-size model that punches above its weight on structured problem solving and code.

Params32.77B
Context32k
qwen2.5-code-32b-instruct
Qwen
CODING

Code generation and completion specialist — fluent across languages, built for agents and refactors.

Params32.77B
Context32k
codestral-22b-v0.1
Mistral
CODING

Code & fill-in-the-middle model tuned for autocomplete, tests, and inline completions.

Params22.25B
Context32k
qwen2.5-coder-7b
Qwen
CODING

A smaller coding model for lower-latency autocomplete and inline suggestions.

Params7B
Context32k
command-r-35b
Cohere
GENERALIST

A well-rounded conversational generalist with strong tool use and retrieval — steerable for open-ended chat, including creative and role-play prompts.

Params34.98B
Context8k
mixtral-8x7b-instruct-v0.1
Mistral
GENERALIST

Sparse mixture-of-experts generalist — a flexible everyday choice for chat, stories, and role-play.

Params46.7B
Context32k
internlm-2.5-7b
InternLM
COMPACT

An efficient 7B model balancing reasoning quality with fast, cheap inference.

Params7.74B
Context32k
llama-3.2-1b
Meta Llama
COMPACT

Ultra-light model for edge tasks and high-volume interactive loops.

Params1B
Context32k
qwen3-0.6b
Qwen
FASTEST

A tiny, sub-billion-parameter model for classification, routing, and high-throughput loops.

Params752M
Context32k
See the full open catalog ↓

One API, every model

Switch between any model in the catalog by changing a single string — no new endpoints, no new billing tier, ever.

The open catalog

Choose your own inference: Llama, DeepSeek, Qwen, Mistral, and more — any open-weight model up to 70B, served hot. Don't see one? Request it.

Community-driven catalog

Don't see the model you want? Request it and vote. The most-requested open models get fast-tracked and go live within days.

The open catalog

Choose your own inference

Open-weight models across text, vision, image, video, and voice — every one up to 70B, all behind the same API. Switch by changing one string — no new billing tier, ever. Every model in the catalog is listed below; search to jump to one.

Model Family Category Context Size Price

Loading the live catalog…

Don't see your model?

Request it — the catalog is community-driven

New open-weight releases land fast. If a model you want is missing, request it and vote. Popular requests typically go live within days, included in every seat like everything else.

+1 vote

The most-requested model each month gets fast-tracked.

Request a model
Why choice matters

Open models, unmetered access

No lock-in

Open weights mean your prompts and pipelines aren't hostage to one vendor's roadmap. Swap models with a string change — including away from ours.

Tuned for serving

Every catalog model is optimized for throughput and reliability on our infrastructure, so you feel speed — not cold starts.

One API, every model

Ours or the catalog's — switch by changing a single string. No new endpoints, no new auth, and never a token meter anywhere.

Any model. No meter.

From our house 70B to the open catalog's best — pick the right brain for the job and never think about the token bill.