Your inference.
Your choice of models.
Pick any open-weight model up to 70B from the catalog — text, vision, image, video, and voice. Same API, same flat price. You choose what thinks for you.
Browse by what you're building
Filter the catalog by workload. Every card shows parameters and context window — and every model is included in your seat.
Reasoning-first. Emits explicit chain-of-thought for math, logic, and hard multi-step problem solving.
The open generalist — strong reasoning, chat, and instruction-following with a large context window for RAG.
Multilingual reasoning tuned for coding and agentic tool use, at a fraction of a 70B footprint.
Reasoning-tuned mid-size model that punches above its weight on structured problem solving and code.
Code generation and completion specialist — fluent across languages, built for agents and refactors.
Code & fill-in-the-middle model tuned for autocomplete, tests, and inline completions.
A smaller coding model for lower-latency autocomplete and inline suggestions.
A well-rounded conversational generalist with strong tool use and retrieval — steerable for open-ended chat, including creative and role-play prompts.
Sparse mixture-of-experts generalist — a flexible everyday choice for chat, stories, and role-play.
An efficient 7B model balancing reasoning quality with fast, cheap inference.
Ultra-light model for edge tasks and high-volume interactive loops.
A tiny, sub-billion-parameter model for classification, routing, and high-throughput loops.
One API, every model
Switch between any model in the catalog by changing a single string — no new endpoints, no new billing tier, ever.
The open catalog
Choose your own inference: Llama, DeepSeek, Qwen, Mistral, and more — any open-weight model up to 70B, served hot. Don't see one? Request it.
Community-driven catalog
Don't see the model you want? Request it and vote. The most-requested open models get fast-tracked and go live within days.
Choose your own inference
Open-weight models across text, vision, image, video, and voice — every one up to 70B, all behind the same API. Switch by changing one string — no new billing tier, ever. Every model in the catalog is listed below; search to jump to one.
| Model | Family | Category | Context | Size | Price |
|---|
Loading the live catalog…
Request it — the catalog is community-driven
New open-weight releases land fast. If a model you want is missing, request it and vote. Popular requests typically go live within days, included in every seat like everything else.
Open models, unmetered access
No lock-in
Open weights mean your prompts and pipelines aren't hostage to one vendor's roadmap. Swap models with a string change — including away from ours.
Tuned for serving
Every catalog model is optimized for throughput and reliability on our infrastructure, so you feel speed — not cold starts.
One API, every model
Ours or the catalog's — switch by changing a single string. No new endpoints, no new auth, and never a token meter anywhere.
Any model. No meter.
From our house 70B to the open catalog's best — pick the right brain for the job and never think about the token bill.