From the land that gave zero to the world
An attempt at solving the price-per-intelligence metric.
Frontier-class coding models at a fraction of direct API cost — through prompt-cache engineering, context management, and smart routing. Every technique disclosed, nothing hidden.
Three endpoints, one API
glm-5.31M contextGLM-5.3
Frontier coding intelligence, smart-routed. Heavy reasoning runs on the full model; execution on the Flash tier.
qwen-3.81M contextQwen-3.8
The Qwen frontier endpoint with the same routing discipline — planning gets the big model, everything else stays fast.
thetaFast tiertheta
Our fast tier. Today a smart-routed ensemble; evolving into our own domain-tuned small model (research in progress).
What we're building
Inference platform
OpenAI- and Anthropic-compatible endpoints. Point Claude Code, Crush, OpenCode, or any SDK at us and ship.
Academics
A free, interactive series on how LLMs actually work — tokenization to KV-cache economics — built for students.
Research
Building toward domain-specific small models for organizations that handle sensitive data — trained where the data lives.
How we keep prices this low — openly
Prompt caching (we engineer prefixes so cache hits stay high), context management, and smart routing between model tiers. Read the full breakdown on the plans page — including the calculator that shows exactly what you'd pay at direct API rates. We may use API traffic to train our own models; you can opt out any time from your dashboard.