From the land that gave zero to the world

An attempt at solving the price-per-intelligence metric.

Frontier-class coding models at a fraction of direct API cost — through prompt-cache engineering, context management, and smart routing. Every technique disclosed, nothing hidden.

Three endpoints, one API

glm-5.31M context

GLM-5.3

Frontier coding intelligence, smart-routed. Heavy reasoning runs on the full model; execution on the Flash tier.

qwen-3.81M context

Qwen-3.8

The Qwen frontier endpoint with the same routing discipline — planning gets the big model, everything else stays fast.

thetaFast tier

theta

Our fast tier. Today a smart-routed ensemble; evolving into our own domain-tuned small model (research in progress).

What we're building

Inference platform

OpenAI- and Anthropic-compatible endpoints. Point Claude Code, Crush, OpenCode, or any SDK at us and ship.

Academics

A free, interactive series on how LLMs actually work — tokenization to KV-cache economics — built for students.

Research

Building toward domain-specific small models for organizations that handle sensitive data — trained where the data lives.

How we keep prices this low — openly

Prompt caching (we engineer prefixes so cache hits stay high), context management, and smart routing between model tiers. Read the full breakdown on the plans page — including the calculator that shows exactly what you'd pay at direct API rates. We may use API traffic to train our own models; you can opt out any time from your dashboard.