Cost estimation, model routing, and compilation across major LLM providers. PCP does not call provider APIs: all support here is deterministic catalog and routing support.
Every LLM API call costs money: and the currency is tokens.
When you send a prompt to an AI model, it doesn't read words: it reads tokens. A token is roughly ¾ of a word. The sentence "Write a blog post about AI trends" is 7 words and about ~9 tokens.
You pay for input tokens (what you send) and output tokens (what the model replies). Output tokens typically cost 3-5× more than input. That's why a long, rambling prompt that generates a long response costs significantly more than a focused one.
A developer asks an AI to review 200 lines of code. Here's the token breakdown.
"Review this authentication module for security issues. Here's the code…"
+ 200 lines of code pasted below
| Model | Tier | Relative efficiency |
|---|---|---|
| Claude Opus 5 | Top | 1× (baseline) |
| Claude Sonnet 5 | Mid | 2.5× more efficient |
| Claude Haiku 4.5 | Small | 5× more efficient |
| Gemini 2.5 Flash-Lite | Small | 58× more efficient |
Same token shape, different model price: the router makes that tradeoff visible before you commit
The insight: Most teams send every prompt to their strongest model: even simple tasks that a smaller, faster model handles equally well. The router analyzes each prompt's complexity and risk, then picks the right-sized model automatically. You get the same quality output with dramatically better token efficiency.
Two levers: use fewer tokens per call, and route each call to the right-sized model.
The multi-stage compression pipeline intelligently strips boilerplate, collapses redundancy, and removes irrelevant content: while protecting code blocks, tables, and important structure. Typical reduction: 20-40% fewer tokens per call.
Not every prompt needs the most powerful model. The router classifies your task's complexity and risk, then picks the right-sized model for the job. A simple code review doesn't need Opus-level reasoning: Haiku handles it just as well, 19× more efficiently.
Without optimizer
With optimizer
The router picks from current default models and estimates cost across the full catalog
Claude family: best for structured reasoning and XML compilation
| Model | Input | Output | Tier |
|---|---|---|---|
| Claude Haiku 4.5 | $1.00 | $5.00 | Small |
| Claude Sonnet 5 | $2.00 | $10.00 | Mid |
| Claude Opus 5 | $5.00 | $25.00 | Top |
| Claude Fable 5 | $10.00 | $50.00 | Top |
per 1M tokens
GPT family: broadest ecosystem and integration support
| Model | Input | Output | Tier |
|---|---|---|---|
| GPT-5.6 Luna | $0.10 | $0.60 | Small |
| GPT-5.6 Terra | $1.00 | $6.00 | Mid |
| GPT-5.6 Sol | $2.00 | $10.00 | Top |
| GPT-4o Mini | $0.15 | $0.60 | Legacy estimate |
| GPT-4o | $2.50 | $10.00 | Legacy estimate |
| o1 | $15.00 | $60.00 | Legacy estimate |
per 1M tokens
Gemini family: competitive pricing with strong multimodal capabilities
| Model | Input | Output | Tier |
|---|---|---|---|
| Gemini 3.7 Flash | $0.75 | $3.75 | Mid |
| Gemini 3.6 Flash | $0.75 | $3.75 | Costed |
| Gemini 3.5 Flash | $1.50 | $9.00 | Costed |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | Costed |
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | Costed |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | Small |
| Gemini 2.5 Flash | $0.30 | $2.50 | Mid |
| Gemini 2.5 Pro | $1.25 | $10.00 | Top |
per 1M tokens
Search-grounded models: auto-selected for research-intent prompts
| Model | Input | Output | Tier |
|---|---|---|---|
| Sonar | $1.00 | $1.00 | Small |
| Sonar Pro | $3.00 | $15.00 | Mid |
| Sonar Reasoning Pro | $2.00 | $8.00 | Top |
per 1M tokens
A 2-step deterministic pipeline picks the right model for every prompt. Zero LLM calls: pure rules.
Routing is shaped by 5 optimization profiles: frozen presets that bundle budget, latency, and quality preferences:
# Route a prompt to the right model pcp route "Analyze quarterly sales data trends" --json # Get cost estimates across all providers pcp cost "Build a REST API for user management" --json # Full pre-flight: classify + score + route in one call pcp preflight "Review security of auth module" --json
One prompt, three compilation targets. Pick the format that matches your LLM.
<role>Senior engineer</role> <goal>Refactor auth module</goal> <constraints> No breaking changes Keep under 200 lines </constraints>
[SYSTEM] You are a senior engineer. Your goal is to refactor the auth module. [USER] Refactor the auth module. No breaking changes. Keep under 200 lines.
## Role Senior engineer ## Goal Refactor auth module ## Constraints - No breaking changes - Keep under 200 lines
Pricing data is versioned in the engine as 2026-08. Every claimed model appears in automated tests for cost estimation and routing consistency. Prices use standard short-context text token rates where applicable and exclude cache, batch, flex, fast/priority, request, search, citation, image, audio, and vendor-specific add-on fees.
Savings are calculated against the current GPT-5.6 Terra baseline so you can see exactly how much cheaper or more expensive a routed model would be before you commit.
Model routing and cost estimation are always free (not metered). No API keys required.
Route your next prompt to the right model at the right price. Free to use, no account required.
Get started