Multi-Provider Support

21 costed models across 4 providers

Cost estimation, model routing, and compilation across major LLM providers. PCP does not call provider APIs: all support here is deterministic catalog and routing support.

4
Providers
21
Costed Models
3
Output Formats

What are tokens?

Every LLM API call costs money: and the currency is tokens.

When you send a prompt to an AI model, it doesn't read words: it reads tokens. A token is roughly ¾ of a word. The sentence "Write a blog post about AI trends" is 7 words and about ~9 tokens.

Your prompt (input tokens)
Write1 a1 blog1 post1 about1 AI1 trends2
= 8 tokens input

You pay for input tokens (what you send) and output tokens (what the model replies). Output tokens typically cost 3-5× more than input. That's why a long, rambling prompt that generates a long response costs significantly more than a focused one.

See how tokens add up

A developer asks an AI to review 200 lines of code. Here's the token breakdown.

📝

The prompt

"Review this authentication module for security issues. Here's the code…" + 200 lines of code pasted below
~1,500 input tokens
~750 output tokens

Same tokens, different efficiency

Model Tier Relative efficiency
Claude Opus 5 Top 1× (baseline)
Claude Sonnet 5 Mid 2.5× more efficient
Claude Haiku 4.5 Small 5× more efficient
Gemini 2.5 Flash-Lite Small 58× more efficient

Same token shape, different model price: the router makes that tradeoff visible before you commit

The insight: Most teams send every prompt to their strongest model: even simple tasks that a smaller, faster model handles equally well. The router analyzes each prompt's complexity and risk, then picks the right-sized model automatically. You get the same quality output with dramatically better token efficiency.

More value from every token

Two levers: use fewer tokens per call, and route each call to the right-sized model.

🗜️

Smart compression

The multi-stage compression pipeline intelligently strips boilerplate, collapses redundancy, and removes irrelevant content: while protecting code blocks, tables, and important structure. Typical reduction: 20-40% fewer tokens per call.

🧭

Intelligent routing

Not every prompt needs the most powerful model. The router classifies your task's complexity and risk, then picks the right-sized model for the job. A simple code review doesn't need Opus-level reasoning: Haiku handles it just as well, 19× more efficiently.

Same code review task, optimized

Without optimizer

Model: Claude Opus 5 (manual pick)
Input: 1,500 tokens (full context)
Output: 750 tokens
Tokens used: 2,250 on a top-tier model

With optimizer

Model: Claude Haiku 4.5 (auto-routed)
Input: 1,050 tokens (30% compressed)
Output: 750 tokens
Tokens used: 1,800 on the right-sized model 20% fewer

Providers & Models

The router picks from current default models and estimates cost across the full catalog

Anthropic

Claude family: best for structured reasoning and XML compilation

Model Input Output Tier
Claude Haiku 4.5 $1.00 $5.00 Small
Claude Sonnet 5 $2.00 $10.00 Mid
Claude Opus 5 $5.00 $25.00 Top
Claude Fable 5 $10.00 $50.00 Top

per 1M tokens

OpenAI

GPT family: broadest ecosystem and integration support

Model Input Output Tier
GPT-5.6 Luna $0.10 $0.60 Small
GPT-5.6 Terra $1.00 $6.00 Mid
GPT-5.6 Sol $2.00 $10.00 Top
GPT-4o Mini $0.15 $0.60 Legacy estimate
GPT-4o $2.50 $10.00 Legacy estimate
o1 $15.00 $60.00 Legacy estimate

per 1M tokens

Google

Gemini family: competitive pricing with strong multimodal capabilities

Model Input Output Tier
Gemini 3.7 Flash $0.75 $3.75 Mid
Gemini 3.6 Flash $0.75 $3.75 Costed
Gemini 3.5 Flash $1.50 $9.00 Costed
Gemini 3.5 Flash-Lite $0.30 $2.50 Costed
Gemini 3.1 Flash-Lite $0.25 $1.50 Costed
Gemini 2.5 Flash-Lite $0.10 $0.40 Small
Gemini 2.5 Flash $0.30 $2.50 Mid
Gemini 2.5 Pro $1.25 $10.00 Top

per 1M tokens

Perplexity

Search-grounded models: auto-selected for research-intent prompts

Model Input Output Tier
Sonar $1.00 $1.00 Small
Sonar Pro $3.00 $15.00 Mid
Sonar Reasoning Pro $2.00 $8.00 Top

per 1M tokens

How Routing Works

A 2-step deterministic pipeline picks the right model for every prompt. Zero LLM calls: pure rules.

  1. Complexity + risk analysis determines a default tier. Simple factual tasks route to small, analytical tasks to mid, and multi-step reasoning or high-risk prompts to top.
  2. Budget and latency overrides shift the tier up or down. High budget sensitivity pushes toward cheaper models; low latency sensitivity allows larger models.
  3. Target preference selects the provider. Setting target to claude, openai, google, or perplexity picks that provider; generic picks Anthropic as the deterministic default.
  4. Research intent detection can route to Perplexity. A strict word-boundary regex identifies research, search, and citation prompts and recommends Perplexity's search-grounded models.

Routing is shaped by 5 optimization profiles: frozen presets that bundle budget, latency, and quality preferences:

cost_minimizer balanced quality_first creative enterprise_safe

Try it from the CLI

# Route a prompt to the right model
pcp route "Analyze quarterly sales data trends" --json

# Get cost estimates across all providers
pcp cost "Build a REST API for user management" --json

# Full pre-flight: classify + score + route in one call
pcp preflight "Review security of auth module" --json

Output Formats

One prompt, three compilation targets. Pick the format that matches your LLM.

Claude XML

<role>Senior engineer</role>
<goal>Refactor auth module</goal>
<constraints>
  No breaking changes
  Keep under 200 lines
</constraints>

OpenAI

[SYSTEM]
You are a senior engineer.
Your goal is to refactor the
auth module.

[USER]
Refactor the auth module.
No breaking changes.
Keep under 200 lines.

Generic Markdown

## Role
Senior engineer

## Goal
Refactor auth module

## Constraints
- No breaking changes
- Keep under 200 lines

Pricing You Can Trust

Pricing data is versioned in the engine as 2026-08. Every claimed model appears in automated tests for cost estimation and routing consistency. Prices use standard short-context text token rates where applicable and exclude cache, batch, flex, fast/priority, request, search, citation, image, audio, and vendor-specific add-on fees.

Savings are calculated against the current GPT-5.6 Terra baseline so you can see exactly how much cheaper or more expensive a routed model would be before you commit.

Model routing and cost estimation are always free (not metered). No API keys required.

Try the model router

Route your next prompt to the right model at the right price. Free to use, no account required.

Get started