439 models from every major provider

Your AI is costing more than you realize

Most teams are overpaying — running premium models where a cheaper one matches quality, with no visibility into spend or governance. See your own exposure in seconds.

OpenAIGPT-5.6 Sol·
OpenAIGPT-5.6 Luna·
OpenAIGPT-5.6 Terra·
OpenAIGPT-5.5·
OpenAIGPT-5.4 mini·
AnthropicClaude Fable 5·
AnthropicClaude Opus 4.8·
AnthropicClaude Sonnet 5·
AnthropicClaude 4.5 Haiku·
GoogleGemini 3.5 Flash·
GoogleGemini 3.1 Pro·
GoogleGemini 3.1 Flash-Lite·
xAIGrok 4.5·
xAIGrok 4.3·
Z AIGLM-5.2·
DeepSeekDeepSeek V4 Pro·
DeepSeekDeepSeek V4 Flash·
MiniMaxMiniMax-M3·
MoonshotKimi K2.6·
MoonshotKimi K3·
AlibabaQwen3.7 Max·
MetaLlama 4 Maverick·
OpenAIGPT-5.6 Sol·
OpenAIGPT-5.6 Luna·
OpenAIGPT-5.6 Terra·
OpenAIGPT-5.5·
OpenAIGPT-5.4 mini·
AnthropicClaude Fable 5·
AnthropicClaude Opus 4.8·
AnthropicClaude Sonnet 5·
AnthropicClaude 4.5 Haiku·
GoogleGemini 3.5 Flash·
GoogleGemini 3.1 Pro·
GoogleGemini 3.1 Flash-Lite·
xAIGrok 4.5·
xAIGrok 4.3·
Z AIGLM-5.2·
DeepSeekDeepSeek V4 Pro·
DeepSeekDeepSeek V4 Flash·
MiniMaxMiniMax-M3·
MoonshotKimi K2.6·
MoonshotKimi K3·
AlibabaQwen3.7 Max·
MetaLlama 4 Maverick·

What will it cost?

439 models · 63 providers · 70 updated recently · Synced Sep 15, 2026, 6:00 PM

Already using a model?

See how much you could save by switching to a cheaper alternative via Xilos Cloud.

QUICK START

Featured Scenarios

Pre-built usage patterns to quickly estimate costs for common applications.

Understand how AI models actually work

The fundamentals — no fluff, so you can make an informed decision.

What are tokens?

Models don't read words — they read tokens, small chunks of text (about 4 characters, or ¾ of a word). Providers bill per million tokens, so the length of your prompts and responses directly sets your cost. Simple requests rarely need the most expensive model.

What is routing?

Routing sends each request to the model best suited to that task — a fast, cheap model for simple queries, a stronger one for hard problems. Instead of paying one premium model for everything, you pay each model only for the work it's good at. Xilos does this automatically at request time.

What happens to my prompt after I hit send?

Your prompt travels to a provider, where a model generates a response token by token. Everything in it — context, examples, data — is sent as-is to that provider. Having visibility into what's sent, what it costs, and where it goes is the layer most teams are missing.

What is private data?

Names, emails, phone numbers, credentials, source code, financials — anything that shouldn't leave your control. If it's in your prompt, a public API sees it. Xilos detects and blocks PII before it reaches a provider, and can route sensitive work to governed or self-hosted models.

XILOS ROUTING EXAMPLE

The best model depends on the task. Xilos chooses at runtime.

One workload, distributed across models to balance quality, speed, and estimated cost — instead of paying a single premium model for everything.

One-model setup

Claude Opus 4.8

Estimated monthly cost

$6,850

Customer conversation

300M in · 60M out tokens/mo

3,000/mo

same model

Intent classification

60M in · 3M out tokens/mo

375/mo

same model

Knowledge retrieval

150M in · 40M out tokens/mo

1,750/mo

same model

Complex escalation

40M in · 40M out tokens/mo

1,200/mo

same model

Sensitive data handling

30M in · 15M out tokens/mo

525/mo

same model

With Xilos routing

5 models, chosen per task

Estimated monthly cost

$2,448

Customer conversation

GPT-5.6 Luna (General-purpose)

$540/mo

Intent classification

Gemini 3.1 Flash-Lite (Fast & inexpensive)

$5/mo

Knowledge retrieval

Gemini 3.1 Pro (Long-context)

$388/mo

Complex escalation

Claude Opus 4.8 (Strong reasoning)

$1,200/mo

Sensitive data handling

Claude Sonnet 5 (Governed / approved)

$315/mo

Estimated $4,402/mo saved (~64%) without changing output quality
How this estimate works

Model pricesare Spotlight's current per-million-token rates from the public model data on this site — not coupon pricing.

Token volumes are illustrative monthly assumptions per task (shown next to each row) and are meant to be representative, not a forecast.

Routing assumes each task is sent to the model class listed; Xilos classifies and routes at request time with negligible per-request overhead.

Actual savings depend on your exact workload, volumes, and routing rules — results will vary.

Popular Comparisons

Side-by-side analysis of leading models.