Quality-first routing, end to end.
From picking models to explaining every decision — everything you need to run many LLMs as one reliable, cost-efficient endpoint.
A quality guard you control.
LinuxAir predicts each model's quality for the exact prompt, then only considers models within your tolerance of the best. Among those, the best value wins.
- Three modes — Best quality, Best value, Economy — or a custom tolerance
- Hard prompts automatically tighten the tolerance
- Uncertainty-aware: thinly measured models aren't trusted blindly
- Optional minimum-quality floor and speed preference
Pick models. Prices fill themselves in.
Connect a provider and choose from its live model list. LinuxAir fills in input and output prices and context windows from a public price feed, and refreshes them daily.
- Add several models in one go
- Local models (Ollama) are free; estimates are clearly labelled
- Models already measured by the shared catalogue go live instantly
Weak answers get caught.
When LinuxAir picks a measurably weaker model, predicts low quality, or is unsure, a judge model scores the answer. If it falls short, the request is retried once on the strongest model you haven't tried — and the result feeds back into learning.
- Automatic failover to the next-best model on provider errors
- Escalations are logged with the judge score
The weak result lowers the small model's score for this kind of prompt.
Gets better on your prompts, not a benchmark's.
Thumbs up and down — from the dashboard or the API — and implicit retries (the same prompt re-sent shortly after) update each model's per-task scores. Live evidence is blended with the probe results in proportion to how much of it exists.
Illustrative — each model has its own profile.
The cheapest call is the one you don't make.
Near-identical prompts are answered from a per-workspace cache: no provider call, no tokens, no credits, and a reply in milliseconds.
- Matches on meaning, not exact text — tune the similarity threshold
- Scoped per workspace, with a TTL and the same model settings
- Bypass per request with one flag
Caps, not surprises.
Set a daily and monthly budget for the workspace and a daily cap per API key. Routing pauses at the cap and an alert lands in your inbox well before it.
- Measured against real per-request provider cost
402 budget_exceededinstead of a runaway bill- Per-key caps for contractors and test environments
Proof, not vibes.
Build an eval set from prompts you actually sent, score every model on it, and quietly shadow-test alternatives against live traffic before you switch anything.
- Mine representative prompts from your own requests
- Judge-scored results per model and per task type
- Shadow runs happen after the real answer — latency is untouched
- Win rates and cost deltas per pairing
Sensitive data stays yours.
Mask emails, phone numbers, card numbers, national IDs, IPs, secrets and your own patterns before a prompt leaves — then put the real values back in the answer. Invite colleagues with the right level of access.
- Placeholders like
[EMAIL_1], restored in the response - Card numbers checked with Luhn to avoid false positives
- Owner, Developer and Viewer roles enforced on the server
See the savings and every decision.
The dashboard shows spend against what your strongest model would have cost, traffic share by model, a model-by-task quality heatmap and the mix of prompts you send. Every request keeps its full decision trail.
Also included.
OpenAI-compatible gateway
Chat completions with streaming, tools and JSON mode; a models endpoint; a feedback endpoint.
API keys with policies
Per-key price sensitivity and rate limits. Keys are stored as hashes and shown only once.
Model aliases
auto, la/quality, la/balanced, la/economy — or pin a model and still get failover and logging.
Credit wallet
Top up with Razorpay (UPI, cards, netbanking) or Stripe. Idempotent crediting, a clear ledger.
Playground
See the decision update live as you type, then run it for real against your providers.
Self-hostable
Run the whole platform on your own infrastructure, with local embedding and judge models if you want.
Batch API
Up to 500 requests per call, processed in the background at a reduced platform fee.
Embeddings passthrough
/v1/embeddings on the same base URL and key.
Capability-aware routing
Tool calling and JSON mode are probed per model and enforced per request.
Metrics & exports
Prometheus gauges and a full usage CSV, plus drift alerts when a model changes.
Public leaderboard
Compare models on the same probe set before you even sign up.
Stop overpaying for easy prompts.
2,000 free credits on signup. Bring your own keys, change one base URL, and see the savings on your own traffic.