OmniRoute
Free MIT AI gateway: one OpenAI-compatible endpoint across 359 providers and 1200+ models, with quota-aware fallback, compression, and MCP tools.

Dhanji Bhagat
Founder, Emiote
Fully hosted platform. Automated backups and SLA.
$0 to start on free tiers; paid provider usage billed per token by each provider
Private compute. Zero seat taxes; team runs ops.
$0 local (SQLite on your machine) or $6-12/mo VPS estimate for the base Docker profile plus model spend
OmniRoute is a free MIT-licensed AI gateway from diegosouzapw. It exposes hundreds of providers and language models through one OpenAI-compatible endpoint on your machine. Quota-aware combos add automatic fallback across free and paid tiers, and a dashboard tracks budgets, keys, and usage.
Scope and currency
This is an architecture evaluation, not a deployment diary. On 2026-09-22 we reviewed the public repository (diegosouzapw/OmniRoute, release v3.8.51), its README, package.json, Dockerfile set, docker-compose.yml, LICENSE, the public/ brand assets, the skills/ and docs/ trees, and the live site omniroute.online. We have not run OmniRoute in production. Stars (69.1k) and forks (9.8k) are observed values on that date and will move, as will provider counts. Editorial review: 2026-09-22.
Contrast with our deployment notes (self-hosted Postgres): those record lived ops. This page is what the repository and docs imply. Same posture as Pascal Editor.
1. What It Replaces and Why It Matters
Coding agents burn tokens across many vendors. The default setup is one API key per provider, one SDK per shape, and one quota surprise per month. When a free tier dries up mid-session, the agent stops and the human re-plumbs keys. Four products cover the practical alternatives today: OpenRouter for hosted aggregation, LiteLLM proxy for self-hosted routing, Portkey for gateway controls with observability, and LibreChat for a self-hosted chat front end over many back ends.
OmniRoute takes the self-hosted aggregator position and aims it at agent CLIs. One endpoint (/v1 on your machine) speaks OpenAI-compatible chat, responses, embeddings, OCR, and audio shapes to every connected tool: Claude Code, Codex CLI, Cursor, OpenCode, Cline, Copilot CLI, and around twenty more named integrations. Behind that endpoint sit combos (ordered model chains with automatic fallback), quota-aware scheduling across pooled keys, and a compression stack that rewrites prompts to spend fewer tokens. The pitch is continuity: coding sessions keep running on free and cheap capacity instead of halting at the first exhausted key.
Who should consider switching: solo developers and small teams whose agent bills come from several vendors, operators who already self-host on a VPS, and tinkerers who want every model behind one local address. Who should stay: teams that need vendor support contracts, regulated shops that must keep inference inside a named region with audit attestations, and anyone whose tools only speak a proprietary first-party API the gateway cannot front.
2. Architecture and Container Topology
OmniRoute is a TypeScript monorepo (Next.js 16, React 19, Tailwind 4, Zod 4) with a Node runtime floor of 22.22 or 24 and an npm-published CLI (bin/omniroute.mjs). State rests on better-sqlite3 in WAL mode plus LowDB, with 122 modules and 178 migrations claimed in the README, SQLite FTS5 with int8 vectors and typed decay for memory, and an optional Qdrant sidecar past one million points. Streaming runs over SSE plus a WebSocket channel. The request path looks like this:
flowchart TD
subgraph Tools["Agent Tools"]
CLI["Coding CLIs<br/>(Claude, Codex, OpenCode, Cursor)"]
API["Any OpenAI-compatible client<br/>(Base URL /v1)"]
MCPc["MCP / A2A consumers<br/>(110 tools, 6 skills)"]
end
subgraph Gateway["OmniRoute Gateway (port 20128/20129)"]
Router["Combo router<br/>(19 strategies, auto scoring)"]
Compress["Compression stack<br/>(12 engines, RTK + Caveman)"]
Quota["Quota scheduler<br/>(pooled keys, telemetry)"]
Guard["Guardrails + resilience<br/>(breakers, cooldowns, lockouts)"]
end
subgraph State["State and Sidecars"]
SQLite["SQLite WAL + LowDB<br/>(./data or /app/data)"]
Redis["Redis 8 Alpine<br/>(rate limiter backend)"]
Opt["Opt-in sidecars<br/>(Qdrant, Bifrost, CLIProxyAPI)"]
end
subgraph Providers["359 Providers"]
Free["Free tiers<br/>(150+ catalog-marked)"]
Paid["Paid APIs<br/>(per-token billing)"]
Local["Local and CLI-backed<br/>(Ollama-style, CLIs)"]
end
CLI --> Router
API --> Router
MCPc --> Router
Router --> Compress
Compress --> Quota
Quota --> Guard
Guard --> Free
Guard --> Paid
Guard --> Local
Router --> SQLite
Router --> Redis
Router --> Opt
Routing depth is the differentiator. Nineteen strategies cover priority, fill-first, weighted, round-robin, least-used, cost-optimized, headroom, reset-window, reset-aware, context-relay, context-optimized, cache-optimized, and several named auto modes, with a 16-factor scoring pass behind the auto combo. Resilience runs a four-tier cascade (subscription, then API, then cheap, then free) guarded by three layers: provider circuit breakers, connection cooldowns, and model lockouts. Compression stacks twelve engines (Session-Dedup, CCR, Lite, RTK, Responses Tool Output, Headroom, Relevance, Caveman, Aggressive, LLMLingua-2, Ultra, OmniGlyph) with presets from Lite around 15 percent to stacked RTK plus Caveman at 78 to 95 percent. Those percentages are project-claimed ranges from the repository, repeated here as claims rather than measured results.
The fallback sequence an agent experiences is mechanical:
sequenceDiagram
autonumber
participant Agent as Coding Agent
participant GW as OmniRoute /v1
participant Q as Quota Scheduler
participant P1 as Primary Provider
participant P2 as Fallback Provider
Agent->>GW: chat/completions (model: auto)
GW->>Q: Select candidate by strategy + headroom
Q-->>GW: Provider A with remaining quota
GW->>P1: Forward with compressed prompt
P1-->>GW: 429 quota exhausted
GW->>Q: Mark cooldown, pick next candidate
Q-->>GW: Provider B
GW->>P2: Retry same payload
P2-->>GW: 200 with usage
GW-->>Agent: Response + X-OmniRoute cost headers
Agent integration runs three ways: OpenAI-compatible base URL plus dashboard-issued key for any tool, omniroute run and omniroute configure helpers for named CLIs, and native MCP (omniroute --mcp, /api/mcp/stream, /api/mcp/sse) plus A2A JSON-RPC with SSE at /.well-known/agent.json. TLS stealth (wreq-js JA3/JA4 fingerprints, three-level proxy support) exists for providers that fingerprint clients. The dashboard serves quota telemetry, free-tier budgets, provider toggles, key management, and compression controls from the same process.
3. Visual Tour and Interface Workflow
Interface proof comes from two authentic sources: a live capture of omniroute.online taken 2026-09-22, and the official dashboard screenshot shipped in the repository docs.

The homepage leads with router positioning (352 providers on the live page against 359 in the release notes, a sign these counts move), OpenAI compatibility, automatic fallback, and the star count beside a self-host CTA. A banner notes the project has joined Cheaper Inference with a hosted option alongside self-hosting, a commercial direction worth watching.

The dashboard is the working surface: a provider grid with connection states, a dedicated free-tier section (some entries need only a signup, others need no credentials at all), an OAuth section for CLI-backed logins, quota tracking, combo groups, compression engines, CLI tool wiring, endpoints, webhooks, and proxy controls in the rail. This matches the architecture above: catalog, keys, budgets, and routing policy in one place.
The operator workflow follows one loop: connect providers in the dashboard, group models into combos with fallback order, point tools at the local base URL with a dashboard key, and watch quota telemetry decide the next hop. Free-tier standing is visible at /dashboard/free-tiers with used and remaining figures per pool.
4. Total Cost of Ownership (TCO)
TCO splits into infrastructure you operate and tokens you consume. The license is MIT, so software cost is zero in every column.
| Dimension | Commercial / Hosted Option | Self-Hosted / Local |
|---|---|---|
| License | N/A, aggregator seats vary by vendor | $0 (MIT, diegosouzapw) |
| Compute | Included in hosted price | $0 local; $6-12/mo VPS estimate for the base profile |
| Storage | Included, limits undisclosed | SQLite file plus Redis volume; small at solo scale |
| API / Model Usage | Per-token billing per provider | Same billing, plus free tiers first: ~1.62B/mo steady project estimate, ~2.22B first month with signup credits |
| Maintenance | Vendor-managed | Your responsibility: updates, key rotation, combo retuning as tiers change |
| Data Ownership | Aggregator sees request metadata | Your disk (./data or /app/data); prompts still reach each provider |
| Annual Cost | Provider spend plus any aggregator fee | Rough operating estimate: $0 local or $72-144/yr VPS, plus token spend |
| Operational Burden | Low | Low for one user, Medium once many providers and combos are live |
Annual math is monthly cost times 12. The token headline needs care: it is pool-deduped math from the project catalog (17 recurring pools with published monthly budgets plus five per-model Groq caps, shared pools counted once), re-audited every two weeks, and it moves both ways as providers open or close tiers. Quotas behind regional identity checks (ModelScope, around 6M) sit outside the headline, and 13 providers carry avoid marks in the terms-risk catalog. Treat the headline as a planning input with a date stamp, never as guaranteed capacity. Engineering labor is separate: expect recurring small sessions retuning combos as free tiers shift rather than a single setup afternoon. Self-hosting removes seat fees while infrastructure, key management, and your update time stay real.
5. The Bad: What to Know Before Adopting
Every limitation below comes from the repository, its compose file, or the live site. None is invented.
Free capacity moves both ways. The project states plainly that ended tiers drop the headline and new tiers raise it. A combo tuned this month can underperform next month with no code change on your side. Budget dashboards accordingly and pin expectations to dated snapshots.
Counts disagree across surfaces. The release notes claim 359 providers, the growth table shows 357 at v3.8.50, and the live site showed 352 on capture day. Denominators differ (registered, catalog-marked, chat pairs, raw IDs), and the project documents several of them. Quote one figure with its source and date instead of mixing them.
Compression percentages are vendor claims. The 15 to 95 percent range with an 89 percent average comes from project benchmarks and eval harnesses in the repo. No independent measurement is cited here because none was performed in this review. Compression also rewrites prompts, so quality-sensitive workloads need before and after evals on your own tasks.
The default bind posture needs attention. Ports default to loopback, which is good, but .env.example ships REQUIRE_API_KEY=false, and the compose comments warn that binding 0.0.0.0 exposes an anonymous /v1 proxy to the network. Redis ships without a password behind the same loopback default. Confirm key enforcement or a reverse proxy with auth before any non-loopback bind.
The CLI profile mounts the Docker socket. The compose file documents this plainly: the socket lets the in-container updater recreate the stack, and equally lets any container process drive your host daemon. The file restricts this profile to trusted single-tenant workstations with loopback-only ports. Follow that restriction exactly.
Web-cookie providers need the heavy profile. The base image has no Chromium, so cookie-backed providers fail until the web profile (Playwright plus browser) is up. Cookie auth also means session expiry and re-login chores that API keys avoid.
Scale of the codebase is its own cost. Around nine thousand commits, thousands of test files, dozens of quality gates, native modules, and a narrow Node window (22.22 to 23, or 24 to 27) make upgrades heavier than the one-container image suggests. Optional sidecars (Qdrant, Bifrost, CLIProxyAPI, codex-app-server) each add their own storage and credentials to back up.
Commercial direction just shifted. The live site now presents Cheaper Inference hosting beside self-hosting. That can be healthy (funded maintenance), and it can reorder priorities toward the hosted path. Watch release notes for where features land first before committing infrastructure.
Security notes for the gateway service
The repository shows serious security scaffolding: AES-256-GCM secret handling, OAuth2 PKCE plus JWT plus scoped MCP keys, DOMPurify, and checked-in configs for secret scanning, container scanning, and workflow hardening. The remaining threat model is operational and credential-shaped: provider OAuth tokens and API keys concentrate in the gateway data volume and sidecar stores, the CDP proxy token guards a live browser session and must stay secret, the codex capability token file ships with 600 permissions that backups must preserve, and any model endpoint an agent can reach deserves the same posture as a code-execution service. Keep binds on loopback, rotate dashboard and provider keys on a schedule, and never paste keys into issues, transcripts, or screenshots.
6. Quickstart and Deployment
Shortest verified paths, quoted from current project sources. Copy the example environment first when using compose:
cp .env.example .env
Global install and first run:
npm install -g omniroute
omniroute
The dashboard answers at http://localhost:20128 and the API at http://localhost:20128/v1. Verify the endpoint:
curl http://localhost:20128/v1/models -H "Authorization: Bearer YOUR_KEY"
Zero-config first call with automatic model selection:
curl http://localhost:20128/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"auto","messages":[{"role":"user","content":"Hello!"}]}'
Docker base profile (loopback ports 20128 dashboard, 20129 API, 20132 live socket, data in a named volume):
docker run -d --name omniroute --restart unless-stopped --stop-timeout 40 -p 127.0.0.1:20128:20128 -v omniroute-data:/app/data diegosouzapw/omniroute:latest
Compose with profiles (base, web for cookie providers, cli, host, memory for Qdrant, bifrost, cliproxyapi):
docker compose --profile base up -d
Source development requires Node 22.22+ or 24.x, then npm install with PORT=20128 npm run dev. Connect any tool with base URL http://localhost:20128/v1, a dashboard-issued key, and model auto. Named helpers (omniroute run claude, omniroute configure codex) cover the common CLIs without hand-written JSON.
7. Recommendation: Who Should Use This
Good fit: developers running several coding CLIs who want one local address, one key ring, and fallback that survives exhausted free tiers. The quota telemetry plus deduped budget math is the honest part competitors often skip, and the MCP plus A2A surface makes the gateway legible to agents instead of humans only.
Bad fit: teams that need contractual support, regional processing guarantees, or a managed compliance story. Also a poor fit for anyone unwilling to rotate keys and retune combos as free tiers churn. The moving headline is a feature for tinkerers and a liability for fixed budgets.
Direct advice: adopt OmniRoute as your routing and fallback layer while keeping spend alerts at each provider. Do not adopt it as a capacity guarantee. The catalog earns the first role on evidence; no aggregator can promise the second on free tiers.
8. ReframeHub Insight
The deeper lesson is metering shared pools honestly instead of summing marketing maximums. Most free-tier roundups add every headline figure and print a fantasy total. This project dedupes shared pools, separates first-month credits from steady state, quarantines identity-gated quotas, and marks providers to avoid, then recomputes on a schedule and admits the number falls. That pattern (dedupe, date-stamp, separate one-offs, publish declines) ports to any cost comparison where vendors overlap: cloud credits, API trials, and SaaS trial tiers all inflate the same way.
9. What I Would Change
Opinion, stated as opinion. I would freeze one canonical provider count with a documented denominator instead of letting four figures circulate. I would ship REQUIRE_API_KEY=true as the default with a guided first-run key ceremony, since secure defaults beat warning comments. I would split the web-cookie stack into its own documentedpliance boundary with session-expiry runbooks. None of these questions the core design, which reads as careful and unusually candid about its own limits.
Evaluating self-hosted model routing or agent token spend?
Reframe ($199) audits your gateway topology, fallback chains, and free-tier math for OmniRoute against hosted aggregator seats. Diagnosis only.
Fixed $199 fee · 100% vendor-neutral review · 3-day delivery guarantee
