Product · AI API gateway
One API, one key and one prepaid balance for text, image and video models from every major provider.
Introduction
Teams building AI products end up integrating provider after provider: a different SDK, key, bill and rate limit for every model they want to try. Routehook puts one OpenAI-compatible endpoint in front of all of them. Chat, embeddings, images and video go through a single base URL; long-running generations run as queued jobs with webhooks and failover; and every request is charged from a prepaid balance on an append-only ledger, so a customer can never be billed twice or spend past their balance.
Details
The console a customer runs their AI traffic from, rebuilt here in code from its own screens. Pick a page, or click the sidebar.
One catalogue of text, image, video, audio, embedding and rerank models, each with the customer’s own price and how much they have used it in the last 30 days.
Recreated from the dashboard’s interface. Labels and columns are Routehook’s own; the account, requests, prices and numbers are samples, and the images are CC0 photographs standing in for model output.
The gateway behaviour behind the dashboard, at full size.
Gateway
The gateway speaks three wire formats, streaming or not, so an existing OpenAI or Anthropic client works by swapping two values.
Response
Every response carries what it cost, which provider answered, how many attempts it took and whether a fallback stepped in.
Ledger
The worst-case cost is held before a request leaves, and a database constraint refuses any hold the balance cannot cover. Afterwards the hold settles to the real cost.
Queue
An image or video request returns a job at once. It can be polled, followed or cancelled, and a signed webhook fires when it finishes — retried with backoff if the receiver is down.
Limits
One atomic script in Redis keeps a bucket per key, so every API process sees the same count.
Errors
Every failure maps to one stable code, so a client can tell a bad key from an empty balance from a provider that is down.
Notifications
Each alert can reach the dashboard, the inbox, both or neither, and the choice saves as it is made.
One /v1 endpoint for chat, completions, embeddings, rerank, audio, images and video, with adapters for Anthropic, Gemini, OpenAI-compatible providers and the image, video and audio model hosts.
A prepaid balance on an append-only ledger with hold-then-settle charging, usage per key and a per-request cost log. Stripe and crypto top-ups via webhooks.
Server-side ad conversion tracking for Meta, Reddit, X, Google Ads and GA4, plus an MCP server with OAuth for agent access.
The running parts
The hard parts
Before a request leaves, its worst-case cost is reserved as a hold, and a database constraint refuses any hold the balance cannot cover. Afterwards the hold settles to the real cost. A sweeper releases holds a crash left behind, and a reconciler charges jobs that finished but never settled.
A database trigger rejects every update and delete on ledger entries, and it is the only thing allowed to write the settled balance. The application cannot quietly change history even with a bug.
A stream is committed to a provider at its first byte. Before that, a bad request stops the chain, an auth failure skips only that provider, and a circuit breaker takes a failing provider out of rotation and lets a single probe test it again.
Limits run as one atomic script in Redis with a bucket per key, so every API process sees the same count and a burst cannot slip between two of them.
Workers claim jobs straight from PostgreSQL with row locks that skip what another worker already holds. Outgoing webhooks are signed, retried with backoff, and checked against the resolved address so they cannot be pointed at internal systems.
Start here
Tell us what is slow, manual or missing. You will hear back within one working day — then get a written scope and a fixed quote, or an honest no.