ghostrouter
SERVICES · WHAT CLIENTS HIRE US FOR

Four things we do. All of them shrink the bill.

Almost everyone starts in the same place: a fixed-fee Routing Review, which tells you what you are spending and what you would save before you commit to anything. Most people then hand us the running of it — Managed AI Routing is the flagship, and it is where the 40–60% usually comes from. The last two are depth rather than breadth: we run the judgement layer for companies with AI agents in production, and we build the private, redundant connections for companies that cannot have their AI traffic crossing the open internet. You can hire any one of them on its own.

01 / AUDIT START HERE · FIXED FEE

Routing Review

A fixed-fee audit of what your company spends on AI. We read a sample of your traffic and show you, workload by workload, exactly what you'd save with smarter routing — before you commit to anything.

  • · Traffic-shape report: every AI workload you run, classified by what it actually needs
  • · Line-by-line savings estimate vs your current bill
  • · A draft routing policy (runbook) you keep, whether or not you hire us
  • · One review call with routing engineering; fixed fee, no commitment

Full details

02 / FLAGSHIP

Managed AI Routing

Our flagship. Point your software at GhostRouter instead of directly at an AI provider — usually one line of config — and every request gets sent to the cheapest model that can do the job properly. Your product behaves the same; your invoice shrinks. Typical clients cut model spend 40–60% in the first billing period.

  • · Onboarding: SDK drop-in or base-URL swap, first routed request in under an hour
  • · A custom routing policy written for your traffic, versioned like code, changed only by an approved pull request
  • · Automatic escalation when a cheap model gets it wrong, and failover across 11 providers when one goes dark
  • · Monthly savings report + client portal showing usage, invoices, and payment

Full details

03 / DEPTH

AI Agent Operations

If you run AI agents in production, we run their judgement layer: every answer is checked before it's trusted, retried when it's wrong, and handed to a human when it matters. You get agents you can leave unattended without waking up to a mess.

  • · Verification rules for every agent output (schema checks + quality judge)
  • · A tuned escalation ladder: retry counts, tier climbs, cost guards per workflow
  • · Human-review queue for the decisions that should never be fully automatic
  • · A complete audit log: every decision, who made it, and the policy clause that authorised it

Full details

04 / DEPTH

Private Networking & Failover

Private, encrypted connections between your systems, your AI traffic, and your providers — with backup paths built and tested before you ever need them. No public exposure, no scrambling during an outage.

  • · Private overlay network between your services: mutually authenticated links, no public ingress
  • · Private relay hops you choose, with per-link latency reporting
  • · Automatic failover in under a second when a path degrades
  • · An audit log entry for every route change, automatic or manual

Full details


SERVICE 01 · ROUTING REVIEW

We tell you what you'd save before you hire us.

You send us a sample of the AI requests your software already makes — a week is usually plenty, and we can work from your provider's own export if that is easier. We sort them into the jobs they actually are: pulling fields out of documents, rewriting text, answering a customer, deciding something genuinely hard. Then we price each group twice: what you paid, and what it would have cost sent to the cheapest model that handles that group correctly.

The difference between those two numbers is the entire reason this company exists. We put it in writing, with the working shown, and you keep the document either way.

The fee is fixed and quoted up front. There is no discovery phase that quietly becomes a project, and there is nothing to sign at the end of it unless you want to.

WHAT YOU GET

traffic-shape report
Every AI workload you run, classified by what it actually needs rather than by which team built it.
savings estimate
Line by line against your current bill, with the assumptions written next to the arithmetic.
draft runbook
A readable routing policy for your traffic. Yours to keep and implement yourself if you prefer.
review call
One working session with the engineers who wrote it. Fixed fee, no commitment.

SAMPLE · WHAT A TRAFFIC-SHAPE REPORT LOOKS LIKE

Illustrative extract · workload names are the client's own, figures are per calendar month
Workload Requests What it needs Paying for Should run on Est. saving
invoice-extract1,240,000fixed-schema extractionfrontiersmall + schema check−91%
support-triage430,000short classificationfrontiersmall−88%
report-summary96,000synthesis, some judgementfrontiermid−54%
contract-review3,100genuine reasoning, high stakesfrontierfrontier · unchanged
translate-ui210,000translationmidsmall−62%
nightly-etl2,900,000mechanical transformmidsmall · batch lane−71%

Note the fourth row. Roughly one workload in five genuinely needs the expensive model, and we say so in the report — a review that recommends moving everything down a tier is a sales document, not an audit.


SERVICE 03 · AI AGENT OPERATIONS

Agents you can leave running overnight.

An agent that is right 95% of the time is not 95% useful — it is a thing somebody has to watch. We run the layer that does the watching: every answer an agent produces is checked before anything downstream trusts it, wrong answers are retried, stubborn ones climb to a stronger model, and the small number that still don't resolve go to a person instead of into your database.

The check is two-part and neither part is a vibe. First the answer has to be the right shape — the fields you asked for, the types you asked for, the totals adding up. Then, where it matters, a second model reads the answer against the request and scores whether it actually did the job. Only answers that pass both are returned.

What happens after a failure is yours to set, per workflow. A customer-facing agent might retry once and climb immediately. A nightly batch job might retry three times and never climb at all, because nobody is waiting and the cheap model gets there eventually. These numbers live in your policy file, not in our heads.

Some decisions should never be fully automatic, and we would rather say so out loud than quietly automate them. Those route to a human-review queue with the agent's working attached, so the person deciding can see what the machine thought and why it wasn't trusted.

Everything is written down as it happens: the decision, the model that made it, the verification result, the retry count, and the clause in your policy that authorised the outcome. When somebody asks in six months why an agent did something, the answer is a query, not an investigation.

THE LOOP, ONCE PER AGENT ANSWER

answer verify escalate return schema + judge PASS FAIL ×n HUMAN REVIEW EVERY BRANCH WRITES ONE AUDIT ROW

DEEP DIVES · THE TECHNICAL STORY

If you want the mechanism, it lives here.

The two services below have pages of their own, written for the person who will actually be responsible for this: how a request is scored, what happens when a cheap model gets it wrong, what a policy file looks like, and how traffic keeps moving when a provider or a path goes dark.

DEEP DIVE

Managed AI Routing

The whole path a request takes: the three jobs run on every call, the escalation ladder that decides when cheap is not good enough, the policy file you own and review, and a console you can type into yourself.

The cheapest correct path, run for you

DEEP DIVE

Private Networking & Failover

The layer underneath: a private overlay between your services with no public ingress, relay hops you nominate, and standby paths that are kept warm and proven long before an incident asks for them.

Paths that exist before you need them

GET STARTED

Start with the review. Decide afterwards.

Tell us roughly what your software does with AI and we will tell you what a Routing Review costs and how long it takes. Existing clients: usage, invoices and payment live in the portal.