Agent-fleet cost intelligence · early access

Watch every dollar your agent fleet burns.

Your coding agents run up a bill you only see three weeks later on the invoice. OmnisVigil turns your OmnisRouter receipts into an honest ledger: cost per commit, per repo, per team, live. It catches the looping agent quietly burning $1,382 an hour and stops it before the invoice does, and it proves with an open benchmark which work never needed a frontier model. Team-level, and never a per-developer scoreboard.

See it on your own bill How it works Built on OmnisRouter receipts Flat fee · BYOK

Illustrative view. The $1,382-in-one-hour looping agent is a real, publicly reported incident.

$1,382/hr
the looping agent OmnisVigil flags in minute two, not three weeks later on the invoice
60%+
of coding-agent requests a cheaper model handles, measured by OmnisBench, not guessed
0%
of your spend we take a cut of. Flat fee. We only win when your bill shrinks
100%
of spend attributed to a commit, repo and team. Content-free. Never a per-dev scoreboard
The bill nobody saw coming

Agentic coding detached the cost from the sticker

A $20 seat anchors what you expect to pay. Then an agent reloads the whole codebase every turn, runs unattended, loops on a bad tool call, and the meter runs the entire time. Uber burned its 2026 AI budget in four months. One team was charged $1,382 in a single hour when an agent looped tagging a backlog. The pain in the wild is always the same two words: bill shock. And it lands after the money is gone, because the tools show you one aggregate number and no daily cap.

It's unpredictable

Some days $40, some days under $10, then a single session that dwarfs the month. You can't forecast it and you can't explain it after the fact.

It loops and burns

Agents retry, explore dead ends, and re-bill the full context every turn. Fifty thousand tokens to do five thousand tokens of work, and nothing stops it.

You can't see inside it

Total tokens, total cost, maybe per model. Never per commit, per repo, per team, or per agent run. So nobody can say what anything cost.

The honest ledger

Every dollar, attributed to the work that spent it

OmnisVigil reads the receipt OmnisRouter already writes for every request and rolls it up into the numbers a team lead needs: cost per commit, per repo, per team, per agent session. No prompt content, no code, no per-developer league table. Only the money: where it went, and whether it was worth it, in a view you can hand to finance without flinching.

# the receipt OmnisRouter writes, rolled up by OmnisVigil
commit  a1f9c2e  payments-api   3 requests   $0.31
  routed  gpt-5-nano   refactor tests      $0.004   saved vs frontier -$0.12
  routed  gemini-flash summarise diff      $0.001   saved vs frontier -$0.03
  escalated opus-5   design the migration   $0.30   needed the frontier
# and the alert that matters, in real time
14:03  RUNAWAY  payments-api agent  1.3B tokens/hr  $1,382  cap hit → paused

The escalation line is the honesty. When a change genuinely needed the frontier model, the ledger says so. When it didn't, you see exactly what the cheaper model did and what it saved. Nobody's hiding the number to protect a margin.

Budgets and a kill-switch

Hard caps, live alerts, and a cut-off

Optimising the average call is not enough when one runaway session can cost more than a month. OmnisVigil watches spend as it happens and gives you the controls the seat-based tools gate behind enterprise: hard caps per team, per agent and per task, an anomaly alert the moment a session goes vertical, and a cut-off that stops the loop before it stops your budget.

Hard caps

Set a daily or monthly ceiling per team, repo, agent or task. A spike trips it mid-day and pauses the offender, so you find out in minutes, not on the invoice.

Runaway detection

A looping or retrying agent shows up as a session going vertical against its own history. OmnisVigil flags it live and can pause it automatically.

One cut-off

Stop a single agent session on the spot and keep its log. The work you meant to run keeps running. The loop that was burning money does not.

Proven cheaper, not cheaper-feeling

The only cost layer whose incentive is to shrink your bill

Look at how the tools around this problem make money. OpenRouter passes provider pricing through but charges a fee on the credits you buy, around 5%. Requesty adds a flat 5% markup. Both earn more as your bill grows, so "we save you money" runs against their own revenue. Observability tools like Langfuse and Helicone log what you spent and stop there. OmnisVigil charges a flat fee, runs on your own keys, and takes nothing off the top. It doesn't ask you to trust a savings claim either: it points at the open OmnisBench benchmark that measures which model each kind of work needs, so "cheaper" is a number you can reproduce.

vs marketplaces

OpenRouter, Requesty. Genuinely useful for one bill across many models, and you can keep using them. But they take a percentage of spend, so shrinking your bill works against them. OmnisVigil sits alongside and does the part they won't.

vs observability

Langfuse, Helicone. Strong at logging and tracing what you spent. They don't act on it: no cheapest-capable routing, no hard cap, no cut-off. OmnisVigil closes the loop from measured spend to cheaper-model action.

vs quality routers

Not Diamond, Martian. They claim big savings on their own benchmarks. The independent RouterArena ranks Not Diamond near the bottom of commercial routers, for frequently choosing expensive models. OmnisVigil's savings ride on the open, reproducible OmnisBench instead.

Credit where it's due: each of these is good at what it was built for, and OmnisVigil sits next to them, not on top of them. The difference is the one thing a spend-percentage business can't offer, an incentive to make the number go down.

Under the hood

It sits on receipts you already have

Point your agents at OmnisRouter, and every request comes back with a receipt: the model, the cost, and what a cheaper-capable model would have cost. OmnisVigil is the layer that reads those receipts and turns them into a fleet view, with the caps and alerts on top.

Content-free by construction

OmnisVigil sees costs, models and metadata, never prompts, code or keys. The receipt is designed to carry the money, not the work.

Attributed to the work

Commit, repo, team and agent session come from the request context, so spend rolls up the way an engineering org is shaped.

Self-host or hosted

Run the whole thing yourself next to OmnisRouter, or use the hosted dashboard for the team view. Your traffic and keys stay yours either way.

Team-level, never a per-developer scoreboard

OmnisVigil reports spend by team, repo and workload, so a lead can govern a budget and defend it upward. It will not rank individual engineers by tokens or cost. Surveillance dressed as FinOps is a different product, and not one we'll build. This is a line, not a setting.

See it on your own bill

OmnisVigil is in early access. Send us a month of your agent usage and we'll reprice it, read-only, and show you what routing would have saved and where the runaways were. If the number isn't real, we'll tell you. That's the whole pitch.

Questions

Frequently asked

Straight answers to what teams ask about agent-fleet cost.

What is OmnisVigil, in one line?
It's a cost-intelligence layer for teams running AI coding agents. It reads the per-request receipts OmnisRouter writes and turns them into a live ledger of spend by commit, repo and team, with hard budget caps, runaway detection and a cut-off.
Does it see our code or prompts?
No. The receipt carries costs, model choices and metadata, never prompt content, code or provider keys. OmnisVigil is content-free by construction, which is what lets you hand the numbers to finance without a data-exposure conversation.
Is this a per-developer productivity tracker?
No, and it never will be. OmnisVigil reports at team, repo and workload level so a lead can govern a budget. It does not rank individual engineers. That's a deliberate line, not a missing feature.
How is it different from an LLM observability tool?
Observability tools log what you spent. OmnisVigil acts on it: it caps spend, catches the runaway in real time, attributes cost to the actual work, and points at an open benchmark for which model each job needed. It closes the loop from measured spend to cheaper-model action.
Why should we trust the savings number?
Because we don't take a cut of your spend, so we have no reason to inflate it, and because the "cheaper model is good enough" call is backed by the open OmnisBench benchmark you can reproduce, not an unpublished internal claim.
Do we have to self-host?
You can. Run OmnisVigil next to OmnisRouter on your own infrastructure, or use the hosted dashboard for the team view while your traffic and keys stay on your side. The routing layer, OmnisRouter, is open source and Apache-2.0.
Read it honestly

Where it stands

Early access, on purpose

OmnisVigil is being built in the open with a handful of teams. The fastest way in is to let us reprice a month of your usage, so the first number you see is your own.

It needs the receipts

The ledger is only as good as the request context. Run your agents through OmnisRouter and every call carries the cost, the model and the cheaper-capable delta that OmnisVigil rolls up.

Savings are measured, then modelled

Coding and math savings come from real OmnisBench runs. Where we estimate, we label it an estimate. We would rather show you a smaller honest number than a big unprovable one.

Dashed prices, dated

Every cost figure comes from a committed, dated pricing snapshot, shared with OmnisRouter and OmnisBench, so a receipt reproduces even after list prices move.