Cost per model, live
Spend, tokens and cache savings resolved per model and project, with prices refreshed for you.
Swap one baseURL. Get cost, latency and full agent traces for every provider you already ship on.
of config to change. The SDK call stays exactly as it is.
providers routed through the same endpoint today.
streaming deadline, then the partial trace is still written.
licensed. Run the Docker image yourself whenever you want.
What lands in the dashboard
Spend, tokens and cache savings resolved per model and project, with prices refreshed for you.
Spikes, error bursts and latency drift come with the model and project that caused them.
Pin a version per request, then compare cost and score on real traffic.
One header turns on exact-match caching. Hits skip the provider and bill zero.
x-spanlens-cache: 3600Grade with an API key, gate the merge, queue the unsure cases for a human.
$ spanlens eval run tone
38 cases · 2 flagged
score 0.91 · gate passedAgent traces
Nested spans across tools, retrievals and model calls, with parallel branches kept intact. The slow one is obvious at a glance.
Open a sample traceThe proxy speaks each provider natively, so streaming, tools and structured output keep working untouched.
chat, responses and embeddings, streaming included
Bring the whole team on any plan. The repo stays MIT, so self-hosting is always the free exit.
For the first integration.
For features already in production.
For several products and tighter control.
For volume, SSO and a signed contract.
Self-hosted stays free. Unlimited requests, unlimited history, your own Docker image.
docker run spanlens/serverThe first request you send shows up with cost, tokens and latency attached.