AI cost observability
Inferly: knowing what an AI feature costs while it is running
Inferly is a developer tool Viva Studio designed, built and runs. Its audience reads documentation closely and notices sloppy work, which made it the most exacting product in this portfolio to get right.

The problem
Teams ship an AI feature and then find out what it cost when the invoice arrives a month later. By then the spike has been running for weeks, and there is no per-call record to work out which feature, which model, or which customer caused it.
Who it is for
Engineering teams who have added AI features to a product and want per-call cost, usage and error visibility in real time.
What we built
The features that make up the product as it runs today.
- Telemetry arrives through a single HTTP endpoint — a developer can start sending data in the time it takes to write one request.
- No SDK to install, so nothing new enters the customer's dependency tree.
- Each call records the model, token counts, latency and whether it succeeded.
- A dated history of provider pricing means that when a model's rate changes, last month's figures still add up.
- Spend, usage and error charts update straight away rather than overnight.
- Spend and error alerts push to Slack or any webhook, so the team hears about a spike from Inferly rather than from finance.
- Prompts never leave the customer's application — Inferly receives the numbers about a call, never its contents.
Stack and integrations
Built with
- Next.js
- TypeScript
- Postgres on Supabase
Talks to
- Slack, and any webhook endpoint, for spend and error alerts.
The hard parts we solved
Any studio can list features. These are the parts that were genuinely difficult.
Pricing that changes under you
Model rates move. If yesterday's calls are re-priced at today's rate, every historical figure quietly becomes wrong. Storing a dated history of provider pricing and pricing each call against the rate in force when it happened is what makes last month's number still true next month.
Earning trust with production data
The first question a technical buyer asks is what you do with their prompts. The answer had to be structural rather than a policy: Inferly only ever receives the numbers about a call. That constraint shaped the API rather than being documented around it.
An API developers will not resent
One endpoint and no SDK is a harder design than a client library, because every affordance a library would give you has to be obvious from the request shape alone. The audience notices the difference.
See it running, then talk to us about yours
Inferly is live and you can open it right now. If your project needs something similar, that is the conversation to have.