Early access

Your provider invoice is one line.
Your AI costs six.

ai-tally meters what an AI feature actually burns across every layer, attributes that spend to the feature and to your own customers, and joins it to revenue. So you can answer what your AI costs all-in, and whether it pays for itself.

OpenTelemetry-native Apache-2.0 core No prompt text stored
Trailing 30 days, one tenant Example data
What the providers billed you
$16,390
OpenAI $10,480 Anthropic $5,180 Google $730
What the features actually cost
$35,440
LLM Compute Vector DB Tool calls Egress Embeddings
The model bill was 46% of the AI bill. The other 54% never appeared on a provider dashboard.

The problem

A provider dashboard tells you what you spent. It cannot tell you on what.

Tokens are the part you can already see, and they are less than half of it. Everything a finance conversation actually turns on is missing from the invoice.

  • Q. Which feature spent it? One API key serves the support copilot, doc search, and three internal jobs. The invoice is a single number for all of them.
  • Q. Which of your customers spent it? Direct spend per account plus each account's share of shared compute and egress, so you can find the customer whose seat price does not cover their inference.
  • Q. What did the layers around the model add? Vector search, tool calls, embeddings, GPU hours, egress. Real ingest paths, not an estimate multiplier.
  • Q. Did any of it come back? Cost per conversion and gross margin per provider, joined on a hashed user id against your own revenue events.
  • Q. How much of it bought nothing at all? Runs you were billed for that ended failed, work a retry made redundant, spend on a model two sizes larger than the job needs.

All-in cost

Every layer, one explorer.

Pivot the same month by layer, feature, model, or customer account. Nothing is modelled or marked up: tokens, tools and embeddings come off instrumented spans, vector cost off wrapped clients, compute and egress off daily cloud-billing connectors.

Cost explorer · last 30 days Example data

Recoverable cost

Five ways to pay for AI that returns nothing.

Each detector names where the waste is, bounds what is recoverable, and drills through to the runs behind it. Findings are hypotheses with evidence, not verdicts, and two of the five below refuse to put a number on themselves.

Paid for nothing 1,204 billed runs ended failed or abandoned. The tokens were charged; no result was produced. Directly observed, so the recoverable figure is exact. $2,840High confidence
Wrong-sized model On doc search, claude-sonnet-4 ties gpt-4o inside the eval confidence interval at $0.04169 per call against $0.04984. Cheaper candidates that lose the eval are not counted, however large their saving looks. $1,010Medium
Duplicated work Error-then-retry only: a run that failed and was superseded by a later same-shape success. A plain repeat is not claimed, because without message bodies it is indistinguishable from real multi-turn use. $1,190Medium
No measured return The onboarding agent shows spend with no attributed value. Top-of-funnel work and revenue that simply is not wired yet look identical from telemetry, so no recoverable amount is claimed. Unbounded
Structural inefficiency Context bloat and runaway loops are judged against each feature's own median, never a global average. The report writer settled 31 runs in this window; the floor is 50. Below floor
Bounded recoverable · 30 days $5,040 2 findings could not be bounded and are shown blank, not zero

Forecast

Where the month lands, and the day you cross budget.

A day-of-week-weighted median projection with an 80% confidence cone. It will not draw a line below a 14-day settled-history floor: a volatile number early in the month is worse than no number at all.

Month-end projection · September Example data
$40K $20K $10K $0 BUDGET $30,000 SEP 26 · BREACH SEP 1 SEP 8 · TODAY SEP 30
Settled to date$9,280
Projected month-end$34,800
80% cone$31.2K – $38.9K
Budget breachSep 26
The $1,160 daily median is drawn from 30 days of settled history, clear of the 14-day floor. Eight days in, that puts month-end at $34,800 against a $30,000 budget, and it reconciles with the $35,440 trailing-30-day total above: the month is tracking flat, not accelerating.

Compare

Are you on the right model? Decided on your traffic, not a leaderboard.

Opt-in sampling captures real requests with their resolved context, replays them against candidate providers under a daily budget cap, and scores the results with a pairwise judge. Win rates carry Wilson 95% intervals, so a tie reads as a tie.

Doc search · 1,240 replayed calls · 30 days Example data
gpt-4o In production
Cost per call $0.04984 · $6,180 / mo over 124,000 calls
BASELINE
claude-sonnet-4
Cost per call $0.04169 · $5,170 / mo, saves $1,010
46–56% WIN
gemini-2.5-flash
Cost per call $0.00900 · $1,116 / mo, saves $5,064
39–49% WIN
1% of doc-search traffic, held under the daily replay budget cap. Only claude-sonnet-4 ties: gemini's interval excludes 50%, so its $5,064 is not claimable as savings. A candidate with no judged eval pass renders for quality, never a placeholder percentage.

Cost per customer

Which of your customers pay for themselves.

Direct spend lands on an account by hashed account id. The shared layers cannot be attributed that way, so they are allocated, and the rule is named on screen rather than buried: pro-rata on direct spend. You are told which half of each number was measured and which was derived.

Cost per customer · last 30 days Example data
AccountDirectAllocatedAI costMargin
Northwind Ltd$4,132$1,988$6,120−13%
Brightline$3,362$1,618$4,98072%
Kestrel Health$2,188$1,052$3,24066%
Vantage Group$1,627$783$2,410
138 other accounts$12,621$6,069$18,69085%
142 accounts$23,930$11,510$35,44079%
Allocated is the $11,510 of shared compute and egress spread pro-rata on direct spend. Northwind bills $5,400 a month and costs $6,120 to serve, so it runs at −13%. The blended 79% covers only accounts with revenue mapped; Vantage is excluded rather than counted as zero.

Also on the spine

One dataset, so the answers agree with each other.

Every surface reads the same spans and the same attribution join, which is why cost per customer reconciles with cost per feature and both reconcile with the invoice.

Reconciliation

Metered spend against the invoice

A nightly job checks what was metered against what the provider actually billed, so the dashboard and the bill do not quietly drift apart.

Attribution

Cost per conversion, per provider

Joined on a hashed user id against Stripe webhooks or your own revenue API. Margin per provider, not just spend per provider.

Unit economics

CAC, LTV, payback

The AI cost of serving a customer folded into the unit economics you already report, so the AI line stops being a separate conversation.

Agent runs

Why this run cost 50× median

Per-agent run distribution at p50 and p99, with the pathological tail surfaced rather than averaged away.

Budgets & guardrails

Limits that observe before they enforce

Per-tenant and per-scope monthly budgets. Guardrails default to recording what would have fired; they never hard-kill customer state.

Connectors

Compute and egress, from the source

AWS Cost Explorer, GCP Billing, Vercel and Cloudflare on a daily cadence, reconciled against metered spend.

The rule

This is a real value. It means we do not know.

A cost tool that fabricates is worse than no cost tool, because you will act on it. So a number that cannot be defended is not rounded down, not estimated, and never rendered as zero. It is rendered blank, with the reason on hover.

  • A blank, not a zero A p95 built from 40 spans is . A quality cell with no eval behind it is . A failed job records that it failed and emits nothing.
  • No bodies in telemetry Counts, hashes and mapped events. Prompts, completions and retrieved text are dropped at the gateway before storage. That is the contract, not a setting.
  • Hashed identities User and account ids are HMAC-SHA256'd under a per-tenant key, so they cannot be reversed or joined across tenants. Credentials are stored as KMS references, never raw.
  • Billing decoupled from sampling A head-time meter counts every trace before the sampling decision, so cost figures stay exact no matter how far analytics sampling is turned down.

Wiring it up

Two lines, or none.

Built on OpenTelemetry gen_ai.* conventions, with extensions namespaced under the same. If you already emit OTel, most of this is a config change.

Python SDK Instrumented path
# tag the span with the feature and the customer it serves.
# cost is resolved from the pricing catalog at ingest.
import tally

tally.init(feature="doc-search", account_id=user.org_id)
Go edge proxy Language-agnostic metering in the request path. Point your base URL at the proxy; no application code changes.
Cloud connectors Read-only AWS, GCP, Vercel and Cloudflare billing, so compute and egress land next to the token spend.
Revenue source A Stripe webhook or the generic revenue API turns cost per feature into margin per feature.

Get started

Find out what your AI actually costs.

Early access is open to teams running AI in production. Bring one feature and a month of traffic; you will see the all-in number the same afternoon.

  1. Create your account Email or SSO. Takes a minute, no card.
  2. Name your organization Your workspace and its tenant key are provisioned with it.
  3. Send your first span Two SDK lines, or point a base URL at the edge proxy. Cost lands as it arrives.