AXIS Launch List your app
Investors & Fundraising editorialmetricsunit-economics

Gross margin math for AI apps: compute cost per dollar of revenue

Started by AXIS Editorial

Read two of the revenue-quality thread — margin, honestly computed — keeps generating enough intake follow-ups to earn its own editorial. This is the full worksheet: what counts, what doesn't, the traps specific to AI apps, and what the number is for. Investors run this math on you; the point of the thread is that you run it first.

The computation

Gross margin = (revenue − cost of revenue) / revenue, where cost of revenue for an AI app honestly includes:

  • Model API spend — all providers, all calls that serve customers (including the retries, the tool-use loops, and the evals you run per-request).
  • Inference-adjacent infrastructure — GPU serving if you self-host, embedding computation, vector store costs, retrieval infrastructure, response caching layers.
  • The per-customer slice of the stack — hosting, storage, and bandwidth that scale with usage (not your fixed dev tooling).
  • Third-party per-unit costs — transcription minutes, OCR pages, telephony, anything metered that a customer request consumes.
  • Human-in-the-loop labor, if serving requests requires it — the most-omitted line; if a contractor reviews outputs before delivery, that's cost of revenue, not operations.

What stays out: your salary and development time (below the line), customer acquisition (its own economics), one-time model experimentation, and fixed costs that don't scale with serving customers.

The traps, in the order makers fall into them

  1. Blended-away power users. A 75% blended margin can hide a top decile of users you serve at a loss. Compute margin by cohort and by plan — the flat-rate unlimited plan is where AI-app margins go to die, and investors specifically hunt for it.
  2. The free tier riding in operations. Free-user compute is real spend; account it as acquisition cost deliberately or it silently erodes the margin line you present.
  3. Yesterday's model prices. Provider price cuts are a tailwind, but margin computed on last quarter's pricing with this quarter's usage mix misleads in both directions. Recompute on the trailing-90 window (the mechanics reply in the revenue-quality thread has the formula).
  4. Growth-correlated degradation. The dangerous pattern: heavier users are your best revenue and your worst margins, so growth makes the blend worse — margin falling as MRR rises. The trendline exposes it; a snapshot hides it. Chart both lines or you're flying one instrument short.

What the number is for

Below roughly 50%, you're operationally a services-ish business wearing a SaaS valuation costume, and both investors and acquirers will re-dress you — pricing, caching, and model routing are the standard repairs, and "here's margin before and after our routing work" is among the strongest slides a technical founder can show. Above 80% with the honest computation, say so loudly and show the work: in a category where diligence keeps finding costume margins, demonstrated margin integrity is itself a differentiator — the arithmetic version of the verified-revenue asymmetry this platform is built on.

Post your number with its computation below — including the parts you're unsure belong. Sorting the edge cases in public is the cheapest margin consulting available.

Replies (4)

AXIS Editorial

Follow-up from maker intake: "Where do prompt-caching savings and batch discounts belong in the computation — and is optimizing them worth my time at $4k MRR?"

They belong in cost of revenue at realized rates — you spend what you spend after caching hits and batch pricing, so the honest number already includes them; what you shouldn't do is present the optimized rate on unoptimized volume ('we'd save X if' is a plan, not a margin). Worth your time: run the ratio check first — model spend as a share of revenue. Above ~20%, an afternoon of caching (system-prompt and context reuse where your provider supports it) and routing (cheap models for cheap subtasks, expensive ones where quality pays) routinely moves margin by whole points, which compounds into valuation at the multiples this ecosystem trades on. Below ~10%, your margin lever is elsewhere (pricing, plan structure) and compute golf is procrastination with graphs. The trendline from the thread tells you which case you're in.

AXIS Editorial

Follow-up from maker intake: "How do I fix the unlimited-plan trap without a churn-inducing repricing? Half my MRR is on a flat 'unlimited' tier I priced before understanding my costs."

The staged repair, per founders who've survived it: (1) instrument first — per-user compute cost by month, so you know whether 'the trap' is five users or fifty (usually it's shockingly few); (2) grandfather visibly — existing unlimited users keep terms, which converts the repricing story from betrayal to respect, and cap the new tier honestly ('fair use' ceilings with named numbers beat vague ones); (3) for the outlier five, the direct conversation works better than founders expect — 'you're our heaviest user; here's a plan priced for how you actually use it' lands as attention, not punishment, and heavy users are heavy because the product matters to them; (4) present the whole arc to investors as the margin-repair slide the thread mentions — a founder who found the trap, measured it, and engineered out of it without a churn spike has demonstrated exactly the unit-economics judgment the margin question was probing for. The version that fails: silent limits appearing in the fine print, discovered by your best users in public.

Jonathan (AXIS Launch)

Verification note from the platform side, because this thread's numbers feed our badges: Metrics Verified on margin means we checked the computation against source data — provider invoices and revenue records — on the stated date, using substantially the worksheet above. Two patterns from checks so far: the honest-mistake class (free-tier compute in operations, trap 2 — we recompute together, the corrected number gets verified, everyone's fine), and the one that concerns me more: makers who've never seen their own per-cohort numbers and are verifying blind — the badge process becoming their first real margin computation. Better late than at diligence, but the worksheet is free and the trendline needs history you can only start collecting now. Run it before you need it verified — the verification should be confirmation, not discovery.

AXIS Editorial

Follow-up from maker intake: "My margin is honestly 45% right now because I route everything to a frontier model for quality. Is sub-50% disqualifying, per the thread's threshold?"

Not disqualifying — undiagnosed sub-50% is the problem; yours comes with a stated cause and therefore a roadmap, which changes the conversation entirely. The frame investors respond to: 45% is a choice (quality-maximal routing during product-market-fit search) with a demonstrated path up — and 'demonstrated' is the operative word, so build the evidence: run your eval suite against cheaper models on your real workload's subtasks, find where quality holds (it usually holds on more than founders fear — classification, extraction, and reformatting rarely need frontier pricing; judgment-heavy generation might), ship routing for the safe subtasks, and chart the margin walk — 45% to 60-something with quality metrics flat is a routine outcome of that exercise, and the before/after chart is the strongest-slide candidate the thread describes. What would be disqualifying: 45% presented as 'AI apps just cost this' — because the investor across the table has seen the routing repair work too many times to accept the shrug.

Threads are permanent — locked, not deleted, once resolved. New posts go through a submission form and are published by moderators, usually within 1 business day.