Cost per AI action / inference COGS
The new COGS line that erodes AI-native unit economics.
- Formula
- (Inference + model API + compute cost) / active users (or / billable action)
- Unit
- $
- Models
- SaaS, Usage-based
No public benchmark exists for this metric yet. This is where Omega Point's proprietary data will fill in.
What it is
Cost per AI action / inference COGS measures the direct variable cost the business incurs each time an AI-powered feature executes on behalf of a user. The formula: (inference costs + model API fees + dedicated compute costs) divided by either active users per period or billable AI actions per period — track both denominators because they reveal different things.
How to calculate it
Pull all costs that exist solely because AI is running: cloud-provider inference compute, third-party model API charges (e.g., per-token or per-call pricing from foundation-model providers), and any dedicated GPU or accelerator capacity provisioned for model serving. Exclude general engineering infrastructure that would exist regardless of the AI feature. Divide the resulting total by active users (to get a per-user cost load) and separately by discrete AI actions completed (to get a per-call cost). The ratio of per-action cost to per-action price is the margin signal operators need to watch most closely.
Why it matters
This is the new COGS line that determines whether an AI-native product earns SaaS-like economics or services-like economics. Traditional SaaS has near-zero marginal cost per user action; AI-native products do not. Every additional use of an AI feature incurs real compute spend. If inference costs are not tracked with the same rigor as cloud infrastructure, gross margin will be systematically overstated until the problem is impossible to ignore — which is typically at the worst possible moment (rapid growth, fundraise, or enterprise contract negotiation).
How to read it
There is no published benchmark for cost per AI action or AI inference COGS. This metric is category-specific: a text-generation call is priced very differently from a multimodal reasoning call or a real-time voice interaction. Model providers change pricing frequently; the cost curve for inference has moved sharply downward over the past two years and continues to do so, so any figure cited today will be stale within months.
Because the cost landscape is this volatile and this model-specific, Omega Point does not propose an estimate band here — doing so would produce false precision. Instead, the right comparison is internal: track your own cost per action weekly, plot it against the price you charge per action (or the implicit value delivered per action in a flat-subscription model), and monitor the ratio. A healthy AI business should show that ratio improving over time as pricing power or efficiency gains outpace inference cost. The secondary comparison is against your own blended gross margin: if inference COGS is invisible in your P&L, your gross margin is not what you think it is.