A blended gross margin can hide how different AI features really are. Text chat may be highly profitable, voice may carry moderate variable cost, and video generation can consume orders of magnitude more compute per user action. If all of those workloads sit inside one subscription and one margin line, management may not know which feature is creating value and which one is quietly consuming it.

Start with feature-level revenue attribution

Identify which revenue is tied to text, voice, image, video or agent usage. This is straightforward for credit-based products and harder for one-price subscriptions.

For subscriptions, allocate revenue using usage, stated plan value or another consistent internal method.

Measure variable serving cost directly

Track model APIs, GPU time, storage, bandwidth and other costs that scale with each modality. Video often includes generation plus encoding and delivery, while voice can include ASR and TTS.

Do not assume one provider invoice can be allocated accurately without workload tagging.

Text can subsidize expensive media

A subscription with heavy text usage and occasional image generation may have strong blended margin. But a small segment of users can consume many videos and dramatically reduce contribution.

Segment margin by usage intensity as well as modality.

Credits can make cost differences visible

If one video costs 100 credits and one chat costs one credit, the product communicates that the underlying resources differ. Internal credit economics should still be reviewed as model prices change.

Credit weights can become outdated quickly.

Voice has several cost layers

Real-time voice may require speech recognition, language model inference and speech synthesis on every turn. Streaming infrastructure adds another cost and latency layer.

Measure voice sessions by minute, turn or successful conversation rather than only tokens.

Video needs a different capacity model

Video workloads can be bursty, GPU-intensive and sensitive to resolution or duration. Reserved GPU capacity may improve unit cost but create idle-expense risk.

Margin reporting should include the real cost of unused reserved capacity where appropriate.

Caching affects modalities differently

Text prompts and embeddings may benefit from caching more than fully personalized video. Measure actual cache-hit savings rather than applying one assumption across the product.

Infrastructure optimizations should show up in the modality P&L.

Free usage can distort blended margin

If free users generate expensive images while paid users mostly chat, a blended gross-margin calculation may obscure acquisition cost inside product COGS.

Separate promotional usage from paid serving economics.

Price expensive features intentionally

A premium video feature may justify higher credit cost, usage limits or a separate plan. Do not rely on low-cost text usage to subsidize unlimited video unless that subsidy is deliberate.

Product pricing should reflect the value and marginal cost of each modality.

Track margin after refunds and creator share

For creator AI, a paid video may also owe a creator royalty and payment fee. Contribution margin should include the full variable economics, not only model inference.

This is especially important when comparing monetization formats.

Use cohorts to see modality adoption

New customers may start with chat and later adopt voice or video. Cohort analysis can show whether higher engagement improves or worsens economics over time.

Expansion is not always positive if the most popular feature is deeply unprofitable.

Review provider changes by modality

A new text model may cut cost while a video provider increases price. One blended infrastructure rate would hide the shift.

Our article on AI startup gross-margin sensitivity provides a framework for provider price scenarios.

Use modality margin when setting product limits

If one feature consistently destroys margin for a small group of heavy users, the response may be plan design rather than a company-wide price increase. Usage caps, premium tiers or lower-cost model routing can target the real problem.

Feature-level economics give product teams more precise levers than a blended gross-margin alarm.

Track support cost by modality too

Video and voice products may create more failed-generation tickets, refund requests or moderation work than text. These variable operational costs can materially change contribution margin.

A modality P&L becomes more useful when it captures the full cost to serve, not just infrastructure.

One company can contain several different AI businesses

Text, voice, image and video can have different pricing, retention and infrastructure economics. Separate modality P&Ls make those differences visible so management can invest in growth without accidentally scaling the wrong margin profile.