Cash forecasting is unusually sensitive for AI startups because infrastructure expense can move faster than headcount. A product launch, model migration or utilization drop can change monthly burn even when revenue and staffing are stable.
Separate fixed and variable compute
API usage is largely variable, while reserved GPU capacity and committed cloud contracts behave more like fixed or semi-fixed costs. Forecasting should model them separately.
Build scenarios for model pricing
Do not assume current token or inference prices stay constant. Model a base case, price-decline case and adverse case where a key provider raises effective cost or discounts expire.
Utilization matters as much as unit price
Owned or reserved capacity can look cheap per theoretical GPU hour while being expensive per productive inference hour when utilization is low. Forecast both reserved capacity and actual workloads.
Map usage growth to cash, not just revenue
If users consume more AI than pricing captures, growth can worsen burn. Usage cohorts should connect product activity to direct model and infrastructure cost.
The commitment risk is discussed in AI Startup Compute Commitments.
Include working-capital timing
Annual prepayments may improve short-term cash while deferred revenue remains. Conversely, cloud vendors may require deposits or prepaid capacity. AI Startup Working Capital explains why cash quality can differ from reported growth.
Stress-test concentration
One large customer can drive both revenue and infrastructure demand. Model what happens if that customer contracts after capacity has already been committed.
Bottom line
An AI startup cash forecast should connect customer usage, model prices, reserved capacity and payment timing. A simple headcount-plus-cloud budget is not enough when infrastructure economics can change materially within a quarter.