GPU Inference Cost

Estimate self-hosted AI inference cost from GPU price, token throughput, utilization and workload volume.

How this calculator works

Enter gpu cost per hour ($), sustained tokens per second, utilization (%), monthly output tokens. Select Calculate to apply the displayed formula and review the labeled results.

Formula / method

Cost per 1M tokens = hourly GPU cost ÷ (tokens/second × utilization × 3,600) × 1,000,000

Worked example

A $2.50/hour GPU sustaining 100 tokens/s at 70% utilization costs about $9.92 per million output tokens.

Assumptions and limitations

Throughput is sustained output-token throughput and the entered utilization captures idle capacity.

CPU, memory, storage, networking, replicas, prompt processing, operations, failures and reserved-instance pricing are excluded.