PLAN THE COST OF GETTING WORK DONE
Same workload.
Different economics.
What will the AI cost? What will it cost to run the service? Compare both with assumptions you can change.
START WITH A BUSINESS TASK
What does a workflow actually use?
Choose an example to explore what the model reads and writes. These invented budgets are teaching examples, not measured usage or quality benchmarks.
New to tokens? Start here.
Tokens are pieces of text that a model reads or produces. They may be words, parts of words, punctuation or other text fragments. There is no fixed words-to-tokens conversion: language, formatting and the model's tokenizer all matter.
Instructions + the customer's text + documents, history or tool results. If you send the same text again in another call, it counts again.
The reply, summary or structured result. Some models also bill for reasoning tokens; include those in a production estimate.
One business task can use several model calls. One attempt is the complete sequence; a retry repeats that whole sequence in this calculator. Measure real requests with the provider's tokenizer and usage reports before budgeting. Document examples assume text is already available; OCR, file parsing and external data access are excluded.
Compare the AI usage bill
AI text charges only. Use Workflow cost to include running costs.
NO BLACK BOX
Know what goes
into the number.
Lower cost is not a quality ranking. Use your own evaluation data before making a production decision.
How the AI usage bill is calculated
Per call = ((total input − cached input) × input rate + cached input × cache-read rate + output × output rate) ÷ 1,000,000. Cached input is a subset of total input, never counted twice.
Per task = per call × calls per attempt. Monthly API bill = per task × monthly tasks. API view assumes one attempt and excludes retries, tool fees and human work.
Workflow cost: all attempts, all outcomes
For per-attempt success p and maximum attempts n = 1 + retries, expected attempts = 1 + (1−p) + … + (1−p)ⁿ⁻¹. AI resolution probability = 1−(1−p)ⁿ.
Escalation probability = remaining failure probability × escalation share. Final resolution probability = AI resolution probability + escalation probability × human success rate. These outcomes do not overlap.
Base task subtotal = expected attempts × (API cost per attempt + tool fees per attempt) + escalation probability × human cost. Known monthly cost = base task subtotal × tasks + additional per-attempt usage × expected attempts × tasks + (initial setup ÷ amortization months + fixed monthly budgets) × allocated share. Known cost per resolved task = known monthly cost ÷ expected monthly resolutions; all failed work is included. No value is shown when volume or final resolution is zero.
Success is defined for the complete multi-call attempt, not each call. Attempts are assumed independent with unchanged cost and success probability, and stop when successful. Correlated failures, longer retry prompts, latency, parallel calls and staffing capacity are not modelled. Additional operating costs use your entered assumptions. Fixed and amortized costs remain at zero task volume; unit costs are then unavailable. Blank cost categories make the estimate partial. Completeness covers the entered categories only, not every possible business expense.
Rate scope & exclusions
Text-token standard API list prices in USD, per million tokens. No batch, priority, flex, volume discounts, taxes, currency conversion or free tiers. Use a consistent billable-token estimate; actual tokenization differs by provider. Include billed reasoning tokens in output if relevant.
Long-context surcharges, cache-write/storage costs, grounding/search charges and other modality costs are excluded. Cached tokens mean eligible cache reads only. Google explicit caching can add storage costs. Enter tool fees separately. Edit prices to match your own contract and context tier.
Official prices & verification
Prices are a dated snapshot, not a live feed. Confirm model availability and context-tier eligibility on the official pages before budgeting.