- Baseline
- $0.01779
- Candidate
- $0.01627
- Delta
- -$0.00152
Step 1 Provider and Model
iSelect the pricing row used for baseline and candidate cache scenarios.Step 2 Quick Mode
iSet baseline usage and cache assumptions before advanced tuning.Estimate whether cache quality and freshness work can pay back this quarter.
Optional Advanced assumptions
iTune retrieval and infra assumptions after Quick Mode is calibrated.Show advanced inputs
Scenario actions
Copy scenario URL
Paste into ChatGPT or Claude, or share with a teammate.
Save and track this scenario
Track pricing drift on this scenario and get an email if the latest result changes.
How tracking works
After you click Save and track, we carry this exact calculator state into the tracked-scenarios page so you can sign in and confirm the save.
We save your assumptions and the pricing snapshot used for this result.
When a newer pricing snapshot lands, we recompute the same scenario, show what changed, and email you if the latest result moved.
1 tracked scenario free, then $12/mo or $120/yr for up to 25 tracked scenarios.
Headline metric
Cache plan saves moneySavings per user / month: $0.1822
Hit-rate lift: 7.00 points (18.0% to 25.0%).
Savings per user / month
$0.1822Monthly savings
$118.45Cache savings delta
+$0.1822Equivalent free requests
10.2Totals
iBaseline vs candidate totals under the same cache-hit assumptions.- Baseline
- $2.1348
- Candidate
- $1.9526
- Delta
- -$0.1822
- Baseline
- 95.6%
- Candidate
- 96.0%
- Delta
- +0.4%
- Baseline
- $2.1348
- Candidate
- $1.9526
- Delta
- -$0.1822
| Metric | Baseline | Candidate | Delta |
|---|---|---|---|
| Cost per request | $0.01779 | $0.01627 | -$0.00152 |
| Cost per user/month | $2.1348 | $1.9526 | -$0.1822 |
| Gross margin % | 95.6% | 96.0% | +0.4% |
| Break-even price | $2.1348 | $1.9526 | -$0.1822 |
Component Breakdown
iBaseline and candidate components are computed independently, then differenced.- Baseline
- $0.114
- Candidate
- $0.114
- Delta
- $0
- Baseline
- $0.0396
- Candidate
- $0.0396
- Delta
- $0
- Baseline
- $2.4
- Candidate
- $2.4
- Delta
- $0
- Baseline
- $0
- Candidate
- $0
- Delta
- $0
- Baseline
- $0.0018
- Candidate
- $0.0018
- Delta
- $0
- Baseline
- $-0.4686
- Candidate
- $-0.6508
- Delta
- -$0.1822
- Baseline
- $0.048
- Candidate
- $0.048
- Delta
- $0
| Component | Baseline | Candidate | Delta |
|---|---|---|---|
| GenerationiModel input/output token spend for requests. | $0.114 | $0.114 | $0 |
| RetrievaliExtra model input spend from retrieved context chunks. | $0.0396 | $0.0396 | $0 |
| RerankingiReranker cost based on docs scored per request. | $2.4 | $2.4 | $0 |
| Embeddings IngestioniAmortized per-user share of the fixed monthly corpus embedding refresh cost. | $0 | $0 | $0 |
| Vector DbiVector database query cost across all requests. | $0.0018 | $0.0018 | $0 |
| CacheiSavings from cache hits. Negative means lower total cost. | $-0.4686 | $-0.6508 | -$0.1822 |
| InfraiNon-model infra overhead per request. | $0.048 | $0.048 | $0 |
Sensitivity RankingiDelta in total cost if one variable increases by 10%.
| Variable | Cost delta % |
|---|---|
| Requests Per User MonthiUser activity level per month. | 10.00% |
| Rerank DocsiDocs reranked per request. | 9.22% |
| Cache Hit RateiFraction of requests served by cache. | -3.33% |
| Output TokensiGenerated tokens per request. | 0.32% |
| Retrieved ChunksiRetrieved chunk count per request. | 0.15% |
| Tokens Per ChunkiAverage chunk size in tokens. | 0.15% |
| Input TokensiPrompt-side tokens per request. | 0.12% |
| Vector Queries Per RequestiVector query count per request. | 0.01% |
| Monthly Active UsersiActive-user estimate used to amortize fixed monthly embedding refresh. | -0.00% |
Assumptions and Units
iExplicit assumptions keep this comparison reproducible.- CurrencyUSD
- Token unittoken
- Pricing snapshot2026-07-20
- Selected model rowOpenAI/GPT-5 Mini
- Comparison ruleOnly cache hit rate changes; other usage assumptions stay shared
- Volume basisMonthly savings and fixed monthly terms use monthly active users as the denominator
Recommended Next Step
iUse this section to turn cache deltas into the next implementation checks.Validate provider constraints and cache implementation path, then re-check assumptions.
Compare infra providers
View Infra RecommendationsSources and Snapshot
iPricing comes from the current dated snapshot.Active Pricing Row
Candidate
OpenAI / GPT-5 Mini
- Input tokens$0.25 / 1M
- Output tokens$2 / 1M
Shared retrieval defaults
- Embedding input$0.02 / 1M
- Rerank docs$1 / 1K
- Snapshot date: 2026-07-20
- Source links and update notes: Pricing Snapshot Reference