Pricing Analysis
*Costs vary with use case and volume. Every model here is priced as published, with no memory of its own.
Five ways of asking the same question, priced against the shelf. The document is the context in every one of them: we ingest it once and are never sent it again, while a frontier model is handed it - and the whole transcript - on every turn. Assumptions are stated under the grid.
01
One person, one message
No cachingWe are never sent the document. Every model here is sent all of it, to answer one question. The cheapest of them, DeepSeek V4.1 Flash, costs 6.7× ours.
02
One person, one conversation
No cachingEach message re-sends the document and the transcript. We are sent neither, so our line is flat. After 12 messages: $8.56 for Opus 5, $1.71 for Haiku 4.5, $0.038 for us.
03
One person, one conversation
Prompt cache warmThe cache holds the document for five minutes at a tenth of the price. It saves $6.80 on Opus. The line still climbs, because the cache cannot hold the transcript: $1.76 against $0.038, 46× the cost.
04
A month of support, 10 messages a day each
No cachingThe curves bend because the transcript grows and is re-sent every turn. Ours is straight. Change the crowd size: the axis grows, the gap does not.
05
A month of support, 10 messages a day each
Prompt cache warmThe cache saves real money and the curves still bend. It is cold again every morning, so every person pays to fill it once a day. The transcript it cannot hold is the part that grows.
Talk to us
The grid above prices a conversation. Pricing an integration means your endpoints, your volumes and your retention, so we would rather work it out against your numbers than publish one we would have to walk back.