Prompt caching: how ZIA and the other Prudai products use it
Prompt caching means a language model keeps the fixed opening part of a prompt and reuses it for the next question. ZIA, Prudai's product for healthcare, runs on the same platform as LEO and VERA. That platform builds every prompt so that the fixed part comes first and everything that changes per question comes last. This saves processing, and according to Anthropic it usually also shortens the wait for the first word of the answer (TTFT).
In short
- A language model remembers nothing between two questions. With every question it receives the whole prompt again: instructions, tools, earlier messages and the new question.
- Prompt caching lets the model provider reuse the already processed start of that prompt. This only works if that start is exactly the same as last time.
- That is why the Prudai platform puts the fixed part first and anything that changes per question, such as the date and time, last.
- For Claude models the platform places the cache markers itself. Gemini models cache automatically.
- The same structure applies to ZIA, LEO, VERA and the other products. Only the instructions and tools differ per product.
What is prompt caching?
Prompt caching is the reuse of the already processed start of a prompt. A prompt is everything a language model is given to read in a single request. How a language model splits text into tokens and processes it is explained in What is generative AI?.
A language model has no memory between two requests. When you ask your fifth question in a conversation, the model also receives the four earlier questions and answers again, plus all instructions. The longer the conversation, the longer the prompt.
With prompt caching, the provider keeps the processed start of such a prompt for a short time. If the next request starts with exactly the same content, that part does not need to be processed again. Only the new part still needs processing.
Why does prompt caching give a faster first answer (TTFT)?
TTFT stands for time to first token: the time between sending your question and the first word of the answer. Before that first word appears, the model has to process the entire prompt. A long prompt therefore means a longer wait.
If a large part of the prompt comes from the cache, that work disappears. Anthropic states in its documentation that you will generally see an improved TTFT for long documents. Costs go down too: at Anthropic, tokens read from the cache cost a tenth of the normal input price.
Exactly how much faster it gets depends on the model, the length of the conversation and the question itself. That is why we do not quote a speed figure of our own here.
Why does the order of the prompt matter?
A cache only works if the start of the prompt is identical, byte for byte. If a single character near the start changes, for example the current time, everything after it counts as new. So the rule is simple: what is fixed goes first, what changes goes last.
This is how the Prudai platform builds a prompt:
| Part of the prompt | What it contains | Does it change? | From the cache? |
|---|---|---|---|
| 1. Tools | What the product can look up or do. For ZIA, for example, searching information from the Dutch Healthcare Authority (NZa). | No, fixed at the first question of the conversation | Yes |
| 2. Instructions | Language policy and the product instructions, for ZIA focused on care | No | Yes |
| 3. Fixed context | Organisation, preferences, project scope and the list of knowledge bases | No, not within a conversation | Yes |
| 4. Earlier messages | The questions and answers so far | Grows, but the older part stays the same | Yes |
| 5. The new question | Your question, with whatever changes per turn: date and time, the conversation's notes and tasks, an open document, files you mention | Yes, every time | No |

Five-step diagram of how the Prudai platform builds a prompt: first the tools, then the instructions, then the fixed context, then the earlier messages, and finally the new question with whatever changes per turn. Everything before the new question can come from the cache.
How it works at Prudai
ZIA supports professionals in healthcare and the social domain. Staff often ask several questions in a row about the same topic. The first part of the prompt is then the same for every question: the tools, the ZIA instructions and the organisation's fixed context. That is exactly the part that can come from the cache.
One structure for all products. ZIA, LEO, VERA, BEVER, IRMA and MAIA use the same code to assemble a prompt. The products differ in instructions, tools and sources, but not in the order. An improvement to that structure therefore works for every product straight away.
Tools are fixed per conversation. At the first question, the platform records which tools the conversation gets. After that, every question carries exactly the same compact description of those tools. If the model needs a tool's full description, it fetches it. That description lands at the end of the prompt, so the fixed start stays the same.
Whatever changes per turn travels with the new question. The date and time, the conversation's notes and an open document belong to a single turn. So they do not sit at the start, but in the last message.
Also for step-by-step lookups. When ZIA looks something up in several steps, each step builds on the previous one. The platform marks the latest lookup result, so the next step can read the earlier part from the cache.
We measure whether it works. For every model call, the platform records how many tokens came from the cache and how many were newly written to it. It also keeps a fingerprint of the fixed start. That way we see immediately when a change quietly disturbs that start. Automated tests guard the order of the prompt.
This matches how we organise information in general: one fact in one fixed place. Read more in One fact, one place: information architecture for AI agents.
What does a care professional notice of prompt caching?
Staff do not need to configure anything. Prompt caching also does not change what the model reads: it gets the same prompt, only part of it does not need to be processed again. The answer is therefore no different because of the cache.
The platform is designed for reuse within one conversation. Asking follow-up questions in the same conversation benefits most. How ZIA fits with information security in Dutch healthcare is covered in NEN 7510 in practice.
For the specialist
Claude via Google Cloud Vertex AI. Vertex AI does not cache Claude automatically: the caller sets cache_control: {"type": "ephemeral"} on a block. Anthropic builds the cacheable part in the order tools, system, messages, up to and including the marked block. A single request may carry at most four markers.
The platform uses three, in fixed positions:
- On the first user message. This covers the tools, the instructions and the fixed context.
- On the second-to-last user message. This covers the earlier messages.
- On the latest tool result, during step-by-step lookups.
The last user message gets no marker. It is built differently at the next turn, so a cache entry for it would never be read again. A marker that moves every turn costs a new cache write each time it moves. That is why the markers sit in fixed positions and the fourth is deliberately left free.
Pricing and lifetime according to Anthropic (documentation checked on 6 October 2026): writing to the cache costs 1.25 times the normal input price, reading from it 0.1 times. A cache entry lives for five minutes by default and is refreshed at no extra cost each time it is used. There is also a one-hour option, at a higher write price. The platform uses the default. Per marker, Anthropic looks back at most twenty blocks for an earlier cache entry.
Gemini via Vertex AI. Gemini models use implicit caching: it is on by default and needs no markers. Google advises placing large, recurring content at the start of the prompt. The fixed order already does that. Here too, the provider reports per request how many tokens came from the cache.
Frequently asked questions
Does prompt caching change ZIA's answer?
No. The model receives exactly the same prompt as without a cache. Only the already processed start does not need to be computed again.
Do I need to configure anything as a user?
No. The platform handles the prompt structure and the cache markers itself. You may notice it in the waiting time, not in how ZIA works.
Does this only apply to ZIA?
No. ZIA, LEO, LEO Fiscaal, VERA, BEVER, IRMA and MAIA use the same prompt structure. They differ in instructions, tools and sources.
Why not simply cache the whole prompt?
Because the last part is different for every question. A cache entry for something that will never return in exactly that form costs a write but brings nothing back. That is why whatever changes sits at the end, outside the cache.
Sources
- Anthropic, Prompt caching (documentation, checked 6 October 2026)
- Google Cloud, Prompt caching for Claude on Vertex AI (documentation, checked 6 October 2026)
- Google Cloud, Context caching on Vertex AI (documentation, checked 6 October 2026)
- Google, Context caching in the Gemini API (documentation, checked 6 October 2026)
Would you like to know how ZIA supports professionals in healthcare and the social domain? See ZIA or read about the Nedap ONS integration.
Updated on 6 October 2026
Photo: ColossusCloud via Pixabay
