Since June 1, 2026, GitHub bills every token in a Copilot session, and cached tokens cost a fraction of fresh ones. GitHub Copilot prompt caching works in opposite ways on Anthropic and OpenAI models, so a habit that saves money on one does nothing on the other. Teams that run both providers need to know which rule applies where.
- Why does one habit not fit both providers? See what makes each cache expire.
- What does a broken cache cost? Understand the price of one miss on each model.
- How do you tell whether your cache works? Learn where GitHub shows it.
Each Provider Needs Its Own Habit
Running OpenAI models? Keep the session alive and come back inside the cache window. After a gap of 40 to 60 minutes, the VS Code team's extended retention made the GPT-5.4 cache hit rate 10.19 times higher than before (+919%). Do not start a new conversation because one answer was weak. Follow up instead.
Running Claude models? Keep the tool set and system instructions unchanged between turns, and keep turns coming. Anthropic's cache lives five minutes and renews each time it is used, but a change to the tool definitions resets all of it.
These platform habits sit on top of the workflow habits that cut token waste that you control directly: shorter output, leaner context, tighter instructions.
GitHub Copilot Prompt Caching Pays Only While the Prefix Holds
Every turn in an agentic session sends the conversation so far, every tool definition, the repository context and the system instructions. Most of that repeats from one turn to the next. Providers call the repeating start the prompt prefix. When a request shares an identical prefix, the provider reuses the model state it already computed instead of processing it again, and the VS Code team says cached tokens can be up to 10 times cheaper.
The condition is in the word identical. A cache pays off only while the prefix stays the same, and the two providers lose it for different reasons.
Anthropic Prompt Caching Breaks on Changes, OpenAI's on Silence
OpenAI's cache runs out when you stop sending requests. Anthropic prompt caching holds as long as the start of the prompt does not change, and runs out only after five minutes without use.

Figure 1. Two ways a prompt cache breaks in GitHub Copilot. View full-size image
OpenAI: extend retention to survive pauses
OpenAI caches the prompt prefix automatically in fast GPU memory. By default that cache is dropped after about 5 to 10 minutes of inactivity, and in some cases up to an hour. VS Code asks the API for 24-hour retention, which moves the cache to roomier GPU-local storage for up to 24 hours, so a lunch break no longer resets it. The team measured the relative increase in cache hit rate by the gap between requests:
| Gap between requests | GPT-5.2 |
GPT-5.3-Codex |
GPT-5.4 |
| 10–20 min | +13% | +32% | +10% |
| 20–30 min | +135% | +142% | +137% |
| 30–40 min | +301% | +203% | +679% |
| 40–60 min | +338% | +279% | +919% |
Source: Ryan Caldwell and Bhavya U, "Improving token efficiency for GitHub Copilot in VS Code," code.visualstudio.com, June 17, 2026. Relative changes, not percentage points.
The gain grows with the pause. After 10 to 20 minutes it is between 10% and 32%, and after 40 to 60 minutes between 279% and 919%, depending on the model. Without extended retention, that same return is a cold start.
GPT-5.6 and GPT-6 change the OpenAI rules
Those measurements cover GPT-5.2 to GPT-5.5. Newer models cache differently. For GPT-5.6 and later, OpenAI's documentation says a cached prefix stays eligible for 30 minutes after its latest write or reuse, and 30 minutes is also the only lifetime setting. GPT-6 Astra became generally available in Copilot on September 4, 2026. GitHub has published no Copilot-specific cache measurements for these models, so the safe habit on them is to return within 30 minutes of the last request.
Claude: protect the prompt structure
Claude prompt caching works the other way round. Nothing caches unless the caller marks it, so VS Code places up to four explicit cache breakpoints: after the tool definitions, after the system prompt, and on the two most recent cacheable messages. The older of the two message anchors is a safety net. If the newest one misses, because a slow tool call let its cache lapse or the content drifted, the older anchor still serves a hit, and the team typically gives up one exchange instead of the whole conversation.
That design reaches a cache hit rate of about 94% in agentic workloads, so only a small share of each request's input is processed again. It still has two weak points. Anthropic's documentation says the prefix is built in the order tools, system prompt, messages, and a change at one level resets that level and everything after it, so a change to tool definitions resets the whole cache.
And the cache lives five minutes, renewed at no extra cost on each use, so a longer pause forces a rewrite. That makes the enabled tool set, including MCP servers, a cost decision and not only a developer preference.
Source: Ryan Caldwell and Bhavya U, VS Code team, June 2026.
Deferred Tool Loading Trims Tokens and Leaves the Cache Alone
Tool search sends the model only the name and description of each tool, and loads the full schema when the model calls it. Before, every tool's full definition went out with every request, and Copilot agents can reach dozens of tools. OpenAI GPT-5.4 and newer support this natively through a defer_loading flag. For Anthropic models, VS Code first used Anthropic's server-side search, then moved the search to the client and matched tools by intent with Copilot's embedding model instead of by keyword.
| Model | Tokens per turn | Median user, whole session |
GPT-5.4 |
−9.8% | −9.0% |
GPT-5.5 |
−8.6% | −10.9% |
| Claude, server-side | −11.1% | −18.0% |
Total tokens, p50. OpenAI measured over four days, Anthropic over seven.
The client-side version added roughly 2% lower latency on Claude Opus 4.6 and cut user error rates by 4% on Claude Sonnet 4.6, according to the VS Code team's write-up. It also protects the cache. Deferred tools sit outside the cached prefix, so loading one does not rewrite the prefix.
WebSocket Transport Makes OpenAI Turns 12 to 19% Faster
An agentic turn is a chain of requests: generate, call a tool, wait, generate again. Over HTTP, each step is a separate API request. The Responses API WebSocket mode keeps one connection open for the whole chain. In VS Code, median time to first token fell 19.5% on GPT-5.3-Codex and 16.4% on GPT-5.4, and median time to complete a turn fell 13.6% and 11.7%. On GPT-5.4, active users rose 2.2% and two-day engagement 3.1%.
GitHub made WebSocket the default for GPT-5.2 and newer across Copilot products, including VS Code, Copilot CLI and the GitHub app, so there is nothing to configure. The saving is time. The money still depends on the cache.
GitHub's Usage Report Shows Whether Your Cache Works
Since August 11, 2026, the AI usage report lists input, output, cache read and cache write tokens for each model next to the AI credits they consumed. Download it from the AI usage page in your billing settings, as an admin on Copilot Business or Copilot Enterprise or as an individual user, per GitHub's changelog.
The cache columns matter because reading and writing the cache are priced differently. On the models below, a cache read costs 5 to 10 percent of the input price. Claude models and GPT-6 Astra also charge 25 percent more than the input price to write the cache, while GPT-5.4 and GPT-5.5 charge nothing for it.
| Model | Input | Cached input | Cache write |
GPT-5.4 |
$2.50 | $0.25 | none |
GPT-5.5 |
$5.00 | $0.50 | none |
GPT-6 Astra |
$10.00 | $1.00 | $12.50 |
Claude Sonnet 4.6 |
$3.00 | $0.30 | $3.75 |
Claude Opus 5.5 |
$4.00 | $0.20 | $5.00 |
USD per million tokens, from GitHub's rate card, checked October 2, 2026.
A break is therefore not free. A prefix written once and read nine times costs about 2.15 times its plain input price instead of 10 times, and each extra break adds another write at 1.25 times the input price. GitHub does not report a hit rate. A rough check is the share of cache read tokens among all input-side tokens for one model, compared month over month, after you confirm how your export counts the input column. A falling share after a workflow change points to idle gaps, model switches or tool changes. The VS Code team's 94% for Claude is its own harness measurement for agentic workloads, so treat it as a reference point and not a target.
The VS Code team says it is working on flagging actions that quietly raise cost, such as resuming after a long pause with an expired cache or changing reasoning effort mid-session. Until then, count both as cache breaks. Caching is one lever in a wider AI cost management practice for Microsoft environments.
Work with Precio Fishbone
To make GitHub Copilot prompt caching work across teams that mix Anthropic and OpenAI models, a practical first step is a shared default for models and enabled tools. Book a free consultation with our AI team, or email me at par.johansson@preciofishbone.se.
Book a free consultationFrequently Asked Questions
How does GitHub Copilot prompt caching differ between Anthropic and OpenAI?
OpenAI caches the prompt prefix automatically and drops it after idle time. Anthropic caching uses up to four explicit breakpoints that hold while the tools, system prompt and earlier messages stay unchanged and turns arrive within five minutes.
How long does Claude prompt caching keep a cache?
Five minutes by default, renewed at no extra cost each time the cache is used, per Anthropic's documentation. A one-hour option costs twice the base input price; we found no GitHub statement on whether Copilot uses it.
Does extended cache retention on OpenAI models cost extra?
GitHub's rate card lists no cache-write charge for GPT-5.4 or GPT-5.5, and the VS Code team's write-up mentions no extra cost. GPT-5.6 and GPT-6 models do charge to write the cache.
Does switching models mid-session break the cache?
Yes. Each model keeps its own cache, so the next request on a new model pays the full input price for the whole conversation. Finish the task on one model, or start a fresh session for the stronger one.
Do these gains apply outside VS Code?
WebSocket transport is the default for GPT-5.2 and newer across Copilot products, including Copilot CLI and the GitHub app. The retention, breakpoint and tool search results were measured in VS Code, so other clients may differ.