Observed 30-day snapshot
Prompt caching can avoid repeatedly processing a stable prefix such as repository context, instructions, or tool definitions. This analysis ranks normalized model families by the share of their observed prompt-side tokens reported as cache reads. It is a usage-composition measurement—not a claim about provider-side hit logic, latency, or money saved.
What the data says
Cohort cached-input share
95.2%
Weighted by prompt-side tokens across eligible models.
Eligible-model median
93.9%
Each eligible model receives equal weight.
Observed cache reads
654.97B
Reported tokens served from an existing prompt cache.
Models meeting floor
56
At least 2 developers and 100.0K prompt-side tokens.
gpt-reserve
3 developers · 1.41B prompt-side tokens
claude-sonnet-5
76 developers · 48.66B prompt-side tokens
claude-opus-5
89 developers · 182.07B prompt-side tokens
grok-bot-cua
8 developers · 686.7M prompt-side tokens
gpt-5.6-sol
86 developers · 236.81B prompt-side tokens
gpt-6-astra
63 developers · 25.23B prompt-side tokens
gpt-daybreak-blue-latest
6 developers · 11.00B prompt-side tokens
claude-opus-4-6
7 developers · 2.46B prompt-side tokens
kimi-k3
9 developers · 299.0M prompt-side tokens
k3
3 developers · 1.04B prompt-side tokens
We calculate each model's cached-input share as cache-read tokens ÷ (uncached input + cache-write + cache-read tokens). A result of 75% means three quarters of the measured prompt-side tokens were reported as reads from an existing cache. Output tokens are outside that denominator.
A high cached-input share can indicate that long, stable prefixes are being reused. It can also reflect a particular agent architecture, session length, provider reporting format, or a small number of very large cached contexts. The ratio does not tell us whether the reused context was relevant or whether a shorter prompt would have been more efficient.
Savings depend on the provider's dated cache-read and cache-write prices, cache lifetime, write frequency, and ordinary input rate. Some providers charge extra to create a cache. Use the linked savings calculator with your own rates before turning this observed share into a budget estimate.
Window. The page requests the current 30-day aggregate. The snapshot timestamp is displayed above. Counts come from daily usage rows belonging to listed developers and exclude rows held from public rankings.
Normalization. Date-suffixed and gateway-prefixed raw model identifiers are merged into normalized model families before token buckets and distinct developers are counted.
Sample floors and coverage. A model must have at least 2 distinct participating developers and 100,000 combined uncached-input, cache-write, and cache-read tokens. The API returns at most 60 models ranked by total tokens, so ratio leaders are leaders among returned rows meeting the floor. These rules reduce one-person and tiny-denominator artifacts; they do not make the cohort statistically representative.
Reporting limits. Token categories are taken from agent/provider usage records and may not be reported consistently by every integration. A zero cache-read value can mean no observed cache reuse or unavailable cache telemetry; this aggregate cannot distinguish those cases.
Cohort limit. Public participation is opt-in. Results describe this cohort's coding sessions, not market share or a survey of all AI developers. They should be treated as directional observed behavior.
Aggregated model, agent, trend, skill, and tool-call totals from participating developers.
Explains what the CLI aggregates, what it does not collect, and the limits of the public cohort.
| 89 |
| grok-bot-cua | 97.4% | 669.0M | 17.6M | 37.1K | 2.8M | 8 |
| gpt-5.6-sol | 97.3% | 230.44B | 35.2M | 6.34B | 782.1M | 86 |
| gpt-6-astra | 97.2% | 24.53B | 23.1M | 673.8M | 80.1M | 63 |
| gpt-daybreak-blue-latest | 97.2% | 10.69B | 0 | 312.2M | 40.2M | 6 |
| claude-opus-4-6 | 97.0% | 2.39B | 73.1M | 46.6K | 19.9M | 7 |
| kimi-k3 | 97.0% | 289.9M | 0 | 9.1M | 462.5K | 9 |
| k3 | 96.8% | 1.01B | 0 | 33.3M | 6.0M | 3 |
| gpt-5.6-terra | 96.8% | 17.95B | 50.3K | 596.2M | 62.3M | 84 |
| claude-fable-5-1 | 96.5% | 12.36B | 438.9M | 8.8M | 65.9M | 31 |
| gpt-5.6-sol-high | 96.3% | 1.83B | 59.0M | 11.7M | 5.9M | 2 |
| muse-spark-1.3-contributor-free | 96.3% | 2.11B | 0 | 81.6M | 3.1M | 15 |
| claude-sonnet-5-thinking-high | 96.1% | 540.9M | 18.7M | 3.2M | 2.0M | 3 |
| model-d0038e30b975 | 96.0% | 409.8M | 0 | 17.1M | 1.7M | 2 |
| deepseek-v4-pro | 96.0% | 500.3M | 5.0M | 15.9M | 1.5M | 11 |
| claude-opus-5-thinking-high | 95.9% | 1.29B | 46.1M | 8.3M | 6.1M | 14 |
| gpt-5.6-luna | 95.9% | 13.21B | 15.3M | 554.2M | 51.0M | 64 |
| claude-opus-4-8 | 95.7% | 14.92B | 656.0M | 7.8M | 61.5M | 53 |
| deepseek-v4-flash | 95.7% | 4.97B | 0 | 223.9M | 12.9M | 12 |
| composer-2.5 | 95.3% | 1.15B | 0 | 57.6M | 7.0M | 13 |
| x-preview-f-free | 95.0% | 7.01B | 0 | 370.2M | 21.0M | 34 |
| minimax-m3 | 94.5% | 398.5M | 0 | 23.1M | 1.2M | 3 |
| deepseek-v4-flash-free | 94.4% | 2.49B | 0 | 148.6M | 17.4M | 23 |
| gpt-5.3-codex-spark | 94.4% | 322.9M | 0 | 19.3M | 2.2M | 13 |
| mimo-v2.5-free | 93.9% | 403.5M | 0 | 26.1M | 2.4M | 13 |
| muse-spark-1.2-contributor | 93.9% | 783.4M | 0 | 50.7M | 1.3M | 3 |
| cursor-grok-4.6-medium-fast | 93.8% | 511.8M | 0 | 33.7M | 2.5M | 9 |
| claude-sonnet-4-6 | 93.6% | 1.57B | 100.8M | 6.0M | 10.5M | 17 |
| claude-haiku-4-5-20251001 | 93.5% | 1.58B | 109.0M | 511.2K | 20.4M | 36 |
| cursor-grok-4.6-xhigh-fast | 93.5% | 1.63B | 2.7M | 111.6M | 13.4M | 10 |
| cursor-grok-4.6-xhigh | 93.4% | 2.01B | 226.8K | 141.6M | 14.3M | 9 |
| composer-2.5-fast | 93.1% | 598.5M | 0 | 44.5M | 5.2M | 32 |
| muse-spark-1.2-contributor-free | 92.7% | 2.58B | 0 | 203.2M | 5.0M | 22 |
| glm-5.3 | 92.6% | 826.0M | 0 | 65.6M | 2.2M | 11 |
| cursor-grok-4.6-high | 92.2% | 4.47B | 896.0K | 375.1M | 28.5M | 26 |
| deepseek-v4-flash-vision-exp | 92.1% | 414.0M | 0 | 35.6M | 816.1K | 4 |
| default | 91.6% | 4.43B | 3.3M | 400.9M | 25.8M | 28 |
| cursor-grok-4.6-low | 91.6% | 274.0M | 0 | 25.0M | 1.3M | 6 |
| ox-alpha | 91.6% | 1.14B | 0 | 104.3M | 7.0M | 12 |
| claude-fable-5 | 91.5% | 32.49B | 1.02B | 2.01B | 285.5M | 43 |
| cursor-grok-4.6-high-fast | 90.9% | 5.32B | 1.2M | 528.7M | 43.8M | 30 |
| claude-opus-4-7 | 89.5% | 1.02B | 115.0M | 5.0M | 3.1M | 9 |
| big-pickle | 89.4% | 392.6M | 0 | 46.4M | 4.0M | 21 |
| gpt-5.5 | 89.3% | 5.37B | 0 | 641.9M | 12.4M | 78 |
| muse-spark-1.3-contributor | 88.9% | 288.4M | 0 | 36.0M | 375.8K | 4 |
| grok-bot-default | 88.4% | 2.93B | 2.0M | 382.1M | 20.9M | 18 |
| cursor-grok-4.6-medium | 84.9% | 1.18B | 697.3K | 209.2M | 9.7M | 22 |
| glm-5.3-flash | 83.2% | 885.3M | 0 | 179.2M | 4.6M | 13 |
| cursor-grok-4.5-high | 82.3% | 247.4M | 404.4K | 52.9M | 2.1M | 15 |
| grok-bot-automation | 80.5% | 574.7M | 2.9M | 136.5M | 7.1M | 14 |
| claude-opus-5-thinking | 77.6% | 1.70B | 27.2M | 464.3M | 5.1M | 2 |
| nemotron-3-ultra-free | 71.7% | 248.7M | 0 | 97.9M | 848.9K | 9 |
| gemini-3.8-flash | 20.0% | 151.5M | 0 | 604.2M | 735.0K | 4 |
| gemini-3.7-flash | 10.2% | 905.8M | 0 | 7.99B | 3.6M | 2 |