Observed 30-day snapshot
Prompt caching can avoid repeatedly processing a stable prefix such as repository context, instructions, or tool definitions. This analysis ranks normalized model families by the share of their observed prompt-side tokens reported as cache reads. It is a usage-composition measurement—not a claim about provider-side hit logic, latency, or money saved.
What the data says
Cohort cached-input share
95.4%
Weighted by prompt-side tokens across eligible models.
Eligible-model median
95.0%
Each eligible model receives equal weight.
Observed cache reads
684.07B
Reported tokens served from an existing prompt cache.
Models meeting floor
56
At least 2 developers and 100.0K prompt-side tokens.
claude-opus-4-7
3 developers · 605.0M prompt-side tokens
glm-5.2-high
4 developers · 230.8M prompt-side tokens
model-ba4f0d6a6174
3 developers · 384.9M prompt-side tokens
claude-sonnet-5
73 developers · 38.61B prompt-side tokens
claude-opus-5
81 developers · 147.37B prompt-side tokens
big-pickle
11 developers · 426.0M prompt-side tokens
gpt-5.6-sol
78 developers · 328.05B prompt-side tokens
claude-opus-4-6
5 developers · 1.71B prompt-side tokens
glm-5.3
4 developers · 304.8M prompt-side tokens
claude-opus-5-thinking-high
6 developers · 1.03B prompt-side tokens
We calculate each model's cached-input share as cache-read tokens ÷ (uncached input + cache-write + cache-read tokens). A result of 75% means three quarters of the measured prompt-side tokens were reported as reads from an existing cache. Output tokens are outside that denominator.
A high cached-input share can indicate that long, stable prefixes are being reused. It can also reflect a particular agent architecture, session length, provider reporting format, or a small number of very large cached contexts. The ratio does not tell us whether the reused context was relevant or whether a shorter prompt would have been more efficient.
Savings depend on the provider's dated cache-read and cache-write prices, cache lifetime, write frequency, and ordinary input rate. Some providers charge extra to create a cache. Use the linked savings calculator with your own rates before turning this observed share into a budget estimate.
Window. The page requests the current 30-day aggregate. The snapshot timestamp is displayed above. Counts come from daily usage rows belonging to listed developers and exclude rows held from public rankings.
Normalization. Date-suffixed and gateway-prefixed raw model identifiers are merged into normalized model families before token buckets and distinct developers are counted.
Sample floors and coverage. A model must have at least 2 distinct participating developers and 100,000 combined uncached-input, cache-write, and cache-read tokens. The API returns at most 60 models ranked by total tokens, so ratio leaders are leaders among returned rows meeting the floor. These rules reduce one-person and tiny-denominator artifacts; they do not make the cohort statistically representative.
Reporting limits. Token categories are taken from agent/provider usage records and may not be reported consistently by every integration. A zero cache-read value can mean no observed cache reuse or unavailable cache telemetry; this aggregate cannot distinguish those cases.
Cohort limit. Public participation is opt-in. Results describe this cohort's coding sessions, not market share or a survey of all AI developers. They should be treated as directional observed behavior.
Aggregated model, agent, trend, skill, and tool-call totals from participating developers.
Explains what the CLI aggregates, what it does not collect, and the limits of the public cohort.
| 3 |
| claude-sonnet-5 | 97.6% | 37.68B | 910.4M | 10.3M | 129.5M | 73 |
| claude-opus-5 | 97.6% | 143.80B | 3.43B | 132.0M | 519.2M | 81 |
| big-pickle | 97.4% | 415.0M | 0 | 11.0M | 1.4M | 11 |
| gpt-5.6-sol | 97.4% | 319.38B | 482.6K | 8.67B | 995.4M | 78 |
| claude-opus-4-6 | 97.1% | 1.66B | 48.8M | 118.3K | 12.7M | 5 |
| glm-5.3 | 97.1% | 295.8M | 0 | 8.9M | 417.9K | 4 |
| claude-opus-5-thinking-high | 97.0% | 1.00B | 24.5M | 6.9M | 4.1M | 6 |
| hy3:free | 97.0% | 137.3M | 0 | 4.3M | 1.9M | 2 |
| deepseek-v4-flash | 96.9% | 5.29B | 81.9K | 170.6M | 14.6M | 19 |
| claude-opus-4-8 | 96.8% | 35.99B | 1.16B | 17.1M | 140.2M | 65 |
| k3 | 96.8% | 1.40B | 0 | 46.5M | 7.4M | 2 |
| glm-5.3-flash | 96.6% | 261.8M | 0 | 9.1M | 516.8K | 3 |
| claude-sonnet-4-6 | 96.5% | 224.4M | 8.1M | 11.6K | 1.8M | 14 |
| deepseek-v4-pro | 96.4% | 438.9M | 0 | 16.2M | 1.0M | 12 |
| kimi-k3 | 96.4% | 343.4M | 0 | 12.9M | 702.3K | 8 |
| gpt-5.6-terra | 96.3% | 16.43B | 0 | 627.8M | 57.2M | 77 |
| gpt-5.4 | 96.3% | 339.6M | 0 | 13.0M | 1.8M | 12 |
| claude-sonnet-5-thinking-high | 96.3% | 292.3M | 9.0M | 2.3M | 1.6M | 3 |
| codex-auto-review | 96.0% | 6.40B | 0 | 264.4M | 18.5M | 16 |
| gpt-5 | 95.9% | 6.90B | 0 | 291.2M | 13.4M | 10 |
| gpt-daybreak-blue-latest | 95.7% | 2.46B | 0 | 109.2M | 9.7M | 4 |
| gpt-5.3-codex-spark | 95.5% | 447.5M | 0 | 21.2M | 2.4M | 11 |
| deepseek-v4-flash-free | 95.4% | 5.95B | 0 | 283.9M | 35.9M | 28 |
| composer-2.5 | 95.2% | 622.4M | 0 | 31.2M | 3.7M | 6 |
| gpt-5.6-sol-high | 95.0% | 691.1M | 30.0M | 6.6M | 2.8M | 3 |
| x-preview-f-free | 95.0% | 6.81B | 0 | 360.6M | 20.4M | 28 |
| muse-spark-1.2-contributor-free | 94.8% | 436.5M | 0 | 23.9M | 795.7K | 11 |
| cursor-grok-4.6-xhigh-fast | 94.1% | 946.7M | 2.3M | 57.4M | 6.6M | 5 |
| cursor-grok-4.6-medium-fast | 93.9% | 173.9M | 0 | 11.3M | 753.2K | 4 |
| composer-2.5-fast | 93.7% | 884.1M | 0 | 59.5M | 8.3M | 23 |
| claude-haiku-4-5-20251001 | 93.5% | 1.40B | 96.6M | 432.4K | 16.5M | 40 |
| model-237bf97a0500 | 93.5% | 201.7M | 0 | 14.1M | 749.8K | 3 |
| gpt-5.5 | 93.2% | 10.47B | 0 | 762.7M | 26.1M | 69 |
| cursor-grok-4.6-xhigh | 92.6% | 850.6M | 93.6K | 68.2M | 6.7M | 3 |
| default | 92.5% | 5.87B | 9.5M | 463.9M | 35.3M | 20 |
| claude-fable-5 | 92.5% | 39.26B | 1.19B | 2.02B | 312.9M | 51 |
| deepseek-v4-flash-vision-exp | 92.4% | 377.5M | 0 | 31.0M | 775.0K | 2 |
| cursor-grok-4.6-low | 91.7% | 269.8M | 0 | 24.4M | 1.3M | 2 |
| ox-alpha | 91.6% | 1.12B | 0 | 103.2M | 6.9M | 10 |
| cursor-grok-4.5-high | 91.4% | 580.5M | 404.4K | 54.4M | 4.2M | 9 |
| gpt-5.6-luna | 91.2% | 15.08B | 19.3M | 1.43B | 50.4M | 57 |
| cursor-grok-4.6-high-fast | 90.8% | 2.43B | 436.8K | 247.5M | 19.4M | 15 |
| cursor-grok-4.6-high | 90.2% | 1.15B | 466.6K | 125.1M | 7.5M | 15 |
| cursor-grok-4.5-medium | 89.9% | 146.9M | 0 | 16.6M | 1.2M | 4 |
| glm-5.2 | 89.1% | 1.86B | 0 | 226.6M | 16.3M | 9 |
| cursor-grok-4.5-high-fast | 88.3% | 1.42B | 147.5K | 188.2M | 13.8M | 20 |
| gpt-5.3-codex | 87.7% | 128.6M | 0 | 18.1M | 830.3K | 2 |
| gemini-3.6-flash-high | 81.0% | 146.4M | 0 | 34.2M | 680.1K | 2 |
| cursor-grok-4.6-medium | 79.2% | 343.7M | 611.1K | 89.8M | 4.3M | 6 |
| nemotron-3-ultra-free | 72.7% | 109.5M | 0 | 41.1M | 304.5K | 8 |
| gemini-3.6-flash | 56.3% | 971.9M | 0 | 755.7M | 2.5M | 2 |
| step-3.7-flash:free | 12.2% | 37.3M | 0 | 267.7M | 2.2M | 3 |
| gemini-3.7-flash | 6.2% | 527.3M | 0 | 7.93B | 3.3M | 2 |