Market Prices

BTC Bitcoin
$75,927.3 -2.11%
ETH Ethereum
$2,405.13 -3.47%
SOL Solana
$97.41 -3.85%
BNB BNB Chain
$714.9 -0.76%
XRP XRP Ledger
$1.31 -7.33%
DOGE Dogecoin
$0.0804 -3.29%
ADA Cardano
$0.1961 -4.15%
AVAX Avalanche
$7.33 -2.42%
DOT Polkadot
$0.9552 -3.59%
LINK Chainlink
$10.84 -5.33%

Event Calendar

{{年份}}
18
03
unlock Sui Token Unlock

Team and early investor shares released

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

28
03
unlock Arbitrum Token Unlock

92 million ARB released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

12
05
halving BCH Halving

Block reward halving event

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

💡 Smart Money

0xa90c...62bd
Arbitrage Bot
+$3.6M
63%
0x1ac4...6a9e
Arbitrage Bot
+$3.7M
76%
0x6466...18e9
Experienced On-chain Trader
-$2.0M
69%

🧮 Tools

All →

Cache Hits and Liquidity Traps: What ZCode’s 98.6% DeepSeek Hit Rate Really Tells Us About AI Agent Economics

CryptoRay
Macro

The 98.6% Illusion

Dax Raad, co-founder of OpenCode, published a 48-hour slice of client-side cache hit rates for DeepSeek traffic. The number that broke his brain was not his own product’s respectable 97.86%. It was Zhipu’s ZCode, an Agentic Development Environment you have probably never heard of, sitting at 98.60%. Claude Code / CLI, the incumbent that every VC pitch deck now benchmarks against, dragged in at 89.31%. Raad said, verbatim: “I don’t know what ZCode is, but it’s doing a really good job.”

But here is the trap. Reading that tweet, you are supposed to conclude that ZCode has built some kind of superior inference-layer engineering. That its client-side caching mechanism has cracked a deeper code. That this twenty-person team in Beijing has somehow surpassed the combined forces of Anthropic and the open-source tinkerers. The follow-up math appears to confirm it: DeepSeek charges roughly 50x more for cache misses than hits. At ZCode’s hit rate, input costs are only about 27% of what Claude Code users pay, given the same token structure. As a macro watcher who spent two decades auditing smart contracts and stress-testing DeFi collateral, I see something else entirely. I see a classic failure-mode misread. A market that confuses a derived metric with an underlying property. A moment where everyone is about to fund the wrong horse.

Context: The Token Tariff and the New Compute Regime

DeepSeek’s pricing model is not an afterthought; it is a deliberate monetary policy. For input tokens that hit the prefix cache, DeepSeek charges a negligible rate — often USD 0.014 per million tokens in recent disclosed tiers. For misses, that rate jumps to roughly USD 0.7 per million, a fiftyfold cliff. This is not a transparency move. It is a liquidity mechanism, precisely analogous to how Ethereum’s EIP-1559 burns base fees during congestion, or how rollup DA layers charge per-calldata-byte. The cache is the new mempool. The hit rate is the new blockspace utilization. And the client — ZCode, OpenCode, Claude Code — is just a node that decides which transactions get bundled into the cache’s favor.

ZCode is an ADE built by Zhipu, the AI company behind the GLM-5.2 model family. In crypto terms, think of it as a self-custody wallet that defaults to a particular chain (GLM-5.2) but also supports RPC providers for other chains. It has a plugin architecture, a terminal, a diff viewer, and a rather excessive amount of dark-mode polish. The team is small, self-funded, and has not issued a token. That alone should make every DAO treasury manager suspicious. The fact that it beat Claude Code on a 48-hour window measured by a competitor — not by Zhipu — is even more suspicious. But the data is what it is. The real question is not who won. The real question is what the metric actually measures.

Core: The Cost Surface and Its Hidden Volatility

Let me perform the kind of failure-mode stress test I used on MakerDAO’s stability fees in 2020. You have a cost function: C = h (1) + (1-h) 50, where h is the cache hit rate, and 1 is the unit cost of a cache-hit token, 50 the unit cost of a miss. At ZCode’s h=0.986, C=1.686 tokens per unit of input. At Claude Code’s h=0.8931, C=6.238. That gives the celebrated 27% cost advantage. But what happens when h is not static? What happens when the traffic distribution shifts?

Cache hit rates are not an inherent property of the client. They are a dynamic function of the workload shape. If all your developers are writing boilerplate — imports, standard CRUD endpoints, common test scaffolding — the prefix cache will be hammered with identical token sequences. Hit rates climb. If your team is working on a novel protocol, a unique Solidity library, or an obscure system in a 30-year-old codebase, every prompt is a cold start. Hit rates collapse. This is exactly why I refuse to evaluate rollups by their gross data posting volume. A rollup that posts 10 MB per day because it is running a Twitter clone and a rollup that posts 10 MB per day because it is settling cross-chain liquidity swaps are not equivalent. The DA market gets the same price per byte, but the second rollup is paying for security; the first is paying for a diary.

Dax Raad’s 48-hour sample is diary traffic. The default behavior of most coding agents is to call the same five utility functions, re-read open files, and repeat error messages verbatim. ZCode’s 98.60% hit rate is not an extraordinary engineering achievement; it is a normal distribution of mainstream development tasks fed into a system that aggressively caches the first hundred tokens of the user prompt. OpenCode V2’s 97.86% is equally impressive, and equally misleading. The spread between 98.6% and 89.31% tells you more about the user base than the software. Claude Code users are disproportionately experienced builders who work on heterogeneous, legacy-heavy projects. ZCode users are, right now, mostly GLM-5.2 early adopters who are testing standard Python and TypeScript exercises. In a zero-to-one market, every user is a tutorial.

Now let’s translate this to blockchain infrastructure, because that is where the real cost will land. Within five years, AI agents will be transacting on-chain. They will need crypto wallets, gas management, and verifiable compute. DeepSeek’s cache is a form of off-chain state. The hit rate is the efficiency of that state. But in DeFi, we learned the hard way that off-chain efficiency is often just deferred on-chain catastrophe. Celsius had an impressive “borrow against yield” hit rate until the market moved. Three Arrows Capital had a pristine “counterparty selection” hit rate until Luna depegged. In 2021, I published a breakdown showing that 85% of NFT floor prices were supported by wash-trading bots. The floor price was perfectly stable. The stability was a lie.

A 98.6% cache hit rate, sustained over 48 hours, is the AI equivalent of a wash-traded floor. It is real for the person who experiences it, but it is not a foundation for a business model. Let me prove it with a simple extrapolation. Suppose ZCode’s hit rate remains 98.6% across a month, but then a new trending framework is released, and every one of its users suddenly switches to a new prompt format. The hit rate drops to 80% overnight. The cost multiplier goes from 1.686 to 0.80 + 0.20*50 = 10.8. That is a 6.4x increase in input cost. The developer does not change anything. The developer just upgrades their framework. This is precisely the kind of “liquidity shock” that has killed more crypto projects than hacks have.

My experience auditing early Ethereum smart contracts in 2017 taught me to look for reentrancy. Here, the reentrancy is in the cache. A cache hit is a snapshot of prior state. When the state changes, the cache becomes a stale liability. Smart contracts handle this with checks-effects-interactions. AI clients handle it with a hash prefix and a sliding window. But the principle is the same: if you cache too eagerly, you execute against the wrong data. ZCode’s 98.6% is not a sign that ZCode is smart; it is a sign that ZCode is conservative. It probably flushes its cache aggressively, which makes its hit rate appear high because the cached prefixes are very short. Claude Code, by contrast, tries to maintain longer context windows, which increases the risk of cache invalidation and lowers the nominal hit rate. The “loss” in hit rate is actually a version of insurance. Claude Code is paying a 6.2x cost to maintain a more flexible state model.

To test this hypothesis, I would need longer-term data across multiple model families. Dax Raad shared a 48-hour window. That is a single block. In on-chain analysis, we never trade a single block. We look at volume-weighted average over days, and even that is subject to manipulation. In 2020, I led a team that stress-tested MakerDAO against a 40% ETH drop. We simulated the liquidation cascade, and found that 15% of collateral would be wiped out within hours. The stability fees looked great at normal volatility. The system was designed for bull markets. ZCode’s cache hit rate is designed for bull markets too — bull markets for AI tooling, where every user is a new developer following the same five Stack Overflow patterns. The moment the bull market ends, the developers diversify. They solve harder problems. They use domain-specific languages. The hit rate drops. The cost skyrockets.

But the more insidious issue is the incentive structure. DeepSeek’s 50x penalty for cache misses is a systemically risky cliff. It rewards clients for gaming the cache prefix. A client can artificially increase its hit rate by truncating user prompts to the first 100 tokens, or by maintaining a local lookup table that substitutes a generic prompt for the actual one. This is the equivalent of a rollup that posts only signatures to the DA layer, claiming “validium” while actually being a glorified permissioned database. In 2022, I spent three months tracing the opaque lending flows between Luna and UST. I saw how a stablecoin with a 99% “stability rate” could still propagate systemic risk through centralized exchanges. The same mechanism is happening here. A high cache hit rate is not a measure of robustness. It is a measure of how much of your workload is boring.

Let’s do the actual capital mathematics. If a developer runs 1 million prompt input tokens per day, a hit-rate of 98.6% costs approximately 1.686 0.014 USD = 0.0236 USD per million tokens. Claude Code at 89.31% costs 6.238 0.014 = 0.0873 USD per million. On a per-day basis, the difference is about 0.064 USD per million tokens. That is nothing. For a serious engineering team running 10 billion tokens per day, the daily difference is 640 USD. Over a year, 233,600 USD. That is a real cost, and it matters. But the variance matters more. If the hit rate of ZCode drops from 98.6% to 90% — a move that can happen with a single framework change — the cost per day for that same team jumps to (0.9 + 0.150)0.014 = 0.0798 USD per million, or 798 USD per 10 billion. Annualized, 291,270 USD. The difference between ZCode and Claude Code narrows to zero, and ZCode loses its entire advantage. The 27% cost edge is not an edge at all; it is a call option on workload homogeneity. And options decay.

This is why I insist on measuring “liquidity-adjusted” metrics. In crypto, we calculate effective fee revenue by subtracting wash trading. In AI, we need to calculate effective token cost by subtracting cache gaming. Dax Raad’s data is self-reported from the client side. That means it is exactly the kind of metric that a developer can game. You can force a cache hit by sending the same prompt twice. You can inflate your hit rate by setting a very short cache window, so only the first 10 tokens hit, and the rest is a “miss” — but a miss with a 10-token prefix is cheap, so the client can claim a high hit rate while actually paying more. The metric is unverifiable. It is not on-chain. It is a screenshot of a liquidity pool with no proof of reserves.

Cache Hits and Liquidity Traps: What ZCode’s 98.6% DeepSeek Hit Rate Really Tells Us About AI Agent Economics

Contrarian: The Decoupling Myth

Now for the contrarian angle that will get me hate mail from two sides. The crypto-native thought leadership is already working overtime to frame this as evidence that “AI agents will disrupt coding” or “China is winning the AI race.” Both are wrong. The real story is that client-side caching is a solved problem. It has been solved for three decades in HTTP through CDNs and browser caches. The 98.6% hit rate is not innovation. It is the same as your local DNS resolver caching a popular domain. What is new is that DeepSeek has monetized the cache as a first-class pricing tier, and that has created a new kind of “gas war” — but only for the 1.4% of traffic that misses. The great drama of the 50x penalty is theater. It convinces users that the cache is magic. It convinces investors that cache-optimized clients have proprietary technology. The reality is that any client can shuffle prompts around to achieve high hit rates, just as any exchange can report high volume by paying market makers for footprint.

The decoupling thesis is even more seductive. Some analysts argue that high cache hit rates decouple AI costs from model quality. That is false. The cache hit rate is a product of the traffic, not the model. ZCode’s GLM-5.2 might be a terrible code generator; its users might be caching prompts precisely because they keep failing and repeating the same request. In fact, a low hit rate could indicate a high-quality model that solves novel problems, forcing the client to send new prompts each time. Claude Code’s 89.31% might be a feature, not a bug. It means Claude Code’s users are tackling fresh tasks, not re-submitting the same broken request. In crypto terms, a node that processes 100% of transactions through its own cache is a sequencer that never experiences an order flow imbalance. That is not a sign of health. It is a sign of censorship.

My audits of DeFi protocols have repeatedly proven that over-optimized metrics hide systemic risk. In 2021, NFT wash trading was not a bug in the market; it was the market’s attempt to create a fake liquidity cushion. ZCode’s high hit rate is a similar cushion. It exists because the user base is homogeneous. But as soon as ZCode becomes popular, its user base will diversify. The hit rate will fall. The cost will rise. And the product will be exposed as having no fundamental advantage over OpenCode or Claude Code beyond a good dark theme.

The “50x penalty” itself is a regulatory red flag. Imagine a bank that charges you 50x more for a withdrawal that requires a human teller. That is not a fee schedule; it is a behavioral enforcement tool. DeepSeek is using the cache hit rate to force a specific prompt structure on its users. It rewards you for repeating yourself and penalizes you for exploring. That is the opposite of what an AGI-native coding environment should do. It is, in fact, the same kind of “theater” as KYC in most crypto projects. The compliance cost is passed entirely to honest users, while the whales — the ones with enough compute to pre-warm the cache — get a free pass. In this case, the honest users are the developers solving novel problems. They pay 50x for the privilege of being original. ZCode’s 98.6% hit rate is just a reflection that its users are not being original. That is not a success story. That is a warning signal.

Takeaway: Cycle Positioning, Not Scoreboard Watching

Let’s step back to the macro picture. The AI coding tool market is following the same adoption curve as every crypto sub-sector: an initial wave of overhyped metrics, a false sense of leadership, and then a brutal consolidation where the only survivors are those with genuine moats. In crypto, the moat was never the block size; it was the liquidity network. In AI coding, the moat will not be the cache hit rate; it will be the ability to handle heterogeneous workloads without punishing users for novelty. The current scoreboard — ZCode 98.6%, OpenCode 97.86%, Claude Code 89.31% — is a snapshot of a bull market in tutorial-driven development. It will invert in a bear market.

My advice to institutional allocators is simple. Do not fund an AI client because it has a high cache hit rate. That is the same mistake as funding a rollup because it posts high data volume without checking whether the data is meaningful. Instead, ask for the miss distribution. Ask for the hit rate across different programming languages, across greenfield versus maintenance tasks, across access to multiple model families. If a client cannot disaggregate its hit rate, it is hiding something. And if you are an engineer using ZCode, enjoy the 27% cost saving while it lasts. But store your prompts. Back them up. Because the day you start building something new, the cache will not save you.

Chaos is just data that has not been stress-tested yet. This 48-hour cache hit-rate snapshot is precisely that: un-stress-tested chaos. Do not confuse it with signal. In the next six months, someone will publish a 30-day sample, and the numbers will shift. The market will panic. The market will chase a new leader. I will still be here, reading the raw logs, checking the miss frequencies, and wondering why we keep building call options on homogeneous workloads when the future is heterogeneous by definition. The only sustainable deployment is one that treats every prompt as a cold start. The only sustainable cache is the one that understands it is a temporary liquidity cushion, not a business model.

ZCode wins the 48-hour race. I am not impressed. I want to see how it survives the 48-hour race where everyone stops copying the same OpenZeppelin boilerplate and starts writing original, unpublished code. That is the real test. And it is the test that every metric so far has failed.

Fear & Greed

51

Neutral

Market Sentiment

Altseason Index

42

Bitcoin Season

BTC Dominance Altseason

Market Cap

All →
# Coin Price
1
Bitcoin BTC
$75,927.3
1
Ethereum ETH
$2,405.13
1
Solana SOL
$97.41
1
BNB Chain BNB
$714.9
1
XRP Ledger XRP
$1.31
1
Dogecoin DOGE
$0.0804
1
Cardano ADA
$0.1961
1
Avalanche AVAX
$7.33
1
Polkadot DOT
$0.9552
1
Chainlink LINK
$10.84

🐋 Whale Tracker

🔴
0xf5fc...b301
2m ago
Out
18,897 BNB
🔵
0x0049...a17a
3h ago
Stake
16,120 BNB
🔵
0xf2bb...ba3e
12m ago
Stake
1,533 ETH