For twenty years we asked how fast the CDN could deliver a response. The better question now is how much of the response the application should ever have to compute.
The CDN is becoming part of the application — and AI is accelerating that shift
For most of its history, a CDN sat in a well-understood place on the architecture diagram:
User → CDN → Origin → Database
It cached static objects, terminated TLS closer to the user, absorbed bursts, and kept the origin from falling over. It made things faster. It did not make architectural decisions.
That framing has quietly stopped being accurate. Applications are API-driven, personalized and increasingly AI-powered, and the dominant cost is no longer bandwidth it is compute. Application CPU, database connections, GPU time and tokens. Once that is true, the interesting question changes from how fast can we deliver this to how far into the stack does this request actually need to travel.
1. Caching is a compute-avoidance strategy, not a latency optimization
Start with something ordinary. A product image is obviously a CDN workload. But the same reasoning extends to application responses:
GET /products/123
Accept: application/json
If that response is identical for a large population of users, it can be cached at the edge under a well-chosen cache key. The arithmetic is not subtle:
without caching: 10,000 requests → 10,000 origin requests
with caching: 10,000 requests → 1 origin request + 9,999 hits
The application did not get faster. The application did less work. That distinction matters, because what you avoid is the expensive part: request parsing, serialization, ORM overhead, connection-pool pressure, query planning, and the container capacity you provisioned to survive peak.
The engineering, though, lives entirely in the cache key and this is where most “we put a CDN in front of it” projects quietly fail:

- Normalize before you key. Sort query parameters, strip tracking parameters (utm_*, fbclid, gclid), lowercase the path where your routing is case-insensitive. An unnormalized key turns one cacheable object into thousands of near-duplicates and tanks your hit ratio.
- Vary is a multiplier, not a switch. Vary: Accept-Encoding is fine. Vary: User-Agent fragments your cache across effectively unbounded cardinality. If you need device-class variation, normalize to a small enumerated set at the edge and vary on that.
- Invalidate by tag, not by URL. Surrogate keys (Fastly), cache tags (Cloudflare, Akamai) let a single product update purge every collection page, search facet and API response that embedded it. Without them, teams default to short TTLs which is just paying for the miss in advance.
- stale-while-revalidate decouples freshness from latency. A 60-second TTL with a 600-second SWR window means the origin sees one revalidation per minute per object rather than a thundering herd at every expiry, and users never wait on it. Pair it with stale-if-error and the edge becomes a availability control as well.
- Microcaching is underrated. A 1–5 second TTL on a hot, “uncacheable” dynamic endpoint sounds pointless until you do the math: at 2,000 RPS, a 2-second TTL removes 99.95% of origin traffic. Nobody notices two seconds of staleness on a trending feed.
- Collapse the misses. Tiered caching and request coalescing mean that a cold object requested simultaneously by 500 edge locations results in one origin fetch, not 500.
None of this is new. What is new is the reason to care: at current compute prices, these are cost-control mechanisms that happen to also reduce latency.
2. Decide workload placement up front, not after the fact
The common pattern is build the application, then put a CDN in front of it. The better one is to decide, per workload, which layer is allowed to answer.
Where each workload should be answered: CDN, edge, application, database or AI layer
The point is not to cache everything. “Show me my current bank balance” is user-specific and time-sensitive; it has to traverse authentication, the account service and the database, and any cache hit there is a correctness bug. “Book seat 14A and take payment” must reach the transactional path it is not idempotent and it is not cacheable at any TTL.
But between those two extremes sits an enormous middle ground that most teams send to origin out of habit: the application shell, the product catalog, category and search-facet pages, pricing for anonymous users, configuration and feature-flag payloads, public API responses. A useful structural pattern here is edge composition serve a cached, cacheable shell to everyone and inject the personalized fragment separately, either through a second authenticated call or through an edge function that stitches identity into the response at the PoP. You get one cache entry serving millions of users instead of one per user.
The objective is not maximum caching. It is correct placement.
3. AI does not extend the cost model. It breaks it.
A conventional request has a predictable cost envelope. An AI request does not.
Traditional request path versus AI request path, and where cost concentrates
Three things change at once:
Cost stops being per-request and becomes per-token. A cached JSON response costs effectively nothing to serve again. A regenerated LLM response costs the full prefill of its context plus every decoded output token, every single time. Worse, the context is usually the expensive part . A RAG prompt carrying eight retrieved chunks can be 10–50x the length of the user’s actual question.
Latency becomes multi-modal. Total latency is now the sum of embedding, ANN search, prompt assembly, prefill and decode. Time-to-first-token is dominated by prefill, which scales with context length; total time scales with output length. Your p99 is no longer a database outlier, it is a long answer.
The pipeline has more failure modes than the thing it replaced. Vector store, embedding model, retrieval ranking, model provider, rate limits, context-window overflow.
Now consider a support assistant. "What is your refund policy?" is a stable, tenant-wide answer. Invoking a model for it thousands of times a day, identically is pure waste. "Why was my refund rejected for order #1234?" is genuinely personalized and genuinely needs retrieval and reasoning.
The principle that falls out of this is simple to state and surprisingly hard to enforce:
Do not invoke inference when the answer can be served from cache, retrieved from a system of record, or precomputed.
A practical corollary: separate prompt caching from semantic caching. Prompt caching (provider-side KV-cache reuse of a shared prefix) cuts prefill cost when many requests share a long system prompt or document. Semantic caching skips the model entirely. They compose well, and teams routinely conflate them.
4. From cache keys to semantic cache keys
Traditional caching asks a syntactic question: is this URL, with this request context, cacheable? AI introduces a semantic one.
These three prompts are different strings with one intent:
“What’s your refund policy?” “Can I get my money back if I cancel?” “What are the rules for getting a refund?”
A URL-keyed cache sees three distinct requests and pays for three generations. A semantic cache embeds the normalized prompt, runs an approximate-nearest-neighbour lookup against previously answered prompts, and reuses the stored response when similarity clears a threshold.
Semantic cache decision flow and the guardrails the cache key must carry
The economics are compelling — an embedding call is orders of magnitude cheaper than a generation — but the failure mode is qualitatively different from a traditional cache miss. A stale object is late. A false semantic hit is confidently wrong, and it looks exactly like a correct answer.
That means the guardrails are not optional:
- Scope is part of the key. Tenant, user entitlement, locale, and plan tier. “What’s my refund policy” means different things to a B2B account and a consumer one. A semantic cache that ignores scope is a cross-tenant data leak with extra steps.
- Version the corpus. If the underlying policy document changes, every cached answer derived from it is invalid. Tag cache entries with source-document versions and purge by tag — the same discipline as surrogate keys, applied to generated content.
- Version the prompt and model. Change the system prompt or swap the model and your cached responses no longer represent what the pipeline would produce. Include a hash of both in the key.
- Tune the threshold against labelled data, and fail open. τ is a precision/recall dial. Too low and you serve wrong answers; too high and you never hit. Log near-misses to find where the boundary actually sits, and when in doubt, fall through to the model.
- Never cache what you cannot re-derive. Anything user-specific, time-sensitive or transactional stays out.
Done carefully, the cache stops storing content and starts avoiding computation. That is a different kind of infrastructure.
5. The edge becomes a decision point
Put those layers together and the request path stops being a straight line. It becomes a sequence of escalating questions, each one more expensive to answer than the last.
The edge as a decision point: cache, edge function, application, data layer, inference
The objective is one sentence: resolve the request at the earliest layer that can correctly answer it. Not the fastest layer, not the closest — the earliest one that is still correct. Correctness is the binding constraint; cost and latency are what you optimize inside it.
6. AI traffic is also breaking assumptions the CDN itself was built on
This is the part that gets least attention and is arguably the most interesting, because it is not about what you build on top of a CDN it is about the cache algorithms underneath it.
Cloudflare, working with researchers at ETH Zürich, published data in April 2026 showing that roughly a third of traffic across their network is automated, and that AI crawlers account for about 80% of self-identified AI bot traffic. The behavioural profile is nothing like a human’s:
- High unique-URL ratio. Human traffic is Zipfian a small set of popular pages carries most requests, which is exactly the distribution LRU was designed for. AI crawlers perform sequential full-site scans, and in modelling of iterative RAG loops the unique-access ratio sits between 70% and 100%.
- No session reuse. Crawlers do not use browser caching, and multiple independent instances each appear as a fresh visitor requesting the same content.
- Crawling inefficiency. A meaningful fraction of fetches from popular crawlers end in 404s or redirects.
The consequence is cache pollution. Long-tail objects that would previously have been evicted get pulled in repeatedly, displacing the popular content human users depend on. Measured hit rate at a single CDN node drops when AI crawler traffic is included, and the standard mitigations — prefetching, cache speculation — get less effective, not more, because there is no locality to predict.
The proposed responses are genuinely architectural. First, replacing LRU for mixed traffic: early experiments suggest eviction policies like SIEVE or S3-FIFO let human traffic hold its hit rate whether or not crawlers are present. Second, and more radically, splitting the cache by traffic class — human traffic served from latency-optimized edge caches, interactive AI traffic (RAG, live summarization) from higher-capacity tiers that tolerate slightly more latency, and bulk training crawls pushed to deep origin-side tiers or queue-admitted and deferred when the infrastructure is under load.
That is a real-world example of a production system where the eviction policy is now a function of who is asking.
7. What the platforms have actually shipped
This is not speculative roadmap material. All three major platforms moved in 2026.
What Akamai, AWS CloudFront and Cloudflare shipped in 2026
Akamai — inference placement as a routing problem. In March 2026 Akamai launched AI Grid intelligent orchestration as part of its Inference Cloud, described as the first global-scale implementation of NVIDIA’s AI Grid reference design. It routes inference across more than 4,400 edge locations plus regional and core GPU capacity, with a workload-aware control plane brokering requests in real time against cost-per-token, time-to-first-token and throughput. Semantic caching and WebAssembly-based serverless compute (EdgeWorkers, Akamai Functions) sit at the edge tier; multi-thousand-GPU Blackwell clusters handle heavy multi-modal and post-training work at the core. The architectural claim is that where inference executes is now a runtime decision, not a deployment-time one.
AWS — the edge as an AI traffic control plane. AWS WAF shipped an AI activity dashboard in February 2026, with Bot Control detection now covering more than 650 distinct bots and agents across categories like AI search crawlers, data collectors and LLM training crawlers. In June it went further with AI traffic monetization: a Monetize rule action that returns HTTP 402 with pricing and accepted payment networks, verifies the agent’s signed payment authorization at the CloudFront edge, then fetches and serves the content — settlement handled by a third-party facilitator. Whatever you think of micropayments, the significant part is architectural: identify, classify, rate-limit, charge and serve, all before the origin is involved.
Cloudflare — convergence of CDN, edge compute and AI gateway. AI Gateway provides caching, rate limiting, model routing, fallback and observability in front of providers including Workers AI, OpenAI, Anthropic and Google, with per-request cache control via headers like cf-aig-cache-ttl. Workers AI runs models on the network itself, Vectorize provides the vector store, and AI Crawl Control and Pay Per Crawl handle the policy side. The stack has effectively merged: CDN + edge compute + AI gateway + inference, behind one control plane.
8. The progression
CDN 1.0 through CDN 5.0
Not every application needs CDN 5.0. Most need CDN 2.0 done properly, which is a more useful observation than it sounds — plenty of teams running on modern platforms have never tuned a cache key. The point of the progression is not a maturity model to climb. It is that the boundary between CDN, application and AI infrastructure has become fluid, and an architecture that treats those as three separately-owned tiers will leave both latency and money on the table.
9. The questions worth asking at design time
The architect’s checklist
Run these in order, and stop at the first layer that can answer correctly:
- Can the edge answer this? If yes, it never reaches the origin.
- Can it be cached or precomputed? If yes, do not recompute it per request.
- Does it need application logic? If not, keep it at the edge.
- Does it actually need inference? If not, do not invoke a model.
- If inference is required, where should it run? Edge, regional or core — driven by latency budget, model size and GPU memory footprint.
- Who is consuming this? Human, crawler, or agent — and should they get the same response, the same cache tier, and the same price?
- What can move closer to the user without becoming incorrect? Correctness first, then cost, then latency.
The goal was never to push everything to the edge. It is to minimize unnecessary movement and computation across the whole architecture.
The CDN of the AI era is not a faster path to the application. It is becoming part of the application — deciding what gets delivered, what gets cached, what reaches the origin, what traffic gets controlled or charged, and increasingly, where computation and inference should happen at all.
The blog is written by Swapnil Vaidya (Lead Cloud Solutions Architect @ Cloud.in)
No comments:
Post a Comment