# Three layers of caching in front of a container that scales to zero

The goal for this site was always "fast, cheap, and mine." Cheap meant letting the blog's API scale to zero. Fast, on top of that, took three layers of caching, each one added because the previous one leaked somewhere. And the first two, it turned out, made the blog fast without making it cheap.

## What it costs

Everything behind koorevaar.com sits in one Azure resource group. The pages themselves are served by Cloudflare Workers on the free plan; Azure runs the APIs and demos behind them. Here's one month of Azure costs, August 25 to September 25, per resource:

| Resource | What it is | Cost |
|---|---|---|
| blog-service-api | Container App: the blog's API | €7.28 |
| cosmos-koorevaar | Cosmos DB: articles, comments, the AI assistant's vector store | €4.37 |
| hangman-game | Container App: the hangman game | €2.13 |
| text-intelligence-api, llm-adapter | Container Apps: the sentiment analysis and language detection demo | €1.69 |
| lab-* (39 of them) | Container Apps: one short-lived CKA practice lab per session, so far only my own test runs | €1.22 |
| Blog Function app | Flex Consumption: comment moderation and the Change Feed syncs, with storage and Log Analytics | €0.92 |
| cka-lab-api | Container App: starts and stops the practice labs (test runs too) | €0.67 |
| capacity-advisor, youtube-channel-guard, the rest | | €0.23 together |
| **Total** | | **€18.51** |

The LLM calls aren't in the table, because they don't run on Azure. Comment moderation, the AI assistant (embedding the question, then generating the answer), the sentiment and language detection demo, the hangman game and the CV request in the contact form all call Cloudflare Workers AI, which gives 10,000 neurons a day for free. Together they have never reached that limit. That's not only because traffic is low: the Workers in front of those calls have rate limiters, and comment moderation has its own daily scoring quota, so a bot or a burst of traffic can't burn through the free allocation.

Every Container App here is the smallest size, 0.25 vCPU and 0.5 GiB of memory. Most demos scale down to zero replicas when nobody uses them, and stop costing anything until the next visitor. You can see it in the table: they cost somewhere between a few cents and a euro each, roughly in proportion to how much people actually use them. The capacity advisor cost seven cents this month.

The hangman game is the exception. It normally keeps one replica running at all times (`min-replicas 1`), and a full month of that costs about €7. It scaled to zero for part of this period, which is why it shows €2.13 here, and even then €1.12 of it is billed at the idle rate: a replica waiting for a player.

Cosmos DB runs on the account's free tier: 1,000 RU/s and 25 GB. Articles, comments and the Change Feed lease containers share one database with exactly 1,000 RU/s, so the blog's data costs nothing. When I added comments, fitting inside that pool without new billable capacity was an explicit design constraint. Nearly all of the €4.37 is the AI assistant's vector store, which has to live in its own database with its own small autoscale throughput. That one is billed, knowingly.

Which leaves the odd one out. The blog's API is more than half of all Container Apps spend: €7.28, about what the hangman game costs in a month when it never scales down at all. The blog API does scale to zero, for a blog that changes a few times a week, and in this period that saved practically nothing. The rest of this post is about why.

## The constraint

The blog's API is a small .NET app, the same 0.25 vCPU and 0.5 GiB as everything else, with `min-replicas 0` and `max-replicas 1`. Scale to zero is supposed to keep it cheap, and the price is the cold start. I measured it on 2026-09-04: about 30 seconds from the first request to the first response.

Most visitors arrive from a LinkedIn link, on an article nobody has opened for a while. Thirty seconds of blank page is long enough for most people to close the tab.

The obvious fix is `min-replicas 1`: no scale to zero, no cold start, and a constant bill for a replica that sits idle most of the day. Every app here is the same size, so the hangman game already shows what that bill would be: about €7 a month. Simple, and a perfectly reasonable choice. I wanted to see how far caching could get me first.

## Layer 1: stale-while-revalidate, not a TTL

All blog reads already pass through `api-proxy`, a Cloudflare Worker that calls the API, renders the Markdown to HTML and hands it to the UI Worker. That's the natural place for a cache.

A plain TTL cache doesn't solve this problem. It moves it: when the entry expires, the next reader waits on origin, and after a quiet spell origin means a cold start. Somebody still pays.

So the cache works the other way around:

1. If there's a cached copy, serve it immediately. Always, no matter how old it is.
2. If that copy is older than `REVALIDATE_INTERVAL_MS`, start a background refresh with `ctx.waitUntil()`. The response has already gone out, so no reader waits for it.
3. The refresh replaces the cached entry only if origin answered with a good response. A timeout, a 429 or a 500 leaves the old copy in place, and the next request past the interval just tries again.

Cloudflare's Cache API (`caches.default`) has no notion of "never expire, just rate-limit the re-check", so the Worker stamps each cached response with its own `X-Swr-Cached-At` header and compares against that. The interval is a limit on how often to bother origin, not an expiry date. An entry never goes cold from age alone.

The first design note said about three minutes, to be checked against a real measurement. With the cold start measured at 30 seconds, and paid only in the background where no reader sees it, I set it to one minute: fresher content after an edit, at the cost of waking the container a bit more often. Keep that last part in mind.

## Trade-offs I accepted instead of solving

**View counts undercount.** The API increments `viewCount` on every real GET. With the cache in front, only the background refreshes reach it: at most one per minute per route, however many readers came by in between. During a burst of traffic the counter is too low. I decided that's fine for a personal blog, not a correctness bug worth building around.

**No purge on write.** When I publish or edit an article, nothing actively clears the cache. The next refresh picks it up within about a minute of the next visit. For a blog where I'm the only author, that's short enough that purge logic would be complexity for nothing.

**Comments are not cached at all.** Someone who just posted a comment expects to see it on reload, comments aren't what people click through from LinkedIn, and they're low traffic anyway.

## Layer 2: a durable fallback, because the Cache API is per colo

Cloudflare's Cache API isn't one global cache. Every data center (colo) has its own. A popular article can be cached in Amsterdam and completely unknown in Frankfurt, and a reader routed to Frankfurt gets a cache miss, which means a cold start.

This happened in production. An article I had edited heavily, clicked from the homepage, landed on a colo that had never served that slug. The full cold start, on exactly the kind of visit this whole design was meant to protect.

The fix was a second, durable layer in Cloudflare KV, which is global rather than per colo:

- Every successful origin fetch also writes a snapshot to KV: the latest 10 articles for the list (without content, since the list never shows it), and one entry per article for detail pages.
- On a cache miss, the Worker races the real origin fetch against a 2 second timer. If origin answers in time, that's the response. If not, the reader gets the KV snapshot right away, and the origin fetch keeps running in the background to fill the cache for the next visitor.
- A colo that nobody visits for a while can hold a very old copy of the list. When such a copy is past the interval, the Worker first checks whether the KV snapshot is newer and serves that instead.

The one case left is a truly first-ever request: a brand new article, no cache anywhere, no snapshot yet. That one still waits for origin. In practice that visitor is usually me, checking my own new post.

## What the first two layers don't fix: the bill

The first two layers hide the cold start from readers. They do nothing to stop the container from starting.

Every background refresh is a real request to origin, and if it arrives while the container is asleep, it wakes it. Once awake, it stays up for a few minutes before scaling back down, and all of that is billed. With several colos, crawlers and bots all hitting the blog, a refresh here and a refresh there were enough to wake the container over and over, even when the content hadn't changed in days. Nobody waited for it anymore, but it ran anyway, re-confirming things I already knew. That's the €7.28 at the top of the table.

## Layer 3: push on write, instead of asking on read

The fix is to stop reads from asking at all. Content only changes when I write, so the write should push. A Cosmos DB Change Feed trigger reacts to every article write and pushes the new state straight into KV, and the Worker reads from KV without asking the container anything. The post counter was the first to move on September 16 ([Three Versions of a Blog-Post Counter](https://blog.koorevaar.com/articles/three-versions-of-a-blog-post-counter) has the whole story), and it never touches the container anymore. The article list followed on September 23, at least on paper: the push had a bug that made it fail on every run until I fixed it on September 26, and the list still re-checks origin in the background, so for now the push only keeps KV fresh. Either way, almost all of that €7.28 is from before. Making the list rely on the push alone is the next step, then individual article pages, and they bring a problem I haven't solved yet: once reads no longer reach the API, `viewCount` stops counting entirely, not just undercounting.

The push isn't free either. It moves work from the Container App to an Azure Function: the same Flex Consumption Function app that already moderates comments, so there's no new app to host. That whole Function app cost €0.92 over the same month, most of it Log Analytics, and it only runs when something is actually written. Whether the pushes save more Container App time than they cost is the comparison I still have to make, once there's a full month after the change.

One more thing that moved rather than disappeared: cost. The KV free tier allows 1,000 writes a day, shared with other side projects on the same Cloudflare account. A bulk edit of every article recently triggered a "50% of daily KV writes used" warning, because every changed article fans out into several KV writes. Every writer now compares against what's already stored and skips the write when nothing a reader would notice has changed. Before any bulk operation, I estimate the KV writes first.

## Where it stands

- The blog API still scales to zero. `min-replicas 1`, like the hangman game has, stays on the table if the caching layers ever cost more attention than a small monthly bill would.
- No blog reader waits for a cold start, except the very first visitor to a brand new article.
- View counts are knowingly too low.
- The counter never touches the container anymore; the list and article pages are next. That's the part that should bring the biggest line on the bill down. Next month's export will tell whether it did, and whether the Function doing the pushing costs less than the container time it saves.

None of this is best practice for its own sake. It started as a cost constraint and a 30 second number, and each of the three layers came from a specific case where the previous one leaked.

---

*Co-authored with Claude.*
