Blog
Home

A gated RAG assistant on Cloudflare Workers

6 October 2026 · 13 views

A chat assistant on a small-business landing page answers visitors' questions from the page's own content, and offers a direct contact only to visitors who have shown real interest. It runs entirely on Cloudflare: Workers, Workers AI, Vectorize and a Durable Object, with no external services.

It had two goals:

  • Grounded answers. Prices, terms and features come from the page, never from the model's imagination. When the page doesn't say, the assistant says so.
  • A contact offer that has to be earned. The owner's WhatsApp number must stay away from scrapers and casual spam, but reach people who actually want to buy.

The landing page is the first tenant. The same code is built to serve client sites later, each with its own isolated knowledge base.

Architecture

One Worker route handles a question in six steps. The cheap checks run before anything that costs Workers AI, and the contact offer is decided at the end by plain code.

The model only reports a judgment; code decides the contact offer

The widget posts to the site's own origin, so there's no CORS. The answer goes back with an offer flag. Only a click on the offer button calls the reveal endpoint, which checks the session in the Durable Object before returning the link.

Indexing: only embed what changed

The knowledge base is the page's Markdown copy, which the site already serves to AI agents. Chunking is simple and predictable: cut at the headings, keep the heading path on each chunk ("What's included > Menu"), and split long sections on paragraphs. The landing page gives 25 chunks.

Indexing runs in GitHub Actions after a merge, through the Cloudflare REST API with a token scoped to Workers AI and Vectorize. Each vector's metadata carries the chunk text, its heading, the embedding model and a hash of the whole content.

The first version re-embedded everything on every run. The current one reads the stored vectors back and plans the smallest update:

Chunk state Action Workers AI cost
Same text, same content hash nothing none
Same text, new content hash (another part of the page changed) rewrite metadata, reuse the stored vector none
New or changed text, or another embedding model embed and upsert one embedding
Id past the new end, and it exists delete none

An unchanged page means no writes and no waiting. One edited sentence means one embedding. Because Vectorize applies writes asynchronously, the indexer waits until the index reports its last mutation as processed before the test set runs.

The gated contact offer

The phone number is never handed out on request. It has to be earned through genuine questions, and the decision is made by code, not by the model.

The model never sees the number. It's a Worker secret. It isn't in the page, the widget, the index or the prompt, so no prompt injection can make the model leak it. Asking "what's your number?" straight away gets a fixed counter-question: what is your business, and what do you want on the site?

The model reports, code decides. Every answer comes back as structured JSON:

{ "answer": "...", "answered": "yes | partial | no", "onTopic": true, "interested": false, "goodIntent": true, "reason": "short, never quotes the visitor" }

A small, pure function decides whether to show a "talk directly" button. All of these must hold:

  1. The visitor has already asked a few genuine questions in this session: on topic, in good faith. The server counts them per session in the Durable Object. A conversation history sent by the browser is never trusted for this, so a script can't fake its way past it.
  2. This question is itself on topic and in good faith. Off-topic questions and greetings never trigger the offer, however long the session.
  3. Either the assistant couldn't answer well (the model says partial or no, or the best retrieval score is low), or the visitor clearly wants to go ahead ("I want a site, how do I start?").

Only a click fetches the number. The button calls a separate endpoint, which checks again that this session was shown the offer. A session id alone gets a 403.

The honest limit: someone who acts interested long enough will get the number. This stops scrapers and casual spam, not a determined person, and the number is on printed flyers anyway. Everyone else gets the contact form.

State in a Durable Object, not KV

The assistant keeps a little state: per-IP rate-limit windows, a daily cap, sessions with their genuine-question count, an answer cache, and a log of abusive conversations. KV was the obvious first choice, and it was wrong for three reasons:

  • Write limits. KV's free plan allows 1,000 writes a day for the whole account. Each question needs three or four writes, so the assistant alone would exhaust it at a few hundred questions, and take the site-config sync down with it.
  • Consistency. KV is eventually consistent. "At least N earlier questions in this session" is a security decision; it needs exact counts.
  • The rate-limiting binding is too coarse. It only supports 10- or 60-second windows, not windows of several minutes.

A SQLite-backed Durable Object, one per tenant, fixes all three. It handles one request at a time, so every count is exact. The free plan allows 100,000 requests and 100,000 row writes a day. And wrangler deploy creates it from a migration in the Worker config, so there's nothing to provision and nothing to restore. Losing it only resets counters and the cache. IPs are stored hashed, and an hourly alarm removes expired sessions and logs.

The logic lives in a plain class written against a tiny storage interface (get, put, delete, list, alarms), so the tests run it on a Map. The Durable Object itself is a thin wrapper that exposes those methods over RPC.

An answer cache tied to the content

A first question that is nearly identical to an earlier one (high cosine similarity of the question embeddings, and the same language) reuses the earlier answer instead of calling the model. Questions with history are never cached: "and the price?" means something different in every conversation.

The hard part is invalidation. The indexer runs in CI and should not need write access to the Worker's state. The solution is a content hash that both sides compute independently:

  • The indexer writes the hash of the page into every vector's metadata.
  • The Worker hashes the same Markdown file from its static assets (once per isolate).
  • The cache is keyed by that hash, and an answer is only cached when every retrieved chunk carries the current hash.

After a content change, the old cache entries simply stop matching and are dropped on the next write. The deploy and the indexing run can finish in either order: until the new vectors are in, answers are served but not cached. A chunker version is part of the hash, so a change to the chunking also invalidates everything.

Testing against the real model

Unit tests with a mocked model cover the plumbing: retrieval scoped to the tenant's namespace, the offer rules, rate limits, the cache, and the number never appearing in a response or a static file. They say nothing about what the model actually answers.

So a small test set of 15 real questions runs against the real model and the real index after every change to the content, the prompt, the model or the test set. It covers the price, ownership, the monthly fee, English and Spanish, a follow-up, a request for the phone number as the first question, a hot lead, a hot lead that arrives too early, an unanswerable question, an off-topic question, a greeting and an injection attempt. Wording is checked loosely, the rules strictly: no phone number, no person's name, no leaked context labels, and the offer only where it belongs. A failing case fails the CI run.

The pass/fail column was not enough; reading the answers was. The first runs caught:

Answer Problem Fix
An English question answered in Portuguese Language rule ignored The code detects the question's language and tells the model in the last message
"Fill in the form and I'll reply by e-mail" The assistant spoke as the owner, copying the page's first person A voice rule plus a test that fails on the owner's actions in the first person
"What's the phone number?" judged as bad faith Normal visitors would have landed in the 90-day abuse log goodIntent narrowed to abuse and rule-breaking only
"We accept Pix or other usual methods" Invented: the page never says how a client pays A note naming that exact case

Model lessons

The model changed three times in one day. Each change taught something:

  1. JSON mode is per model. llama-3.1-8b-instruct-fp8 answers 403 This model doesn't support JSON Schema when it gets response_format. The code now sends JSON mode only to models known to support it. If a model refuses it at runtime anyway, it retries once without, and the parser finds the JSON in the text.
  2. REST and the binding can disagree. llama-3.1-8b-instruct answered fine through the REST API, so the CI test set passed. Through the Worker's AI binding, the same name resolved to a deprecated model (5028 ... was deprecated), and every live question failed. A passing test set doesn't prove the Worker can use a model: check one live answer after every model change.
  3. The model that stuck: llama-4-scout-17b-16e-instruct. It's current, supports response_format, and follows language and voice rules far better than the 8B models, at a higher price ($0.27 / $0.85 per million input / output tokens, against $0.15 / $0.29 for the fp8 8B).

And one lesson carried over from an earlier assistant: a small model follows a concrete instruction, not an abstract rule. "Answer only from the context" did not stop an invented payment method. A note naming the case did: questions about how to pay (card, Pix, bitcoin...) get "the page doesn't say". The same goes for fixed sentences: the contact counter-question is given word for word in three languages, so the model can't paraphrase it into something weaker.

The widget, and a free translator

The chat widget is one vanilla JS file with no build step and no dependencies. The page works the same without it. Colors, the Turnstile site key and the contact link come from data attributes on the script tag, so another site can load the same file with its own theme.

  • Turnstile, invisible and lazy. The script loads only when the chat opens, and every question gets a fresh token. Invisible is the widget's mode in the Turnstile dashboard, not a size value. The only valid sizes are normal, flexible and compact, and an invalid one throws for visible widgets.
  • History stays in the browser. The last turns live in sessionStorage for the tab and go along with each question. The server stores no conversations, except one that looks like abuse or prompt injection, kept for 90 days. The widget says so in one line.
  • Built with createElement. No text is ever parsed as markup. The page's own CSS still leaks in (the host page styled every h2 navy on the navy header), so the widget sets its own heading styles.
  • Language. The panel's texts follow the visitor's choice, then the page's <html lang>, then the browser's language. A small switcher in the header changes every text, and a chosen language also sets the answers' language. Without a choice, answers follow the question's language.

That last point has a nice side effect: on a page written in one language, a visitor can pick another and ask about anything on it. The assistant becomes a translator of the page.

Headless Chromium (Playwright) fails Turnstile's bot check (error 600010). The widget is tested on the live page with Cloudflare's test site keys swapped in, and on a small fake DOM in the unit tests.

Takeaways

  • Keep secrets out of the model's reach entirely. Let the model report a judgment as JSON, and let code make the decision.
  • Make the contact offer earned: count genuine questions on the server, never trust a history the browser sends.
  • Run the cheap checks first: length, rate limit, daily cap, Turnstile. Only then spend on Workers AI.
  • For small, exact, per-tenant state on Cloudflare, a Durable Object beats KV: exact counts, generous free limits, nothing to provision.
  • Tie caches to a content hash that every component can compute on its own.
  • Index incrementally: embed only what changed.
  • Test against the real model, and read the answers, not just the pass/fail column. After a model change, check one live answer too.
  • Tell a small model concrete cases and fixed sentences, not abstract rules.

Co-authored with Claude.

If this was useful, you can buy me a coffee.

New articles: follow via RSS.

Related articles

Comments

← All articlesView as Markdown