AI Protection on Cloudflare Workers
AI Protection checks requests before your application invokes a model. Local rules evaluate trusted server context; WebDecoy provides remote bot detection and shared quotas. The Workers adapter integrates into your existing AI endpoint.
Supported features
Section titled “Supported features”| Feature | Workers status |
|---|---|
| Request admission and remote bot detection | Experimental |
| Application-defined local rules | Experimental |
| Shared account/session quotas | Experimental; requires compatible WebDecoy service |
| Asynchronous decision reporting | Best-effort, attached to the request lifecycle |
| Streaming response passthrough | Original response body passes through |
| Concurrency leases | Not supported yet |
| Spending budgets and final model usage settlement | Not supported yet |
The adapter does not provide prompt-injection filtering, model/tool authorization or proof that a caller is human. Keep your application’s authentication, permissions, origin checks and input limits.
1. Prepare your property and SDK
Section titled “1. Prepare your property and SDK”-
Open AI Protection and select your property.
-
Create an API key scoped to that property with Write Detections permission in Settings → API keys. Copy the property ID from the setup screen.
-
Obtain the pilot SDK package containing the Workers adapter and install it in your application. For a provided tarball:
Terminal window npm install /path/to/webdecoy-ai-protection-<version>.tgz
Start in observation mode. Cloud detection enforcement also requires an entitled account and the property’s cloud mode to permit enforcement. Each local rule and shared quota has its own mode; an enforced local rule can deny a request even while cloud detection is observing.
2. Configure your Worker
Section titled “2. Configure your Worker”Use Node compatibility for the SDK’s crypto, Buffer and IP-validation APIs. The tested Wrangler configuration is:
name = "my-ai-application"main = "src/worker.js"compatibility_date = "2026-01-01"compatibility_flags = ["nodejs_compat"]
[vars]WEBDECOY_URL = "https://ai-protection.webdecoy.com"WEBDECOY_PROPERTY_ID = "your-property-id"Store WEBDECOY_KEY and WEBDECOY_SUBJECT_SECRET as Worker secrets. Generate a
random subject secret of at least 32 characters and keep it consistent across
instances that share quota subjects.
npx wrangler secret put WEBDECOY_KEYnpx wrangler secret put WEBDECOY_SUBJECT_SECRETFor local development, use a gitignored .dev.vars file instead. API keys and
subject secrets must stay in the Worker; never include them in browser code.
3. Protect your AI endpoint
Section titled “3. Protect your AI endpoint”Create the protection object inside fetch() for each incoming request.
The returned wrapper is bound to that request, so protect() takes a callback
and trusted context without another Request argument.
import {createWorkerAIProtection} from '@webdecoy/ai-protection/workers';
export default { async fetch(request, env, ctx) { // Your existing auth, permission, origin and bounded input checks. const user = await authenticateAndValidate(request);
const protect = createWorkerAIProtection(request, { webdecoyUrl: env.WEBDECOY_URL, webdecoyKey: env.WEBDECOY_KEY, propertyId: env.WEBDECOY_PROPERTY_ID, subjectSecret: env.WEBDECOY_SUBJECT_SECRET, scopeId: 'support-chat', route: '/api/chat', // Fixed route template, without user identifiers. protectionMode: 'observe', resolveClientIP: trustedClientIP, accountQuota: { ruleId: 'chat', limit: 20, windowSeconds: 60, mode: 'observe', subject: context => ({accountId: context.accountId}) } }, ctx);
return protect( () => callModelAndReturnResponse(request.signal), {accountId: user.id} ); }};authenticateAndValidate, trustedClientIP and callModelAndReturnResponse are
your application functions, not SDK exports. The model callback must return a
standard Response and honor your cancellation signal. The callback runs only
when admission allows the request. Use authenticated server state for quota
subjects and local rules; never trust a user-supplied account ID or plan.
On direct Cloudflare ingress, your resolver can return
request.headers.get('cf-connecting-ip'). Worker-to-Worker requests have different
header behavior: define a trust contract for that path instead of assuming the
header identifies the original user. Return null when no trustworthy address
is available; cloud detection is then skipped with degraded coverage. See
Cloudflare’s client-IP header behavior.
4. Verify the installation
Section titled “4. Verify the installation”- Send an authenticated test request through your deployed Worker in observation mode.
- Open AI Protection for the same property and refresh. Check the stored AI verdict and the Application decisions report; a successful model response alone does not prove protection is connected.
- Confirm the expected client IP and check that requests are not degraded by a missing IP, unavailable account binding or detector outage.
- Test a local denial and an enforced shared quota in staging. Confirm the model callback is not invoked for denied requests and reports persist.
- Exercise simultaneous requests and a detector outage before enabling cloud enforcement. Detector failures allow requests by default; a closed quota failure mode is a separate, explicit availability choice.
Shared quotas live in WebDecoy’s service; they do not depend on persistent Worker memory. Schema-2 idempotent quota recovery is optional and requires a compatible, migrated service. Persist server-generated operation IDs when recovery must survive request loss.
Request lifecycle and streaming
Section titled “Request lifecycle and streaming”Do not cache the protection object in module scope or reuse it across incoming requests. Each instance owns its account lookup, observations and pending reports. This avoids sharing request-owned promises or an execution context between callers; it also means each incoming request performs a fresh account lookup.
The adapter automatically attaches reports using task => ctx.waitUntil(task).
No separate lifecycle hook is needed. Reporting is best-effort and bounded
(1 second by default, at most 10 seconds); it is not durable delivery.
Cloudflare permits up to 30 seconds of extension after response completion or
client disconnect. See Cloudflare’s execution context.
Streaming bodies pass through unchanged. Admission reports describe callback and response creation, not stream completion or final token usage. Propagate cancellation to your model provider and validate disconnect behavior in staging. Concurrency limits, spending budgets and linked-account usage reporting that requires budget telemetry are not supported by this Workers adapter yet.
For explicit decisions, call protect.check(trustedContext), then
protect.report(decision, outcome) once. protect.flush() waits for this
request’s pending reports only. Never interpret missing reports as proof that
requests were safe.
Related setup
Section titled “Related setup”- Cloudflare edge sensor for broader request visibility.
- API keys for property-scoped credentials.
- AI traffic monitoring for crawler analytics.
The managed edge-sensor installer does not install this application adapter. Integrate the adapter into your existing Worker; do not replace its AI handler with an edge-sensor Worker.