Skip to content

AI Protection on Cloudflare Workers

AI Protection checks requests before your application invokes a model. Local rules evaluate trusted server context; WebDecoy provides remote bot detection and shared quotas. The Workers adapter integrates into your existing AI endpoint.

Feature Workers status
Request admission and remote bot detection Experimental
Application-defined local rules Experimental
Shared account/session quotas Experimental; requires compatible WebDecoy service
Asynchronous decision reporting Best-effort, attached to the request lifecycle
Streaming response passthrough Original response body passes through
Concurrency leases Not supported yet
Spending budgets and final model usage settlement Not supported yet

The adapter does not provide prompt-injection filtering, model/tool authorization or proof that a caller is human. Keep your application’s authentication, permissions, origin checks and input limits.

  1. Open AI Protection and select your property.

  2. Create an API key scoped to that property with Write Detections permission in Settings → API keys. Copy the property ID from the setup screen.

  3. Obtain the pilot SDK package containing the Workers adapter and install it in your application. For a provided tarball:

    Terminal window
    npm install /path/to/webdecoy-ai-protection-<version>.tgz

Start in observation mode. Cloud detection enforcement also requires an entitled account and the property’s cloud mode to permit enforcement. Each local rule and shared quota has its own mode; an enforced local rule can deny a request even while cloud detection is observing.

Use Node compatibility for the SDK’s crypto, Buffer and IP-validation APIs. The tested Wrangler configuration is:

name = "my-ai-application"
main = "src/worker.js"
compatibility_date = "2026-01-01"
compatibility_flags = ["nodejs_compat"]
[vars]
WEBDECOY_URL = "https://ai-protection.webdecoy.com"
WEBDECOY_PROPERTY_ID = "your-property-id"

Store WEBDECOY_KEY and WEBDECOY_SUBJECT_SECRET as Worker secrets. Generate a random subject secret of at least 32 characters and keep it consistent across instances that share quota subjects.

Terminal window
npx wrangler secret put WEBDECOY_KEY
npx wrangler secret put WEBDECOY_SUBJECT_SECRET

For local development, use a gitignored .dev.vars file instead. API keys and subject secrets must stay in the Worker; never include them in browser code.

Create the protection object inside fetch() for each incoming request. The returned wrapper is bound to that request, so protect() takes a callback and trusted context without another Request argument.

import {createWorkerAIProtection} from '@webdecoy/ai-protection/workers';
export default {
async fetch(request, env, ctx) {
// Your existing auth, permission, origin and bounded input checks.
const user = await authenticateAndValidate(request);
const protect = createWorkerAIProtection(request, {
webdecoyUrl: env.WEBDECOY_URL,
webdecoyKey: env.WEBDECOY_KEY,
propertyId: env.WEBDECOY_PROPERTY_ID,
subjectSecret: env.WEBDECOY_SUBJECT_SECRET,
scopeId: 'support-chat',
route: '/api/chat', // Fixed route template, without user identifiers.
protectionMode: 'observe',
resolveClientIP: trustedClientIP,
accountQuota: {
ruleId: 'chat', limit: 20, windowSeconds: 60, mode: 'observe',
subject: context => ({accountId: context.accountId})
}
}, ctx);
return protect(
() => callModelAndReturnResponse(request.signal),
{accountId: user.id}
);
}
};

authenticateAndValidate, trustedClientIP and callModelAndReturnResponse are your application functions, not SDK exports. The model callback must return a standard Response and honor your cancellation signal. The callback runs only when admission allows the request. Use authenticated server state for quota subjects and local rules; never trust a user-supplied account ID or plan.

On direct Cloudflare ingress, your resolver can return request.headers.get('cf-connecting-ip'). Worker-to-Worker requests have different header behavior: define a trust contract for that path instead of assuming the header identifies the original user. Return null when no trustworthy address is available; cloud detection is then skipped with degraded coverage. See Cloudflare’s client-IP header behavior.

  1. Send an authenticated test request through your deployed Worker in observation mode.
  2. Open AI Protection for the same property and refresh. Check the stored AI verdict and the Application decisions report; a successful model response alone does not prove protection is connected.
  3. Confirm the expected client IP and check that requests are not degraded by a missing IP, unavailable account binding or detector outage.
  4. Test a local denial and an enforced shared quota in staging. Confirm the model callback is not invoked for denied requests and reports persist.
  5. Exercise simultaneous requests and a detector outage before enabling cloud enforcement. Detector failures allow requests by default; a closed quota failure mode is a separate, explicit availability choice.

Shared quotas live in WebDecoy’s service; they do not depend on persistent Worker memory. Schema-2 idempotent quota recovery is optional and requires a compatible, migrated service. Persist server-generated operation IDs when recovery must survive request loss.

Do not cache the protection object in module scope or reuse it across incoming requests. Each instance owns its account lookup, observations and pending reports. This avoids sharing request-owned promises or an execution context between callers; it also means each incoming request performs a fresh account lookup.

The adapter automatically attaches reports using task => ctx.waitUntil(task). No separate lifecycle hook is needed. Reporting is best-effort and bounded (1 second by default, at most 10 seconds); it is not durable delivery. Cloudflare permits up to 30 seconds of extension after response completion or client disconnect. See Cloudflare’s execution context.

Streaming bodies pass through unchanged. Admission reports describe callback and response creation, not stream completion or final token usage. Propagate cancellation to your model provider and validate disconnect behavior in staging. Concurrency limits, spending budgets and linked-account usage reporting that requires budget telemetry are not supported by this Workers adapter yet.

For explicit decisions, call protect.check(trustedContext), then protect.report(decision, outcome) once. protect.flush() waits for this request’s pending reports only. Never interpret missing reports as proof that requests were safe.

The managed edge-sensor installer does not install this application adapter. Integrate the adapter into your existing Worker; do not replace its AI handler with an edge-sensor Worker.