AI Protection
Protect the routes in your application that call an AI model. WebDecoy checks incoming requests before inference, applies your application rules, and reports whether the application allowed or denied the request. Optional shared quotas, concurrency limits, and model budgets help control repeated and expensive use across application replicas.
Alpha release: the SDKs and hosted runtime are available for pilot integrations. APIs may change; pin your package version and start in observation mode. Real-world detection accuracy and model-cost savings have not been established.
Set up AI Protection → · Open AI Protection in the dashboard
Where it fits
Section titled “Where it fits”Use AI Protection for an authenticated chat endpoint, a generation feature, or an AI-backed API where automated requests can consume your model allowance or application capacity.
The integration runs in your backend. Your application continues to call its existing model provider directly. Application-defined rules run locally; request metadata goes to WebDecoy for cloud detection. Model prompts and responses are not required for these checks.
Request → your authentication and input validation → local rules / optional shared quota / WebDecoy detection → optional concurrency lease and per-attempt budget → your model provider → outcome and usage reports → AI Protection dashboardAn SDK installation and route integration are required. Saving dashboard settings does not install protection into your application.
Available in the Alpha
Section titled “Available in the Alpha”| Capability | Customer value | Integration needed |
|---|---|---|
| Request admission | Evaluate automated traffic before starting inference | Wrap the AI route or enforce the SDK decision |
| Local application rules | Apply your own entitlement and input policies | Rules using authenticated server context |
| Shared account/session quotas | Limit repeated requests across replicas | Configure the quota and trusted identity callback |
| Shared concurrency | Limit simultaneous AI work | Use the concurrency wrapper and track completion of the full stream |
| Per-attempt model budgets | Reserve a maximum allowance before a model call; reconcile confirmed usage | Wrap each model attempt, supply price and token bounds, report final usage |
| Outcome and usage reporting | Inspect decisions, degraded coverage, and model attempts | Keep reporting and completion hooks active |
See controls and verification for behavior during outages, disconnects, and missing usage.
Choose an SDK
Section titled “Choose an SDK”| Runtime | Alpha version | Guide |
|---|---|---|
| Node.js 22.22.3+ / Next.js | 0.1.0-alpha.2 |
Node / Next.js |
| Go 1.26.1+ | v0.1.0-alpha.2 |
Go |
| Python 3.11+ / FastAPI | 0.1.0a1 |
Python / FastAPI |
All three SDKs are public and Apache-2.0 licensed. The hosted detection service is separate. Framework features vary: Python currently supports asyncio and FastAPI/Starlette; its browser-evidence and MCP adapters are not implemented. The Node SDK requires a server runtime and does not support edge runtimes.
Scope of protection
Section titled “Scope of protection”AI Protection handles request admission and runtime resource controls. It does not inspect model text for prompt injection, certify model output safety, or guarantee a provider bill. Your application remains responsible for authentication, authorization, safe tool execution, provider limits, and cancellation.
AI crawler detection covers traffic such as crawlers visiting your site. AI Protection covers requests to your own AI-powered application features.