Envoy AI Gateway Plugin for Waldur Site Agent
This plugin integrates Envoy AI Gateway with Waldur so that LLM inference access can be sold, provisioned, metered, and billed through the Waldur marketplace. It provisions per-customer API keys from marketplace orders and reports usage back to Waldur for metering and billing. Token/cost budgets are enforced by Waldur (report usage -> mastermind pauses the resource, or a single key, when a limit is reached -> the agent blocks the keys), not by the gateway.
The plugin is Kubernetes-native: API keys live in Kubernetes Secrets that the Envoy Gateway
SecurityPolicy reads.
At a glance
| Entry point | Group | Role |
|---|---|---|
envoy |
waldur_site_agent.backends |
order processing, membership sync |
envoy-usage |
waldur_site_agent.backends |
reporting |
Modes: order_process, membership_sync, report, event_process (for API key commands).
| Operation | Behaviour |
|---|---|
| Create resource | Registers the resource and its gateway endpoint; the agent then issues its API keys |
| Terminate resource | Removes every key of the resource from both Secrets |
| Update limits | No-op — Waldur enforces limits by pausing the resource from reported usage |
| Pause / restore | Moves the resource's keys to the blocked Secret and back |
| Downscale | Same as pause — keys are blocked; a key has no partial-capacity state |
| Usage reporting | envoy: No-op. envoy-usage: usage from the warehouse API |
Features
- API key lifecycle: provision, pause, restore, and terminate gateway API keys directly from Waldur marketplace orders
- Per-key governance: request, rotate, pause, resume and delete one key at a time from the portal. The resource's other keys keep serving
- Two-secret pause model: suspend access at the authentication layer by moving a key between an active and a blocked Secret — no key regeneration on restore
- Usage reporting: report token usage from a pluggable usage warehouse back to Waldur for metering and billing, per resource and, where the usage shipper records it, per key
- In-cluster or external: run inside the cluster (in-cluster config) or against a kubeconfig context for local development
Overview
The plugin exposes two backends that are normally paired on a single composed offering:
- Management backend (
envoy): theorder_processing_backend. Owns the API key lifecycle: provision, pause, restore and terminate for the whole resource, and Waldur's per-key commands. - Usage reporting backend (
envoy-usage): thereporting_backend. Reads token usage from a usage warehouse and submits it to Waldur, per resource and, where the rows name a key, per key.
Token pricing itself lives in Waldur: the reporting backend submits raw token counts, and the
offering plan prices them (e.g. €/token). Everything is keyed off the resource's backend_id:
its keys are the Secret entries <backend_id>-<n>, and its usage is recorded under the
backend_id itself.
For deployments that price usage upstream instead (per-model rates, discounts), the sibling
cscs-dwdi-inference reporting backend is the cost-model alternative: it reports a pre-priced
token_cost component that Waldur records as-is rather than raw token counts. Pair it in place of
envoy-usage when the warehouse — not Waldur — owns pricing.
Backend Types
Management Backend (envoy)
Envoy Gateway's SecurityPolicy authenticates requests against API keys stored as clientID: key
entries in a Kubernetes Secret. This backend manages those entries.
Key lifecycle — to pause and restore a key without regenerating it, two Secrets are kept and entries move between them:
- create (order): register the resource under the backend id the order processor generates
(
{allocation_prefix}{resource slug}) and surface the gateway endpoint on it. The agent then generates the resource's keys, adds each as a<backend_id>-<n>: keyentry to the active Secret, and reports it to Waldur. - pause: move every key of the resource active → blocked — authentication now fails with 401.
- restore: move the resource's keys blocked → active, except keys paused on their own.
- terminate: remove every key of the resource from both Secrets.
Per-key lifecycle. Each key is one Secret entry, <backend_id>-<n>, so the backend declares
supports_resource_api_key_lifecycle and carries out Waldur's commands for a single key. The agent
applies each command to the Secrets first and then acknowledges it to Waldur:
| Command | Acknowledgement | What happens in the Secrets |
|---|---|---|
| create | set_key |
A new entry under the next free <backend_id>-<n>; blocked on a paused resource |
| rotate | set_key |
The value is overwritten in place; a key left paused by an erred pause serves again |
| pause | set_paused |
The entry moves to the blocked Secret as <client_id>.paused |
| resume | set_ok |
The entry moves back to the active Secret, or to a plain blocked entry on a paused resource |
| update | set_ok |
Nothing, as Waldur enforces limits; a key left paused by an erred pause serves again |
| delete | set_deleted |
The entry is removed from both Secrets |
A new key never reuses a client_id Waldur has held for the resource, including a deleted key's, because Waldur attributes usage by client_id. A key resumed on a paused resource comes back with the other keys when the resource is restored.
Waldur accepts these commands only on an offering with the enable_api_key_provisioning plugin
option set; without it, keys can still be rotated. A key command is not an order, but the agent
picks it up alongside the orders:
- In
order_processmode, on the next cycle, after the orders (everyWALDUR_SITE_AGENT_ORDER_PROCESS_PERIOD_MINUTES, 5 by default), however recently Waldur issued it. - In
event_processmode withstomp_enabled: true, as soon as Waldur sends it. If the message is lost, the reconciliation sweep applies the command once it has been pending for 30 minutes, on the sweep's next run (everyWALDUR_SITE_AGENT_RECONCILIATION_PERIOD_MINUTES, 60 by default).
As with orders, order_process skips an offering with stomp_enabled: true.
Resource pause and key pause. Pausing a whole resource and pausing one key are independent,
and the name of an entry in the blocked Secret records which one holds it back. The gateway only
reads the active Secret; the blocked Secret is where a switched-off key's value waits, so it can
serve again without being regenerated. Take a resource abc with the keys abc-1, abc-2 and
abc-3:
1 2 3 4 | |
<client_id>.paused— the key is paused on its own. Only a resume, rotate or update brings it back.<client_id>— the key is held back because its resource is paused. Restoring the resource moves these entries, and only these, back to the active Secret. This matters because the membership sync restores every resource that is not paused on every cycle: without the suffix, a key paused on its own would serve again within one cycle.<backend_id>.resource-paused— records the resource pause itself; it is a flag, not a key. A key added to the resource checks it to decide whether it lands active or blocked. The keys alone cannot answer that when the resource has no keys left, or when every key left is paused on its own.
No additional Secret is needed.
Rotate and update settle a key as OK in Waldur, so after either one the key serves, unless its resource is paused. Waldur allows both on an Erred key, and an Erred key can still be held by its own pause when the pause was applied but its acknowledgement failed.
Limits and models. A key's limits work like the resource's limits. The plugin stores and enforces nothing: the reporting backend reports each key's usage, and Waldur pauses the key once its usage reaches a limit. The gateway has no per-key model rule, so a command that would restrict a key to some models fails with an error rather than leaving an unrestricted key that Waldur shows as restricted. Leave a key's model list empty on this backend.
Kubernetes objects touched: Secrets (get/patch) in the configured namespace.
Limit enforcement — the gateway does not cap usage. The reporting backend submits usage to
Waldur; when a LIMIT component's reported usage reaches its limit and the offering sets
plugin_options.action_on_usage_limit: pause, mastermind pauses the resource and the agent blocks
its keys (active -> blocked Secret). When usage drops back below the limit, the resource is
unpaused and the keys are restored. A key's own limit works the same way for that key alone (see
Limits and models above).
Usage Reporting Backend (envoy-usage)
A read-only backend that reports token usage from a usage warehouse — any HTTP service that exposes the endpoints below (for example a small collector that aggregates the gateway's per-request usage metrics). Usage is reported per resource and, when the rows name the key they came through, per key as well (see Per-key usage).
API endpoints used:
GET /usage-month?from=YYYY-MM&to=YYYY-MM&client_id=...— per-client_idusage rows for a month range; response shape{"usage": [{"client_id": "...", "input_tokens": N, "output_tokens": N}]}, where a row may also carrykey_client_idGET /health— liveness check used byping()
The backend maps the warehouse's token fields onto the offering's components, reporting only the
components the offering defines. It implements get_usage_report_for_period(), so historical
usage can be bulk-loaded with the core waldur_site_load_historical_usage command.
Configuration
The two backends are typically combined on one composed offering and share a single
backend_settings block. The keys for each backend are listed separately below, followed by a
combined example.
Management backend settings (envoy)
| Setting | Required | Default | Description |
|---|---|---|---|
namespace |
yes | — | Namespace holding the api-key Secrets |
gateway_url |
yes | — | Public base URL of the gateway, used for the resource endpoint |
apikey_secret |
no | envoy-ai-gateway-apikeys |
Secret the SecurityPolicy reads keys from |
blocked_secret |
no | <apikey_secret>-blocked |
Secret holding paused keys (resource or key paused) |
kubeconfig_path |
no | in-cluster | Path to a kubeconfig; omit to use in-cluster config |
kube_context |
no | — | kubeconfig context (local/dev); without it and kubeconfig_path: in-cluster |
Usage reporting backend settings (envoy-usage)
| Setting | Required | Default | Description |
|---|---|---|---|
api_url |
yes | — | Usage warehouse base URL |
api_token |
no | — | Bearer token for the warehouse API |
Composed offering (both backends)
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 | |
The input_tokens / output_tokens components are the usage meters the reporting backend fills
in; accounting_type: usage here is the agent-side metering type. Enforcement is configured on the
Waldur offering, not in this file: for a resource to auto-pause, the matching offering
component must be billing_type: LIMIT and the offering must set
plugin_options.action_on_usage_limit: pause — only LIMIT components are checked against their
limit. When reported usage reaches the limit, Waldur pauses the resource and the agent blocks its
keys. Blocking runs in the membership-sync loop, so membership_sync_backend: envoy must be set
(as above) — without it the agent never blocks the keys and the limit is not enforced.
Prerequisites
- Both Secrets must pre-exist. Create empty
apikey_secretand<apikey_secret>-blockedSecrets innamespaceas part of the deployment; the plugin patches entries into them but does not create the Secrets themselves. - RBAC. The agent's ServiceAccount needs the permissions below.
Kubernetes RBAC
1 2 3 4 5 6 7 8 9 | |
Usage Reporting
The envoy-usage backend is read-only. It implements usage reporting and resource pull, but does
not create, pause, restore, or otherwise manage resources (those belong to the envoy backend).
Usage is reported for the current month via _get_usage_report() and for arbitrary past months
via get_usage_report_for_period(), which the historical loader uses.
Per-key requests, per-resource billing
A resource owns several API keys, but usage is attributed per resource — otherwise rotating a key would split a tenant's bill in two.
The gateway forwards x-client-id per key (<resource_backend_id>-<n>), and the rollup
happens at the usage-shipper (a Vector pipeline): its remap transform strips the
-<n> suffix, so usage lands under the resource's backend_id:
1 2 | |
envoy-usage then queries the warehouse by resource_backend_id unchanged — no key
enumeration. Pausing, restoring and terminating the resource are the opposite: they act on the
gateway Secret rather than on usage, so they fan out to every client-id the resource owns.
Waldur's per-key commands act on the one client-id they name.
Per-key usage
Waldur enforces a key's limits against the usage reported for that key. The shipper supplies it by
keeping the unstripped client-id on each row as key_client_id, beside the stripped
client_id:
1 2 3 | |
1 2 | |
Billing does not change. The resource total is still the sum of every row for the resource's
client_id, whether or not the row names a key. Rows that name one of the resource's keys are also
summed per key, so the per-key figures add up to that total. The agent reports every key's
current-month usage with report_usage, including deleted keys, in the same report cycle
that submits the resource total. A key with no rows this month is
reported as zero, so last month's figure does not keep it over its limit. Rows naming a key Waldur
does not hold are logged and not reported, so the per-key figures then add up to less than the
total.
A resource that has non-zero usage on a row naming none of its keys is not reported per key for
that month. Splitting only part of the total across keys would under-count against their limits.
This happens with rows shipped before the shipper recorded keys. It is also skipped for an
offering that does not have Waldur's enable_api_key_provisioning plugin option set.
Installation
The plugin is discovered automatically once waldur-site-agent-envoy-ai-gateway is installed
alongside waldur-site-agent.
UV Workspace Installation
1 2 | |
Manual Installation
1 2 | |
Testing
Run the tests from inside the plugin directory; the plugin's entry points only resolve there.
1 2 3 4 5 6 7 | |
The suite covers the resource-level key lifecycle (provision/pause/restore/terminate), each per-key command and how it interacts with a resource pause, the usage warehouse client, and usage report mapping, per resource and per key. The Kubernetes and HTTP clients are mocked or replaced by an in-memory Secret store, so no live cluster or warehouse is required.
Troubleshooting
Keys are provisioned but requests are rejected
- Confirm the
SecurityPolicyreads the Secret named byapikey_secret. - Check the key landed in the active Secret, not the blocked one (
kubectl get secret ...). - Verify the request sends the API key the order surfaced.
One key is rejected while the resource's other keys work
- The key is paused on its own: the blocked Secret holds it as
<client_id>.paused. Resume it from Waldur; restoring the resource does not bring it back. - The key is deleted, or its entry was never written. Check the key's state in Waldur: a command that failed leaves the key Erred with the error attached.
A new key is blocked
- The resource is paused in Waldur, or the blocked Secret still holds its
<backend_id>.resource-pausedmarker. A resource that is not paused has the marker cleared on the next membership-sync cycle.
A per-key command stays pending in Waldur
- Confirm the agent runs in
order_processorevent_processmode;reportandmembership_syncdo not carry out key commands. - In
order_processmode, the command is applied on the next cycle (5 minutes by default). That mode skips an offering withstomp_enabled: true, so such an offering needsevent_process. - In
event_processmode, a command whose STOMP message was lost waits for the reconciliation sweep: 30 minutes pending, then the sweep's next run (hourly by default). - The agent log names each command it applies:
Applying API key <action> to key <uuid>.
Usage reports are empty or zero
- Check the warehouse is reachable: the backend's
ping()hitsGET /health. - Confirm the warehouse returns rows keyed by the same
client_idthe gateway authenticates with (the resourcebackend_id). - Confirm the offering's
backend_componentsare named to match the warehouse fields (input_tokens,output_tokens).
Per-key usage is not reported
- The Waldur offering must set the
enable_api_key_provisioningplugin option. - The warehouse rows must carry
key_client_id(see Per-key usage). A resource with non-zero usage on a row that names none of its keys is reported only as a total; the agent logsUsage of <backend_id> is not attributed to its keys.
RBAC errors on Secrets
- Grant
get/patchon Secrets innamespace(see Kubernetes RBAC).
Development
Project Structure
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 | |
Key Classes
EnvoyAIGatewayBackend(envoy): management backend — provisions and blocks/unblocks a resource's keys, and mints, pauses, resumes, updates and deletes a single keyEnvoyAIGatewayClient: Kubernetes client for api-key SecretsEnvoyUsageReportingBackend(envoy-usage): reports token usage to Waldur, per resource and per keyEnvoyUsageClient: HTTP client for the usage warehouse
Registered Entry Points
Each backend registers under the same name across three groups (waldur_site_agent.backends,
waldur_site_agent.component_schemas, waldur_site_agent.backend_settings_schemas):
envoy— management backend + its schemasenvoy-usage— reporting backend + its schemas
Extension Points
- Different usage warehouse: implement the two HTTP endpoints, or subclass
EnvoyUsageClientto match another warehouse's API. - Cost-based reporting: report a priced
token_costcomponent instead of raw token counts if pricing should live outside Waldur — the siblingcscs-dwdi-inferencebackend implements exactly this model and can be used as thereporting_backendin place ofenvoy-usage.