Hybrid Deployment
Run MCP tool execution on your own infrastructure while keeping management on Willow SaaS. The run service is deployed on-prem — the admin app, connect, and db-service stay on SaaS. You can optionally also self-host the Slack background agent in your cluster.
Why Hybrid?
- Data stays on-prem — tool calls (API keys, database queries, internal data) execute inside your network and never leave it
- Zero management overhead — updates, database, SSO, and admin UI are all managed by Willow
- Compliance — satisfies data residency and network isolation requirements without a full on-prem deployment
- Simple operations — one stateless pod to run, no database to manage
How It Works
┌──────────────────────────────────────────────────────────┐
│ Willow SaaS │
│ │
│ ┌──────────┐ ┌─────────────┐ ┌─────────────────┐ │
│ │ Admin │──▶│ db-service │◀──│ connect │ │
│ │ App │ └──────┬──────┘ └──┬──────────────┘ │
│ └──────────┘ │ │ │
│ default run gateway /api/on-prem- │
│ (SaaS → on-prem) db-service/* │
│ │ (on-prem → SaaS) │
└─────────────────────────┼──────────────┼─────────────────┘
│ ▲
▼ │
┌─────────────────────────┼──────────────┼─────────────────┐
│ Your Kubernetes cluster│ │ │
│ ┌────┴──────────────┴───┐ │
│ │ run │ │
│ MCP clients ────▶ │ (tool execution, │ │
│ (Claude, Cursor) │ MCP protocol) │ │
│ └──────────────────────┘ │
└──────────────────────────────────────────────────────────┘
On-prem run → SaaS: run reaches db-service through connect's authenticated proxy (/api/on-prem-db-service/*), using the gateway auth secret.
SaaS → on-prem run (recommended): db-service calls run for live tool listing, tool testing, guard evaluation, condition lookups, and MCP server setup, using the org's gateway URL. Tool execution and runtime guard enforcement never depend on it, so the platform keeps working without it — but most of the admin experience degrades. See What You Lose Without Inbound before deciding to block it, and Isolated Mode if you must.
MCP clients → on-prem run: Users connect directly to the on-prem run endpoint. That is the difference from SaaS: there, clients talk to Willow's hosted gateway; here they talk to your hostname. Cursor, Claude, and similar clients come from wherever the user is, so the hostname is typically public — you do not need a WAF, a list of client CIDRs, or an extra edge token. Access control is MCP OAuth (your IdP), the same as SaaS. OAuth is handled by the SaaS connect service by default; you can optionally serve the OAuth flow from your own run gateway so the code exchange and token encryption stay inside your network — see On-Prem OAuth (Run-Hosted Connect). If you do know the source IPs, an optional IP Access Filter can restrict the gateway further.
Before You Start
Read this whole section before doing anything. The steps that follow assume the infrastructure in Phase 1 already exists, and the order matters: the Willow-side gateway registration (Step 5) needs the hostname you choose in Step 1, and the Helm install (Step 7) needs the secret produced by Step 5.
Who you need
A hybrid deployment is not a one-person job. Line up these people before you start — the "Owner" line on each step below refers back to this table.
| Role | What they need access to | Which steps |
|---|---|---|
| Platform / Kubernetes engineer | kubectl and helm against the target cluster, permission to create namespaces, secrets, and ingresses | Steps 1, 4, 6, 7, 8 |
| DNS admin | Ability to create records on your domain | Step 2 |
| Network / firewall admin | Egress and ingress rules for the cluster | Step 3 |
| Willow org admin | An admin account in the Willow admin app with the org:edit scope | Steps 5, 8 |
| AWS / IAM admin (optional) | KMS and IAM in your AWS account | Only for Write-Only KMS |
| Slack workspace admin (optional) | Create and install a Slack app | Only for the Slack background agent |
Step 5 is done in the Willow admin app and requires the org:edit scope — viewing the page only needs org:read, but saving a gateway needs org:edit. If you can see Gateway Settings but saving fails with a permission error, ask an existing Willow admin to grant org:edit under Admin → Admin Users, or to run Step 5 for you.
What must already exist
None of these are created by Willow. Confirm all four before Step 1:
-
A Kubernetes cluster (v1.23+) you can install a Helm chart into. For production, size the node group for
runfirst — see Production Sizing and Hardening -
An ingress controller running in that cluster, with an external IP or hostname. Everything in Phase 1 depends on it, so verify it now:
kubectl get ingressclass # which controllers exist at allkubectl get svc -A | grep -Ei 'ingress|nginx' # find its namespace and address# Look for TYPE=LoadBalancer and a populated EXTERNAL-IPEvery later command in this guide is written against a self-installed ingress-nginx in the
ingress-nginxnamespace — adjust the namespace to whatever the command above reports. The AKS application routing add-on, for example, runs inapp-routing-system.If a fresh cluster has no controller yet, install one — the Willow chart does not:
helm repo add ingress-nginx https://kubernetes.github.io/ingress-nginx && helm repo updatehelm upgrade --install ingress-nginx ingress-nginx/ingress-nginx \--namespace ingress-nginx --create-namespace -
Control over a domain where you can create DNS records
-
An organization on Willow SaaS, and an admin account on it (see the permission note above)
You do not need a database, a Willow license file, or any Willow-side infrastructure. The run service is stateless.
The order of operations
The four phases must happen in this order. The dependency in the right-hand column is the reason why.
| Phase | What happens | Why it can't move earlier |
|---|---|---|
| 1. Prepare your infrastructure (Steps 1–4) | Hostname, DNS, firewall, TLS | Your ingress controller already has an IP, so all of this can be done before Willow is involved at all |
| 2. Register the gateway in Willow (Step 5) | Willow learns your run URL and issues an auth secret | Needs the final hostname from Step 1 |
| 3. Install run in your cluster (Steps 6–7) | values.yaml, helm upgrade --install | Needs the auth secret from Step 5, which is shown once |
| 4. Verify end to end (Step 8) | Pods, SaaS connectivity, OAuth discovery, a real MCP client | Needs everything above |
Optional add-ons (Isolated Mode, Write-Only KMS, the Slack background agent, run-hosted OAuth) come after a working baseline. Don't combine them with the first install — it makes failures much harder to isolate.
Checklist
Work top to bottom. Each item links to the step that completes it.
The boxes below are read-only on this page. Use Copy Page at the top to copy this guide as Markdown into your ticket tracker or notes — the checklist pastes in as a working task list you can mark done as you go. The Copy for Agent option in the same menu hands the whole guide to a coding agent that will walk you through it.
Phase 1 — Prepare your infrastructure
- Ingress controller confirmed to have an external IP (Prerequisites)
- Run hostname chosen and written down (Step 1)
- DNS record created and resolving (Step 2)
- Outbound HTTPS to Willow SaaS and your tool APIs allowed (Step 3)
- Inbound from your MCP clients allowed (Step 3)
- Decision made on inbound from Willow SaaS, and the egress IP allowlisted if you're allowing it (Step 3)
- Hostname confirmed reachable from outside your own network (Step 3)
- Certificate option chosen, and issuable given your current firewall state (Step 4)
- If using cert-manager: cert-manager installed and the
ClusterIssuerreportingREADY: True— neither is created by the Willow chart (Step 4) - Publicly-trusted TLS certificate terminating on the run hostname (Step 4)
- Streaming (SSE) confirmed unbuffered with a generous idle timeout (Step 4)
Phase 2 — Register the gateway in Willow
- External Run Service gateway added in Settings → Gateway Settings (Step 5)
- One-time gateway auth secret saved to your secret store (Step 5)
- External gateway marked Set as Default (Step 5)
Phase 3 — Install run in your cluster
-
values.yamlwritten, with onlyrunenabled (Step 6) - Auth secret supplied as a Kubernetes secret, not plaintext (Step 6)
- Chart installed (Step 7)
Phase 4 — Verify end to end
-
runpod1/1 Running(Step 8) - run → SaaS db-service reachable (Step 8)
- OAuth discovery returns the right resource and authorization server (Step 8)
- Gateway badge shows Healthy (or Isolated by choice) (Step 8)
- Tools tab loads in the admin app (Step 8)
- A real MCP client can list and call a tool (Step 8)
Phase 1 — Prepare your infrastructure
Nothing in this phase touches Willow. It is all your own cluster, DNS, and network.
Step 5 only needs the hostname you choose in Step 1 — it does not probe it at save time, and the gateway's health badge reads Unreachable until Step 7 regardless. Finish this phase first anyway: TLS is a hard requirement before the Step 8 checks can pass, and doing it now means you aren't debugging DNS, firewall, and certificates at the same time as a new Helm release. The one option that genuinely has to come later is a certificate issued from the annotation on the Willow ingress (Step 4), since that ingress doesn't exist until Step 7.
Step 1 — Choose the run hostname
Owner: Platform engineer · Needs: nothing · Produces: the hostname used in every step below
Pick the public hostname your MCP clients (and, unless you go isolated, Willow SaaS) will call. Throughout this guide it is written as willow.<YOUR_DOMAIN>, which matches the Helm chart's default: the chart builds the run hostname as <deployments.run.ingress.subdomain>.<global.domain.host>, and subdomain already defaults to willow.
So if global.domain.host is example.com and you leave the subdomain alone, run is served at https://willow.example.com. Write your hostname down — Steps 2, 4, 5, 6, and 8 all use it.
The hostname is baked into the run service's BASE_URL (used for OAuth discovery) and stored on the Willow side as your gateway URL. Changing it later means editing values.yaml, re-running Helm, and updating the gateway in the admin app. Choose one you can keep.
Verify: you have a single hostname written down, and it's inside a domain you can create records on.
Step 2 — Point DNS at your ingress
Owner: DNS admin · Needs: Step 1 · Produces: a resolving hostname
Create an A record (or CNAME, if your ingress exposes a hostname) from your run hostname to your ingress controller's external address:
kubectl get svc -n ingress-nginx # adjust the namespace for your controller
willow.<YOUR_DOMAIN> → <ingress-external-ip-or-hostname>
This is genuinely step two, not something you retrofit at the end: your ingress controller is a prerequisite that already has an address, so DNS can resolve before Willow's run service exists. It has to, because the TLS certificate in Step 4 is usually validated over DNS or HTTP against this record.
Verify:
dig +short willow.<YOUR_DOMAIN>
# Expect: your ingress controller's external IP
If it fails
| Symptom | Cause | Fix |
|---|---|---|
dig returns nothing | Record not created, or not propagated yet | Re-check the record; wait out the TTL. Willow can't help here — this is entirely your DNS provider |
EXTERNAL-IP stays <pending> | Your cloud didn't provision a load balancer for the ingress controller | Fix the ingress controller before continuing. Nothing downstream will work |
| Resolves to the wrong IP | Pointed at a node IP or an old load balancer | Point it at the ingress controller's LoadBalancer address, not a node |
| Resolves to your CDN's IPs even though this record isn't proxied | A wildcard record (*.<YOUR_DOMAIN>) is answering, and it is proxied | Wildcards are shadowed by an exact match, so the fix is usually cache, not the record: re-check with dig +short willow.<YOUR_DOMAIN> @1.1.1.1 to bypass your resolver. If an authoritative lookup still returns CDN addresses, the exact record is missing or misspelled |
dig disagrees with what your browser or curl reaches | Stale resolver or OS cache from before the record existed | Query an authoritative resolver directly (dig +short willow.<YOUR_DOMAIN> @1.1.1.1) rather than trusting the local cache |
If your DNS provider has a proxied wildcard for the domain, willow.<YOUR_DOMAIN> can end up behind a caching/buffering CDN without anyone creating a record for it. That breaks MCP streaming in a way that looks like a Willow problem — see the SSE warning in Step 4. Create an explicit, unproxied record for the run hostname rather than relying on the wildcard.
Step 3 — Open the network paths
Owner: Network / firewall admin · Needs: Step 1 · Produces: the traffic paths run depends on
The on-prem run service talks to Willow SaaS in both directions. Allow the following.
Outbound from your cluster (HTTPS / 443):
- To Willow SaaS (
*.withwillow.ai) — db-service proxy and OAuth discovery. Required. - To the third-party APIs your tools call (GitHub, Slack, Jira, etc.). Required, or those tools fail.
- If you enable AWS KMS integration: to the AWS KMS endpoint (
kms.<region>.amazonaws.com).
Inbound to willow.<YOUR_DOMAIN> (HTTPS / 443):
-
From your end users (MCP clients such as Claude and Cursor) — required. Those clients are not a CIDR you can list up front. OAuth is the access control; a WAF is optional and does not replace it.
-
From Willow SaaS — recommended. Allow the Willow SaaS egress IP for your org's region:
*.withwillow.ai(US):3.130.252.122*.eu.withwillow.ai(EU):3.120.156.158
This is a single, stable NAT egress IP per region — one static address, not a range, and it does not change with releases or scaling. Allowing it is the difference between a fully working admin console and a partly manual one; see What You Lose Without Inbound. If your policy genuinely forbids it, see Isolated Mode.
-
From your managed background agent platform, if you use one. Agents on the Claude, AWS Bedrock AgentCore, Cursor, or Willow Agents platforms run outside your network and connect to
runas ordinary MCP clients, so they need inbound of their own. Allowlisting Willow's NAT IP does not cover them, because the traffic originates from the platform's cloud rather than from Willow SaaS. If you cannot open inbound for a vendor cloud, run agents inside your own cluster with the agent harness.
Verify — do this now, not after installing. Your ingress controller is already running, so it should accept a TLS connection on your hostname even though Willow isn't deployed yet. A 404 or a certificate warning here is a pass: it proves packets reach the cluster.
Run this from a machine outside your own network — a home connection, or a phone tether with Wi-Fi off:
curl -sS -o /dev/null -w 'http=%{http_code}\n' --max-time 5 -k \
https://willow.<YOUR_DOMAIN>/
-k is deliberate here and only here: no certificate exists yet, and this check is about whether packets reach the cluster at all. Every later check (Step 4 and Step 8) must run without -k, because from then on the certificate is exactly what you're testing.
This is the most common way a hybrid deployment loses a day. A laptop on the corporate network, a VPN connection, or a shell inside the cluster all reach the endpoint by a path that neither your users nor Willow SaaS have. "It works for me" from inside is not evidence that the endpoint is public — and DNS resolving publicly doesn't mean the port is open, because the record and the firewall are independent.
If it fails
| Symptom | Cause | Fix |
|---|---|---|
Connection timed out | Packets are being silently dropped — a firewall, security group, or a load balancer with no rule for 443 | Inbound network has the per-cloud checks for AKS, EKS, and GKE |
Connection refused, immediately | The address is reachable but nothing is listening on 443 | Usually a missing load balancer rule; same section as above |
Could not resolve host | DNS record missing | Revisit Step 2 |
Certificate warning, or 404 | This is a pass. Traffic reaches the cluster | Continue to Step 4 |
A machine that is reachable with nothing listening replies with a TCP reset, and you see connection refused in milliseconds. A timeout means the packets vanished, which is what firewalls do by default. If you see a timeout, the problem is in your network, not in your cluster — and no amount of Kubernetes debugging will find it.
Step 4 — Terminate TLS on the run hostname
Owner: Platform engineer · Needs: Step 2 · Produces: a valid HTTPS endpoint that streams
Willow SaaS (tool listing / guard evaluation) and MCP OAuth clients both require the run endpoint to be served over valid, publicly-trusted HTTPS. A self-signed certificate will cause SaaS→run calls and OAuth discovery to fail. The Helm chart does not configure TLS for you — deployments.run.ingress.tls is empty by default, so you must supply it.
The chart defaults deployments.run.ingress.className to "nginx". That must match an IngressClass that exists in your cluster, or the ingress is created and then ignored by every controller — nothing is programmed, and the hostname returns 404 with no error in the Willow release to explain it.
kubectl get ingressclass
If you're on AKS using the application routing add-on, the class is webapprouting.kubernetes.azure.com, certificates come from Azure Key Vault instead of cert-manager, and the TLS secretName must be exactly keyvault-run. See Match your ingress controller for the values, and skip the cert-manager options below.
If the cluster already has a self-installed ingress-nginx, do not enable that add-on just to attach a Key Vault certificate — it creates a second NGINX controller and a second public IP. Export the cert and install it as a Kubernetes TLS secret (the first option in the table below) instead.
Terminate TLS at your ingress with a real certificate. You have four ways to get one, and the right choice depends on whether inbound from the internet is open yet:
| Option | Use when | Works before inbound is open? |
|---|---|---|
| A certificate your organization already owns | You have a wildcard for the domain, or a procurement process | Yes |
| cert-manager with a DNS-01 challenge | You want auto-renewal and can use your DNS provider's API | Yes |
| cert-manager with an HTTP-01 challenge | You want auto-renewal and the hostname is already public | No |
| A certificate on your cloud load balancer | You already terminate TLS at an ALB, Application Gateway, or GCLB | Yes |
Let's Encrypt has to reach your hostname on port 80 to validate an HTTP-01 challenge. If Step 3 isn't finished, the challenge never validates and the Certificate sits at Ready: False with no useful error — which reads as "cert-manager is broken" when it isn't. Either finish Step 3 first, or use a certificate you already have.
To install one you already have, the full chain and key go into a secret whose name matches the secretName below, in the same namespace as the Willow release:
kubectl create namespace <namespace> --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret tls willow-run-tls \
--namespace <namespace> \
--cert=fullchain.pem --key=privkey.pem
Note that a wildcard for *.example.com covers willow.example.com but not willow.external.example.com — wildcards match a single label.
For choosing between the four options, verifying what's actually being served, and debugging a stuck cert-manager Certificate, see TLS and Inbound Reachability.
If you're using cert-manager, install it and create the issuer first
The cert-manager.io/cluster-issuer annotation below only names a ClusterIssuer that must already exist in the cluster. If it doesn't, no Certificate is issued, the ingress keeps serving your controller's default certificate, and the only clue is an Issuer not found event on the Certificate object — nothing in the Willow release reports it.
helm repo add jetstack https://charts.jetstack.io && helm repo update
helm upgrade --install cert-manager jetstack/cert-manager \
--namespace cert-manager --create-namespace \
--set crds.enabled=true
Then create the issuer the annotation refers to. This is the HTTP-01 form — it only works once Step 3 has opened inbound on port 80. For DNS-01, substitute your DNS provider's solver.
# clusterissuer.yaml
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
name: letsencrypt-prod # must match the annotation in values.yaml
spec:
acme:
email: <YOUR_EMAIL> # Let's Encrypt expiry notices go here
server: https://acme-v02.api.letsencrypt.org/directory
privateKeySecretRef:
name: letsencrypt-prod
solvers:
- http01:
ingress:
class: nginx # must match your IngressClass from above
kubectl apply -f clusterissuer.yaml
kubectl get clusterissuer letsencrypt-prod # READY must be True before Step 7
With cert-manager and the issuer in place, add this to the values.yaml you write in Step 6:
deployments:
run:
ingress:
annotations:
cert-manager.io/cluster-issuer: "letsencrypt-prod"
tls:
- hosts:
- "willow.<YOUR_DOMAIN>"
secretName: "willow-run-tls"
Helm merges annotation keys with the chart's defaults, so adding cert-manager.io/cluster-issuer keeps the chart's streaming annotations (below) in place.
The MCP protocol uses a streaming server→client channel: the client opens a GET /mcp Server-Sent Events stream (content-type: text/event-stream) that stays open and must be flushed to the client immediately. If any proxy in front of run buffers the response, or times out the idle stream, MCP clients hang on a cold connection and fail, even though tool listing itself completes in ~1s on the server. Requests made with plain curl tools/list look fast because they don't hold the SSE stream open — so this is easy to misdiagnose as a server problem.
The chart already sets sane defaults on the run Ingress — proxy-read-timeout: 300, proxy-send-timeout: 300, and proxy-buffering: "off". That covers a plain ingress-nginx setup out of the box. The failures come from the layers the chart cannot see, or from overriding those keys:
-
NGINX Ingress: raise the timeouts further if your clients hold idle streams longer than 5 minutes, and never turn buffering back on for this route:
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"nginx.ingress.kubernetes.io/proxy-buffering: "off" -
Istio / Envoy: Envoy streams by default — do not enable a response
bufferfilter on the run route, and set a generous routetimeout(e.g.0sto disable, or a large value) so long-lived streams aren't cut. -
Cloud load balancers (AWS ALB/NLB, Azure Load Balancer, Azure Application Gateway, GCLB, etc.): raise the idle timeout well above your client's timeout — the AWS ALB default is 60s, Azure Load Balancer defaults to 4 minutes, and Azure Application Gateway enforces its own backend request timeout — and don't front the run hostname with a buffering layer. This is the most common cause, because the chart's ingress annotations have no effect on an LB in front of it.
On AKS the idle timeout is set with an annotation on the ingress controller's
Service, not on the Willow ingress (maximum 30 minutes):kubectl annotate svc ingress-nginx-controller -n ingress-nginx --overwrite \service.beta.kubernetes.io/azure-load-balancer-tcp-idle-timeout="30"If you installed the controller with Helm, set the same key under its
controller.service.annotationsso a chart upgrade doesn't drop it. -
Cloudflare / other CDNs: the run hostname must not be proxied through a caching/buffering layer. On Cloudflare that means the record is DNS-only (grey cloud), not proxied (orange cloud) — or, if you must proxy it, exempt
text/event-streamfrom buffering.
Verify (after Step 7, once run is actually serving — come back for this one):
curl -sSI https://willow.<YOUR_DOMAIN>/healthz # valid cert, HTTP 200
curl -sS -i -N -H "Accept: text/event-stream" https://willow.<YOUR_DOMAIN>/mcp
The /mcp request returns 401 with a JSON body within a second or two, and a www-authenticate: Bearer ... header pointing at /.well-known/oauth-protected-resource/mcp. That is the pass: /mcp requires authentication, so an anonymous request is rejected before any stream opens.
An unauthenticated /mcp request never produces an SSE stream, so there is nothing here to hold open past 60s. A fast 401 rules out a proxy that stalls headers, but it cannot detect a buffering layer or a low idle timeout — only an authenticated, long-lived stream can. Configure the timeouts above from your infrastructure's settings, and treat check 8 of Step 8 (a real MCP client making a cold connection) as the actual streaming test.
If /mcp hangs instead of returning 401 promptly, something in front of run is buffering the response — that is a failure.
If it fails
| Symptom | Cause | Fix |
|---|---|---|
curl: (60) SSL certificate problem | Self-signed or incomplete chain | Issue a publicly-trusted certificate. A private CA is not enough — MCP clients and Willow SaaS both reject it |
| Ingress serves the default/fake certificate | ingress.tls not set, so the controller has no certificate for this host | Add the tls block above and re-run Helm |
cert-manager Certificate stuck Ready: False | ACME challenge can't validate | Check DNS from Step 2 resolves publicly, and that HTTP-01 traffic reaches the ingress |
/mcp returns 401 immediately | This is a pass. /mcp requires authentication | Continue — the real streaming test is a live MCP client |
/mcp hangs with no response at all | A proxy is buffering the response | Disable response buffering on the run route |
| An MCP client connects, then drops at ~60s | Read/idle timeout too low — usually a cloud LB, not the ingress | Raise the LB idle timeout; the chart's ingress annotations don't apply to it |
cert-manager Certificate reports Issuer not found | The ClusterIssuer named in the annotation doesn't exist | Create it — see above |
Phase 2 — Register the gateway in Willow SaaS
Step 5 — Register the gateway in Willow
Owner: Willow org admin (needs org:edit) · Needs: Step 1 · Produces: the AUTH_SECRET that Step 6 requires
This tells Willow SaaS where your run service lives and establishes the shared secret both sides authenticate with.
- Log in to the Willow admin app (
app.withwillow.ai, orapp.eu.withwillow.aiin the EU). - In the sidebar, go to Admin → Settings.
- Expand the Gateway Settings section. You can also jump straight to
/admin/settings#gateway. - Select Add Gateway.
- Choose External Run Service — described as "Route requests to your own self-hosted run service".
- In External Run URL, enter the hostname from Step 1, including the scheme and no path:
https://willow.<YOUR_DOMAIN>. - Select Add Gateway to save.
- A Gateway Auth Secret dialog appears with a one-time secret. Copy it now and put it straight into your secret store — it is not shown again. Its "Next steps" list also names
KMS_KEY_IDand AWS credentials: those belong to the optional Write-Only KMS section, and a first install is complete without them. It also says to setAUTH_SECRETas an environment variable — supply it as a Kubernetes secret instead, per Step 6. - Back on the gateway list, find your new external gateway card and select Set as Default. It should then show a Default Gateway badge.
Verify: the external gateway card is listed under Gateway Settings, shows Run: https://willow.<YOUR_DOMAIN>, and carries a Default Gateway badge. Its health badge will say Unreachable until Step 7 — that's expected, nothing is deployed yet.
Run URL has no /mcp — Willow-hosted ones doAn external gateway's Run line shows exactly the URL you entered, while Willow-hosted gateway cards show theirs with /mcp appended (https://<org>.mcp-s.com/mcp). Both are correct: Willow derives the hosted MCP path itself, and for an external gateway it appends the path when handing URLs to clients. Don't "fix" the inconsistency by adding /mcp to your entry — the field rejects a path.
If it fails
| Symptom | Cause | Fix |
|---|---|---|
| No Gateway Settings section in Settings | Your org isn't a SaaS org, or you're signed in to a full on-prem admin app | Gateway Settings only exists on Willow SaaS. On a self-hosted admin app the run URL comes from Helm, not the UI |
| No Admin → Settings in the sidebar at all | Missing the org:read scope | Ask an existing Willow admin to grant it under Admin → Admin Users |
| Saving fails with a permission error | Missing the org:edit scope | Same — org:read shows the page, org:edit saves it |
| No External Run Service option in the Add Gateway dialog | You're not on a SaaS org | As above |
| "Please enter a valid URL" | Missing scheme, or a path was included | Use exactly https://willow.<YOUR_DOMAIN> — no trailing slash, no /mcp |
| You closed the secret dialog before copying | The secret is shown only once | Use Rotate Secret on the gateway card to issue a new one, then use that value in Step 6 |
| No On-Prem item in the sidebar afterwards | The external gateway isn't the default | Select Set as Default on its card (step 9) |
Phase 3 — Install run in your cluster
Step 6 — Create values.yaml
Owner: Platform engineer · Needs: Step 5 · Produces: the config Helm installs
The chart creates the run Ingress from deployments.run.ingress below. You do not write a separate Ingress manifest — a second one will fight the object Helm manages.
Replace the placeholders:
<ORG_SLUG>— your organization slug (visible in your admin app URL, or under Settings → General)<YOUR_DOMAIN>— the domain from Step 1
withwillow.aiDB_SERVICE_URL and CONNECT_URL below use <ORG_SLUG>.withwillow.ai, which is wrong for a large share of orgs: EU orgs use <ORG_SLUG>.eu.withwillow.ai, and orgs created before the rename are served on <ORG_SLUG>.mcp-s.com. Your gateway card in Settings → Gateway Settings shows the connect hostname Willow actually uses for you — the Connect line. Whatever it is, the same apex goes in both variables.
A wrong apex doesn't fail at install time. It surfaces later as a timeout on Step 8 check 3, or as MCP clients being unable to authenticate (check 4).
global:
domain:
host: "<YOUR_DOMAIN>"
org: "<ORG_SLUG>"
OPENAI_API_KEY: ""
deployments:
app:
enabled: false
connect:
enabled: false
db-service:
enabled: false
run:
enabled: true
ingress:
# Public host = <subdomain>.<domain>. "willow" is the chart default, so
# this line only needs changing if you picked a different hostname in
# Step 1.
subdomain: "willow"
# TLS from Step 4 — the chart ships no certificate config.
# This is the self-installed ingress-nginx + cert-manager form. On the AKS
# application routing add-on, use the Key Vault form below instead.
annotations:
cert-manager.io/cluster-issuer: "letsencrypt-prod"
tls:
- hosts:
- "willow.<YOUR_DOMAIN>"
secretName: "willow-run-tls"
# The gateway auth secret from Step 5. See the warning below.
secretName: "willow-secrets"
env:
PORT: "3000"
LOG_LEVEL: "info"
DB_SERVICE_URL: "https://<ORG_SLUG>.withwillow.ai/api/on-prem-db-service"
CONNECT_URL: "https://<ORG_SLUG>.withwillow.ai"
The add-on ignores className: nginx and takes its certificate from Azure Key Vault, so drop the cert-manager annotation entirely and use the required keyvault-<ingress-name> secret name (Match your ingress controller):
deployments:
run:
ingress:
className: "webapprouting.kubernetes.azure.com"
annotations:
kubernetes.azure.com/tls-cert-keyvault-uri: "<KEY_VAULT_CERT_URI>"
tls:
- hosts:
- "willow.<YOUR_DOMAIN>"
secretName: "keyvault-run" # must be exactly this for the run service
Leaving the cert-manager annotation in place here fails silently: no Certificate is issued and the add-on serves its own default certificate.
Create the release namespace, then the secret referenced above with the value from Step 5. The namespace command is safe to rerun if you already created it for a TLS secret in Step 4. Keep the single quotes — a generated secret can begin with -, which an unquoted --from-literal value hands to kubectl as a flag:
kubectl create namespace <namespace> --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret generic willow-secrets \
--namespace <namespace> \
--from-literal=AUTH_SECRET='<gateway-secret-from-step-5>'
You can set deployments.run.env.AUTH_SECRET directly, and the quickest tests often do. But values.yaml usually ends up in Git, and the chart renders env into a ConfigMap — so the secret would sit in plaintext in your cluster and your repo. Use secretName as shown above; secrets are loaded after the ConfigMap, so a value from the secret wins.
If AUTH_SECRET isn't supplied by either your secret or run.env, the chart falls back to the shipped default for global.dbAuthSecret. The pod then starts perfectly healthy and every call to SaaS db-service returns 401. There is no crash and no startup error to alert you.
That's why Step 8 checks db-service connectivity explicitly rather than trusting pod status.
Configuration reference
| Value | Purpose |
|---|---|
global.domain.host | Your domain. Combined with the run ingress.subdomain to build the run ingress hostname (<subdomain>.<domain>) and the service's BASE_URL. |
global.org | Your org slug. Must match the SaaS org exactly. Used for org identification across all run operations. |
global.OPENAI_API_KEY | Optional. Required only if you use AI-powered guardrails. |
deployments.run.ingress.subdomain | First label of the public run hostname. Chart default: willow. |
deployments.run.ingress.tls | TLS certificate for the run host. Empty by default — you must set it (see Step 4). |
deployments.run.secretName | Kubernetes secret holding AUTH_SECRET. Overrides the ConfigMap. |
DB_SERVICE_URL | How run reaches db-service — proxied through connect at /api/on-prem-db-service/*. |
CONNECT_URL | The SaaS connect URL for your org. Used for MCP OAuth discovery so clients can authenticate. |
The following are generated automatically from your configuration — do not set them manually:
BASE_URL— derived from the run ingress host,<deployments.run.ingress.subdomain>.<global.domain.host>ON_PREM—true, from the chart'sglobal.onPremdefaultORG— derived fromglobal.org
Verify without installing anything:
helm template willow willow/webrix-helm -f values.yaml \
| grep -E "BASE_URL|CONNECT_URL|DB_SERVICE_URL|host:"
Confirm BASE_URL is https://willow.<YOUR_DOMAIN> and that only run resources are rendered.
Step 7 — Install the chart
Owner: Platform engineer · Needs: Step 6 · Produces: a running run pod
helm repo add willow https://webrix-ai.github.io/webrix-helm
helm repo update
helm upgrade --install willow willow/webrix-helm \
--namespace <namespace> \
--create-namespace \
-f values.yaml \
--wait
The chart is also published as an OCI artifact to oci://ghcr.io/webrix-ai/charts. If your organization standardizes on OCI registries, skip helm repo add and reference the chart directly by its oci:// URL — it's the exact same chart:
helm upgrade --install willow oci://ghcr.io/webrix-ai/charts/webrix-helm \
--namespace <namespace> \
--create-namespace \
-f values.yaml \
--wait
Pending, check node capacity before anything elserun requests 300m CPU / 500Mi (chart 1.0.58 and later; earlier charts requested a full core, which frequently would not fit). A CPU request is a reservation that must be satisfied on a single node, so an ingress controller, cert-manager, and your cloud's system pods all compete for the same allocatable pool. helm --wait simply times out — nothing crashes.
kubectl describe node <node> | sed -n '/Allocated resources/,/Events/p'
If you raise the request for production throughput, keep it in step with autoscaling.targetCPUUtilizationPercentage: the HPA measures usage as a percentage of the request, so an oversized request means the target is never reached and run never scales out. Production Sizing and Hardening has load-tested requests, HPA settings, and the node group to put run on.
Verify:
helm status willow -n <namespace> # STATUS: deployed
kubectl get pods -n <namespace> # run-xxx 1/1 Running
If it fails
| Symptom | Cause | Fix |
|---|---|---|
ImagePullBackOff | Cluster has no access to quay.io/webrix — often because egress allowlisted quay.io but not its CDN (cdn.quay.io, cdn01.quay.io, …) — or the pull secret exists but isn't referenced | Allow the CDN hosts too, then create the secret and list it in global.imagePullSecrets — see Custom Image Pull Secrets |
Pod Pending, Insufficient cpu | No single node has enough allocatable CPU left for run's request | See the note above; kubectl describe pod <pod> -n <namespace> names the resource that didn't fit |
another operation (install/upgrade/rollback) is in progress | A previous helm run was interrupted and the release is stuck pending-install | helm status willow -n <namespace> to confirm, then helm uninstall willow -n <namespace> and reinstall (nothing is lost — run is stateless) |
Extra pods for app, connect, db-service | Those deployments weren't disabled | Set all three to enabled: false — hybrid runs run only |
--wait times out | Probes never pass | kubectl logs -n <namespace> -l app=run and continue with Step 8 |
Phase 4 — Verify end to end
Step 8 — Verify end to end
Owner: Platform engineer, plus the Willow org admin for the last two checks · Needs: Step 7
Run these in order. Each one isolates a different link in the chain, so the first failure tells you where to look.
1. Pod health
kubectl get pods -n <namespace>
# Expected: run-xxx 1/1 Running
2. Logs
kubectl logs -n <namespace> -l app=run --tail=50
Look for the server listening on port 3000 and 200 responses on /healthz and /readyz. The first lines are usually Fastify deprecation warnings (FSTDEP023, FSTDEP024) logged at level 50 — they are harmless and expected, not a sign of an unhealthy pod.
/readyz returns 200 unconditionally — it does not check SaaS connectivity, the auth secret, or your gateway registration. 1/1 Running means the process started, nothing more. Checks 3 through 6 are the ones that matter.
3. run → SaaS db-service
kubectl exec -n <namespace> deploy/run -- \
wget -qO- --header="Authorization: <AUTH_SECRET>" \
"<CONNECT_URL>/api/on-prem-db-service/healthz"
Replace <CONNECT_URL> with the value you set in Step 6. A JSON response confirms run can reach db-service through the SaaS proxy. Use wget — the run image intentionally ships without curl.
4. MCP OAuth discovery
curl https://willow.<YOUR_DOMAIN>/.well-known/oauth-protected-resource
Verify:
resourcepoints to your on-prem run URL —https://willow.<YOUR_DOMAIN>/authorization_serverspoints to SaaS connect — yourCONNECT_URLfrom Step 6
run derives them from BASE_URL and CONNECT_URL through URL normalization, so a bare origin comes back as https://willow.<YOUR_DOMAIN>/. That is correct output, not a misconfiguration — compare the hostname, not the exact string. Set BASE_URL/CONNECT_URL without a trailing slash as shown in Step 6 and let run normalize them.
run answers both /.well-known/oauth-protected-resource (checked above) and the per-resource form /.well-known/oauth-protected-resource/mcp. The www-authenticate header on a 401 from /mcp points at the /mcp form, which is what MCP clients actually fetch, and its resource is the full https://willow.<YOUR_DOMAIN>/mcp. Curl both if a client fails to authenticate while the check above looks fine:
curl https://willow.<YOUR_DOMAIN>/.well-known/oauth-protected-resource/mcp
5. Streaming — run the SSE check from Step 4 now that there's a server behind the hostname.
6. Gateway health in the admin app — in Settings → Gateway Settings, your external gateway card should show Healthy. If you chose isolated mode it will show Isolated instead, which is also a pass.
If it says Unreachable or Not the run service, reproduce Willow's exact check yourself before changing anything. Willow performs one GET https://<your-gateway-url>/healthz with a 5-second timeout, and it does not follow redirects:
curl -sS -w '\nhttp=%{http_code} redirect=%{redirect_url}\n' \
--max-time 5 --max-redirs 0 \
https://willow.<YOUR_DOMAIN>/healthz
Run it from outside your network. A pass is http=200 with a body of {"status":"ok"} — a 200 with any other body means something other than the run pod answered. The result tells you which layer is at fault, and the failure table below maps each one to a fix. Don't add -k or -L — they suppress exactly the certificate and redirect problems you're looking for.
Willow probes whatever URL is on the gateway card. If the chart serves willow.example.com but the card says willow2.example.com, the check fails at DNS or with a 404 from the ingress default backend, and the badge reads Unreachable for a reason that has nothing to do with your network. Compare the card against kubectl get ingress -n <namespace> before investigating anything else.
7. Tools tab — in the admin app, go to MCP Servers (/integrations), open any server's Edit page (/integrations/<slug>/edit), and select Tools from the section dropdown. Tools should load without errors.
8. A real MCP client — connect Claude or Cursor to your run endpoint, list tools, and call one. This is the only check that exercises the full path your users take.
If it fails
| Failing check | Symptom | Cause | Fix |
|---|---|---|---|
| 3 | 401 | Auth secret mismatch between the cluster and Willow | Confirm the AUTH_SECRET in your secret equals the Step 5 value. If in doubt, Rotate Secret in Gateway Settings and update the Kubernetes secret |
| 3 | 404 | DB_SERVICE_URL missing the /api/ prefix | Must be your CONNECT_URL followed by /api/on-prem-db-service |
| 3 | Timeout / DNS failure | Outbound to *.withwillow.ai blocked | Revisit Step 3 |
| 4 | Wrong authorization_servers | Wrong CONNECT_URL | Must be the Connect host on your gateway card in Settings → Gateway Settings |
| 4 | Wrong resource | BASE_URL doesn't match your real hostname | Check ingress.subdomain and global.domain.host in Step 6 |
| 4 | Connection refused / cert error | DNS, ingress, or TLS | Revisit Step 2 and Step 4 |
| 6 | Unreachable | No inbound from Willow SaaS to run | Allowlist the region's egress IP from Step 3, or turn on Isolated gateway if that's deliberate |
| 6 | Unreachable, and check 4 also times out | Inbound is blocked for everyone, not just Willow | Inbound network — per-cloud checks for AKS, EKS, and GKE |
| 6 | Unreachable, but curl from outside returns 200 | The gateway URL registered in Willow doesn't match the hostname the ingress serves | Compare the URL on the gateway card against kubectl get ingress -n <namespace>; they must be identical |
| 6 | Unreachable, and curl reports a certificate error | Certificate missing, self-signed, or from a private CA | Choosing a certificate |
| 6 | Unreachable, and curl returns 404 | Ingress has no rule for this hostname | Hostname mismatches |
| 6 | Unreachable, and curl returns 503 | Ingress is up but has no Ready pod behind it | No healthy backend — usually ImagePullBackOff |
| 6 | Unreachable, and curl returns 301 or 302 | A redirect is in the path. Willow does not follow redirects | Remove the HTTP-to-HTTPS redirect or auth layer on /healthz |
| 6 | Auth issue | Gateway is reachable but the secret doesn't match | Same fix as check 3's 401 |
| 6 | Not the run service, and curl returns 200 with an empty or HTML body | DNS and TLS are fine, but the ingress is answering instead of the run pod | Add or fix the ingress rule so this host routes to the run Service. /healthz must return {"status":"ok"} |
| 6 | Not the run service, but curl /healthz returns {"status":"ok"} | The host routes to app or connect. They serve an identical /healthz, but not /healthz/validate | Point this host's ingress rule at the run Service. Confirm with curl https://willow.<YOUR_DOMAIN>/healthz/validate, which must not return 404 |
| 6 | No card at all | Gateway never saved | Redo Step 5 |
| 7 | Tools tab empty or errors | Usually the same missing inbound as check 6 | See What You Lose Without Inbound |
| 8 | Client gets SaaS URLs, not yours | External gateway isn't the default | Set as Default on the gateway card (Step 5, step 9) |
| 8 | Cold start hangs ~60s then times out | SSE buffered or timed out upstream | See Step 4 |
Once all eight pass, the baseline deployment is done. Only now move on to the optional sections below.
Reference — inbound from Willow SaaS
Read this when you're making the decision in Step 3, or when a gateway shows Unreachable in Step 8.
What You Lose Without Inbound
Blocking SaaS→run inbound is safe for your data path and costly for everything else. Nothing below affects a tool call made by a real user through an MCP client — but all of it is part of how admins build, test, and maintain the platform day to day.
| What | Impact without inbound |
|---|---|
| Live tool listing (integration Tools tab) | Falls back to the tools last synced on the integration. New or renamed tools on an upstream MCP server won't appear until someone re-syncs from an in-network browser. |
| Tool refresh / re-sync | No longer happens server-side. An admin has to open the MCP server in the admin app from a browser inside your network so the browser can collect directly and push results back. |
| Testing a tool (Test Run in the admin console) | Unavailable. Admins cannot verify a connector, its credentials, or a parameter mapping from the console, and have to test through a real MCP client instead. |
| Guard playground (build-time guard testing) | Unavailable. Guard rules must be written without a dry run and validated only once they are live. Runtime enforcement is unaffected. |
| Condition option pickers | Access-rule and policy conditions that populate their values live from the connector cannot load options. Admins must type values by hand and get no validation that they exist. |
| Adding a new MCP server | The setup wizard's reachability probe and OAuth discovery both run through run. The server still saves, but it finishes with a "not reachable from gateway" warning, collects no tools, and cannot auto-detect its auth requirements. |
| Gateway health | The admin app cannot report whether your gateway is up. You lose Willow-side monitoring of your own runtime and find out about an outage from your users. |
| Managed background agents | Agents hosted outside your network cannot reach the gateway at all. See the note below — this one is not solved by allowlisting Willow's NAT IP. |
Willow registers your gateway URL with the agent platform at sync time, and the agent runtime then calls it directly as an MCP client. For agents on Claude, AWS Bedrock AgentCore, Cursor, or Willow Agents, that runtime lives in the platform's cloud, not in Willow SaaS — so Willow's NAT egress IP does not cover it, and an ingress locked to that IP will still block every agent tool call.
If you want background agents without opening inbound to a vendor cloud, deploy the agent harness in your own cluster. Agents then run beside run, and the harness can even be configured to poll Willow rather than receive inbound.
Isolated Mode (no inbound from Willow SaaS)
Isolated mode exists for organizations whose policy forbids any inbound from a SaaS provider. It is a permanently degraded configuration: read What You Lose Without Inbound first. If your objection is to opening a broad range, note that the allowlist is a single static IP — most customers who start in isolated mode end up allowlisting it later.
If you truly cannot allow it, open inbound to willow.<YOUR_DOMAIN> only from your MCP clients and skip the Willow SaaS egress allowlist entirely. This is a network posture, not a separate build — the Helm values are identical to Step 6.
What still works:
- MCP clients → run: tool discovery and tool execution
- Runtime guard enforcement on every tool call — guards are evaluated locally inside
run, so blocking/redaction/warnings all apply - MCP OAuth authentication — driven by SaaS connect; the end user's browser reaches connect, so no SaaS→run inbound is needed
- run → SaaS (db-service proxy) for config, token exchange, and logging — this is outbound from your cluster
- Audit logging of every tool call
Mark the gateway as isolated. In Settings → Gateway Settings, turn on the Isolated gateway switch on the external gateway's card. This is optional but recommended once you have committed to the posture: Willow then never opens a connection to your gateway, so admin screens stop waiting on calls that will time out and instead offer the in-network browser flow immediately. The Unreachable health badge is replaced by an Isolated badge, since an unreachable warning would otherwise read as a fault when it is the configured behavior.
Turning it off again is a toggle, not a redeploy — allowlist the region's egress IP (see Step 3) and switch Isolated gateway off to restore every feature in the table above.
Background agents in isolated mode. The agent harness normally expects Willow to call it, which isolated mode rules out the same way it rules out SaaS→run. Install it with willow.mode: pull instead: the control plane collects queued work over the same outbound path and gateway secret run already uses, needs no ingress, and requires no new network permission. See Pull delivery.
Slack Background Agent (optional)
Only start this once Step 8 passes.
The Slack background agent lets background agents be triggered from Slack channels and stream their replies back into the thread. On hybrid you have two choices:
- Let Willow host it (default, simplest) — nothing to deploy. Configure each agent's Slack app in the admin app and skip the rest of this section; follow Set up Slack for background agents.
- Self-host it in your cluster — run the
slack-background-agentservice next torunso the Slack connection and message routing stay in your network. Continue below.
Who you need: the platform engineer from Phase 3, plus a Slack workspace admin to create the Slack app and issue the app-level token.
How it fits the hybrid model
The agent connects to Slack over Socket Mode (an outbound WebSocket — no ingress or public URL) and reaches the SaaS db-service through the same connect proxy run uses. run forwards the reply_to_slack_thread tool to it in-cluster at http://slack-background-agent, which the chart wires automatically when the deployment is enabled.
┌──────────────────────────────────────────────────────────┐
│ Willow SaaS (admin app · connect · db-service) │
└───────────────▲──────────────────────────▲───────────────┘
/api/on-prem-│db-service (auth secret) │
│ │
┌───────────────┼──────────────────────────┼───────────────┐
│ Your cluster │ │ │
│ ┌──────┴───────┐ in-cluster ┌───┴──────────────┐ │
│ │ run │─────────────▶│ slack-background- │ │
│ │ (gateway) │◀─ replies ───│ agent │ │
│ └──────────────┘ /agent-reply└────────┬─────────┘ │
└───────────────────────────────────────────────┼───────────┘
▼ Socket Mode (outbound wss)
Slack
Step S1 — Allow outbound to Slack
Owner: Network admin
In addition to Step 3, allow outbound from the cluster:
- To Slack (
https://slack.comand the Socket Mode WebSocketwss://wss-primary.slack.com/*.slack.com, HTTPS/443) — required. - To Willow SaaS (
*.withwillow.ai) — already allowed forrun; the agent reuses the same db-service proxy.
No inbound is needed for the Slack agent — Socket Mode is outbound-only.
Step S2 — Create the deploy secret
Owner: Platform engineer
At deploy time the agent needs only the same gateway auth secret run uses — it authenticates to the SaaS db-service with it. The per-agent Slack app-level tokens (xapp-…, scope connections:write) are not deploy-time secrets: by default the agent loads them per org from db-service (Admin → Slack Agent), polled roughly every 60s. So create a secret with just AUTH_SECRET:
kubectl create secret generic slack-background-agent-secret \
--namespace <namespace> \
--from-literal=AUTH_SECRET='<gateway-secret-from-step-5>'
AUTH_SECRET hereThe agent authenticates to the SaaS db-service with the same gateway secret as run. If you leave it out, the chart falls back to its placeholder default and every db-service call from the agent returns 401 — with a healthy-looking pod, exactly as described in Step 6.
If you run a single shared Slack app instead of per-agent installs, add SLACK_APP_TOKEN (xapp-…, scope connections:write) to the same secret. Leave it unset to use the default per-org lookup from db-service.
--from-literal=SLACK_APP_TOKEN='xapp-...'
Step S3 — Add it to your values.yaml
Owner: Platform engineer
Extend the Step 6 values.yaml with the slack-background-agent deployment. It uses the same gateway secret as run to authenticate to db-service, and the same DB_SERVICE_URL:
deployments:
# ... app/connect/db-service disabled, run enabled (from Step 6) ...
slack-background-agent:
enabled: true
# Required on this instance if it should serve agents pinned to a specific
# named gateway (Step S4) — without it, those agents get no Socket Mode
# connection from anyone, not even your org's default instance. Remove this
# line entirely on your org's single default instance, which already claims
# every un-pinned agent automatically.
gatewayId: "<gateway-id-from-step-s4>" # copy from the admin app
env:
DB_SERVICE_URL: "https://<ORG_SLUG>.withwillow.ai/api/on-prem-db-service"
LOG_LEVEL: "info"
# Supplies AUTH_SECRET (Step S2).
secretName: "slack-background-agent-secret"
| Value | Purpose |
|---|---|
AUTH_SECRET | Same gateway secret as run — authenticates the agent to the SaaS db-service proxy. Supply via a secret. |
DB_SERVICE_URL | Same connect-proxied db-service URL as run (/api/on-prem-db-service). |
gatewayId | Injected as GATEWAY_ID. Required on any instance that should serve agents pinned to a specific named gateway from Step S4 — without it, those agents get no Socket Mode connection from anyone. Leave unset only on your org's single default instance, which already claims every un-pinned agent. |
When deployments.slack-background-agent.enabled is true, the chart handles the in-cluster wiring for you — do not set these manually:
SLACK_BG_AGENT_URLonrun→http://slack-background-agent, soruncan forward replies.ORG→ derived fromglobal.org, so only your org's agents connect.
Re-run the same helm upgrade --install from Step 7 to apply.
Step S4 — Register and configure in the admin app
Owner: Willow org admin
-
Set your org's default Slack agent service URL — required. Go to Build → Background Agents, open the settings (gear) icon to reach Background Agent Settings, find Slack under Channels configurations, and select Set up (or Manage). Enter your self-hosted instance's reachable URL in Slack agent service URL (marked required) and save.
Don't skip this — hybrid defaults to SaaS otherwisedb-service decides whether the shared platform-wide
slack-background-agentinstance should leave your org alone by checking whether this URL is stored. Skip it and your org is indistinguishable from a SaaS org: Willow's shared instance keeps claiming your un-pinned agents' Socket Mode connections (or races your own deployment for them) instead of leaving them to your own instance from Step S3. This field is separate from — and required regardless of — the named Slack agent gateways in step 2 below, which only matter once you run more than one self-hosted instance. -
(Optional) Register additional named gateways. Only needed if you run more than one self-hosted
slack-background-agentinstance for this org. In the same dialog, add each extra instance under Slack agent gateways with a name and its reachable URL, then Save gateways. Willow uses this to pin specific agents to a specific instance.- Each saved gateway shows its Gateway ID with a copy button next to it — copy it into that instance's
gatewayIdvalue from Step S3. A single default instance (step 1) needs none of this.
- Each saved gateway shows its Gateway ID with a copy button next to it — copy it into that instance's
-
Set up each agent's Slack app (per-agent OAuth install, triggers) — see Set up Slack for background agents. On hybrid with more than one registered gateway, pick which one serves each agent from the Slackbot App section.
Both are hidden unless your org already has an External Run Service gateway from Step 5 — that's what makes the org "hybrid" in the first place.
Step S5 — Verify
kubectl get pods -n <namespace>
# Expected: slack-background-agent-xxx 1/1 Running
kubectl logs -n <namespace> -l app=slack-background-agent --tail=50
Look for a successful Socket Mode connection and the periodic load of per-agent Slack apps from db-service (roughly every 60s).
If it fails
| Symptom | Cause | Fix |
|---|---|---|
| Pod crashloops with an env error | AUTH_SECRET missing | Confirm slack-background-agent-secret exists and carries AUTH_SECRET, or supply it from a shared secret |
| Never receives channel events | Socket Mode not connecting (no per-agent token configured yet, bad xapp- token), or outbound to Slack blocked | Configure each agent's Slack app in the admin app so its xapp- token (scope connections:write) loads from db-service; re-check Step S1 |
| Connects, but replies never appear | run can't reach the agent, or the gateway isn't registered | Confirm both pods are in the same namespace so run resolves http://slack-background-agent, and complete Step S4 |
| Slack agent gateways section missing | No external gateway on the org | Complete Step 5 first |
The agent normally resolves each workspace's bot token from db-service per agent. To pin one static workspace and skip that lookup, add SLACK_BOT_TOKEN (xoxb-…) to slack-background-agent-secret.
TLS / Custom CA
If your network uses TLS inspection with a private certificate authority, add the CA certificate so run trusts SaaS endpoints:
global:
caCertificate: |
-----BEGIN CERTIFICATE-----
MIIDxTCCAq2gAwIBAgI...
-----END CERTIFICATE-----
The chart mounts the certificate and sets NODE_EXTRA_CA_CERTS automatically.
This is about run trusting outbound connections through your inspection proxy. It is unrelated to the certificate you serve on the run hostname — that still has to be publicly trusted, per Step 4.
Advanced
Rotating the Gateway Auth Secret
Owner: Willow org admin + platform engineer, together
The secret is shared, so rotation is a two-sided change with a brief window where the two sides disagree. Plan it as one operation:
-
In Settings → Gateway Settings, select Rotate Secret on the external gateway card and copy the new value.
-
Update the Kubernetes secret:
kubectl create secret generic willow-secrets \--namespace <namespace> \--from-literal=AUTH_SECRET='<new-secret>' \--dry-run=client -o yaml | kubectl apply -f - -
Restart run so it picks the new value up:
kubectl rollout restart deploy/run -n <namespace> -
Re-run connectivity check 3 from Step 8.
Between steps 1 and 3, run → SaaS calls return 401. Rotate during a quiet window.
Egress Allowlist and Host Discovery
Your firewall decides which upstreams a tool call can reach, so Willow can export the list of hosts to allow: Integrations → the ⋮ menu → Export Egress Allowlist. Nothing to configure — it reads your stored integration config.
Stdio MCP servers are the gap. They store a package to run (npx -y example-mcp), not a host to reach, so their upstreams live inside the package and the export can only mark them unknown. To close that, let run record the DNS lookups each sandbox makes:
deployments:
run:
env:
# Your in-cluster resolver — what sandboxes already resolve through.
# kubectl get svc -n kube-system kube-dns -o jsonpath='{.spec.clusterIP}'
SANDBOX_DNS_UPSTREAM: "10.96.0.10"
Requires chart 1.0.59 or later, which grants run the NET_BIND_SERVICE capability it needs to listen on port 53. If you filter the sandbox network's egress, you also need a UDP/53 rule to that network's own gateway — the RFC1918 drop otherwise covers it.
Each sandbox gets the listener as its first resolver and yours as its second, so DNS keeps working if the listener stops answering. If it cannot bind at all, run logs the reason, adds no --dns flag, and sandbox DNS is unchanged. Removing SANDBOX_DNS_UPSTREAM is the off switch.
On hybrid, run reports observed hostnames to Willow SaaS, because that is where your organization's configuration lives. The queries, the answers, and the tool traffic itself stay inside your network.
Full reference — how to read the Provenance column, what is deliberately not recorded, the metrics to watch, and troubleshooting: Egress Allowlist.
Custom Image Pull Secrets
If your cluster doesn't already have access to the quay.io/webrix registry, create an image pull secret. Restricted egress must allow quay.io and its CDN (cdn.quay.io, cdn01.quay.io, …) — image pulls redirect there, so allowlisting quay.io alone still produces ImagePullBackOff.
kubectl create secret docker-registry willow-registry \
--namespace <namespace> \
--docker-server=quay.io \
--docker-username=<robot-username> \
--docker-password=<robot-token> \
--docker-email=unused@withwillow.ai
global.imagePullSecrets is [] in the chart defaults, so no pull secret is attached to any pod until you list one. Creating willow-registry and stopping there leaves the pod in ImagePullBackOff with no indication that the credentials you just created are being ignored. Add it to your values.yaml:
global:
imagePullSecrets:
- name: willow-registry # any name you like — it just has to match the secret
If you don't have quay.io/webrix credentials yet, ask your Willow contact — they are issued per customer, not self-served.
AWS KMS Integration (Write-Only KMS)
Most hybrid deployments don't need this. Enable it only if your security policy requires that Willow SaaS can never decrypt your secrets, and only after Step 8 passes. Adding KMS to a deployment that isn't working yet makes both problems harder to diagnose.
Write-Only KMS is implemented against AWS KMS, and the run service authenticates to it with EKS IRSA. There is no Azure Key Vault or Google Cloud KMS path, and no AKS Workload Identity or GKE Workload Identity equivalent. If your cluster runs on AKS or GKE you can still use an AWS KMS key by supplying static IAM credentials to the run pod instead of IRSA, but talk to your Willow contact first — the rest of this section assumes EKS.
The one-time Gateway Auth Secret dialog in Step 5 lists KMS_KEY_ID and AWS credentials alongside AUTH_SECRET. Those belong to this optional section — a first hybrid install is complete without them.
By default, the SaaS db-service decrypts tokens and returns them to your on-prem run over the authenticated channel. With Write-Only KMS, you bring your own AWS KMS key: Willow SaaS encrypts secrets with it but can never decrypt them — only your on-prem run service can. Secrets stay encrypted at rest in SaaS Postgres in the existing EncryptedPayload format, and plaintext exists only inside your cluster.
Who you need: an AWS / IAM admin for K1–K3, the Willow org admin for K4, and the platform engineer for K5.
The permission split
The two sides get different permissions on the same customer-owned key. This is the part that's easy to get wrong — your run service does not need encrypt permissions.
| Who | Where it runs | Permissions on your KMS key |
|---|---|---|
| Your run service | Your EKS (via IRSA) | kms:Decrypt (+ kms:DescribeKey) — decrypt only |
| Willow SaaS db-service | Willow cloud | kms:GenerateDataKey + kms:Encrypt (+ kms:DescribeKey) — encrypt only, never Decrypt |
When an OAuth token expires, your run service decrypts the refresh token locally, refreshes it with the provider directly, and sends the new tokens back to the SaaS db-service to re-encrypt and store — so run never needs GenerateDataKey.
Step K1 — Create a KMS key in your AWS account
Create a symmetric encryption key (or reuse an existing one) in the same AWS region you'll configure on the run service. Note its key ARN, e.g. arn:aws:kms:us-east-1:<YOUR_ACCOUNT_ID>:key/<KEY_ID>.
Step K2 — Create the run IRSA role (Decrypt only)
Create an IAM role trusted by your EKS cluster's OIDC provider and bound to the run service account. Use a trust policy like this — replace <OIDC_ID> with your cluster's OIDC provider ID (EKS → your cluster → Overview → OpenID Connect provider URL), and set the sub to match the run service account (see the note below):
{
"Version": "2012-10-17",
"Statement": [
{
"Effect": "Allow",
"Principal": {
"Federated": "arn:aws:iam::<YOUR_ACCOUNT_ID>:oidc-provider/oidc.eks.<YOUR_REGION>.amazonaws.com/id/<OIDC_ID>"
},
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"oidc.eks.<YOUR_REGION>.amazonaws.com/id/<OIDC_ID>:aud": "sts.amazonaws.com",
"oidc.eks.<YOUR_REGION>.amazonaws.com/id/<OIDC_ID>:sub": "system:serviceaccount:<namespace>:willow-run"
}
}
}
]
}
Attach a decrypt-only permissions policy:
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "RunDecryptOnly",
"Effect": "Allow",
"Action": ["kms:Decrypt", "kms:DescribeKey"],
"Resource": "arn:aws:kms:<YOUR_REGION>:<YOUR_ACCOUNT_ID>:key/<KEY_ID>"
}
]
}
sub must match the run service account nameThe sub condition must equal system:serviceaccount:<namespace>:<run-service-account>. The run service account name is whatever you set in deployments.run.serviceAccount.name (Step K5) — e.g. willow-run. If you leave serviceAccount.name unset, the chart names it run-sa, so the sub must then be system:serviceaccount:<namespace>:run-sa. A mismatch here is the most common cause of sts:AssumeRoleWithWebIdentity and decrypt (AccessDenied) failures.
Step K3 — Grant Willow SaaS encrypt access
Does a Willow account need access to my key? Yes — but only to encrypt. There are two ways to grant it; pick one.
Option A — Provide scoped IAM credentials to Willow (no cross-account access needed).
In the admin app's External KMS dialog (Step K4), you enter an AWS access key / secret for an IAM user in your own account that has kms:GenerateDataKey + kms:Encrypt on the key. Willow stores these encrypted and uses them to encrypt. With this option, no Willow AWS account touches your key directly.
Option B — Cross-account key policy.
Grant Willow's SaaS principal (AWS account 992382826040) encrypt access via your key policy. Use this if you'd rather not hand over static IAM credentials.
{
"Sid": "AllowWillowSaaSEncryptOnly",
"Effect": "Allow",
"Principal": {
"AWS": "arn:aws:iam::992382826040:role/<WILLOW_SAAS_PRINCIPAL>"
},
"Action": ["kms:GenerateDataKey", "kms:Encrypt", "kms:DescribeKey"],
"Resource": "*"
}
The exact IAM principal ARN on the Willow side depends on the SaaS configuration. Ask your Willow contact for the precise ARN before applying Option B — don't assume it.
Also make sure your key policy allows your own run role to decrypt (from Step K2), since the key is in your account:
{
"Sid": "AllowRunDecryptOnly",
"Effect": "Allow",
"Principal": {
"AWS": "arn:aws:iam::<YOUR_ACCOUNT_ID>:role/<YOUR_RUN_IRSA_ROLE>"
},
"Action": ["kms:Decrypt", "kms:DescribeKey"],
"Resource": "*"
}
Step K4 — Configure External KMS in the admin app
Owner: Willow org admin
- Go to Admin → Settings → Gateway Settings.
- On your external gateway's card, select Configure KMS. The External KMS Configuration dialog opens.
- Enter the KMS Key ARN — the full ARN from Step K1, not the bare key ID. It must match the
KMS_KEY_IDyou set in Step K5. - Enter the AWS Region.
- If you chose Option A, enter the AWS Access Key ID and AWS Secret Access Key for the IAM user with
GenerateDataKey+Encrypt. - Select Save KMS Config.
The card then shows a KMS badge, and the button becomes Edit KMS. This tells the SaaS db-service to route encryption for this org through your key.
Step K5 — Point the run service at the key
Owner: Platform engineer
Add the KMS settings and IRSA service account annotation to your values.yaml:
deployments:
run:
serviceAccount:
create: true
# SA name. Must match the trust policy `sub` from Step K2
# (system:serviceaccount:<namespace>:willow-run). If omitted, the chart
# names it "run-sa".
name: willow-run
annotations:
# The decrypt-only IRSA role from Step K2
eks.amazonaws.com/role-arn: "arn:aws:iam::<YOUR_ACCOUNT_ID>:role/<YOUR_RUN_IRSA_ROLE>"
env:
# Enables local KMS decryption on the run service
KMS_KEY_ID: "arn:aws:kms:<YOUR_REGION>:<YOUR_ACCOUNT_ID>:key/<KEY_ID>"
AWS_REGION: "<YOUR_REGION>"
Do not set ENCRYPTION_KEY (that selects the static-key on-prem mode instead of KMS), and do not set AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY — let IRSA supply credentials to the run pod.
Once KMS_KEY_ID is set, the run service automatically requests encrypted payloads from the SaaS db-service and decrypts them locally — no additional flag is required.
Verify: re-run Helm, then confirm the pod picked up IRSA and can use the key:
kubectl get pod -n <namespace> -l app=run \
-o jsonpath='{.items[0].spec.serviceAccountName}{"\n"}'
# Must equal the `sub` you used in Step K2
kubectl exec -n <namespace> deploy/run -- env | grep AWS_
# Expect AWS_ROLE_ARN and AWS_WEB_IDENTITY_TOKEN_FILE, injected by IRSA
Then connect an integration whose token is stored encrypted and call one of its tools — a successful call proves the round trip (SaaS encrypts with your key, run decrypts locally).
If it fails
| Symptom | Cause | Fix |
|---|---|---|
AssumeRoleWithWebIdentity AccessDenied | Trust policy sub doesn't match the pod's service account | Set the sub to system:serviceaccount:<namespace>:<sa>, where <sa> is deployments.run.serviceAccount.name (or run-sa if unset) |
AWS_ROLE_ARN not in the pod env | Annotation missing, or on the wrong service account | Confirm the eks.amazonaws.com/role-arn annotation is under deployments.run.serviceAccount.annotations |
Decrypt fails with AccessDenied | Key policy doesn't grant your run role, or the region is wrong | Check the AllowRunDecryptOnly statement and that AWS_REGION matches the key's region |
| Willow can't encrypt | Option A credentials lack GenerateDataKey/Encrypt, or Option B principal is wrong | Re-check Step K3; confirm the exact Willow principal ARN with your Willow contact |
| Key mismatch errors | The ARN in the admin app differs from KMS_KEY_ID | They must be identical, and both full ARNs |
Using a shared Kubernetes Secret across services
You can also set global.secretName to share one secret across all deployed services rather than naming it per service, or use sealedSecrets for encrypted secret management. Secret values always override the ConfigMap, so anything supplied this way wins over global.dbAuthSecret.
Troubleshooting index
Each step above has its own troubleshooting table — start there, since it tells you what else to have already verified. This index maps a symptom to the step that owns it.
| Symptom | Go to |
|---|---|
| Can't find Configure Gateway / On-Prem in the admin app | Step 5 — the path is Settings → Gateway Settings, and On-Prem only appears once your external gateway is the default |
| Users are handed SaaS MCP URLs instead of your on-prem one | Step 5 — Set as Default was skipped |
| 401 from the SaaS db-service proxy | Step 8 check 3 — auth secret mismatch, possibly the chart's placeholder default |
| 404 on tool listing | Step 8 check 3 — DB_SERVICE_URL missing the /api/ prefix |
| Gateway shows Unreachable; Tools tab won't list live | Step 3 — allowlist the region's egress IP (3.130.252.122 US, 3.120.156.158 EU), or adopt Isolated Mode deliberately. 3.136.98.54 from older docs is stale |
| Hostname resolves but connections time out from outside | Inbound network — per-cloud checks for AKS, EKS, and GKE |
| Works from inside your network, not from outside | Step 3 — a public DNS record does not mean the port is open |
| Ingress created but never programmed; hostname 404s | Match your ingress controller — className doesn't match any IngressClass |
| Using the AKS application routing add-on | Match your ingress controller — different class, and certificates come from Key Vault |
| Which certificate do I need, and how do I install it? | Choosing a certificate |
cert-manager Certificate reports Issuer not found | Step 4 — the chart doesn't install cert-manager or create the ClusterIssuer |
cert-manager Certificate stuck at Ready: False | Verifying certificates — HTTP-01 can't validate while inbound is closed |
| Run hostname resolves to CDN addresses you didn't configure | Step 2 — a proxied wildcard record, or a stale resolver cache |
curl https://willow.<domain>/mcp returns 401 | Step 4 — expected; /mcp requires authentication |
| Ingress serves a "fake certificate" | Choosing a certificate — the TLS secret is missing or misnamed |
Endpoint returns 503 | No healthy backend — usually ImagePullBackOff |
Endpoint returns 404 from the ingress | Hostname mismatches |
| Two Willow releases installed, unclear which is live | Hostname mismatches — helm list -A |
| Test Run, guard playground, condition pickers, or tool refresh fail | What You Lose Without Inbound — same missing inbound, and these have no fallback |
| Background agent tool calls fail but end-user clients work | Step 3 — the agent runtime calls from the platform's cloud, which Willow's NAT IP doesn't cover |
| MCP clients can't authenticate | Step 8 check 4 — wrong CONNECT_URL |
OAuth discovery resource has a trailing slash you didn't configure | Step 8 check 4 — expected; run normalizes BASE_URL. Compare hostnames, not exact strings |
In-cluster wget against the ingress IP returns 404 | Inbound network step 3 — the check needs a Host header, or the ingress matches no rule |
| MCP clients can't connect at all | Step 2 and Step 4 — DNS, ingress, or certificate |
Cold start hangs ~60s then times out (Daemon request timeout, MCP -32001); warm calls succeed | Step 4 — SSE buffered or idle-timed-out, usually by a cloud LB |
ImagePullBackOff after creating the pull secret | Custom Image Pull Secrets — global.imagePullSecrets is empty by default, so the secret is ignored until you list it |
ImagePullBackOff after allowlisting quay.io | Custom Image Pull Secrets — image pulls redirect to the quay CDN (cdn.quay.io, cdn01.quay.io, …) |
Pod Pending with Insufficient cpu | Step 7 — no single node has enough allocatable CPU left beside your ingress controller and system pods |
ImagePullBackOff or pod Pending | Step 7 |
| Write-Only KMS on AKS or GKE | AWS KMS Integration — AWS-only today, and the documented path assumes EKS IRSA |
| Wrong org's data appears | Step 6 — global.org must be the exact org slug |
| Extra pods running | Step 7 — disable app, connect, db-service |
| Secrets fail to decrypt on run (KMS mode) | Step K5 |
slack-background-agent problems | Step S5 |
If you're still stuck, contact Willow support with your values.yaml (secrets redacted), kubectl get pods -n <namespace>, kubectl logs -n <namespace> -l app=run --tail=200, and which numbered check in Step 8 first failed.