Projects

Shared ingress on Envoy Gateway

Every new customer needs a load balancera route.

The evaluation, routing design and rollout were mine.

Our product runs one server per customer, about 270 across five AWS regions. Customers' devices connect on fourteen ports, most of them custom TCP, not HTTP.

Why it was needed

Each server had its own AWS load balancer, so cost and DNS records grew with every customer, and an AWS account limit was close. Sharing one was hard: every customer uses the same port numbers, and customers' firewalls allow our IP address, so it could not change.

AWS region · one of fiveEKS clusterTLS + SNIManaged devices14 ports · TLSShared NLBone Elastic IPEnvoy GatewaySNI routing · passthroughCustomer serverown TLS certificateCustomer serverown TLS certificate… one per customer~270 across 5 regions
One Elastic IP and load balancer per region. Envoy routes each connection by the server name in its TLS handshake and passes it through, so certificates stay on each server.

How it was done

  1. Tested the main ingress options against two requirements, and wrote down why each one failed.

  2. Routed by server name with TLSRoute, because TCPRoute cannot tell customers apart on the same port.

  3. Proved all fourteen ports on staging with our architect's upstream NATS fix, so every protocol sends a name to route on.

  4. Rolled out in batches over three months behind one Elastic IP, with a switch back to per-customer load balancers.

Problems on the way

  • The first batch of 20 customers could not reach their servers.

    Envoy's default circuit breaking on the shared gateway. Fix: Turned it off with a BackendTrafficPolicy before the next batch.

  • Half of the second batch had not updated their DNS.

    Their own domains still pointed at the old load balancer. Fix: Checked each domain first and moved the late ones one by one.

  • Months later, one stale route took down port 443 for a region.

    Two routes claimed one hostname, so Envoy rejected the whole listener. Fix: Alerts on the gateway's rejections and on applications left out of sync.

Sharing moved the risk. It did not remove it.

Results

  • N → 1load balancers per region, behind one Elastic IP
  • 4options rejected, each for a written reason
  • 4production clusters moved in batches