Back to overview
Resolved

Postmortem: ElastiCache Routing Misconfiguration & Multi-Region Outage

Sep 3, 2026 at 2:57pm UTC
Affected services
Dashboard & API

Resolved
Sep 3, 2026 at 2:57pm UTC

Postmortem: ElastiCache Routing Misconfiguration & Multi-Region Outage

Date: September 2, 2026

Incident Duration: ~6 hours 30 minutes
Impact: Partial service degradation across EU and US-West regions due to primary cache inaccessibility.

Executive Summary

On September 2, 2026, a scheduled ElastiCache service upgrade triggered a failover of Pangolin’s primary cache node to a replica located within a newly provisioned VPC subnet. A routing table misconfiguration prevented traffic from the EU and US-West regions from reaching this new subnet. The sudden loss of cache access caused thread pool starvation and cascading service failures across downstream applications. Service was fully restored after updating cross-region routing rules and recycling connection pools across affected clusters.

Root Cause Analysis

  1. Subnet Provisioning Gap: The maintenance event failed over the ElastiCache primary to a replica inside a newly provisioned subnet.
  2. Missing Routing Entries: Cross-region VPC peering routes and security group rules for EU and US-West were not updated to include the cidr block of the newly provisioned subnet.
  3. Cascading Failure: Downstream microservices lacked resilient circuit breakers for cache unavailability. Unhandled connection timeouts quickly exhausted application worker threads, causing cascading service crashes.
  4. Stale Connection Persistence: Even after routing was restored, persistent TCP connection pools in certain client nodes continued trying to reach the old unreachable endpoints until services were manually restarted.

Impact Assessment

  • Regions Affected: EU and US-West regions experienced severe latency and request failures. US-East remained largely functional due to direct local routing.
  • Cache Layer: Read/write operations to the primary ElastiCache cluster failed completely for impacted regions.
  • Client/Node Connections: Certain client services required explicit service restarts to drop stale sockets and re-establish active connections to the updated cluster endpoint.

Corrective Actions & Preventative Measures

Completed

  • Route Remediation: Updated all cross-region VPC routing tables and security groups to allow full traffic access to the new ElastiCache subnet.
  • Connection Refresh: Executed targeted pod and service restarts across EU and US-West deployments to clear dead TCP sessions.
  • Customer Advisory: Published status updates advising self-hosted nodes and client connections to restart lingering services.

Planned

  • Infrastructure as Code (IaC) Enforcement: Update CDK subnet modules to automatically update all cross-region peering route tables whenever new subnets are created.
  • Pre-Flight Failover Audits: Implement automated integration tests that validate cross-region connectivity across all standby cache nodes prior to scheduling upgrades.
  • Graceful Cache Degradation: Implement circuit breakers and fallback mechanisms so that cache connectivity loss degrades performance gracefully rather than causing total service failures.