Postmortem: Incident: Storefront errors and degraded performance
Date: September 13, 2026
Duration: 12:09–12:47 UTC and 13:18–13:26 UTC, approximately 45 minutes of customer impact
Summary
Some customers and end users experienced intermittent storefront errors and degraded performance during two periods on September 13. Services were restored, additional traffic protections were implemented, and the platform is operating normally.
What happened
Our infrastructure experienced an unusual spike in incoming traffic that exceeded the platform’s normal operating range. The increased load placed pressure on shared platform resources, causing intermittent errors and degraded performance across storefronts.
Timeline of events
- 12:09 UTC: Customer-visible degradation began.
- 12:42 UTC: Additional traffic controls were applied.
- 12:47 UTC: Services recovered from the first period of impact.
- 13:18 UTC: A second period of degradation began.
- 13:22 UTC: Traffic controls were further enhanced.
- 13:26 UTC: Services were restored.
- 16:05 UTC: Additional protections were confirmed in place, and the platform was still operating normally.
Root cause
An abnormal traffic pattern generated a sudden volume of incoming requests beyond the platform’s normal operating range. Processing these requests placed excessive pressure on shared platform resources, which led to degraded storefront performance and intermittent errors.
What we did
Our engineering team investigated the traffic pattern, applied traffic controls, and restored services. We then enabled additional protections to allow legitimate traffic through while limiting disruptive requests and monitored to confirm platform stability.
What we’re doing to prevent recurrence
We are strengthening the platform in four areas:
- Expanding alerting for abnormal traffic and elevated platform resource usage.
- Strengthening traffic handling and protection across storefronts.
- Improving storefront caching and application efficiency to reduce the resource cost of unexpected traffic.
- Increasing service isolation so that traffic pressure on one part of the platform is less likely to affect other surfaces.
Customer impact
Affected surfaces:
- Admin Portal
- Storefront
- API - V1
- API - V2
User-facing experience: Intermittent errors, degraded performance, and temporary unavailability
Data integrity: No data loss, no security breach, no impact to customer accounts or stored content
Thank you for your patience while our engineering team worked through this. We hold a high standard for platform reliability, and we are applying what we learned here.