
CircleCI Status
Real-time updates of CircleCI issues and outages
CircleCI status is Operational
CircleCI Docker Jobs
CircleCI Machine Jobs
CircleCI Pipelines & Workflows
CircleCI API
CircleCI UI
CircleCI Notifications & Status Updates
Lior Grossman
Maker of StatusSight
I built StatusSight. Here's what I'm building next with AI.
In Lior Builds, my free newsletter, I share one story from a product I built with AI, plus short notes on the tools and models worth using: what worked, what did not, and the numbers behind it. Every week or two.
Active Incidents
Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.
Identified: A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.
Identified: Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.
Identified: Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
There is an increased wait time for jobs on Docker Gen2.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're still seeing spikes in task wait times but on average wait times are looking better for all Docker Gen2.
Identified: Note: delays will be more noticeable for people running on Docker xlarge and above.
Identified: There is an increased wait time for jobs on Docker Gen2.
We are investigating a rise in infra fails for customer jobs.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're seeing signs of recovery. Docker Gen2 wait times are slightly elevated and we're monitoring.
Investigating: We are investigating a rise in infra fails for customer jobs.
Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: The incident has been resolved. Thank you for your patience
Monitoring: We are seeing signs of recovery and monitoring the situation.
Investigating: Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Postmortem: ## Summary
On October 6, 2026, from 15:30 to 20:30 UTC, CircleCI customers experienced delays creating pipelines and starting workflows and jobs. From 15:48 to 16:38 UTC, no workflows or jobs started, and pipelines triggered from 16:02 UTC failed. Shorter, intermittent delays began earlier, at 13:30 UTC. Data in the CircleCI UI, notifications, and commit status updates to version control providers were also delayed.
The incident started in the database behind our workflow orchestration service, which coordinates every pipeline, workflow and job on CircleCI. Routine maintenance on the database's largest tables wrote transaction logs faster than the database could archive them, and the logs filled the storage volume set aside for them. To keep the database from becoming read-only, our engineers moved the logs to the database's main storage volume. That move took the database offline for 50 minutes. When the database came back, it processed work at about half its normal rate until our engineers changed a database setting at 18:19 UTC. We then worked through the backlog, and job start times returned to normal by 20:30 UTC.
Customers whose jobs failed during this window may rerun them.
The original status page can be found here.
Background
When you trigger a pipeline, CircleCI hands it to a workflow orchestration service. That service tracks the state of every workflow and job, decides when each job is ready to run, and passes ready jobs to our execution fleet. It also drives the job data you see in the UI and the notifications and commit statuses we send when work finishes.
The orchestration service stores its state in a database. The database writes every change to a transaction log before applying it, and it keeps those logs on a dedicated storage volume. It archives the logs continuously so it can reuse that space. If archiving falls behind and the volume fills, the database stops accepting writes, and no new work can move forward.
The database also runs a periodic maintenance process (autovacuums) on each table to keep its internal records valid over time. On a very large table, this process reads every page of the table and writes a large volume of transaction logs.
What Happened
(All times UTC)
From 10:02 on October 6, our monitoring raised short-lived alerts about errors in the orchestration service. Each alert cleared on its own within about 15 minutes. Our engineers had seen similar brief alerts before, and the system showed no other signs of stress, so they followed the runbook each time, and each alert cleared on its own partway through and also started work to make the alerts less sensitive.
Around 12:40, along with the regular workload which generates its own transaction logs, the maintenance process on several of the database's largest tables began writing additional transaction logs faster than the database could archive them, and the log volume started to fill.
At 13:30, the orchestration service began to slow down. Through 15:30, some pipelines and jobs took longer to start, in short bursts. At 13:35, our monitoring alerted again, this time alongside related alerts from several other services. Infrastructure Engineers followed the alert runbooks, clearing up some of the alerts. The alerts returned again at around 14:30 and after investigating further, we declared an incident at 14:54. We should have recognized the combination of alerts sooner, and we have covered how we are fixing that below.
After declaring the incident, our engineers traced the slowdown to the maintenance process (anti-wraparound autovacuum). The database restarts this process automatically if it is cancelled, so our engineers changed its settings to help it finish faster. By 15:30, the delays were continuous, and some jobs began to fail.
At 15:43, the transaction log volume was close to full. If it filled, the database would stop accepting writes and no work could run. Our engineers decided to move the transaction logs onto the database's main storage volume, which had plenty of free space. The move required the database to go offline, and we could not predict in advance how long that would take.
The move started at 15:47. From 15:48 to 16:38, the orchestration service could not process any work. Customers could still trigger pipelines, but no workflows or jobs started, and from 16:02 newly triggered pipelines failed. We posted to our status page at 15:58 and raised it to a major outage at 16:20. While the database was offline, our engineers prepared a replacement database as a fallback. The original database came back first, so we kept it in service.
At 16:38, the database came back online and work started flowing again. With transaction logs and regular data now sharing the same storage, the database ran more slowly than before. From 16:38 to 18:19, the orchestration service processed work at about half its normal rate, and jobs waited tens of minutes to start, up to about 50 minutes at the longest. During that time, the database had to finish archiving its backlog of transaction logs before we could move them back to a dedicated volume. We also investigated and implemented several database parameter changes to stop further maintenance runs from starting and prepared ways to reduce the work reaching the service. Along with the replacement database, we also prepared a complete stack of the service during that time so that we could move to it at the risk of data loss and opted against that move.
At 18:19, our engineers changed a database setting to cut the time each write spent waiting on storage. Processing speed recovered, and the backlog began to clear. To clear it faster, we added database and server capacity to the systems that hand jobs to our execution fleet. By about 19:55, new jobs were again starting on time.
The backlog then reached our execution fleet. From 19:49 to 20:30, some Docker jobs on larger resource classes waited up to about 10 minutes to start while capacity scaled up. By 20:30, job start times were back to normal. We moved the status page to monitoring at 20:32 and resolved the incident at 21:19.
About 3,500 jobs failed to start during the incident. Customers may rerun these jobs. We found no evidence that CircleCI ran any job more than once. Customers who reran a workflow while the original run was still delayed may have seen both runs complete. Annual plan customers can work with their account team to review usage.
Future Prevention and Process Improvement
We are taking the following steps to prevent a recurrence and improve our response time:
We are adding alerts on the conditions that caused this incident. Our monitoring caught the slowdown, but we had no alert on how full the transaction log volume was or on how far archiving had fallen behind. We are adding alerts on both, along with alerts on the database maintenance that drives them, so we can act earlier. We are also revisiting all our alerts to ensure noisy or sensitive alerts don’t hide real issues.
We are tuning how maintenance process (autovacuum) runs on our database. We are looking at the frequency and the aggressiveness with which the autovacuums run on our database. We are also tracking the size of these tables over time so they stay within safe limits.
We are removing the single transaction log bottleneck. All of the orchestration service's writes currently go through one database and one transaction log, so when that log fell behind, every pipeline, workflow and job slowed down with it. We are evaluating two approaches: splitting the service across multiple database instances, each with its own transaction log, and moving it to a distributed database built to spread writes across many nodes. Either approach would limit the effect of a backlog to part of the workload.
We are making our systems wait for the orchestration service to recover. During the incident, some jobs failed because our systems gave up after retrying. We are working on changing that behavior.
We are improving our status page updates during long incidents. Customers told us our early updates did not give enough detail. We are updating our guidance so updates include specific impact and timing sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve this incident. Please reach out to our support team with any questions or concerns.
Resolved: Between approximately 15:30 UTC and 20:30 UTC on October 6, 2026, customers experienced delays and failures affecting pipelines, workflows, and jobs.
From approximately 16:02 UTC to 16:39 UTC, pipelines were failing and no new workflows or jobs could start. After service resumed, workflows and jobs continued to start with delays until approximately 20:00 UTC. Start times averaged up to approximately 40 minutes at their peak, and a very small number of workflows waited more than 1 hour. Outbound notifications and webhooks were also delayed until approximately 19:00 UTC, typically by about a minute.
As the backlog of delayed work cleared, the resulting increase in demand placed pressure on our execution capacity, causing further delays for Docker and Linux jobs. From approximately 19:45 UTC to 20:30 UTC, these jobs waited in the queue longer than usual before starting. At the peak, around 20:05 UTC, waits reached up to approximately 8 minutes for 2xlarge gen2, 6 minutes for xlarge gen2, and less than 5 minutes for a few other resource classes.
The issue has been resolved, and all affected functionality has returned to normal. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Workflows and jobs are starting normally, and wait times for Docker and Linux jobs have returned to normal. Our engineers are continuing to monitor this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: The backlog of delayed work has been cleared, and workflows are starting again with no delays in notifications. However, customers using Docker and Linux jobs across multiple resource classes, including medium, large, and 2xlarge gen2, may now see jobs waiting in the queue for few minutes before they start. Our engineers are continuing to work on mitigating this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Delays for workflows and jobs to start are continuing to decrease. Most are now starting within approximately 10 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour. If your job failed to start, please rerun the job. Delays in outbound notifications have improved.
The remaining delay is due to a backlog of work that built up during the incident, and that backlog has reduced by more than half. We expect the backlog to be cleared in approximately 15 minutes. We will share another update once it has cleared, or within the next 30 minutes at the latest.
Identified: Customers continue to experience delays. Most are now starting within approximately 20 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour.
Our engineers have made changes that have improved service stability, and the remaining delay is due to a backlog of work that built up during the incident. The backlog is decreasing, and we are focused on clearing it faster. We will share more information as progress continues.
We thank you for your patience while our engineers work through the backlog. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays before workflows and jobs start, in some cases up to approximately 30 minutes. This is improved from earlier in the incident, but delays remain. Delays in outbound notifications have improved.
We thank you for your patience while our engineers continue to work toward restoring normal job start times. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays between jobs in a workflow, as well as delays in outbound notifications.
We thank you for your patience while our engineers continue to investigate and work on mitigating the issue. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: While Pipelines, workflows, and jobs are running again, customers may continue to see longer than usual delays between jobs in a workflow, and outbound notifications are delayed by approximately 1 minute. Notifications and status updates to your version control system are beginning to go out again.
Our engineers are continuing to work toward a full recovery.
Identified: As of 16:39 UTC, we are starting to see recovery for new pipelines and workflows being submitted, and jobs are beginning to run again. Our engineers are continuing to work on the recovery and will share more information as soon as we have it.
Identified: As of 16:02 UTC, pipelines are failing entirely, so no new workflows or jobs are running at this point. This is in addition to the delays and failed jobs reported earlier. Our engineers are continuing to work on mitigating the issue, and we will share more information as soon as we have it.
Investigating: Since 15:47 UTC, customers may also see failed jobs, in addition to the delays previously reported. Our engineers are continuing to investigate and mitigate the issue.
Investigating: What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Recently Resolved Incidents
There is a delay in our data infrastructure causing delays before build data and some notifications become visible. We are working on a mitigation.
Resolved: Between 11:15 UTC and 23:47 UTC on October 9, customers experienced delays and failures across pipelines, workflows and jobs, along with delays in the CircleCI UI, API and notifications.
From 11:15 to about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and Slack and email notifications were not sent. From about 13:00 to 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
From about 15:15 to 20:48 UTC, workflows and jobs started with increasing delays, of up to about an hour, and the CircleCI UI and API fell behind by up to about 2 hours. During this time, some jobs failed to start and some in-progress workflows failed.
From 20:48 to about 22:10 UTC, we temporarily paused starting new workflows and jobs so our systems could recover. New pipelines were accepted and queued during this time. From about 22:10 to 23:47 UTC, queued workflows and jobs resumed and were processed. Some of them had waited up to about 2 hours.
The issue has been resolved and all affected functionality has returned to normal. Usage charge data on the Project Usage page is not yet up to date and will be updated once our data backfills complete. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our engineers worked on implementing a fix.
Monitoring: Workflows and jobs are starting and running normally, and the CircleCI UI and API are up to date. We are continuing to monitor our systems closely to make sure they remain stable.
We will provide another update by 02:00 UTC.
Identified: Workflows and jobs have continued to start normally, within seconds, and the CircleCI UI and API remain up to date. Slack and email notifications delayed earlier in the incident are still being delivered.
Usage charge data on the Project Usage page is stale.
We are continuing to monitor our systems closely and will provide another update by 01:15 UTC.
Identified: Since our last update, the backlog of queued workflows and jobs has cleared, as of about 23:47 UTC. Workflows and jobs are now starting within seconds, and the CircleCI UI and API are up to date. We are monitoring our systems closely to make sure they remain stable. We will provide another update by 00:30 UTC.
Identified: We are still working through a significant backlog of queued workflows and jobs. Wait times are improving but they remain longer than normal and some jobs have waited over an hour. The CircleCI UI and API remain up to date. Please avoid retriggering pipelines or workflows while the backlog clears, as this may create duplicates. We will provide another update by 00:00 UTC.
Identified: A large number of queued workflows and jobs have started, and the CircleCI UI and API have caught up and are showing updates normally with few delays. We are continuing to work through the remaining queued work. Jobs from the queue may have waited about an hour or more to start, and some Linux jobs are waiting up to about 10 minutes for capacity.
Please continue to avoid retriggering while queued work continues to process.
We will provide another update by 23:30 UTC.
Identified: Since our last update, we have begun gradually resuming work, and some queued workflows and jobs are starting again. Most queued work is still waiting, and jobs that do start may have waited over an hour. We are increasing the rate carefully to avoid overloading our systems. Delayed updates continue to appear in the CircleCI UI and API.
Please continue to avoid retriggering pipelines or workflows, as this may create duplicates as queued work resumes.
We will provide another update by 23:00 UTC.
Identified: Since our last update, most of the pipelines that were delayed earlier in the incident have now been processed, and their workflows are queued to start. New workflows and jobs remain paused while we prepare to resume work safely. Delayed updates continue to appear in the CircleCI UI and API.
Please continue to avoid retriggering pipelines or workflows, as this may create duplicates once work resumes.
We will provide another update by 22:30 UTC.
Identified: Since about 20:48 UTC, new workflows and jobs are not starting. We have temporarily paused starting new work so that our systems can catch up on work already in progress, including delayed updates to the CircleCI UI and API. Older updates are now appearing in the UI and API as this backlog clears. New pipelines are still being accepted and queued
Please avoid retriggering pipelines or workflows during this time, as this may create duplicates once work resumes.
We will provide another update by 22:00 UTC.
Identified: Jobs are starting faster and fewer are failing: most jobs are now starting within about 4 minutes, down from about 6 minutes, and the longest waits have dropped from about 30 minutes to about 20 minutes.
Pipelines from earlier in the incident are now being worked through.
The CircleCI UI is still behind, and some older updates may take longer to appear. Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI.
We will provide another update by 20:30 UTC.
Identified: Since our last update, jobs are starting faster: most are now starting within about 6 minutes, down from about 8 minutes. The CircleCI UI is catching up but still about an hour behind. New pipelines are being processed as they arrive.
Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI.
We will provide another update by 20:00 UTC.
Identified: What's changed since our last update
- Jobs are starting faster. Most are now starting within about 8 minutes, down from about 15 minutes, and the longest waits have dropped from over an hour to about 45 minutes.
- The CircleCI UI is beginning to catch up, but it is still behind. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all.
- Pipelines are being processed faster, but new pipelines are still delayed.
Still ongoing
- Some workflows and jobs are still failing.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is stale.
Next update We will provide another update by 19:30 UTC.
Identified: What's changed since our last update
- The CircleCI UI is about an hour behind, and up to about 90 minutes for some updates. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all.
- Jobs that are running are starting faster. Most are now starting within about 15 minutes, down from about 28 minutes, though some are still waiting over an hour.
- New pipelines are increasingly delayed.
- Some workflows are still failing. Workflows that could not be processed in time are being marked as failed.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
Still ongoing: Usage charge data on the Project Usage page is stale.
Next update We will provide another update by 19:00 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What to expect
- Since about 17:00 UTC, we are seeing an increase in delays. Most jobs are now taking about 25 minutes to start, and some are waiting up to about an hour.
- Since about 17:15 UTC, some in-progress workflows have failed and a small number of jobs are failing to start.
- New pipelines are being queued and are taking longer to start.
- The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Our engineers continue to work to resolve the issue at hand.
Next update We will provide another update by 18:30 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Since about 17:00 UTC, delays have increased significantly. Most jobs are now taking about 12 minutes to start, and some are waiting up to about 30 minutes.
- Since about 17:15 UTC, some jobs are failing to start.
- Pipelines are being processed more slowly, so some may take longer to start.
- The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 18:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs are holding steady. Most jobs are starting within about 4 minutes, and some are waiting up to about 8 minutes, down from about 25 minutes.
- The CircleCI UI remains about 20 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 17:30 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs have stopped increasing and are beginning to ease. Most jobs are starting within about 4 minutes, and some are waiting up to about 25 minutes.
- The CircleCI UI is now about 20 minutes behind, down from about 30 minutes. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 17:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Since about 15:15 UTC, workflows and jobs have been starting with increasing delays. Most jobs are currently starting within about 4 minutes, and some are waiting up to about 20 minutes.
- The CircleCI UI remains about 30 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 16:30 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What's impacted Customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Workflows and jobs are running on time.
- The CircleCI UI is about 30 minutes behind. New workflows and status changes may take that long to appear.
- Our engineers are actively working to speed up processing so the UI can catch up.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce this delay.
Next update We will provide another update by 16:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- The backlog of pipelines and workflows from the incident has cleared.
- Since about 14:45 UTC, most workflows and jobs are starting within about 15 seconds, and start times are continuing to return to normal.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers monitor the recovery.
Next update We will provide another update by 15:20 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- The backlog of workflows from pipelines triggered during the incident cleared at about 14:40 UTC.
- Most jobs are now starting within about 30 seconds, and some are waiting up to about 3 minutes.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers continue to work on the underlying issue.
Next update We will provide another update by 15:10 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs are improving. Most jobs are now starting within about 3 minutes, down from about 10 minutes at 13:50 UTC.
- Some workflows from pipelines triggered during the incident are still waiting up to about 40 minutes to start.
- Working through the backlog is taking slightly longer than expected. We now expect to clear it by about 14:45 UTC.
- Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 15:00 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- Pipelines are being processed again, including those triggered during the incident.
- Since about 13:25 UTC, workflows and jobs have been starting with increasing delays as we work through the backlog. Most jobs are currently starting within about 10 minutes, and some are waiting up to about 25 minutes.
- We expect to work through this backlog by about 14:35 UTC.
- Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 14:30 UTC.
Identified: Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications.
What can you expect:
- Since about 13:00 UTC, pipelines are being processed again and jobs are starting.
- Most jobs are currently starting within about 4 minutes, and some are waiting up to about 10 minutes. Pipelines triggered during the incident are still being processed and may take longer to start.
- Outbound webhooks and notifications have resumed.
- Pipelines are still being processed from the backlog, so retriggering may result in duplicate pipelines.
Thank you for your patience while our engineers work to resolve this.
Next update: We will provide another update by 14:00 UTC.
Identified: What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications.
What can you expect Since 11:15 UTC:
- Pipelines are still being created while this is ongoing but not yet processed, so retriggering may result in duplicate pipelines.
- Outbound webhooks and transactional notifications (Slack and email) are not being sent.
- Pipeline statuses may not be visible.
- Usage charge data on the Project Usage page will be stale.
Thank you for your patience while our engineers work to resolve this.
Next update: We will provide another update by 13:30 UTC.
Identified: There is a delay in our data infrastructure causing delays before build data and some notifications become visible. We are working on a mitigation.
CircleCI Outage Survival Guide
You're checking on CircleCI - but is your own site converting?
Outages aren't the only thing that costs you visitors.
Get a visual audit that shows exactly where your site loses conversions.
Free, takes 2 minutes.
CircleCI Components
CircleCI Docker Jobs
There is an increased wait time for jobs on Docker Gen2.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're still seeing spikes in task wait times but on average wait times are looking better for all Docker Gen2.
Identified: Note: delays will be more noticeable for people running on Docker xlarge and above.
Identified: There is an increased wait time for jobs on Docker Gen2.
CircleCI Machine Jobs
Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: The incident has been resolved. Thank you for your patience
Monitoring: We are seeing signs of recovery and monitoring the situation.
Investigating: Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.
Identified: A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.
Identified: Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.
Identified: Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
CircleCI macOS Jobs
CircleCI Windows Jobs
CircleCI Pipelines & Workflows
What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Postmortem: ## Summary
On October 6, 2026, from 15:30 to 20:30 UTC, CircleCI customers experienced delays creating pipelines and starting workflows and jobs. From 15:48 to 16:38 UTC, no workflows or jobs started, and pipelines triggered from 16:02 UTC failed. Shorter, intermittent delays began earlier, at 13:30 UTC. Data in the CircleCI UI, notifications, and commit status updates to version control providers were also delayed.
The incident started in the database behind our workflow orchestration service, which coordinates every pipeline, workflow and job on CircleCI. Routine maintenance on the database's largest tables wrote transaction logs faster than the database could archive them, and the logs filled the storage volume set aside for them. To keep the database from becoming read-only, our engineers moved the logs to the database's main storage volume. That move took the database offline for 50 minutes. When the database came back, it processed work at about half its normal rate until our engineers changed a database setting at 18:19 UTC. We then worked through the backlog, and job start times returned to normal by 20:30 UTC.
Customers whose jobs failed during this window may rerun them.
The original status page can be found here.
Background
When you trigger a pipeline, CircleCI hands it to a workflow orchestration service. That service tracks the state of every workflow and job, decides when each job is ready to run, and passes ready jobs to our execution fleet. It also drives the job data you see in the UI and the notifications and commit statuses we send when work finishes.
The orchestration service stores its state in a database. The database writes every change to a transaction log before applying it, and it keeps those logs on a dedicated storage volume. It archives the logs continuously so it can reuse that space. If archiving falls behind and the volume fills, the database stops accepting writes, and no new work can move forward.
The database also runs a periodic maintenance process (autovacuums) on each table to keep its internal records valid over time. On a very large table, this process reads every page of the table and writes a large volume of transaction logs.
What Happened
(All times UTC)
From 10:02 on October 6, our monitoring raised short-lived alerts about errors in the orchestration service. Each alert cleared on its own within about 15 minutes. Our engineers had seen similar brief alerts before, and the system showed no other signs of stress, so they followed the runbook each time, and each alert cleared on its own partway through and also started work to make the alerts less sensitive.
Around 12:40, along with the regular workload which generates its own transaction logs, the maintenance process on several of the database's largest tables began writing additional transaction logs faster than the database could archive them, and the log volume started to fill.
At 13:30, the orchestration service began to slow down. Through 15:30, some pipelines and jobs took longer to start, in short bursts. At 13:35, our monitoring alerted again, this time alongside related alerts from several other services. Infrastructure Engineers followed the alert runbooks, clearing up some of the alerts. The alerts returned again at around 14:30 and after investigating further, we declared an incident at 14:54. We should have recognized the combination of alerts sooner, and we have covered how we are fixing that below.
After declaring the incident, our engineers traced the slowdown to the maintenance process (anti-wraparound autovacuum). The database restarts this process automatically if it is cancelled, so our engineers changed its settings to help it finish faster. By 15:30, the delays were continuous, and some jobs began to fail.
At 15:43, the transaction log volume was close to full. If it filled, the database would stop accepting writes and no work could run. Our engineers decided to move the transaction logs onto the database's main storage volume, which had plenty of free space. The move required the database to go offline, and we could not predict in advance how long that would take.
The move started at 15:47. From 15:48 to 16:38, the orchestration service could not process any work. Customers could still trigger pipelines, but no workflows or jobs started, and from 16:02 newly triggered pipelines failed. We posted to our status page at 15:58 and raised it to a major outage at 16:20. While the database was offline, our engineers prepared a replacement database as a fallback. The original database came back first, so we kept it in service.
At 16:38, the database came back online and work started flowing again. With transaction logs and regular data now sharing the same storage, the database ran more slowly than before. From 16:38 to 18:19, the orchestration service processed work at about half its normal rate, and jobs waited tens of minutes to start, up to about 50 minutes at the longest. During that time, the database had to finish archiving its backlog of transaction logs before we could move them back to a dedicated volume. We also investigated and implemented several database parameter changes to stop further maintenance runs from starting and prepared ways to reduce the work reaching the service. Along with the replacement database, we also prepared a complete stack of the service during that time so that we could move to it at the risk of data loss and opted against that move.
At 18:19, our engineers changed a database setting to cut the time each write spent waiting on storage. Processing speed recovered, and the backlog began to clear. To clear it faster, we added database and server capacity to the systems that hand jobs to our execution fleet. By about 19:55, new jobs were again starting on time.
The backlog then reached our execution fleet. From 19:49 to 20:30, some Docker jobs on larger resource classes waited up to about 10 minutes to start while capacity scaled up. By 20:30, job start times were back to normal. We moved the status page to monitoring at 20:32 and resolved the incident at 21:19.
About 3,500 jobs failed to start during the incident. Customers may rerun these jobs. We found no evidence that CircleCI ran any job more than once. Customers who reran a workflow while the original run was still delayed may have seen both runs complete. Annual plan customers can work with their account team to review usage.
Future Prevention and Process Improvement
We are taking the following steps to prevent a recurrence and improve our response time:
We are adding alerts on the conditions that caused this incident. Our monitoring caught the slowdown, but we had no alert on how full the transaction log volume was or on how far archiving had fallen behind. We are adding alerts on both, along with alerts on the database maintenance that drives them, so we can act earlier. We are also revisiting all our alerts to ensure noisy or sensitive alerts don’t hide real issues.
We are tuning how maintenance process (autovacuum) runs on our database. We are looking at the frequency and the aggressiveness with which the autovacuums run on our database. We are also tracking the size of these tables over time so they stay within safe limits.
We are removing the single transaction log bottleneck. All of the orchestration service's writes currently go through one database and one transaction log, so when that log fell behind, every pipeline, workflow and job slowed down with it. We are evaluating two approaches: splitting the service across multiple database instances, each with its own transaction log, and moving it to a distributed database built to spread writes across many nodes. Either approach would limit the effect of a backlog to part of the workload.
We are making our systems wait for the orchestration service to recover. During the incident, some jobs failed because our systems gave up after retrying. We are working on changing that behavior.
We are improving our status page updates during long incidents. Customers told us our early updates did not give enough detail. We are updating our guidance so updates include specific impact and timing sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve this incident. Please reach out to our support team with any questions or concerns.
Resolved: Between approximately 15:30 UTC and 20:30 UTC on October 6, 2026, customers experienced delays and failures affecting pipelines, workflows, and jobs.
From approximately 16:02 UTC to 16:39 UTC, pipelines were failing and no new workflows or jobs could start. After service resumed, workflows and jobs continued to start with delays until approximately 20:00 UTC. Start times averaged up to approximately 40 minutes at their peak, and a very small number of workflows waited more than 1 hour. Outbound notifications and webhooks were also delayed until approximately 19:00 UTC, typically by about a minute.
As the backlog of delayed work cleared, the resulting increase in demand placed pressure on our execution capacity, causing further delays for Docker and Linux jobs. From approximately 19:45 UTC to 20:30 UTC, these jobs waited in the queue longer than usual before starting. At the peak, around 20:05 UTC, waits reached up to approximately 8 minutes for 2xlarge gen2, 6 minutes for xlarge gen2, and less than 5 minutes for a few other resource classes.
The issue has been resolved, and all affected functionality has returned to normal. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Workflows and jobs are starting normally, and wait times for Docker and Linux jobs have returned to normal. Our engineers are continuing to monitor this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: The backlog of delayed work has been cleared, and workflows are starting again with no delays in notifications. However, customers using Docker and Linux jobs across multiple resource classes, including medium, large, and 2xlarge gen2, may now see jobs waiting in the queue for few minutes before they start. Our engineers are continuing to work on mitigating this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Delays for workflows and jobs to start are continuing to decrease. Most are now starting within approximately 10 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour. If your job failed to start, please rerun the job. Delays in outbound notifications have improved.
The remaining delay is due to a backlog of work that built up during the incident, and that backlog has reduced by more than half. We expect the backlog to be cleared in approximately 15 minutes. We will share another update once it has cleared, or within the next 30 minutes at the latest.
Identified: Customers continue to experience delays. Most are now starting within approximately 20 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour.
Our engineers have made changes that have improved service stability, and the remaining delay is due to a backlog of work that built up during the incident. The backlog is decreasing, and we are focused on clearing it faster. We will share more information as progress continues.
We thank you for your patience while our engineers work through the backlog. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays before workflows and jobs start, in some cases up to approximately 30 minutes. This is improved from earlier in the incident, but delays remain. Delays in outbound notifications have improved.
We thank you for your patience while our engineers continue to work toward restoring normal job start times. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays between jobs in a workflow, as well as delays in outbound notifications.
We thank you for your patience while our engineers continue to investigate and work on mitigating the issue. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: While Pipelines, workflows, and jobs are running again, customers may continue to see longer than usual delays between jobs in a workflow, and outbound notifications are delayed by approximately 1 minute. Notifications and status updates to your version control system are beginning to go out again.
Our engineers are continuing to work toward a full recovery.
Identified: As of 16:39 UTC, we are starting to see recovery for new pipelines and workflows being submitted, and jobs are beginning to run again. Our engineers are continuing to work on the recovery and will share more information as soon as we have it.
Identified: As of 16:02 UTC, pipelines are failing entirely, so no new workflows or jobs are running at this point. This is in addition to the delays and failed jobs reported earlier. Our engineers are continuing to work on mitigating the issue, and we will share more information as soon as we have it.
Investigating: Since 15:47 UTC, customers may also see failed jobs, in addition to the delays previously reported. Our engineers are continuing to investigate and mitigate the issue.
Investigating: What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
We are investigating a rise in infra fails for customer jobs.
Postmortem: ## Summary
Between October 1 and October 8, 2026, CircleCI customers experienced five separate incidents in which jobs waited longer than normal to start, or failed to start. Three of the incidents were driven by limited instance availability from our cloud provider. One was caused by a configuration issue in our scheduling infrastructure, and one by memory limits on an internal service, which we raised during the incident.
- October 1, 07:57 to 10:24 UTC: Docker Gen 2 jobs waited up to 30 minutes to start. A configuration issue prevented the Gen 2 scheduling servers from using their high-performance local disks. During a burst in traffic, the fallback disks saturated and job placement slowed.
- October 7, 13:27 to 14:06 UTC: Machine jobs waited longer than normal because we could not acquire enough virtual machines from our cloud provider.
- October 7, 14:24 to 15:21 UTC: Some jobs failed with an infrastructure error, and workflows were delayed. The service that gives each job its configuration at startup ran out of memory and restarted.
- October 7, 18:49 to 20:05 UTC: Docker Gen 2 jobs waited longer than normal because we could not acquire enough instances from our cloud provider.
- October 8, 12:48 to 16:00 UTC: Linux machine jobs and remote Docker jobs waited up to 50 minutes to start. Our cloud provider did not have enough of the instance type we request, and a defect in the logic that picks a different instance type kept us from using alternatives.
Jobs that waited during these incidents were delayed, not lost. Customers whose jobs failed with an infrastructure error during the October 7 incident from 14:24 to 15:21 UTC can rerun them.
The original status pages can be found below:
- October 1 - Delays starting Gen 2 Docker Jobs
- October 7 - Delay on starting Machine Job Tasks
- October 7 - Elevated level of infra fails on customer jobs
- October 7 - Increased task wait times for Docker Gen2
- October 8 - Elevated wait times for machine jobs
Background
Every job that runs on CircleCI needs compute. Where that compute comes from depends on the executor.
Docker jobs run on clusters of servers that schedule work onto a fleet of instances. The scheduler is the software on those servers that decides which instance runs each job. Gen 2 Docker jobs ran on one such cluster at the time of the October 1 incident. Gen 1 Docker jobs are spread across several clusters.
Machine jobs, including remote Docker jobs, get a dedicated virtual machine. A remote Docker job is handled as a Linux machine job from submission through provisioning. Our machine provisioning service requests each machine from our cloud provider in one of the two regions we normally use for machine jobs. When the provider cannot supply the instance type we request, the service falls back: it requests a different instance type, or a different region, from a list we maintain. A job waits until a machine is available.
When any job starts, it asks a job configuration service for its configuration. That service runs as a set of pods, which are small copies of the same program. If too many pods stop, the remaining pods carry more load, and jobs cannot start until the service recovers.
These paths depend on three kinds of capacity: the throughput of our scheduling systems, the supply of instances from our cloud provider, and the memory and number of copies of the services that start jobs. A shortfall in any of them makes jobs wait or fail.
What Happened
(All times UTC)
October 1: Docker Gen 2 jobs delayed up to 30 minutes
At 07:44, disk activity on the servers that schedule Gen 2 Docker jobs rose to 90% of capacity and stayed there. At 07:57, the number of Docker jobs waiting to start began to grow, and wait times rose with it. An automated alert paged the on-call engineer at 08:08. We declared an incident at 08:42.
At 08:43, engineers identified high disk latency on the scheduling servers. They first believed the disks were too slow for the load, so they prepared a change to a faster disk type. At 09:01, the team agreed on two workstreams: move the servers to faster disks, and add a second Gen 2 cluster so that jobs would split across two clusters.
At 09:41, engineers found the root cause. Each scheduling server has fast disks attached directly to it, and the scheduler is meant to keep its working data there. These servers run on a security-hardened operating system image, and on that image the fast disks were not being used. The scheduler wrote its data to slower network-attached disks instead. The script that sets up the fast disks failed without an error, and a gap in logging on these servers hid the failure. Gen 2 traffic had also become increasingly bursty. The slower disks saturated during those bursts, and the scheduler fell behind.
Wait times peaked at 09:54, when jobs on the Gen 2 medium resource class waited up to 30 minutes. At 09:55, engineers applied a fix that set up the fast disks correctly. At 10:11, the scheduler began placing jobs faster, which confirmed the servers could keep up with new jobs. Wait times for Gen 2 medium fell to 13 minutes at 10:14, 7 minutes at 10:17, and 1 minute 30 seconds at 10:20. We moved the status page to monitoring at 10:24 and resolved the incident at 10:45.
At 11:20, the second Gen 2 cluster began running jobs.
October 7, 13:27 to 14:06: Elevated wait times for machine jobs
We declared an incident at 13:27 after seeing elevated queueing and wait times. We posted to the status page at 13:31. At 13:31, engineers saw that we were struggling to acquire virtual machines for machine jobs, and that the number of jobs waiting to start was high mainly for machine jobs. At 13:34, wait times were elevated across executors.
At 13:48, wait times began to recover as more machines became available, and we resolved the incident at 14:06.
October 7, 14:24 to 15:21: Job failures and workflow delays
Shortly after the previous incident was marked as resolved, we declared a new incident at 14:24 due to a sharp drop in the number of jobs submitted to our execution systems. At 14:26, engineers saw that pods of the job configuration service were restarting. Pods that restart leave fewer copies of the service to answer requests, and jobs that could not retrieve their configuration after repeated attempts failed with an infrastructure error. Workflows that were already running waited for updates from jobs that had not started.
At 14:35, engineers increased the amount of available memory to the pods and increased the number of pods, which gave the service more copies to share the load. We also increased the amount of memory available Job submissions began recovering, and the number of messages flowing through our workflow system returned to normal levels. At 14:44, we moved the status page to monitoring.
Pods in the job configuration service reached their memory limits and restarted. We increased memory allocations and pod counts, and the service recovered. At 15:13, we increased the number of instances in our Gen 2 Docker clusters so we could work through the backlog of waiting jobs quickly. We resolved the incident at 15:21.
October 7, 18:49 to 20:05: Elevated wait times for Docker Gen 2 jobs
We declared a new incident at 18:49 after Docker Gen 2 wait times rose. At 18:51, engineers confirmed that a shortage of available instances kept us from acquiring enough of them, which raised wait times. Wait times differed by resource class size.
At 19:52, average wait times were lower than an hour earlier as more capacity became available, though brief spikes continued, and we moved the status page to monitoring. We resolved the incident at 20:05.
October 8: Linux machine and remote Docker jobs wait up to 50 minutes
We declared an incident at 12:48 after seeing a large number of jobs waiting to start and long wait times for machine and remote Docker jobs. At 12:49, engineers saw that our cloud provider could not supply enough instances. Delays continued to climb as we investigated.
At 13:13, we began examining the logic that chooses a different instance type or region when our first choice is unavailable. At 13:24, we suspected that a defect in that logic was stopping us from using other instance types in our primary region, so we disabled our fallback region and confirmed the cause of the defect. Engineers began working on a fix and deployed it at 14:49. The service began requesting other instance types and started machines in large numbers. At 14:58, the queue of jobs waiting for a machine dropped steadily. Average wait times across resource classes continued to drop over the next hour as we worked through the backlog of waiting jobs.
At 16:00, the queue of jobs waiting for a machine cleared and the number of waiting jobs returned to normal. We moved the status page to monitoring at 16:09 and resolved the incident at 16:35.
Future Prevention and Process Improvement
We are taking the following steps to strengthen the resilience of job start times, including supply diversification, fallback logic, and earlier detection.
We are expanding our compute regions. In the October 7 and October 8 incidents, our cloud provider could not supply enough instances in the regions we used. We are adding new regions to our fallback logic in order handle spikes for demand that outpace our preferred regions.
We are fixing the logic that picks a different instance type. On October 8, a defect kept the machine provisioning service from using other instance types in the same region. We deployed a fix during the incident and are now reviewing and refactoring this logic as a whole, with tests, so that a shortage of one instance type moves work to another.
We are expanding the instance types we offer and securing additional capacity. We are working with our cloud provider to reserve additional capacity in our existing regions, and we are adding instance types that jobs can run on. Industry demand for high-performance compute is growing. We are diversifying where and how we source capacity so our customers are insulated from supply variation.
We added a second Gen 2 Docker cluster so traffic bursts are distributed across two clusters. We are also adding alerts on scheduling server disk saturation and on how fast the scheduler places jobs, and an alert that fires when a server starts without its fast disks in use. We made the disk setup script fail with an error, and we restored log shipping from these servers so that a similar problem is visible sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents. Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're seeing signs of recovery. Docker Gen2 wait times are slightly elevated and we're monitoring.
Investigating: We are investigating a rise in infra fails for customer jobs.
CircleCI API
There is a delay in our data infrastructure causing delays before build data and some notifications become visible. We are working on a mitigation.
Resolved: Between 11:15 UTC and 23:47 UTC on October 9, customers experienced delays and failures across pipelines, workflows and jobs, along with delays in the CircleCI UI, API and notifications.
From 11:15 to about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and Slack and email notifications were not sent. From about 13:00 to 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
From about 15:15 to 20:48 UTC, workflows and jobs started with increasing delays, of up to about an hour, and the CircleCI UI and API fell behind by up to about 2 hours. During this time, some jobs failed to start and some in-progress workflows failed.
From 20:48 to about 22:10 UTC, we temporarily paused starting new workflows and jobs so our systems could recover. New pipelines were accepted and queued during this time. From about 22:10 to 23:47 UTC, queued workflows and jobs resumed and were processed. Some of them had waited up to about 2 hours.
The issue has been resolved and all affected functionality has returned to normal. Usage charge data on the Project Usage page is not yet up to date and will be updated once our data backfills complete. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our engineers worked on implementing a fix.
Monitoring: Workflows and jobs are starting and running normally, and the CircleCI UI and API are up to date. We are continuing to monitor our systems closely to make sure they remain stable.
We will provide another update by 02:00 UTC.
Identified: Workflows and jobs have continued to start normally, within seconds, and the CircleCI UI and API remain up to date. Slack and email notifications delayed earlier in the incident are still being delivered.
Usage charge data on the Project Usage page is stale.
We are continuing to monitor our systems closely and will provide another update by 01:15 UTC.
Identified: Since our last update, the backlog of queued workflows and jobs has cleared, as of about 23:47 UTC. Workflows and jobs are now starting within seconds, and the CircleCI UI and API are up to date. We are monitoring our systems closely to make sure they remain stable. We will provide another update by 00:30 UTC.
Identified: We are still working through a significant backlog of queued workflows and jobs. Wait times are improving but they remain longer than normal and some jobs have waited over an hour. The CircleCI UI and API remain up to date. Please avoid retriggering pipelines or workflows while the backlog clears, as this may create duplicates. We will provide another update by 00:00 UTC.
Identified: A large number of queued workflows and jobs have started, and the CircleCI UI and API have caught up and are showing updates normally with few delays. We are continuing to work through the remaining queued work. Jobs from the queue may have waited about an hour or more to start, and some Linux jobs are waiting up to about 10 minutes for capacity.
Please continue to avoid retriggering while queued work continues to process.
We will provide another update by 23:30 UTC.
Identified: Since our last update, we have begun gradually resuming work, and some queued workflows and jobs are starting again. Most queued work is still waiting, and jobs that do start may have waited over an hour. We are increasing the rate carefully to avoid overloading our systems. Delayed updates continue to appear in the CircleCI UI and API.
Please continue to avoid retriggering pipelines or workflows, as this may create duplicates as queued work resumes.
We will provide another update by 23:00 UTC.
Identified: Since our last update, most of the pipelines that were delayed earlier in the incident have now been processed, and their workflows are queued to start. New workflows and jobs remain paused while we prepare to resume work safely. Delayed updates continue to appear in the CircleCI UI and API.
Please continue to avoid retriggering pipelines or workflows, as this may create duplicates once work resumes.
We will provide another update by 22:30 UTC.
Identified: Since about 20:48 UTC, new workflows and jobs are not starting. We have temporarily paused starting new work so that our systems can catch up on work already in progress, including delayed updates to the CircleCI UI and API. Older updates are now appearing in the UI and API as this backlog clears. New pipelines are still being accepted and queued
Please avoid retriggering pipelines or workflows during this time, as this may create duplicates once work resumes.
We will provide another update by 22:00 UTC.
Identified: Jobs are starting faster and fewer are failing: most jobs are now starting within about 4 minutes, down from about 6 minutes, and the longest waits have dropped from about 30 minutes to about 20 minutes.
Pipelines from earlier in the incident are now being worked through.
The CircleCI UI is still behind, and some older updates may take longer to appear. Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI.
We will provide another update by 20:30 UTC.
Identified: Since our last update, jobs are starting faster: most are now starting within about 6 minutes, down from about 8 minutes. The CircleCI UI is catching up but still about an hour behind. New pipelines are being processed as they arrive.
Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI.
We will provide another update by 20:00 UTC.
Identified: What's changed since our last update
- Jobs are starting faster. Most are now starting within about 8 minutes, down from about 15 minutes, and the longest waits have dropped from over an hour to about 45 minutes.
- The CircleCI UI is beginning to catch up, but it is still behind. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all.
- Pipelines are being processed faster, but new pipelines are still delayed.
Still ongoing
- Some workflows and jobs are still failing.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is stale.
Next update We will provide another update by 19:30 UTC.
Identified: What's changed since our last update
- The CircleCI UI is about an hour behind, and up to about 90 minutes for some updates. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all.
- Jobs that are running are starting faster. Most are now starting within about 15 minutes, down from about 28 minutes, though some are still waiting over an hour.
- New pipelines are increasingly delayed.
- Some workflows are still failing. Workflows that could not be processed in time are being marked as failed.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
Still ongoing: Usage charge data on the Project Usage page is stale.
Next update We will provide another update by 19:00 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What to expect
- Since about 17:00 UTC, we are seeing an increase in delays. Most jobs are now taking about 25 minutes to start, and some are waiting up to about an hour.
- Since about 17:15 UTC, some in-progress workflows have failed and a small number of jobs are failing to start.
- New pipelines are being queued and are taking longer to start.
- The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Our engineers continue to work to resolve the issue at hand.
Next update We will provide another update by 18:30 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Since about 17:00 UTC, delays have increased significantly. Most jobs are now taking about 12 minutes to start, and some are waiting up to about 30 minutes.
- Since about 17:15 UTC, some jobs are failing to start.
- Pipelines are being processed more slowly, so some may take longer to start.
- The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 18:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs are holding steady. Most jobs are starting within about 4 minutes, and some are waiting up to about 8 minutes, down from about 25 minutes.
- The CircleCI UI remains about 20 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 17:30 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs have stopped increasing and are beginning to ease. Most jobs are starting within about 4 minutes, and some are waiting up to about 25 minutes.
- The CircleCI UI is now about 20 minutes behind, down from about 30 minutes. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 17:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Since about 15:15 UTC, workflows and jobs have been starting with increasing delays. Most jobs are currently starting within about 4 minutes, and some are waiting up to about 20 minutes.
- The CircleCI UI remains about 30 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 16:30 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What's impacted Customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Workflows and jobs are running on time.
- The CircleCI UI is about 30 minutes behind. New workflows and status changes may take that long to appear.
- Our engineers are actively working to speed up processing so the UI can catch up.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce this delay.
Next update We will provide another update by 16:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- The backlog of pipelines and workflows from the incident has cleared.
- Since about 14:45 UTC, most workflows and jobs are starting within about 15 seconds, and start times are continuing to return to normal.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers monitor the recovery.
Next update We will provide another update by 15:20 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- The backlog of workflows from pipelines triggered during the incident cleared at about 14:40 UTC.
- Most jobs are now starting within about 30 seconds, and some are waiting up to about 3 minutes.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers continue to work on the underlying issue.
Next update We will provide another update by 15:10 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs are improving. Most jobs are now starting within about 3 minutes, down from about 10 minutes at 13:50 UTC.
- Some workflows from pipelines triggered during the incident are still waiting up to about 40 minutes to start.
- Working through the backlog is taking slightly longer than expected. We now expect to clear it by about 14:45 UTC.
- Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 15:00 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- Pipelines are being processed again, including those triggered during the incident.
- Since about 13:25 UTC, workflows and jobs have been starting with increasing delays as we work through the backlog. Most jobs are currently starting within about 10 minutes, and some are waiting up to about 25 minutes.
- We expect to work through this backlog by about 14:35 UTC.
- Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 14:30 UTC.
Identified: Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications.
What can you expect:
- Since about 13:00 UTC, pipelines are being processed again and jobs are starting.
- Most jobs are currently starting within about 4 minutes, and some are waiting up to about 10 minutes. Pipelines triggered during the incident are still being processed and may take longer to start.
- Outbound webhooks and notifications have resumed.
- Pipelines are still being processed from the backlog, so retriggering may result in duplicate pipelines.
Thank you for your patience while our engineers work to resolve this.
Next update: We will provide another update by 14:00 UTC.
Identified: What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications.
What can you expect Since 11:15 UTC:
- Pipelines are still being created while this is ongoing but not yet processed, so retriggering may result in duplicate pipelines.
- Outbound webhooks and transactional notifications (Slack and email) are not being sent.
- Pipeline statuses may not be visible.
- Usage charge data on the Project Usage page will be stale.
Thank you for your patience while our engineers work to resolve this.
Next update: We will provide another update by 13:30 UTC.
Identified: There is a delay in our data infrastructure causing delays before build data and some notifications become visible. We are working on a mitigation.
CircleCI UI
There is a delay in our data infrastructure causing delays before build data and some notifications become visible. We are working on a mitigation.
Resolved: Between 11:15 UTC and 23:47 UTC on October 9, customers experienced delays and failures across pipelines, workflows and jobs, along with delays in the CircleCI UI, API and notifications.
From 11:15 to about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and Slack and email notifications were not sent. From about 13:00 to 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
From about 15:15 to 20:48 UTC, workflows and jobs started with increasing delays, of up to about an hour, and the CircleCI UI and API fell behind by up to about 2 hours. During this time, some jobs failed to start and some in-progress workflows failed.
From 20:48 to about 22:10 UTC, we temporarily paused starting new workflows and jobs so our systems could recover. New pipelines were accepted and queued during this time. From about 22:10 to 23:47 UTC, queued workflows and jobs resumed and were processed. Some of them had waited up to about 2 hours.
The issue has been resolved and all affected functionality has returned to normal. Usage charge data on the Project Usage page is not yet up to date and will be updated once our data backfills complete. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our engineers worked on implementing a fix.
Monitoring: Workflows and jobs are starting and running normally, and the CircleCI UI and API are up to date. We are continuing to monitor our systems closely to make sure they remain stable.
We will provide another update by 02:00 UTC.
Identified: Workflows and jobs have continued to start normally, within seconds, and the CircleCI UI and API remain up to date. Slack and email notifications delayed earlier in the incident are still being delivered.
Usage charge data on the Project Usage page is stale.
We are continuing to monitor our systems closely and will provide another update by 01:15 UTC.
Identified: Since our last update, the backlog of queued workflows and jobs has cleared, as of about 23:47 UTC. Workflows and jobs are now starting within seconds, and the CircleCI UI and API are up to date. We are monitoring our systems closely to make sure they remain stable. We will provide another update by 00:30 UTC.
Identified: We are still working through a significant backlog of queued workflows and jobs. Wait times are improving but they remain longer than normal and some jobs have waited over an hour. The CircleCI UI and API remain up to date. Please avoid retriggering pipelines or workflows while the backlog clears, as this may create duplicates. We will provide another update by 00:00 UTC.
Identified: A large number of queued workflows and jobs have started, and the CircleCI UI and API have caught up and are showing updates normally with few delays. We are continuing to work through the remaining queued work. Jobs from the queue may have waited about an hour or more to start, and some Linux jobs are waiting up to about 10 minutes for capacity.
Please continue to avoid retriggering while queued work continues to process.
We will provide another update by 23:30 UTC.
Identified: Since our last update, we have begun gradually resuming work, and some queued workflows and jobs are starting again. Most queued work is still waiting, and jobs that do start may have waited over an hour. We are increasing the rate carefully to avoid overloading our systems. Delayed updates continue to appear in the CircleCI UI and API.
Please continue to avoid retriggering pipelines or workflows, as this may create duplicates as queued work resumes.
We will provide another update by 23:00 UTC.
Identified: Since our last update, most of the pipelines that were delayed earlier in the incident have now been processed, and their workflows are queued to start. New workflows and jobs remain paused while we prepare to resume work safely. Delayed updates continue to appear in the CircleCI UI and API.
Please continue to avoid retriggering pipelines or workflows, as this may create duplicates once work resumes.
We will provide another update by 22:30 UTC.
Identified: Since about 20:48 UTC, new workflows and jobs are not starting. We have temporarily paused starting new work so that our systems can catch up on work already in progress, including delayed updates to the CircleCI UI and API. Older updates are now appearing in the UI and API as this backlog clears. New pipelines are still being accepted and queued
Please avoid retriggering pipelines or workflows during this time, as this may create duplicates once work resumes.
We will provide another update by 22:00 UTC.
Identified: Jobs are starting faster and fewer are failing: most jobs are now starting within about 4 minutes, down from about 6 minutes, and the longest waits have dropped from about 30 minutes to about 20 minutes.
Pipelines from earlier in the incident are now being worked through.
The CircleCI UI is still behind, and some older updates may take longer to appear. Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI.
We will provide another update by 20:30 UTC.
Identified: Since our last update, jobs are starting faster: most are now starting within about 6 minutes, down from about 8 minutes. The CircleCI UI is catching up but still about an hour behind. New pipelines are being processed as they arrive.
Jobs may already have run even if the UI does not show them yet, so please avoid retriggering workflows that don't appear in the UI.
We will provide another update by 20:00 UTC.
Identified: What's changed since our last update
- Jobs are starting faster. Most are now starting within about 8 minutes, down from about 15 minutes, and the longest waits have dropped from over an hour to about 45 minutes.
- The CircleCI UI is beginning to catch up, but it is still behind. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all.
- Pipelines are being processed faster, but new pipelines are still delayed.
Still ongoing
- Some workflows and jobs are still failing.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is stale.
Next update We will provide another update by 19:30 UTC.
Identified: What's changed since our last update
- The CircleCI UI is about an hour behind, and up to about 90 minutes for some updates. Workflows and jobs may already have run, or may be running, even if the UI shows them as not started or does not show them at all.
- Jobs that are running are starting faster. Most are now starting within about 15 minutes, down from about 28 minutes, though some are still waiting over an hour.
- New pipelines are increasingly delayed.
- Some workflows are still failing. Workflows that could not be processed in time are being marked as failed.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
Still ongoing: Usage charge data on the Project Usage page is stale.
Next update We will provide another update by 19:00 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What to expect
- Since about 17:00 UTC, we are seeing an increase in delays. Most jobs are now taking about 25 minutes to start, and some are waiting up to about an hour.
- Since about 17:15 UTC, some in-progress workflows have failed and a small number of jobs are failing to start.
- New pipelines are being queued and are taking longer to start.
- The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Our engineers continue to work to resolve the issue at hand.
Next update We will provide another update by 18:30 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers triggering pipelines and running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Since about 17:00 UTC, delays have increased significantly. Most jobs are now taking about 12 minutes to start, and some are waiting up to about 30 minutes.
- Since about 17:15 UTC, some jobs are failing to start.
- Pipelines are being processed more slowly, so some may take longer to start.
- The CircleCI UI is more than an hour behind. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 18:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What’s impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs are holding steady. Most jobs are starting within about 4 minutes, and some are waiting up to about 8 minutes, down from about 25 minutes.
- The CircleCI UI remains about 20 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don’t yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 17:30 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog. Since about 15:15 UTC, workflows and jobs have again been starting with delays.
What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs have stopped increasing and are beginning to ease. Most jobs are starting within about 4 minutes, and some are waiting up to about 25 minutes.
- The CircleCI UI is now about 20 minutes behind, down from about 30 minutes. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 17:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What's impacted Customers running workflows and jobs, customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Since about 15:15 UTC, workflows and jobs have been starting with increasing delays. Most jobs are currently starting within about 4 minutes, and some are waiting up to about 20 minutes.
- The CircleCI UI remains about 30 minutes behind on average, and some updates are taking longer. New workflows and status changes may take that long to appear.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce these delays.
Next update We will provide another update by 16:30 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What's impacted Customers viewing pipelines, workflows and jobs in the CircleCI UI, and customers viewing usage data on the Project Usage page.
What can you expect
- Workflows and jobs are running on time.
- The CircleCI UI is about 30 minutes behind. New workflows and status changes may take that long to appear.
- Our engineers are actively working to speed up processing so the UI can catch up.
- Please avoid retriggering workflows that don't yet appear in the UI, as this may create duplicates.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to reduce this delay.
Next update We will provide another update by 16:00 UTC.
Monitoring: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- The backlog of pipelines and workflows from the incident has cleared.
- Since about 14:45 UTC, most workflows and jobs are starting within about 15 seconds, and start times are continuing to return to normal.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers monitor the recovery.
Next update We will provide another update by 15:20 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent. Between about 13:00 UTC and 14:40 UTC, workflows and jobs started with delays while we worked through the backlog.
What’s impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- The backlog of workflows from pipelines triggered during the incident cleared at about 14:40 UTC.
- Most jobs are now starting within about 30 seconds, and some are waiting up to about 3 minutes.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers continue to work on the underlying issue.
Next update We will provide another update by 15:10 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- Delays starting workflows and jobs are improving. Most jobs are now starting within about 3 minutes, down from about 10 minutes at 13:50 UTC.
- Some workflows from pipelines triggered during the incident are still waiting up to about 40 minutes to start.
- Working through the backlog is taking slightly longer than expected. We now expect to clear it by about 14:45 UTC.
- Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 15:00 UTC.
Identified: Since 11:15 UTC, customers have experienced disruption to pipelines and notifications. Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted Customers running workflows and jobs, and customers viewing usage data on the Project Usage page.
What can you expect
- Pipelines are being processed again, including those triggered during the incident.
- Since about 13:25 UTC, workflows and jobs have been starting with increasing delays as we work through the backlog. Most jobs are currently starting within about 10 minutes, and some are waiting up to about 25 minutes.
- We expect to work through this backlog by about 14:35 UTC.
- Workflows from pipelines triggered during the incident are still waiting to start, so retriggering may result in duplicate workflows.
- Usage charge data on the Project Usage page is currently stale.
Thank you for your patience while our engineers work to resolve this.
Next update We will provide another update by 14:30 UTC.
Identified: Between 11:15 UTC and about 13:00 UTC, pipelines were created but not processed, and outbound webhooks and transactional notifications (Slack and email) were not sent.
What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications.
What can you expect:
- Since about 13:00 UTC, pipelines are being processed again and jobs are starting.
- Most jobs are currently starting within about 4 minutes, and some are waiting up to about 10 minutes. Pipelines triggered during the incident are still being processed and may take longer to start.
- Outbound webhooks and notifications have resumed.
- Pipelines are still being processed from the backlog, so retriggering may result in duplicate pipelines.
Thank you for your patience while our engineers work to resolve this.
Next update: We will provide another update by 14:00 UTC.
Identified: What's impacted: Customers triggering pipelines, and customers relying on outbound webhooks and email and Slack notifications.
What can you expect Since 11:15 UTC:
- Pipelines are still being created while this is ongoing but not yet processed, so retriggering may result in duplicate pipelines.
- Outbound webhooks and transactional notifications (Slack and email) are not being sent.
- Pipeline statuses may not be visible.
- Usage charge data on the Project Usage page will be stale.
Thank you for your patience while our engineers work to resolve this.
Next update: We will provide another update by 13:30 UTC.
Identified: There is a delay in our data infrastructure causing delays before build data and some notifications become visible. We are working on a mitigation.
CircleCI Artifacts
CircleCI Runner
CircleCI Webhooks
CircleCI Insights
CircleCI Releases
CircleCI Notifications & Status Updates
What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Postmortem: ## Summary
On October 6, 2026, from 15:30 to 20:30 UTC, CircleCI customers experienced delays creating pipelines and starting workflows and jobs. From 15:48 to 16:38 UTC, no workflows or jobs started, and pipelines triggered from 16:02 UTC failed. Shorter, intermittent delays began earlier, at 13:30 UTC. Data in the CircleCI UI, notifications, and commit status updates to version control providers were also delayed.
The incident started in the database behind our workflow orchestration service, which coordinates every pipeline, workflow and job on CircleCI. Routine maintenance on the database's largest tables wrote transaction logs faster than the database could archive them, and the logs filled the storage volume set aside for them. To keep the database from becoming read-only, our engineers moved the logs to the database's main storage volume. That move took the database offline for 50 minutes. When the database came back, it processed work at about half its normal rate until our engineers changed a database setting at 18:19 UTC. We then worked through the backlog, and job start times returned to normal by 20:30 UTC.
Customers whose jobs failed during this window may rerun them.
The original status page can be found here.
Background
When you trigger a pipeline, CircleCI hands it to a workflow orchestration service. That service tracks the state of every workflow and job, decides when each job is ready to run, and passes ready jobs to our execution fleet. It also drives the job data you see in the UI and the notifications and commit statuses we send when work finishes.
The orchestration service stores its state in a database. The database writes every change to a transaction log before applying it, and it keeps those logs on a dedicated storage volume. It archives the logs continuously so it can reuse that space. If archiving falls behind and the volume fills, the database stops accepting writes, and no new work can move forward.
The database also runs a periodic maintenance process (autovacuums) on each table to keep its internal records valid over time. On a very large table, this process reads every page of the table and writes a large volume of transaction logs.
What Happened
(All times UTC)
From 10:02 on October 6, our monitoring raised short-lived alerts about errors in the orchestration service. Each alert cleared on its own within about 15 minutes. Our engineers had seen similar brief alerts before, and the system showed no other signs of stress, so they followed the runbook each time, and each alert cleared on its own partway through and also started work to make the alerts less sensitive.
Around 12:40, along with the regular workload which generates its own transaction logs, the maintenance process on several of the database's largest tables began writing additional transaction logs faster than the database could archive them, and the log volume started to fill.
At 13:30, the orchestration service began to slow down. Through 15:30, some pipelines and jobs took longer to start, in short bursts. At 13:35, our monitoring alerted again, this time alongside related alerts from several other services. Infrastructure Engineers followed the alert runbooks, clearing up some of the alerts. The alerts returned again at around 14:30 and after investigating further, we declared an incident at 14:54. We should have recognized the combination of alerts sooner, and we have covered how we are fixing that below.
After declaring the incident, our engineers traced the slowdown to the maintenance process (anti-wraparound autovacuum). The database restarts this process automatically if it is cancelled, so our engineers changed its settings to help it finish faster. By 15:30, the delays were continuous, and some jobs began to fail.
At 15:43, the transaction log volume was close to full. If it filled, the database would stop accepting writes and no work could run. Our engineers decided to move the transaction logs onto the database's main storage volume, which had plenty of free space. The move required the database to go offline, and we could not predict in advance how long that would take.
The move started at 15:47. From 15:48 to 16:38, the orchestration service could not process any work. Customers could still trigger pipelines, but no workflows or jobs started, and from 16:02 newly triggered pipelines failed. We posted to our status page at 15:58 and raised it to a major outage at 16:20. While the database was offline, our engineers prepared a replacement database as a fallback. The original database came back first, so we kept it in service.
At 16:38, the database came back online and work started flowing again. With transaction logs and regular data now sharing the same storage, the database ran more slowly than before. From 16:38 to 18:19, the orchestration service processed work at about half its normal rate, and jobs waited tens of minutes to start, up to about 50 minutes at the longest. During that time, the database had to finish archiving its backlog of transaction logs before we could move them back to a dedicated volume. We also investigated and implemented several database parameter changes to stop further maintenance runs from starting and prepared ways to reduce the work reaching the service. Along with the replacement database, we also prepared a complete stack of the service during that time so that we could move to it at the risk of data loss and opted against that move.
At 18:19, our engineers changed a database setting to cut the time each write spent waiting on storage. Processing speed recovered, and the backlog began to clear. To clear it faster, we added database and server capacity to the systems that hand jobs to our execution fleet. By about 19:55, new jobs were again starting on time.
The backlog then reached our execution fleet. From 19:49 to 20:30, some Docker jobs on larger resource classes waited up to about 10 minutes to start while capacity scaled up. By 20:30, job start times were back to normal. We moved the status page to monitoring at 20:32 and resolved the incident at 21:19.
About 3,500 jobs failed to start during the incident. Customers may rerun these jobs. We found no evidence that CircleCI ran any job more than once. Customers who reran a workflow while the original run was still delayed may have seen both runs complete. Annual plan customers can work with their account team to review usage.
Future Prevention and Process Improvement
We are taking the following steps to prevent a recurrence and improve our response time:
We are adding alerts on the conditions that caused this incident. Our monitoring caught the slowdown, but we had no alert on how full the transaction log volume was or on how far archiving had fallen behind. We are adding alerts on both, along with alerts on the database maintenance that drives them, so we can act earlier. We are also revisiting all our alerts to ensure noisy or sensitive alerts don’t hide real issues.
We are tuning how maintenance process (autovacuum) runs on our database. We are looking at the frequency and the aggressiveness with which the autovacuums run on our database. We are also tracking the size of these tables over time so they stay within safe limits.
We are removing the single transaction log bottleneck. All of the orchestration service's writes currently go through one database and one transaction log, so when that log fell behind, every pipeline, workflow and job slowed down with it. We are evaluating two approaches: splitting the service across multiple database instances, each with its own transaction log, and moving it to a distributed database built to spread writes across many nodes. Either approach would limit the effect of a backlog to part of the workload.
We are making our systems wait for the orchestration service to recover. During the incident, some jobs failed because our systems gave up after retrying. We are working on changing that behavior.
We are improving our status page updates during long incidents. Customers told us our early updates did not give enough detail. We are updating our guidance so updates include specific impact and timing sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve this incident. Please reach out to our support team with any questions or concerns.
Resolved: Between approximately 15:30 UTC and 20:30 UTC on October 6, 2026, customers experienced delays and failures affecting pipelines, workflows, and jobs.
From approximately 16:02 UTC to 16:39 UTC, pipelines were failing and no new workflows or jobs could start. After service resumed, workflows and jobs continued to start with delays until approximately 20:00 UTC. Start times averaged up to approximately 40 minutes at their peak, and a very small number of workflows waited more than 1 hour. Outbound notifications and webhooks were also delayed until approximately 19:00 UTC, typically by about a minute.
As the backlog of delayed work cleared, the resulting increase in demand placed pressure on our execution capacity, causing further delays for Docker and Linux jobs. From approximately 19:45 UTC to 20:30 UTC, these jobs waited in the queue longer than usual before starting. At the peak, around 20:05 UTC, waits reached up to approximately 8 minutes for 2xlarge gen2, 6 minutes for xlarge gen2, and less than 5 minutes for a few other resource classes.
The issue has been resolved, and all affected functionality has returned to normal. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Workflows and jobs are starting normally, and wait times for Docker and Linux jobs have returned to normal. Our engineers are continuing to monitor this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: The backlog of delayed work has been cleared, and workflows are starting again with no delays in notifications. However, customers using Docker and Linux jobs across multiple resource classes, including medium, large, and 2xlarge gen2, may now see jobs waiting in the queue for few minutes before they start. Our engineers are continuing to work on mitigating this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Delays for workflows and jobs to start are continuing to decrease. Most are now starting within approximately 10 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour. If your job failed to start, please rerun the job. Delays in outbound notifications have improved.
The remaining delay is due to a backlog of work that built up during the incident, and that backlog has reduced by more than half. We expect the backlog to be cleared in approximately 15 minutes. We will share another update once it has cleared, or within the next 30 minutes at the latest.
Identified: Customers continue to experience delays. Most are now starting within approximately 20 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour.
Our engineers have made changes that have improved service stability, and the remaining delay is due to a backlog of work that built up during the incident. The backlog is decreasing, and we are focused on clearing it faster. We will share more information as progress continues.
We thank you for your patience while our engineers work through the backlog. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays before workflows and jobs start, in some cases up to approximately 30 minutes. This is improved from earlier in the incident, but delays remain. Delays in outbound notifications have improved.
We thank you for your patience while our engineers continue to work toward restoring normal job start times. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays between jobs in a workflow, as well as delays in outbound notifications.
We thank you for your patience while our engineers continue to investigate and work on mitigating the issue. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: While Pipelines, workflows, and jobs are running again, customers may continue to see longer than usual delays between jobs in a workflow, and outbound notifications are delayed by approximately 1 minute. Notifications and status updates to your version control system are beginning to go out again.
Our engineers are continuing to work toward a full recovery.
Identified: As of 16:39 UTC, we are starting to see recovery for new pipelines and workflows being submitted, and jobs are beginning to run again. Our engineers are continuing to work on the recovery and will share more information as soon as we have it.
Identified: As of 16:02 UTC, pipelines are failing entirely, so no new workflows or jobs are running at this point. This is in addition to the delays and failed jobs reported earlier. Our engineers are continuing to work on mitigating the issue, and we will share more information as soon as we have it.
Investigating: Since 15:47 UTC, customers may also see failed jobs, in addition to the delays previously reported. Our engineers are continuing to investigate and mitigate the issue.
Investigating: What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.