
CircleCI Status
Real-time updates of CircleCI issues and outages
CircleCI status is Operational
CircleCI Docker Jobs
CircleCI Machine Jobs
CircleCI Pipelines & Workflows
CircleCI Notifications & Status Updates
You're checking on CircleCI - but is your own site converting?
Outages aren't the only thing that costs you visitors.
Get a visual audit that shows exactly where your site loses conversions.
Free, takes 2 minutes.
Active Incidents
Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.
Identified: A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.
Identified: Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.
Identified: Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
There is an increased wait time for jobs on Docker Gen2.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're still seeing spikes in task wait times but on average wait times are looking better for all Docker Gen2.
Identified: Note: delays will be more noticeable for people running on Docker xlarge and above.
Identified: There is an increased wait time for jobs on Docker Gen2.
We are investigating a rise in infra fails for customer jobs.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're seeing signs of recovery. Docker Gen2 wait times are slightly elevated and we're monitoring.
Investigating: We are investigating a rise in infra fails for customer jobs.
Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: The incident has been resolved. Thank you for your patience
Monitoring: We are seeing signs of recovery and monitoring the situation.
Investigating: Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Postmortem: ## Summary
On October 6, 2026, from 15:30 to 20:30 UTC, CircleCI customers experienced delays creating pipelines and starting workflows and jobs. From 15:48 to 16:38 UTC, no workflows or jobs started, and pipelines triggered from 16:02 UTC failed. Shorter, intermittent delays began earlier, at 13:30 UTC. Data in the CircleCI UI, notifications, and commit status updates to version control providers were also delayed.
The incident started in the database behind our workflow orchestration service, which coordinates every pipeline, workflow and job on CircleCI. Routine maintenance on the database's largest tables wrote transaction logs faster than the database could archive them, and the logs filled the storage volume set aside for them. To keep the database from becoming read-only, our engineers moved the logs to the database's main storage volume. That move took the database offline for 50 minutes. When the database came back, it processed work at about half its normal rate until our engineers changed a database setting at 18:19 UTC. We then worked through the backlog, and job start times returned to normal by 20:30 UTC.
Customers whose jobs failed during this window may rerun them.
The original status page can be found here.
Background
When you trigger a pipeline, CircleCI hands it to a workflow orchestration service. That service tracks the state of every workflow and job, decides when each job is ready to run, and passes ready jobs to our execution fleet. It also drives the job data you see in the UI and the notifications and commit statuses we send when work finishes.
The orchestration service stores its state in a database. The database writes every change to a transaction log before applying it, and it keeps those logs on a dedicated storage volume. It archives the logs continuously so it can reuse that space. If archiving falls behind and the volume fills, the database stops accepting writes, and no new work can move forward.
The database also runs a periodic maintenance process (autovacuums) on each table to keep its internal records valid over time. On a very large table, this process reads every page of the table and writes a large volume of transaction logs.
What Happened
(All times UTC)
From 10:02 on October 6, our monitoring raised short-lived alerts about errors in the orchestration service. Each alert cleared on its own within about 15 minutes. Our engineers had seen similar brief alerts before, and the system showed no other signs of stress, so they followed the runbook each time, and each alert cleared on its own partway through and also started work to make the alerts less sensitive.
Around 12:40, along with the regular workload which generates its own transaction logs, the maintenance process on several of the database's largest tables began writing additional transaction logs faster than the database could archive them, and the log volume started to fill.
At 13:30, the orchestration service began to slow down. Through 15:30, some pipelines and jobs took longer to start, in short bursts. At 13:35, our monitoring alerted again, this time alongside related alerts from several other services. Infrastructure Engineers followed the alert runbooks, clearing up some of the alerts. The alerts returned again at around 14:30 and after investigating further, we declared an incident at 14:54. We should have recognized the combination of alerts sooner, and we have covered how we are fixing that below.
After declaring the incident, our engineers traced the slowdown to the maintenance process (anti-wraparound autovacuum). The database restarts this process automatically if it is cancelled, so our engineers changed its settings to help it finish faster. By 15:30, the delays were continuous, and some jobs began to fail.
At 15:43, the transaction log volume was close to full. If it filled, the database would stop accepting writes and no work could run. Our engineers decided to move the transaction logs onto the database's main storage volume, which had plenty of free space. The move required the database to go offline, and we could not predict in advance how long that would take.
The move started at 15:47. From 15:48 to 16:38, the orchestration service could not process any work. Customers could still trigger pipelines, but no workflows or jobs started, and from 16:02 newly triggered pipelines failed. We posted to our status page at 15:58 and raised it to a major outage at 16:20. While the database was offline, our engineers prepared a replacement database as a fallback. The original database came back first, so we kept it in service.
At 16:38, the database came back online and work started flowing again. With transaction logs and regular data now sharing the same storage, the database ran more slowly than before. From 16:38 to 18:19, the orchestration service processed work at about half its normal rate, and jobs waited tens of minutes to start, up to about 50 minutes at the longest. During that time, the database had to finish archiving its backlog of transaction logs before we could move them back to a dedicated volume. We also investigated and implemented several database parameter changes to stop further maintenance runs from starting and prepared ways to reduce the work reaching the service. Along with the replacement database, we also prepared a complete stack of the service during that time so that we could move to it at the risk of data loss and opted against that move.
At 18:19, our engineers changed a database setting to cut the time each write spent waiting on storage. Processing speed recovered, and the backlog began to clear. To clear it faster, we added database and server capacity to the systems that hand jobs to our execution fleet. By about 19:55, new jobs were again starting on time.
The backlog then reached our execution fleet. From 19:49 to 20:30, some Docker jobs on larger resource classes waited up to about 10 minutes to start while capacity scaled up. By 20:30, job start times were back to normal. We moved the status page to monitoring at 20:32 and resolved the incident at 21:19.
About 3,500 jobs failed to start during the incident. Customers may rerun these jobs. We found no evidence that CircleCI ran any job more than once. Customers who reran a workflow while the original run was still delayed may have seen both runs complete. Annual plan customers can work with their account team to review usage.
Future Prevention and Process Improvement
We are taking the following steps to prevent a recurrence and improve our response time:
We are adding alerts on the conditions that caused this incident. Our monitoring caught the slowdown, but we had no alert on how full the transaction log volume was or on how far archiving had fallen behind. We are adding alerts on both, along with alerts on the database maintenance that drives them, so we can act earlier. We are also revisiting all our alerts to ensure noisy or sensitive alerts don’t hide real issues.
We are tuning how maintenance process (autovacuum) runs on our database. We are looking at the frequency and the aggressiveness with which the autovacuums run on our database. We are also tracking the size of these tables over time so they stay within safe limits.
We are removing the single transaction log bottleneck. All of the orchestration service's writes currently go through one database and one transaction log, so when that log fell behind, every pipeline, workflow and job slowed down with it. We are evaluating two approaches: splitting the service across multiple database instances, each with its own transaction log, and moving it to a distributed database built to spread writes across many nodes. Either approach would limit the effect of a backlog to part of the workload.
We are making our systems wait for the orchestration service to recover. During the incident, some jobs failed because our systems gave up after retrying. We are working on changing that behavior.
We are improving our status page updates during long incidents. Customers told us our early updates did not give enough detail. We are updating our guidance so updates include specific impact and timing sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve this incident. Please reach out to our support team with any questions or concerns.
Resolved: Between approximately 15:30 UTC and 20:30 UTC on October 6, 2026, customers experienced delays and failures affecting pipelines, workflows, and jobs.
From approximately 16:02 UTC to 16:39 UTC, pipelines were failing and no new workflows or jobs could start. After service resumed, workflows and jobs continued to start with delays until approximately 20:00 UTC. Start times averaged up to approximately 40 minutes at their peak, and a very small number of workflows waited more than 1 hour. Outbound notifications and webhooks were also delayed until approximately 19:00 UTC, typically by about a minute.
As the backlog of delayed work cleared, the resulting increase in demand placed pressure on our execution capacity, causing further delays for Docker and Linux jobs. From approximately 19:45 UTC to 20:30 UTC, these jobs waited in the queue longer than usual before starting. At the peak, around 20:05 UTC, waits reached up to approximately 8 minutes for 2xlarge gen2, 6 minutes for xlarge gen2, and less than 5 minutes for a few other resource classes.
The issue has been resolved, and all affected functionality has returned to normal. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Workflows and jobs are starting normally, and wait times for Docker and Linux jobs have returned to normal. Our engineers are continuing to monitor this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: The backlog of delayed work has been cleared, and workflows are starting again with no delays in notifications. However, customers using Docker and Linux jobs across multiple resource classes, including medium, large, and 2xlarge gen2, may now see jobs waiting in the queue for few minutes before they start. Our engineers are continuing to work on mitigating this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Delays for workflows and jobs to start are continuing to decrease. Most are now starting within approximately 10 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour. If your job failed to start, please rerun the job. Delays in outbound notifications have improved.
The remaining delay is due to a backlog of work that built up during the incident, and that backlog has reduced by more than half. We expect the backlog to be cleared in approximately 15 minutes. We will share another update once it has cleared, or within the next 30 minutes at the latest.
Identified: Customers continue to experience delays. Most are now starting within approximately 20 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour.
Our engineers have made changes that have improved service stability, and the remaining delay is due to a backlog of work that built up during the incident. The backlog is decreasing, and we are focused on clearing it faster. We will share more information as progress continues.
We thank you for your patience while our engineers work through the backlog. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays before workflows and jobs start, in some cases up to approximately 30 minutes. This is improved from earlier in the incident, but delays remain. Delays in outbound notifications have improved.
We thank you for your patience while our engineers continue to work toward restoring normal job start times. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays between jobs in a workflow, as well as delays in outbound notifications.
We thank you for your patience while our engineers continue to investigate and work on mitigating the issue. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: While Pipelines, workflows, and jobs are running again, customers may continue to see longer than usual delays between jobs in a workflow, and outbound notifications are delayed by approximately 1 minute. Notifications and status updates to your version control system are beginning to go out again.
Our engineers are continuing to work toward a full recovery.
Identified: As of 16:39 UTC, we are starting to see recovery for new pipelines and workflows being submitted, and jobs are beginning to run again. Our engineers are continuing to work on the recovery and will share more information as soon as we have it.
Identified: As of 16:02 UTC, pipelines are failing entirely, so no new workflows or jobs are running at this point. This is in addition to the delays and failed jobs reported earlier. Our engineers are continuing to work on mitigating the issue, and we will share more information as soon as we have it.
Investigating: Since 15:47 UTC, customers may also see failed jobs, in addition to the delays previously reported. Our engineers are continuing to investigate and mitigate the issue.
Investigating: What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Recently Resolved Incidents
No recent incidents
CircleCI Outage Survival Guide
CircleCI Components
CircleCI Docker Jobs
There is an increased wait time for jobs on Docker Gen2.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're still seeing spikes in task wait times but on average wait times are looking better for all Docker Gen2.
Identified: Note: delays will be more noticeable for people running on Docker xlarge and above.
Identified: There is an increased wait time for jobs on Docker Gen2.
CircleCI Machine Jobs
Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: The incident has been resolved. Thank you for your patience
Monitoring: We are seeing signs of recovery and monitoring the situation.
Investigating: Customers may be experiencing elevated wait times for Machine Job Tasks, we are investigating the root cause and will update when we know more.
Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: Between 12:40 UTC and 16:00 UTC on October 8, customers using Linux machine jobs and remote Docker experienced elevated wait times. The issue has been resolved and wait times have returned to normal. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Wait times for customers using Linux machine jobs and remote Docker have returned to normal. We are monitoring to confirm wait times remain stable while we continue to add capacity. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 90 seconds, with the longest waits up to about 6 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:30 UTC.
Identified: Wait times continue to decrease, but customers using Linux machine jobs and remote Docker are still experiencing delays. Wait times currently average about 6 minutes. The longest waits, up to about 20 minutes, are on the 2xlarge, arm.2xlarge and gpu.nvidia.small resource classes. We are working to add capacity as quickly as possible. We will provide another update by 16:00 UTC.
Identified: A fix has been deployed and wait times are decreasing, but customers using Linux machine jobs and Remote Docker are still experiencing delays. Wait times currently average about 11 minutes, with the longest waits exceeding 35 minutes on some resource classes. We are working to add capacity as quickly as possible. We will provide another update by 15:30 UTC.
Identified: Customers using Linux machine jobs and remote Docker are experiencing elevated wait times. Wait times have started to decrease and now average about 20 minutes, with the longest waits exceeding 40 minutes on some resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 15:00 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 40 minutes, with the longest waits exceeding 50 minutes. Most Linux machine resource classes are affected, including medium, large, xlarge, 2xlarge and their Arm equivalents. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:30 UTC.
Identified: Customers using Linux machine jobs are experiencing elevated wait times, averaging about 11 minutes, with the longest waits exceeding 30 minutes on the medium, arm.medium and arm.large resource classes. Our engineers have identified the issue and are working on a fix. We will provide another update by 14:00 UTC.
Identified: Customers may be experiencing elevated wait times for machine jobs. We are working to resolve this.
CircleCI macOS Jobs
CircleCI Windows Jobs
CircleCI Pipelines & Workflows
What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Postmortem: ## Summary
On October 6, 2026, from 15:30 to 20:30 UTC, CircleCI customers experienced delays creating pipelines and starting workflows and jobs. From 15:48 to 16:38 UTC, no workflows or jobs started, and pipelines triggered from 16:02 UTC failed. Shorter, intermittent delays began earlier, at 13:30 UTC. Data in the CircleCI UI, notifications, and commit status updates to version control providers were also delayed.
The incident started in the database behind our workflow orchestration service, which coordinates every pipeline, workflow and job on CircleCI. Routine maintenance on the database's largest tables wrote transaction logs faster than the database could archive them, and the logs filled the storage volume set aside for them. To keep the database from becoming read-only, our engineers moved the logs to the database's main storage volume. That move took the database offline for 50 minutes. When the database came back, it processed work at about half its normal rate until our engineers changed a database setting at 18:19 UTC. We then worked through the backlog, and job start times returned to normal by 20:30 UTC.
Customers whose jobs failed during this window may rerun them.
The original status page can be found here.
Background
When you trigger a pipeline, CircleCI hands it to a workflow orchestration service. That service tracks the state of every workflow and job, decides when each job is ready to run, and passes ready jobs to our execution fleet. It also drives the job data you see in the UI and the notifications and commit statuses we send when work finishes.
The orchestration service stores its state in a database. The database writes every change to a transaction log before applying it, and it keeps those logs on a dedicated storage volume. It archives the logs continuously so it can reuse that space. If archiving falls behind and the volume fills, the database stops accepting writes, and no new work can move forward.
The database also runs a periodic maintenance process (autovacuums) on each table to keep its internal records valid over time. On a very large table, this process reads every page of the table and writes a large volume of transaction logs.
What Happened
(All times UTC)
From 10:02 on October 6, our monitoring raised short-lived alerts about errors in the orchestration service. Each alert cleared on its own within about 15 minutes. Our engineers had seen similar brief alerts before, and the system showed no other signs of stress, so they followed the runbook each time, and each alert cleared on its own partway through and also started work to make the alerts less sensitive.
Around 12:40, along with the regular workload which generates its own transaction logs, the maintenance process on several of the database's largest tables began writing additional transaction logs faster than the database could archive them, and the log volume started to fill.
At 13:30, the orchestration service began to slow down. Through 15:30, some pipelines and jobs took longer to start, in short bursts. At 13:35, our monitoring alerted again, this time alongside related alerts from several other services. Infrastructure Engineers followed the alert runbooks, clearing up some of the alerts. The alerts returned again at around 14:30 and after investigating further, we declared an incident at 14:54. We should have recognized the combination of alerts sooner, and we have covered how we are fixing that below.
After declaring the incident, our engineers traced the slowdown to the maintenance process (anti-wraparound autovacuum). The database restarts this process automatically if it is cancelled, so our engineers changed its settings to help it finish faster. By 15:30, the delays were continuous, and some jobs began to fail.
At 15:43, the transaction log volume was close to full. If it filled, the database would stop accepting writes and no work could run. Our engineers decided to move the transaction logs onto the database's main storage volume, which had plenty of free space. The move required the database to go offline, and we could not predict in advance how long that would take.
The move started at 15:47. From 15:48 to 16:38, the orchestration service could not process any work. Customers could still trigger pipelines, but no workflows or jobs started, and from 16:02 newly triggered pipelines failed. We posted to our status page at 15:58 and raised it to a major outage at 16:20. While the database was offline, our engineers prepared a replacement database as a fallback. The original database came back first, so we kept it in service.
At 16:38, the database came back online and work started flowing again. With transaction logs and regular data now sharing the same storage, the database ran more slowly than before. From 16:38 to 18:19, the orchestration service processed work at about half its normal rate, and jobs waited tens of minutes to start, up to about 50 minutes at the longest. During that time, the database had to finish archiving its backlog of transaction logs before we could move them back to a dedicated volume. We also investigated and implemented several database parameter changes to stop further maintenance runs from starting and prepared ways to reduce the work reaching the service. Along with the replacement database, we also prepared a complete stack of the service during that time so that we could move to it at the risk of data loss and opted against that move.
At 18:19, our engineers changed a database setting to cut the time each write spent waiting on storage. Processing speed recovered, and the backlog began to clear. To clear it faster, we added database and server capacity to the systems that hand jobs to our execution fleet. By about 19:55, new jobs were again starting on time.
The backlog then reached our execution fleet. From 19:49 to 20:30, some Docker jobs on larger resource classes waited up to about 10 minutes to start while capacity scaled up. By 20:30, job start times were back to normal. We moved the status page to monitoring at 20:32 and resolved the incident at 21:19.
About 3,500 jobs failed to start during the incident. Customers may rerun these jobs. We found no evidence that CircleCI ran any job more than once. Customers who reran a workflow while the original run was still delayed may have seen both runs complete. Annual plan customers can work with their account team to review usage.
Future Prevention and Process Improvement
We are taking the following steps to prevent a recurrence and improve our response time:
We are adding alerts on the conditions that caused this incident. Our monitoring caught the slowdown, but we had no alert on how full the transaction log volume was or on how far archiving had fallen behind. We are adding alerts on both, along with alerts on the database maintenance that drives them, so we can act earlier. We are also revisiting all our alerts to ensure noisy or sensitive alerts don’t hide real issues.
We are tuning how maintenance process (autovacuum) runs on our database. We are looking at the frequency and the aggressiveness with which the autovacuums run on our database. We are also tracking the size of these tables over time so they stay within safe limits.
We are removing the single transaction log bottleneck. All of the orchestration service's writes currently go through one database and one transaction log, so when that log fell behind, every pipeline, workflow and job slowed down with it. We are evaluating two approaches: splitting the service across multiple database instances, each with its own transaction log, and moving it to a distributed database built to spread writes across many nodes. Either approach would limit the effect of a backlog to part of the workload.
We are making our systems wait for the orchestration service to recover. During the incident, some jobs failed because our systems gave up after retrying. We are working on changing that behavior.
We are improving our status page updates during long incidents. Customers told us our early updates did not give enough detail. We are updating our guidance so updates include specific impact and timing sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve this incident. Please reach out to our support team with any questions or concerns.
Resolved: Between approximately 15:30 UTC and 20:30 UTC on October 6, 2026, customers experienced delays and failures affecting pipelines, workflows, and jobs.
From approximately 16:02 UTC to 16:39 UTC, pipelines were failing and no new workflows or jobs could start. After service resumed, workflows and jobs continued to start with delays until approximately 20:00 UTC. Start times averaged up to approximately 40 minutes at their peak, and a very small number of workflows waited more than 1 hour. Outbound notifications and webhooks were also delayed until approximately 19:00 UTC, typically by about a minute.
As the backlog of delayed work cleared, the resulting increase in demand placed pressure on our execution capacity, causing further delays for Docker and Linux jobs. From approximately 19:45 UTC to 20:30 UTC, these jobs waited in the queue longer than usual before starting. At the peak, around 20:05 UTC, waits reached up to approximately 8 minutes for 2xlarge gen2, 6 minutes for xlarge gen2, and less than 5 minutes for a few other resource classes.
The issue has been resolved, and all affected functionality has returned to normal. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Workflows and jobs are starting normally, and wait times for Docker and Linux jobs have returned to normal. Our engineers are continuing to monitor this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: The backlog of delayed work has been cleared, and workflows are starting again with no delays in notifications. However, customers using Docker and Linux jobs across multiple resource classes, including medium, large, and 2xlarge gen2, may now see jobs waiting in the queue for few minutes before they start. Our engineers are continuing to work on mitigating this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Delays for workflows and jobs to start are continuing to decrease. Most are now starting within approximately 10 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour. If your job failed to start, please rerun the job. Delays in outbound notifications have improved.
The remaining delay is due to a backlog of work that built up during the incident, and that backlog has reduced by more than half. We expect the backlog to be cleared in approximately 15 minutes. We will share another update once it has cleared, or within the next 30 minutes at the latest.
Identified: Customers continue to experience delays. Most are now starting within approximately 20 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour.
Our engineers have made changes that have improved service stability, and the remaining delay is due to a backlog of work that built up during the incident. The backlog is decreasing, and we are focused on clearing it faster. We will share more information as progress continues.
We thank you for your patience while our engineers work through the backlog. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays before workflows and jobs start, in some cases up to approximately 30 minutes. This is improved from earlier in the incident, but delays remain. Delays in outbound notifications have improved.
We thank you for your patience while our engineers continue to work toward restoring normal job start times. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays between jobs in a workflow, as well as delays in outbound notifications.
We thank you for your patience while our engineers continue to investigate and work on mitigating the issue. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: While Pipelines, workflows, and jobs are running again, customers may continue to see longer than usual delays between jobs in a workflow, and outbound notifications are delayed by approximately 1 minute. Notifications and status updates to your version control system are beginning to go out again.
Our engineers are continuing to work toward a full recovery.
Identified: As of 16:39 UTC, we are starting to see recovery for new pipelines and workflows being submitted, and jobs are beginning to run again. Our engineers are continuing to work on the recovery and will share more information as soon as we have it.
Identified: As of 16:02 UTC, pipelines are failing entirely, so no new workflows or jobs are running at this point. This is in addition to the delays and failed jobs reported earlier. Our engineers are continuing to work on mitigating the issue, and we will share more information as soon as we have it.
Investigating: Since 15:47 UTC, customers may also see failed jobs, in addition to the delays previously reported. Our engineers are continuing to investigate and mitigate the issue.
Investigating: What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
We are investigating a rise in infra fails for customer jobs.
Postmortem: ## Summary
Since October 1, CircleCI customers have experienced multiple incidents which have caused delays and failures in customer pipeline execution. There are two (unrelated) causes, both of which the CircleCI engineering team is actively mitigating:
- Available Cloud Computing Capacity
- Internal Infrastructure
Available Cloud Computing Capacity
Problem: Demand for high-performance cloud instances is rising extremely rapidly across the industry, which puts pressure on the cloud provider instance types that we use to execute customer jobs.
Incidents:
Mitigations and Resolutions:
- We are expanding our set of cloud computing regions to include additional regions with available high-performance instances.
- We are also working to secure additional guaranteed capacity from our cloud computing partners in our existing regions.
- We are further expanding the set of instance types we can offer to customers.
- We will publish a detailed Incident Report on these incidents on October 9, 2026.
Internal Infrastructure
Problem: We assign customer workflows over multiple compute providers via a workflow orchestration service. This service is currently gated by the write throughput of a single database.
Incidents:
Mitigations and Resolutions:
- By October 23, 2026, we plan to implement multiple redundant workflow orchestration service instances, each with its own database, which will enable us split the orchestration load between them.
Our Commitment
For avoidance of doubt,
- These particular incidents are not related to ongoing outages at Github
- We are not currently migrating our infrastructure or services
- We are not currently under cyberattack
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve these incidents and works to prevent future incidents.
Please reach out to our support team with any questions or concerns.
Resolved: This incident has been resolved.
Monitoring: We're seeing signs of recovery. Docker Gen2 wait times are slightly elevated and we're monitoring.
Investigating: We are investigating a rise in infra fails for customer jobs.
CircleCI API
CircleCI UI
CircleCI Artifacts
CircleCI Runner
CircleCI Webhooks
CircleCI Insights
CircleCI Releases
CircleCI Notifications & Status Updates
What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.
Postmortem: ## Summary
On October 6, 2026, from 15:30 to 20:30 UTC, CircleCI customers experienced delays creating pipelines and starting workflows and jobs. From 15:48 to 16:38 UTC, no workflows or jobs started, and pipelines triggered from 16:02 UTC failed. Shorter, intermittent delays began earlier, at 13:30 UTC. Data in the CircleCI UI, notifications, and commit status updates to version control providers were also delayed.
The incident started in the database behind our workflow orchestration service, which coordinates every pipeline, workflow and job on CircleCI. Routine maintenance on the database's largest tables wrote transaction logs faster than the database could archive them, and the logs filled the storage volume set aside for them. To keep the database from becoming read-only, our engineers moved the logs to the database's main storage volume. That move took the database offline for 50 minutes. When the database came back, it processed work at about half its normal rate until our engineers changed a database setting at 18:19 UTC. We then worked through the backlog, and job start times returned to normal by 20:30 UTC.
Customers whose jobs failed during this window may rerun them.
The original status page can be found here.
Background
When you trigger a pipeline, CircleCI hands it to a workflow orchestration service. That service tracks the state of every workflow and job, decides when each job is ready to run, and passes ready jobs to our execution fleet. It also drives the job data you see in the UI and the notifications and commit statuses we send when work finishes.
The orchestration service stores its state in a database. The database writes every change to a transaction log before applying it, and it keeps those logs on a dedicated storage volume. It archives the logs continuously so it can reuse that space. If archiving falls behind and the volume fills, the database stops accepting writes, and no new work can move forward.
The database also runs a periodic maintenance process (autovacuums) on each table to keep its internal records valid over time. On a very large table, this process reads every page of the table and writes a large volume of transaction logs.
What Happened
(All times UTC)
From 10:02 on October 6, our monitoring raised short-lived alerts about errors in the orchestration service. Each alert cleared on its own within about 15 minutes. Our engineers had seen similar brief alerts before, and the system showed no other signs of stress, so they followed the runbook each time, and each alert cleared on its own partway through and also started work to make the alerts less sensitive.
Around 12:40, along with the regular workload which generates its own transaction logs, the maintenance process on several of the database's largest tables began writing additional transaction logs faster than the database could archive them, and the log volume started to fill.
At 13:30, the orchestration service began to slow down. Through 15:30, some pipelines and jobs took longer to start, in short bursts. At 13:35, our monitoring alerted again, this time alongside related alerts from several other services. Infrastructure Engineers followed the alert runbooks, clearing up some of the alerts. The alerts returned again at around 14:30 and after investigating further, we declared an incident at 14:54. We should have recognized the combination of alerts sooner, and we have covered how we are fixing that below.
After declaring the incident, our engineers traced the slowdown to the maintenance process (anti-wraparound autovacuum). The database restarts this process automatically if it is cancelled, so our engineers changed its settings to help it finish faster. By 15:30, the delays were continuous, and some jobs began to fail.
At 15:43, the transaction log volume was close to full. If it filled, the database would stop accepting writes and no work could run. Our engineers decided to move the transaction logs onto the database's main storage volume, which had plenty of free space. The move required the database to go offline, and we could not predict in advance how long that would take.
The move started at 15:47. From 15:48 to 16:38, the orchestration service could not process any work. Customers could still trigger pipelines, but no workflows or jobs started, and from 16:02 newly triggered pipelines failed. We posted to our status page at 15:58 and raised it to a major outage at 16:20. While the database was offline, our engineers prepared a replacement database as a fallback. The original database came back first, so we kept it in service.
At 16:38, the database came back online and work started flowing again. With transaction logs and regular data now sharing the same storage, the database ran more slowly than before. From 16:38 to 18:19, the orchestration service processed work at about half its normal rate, and jobs waited tens of minutes to start, up to about 50 minutes at the longest. During that time, the database had to finish archiving its backlog of transaction logs before we could move them back to a dedicated volume. We also investigated and implemented several database parameter changes to stop further maintenance runs from starting and prepared ways to reduce the work reaching the service. Along with the replacement database, we also prepared a complete stack of the service during that time so that we could move to it at the risk of data loss and opted against that move.
At 18:19, our engineers changed a database setting to cut the time each write spent waiting on storage. Processing speed recovered, and the backlog began to clear. To clear it faster, we added database and server capacity to the systems that hand jobs to our execution fleet. By about 19:55, new jobs were again starting on time.
The backlog then reached our execution fleet. From 19:49 to 20:30, some Docker jobs on larger resource classes waited up to about 10 minutes to start while capacity scaled up. By 20:30, job start times were back to normal. We moved the status page to monitoring at 20:32 and resolved the incident at 21:19.
About 3,500 jobs failed to start during the incident. Customers may rerun these jobs. We found no evidence that CircleCI ran any job more than once. Customers who reran a workflow while the original run was still delayed may have seen both runs complete. Annual plan customers can work with their account team to review usage.
Future Prevention and Process Improvement
We are taking the following steps to prevent a recurrence and improve our response time:
We are adding alerts on the conditions that caused this incident. Our monitoring caught the slowdown, but we had no alert on how full the transaction log volume was or on how far archiving had fallen behind. We are adding alerts on both, along with alerts on the database maintenance that drives them, so we can act earlier. We are also revisiting all our alerts to ensure noisy or sensitive alerts don’t hide real issues.
We are tuning how maintenance process (autovacuum) runs on our database. We are looking at the frequency and the aggressiveness with which the autovacuums run on our database. We are also tracking the size of these tables over time so they stay within safe limits.
We are removing the single transaction log bottleneck. All of the orchestration service's writes currently go through one database and one transaction log, so when that log fell behind, every pipeline, workflow and job slowed down with it. We are evaluating two approaches: splitting the service across multiple database instances, each with its own transaction log, and moving it to a distributed database built to spread writes across many nodes. Either approach would limit the effect of a backlog to part of the workload.
We are making our systems wait for the orchestration service to recover. During the incident, some jobs failed because our systems gave up after retrying. We are working on changing that behavior.
We are improving our status page updates during long incidents. Customers told us our early updates did not give enough detail. We are updating our guidance so updates include specific impact and timing sooner.
Customer experience is our top priority, and we commit to continually improving the reliability of our systems to match the trust that our customers place in us. We thank our customers for their patience while our team worked to resolve this incident. Please reach out to our support team with any questions or concerns.
Resolved: Between approximately 15:30 UTC and 20:30 UTC on October 6, 2026, customers experienced delays and failures affecting pipelines, workflows, and jobs.
From approximately 16:02 UTC to 16:39 UTC, pipelines were failing and no new workflows or jobs could start. After service resumed, workflows and jobs continued to start with delays until approximately 20:00 UTC. Start times averaged up to approximately 40 minutes at their peak, and a very small number of workflows waited more than 1 hour. Outbound notifications and webhooks were also delayed until approximately 19:00 UTC, typically by about a minute.
As the backlog of delayed work cleared, the resulting increase in demand placed pressure on our execution capacity, causing further delays for Docker and Linux jobs. From approximately 19:45 UTC to 20:30 UTC, these jobs waited in the queue longer than usual before starting. At the peak, around 20:05 UTC, waits reached up to approximately 8 minutes for 2xlarge gen2, 6 minutes for xlarge gen2, and less than 5 minutes for a few other resource classes.
The issue has been resolved, and all affected functionality has returned to normal. Customers whose jobs failed during this window may rerun affected jobs. We thank you for your patience while our team worked on implementing a fix.
Monitoring: Workflows and jobs are starting normally, and wait times for Docker and Linux jobs have returned to normal. Our engineers are continuing to monitor this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: The backlog of delayed work has been cleared, and workflows are starting again with no delays in notifications. However, customers using Docker and Linux jobs across multiple resource classes, including medium, large, and 2xlarge gen2, may now see jobs waiting in the queue for few minutes before they start. Our engineers are continuing to work on mitigating this issue. If your job failed to start, please rerun the job.
We thank you for your patience while our engineers work on this. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Delays for workflows and jobs to start are continuing to decrease. Most are now starting within approximately 10 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour. If your job failed to start, please rerun the job. Delays in outbound notifications have improved.
The remaining delay is due to a backlog of work that built up during the incident, and that backlog has reduced by more than half. We expect the backlog to be cleared in approximately 15 minutes. We will share another update once it has cleared, or within the next 30 minutes at the latest.
Identified: Customers continue to experience delays. Most are now starting within approximately 20 minutes, improved from approximately 40 minutes earlier in the incident. A very small number of workflows may have waited more than 1 hour.
Our engineers have made changes that have improved service stability, and the remaining delay is due to a backlog of work that built up during the incident. The backlog is decreasing, and we are focused on clearing it faster. We will share more information as progress continues.
We thank you for your patience while our engineers work through the backlog. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays before workflows and jobs start, in some cases up to approximately 30 minutes. This is improved from earlier in the incident, but delays remain. Delays in outbound notifications have improved.
We thank you for your patience while our engineers continue to work toward restoring normal job start times. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: Customers continue to experience longer than usual delays between jobs in a workflow, as well as delays in outbound notifications.
We thank you for your patience while our engineers continue to investigate and work on mitigating the issue. We will share another update within the next 30 minutes, or sooner if more information becomes available.
Identified: While Pipelines, workflows, and jobs are running again, customers may continue to see longer than usual delays between jobs in a workflow, and outbound notifications are delayed by approximately 1 minute. Notifications and status updates to your version control system are beginning to go out again.
Our engineers are continuing to work toward a full recovery.
Identified: As of 16:39 UTC, we are starting to see recovery for new pipelines and workflows being submitted, and jobs are beginning to run again. Our engineers are continuing to work on the recovery and will share more information as soon as we have it.
Identified: As of 16:02 UTC, pipelines are failing entirely, so no new workflows or jobs are running at this point. This is in addition to the delays and failed jobs reported earlier. Our engineers are continuing to work on mitigating the issue, and we will share more information as soon as we have it.
Investigating: Since 15:47 UTC, customers may also see failed jobs, in addition to the delays previously reported. Our engineers are continuing to investigate and mitigate the issue.
Investigating: What's impacted Customers using CircleCI pipelines and workflows, including UI data display, outbound notifications, and outbound webhooks.
What can you expect Delays in pipelines being created and in workflows and jobs starting. Job, workflow, and pipeline data may take up to 15 minutes to appear in the UI. Outbound notifications and webhooks are delayed by up to 1 minute.
Next update We will provide another update within 30 minutes or as we have more information to share.