Buildkite logo

Buildkite Status Page

CI/CD & Build · monitored by Alert24

All Systems Operational

Is Buildkite down right now?

No — Buildkite is up. All systems operational as of Aug 28, 1:46 AM UTC.

Current Status

All Systems Operational

View Buildkite status page ↗

Components

REST API
Operational
Web
Operational
AWS ec2-us-east-1
Operational
GitHub
Operational
GitHub Commit Status Notifications
Operational
Web
Operational
Hosted Agents
Operational
Web
Operational
Email Notifications
Operational
Agent API
Operational
AWS elasticache-us-east-1
Operational
GitHub API Requests
Operational
Ingestion
Operational
MacOS
Operational
Package Managers - API
Operational
Remote MCP Server
Operational
Slack Notifications
Operational
AWS elb-us-east-1
Operational
REST API
Operational
GitHub Webhooks
Operational

Recent Incidents

Test Engine ingestion delayed, catching up

minor

Aug 27, 2026 · resolved Aug 27

Processing of uploaded test execution data has caught up and is back to normal.

Issues identified with buildkite-agent registration

minor

Aug 26, 2026 · resolved Aug 26

Mitigation strategies worked as expected and now agent registration is back to normal. All systems are stable now.

Buildkite service disruption

major

Aug 25, 2026 · resolved Aug 26

## Service Impact Customers experienced a site-wide disruption affecting the Buildkite web interface, REST API, Agent API, job queue, and notifications. The disruption began at 22:47 UTC and continued until 23:16 UTC. During this period, all customers were unable to create new builds via the UI or REST API, and job dispatch was delayed. We continued to receive webhooks from source control providers, but processing of those webhooks was delayed. New build processing resumed at 23:10 UTC and the accumulated backlog of work was processed by 23:16 UTC. Some builds that were already in progress at the beginning of the impact period remained temporarily stuck until services recovered. Some customers also experienced errors in the web UI, and on receiving webhooks from SCM providers. This impacted <1.5% of all requests during the impact period. ## Incident Summary **Background** Buildkite services running in one of our production Kubernetes clusters depend on an internal DNS service CoreDNS to locate databases, queues, and other application components. Over the last four months, we have been migrating our production workloads from AWS ECS to this AWS EKS cluster. The cluster has been steadily increasing in size during the course of this migration. Additionally, Buildkite recently moved time-sensitive notification jobs from a general-purpose pool of background workers into a new low-latency worker pool. To ensure sufficient capacity for both pools, we initially configured each with the same high maxReplica count as the original shared pool, with the intention of reviewing and adjusting the limits for both pools downwards at a later date. **Trigger: A sudden increase in demand on CoreDNS** At 22:44 UTC, an application deploy created a surge in application Pod volume, which caused our EKS cluster to scale out. This surge consumed the available headroom on already-deployed Nodes, which limited applications’ capacity to autoscale promptly. Some background workers \(including the aforementioned notification workers\) were also attempting to scale out at this time. The headroom shortage and subsequent cluster autoscaling delayed provisioning of the compute requested by those services. Since there had been no change in the metrics that triggered the services to scale up, those services requested even more Pods. Normally the impact of such runaway autoscaling would be limited by the services’ configured maximums. However, as mentioned previously the maximums for these services had been set higher than usual. The application deployment and runaway autoscaling combined to trigger an unusually high rate of change to applications, network endpoints, and cluster nodes. **Why CoreDNS failed** The cluster's CoreDNS service was running at a fixed size and did not automatically scale with the size or rate of change of the cluster. During post-incident analysis we discovered some bugs and gaps in our monitoring of CoreDNS, particularly around query volume and duration. * We found a defect in a key monitoring query which had masked an upward trend in query duration, correlated with increasing cluster size. * We also found that DNS query volume was not sufficiently monitored, and had been trending upwards as we migrated more workloads into EKS. These defects masked the upward trend in latency on the CoreDNS side; the client-side latency increase was offset by recent performance gains in our applications, and so did not catch our attention. Hence, we had not correctly prioritised our planned implementation of autoscaling for CoreDNS. **What happened when CoreDNS failed** During the impact period CoreDNS’s completed query rate remained stable, but query processing time increased from under 1ms to approximately 780ms. Pending requests accumulated in memory, until all three original CoreDNS pods exceeded their allowed memory limits and were restarted by Kubernetes. Continued demand and DNS retries prevented the service from recovering after restarting. CoreDNS unavailability caused failures across APIs, job dispatch, and notifications for all customers. Retries and delayed work increased the load during recovery. As part of the migration from ECS to EKS, database queries for a subset of customers were routed to PgBouncer instances running in the EKS cluster. During the impact period, these customers’ requests to webhooks and the Web UI received error responses. It was less than 1.5% of all requests during this incident that returned errors. **How we responded** We paused further application deployments, deployed more CoreDNS service replicas, raised the memory available to each CoreDNS replica, and expanded the node pool available to run them. The new set of CoreDNS Pods came into service by 23:12 UTC. After that, DNS errors fell rapidly. Customer-facing services processed the accumulated backlog and recovered fully by 23:16 UTC. ## Changes we're making * **Increase immediate CoreDNS headroom.** We increased CoreDNS from three to twelve pods, raised each pod’s memory limit from 512 MiB to 2 GiB, and expanded the system node pool available to run them. * **Enable automatic CoreDNS capacity scaling.** We will enable EKS-managed CoreDNS autoscaling, with a tested minimum replica count and sufficient capacity to distribute those replicas across nodes and availability zones. * **Detect both rapid degradation and declining headroom.** We will add direct alerts for CoreDNS query duration, query volume, goroutine growth, memory pressure, OOM restarts, and available replicas. We will also monitor longer-term trends as part of capacity planning. * **Apply the same standard to other cluster-critical services.** We will identify and remediate critical services that lack tested capacity, direct alerting, failure-domain distribution, and either safe autoscaling or documented static headroom.

Unexpected agent disconnection for some customers

minor

Aug 25, 2026 · resolved Aug 25

This incident has been resolved.

Buildkite service disruption

minor

Aug 17, 2026 · resolved Aug 17

This incident has been resolved.

Get alerted when Buildkite goes down

Alert24 monitors Buildkite and 3,700+ other cloud and SaaS providers. When an outage is detected, it updates your status page automatically and pages your on-call team. No manual updates at 2 AM.

Start free — no credit card

Buildkite status — frequently asked questions

Is Buildkite down right now?

No — Buildkite is up. All systems operational as of Aug 28, 1:46 AM UTC.

What is Buildkite's current status?

Buildkite: All Systems Operational. Alert24 checks Buildkite's status page continuously and can notify you the moment it changes.

How do I get alerted when Buildkite goes down?

Alert24 monitors Buildkite and 3,700+ other cloud and SaaS providers. When an outage is detected it updates your status page automatically and pages your on-call team — no manual checks. Start free at alert24.net.

More CI/CD & Build status pages