Harness logo

Harness Status Page

CI/CD & Build · monitored by Alert24

harness.io
All Systems Operational

Is Harness down right now?

No — Harness is up. All systems operational as of Aug 28, 2:49 AM UTC.

Current Status

All Systems Operational

View Harness status page ↗

Components

Continuous Delivery (CD) - FirstGen - EOS
Operational
Continuous Delivery (CD) - FirstGen - EOS
Operational
Continuous Delivery (CD) - FirstGen - EOS
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Software Engineering Insights FirstGen (fka Propelo) - EU
Operational
US - app.traceable.ai / api.traceable.ai
Operational
SDK API
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Continuous Delivery (CD) - FirstGen - EOS
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Continuous Delivery (CD) - FirstGen - EOS
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Continuous Delivery - Next Generation (CDNG)
Operational
Software Engineering Insights FirstGen (fka Propelo) - US
Operational
Cloud Cost Management (CCM)
Operational
Chaos Engineering
Operational
EU - app.eu.traceable.ai / api.eu.traceable.ai
Operational

Recent Incidents

Pipelines are failing for harness IACM customers

major

Aug 26, 2026 · resolved Aug 26

This incident has been resolved.

Feature Management & Experimentation (FME) user interface unavailable

major

Aug 24, 2026 · resolved Aug 24

## Summary * Starting at **23:42 UTC** on August 23, 2026, several FME customers reported failures loading the FME UI. * FME UI Artifacts served from the CDN expired due to a retention policy, causing FME UI to fail to load. * Any flag request changes through the API, change delivery, and the data pipeline continued to work with no interruption. ## Root Cause * The FME UI is served from a CDN. The UI artifacts got evicted due to a retention policy, causing the UI to fail to load for all users. ## Impact * The FME UI was unable to load for all users across all production environments. ### What was not impacted? * SDK functionality and runtime flag evaluation * Admin API calls * Customer flag configuration data * No data loss occurred ## Remediation * FME UI got restored in the CDN through a deployment * Recovery confirmed across all production environments before closing the incident. ## Action Items * Improve the asset retention policy so that the currently active version is never subject to eviction.

All modules are running slow in Prod1/2/3/4 due to cloud provider incident

major

Aug 20, 2026 · resolved Aug 20

# Summary On 20 August 2026, beginning at approximately 15:00 UTC, the Harness platform experienced widespread performance degradation across all production environments. Pipeline executions that normally complete in around two minutes took seven to ten minutes. Continuous Delivery, Continuous Integration, pipeline orchestration, and Feature Management & Experimentation were all affected. Google Cloud Platform experienced a multi-product incident in the us-west1 region affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk I/O. Harness production infrastructure runs on persistent disks in that region. The degradation raised database operation latency from approximately 2 ms to over 10 ms at the 95th percentile, which in turn caused message-queue processing lag and propagated to every service that depends on timely database access. ‌ # Impact This was a degradation, not an outage. Pipelines continued to execute and complete successfully throughout; they were slow rather than failing. No data was lost, and no customer work was dropped as a result of this incident. # **Root cause** Harness production infrastructure in the affected environments runs on Google Cloud Platform persistent disks in the us-west1 region. When that storage layer degraded, the effect propagated through the platform in a predictable chain: **Persistent-disk I/O degradation in us-west1.** Google Cloud Platform experienced a multi-product incident affecting Bigtable, Compute Engine, Google Kubernetes Engine, and persistent-disk performance. This was an infrastructure failure in the provider’s environment, outside Harness’s control. # **Preventive actions** Although Harness cannot prevent a cloud provider infrastructure failure. The actions below are aimed at detecting one faster and being better positioned to act on it. | **Action** | | --- | | Continue routine pre-testing of targeted cross-region database failovers, as performed during this incident, to keep failover readiness verified rather than assumed | | Assess full-stack multi-region failover readiness for future scenarios in which cross-region latency would be unacceptable |

FME API write operations started returning 499 errors

minor

Aug 20, 2026 · resolved Aug 20

### Summary On August 20, 2026, between 10:24 and 14:55 UTC, a subset of FME writes failed. Writes made from the FME UI and writes made with Harness access tokens \(PATs and SATs\) were not affected. Runtime flag evaluation continued to work normally. The issue was mitigated by reverting a recent authentication change in a shared governance service, and affected writes returned to normal by 14:55 UTC. Status: [https://status.harness.io/incidents/rhthgm7d5dkz](https://status.harness.io/incidents/rhthgm7d5dkz) ### Root Cause A change in how a shared governance service authenticated inbound calls resulted in some FME writes being rejected. Those writes used service-to-service credentials that the governance service could no longer verify after the change. FME surfaces a governance failure to the client as HTTP 499, the same status used when a governance policy intentionally denies a change. Because 499 is a valid, expected response in that deny path, the failures did not look like an outage on our alerts, and the incident was identified from customer reports rather than internal detection. ### Impact * A subset of FME writes failed during the window, primarily those made using legacy Split API keys or change request scheduling. * Writes made from the FME UI were not impacted. * Writes using Harness access tokens \(PATs and SATs\) were not impacted. * Runtime flag evaluation continued normally. * No data loss occurred. Failed writes did not apply. ‌ ### Remediation Reverted the governance-service authentication change. Affected writes returned to normal immediately. ### Action Items To prevent such issues from happening again, * Harness will return a distinct error \(not 499\) when a write fails because governance could not be evaluated, so it is not confused with an intentional policy denial. * Add alerting on the governance evaluation call itself, rather than relying on the client-facing status code. * Expand authentication support for policy evaluations. * Expand automated coverage for additional write scenarios.

Data ingestion is delayed on Traceable US production

major

Aug 19, 2026 · resolved Aug 19

**Summary** On 19 August 2026 between 12:35 and 17:29 UTC, the Harness Application Security service experienced a significant disruption affecting both the customer-facing console and the data ingestion pipeline in the SaaS Production and US1 regions. ‌ **Root Cause** The internal configuration service that supplies runtime settings to nearly every other component became overloaded and entered a repeated restart cycle. Because so many services depend on it, the effects were broad: console pages such as protection policies, posture views, activity logs, API inventory, and custom policy failed to load or timed out, and downstream processing stalled while waiting for configuration it could not obtain. # **Customer impact** | **Dimension** | **Detail** | | --- | --- | | Console \(UI\) impact | Multiple pages failed to load or timed out, including protection policies, posture event pages and posture views inside dashboards and insight pages, activity log queries, API inventory screens, custom policy, and sensitive-data views and widgets. | | Ingestion impact | Security telemetry processing degraded severely and, in some paths, stopped entirely. Consumer lag grew across normalisation, grouping, anomaly detection, generation, and related processing stages. | | Data loss | A subset of telemetry ingested during the disruption was permanently dropped. | ‌ **Mitigation** Several intermediate mitigations additional CPU and memory, relaxed health-check thresholds, a database restart, and a larger connection pool ameliorated the issue. Disabling the new feature in both affected regions restored throughput sharply and durably. The incident was resolved at 17:29 UTC. ‌ # **Preventive actions** The following actions are committed and tracked internally to completion. The feature that triggered this incident remains disabled and will not be re-enabled until the work below is complete and validated. | **Action** | | --- | | | | OPtimize the code by tuning parameters such as cache eviction and retention , evaluate cursor-based pagination for bulk rule retrieval as rule counts grow | | Add a purpose-built database index for the service-scoping access pattern | | Remediate pipeline recovery semantics so consumers replay safely after position-marker loss instead of skipping backlog | | Mandate staged rollout for configuration overrides that alter downstream request patterns: low-volume cluster, then mid-volume, then high-volume | | Add backpressure and concurrency protection to the configuration service: circuit breaking, bounded queues, and timeout isolation | | Enhance observability by Instrumenting more detailed metrics |

Get alerted when Harness goes down

Alert24 monitors Harness and 3,700+ other cloud and SaaS providers. When an outage is detected, it updates your status page automatically and pages your on-call team. No manual updates at 2 AM.

Start free — no credit card

Harness status — frequently asked questions

Is Harness down right now?

No — Harness is up. All systems operational as of Aug 28, 2:49 AM UTC.

What is Harness's current status?

Harness: All Systems Operational. Alert24 checks Harness's status page continuously and can notify you the moment it changes.

How do I get alerted when Harness goes down?

Alert24 monitors Harness and 3,700+ other cloud and SaaS providers. When an outage is detected it updates your status page automatically and pages your on-call team — no manual checks. Start free at alert24.net.

More CI/CD & Build status pages