Summary
On July 24, 2026, at 20:22 UTC, our engineering team began investigating an issue affecting customers in our ca1: Cloud — Canada region. Customers with affected workspaces were unable to open their models. As a precaution, while we validated the scope of the issue, we restricted access to the region behind a maintenance page at 02:18 UTC on July 25. General access to CA1 was restored at 17:24 UTC on July 25. Bring Your Own Key (BYOK) workspaces remained offline for additional safeguards, and were re-enabled progressively as those safeguards were validated. The incident was fully resolved on July 31, 2026, at 13:47 UTC.
Root cause
The disruption was caused by a defect in a third-party component used within our BYOK service. The defect only surfaced under a very specific combination of events occurring in a particular order on the same host. When a BYOK workspace was unloaded, the component failed tofully clear one of its local resources, leaving behind a stale reference. When a BYOK workspace was subsequently loaded onto the same host, the component attempted to clean up that stale reference before proceeding. During this step, it incorrectly executed a removal that extended beyond the stale reference and deleted files that were still in active use.
The affected files were captured in our regular backups. Once our engineering team identified the source of the activity, isolated it, and applied protective controls to stop any further impact, restoration became a controlled process of returning each affected file to its most recent backup.
Recovery
Our engineering team identified the issue and isolated it at its source, then worked systematically to restore the affected files. This allowed us to bring non-BYOK workspaces back online in a controlled sequence, and general access to CA1 was restored at 17:24 UTC on July 25.
We deliberately kept BYOK workspaces offline while our team worked with the third-party vendor to reproduce the trigger in a controlled, non-production environment. This reproduction gave us the diagnostic evidence the vendor needed to build a fix and to confirm the exact cause. It also allowed us to develop and validate our own temporary safeguards — targeted changes to how BYOK workspaces are scheduled — that eliminated the specific combination of conditions required to trigger the defect. These safeguards act as compensating controls to bring BYOK workspaces back online while the third party completes the permanent fix to the underlying component. We re-enabled BYOK workspaces once those safeguards were validated, ensuring the trigger conditions couldn't recur. The incident was fully resolved on July 31, 2026, at 13:47 UTC.
Corrective and preventative actions
Our corrective actions follow two complementary tracks. The first removes the specific combination of conditions required to trigger the defect, using controls we have developed and deployed ourselves as compensating safeguards. The second is the permanent fix to the underlying component itself, which the third party is delivering. Together, these tracks address both the trigger and the defect, so that neither can produce another incident of this kind. We are implementing the following actions to prevent recurrence:
We have deployed changes that prevent the specific combination of conditions required to trigger the defect. This is a temporary but effective control that removes the trigger today, ahead of the permanent fix.
We are working with the third party to deploy their validated fix. This removes the defect at its source and closes the underlying cause of this incident.
We are strengthening how we validate BYOK third-party components in non-production before they reach production, including reproducing a wider range of workspace lifecycle scenarios and event sequences. This gives us stronger assurance to surface these types of issues in non-production and are addressed before they can affect customers.
We have deployed dedicated alerting on the specific event pattern that triggered this incident and are actively reviewing additional file-level alerting. These alerts provide an additional safety net and earlier warning, enabling faster preventative action before customers are affected.
Closing
We apologize for any impact this issue may have had on your business operations. We are continuously strengthening our systems and procedures to ensure we avoid future disruptions to your business and users.
If you have further questions or concerns, please visit our Support website. We appreciate your patience during this incident and value the trust you place in Anaplan.