Root Cause
The incident was caused by a deployment of one service that ran a database index creation. The team has identified an issue with the way we perform index creations as database migrations.
During a deployment the pods with newest image of the service attempted to create an index on a big table. The team has identified that the index creation took 7 minutes. However, during index creation pods were not responsive, and as a result, Kubernetes deemed them as unhealthy pods and attempted to retry those pods after 5 minutes.
As a result, because the pod got killed before the index creation was fully completed, the database transaction was rolled back. Then, subsequent pods attempted to create the index again, dying after 5 minutes.
This lead to the database being in unhealthy state, and the service was down.
The team has rolled back the deployment, and killed all replicas that attempted to create the index. Thanks to that, the service and the database was in healthy state again.
Why safety measures did not help
Cycode provides a safety mechanism that unblocks all Pull Request scans after a specific period of time, giving each scan a maximum duration before the Pull Request is unblocked. However, because the service that is responsible for triggering and completing scans, as well as this safety net, was down, the process couldn't behave as expected. We acknowledge this gap and are working on strengthening this area of our system.
Action items
The team is actively investigating enhancements and new safety protocols that can be put in place in order to have another safety net preventing Pull Request scans being stuck in case of any incident.
The team is investigating changes to the index creation process.
Resolved
Root Cause
The incident was caused by a deployment of one service that ran a database index creation. The team has identified an issue with the way we perform index creations as database migrations.
During a deployment the pods with newest image of the service attempted to create an index on a big table. The team has identified that the index creation took 7 minutes. However, during index creation pods were not responsive, and as a result, Kubernetes deemed them as unhealthy pods and attempted to retry those pods after 5 minutes.
As a result, because the pod got killed before the index creation was fully completed, the database transaction was rolled back. Then, subsequent pods attempted to create the index again, dying after 5 minutes.
This lead to the database being in unhealthy state, and the service was down.
The team has rolled back the deployment, and killed all replicas that attempted to create the index. Thanks to that, the service and the database was in healthy state again.
Why safety measures did not help
Cycode provides a safety mechanism that unblocks all Pull Request scans after a specific period of time, giving each scan a maximum duration before the Pull Request is unblocked. However, because the service that is responsible for triggering and completing scans, as well as this safety net, was down, the process couldn't behave as expected. We acknowledge this gap and are working on strengthening this area of our system.
Action items
The team is actively investigating enhancements and new safety protocols that can be put in place in order to have another safety net preventing Pull Request scans being stuck in case of any incident.
The team is investigating changes to the index creation process.
Resolved
System should be back to being fully operational.
Investigating
All scan types except for SAST are fully operational. SAST continues to stabilize and will soon be fully stable.
Investigating
The root cause has been resolved.
The system began to stabilize itself, and all the scans are starting to get processed with regular performance.
Investigating
The team has identified the root cause of the issue and is working on the solution.
Investigating
Customers may experience degraded performance in scans. Pull request and CLI scans may be affected.