Incident
Long response times
Affected: TwentyThree Platform
Resolved### Summary We are following up on the outage **on January 20th 2022** starting at 08:45 UTC. The outage was caused by a load balancer not being in a correct state and where some requests to the platform would time out in the interval between 08:45 UTC and 09:29 UTC. We have mitigated this issue in the future and increased the resilience of our platform in the process. We know that issues such as this one significantly impacts our customers – and want to use this occasion to apologize for the problems caused. In addition, we are committed to using the findings to improve our service further. ### Timeline At 08:45 UTC on Thursday, January 20th, our load balancer stopped responding to a portion of incoming requests. The load balancer operated in 3 different geographical locations or zones, and the traffic timed out when the requests reached a specific location. As a result, roughly a third of incoming requests timed out. Affected services were TwentyThree Platform, Analytics, Webinars, and FTP. At 08:45, our monitoring infrastructure registered the issues and alerted on-duty personnel to respond to the problem. At 08:48, the responding team acknowledged the problem, and work commenced to identify and resolve the issue. At 09:09, a notice about the ongoing issue was posted on the TwentyThree status page. At 09:23, we identified a load balancer as the issue, and we started investigating why the load balancer had stopped responding. We escalated the problem to our cloud provider AWS. At 09:29, we deployed changes to our infrastructure to avoid using the faulty zone on our core system. After these changes, the TwentyThree Platform and Webinars became fully operational. However, our Analytics and FTP services were still suffering from timeouts on roughly one-third of incoming. At 11:00, we deployed a new load balancer. After this change, all services apart from our FTP service became fully operational. At 11:40, we deployed changes to the FTP service to become fully operational. ### What happened to cause the incident? The TwentyThree platform runs on the Amazon Web Services cloud. We use an AWS-provisioned load balancer to handle incoming traffic, and the load balancer is deployed in three different availability zones. The implicit behavior of this load balancer is to accept traffic in availability zones that appear healthy. The TwentyThree platform is deployed in two availability zones, and thus the load balancer would accept traffic in these availability zones. The incident happened when AWS stopped registering some of our internal services as "healthy" despite them still accepting traffic and reporting healthy. However, because they appeared "unhealthy" to the load balancer, it seemed as if none of the availability zones could accept traffic. This triggered "fail open" behavior, making the load balancer accept traffic in all availability zones, regardless of health status. This led to timeouts when the load balancer started accepting traffic in an availability zone without any resources to handle that traffic. ### What will TwentyThree do to ensure the problem doesn't re-occur? Based on the findings above, we have changed the settings of the load balancer to accept traffic in all availability zones and to route it to appropriate resources, even if they are in a different availability zone. These settings changes will prevent similar incidents from happening in the future.
ResolvedThe issue has been fully resolved. We started seeing problems delivering traffic through a load balancer in one of our AWS availability zones at 8:48am UTC, which can have caused the service to be unavailable to customers in 30 second intervals in the minutes afterwards. After identifying the problem, we moved traffic away from the faulty load balancer, which was fully completed at 9:30am UTC.
MonitoringWe have identified on a network load balancer run by our hosting partner, AWS. We have removed the faulty component from use and we're seeing full recovery. This thread continue to be updated.
InvestigatingWe are investigating an issue on the platform causing slow loading. We will update this thread when we know more.