ResolvedBetween 14:37 UTC and 14:41 UTC on the 22nd of May we observed a loss of video transmission on a number of webinars live at the time.
The issue was caused by the introduction of a new feature around the mixing of webinar streams. The introduction of this feature, although in beta and not accessible to the majority of our users, required an overhaul of the inner workings controlling how images and video are displayed on screen. This is what ultimately introduced the issue that some users experienced.
The issue would've been felt by users performing a specific sequence of mixing actions during the 4 minute response window, or after this period, provided they had been in the webinar room during that timespan. A partial amount of webinars managed to be manually recovered and resumed without issues, but unfortunately, for some of the affected users, it resulted in the webinars ending up in an unrecoverable state.
During the whole process, both video and audio transmission managed to reach our servers, as well as other speakers and producers, but only audio ended up reaching the webinar audience. As such, integral recording of all the webinar sources are available and can be recovered using our RawMode feature.
Timeline:
At 14:37 UTC we received the first report of the issue
At 14:38 UTC, post initial diagnostic steps the issue got escalated to engineering
Between 14:39 UTC and 14:40 UTC we attempted to recover the affected webinars while attempting to identify the cause of the issue
At 14:40 UTC we identified the recently made release as the primary suspect for the outage
At 14:41 UTC we rolled back on the introduction of this feature, ensuring the issue wouldn't spread further
We spent the following minutes attempting to restore the state of affected webinars with partial success.
ResolvedBetween 6:41 UTC and 7:20 UTC on the 13th of May we observed an issue preventing speakers from entering webinar rooms.
This issue was caused by a new feature introduced to the WebinarTray that through a development oversight, caused an error to happen, specificaly for speakers, when entering the webinar room. Neither webinar producers, nor audience members were affeted by this issue.
Timeline:
* At 8:41 we received the first report of users having issues going into a webinar room;
* At 8:46 and in the following minutes we received more reports of users having issues going into webinar rooms;
* At 9:00 the issue got escalated to engineering;
* At 9:04 the cause of the issue was identified;
* Between 9:06 and 9:10 we tried to mitigate the issue for existing, live, or soon to be live webinars;
* At 9:14 a resolution for the issue was developed and a patch created;
* At 9:19 a patch got released fixing the issue;
* At 9:20 access to the webinar room was back to nominal conditions and the issue fully resolved.
Moving forward we'll work on improving our QA processes to make sure that issues like this one don't escape our tests and are caught before reaching our users.
ResolvedBetween 10:13 and 10:18 UTC the TwentyThree platform responded with high latency with a portion of requests taking tens of seconds to finish.
Our API received traffic with unusual patterns, generating a lot of stress on our main database. Which led to database queries stalling which in turn led to increased latency. We are working on optimizing the API to avoid this issue in the future.
1 May 2024
No incidents reported.
Availability is (seconds up + seconds unmonitored) รท all seconds in the Europe/Copenhagen calendar month; only measured downtime counts against it, Maintenance included.