Jul 15, 23:33 UTC
Resolved- We have re-enabled the affected events as those events no longer appear to be a major source of issues.
Feb 17, 22:14 UTC
Monitoring- Service has been stable for 24 hours.
Events will continue to be disabled until we engineer an infra or software solution.
Edit (4th Mar): Someone is assigned and actively working on a software solution.
Feb 16, 21:14 UTC
Investigating- Aware some events are being dropped, currently adding event replicas to fix this.
Update: Now being rolled out to cluster.
Update: Rolled back event replicas, ended up causing more issues.
Feb 16, 21:10 UTC
Monitoring- Just had a major face palm moment regarding the database, service should be a bit more stable now.
More infra changes to follow. 🙂
Feb 16, 18:53 UTC
Identified- Trying to ease up the burden on parts of the system, there will be rolling restarts of services which ), but they may increase load. This process should complete in about 15 minutes, up to 30 minutes at most.
Update: It's dropping enough users in each batch that the database is tanking performance to zero, we'll be re-scaling this afterwards.
Update: Preparing to re-scale the database.
Feb 13, 18:20 UTC
Identified- Just saw a burst of users coming online, attempting to re-scale.
Update: Scaled, monitoring performance, we may have to go
Update: Trying to push throughput even further
Update: We're at an architectural limit, going to temporarily disable some events (incl. typing indicators and user updates) to help ease congestion while an actual fix is being put together
Feb 12, 17:03 UTC
Identified- I suspect we are hitting limits with our message pubsub, a solution is being put together.
Update: Scaled vertically for now.
Feb 11, 20:20 UTC
Monitoring- Production services are now scaled up.
There is a possibility of hitting further bottlenecks but we should be okay for a moment.
Will continue to monitor and improve the deployment pattern.
Feb 11, 19:26 UTC
Identified- Ordered more servers, waiting for fulfillment. Service is generally stable right now, but more load is expected either today or tomorrow peak hours.
Feb 11, 19:17 UTC
Identified- Single-node cluster deployment was successful, now scaling it up.
Feb 11, 18:30 UTC
Identified- Deployment is taking a little longer than expected, but I would expect this to take less than an hour to resolve.
Feb 11, 16:58 UTC
Identified- We are currently working on scaling up our services.