Originally published in 2024. This is a historical incident account.
In DevOps, an apparently stable environment can change in hours. This account describes an outage that lasted three days and the decision to accelerate a planned move to AWS with Terraform, Ansible, and a focused recovery effort.
The incident
Users could not connect, monitoring alerts were firing, and SSH access was unavailable. After investigation, the hosting provider’s unannounced regional migration had left root and backup disks incorrectly configured. The server existed but was unusable.
With no clear repair timeline, waiting was no longer the best option. The AWS migration had already been planned, so the team used that preparation to begin rebuilding.
Rebuilding on AWS
Terraform defined the immediate building blocks:
- EC2 instances for Dockerized backend and frontend services
- RDS for the database
- S3 for backups and assets
- IAM roles to constrain permissions
Ansible then automated the server setup: Docker installation, environment configuration, and basic security hardening. Repeatable playbooks made it possible to rebuild an instance quickly when it needed adjustment.
Data recovery and validation
The database recovery was the most difficult step. The available backups were incomplete and poorly organized, and legacy character-set and collation issues required debugging before the recovered data could be loaded into RDS.
Before switching DNS, the checks covered service connectivity, configured variables and secrets, database connections, and user-facing behavior. By the end of the third day, the environment was operating on AWS.
Lessons
The incident reinforced a few durable practices: document infrastructure and recovery procedures, test disaster recovery rather than only keeping backups, evaluate provider support and reliability, and begin migrations before an incident forces the schedule.