← Blog

Originally published in 2024. This is a historical incident account.

In DevOps, an apparently stable environment can change in hours. This account describes an outage that lasted three days and the decision to accelerate a planned move to AWS with Terraform, Ansible, and a focused recovery effort.

The incident

Users could not connect, monitoring alerts were firing, and SSH access was unavailable. After investigation, the hosting provider’s unannounced regional migration had left root and backup disks incorrectly configured. The server existed but was unusable.

With no clear repair timeline, waiting was no longer the best option. The AWS migration had already been planned, so the team used that preparation to begin rebuilding.

Rebuilding on AWS

Terraform defined the immediate building blocks:

Ansible then automated the server setup: Docker installation, environment configuration, and basic security hardening. Repeatable playbooks made it possible to rebuild an instance quickly when it needed adjustment.

Data recovery and validation

The database recovery was the most difficult step. The available backups were incomplete and poorly organized, and legacy character-set and collation issues required debugging before the recovered data could be loaded into RDS.

Before switching DNS, the checks covered service connectivity, configured variables and secrets, database connections, and user-facing behavior. By the end of the third day, the environment was operating on AWS.

Lessons

The incident reinforced a few durable practices: document infrastructure and recovery procedures, test disaster recovery rather than only keeping backups, evaluate provider support and reliability, and begin migrations before an incident forces the schedule.

Publicação original