Incident Details
Post-maintenance issues with sa-east-br monitoring region (Regions) Partial Outage Resolved
Regions • Started UTC • resolved UTC • Duration
- Investigating
Our
sa-east-brregion is having some issues after the scheduled maintenance event.We're investigating the issue.
- Identified
We've identified the issue and are working on a fix.
Impact: Only the
sa-east-brregion is affected
Current status: Identified.
Next update: 15min - Monitoring
A fix has been implemented and we're monitoring the results.
sa-east-brregion is again working properly.
We're monitoring the region for any issues.
- Resolved Latest
This incident has been resolved.
Postmortem coming up.
📘 Postmortem
Summary
Following scheduled hosting-provider maintenance on our São Paulo (sa-east-br) monitoring node, the node rebooted but failed to reconnect to the main platform. It showed as Offline on the public Monitoring Regions page for about an hour. No customer monitor was ever incorrectly reported down — checks continued to be confirmed by our other regions throughout.
Impact
No customer-facing incident: StatusPage.me requires agreement across a majority of independent monitoring regions before declaring a monitor down, so one region being unreachable did not affect any customer's status page.
Visible impact was limited to our public Monitoring Regions page, which showed sa-east-br as Offline for the duration of the outage, plus a residual display lag after the fix (see Root Cause / Lessons Learned).
Root Cause
The maintenance reboot moved the sa-east-br node onto different underlying host hardware. Network connectivity came back up and looked healthy, but DNS resolution needed for the node to reach our main platform was silently broken as a side effect of the provider's infrastructure change. Because basic reachability looked fine, the issue wasn't obvious at a glance and required investigation to isolate.
Remediation
- Identified the DNS misconfiguration on the sa-east-br node and corrected it; confirmed normal reporting resumed immediately.
- Hardening monitoring nodes to fall back to alternate DNS resolution automatically, so a single provider-side DNS fault can't take a node offline again.
- Adding direct alerting for when a monitoring region stops reporting in, instead of relying on visual inspection of the public page.
- Reducing the refresh interval on the public Monitoring Regions page so recovery is reflected sooner (currently hourly).
Timeline
- Hosting provider completes scheduled maintenance on sa-east-br node; node reboots
- Node comes back up on new host hardware; DNS resolution silently broken, node unable to reach main platform
- sa-east-br shows Offline on public Monitoring Regions page
- Team begins investigation
- DNS misconfiguration identified and corrected
- Node reconnects; region resumes normal reporting
- Public Monitoring Regions page reflects recovery (delayed by hourly refresh cycle)
Published 21.07.2026. 11:26:26