Resolved -
Working with our network hardware vendor, we have deployed a global change across all compute sites to improve the reliability of the connection between our Internet/backbone edge routers and our data center core networks.
This change addresses the blackhole condition that caused the overnight outage. With it in place, a malformed route entering the fabric will no longer prevent traffic forwarding - sites will continue to forward via their local edge routers rather than losing reachability. We considered the previous behavior a critical flaw and have prioritized this fix accordingly. All sites, including the unaffected data centers, had this configuration change completed successfully earlier today.
We continue to investigate the underlying condition that caused the default route from MIA1 to be advertised/readvertised with incorrect attributes, and will publish a full RFO once that analysis is complete.
Please note: a separate, unrelated incident affecting our global traffic routing systems remains open. That work involves a distinct vendor defect and route scaling improvements, and is not connected to the outage described here.
Aug 12, 18:29 UTC
Monitoring -
(This status is split off from a previous issue that was originally believed to have been related.)
Customers at LON1, AMS1, AMS2, AMS3, DUB1, DUB2, FRA2, SGP1, SGP2, TYO1, TYO2 and TYO3 experienced loss of reachability to Internet destinations, as well as to internal Teraswitch backbone destinations between affected sites. Other North American sites were not affected.
MIA1 (Miami, FL) was intentionally removed from the backbone as part of our response. MIA1 remains reachable and in service, but is currently operating without full backbone connectivity, traffic to and from other Teraswitch sites is routed over the public Internet rather than our private backbone. Customers relying on MIA1 for private backbone or inter-site connectivity should expect changed latency and path characteristics until MIA1 is reintegrated.
Cause: Teraswitch uses a default route (0.0.0.0/0) internally to signal that an edge router is able to forward traffic to the Internet. Each site normally prefers the default originated by its own local edge routers, with route attributes distinguishing a local origination from a remote one.
A default route originated at MIA1 was propagated with its metric and communities stripped and an AS-path containing only our own ASN. At this time, our understanding is that this route never existed within MIA1's routers as a valid route. A route reflector at our AMS2 site propagated this altered route into our EU and APAC markets. Receiving edge routers interpreted it as locally originated and preferred it over their own valid local default. Those routers then advertised the route to the downstream data center core, which rejected it as invalid. With no acceptable default present, the affected site fabrics stopped forwarding traffic to their own edge routers, resulting in the loss of reachability observed.
Resolution: Engineers identified the malformed route within 10 minutes of onset and removed MIA1 from the backbone to halt further propagation. Affected sites reconverged on their local default routes and service was restored at 04:16:15 UTC.
Next steps: The underlying defect, whether in the MIA1 edge routers or in the route reflector, has not yet been identified. We are working to reproduce the condition and have engaged our vendor. In the interim we are implementing policy to enforce minimal required attributes on internally originated default routes so that a malformed advertisement cannot be preferred over a valid local one. MIA1 will remain off the backbone until the defect is understood and mitigations are verified; we will post an update when it is reintegrated.
A full RFO will follow.
Aug 12, 05:18 UTC