3DEXPERIENCE platform on the Cloud unavailability event, January 5th 2022
On January 5 2022, from 01:35PM until 04:07PM UTC, Cloud users (Public Cloud, Academic and Private Cloud), worldwide, were not able to access or use their 3DEXPERIENCE platforms.
This incident was immediately detected by the Cloud Operations team and managed until resolution with IaaS teams.
Summary
On January 5 2022, from 01:35PM until 04:07PM UTC, Cloud users (Public Cloud, Academic and Private Cloud), worldwide, were not able to access or use their 3DEXPERIENCE platforms.
This incident was immediately detected by the Cloud Operations team and managed until resolution with IaaS teams.
Symptom
- Unable to access the 3DEXPERIENCE platform (errorMessage: “Internal Error”)
- Unable to use Web Applications for users that were already connected
- Unable to save/open/search from Native Apps
Cause
Context:
Network communication between 3DEXPERIENCE platform services relies on DNS (Domain Name System) for names resolution (e.g. eu1-ds-iam.3dexperience.com).
Summary:
The DNS resolution service, provided by 3DS OUTSCALE, went down, leading to the 3DEXPERIENCE platforms becoming unavailable to users. This service is protected by a cluster of firewalls (working in an active / passive mode), relying on several hardware devices to ensure redundancy and to provide high availability.
Earlier in the day, one of these devices had a hardware failure and became unavailable. Thanks to the redundancy mechanism, DNS resolution service was still operational and this first event did not introduce any service interruption.
Part of our standard operating procedures, replacement of defective hardware was engaged on site by Datacenter operators: this procedure had no track record of failure. Despite the application of this procedure recommended by the hardware manufacturer, the new firewall integrated in the cluster unexpectedly became the primary member, and synchronized its default and blank configuration to the other ones.
Consequently, all firewalls became inoperative, preventing all network communications to reach the DNS servers and then causing all name resolution requests to fail. Our first analyses identified a software issue in the new firewall, integrated in the cluster.
Remediation
Our engineering team was sent on site to restore the configuration of all firewalls. After this action, DNS resolution service as well as redundancy were successfully restored.
Prevention
We are taking multiple steps with this hardware manufacturer (as well as with our other hardware manufacturers) in order to prevent future occurrence of such an issue, which had never happened before.
In closing
Finally, we sincerely apologize for the inconvenience this unprecedented event may have caused you. We know how critical the 3DEXPERIENCE platform is to our users and their businesses. We will make sure to learn from this event in order to maintain our customers’ trust and to continue improving on availability of our online services even further.