Why Did ChatGPT, Claude & Microsoft Services Stop Working?
News Desk
A series of technology disruptions temporarily affected some of the world’s most widely used digital services this week, leaving users struggling to access Microsoft 365 tools as well as major artificial intelligence platforms including ChatGPT, Claude and Grok.
The incidents highlighted how even highly distributed cloud-based services can experience sudden disruptions, and how a problem in one critical layer, such as authentication or infrastructure, can ripple across multiple products.
Microsoft outage traced to authentication configuration
Microsoft’s disruption began on August 31 and affected several Microsoft 365 services, with Exchange Online among the most prominent.
Microsoft identified an issue involving a core authentication configuration used by multiple Microsoft 365 services. The disruption affected functions including Outlook access, email attachments, mail flow, search and other Exchange Online operations. Other services reported as affected included Teams, OneDrive, SharePoint, Microsoft 365 Copilot, Microsoft Purview, Universal Print and Defender XDR.
Microsoft’s recovery process involved deploying a targeted fix across affected infrastructure and monitoring service performance. The company reported progressively improving availability, although some users could continue to experience intermittent problems or delays while remaining systems recovered.
One important consequence was the effect on email delivery. Microsoft warned that organizations could experience delays as previously queued messages worked their way through the system.
The incident demonstrated the importance of authentication infrastructure in modern cloud services. Although users may interact with individual products such as Outlook or Teams, many of those products depend on shared backend systems for identity verification and access.
ChatGPT, Claude and Grok face separate disruptions
The Microsoft incident came amid a separate wave of outages affecting major AI platforms on September 3.
OpenAI’s status page recorded elevated errors across ChatGPT and Codex before the company confirmed that the problem had been resolved. OpenAI also reported another ChatGPT incident earlier that day involving the web interface, where some users experienced blank responses because misconfigured DDoS protection rules blocked certain files required by the website.
Anthropic also reported elevated errors affecting several Claude models. According to the company’s status page, the incident affected Mythos 5.1, Fable 5.1 and Opus 5, while Opus 4.8 and other models also experienced elevated error rates during parts of the disruption. Anthropic subsequently deployed a fix and reported that the impact had ended.
Grok, operated by xAI, also experienced a disruption around the same period. Reports indicated that the service encountered problems across its web and mobile platforms, with the company later attributing the incident to an issue at its Memphis computing facility.
Were the outages connected?
The simultaneous timing of the AI disruptions naturally raised questions about whether a common technical problem was responsible.
However, there is currently no confirmed evidence that the ChatGPT, Claude and Grok incidents were caused by the Microsoft outage or by a single shared infrastructure failure.
Reporting on the incidents indicates that the companies identified different technical problems. OpenAI’s September 3 incidents involved its own systems, Anthropic described an infrastructure issue affecting Claude, while xAI linked Grok’s disruption to its Memphis computing infrastructure.
That distinction is important because simultaneous outages do not necessarily mean that the affected companies share the same underlying failure.
A reminder of the hidden infrastructure behind everyday apps
For ordinary users, an outage is often visible simply as an error message, a failed login or an application that refuses to load.
Behind those screens, however, modern digital platforms rely on layers of authentication systems, application programming interfaces, cloud infrastructure, content-delivery networks, databases and computing facilities.
The Microsoft incident illustrated how an authentication configuration can affect numerous services at once. The AI outages, meanwhile, showed that even individual AI providers can encounter failures in different parts of their infrastructure.
The disruptions also underline a broader reality of the AI era: access to increasingly important digital tools depends on complex systems that users rarely see.
As Microsoft and the AI companies restored their services, most affected platforms returned to normal operation. But the incidents served as a reminder that reliability is not simply about keeping an application online, it also depends on the health of the many interconnected systems operating behind it.
For businesses, researchers, students and millions of everyday users who increasingly depend on cloud and AI services, even a short outage can quickly become a significant disruption.
