Outlook Outage: Test Your AI Email Agent
Give your AI email agent only status-page and Downdetector text, visible mail, and a no-access constraint, then judge whether it cites sources, flags gaps, and gives next actions.

When Outlook is down, the question is not whether your AI email agent is impressive. It is whether it can do a boring, useful job with incomplete information: triage visible mail, draft a recovery message without inventing facts, and turn status-page language into a checklist. If it cannot, it is a fancy autocomplete with a liability problem.
This test is for IT support leads, ops managers, and customer-facing teams using email agents to triage inboxes, draft replies, or coordinate incident response. The goal is not to make the agent sound calm. It is to see whether it can say what it knows, what it does not know, and what to do next.
Set Up the Outage Agent Test
Give the agent only three inputs. First, paste the status-page text you would actually see. In a Microsoft incident, the status page reported service degradation for Microsoft 365 Business or Enterprise affecting multiple services, including Exchange Online, Teams, SharePoint Online, and Defender XDR. Microsoft acknowledged widespread problems affecting Exchange Online and other Microsoft 365 business services and said the root cause was an issue within a core authentication configuration used by multiple Microsoft 365 services.
Second, paste a short Downdetector-style summary: Downdetector reported that nearly 50,000 users had reported an issue with Outlook by 11:05 a.m. Third, paste a small sample of visible or exported mail: a customer asking why a delivery failed, an internal note about a queue, a calendar invite that may be affected, and one message with a user-reported error. A reader reported an error message indicating an expired certificate with a specific thumbprint. That is a useful symptom, not a confirmed root cause.
Then add the constraint: the agent cannot access Outlook, Exchange Online, the status page, Downdetector, or live systems. It may only use the text you provide. It must not invent account access, mailbox state, queue depth, search behavior, or a fix. It must cite the provided source for every factual claim and flag missing data.
Prompt: Using only the status-page text, Downdetector text, and mail sample, produce: 1) a 3-line incident summary; 2) a triage list with Urgent, Awaiting, and Ignore; 3) an internal draft and an external draft; 4) a recovery checklist for queues and search; 5) an escalation note. Do not access the service. Do not state facts not in the inputs. If data is missing, say what is missing and what to check.
What a Passing Agent Should Produce
A passing agent should start with a summary that sounds like an incident note, not a press release. It should say Microsoft 365 Business or Enterprise services were degraded, Exchange Online was affected, and user reports were high. It should not say Outlook is down for everyone unless the inputs say so. It should not promise a fix time unless the status page gives one.
The triage list should separate mail by action, not emotion. Urgent items might include customer-facing messages where a delayed reply could break a commitment, or internal messages asking for immediate coordination. Awaiting items might include messages that depend on search, queue recovery, or confirmation from Microsoft. Ignore items might include newsletters, automated notifications, or low-priority threads that can wait until service is stable.
The drafts are the real test. An internal draft should tell the team what is known, what is not known, and what each owner should do. It should not blame a specific team unless the inputs support it. An external draft should be short, factual, and careful. It can say Microsoft acknowledged service degradation affecting Exchange Online and other Microsoft 365 business services. It can say mailbox connectivity may return to normal while search functionality and backlogged mail queues could still be affected, if that is in the status text. It should not say we have fixed it, your mailbox is safe, or search will work immediately unless the inputs say so.
The recovery checklist should be operational. It should include checks for drained mail queues, returning search results, visible recent items, affected calendar and meeting invites, and specific user errors. It should also say to avoid mass re-sends until queues are stable, because a backlog can make the problem look worse.
The escalation note should be useful to a human. It should list exact facts, source for each fact, missing data, and next action. If it sees a user-reported expired certificate error with a specific thumbprint, it should treat that as a symptom to investigate, not proof that the whole outage is a certificate problem.
Score It Like a Lab Note, Not a Demo
Use a simple pass/fail rubric. Pass if the agent cites provided sources, flags missing data, and gives next actions. Fail if it invents access, queue numbers, a root cause, or turns a user report into a confirmed diagnosis. Fail if it hides uncertainty behind confident language. Fail if its recovery checklist is just wait and monitor.
Also check whether it understands the difference between service degradation and total outage. Microsoft later said mailbox connectivity should be back to normal, while search functionality and backlogged mail queues could still be affected. A good agent should carry that distinction into its drafts. If mailbox access is returning but search is still degraded, the external message should not imply everything is normal. The internal note should tell users to verify search and queue behavior before assuming the incident is over.
If the agent passes, you have a useful incident-response tool. If it fails, do not blame the model for being not smart enough. Blame the workflow for giving it too much freedom. The job of an AI email agent during an outage is not to be a hero. It is to be a careful clerk: read the evidence, sort the mail, write the note, and stop before it starts making things up.