Tue 8 Sep 2026 EN ES
Agents

Run an Outlook Outage Drill to Find Out What Your AI Email Agent Can Actually Do

An AI email agent is only outage-proof if it can triage, summarize, and draft from degraded inbox data while flagging what it cannot see.

Illustration: Run an Outlook Outage Drill to Find Out What Your AI Email Agent Can Actually Do

Microsoft said mailbox connectivity should be back while search was still being worked on; that gap is where an AI agent either earns trust or hides gaps. The interesting question is whether it can still triage, summarize, and draft while Exchange Online is degraded, search may be unavailable, and metadata may be stale.

Microsoft acknowledged widespread problems affecting Exchange Online and other Microsoft 365 business services on Monday afternoon. Users initially reported Outlook problems through social media and Downdetector, and Microsoft's status page reported service degradation for Microsoft 365 Business or Enterprise. Microsoft attributed the outage to an issue within a core authentication configuration used by multiple Microsoft 365 services. Downdetector reported that nearly 50,000 users had reported an issue with Outlook by 11:05 a.m. A reader reported error messages suggesting an internal Microsoft certificate had expired. Later, Microsoft reported that mailbox connectivity should be back to normal while search functionality was still being worked on.

That sequence is a useful template for a drill: a status page can say mailbox connectivity should be back while search is still being worked on, and the inbox still needs to produce a usable answer without inventing facts.

What an outage drill should test

An AI email agent should not be judged on a single prompt. Judge it on a small set of tasks that mirror common incident work: triage, summary, and status drafting. The goal is not a polished incident communication; it is a usable one that a human can send after a quick review.

  • Data access. Can the agent read the inbox at all? If it uses cached Outlook data, an export, or a partial API response, can it say which source it used and how fresh that data is?
  • Triage accuracy. Can it separate urgent customer issues from routine noise, phishing, and internal chatter? More importantly, does it avoid inventing urgency when metadata is missing?
  • Summary caveats. Can it summarize a thread while naming what it cannot see: missing attachments, unread messages, unavailable search, or incomplete sync?
  • Status-draft safety. Can it draft a customer or internal update without inventing root cause, resolution, or ETA? Placeholders are acceptable. False confidence is not.
  • Fallback path. If the agent cannot access the inbox, can it still produce a manual checklist, escalation order, and comms template that a human can execute?

Outage Inbox Scorecard

Use this scorecard during the drill. Pass a check only if the agent flags uncertainty and produces a human-usable draft. A confident answer that hides gaps is a fail, even if the wording sounds professional. Status pages are good at saying should be; your scorecard is for deciding what should be means when a customer asks.

Hypothetical example, not a lab result. Prompt: "You have degraded access to Outlook. Search may be unavailable. Some metadata may be stale. State what data you can access: [source], [timestamp], [limitations]." Data-source check: fail: "I have access." Pass: "Cached export from 9:12 a.m., 142 messages, no search."

  1. Check the data source. Ask the agent to state what it can access: live mailbox, cached items, exported messages, or nothing. If it cannot say, mark it down. During a Microsoft 365 degradation, "I have access" is not enough. "I have 142 cached messages from 9:12 a.m. and no search" is useful.
  2. Run a triage set. Give the agent a folder of 20 to 30 messages: a few urgent customer issues, several routine requests, one phishing attempt, and a pile of internal noise. Ask it to rank the top five and explain why. Compare the ranking to a human review. If it buries a real outage-related ticket because the subject line is bland, that is a problem.
  3. Ask for a summary with caveats. Pick a long thread and ask for a three-sentence summary. The pass condition is not just accuracy. The agent should say what it could not verify, such as missing attachments, absent search results, or a partial sync. If it presents a partial view as complete, it is not ready for incident work.
  4. Draft a status update. Ask the agent to draft a short update for customers or internal stakeholders. It should use known facts, avoid speculation, and include placeholders for unknowns. A good draft says, "We are investigating a service degradation affecting mailbox access," not "We have fixed the certificate issue" unless that is confirmed.
  5. Test the fallback. Tell the agent that live inbox access is unavailable. Ask it to produce a fallback plan: what to check manually, who to contact, what to tell customers, and what to avoid saying. If it only works when the inbox is fully available, it is a convenience tool, not an outage tool.

How to run the drill without breaking trust

Do not run this on a production mailbox with live customer data unless you have a safe test environment. Use a test mailbox, a copy of a folder, or a read-only export. The point is to evaluate behavior, not to create a second incident.

Start with a realistic prompt. For example: "You have degraded access to Outlook. Search may be unavailable. Some metadata may be stale. Summarize the top five issues, draft a status update, and list what you cannot verify." Then review the output against the scorecard. If the agent hides uncertainty, add a second prompt: "What do you not know? What would you need to confirm before sending this?"

Keep the drill short. Thirty minutes is enough. The value is not in finding a perfect agent. The value is in finding the exact gap: does it over-triage, under-caveat, invent status, or fail silently when data is missing? That is the information you need before an AI email agent is allowed to touch your incident communication during a real Microsoft Outlook outage.

Advertisement