Use This 5-Point Outlook-Down Test to See If Your AI Email Agent Can Keep Email Moving
If your AI email agent only works when the cloud is calm, it is a toy, not an outage workflow.

Why an outage is the real acceptance test
Lab note: the recent multi-hour Microsoft Outlook and Exchange Online outage is the moment an AI email agent stops being a demo and becomes an acceptance test. The observable symptoms were email delays, failures, and authentication issues arriving at once. Most agents are judged in the same lazy way: send a few sample messages, read the drafts, nod at the summaries. That is not evaluation. An agent that cannot draft, triage, or route mail through degraded paths is mostly a fancy autocomplete with a badge.
Vendor language will tell you the agent is intelligent, adaptive, and always on. Those words are not evidence. Evidence is behavior under constraint: what it drafts when the mailbox is slow, what it routes when the destination is wrong, and what it tells a human when it is unsure. If the vendor cannot show degraded-state behavior, ask for the failure modes, not the feature list.
The 5-point mail-down test
Run this as a no-cause drill. Do not announce it as a game. Pick a normal business day, choose a small group of real inboxes, and ask the agent to keep email moving while the primary path is degraded. The goal is not to prove the vendor wrong. The goal is to learn what your team can actually do when the usual assumptions break.
- Drafting under degraded state. Pass if the agent drafts from a queue of incoming mail while the mailbox is slow, partially unavailable, or only reachable through a fallback, and it marks what it could not verify. Fail if it pretends a failed lookup was a success. A useful agent should say I could not confirm the recipient, thread, or attachment when that is true. If it cannot say that, it is not ready for real work.
- Triage without full search. Pass if the agent can sort urgent, routine, and noise based on sender, subject, keywords, and recent context without a full mailbox search. Fail if it freezes when search is unavailable. Search is often the first thing to go, so an agent that depends on full search to decide what matters will fail exactly when your team needs it most.
- Fallback routing. Pass if the agent knows where mail goes when the normal destination is unavailable: a shared queue, a backup mailbox, a support alias, or a human inbox, and it records the reason for the reroute. Fail if it reroutes without a reason. A reroute without a reason is just a mystery later.
- Auth and dependency map. Pass if, before the drill, you document what the agent needs to sign in, what tokens it uses, and which services it calls, and during the drill the agent surfaces silent failures plainly. Fail if a draft looks fine but cannot send, a summary cannot open attachments, or a routing rule cannot verify identity, and nobody notices. The agent should make those failures visible, not hide them.
- Human escalation. Pass if the agent knows when to stop and hand the message to a person: ambiguous requests, high-stakes senders, legal or security language, and anything that cannot be verified. Fail if it keeps acting when the safe move is to stop. Escalation is not a bug. It is the feature that keeps the workflow from becoming a liability.
How to run the drill without turning into a war room
Keep the scope small. Use a few inboxes, a few hours, and a clear stop condition. Capture what the agent did, what it could not do, and what a human had to fix. Do not judge the vendor by one bad draft. Judge the system by whether email kept moving and whether the team could explain why.
After the drill, write a one-page result. List the failures, the fallbacks that worked, and the next change you will make. If the agent cannot draft under degraded state, fix that before adding more features. If it can draft but cannot route, fix routing. If it can route but cannot escalate, fix escalation. The best outcome is boring: the mail client is down, the agent is doing its job, and nobody has to guess what happened.
Return to the outage evidence before you close the finding. The evidence is not a single summary; it is the sequence of what came back first. If the mailbox is reachable again while search is still degraded, the drill should separate restored send/receive from restored search and report the result as a one-page finding. An AI email agent should be measured like any other operational tool: by the work it completes when conditions are bad, not by the demo it gives when everything is green. If your current setup cannot pass a simple mail-down test, that is not a failure of imagination. It is a finding. Now you know what to fix before the next outage finds you.