Business internet outages are operational events, not just IT tickets. Payroll, point-of-sale, voice over IP, cloud applications, and site-to-site VPNs often ride the same path — so the first minutes matter for both restoration and communication. A playbook does not prevent outages; it reduces confusion about who does what, what to tell the provider, and how to capture the record you will need for SLA review and future sourcing decisions.
Before an outage: design and documentation
Response quality depends on preparation. Maintain a single-page service sheet per location: primary circuit type and ID, backup path if any, provider support numbers, account numbers, escalation contacts, and which applications are business-critical on that link. Store it where on-call staff can reach it without digging through email.
Failover and backup paths
Backup connectivity only helps if it fails over cleanly. That means tested routing — automatic or runbook-driven — and clarity about which apps tolerate asymmetric paths and which do not. A secondary DSL or LTE link may keep email flowing while large file transfers or real-time voice struggle; knowing that in advance sets honest expectations during an incident. For design options, see the backup internet guide.
Roles and communication
Name who opens provider tickets, who approves failover, and who communicates to staff and customers. Ambiguity during an outage produces duplicate tickets, conflicting messages, and changes nobody documented. A simple rule helps: one incident owner, one external voice, one ticket thread per provider.
During an outage: triage and escalation
Start by confirming the problem is upstream. Local power loss, a rebooting firewall, or a misconfigured DHCP scope can look like a carrier outage. Once upstream failure is likely, open a provider ticket immediately — waiting to “see if it comes back” burns SLA clocks and delays dispatch.
Escalation should follow the path in your agreement: tier-one for logging, technical teams for diagnosis, account or service management when restoration windows slip or impact is severe. Each escalation is more effective with ticket number, circuit ID, start time, business impact statement, and what you have already tried locally.
Incident response sequence
- 01Confirm scope: one site, one circuit, or provider-wide — check provider status pages and internal monitoring.
- 02Rule out local failure: power, CPE, firewall, DNS, and Wi‑Fi before opening a carrier ticket.
- 03Open a provider ticket with account ID, circuit ID, service address, and exact start time.
- 04Notify internal stakeholders with expected update cadence and interim workarounds.
- 05Activate failover or backup path if designed — document switch time and any app impact.
- 06Log every update from the provider with timestamp and promised restoration window.
- 07Escalate through contract escalation paths when SLA thresholds or business impact warrant it.
- 08Communicate externally only through approved channels — avoid speculative ETAs.
- 09After restoration, verify applications, run speed/latency checks, and close the ticket with root cause.
- 10Schedule a post-incident review: what worked, what failed, and what to change in design or contract.
After restoration: review and procurement follow-up
Close the loop with a short post-incident review within a week. Capture root cause as the provider stated it, total downtime, whether failover behaved as expected, and whether SLA credits apply under your agreement. Compare the event against your contract commitments — response time, restoration time, and credit mechanics — without assuming credits apply until you read the clause.
Recurring outages at the same site are a signal to revisit architecture and sourcing: diverse paths, different last-mile technologies, or a provider change at the next renewal window. That review belongs in your renewal calendar, not in a panic switch the day service returns. The technology renewal guide covers how to time that conversation.
What to avoid
Multiple untracked tickets. One incident, one primary ticket — append updates rather than opening parallel cases that confuse provider teams.
Promising customer ETAs from tier-one guesses. Communicate what you know and when you will update next; avoid relaying unconfirmed restoration times.
Skipping the post-incident record. Without timestamps and root cause, SLA review and sourcing decisions lack evidence — and the same failure mode repeats.
Reviewed by the SwitchU procurement desk — last reviewed July 2026.