NOC OPERATIONS

SLA Availability and Downtime: Read the Fine Print

How availability percentages translate into minutes, what usually counts as downtime, and how to avoid misleading SLA reports.

Network operations dashboard illustration

A percentage hides the operational impact

Availability is usually calculated as available time divided by total measured time. The difference between 99.9% and 99.99% looks small on paper, but over a 30-day month it changes the downtime budget from 43.2 minutes to about 4.32 minutes.

That is why leadership and customers often understand an SLA better when the percentage is shown alongside the equivalent downtime.

Contract language decides what counts

The arithmetic is easy; the scope is not. Many agreements exclude planned maintenance, customer-caused outages, power failures at customer premises, force majeure or events outside the provider’s control. Other agreements define availability at a particular handoff or require a minimum duration before an event is counted.

For defensible reporting, tie every exclusion to the actual contract language. A spreadsheet that silently removes incidents without an auditable reason is hard to defend during a dispute.

Measure customer impact, not ticket lifecycle

An incident ticket may open before service is affected or remain open after traffic is restored. If the SLA defines availability by service impact, use the impact interval rather than the ticket’s administrative open/close duration.

This distinction matters when teams keep tickets open for monitoring, root-cause work or customer communication. Operational process time should not automatically become customer downtime.

Redundancy changes risk, not the formula

A redundant design can reduce the probability or duration of outages, but it does not automatically guarantee an availability number. Shared fiber routes, common power, software defects and maintenance collisions can defeat nominal redundancy.

Track redundancy risks as part of availability management. A service can be “up” today while operating on a single remaining path that leaves the next failure customer-affecting.

Build reporting from evidence

Keep timestamps, monitoring events, incident notes, maintenance records and customer-impact evidence together. When an SLA report is generated later, the reviewer should be able to trace each downtime entry back to a source.

Automation can reduce the manual hunt, but good real-time ticket notes are still essential. The best monthly report starts with accurate documentation while the incident is happening.

← Back to all guidesBrowse network tools →
Use these guides as engineering references, not as a substitute for your network design standards, current vendor documentation or production change review.