Network Incident Severity Metric for SLA Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Network administrators face challenges in predicting and preventing Service Level Agreement (SLA) violations due to the difficulty in determining the impact of network incidents on service level objectives until after the SLA has been violated, making it hard to maintain network service levels.
Innovation Solution
A method and system that identify network incidents over specific measurement periods, determine remaining incidence tolerance limits, generate severity metric values based on aggregate impact characteristics, and select incidents for remediation to prevent SLA violations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If network administrators monitor network incidents to maintain service levels, then network reliability is improved, but the complexity of detecting and measuring incident impact increases
Solution Approach 1:
The system performs preliminary actions by establishing service level objectives and incident tolerance limits before incidents occur. It proactively monitors and identifies incidents during the measurement period, calculating their cumulative impact in advance to predict potential SLA violations before they happen, rather than reacting after violations occur.
Solution Approach 2:
The system implements feedback by continuously monitoring network incidents, calculating their aggregate impact against predefined tolerance limits, and providing information about remaining tolerance capacity. This feedback loop enables administrators to understand the current state of service level compliance and take corrective actions before SLAs are violated.
2Reliability
If the system identifies and responds to network incidents proactively, then SLA violation prevention is improved, but the device complexity increases
Solution Approach 1:
The system segments the measurement period into identifiable incident events, each with its own impact characteristics. It divides the overall SLA compliance assessment into individual incident evaluations, allowing systematic tracking of cumulative impact against tolerance limits without requiring complex monolithic analysis.
Solution Approach 2:
The system changes parameters by defining specific measurable attributes for incidents (impact characteristics, tolerance limits, severity metrics) and transforming raw incident data into standardized parameters that can be systematically compared and aggregated. This parameterization simplifies the complexity of evaluating diverse incident types against SLA requirements.
Data Source
Figure 1
Figure 2A~2B
Figure 3A~3B
AI summary
Described are methods and system for maintaining network service levels. In general, the system identifies, using records of network incidents, a first plurality of network incidents occurring over a first portion of a measurement period and a second plurality of network incidents occurring over a subsequent second portion of the measurement period. The system then determines a plurality of remaining incidence tolerance limits based on an impact of the first and second pluralities of network incidents on corresponding sets of incidence tolerance limits for the measurement period, generates severity metric values for at least a subset of the second network incidents based on aggregate impact characteristics of one or more of the second plurality of network incidents weighted by remaining incidence tolerance limits associated with each of the second network incidents in the subset of the second network incidents, and selects one or more network incidents for remediation.