Network Incident Severity Metric for SLA Compliance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Network administrators face challenges in predicting and preventing Service Level Agreement (SLA) violations due to the difficulty in determining the impact of network incidents on service level objectives until after the SLA has been violated, making it hard to maintain network service levels.

Innovation Solution

A method and system that identify network incidents over specific measurement periods, determine remaining incidence tolerance limits, generate severity metric values based on aggregate impact characteristics, and select incidents for remediation to prevent SLA violations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If network administrators monitor network incidents to maintain service levels, then network reliability is improved, but the complexity of detecting and measuring incident impact increases

Engineering Contradiction:
Improvenetwork service levelVSAvoidincident impact
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The system performs preliminary actions by establishing service level objectives and incident tolerance limits before incidents occur. It proactively monitors and identifies incidents during the measurement period, calculating their cumulative impact in advance to predict potential SLA violations before they happen, rather than reacting after violations occur.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements feedback by continuously monitoring network incidents, calculating their aggregate impact against predefined tolerance limits, and providing information about remaining tolerance capacity. This feedback loop enables administrators to understand the current state of service level compliance and take corrective actions before SLAs are violated.

Inventive Principle:
Principle #23Feedback

2Reliability

If the system identifies and responds to network incidents proactively, then SLA violation prevention is improved, but the device complexity increases

Engineering Contradiction:
ImproveSLA complianceVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the measurement period into identifiable incident events, each with its own impact characteristics. It divides the overall SLA compliance assessment into individual incident evaluations, allowing systematic tracking of cumulative impact against tolerance limits without requiring complex monolithic analysis.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes parameters by defining specific measurable attributes for incidents (impact characteristics, tolerance limits, severity metrics) and transforming raw incident data into standardized parameters that can be systematically compared and aggregated. This parameterization simplifies the complexity of evaluating diverse incident types against SLA requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3154224B1Systems and methods for maintaining network service levels
Publication Date: 2019.08.21 GOOGLE LLC
  • EP3154224B1 patent drawingFigure 1
  • EP3154224B1 patent drawingFigure 2A~2B
  • EP3154224B1 patent drawingFigure 3A~3B

AI summary

Described are methods and system for maintaining network service levels. In general, the system identifies, using records of network incidents, a first plurality of network incidents occurring over a first portion of a measurement period and a second plurality of network incidents occurring over a subsequent second portion of the measurement period. The system then determines a plurality of remaining incidence tolerance limits based on an impact of the first and second pluralities of network incidents on corresponding sets of incidence tolerance limits for the measurement period, generates severity metric values for at least a subset of the second network incidents based on aggregate impact characteristics of one or more of the second plurality of network incidents weighted by remaining incidence tolerance limits associated with each of the second network incidents in the subset of the second network incidents, and selects one or more network incidents for remediation.