Network Failure Troubleshooting via Historical Alarm Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data centers face challenges in efficiently resolving network failures due to the complexity and heterogeneity of network infrastructure, leading to prolonged troubleshooting times and potential service downtime, especially as they scale to include thousands of devices.

Innovation Solution

A system that identifies potential troubleshooting options and prioritizes network failures by mapping alarm symptoms to historical data, providing operators with a list of recommended actions and their probabilities of success, and automatically resolves issues when possible, using a data-driven approach to improve accuracy over time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If commodity network infrastructure devices are used to decrease capital costs, then device cost is reduced, but device reliability deteriorates

Engineering Contradiction:
Improvedevice costVSAvoiddevice reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system implements feedback by continuously monitoring network device performance and automatically adjusting configurations based on observed behavior and historical data, compensating for the lower inherent reliability of commodity devices through dynamic adaptive control

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables self-service by allowing network devices to automatically diagnose and resolve their own issues using embedded sensors, historical performance data, and machine learning models, reducing dependency on operator intervention and compensating for device reliability limitations

Inventive Principle:
Principle #25Self-service

2Productivity

If data centers scale to include thousands of devices, then service capacity is increased, but troubleshooting complexity increases

Engineering Contradiction:
Improveservice capacityVSAvoidtroubleshooting complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system applies segmentation by dividing the large-scale network into smaller monitored zones and grouping devices by type, function, and performance characteristics, allowing operators to isolate and troubleshoot specific segments rather than managing the entire complex system as a whole

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary automated monitoring and analysis platform that sits between the scaled network devices and human operators, translating complex device data into simplified alerts and recommendations, thereby reducing the perceived complexity for operators while maintaining the ability to manage thousands of devices

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If operators use trial-and-error approach to address alarms, then flexibility in problem-solving is maintained, but troubleshooting time increases

Engineering Contradiction:
Improveproblem-solving flexibilityVSAvoidtroubleshooting time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system implements preliminary action by pre-calculating and storing optimal troubleshooting sequences based on historical data and device characteristics, so when an alarm occurs, operators receive ready-formulated action plans rather than needing to derive solutions in real-time, significantly reducing troubleshooting time while maintaining flexibility through adaptive refinement

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system substitutes the manual trial-and-error mechanical process with an automated information processing system that uses machine learning models and historical data to generate optimized troubleshooting recommendations, replacing the iterative physical trial-and-error approach with a data-driven decision support system

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11057266B2Identifying troubleshooting options for resolving network failures
Publication Date: 2021.07.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11057266B2 patent drawing
  • US11057266B2 patent drawing
  • US11057266B2 patent drawing

AI summary

Described herein are various technologies pertaining to providing assistance to an operator in a data center with respect to failures in the data center. An alarm is received, and a failing device is identified based upon content of the alarm. Failure conditions of the alarm are mapped to a failure symptom that may be exhibited by the failing device, and troubleshooting options previously employed to mitigate the failure symptom are retrieved from historical data. Labels are respectively assigned to the troubleshooting options, where a label is indicative of a probability that a troubleshooting option to which the label has been assigned will mitigate the failure symptom.