Network Failure Troubleshooting via Historical Alarm Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data centers face challenges in efficiently resolving network failures due to the complexity and heterogeneity of network infrastructure, leading to prolonged troubleshooting times and potential service downtime, especially as they scale to include thousands of devices.
Innovation Solution
A system that identifies potential troubleshooting options and prioritizes network failures by mapping alarm symptoms to historical data, providing operators with a list of recommended actions and their probabilities of success, and automatically resolves issues when possible, using a data-driven approach to improve accuracy over time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If commodity network infrastructure devices are used to decrease capital costs, then device cost is reduced, but device reliability deteriorates
Solution Approach 1:
The system implements feedback by continuously monitoring network device performance and automatically adjusting configurations based on observed behavior and historical data, compensating for the lower inherent reliability of commodity devices through dynamic adaptive control
Solution Approach 2:
The system enables self-service by allowing network devices to automatically diagnose and resolve their own issues using embedded sensors, historical performance data, and machine learning models, reducing dependency on operator intervention and compensating for device reliability limitations
2Productivity
If data centers scale to include thousands of devices, then service capacity is increased, but troubleshooting complexity increases
Solution Approach 1:
The system applies segmentation by dividing the large-scale network into smaller monitored zones and grouping devices by type, function, and performance characteristics, allowing operators to isolate and troubleshoot specific segments rather than managing the entire complex system as a whole
Solution Approach 2:
The system introduces an intermediary automated monitoring and analysis platform that sits between the scaled network devices and human operators, translating complex device data into simplified alerts and recommendations, thereby reducing the perceived complexity for operators while maintaining the ability to manage thousands of devices
3Ease of operation
If operators use trial-and-error approach to address alarms, then flexibility in problem-solving is maintained, but troubleshooting time increases
Solution Approach 1:
The system implements preliminary action by pre-calculating and storing optimal troubleshooting sequences based on historical data and device characteristics, so when an alarm occurs, operators receive ready-formulated action plans rather than needing to derive solutions in real-time, significantly reducing troubleshooting time while maintaining flexibility through adaptive refinement
Solution Approach 2:
The system substitutes the manual trial-and-error mechanical process with an automated information processing system that uses machine learning models and historical data to generate optimized troubleshooting recommendations, replacing the iterative physical trial-and-error approach with a data-driven decision support system
Data Source
AI summary
Described herein are various technologies pertaining to providing assistance to an operator in a data center with respect to failures in the data center. An alarm is received, and a failing device is identified based upon content of the alarm. Failure conditions of the alarm are mapped to a failure symptom that may be exhibited by the failing device, and troubleshooting options previously employed to mitigate the failure symptom are retrieved from historical data. Labels are respectively assigned to the troubleshooting options, where a label is indicative of a probability that a troubleshooting option to which the label has been assigned will mitigate the failure symptom.


