Automated Root Cause Analysis for Network Service Degradation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network management systems require manual expertise and complete data sets to perform Root Cause Analysis (RCA) for network service degradation, which is inefficient and often unavailable in multi-vendor, multi-layer networks, especially when incomplete data or lack of domain expertise is present.
Innovation Solution
The system automatically performs RCA using Performance Monitoring data, path alarms, service alarms, network topology, and configuration logs, even with incomplete data, by generating derived alarms and employing machine learning techniques to identify root causes without requiring domain expertise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual Root Cause Analysis is performed by domain experts using complete data sets, then diagnostic accuracy is improved, but operational complexity and time consumption increase
Solution Approach 1:
The system performs preliminary clustering of service alarms based on timing characteristics before root cause identification. By pre-grouping alarms that occur within specific time windows and identifying their common temporal patterns, the system prepares structured data that accelerates the subsequent root cause analysis, reducing the time experts need to manually correlate scattered alarm data while maintaining diagnostic accuracy
Solution Approach 2:
The patent introduces timing characteristics and clustered alarm groups as intermediary structures between raw alarm data and root cause identification. These intermediaries organize and pre-process the data, making it more accessible and interpretable for both automated systems and domain experts, thereby reducing the time required for analysis without sacrificing diagnostic precision
2Measurement precision
If manual Root Cause_analysis is performed by domain experts, then diagnostic accuracy is improved, but device complexity and expertise requirements increase
Solution Approach 1:
The system performs self-service by automatically clustering service alarms based on their timing characteristics and identifying potential root causes without requiring continuous human intervention. The automated clustering algorithm independently processes alarm data, groups related events, and presents organized results, reducing the burden on domain experts to manually perform data correlation while maintaining diagnostic accuracy
Solution Approach 2:
The patent replaces the mechanical process of manual expert analysis with an automated computational system that uses timing-based clustering algorithms. This substitution transforms the complex human expert process into a systematic automated procedure that can handle large volumes of alarm data, reducing the need for extensive domain expertise while preserving diagnostic capabilities
3Reliability
If complete data sets are required for Root Cause_analysis, then diagnostic reliability is improved, but adaptability to incomplete data scenarios worsens
Solution Approach 1:
The system applies partial action by performing clustering analysis on the subset of alarm data that exhibits timing correlations, rather than requiring complete analysis of all available data. By focusing computational resources on identifying and clustering the most relevant temporally-related alarms, the system achieves reliable root cause identification even when some data is missing or incomplete, improving adaptability without sacrificing diagnostic reliability
Data Source
AI summary
Systems and methods are provided for analyzing one or more root causes of service degradation events in a network or other environment. A method, according to one implementation, includes a step of monitoring a plurality of overlying services offered in an underlying infrastructure having a plurality of resources arranged with a specific topology. In response to detecting a negative impact on the overlying services during a predetermined time window and based on an understanding of the specific topology, the method further includes the step of identifying suspect components from the plurality of resources in the underlying infrastructure. The method also includes the step of obtaining status information with respect to the suspect components to determine a root cause of the negative impact on the overlying services.


