Machine Learning Root Cause Identification in Multi-Layer Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional performance monitoring and root cause analysis in networks are manual, time-consuming, and costly, and existing machine learning solutions struggle to identify root causes in complex networks with multiple affected elements, especially when data is non-stationary and distribution is not normally distributed, and they fail to scale with large networks.
Innovation Solution
A system and method using Machine Learning (ML) to analyze communications networks of known topology by obtaining Performance Monitoring (PM) metrics and alarms, identifying problematic components and root causes through a processing device configured to utilize ML processes, treating each Network Element independently and considering network topology to create scalable algorithms, and employing object recognition techniques to analyze multi-dimensional time-series data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual performance monitoring and root cause analysis are used, then human expertise can interpret complex network failures, but the process is time-consuming, expensive, and not scalable to large networks
Solution Approach 1:
The patent replaces manual mechanical analysis by human experts with an automated machine learning system that uses deep learning models to analyze network performance data, identify patterns, and determine root causes automatically, eliminating the time-consuming manual process while maintaining or improving accuracy
Solution Approach 2:
The patent introduces an intermediary machine learning system that acts as a bridge between raw network performance data and human decision-makers, processing and interpreting complex multi-dimensional data to provide actionable insights without requiring direct human analysis of the raw data
2Ease of manufacture
If rule-based engines with hard-coded thresholds are used for anomaly detection, then known failures can be detected, but the approach cannot find failures spanning multiple network elements and requires extensive expert knowledge to define rules
Solution Approach 1:
The patent transforms the detection approach by changing from fixed threshold parameters to dynamic pattern recognition parameters learned by deep learning models from historical data, allowing the system to adapt to various failure modes without reprogramming rules
Solution Approach 2:
The patent creates a universal deep learning model that can detect multiple types of failures across different network elements simultaneously, replacing the need for separate rule-based detectors for each failure type and making the system versatile without requiring extensive rule configuration
3Device complexity
If conventional anomaly detection approaches are used, then simple threshold crossings can be identified, but the approach presumes failures occur at single points and cannot detect cascading failures across multiple network elements
Solution Approach 1:
The patent adds temporal and topological dimensions to the analysis by examining performance data across multiple time points and network elements simultaneously, allowing the detection of cascading failure patterns that span across the network topology rather than isolated single-point failures
4Productivity
If Machine Learning methodologies are introduced for pattern detection, then automated fault detection and forecasting can be enabled, but the computational complexity and data requirements increase
Solution Approach 1:
The patent segments the network into manageable components and applies localized analysis to each segment, processing performance data in distributed fashion across network elements, which reduces the computational burden on any single system while maintaining overall detection capability
Data Source
AI summary
Systems and methods for detecting patterns in data from a time-series are provided. According to some implementations, the systems and methods may use network topology information combined with object recognition techniques to detect patterns. One embodiment of a method includes the steps of obtaining information defining a topology of a multi-layer network having a plurality of Network Elements (NEs) and a plurality of links interconnecting the NEs and receiving Performance Monitoring (PM) metrics and one or more alarms from the multi-layer network. Based on the information defining the topology, the PM metrics, and the one or more alarms, the method also includes the step of utilizing a Machine Learning (ML) process to identify a problematic component from the plurality of NEs and links and to identify a root cause associated with the problematic component.


