Root Cause Analysis in Sensor-Actuator Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional root cause analysis in sensor-actuator fabrics is cumbersome, error-prone, and difficult to scale due to the massive amount of data from thousands of time-series metrics, often requiring manual processing which is inefficient and prone to errors.
Innovation Solution
Automated root cause analysis is performed by identifying a network dependency group, receiving time-series performance data, and using a multiscale local subspace algorithm to detect statistically significant changes and determine the root cause of faults within the network, allowing for real-time detection and impact assessment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual root cause analysis is performed on thousands of time-series metrics, then comprehensive fault analysis can be achieved, but the process becomes cumbersome, error-prone, and difficult to scale
Solution Approach 1:
The patent replaces manual mechanical analysis processes with automated computational algorithms. Specifically, it uses machine learning models and automated data processing systems to analyze time-series metrics from sensor-actuator fabrics, substituting human operators with computational systems that can process thousands of metrics simultaneously without error or fatigue.
Solution Approach 2:
The patent creates simplified representations or models of the complex sensor-actuator fabric system. By generating synthetic data copies and using digital twins or simulation models, the system can analyze fault patterns without directly processing all原始 data, reducing computational complexity while maintaining analysis accuracy.
2Loss of information
If manual root cause analysis is used, then detailed fault investigation is possible, but the process is time-consuming and difficult to scale
Solution Approach 1:
The patent implements preliminary automated preprocessing and filtering of time-series data before full analysis. It pre-identifies potential fault patterns, pre-processes sensor data into standardized formats, and pre-ranks metrics by relevance, so that when faults occur, the system can immediately analyze pre-prepared data structures rather than processing raw data from scratch.
Solution Approach 2:
The patent divides the analysis process into discrete automated segments or modules. Each module handles specific aspects of fault analysis (data collection, preprocessing, pattern recognition, root cause identification), allowing parallel processing and reducing overall analysis time while maintaining comprehensive fault detection through systematic coverage of all metrics.
3Reliability
If comprehensive monitoring of all time-series metrics is implemented, then complete fault coverage is achieved, but the data volume becomes massive and overwhelming
Solution Approach 1:
The patent extracts and isolates only the most relevant features and metrics from the massive time-series data. It uses feature extraction techniques to identify and pull out critical signal characteristics, fault indicators, and key performance parameters, discarding redundant information while preserving all essential fault detection capabilities.
Solution Approach 2:
The patent implements selective monitoring that focuses on critical subsets of metrics rather than uniformly processing all data. It applies partial action by concentrating computational resources on high-risk or high-impact metrics identified through risk assessment, while using lighter processing for less critical parameters, achieving reliable fault detection with reduced overall data processing.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
In one embodiment, the techniques herein provide that a node may receive indicia of a fault state in one or more components of a computer network. Based on the indicia, the node may then identify a network dependency group including a plurality of network components that are hierarchically associated with the one or more components. The node may then receive, from a database, a time series of performance data values corresponding to the network dependency group, wherein the time series comprises performance data values from before and after the onset of the fault state. The node may then identify altered performance data values in the time series comprising values which differ before and after onset of the fault state, and then determine a root cause of the fault state by identifying one or more particular components within the network dependency group that are associated with the altered performance data values.