Automated Anomaly Detection in Distributed Computing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In distributed computing systems, the vast amount of metric data makes it difficult for IT administrators to manually monitor performance issues in real time, leading to potential service interruptions and significant costs when issues like server application failures occur.
Innovation Solution
Automated processes and systems that collect and analyze historical metrics to compute time-dependent system indicators, train state classifiers, and detect abnormal behavior, generating alerts and recommendations for remedial actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual monitoring of metrics is used, then IT administrators can understand system performance, but the enormous number of metric data streams makes it very difficult to manually monitor and detect performance issues in real time
Solution Approach 1:
The system performs automated anomaly detection using machine learning models that continuously learn from historical data and detect abnormalities without human intervention. The automated anomaly detection system processes metric data streams independently, eliminating the need for manual analysis while reducing the complexity burden on IT administrators.
Solution Approach 2:
The patent replaces manual monitoring processes with automated machine learning-based detection systems. The mechanical process of manually analyzing numerous metric streams is substituted with computational algorithms that automatically identify anomalies, thereby reducing operational complexity while maintaining monitoring effectiveness.
2Reliability
If real-time response to performance issues is implemented, then service interruptions can be prevented, but the enormous number of metrics makes it difficult to detect and respond in real time
Solution Approach 1:
The system performs preliminary actions by continuously training machine learning models on historical data and establishing baseline behaviors before anomalies occur. The system pre-processes and stores metric data patterns, enabling rapid detection and response when abnormalities are detected, thus preventing service interruptions while reducing response time.
Solution Approach 2:
The automated anomaly detection system implements feedback mechanisms where detected anomalies trigger immediate responses and the system continuously learns from detected patterns to improve detection accuracy over time. This feedback loop enables real-time response to performance issues while adapting to changing system behaviors.
3Extent of automation
If automated anomaly detection is implemented, then real-time detection of abnormal behavior is enabled, but the system complexity increases with metrics collection and processing
Solution Approach 1:
The machine learning model serves multiple functions: it detects anomalies, classifies abnormal behaviors, and provides recommendations for remedial actions. By making the detection system multi-functional, the patent reduces the need for separate specialized systems, thereby managing complexity while enabling extensive automation.
Solution Approach 2:
The patent introduces an intermediary layer of machine learning models that sit between metric collection and automated response systems. These models process and interpret complex metric patterns, simplifying the overall system architecture by providing a unified interpretation layer that handles multiple detection and response functions.
Data Source
AI summary
Automated processes and systems for detecting abnormally behaving objects of a distributed computing system are described. Processes and systems obtain metrics that are generated in a historical time window and are associated with an object of the distributed computing system. Processes and system use the metrics to compute a time-dependent system indicator over the historical time window. Each value of the system indicator corresponds to a point in time of the historical time window when the object was in a normal or an abnormal state. Processes and systems use the normal and abnormal states of the system indicator in the historical time window to train a state classifier that is used to detect run-time abnormal behavior of the object. When the state classifier identifies abnormal behavior of the object, an alert is generated, indicating the abnormal behavior of the object.


