Anomaly Detection in Distributed Systems via Clustered Trace Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing complexity of distributed computing systems has led to management and administration challenges, including inefficiencies and scaling issues, making traditional approaches to automated management and administration impractical due to high computational overheads and complexity.
Innovation Solution
The use of distributed-computer-system metrics, call traces, and attribute values to identify attribute dimensions related to anomalous behaviors in distributed-computer-system components through decision-tree-related analyses, facilitating the resolution of these anomalies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If traditional automated management and administration approaches are used in distributed computing systems, then management functionalities can be automated, but computational overheads become excessively high and complexity increases to impractical levels
Solution Approach 1:
The patent extracts only the essential information needed for anomaly detection from the complex distributed system data. By focusing on key metrics, traces, and attribute values rather than processing all system data, the solution reduces computational overhead while maintaining effective anomaly detection capabilities
Solution Approach 2:
The patent segments the anomaly detection process into distinct phases: data collection (metrics, traces, attribute values), data processing (partitioning and clustering), and analysis (decision-tree-related analyses). This segmentation allows each phase to be optimized independently and reduces overall system complexity
2Extent of automation
If traditional automated management and administration approaches are used in distributed computing systems, then management functionalities can be automated, but computational overheads become excessively high
Solution Approach 1:
The patent extracts only the essential information needed for anomaly detection from the complex distributed system data. By focusing on key metrics, traces, and attribute values rather than processing all system data, the solution reduces computational overhead while maintaining effective anomaly detection capabilities
Solution Approach 2:
The patent applies partial action by processing only the necessary subset of system data (metrics, traces, and attribute values) rather than all available data. This selective processing approach reduces computational resource consumption while maintaining sufficient information for effective anomaly detection
3Measurement precision
If comprehensive analysis of all system data is performed to detect anomalies, then detection accuracy improves, but time required for analysis increases
Solution Approach 1:
The patent extracts only the essential information needed for anomaly detection from the complex distributed system data. By focusing on key metrics, traces, and attribute values rather than processing all system data, the solution reduces computational overhead while maintaining effective anomaly detection capabilities
Solution Approach 2:
The patent performs preliminary data processing by collecting and organizing metrics, traces, and attribute values into structured formats before anomaly detection. This preorganization of data in a partitioned and clustered manner enables faster analysis while maintaining detection accuracy
Data Source
AI summary
The current document is directed to methods and systems that employ distributed-computer-system metrics collected by one or more distributed-computer-system metrics-collection services, call traces collected by one or more call-trace services, and attribute values for distributed-computer-system components to identify attribute dimensions related to anomalous behavior of distributed-computer-system components. In a described implementation, nodes correspond to particular types of system components and node instances are individual components of the component type corresponding to a node. Node instances are associated with attribute values and node are associated with attribute-value spaces defined by attribute dimensions. A set of call traces is partitioned, by clustering. Using attribute values and call traces, attribute dimensions that are likely related to particular anomalous behaviors of distributed-computer-system components are determined by decision-tree-related analyses for each partition and are reported to one or more computational entities to facilitate resolution of the anomalous behaviors.


