Hierarchical Anomaly Detection and Resolution in Cloud Infrastructure
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud computing systems face challenges in automatically detecting and resolving anomalies, leading to potential service level agreement (SLA) violations due to manual intervention, inefficiencies in resource usage, and delayed corrective actions, which can result in unsatisfactory quality of service (QoS) and increased costs.
Innovation Solution
An anomaly detection and resolution system (ADRS) that automatically detects and corrects anomalies by implementing anomaly classification, using defined and undefined anomaly categories, and employing a hierarchical approach with anomaly detection and resolution components (ADRCs) to manage anomalies locally within components, reducing the need for centralized management and minimizing human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual administration and monitoring tools are used to manage cloud data centers, then operational flexibility and human judgment can be applied, but the cost of cloud services increases and service quality deteriorates due to delayed anomaly detection and correction
Solution Approach 1:
The system implements self-service through autonomous anomaly detection and resolution capabilities. ADRCs automatically monitor system state, detect anomalies by comparing against learned normal behavior patterns, and execute corrective actions without human intervention. This eliminates the need for manual administration while maintaining or improving service quality through faster, more consistent anomaly response.
Solution Approach 2:
The patent replaces manual mechanical operations with automated computational systems. Machine learning models substitute human judgment in analyzing system logs and metrics, while automated rule engines replace manual decision-making processes. This substitution reduces operational complexity while enhancing reliability through consistent, scalable automation.
2Loss of time
If centralized anomaly management is implemented, then comprehensive system-wide monitoring is achieved, but response time increases and human intervention is required more frequently
Solution Approach 1:
The system segments anomaly management by deploying distributed ADRCs at multiple levels within the cloud infrastructure hierarchy. Each ADRC operates autonomously within its local context, detecting and resolving anomalies immediately without requiring centralized coordination. This segmentation dramatically reduces response time while maintaining high automation levels through localized decision-making.
Solution Approach 2:
ADRCs perform preliminary actions by continuously learning normal system behavior patterns through machine learning and establishing baseline expectations. When anomalies occur, corrective actions are pre-defined and automatically executed based on learned patterns, enabling rapid response without human intervention or centralized approval processes.
3Measurement precision
If machine learning algorithms analyze log files to detect anomalies, then automated detection capability is improved, but false positives increase and programmatic corrective action becomes difficult due to broad and noisy error detection
Solution Approach 1:
The system applies local quality by tailoring anomaly detection to specific cloud infrastructure components and contexts. Each ADRC learns normal behavior patterns specific to its local environment and component type, rather than applying generic anomaly detection across the entire system. This contextualized approach reduces false positives while maintaining high detection precision for relevant anomalies.
Solution Approach 2:
The system implements feedback mechanisms where ADRCs continuously learn from detected anomalies and their outcomes. Detection accuracy improves over time as the machine learning models are refined based on feedback from actual system behavior and corrective action results. This feedback loop enhances both detection precision and reliability by eliminating false positives and improving anomaly classification.
Data Source
AI summary
An anomaly detection and resolution system (ADRS) is disclosed for automatically detecting and resolving anomalies in computing environments. The ADRS may be implemented using an anomaly classification system defining different types of anomalies (e.g., a defined anomaly and an undefined anomaly). A defined anomaly may be based on bounds (fixed or seasonal) on any metric to be monitored. An anomaly detection and resolution component (ADRC) may be implemented in each component defining a service in a computing system. An ADRC may be configured to detect and attempt to resolve an anomaly locally. If the anomaly event for an anomaly can be resolved in the component, the ADRC may communicate the anomaly event to an ADRC of a parent component, if one exists. Each ADRC in a component may be configured to locally handle specific types of anomalies to reduce communication time and resource usage for resolving anomalies.


