Hierarchical Anomaly Detection and Resolution in Cloud Infrastructure

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Cloud computing systems face challenges in automatically detecting and resolving anomalies, leading to potential service level agreement (SLA) violations due to manual intervention, inefficiencies in resource usage, and delayed corrective actions, which can result in unsatisfactory quality of service (QoS) and increased costs.

Innovation Solution

An anomaly detection and resolution system (ADRS) that automatically detects and corrects anomalies by implementing anomaly classification, using defined and undefined anomaly categories, and employing a hierarchical approach with anomaly detection and resolution components (ADRCs) to manage anomalies locally within components, reducing the need for centralized management and minimizing human intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual administration and monitoring tools are used to manage cloud data centers, then operational flexibility and human judgment can be applied, but the cost of cloud services increases and service quality deteriorates due to delayed anomaly detection and correction

Engineering Contradiction:
Improveservice qualityVSAvoidoperational complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system implements self-service through autonomous anomaly detection and resolution capabilities. ADRCs automatically monitor system state, detect anomalies by comparing against learned normal behavior patterns, and execute corrective actions without human intervention. This eliminates the need for manual administration while maintaining or improving service quality through faster, more consistent anomaly response.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical operations with automated computational systems. Machine learning models substitute human judgment in analyzing system logs and metrics, while automated rule engines replace manual decision-making processes. This substitution reduces operational complexity while enhancing reliability through consistent, scalable automation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of time

If centralized anomaly management is implemented, then comprehensive system-wide monitoring is achieved, but response time increases and human intervention is required more frequently

Engineering Contradiction:
Improveanomaly response timeVSAvoidautomation level
Core Design Contradiction:
Loss of timeVSExtent of automation

Solution Approach 1:

The system segments anomaly management by deploying distributed ADRCs at multiple levels within the cloud infrastructure hierarchy. Each ADRC operates autonomously within its local context, detecting and resolving anomalies immediately without requiring centralized coordination. This segmentation dramatically reduces response time while maintaining high automation levels through localized decision-making.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

ADRCs perform preliminary actions by continuously learning normal system behavior patterns through machine learning and establishing baseline expectations. When anomalies occur, corrective actions are pre-defined and automatically executed based on learned patterns, enabling rapid response without human intervention or centralized approval processes.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If machine learning algorithms analyze log files to detect anomalies, then automated detection capability is improved, but false positives increase and programmatic corrective action becomes difficult due to broad and noisy error detection

Engineering Contradiction:
Improveanomaly detection precisionVSAvoiddetection reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system applies local quality by tailoring anomaly detection to specific cloud infrastructure components and contexts. Each ADRC learns normal behavior patterns specific to its local environment and component type, rather than applying generic anomaly detection across the entire system. This contextualized approach reduces false positives while maintaining high detection precision for relevant anomalies.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system implements feedback mechanisms where ADRCs continuously learn from detected anomalies and their outcomes. Detection accuracy improves over time as the machine learning models are refined based on feedback from actual system behavior and corrective action results. This feedback loop enhances both detection precision and reliability by eliminating false positives and improving anomaly classification.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10853161B2Automatic anomaly detection and resolution system
Publication Date: 2020.12.01 ORACLE INT CORP
  • US10853161B2 patent drawing
  • US10853161B2 patent drawing
  • US10853161B2 patent drawing

AI summary

An anomaly detection and resolution system (ADRS) is disclosed for automatically detecting and resolving anomalies in computing environments. The ADRS may be implemented using an anomaly classification system defining different types of anomalies (e.g., a defined anomaly and an undefined anomaly). A defined anomaly may be based on bounds (fixed or seasonal) on any metric to be monitored. An anomaly detection and resolution component (ADRC) may be implemented in each component defining a service in a computing system. An ADRC may be configured to detect and attempt to resolve an anomaly locally. If the anomaly event for an anomaly can be resolved in the component, the ADRC may communicate the anomaly event to an ADRC of a parent component, if one exists. Each ADRC in a component may be configured to locally handle specific types of anomalies to reduce communication time and resource usage for resolving anomalies.