Anomaly Correlation System for Cloud Change Impact Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Managing complex cloud computing systems is challenging due to the interconnected nature of components, making it difficult to determine the impact of changes across different layers, leading to potential failures and requiring extensive testing that can be time-consuming and inaccurate.

Innovation Solution

An anomaly correlation system that receives telemetry data on change and failure events across multiple computing layers, detects anomalies, and determines correlations between change events and failure anomalies, providing recommendations to mitigate these issues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional testing methods (simulation or user testing) are used to detect negative impacts of changes, then measurement precision of change impact is improved, but loss of time increases significantly

Engineering Contradiction:
Improveaccuracy of change impact detectionVSAvoidtesting duration
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system proactively monitors and detects anomalies before they manifest as failures by continuously analyzing telemetry data from multiple layers. Change events are tracked and correlated with system state in advance, allowing potential issues to be identified through anomaly detection algorithms before they cause actual problems, thus eliminating the need for time-consuming post-change testing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces conventional mechanical testing approaches (simulation environments and user beta testing) with an automated electronic monitoring system that uses telemetry data collection, anomaly detection algorithms, and automated correlation analysis. This substitution enables continuous real-time monitoring without human intervention, dramatically reducing testing time while maintaining or improving detection accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If user testing (beta testing) is conducted to accurately measure failure statistics, then measurement precision is improved, but device complexity and operational disruption increase

Engineering Contradiction:
Improvefailure statistics accuracyVSAvoidtesting infrastructure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-diagnosis and self-monitoring by automatically collecting telemetry data, detecting anomalies, and correlating them with change events without requiring external user testing infrastructure. The cloud computing system monitors itself through embedded agents and telemetry collection mechanisms, eliminating the need for separate beta testing environments and reducing overall system complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The monitoring system serves multiple functions simultaneously: it collects telemetry data for operational monitoring, detects anomalies for quality assurance, correlates events for root cause analysis, and provides feedback for continuous improvement. This multi-functional approach consolidates what would otherwise require separate testing infrastructures into a single integrated system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If extensive testing is performed to ensure reliability, then reliability is improved, but productivity decreases due to time consumption

Engineering Contradiction:
Improvesystem stabilityVSAvoidchange implementation speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements continuous feedback loops where telemetry data is constantly monitored, anomalies are detected in real-time, and correlations with change events are automatically established. This feedback mechanism provides immediate visibility into system health and change impact, allowing rapid response to issues while maintaining high reliability, thus enabling faster change implementation without sacrificing stability.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

By proactively detecting anomalies before they become failures and continuously monitoring system state, the system ensures reliability through prevention rather than through extensive post-change testing. This preliminary detection approach allows changes to be implemented quickly with confidence that issues will be caught early, thereby maintaining both high reliability and high productivity.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If the cloud computing system monitors all change events across multiple layers, then measurement precision of cross-layer impact is improved, but use of energy and processing expenses increase

Engineering Contradiction:
Improvecross-layer correlation accuracyVSAvoidprocessing energy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The monitoring system is divided into segment-specific agents that collect telemetry data only from their local layer or component. Each agent monitors its own segment and sends relevant data to a central correlation system. This segmentation reduces the processing load on any single component while maintaining comprehensive cross-layer monitoring capability, as each segment processes only its local data rather than all system data.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system extracts only the essential and relevant telemetry data from each layer that is necessary for cross-layer correlation analysis. Rather than processing all generated data, the monitoring agents filter and extract key metrics and events that are most likely to indicate cross-layer impacts. This extraction approach maintains measurement precision while significantly reducing processing energy consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240378101A1Detecting and mitigating cross-layer impact of change events on a cloud computing system
Publication Date: 2024.11.14 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20240378101A1 patent drawing
  • US20240378101A1 patent drawing
  • US20240378101A1 patent drawing

AI summary

The present disclosure relates to systems, methods, and computer-readable media for identifying anomalies of failure events on a cloud computing system and determining cross-component and cross-layer correlation between change events that occur on the cloud computing system and the failure events associated with the anomalies. In particular, this disclosure describes a system that receives telemetry related to change events and failure events across any number of computing layers of a distributed computing environment (e.g., a cloud computing system) and detects anomalies based on counts of failure events that are manifested over discrete periods of time. Based on these detected anomalies, the anomaly correlation system can determine cross-layer and cross-component correlations between selective change events and the detected anomalies of failure events. The anomaly correlation system may further generate and provide recommendations related to mitigating or otherwise addressing the anomalies based on the determined correlations.