Cloud Root Cause Analysis via ML Prediction and Multi-Dimensional Signal Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In complex cloud environments, existing root cause analysis systems face challenges in detecting and analyzing faults due to large-scale, multi-tenant setups with complex dependencies, high-frequency accidents, and the need for multi-dimensional signal processing, often relying on manual supervision and limited to specific layers of the cloud infrastructure.

Innovation Solution

A root cause analysis system that includes an infra controller for searching data source endpoints, a monitoring module for real-time data collection, a prediction and localization module using machine learning-based models for anomaly prediction and root cause identification, and a treatment module for preliminary recovery, enabling multi-dimensional signal processing and large data handling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If manual supervision is used for root cause analysis, then system complexity is reduced, but productivity and fault detection speed deteriorate

Engineering Contradiction:
Improvesystem complexityVSAvoidfault detection speed
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The system implements automated self-service through machine learning models that autonomously perform anomaly detection, root cause identification, and preliminary remediation actions without requiring manual supervision, thereby maintaining low operational complexity while significantly improving fault detection speed and productivity

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical supervision with automated electronic systems including monitoring agents, data collectors, and machine learning-based prediction models that automatically analyze cloud infrastructure data, detect anomalies, and identify root causes, substituting human effort with intelligent automated mechanisms

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If existing monitoring solutions are applied to cloud environments, then ease of operation is maintained, but measurement precision and root cause analysis accuracy deteriorate

Engineering Contradiction:
Improveease of operationVSAvoidroot cause analysis accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system changes the parameters of analysis by incorporating multiple signal dimensions (metrics, logs, traces) and using machine learning models that dynamically adjust analysis parameters based on the specific cloud environment and workload characteristics, thereby improving measurement precision while maintaining ease of operation through automated parameter tuning

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent adds another dimension to monitoring by implementing multi-dimensional signal processing that analyzes cloud infrastructure data across multiple layers (infrastructure, platform, application) and multiple signal types simultaneously, enabling more accurate root cause analysis while maintaining operational simplicity through integrated multi-dimensional analysis

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If multi-dimensional signal processing is implemented, then root cause analysis accuracy is improved, but device complexity and data processing requirements worsen

Engineering Contradiction:
Improveroot cause analysis accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies segmentation by dividing the complex multi-dimensional analysis into distinct modular components: monitoring agents deployed in clusters, data collectors that aggregate data, prediction models that analyze patterns, and root cause analyzers that identify sources. This modular segmentation reduces overall device complexity while maintaining high root cause analysis accuracy through specialized processing at each stage

Inventive Principle:
Principle #1Segmentation

4Ease of operation

If existing solutions are used in mixed cloud environments, then ease of operation is maintained, but reliability and fault detection capability worsen

Engineering Contradiction:
Improveease of operationVSAvoidfault detection capability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent implements universality by designing a multi-functional monitoring system that can operate across diverse mixed cloud environments (public cloud, private cloud, hybrid cloud) and handle various workload types (containers, virtual machines, bare metal) through a unified architecture, thereby improving reliability and fault detection capability while maintaining ease of operation through environment-agnostic design

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240202063A1Root cause analysis method, apparatus and system in cloud environment
Publication Date: 2024.06.20 FOUND OF SOONGSIL UNIV IND COOP
  • US20240202063A1 patent drawing
  • US20240202063A1 patent drawing
  • US20240202063A1 patent drawing

AI summary

A root cause analysis system for analyzing a root cause in a cloud environment comprises: an infra controller searching a data source endpoint, and bringing address and port information of the data source endpoint when a monitoring agent is installed in a plurality of clusters; a monitoring module registering the address and port information of the data source endpoint according to a request of the intra controller, and collecting data from the monitoring agent in real time; a prediction and localization module predicting an abnormal accident in the plurality of clusters by inputting the data collected in real time into a machine learning-based prediction model, and searching a root cause of the abnormal accident by using a feature score and a log score for a metric of the abnormal accident; and a remediation (treatment) module performing a preliminary recovery process according to the searched root cause.