Cloud Service Health Model for Correlating Resource and Service Status
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cloud resource monitoring systems struggle to provide a holistic, adaptable, and efficient health monitoring solution that correlates the health of individual resources to broader service health determinations, often leading to overwhelming alerts and manual analysis due to varying owner perspectives on health.
Innovation Solution
A directed graph health model is implemented to automatically correlate cloud resource health to broader service health, using anomaly detection and machine learning to determine health categories, adaptable to different owners' perspectives, and reducing alert noise through customizable thresholds and actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis is used to determine service health, then health determination can be made, but it requires significant time and human resources
Solution Approach 1:
The system enables self-service health monitoring by automatically collecting metrics from cloud resources, applying anomaly detection algorithms, and generating health determinations without human intervention. The health model autonomously processes data from multiple sources and correlates resource health to service health levels.
Solution Approach 2:
Manual analysis is replaced with automated computational systems including anomaly detection algorithms and machine learning models. These systems process metrics and determine health states automatically, substituting human analytical work with computational processes.
2Loss of information
If comprehensive monitoring of all cloud resources is implemented, then complete health visibility is achieved, but alert noise becomes overwhelming
Solution Approach 1:
The system merges health information from multiple individual cloud resources into a consolidated service health level. By aggregating resource health data and correlating it through the health model, the system provides comprehensive visibility while presenting a unified health status that avoids overwhelming alert noise.
Solution Approach 2:
The service health level acts as an intermediary between individual resource metrics and end-user concerns. Instead of presenting raw metrics from each resource, the system translates them into meaningful service health determinations that are easier to interpret and act upon.
3Productivity
If a standardized health model is applied to all entities, then scalability is improved, but adaptability to specific owner perspectives is reduced
Solution Approach 1:
The health model incorporates dynamic customization capabilities that allow owners to adjust parameters, thresholds, and priorities based on their specific perspectives. The system transitions from a static standardized model to a dynamic one that can be adapted while maintaining the underlying standardized structure for scalability.
Solution Approach 2:
The system applies local quality by allowing customization at the owner level while maintaining a standardized core model. Each owner can configure specific parameters and thresholds relevant to their needs, while the overall health model structure remains consistent across the platform, enabling both scalability and adaptability.
4Measurement precision
If detailed metrics are collected from each cloud resource, then monitoring precision is improved, but system complexity increases
Solution Approach 1:
The monitoring system is segmented into distinct functional components: metric collection modules for individual resources, an anomaly detection layer, a health correlation engine, and a reporting layer. This segmentation allows detailed metric collection while managing complexity through modular architecture, where each component handles specific tasks independently.
Data Source
AI summary
The techniques described herein automatically correlate the health of cloud resources to a broader health determination for an entity executing within, or supported by, a distributed computing environment. In contrast to the typical manual analysis that is required to make a broader health determination for a specific entity, the techniques generate and use a standard health model that can be applied, or scaled, to detect unhealthy scenarios across a variety of different entities with different owners (e.g., different tenants and/or different cloud resource providers). Furthermore, to meet varying owner perspectives on health, the techniques include a layer on top of the standard health model that enables an owner to provide input that customizes the standard health model for their own entity.


