Cloud Outage Detection via Synthetic Measurements and Aggregated Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional error monitoring solutions for cloud computing systems are limited, as they rely on manual processes that are not cost-effective and only monitor and log information locally, failing to efficiently detect outages across distributed systems.
Innovation Solution
Implementing a management application that processes synthetic measurements and anonymized usage data through a shared aggregator to generate aggregated data, analyzed via a decision tree to correlate outages, with a confidence value assigned and alerts generated for stakeholders.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual processes are used for installation and configuration support of cloud computing assets, then human components can provide installation and configuration support, but it is not cost effective
Solution Approach 1:
The system enables self-service through automated monitoring and alerting mechanisms. The health monitoring system automatically detects issues, generates alerts, and notifies relevant personnel without requiring manual intervention for routine monitoring tasks, thereby reducing operational costs while maintaining support quality
Solution Approach 2:
Manual monitoring and detection processes are replaced with automated electronic health monitoring systems that continuously collect metrics, analyze system state, and generate alerts automatically, eliminating the need for manual labor in routine monitoring while improving cost-effectiveness
2Reliability
If individual components monitor health related metrics locally, then local monitoring is performed, but information is consumed and processed locally without centralized correlation
Solution Approach 1:
The patent merges distributed local monitoring data into a centralized health monitoring system. Individual component metrics are aggregated and correlated centrally, enabling comprehensive system-wide outage detection that leverages both local monitoring reliability and centralized information correlation capabilities
Solution Approach 2:
A centralized health monitoring system acts as an intermediary between individual component monitors and final outage detection. This intermediary collects, aggregates, and correlates data from distributed sources, enabling comprehensive analysis that neither local nor purely centralized approaches could achieve alone
3Ease of operation
If conventional error monitoring solutions are used, then local monitoring and logging is performed, but outage detection across distributed systems is not efficient
Solution Approach 1:
The system implements feedback mechanisms where health metrics from individual components are continuously collected, analyzed, and used to generate system-wide outage detection insights. Alert information flows back to stakeholders, enabling rapid response and improving overall outage detection efficiency across the distributed system
Data Source
AI summary
Outage detection in a cloud based service is provided using synthetic measurements and anonymized usage data of the cloud based service. Synthetic measurements and usage data are processed through a shared aggregator to generate aggregated data. The synthetic measurements and the usage data are analyzed through a decision tree to correlate an outage based on the synthetic measurements and the usage data. A confidence value is assigned to the outage. An alert is generated that includes information associated with the outage and the confidence value.


