Large-Scale Event Detector for Cloud Hosts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computing systems lack effective mechanisms for proactive detection of large-scale events such as power, thermal, or network failures, leading to delayed awareness and prolonged downtime, which can severely impact businesses relying on these systems.

Innovation Solution

A large-scale event detection system that monitors host responsiveness, assigns weight values based on criticality, and filters out unrelated hosts to determine the occurrence and probable root cause of events, enabling timely notification and action.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional monitoring methods are used to detect computing system failures, then the system complexity is low, but the detection speed and reliability are insufficient leading to delayed awareness of large-scale events

Engineering Contradiction:
Improvedetection reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the monitoring task by dividing hosts into different groups based on their criticality weights. Instead of treating all hosts uniformly, the system creates weighted categories (e.g., critical, important, standard hosts) and applies different monitoring thresholds and analysis methods to each segment. This segmentation enables more reliable detection of large-scale events by focusing resources on critical hosts while still monitoring the overall system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary component - the large-scale event detector - that sits between the individual host monitoring and the final event determination. This intermediary aggregates host status information, applies weighted calculations, and filters out unrelated unresponsiveness before determining whether a large-scale event has occurred. This intermediary layer improves detection reliability by providing a systematic approach to analyzing complex host data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If all host unresponsiveness is treated equally, then the monitoring process is simple, but the ability to identify true large-scale events is reduced due to noise from unrelated hosts

Engineering Contradiction:
Improveevent detection precisionVSAvoidmonitoring complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by assigning different weights to different hosts based on their criticality to the customer's business. Critical hosts receive higher weights and their unresponsiveness has a greater impact on the large-scale event determination. This local differentiation allows the system to focus on hosts that matter most, improving detection precision by reducing noise from non-critical hosts while maintaining a manageable monitoring complexity through automated weight-based classification.

Inventive Principle:
Principle #3Local quality

3Loss of time

If proactive detection of large-scale events is implemented, then the awareness time is improved, but the difficulty of determining event occurrence increases due to multiple factors

Engineering Contradiction:
ImprovedowntimeVSAvoidevent determination difficulty
Core Design Contradiction:
Loss of timeVSDifficulty of detecting and measuring

Solution Approach 1:

The patent implements preliminary action by pre-calculating and storing weight values for each host based on their criticality to the customer's business. These weights are determined in advance and stored in a data structure, so when host unresponsiveness occurs, the system can quickly retrieve and apply the pre-computed weights without performing complex analysis in real-time. This preliminary preparation significantly reduces the time needed for event determination while maintaining accurate detection of large-scale events.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10176033B1Large-scale event detector
Publication Date: 2019.01.08 AMAZON TECH INC
  • US10176033B1 patent drawing
  • US10176033B1 patent drawing
  • US10176033B1 patent drawing

AI summary

A system and method for detecting the occurrence of an event causing multiple hosts to be unresponsive. The system and method including, for a set of hosts providing services to one or more customers of a computing resource service provider, determining one or more subsets of hosts that are unresponsive, determining whether the one or more subsets of hosts that are unresponsive meet a set of criteria for an occurrence of an large-scale event affecting multiple hosts, based at least in part on a determination that the set of criteria is met, initiating a remediation action.