Cluster Processing for Identifying Failing Host Sets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In network computing and storage systems, identifying and addressing device malfunctions or potential failures at scale can be challenging due to the large number of devices managed by computing resource service providers, often leading to unnecessary removal of hosts from production without resolving the underlying issues.

Innovation Solution

A method and system that utilize dimensional algorithms for cluster processing of resource host attribute data to identify related groups of failing or troubled network nodes, allowing for prioritization and rehabilitation of shared failure modes, thereby reducing capital expenditure and improving resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional monitoring methods are used to identify device failures in large-scale networks, then individual device problems can be detected, but the scale and complexity of network-wide events make identification difficult and time-consuming

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidtime to identify failures
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the large-scale network into smaller clusters based on device attributes and failure patterns. By dividing the monitoring space into manageable segments, the system can efficiently identify and analyze failures within each cluster rather than searching through the entire network, thus maintaining detection accuracy while reducing time loss.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dimensional analysis by examining failures across multiple attributes simultaneously (device type, location, time, failure mode). This multi-dimensional approach transforms the problem from linear device-by-device checking to pattern recognition across attribute spaces, dramatically reducing identification time while improving precision through comprehensive analysis.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If hosts are removed from production when failures are detected, then service reliability is maintained, but unnecessary removals occur and capital expenditure increases

Engineering Contradiction:
Improveservice reliabilityVSAvoidcapital expenditure
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent performs preliminary analysis of failure patterns and correlations before making removal decisions. By pre-identifying which failures are isolated versus which indicate broader issues, the system can take preliminary actions to address root causes and prevent unnecessary host removals, thereby maintaining reliability while reducing capital expenditure.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms that continuously monitor failure patterns and adjust removal decisions based on learned correlations. When the system detects that certain failure modes are transient or self-correcting, feedback loops prevent unnecessary removals. When genuine systemic issues are detected, feedback ensures appropriate hosts are removed, optimizing both reliability and cost.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If comprehensive monitoring of all device attributes is performed, then accurate failure pattern recognition is achieved, but data processing complexity and resource requirements increase

Engineering Contradiction:
Improvefailure pattern recognition accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and focuses on the most critical attributes and failure patterns that are most indicative of systematic issues. By taking out only the essential data elements needed for correlation analysis rather than processing all possible attributes, the system achieves accurate pattern recognition while reducing data processing complexity and resource requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different levels of monitoring intensity to different device groups based on their failure patterns and criticality. High-value attributes are monitored closely for devices showing signs of systematic failure, while standard monitoring is maintained for stable devices. This local quality approach maintains recognition accuracy for problematic areas while reducing overall processing complexity.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10592328B1Using cluster processing to identify sets of similarly failing hosts
Publication Date: 2020.03.17 AMAZON TECH INC
  • US10592328B1 patent drawing
  • US10592328B1 patent drawing
  • US10592328B1 patent drawing

AI summary

Data, attributes, and metrics from unavailable resource hosts may be collected and used for cluster analysis in order to correlate the different hosts and group similar hosts into clusters. The clusters may be ranked based on the collected information and used to provide a simple way to identify shared failure modes among the unavailable hosts. By identifying the hosts of each cluster, shared failures can be corrected for large groups of hosts at the same time, enabling the hosts to return to operational states.