Cluster Processing for Identifying Failing Host Sets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In network computing and storage systems, identifying and addressing device malfunctions or potential failures at scale can be challenging due to the large number of devices managed by computing resource service providers, often leading to unnecessary removal of hosts from production without resolving the underlying issues.
Innovation Solution
A method and system that utilize dimensional algorithms for cluster processing of resource host attribute data to identify related groups of failing or troubled network nodes, allowing for prioritization and rehabilitation of shared failure modes, thereby reducing capital expenditure and improving resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional monitoring methods are used to identify device failures in large-scale networks, then individual device problems can be detected, but the scale and complexity of network-wide events make identification difficult and time-consuming
Solution Approach 1:
The patent segments the large-scale network into smaller clusters based on device attributes and failure patterns. By dividing the monitoring space into manageable segments, the system can efficiently identify and analyze failures within each cluster rather than searching through the entire network, thus maintaining detection accuracy while reducing time loss.
Solution Approach 2:
The patent introduces dimensional analysis by examining failures across multiple attributes simultaneously (device type, location, time, failure mode). This multi-dimensional approach transforms the problem from linear device-by-device checking to pattern recognition across attribute spaces, dramatically reducing identification time while improving precision through comprehensive analysis.
2Reliability
If hosts are removed from production when failures are detected, then service reliability is maintained, but unnecessary removals occur and capital expenditure increases
Solution Approach 1:
The patent performs preliminary analysis of failure patterns and correlations before making removal decisions. By pre-identifying which failures are isolated versus which indicate broader issues, the system can take preliminary actions to address root causes and prevent unnecessary host removals, thereby maintaining reliability while reducing capital expenditure.
Solution Approach 2:
The patent implements feedback mechanisms that continuously monitor failure patterns and adjust removal decisions based on learned correlations. When the system detects that certain failure modes are transient or self-correcting, feedback loops prevent unnecessary removals. When genuine systemic issues are detected, feedback ensures appropriate hosts are removed, optimizing both reliability and cost.
3Measurement precision
If comprehensive monitoring of all device attributes is performed, then accurate failure pattern recognition is achieved, but data processing complexity and resource requirements increase
Solution Approach 1:
The patent extracts and focuses on the most critical attributes and failure patterns that are most indicative of systematic issues. By taking out only the essential data elements needed for correlation analysis rather than processing all possible attributes, the system achieves accurate pattern recognition while reducing data processing complexity and resource requirements.
Solution Approach 2:
The patent applies different levels of monitoring intensity to different device groups based on their failure patterns and criticality. High-value attributes are monitored closely for devices showing signs of systematic failure, while standard monitoring is maintained for stable devices. This local quality approach maintains recognition accuracy for problematic areas while reducing overall processing complexity.
Data Source
AI summary
Data, attributes, and metrics from unavailable resource hosts may be collected and used for cluster analysis in order to correlate the different hosts and group similar hosts into clusters. The clusters may be ranked based on the collected information and used to provide a simple way to identify shared failure modes among the unavailable hosts. By identifying the hosts of each cluster, shared failures can be corrected for large groups of hosts at the same time, enabling the hosts to return to operational states.


