Failure Area Identification via Probe Group Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud computing environments, storage administrators face challenges in identifying unhealthy regions or availability zones, as existing systems cannot detect failure areas outside the created storage nodes or volumes, hindering the creation of safe volume groups.
Innovation Solution
A method and system that create probe groups for error detection, correlate errors across multiple volume groups, and identify common error zones by analyzing error information and correlation rules, allowing for the identification of failure areas such as regions, availability zones, or paths.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If storage administrators create volume groups in cloud environment, then data replication and disaster recovery capability is improved, but ability to identify failure areas outside storage nodes is worsened
Solution Approach 1:
The system segments the cloud infrastructure into hierarchical area units (regions, availability zones, storage nodes) and creates probe groups that span multiple segments. By distributing probe volumes across different availability zones and regions, the system can identify failure areas at various hierarchical levels through correlated error analysis of these segmented probe groups.
Solution Approach 2:
The system introduces probe groups as intermediary entities between storage administrators and the underlying infrastructure. These probe groups act as mediators that actively test and report on the health of regions and availability zones, enabling indirect detection of failure areas that would otherwise be invisible to standard storage node monitoring.
2Measurement precision
If storage nodes detect their own errors, then local error detection capability is improved, but identification of external failure areas is worsened
Solution Approach 1:
The system implements feedback loops where probe groups continuously monitor infrastructure health and report errors back to the failure area identification system. Error information from multiple probe groups is collected, correlated, and processed to generate feedback about external failure areas, which then informs storage administrator decisions about where to create safe volume groups.
Solution Approach 2:
The system performs preliminary actions by proactively creating probe groups and executing error detection operations before actual failures impact production systems. By提前 deploying probe volumes across different availability zones and regions, the system预先 identifies potential failure areas, enabling storage administrators to avoid these areas when creating volume groups.
3Measurement precision
If error correlation analysis is performed across multiple probe groups, then failure area identification accuracy is improved, but system complexity is worsened
Solution Approach 1:
The system applies universal error correlation rules that work across all probe groups regardless of their specific location or configuration. By using a standardized approach to error correlation that can be applied universally to any combination of probe groups, the system achieves high failure area identification accuracy without requiring complex custom analysis logic for each specific scenario.
Data Source
AI summary
Aspects of the present disclosure involve an innovative method for detecting error zones from a plurality of volume groups. The method may include creating a plurality of probe groups for error detection; detecting a new error associated with the plurality of probe groups and the plurality of volume groups; retrieving error information associated with the new error, wherein the error information comprises an error source, an error type, and an error time; retrieving an error correlation rule associated with the error information; determining if the error correlation rule is satisfied by the error information and information of other known errors; and identifying a common zone based on the error information and the information of the other known errors as an error zone.


