Distributed Management Component for Cloud Storage Failure Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In virtualized, shared infrastructure, identifying and diagnosing device failures across physically distant cluster nodes with different hardware and software configurations is challenging due to the complexity and variability of storage devices.
Innovation Solution
A distributed management component with a notification system is implemented, including a polling component, support client, scheduler, collector, and delivery component, to manage automated support and diagnose device failures by aggregating and delivering diagnostic information across the network, and translating OEM product labels into company product labels for replacement identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a distributed storage system uses multiple cluster nodes with different hardware and software configurations located at physical distances, then the system scalability and geographic distribution are improved, but the difficulty of identifying and diagnosing device failures increases
Solution Approach 1:
The system performs preliminary actions by proactively collecting diagnostic information from storage devices before failures occur. The diagnostic information collection module continuously gathers data about device status, performance metrics, and configuration details, preparing comprehensive diagnostic records that enable rapid failure identification and diagnosis when issues arise, thus resolving the contradiction between system scalability and failure diagnosis difficulty.
2Reliability
If automated support systems collect and transmit diagnostic information across geographically dispersed nodes, then the system reliability is improved, but the network bandwidth consumption and data transmission time increase
Solution Approach 1:
The system extracts and transmits only the essential diagnostic information needed for failure diagnosis rather than collecting and transmitting all possible data. The diagnostic information collection module identifies and extracts critical parameters such as device status, error logs, and performance metrics, reducing network bandwidth consumption and data transmission time while maintaining system reliability through automated support.
3Ease of repair
If the system translates OEM product labels to company product labels for replacement identification, then the ease of repair is improved, but the processing time and computational resources increase
Solution Approach 1:
The system performs preliminary translation of OEM product labels to company product labels and stores the mappings in advance. The diagnostic information collection module pre-processes product label information and creates a lookup table of equivalent products, enabling rapid replacement identification during repair operations without incurring processing delays at the time of failure, thus resolving the contradiction between ease of repair and processing time.
Data Source
AI summary
An OEM product label may be associated with a plurality of different OEM products; an identity of the manufacturer of the product associated with the label is determined based on the context of the product.


