Disk Array Failure Prediction and Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems face challenges in preventing data unavailable (DU) and data lost (DL) events, which can lead to user dissatisfaction and increased support pressure, as existing methods are inadequate in detecting potential issues proactively and providing effective solutions, especially for hardware-related problems and large-capacity drives.
Innovation Solution
A method and device that collect data from disk arrays, analyze it for potential failure events using a knowledge database and matching algorithm, and generate reports to alert users of impending issues, enabling timely action to prevent DU and DL events by determining appropriate actions based on collected data and historical failure events.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If two storage pools are configured to work in active mode for high reliability, then system availability is improved, but the risk of simultaneous failure causing data unavailability increases
Solution Approach 1:
The system performs preliminary analysis of collected data to determine potential failure events before they occur. By proactively identifying disks at risk of failure and taking preventive actions (such as data migration or replacement), the system avoids the harmful effect of simultaneous failures that would cause data unavailability, while maintaining the high availability architecture.
2Device complexity
If conventional RAID configuration is used without proactive failure detection, then system complexity is reduced, but data unavailable and data lost rates increase
Solution Approach 1:
The system implements a feedback mechanism by continuously collecting data from storage devices, analyzing this data to determine potential failure events, and providing alerts or automated responses. This feedback loop enables proactive detection of failing disks and triggers preventive actions, significantly reducing data unavailable and data lost rates while maintaining relatively simple conventional RAID configurations.
3Reliability
If proactive failure detection and prevention mechanisms are implemented, then data unavailable and data lost rates are reduced, but system complexity and data processing requirements increase
Solution Approach 1:
The system employs self-service mechanisms where storage devices automatically provide diagnostic information and health status data without requiring external intervention. The analysis system processes this self-provided data to identify potential failures and trigger preventive actions, reducing the need for complex manual monitoring and intervention systems while achieving low data lost rates.
Data Source
AI summary
Techniques involve avoiding a potential failure event on a disk array. Along these lines, data collected for a disk array are obtained. It is determined, based on the collected data, whether a potential failure event is to occur on the disk array. In response to determining that the potential failure event is to occur on the disk array, an action to be taken for the disk array is determined, to avoid occurrence of the potential failure event.


