Proactive Disk Failure Prediction in RAID Storage Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage management systems often operate reactively, leading to performance degradation or failure, especially due to unpredictable hard disk failures that result in data loss, as they lack proactive prediction and management capabilities.
Innovation Solution
A storage management computing device that proactively predicts disk failure in a RAID group by obtaining performance data, comparing it against threshold values, classifying drives as failed or non-defective, and temporarily suspending operations on failed drives to prevent errors, while also storing classification data for future predictions and backing up predicted failing drives to prevent data loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If reactive management is used, then human resource requirements are reduced, but data loss and performance degradation occur
Solution Approach 1:
The system performs preliminary actions by proactively predicting disk failures before they occur. The storage management computing device analyzes performance data, identifies drives at risk of failure, and suspends operations on these drives in advance, preventing data loss and maintaining continuous accessibility without waiting for actual failures.
2Reliability
If proactive prediction is implemented, then data loss is prevented, but system complexity increases
Solution Approach 1:
The storage management computing device performs self-service by automatically analyzing performance data, predicting failures, and suspending operations without human intervention. The system monitors its own storage infrastructure, identifies at-risk drives through pattern recognition, and takes corrective actions autonomously, reducing the need for complex manual management processes.
Solution Approach 2:
The system implements feedback mechanisms by continuously collecting performance data from storage drives, analyzing trends to predict failures, and using these predictions to adjust operations. The feedback loop enables the system to learn from historical data and improve its prediction accuracy over time, managing complexity through adaptive intelligence rather than rigid complex rules.
3Reliability
If operations are suspended on failing drives, then data loss is prevented, but productivity decreases
Solution Approach 1:
Operations are suspended in advance on drives predicted to fail, preventing data loss before it occurs. By identifying at-risk drives through proactive analysis and suspending operations beforehand, the system avoids the complete productivity loss that would result from actual failures and emergency recovery operations.
Solution Approach 2:
The system temporarily discards operations on predicted failing drives to prevent data loss, then recovers productivity by redirecting operations to healthy drives in the RAID group. This approach maintains overall storage capacity and performance while protecting critical data, allowing the suspended drives to be restored or replaced without complete system downtime.
Data Source
AI summary
A method, non-transitory computer readable medium, and device that assists with proactive prediction of disk failure in a RAID group includes obtaining performance data for a plurality of storage drives. The obtained performance data is compared with a stored classification data to predict one or more storage drives of the plurality of storage drives failing within a time period. The data present in the one or more storage drives predicted to fail based on the comparison is copied on to one or more secondary storage drives. A notification including a list of the one or more storage drives predicted to fail is sent upon the copying the data on to the one or more secondary storage drives.


