Storage Controller Scheduling for RAID Multi-Dead Prevention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In RAID systems with redundancy, predicting and preventing a multi-dead state where concurrent malfunctions of multiple storage devices lead to data unrecoverability, existing methods often result in further malfunctions during replacement, causing data loss.
Innovation Solution
A storage controlling device and method that determines replacement timings for storage devices based on SMART information, prioritizing devices with rapidly increasing reallocated sectors counts and adjusting replacement sequences to avoid concurrent malfunctions, using a processor to analyze SMART data and output replacement information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multiple storage devices are replaced simultaneously based on SMART information, then the risk of multi-dead state is reduced, but additional malfunctions may occur during replacement causing data loss
Solution Approach 1:
The system performs preliminary actions by determining replacement timings for multiple storage devices in advance based on SMART information, analyzing the temporal relationships between predicted malfunctions, and establishing a coordinated replacement schedule before any malfunctions occur. This preliminary planning prevents the harmful effect of concurrent malfunctions during replacement operations.
Solution Approach 2:
The system continuously monitors SMART information from storage devices and uses this feedback to dynamically adjust replacement timing determinations. By analyzing trends in SMART data, the system can predict future malfunctions and modify the replacement schedule to avoid concurrent replacements, thereby preventing additional malfunctions during the replacement process.
2Reliability
If replacement timing is determined solely based on predicted malfunction time, then data recovery is maximized, but concurrent replacements may cause system instability
Solution Approach 1:
The system determines replacement timings for multiple storage devices in advance by analyzing SMART information and predicting malfunction times. It then performs a preliminary assessment to identify whether these predicted malfunctions would occur concurrently, and adjusts the replacement schedule beforehand to prevent concurrent replacements, thereby maintaining system stability while preserving data recovery capabilities.
Solution Approach 2:
The system takes preliminary anti-action by identifying potential concurrent malfunction scenarios through SMART analysis and proactively adjusting replacement timings to prevent such concurrency. This preliminary countermeasure eliminates the risk of system instability before it can occur during the replacement process.
3Reliability
If all storage devices are replaced before predicted malfunction, then data loss is prevented, but replacement operations become excessive and costly
Solution Approach 1:
The system changes the parameter of replacement timing from a uniform early replacement approach to a differentiated schedule based on individual storage device characteristics and predicted malfunction times derived from SMART information. By analyzing the remaining life and degradation trends of each device, the system optimizes replacement timing to prevent data loss only when necessary, avoiding excessive replacements and reducing resource waste.
Data Source
AI summary
A storage controlling device including a memory and a processor configured to obtain information on each of a plurality of remaining lives of each of a plurality of storage devices included in a redundancy storage system, determine each of a plurality of timings for replacement of each of the plurality of storage devices so that a number of the timings for replacement included in a predetermined time range is less than a predetermined number, each of a plurality of timings for replacement being determined to be earlier than each of the plurality of timings that malfunctions occur in each of the plurality of storage devices corresponding to each of a plurality of timings for replacement, each of the plurality of timings that malfunctions occur being specified based on the obtained information, and output information that indicates at least one of the plurality of determined timings for replacement.


