Dynamic Error Sensitivity in Mapped RAID Storage Arrays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage appliances using RAID technology face performance limitations and risk of data loss due to the speed bottleneck in rebuilding processes and simultaneous drive failures, especially when multiple drives experience error rates leading to failure at the same time.
Innovation Solution
The method involves adjusting error sensitivity settings in a mapped RAID pool to minimize the likelihood of additional drive failures during rebuild or proactive copying processes, by dynamically changing error weights based on the state of drives within the pool, thereby reducing the risk of data loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If error sensitivity setting is increased to detect failing drives earlier, then proactive copying can be performed earlier, but the likelihood of false positives increases causing unnecessary interruptions
Solution Approach 1:
The system dynamically changes the error sensitivity parameter based on the operational context. When a drive is being proactively copied or rebuilt, the error sensitivity threshold is adjusted to prevent false positives that would trigger unnecessary interruptions. The controller monitors error rates and compares them against dynamically adjusted thresholds rather than fixed thresholds, allowing early detection of true failures while tolerating temporary error spikes during critical operations.
2Reliability
If multiple drives are monitored with high error sensitivity simultaneously, then more failing drives can be detected, but the risk of multiple simultaneous failures increases data loss probability
Solution Approach 1:
The system performs preliminary proactive copying of data from drives showing early signs of failure before they completely fail. By detecting error rate increases and initiating copying operations in advance, the system ensures data is secured on healthy drives before the failing drive becomes completely unavailable. This preliminary action prevents data loss even if multiple drives fail simultaneously, as the copying process is already underway or completed for at-risk drives.
Solution Approach 2:
The system creates a protective cushion by maintaining copies of critical data on multiple healthy drives before failures occur. When error rates indicate impending failure, the system prioritizes copying data from those drives to create redundant copies that serve as a cushion against potential data loss. This cushioning effect ensures that even if multiple drives fail simultaneously, the data remains protected through pre-established redundancy.
3Productivity
If proactive copying is performed on drives with moderate error rates, then performance bottleneck is avoided, but the drive may fail during copying process
Solution Approach 1:
The system dynamically adjusts error sensitivity thresholds based on the current operational state of drives. During proactive copying operations, the error threshold is raised to allow the copying process to complete without being interrupted by temporary error spikes. The controller monitors whether error rates are increasing or stable, and only triggers interruptions when errors are genuinely increasing, allowing moderate error rates to be tolerated during critical copying operations.
4Productivity
If error sensitivity is reduced during rebuild operations, then false positives are minimized, but actual failures may be missed
Solution Approach 1:
The system applies different error sensitivity settings to different drives based on their individual states and operational context. Drives currently undergoing proactive copying or rebuild operations receive elevated error thresholds to prevent interruptions, while other drives in the array maintain normal monitoring sensitivity. The controller identifies which drives are in critical operations and applies localized parameter adjustments only to those drives, maintaining high detection sensitivity for drives not currently being operated on.
Data Source
AI summary
A method is performed by an extent pool manager running on a data storage device. It is configured to manage assignment of disk extents provided by a pool of drives to a set of mapped RAID extents. The method includes (a) receiving an indication that a particular drive has triggered an end-of-life (EOL) condition based on an error count of that drive and a standard sensitivity setting, (b) in response to receiving the indication, changing a sensitivity setting of other drives to be less sensitive than the standard sensitivity setting, and (c) remapping disk extents from the particular drive to the other drives of the pool while the other drives continue operation using the changed sensitivity setting. An apparatus, system, and computer program product for performing a similar method are also provided.


