RAID Data Redistribution via Drive Health Metrics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage appliances using RAID technology face performance limitations during rebuild processes, which can lead to data loss if additional drives fail before the rebuild is completed, and are vulnerable to simultaneous failures within a Mapped RAID pool.
Innovation Solution
The system collects physical state information from drives to estimate failure probabilities and rearranges data distribution to minimize the likelihood of data unavailability or loss by redistributing data across multiple drives, using techniques such as rearrangement managers to calculate and adjust failure probabilities and data skew.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is rebuilt by copying reconstructed data to a spare drive, then data redundancy is restored, but the rebuild speed is bottlenecked by the write throughput of the spare drive
Solution Approach 1:
The patent segments the rebuild process by distributing reconstructed data across multiple drives simultaneously rather than writing to a single spare drive. This is achieved by identifying multiple candidate drives from the drive pool and parallelizing the write operations, thereby overcoming the write throughput bottleneck of a single drive while restoring data redundancy.
Solution Approach 2:
The patent enables multiple drives to serve the dual function of storing reconstructed data and maintaining data availability. By selecting candidate drives from the existing drive pool rather than requiring dedicated spare drives, the system allows functional drives to universally serve both data storage and redundancy purposes, improving rebuild speed without sacrificing reliability.
2Productivity
If Mapped RAID distributes data across RAID extents made up of disk extents from various physical drives, then rebuild performance is improved by distributing write operations, but the system becomes vulnerable to simultaneous failures of multiple drives
Solution Approach 1:
The patent implements a feedback mechanism by continuously monitoring drive health metrics and failure probabilities. Before selecting candidate drives for receiving reconstructed data, the system evaluates current drive conditions and adjusts the selection to avoid drives with high failure probabilities. This feedback loop ensures that parallel rebuild operations do not inadvertently increase vulnerability to simultaneous failures.
Solution Approach 2:
The patent performs preliminary assessment of drive health and failure probabilities before initiating the rebuild process. By pre-evaluating the reliability of candidate drives and selecting those with lower failure risks, the system proactively prevents the exacerbation of simultaneous failure vulnerabilities while maintaining high rebuild performance through parallel operations.
3Quantity of substance
If the system uses a single-fault-tolerant scheme like RAID-5, then storage capacity is maximized, but data loss occurs if a second drive fails before the first drive is rebuilt
Solution Approach 1:
The patent performs preliminary selection of candidate drives based on health metrics and failure probabilities before the rebuild process begins. By pre-identifying drives with low failure risks and ensuring they are not currently experiencing issues, the system maintains single-fault tolerance while minimizing the risk of data loss during rebuild operations.
Solution Approach 2:
The patent continuously monitors drive health during the rebuild process and dynamically adjusts the rebuild operation based on real-time feedback. If a drive shows signs of failure or unexpected issues arise, the system can pause or redirect the rebuild operation to prevent data loss, thereby maintaining the benefits of RAID-5 capacity while reducing data loss risk.
Data Source
AI summary
Techniques are presented for maintaining data distributed across a plurality of storage drives (drives) in a robust manner. A method includes (a) collecting physical state information from each drive of the plurality of drives, (b) generating a predicted failure probability of each drive based on the collected physical state information from that drive, the predicted failure probability indicating a likelihood that that drive will fail within a predetermined period of time, and (c) rearranging a distribution of data across the plurality of drives to minimize a probability of DU/DL. Systems, apparatuses, and computer program products for performing similar methods are also provided.


