Parity Declustered RAID Distributed Hot Sparing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional RAID systems face performance degradation and prolonged recovery times during disk failures due to limited I/O bandwidth and reliance on a dedicated hot sparing drive, which can lead to data loss if additional failures occur before spare disks are replaced.
Innovation Solution
Implementing a parity declustered RAID organization with distributed hot sparing, where each physical drive reserves hot spare space for data reconstruction, allowing concurrent rebuilding across multiple drives and optimizing spare space allocation to reduce bottlenecks and minimize storage waste.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a dedicated hot sparing drive is used for data reconstruction, then data redundancy is maintained, but I/O bandwidth is limited and recovery time is prolonged
Solution Approach 1:
The patent divides the hot sparing function from a single dedicated drive and distributes it across multiple drives. Each drive maintains its own hot spare space, enabling parallel reconstruction operations. This segmentation of the reconstruction function across multiple drives simultaneously improves reliability maintenance while reducing recovery time through concurrent operations.
Solution Approach 2:
The patent transitions from a single-dimension reconstruction approach (one dedicated hot spare drive) to a multi-dimensional approach where multiple drives participate in reconstruction simultaneously. By adding the dimension of parallel reconstruction across multiple drives, the system maintains data redundancy while significantly reducing the time required for recovery.
2Reliability
If all surviving disks are required for providing redundant data, then data integrity is ensured, but I/O bandwidth is limited during reconstruction
Solution Approach 1:
The patent segments the reconstruction process so that not all surviving drives need to participate simultaneously. By dividing the reconstruction workload and allowing selective participation of drives, the system ensures data integrity through redundancy while freeing up I/O bandwidth for other operations. This segmentation enables parallel processing and reduces the bottleneck effect.
Solution Approach 2:
The patent implements partial action by allowing reconstruction to proceed with a subset of surviving drives rather than requiring all of them. This partial participation in the reconstruction process maintains data integrity while preserving I/O bandwidth for other system operations, thereby improving overall productivity during the reconstruction phase.
3Productivity
If hot spare disk is pre-configured as replacement disk, then replacement speed is improved, but reconstruction time remains long due to limited bandwidth
Solution Approach 1:
The patent segments the reconstruction task across multiple drives instead of concentrating it on a single hot spare drive. This segmentation allows the reconstruction process to utilize the combined bandwidth of multiple drives simultaneously, maintaining the fast replacement capability while dramatically reducing the overall reconstruction time through parallel processing.
Solution Approach 2:
The patent merges the reconstruction function across multiple drives rather than relying on a single dedicated hot spare. By combining the resources and bandwidth of multiple drives for the reconstruction process, the system achieves both fast replacement (through pre-configured hot spares) and reduced reconstruction time (through aggregated bandwidth).
Data Source
AI summary
A network storage server implements a method to maintain a parity declustered RAID organization with distributed hot sparing. The parity declustered RAID organization, which provides data redundancy for a network storage system, is configured as a RAID organization with a plurality of logical drives. The RAID organization is then distributed in a parity declustered fashion to a plurality of physical drives in the network storage system. The RAID organization also has a spare space pre-allocating on each of the plurality of physical drives. Upon failure of one of the plurality of physical drives, data stored in the failed physical drives can be reconstructed and stored to spare space of the surviving physical drives. After reconstruction, the plurality of logical drives remains parity-declustered on the plurality of surviving physical drives.


