RAID Cluster Splitting With Distributed Spares for Parallel Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing RAID systems face inefficiencies in distributing spare capacity and managing drive failures, particularly in scaling up and splitting clusters, which can hinder fast and parallel data recovery.
Innovation Solution
A method and apparatus for creating and distributing spare capacity by forming drive clusters with W+1 drives, organizing them into submatrices, and distributing spare cells across multiple drives, allowing for predictable and efficient recovery from drive failures through rotation and relocation of RAID members.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If spare capacity is concentrated in a single cluster, then management is simplified, but data recovery speed decreases due to sequential reconstruction
Solution Approach 1:
The patent divides the spare capacity into multiple segments distributed across different clusters. Each cluster maintains its own spare capacity independently, allowing parallel reconstruction operations to occur simultaneously in multiple clusters, thereby increasing overall data recovery speed while maintaining manageable complexity through modular organization
Solution Approach 2:
The patent transitions from a single-dimension spare capacity model (one cluster) to a multi-dimensional model where spare capacity exists across multiple clusters simultaneously. This dimensional expansion enables parallel recovery operations across different spatial dimensions (clusters), fundamentally increasing recovery throughput without proportionally increasing management complexity
2Quantity of substance
If the drive cluster is scaled up continuously, then storage capacity increases, but spare capacity distribution becomes inefficient
Solution Approach 1:
The patent segments the continuously growing drive cluster into multiple smaller clusters, each maintaining its own spare capacity. This segmentation allows each sub-cluster to efficiently utilize its spare capacity for local reconstructions, preventing the inefficiency that would occur in a single large cluster where spare capacity would be too dispersed to be effective
Solution Approach 2:
The patent implements a dynamic cluster structure that can be split and reorganized as storage capacity needs grow. Rather than continuously expanding a single static cluster, the system dynamically creates and manages multiple clusters that can be independently optimized, maintaining high spare capacity utilization efficiency throughout the scaling process
3Quantity of substance
If all drives are used for data storage, then storage density maximizes, but fault tolerance decreases
Solution Approach 1:
The patent applies local quality by dedicating specific drives within each cluster to serve as spare capacity rather than using them for general data storage. This creates localized zones of redundancy within the storage system, ensuring that fault tolerance is maintained in specific critical areas while maximizing overall storage density through efficient use of remaining drives
Solution Approach 2:
The patent implements beforehand cushioning by pre-allocating spare capacity drives in each cluster before failures occur. These spare drives act as a cushion or buffer that can immediately absorb and recover from drive failures, maintaining fault tolerance without requiring real-time resource allocation decisions during failure events
Data Source
AI summary
A minimal cluster of W+1 drives for RAID width W includes at least 2*W same-size sequentially indexed cells on each drive. More specifically, each drive is organized into W*(LCM of 2, 3, . . . N, N+1) same-size cells, where N is a predetermined number to support even distribution of spares over the first N cluster splits. Spare cells of aggregate capacity equivalent to a single drive are distributed such that there are W spare cells in each W-cell-wide submatrix. In the first cell index of each of the submatrices, a protection group is created in sequentially indexed drives starting with the lowest indexed drive. Also in each submatrix, W spare cells and (W−1) additional protection groups are distributed in cell-index-wise consecutive runs of a spare cell followed by sequentially ordered members of the additional protection groups. Spare capacity is widely distributed by adding drives to the cluster, creating spare cells, splitting the cluster into first and second clusters, and distributing the spare cells across W drives in each of the first and second clusters.


