RAID Spare PE Reallocation for Faster Array Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional data storage systems require large and complex disk farms for scale-up architectures, leading to inefficient spare physical extent (PE) allocation during disk failures, which prolongs degraded RAID array operation.
Innovation Solution
An enhanced spare allocation scheme that reallocates spare PEs across RAID arrays to adhere to the 'array row rule', ensuring efficient utilization and minimizing reconstruction wait times by exchanging spare PEs between arrays.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional spare PE allocation is used in RAID systems, then disk farm capacity can be scaled up, but spare PE utilization becomes inefficient and reconstruction wait time increases
Solution Approach 1:
The patent implements dynamic spare PE allocation where the system continuously monitors disk failure events and reallocates spare PEs from arrays that do not need them immediately to arrays that require reconstruction. This dynamic adjustment allows the spare PE pool to be flexibly assigned based on real-time needs, reducing waiting time while maintaining overall system capacity.
Solution Approach 2:
The patent introduces a cross-array dimension to spare PE management by allowing spare PEs to be shared and reallocated across multiple RAID arrays simultaneously. Instead of dedicating spare PEs to single arrays, the system manages spare PEs at a system-wide level, enabling one spare PE to serve multiple arrays sequentially or concurrently, thus improving utilization efficiency and reducing reconstruction delays.
2Reliability
If spare PEs are dedicated to individual RAID arrays, then array reliability is maintained, but spare PE utilization efficiency decreases
Solution Approach 1:
The patent makes spare PEs universal by allowing them to serve multiple RAID arrays rather than being dedicated to a single array. A spare PE can be allocated to different arrays based on failure events, enabling the same physical resource to fulfill reconstruction needs across multiple logical units. This multi-functionality maintains reliability for each array while dramatically improving overall spare PE utilization efficiency.
Solution Approach 2:
The system implements self-service through automated spare PE reallocation where the RAID management system automatically detects disk failures, identifies available spare PEs from other arrays, and performs reallocation without manual intervention. This self-service mechanism ensures that reliability requirements are met while maximizing the productive use of spare PEs across the entire disk farm.
Data Source
AI summary
A physical extent manager (PEM) receives a first message indicating a first PE of a first redundant array of independent disks (RAID) array of a disk has failed. The PEM determines a plurality of spare physical extents (PEs) available for reconstruction, the plurality of spare PEs including a first spare PE and a second spare PE. The PEM determines that the first spare PE has previously been assigned to a second PE of a second RAID array for reconstruction. The PEM determines that the first PE is able to use the first spare PE for reconstruction as the first spare PE meets the array row rule for the first PE and that the second PE is able to use the second spare PE for reconstruction as the second spare PE meets the array row rule for the second PE.


