Storage Rebuild Destination Selection via Availability Counting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems face challenges in efficiently selecting a rebuild destination for failed extents in RAID-based storage systems, leading to inadequate utilization of reserved storage space and inefficient data recovery processes.
Innovation Solution
A method is introduced to detect failed stripes and normal storage devices, determining an availability count for each normal storage device to assess its capacity for rebuilding failed stripes, and selecting a destination based on this count to ensure effective utilization of reserved space.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional RAID rebuild selection methods are used, then data recovery can be performed, but the utilization of reserved storage space is inadequate and data recovery efficiency is low
Solution Approach 1:
The patent introduces a counting mechanism that tracks the number of failed stripes each normal storage device can accommodate. This parameter-based approach transforms the rebuild selection from a simple availability check to an optimized allocation process, where devices are selected based on their remaining capacity to accept failed stripes, thereby improving both space utilization and recovery efficiency
Solution Approach 2:
The patent performs preliminary assessment of each normal storage device's capacity to accept failed stripes before initiating the rebuild process. By pre-calculating the first count (number of allowed failed stripes) for each device, the system can make informed selection decisions that maximize reserved space utilization while ensuring efficient data recovery
2Ease of operation
If multiple failed stripes are rebuilt to the same normal storage device, then the rebuild process is simplified, but the distribution of data across storage devices becomes uneven and reliability decreases
Solution Approach 1:
The patent applies different selection criteria to different normal storage devices based on their individual capacity to accept failed stripes. Each device is evaluated locally with its own first count parameter, allowing the system to distribute failed stripes across multiple devices in an uneven but controlled manner, balancing operational simplicity with data distribution reliability
Solution Approach 2:
The patent implements a dynamic selection process where the first count for each normal storage device is updated as failed stripes are progressively rebuilt. This dynamic adjustment allows the system to adaptively balance the distribution of data across storage devices during the rebuild process, preventing any single device from becoming overloaded while maintaining operational efficiency
Data Source
AI summary
In techniques for selecting a rebuild destination in a storage system, a failed stripe group associated with a failed extent group in a failed storage device among storage devices is detected. A group of normal storage devices other than the failed storage device is determined. Regarding a normal storage device in the group of normal storage devices, a first count for the normal storage device is obtained, the first count representing a number of failed stripes which are allowed to be rebuilt to the normal storage device in the failed stripe group. Based on the first count, a destination storage device is selected from the group of normal storage devices for rebuilding a failed stripe in the failed stripe group. During rebuild, a destination for rebuilding the failed stripe may be effectively selected, and extents in reserved space in the storage system may be more fully utilized.


