Storage Rebuild Destination Selection via Availability Counting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current storage systems face challenges in efficiently selecting a rebuild destination for failed extents in RAID-based storage systems, leading to inadequate utilization of reserved storage space and inefficient data recovery processes.

Innovation Solution

A method is introduced to detect failed stripes and normal storage devices, determining an availability count for each normal storage device to assess its capacity for rebuilding failed stripes, and selecting a destination based on this count to ensure effective utilization of reserved space.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional RAID rebuild selection methods are used, then data recovery can be performed, but the utilization of reserved storage space is inadequate and data recovery efficiency is low

Engineering Contradiction:
Improvedata recovery efficiencyVSAvoidutilization of reserved storage space
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent introduces a counting mechanism that tracks the number of failed stripes each normal storage device can accommodate. This parameter-based approach transforms the rebuild selection from a simple availability check to an optimized allocation process, where devices are selected based on their remaining capacity to accept failed stripes, thereby improving both space utilization and recovery efficiency

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs preliminary assessment of each normal storage device's capacity to accept failed stripes before initiating the rebuild process. By pre-calculating the first count (number of allowed failed stripes) for each device, the system can make informed selection decisions that maximize reserved space utilization while ensuring efficient data recovery

Inventive Principle:
Principle #10Preliminary action

2Ease of operation

If multiple failed stripes are rebuilt to the same normal storage device, then the rebuild process is simplified, but the distribution of data across storage devices becomes uneven and reliability decreases

Engineering Contradiction:
Improverebuild process simplicityVSAvoiddata distribution balance
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies different selection criteria to different normal storage devices based on their individual capacity to accept failed stripes. Each device is evaluated locally with its own first count parameter, allowing the system to distribute failed stripes across multiple devices in an uneven but controlled manner, balancing operational simplicity with data distribution reliability

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements a dynamic selection process where the first count for each normal storage device is updated as failed stripes are progressively rebuilt. This dynamic adjustment allows the system to adaptively balance the distribution of data across storage devices during the rebuild process, preventing any single device from becoming overloaded while maintaining operational efficiency

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11500726B2Method, device, and program product for selecting rebuild destination in storage system
Publication Date: 2022.11.15 EMC IP HLDG CO LLC
  • US11500726B2 patent drawing
  • US11500726B2 patent drawing
  • US11500726B2 patent drawing

AI summary

In techniques for selecting a rebuild destination in a storage system, a failed stripe group associated with a failed extent group in a failed storage device among storage devices is detected. A group of normal storage devices other than the failed storage device is determined. Regarding a normal storage device in the group of normal storage devices, a first count for the normal storage device is obtained, the first count representing a number of failed stripes which are allowed to be rebuilt to the normal storage device in the failed stripe group. Based on the first count, a destination storage device is selected from the group of normal storage devices for rebuilding a failed stripe in the failed stripe group. During rebuild, a destination for rebuilding the failed stripe may be effectively selected, and extents in reserved space in the storage system may be more fully utilized.