Cross-RAID Idle Space Allocation for Failed Stripe Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing storage systems face challenges in effectively managing failed stripes due to depletion of reserved idle spaces, leading to insufficient capacity for data reconstruction and potential data loss when multiple storage devices fail.

Innovation Solution

A method and device that determine a failed stripe in a storage system, identify idle space for reconstruction across different sets of storage devices in a redundant array of independent disks, and reconstruct the failed stripe to ensure data integrity and availability, even when initial idle spaces are insufficient, by utilizing idle spaces in a second set of storage devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If reserved idle parts are set in each storage device to enable reconstruction operation, then data reliability is improved, but idle space becomes depleted during system operation

Engineering Contradiction:
Improvedata reliabilityVSAvoididle space
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent introduces a new dimension of idle space management by allowing idle space to be allocated across different RAID groups. When a storage device fails, the system can borrow idle space from other RAID groups to perform reconstruction, thereby resolving the contradiction between maintaining data reliability and preserving sufficient idle space for future failures.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent makes idle space universal across different RAID groups, allowing it to serve multiple purposes: maintaining data reliability within the local RAID group and providing reconstruction capacity across the entire storage system. This multi-functionality of idle space resolves the contradiction by enabling the same resource to serve both reliability and capacity needs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If idle space is depleted in the first set of storage devices, then reconstruction operation cannot be performed, but data loss occurs

Engineering Contradiction:
Improvedata integrityVSAvoidreconstruction capability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent introduces a cross-RAID group idle space allocation mechanism as an intermediary, allowing the system to perform reconstruction operations even when local idle space is depleted. The controller can allocate idle space from other RAID groups to the affected RAID group, enabling reconstruction to proceed and preventing data loss.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent establishes a preliminary mechanism for allocating idle space across RAID groups before failures occur. The system pre-configures the ability to borrow idle space from other RAID groups, so when a failure occurs and local idle space is insufficient, the reconstruction operation can immediately proceed using the pre-arranged idle space allocation.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If multiple storage devices fail simultaneously, then reserved idle space is insufficient, but system reliability deteriorates

Engineering Contradiction:
Improvefault toleranceVSAvoidfailure handling capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent extends the idle space management from a single RAID group dimension to a system-wide dimension spanning multiple RAID groups. This allows the system to handle multiple simultaneous failures by allocating idle space from unaffected RAID groups to affected ones, thereby maintaining fault tolerance and improving adaptability to various failure scenarios.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11422909B2Method, device, and storage medium for managing stripe in storage system
Publication Date: 2022.08.23 EMC IP HLDG CO LLC
  • US11422909B2 patent drawing
  • US11422909B2 patent drawing
  • US11422909B2 patent drawing

AI summary

When managing stripes in a storage system, based on a determination that a failed storage device appears in first storage devices, a failed stripe involving the failed storage device is determined in a first redundant array of independent disks (RAID). An idle space that can be used to reconstruct the failed stripe is determined in the first storage devices. The failed stripe is reconstructed to second storage devices in the storage system based on a determination that the idle space is insufficient to reconstruct the failed stripe, the second storage devices being storage devices in a second RAID. An extent in the failed stripe is released in the first storage devices. Accordingly, it is possible to reconstruct a failed stripe as soon as possible to avoid data loss, and further to provide more idle spaces in the first storage devices for future reconstruction.