RAID Array Reconfiguration Using ECC and Parity After Drive Failures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

RAID systems face data loss when the number of defective storage elements falls below a certain threshold, as they are not designed to handle reduced numbers of functional storage elements effectively, leading to a lack of redundancy and increased risk of data loss.

Innovation Solution

The system reconfigures an array of storage elements by generating and storing new ECC blocks and parity data on fewer storage elements, using available data and parity information to regenerate missing data and ensure data integrity even when some storage elements are unavailable.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If RAID systems use a fixed number of storage elements for data protection, then data redundancy is maintained under normal conditions, but data loss occurs when the number of defective storage elements falls below the threshold

Engineering Contradiction:
Improvedata protection capabilityVSAvoidadaptability to reduced storage elements
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic reconfiguration of the RAID array by allowing the system to transition from a static fixed-threshold protection model to a dynamic model where the number of active storage elements can be reduced while maintaining protection capabilities. The controller dynamically adjusts the array configuration by identifying defective elements, regenerating their data using parity information, and redistributing data across fewer remaining elements, thereby adapting the system's structure in response to degradation conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key parameters including the number of active storage elements, the distribution of data and parity across elements, and the protection threshold. By modifying these parameters dynamically, the system can operate with fewer storage elements while maintaining data protection. The controller calculates new parameter settings that ensure the remaining elements can still provide the required redundancy and fault tolerance.

Inventive Principle:
Principle #35Parameter changes

2Duration of action of stationary object

If storage elements are removed from the array to extend device lifespan, then operational duration is extended, but data redundancy and reliability deteriorate

Engineering Contradiction:
Improvedevice lifespanVSAvoiddata redundancy
Core Design Contradiction:
Duration of action of stationary objectVSReliability

Solution Approach 1:

The patent applies the discarding and recovering principle by systematically retiring defective storage elements from the array while recovering their data through regeneration using parity information. The controller identifies elements that have reached their lifespan or are defective, extracts their data using parity calculations, and redistributes it across remaining elements. This allows the system to discard worn elements and extend overall array lifespan while maintaining redundancy through intelligent data redistribution and continued parity protection.

Inventive Principle:
Principle #34Discarding and recovering

3Reliability

If the RAID system maintains a fixed protection threshold, then data integrity is ensured under normal conditions, but the system cannot adapt when storage elements become unavailable

Engineering Contradiction:
Improvedata integrityVSAvoidadaptability to element unavailability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system transitions from static fixed-threshold protection to dynamic adaptive protection. The controller continuously monitors storage element availability and dynamically adjusts the protection mechanism by regenerating data from unavailable elements using parity information and redistributing it across available elements. This dynamic adaptation ensures data integrity is maintained even as the system composition changes due to element failures or retirements.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The RAID system performs self-service by automatically detecting unavailable storage elements, regenerating their data using stored parity information, and redistributing the recovered data across remaining elements without external intervention. This self-healing capability allows the system to maintain data integrity and adapt to element unavailability autonomously, preserving protection capabilities throughout the array's operational lifecycle.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS8738991B2Apparatus, system, and method for reconfiguring an array of storage elements
Publication Date: 2014.05.27 SANDISK TECHNOLOGIES LLC
  • US8738991B2 patent drawing
  • US8738991B2 patent drawing
  • US8738991B2 patent drawing

AI summary

Apparatuses, systems, and methods are disclosed for reconfiguring an array of storage elements. A storage element error module is configured to determine that one or more storage elements in an array of storage elements are in error. An array of storage elements stores a first ECC block and first parity data generated from the first ECC block. A data reconfiguration module is configured to generate a second ECC block comprising at least a portion of data of a first ECC block. A new configuration storage module is configured to store a second ECC block and associated second parity data on fewer storage elements than a number of storage elements in an array.