RAID Pseudo-Bad Indicators for Unrecoverable Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current RAID systems face disruptions and inability to recover data when encountering unrecoverable errors, leading to data corruption and client access issues due to the need for drastic recovery actions like file system consistency checks.

Innovation Solution

The implementation of pseudo-bad indicators within RAID arrays to mark and protect data blocks with unrecoverable errors, allowing the system to continue serving requests and ensure reliable data recovery by leveraging existing RAID protection mechanisms and redundancy techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If RAID systems use traditional error recovery mechanisms, then data protection is provided, but system availability is reduced due to disruptions and inability to recover data when encountering unrecoverable errors

Engineering Contradiction:
Improvedata protectionVSAvoidsystem availability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-marking data blocks with pseudo-bad indicators before they become corrupted, and pre-establishing protection mechanisms (parity or mirroring) for these indicators. This allows the system to proactively identify and isolate unrecoverable errors without disrupting ongoing operations, thereby maintaining system availability while ensuring data protection.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces pseudo-bad indicators as intermediary elements between the data blocks and the RAID protection mechanisms. These indicators serve as mediators that communicate the status of data blocks to the RAID system, enabling the system to handle unrecoverable errors gracefully without halting operations, thus resolving the contradiction between data protection and system availability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If RAID systems perform file system consistency checks to recover from errors, then data integrity is restored, but client access is disrupted and time is lost

Engineering Contradiction:
Improvedata integrityVSAvoidclient access disruption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the error recovery process into two independent parts: (1) marking individual corrupted data blocks with pseudo-bad indicators, and (2) protecting these indicators using RAID parity or mirroring. This segmentation allows the system to recover data integrity at the block level without requiring time-consuming file system consistency checks, thereby maintaining client access while ensuring data integrity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates copies of the pseudo-bad indicators using RAID protection mechanisms (parity blocks or mirrored blocks). These copies allow the system to preserve information about corrupted blocks without affecting the actual data blocks, enabling integrity restoration without disrupting client access or consuming additional time.

Inventive Principle:
Principle #26Copying

3Reliability

If RAID systems mark data blocks with unrecoverable errors, then data corruption is prevented, but the marked blocks cannot be served to clients

Engineering Contradiction:
Improvedata corruption preventionVSAvoiddata block accessibility
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent applies local quality by treating marked data blocks differently from healthy blocks. The pseudo-bad indicators enable the system to locally identify and isolate corrupted blocks, allowing normal operation to continue on unaffected blocks while preventing corrupted blocks from being served to clients. This maintains data corruption prevention without unnecessarily restricting access to the entire storage system.

Inventive Principle:
Principle #3Local quality

4Loss of information

If RAID systems use traditional error logging mechanisms, then unrecoverable errors are tracked, but bandwidth is consumed and processing is delayed

Engineering Contradiction:
Improveerror trackingVSAvoidprocessing speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent merges the error tracking function with the existing RAID protection infrastructure. By integrating pseudo-bad indicators into the standard RAID parity or mirroring mechanism, the system combines error tracking with data protection operations, eliminating the need for separate error logging processes. This reduces bandwidth consumption and processing delays while maintaining effective error tracking.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS8417987B1Mechanism for correcting errors beyond the fault tolerant level of a raid array in a storage system
Publication Date: 2013.04.09 NETAPP INC
  • US8417987B1 patent drawing
  • US8417987B1 patent drawing
  • US8417987B1 patent drawing

AI summary

Embodiments of the present invention provide novel, reliable and efficient technique for tracking, tolerating and correcting unrecoverable errors (i.e., errors that cannot be recovered by the existing RAID protection schemes) in a RAID array by reducing the need to perform drastic recovery actions, such as a file system consistency check, which typically disrupts client access to the storage system. Advantageously, ability to tolerate and correct errors in the RAID array beyond the fault tolerance level of the underlying RAID technique increases resiliency and availability of the storage system.