RAID Pseudo-Bad Indicators for Unrecoverable Error Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current RAID systems face disruptions and inability to recover data when encountering unrecoverable errors, leading to data corruption and client access issues due to the need for drastic recovery actions like file system consistency checks.
Innovation Solution
The implementation of pseudo-bad indicators within RAID arrays to mark and protect data blocks with unrecoverable errors, allowing the system to continue serving requests and ensure reliable data recovery by leveraging existing RAID protection mechanisms and redundancy techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If RAID systems use traditional error recovery mechanisms, then data protection is provided, but system availability is reduced due to disruptions and inability to recover data when encountering unrecoverable errors
Solution Approach 1:
The patent applies preliminary action by pre-marking data blocks with pseudo-bad indicators before they become corrupted, and pre-establishing protection mechanisms (parity or mirroring) for these indicators. This allows the system to proactively identify and isolate unrecoverable errors without disrupting ongoing operations, thereby maintaining system availability while ensuring data protection.
Solution Approach 2:
The patent introduces pseudo-bad indicators as intermediary elements between the data blocks and the RAID protection mechanisms. These indicators serve as mediators that communicate the status of data blocks to the RAID system, enabling the system to handle unrecoverable errors gracefully without halting operations, thus resolving the contradiction between data protection and system availability.
2Reliability
If RAID systems perform file system consistency checks to recover from errors, then data integrity is restored, but client access is disrupted and time is lost
Solution Approach 1:
The patent segments the error recovery process into two independent parts: (1) marking individual corrupted data blocks with pseudo-bad indicators, and (2) protecting these indicators using RAID parity or mirroring. This segmentation allows the system to recover data integrity at the block level without requiring time-consuming file system consistency checks, thereby maintaining client access while ensuring data integrity.
Solution Approach 2:
The patent creates copies of the pseudo-bad indicators using RAID protection mechanisms (parity blocks or mirrored blocks). These copies allow the system to preserve information about corrupted blocks without affecting the actual data blocks, enabling integrity restoration without disrupting client access or consuming additional time.
3Reliability
If RAID systems mark data blocks with unrecoverable errors, then data corruption is prevented, but the marked blocks cannot be served to clients
Solution Approach 1:
The patent applies local quality by treating marked data blocks differently from healthy blocks. The pseudo-bad indicators enable the system to locally identify and isolate corrupted blocks, allowing normal operation to continue on unaffected blocks while preventing corrupted blocks from being served to clients. This maintains data corruption prevention without unnecessarily restricting access to the entire storage system.
4Loss of information
If RAID systems use traditional error logging mechanisms, then unrecoverable errors are tracked, but bandwidth is consumed and processing is delayed
Solution Approach 1:
The patent merges the error tracking function with the existing RAID protection infrastructure. By integrating pseudo-bad indicators into the standard RAID parity or mirroring mechanism, the system combines error tracking with data protection operations, eliminating the need for separate error logging processes. This reduces bandwidth consumption and processing delays while maintaining effective error tracking.
Data Source
AI summary
Embodiments of the present invention provide novel, reliable and efficient technique for tracking, tolerating and correcting unrecoverable errors (i.e., errors that cannot be recovered by the existing RAID protection schemes) in a RAID array by reducing the need to perform drastic recovery actions, such as a file system consistency check, which typically disrupts client access to the storage system. Advantageously, ability to tolerate and correct errors in the RAID array beyond the fault tolerance level of the underlying RAID technique increases resiliency and availability of the storage system.


