Cooperative SSD Error Recovery for High Bit-Error Flash
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed computing systems face challenges with high latency due to the use of advanced flash memory technologies like TLC and QLC flash, which are prone to higher bit-error rates, leading to increased complexity and cost in error-correcting code methods.
Innovation Solution
A cooperative error correction approach where data storage drives and a storage host work together to correct errors, using iterative processes between intra-drive and inter-drive ECC schemes, allowing the host to manage latency and select appropriate error correction processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If robust error-correcting code methods (such as LDPC code) are employed to correct higher bit-error rates in TLC flash and QLC flash, then error correction capability is improved, but latency increases and circuit complexity increases leading to higher cost and power consumption
Solution Approach 1:
The patent segments the error correction function into two parts: intra-drive ECC implemented in the SSD controller and inter-drive ECC implemented in the storage host. This segmentation allows the SSD to use simpler, faster correction methods while the host handles more complex correction algorithms, thereby reducing the circuit complexity and power consumption in the SSD while maintaining overall error correction capability.
Solution Approach 2:
The patent introduces an intermediary error correction mechanism that operates between the SSD and the storage host. When intra-drive ECC fails to correct errors, the system uses inter-drive ECC algorithms in the host as a mediator to recover data by combining information from multiple drives. This intermediary approach reduces the need for complex ECC circuits within the SSD itself.
2Reliability
If robust error-correcting code methods (such as LDPC code) are employed to correct higher bit-error rates in TLC flash and QLC flash, then error correction capability is improved, but power consumption increases
Solution Approach 1:
The patent segments the error correction function into two parts: intra-drive ECC implemented in the SSD controller and inter-drive ECC implemented in the storage host. This segmentation allows the SSD to use simpler, faster correction methods while the host handles more complex correction algorithms, thereby reducing the circuit complexity and power consumption in the SSD while maintaining overall error correction capability.
Solution Approach 2:
The patent employs a hierarchical error correction strategy where simpler, lower-power intra-drive ECC is used first for routine error correction. Only when this fails does the system engage the more power-intensive inter-drive ECC processes. This approach minimizes overall power consumption by using the more expensive computational resources only when absolutely necessary.
3Productivity
If advanced flash memory technologies (TLC flash, QLC flash) are used to improve storage capacity and performance, then storage density and speed are improved, but bit-error rates increase
Solution Approach 1:
The patent merges multiple error correction mechanisms into a unified system: intra-drive ECC for basic correction within each drive, and inter-drive ECC that combines information across multiple drives. This combination allows the system to handle the higher bit-error rates of advanced flash memory technologies while maintaining the performance benefits of TLC and QLC flash.
Solution Approach 2:
The patent implements prior cushioning by pre-establishing redundant data representations across multiple drives through inter-drive ECC encoding. When errors occur in advanced flash memory, this pre-established redundancy provides a cushion that enables error recovery without requiring complex real-time correction algorithms in the SSD controller.
4Reliability
If complex error-correcting code algorithms are implemented in the SSD controller, then error correction capability is improved, but latency increases
Solution Approach 1:
The patent segments the error correction function into two parts: intra-drive ECC implemented in the SSD controller and inter-drive ECC implemented in the storage host. This segmentation allows the SSD to use simpler, faster correction methods while the host handles more complex correction algorithms, thereby reducing the circuit complexity and power consumption in the SSD while maintaining overall error correction capability.
Solution Approach 2:
The patent applies preliminary action by pre-calculating and storing parity information across multiple drives during the data writing phase. When errors occur during reading, this pre-prepared information enables faster error correction without requiring complex real-time computations in the SSD controller, thereby reducing latency.
Data Source
AI summary
In a network storage device that includes a plurality of data storage drives, error correction and/or recovery of data stored on one of the plurality of data storage drives is performed cooperatively by the drive itself and by a storage host that is configured to manage storage in the plurality of data storage drives. When an error-correcting code (ECC) operation performed by the drive cannot correct corrupted data stored on the drive, the storage host can attempt to correct the corrupted data based on parity and user data stored on the remaining data storage drives. In some embodiments, data correction can be performed iteratively between the drive and the storage host. Furthermore, the storage host can control latency associated with error correction by selecting a particular error correction process.


