Cooperative SSD Error Recovery for High Bit-Error Flash

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed computing systems face challenges with high latency due to the use of advanced flash memory technologies like TLC and QLC flash, which are prone to higher bit-error rates, leading to increased complexity and cost in error-correcting code methods.

Innovation Solution

A cooperative error correction approach where data storage drives and a storage host work together to correct errors, using iterative processes between intra-drive and inter-drive ECC schemes, allowing the host to manage latency and select appropriate error correction processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If robust error-correcting code methods (such as LDPC code) are employed to correct higher bit-error rates in TLC flash and QLC flash, then error correction capability is improved, but latency increases and circuit complexity increases leading to higher cost and power consumption

Engineering Contradiction:
Improveerror correction capabilityVSAvoidcircuit complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the error correction function into two parts: intra-drive ECC implemented in the SSD controller and inter-drive ECC implemented in the storage host. This segmentation allows the SSD to use simpler, faster correction methods while the host handles more complex correction algorithms, thereby reducing the circuit complexity and power consumption in the SSD while maintaining overall error correction capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary error correction mechanism that operates between the SSD and the storage host. When intra-drive ECC fails to correct errors, the system uses inter-drive ECC algorithms in the host as a mediator to recover data by combining information from multiple drives. This intermediary approach reduces the need for complex ECC circuits within the SSD itself.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If robust error-correcting code methods (such as LDPC code) are employed to correct higher bit-error rates in TLC flash and QLC flash, then error correction capability is improved, but power consumption increases

Engineering Contradiction:
Improveerror correction capabilityVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by stationary object

Solution Approach 1:

The patent segments the error correction function into two parts: intra-drive ECC implemented in the SSD controller and inter-drive ECC implemented in the storage host. This segmentation allows the SSD to use simpler, faster correction methods while the host handles more complex correction algorithms, thereby reducing the circuit complexity and power consumption in the SSD while maintaining overall error correction capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs a hierarchical error correction strategy where simpler, lower-power intra-drive ECC is used first for routine error correction. Only when this fails does the system engage the more power-intensive inter-drive ECC processes. This approach minimizes overall power consumption by using the more expensive computational resources only when absolutely necessary.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Productivity

If advanced flash memory technologies (TLC flash, QLC flash) are used to improve storage capacity and performance, then storage density and speed are improved, but bit-error rates increase

Engineering Contradiction:
Improvestorage performanceVSAvoidbit-error rate
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent merges multiple error correction mechanisms into a unified system: intra-drive ECC for basic correction within each drive, and inter-drive ECC that combines information across multiple drives. This combination allows the system to handle the higher bit-error rates of advanced flash memory technologies while maintaining the performance benefits of TLC and QLC flash.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements prior cushioning by pre-establishing redundant data representations across multiple drives through inter-drive ECC encoding. When errors occur in advanced flash memory, this pre-established redundancy provides a cushion that enables error recovery without requiring complex real-time correction algorithms in the SSD controller.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

4Reliability

If complex error-correcting code algorithms are implemented in the SSD controller, then error correction capability is improved, but latency increases

Engineering Contradiction:
Improveerror correction capabilityVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the error correction function into two parts: intra-drive ECC implemented in the SSD controller and inter-drive ECC implemented in the storage host. This segmentation allows the SSD to use simpler, faster correction methods while the host handles more complex correction algorithms, thereby reducing the circuit complexity and power consumption in the SSD while maintaining overall error correction capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-calculating and storing parity information across multiple drives during the data writing phase. When errors occur during reading, this pre-prepared information enables faster error correction without requiring complex real-time computations in the SSD controller, thereby reducing latency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9946596B2Global error recovery system
Publication Date: 2018.04.17 KIOXIA CORP
  • US9946596B2 patent drawing
  • US9946596B2 patent drawing
  • US9946596B2 patent drawing

AI summary

In a network storage device that includes a plurality of data storage drives, error correction and/or recovery of data stored on one of the plurality of data storage drives is performed cooperatively by the drive itself and by a storage host that is configured to manage storage in the plurality of data storage drives. When an error-correcting code (ECC) operation performed by the drive cannot correct corrupted data stored on the drive, the storage host can attempt to correct the corrupted data based on parity and user data stored on the remaining data storage drives. In some embodiments, data correction can be performed iteratively between the drive and the storage host. Furthermore, the storage host can control latency associated with error correction by selecting a particular error correction process.