Runtime DRAM Row Repair in SSDs Using Accumulated Error Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Solid state drives (SSDs) based on flash memory face reliability and lifespan issues due to faults in dynamic random access memory (DRAM) devices, which can lead to operational failures after manufacturing, especially when faults occur post-manufacturing and during user deployment.

Innovation Solution

A storage device and operation method that includes a nonvolatile memory device, a DRAM device, and a storage controller with a DRAM error correction unit and a repair manager to detect and perform runtime repairs on fail rows in the DRAM device, ensuring continued operation and improved reliability and lifespan.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional fault repair schemes are used during manufacturing, then manufacturing precision is improved, but the device cannot handle faults that occur after manufacturing and during operation

Engineering Contradiction:
Improvefault repair during manufacturingVSAvoidoperational reliability after manufacturing
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent implements preliminary error detection and information collection mechanisms during normal operation. The storage controller continuously monitors and accumulates error information in a dedicated memory area before actual failures occur, enabling proactive repair preparation rather than reactive fault handling

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent allocates reserved row areas and error information storage spaces in advance. When faults are detected, the system can immediately switch to reserved resources without interruption, cushioning the impact of failures on operational reliability

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

2Productivity

If the storage device operates continuously without repair operations, then productivity is maintained, but faults accumulate and lead to operational failure

Engineering Contradiction:
Improvecontinuous operation capabilityVSAvoidoperational stability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent enables continuous error information collection and processing during normal storage operations. The repair manager and error correction units operate concurrently with data read/write operations, maintaining productivity while continuously monitoring for faults through accumulated error information

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The patent implements a feedback mechanism where error information is continuously collected during operation, processed by the repair manager, and used to trigger repair operations. This closed-loop feedback system maintains reliability by responding to accumulated error data while preserving continuous operation capability

Inventive Principle:
Principle #23Feedback

3Reliability

If runtime repair operations are performed, then reliability is improved, but device complexity increases due to additional error correction units and repair management mechanisms

Engineering Contradiction:
Improveoperational reliability through runtime repairVSAvoidcomplexity of error correction and repair management
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges the error correction unit and repair manager functions within the existing storage controller architecture. By integrating these repair functions into the controller rather than adding separate external components, the system achieves runtime repair capability while minimizing the increase in overall device complexity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The storage controller is designed with multi-functionality, serving both as the primary control unit for normal operations and as the error correction and repair management system. This universal design reduces device complexity by eliminating the need for separate dedicated repair hardware

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11302414B2Storage device that performs runtime repair operation based on accumulated error information and operation method thereof
Publication Date: 2022.04.12 SAMSUNG ELECTRONICS CO LTD
  • US11302414B2 patent drawing
  • US11302414B2 patent drawing
  • US11302414B2 patent drawing

AI summary

A storage device including a nonvolatile memory device, a dynamic random access memory (DRAM) device, and a storage controller, an operation method of the storage device including performing an access operation on the DRAM device, collecting accumulated error information about the DRAM device based on the access operation, detecting a fail row of the DRAM device based on the accumulated error information, and performing a runtime repair operation on the detected fail row.