Storage Controller Fault Recovery and Life Extension

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-end storage devices with high capacity and performance are costly to replace when defective, and existing technologies do not effectively manage faults to prevent device failure and extend the life cycle.

Innovation Solution

A storage device with a controller that detects faults, notifies the host device of recovery schemes, and performs recovery operations to maintain functionality even when a reserved area is depleted, preventing the device from entering a fail state.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a high-end storage device is used with high capacity and performance, then storage performance and capacity are improved, but replacement cost increases significantly when defective

Engineering Contradiction:
Improvestorage performance and capacityVSAvoidreplacement cost
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies preliminary action by performing self-diagnosis and fault detection before the storage device completely fails. The controller proactively identifies potential failures in memory cells, storage areas, or other components, and executes recovery operations (such as data migration to reserved areas or remapping) in advance. This prevents complete device failure and extends operational life, reducing the frequency and cost of replacements while maintaining high storage performance and capacity.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If fault detection and recovery mechanisms are implemented, then device reliability is improved, but device complexity increases

Engineering Contradiction:
Improvedevice reliabilityVSAvoiddevice complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service by enabling the storage device to autonomously detect faults, diagnose issues, and execute recovery operations without external intervention. The controller continuously monitors the health of memory cells, storage areas, and other components, and automatically performs recovery actions such as data migration to reserved areas, remapping of failed blocks, or activation of backup components. This self-diagnosis and self-recovery capability improves device reliability while minimizing the need for complex external monitoring and management systems.

Inventive Principle:
Principle #25Self-service

3Productivity

If the storage device continues to operate after a fault occurs, then productivity is maintained, but the risk of complete failure increases

Engineering Contradiction:
Improveoperational continuityVSAvoidfailure risk
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies beforehand cushioning by allocating reserved areas (spare memory cells, reserved storage areas, or backup components) that are prepared in advance to compensate for potential failures. When a fault is detected in operational memory cells or storage areas, the controller automatically migrates data to these pre-configured reserved areas or remaps logical addresses to healthy physical locations. This cushioning mechanism allows the storage device to continue operating at full productivity while isolating and containing the fault, preventing complete failure even under continued operation.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

Data Source

PatentUS10372558B2Storage device, an operating method of the storage device and an operating method of a computing system including the storage device and a host device
Publication Date: 2019.08.06 SAMSUNG ELECTRONICS CO LTD
  • US10372558B2 patent drawing
  • US10372558B2 patent drawing
  • US10372558B2 patent drawing

AI summary

An operating method of a storage device that includes a nonvolatile memory device and a controller configured to control the nonvolatile memory device, the method including: detecting, by the controller, a fault of the nonvolatile memory device or the controller, notifying, by the controller, a host device of the fault, notifying, by the controller, the host device of one or more recovery schemes for recovering the fault, and recovering, by the controller, the fault in response to a recovery scheme selected by the host device.