Storage Area Retirement via ECC Analysis for SSD Degradation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Storage devices, such as SSDs, experience degradation over time due to frequent writes and erases, leading to failed memory cells and blocks, which can affect data integrity and storage performance, as existing technologies lack effective methods to identify and retire failing portions proactively.

Innovation Solution

A storage system with a controller that analyzes individual storage areas, determines failed portions, and retires them using error correction codes (ECC) to ensure data integrity and maintain performance by flagging and terminating I/O operations to failing areas, while periodically testing remaining areas for continued functionality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If storage areas are continuously used for data storage, then storage capacity is maximized, but degradation and failures occur leading to data integrity issues

Engineering Contradiction:
Improvedata integrityVSAvoidstorage area lifespan
Core Design Contradiction:
ReliabilityVSDuration of action of stationary object

Solution Approach 1:

The controller performs periodic testing of storage areas before failures occur. The system proactively identifies degraded storage areas through read operations and ECC analysis, retiring them before they can cause data integrity issues. This preliminary detection and retirement mechanism prevents failures rather than reacting to them after occurrence.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The storage device performs self-diagnosis and self-management through automated testing and retirement processes. The controller continuously monitors storage area health, identifies failures, and retires problematic areas without external intervention. This self-service approach maintains data integrity by automatically managing the storage media lifecycle.

Inventive Principle:
Principle #25Self-service

2Productivity

If all storage areas are actively used, then storage performance is optimized, but failing areas can compromise overall system reliability

Engineering Contradiction:
Improvestorage performanceVSAvoidsystem reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The storage device divides storage areas into individually manageable units that can be tested and retired independently. When a failure is detected in one storage area, only that specific area is retired while others continue to operate. This segmentation allows the system to maintain high productivity by keeping functional areas active while isolating and removing only the failed segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system applies different operational states to different storage areas based on their individual health status. Functional areas continue to be actively used for data storage, while failed areas are retired and excluded from use. This local quality approach ensures that the failure of one area does not compromise the reliability of the entire system, maintaining optimal performance through selective management.

Inventive Principle:
Principle #3Local quality

3Reliability

If storage areas are tested frequently, then failures are detected early improving reliability, but testing overhead reduces storage productivity

Engineering Contradiction:
Improvefailure detection accuracyVSAvoidstorage throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The controller performs testing of storage areas periodically rather than continuously or on every access operation. This periodic testing approach balances reliability and productivity by detecting failures at appropriate intervals without constantly interfering with data operations. The system schedules tests to minimize impact on storage throughput while maintaining adequate monitoring of storage area health.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS11443826B2Storage area retirement in a storage device
Publication Date: 2022.09.13 SEAGATE TECH LLC
  • US11443826B2 patent drawing
  • US11443826B2 patent drawing
  • US11443826B2 patent drawing

AI summary

Systems and methods presented herein provide for testing degradation in a storage device. In one embodiment, a storage controller is operable to test individual portions of a first of the plurality of storage areas of the storage device by: analyzing individual portions of the first storage area; determining that one or more of the individual portions of the first storage area have failed; and retire the failed one or more portions of the first storage area. The storage controller is further operable to write to the first storage area using an error correction code (ECC), and to test the remaining portions of the first storage area to determine whether the first storage area should be retired in response to writing to the first storage area.