Error Correction Metric for Storage Device Health Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data storage systems face challenges in precisely identifying poorly performing data storage devices (DSDs) that exhibit high tail latency, as existing methods rely on resource-intensive approaches and fail to distinguish between internal errors and external latency sources, making it difficult to detect and address these issues in real-time.

Innovation Solution

The implementation of a reliability engine within the data storage system that monitors the health of DSDs by analyzing recovery logs and log sense counters, computing metrics such as Full Recoveries Per Hour (FRPH) and Quality of Service (QoS), to proactively detect and perform in-situ repairs on problematic devices, thereby reducing latency and maintaining system performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If resource-intensive approaches are used to identify poorly performing DSDs, then measurement precision is improved, but productivity deteriorates

Engineering Contradiction:
Improveidentification accuracyVSAvoidsystem throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts the identification of poorly performing DSDs from resource-intensive comprehensive monitoring approaches and isolates it to specific error correction metrics (QEC margins) that can be evaluated with minimal overhead. This allows precise identification of problematic devices without requiring full system-wide resource consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses disposable error correction code margins as a low-cost indicator for identifying failing drives. Instead of implementing expensive, continuous comprehensive monitoring systems, the approach leverages the existing error correction capacity as a disposable metric that naturally degrades and signals drive failure risk without requiring additional resources.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Reliability

If comprehensive monitoring of all DSDs is implemented, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improvesystem reliabilityVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential reliability monitoring function from complex comprehensive monitoring systems and isolates it to tracking error correction code margins. This extracted metric provides sufficient reliability information without requiring the complexity of monitoring multiple parameters simultaneously.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The error correction code margin metric serves multiple functions simultaneously: it indicates drive health status, predicts impending failures, and identifies poorly performing devices. This universal metric replaces the need for multiple specialized monitoring systems, reducing overall device complexity while maintaining reliability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of time

If real-time detection of poorly performing DSDs is achieved, then loss of time is reduced, but measurement precision requirements increase

Engineering Contradiction:
Improvedetection timeVSAvoidmetric accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent implements preliminary monitoring of error correction code margins that continuously track drive degradation before actual failures occur. This preliminary action allows the system to identify poorly performing DSDs in advance, reducing detection time while the accumulated precision data from continuous monitoring meets the accuracy requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes a feedback loop where error correction margin measurements are continuously taken, analyzed, and used to update the identification of poorly performing DSDs. This feedback mechanism enables real-time detection as the metrics naturally update with each error correction event, providing both timely detection and accumulating measurement precision.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11237893B2Use of error correction-based metric for identifying poorly performing data storage devices
Publication Date: 2022.02.01 WESTERN DIGITAL TECHNOLOGIES INC
  • US11237893B2 patent drawing
  • US11237893B2 patent drawing
  • US11237893B2 patent drawing

AI summary

An approach to identifying poorly performing data storage devices (DSDs) in a data storage system, such as hard disk drives (HDDs) and/or solid-state drives (SSDs), involves retrieving and evaluating a respective set of log pages, such as SCSI Log Sense counters, from each of multiple DSDs. Based on each respective set of log pages, a value for a Quality of Service (QoS) metric is determined for each respective DSD, where each QoS value represents an average percentage of bytes processed without the respective DSD performing an autonomous error correction. In response to a particular DSD reaching a predetermined threshold QoS value, an in-situ repair may be determined for the particular DSD or the particular DSD may be added to a list of candidate DSDs for further examination, which may include an FRPH examination for suitably configured DSDs.