Error Correction Metric for Storage Device Health Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face challenges in precisely identifying poorly performing data storage devices (DSDs) that exhibit high tail latency, as existing methods rely on resource-intensive approaches and fail to distinguish between internal errors and external latency sources, making it difficult to detect and address these issues in real-time.
Innovation Solution
The implementation of a reliability engine within the data storage system that monitors the health of DSDs by analyzing recovery logs and log sense counters, computing metrics such as Full Recoveries Per Hour (FRPH) and Quality of Service (QoS), to proactively detect and perform in-situ repairs on problematic devices, thereby reducing latency and maintaining system performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If resource-intensive approaches are used to identify poorly performing DSDs, then measurement precision is improved, but productivity deteriorates
Solution Approach 1:
The patent extracts the identification of poorly performing DSDs from resource-intensive comprehensive monitoring approaches and isolates it to specific error correction metrics (QEC margins) that can be evaluated with minimal overhead. This allows precise identification of problematic devices without requiring full system-wide resource consumption.
Solution Approach 2:
The patent uses disposable error correction code margins as a low-cost indicator for identifying failing drives. Instead of implementing expensive, continuous comprehensive monitoring systems, the approach leverages the existing error correction capacity as a disposable metric that naturally degrades and signals drive failure risk without requiring additional resources.
2Reliability
If comprehensive monitoring of all DSDs is implemented, then reliability is improved, but device complexity increases
Solution Approach 1:
The patent extracts the essential reliability monitoring function from complex comprehensive monitoring systems and isolates it to tracking error correction code margins. This extracted metric provides sufficient reliability information without requiring the complexity of monitoring multiple parameters simultaneously.
Solution Approach 2:
The error correction code margin metric serves multiple functions simultaneously: it indicates drive health status, predicts impending failures, and identifies poorly performing devices. This universal metric replaces the need for multiple specialized monitoring systems, reducing overall device complexity while maintaining reliability.
3Loss of time
If real-time detection of poorly performing DSDs is achieved, then loss of time is reduced, but measurement precision requirements increase
Solution Approach 1:
The patent implements preliminary monitoring of error correction code margins that continuously track drive degradation before actual failures occur. This preliminary action allows the system to identify poorly performing DSDs in advance, reducing detection time while the accumulated precision data from continuous monitoring meets the accuracy requirements.
Solution Approach 2:
The patent establishes a feedback loop where error correction margin measurements are continuously taken, analyzed, and used to update the identification of poorly performing DSDs. This feedback mechanism enables real-time detection as the metrics naturally update with each error correction event, providing both timely detection and accumulating measurement precision.
Data Source
AI summary
An approach to identifying poorly performing data storage devices (DSDs) in a data storage system, such as hard disk drives (HDDs) and/or solid-state drives (SSDs), involves retrieving and evaluating a respective set of log pages, such as SCSI Log Sense counters, from each of multiple DSDs. Based on each respective set of log pages, a value for a Quality of Service (QoS) metric is determined for each respective DSD, where each QoS value represents an average percentage of bytes processed without the respective DSD performing an autonomous error correction. In response to a particular DSD reaching a predetermined threshold QoS value, an in-situ repair may be determined for the particular DSD or the particular DSD may be added to a list of candidate DSDs for further examination, which may include an FRPH examination for suitably configured DSDs.


