Hard Disk Failure Prediction via Initial Non-Zero Medium Error Counts

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting hard disk failure, such as using reallocated sectors, are incomplete as not all medium errors result in sector reallocation, necessitating the exploration of additional predictors like initial non-zero medium error counts (NMECs).

Innovation Solution

A system and method that collect and analyze data on initial non-zero medium error counts (NMECs) from multiple hard disks across various manufacturers and models to generate a probability of failure, utilizing a database framework with APIs to process customer reports and historical data, thereby improving the accuracy of failure prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If reallocated sectors are used as predictors, then disk failure prediction is possible, but the prediction is incomplete because not all medium errors result in sector reallocation

Engineering Contradiction:
Improveprediction accuracyVSAvoidfailure prediction completeness
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the medium error detection into multiple independent predictor components: reallocated sector counts, pending sector counts, and initial non-zero medium error counts (NMECs). By dividing the prediction system into these separate segments, each capturing different aspects of disk degradation, the system achieves more complete coverage of failure modes without requiring all errors to result in reallocation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a multi-functional prediction system that uses multiple predictor types serving different functions. Reallocated sectors capture one aspect of degradation, pending sectors capture another, and NMECs capture immediate medium errors before reallocation occurs. This universal approach allows the system to detect various failure modes through different mechanisms, improving overall prediction completeness.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If only reallocated sectors are monitored, then the monitoring system remains simple, but early failure detection is limited

Engineering Contradiction:
Improvemonitoring system complexityVSAvoidearly failure detection capability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements preliminary action by monitoring initial non-zero medium error counts (NMECs) before sector reallocation occurs. This early detection mechanism captures medium errors at their initial stage, providing advance warning of potential failures. By acting preliminarily—detecting errors before they propagate to reallocation—the system enhances early failure detection without requiring complex real-time analysis of ongoing reallocation processes.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent establishes feedback loops that continuously monitor multiple predictor metrics (reallocated sectors, pending sectors, NMECs) and use this information to update failure probability assessments. This feedback mechanism allows the system to adapt to changing disk conditions and provide progressively more accurate predictions as degradation progresses, improving reliability while maintaining manageable complexity through systematic information processing.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9612896B1Prediction of disk failure
Publication Date: 2017.04.04 DELL EMC
  • US9612896B1 patent drawing
  • US9612896B1 patent drawing
  • US9612896B1 patent drawing

AI summary

Systems and methods are disclosed for predicting failure of a hard disk in a storage system. Embodiments are disclosed that predict failure of at least one hard disk in a storage system having a plurality hard disks. A data center reports to a data collection center than a hard disk has reported an initial non-zero medium error count (NMEC). The data collection center stores historic data of initial NMEC for many hard disks, and subsequent failure of those hard disks. From the historic data, the data collection center can report to the data center a prediction of when a hard disk reporting an initial NMEC may fail. Different models of hard disks fail at different times relative to a reported initial NMEC. The data collection center can track historic hard disk data by manufacturer, model of hard disk, and by model of storage system and thus can predict, by hard disk model, a probability of failure of a hard disk.