Disk Fault Isolator for False Failure Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current software systems for detecting faulty disk drives in data storage systems often result in high false-failure rates, as they rely on static models that fail to account for the dynamic behavior of hardware, software, and operating systems, leading to incorrect identification of faulty components.

Innovation Solution

A disk fault isolator mechanism that evaluates events leading up to disk replacement, using pattern matching algorithms and real-time telemetry data to distinguish between true failures and false alarms, and categorizes disk failures into specific categories based on historical patterns, thereby improving the accuracy of fault detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a static model is used to detect faulty disk drives, then the detection process is simple and fast, but the false-failure rate is high

Engineering Contradiction:
Improvefault detection accuracyVSAvoiddetection system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from a static failure detection model to a dynamic one by continuously monitoring multiple parameters (error rates, performance metrics, environmental conditions) over time. The system adapts to changing system states and uses temporal patterns to distinguish true failures from transient issues, thereby reducing false positives while maintaining manageable complexity through automated dynamic analysis.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms by continuously collecting telemetry data from disk drives and comparing it against learned patterns from historical failure data. The detection algorithm adjusts its thresholds and parameters based on feedback from actual system behavior and verified failure cases, improving accuracy over time while managing complexity through iterative refinement rather than overly complex initial rules.

Inventive Principle:
Principle #23Feedback

2Adaptability or versatility

If a static error threshold model is used, then the detection process is simple, but it cannot adapt to dynamic behavior changes in hardware and software

Engineering Contradiction:
Improveadaptability to dynamic behaviorVSAvoiddetection algorithm complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic adaptability by continuously monitoring multiple parameters (error rates, performance metrics, environmental conditions) over time and adjusting detection thresholds based on observed system behavior patterns. The system evolves its detection criteria to accommodate hardware aging, software updates, and changing workloads without requiring manual reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The detection system performs self-service by automatically learning from historical failure data and telemetry information, adjusting its own detection parameters and thresholds without external intervention. The system autonomously adapts to new failure patterns and system configurations, reducing the need for manual tuning while managing complexity through self-optimization.

Inventive Principle:
Principle #25Self-service

3Productivity

If software frequently replaces disk drives based on static models, then faulty components are removed from service quickly, but unnecessary replacements increase due to false failures

Engineering Contradiction:
Improvecomponent replacement speedVSAvoidreplacement accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary analysis by continuously monitoring multiple parameters and identifying potential failures before they cause system breakdown. By detecting early warning signs and predicting failures in advance, the system can schedule replacements during maintenance windows rather than performing emergency replacements, improving both accuracy and operational efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The replacement decision system uses feedback from verified failure cases and telemetry data to continuously refine its detection accuracy. By learning from actual replacement outcomes and comparing predicted versus actual failures, the system improves its ability to distinguish true failures from false positives, reducing unnecessary replacements while maintaining rapid response to genuine failures.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10223224B1Method and system for automatic disk failure isolation, diagnosis, and remediation
Publication Date: 2019.03.05 DELL EMC
  • US10223224B1 patent drawing
  • US10223224B1 patent drawing
  • US10223224B1 patent drawing

AI summary

According to one embodiment, a test result of a first disk that was removed from a storage system and tested at a remote testing facility is received. A data analysis is performed on operational statistics data associated with the first disk based on one or more predetermined data patterns, where the operational statistics data was periodically collected from the storage system during operations of the storage system. A failure category of the first disk is determined based on the data analysis by comparing the operational statistics data against the predetermined data patterns. At least one of the data patterns is adjusted for subsequent determination of failure categories in view of an analysis result of the analysis, the failure category, and the testing result received from the testing facility.