Remote Storage Degradation Detection via Host Telemetry

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for detecting hard disk drive failures in enterprise systems suffer from high missed-alarm and false-alarm probabilities, leading to potential massive data loss, and existing redundancy solutions are inadequate for preventing data loss due to increased rebuild times and susceptibility to secondary failures, especially in remote storage devices where health information is not readily available.

Innovation Solution

A system that monitors telemetry performance parameters from a host computer system to detect degradation in remote storage devices using a non-linear non-parametric regression technique, such as multivariate state estimation, to generate predicted values and identify deviations, allowing for preemptive actions like replacing or failing over to a redundant device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If internal counter-type variables (read retries, write retries, seek errors, dwell time) are analyzed to detect disk drive failures, then failure detection capability is improved, but missed-alarm probability increases to 50% and false-alarm probability increases to 1%

Engineering Contradiction:
Improvefailure detection capabilityVSAvoidalarm accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent combines multiple internal counter-type variables (read retries, write retries, seek errors, dwell time) into a unified monitoring system that tracks their correlations and patterns. By merging these variables into a composite health assessment model, the system achieves more accurate failure prediction while reducing both missed-alarm and false-alarm probabilities compared to analyzing individual variables separately.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms static counter values into dynamic performance metrics by monitoring their rates of change, correlations, and patterns over time. This parameter transformation enables the system to detect early signs of degradation before actual failures occur, improving detection accuracy while maintaining low alarm probabilities.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If SMART variables are monitored to identify hard disk drive failures, then failure identification capability is improved, but approximately 50% of imminent failures remain undetected

Engineering Contradiction:
Improvefailure identification capabilityVSAvoidfailure detection completeness
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the monitoring approach into multiple layers: internal counter-type variables within the disk drive, SMART variables from the drive controller, and host system performance metrics. By dividing the monitoring function across these segments and analyzing their interrelationships, the system achieves more complete failure detection than monitoring any single segment alone.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements feedback mechanisms that continuously monitor the relationship between host system performance metrics and storage device health indicators. This feedback loop enables the system to adjust detection thresholds and identify patterns that precede failures, improving detection completeness while reducing false alarms.

Inventive Principle:
Principle #23Feedback

3Reliability

If redundant arrays of inexpensive disks (RAID) are used to prevent catastrophic data loss, then data protection capability is improved, but rebuild time increases dramatically and susceptibility to secondary failures remains high

Engineering Contradiction:
Improvedata protection capabilityVSAvoidrebuild time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by continuously monitoring multiple failure indicators and predicting drive failures before they occur. This early warning system allows administrators to proactively replace at-risk drives during scheduled maintenance windows rather than during unexpected failures, thereby avoiding the need for time-critical RAID rebuild operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements beforehand cushioning by detecting degradation patterns that precede actual failures. This early detection creates a buffer period during which preventive replacement can be performed, cushioning the system against the need for emergency rebuild operations and reducing the window of vulnerability to secondary failures.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

4Quantity of substance

If remote storage devices are used to store data, then storage capacity and accessibility are improved, but health monitoring capability deteriorates as information about device health becomes unavailable

Engineering Contradiction:
Improvestorage capacityVSAvoidhealth monitoring capability
Core Design Contradiction:
Quantity of substanceVSDifficulty of detecting and measuring

Solution Approach 1:

The patent uses the host computer system as an intermediary to monitor remote storage device health. By capturing and analyzing performance metrics from the host's perspective (I/O response times, error rates, throughput), the system indirectly monitors remote device health without requiring direct access to internal drive sensors or SMART data, thus overcoming the remoteness barrier.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent makes the host computer system perform multiple functions: it serves as both the computational platform and the monitoring system for remote storage devices. By universally utilizing existing host resources (CPU, memory, I/O controllers) for dual purposes of data processing and health monitoring, the system eliminates the need for separate monitoring infrastructure while maintaining comprehensive oversight of remote devices.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7769562B2Method and apparatus for detecting degradation in a remote storage device
Publication Date: 2010.08.03 ORACLE AMERICAN INC
  • US7769562B2 patent drawing
  • US7769562B2 patent drawing
  • US7769562B2 patent drawing

AI summary

A system that monitors telemetry from a host computer system to detect degradation in a remote storage device. During operation, the system monitors performance parameters from a host computer system which accesses the remote storage device, wherein the performance parameters relate to the interactions between the host computer system and the remote storage device. The system then determines whether the monitored performance parameters have deviated from predicted values for the performance parameters. If so, the system generates a signal indicating that the remote storage device has degraded.