Remote Storage Degradation Detection via Host Telemetry
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for detecting hard disk drive failures in enterprise systems suffer from high missed-alarm and false-alarm probabilities, leading to potential massive data loss, and existing redundancy solutions are inadequate for preventing data loss due to increased rebuild times and susceptibility to secondary failures, especially in remote storage devices where health information is not readily available.
Innovation Solution
A system that monitors telemetry performance parameters from a host computer system to detect degradation in remote storage devices using a non-linear non-parametric regression technique, such as multivariate state estimation, to generate predicted values and identify deviations, allowing for preemptive actions like replacing or failing over to a redundant device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If internal counter-type variables (read retries, write retries, seek errors, dwell time) are analyzed to detect disk drive failures, then failure detection capability is improved, but missed-alarm probability increases to 50% and false-alarm probability increases to 1%
Solution Approach 1:
The patent combines multiple internal counter-type variables (read retries, write retries, seek errors, dwell time) into a unified monitoring system that tracks their correlations and patterns. By merging these variables into a composite health assessment model, the system achieves more accurate failure prediction while reducing both missed-alarm and false-alarm probabilities compared to analyzing individual variables separately.
Solution Approach 2:
The patent transforms static counter values into dynamic performance metrics by monitoring their rates of change, correlations, and patterns over time. This parameter transformation enables the system to detect early signs of degradation before actual failures occur, improving detection accuracy while maintaining low alarm probabilities.
2Reliability
If SMART variables are monitored to identify hard disk drive failures, then failure identification capability is improved, but approximately 50% of imminent failures remain undetected
Solution Approach 1:
The patent segments the monitoring approach into multiple layers: internal counter-type variables within the disk drive, SMART variables from the drive controller, and host system performance metrics. By dividing the monitoring function across these segments and analyzing their interrelationships, the system achieves more complete failure detection than monitoring any single segment alone.
Solution Approach 2:
The patent implements feedback mechanisms that continuously monitor the relationship between host system performance metrics and storage device health indicators. This feedback loop enables the system to adjust detection thresholds and identify patterns that precede failures, improving detection completeness while reducing false alarms.
3Reliability
If redundant arrays of inexpensive disks (RAID) are used to prevent catastrophic data loss, then data protection capability is improved, but rebuild time increases dramatically and susceptibility to secondary failures remains high
Solution Approach 1:
The patent applies preliminary action by continuously monitoring multiple failure indicators and predicting drive failures before they occur. This early warning system allows administrators to proactively replace at-risk drives during scheduled maintenance windows rather than during unexpected failures, thereby avoiding the need for time-critical RAID rebuild operations.
Solution Approach 2:
The patent implements beforehand cushioning by detecting degradation patterns that precede actual failures. This early detection creates a buffer period during which preventive replacement can be performed, cushioning the system against the need for emergency rebuild operations and reducing the window of vulnerability to secondary failures.
4Quantity of substance
If remote storage devices are used to store data, then storage capacity and accessibility are improved, but health monitoring capability deteriorates as information about device health becomes unavailable
Solution Approach 1:
The patent uses the host computer system as an intermediary to monitor remote storage device health. By capturing and analyzing performance metrics from the host's perspective (I/O response times, error rates, throughput), the system indirectly monitors remote device health without requiring direct access to internal drive sensors or SMART data, thus overcoming the remoteness barrier.
Solution Approach 2:
The patent makes the host computer system perform multiple functions: it serves as both the computational platform and the monitoring system for remote storage devices. By universally utilizing existing host resources (CPU, memory, I/O controllers) for dual purposes of data processing and health monitoring, the system eliminates the need for separate monitoring infrastructure while maintaining comprehensive oversight of remote devices.
Data Source
AI summary
A system that monitors telemetry from a host computer system to detect degradation in a remote storage device. During operation, the system monitors performance parameters from a host computer system which accesses the remote storage device, wherein the performance parameters relate to the interactions between the host computer system and the remote storage device. The system then determines whether the monitored performance parameters have deviated from predicted values for the performance parameters. If so, the system generates a signal indicating that the remote storage device has degraded.


