Disk Failure Prediction Using Quantile Distribution Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems lack a reliable mechanism to predict single and multi-disk failures in RAID environments, with existing monitoring technologies like S.M.A.R.T. providing inconsistent and vendor-specific indicators, leading to reactive management and potential data loss.
Innovation Solution
The use of quantile distribution techniques to select and analyze diagnostic parameters such as S.M.A.R.T. attributes and SCSI disk return codes from known failed and working disks to identify the most reliable disk failure indicators, allowing for the prediction of disk failures by determining optimal thresholds and calculating failure probabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Difficulty of detecting and measuring
If S.M.A.R.T. monitoring is used to detect disk failures, then failure detection capability is improved, but reliability of failure prediction deteriorates due to inconsistent and vendor-specific indicators
Solution Approach 1:
The patent transforms multiple S.M.A.R.T. attributes and SCSI return codes from different vendors into a unified set of normalized diagnostic parameters. By changing the parameter representation and applying consistent interpretation rules across vendors, the system achieves reliable failure prediction while maintaining the ability to detect failures through multiple indicators.
Solution Approach 2:
The patent creates a universal failure prediction mechanism that works across different disk vendors and RAID configurations. The system evaluates multiple diagnostic parameters (S.M.A.R.T. attributes and SCSI return codes) together, making the prediction system versatile and vendor-agnostic while improving overall reliability through multi-parameter analysis.
2Measurement precision
If multiple diagnostic parameters are monitored to improve prediction accuracy, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent segments the monitoring task by separating data collection from data analysis. The system collects multiple S.M.A.R.T. attributes and SCSI return codes independently, then applies a structured evaluation methodology that assesses each parameter's contribution to failure prediction. This segmentation allows comprehensive monitoring while maintaining manageable system complexity through modular processing.
Solution Approach 2:
The patent evaluates the contribution of each diagnostic parameter individually to determine its effectiveness in predicting failures. By assessing and potentially eliminating parameters that contribute minimally to prediction accuracy, the system achieves high measurement precision with optimized complexity, avoiding the need to monitor all possible parameters equally.
3Ease of operation
If reactive management is used to respond to failures, then ease of operation is maintained, but loss of time increases due to performance degradation before failure detection
Solution Approach 1:
The patent implements preliminary failure prediction by continuously evaluating diagnostic parameters against failure indicators before actual failures occur. The system proactively identifies disks at risk of failure and alerts administrators in advance, enabling preventive replacement before performance degradation occurs. This preliminary action maintains operational simplicity while eliminating the time loss associated with reactive failure detection.
Data Source
AI summary
Techniques for determining a disk failure indicator for predicting disk failures are described herein. According to one embodiment, diagnostic parameters are received which are collected from a set of known working disks and a set of known failed disks of a storage system. For each of the diagnostic parameters, a first quantile distribution representation is generated for the set of known working disks, and a second quantile distribution representation is generated for the set of known failed disks. The first quantile distribution representation and the second quantile distribution representation of each of the diagnostic parameters are then compared to select one or more of the diagnostic parameters as one or more disk failure indicators for predicting future disk failures.


