Adaptive Failure Prediction Modeling for Storage Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage device failure prediction systems, such as SMART systems, suffer from low accuracy due to the complex nature of user environments and difficulties in setting appropriate thresholds, leading to a high number of false positives and unnecessary replacements.
Innovation Solution
An adaptive failure prediction modeling method that generates an initial failure prediction model based on performance data from nominally identical devices, which is updated using real-world data from field deployments and failure notifications, allowing for enhanced failure prediction and model refinement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional SMART systems are used for failure prediction, then the system can detect potential failures, but the accuracy is low due to complex user environments and difficulty in setting appropriate thresholds, leading to high false positives
Solution Approach 1:
The system segments the failure prediction task into multiple components: (1) local monitoring of device parameters by individual storage devices, (2) collection of performance data from multiple devices, (3) centralized generation of failure prediction models using machine learning, and (4) distribution of updated models back to devices. This segmentation allows complex ML operations to be performed centrally while keeping individual devices simple.
Solution Approach 2:
The system introduces an intermediary FPM generation system that acts as a mediator between raw device performance data and failure predictions. This intermediary collects data from multiple devices, applies complex machine learning algorithms to generate improved FPMs, and distributes them back to devices, thereby resolving the contradiction between prediction accuracy and device complexity.
2Reliability
If failure prediction thresholds are set to be sensitive, then more potential failures are detected, but the number of false positives increases leading to unnecessary replacements
Solution Approach 1:
The system performs preliminary actions by collecting and analyzing performance data from multiple devices before generating failure predictions. The FPM generation system uses machine learning to pre-process data from many devices, identify patterns, and create refined failure prediction models that reduce false positives before deployment to individual devices.
Solution Approach 2:
The system implements feedback loops where performance data and failure notifications from field-deployed devices are continuously collected and fed back to the FPM generation system. This feedback enables iterative refinement of failure prediction models, improving reliability while reducing false positives through continuous learning from real-world data.
3Measurement precision
If failure prediction models are updated frequently, then prediction accuracy improves, but the complexity of model management and data transmission increases
Solution Approach 1:
The system employs periodic action by updating failure prediction models at scheduled intervals rather than continuously. The FPM generation system periodically collects accumulated performance data from field-deployed devices, generates updated models, and distributes them back to devices. This periodic approach balances prediction accuracy improvement with reduced data transmission and processing overhead.
Data Source
AI summary
Method and apparatus for predicting data storage device failures using adaptive failure prediction modeling. In some embodiments, monitored parameters are used to predict a potential imminent failure of a first data storage device using a first copy of a first failure prediction model (FPM). Data associated with the predicted potential imminent failure are transferred by the device across a computer network to a host device. The host device generates an updated, second FPM responsive to the transferred data as well as from data from at least a second data storage device transmitted across the computer network having a second copy of the first FPM. A first copy of the updated, second FPM is transferred, via the network, for use by the second data storage device.


