Disk Sector Failure Prediction with Temporal Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional disk failure determination technologies lack the necessary granularity to identify and address failures at the sector level, leading to insufficient fine-grained processing and difficulty in meeting user and administrator needs.
Innovation Solution
A model training method is employed to analyze disk failure data sets, including background medium scan logs, to predict sector set failures and their types, using a machine learning model like random forest to enhance prediction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional disk failure determination technology is used, then the entire disk is monitored for failures, but the granularity is insufficient to identify specific failed sectors
Solution Approach 1:
The patent segments the disk into multiple sector sets and creates separate failure determining models for each sector set. This allows the system to identify failures at the sector level rather than treating the entire disk as a single unit, thereby improving measurement precision while maintaining manageable complexity through modular processing
Solution Approach 2:
The patent introduces a new dimension of analysis by creating temporal sequences of disk failure data sets and processing them through machine learning models. This transforms the monitoring approach from static snapshot analysis to dynamic temporal pattern recognition, enabling finer-grained failure detection without proportionally increasing processing complexity
2Reliability
If fine-grained sector-level failure identification is implemented, then prediction accuracy improves, but computational requirements increase
Solution Approach 1:
By dividing the disk into multiple sector sets and training separate failure determining models for each sector set, the system achieves fine-grained failure prediction accuracy. The segmentation allows computational resources to be distributed across multiple smaller models rather than requiring one complex model to process entire disk data, thereby improving reliability while managing energy consumption
Solution Approach 2:
The patent performs preliminary actions by collecting disk failure data sets over time periods and pre-processing them before model training. This includes acquiring data sets at multiple time points and preparing them in advance, which reduces the computational burden during actual failure prediction and lowers real-time energy consumption while maintaining high prediction accuracy
3Reliability
If machine learning models are trained on temporal disk failure data, then prediction credibility exceeds 98%, but data processing time increases
Solution Approach 1:
The patent applies preliminary action by collecting and pre-processing disk failure data sets over extended time periods before model training. Data sets are acquired at multiple predetermined time points and prepared in advance, which enables the machine learning models to learn from comprehensive temporal patterns and achieve prediction credibility exceeding 98%, while the pre-processing reduces the actual training time by having data ready in structured formats
Solution Approach 2:
The system implements periodic action by collecting disk failure data sets at predetermined time intervals and updating models periodically rather than continuously. This periodic approach allows the system to maintain high prediction credibility through regular data collection and model updates while minimizing the time loss associated with constant processing, as models are trained on accumulated data from multiple time periods
Data Source
AI summary
Embodiments of the present disclosure relate to a model training method, a failure determining method, an electronic device, and a computer program product. The model training method includes: acquiring a plurality of disk failure data sets collected in a first time period; acquiring another disk failure data set that is collected at a predetermined time point after the first time period and indicates failure information of at least one failed sector set; and training a failure determining model based on the plurality of disk failure data sets and the failure information, so that a probability of matching of predicted failure information at a predetermined time point determined by the trained failure determining model based on the plurality of disk failure data sets and the failure information is greater than a first threshold probability. By using the technical solution of the present disclosure, it is possible to predict the failure information that will occur in the sector set included in a disk based on the disk failure data set associated with a failed sector, so that a user or administrator of the disk can know the failure condition that will occur in the sector set of the disk in advance.


