Disk Failure Prediction Methods and Equipment
By acquiring disk feature data and using a preset model to predict future failure probabilities and determine failure causes, the problem of insufficient interpretability of disk failures caused by static threshold monitoring is solved. This enables dynamic prediction of disk status and explanation of failure causes, thus improving the user experience.
Patent Information
- Application Number
- CN202511323719.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-09-16
AI Technical Summary
The existing technology, which monitors disk attribute parameters through static thresholds, results in insufficient interpretability of disk failures, cannot dynamically adapt to different working environments and load conditions, and cannot determine the cause of failures.
By acquiring the target feature data of the target disk, the probability of future failures is predicted using a preset disk failure prediction model. When the predicted probability exceeds a threshold, the feature contribution and failure cause are determined, and a preset failure analysis model is used to analyze the failure cause.
It enables the prediction of future disk failure probabilities and the explanation of failure causes, improving the interpretability of disk failures. Users can understand the operating status and potential risks in advance, thus improving the user experience.
Smart Images

Figure CN120832278B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to disk failure prediction methods and devices. Background Technology
[0002] Driven by the rapid development of next-generation information technologies such as the Internet of Things and cloud computing, the global data volume is experiencing explosive growth. Disks bear the responsibility of persistently storing massive amounts of data, and their operational stability directly affects the reliability and data security of the entire computer system. Disk failure can lead to anything from data read / write delays and service interruptions to critical data loss and system crashes.
[0003] Currently, disk fault identification typically involves monitoring disk attribute parameters using static thresholds, triggering fault alarms only when parameters abnormally exceed the thresholds. This results in insufficient interpretability of disk faults. Summary of the Invention
[0004] This application provides a disk failure prediction method and device to at least solve the problem in related technologies where disk attribute parameters are monitored by static thresholds, and fault alarms are triggered only when the parameters abnormally exceed the thresholds, resulting in insufficient interpretability of disk failures.
[0005] Firstly, this application provides a disk failure prediction method, the method comprising:
[0006] Obtain target feature data for the target disk; the target feature data includes multiple feature data; the target feature data reflects the current status of the target disk; the target feature data includes the disk type of the target disk;
[0007] A preset disk failure prediction model is used to predict the target disk based on the target disk's characteristic data, and the target prediction probability is output; the target prediction probability is the probability that the target disk will fail within a future target time period.
[0008] If the target prediction probability is determined to be greater than the first preset threshold, the feature contribution of each feature data is determined based on the preset disk failure prediction model and the target feature data.
[0009] The cause of the target fault is determined by using a preset fault analysis model and based on the target feature data and the feature contribution of each feature data, and then sent to the client.
[0010] Secondly, this application also provides a disk failure prediction device, comprising:
[0011] The acquisition module is used to acquire target feature data of the target disk; the target feature data includes multiple feature data; the target feature data reflects the current status of the target disk; the target feature data includes the disk type of the target disk;
[0012] The prediction module is used to predict the target disk using a preset disk failure prediction model and based on the target disk's characteristic data, and outputs the target prediction probability; the target prediction probability is the probability that the target disk will fail within a future target time period.
[0013] The determination module is used to determine the feature contribution of each feature data based on the preset disk failure prediction model and the target feature data if the predicted probability of the target is greater than the first preset threshold.
[0014] The determination module is also used to determine the cause of the target fault by using a preset fault analysis model and based on the target feature data and the feature contribution of each feature data, and then send it to the client.
[0015] Thirdly, this application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the disk failure prediction method provided in the first aspect.
[0016] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the disk failure prediction method provided in the first aspect.
[0017] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the disk failure prediction method provided in the first aspect.
[0018] The disk failure prediction method and device provided in this application automatically acquires target feature data of the target disk, inputs the target feature data of the target disk into a preset disk failure prediction model for prediction, and outputs the target prediction probability. This enables the prediction of the probability of the target disk failing within a future target time period. When the target prediction probability is determined to be greater than a first preset threshold, the feature contribution of each feature data is determined based on the preset disk failure prediction model and the target feature data. Furthermore, a preset failure analysis model is used, and the target failure cause is determined based on the target feature data and the feature contribution of each feature data. This result is then sent to the client, thus predicting the cause of disk failure. Compared to determining the health status of a disk by checking whether its own attribute parameters exceed a threshold, predicting the probability of the target disk failing within a future target time period based on the target disk's target feature data, as well as predicting the cause of the target failure, allows users to understand the operating status of the target disk in advance and the cause of failure when the failure probability exceeds the threshold. This improves the interpretability of disk failures and helps to enhance the user experience. Attached Figure Description
[0019] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is an application scenario diagram of the disk failure prediction method provided in the embodiments of this application;
[0021] Figure 2 A schematic flowchart of a disk failure prediction method provided in an embodiment of this application;
[0022] Figure 3 A schematic flowchart of a disk failure prediction method provided in another embodiment of this application;
[0023] Figure 4 This is a schematic diagram of the structure of a disk failure prediction device provided in an embodiment of this application;
[0024] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0026] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0027] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] In today's era of booming artificial intelligence technology, disks, as a critical medium for data storage, bear the responsibility of persistently storing massive amounts of data. Their operational stability and reliability directly affect the normal operation and data security of the entire computer system. Faced with ever-increasing data processing demands, disks not only need to provide efficient data read / write services but also must ensure stability during long-term continuous operation. Disk failures can range from causing data read / write delays and service interruptions, affecting the continuity of business processes, to potentially leading to the loss of critical data and system crashes. Therefore, timely and accurate identification of disk health status and early warning of potential failure risks are crucial for ensuring data security and maintaining system stability. Currently, disk failure identification mainly relies on setting fixed thresholds for disk attribute parameters (such as disk self-monitoring, analysis, and reporting technology information) and assessing disk health by monitoring whether these parameters exceed preset thresholds. Because this is based on static thresholds, it cannot dynamically adapt to changes in different working environments and load conditions. Furthermore, it can only identify the current disk failure but cannot determine the underlying cause, resulting in poor interpretability of disk failures.
[0029] Therefore, when facing the aforementioned technical problems, instead of simply monitoring whether the disk's attribute parameters exceed a preset threshold to identify whether a disk has failed, the system acquires the target disk's characteristic data and uses a preset disk failure prediction model to predict the probability of the target disk failing within a future target time period based on this data. It outputs the target prediction probability, and when the target prediction probability exceeds a first preset threshold, it determines the feature contribution of each feature data point based on the preset disk failure prediction model and the target characteristic data. Furthermore, it employs a preset failure analysis model and, based on the target characteristic data and the feature contribution of each feature data point, determines the cause of the target failure and sends it to the client. This allows users to understand the target disk's operating status in advance and the cause of failures where the probability exceeds the threshold, thereby enhancing the explainability of disk failures. Users can understand the potential risks of the disk and the root causes of failures in advance, improving the user experience.
[0030] Figure 1 This diagram illustrates an application scenario of the disk failure prediction method provided in this application. For example... Figure 1 As shown, the application scenario provided in this embodiment includes a server device 10 and a client 20. The disk failure prediction method is applied to the server device 10. The server device 10 obtains target feature data of the target disk from a preset database, and uses a preset disk failure prediction model to predict the target disk based on the target feature data, outputting the target prediction probability. The target feature data includes multiple feature data, which reflect the current state of the target disk and include the disk type of the target disk. Further, if the server device 10 determines that the target prediction probability is greater than a first preset threshold, it determines the feature contribution of each feature data based on the preset disk failure prediction model and the target feature data, and uses a preset failure analysis model to determine the target failure cause based on the target feature data and the feature contribution of each feature data, and sends this result to the client 20. The client 20 displays the received target prediction probability and target failure cause.
[0031] Figure 2 This is a flowchart illustrating a disk failure prediction method provided in an embodiment of this application, as shown below. Figure 2 As shown. The disk failure prediction method provided in this embodiment is applied to server-side devices. The disk failure prediction method provided in this embodiment specifically includes the following steps:
[0032] S201: Obtain target feature data of the target disk.
[0033] The target feature data includes multiple features. These features reflect the current state of the target disk and include its disk type.
[0034] The disk types include mechanical hard disks and solid-state disks.
[0035] The target disk is the disk for which fault prediction is to be performed.
[0036] For example, the target feature data of the target disk includes real-time parameter data of items 5, 187, 188, 197 and 198 of the Self-Monitoring, Analysis, and Reporting Technology (SMART) information corresponding to the target disk.
[0037] The target characteristic data of the target disk also includes the disk's real-time temperature, current vibration peak, current disk read / write operations per second, the number of disk-related errors in the system log in the past hour, the growth rate of the number of remapped sectors in the past 7 days, the number of sectors to be repaired in the past day per hour, the disk temperature rise rate in the past hour, the average change of the vibration peak value of the disk tray accelerometer in the past 3 days, the trend of the number of read / write operations per second in the past 12 hours, and the growth rate of the error log count in the past 12 hours.
[0038] The target disk's characteristic data also includes the peak difference of disk temperature over the past 24 hours, the ratio of peak to mean vibration peak over the past hour, the standard deviation of read / write operations per second over the past 6 hours, and the standard deviation of fan speed over the past 8 hours. The target disk's characteristic data also includes the disk operating temperature within a preset time range, the disk's instantaneous vibration acceleration, the cooling fan speed, the target disk's read / write operations per second, the average temperature when the read / write operations per second exceeded a preset read / write threshold over the past hour, and the average response time when the temperature exceeded a preset temperature threshold over the past hour.
[0039] Optionally, the preset read / write threshold and preset temperature threshold can be set independently according to requirements, and are not limited in this embodiment.
[0040] Among them, the 5th item is the number of remapped sectors, the 187th item is the count of unrepairable errors, the 188th item is the number of command timeouts, the 197th item is the number of sectors currently awaiting repair, and the 198th item is the number of unrepairable sectors detected offline.
[0041] The remapped sectors count records the number of bad sectors that have been remapped. When the disk detects a damaged sector, it marks it as a bad sector and moves the data to a reserved spare sector; this process is called sector remapping. The unrepairable error count records the number of read / write errors that cannot be corrected through retries or other means. The command timeout count indicates the number of times the disk timed out while executing a command. The current unrepairable sectors count records the number of sectors that have been detected as having problems but have not yet been remapped. The offline unrepairable sectors count refers to the number of uncorrectable sectors found during offline disk testing.
[0042] Specifically, in this embodiment, the server device obtains the target feature data of the target disk from a preset database.
[0043] Optionally, the preset time range can be 1 hour, 1 day, 3 days, etc., and this embodiment does not limit it.
[0044] S202: Use a preset disk failure prediction model and make a prediction on the target disk based on the target disk's target feature data, and output the target prediction probability.
[0045] The target prediction probability is the probability that the target disk will fail within a target time period in the future.
[0046] Understandably, assuming the target time period can be 72 hours, the target prediction probability refers to the probability that the target disk will fail within the next 72 hours, starting from the time when the latest target feature data is obtained from the target disk.
[0047] The target time period corresponds to the time period that the preset disk failure prediction model can predict.
[0048] Specifically, in this embodiment, the server device inputs the target feature data of the target disk into a preset disk failure prediction model, so that the preset disk failure prediction model predicts the target disk based on the target feature data of the target disk and outputs the target prediction probability.
[0049] S203: If the target prediction probability is determined to be greater than the first preset threshold, the feature contribution of each feature data is determined based on the preset disk failure prediction model and the target feature data.
[0050] Optionally, the first preset threshold can be 60%, or it can be set independently according to needs; this embodiment does not impose any limitations.
[0051] Among them, feature contribution refers to the degree of influence of each feature data on the probability of disk failure.
[0052] Specifically, in this embodiment, a preset contribution algorithm is used to calculate the feature contribution of each feature data based on the preset disk fault prediction model and the target feature data of the target disk, thereby obtaining the feature contribution of each feature data.
[0053] Optionally, the preset contribution algorithm can be the Shapley Additive Explanations (SHAP) algorithm. The Shapley Additive Explanations (SHAP) algorithm is based on the Shapley value in game theory and decomposes the prediction results of the machine learning model into the sum of the contribution values of each feature.
[0054] S204: The cause of the target fault is determined by using a preset fault analysis model and based on the target feature data and the feature contribution of each feature data, and then sent to the client.
[0055] Specifically, in this embodiment, the server device inputs the feature contribution of each feature data and the target feature data into a preset fault analysis model, and uses the preset fault analysis model to predict the cause of the target disk failure and output the cause of the target failure.
[0056] Among them, the cause of target failure refers to the reason that causes the predicted probability of generating the target to exceed the first preset threshold.
[0057] Optionally, the preset fault analysis model can be a Bayesian network, or it can be set independently according to requirements. This embodiment does not impose any limitations.
[0058] Specifically, by automatically acquiring target feature data of the target disk, the target feature data of the target disk is input into a preset disk failure prediction model for prediction, thereby outputting the target prediction probability. This enables the prediction of the probability of the target disk failing within a future target time period. When the target prediction probability is determined to be greater than a first preset threshold, the feature contribution of each feature data is determined based on the preset disk failure prediction model and the target feature data. Furthermore, a preset failure analysis model is used, and the cause of the target failure is determined based on the target feature data and the feature contribution of each feature data. This result is then sent to the client, thus predicting the cause of disk failure. Compared to determining the health status of a disk by checking whether its own attribute parameters exceed a threshold, predicting the probability of the target disk failing within a future target time period based on the target disk's feature data, as well as predicting the cause of the target failure, allows users to understand the operating status of the target disk in advance and the cause of failure when the failure probability exceeds the threshold. This enhances the interpretability of disk failures and helps improve the user experience.
[0059] As an optional implementation, based on the above embodiment, obtaining target feature data of the target disk includes:
[0060] Obtain raw data within a first preset time period; the raw data includes parameter data of the target disk, target environment data, and target load data;
[0061] The original data within the first preset time period is scaled using a preset scaling algorithm to obtain scaled original data.
[0062] Feature extraction is performed on the scaled original data to obtain the target feature data of the target disk.
[0063] The target disk parameter data refers to items 5, 187, 188, 197, and 198 of the SMART information corresponding to the target disk, each representing an attribute parameter. The target environment data includes the target disk's operating temperature, vibration acceleration, and system cooling fan speed. The target load data includes the target disk's read / write operations per second, response time, read / write throughput, and error log count. The error log count refers to disk errors recorded by the operating system. Read / write throughput refers to the amount of data transferred per second.
[0064] It is understandable that the SMART information corresponding to a disk contains both raw values and normalized values. The normalized values are calculated from the raw values using an algorithm built into the disk manufacturer. In this embodiment, the normalized value refers to the normalized value corresponding to each attribute parameter in the SMART information.
[0065] The current standard values for each parameter of the target disk refer to the latest current standard values for each parameter of the target disk.
[0066] For example, assuming the current time is February 1st and the first preset time period is 29 days, the current standard value corresponding to each attribute parameter of the target disk refers to the standardized value corresponding to each attribute parameter on February 1st and the standardized value corresponding to each attribute parameter of the target disk in the 29 consecutive days before February 1st.
[0067] Understandably, if 30 days of target feature data are needed, then the first preset time period is 29 days. The first preset time period is determined by the number of days of target feature data required.
[0068] Specifically, in this embodiment, the server device uses a preset interface or preset command to obtain the current standard value corresponding to each parameter of the target disk, and obtains the historical standard value of each parameter of the target disk within a first preset time period from a preset database. Then, it uses the corresponding scaling formula in the preset scaling algorithm to scale the current standard value of each parameter and the historical standard value of the parameter within the first preset time period to obtain the target value of each parameter after scaling, and uses the target value of each parameter after scaling as the parameter data of the scaled target disk.
[0069] Optionally, the default interface or default command is a default setting.
[0070] The scaling algorithm in the preset scaling algorithm is as follows: subtract the minimum value among the standard values corresponding to the parameter to obtain the first difference, and subtract the minimum value among the standard values corresponding to the parameter from the maximum value to obtain the second difference. Then, the ratio of the first difference to the second difference is doubled, and the multiplied value is subtracted from 1 to obtain the target value corresponding to the parameter.
[0071] Understandably, scaling is only applied to the data in the original dataset that requires scaling. Data that does not require scaling is retained in its original state.
[0072] Furthermore, in this embodiment, the parameter data of the scaled target disk is obtained, and corresponding calculations are performed based on the parameter data of the scaled target disk, the target environment data, and the target load data to obtain the target feature data of the target disk.
[0073] It is understandable that, based on the target feature data described earlier, the corresponding data in the original data can be calculated to obtain the corresponding data in the target feature data.
[0074] For example, feature extraction refers to performing corresponding calculations based on the raw data to obtain corresponding feature data. The current vibration peak value refers to the peak value of the real-time vibration of the disk tray accelerometer within the most recent minute. The current disk read / write operations per second (DDoS) refers to the average DDoS operation per second over the most recent minute. The growth rate of the number of remapped sectors over the past 7 days is obtained by subtracting the number of remapped sectors from 7 days ago and dividing by 7. The disk temperature rise rate over the past hour is obtained by subtracting the temperature from 24 hours ago and dividing by 24. The change in the average vibration peak value of the disk tray accelerometer over the past 3 days is obtained by subtracting the average value of the past day from the average value of the past 3 days and dividing by 2. The trend of the number of read / write operations per second (DDoS) over the past 12 hours is obtained by subtracting the average number of read / write operations per second over the past 6 hours from the average number of read / write operations per second over the past 12 hours and dividing by 6. The growth rate of the error log count over the past 12 hours is obtained by subtracting the total number of error logs over the past 6 hours from the total number of error logs over the past 12 hours and dividing by 6.
[0075] Specifically, by scaling the original data on the target disk, the original data is unified to the same scale, thereby ensuring data consistency.
[0076] As an optional implementation, based on any of the above embodiments, a preset disk failure prediction model is used to predict the target disk based on the target disk's target feature data, and the target prediction probability is output, including:
[0077] The target feature data of the target disk is input into the preset disk failure prediction model, and the preset disk failure prediction model is used to predict the target disk and output the target prediction probability.
[0078] Optionally, the preset disk failure prediction model is a convolutional neural network, or it can be a combination of a convolutional neural network and a long short-term memory network, etc., which is not limited in this embodiment.
[0079] Specifically, in this embodiment, the server device inputs the target feature data of the target disk into a preset disk fault prediction model, so that the preset disk fault prediction model makes a prediction on the target disk, thereby obtaining the target prediction probability output by the preset disk fault prediction model.
[0080] Specifically, by inputting the target feature data of the target disk into a preset disk failure prediction model, the probability of the target disk failing within a future target time period can be estimated. Then, the target disk can be processed accordingly based on the predicted probability, thereby ensuring the stability of the system.
[0081] As an optional implementation, based on any of the above embodiments, if the target prediction probability is greater than a first preset threshold, then the feature contribution of each feature data is determined based on the preset disk failure prediction model and the target feature data, including:
[0082] The preset disk failure prediction model is input into the preset feature contribution interpreter for initialization to obtain the initialized preset feature contribution interpreter.
[0083] The target feature data is input into the preset feature contribution interpreter, and the preset feature contribution interpreter is used to calculate the contribution of the feature data in the target feature data, and output the feature contribution corresponding to each feature data.
[0084] Among them, the preset feature contribution interpreter can be the Shapley additivity interpreter.
[0085] Specifically, in this embodiment, a preset disk failure prediction model is input into a preset feature contribution interpreter for initialization, so that the preset feature contribution interpreter learns the internal parameters and decision logic of the preset disk failure prediction model. After the preset feature contribution interpreter is initialized, the target feature data is input into the preset feature contribution interpreter, so that the preset feature contribution interpreter calculates the contribution of each feature data in the target feature data and outputs the feature contribution corresponding to each feature data.
[0086] It is understandable that the preset disk failure prediction model is a pre-trained disk failure prediction model.
[0087] When using the Shapley Additive Interpreter to calculate feature contribution values, the preset disk failure prediction model only needs to be input once (during Shapley Additive Interpreter initialization). Subsequent calculations of the target disk's feature contribution values do not require re-inputting the model. This is because the parameters of the preset disk failure prediction model are fixed. After only one initialization, the Shapley Additive Interpreter "remembers" the decision logic of the preset disk failure prediction model, and subsequent calculations only require inputting new feature vectors to calculate contribution values.
[0088] In this embodiment, the internal parameters of the preset disk failure prediction model are weights, convolution kernels, tree structure, etc.
[0089] Specifically, by quantifying the contribution of each feature data in the target feature data, it is possible to distinguish between key and secondary factors that cause the target prediction probability to exceed the first preset threshold, which helps in the subsequent analysis and processing of the cause of the target failure.
[0090] As an optional implementation, based on any of the above embodiments, a preset fault analysis model is used, and the cause of the target fault is determined based on the target feature data and the feature contribution of each feature data, including:
[0091] Sort the feature contribution of each feature data in descending order to obtain the feature contribution of each feature data after sorting.
[0092] From the feature contribution values of each sorted feature data, obtain the feature contribution values of a preset number of feature data in order from front to back;
[0093] Input the feature contribution of a preset number of feature data and the target feature data into a preset fault analysis model, and use the preset fault analysis model to predict the cause of the target disk failure and output the cause of the target failure.
[0094] Optionally, the preset quantity can be 3 or other positive integers; this embodiment does not impose any limitations.
[0095] Specifically, in this embodiment, the server device sorts the feature contribution of each feature data in descending order to obtain the feature contribution of each feature data after sorting. Then, it sequentially obtains the feature contribution of a preset number of feature data from the feature contribution of each feature data in descending order and inputs the feature contribution of the preset number of feature data and the target feature data into a preset fault analysis model. The preset fault analysis model is used to predict the cause of the target disk failure and output the cause of the target failure.
[0096] The preset fault analysis model is a pre-trained Bayesian network.
[0097] For example, the training process of the preset fault analysis model is as follows: A sample set, a preset structural constraint strategy, and an initial fault analysis model are obtained. The initial fault analysis model is an untrained model. The sample set is divided into a training sample set and a validation sample set. The training sample set and the preset structural constraint strategy are input into the initial fault analysis model for structural learning, enabling the initial fault analysis model to generate a target network structure, thus obtaining the structure-learned initial fault analysis model. The target network structure and the training sample set are then input into the structure-learned initial fault analysis model for parameter learning, enabling the initial fault analysis model to generate a target probability table, thus obtaining the parameter-learned initial fault analysis model. The parameter-learned initial fault analysis model is validated using a validation set. The validation sample set is input into the parameter-learned initial fault analysis model for causal reasoning, calculating each possible fault cause and its corresponding probability. The cause with the highest probability is output, and the fault cause identification accuracy is calculated based on the output results. When the fault cause identification accuracy is greater than a preset accuracy threshold, training ends, and the parameter-learned initial fault analysis model is determined as the preset fault analysis model.
[0098] Optionally, the preset structural constraint strategy is a predefined rule or restriction condition used to guide the construction of the network topology. It can be set independently according to needs, and is not limited in this embodiment.
[0099] Optionally, the preset accuracy threshold can be set independently according to needs, and this embodiment does not impose any limitations.
[0100] The goal of structure learning is to find causal relationships between nodes.
[0101] The sample set includes fault samples and normal samples. The fault samples include disk fault data and fault causes. Specifically, by ranking by contribution and filtering by preset quantity, the features with the greatest impact on the fault can be focused on. This makes the preset fault analysis model's reasoning about the target fault cause more focused on the core contradictions. By determining the target fault cause through the preset fault analysis model, maintenance personnel can resolve the fault in a timely manner and improve maintenance efficiency.
[0102] As an optional implementation, based on any of the above embodiments, the method further includes:
[0103] Determine whether the target prediction probability is greater than the first fault threshold and less than the second fault threshold;
[0104] If the predicted probability of the target is greater than the first fault threshold and less than the second fault threshold, a first prompt message is generated; the first prompt message is used to prompt the data in the target disk to be backed up.
[0105] If the predicted probability of the target is greater than or equal to the second fault threshold, a second prompt message is generated; the second prompt message is used to indicate that the target disk needs to be replaced.
[0106] If the target prediction probability is less than the first fault threshold, then determine whether the target prediction probability is less than the third fault threshold.
[0107] If the predicted probability of the target disk is less than the third fault threshold, a third prompt message is generated to indicate that the target disk is currently in a healthy state.
[0108] Optionally, the first fault threshold may be 50%, the second fault threshold may be 80%, and the third fault threshold may be 10%, etc., but no limitation is made in this embodiment.
[0109] Among them, the second fault threshold is greater than the first fault threshold, and the first fault threshold is greater than the third fault threshold.
[0110] For example, in this embodiment, it is assumed that the first fault threshold is 50%, the second fault threshold is 80%, and the third fault threshold is 10%. After obtaining the target predicted probability, the server device determines whether the target predicted probability is greater than the first fault threshold and less than the second fault threshold. If the target predicted probability is greater than 50% and less than 80%, a first prompt message is generated, which may be: "The disk is at risk of failure. It is recommended to back up all data as soon as possible and pay attention to the disk's operating status." If the target predicted probability is greater than or equal to 80%, a second prompt message is generated, which may be: "The disk is at a serious risk of failure and may fail at any time. Please back up your data immediately and replace the disk." If the target predicted probability is less than 50%, it is determined whether the target predicted probability is also less than 10%. If the target predicted probability is less than 10%, a third prompt message is generated, which may be: "The disk is in good condition. It is recommended to back up your data regularly."
[0111] It is understandable that the first, second, and third prompts can be other textual descriptions to achieve the corresponding prompting purpose.
[0112] The first preset threshold and the first fault threshold can be the same value or different values; this embodiment does not impose any limitation.
[0113] Specifically, the server-side device uses a two-tiered threshold to divide risk ranges. When the predicted probability of the target is in the middle range, a data backup prompt is triggered, which buys time for data migration in case of potential failures. When the predicted probability of the target is greater than the second failure threshold, a disk replacement prompt is triggered directly, thereby achieving zero-latency response to high-risk failures. This can significantly reduce the risk of business interruption caused by sudden disk failures. Furthermore, a third failure threshold is introduced to reconfirm the low-risk status. When the probability is lower than the third failure threshold, the health status is clearly indicated, reducing users' unnecessary attention to normal disks.
[0114] As an optional implementation, based on any of the above embodiments, the method further includes:
[0115] The target feature data of the target disk is input into the preset disk failure prediction model, and the target disk is scored using the preset disk failure prediction model. The target disk health score is then output. The target disk health score is used to identify the health status of the target disk.
[0116] If the target disk health score is determined to be greater than the second preset threshold, the feature contribution of each feature data is determined based on the preset disk failure prediction model and the target feature data.
[0117] The cause of the target fault is determined by using a preset fault analysis model and based on the target feature data and the feature contribution of each feature data, and then sent to the client.
[0118] Specifically, in this embodiment, the server device inputs the target feature data of the target disk into a preset disk fault prediction model, enabling the preset disk fault prediction model to perform a health score on the target disk, thereby obtaining the target disk health score output by the preset disk fault prediction model. The target disk health score is then compared with a second preset threshold. If the target disk health score is determined to be greater than the second preset threshold, the preset disk fault prediction model and the target feature data are input into a preset feature contribution interpreter to determine the feature contribution of each feature data. A preset fault analysis model is then used to determine the cause of the target fault based on the target feature data and the feature contribution of each feature data, and then sent to the client.
[0119] Optionally, the second preset threshold can be set independently according to needs, and this embodiment does not impose any limitations.
[0120] Specifically, a pre-set disk failure prediction model is used to score the health of the target disk, and when the health score of the target disk exceeds the threshold, the prediction of the cause of the target failure is triggered, thereby realizing the contribution calculation only for high-risk disks (scores close to the threshold), thus reducing the consumption of computing resources.
[0121] As an optional implementation, based on any of the above embodiments, before using a preset disk failure prediction model and predicting the target disk based on the target disk's target feature data, the method further includes:
[0122] Obtain the target sample set and the initial disk failure prediction model;
[0123] The target sample set is divided into a target training set, a target validation set, and a target test set;
[0124] The backpropagation algorithm is used to train the initial disk failure prediction model in rounds based on the target training set to obtain the initial disk failure prediction model after each round of training.
[0125] The target disk failure prediction model is determined based on the target validation set and the initial disk failure prediction model after each round of training;
[0126] A preset disk failure prediction model is determined based on the target test set and the target disk failure prediction model.
[0127] The target sample set includes normal disk sample data and faulty disk sample data. Different labels are applied to the normal and faulty disk sample data, and a health score is also assigned to each disk sample data. The initial disk failure prediction model is an untrained disk failure prediction model.
[0128] Optionally, the ratio of the target training set, target validation set, and target test set can be 6:1:3, or other ratios, which are not limited in this embodiment.
[0129] The proportion of the target training set is greater than the proportion of the target validation set and the proportion of the target test set.
[0130] The target training set is used to train the initial disk failure prediction model, the target validation set is used to monitor the overfitting of the initial disk failure prediction model in each training round and to select the target disk failure prediction model, and the target test set is used to perform the final performance evaluation of the target disk failure prediction model.
[0131] Among them, the target disk failure prediction model is the selected model that can be used for final performance evaluation.
[0132] Specifically, in this embodiment, a target sample set and an initial disk failure prediction model are obtained from a preset database, and the target sample set is divided into a target training set, a target validation set, and a target test set according to a preset partitioning ratio. In each round of training, samples from the target training set are input into the initial disk failure prediction model to obtain the initial disk failure prediction output. A preset loss function is used to calculate the loss value between the prediction output and the actual fault label. Using the backpropagation algorithm, starting from the output layer, the gradient of the loss function with respect to the parameters (weights and biases) of each layer is calculated layer by layer, thereby updating the parameters of the initial disk failure prediction model according to the preset gradient descent algorithm to reduce the loss value. After each round of training, the model parameters obtained from the current training are saved to form the initial disk failure prediction model after each round of training. For each training round of the initial disk failure prediction model, samples from the target validation set are input into the initial disk failure prediction model to obtain validation prediction results. Then, based on the validation prediction results, the first recall and validation loss of the initial disk failure prediction model on the target validation set are calculated. If the validation loss is less than a preset loss threshold, the initial disk failure prediction model of the corresponding training round is determined as a candidate disk failure prediction model. There must be at least two candidate disk failure prediction models. It is then determined whether the first recall corresponding to each candidate disk failure prediction model is greater than a preset recall threshold. If the first recall corresponding to each candidate disk failure prediction model is greater than the preset recall threshold, the candidate disk failure prediction model with the smallest validation loss is determined as the target disk failure prediction model. Furthermore, the samples in the target test set are input into the target disk failure prediction model to obtain the test prediction results. Then, based on the test prediction results, the second recall and test loss value of the target disk failure prediction model on the target test set are calculated. When the test loss value is less than the preset test threshold and the second recall value is greater than the preset recall threshold, the target disk failure prediction model is determined as the preset disk failure prediction model. If either the test loss value is less than the preset test threshold or the second recall value is greater than the preset recall threshold is not met, the initial disk failure prediction model is retrained.
[0133] Optionally, the preset loss function can be the cross-entropy loss function, mean squared error, etc., and this embodiment is not limited to any specific loss function.
[0134] Optionally, the preset recall rate threshold can be 95%, or it can be set independently according to needs. This embodiment does not impose any limitations.
[0135] The first recall rate refers to the recall rate determined based on the target validation set. The second recall rate refers to the recall rate determined based on the target test set.
[0136] The recall rate refers to the proportion of disks that were successfully alerted three time periods before the failure occurred. The recall rate is the ratio of the actual number of disks that were alerted within the third time period to the total number of disks predicted.
[0137] Optionally, the preset gradient descent algorithm can be a stochastic gradient descent algorithm, etc., which is not limited in this embodiment.
[0138] Stochastic gradient descent is an iterative algorithm used to solve optimization problems.
[0139] Backpropagation is a supervised learning algorithm used to train neural networks. Its core purpose is to adjust the weight parameters in the neural network so that the network output is as close as possible to the actual target value, thereby minimizing the loss function.
[0140] Specifically, by dividing the target sample set into a target training set, a target validation set, and a target test set, and training the initial disk failure prediction model multiple times using the backpropagation algorithm, the target validation set is used to monitor and filter the initial disk failure prediction model after each training round during the training process. This allows for the timely detection of overfitting trends in the initial disk failure prediction model after each training round, and the selection of the model with better generalization performance as the target disk failure prediction model. Furthermore, the target test set is used to perform a final evaluation of the target disk failure prediction model, determining the preset disk failure prediction model. This improves the reliability of the preset disk failure prediction model and ensures that the preset disk failure prediction model has strong generalization ability.
[0141] As an optional implementation, based on any of the above embodiments, obtaining the target sample set includes:
[0142] Acquire historical data from multiple training disks within a second preset time period; the historical data includes parameter data, historical environment data, and historical load data for each training disk.
[0143] A preset scaling algorithm is used, and historical data within a second preset time period is scaled to obtain scaled historical data.
[0144] Feature extraction is performed on the scaled historical data to obtain training feature data for each training disk;
[0145] Determine whether each training disk experiences a failure within a third preset time period; the third preset time period refers to the period following and continuing from the second preset time period.
[0146] If any training disk fails within the third preset time period, a first label is generated and added to the training feature data of the corresponding training disk.
[0147] If no failure occurs in any training disk within the third preset time period, a second label is generated and added to the training feature data of the corresponding training disk.
[0148] Obtain the historical health score for each training disk;
[0149] A third label is generated based on the historical health score of each training disk;
[0150] Determine the disk type of each training disk;
[0151] A fourth label is generated based on the disk type of each training disk;
[0152] Obtain the disk identification information corresponding to each training disk;
[0153] The disk identification information, training feature data and corresponding labels of each training disk are determined as the corresponding training disk samples.
[0154] The training disk samples including the first label and the training disk samples including the second label are sampled according to a preset ratio, and the sampled training disk samples are determined as the target sample set.
[0155] Each training disk corresponds to one disk.
[0156] Optionally, the second preset time period can be 30 days or other numbers of days; this embodiment does not impose any limitation.
[0157] Optionally, the third preset time period can be 3 days or other numbers of days; this embodiment does not impose any limitation.
[0158] Specifically, in this embodiment, the server device acquires historical data from multiple training disks within a second preset time period. The parameter data of each training disk within the second preset time period is scaled using a preset scaling algorithm to obtain scaled historical data. Feature extraction is performed on the scaled historical data to obtain training feature data for each training disk. It is determined whether each training disk experienced a failure within a third preset time period. If a failure occurred, a first label is generated and added to the corresponding training feature data. If no failure occurred within the third preset time period, a second label is generated and added to the corresponding training feature data. Further, historical health scores for each training disk are obtained from a preset database and used as a third label. The disk type of each training disk is determined and used as a fourth label. Obtain disk identification information corresponding to each training disk from the preset database, determine the disk identification information corresponding to each training disk, the training feature data within the second preset time period, and the corresponding label as the corresponding training disk sample, and sample the training disk sample including the first label and the training disk sample including the second label according to the preset ratio, and determine the sampled training disk sample as the target sample set.
[0159] Optionally, the preset ratio can be 1:10, that is, the ratio of training disk samples including the first label to training disk samples including the second label is 1:10.
[0160] For example, the first label can be 1 and the second label can be 0. Assuming the second preset time period is 30 days and the third preset time period is 3 days, and assuming the 30th day is February 1, then within 3 days after February 1, if the disk fails, it will be marked as 1, and if the disk does not fail, it will be marked as 0.
[0161] The disk identification information is used to distinguish different training disks.
[0162] Specifically, a supervised learning framework with clear temporal continuity is constructed through a fault label generation mechanism within a third preset time period, ensuring that the model learns fault precursor features of disk state evolution over time, rather than static attribute associations. Furthermore, stratified sampling of positive and negative samples according to a preset ratio solves the inherent class imbalance problem in disk fault data.
[0163] Figure 3 A flowchart illustrating a disk failure prediction method provided in another embodiment of this application is shown below. Figure 3 As shown. The disk failure prediction method provided in this embodiment specifically includes the following steps:
[0164] S301: Obtain target feature data of the target disk.
[0165] S302: Use a preset disk failure prediction model and predict the target disk based on the target disk's characteristic data, and output the target prediction probability and the target disk health score.
[0166] S303: If the target prediction probability is greater than the first preset threshold or the target disk health score is greater than the second preset threshold, then the preset disk failure prediction model is input into the preset feature contribution interpreter for initialization to obtain the initialized preset feature contribution interpreter.
[0167] S304: Input the target feature data into the preset feature contribution interpreter, and use the preset feature contribution interpreter to calculate the contribution of the feature data in the target feature data, and output the feature contribution corresponding to each feature data.
[0168] S305: Sort the feature contribution of each feature data in descending order, and then obtain the feature contribution of a preset number of feature data in sequence from front to back after sorting.
[0169] S306: Input the feature contribution of a preset number of feature data and the target feature data into the preset fault analysis model, and use the preset fault analysis model to predict the cause of the target disk failure, and output the cause of the target failure.
[0170] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0171] Figure 4 This is a schematic diagram of the structure of a disk failure prediction device provided in an embodiment of this application. Figure 4 As shown, the execution entity of the above-mentioned disk failure prediction method is a disk failure prediction device, which can be implemented by a computer program; it can also be implemented by a medium storing the relevant computer program, such as a USB flash drive and / or optical disc, or it can be implemented by a physical device integrating or installing the relevant computer program, such as an electronic device. The electronic device can be a computer or a server device. The disk failure prediction device provided in this embodiment is located in an electronic device, therefore the disk failure prediction device 40 provided in this embodiment includes: an acquisition module 41, a prediction module 42, and a determination module 43.
[0172] Specifically, the acquisition module 41 is used to acquire target feature data of the target disk. The target feature data includes multiple feature data. The target feature data reflects the current state of the target disk. The target feature data includes the disk type of the target disk. The prediction module 42 is used to predict the target disk using a preset disk failure prediction model and based on the target feature data, and outputs the target prediction probability. The target prediction probability is the probability that the target disk will fail within a future target time period. The determination module 43 is used to determine the feature contribution of each feature data based on the preset disk failure prediction model and the target feature data if the determined target prediction probability is greater than a first preset threshold. The determination module 43 is also used to determine the cause of the target failure using a preset failure analysis model and based on the target feature data and the feature contribution of each feature data, and send it to the client.
[0173] Optionally, the acquisition module 41, when acquiring target feature data of the target disk, is used to acquire raw data within a first preset time period. The raw data includes parameter data of the target disk, target environment data, and target load data. The raw data within the first preset time period is scaled using a preset scaling algorithm to obtain scaled raw data. Feature extraction is performed on the scaled raw data to obtain target feature data of the target disk.
[0174] Optionally, the prediction module 42, when using a preset disk failure prediction model to predict the target disk based on the target feature data of the target disk and outputting the target prediction probability, is used to input the target feature data of the target disk into the preset disk failure prediction model, use the preset disk failure prediction model to predict the target disk, and output the target prediction probability.
[0175] Optionally, the determining module 43, when determining the feature contribution of each feature data based on the preset disk failure prediction model and target feature data if the predicted probability of the target is greater than a first preset threshold, is used to initialize the preset disk failure prediction model by inputting it into a preset feature contribution interpreter to obtain an initialized preset feature contribution interpreter. The target feature data is then input into the preset feature contribution interpreter, and the preset feature contribution interpreter is used to calculate the contribution of each feature data in the target feature data, and output the feature contribution corresponding to each feature data.
[0176] Optionally, the determining module 43, when determining the cause of the target fault using a preset fault analysis model and based on the target feature data and the feature contribution degree of each feature data, sorts the feature contribution degrees of each feature data in descending order to obtain the sorted feature contribution degrees of each feature data. From the sorted feature contribution degrees of each feature data, a preset number of feature data are sequentially obtained from front to back. The feature contribution degrees of the preset number of feature data and the target feature data are input into the preset fault analysis model, and the preset fault analysis model is used to predict the cause of the target disk fault and output the target fault cause.
[0177] Optionally, the disk failure prediction device also includes a generation module.
[0178] Accordingly, the determining module 43 is further configured to determine whether the target predicted probability is greater than a first fault threshold and less than a second fault threshold. The generating module is configured to generate a first prompt message if the target predicted probability is greater than the first fault threshold and less than the second fault threshold. The first prompt message prompts the user to back up the data on the target disk. If the target predicted probability is greater than or equal to the second fault threshold, a second prompt message is generated. The second prompt message prompts the user to replace the target disk. The determining module 43 is further configured to determine whether the target predicted probability is less than a third fault threshold if the target predicted probability is less than the first fault threshold. The generating module is further configured to generate a third prompt message if the target predicted probability is less than the third fault threshold. The third prompt message prompts the user that the target disk is currently in a healthy state.
[0179] Optionally, the disk failure prediction device also includes a scoring module.
[0180] Accordingly, the scoring module is used to input the target feature data of the target disk into a preset disk failure prediction model, and to use the preset disk failure prediction model to score the health of the target disk, and output the target disk health score. The target disk health score is used to identify the health status of the target disk. The determination module 43 is also used to determine the feature contribution of each feature data based on the preset disk failure prediction model and the target feature data if the determined target disk health score is greater than a second preset threshold. The preset failure analysis model is used to determine the cause of the target failure based on the target feature data and the feature contribution of each feature data, and then the result is sent to the client.
[0181] Optionally, the disk failure prediction device also includes a partitioning module and a training model.
[0182] Accordingly, the acquisition module 41 is used to acquire a target sample set and an initial disk failure prediction model before predicting the target disk using a preset disk failure prediction model and based on the target feature data of the target disk. The partitioning module is used to partition the target sample set into a target training set, a target validation set, and a target test set. The model training module is used to train the initial disk failure prediction model in rounds using the backpropagation algorithm and based on the target training set to obtain the initial disk failure prediction model after each round of training. The determination module 43 is used to determine the target disk failure prediction model based on the target validation set and the initial disk failure prediction models after each round of training. Finally, the preset disk failure prediction model is determined based on the target test set and the target disk failure prediction model.
[0183] Optionally, the acquisition module 41, when acquiring the target sample set, is specifically used for: acquiring historical data of multiple training disks within a second preset time period. The historical data includes parameter data, historical environment data, and historical load data for each training disk. A preset scaling algorithm is used, and the historical data within the second preset time period is scaled to obtain scaled historical data. Feature extraction is performed on the scaled historical data to obtain training feature data for each training disk. It is determined whether each training disk experienced a failure within a third preset time period. The third preset time period refers to the time period following and consecutive to the second preset time period. If each training disk experienced a failure within the third preset time period, a first label is generated and added to the corresponding training feature data. If each training disk did not experience a failure within the third preset time period, a second label is generated and added to the corresponding training feature data. The historical health score of each training disk is acquired. A third label is generated based on the historical health score of each training disk. The disk type of each training disk is determined. A fourth label is generated based on the disk type of each training disk. The disk identification information corresponding to each training disk is acquired. The disk identifier information corresponding to each training disk, the training feature data within the second preset time period, and the corresponding label are determined as the corresponding training disk samples. Training disk samples including the first label and training disk samples including the second label are sampled according to a preset ratio, and the sampled training disk samples are determined as the target sample set.
[0184] For a description of the features in the embodiment of the disk failure prediction device, please refer to the relevant description of the embodiment of the disk failure prediction method, which will not be repeated here.
[0185] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 5 As shown, the electronic device 50 provided in the embodiments of this application includes: a memory 52 and a processor 51.
[0186] The memory 52 stores a computer program, and the processor 51 is configured to run the computer program to perform the steps in any of the disk failure prediction method embodiments described above.
[0187] The specific implementation process of processor 51 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0188] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0189] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0190] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0191] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the disk failure prediction method embodiments described above when running.
[0192] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0193] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the disk failure prediction method embodiments described above.
[0194] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the disk failure prediction method embodiments described above.
[0195] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0196] The above provides a detailed description of a device information display method provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method of predicting disk failure, the method comprising: The method comprises: obtaining target feature data of a target disk; the target feature data comprises a plurality of feature data; the target feature data is used to reflect the current state of the target disk; the target feature data comprises a disk type of the target disk; the disk type comprises a mechanical disk and a solid-state disk; using a preset disk failure prediction model and based on the target feature data of the target disk to predict the target disk, and output a target prediction probability; the target prediction probability is the probability of failure of the target disk within a future target time period; inputting the target feature data of the target disk into the preset disk failure prediction model, and using the preset disk failure prediction model to score the health of the target disk, and output a target disk health degree score; the target disk health degree score is used to identify the health status of the target disk; if it is determined that the target prediction probability is greater than a first preset threshold, or if it is determined that the target disk health degree score is greater than a second preset threshold, then based on the preset disk failure prediction model and the target feature data, the feature contribution degree of each feature data is determined; using a preset failure analysis model and based on the target feature data, the feature contribution degree of each feature data to determine a target failure cause, and sending to a client; the using a preset failure analysis model and based on the target feature data, the feature contribution degree of each feature data to determine a target failure cause comprises: sorting the feature contribution degrees of each feature data from large to small to obtain the sorted feature contribution degrees of each feature data; from the sorted feature contribution degrees of each feature data, the feature contribution degrees of a preset number of feature data are obtained from front to back in turn; inputting the feature contribution degrees of the preset number of feature data and the target feature data into the preset failure analysis model, and using the preset failure analysis model to predict the failure cause of the target disk, and output the target failure cause; the method further comprises: determining whether the target prediction probability is greater than a first failure threshold and less than a second failure threshold; if the target prediction probability is greater than the first failure threshold and less than the second failure threshold, generating a first prompt information; the first prompt information is used to prompt to backup the data in the target disk; if the target prediction probability is greater than or equal to the second failure threshold, generating a second prompt information; the second prompt information is used to prompt that the target disk is to be replaced; if the target prediction probability is less than the first failure threshold, determining whether the target prediction probability is less than a third failure threshold; if the target prediction probability is less than the third failure threshold, generating a third prompt information, the third prompt information is used to prompt that the target disk is currently in a healthy state.
2. The magnetic disk failure prediction method according to claim 1, characterized by, the obtaining target feature data of a target disk comprises: obtaining original data within a first preset time period; the original data comprises parameter data, target environment data and target load data of the target disk; scaling the original data in the first preset time period by using a preset scaling algorithm to obtain scaled original data; performing feature extraction on the scaled original data to obtain the target feature data of the target disk.
3. The magnetic disk failure prediction method of claim 1, wherein, The method further comprises the following steps before predicting the target disk based on the target feature data of the target disk by using the preset disk failure prediction model: inputting the target feature data of the target disk into the preset disk failure prediction model, and predicting the target disk by using the preset disk failure prediction model to output a target prediction probability.
4. The magnetic disk failure prediction method of claim 1, wherein, If it is determined that the target prediction probability is greater than a first preset threshold, the method further comprises the following steps: initializing the preset feature contribution degree interpreter by inputting the preset disk failure prediction model into the preset feature contribution degree interpreter to obtain an initialized preset feature contribution degree interpreter; inputting the target feature data into the preset feature contribution degree interpreter, and calculating the contribution degree of the feature data in the target feature data by using the preset feature contribution degree interpreter to output the feature contribution degree corresponding to each feature data.
5. The magnetic disk failure prediction method of claim 1, wherein, The method further comprises the following steps before predicting the target disk based on the target feature data of the target disk by using the preset disk failure prediction model: obtaining a target sample set and an initial disk failure prediction model; dividing the target sample set into a target training set, a target validation set, and a target test set; performing round training on the initial disk failure prediction model based on the target training set by using a back propagation algorithm to obtain an initial disk failure prediction model after each round of training; determining a target disk failure prediction model based on the target validation set and the initial disk failure prediction model after each round of training; determining the preset disk failure prediction model based on the target test set and the target disk failure prediction model.
6. The magnetic disk failure prediction method according to claim 5, characterized by, The method further comprises the following steps before obtaining the target sample set: obtaining historical data of a plurality of training disks in a second preset time period; the historical data includes parameter data, historical environment data, and historical load data of each training disk; performing scaling processing on the historical data in the second preset time period by using a preset scaling algorithm to obtain scaled historical data; performing feature extraction on the scaled historical data to obtain training feature data of each training disk; determining whether each training disk fails in a third preset time period; the third preset time period refers to a time period after the second preset time period and continuous with the second preset time period; if each training disk fails in the third preset time period, generating a first label and adding it to the training feature data corresponding to the training disk; if each training disk does not fail in the third preset time period, generating a second label and adding it to the training feature data corresponding to the training disk; obtaining a historical health score of each training disk; generating a third label based on the historical health scores of the training disks; determining disk types of the training disks; generating a fourth label based on the disk types of the training disks; obtaining disk identification information corresponding to the training disks; determining the disk identification information corresponding to the training disks, the training feature data and the corresponding labels in the second preset time period as corresponding training disk samples; sampling the training disk samples including the first label and the training disk samples including the second label according to a preset ratio, and determining the sampled training disk samples as a target sample set.
7. An electronic device, comprising: comprising: a memory for storing a computer program; a processor for implementing the steps of the disk failure prediction method according to any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Hard disk fault prediction model interpretation method and device
CN111737067A
Fault analysis method and device, electronic equipment and computer readable storage medium
CN117102950A