Hard drive failure prediction methods, electronic devices and storage media

By performing multi-dimensional analysis of hard drive operating status and environmental data, and utilizing feature extraction and fault prediction models, the high cost and false alarm rate of hard drive monitoring methods have been solved. This enables accurate prediction of hard drive health status and consideration of environmental factors, reducing the complexity of operation and maintenance.

CN120849203BActive Publication Date: 2026-01-30INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511357702.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-01-30
Estimated Expiration
2045-09-22

AI Technical Summary

Technical Problem

Existing hard drive monitoring methods based on threshold alarms are costly and have a high false alarm rate, making it difficult to meet the needs of large-scale hard drive management. Analyzing SMART parameters alone cannot accurately capture the real trend of hard drive health status changes, and the impact of environmental factors is not fully considered. Furthermore, there are protocol differences in the definition and representation of SMART parameters, which increases the complexity of operation and maintenance.

Method used

The system collects operating status and environmental data of hard drives of different protocol types over multiple time periods. It then performs normalization and feature extraction using a feature extraction model, determines the predicted failure probability and level of the hard drive using a fault prediction model, adjusts the weight coefficients of the attention layer based on environmental data, and outputs the prediction results.

Benefits of technology

It reduced the false alarm rate, improved operation and maintenance efficiency, built a fault prediction model applicable to different protocol types, eliminated the management difficulties caused by protocol differences, and improved the accuracy and adaptability of prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849203B_ABST
    Figure CN120849203B_ABST
Patent Text Reader

Abstract

This application discloses a hard disk fault prediction method, electronic device, and storage medium, relating to the field of storage technology. The method includes collecting hard disk operating status data and environmental data; normalizing the operating status data using the normalization layer of a feature extraction model to generate feature vectors; extracting features from the feature vectors using the feature layer of the feature extraction model to output time-series features; and determining the weight coefficients of the attention layer in the fault prediction model to output the predicted fault probability, predicted fault level, and predicted fault type of the hard disk. This method addresses the technical problems of high cost and false alarm rate, difficulty in meeting the needs of large-scale hard disk management, insufficient consideration of environmental factors, inability to meet actual operational needs, and high operational complexity. It achieves the technical effects of eliminating management difficulties caused by protocol differences, fully considering the impact of environmental factors on hard disk health, improving prediction accuracy, reducing false alarm rate, and improving operational efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of storage, and in particular, to a hard disk failure prediction method, an electronic device and a storage medium. BACKGROUND

[0002] The reliability and stability of a hard disk are crucial for the safe storage of data and the continuous operation of a business. In the related art, the key performance parameters (such as temperature, read / write error rate, remaining life, etc.) of a hard disk can be dynamically compared with preset safety thresholds, and once an anomaly is detected, an alarm notification is triggered. Based on the SMART (Self-Monitoring Analysis and Reporting Technology) technology, the hard disk can be polled regularly, SMART data can be collected, and the collected data can be input into a threshold comparison process. When a parameter value exceeds a warning threshold, the system triggers a warning. When it exceeds a failure threshold, a failure alarm is triggered, and the operation and maintenance personnel are notified by email or SMS.

[0003] However, in the related art, the threshold alarm-based monitoring method requires operation and maintenance personnel to spend a lot of time and effort to deal with unnecessary alarms, increasing the operation and maintenance cost, and it is difficult to meet the needs of large-scale hard disk management. In addition, the false positive rate is high, which can easily lead to data loss. Separate analysis of a single SMART parameter ignores the timing correlation between different parameters, cannot accurately capture the real trend of the health status of the hard disk, and does not fully consider the influence of environmental factors, cannot adapt to complex and variable actual operating environments, is prone to false positives, and in addition, there are protocol differences in the definition and representation of SMART parameters, which cannot be directly compared and uniformly analyzed, increasing the complexity of operation and maintenance, and improvement is urgently needed. SUMMARY

[0004] The present application provides a hard disk failure prediction method, an electronic device and a storage medium to at least solve the problems in the related art that the threshold alarm-based monitoring method has high cost and false positive rate, is difficult to meet the needs of large-scale hard disk management, separate analysis of a single SMART parameter cannot accurately capture the real trend of the health status of the hard disk, does not fully consider the influence of environmental factors, cannot adapt to complex and variable actual operating environments, is prone to false positives, in addition, there are protocol differences in the definition and representation of SMART parameters, which cannot be directly compared and uniformly analyzed, increasing the complexity of operation and maintenance, etc.

[0005] The application provides a hard disk failure prediction method, comprising: collecting running state data and environment data of hard disks of at least two protocol types in at least two time periods; inputting the running state data and the environment data into a pre-constructed feature extraction model, normalizing the running state data by using a normalization layer of the feature extraction model, generating a feature vector corresponding to the running state data, and extracting features of the feature vector by using a feature layer of the feature extraction model to output time sequence features corresponding to the feature vector; inputting the environment data and the time sequence features into a pre-constructed failure prediction model, determining a weight coefficient of an attention layer in the failure prediction model, and outputting at least one of a predicted failure probability, a predicted failure level and a predicted failure type of the hard disk based on the weight coefficient.

[0006] The application also provides a hard disk failure prediction device, comprising: a collection module configured to collect running state data and environment data of hard disks of at least two protocol types in at least two time periods; a first output module configured to input the running state data and the environment data into a pre-constructed feature extraction model, normalize the running state data by using a normalization layer of the feature extraction model, generate a feature vector corresponding to the running state data, and extract features of the feature vector by using a feature layer of the feature extraction model to output time sequence features corresponding to the feature vector; and a second output module configured to input the environment data and the time sequence features into a pre-constructed failure prediction model, determine a weight coefficient of an attention layer in the failure prediction model, and output at least one of a predicted failure probability, a predicted failure level and a predicted failure type of the hard disk based on the weight coefficient.

[0007] The application also provides an electronic device, comprising: a memory configured to store a computer program; and a processor configured to implement steps of any of the above hard disk failure prediction methods when executing the computer program.

[0008] The application also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement steps of any of the above hard disk failure prediction methods.

[0009] The application also provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement steps of any of the above hard disk failure prediction methods.

[0010] By the present application, the running state data and environment data of at least two protocol types of hard disks in at least two time periods can be input into a pre-constructed feature extraction model at the same time, and then the corresponding feature vectors and time sequence features are obtained, and the weight coefficients of the attention layer in the fault prediction model are determined, so as to output the predicted fault probability, predicted fault level and predicted fault type of the hard disk by using the fault prediction model. Therefore, the technical problem of high cost and false alarm rate of the monitoring method based on threshold alarm, which is difficult to meet the demand of large-scale hard disk management, can be solved. The single SMART parameter cannot accurately capture the real change trend of the health status of the hard disk, and the influence of environmental factors is not fully considered, which cannot adapt to the complex and changeable actual operating environment, and is prone to false alarm. In addition, there are protocol differences in the definition and representation of SMART parameters, which cannot be directly compared and uniformly analyzed, increasing the complexity of operation and maintenance. The technical effects of reducing the false alarm rate of the fault, improving the operation and maintenance efficiency, constructing the fault prediction model suitable for different protocol types, eliminating the management difficulties caused by protocol differences, fully considering the influence of environmental factors on the health status of the hard disk, and improving the accuracy of prediction are achieved. BRIEF DESCRIPTION OF DRAWINGS

[0011] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating laborious work.

[0012] Figure 1 A flowchart of a hard disk fault prediction method according to an embodiment of the present application is provided.

[0013] Figure 2 A structural schematic diagram of a fault prediction model according to an embodiment of the present application is provided.

[0014] Figure 3 A flowchart of the working principle of a hard disk fault prediction method according to an embodiment of the present application is provided.

[0015] Figure 4 A block schematic diagram of a hard disk fault prediction device according to an embodiment of the present application is provided.

[0016] Reference signs:

[0017] Among them, 10-hard disk fault prediction device; 100-acquisition module, 200-first output module, 300-second output module. DETAILED DESCRIPTION

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] The embodiments of this application provide a hard disk failure prediction method, an electronic device, and a storage medium. The method is described in detail below in conjunction with the execution flow of the hard disk failure prediction method.

[0022] Specifically, Figure 1 This is a flowchart of a hard disk failure prediction method provided according to an embodiment of this application.

[0023] like Figure 1 As shown, the hard drive failure prediction method includes the following steps:

[0024] In step S101, operating status data and environmental data of at least two protocol types of hard drives are collected for at least two time periods.

[0025] It is understood that, in the embodiments of this application, the operating status data can be SMART parameter data, which may include, but is not limited to, the number of reassigned sectors, the number of currently unmapped sectors, the number of uncorrectable errors, power-on time, number of start / stop times, etc., and this application does not impose specific limitations; environmental data may include, but is not limited to, temperature data, vibration frequency data, humidity data, power supply fluctuation data, load rate data, etc., and this application does not impose specific limitations.

[0026] Further, the embodiment of the present application can collect the SMART parameters of a SAS (Serial Attached Small Computer System Interface) hard disk through a smartctl (SMART Control Tool) tool, collect the SMART parameters of an NVMe (Non-Volatile Memory Express) hard disk through an nvme-cli (NVMe Command Line Interface) tool, obtain temperature data through an IPMI (Intelligent Platform Management Interface) interface, collect vibration frequency data through a vibration sensor, and collect humidity data through a humidity sensor, thereby monitoring the running state data and the environmental data of the hard disk in real time. The running state data and the environmental data can also be collected by using other manners or tools, and the specific manners can be set by a person skilled in the art according to actual conditions, and the present application does not make specific limitations.

[0027] The IPMI is a standardized hardware management interface specification, which can conveniently obtain various state information of a server hardware, ensure the synchronous collection of the environmental data and the running state data, and provide a comprehensive data basis for subsequent comprehensive analysis.

[0028] In some embodiments, the embodiment of the present application can collect the running state data and the environmental data of the hard disk in at least two time periods in real time. The data of every 7 days can be divided into one time period, or the data of every 30 days can be divided into one time period, and the specific manners can be set by a person skilled in the art according to actual conditions, and the present application does not make specific limitations.

[0029] For example, the embodiment of the present application can collect the SMART parameters of a SAS hard disk through a smartctl tool every 15 minutes, collect the SMART parameters of an NVMe hard disk through an nvme-cli tool, obtain temperature data through an IPMI interface, collect vibration frequency data through a vibration sensor, and collect humidity data through a humidity sensor, thereby collecting the running state data and the environmental data of the hard disk in real time.

[0030] Optionally, in an embodiment of the present application, before the running state data and the environment data are input into the pre-constructed feature extraction model, it further includes: identifying the protocol type of the target hard disk; based on the protocol type and the target running state data of the target hard disk, extracting the wear degree feature, the bad block event count feature, the read error rate feature and the rotation retry count feature of the target hard disk; based on the wear degree feature, the bad block event count feature, the read error rate feature and the rotation retry count feature, determining the target feature vector corresponding to the target running state data, so as to construct the normalization layer of the feature extraction model by using the target feature vector.

[0031] It can be understood that the protocol type in the embodiment of the present application can include SAS protocol and NVMe protocol, and the present application does not make specific limitation, and the embodiment of the present application is to eliminate the parameter difference brought by different protocols to construct the normalization layer of the feature extraction model, and the specific content can be:

[0032] In some embodiments, the embodiment of the present application can calculate the wear degree feature, the bad block event count feature, the read error rate feature and the rotation retry count feature of the target hard disk according to the protocol type and the target running state data of the target hard disk, and then determine the target feature vector corresponding to the target running state data, so as to construct the normalization layer of the feature extraction model by using the target feature vector.

[0033] For example, in the case of SAS protocol, the embodiment of the present application calculates the wear degree feature, the bad block event count feature, the read error rate feature and the rotation retry count feature of the target SAS hard disk according to the target running state data of the target SAS hard disk, and then obtains the target SAS feature vector; in the case of NVMe protocol, the embodiment of the present application calculates the wear degree feature, the bad block event count feature, the read error rate feature and the rotation retry count feature of the target NVMe hard disk according to the target running state data of the target NVMe hard disk, and then obtains the target NVMe feature vector, and then constructs the normalization layer of the feature extraction model based on the target SAS feature vector and the target NVMe feature vector.

[0034] The embodiment of the present application maps the target running state data into unified features (such as wear degree feature, bad block count, read error rate feature, etc.) according to the protocol type of the hard disk, and then constructs the normalization layer in the feature extraction model, which provides a consistent data basis for subsequent multi-dimensional time series feature extraction and eliminates the management difficulties brought by protocol differences

[0035] Optionally, in an embodiment of the present application, before the running state data and the environment data are input into the pre-constructed feature extraction model, further comprising: determining the size and step length of the sliding window based on the target running state data; calculating the wear trend feature and the wear acceleration feature of the target hard disk based on the size, the step length and the wear degree feature; calculating the bad block fluctuation rate feature and the bad block distribution entropy feature of the target hard disk based on the size, the step length and the bad block event count feature; calculating the read error trend feature of the target hard disk based on the size, the step length and the read error rate feature; calculating the spin retry fluctuation rate feature of the target hard disk based on the size, the step length and the spin retry count feature; and determining the target time series feature corresponding to the target running state data based on the wear trend feature, the wear acceleration feature, the bad block fluctuation rate feature, the bad block distribution entropy feature, the read error trend feature and the spin retry fluctuation rate feature, so as to construct the feature layer of the feature extraction model by using the target time series feature.

[0036] It can be understood that, in the embodiments of the present application, the trend features such as the wear trend feature and the read error trend feature are long-term early warning signals of failure, and if the indicators continue to deteriorate, they can be represented by a positive slope, such as wear trend feature = 0.8, indicating that the wear degree feature increases by 0.8 per day; the fluctuation features such as the bad block fluctuation rate feature and the spin retry fluctuation rate feature reflect the stability of the hard disk state, and if the indicators fluctuate more violently, the standard deviation is larger, such as bad block fluctuation rate feature = 10, indicating that the number of bad blocks fluctuates ± 10 within 7 days; the wear acceleration feature is a core indicator of rapid aging failure, and a positive value indicates that the wear is accelerating, such as wear acceleration feature = 0.2, indicating that the wear trend feature increases by 0.2 per day; the bad block distribution entropy feature reflects the dispersion degree of the bad blocks, and the higher the entropy value, the more dispersed the bad blocks, and the higher the risk of failure, so as to prompt the large-area degradation of the medium (such as disc scratching), such as bad block distribution entropy > 3, indicating that the bad blocks are distributed in a scattered manner (not locally concentrated).

[0037] In some embodiments, the embodiments of the present application generate target time series features by using the method of sliding window aggregation, so as to capture the dynamic changes of the running state of the target hard disk, and then construct the feature layer of the feature extraction model.

[0038] Further, the embodiments of the present application determine the size and step length of the sliding window based on the target running state data, and calculate the wear trend feature, the wear acceleration feature, the bad block fluctuation rate feature, the bad block distribution entropy feature, the read error trend feature and the spin retry fluctuation rate feature of the target hard disk in combination with the wear degree feature, the bad block event count feature, the read error rate feature and the spin retry count feature, and then determine the target time series feature corresponding to the target running state data, so as to construct the feature layer of the feature extraction model.

[0039] For example, in this embodiment of the application, the sliding window size can be set to 7 days and the step size to 1 day. Before calculating the wear trend feature, wear acceleration feature and read error trend feature, the trend slope is calculated first. Then, the wear trend feature, wear acceleration feature and read error trend feature are calculated using the trend slope. The wear acceleration feature is the rate of change (second derivative) of the wear trend feature. Therefore, when calculating the wear acceleration feature, the daily wear slope within the window (3-day sub-window) can be calculated first, and then the slope of the slope is calculated, which is the wear acceleration feature.

[0040] Furthermore, in this embodiment of the application, the bad block distribution entropy characteristics of the target hard disk can be obtained by calculating the information entropy and dividing it into 5 levels of statistical bad block distribution entropy.

[0041] Furthermore, in this embodiment of the application, bad block volatility characteristics and rotational retry volatility characteristics of the target hard disk are calculated based on bad block event count characteristics and rotational retry count characteristics, respectively.

[0042] Additionally, it should be noted that the embodiments of this application can incorporate interactive items of environmental data and operating status data to reflect the direct impact of the environment on the hard drive's status. For example, when the temperature change exceeds a certain threshold, a high-temperature accelerated wear characteristic can be generated, and its calculation formula can be: High-temperature accelerated wear characteristic = Wear trend characteristic * When the power supply fluctuation exceeds a certain threshold, a feature indicating aggravated read errors due to power supply fluctuations can be generated. The calculation formula is: Aggravated read error feature due to power supply fluctuations = Read error trend feature * ,in, Indicates the amount of temperature change. This indicates the amount of fluctuation in power supply.

[0043] This application embodiment can employ a sliding window to construct the feature layer of the feature extraction model by combining environmental data and operational status data. The dynamic sliding window mechanism adapts to dynamic changes in hard disk status, avoids feature distortion caused by a fixed window, and integrates multi-dimensional temporal features to reduce the risk of data loss and improve fault prediction accuracy.

[0044] Optionally, in one embodiment of this application, before inputting the running status data and environmental data into the pre-built feature extraction model, the method further includes: determining whether the change value of different environmental parameters in the environmental data is greater than a preset threshold; if the change value is greater than the preset threshold, then allowing the environmental data of the corresponding environmental parameter to be input into the feature extraction model; if the change value is less than or equal to the preset threshold, then prohibiting the input of the environmental data of the corresponding environmental parameter into the feature extraction model.

[0045] It can be understood that in the embodiments of the present application, the environmental parameters can include but are not limited to temperature parameters, vibration frequency parameters, humidity parameters, power supply fluctuation parameters, load rate parameters, etc., and the present application does not make specific limitations.

[0046] Further, in the embodiments of the present application, if the environment of the current hard disk changes greatly, the weight of the historical features is appropriately reduced at this time, that is, the model pays more attention to the feature information at the current time, because in the case of large environmental change, the environmental data at the current time can better reflect the actual running condition of the hard disk. Therefore, the embodiments of the present application can first determine whether the change value of different environmental parameters in the environmental data is greater than a certain threshold, and when it is greater, the environmental data of the corresponding environmental parameter is allowed to be input to the feature extraction model, otherwise, the environmental data of the corresponding environmental parameter is prohibited to be input to the feature extraction model. The certain threshold can be set by a person skilled in the art according to the actual situation, and the present application does not make specific limitations.

[0047] For example, the embodiments of the present application can determine whether the temperature change, the vibration change, the humidity change, the power supply fluctuation change, and the load change are greater than a certain threshold, if the temperature change is greater than a certain threshold, the temperature data in the environmental data is allowed to be input to the feature extraction model; if the temperature change and the vibration change are both greater than a certain threshold, the temperature data and the vibration frequency data in the environmental data are allowed to be input to the feature extraction model; if the temperature change, the vibration change, the humidity change, the power supply fluctuation change, and the load change are all greater than a certain threshold, the temperature data, the vibration frequency data, the humidity data, the power supply fluctuation data, and the load rate data are allowed to be input to the feature extraction model; if the temperature change, the vibration change, the humidity change, the power supply fluctuation change, and the load change are all less than or equal to a certain threshold, the temperature data, the vibration frequency data, the humidity data, the power supply fluctuation data, and the load rate data are prohibited to be input to the feature extraction model.

[0048] The embodiments of the present application comprehensively consider the influence of temperature parameters, vibration frequency parameters, humidity parameters, power supply fluctuation parameters, load rate parameters and other environmental parameters on the hard disk, so that the model can better adapt to the complex and changeable running environment, and improve the accuracy of the environmental sensitive mechanical hard disk fault prediction.

[0049] In step S102, the running state data and the environmental data are input into the pre-constructed feature extraction model, the running state data is normalized by using the normalization layer of the feature extraction model, the feature vector corresponding to the running state data is generated, and the feature layer of the feature extraction model is used to extract the features of the feature vector to output the time sequence features corresponding to the feature vector.

[0050] In some embodiments, the embodiments of the present application can normalize the running state data by using the normalization layer of the feature extraction model, to generate a feature vector corresponding to the running state data.

[0051] For example, the embodiments of the present application can input the collected running state data into the normalization layer of the feature extraction model, convert the running state data into a unified feature vector according to the protocol type of the hard disk, and eliminate the influence of protocol differences on subsequent analysis.

[0052] In some embodiments, the embodiments of the present application can perform feature extraction on the feature vector by using the feature layer of the feature extraction model, to output time sequence features corresponding to the feature vector.

[0053] For example, the embodiments of the present application can perform sliding window aggregation operation on the unified feature vector with a window of 7 days, to calculate time sequence features such as wear trend features, wear acceleration features, and bad block volatility features, and further dynamically reflect the change trend of the health status of the hard disk over time, which has more predictive value than static features.

[0054] In step S103, the environment data and the time sequence features are input into a pre-constructed failure prediction model, a weight coefficient of an attention layer in the failure prediction model is determined, and at least one of a predicted failure probability, a predicted failure level, and a predicted failure type of the hard disk is output based on the weight coefficient.

[0055] It can be understood that in the embodiments of the present application, the failure prediction model can select an LSTM-Attention (Long Short-Term Memory with Attention Mechanism) prediction model, or a Transformer prediction model, which can be set by a person skilled in the art according to actual conditions, and the present application does not make specific limitations.

[0056] Taking the LSTM-Attention prediction model as an example, the structure of the failure prediction model of the embodiments of the present application is shown in Figure 2 which can but not limited to include an input layer, an LSTM layer, an attention layer, and an output layer.

[0057] The input layer is a multi-dimensional time sequence feature sequence processed by the feature extraction model.

[0058] Exemplarily, in the embodiment of the present application, the input of the model can be a multi-dimensional time series feature sequence of nearly 30 days, i.e., dimension (30, 6) (30 time steps, each time step containing 6 time series features: wear trend feature, wear acceleration feature, bad block volatility rate feature, bad block distribution entropy feature, read error trend feature and spin retry volatility rate feature. The specific processing steps are as follows:

[0059] Time series feature standardization: standard deviation standardization is performed on the 30-day time series feature, and the calculation formula can be but is not limited to:

[0060]

[0061] wherein, is the mean of the training set, is the standard deviation of the training set.

[0062] Sequence construction: the 30-day standardized features are spliced in time sequence to form an input tensor with a shape of (batch_size, 30, 6) (wherein, batch_size is the number of hard disk samples for each training).

[0063] Label construction: whether a failure occurs within the next 7 days is taken as a label (1 = failure, 0 = normal), and the training data is labeled in combination with the actual hard disk failure log.

[0064] The LSTM layer has 64 hidden units, and through the synergistic effect of the forget gate, the input gate and the output gate, it effectively captures the long-term dependence relationship in the time series and learns the change rule of the hard disk running state in the time dimension.

[0065] In the embodiment of the present application, the processing mechanism of the LSTM layer on the multi-dimensional time series feature sequence can be as follows:

[0066] The forget gate is used to determine which historical features to retain. For example, when the wear trend feature is continuously positive and the wear acceleration feature increases, the forget gate weight tends to 1, retaining the historical wear degradation information; when the spin retry volatility rate feature suddenly rises but then recovers, the forget gate weight tends to 0, discarding short-term fluctuation noise.

[0067] The input gate is used to update the feature information of the current time step. For example, when the bad block volatility rate feature suddenly rises from 2 to 4 (bad block dispersion), the input gate weight increases, and the key change is included in the hidden state.

[0068] Further, when training based on the multi-dimensional time series feature sequence in the embodiment of the present application, the cross-entropy loss is taken as the optimization objective, and the process can be as follows:

[0069] ​Weight initialization: the input weights corresponding to the wear acceleration feature and the bad block fluctuation rate feature are initialized to 0.2 (and 0.1 for other features), so that the model pays more attention to the high-fault-correlation features.

[0070] Regularization: a random inactivation layer (inactivation probability = 0.2) is added to the LSTM layer to prevent overfitting (e.g., to avoid the model over-relying on the accidental fluctuations of a certain feature)

[0071] The attention layer is used to dynamically adjust the weights of the input features, so that the model can pay more attention to the time periods and features that are more critical to fault prediction.

[0072] After the output layer is processed by the attention layer, the data is input to a fully connected layer with 32 neurons for further feature fusion and mapping, and finally the predicted fault probability is output, and then the predicted fault level and predicted fault type are determined.

[0073] For example, the embodiments of the present application can output the predicted fault probability P and health score for the next 7 days, and then determine the predicted fault level, and output the predicted fault type such as medium aging fault, medium scratch fault, mechanical fault, etc. according to the attention layer feature weight distribution, to assist the operation and maintenance personnel in formulating replacement strategies (such as migrating data in priority for medium fault). The calculation formula of the predicted fault probability can be, but is not limited to:

[0074]

[0075] wherein, represents an activation function, represents a weighted hidden state, represents the learnable parameters of the fully connected layer, and P is between 0 and 1.

[0076] Optionally, in an embodiment of the present application, the weight coefficient of the attention layer in the fault prediction model is determined, including: determining the initial weight coefficient of the attention layer in the fault prediction model based on the time series feature; calculating the interference factor of the attention layer based on the environmental parameter; calculating the weight decay coefficient corresponding to the initial weight coefficient by using the interference factor, so as to adjust the initial weight coefficient by using the weight decay coefficient to obtain the adjusted weight coefficient.

[0077] It can be understood that the embodiments of the present application can first determine the initial weight coefficient of the attention layer in the fault prediction model based on the time series feature.

[0078] ​​Further, the embodiments of the present application add five types of environmental parameters, including temperature, vibration frequency, humidity, load rate, and power supply fluctuation, according to the different characteristics of SAS hard disks and NVMe hard disks. When the change value of the environmental parameters is greater than a certain threshold, the interference factor of the attention layer is calculated based on the environmental parameters, so as to reduce the false positives caused by environmental interference. The expression of the interference factor can be, but is not limited to,

[0079] ,

[0080] wherein, represents the environmental interference factor, which is used to quantify the comprehensive influence degree of vibration and temperature change on the hard disk, represents the vibration change amount, represents the temperature change amount, represents the power supply fluctuation change amount, represents the load change amount, represents the humidity change amount, , , , , respectively represent the weight coefficients of the vibration change amount, the temperature change amount, the power supply fluctuation change amount, the humidity change amount, and the load rate change amount. The weight of each factor is calibrated through experiments, so as to balance the relative importance of vibration and temperature, humidity, load rate, and power supply fluctuation in environmental interference.

[0081] It should be noted that in the embodiments of the present application, has the greatest impact on the mechanical structure of a mechanical hard disk (Hard Disk Drive, abbreviated as HDD), followed by; has a significant impact on the flash life of a solid state disk (Solid State Drive, abbreviated as SSD), so the weight is higher than that of humidity .

[0082] Further, the embodiments of the present application can use the interference factor to correct the initial weight coefficient, and then obtain the adjusted weight coefficient. The content can be:

[0083] (1) Calculate the weight attenuation coefficient: the greater the IF, the smaller the weight attenuation coefficient (maximum attenuation 50%), and the calculation formula can be, but is not limited to:

[0084] ,

[0085] wherein, represents the weight attenuation coefficient, and when IF> 10, the attenuation is 50%.

[0086] (2) Correct the initial weight coefficient, and the calculation formula can be, but is not limited to:

[0087] ,

[0088] wherein, represents an initial weight coefficient, represents an adjusted weight coefficient.

[0089] For example, when the air conditioner in the machine room fails to cause , ,the initial weight coefficient of the attention layer of the history 29 days is attenuated by 40%, and the weight of the current 30th day is raised to 0.4-0.5, avoiding the sudden deterioration caused by the model ignoring the current environment due to the history normal data.

[0090] The embodiments of the present application can adjust the initial weight coefficient of the attention layer according to the interference factor. When the interference factor is large, it means that the environment change of the current hard disk is relatively violent. At this time, the weight of the historical feature is appropriately reduced, which can make the model pay more attention to the feature information at the current moment, better capture the coupling risks of high temperature and high wear, power supply fluctuation and high rotation retry, and adapt to the complex and changeable running environment, and improve the accuracy of the fault prediction of the environment sensitive mechanical hard disk.

[0091] Optionally, in an embodiment of the present application, the weight coefficient of the attention layer in the fault prediction model is determined, comprising: obtaining the feature value of the time sequence feature in at least two time periods; based on the feature value, adjusting the weight coefficient in the corresponding period to obtain the adjusted weight coefficient.

[0092] In some embodiments, the embodiments of the present application can dynamically adjust the corresponding weight coefficient according to the feature value of different time periods, and then obtain the adjusted weight coefficient. For example, if the wear acceleration feature of the 25th-30th day is >0.3 and the bad block distribution entropy feature is >3, the weight coefficient of this period is raised to 0.6-0.8 (other periods <0.2), and the recent rapid degradation feature is focused on; if the read error trend feature of the 10th-15th day rises sharply but recovers subsequently, the weight coefficient of this period is reduced to below 0.1, and the short-term noise influence is weakened.

[0093] For example, the weight coefficient of the attention layer of the embodiments of the present application can dynamically adjust the weight according to the contribution of the time sequence feature to the fault, and the calculation formula can be but not limited to

[0094] ,

[0095] ,

[0096] wherein, is the LSTM hidden state of the th day, and For learnable parameters, This represents the weighted hidden state of the LSTM.

[0097] This application embodiment dynamically adjusts the weight coefficients by taking feature values ​​at different times, enabling the fault prediction model to adapt to the feature importance at different times. This significantly improves the model's ability to capture changes in hard disk health status and its predictive robustness, reduces operation and maintenance costs, and provides key technical support for the intelligent operation and maintenance of large-scale storage systems.

[0098] Optionally, in one embodiment of this application, the method further includes: determining that the hard drive has not failed when the predicted fault level is a first fault level; determining that the hard drive has a fault risk when the predicted fault level is a second fault level and alerting the user to pay attention; and determining that the hard drive is unusable when the predicted fault level is a third fault level and replacing it.

[0099] It is understood that the embodiments of this application can determine the predicted fault level based on the predicted fault probability, and divide the predicted fault level into three levels: the first fault level, i.e., the low-risk level; the second fault level, i.e., the medium-risk level; and the third fault level, i.e., the high-risk level.

[0100] In some embodiments, this application can determine that the hard drive has not failed when the predicted fault level is a first fault level; determine that the hard drive has a fault risk when the predicted fault level is a second fault level and alert the user to pay attention; and determine that the hard drive is unusable when the predicted fault level is a third fault level and replace it.

[0101] For example, in the embodiments of this application when When the health score is between 80 and 100, the predicted failure level is determined to be the first failure level, and the hard drive is in a stable state; when When the health score is between 40 and 79, the predicted fault level is determined to be the second fault level, and monitoring is strengthened; when If the health score is between 0 and 39, the predicted failure level is determined to be the third failure level, with a failure probability of >90% within 7 days, and the hard drive should be replaced immediately.

[0102] The embodiments of this application can adopt different processing strategies according to different fault levels. For example, an alarm can be issued at the second fault level, and the hard drive can be replaced at the third fault level. This avoids misjudgment by applying a one-size-fits-all approach, realizes fine-grained fault classification, balances security and cost, ensures business continuity, and adapts to different scenario requirements.

[0103] Optionally, in an embodiment of the present application, based on the weight coefficient, the predicted failure type of the hard disk is output, comprising: based on the weight coefficient, calculating a first sum value between the wear trend feature and the wear acceleration feature, a second sum value between the bad block fluctuation rate feature and the bad block distribution entropy feature, and a third sum value between the read error trend feature and the rotation retry fluctuation rate feature; comparing the sizes of the first sum value, the second sum value and the third sum value; when the first sum value is greater than the second sum value and the first sum value is greater than the third sum value, determining that the predicted failure type is the medium aging failure type; when the second sum value is greater than the first sum value and the second sum value is greater than the third sum value, determining that the predicted failure type is the medium scratch failure type; and when the third sum value is greater than the first sum value and the third sum value is greater than the second sum value, determining that the predicted failure type is the mechanical failure type.

[0104] In some embodiments, the embodiments of the present application can determine the predicted failure type according to the sum values between the wear trend feature and the wear acceleration feature, the bad block fluctuation rate feature and the bad block distribution entropy feature, and the read error trend feature and the rotation retry fluctuation rate feature, for example, when the first sum value between the wear trend feature and the wear acceleration feature is large, determining that the predicted failure type is the medium aging failure type; when the second sum value between the bad block fluctuation rate feature and the bad block distribution entropy feature is large, determining that the predicted failure type is the medium scratch failure type; and when the third sum value between the read error trend feature and the rotation retry fluctuation rate feature is large, determining that the predicted failure type is the mechanical failure type, as shown in Table 1, wherein Table 1 is a rule table for predicting the failure type in combination with the multi-time sequence feature weight distribution according to an embodiment of the present application.

[0105] Table 1

[0106] Feature weight ratio highest item Predicted failure type Applicable scenario First and value Medium aging failure Solid state disk flash memory cell wear out Second and value Medium scratch failure Mechanical hard disk platter large area scratch Third and value Mechanical failure (head / motor) Mechanical hard disk head positioning error, motor jam

[0107] The embodiments of the present application can determine the predicted failure type by comparing the sum values between the wear trend feature and the wear acceleration feature, the bad block fluctuation rate feature and the bad block distribution entropy feature, and the read error trend feature and the rotation retry fluctuation rate feature, breaking through the limitation of a single feature, improving the classification accuracy, strengthening the contribution of key features, improving the classification robustness, and making the application scenarios more diverse and extensive.

[0108] The working principle of the hard disk failure prediction method proposed in the embodiments of the present application will be introduced below in combination with a specific embodiment.

[0109] wherein, Figure 3 is a flow chart of the working principle of the hard disk failure prediction method according to an embodiment of the present application.

[0110] Step S301: data acquisition layer.

[0111] In the data acquisition layer, the embodiment of the present application can collect the running state data and the environment data of the hard disk in at least two time periods in real time.

[0112] Step S302: feature extraction layer.

[0113] In the feature extraction layer, the embodiment of the present application can normalize the running state data by using the normalization layer of the feature extraction model, generate the feature vector corresponding to the running state data, and extract the feature of the feature vector by using the feature layer of the feature extraction model, to output the time sequence feature corresponding to the feature vector.

[0114] Step S303: model prediction layer.

[0115] In the model prediction layer, the embodiment of the present application can output the predicted failure probability, the predicted failure level and the predicted failure type of the hard disk by using the pre-constructed failure prediction model.

[0116] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment.

[0117] The hard disk failure prediction method provided by the embodiment of the present application can simultaneously input the running state data and the environment data of the hard disk of at least two protocol types in at least two time periods into the pre-constructed feature extraction model, and then obtain the corresponding feature vector and time sequence feature, and determine the weight coefficient of the attention layer in the failure prediction model, so as to output the predicted failure probability, the predicted failure level and the predicted failure type of the hard disk by using the failure prediction model, thereby solving the technical problems that the monitoring method based on threshold alarm has high cost and false alarm rate, and is difficult to meet the demand of large-scale hard disk management; the single SMART parameter analysis cannot accurately capture the real change trend of the hard disk health status, and the influence of the environmental factors is not fully considered, which cannot adapt to the complex and changeable actual operating environment, and is easy to produce false alarm, in addition, there are protocol differences in the definition and representation of the SMART parameter, which cannot be directly compared and uniformly analyzed, increasing the complexity of operation and maintenance, achieving the technical effects of reducing the failure false alarm rate, improving the operation and maintenance efficiency, constructing the failure prediction model suitable for different protocol types, eliminating the management difficulties caused by the protocol differences, fully considering the influence of the environmental factors on the hard disk health status, and improving the prediction accuracy.

[0118] The embodiment of the present application also provides a hard disk failure prediction device.

[0119] Figure 4 A block schematic diagram of the hard disk failure prediction device provided by the embodiment of the present application is shown.

[0120] As shown in Figure 4 The hard disk failure prediction apparatus 10 comprises a collection module 100, a first output module 200 and a second output module 300.

[0121] The collection module 100 is configured to collect running state data and environment data of hard disks of at least two protocol types in at least two time periods.

[0122] The first output module 200 is configured to input the running state data and the environment data into a pre-constructed feature extraction model, normalize the running state data by using a normalization layer of the feature extraction model, generate a feature vector corresponding to the running state data, and extract features of the feature vector by using a feature layer of the feature extraction model to output time sequence features corresponding to the feature vector.

[0123] The second output module 300 is configured to input the environment data and the time sequence features into a pre-constructed failure prediction model, determine a weight coefficient of an attention layer in the failure prediction model, and output at least one of a predicted failure probability, a predicted failure level and a predicted failure type of the hard disk based on the weight coefficient.

[0124] Optionally, in an embodiment of the present application, the apparatus further comprises a judgment module, a first input module and a second input module.

[0125] The judgment module is configured to judge whether a change value of a different environment parameter in the environment data is greater than a preset threshold before the running state data and the environment data are input into the pre-constructed feature extraction model.

[0126] The first input module is configured to allow the environment data of the corresponding environment parameter to be input into the feature extraction model when the change value is greater than the preset threshold.

[0127] The second input module is configured to prohibit the environment data of the corresponding environment parameter from being input into the feature extraction model when the change value is less than or equal to the preset threshold.

[0128] Optionally, in an embodiment of the present application, the second output module 300 comprises a first determination unit, a first calculation unit and a first adjustment unit.

[0129] The first determination unit is configured to determine an initial weight coefficient of the attention layer in the failure prediction model based on the time sequence features.

[0130] The first calculation unit is configured to calculate an interference factor of the attention layer based on the environment parameter.

[0131] The first adjusting unit is configured to calculate a weight attenuation coefficient corresponding to the initial weight coefficient by using the interference factor, and adjust the initial weight coefficient by using the weight attenuation coefficient to obtain an adjusted weight coefficient.

[0132] Optionally, in an embodiment of the present application, the second output module 300 comprises an obtaining unit and a second adjusting unit.

[0133] The obtaining unit is configured to obtain feature values of the timing feature in at least two time periods.

[0134] The second adjusting unit is configured to adjust the weight coefficient in the corresponding time period based on the feature values to obtain an adjusted weight coefficient.

[0135] Optionally, in an embodiment of the present application, the method further comprises a first determining module, a second determining module and a third determining module.

[0136] The first determining module is configured to determine that the hard disk has not failed in a case where the predicted failure level is a first failure level.

[0137] The second determining module is configured to determine that the hard disk has a failure risk and alarm the user to remind the user to pay attention in a case where the predicted failure level is a second failure level.

[0138] The third determining module is configured to determine that the hard disk is unusable and replace it in a case where the predicted failure level is a third failure level.

[0139] Optionally, in an embodiment of the present application, the second output module 300 comprises a second calculating unit, a comparing unit, a second determining unit, a third determining unit and a fourth determining unit.

[0140] The second calculating unit is configured to calculate a first sum value between the wear trend feature and the wear acceleration feature, a second sum value between the bad block fluctuation rate feature and the bad block distribution entropy feature, and a third sum value between the read error trend feature and the rotation retry fluctuation rate feature based on the weight coefficient.

[0141] The comparing unit is configured to compare the sizes of the first sum value, the second sum value and the third sum value.

[0142] The second determining unit is configured to determine that the predicted failure type is a medium aging failure type in a case where the first sum value is greater than the second sum value and the first sum value is greater than the third sum value.

[0143] The third determining unit is configured to determine that the predicted failure type is a medium scratch failure type in a case where the second sum value is greater than the first sum value and the second sum value is greater than the third sum value.

[0144] The fourth determining unit is configured to determine that the predicted failure type is the mechanical failure type when the third sum is greater than the first sum and the third sum is greater than the second sum.

[0145] Optionally, in an embodiment of the present application, the method further comprises: an identifying module, a first calculating module and a first constructing module.

[0146] The identifying module is configured to identify the protocol type of the target hard disk before inputting the running state data and the environment data into the pre-constructed feature extraction model.

[0147] The first calculating module is configured to extract the wear degree feature, the bad block event count feature, the read error rate feature and the spin retry count feature of the target hard disk based on the protocol type and the target running state data of the target hard disk.

[0148] The first constructing module is configured to determine a target feature vector corresponding to the target running state data based on the wear degree feature, the bad block event count feature, the read error rate feature and the spin retry count feature, so as to construct a normalization layer of the feature extraction model by using the target feature vector.

[0149] Optionally, in an embodiment of the present application, the method further comprises: a fourth determining module, a second calculating module, a third calculating module, a fourth calculating module, a fifth calculating module and a second constructing module.

[0150] The fourth determining module is configured to determine the size and the step length of the sliding window based on the target running state data before inputting the running state data and the environment data into the pre-constructed feature extraction model.

[0151] The second calculating module is configured to calculate the wear trend feature and the wear acceleration feature of the target hard disk based on the size, the step length and the wear degree feature.

[0152] The third calculating module is configured to calculate the bad block fluctuation rate feature and the bad block distribution entropy feature of the target hard disk based on the size, the step length and the bad block event count feature.

[0153] The fourth calculating module is configured to calculate the read error trend feature of the target hard disk based on the size, the step length and the read error rate feature.

[0154] The fifth calculating module is configured to calculate the spin retry fluctuation rate feature of the target hard disk based on the size, the step length and the spin retry count feature.

[0155] The second constructing module is configured to determine a target time sequence feature corresponding to the target running state data based on the wear trend feature, the wear acceleration feature, the bad block fluctuation rate feature, the bad block distribution entropy feature, the read error trend feature and the spin retry fluctuation rate feature, so as to construct a feature layer of the feature extraction model by using the target time sequence feature.

[0156] The description of the features in the embodiments of the hard disk failure prediction device can refer to the related description of the embodiments of the hard disk failure prediction method, which will not be repeated here.

[0157] The hard disk failure prediction device provided in the embodiments of the present application can simultaneously input the running state data and environmental data of hard disks of at least two protocol types in at least two time periods into a pre-constructed feature extraction model, and then obtain corresponding feature vectors and time sequence features, and determine the weight coefficients of the attention layer in the failure prediction model, so as to output the predicted failure probability, predicted failure level and predicted failure type of the hard disk by using the failure prediction model. Therefore, the technical problem that the monitoring mode based on threshold alarm has high cost and false alarm rate and is difficult to meet the demand of large-scale hard disk management can be solved. The single SMART parameter cannot accurately capture the real change trend of the health status of the hard disk, and the influence of environmental factors is not fully considered, which cannot adapt to the complex and changeable actual operating environment and is prone to false alarms. In addition, there are protocol differences in the definition and representation of the SMART parameter, which cannot be directly compared and uniformly analyzed, increasing the complexity of operation and maintenance. The technical effects of reducing the failure false alarm rate, improving the operation and maintenance efficiency, constructing a failure prediction model suitable for different protocol types, eliminating the management difficulties caused by protocol differences, fully considering the influence of environmental factors on the health status of the hard disk, and improving the accuracy of prediction are achieved.

[0158] The embodiments of the present application also provide an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above-mentioned hard disk failure prediction method embodiments.

[0159] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any one of the above-mentioned hard disk failure prediction method embodiments when running.

[0160] In an exemplary embodiment, the above-mentioned computer readable storage medium can include but is not limited to a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0161] The embodiments of the present application also provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the steps in any one of the above-mentioned hard disk failure prediction method embodiments.

[0162] The embodiment of the present application further provides another computer program product, comprising a nonvolatile computer readable storage medium, the nonvolatile computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps in any of the above hard disk failure prediction method embodiments.

[0163] Those skilled in the art can further understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in general terms. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0164] The above describes in detail a hard disk failure prediction method provided by the present application. The principles and implementation modes of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A hard disk failure prediction method characterized by, The method comprises the following steps: Collecting running state data and environment data of at least two protocol types of hard disks in at least two time periods; Inputting the running state data and the environment data into a pre-constructed feature extraction model, normalizing the running state data by using a normalization layer of the feature extraction model, generating a feature vector corresponding to the running state data, and extracting features of the feature vector by using a feature layer of the feature extraction model to output time sequence features corresponding to the feature vector; Inputting the environment data and the time sequence features into a pre-constructed fault prediction model, calculating an interference factor of an attention layer by using the environment data, determining a weight coefficient of the attention layer in the fault prediction model according to the interference factor, and outputting a predicted failure probability, at least one of a predicted failure level and a predicted failure type of the hard disk based on the weight coefficient; The determination of the weight coefficient of the attention layer in the fault prediction model comprises: Obtaining feature values of the time sequence features in the at least two time periods; Based on the feature values, adjust the weight coefficient in the corresponding period to obtain the adjusted weight coefficient; Before inputting the running state data and the environment data into the pre-constructed feature extraction model, further comprising: Identifying the protocol type of the target hard disk; Based on the protocol type and the target running state data of the target hard disk, extracting the wear degree feature, the bad block event count feature, the read error rate feature and the spin retry count feature of the target hard disk; Based on the wear degree feature, the bad block event count feature, the read error rate feature and the spin retry count feature, determine the target feature vector corresponding to the target running state data, and use the target feature vector to construct the normalization layer of the feature extraction model; Before inputting the running state data and the environment data into the pre-constructed feature extraction model, further comprising: Based on the target running state data, determine the size and step length of the sliding window; Based on the size, the step length and the wear degree feature, calculate the wear trend feature and the wear acceleration feature of the target hard disk; Based on the size, the step length and the bad block event count feature, calculate the bad block fluctuation rate feature and the bad block distribution entropy feature of the target hard disk; Based on the size, the step length and the read error rate feature, calculate the read error trend feature of the target hard disk; Based on the size, the step length and the spin retry count feature, calculate the spin retry fluctuation rate feature of the target hard disk; Based on the wear trend feature, the wear acceleration feature, the bad block fluctuation rate feature, the bad block distribution entropy feature, the read error trend feature and the spin retry fluctuation rate feature, determine the target time sequence feature corresponding to the target running state data, and use the target time sequence feature to construct the feature layer of the feature extraction model; The output of the predicted failure type of the hard disk based on the weight coefficient comprises: calculating a first sum value between a wear trend feature and a wear acceleration feature, a second sum value between a bad block fluctuation rate feature and a bad block distribution entropy feature, and a third sum value between a read error trend feature and a rotation retry fluctuation rate feature based on the weight coefficients; comparing the first sum value, the second sum value, and the third sum value; determining that the predicted failure type is a media aging failure type when the first sum value is greater than the second sum value and the first sum value is greater than the third sum value; determining that the predicted failure type is a media scratch failure type when the second sum value is greater than the first sum value and the second sum value is greater than the third sum value; determining that the predicted failure type is a mechanical failure type when the third sum value is greater than the first sum value and the third sum value is greater than the second sum value.

2. The method of claim 1, wherein, Before inputting the running state data and the environment data into a pre-constructed feature extraction model, the method further comprises: judging whether a change value of a different environment parameter in the environment data is greater than a preset threshold value; if the change value is greater than the preset threshold value, allowing the environment data of the corresponding environment parameter to be input into the feature extraction model; if the change value is less than or equal to the preset threshold value, prohibiting the environment data of the corresponding environment parameter from being input into the feature extraction model.

3. The method of claim 2, wherein, The method of determining the weight coefficients of the attention layer in the failure prediction model comprises: determining initial weight coefficients of the attention layer in the failure prediction model based on the time series features; calculating an interference factor of the attention layer based on the environment parameters; calculating a weight decay coefficient corresponding to the initial weight coefficients by using the interference factor, so as to adjust the initial weight coefficients by using the weight decay coefficient to obtain adjusted weight coefficients.

4. The method of claim 1, wherein, The method further comprises: in a case where the predicted failure level is a first failure level, determining that the hard disk has not failed; in a case where the predicted failure level is a second failure level, determining that the hard disk has a failure risk and alarming to a user to remind the user to pay attention; in a case where the predicted failure level is a third failure level, determining that the hard disk is unusable and replacing the hard disk.

5. An electronic device, comprising: The method comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the hard disk failure prediction method according to any one of claims 1-4.

6. A computer-readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the hard disk failure prediction method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Storage hard disk remote diagnosis system and method based on Internet of Things

    CN120256178A