Hard disk fault prediction method and device
Through the nonlinear filtering mechanism, the target index values during hard disk operation are processed and classified, which solves the problem of low accuracy in hard disk failure prediction in the existing technology, and achieves accurate prediction of hard disk failures and improves data stability and security.
Patent Information
- Application Number
- CN202411284095.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-06-03
AI Technical Summary
In the prior art, the accuracy of hard disk failure prediction is low, and early warning cannot be made before the failure, resulting in an increase in the risk of data loss.
By obtaining the target index value sequence of different target indicators during hard disk operation, filtering the target index value based on the nonlinear filtering mechanism, eliminating errors and noise, and classifying the filtered target index value based on the reference index value, and outputting a fault warning prompt.
Accurate prediction of hard disk failures is achieved, the risk of data loss is reduced, and the stability and security of hard disk data is improved.
Smart Images

Figure CN120086076A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a hard disk failure prediction method and device. Background Art
[0002] A hard disk drive (HDD) is the most important storage device of a computer. Since most of the data in the computer is stored in the hard disk, the stable operation of the hard disk is a prerequisite for the stable operation of the computer.
[0003] However, due to factors such as long-term operation, environmental factors, or physical wear, the hard disk may also fail. In the existing related technologies, most of them make judgments and give warnings only after the hard disk has actually been damaged, and it is impossible to give early warnings before the failure. Even if the data on the hard disk is repairable, there is still a risk of data loss, which will reduce the stability and security of the hard disk data to a certain extent. Currently, there are also related fault prediction solutions, but most of them directly judge and predict using the operating parameters of the hard disk itself, with low prediction accuracy, or train complex machine learning models for prediction, with high computational complexity, and it is impossible to accurately predict hard disk failures for hard disks of different specifications. Summary of the Invention
[0004] The present invention provides a hard disk failure prediction method and device to solve the problem of low accuracy in hard disk failure prediction in the prior art.
[0005] The present invention provides a hard disk failure prediction method, including the following steps.
[0006] Obtain a target index value sequence of different target indexes during the operation of the hard disk, and each target index corresponds to a target index value sequence.
[0007] Perform filtering processing on the target index values in the target index value sequence based on a non-linear filtering mechanism.
[0008] Classify the filtered target index values in the target index value sequence based on different categories of reference index values to determine the target category where the filtered target index values are located.
[0009] When the target category is the warning reference index value category, output a failure warning prompt for the target index corresponding to the filtered target index value.
[0010] According to the hard disk failure prediction method provided by the present invention, before obtaining the target index value sequence of different target indexes during the operation of the hard disk, the following steps are further included.
[0011] Obtain all source indexes related to the operation of the hard disk.
[0012] Query the historical hard disk failure data, analyze the N source metrics that cause the most failures, and determine the N source metrics as the target metrics, where N is greater than or equal to 1.
[0013] According to a hard disk failure prediction method provided by the present invention, the target metrics include at least one of the number of remapped sectors, the number of errors that cannot be recovered by hardware ECC, the number of operations terminated due to hard disk timeout, the number of unstable sectors that have not been remapped currently, and the total number of uncorrectable errors that occur during read and write sectors.
[0014] According to a hard disk failure prediction method provided by the present invention, the filtering process of the target metric values in the target metric value sequence based on the non-linear filtering mechanism includes the following steps.
[0015] Determine the sliding window size for filtering the target metric value sequence, obtain multiple sliding windows on the target metric value sequence, and each sliding window includes n target metric values, where n is an odd number greater than 1.
[0016] Determine the polynomial used to fit the target metric values within the sliding window. The dependent variable of the polynomial is the target metric value within the sliding window, the independent variable is the number of the target metric value within the sliding window, the order of the polynomial is k - 1, and in the polynomial, the coefficient corresponding to each order of the independent variable is the fitting parameter, where n is greater than or equal to k.
[0017] For each sliding window, construct a fitting equation set based on the target metric values within the current sliding window and the polynomial, solve the fitting equation set by the least squares method to obtain the fitting parameters, and substitute the fitting parameters into the polynomial to obtain multiple fitted target metric values within the current sliding window.
[0018] Determine the fitted target metric value corresponding to the middlemost number within the current sliding window as the filtered target metric value within the sliding window.
[0019] According to a hard disk failure prediction method provided by the present invention, the classification of the filtered target metric values in the target metric value sequence based on different categories of reference metric values to determine the target category where the filtered target metric value is located includes the following steps.
[0020] Calculate the Euclidean distance between the filtered target metric value and the reference metric values of different categories.
[0021] Determine the target category based on the categories to which the K reference metric values with the smallest Euclidean distance belong.
[0022] A hard disk failure prediction method provided by the present invention further includes: when the target category is the category of failure reference index values, outputting a failure warning prompt for the target index corresponding to the filtered target index value.
[0023] The present invention also provides a hard disk failure prediction device, including the following modules.
[0024] An index value sequence acquisition module, configured to acquire target index value sequences of different target indexes during the operation of the hard disk, and each target index corresponds to a target index value sequence.
[0025] A filtering processing module, configured to perform filtering processing on the target index values in the target index value sequence based on a non-linear filtering mechanism.
[0026] A target category determination module, configured to classify the filtered target index values in the target index value sequence based on reference index values of different categories to determine the target category where the filtered target index values are located.
[0027] A failure early warning prompt module, configured to output a failure early warning prompt for the target index corresponding to the filtered target index value when the target category is the category of early warning reference index values.
[0028] The present invention also provides a hard disk backplane, on which a complex programmable logic device or a field programmable gate array is provided, and the complex programmable logic device or the field programmable gate array is connected to the hard disk and is configured to execute the hard disk failure prediction method described in any one of the above.
[0029] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the hard disk failure prediction method described in any one of the above is implemented.
[0030] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the hard disk failure prediction method described in any one of the above is implemented.
[0031] The hard disk failure prediction method and device provided by the present invention perform filtering processing on the target index values in the target index value sequence through a non-linear filtering mechanism, removing errors and noises in the target index value sequence, making the filtered target index values more accurate and reliable, and classifying the filtered target index values in the target index value sequence according to the preset reference index value categories. When the classified target category is the category of early warning reference index values, a failure early warning prompt for the target index corresponding to the filtered target index value is output, thereby realizing accurate prediction of hard disk failures. Description of the Drawings
[0032] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0033] Figure 1 It is a schematic flowchart of the hard disk failure prediction method provided by the present invention.
[0034] Figure 2 It is a schematic diagram of SG filtering in the hard disk failure prediction method provided by the present invention.
[0035] Figure 3 It is a schematic diagram of KNN classification in the hard disk failure prediction method provided by the present invention.
[0036] Figure 4 It is a schematic diagram of an application scenario of the hard disk failure prediction method provided by the present invention.
[0037] Figure 5 It is a schematic structural diagram of the hard disk failure prediction device provided by the present invention.
[0038] Figure 6 It is a schematic structural diagram of the electronic device provided by the present invention. Specific Embodiments
[0039] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0040] In the existing related technologies, hard disks generally have the built-in SMART (Self-Monitoring Analysis and Reporting Technology) function. The SMART function is used to monitor various parameters during the operation of the hard disk, such as temperature, data read error rate, disk read and write performance, number of unexpected power failures, etc., and compare the parameters with the safe operation values preset by the manufacturer for a healthy hard disk. If the current operating parameters exceed the safe operating range, the hard disk will alarm the server monitoring system and perform self-repair. Although the SMART function can provide various parameters during the operation of the hard disk, this function only compares various parameters with the safe range preset by the manufacturer. When the range is exceeded, it judges whether the fault can be self-repaired and alarms when it cannot be repaired. Moreover, the built-in SMART function of the hard disk is not intelligent enough in monitoring the hard disk. When the hard disk frequently has various types of faults that can be self-repaired, the operation of the hard disk is already in an unstable state, yet the SMART function still judges that the hard disk is in good condition. Therefore, using the built-in SMART function cannot timely and accurately predict hard disk faults.
[0041] Another optimization solution is to use machine learning technology to build a machine learning model, simulate and emulate various parameters in SMART, analyze the changes in various parameters when the hard disk runs from a healthy state to a state close to a hard disk failure through a large number of samples, retain the characteristic items related to the hard disk operation failure, remove irrelevant parameters, and finally build a machine learning model for deriving the hard disk failure rate based on the changes in hard disk parameters. However, the computational amount of the machine learning model is extremely large, and the data cannot be fed back in real time. Only the model can be trained in advance, and then the real-time parameters of the hard disk are substituted later for failure prediction. In addition, for hard disks of different specifications, re-training is required to obtain an accurate prediction model for accurate prediction. Therefore, the solution of predicting hard disk faults through the machine learning model cannot accurately predict hard disk faults either, and the computational amount is large.
[0042] In view of the above problems existing in the existing related technologies, the embodiment of the present invention provides a hard disk fault prediction method, and the specific process is as Figure 1 shown, including the following steps S110 to S140.
[0043] Step S110: Obtain the target index value sequences of different target indexes during the operation of the hard disk, and each target index corresponds to a target index value sequence. Among them, the SMART parameters can be used as the target indexes. For a target index, the target index value sequence is a sequence composed of multiple target index values collected at the acquisition frequency of the SMART function during the operation of the hard disk. Each target index corresponds to a target index value sequence. For example, for the hard disk temperature, it corresponds to a temperature value sequence, and the temperature value sequence includes the temperature values of the hard disk at each time acquisition point.
[0044] Step S120: Filter the target index values in the target index value sequence based on a non - linear filtering mechanism. For a relatively complex system such as a hard disk, there are many indicators during the working process. During data acquisition, the corresponding sensors or acquisition programs will inevitably introduce error points and noise, and most of these indicators have non - linear characteristics. Therefore, in this embodiment, a non - linear filtering mechanism is used to filter the target index values in the target index value sequence, so as to retain the original waveform characteristics of the indicators and extract the data characteristics of the indicators themselves as much as possible under the condition of noise interference, and eliminate errors and noise as much as possible, making the filtered target index values more accurate and reliable.
[0045] Step S130: Classify the filtered target index values in the target index value sequence based on different categories of reference index values to determine the target category where the filtered target index values are located. Specifically, the reference index value categories at least include the warning reference index value category. For example, the normal operating temperature of a mechanical hard disk is between 5 - 55 degrees Celsius. For the temperature indicator, the temperature ranges corresponding to the warning reference index value category are 5 - 10 degrees Celsius and 50 - 55 degrees Celsius, while the temperature range corresponding to the safety reference index value category is 5 - 55 degrees Celsius, and the temperature ranges below 5 degrees Celsius and above 55 degrees Celsius are the temperature ranges corresponding to the failure reference index value category. Determine the target category where the current filtered hard disk temperature is located through the temperature ranges corresponding to the above three categories.
[0046] Among them, the warning reference index values, safety reference index values, and failure reference index values can all be selected based on the empirical values in the actual use process, and at the same time, the operation results of some mature models can also be referred to. Moreover, multiple values within their respective category ranges can be selected for each category of reference index values to facilitate subsequent classification processing.
[0047] Step S140: When the target category is the warning reference index value category, output a fault warning prompt for the target index corresponding to the filtered target index value. For example, when the current filtered hard disk temperature reaches the range of 5 - 10 degrees Celsius and 50 - 55 degrees Celsius, determine the target category of the current filtered hard disk temperature as the warning reference index value category. At this time, it can be warned that the temperature is too high, and a prediction result of hard disk failure due to temperature can be made to remind the maintenance personnel to make corresponding treatments, thus realizing the prediction of hard disk failure.
[0048] In the hard disk failure prediction method of this embodiment, the target index values in the target index value sequence are filtered through a non-linear filtering mechanism to eliminate errors and noises in the target index value sequence, making the filtered target index values more accurate and reliable. Then, the filtered target index values in the target index value sequence are classified according to the preset reference index value categories. When the classified target category is the warning reference index value category, a failure warning prompt for the target index corresponding to the filtered target index value is output, thereby realizing accurate prediction of hard disk failures.
[0049] Since there are many indexes during the operation of the hard disk, that is, there are many SMART parameters, and not all of them are related to hard disk failures. If all indexes are used for hard disk failure prediction, it will increase unnecessary computational overhead. Therefore, in some embodiments, before step S110, it further includes: obtaining all source indexes related to the operation of the hard disk, that is, obtaining all SMART parameters, querying the hard disk historical failure data, analyzing the N source indexes with the most failure times, and determining the N source indexes as the target indexes, where N is greater than or equal to 1, and the target indexes and their quantities can be set according to different specifications of the hard disk. The N source indexes with the most failure times have a relatively high correlation with hard disk failures. They can not only accurately predict hard disk failures, but also reduce the number of indexes used in prediction, thereby reducing the amount of calculation and improving the prediction efficiency.
[0050] In some embodiments, the target indexes include at least one of the following: the number of remapped sectors (SMART 5), the number of errors that cannot be recovered by hardware ECC (SMART 187), the number of operations terminated due to hard disk timeout (SMART 188), the number of unstable sectors that have not been remapped currently (SMART 197), and the total number of uncorrectable errors that occur during read and write sectors (SMART 198). Specifically, the five source indexes of SMART 5, SMART 187, SMART 188, SMART 197, and SMART 198 have the highest correlation with hard disk failures. For a healthy hard disk, the index values corresponding to these five indexes should all be 0. When non-zero values appear, it means that the hard disk begins to show abnormalities and may even lead to failures. Preferably, these five source indexes are used as the target indexes.
[0051] In addition, parameters such as SMART 12: power-on cycle count, SMART 190: internal disk platter air flow temperature of the hard disk, SMART 191: frequency of errors caused by external vibrations of the hard disk, SMART 192: number of power failures outside the hard disk, and SMART 194: current operating temperature of the hard disk are also related to hard disk failures, and different parameters can be determined as target indexes according to actual needs.
[0052] In some embodiments, the step S120 specifically uses a Savitzky-Golay (SG) filtering algorithm to filter the target indicator value sequence, and specifically includes the following steps.
[0053] Step 1: Determine the sliding window size for filtering the target index value sequence, and obtain multiple sliding windows on the target index value sequence, each of which includes n target index values, where n is an odd number greater than 1. Figure 2 As shown, the current target indicator value to be processed is , the m points before and after are respectively recorded as y -m ,y -m+1 , …, y 0 , …, y m-1 ,y m Therefore, the sliding window width n=2m+1, and all target index values to be filtered can be traversed by moving the sliding window in sequence.
[0054] Specifically, when the number of target index values contained in the sliding window is large, the filtering effect can be enhanced, but it is also easy to lose the key information of the target index value, resulting in data distortion. In this embodiment, the size of the sliding window is set to 3-7. Figure 2 In the example, the size of the sliding window is set to 5, that is, the sliding window includes 5 target index values. This ensures the filtering effect while avoiding data distortion as much as possible, making the final filtered target index value more accurate.
[0055] Step 2: Determine a polynomial for fitting the target index value in the sliding window, the dependent variable of the polynomial is the target index value in the sliding window, the independent variable is the number of the target index value in the sliding window, the order of the polynomial is k-1, and in the polynomial, the coefficient corresponding to each order independent variable is the fitting parameter, where n is greater than or equal to k, and the polynomial expression is as follows.
[0056] .
[0057] in, y Indicates the target indicator value, x Indicates the target index value of the processing point y The corresponding number in the sliding window, that is, the corresponding number is selected from -m, -m+1, ..., 0, ..., m-1, m, ~ Both represent fitting parameters. If the polynomial order k is set too high, more detailed features of the target index value can be retained, but it may also introduce fitting errors or noise and increase the amount of calculation. The k value can be selected according to the actual situation. It is preferred that k is less than n. For example, when n=5, k=3, while retaining more data details, try to avoid introducing errors and noise.
[0058] Step 3: For each of the sliding windows, based on the target metric values within the current sliding window and the polynomial, construct a fitting equation system, solve the fitting equation system by the least squares method to obtain the fitting parameters, and substitute the fitting parameters into the polynomial to obtain multiple fitting target metric values within the current sliding window.
[0059] Specifically, by performing fitting processing on all the target metric values within a sliding window using the above polynomial, the following fitting equation system can be obtained.
[0060] 。
[0061] Simplify the fitting equation system into an expression: 。
[0062] ; ; 。
[0063] Fitting coefficient A The least squares solution of is:
[0064] Therefore, the Y value after filtering is: 。
[0065] Step 4: Determine the fitting target metric value corresponding to the middle number within the current sliding window as the filtered target metric value within the sliding window. For example, as Figure 2 shown, take the target metric value y 0 with number i = 0 in each sliding window as the filtered target metric value.
[0066] To avoid abnormal SMART function feedback parameters caused by environmental reasons, sensor or hard disk itself reasons, in this embodiment, the SG filtering algorithm is used to filter the target metric value sequence. The SG filtering algorithm is applicable to filtering non-linear data containing various noises and can estimate the true state of the hard disk operation state under the condition of noise interference, making the filtered target metric value more accurate.
[0067] In some embodiments, step S130 may use the K-Nearest Neighbor (KNN) classification algorithm to classify the filtered target metric values, including the following steps.
[0068] Step 1: Calculate the Euclidean distance between the filtered target metric value and the reference metric values of different categories d , and the calculation formula is as follows.
[0069] 。
[0070] Among them, w represents the dimension of the target index value. For one-dimensional target index values such as temperature and number of times (e.g., number of remapped sectors), w = 1, x i represents the reference index value, y i represents the filtered target index value.
[0071] Step 2: Determine the target category based on the categories to which the K reference index values with the smallest Euclidean distance belong. Specifically, as Figure 3 shown, among the K reference index values closest to the target index value, the category to which the largest number of reference index values corresponds is the target category. Among them, the value of K determines the classification accuracy. When the value of K is too small, it is easily affected by abnormal samples and overfitting occurs; when the value of K is too large, the sample balance deteriorates, and samples with a large distance from the target index value will also affect the classification, ultimately causing the model to underfit. In practical applications, it can be determined according to the number of target indexes related to hard disk failures selected and the number of reference index values corresponding to different categories for each target index.
[0072] The KNN algorithm is relatively simple and easy to implement, and can accurately classify the target index value, that is, accurately determine the current state of the hard disk, so as to achieve accurate prediction of hard disk failures.
[0073] In some embodiments, the hard disk failure prediction method further includes: when the target category is the category of fault reference index values, output a fault warning prompt for the target index corresponding to the filtered target index value, that is, after a fault warning, if the fault is not handled in time, the hard disk will probably not work properly and needs to be repaired. Output a fault warning prompt to remind the maintenance personnel to repair or replace the hard disk.
[0074] The SMART function has a self - repair function. When the hard disk frequently experiences various types of self - repairable faults, the operation of the hard disk is already in an unstable state. In addition, the SMART function cannot predict the chain reaction caused by faults. Frequent small controllable faults are very likely to cause uncontrollable large faults, but the SMART function will judge the status as normal according to its logic until an uncontrollable large fault occurs and then gives an alarm. At this time, the hard disk is often unusable and even the data is damaged. In the above - mentioned embodiments, the target index is obtained from the SMART parameters. Therefore, there will be a situation of repeated fault warnings. Although no fault alarm is given, after the hard disk continues to self - repair or is repaired manually and then runs, it will cause the hard disk to directly fail without time for warning prompts in a short time, resulting in service interruption and data loss. To avoid the above problems, in some embodiments, after step S140, that is, after outputting the fault warning prompt for the target index corresponding to the filtered target index value, the number of fault warning prompts corresponding to the target index is counted. When the number of fault warning prompts reaches the preset number threshold, a fault alarm message is output and the hard disk operation is safely stopped. Among them, the preset number threshold can be set according to different target indexes. Different target indexes correspond to different preset number thresholds, which can be specifically determined according to the number of times different target indexes are abnormal after a fault occurs in the historical fault data. The preset number threshold is less than the number of times the target index is abnormal.
[0075] The calculation amount of the hard disk fault prediction method in the above - mentioned embodiments is small. Especially compared with the prediction method of machine learning, the calculation amount is greatly reduced and it can be deployed to run in a Complex Programmable Logic Device (CPLD) or a Field - Programmable Gate Array (FPGA). As Figure 4 shown, it is a schematic diagram of the application scenario of the hard disk fault prediction method in the above - mentioned embodiments. The execution program of the hard disk fault prediction method is deployed in the CPLD or FPGA on the hard disk backplane. When the BIOS system runs, the CPU obtains the SMART parameters (i.e., the source index and its index value) of the hard disk through the PCIE / SATA bus, and then transmits the SMART parameters to the BMC through I2C. The BMC transmits the SMART parameters to the CPLD or FPGA on the hard disk backplane through I2C, and the CPLD or FPGA executes the hard disk fault prediction method in the above - mentioned embodiments.
[0076] Specifically, there is also a warning indicator light (RED LED) on the CPLD or FPGA. When the target category to which the filtered target index value belongs is the category of safety reference index values, the CPLD or FPGA does not perform the lighting operation.
[0077] When the target category to which the filtered target index value belongs is the early warning reference index value category, it means that the operating life of the hard disk is already relatively low. The CPLD or FPGA makes the red light flash at a frequency of 4 Hz to indicate to the user to back up the files in the hard disk in time or directly replace the hard disk.
[0078] When the target category to which the filtered target index value belongs is the fault reference index value category, the CPLD or FPGA makes the red light stay on constantly. The hard disk is very likely unable to work properly and needs to be repaired.
[0079] The hard disk fault prediction device provided by the present invention will be described below. The hard disk fault prediction device described below can be mutually corresponding and referred to the hard disk fault prediction method described above.
[0080] The hard disk fault prediction device according to an embodiment of the present invention, as Figure 5 shown, includes the following modules.
[0081] The target index value sequence acquisition module 510 is used to acquire the target index value sequences of different target indexes during the operation of the hard disk, and each target index corresponds to a target index value sequence.
[0082] The filtering processing module 520 is used to perform filtering processing on the target index values in the target index value sequence based on a non-linear filtering mechanism.
[0083] The target category determination module 530 is used to classify the filtered target index values in the target index value sequence based on reference index values of different categories to determine the target category where the filtered target index values are located.
[0084] The fault early warning prompt module 540 is used to output a fault early warning prompt for the target index corresponding to the filtered target index value when the target category is the early warning reference index value category.
[0085] The hard disk fault prediction device of this embodiment performs filtering processing on the target index values in the target index value sequence through a non-linear filtering mechanism, eliminates errors and noises in the target index value sequence, makes the filtered target index values more accurate and reliable, and classifies the filtered target index values in the target index value sequence according to the preset reference index value categories. When the classified target category is the early warning reference index value category, a fault early warning prompt for the target index corresponding to the filtered target index value is output, thereby realizing accurate prediction of hard disk faults.
[0086] In some embodiments, the hard disk failure prediction device further includes: a target metric determination module, configured to obtain all source metrics related to the operation of the hard disk before obtaining the sequence of target metric values of different target metrics during the operation of the hard disk; query the hard disk historical failure data, analyze the N source metrics that cause the most failures, and determine the N source metrics as the target metrics, where N is greater than or equal to 1.
[0087] In some embodiments, the target metrics include at least one of the number of remapped sectors, the number of errors that cannot be recovered using hardware ECC, the number of operations terminated due to hard disk timeout, the number of unstable sectors that have not been remapped currently, and the total number of uncorrectable errors that occur during read and write sectors.
[0088] In some embodiments, the filtering processing module 520 specifically includes the following modules.
[0089] A sliding window determination module, configured to determine the size of the sliding window for filtering the sequence of target metric values, obtain a plurality of sliding windows on the sequence of target metric values, and each sliding window includes n target metric values, where n is an odd number greater than 1.
[0090] A polynomial determination module, configured to determine a polynomial for fitting the target metric values within the sliding window, where the dependent variable of the polynomial is the target metric values within the sliding window, the independent variable is the number of the target metric values within the sliding window, the order of the polynomial is k - 1, and in the polynomial, the coefficient corresponding to each order of the independent variable is a fitting parameter, where n is greater than or equal to k.
[0091] A fitting and filtering module, configured to, for each sliding window, construct a fitting equation set based on the target metric values within the current sliding window and the polynomial, solve the fitting equation set by the least squares method to obtain the fitting parameters, and substitute the fitting parameters into the polynomial to obtain a plurality of fitted target metric values within the current sliding window.
[0092] A filtered value determination module, configured to determine the fitted target metric value corresponding to the middlemost number within the current sliding window as the filtered target metric value within the sliding window.
[0093] In some embodiments, the target category determination module 530 is specifically configured to calculate the Euclidean distance between the filtered target metric values and the reference metric values of different categories; determine the target category based on the categories to which the K reference metric values with the smallest Euclidean distance belong.
[0094] In some embodiments, the hard disk failure prediction device further includes: a failure warning prompt module, configured to output a failure warning prompt for the target indicator corresponding to the filtered target indicator value when the target category is the failure reference indicator value category.
[0095] The present invention also provides a hard disk backplane, on which a complex programmable logic device or a field programmable gate array is provided. The complex programmable logic device or the field programmable gate array is connected to the hard disk and is configured to execute the hard disk failure prediction method described in any of the above embodiments. Refer to Figure 4 , deploy the execution program of the hard disk failure prediction method in the CPLD or FPGA of the hard disk backplane. When the BIOS system runs, the CPU obtains the SMART parameters (i.e., the source indicators and their indicator values) of the hard disk through the PCIE / SATA bus, and then transmits the SMART parameters to the BMC through I2C. The BMC transmits the SMART parameters to the CPLD or FPGA of the hard disk backplane through I2C, and the CPLD or FPGA executes the hard disk failure prediction method of the above embodiment.
[0096] Figure 6 Illustrates a schematic physical structure diagram of an electronic device, as Figure 6 shown. The electronic device may include: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 complete mutual communication through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute the hard disk failure prediction method, and the method includes the following steps.
[0097] Obtain a target indicator value sequence of different target indicators during the operation of the hard disk, and each target indicator corresponds to a target indicator value sequence.
[0098] Perform filtering processing on the target indicator values in the target indicator value sequence based on a non-linear filtering mechanism.
[0099] Classify the filtered target indicator values in the target indicator value sequence based on different categories of reference indicator values to determine the target category where the filtered target indicator values are located.
[0100] When the target category is the early warning reference indicator value category, output a failure early warning prompt for the target indicator corresponding to the filtered target indicator value.
[0101] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0102] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the hard disk failure prediction method provided by the above-mentioned various methods. The method includes the following steps.
[0103] Obtain the target index value sequences of different target indexes during the operation of the hard disk, and each target index corresponds to a target index value sequence.
[0104] Filter the target index values in the target index value sequence based on a non-linear filtering mechanism.
[0105] Classify the filtered target index values in the target index value sequence based on different categories of reference index values to determine the target category where the filtered target index values are located.
[0106] In the case where the target category is the warning reference index value category, output a fault warning prompt for the target index corresponding to the filtered target index value.
[0107] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the hard disk failure prediction method provided by the above-mentioned various methods. The method includes the following steps.
[0108] Obtain the target index value sequences of different target indexes during the operation of the hard disk, and each target index corresponds to a target index value sequence.
[0109] Filter the target index values in the target index value sequence based on a non-linear filtering mechanism.
[0110] Classify the filtered target index values in the target index value sequence based on reference index values of different categories to determine the target category where the filtered target index values are located.
[0111] When the target category is the early warning reference index value category, output a fault warning prompt for the target index corresponding to the filtered target index value.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hard disk failure prediction method, characterized in that: include: Obtain target indicator value sequences of different target indicators during the operation of the hard disk, where each target indicator corresponds to a target indicator value sequence; Performing filtering processing on the target indicator values in the target indicator value sequence based on a nonlinear filtering mechanism; Classifying the filtered target indicator values in the target indicator value sequence based on reference indicator values of different categories to determine the target category to which the filtered target indicator values belong; In the case where the target category is a warning reference indicator value category, a fault warning prompt of a target indicator corresponding to the filtered target indicator value is output.
2. The hard disk failure prediction method according to claim 1, characterized in that: Before obtaining the target indicator value sequence of different target indicators during the operation of the hard disk, the method further includes: Get all source metrics related to hard drive operation; Query the hard disk historical failure data, analyze the N source indicators that cause the most failures, and determine the N source indicators as the target indicators, where N is greater than or equal to 1.
3. The hard disk failure prediction method according to claim 2, characterized in that: The target indicators include: at least one of the number of remapped sectors, the number of errors that cannot be recovered using hardware ECC, the number of operations terminated due to hard disk timeout, the number of unstable sectors that have not yet been remapped, and the total number of uncorrectable errors in read and write sectors.
4. The hard disk failure prediction method according to claim 1, characterized in that: The filtering process of the target indicator values in the target indicator value sequence based on the nonlinear filtering mechanism includes: Determine a sliding window size for filtering the target indicator value sequence, and obtain a plurality of sliding windows on the target indicator value sequence, each of the sliding windows including n target indicator values, wherein n is an odd number greater than 1; Determine a polynomial for fitting the target index value in the sliding window, wherein the dependent variable of the polynomial is the target index value in the sliding window, the independent variable is the number of the target index value in the sliding window, the order of the polynomial is k-1, and in the polynomial, the coefficient corresponding to each order independent variable is a fitting parameter, wherein n is greater than or equal to k; For each of the sliding windows, a fitting equation group is constructed based on the target index value in the current sliding window and the polynomial, the fitting equation group is solved by the least square method to obtain the fitting parameters, and the fitting parameters are substituted into the polynomial to obtain a plurality of fitting target index values in the current sliding window; The fitting target index value corresponding to the middle number in the current sliding window is determined as the filtered target index value in the sliding window.
5. The hard disk failure prediction method according to claim 1, characterized in that: The classifying the filtered target indicator values in the target indicator value sequence based on reference indicator values of different categories to determine the target category to which the filtered target indicator values belong includes: Calculating the Euclidean distance between the filtered target index value and reference index values of different categories; The target category is determined based on the category to which the K reference index values with the smallest Euclidean distance belong.
6. The hard disk failure prediction method according to any one of claims 1 to 5, characterized in that: Also includes: In the case where the target category is a fault reference indicator value category, a fault alarm prompt of a target indicator corresponding to the filtered target indicator value is output.
7. A hard disk failure prediction device, characterized in that: include: An indicator value sequence acquisition module is used to acquire target indicator value sequences of different target indicators during the operation of the hard disk, and each target indicator corresponds to a target indicator value sequence; A filtering processing module, used for filtering the target indicator values in the target indicator value sequence based on a nonlinear filtering mechanism; A target category determination module, used for classifying the filtered target indicator values in the target indicator value sequence based on reference indicator values of different categories to determine the target category to which the filtered target indicator value belongs; The fault warning prompt module is used to output a fault warning prompt of the target indicator corresponding to the filtered target indicator value when the target category is a warning reference indicator value category.
8. A hard disk backplane, characterized in that: The hard disk backplane is provided with a complex programmable logic device or a field programmable logic gate array, and the complex programmable logic device or the field programmable logic gate array is connected to the hard disk and is used to execute the hard disk failure prediction method as described in any one of claims 1 to 6.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the hard disk failure prediction method according to any one of claims 1 to 6 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the hard disk failure prediction method according to any one of claims 1 to 6 is implemented.