Hard disk fault determination method and device applied to periodic rule

By identifying and utilizing highly contributory features through a machine learning-based model trained on time-series data, the method addresses the issue of inaccurate hard disk failure prediction in existing technologies, enhancing detection accuracy and efficiency.

CN120315918APending Publication Date: 2025-07-15中国邮政储蓄银行股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510324580.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The traditional feature selection method in the prior art is not representative when determining hard disk failure, and cannot be effectively used for accurate detection with the model, resulting in low accuracy and efficiency of hard disk failure detection.

Method used

By determining the characteristic items in the target hard disk running data whose contribution degree is greater than the contribution threshold is the target feature items, and using multiple sets of training data to process the cascading model trained through machine learning, the characteristic items with high contribution to hard disk failure detection are selected, and feature reconstruction and combination are combined with XGBoost and FM models to build a fault determination model.

Benefits of technology

It realizes accurate detection of hard disk failures, improves detection accuracy and efficiency, and solves the problem that traditional feature selection methods cannot effectively combine models for accurate detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120315918A_ABST
    Figure CN120315918A_ABST
Patent Text Reader

Abstract

The invention discloses a hard disk fault determination method and device applied to a periodic rule. The method comprises the following steps: determining a feature item of which the contribution degree is greater than a contribution degree threshold value in operation data of a target hard disk as a target feature item; inputting the target feature item into a fault determination model, and processing the target feature item by using the fault determination model to obtain a fault probability that the target hard disk has a fault; and when the fault probability is greater than a preset probability value, determining that the target hard disk has a fault, or when the fault probability is not greater than the preset probability value, determining that the target hard disk does not have a fault. According to the method and the device, the technical problems that the features selected by a traditional feature selection method are not representative and cannot be effectively combined with a model to accurately detect the hard disk fault when the hard disk fault is determined in the related technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault detection, and in particular, to a method and device for determining hard disk faults applied to periodic rules. Background Art

[0002] At present, there are various types of data storage media. However, the mainstream storage devices for massive data include mechanical hard disks (HDDs) and solid-state drives (SSDs). Considering many factors such as stability and price, mechanical hard disks are still the best choice in various fields at present. A hard disk is a complex system covering the fields of motors, magnetism, mechanics, and fluids, which contains many precision components, has a complex production process, and a high maintenance cost.

[0003] A hard disk is the carrier of data and the guarantee for the stable operation of an information system. Therefore, accurately determining the faults of a hard disk has inestimable value for data backup and system operation. Hard disk fault determination is a key technology, aiming to determine the lifespan of a hard disk and whether it may fail in advance by analyzing important attribute parameters and operating states during the operation of the hard disk. Accurately determining hard disk faults helps to avoid data loss, improve the reliability and availability of the hard disk, and thus plays an important role in data storage and data management. In addition, the reliability and high availability of the system also rely on the healthy operation of the hard disk.

[0004] However, hard disk fault determination still faces some challenges. First, the lifespan of a hard disk is affected by various factors, including usage patterns, environmental conditions, and the quality of the hard disk itself. Therefore, establishing an accurate determination model requires considering the complex interactions of these factors. In addition, since the operating state of a hard disk changes over time and with use, hard disk fault determination belongs to an event with periodic change rules, so the determination model needs to be updated and adjusted in a timely manner. At present, machine learning methods are less applied in the field of hard disk lifespan determination and the training methods are single, without considering periodic changes; furthermore, there are many types of hard disk parameters, and the existing feature selection methods have a single selection method, and feature selection is not effectively combined with model training, that is, the results of feature selection do not necessarily have a positive contribution to the model.

[0005] Aiming at the problem that the features selected by traditional feature selection methods during hard disk fault determination in the above-mentioned related technologies are not representative and cannot effectively combine with the model to accurately detect hard disk faults, no effective solution has been proposed yet. Summary of the Invention

[0006] An embodiment of the present invention provides a method and device for determining hard disk failures applicable to periodic rules, so as to at least solve the technical problem in the related art that the features selected by the traditional feature selection method during hard disk failure determination are not representative and cannot effectively combine with the model to accurately detect hard disk failures.

[0007] According to one aspect of the embodiments of the present invention, a method for determining hard disk failures applicable to periodic rules is provided, including: determining a feature item with a contribution degree greater than a contribution degree threshold in the operation data of the target hard disk as the target feature item, where the target hard disk is the hard disk that needs to be subjected to failure detection, and the contribution degree refers to the importance of the feature item for failure detection of the target hard disk; inputting the target feature item into a failure determination model to process the target feature item by using the failure determination model to obtain a failure probability that the target hard disk has a failure, where the failure determination model is a cascade model obtained by performing periodic training on multiple groups of first training data in a machine learning manner, the multiple groups of first training data are time series data within a preset time period, and each group of the multiple groups of first training data includes: a sample target feature item and a first sample failure probability corresponding to the sample target feature item; in the case where the failure probability is greater than a preset probability value, determining that the target hard disk has the failure, or, in the case where the failure probability is not greater than the preset probability value, determining that the target hard disk does not have the failure.

[0008] Optionally, the contribution degree includes a single contribution degree, and the contribution degree threshold includes a first contribution degree threshold. Determining a feature item with a contribution degree greater than a contribution degree threshold in the operation data of the target hard disk as the target feature item includes: obtaining the current operation data of the target hard disk; parsing the operation data to obtain multiple feature items of the operation data and the field name corresponding to each feature item; determining the first linear coefficient of the sample feature item in the initial failure determination model as the single contribution degree of the feature item according to the field name, where the initial failure determination model is obtained by performing training on multiple groups of second training data in a machine learning manner, and each group of the multiple groups of second training data includes: the sample feature item and a second sample failure probability corresponding to the sample feature item, the sample feature item has the same field name as the feature item, and the single contribution degree refers to the importance of a single feature item for failure detection of the target hard disk; determining the feature item with the single contribution degree greater than the first contribution degree threshold as the target feature item.

[0009] Optionally, the contribution degree includes a combined contribution degree, and the contribution degree threshold includes a second contribution degree threshold. Determining a feature item in the operation data of the target hard disk with a contribution degree greater than the contribution degree threshold as a target feature item includes: determining a combination of any two of the multiple feature items as a feature combination; determining the second linear coefficient of the sample feature combination in the initial fault determination model as the combined contribution degree of the feature combination according to the field names of the two feature items in the feature combination, where the two sample feature items in the sample feature combination have the same field names as the two feature items in the feature combination, and the combined contribution degree refers to the importance of combining the two feature items for fault detection of the target hard disk; determining the feature combination with the combined contribution degree greater than the second contribution degree threshold as a target feature combination, where the second contribution degree threshold is a contribution degree threshold that is the same as or different from the first contribution degree threshold; determining that both of the two feature items in the target feature combination are the target feature items.

[0010] Optionally, the method for determining a hard disk fault applied to periodic patterns further includes: before determining the first linear coefficient of the sample feature item in the initial fault determination model as the contribution degree of the feature item according to the field name, or determining the combined contribution degree of the sample feature combination in the initial fault determination model according to the field names of the two feature items in the feature combination, obtaining multiple pieces of labeled historical operation data of the target hard disk in a preset historical time period according to the time series, where the labeled historical operation data refers to the historical operation data carrying historical fault probability labels; dividing the multiple pieces of labeled historical operation data to obtain a labeled training data set and a labeled validation data set, where the labeled training data set is used for training to obtain the initial fault determination model, and the labeled validation data set is used for performance verification of the initial fault determination model; iteratively training the machine learning model using the labeled training data set until the iteration termination condition is reached, and determining the currently trained machine learning model as the original fault determination model, where the iteration termination condition includes at least one of the following: the number of iterations reaches a preset number of iterations, the prediction error is not greater than the error threshold; using the labeled validation data set to perform performance verification on the original fault determination model to obtain a verification result; in the case where the verification result indicates that the performance score of the original fault determination model is higher than the score threshold, determining the original fault determination model as the initial fault determination model.

[0011] Optionally, after obtaining multiple pieces of labeled historical operation data of the target hard disk within a preset historical time period according to the time series, the hard disk fault determination method applied to periodic rules further includes: removing the incomplete labeled historical operation data among the multiple pieces of labeled historical operation data, where the incomplete labeled historical operation data refers to the labeled historical operation data with missing values; performing one-hot encoding processing on the discrete feature items among the multiple pieces of labeled historical operation data to convert the data verification of the discrete feature items into a preset target format, where the discrete feature items refer to the feature items with a finite number or a countable number of discontinuous values; dividing the multiple pieces of labeled historical operation data according to the sample type to obtain multiple sample data sets, where the sample type includes: fault samples and non-fault samples, and each sample data set includes multiple pieces of labeled historical operation data of the same sample type; using an oversampling algorithm to expand the sample data set with the total number of data less than the preset quantity to obtain a target sample data set, where the total number of data refers to the number of labeled historical operation data in the sample data set; performing normalization processing on the labeled historical operation data in the sample data set and the target sample data set to complete the preprocessing operation of the multiple pieces of labeled historical operation data.

[0012] Optionally, using a labeled training data set to iteratively train a machine learning model until the iteration termination condition is reached, and determining the currently trained machine learning model as the original fault determination model, includes: using an XGBoost model to perform feature reconstruction on multiple pieces of labeled historical operation data to obtain multiple reconstructed feature items, where the XGBoost model is one of the machine learning models; using the multiple feature reconstruction items to iteratively train an FM model until the iteration termination condition is reached, and determining the currently trained FM model as the target FM model, where the FM model is a machine learning model different from the XGBoost model; fusing the XGBoost model and the target FM model to obtain the original fault determination model.

[0013] Optionally, when the verification result indicates that the performance score of the original fault determination model is higher than the score threshold, after determining the original fault determination model as the initial fault determination model, the hard disk fault determination method applied to periodic patterns further includes: determining the sample feature items in the initial fault determination model whose first linear coefficient is greater than the first contribution threshold as the sample target feature items; determining the sample feature items in the sample feature combination in the initial fault determination model whose second linear coefficient is greater than the second contribution threshold as the sample target feature items; performing periodic training on the target FM model using the sample target feature items according to the time series to obtain the fault determination model.

[0014] According to another aspect of the embodiments of the present invention, there is also provided a hard disk fault determination device applied to periodic patterns, including: a first determination unit, configured to determine the feature items in the operation data of the target hard disk whose contribution degree is greater than the contribution threshold as the target feature items, where the target hard disk is the hard disk to be subjected to fault detection, and the contribution degree refers to the importance of the feature items for fault detection of the target hard disk; a first acquisition unit, configured to input the target feature items into the fault determination model to process the target feature items using the fault determination model to obtain the fault probability that the target hard disk has a fault, where the fault determination model is a cascade model obtained by performing periodic training on multiple sets of first training data in a machine learning manner, the multiple sets of first training data are time series data within a preset time period, and each set of the multiple sets of first training data includes: sample target feature items and the first sample fault probability corresponding to the sample target feature items; a second determination unit, configured to determine that the target hard disk has the fault when the fault probability is greater than the preset probability value, or determine that the target hard disk does not have the fault when the fault probability is not greater than the preset probability value.

[0015] Optionally, the first determination unit includes: a first acquisition module, configured to acquire the current operation data of the target hard disk; a second acquisition module, configured to parse the operation data to obtain multiple feature items of the operation data and the field name corresponding to each feature item; a first determination module, configured to determine, according to the field name, that the first linear coefficient of the sample feature item in the initial fault determination model is the single contribution degree of the feature item, where the initial fault determination model is obtained by training through machine learning using multiple groups of second training data, and each group of the multiple groups of second training data includes: the sample feature item and the second sample fault probability corresponding to the sample feature item, the sample feature item has the same field name as the feature item, and the single contribution degree refers to the importance of a single feature item for fault detection of the target hard disk; a second determination module, configured to determine that the feature item with the single contribution degree greater than the first contribution degree threshold is the target feature item.

[0016] Optionally, the first determination unit includes: a third determination module, configured to determine a combination of any two of the multiple feature items as a feature combination; a fourth determination module, configured to determine, according to the field names of the two feature items in the feature combination, that the second linear coefficient of the sample feature combination in the initial fault determination model is the combined contribution degree of the feature combination, where the two sample feature items in the sample feature combination have the same field names as the two feature items in the feature combination, and the combined contribution degree refers to the importance of combining the two feature items for fault detection of the target hard disk; a fifth determination module, configured to determine that the feature combination with the combined contribution degree greater than the second contribution degree threshold is the target feature combination, where the second contribution degree threshold is a contribution degree threshold that is the same as or different from the first contribution degree threshold; a sixth determination module, configured to determine that both of the two feature items in the target feature combination are the target feature items.

[0017] Optionally, the hard disk failure determination device applied to the periodic rule further includes: a second acquisition unit, configured to obtain multiple pieces of labeled historical operation data of the target hard disk within a preset historical time period according to a time series before determining the first linear coefficient of the sample feature item in the initial failure determination model as the contribution degree of the feature item according to the field name, or determining the combined contribution degree of the sample feature combination in the initial failure determination model according to the field names of two feature items in the feature combination, where the labeled historical operation data refers to the historical operation data carrying the historical failure probability label; a third acquisition unit, configured to divide the multiple pieces of labeled historical operation data to obtain a labeled training data set and a labeled verification data set, where the labeled training data set is used to train and obtain the initial failure determination model, and the labeled verification data set is used to verify the performance of the initial failure determination model; a third determination unit, configured to iteratively train the machine learning model by using the labeled training data set until the iterative termination condition is reached, and determine the currently trained machine learning model as the original failure determination model, where the iterative termination condition includes at least one of the following: the number of iterations reaches a preset number of iterations, and the prediction error is not greater than an error threshold; a fourth acquisition unit, configured to verify the performance of the original failure determination model by using the labeled verification data set to obtain a verification result; a fourth determination unit, configured to determine the original failure determination model as the initial failure determination model when the verification result indicates that the performance score of the original failure determination model is higher than a score threshold.

[0018] Optionally, the hard disk fault determination device applied to periodic rules further includes: a culling unit configured to cull the incomplete tag historical operation data among the multiple tag historical operation data of the target hard disk within a preset historical time period according to a time series, where the incomplete tag historical operation data refers to the tag historical operation data with missing values; a conversion unit configured to perform one-hot encoding processing on the discrete feature items among the multiple tag historical operation data to convert the data verification of the discrete feature items into a preset target format, where the discrete feature items refer to the feature items with a finite number or a countable number of discontinuous values; a partitioning unit configured to partition the multiple tag historical operation data according to sample types to obtain multiple sample data sets, where the sample types include: fault samples and non-fault samples, and each sample data set includes multiple tag historical operation data of the same sample type; an expansion unit configured to use an oversampling algorithm to expand the sample data set with the total number of data less than a preset quantity to obtain a target sample data set, where the total number of data refers to the number of tag historical operation data in the sample data set; and a processing unit configured to perform normalization processing on the tag historical operation data in the sample data set and the target sample data set to complete the preprocessing operation of the multiple tag historical operation data.

[0019] Optionally, the third determination unit includes: a third acquisition module configured to perform feature reconstruction on the multiple tag historical operation data by using an XGBoost model to obtain multiple reconstructed feature items, where the XGBoost model is one of the machine learning models; a seventh determination module configured to perform iterative training on an FM model by using the multiple feature reconstruction items until the iterative termination condition is reached, and determine the currently trained FM model as the target FM model, where the FM model is a machine learning model different from the XGBoost model; and a fourth acquisition module configured to fuse the XGBoost model and the target FM model to obtain the original fault determination model.

[0020] Optionally, the hard disk fault determination device applied to periodic rules further includes: an eighth determination module, configured to, when the verification result indicates that the performance score of the original fault determination model is higher than the score threshold, after determining the original fault determination model as the initial fault determination model, determine the sample feature items in the initial fault determination model whose first linear coefficient is greater than the first contribution threshold as the sample target feature items; a ninth determination module, configured to determine the sample feature items in the sample feature combination in the initial fault determination model whose second linear coefficient is greater than the second contribution threshold as the sample target feature items; a fifth acquisition module, configured to perform periodic training on the target FM model by using the sample target feature items according to the time series to obtain the fault determination model.

[0021] According to another aspect of the embodiments of the present invention, there is also provided a hard disk fault determination system applied to periodic rules, and the hard disk fault determination system applied to periodic rules uses any one of the above-mentioned hard disk fault determination methods applied to periodic rules.

[0022] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, and the computer-readable storage medium includes a stored program, wherein the program executes any one of the above-mentioned hard disk fault determination methods applied to periodic rules.

[0023] According to another aspect of the embodiments of the present invention, there is also provided a processor, and the processor is used to run a program, wherein when the program runs, it executes any one of the above-mentioned hard disk fault determination methods applied to periodic rules.

[0024] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including computer instructions, and when the computer instructions are executed by a processor, they execute any one of the above-mentioned hard disk fault determination methods applied to periodic rules.

[0025] In an embodiment of the present invention, it is possible to determine that a feature item in the operation data of a target hard disk with a contribution degree greater than a contribution degree threshold is a target feature item, where the target hard disk is a hard disk that needs to be subjected to fault detection, and the contribution degree refers to the importance of the feature item for fault detection of the target hard disk; input the target feature item into a fault determination model to use the fault determination model to process the target feature item to obtain a fault probability that the target hard disk has a fault, where the fault determination model is a cascaded model obtained by performing periodic training on multiple sets of first training data through machine learning, and the multiple sets of first training data are time series data within a preset time period, and each set of the multiple sets of first training data includes: a sample target feature item and a first sample fault probability corresponding to the sample target feature item; in the case where the fault probability is greater than a preset probability value, it is determined that the target hard disk has a fault, or, in the case where the fault probability is not greater than the preset probability value, it is determined that the target hard disk does not have a fault. Through the above technical solution, the purpose of screening out features with a relatively high contribution degree to hard disk fault detection and then using the corresponding fault determination model to process the screened features to determine whether the hard disk has a fault is achieved, realizing the technical effect of selectively and effectively combining the model to accurately detect the corresponding fault by specifically selecting representative features, improving the accuracy and efficiency of hard disk fault detection, and further solving the technical problem in the related art that the features selected by the traditional feature selection method when determining hard disk faults are not representative and cannot effectively combine the model to accurately detect hard disk faults. Description of the Drawings

[0026] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0027] Figure 1 is a hardware structure block diagram of a mobile terminal for a method for determining hard disk faults applied to periodic rules according to an embodiment of the present invention;

[0028] Figure 2 is a flowchart of a method for determining hard disk faults applied to periodic rules according to an embodiment of the present invention;

[0029] Figure 3 is a flowchart of model training and feature selection according to an embodiment of the present invention;

[0030] Figure 4 is a flowchart of data preprocessing according to an embodiment of the present invention;

[0031] Figure 5 is a schematic diagram of expanding samples according to an embodiment of the present invention;

[0032] Figure 6Schematic diagram of an optional augmented sample according to an embodiment of the present invention;

[0033] Figure 7 Schematic diagram of an FM model according to an embodiment of the present invention;

[0034] Figure 8 Schematic diagram of a cascade model according to an embodiment of the present invention;

[0035] Figure 9 Schematic diagram of a hard disk failure determination device applied to periodic rules according to an embodiment of the present invention. Detailed implementation manners

[0036] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0037] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0038] The following explains some nouns or terms involved in the embodiments of the present invention:

[0039] S.M.A.R.T. (Self-Monitoring, Analysis, and Reporting Technology) is a self-monitoring, analysis, and reporting technology for hard disk drives; it aims to help users and system administrators determine hard disk failures and take appropriate actions to prevent data loss and system interruptions.

[0040] As introduced in the background art, the features selected by the traditional feature selection method in the related art when determining hard disk failures are not representative and cannot effectively combine with the model to accurately detect hard disk failures. To address the above deficiencies, in the embodiments of the present invention, a method and device for determining hard disk failures applicable to periodic patterns are provided.

[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0042] The method embodiments provided in the embodiments of the present invention can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking the operation on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal of a method for determining hard disk failures applicable to periodic patterns according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in Figure 1 processors 102 (the processors 102 may include, but are not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown, or have a different configuration from

[0043] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the hard disk fault determination method applied to periodic patterns in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, that is, the above-mentioned method is implemented. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include the wireless network provided by the communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0044] According to an embodiment of the present invention, a method embodiment of a hard disk fault determination method applied to periodic patterns is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0045] Figure 2 is a flowchart of a hard disk fault determination method applied to periodic patterns according to an embodiment of the present invention, as Figure 2 shown, the method includes the following steps:

[0046] Step S202, determine that the feature items in the running data of the target hard disk with a contribution degree greater than the contribution degree threshold are target feature items, where the target hard disk is the hard disk that needs to be fault-detected, and the contribution degree refers to the importance of the feature item for fault-detecting the target hard disk.

[0047] In this embodiment, first, the current running data of the target hard disk (hereinafter simply referred to as the hard disk) can be obtained, and the contribution degrees of the feature items in the running data are analyzed, so as to select representative feature items to evaluate whether the hard disk has faults, thereby avoiding interference brought by other feature items.

[0048] According to the above embodiments of the present invention, in the above step S202, the contribution degree includes a single contribution degree, and the contribution degree threshold includes a first contribution degree threshold. Determining a feature item in the operation data of the target hard disk whose contribution degree is greater than the contribution degree threshold as a target feature item includes: obtaining the current operation data of the target hard disk; parsing the operation data to obtain multiple feature items of the operation data and the field name corresponding to each feature item; determining the first linear coefficient of the sample feature item in the initial fault determination model as the single contribution degree of the feature item according to the field name, where the initial fault determination model is trained by machine learning using multiple sets of second training data, and each set of the multiple sets of second training data includes: a sample feature item and a second sample fault probability corresponding to the sample feature item, the sample feature item has the same field name as the feature item, and the single contribution degree refers to the importance of a single feature item for fault detection of the target hard disk; determining a feature item whose single contribution degree is greater than the first contribution degree threshold as a target feature item.

[0049] On the one hand, according to the single contribution degree of each feature item in the operation data, if the single contribution degree is greater than the first contribution degree threshold, it can be considered that the feature item corresponding to the single contribution degree is representative and can be used as a target feature item as the data basis for subsequent judgment of whether the hard disk has a fault.

[0050] Specifically, the field name corresponding to each feature item can be determined first, then the first linear coefficient of each sample feature item in the initial fault determination model is obtained, and then matching is performed according to the field name. If the field name of a certain feature item is the same as that of a certain sample feature item, the first linear coefficient of the sample feature item can be used as the single contribution degree of the feature item; the initial fault determination model here is trained by machine learning using multiple sets of sample training data, and the training process of this model will be described in detail below and will not be elaborated here.

[0051] According to the above embodiments of the present invention, in the above step S202, the contribution degree includes a combined contribution degree, and the contribution degree threshold includes a second contribution degree threshold. Determining a feature item in the operation data of the target hard disk whose contribution degree is greater than the contribution degree threshold as a target feature item includes: determining a combination of any two feature items among multiple feature items as a feature combination; determining the second linear coefficient of the sample feature combination in the initial fault determination model as the combined contribution degree of the feature combination according to the field names of the two feature items in the feature combination, where the two sample feature items in the sample feature combination have the same field names as the two feature items in the feature combination, and the combined contribution degree refers to the importance of combining two feature items for fault detection of the target hard disk; determining a feature combination whose combined contribution degree is greater than the second contribution degree threshold as a target feature combination, where the second contribution degree threshold is a contribution degree threshold that is the same as or different from the first contribution degree threshold; determining that both feature items in the target feature combination are target feature items.

[0052] On the other hand, some feature items may not be significant when considered individually, but may have an important impact on determining whether there is a hard disk failure when interacting with other features. Therefore, multiple feature items of the running data can be arbitrarily combined in pairs to obtain multiple feature combinations. Then, according to the field names of the two feature items in the feature combination, they are matched with the field names of the two sample feature items in each sample feature combination in the initial fault determination model. If the match is successful (that is, the field names of the two sample feature items in a certain sample feature combination correspond to the field names of the two feature items in a certain feature combination), the second linear coefficient of the sample feature combination in the initial fault determination model can be used as the combination contribution degree of the feature combination. If the combination contribution degree is greater than the second contribution degree threshold, it can be considered that the two feature items in the corresponding feature combination have a relatively important impact on determining whether there is a hard disk failure when combined, and it can also be considered as representative feature items, that is, they can be used as target feature items.

[0053] It should be noted that the above first contribution degree threshold is mainly for the importance evaluation of a single feature, while the above second contribution degree is mainly for the contribution degree evaluation of a feature combination. Therefore, their value sizes are not directly related to each other, but are set according to the model training results and specific application requirements respectively. However, from the perspective of logic and model optimization, the setting of these two thresholds can follow certain principles to achieve the best feature screening effect, and can be selected according to the actual situation, and no specific restrictions are made here.

[0054] In another alternative embodiment of the present invention, the hard disk failure determination method applied to periodic rules further includes: before determining the first linear coefficient of the sample feature item in the initial failure determination model as the contribution degree of the feature item according to the field name, or determining the combined contribution degree of the sample feature combination in the initial failure determination model according to the field names of two feature items in the feature combination, obtaining multiple pieces of labeled historical operation data of the target hard disk within a preset historical time period according to the time series, where the labeled historical operation data refers to the historical operation data carrying the historical failure probability label; dividing the multiple pieces of labeled historical operation data to obtain a labeled training data set and a labeled verification data set, where the labeled training data set is used to train and obtain the initial failure determination model, and the labeled verification data set is used to verify the performance of the initial failure determination model; iteratively training the machine learning model using the labeled training data set until the iteration termination condition is reached, and determining the currently trained machine learning model as the original failure determination model, where the iteration termination condition includes at least one of the following: the number of iterations reaches the preset number of iterations, and the prediction error is not greater than the error threshold; verifying the performance of the original failure determination model using the labeled verification data set to obtain a verification result; and determining the original failure determination model as the initial failure determination model when the verification result indicates that the performance score of the original failure determination model is higher than the score threshold.

[0055] The following combines Figure 3 to elaborate in detail on the training processes of the initial failure determination model and the failure determination model in the above embodiments of the present invention. Figure 3 is a flowchart of model training and feature selection according to an embodiment of the present invention.

[0056] As Figure 3 shown, multiple pieces of labeled historical operation data (SMART data, that is, labeled historical operation data) within a certain historical time period can be obtained according to the time series as sample data to iteratively train the machine learning model, and an initial failure determination model is obtained. The initial failure determination model is a cascaded model composed of an XGBoost model and an FM model.

[0057] The SMART data here specifically refers to the numerical values of these metrics recorded by the hard disk at different time points. These data are time series data, which means they are arranged in chronological order, and the data at each time point reflects the operating condition of the hard disk at that time point. By analyzing these historical operation data, the changing rules of the hard disk health status can be found, including periodic changes, as well as the decrease in hard disk performance or the increase in the risk of failure under specific conditions.

[0058] Specifically, in the process of selecting feature items, a feature contribution calculation model (i.e., the initial fault determination model in the embodiment of the present invention) can be established based on a mathematical analytical method. The model can generate a physically interpretable feature importance ranking by quantitatively analyzing the interaction strength between first-order linear features and high-order combined features. The ranking result is fed back to the model training system to guide feature engineering optimization and training data enhancement, and a closed-loop learning system of "feature screening-model optimization" is formed through iterative training. The feature importance ranking result can be directly applied to the formulation of hard disk operation and maintenance strategies, and more representative feature items can be selected by focusing on monitoring high-contribution feature indicators to achieve active control of failure rate.

[0059] In a specific embodiment of the present invention, after obtaining multiple label historical operation data of the target hard disk within a preset historical time period according to a time series, the hard disk fault determination method applied to periodic rules also includes: eliminating incomplete label historical operation data from the multiple label historical operation data, wherein the incomplete label historical operation data refers to label historical operation data with missing values; performing unique hot encoding processing on discrete feature items in the multiple label historical operation data to verify and convert the data of the discrete feature items into a preset target format, wherein the discrete feature items refer to feature items with values of a finite or countable number of discontinuous features; dividing the multiple label historical operation data according to sample types to obtain multiple sample data sets, wherein the sample types include: fault samples and non-fault samples, and each sample data set includes multiple label historical operation data of the same sample type; using an upsampling algorithm to expand the sample data set whose total number of data items is less than a preset number to obtain a target sample data set, wherein the total number of data items refers to the number of label historical operation data in the sample data set; normalizing the label historical operation data in the sample data set and the target sample data set to complete the preprocessing operation of the multiple label historical operation data.

[0060] Combine the following Figure 4 The data preprocessing process in the above embodiment of the present invention is described in detail. Figure 4 is a flow chart of data preprocessing according to an embodiment of the present invention. Figure 4 As shown in the figure, data preprocessing specifically includes the following steps: 1) Eliminate samples containing missing data values; 2) Perform one-hot encoding on discrete features, including the working mode, capacity, and model of the hard disk; 3) For each type of minority class sample after one-hot encoding, use an upsampling algorithm based on the minority class support vector to expand the sample; 4) After expanding the sample, normalize the continuous features of all types of samples uniformly, including 255 continuous attributes.

[0061] Generally speaking, in sample data, there may be a situation where the number of faulty samples (minority class) is much less than that of normal samples (majority class). To avoid the imbalance in the types of model training data, the minority class samples can be augmented. The following combines Figure 5 and Figure 6 to elaborate in detail on the process of sample augmentation in the above embodiments of the present invention. Figure 5 is a schematic diagram of augmented samples according to an embodiment of the present invention. Figure 6 is a schematic diagram of optional augmented samples according to an embodiment of the present invention.

[0062] As Figure 5 and Figure 6 shown, based on the minority class support vectors, the embodiments of the present invention propose two different upsampling strategies: interpolation and extrapolation. This algorithm is divided into two stages: finding support vectors and augmenting minority class samples. Specifically as follows: 1) Use the SVM algorithm to find the support vectors of the minority class; 2) Calculate the k-nearest neighbors of each minority class support vector using the Euclidean metric. Assume there is a minority class support vector: if all its k-nearest neighbors are majority class, it is marked as a noise sample (noise), if more than half of its k-nearest neighbors are majority class, it is marked as a dangerous sample (danger), and if less than half of its k-nearest neighbors are minority class, it is marked as a safe sample (safety); 3) For each dangerous sample x i , find its nearest neighbor of the same class sample x j , and use the sample interpolation method to insert new samples between them: x new =x i +rand(0, 1)*{x j -x i )). For each safe sample x i , use the sample extrapolation method to generate new samples on the extension line of the two samples: x new =x i -rand(0, 1)*(x j -x i ), where rand(0, 1) is a random number between 0 and 1. This method expands the minority class samples to the sample space where the density of the majority class is not high, thus solving the problem of insufficient number of faulty hard disk samples for subsequent fault prediction.

[0063] In another specific embodiment of the present invention, the machine learning model is iteratively trained using a labeled training dataset until the iteration termination condition is reached, and the machine learning model obtained by the current training is determined as the original fault determination model, including: using the XGBoost model to perform feature reconstruction on multiple pieces of labeled historical operation data to obtain multiple reconstructed feature items, where the XGBoost model is a model in the machine learning model; using the multiple feature reconstruction items to iteratively train the FM model until the iteration termination condition is reached, and determining the FM model obtained by the current training as the target FM model, where the FM model is a model different from the XGBoost model in the machine learning model; fusing the XGBoost model and the target FM model to obtain the original fault determination model.

[0064] Specifically, the initial fault determination model provided in the embodiment of the present invention is obtained by cascading and combining the XGBoost model and the FM model. Among them, the XGBoost model is used for feature reconstruction, that is, by constructing multiple decision trees, the original SMART data features are converted into more abstract and highly predictive new features. The XGBoost model can automatically learn the complex relationships between features and combine multiple weak learners (decision trees) into a strong learner through an additive model, improving the accuracy and robustness of prediction; the FM model is used to perform feature combination and classification based on the reconstructed features generated by XGBoost. By considering the second-order interaction between features, the FM model can more comprehensively evaluate and utilize the information of feature combinations and capture complex patterns that may contribute to hard disk fault prediction. The introduction of the FM model enables the model to not only utilize the information of individual features but also utilize the interaction between features, improving the prediction accuracy and interpretability of the model.

[0065] The following combines Figure 7 and Figure 8 to detail the construction process of the initial fault determination model (a cascaded model) in the above embodiment of the present invention. Figure 7 is a schematic diagram of the FM model according to the embodiment of the present invention. Figure 8 is a schematic diagram of the cascaded model according to the embodiment of the present invention. The specific construction process is as follows:

[0066] (1) XGBoost Reconstructed Features:

[0067] The process of XGBoost ensemble learning is closely related to the gradient direction of the optimization objective, so it is called gradient boosting. Such methods are all through an additive model: There is no coefficient in front of the basic learner f k here. Among them: For sample x i The corresponding predicted value; K is the number of base learners; F is the function space composed of all base learners: F = {f1, f2,..., f k} (2). Therefore, the XGBoost algorithm is actually a combined optimization problem of multiple functions. For this reason, the algorithm optimization objective is defined Among them: is the sample loss function, used to reduce the model bias; Ω(f k ) is the regularization term of the base learner, preventing overfitting and reducing the model variance; XGBoost learns each decision tree in turn, and all samples in the dataset will fall on the leaf nodes of the decision tree. Therefore, the dataset Each decision tree of XGBoost maps the sample x i to a corresponding leaf node, obtaining the node index value. Therefore, the mapping is defined as: q: x D → {1, 2,..., T} (4), where: D is the dimension of the input sample x; T is the number of leaf nodes of the decision tree; According to the above mapping, the decision tree can be expressed as: f(x) = ω q(x) (5), where: ω ∈ R T , representing the weight vector corresponding to each leaf node of the decision tree; The regularization term in the objective function formula (3) used to control the complexity of the decision tree can be expressed as: Among them: γ is the hyperparameter that controls the number of leaf nodes T, and λ is the hyperparameter that controls the L2 norm of the node weight vector; Similar to the ordinary decision tree, XGBoost also determines the optimal splitting point by calculating the splitting gain on each feature.

[0068] (2) FM trains the classifier:

[0069] As Figure 7 shown, the Factorization Machine (FM model) is an idea for solving multi-dimensional sparse scenarios. On the basis of logistic regression, the FM model considers the interaction between features and proposes feature latent vectors to estimate model parameters. After introducing the second-order polynomial, the LR model can be rewritten as: Among them: N represents the feature dimension; w0 ∈ R; w = {w1, w2,..., w n} ∈ R N ; W ∈ R N×N ; It can be seen from formula (7) that the coefficients w ij of the feature combination terms are all independent, and the number of combination terms is parameters, introducing the idea of matrix factorization in collaborative filtering. As known from linear algebra, when k satisfies certain conditions, for any positive definite real symmetric matrix W, there exists a real matrix V ∈ R K×N , such that W = V T V holds. Since we only care about the relationships between distinct features, we can assume that the diagonal elements of matrix W are large enough, so W is a strictly diagonally dominant matrix, and thus W is positive definite. Therefore, the parameters of the combined terms can be expressed as the inner product of the corresponding latent vectors, as shown below: The second-order polynomial model (7) is rewritten as the FM model: where: v i represents the latent vector of the feature component x i , with a length of K (K ∈ N + , K << N); <·> represents the inner product.

[0070] At this time, the parameters of the combined terms are reduced to KN. Moreover, all feature combination methods containing the feature component x have the opportunity to learn the latent vector v i , and this advantage enables the FM model to handle the problems of high-dimensional data sparsity and insufficient samples well. The time complexity of Equation (9) is: According to the method of completing the square, Equation (9) is transformed as follows: At this time, it can be rewritten as: Therefore, the time complexity of FM after the transformation by the method of completing the square is: O(KN) = K{[N + (N - 1) + 1] + [3N + (N - 1) + 1]} + (K - 1) + 1 (13); It can be seen that the FM model can complete the target task in linear time. For classification tasks, the logit function is generally selected as the loss function:

[0071] (3) Construction of the cascade model:

[0072] As Figure 8 shown, assume that the feature set involved in the data set is C = {c1, c2,... c N}, and N represents the number of features. Therefore: x i = {x1, x2,... x N}, x i ∈ R N(15), and since the work completed by XGBoost is to map each piece of data to the leaf nodes of each subtree, obtaining the corresponding index vector for this data: XGBoost: x i → ω i (16), where: ω i ={ω i1 , ω i2 ,… ω iT}, T is the number of decision trees, and the above formula is the key to feature reconstruction.

[0073] For a trained XGBoost model, assuming that the leaf nodes of the k-th tree are encoded in natural numbers from left to right, denoted as the set: L = {1, 2,… l k}, k ∈ T (17), where: l k is the number of leaf nodes of the current subtree; XGBoost maps the input data x i to the node index vector ω i , that is, the element ω i in ω ik is the encoding of the data x i mapped to the node where the k-th subtree is located, and the vector ω i is regarded as the feature vector reconstructed from the original data by XGBoost; the feature vector ω i reconstructed by XGBoost is the implicit information of the original data x i , and one-hot encoding is performed on it, encoded as At this time the elements in are no longer the encoding values of the leaf nodes, but a sparse vector with a length of l k and only one subscript value is 1, and the other positions are all 0. All vectors are combined into a feature vector Therefore, only the elements in the positions corresponding to the number of decision trees in this vector are 1.

[0074] Therefore, the new vector reconstructed by XGBoost has a dimension equal to the sum of the decision tree leaf nodes: Although N′ >> N, the reconstructed high-dimensional sparse vector does not increase the difficulty of training the FM model, but instead makes the feature combination part of the model easier to perform.

[0075] XGBoost has the function of feature reconstruction, and FM has the function of feature crossing. The embodiments of the present invention will fuse these two models to solve practical problems, and the dimension of the newly generated features is determined by all the decision trees in XGBoost.

[0076] In yet another specific embodiment of the present invention, when the verification result indicates that the performance score of the original fault determination model is higher than the score threshold, after determining the original fault determination model as the initial fault determination model, the hard disk fault determination method applied to the periodic law further includes: determining the sample feature items with the first linear coefficient greater than the first contribution degree threshold in the initial fault determination model as the sample target feature items; determining the sample feature items in the sample feature combination with the second linear coefficient greater than the second contribution degree threshold in the initial fault determination model as the sample target feature items; and performing periodic training on the target FM model using the sample target feature items according to the time series to obtain the fault determination model.

[0077] Specifically, when predicting hard disk faults, the selection of SMART attribute values is crucial for the prediction hit rate. Traditional feature screening methods do not consider the impact of the correlation information between attributes on the accuracy of hard disk fault prediction. In the embodiments of the present invention, the second-order feature items of the FM model provide a new idea for screening the optimal combined features. On the one hand, assuming that the new samples reconstructed by XGBoost are known The dimension is N′, and the second-order FM model of the reconstructed samples is:

[0078] Where: Is the jth feature of the ith reconstructed sample Its value is either 0 or 1, and λ j Is the linear term coefficient of the model; for all linear features, a factor threshold ε (i.e., the first contribution degree threshold) is set, and the features that meet λ j > ε are all selected. Therefore, it is reasonable to believe that the features with a contribution degree greater than ε may be more worthy of reference. Assuming that the set of coefficients of the selected linear features is λ1 = {λ j |λ j > ε}, the corresponding feature set is

[0079] On the other hand, for the second-order polynomial part of the FM model, it is defined as: λ jk =<v j , v k >(20), λ jk Is a second-order term coefficient of the model, which represents And The contribution of these two feature combinations to FM. A factor threshold ζ (i.e., the second contribution degree threshold) is set, and the feature combinations that meet λ jk > ζ And Are all selected; assuming that the finally screened coefficient set that meets the conditions is λ = {λ jk|λ jk > ζ}, for the set element λ jk , there must exist and which are the reconstruction features of two subtrees respectively. Let the original subsets used by XGBoost to build these two subtrees be C j and C k , then the optimal feature combination can be expressed as C j ∪C k . Taking the intersection of all feature combinations that meet the above threshold conditions, we can get: Therefore is the set of feature screening results proposed according to the embodiments of the present invention (i.e., the set of sample target feature items).

[0080] After screening out representative sample target feature items, considering that hard disk failure prediction belongs to an event with periodic change rules but not large change ranges, in the embodiments of the present invention, one year's data can be used to train the XGBoost model to generate reconstructed features; then one quarter's data can be used to train the FM classifier. For the trained model, in the next quarter, the classifier can be directly fine-tuned based on the existing parameters of the FM model to accelerate model convergence, so as to update and optimize the initial fault determination model, and then the fault determination model can be obtained; for events with large change ranges between each period, one quarter's data can be used to train the XGBoost model and one month's data can be used to fine-tune the FM model, which not only reduces the model training difficulty but also improves the model accuracy, and is also applicable to other scenarios with any periodic changes.

[0081] Generally speaking, in the embodiments of the present invention, a separation architecture of a fixed feature extraction module (i.e., the XGBoost model) and an adjustable classifier module (i.e., the FM model) can be used to construct the fault determination model. Among them, the feature extraction module maintains the parameter freezing state in different periodic scenarios, and a periodic adaptation model can be generated by only fine-tuning the classifier module, so as to obtain a multi-period prediction model without increasing the repeated training cost of feature extraction, and balance the contradictory relationship between the model training complexity and the prediction accuracy.

[0082] It should be noted that in the embodiments of the present invention, first, a completed sample data set can be used for the initial training of the cascade model to obtain an initial fault determination model. In this process, the model will learn the weights of all features and the contribution degrees of feature combinations according to the data. According to this contribution degree, feature selection can be further performed on the sample data set to screen out representative sample target feature items. Then, these sample target feature items can be used to perform periodic training and optimization on the FM model in the initial fault determination model in the above-mentioned periodic training manner, so as to obtain a fault prediction model.

[0083] In addition, after the model is trained, the samples in the validation set can be one-hot encoded first, then uniformly normalized within the validation set, and finally input into the cascade model to obtain a prediction result to verify whether the model performance meets the requirements. This verification process is applicable to both the initial fault determination model and the fault determination model.

[0084] Step S204: Input the target feature items into the fault determination model to process the target feature items using the fault determination model to obtain the fault probability that the target hard disk has a fault. Among them, the fault determination model is a cascade model obtained by performing periodic training on multiple sets of first training data through machine learning. The multiple sets of first training data are time series data within a preset time period. Each set of the multiple sets of first training data includes: sample target feature items and the corresponding first sample fault probabilities.

[0085] In this embodiment, the target feature items screened out in the above steps can be input into the trained fault determination model for processing to output the fault probability that the target hard disk has a fault.

[0086] Step S206: Determine that the target hard disk has a fault when the fault probability is greater than a preset probability value, or determine that the target hard disk does not have a fault when the fault probability is not greater than the preset probability value.

[0087] In this embodiment, if the fault probability is greater than the preset probability value, it can be considered that there is a fault in the target hard disk. If the fault probability is not greater than the preset probability value, it can be considered that there is no fault in the target hard disk currently.

[0088] It should be noted that this judgment process can actually be completed inside the fault determination model, that is, the model output result is directly that the target hard disk has a fault or the target hard disk does not have a fault.

[0089] As can be seen from the above, through the technical solution provided by the above embodiments of the present invention, it is possible to determine that the feature items with a contribution degree greater than the contribution degree threshold in the operation data of the target hard disk are target feature items, where the target hard disk is the hard disk that needs to be fault-detected, and the contribution degree refers to the importance of the feature item for fault-detecting the target hard disk; input the target feature items into the fault determination model to use the fault determination model to process the target feature items to obtain the fault probability that the target hard disk has a fault, where the fault determination model is obtained by periodically training using multiple groups of first training data through machine learning. Cascade model, and the multiple groups of first training data are time series data within a preset time period, and each group of the multiple groups of first training data includes: sample target feature items, and the first sample fault probability corresponding to the sample target feature items; in the case where the fault probability is greater than the preset probability value, it is determined that the target hard disk has a fault, or, in the case where the fault probability is not greater than the preset probability value, it is determined that the target hard disk does not have a fault, achieving the purpose of screening out the feature items with a relatively high contribution degree to the hard disk fault detection, and then using the corresponding fault determination model to process the screened feature items to determine whether the hard disk has a fault, realizing the technical effect of selectively and representatively selecting features and effectively combining the model to accurately detect the corresponding fault, and improving the accuracy and efficiency of detecting the hard disk fault.

[0090] Therefore, through the technical solution provided by the above embodiments of the present invention, the technical problem in the related art that the features selected by using the traditional feature selection method during hard disk fault determination are not representative and cannot effectively combine the model to accurately detect the hard disk fault is solved.

[0091] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations, but those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0092] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0093] According to an embodiment of the present invention, there is also provided a device for determining a hard disk failure applied to a periodic pattern for implementing the above method for determining a hard disk failure applied to a periodic pattern. Figure 9 is a schematic diagram of a device for determining a hard disk failure applied to a periodic pattern according to an embodiment of the present invention, as Figure 9 shown. The device includes: a first determination unit 91, a first acquisition unit 93, and a second determination unit 95. The following will describe the device for determining a hard disk failure applied to a periodic pattern in detail.

[0094] The first determination unit 91 is configured to determine a feature item in the operation data of the target hard disk whose contribution degree is greater than the contribution degree threshold as the target feature item, where the target hard disk is the hard disk to be subjected to failure detection, and the contribution degree refers to the importance of the feature item for failure detection of the target hard disk.

[0095] The first acquisition unit 93 is configured to input the target feature item into the failure determination model to process the target feature item by using the failure determination model, so as to obtain the failure probability that the target hard disk has a failure, where the failure determination model is a cascade model obtained by performing periodic training on multiple sets of first training data by means of machine learning. The multiple sets of first training data are time series data within a preset time period, and each set of the multiple sets of first training data includes: a sample target feature item and a first sample failure probability corresponding to the sample target feature item.

[0096] The second determination unit 95 is configured to determine that the target hard disk has a failure when the failure probability is greater than the preset probability value, or determine that the target hard disk does not have a failure when the failure probability is not greater than the preset probability value.

[0097] It should be noted here that the above first determination unit 91, first acquisition unit 93, and second determination unit 95 correspond to steps S202 to S206 in the above embodiments. The examples and application scenarios implemented by the three units and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments.

[0098] As can be seen from the above, in the solution described in the above embodiments of the present invention, the first determination unit can be used to determine that a feature item with a contribution degree greater than the contribution degree threshold in the operation data of the target hard disk is the target feature item, where the target hard disk is the hard disk that needs to be fault-detected, and the contribution degree refers to the importance of the feature item for fault-detecting the target hard disk; then the first acquisition unit inputs the target feature item into the fault determination model to process the target feature item by using the fault determination model to obtain the fault probability that the target hard disk has a fault, where the fault determination model is a cascade model obtained by periodically training using multiple sets of first training data through machine learning, and the multiple sets of first training data are time series data within a preset time period, and each set of the multiple sets of first training data includes: a sample target feature item, and a first sample fault probability corresponding to the sample target feature item; finally, the second determination unit determines that the target hard disk has a fault when the fault probability is greater than the preset probability value, or determines that the target hard disk does not have a fault when the fault probability is not greater than the preset probability value, achieving the purpose of screening out features with a relatively high contribution degree to hard disk fault detection, and then using the corresponding fault determination model to process the screened features to determine whether the hard disk has a fault, realizing the technical effect of selectively and representatively selecting features and effectively combining the model to accurately detect corresponding faults, and improving the accuracy and efficiency of hard disk fault detection.

[0099] Therefore, through the technical solution provided by the above embodiments of the present invention, the technical problem in the related art that the features selected by using the traditional feature selection method when determining the hard disk fault are not representative and cannot effectively combine the model to accurately detect the hard disk fault is solved.

[0100] In an alternative embodiment, the first determination unit includes: a first acquisition module, configured to acquire the current operation data of the target hard disk; a second acquisition module, configured to parse the operation data to obtain multiple feature items of the operation data and the field name corresponding to each feature item; a first determination module, configured to determine the first linear coefficient of the sample feature item in the initial fault determination model as the single contribution degree of the feature item according to the field name, where the initial fault determination model is obtained by training using multiple sets of second training data through machine learning, and each set of the multiple sets of second training data includes: a sample feature item and a second sample fault probability corresponding to the sample feature item, the sample feature item has the same field name as the feature item, and the single contribution degree refers to the importance of a single feature item for fault-detecting the target hard disk; a second determination module, configured to determine that a feature item with a single contribution degree greater than the first contribution degree threshold is the target feature item.

[0101] In an alternative embodiment, the first determination unit includes: a third determination module, configured to determine a combination of any two feature items among a plurality of feature items as a feature combination; a fourth determination module, configured to determine the second linear coefficient of the sample feature combination in the initial fault determination model as the combination contribution degree of the feature combination according to the field names of the two feature items in the feature combination, where the two sample feature items in the sample feature combination have the same field names as the two feature items in the feature combination, and the combination contribution degree refers to the importance of combining the two feature items for fault detection of the target hard disk; a fifth determination module, configured to determine the feature combination with the combination contribution degree greater than the second contribution degree threshold as the target feature combination, where the second contribution degree threshold is a contribution degree threshold that is the same as or different from the first contribution degree threshold; a sixth determination module, configured to determine that both feature items in the target feature combination are target feature items.

[0102] In an alternative embodiment, the apparatus for determining a hard disk fault applicable to periodic rules further includes: a second acquisition unit, configured to acquire multiple pieces of labeled historical operation data of the target hard disk within a preset historical time period according to a time series before determining the first linear coefficient of the sample feature item in the initial fault determination model as the contribution degree of the feature item or determining the combination contribution degree of the sample feature combination in the initial fault determination model according to the field names of the two feature items in the feature combination, where the labeled historical operation data refers to historical operation data carrying historical fault probability labels; a third acquisition unit, configured to divide the multiple pieces of labeled historical operation data to obtain a labeled training data set and a labeled verification data set, where the labeled training data set is used for training to obtain the initial fault determination model, and the labeled verification data set is used for performance verification of the initial fault determination model; a third determination unit, configured to iteratively train the machine learning model using the labeled training data set until the iteration termination condition is reached, and then determine the currently trained machine learning model as the original fault determination model, where the iteration termination condition includes at least one of the following: the number of iterations reaches a preset number of iterations, and the prediction error is not greater than an error threshold; a fourth acquisition unit, configured to perform performance verification on the original fault determination model using the labeled verification data set to obtain a verification result; a fourth determination unit, configured to determine the original fault determination model as the initial fault determination model when the verification result indicates that the performance score of the original fault determination model is higher than a score threshold.

[0103] In an alternative embodiment, the hard disk failure determination device applied to periodic patterns further includes: a rejection unit configured to reject the incomplete tag historical operation data among the multiple tag historical operation data after acquiring multiple tag historical operation data of the target hard disk within a preset historical time period according to a time series, where the incomplete tag historical operation data refers to the tag historical operation data with missing values; a conversion unit configured to perform one-hot encoding processing on the discrete feature items in the multiple tag historical operation data to convert the data verification of the discrete feature items into a preset target format, where the discrete feature items refer to the feature items with a finite number or a countable number of discontinuous values; a partitioning unit configured to partition the multiple tag historical operation data according to the sample type to obtain multiple sample data sets, where the sample type includes: failure samples and non-failure samples, and each sample data set includes multiple tag historical operation data with the same sample type; an expansion unit configured to expand the sample data set with the total number of data less than a preset quantity by using an oversampling algorithm to obtain a target sample data set, where the total number of data refers to the number of tag historical operation data in the sample data set; and a processing unit configured to perform normalization processing on the tag historical operation data in the sample data set and the target sample data set to complete the preprocessing operation of the multiple tag historical operation data.

[0104] In an alternative embodiment, the third determination unit includes: a third acquisition module configured to perform feature reconstruction on the multiple tag historical operation data by using an XGBoost model to obtain multiple reconstructed feature items, where the XGBoost model is a model in machine learning models; a seventh determination module configured to perform iterative training on an FM model by using the multiple feature reconstruction items until an iteration termination condition is reached, and then determine the currently trained FM model as the target FM model, where the FM model is a model different from the XGBoost model in machine learning models; and a fourth acquisition module configured to fuse the XGBoost model and the target FM model to obtain an original failure determination model.

[0105] In an alternative embodiment, the hard disk failure determination device applied to periodic patterns further includes: an eighth determination module configured to, when the verification result indicates that the performance score of the original failure determination model is higher than a score threshold, determine the sample feature items with the first linear coefficient greater than the first contribution threshold in the original failure determination model as sample target feature items after determining the original failure determination model as an initial failure determination model; a ninth determination module configured to determine the sample feature items in the sample feature combination with the second linear coefficient greater than the second contribution threshold in the initial failure determination model as sample target feature items; and a fifth acquisition module configured to perform periodic training on the target FM model by using the sample target feature items according to a time series to obtain a failure determination model.

[0106] According to another aspect of the embodiments of the present invention, there is also provided a hard disk fault determination system applied to periodic rules. The hard disk fault determination system applied to periodic rules uses any one of the above-mentioned hard disk fault determination methods applied to periodic rules.

[0107] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium. The computer-readable storage medium includes a stored program. Among them, the program executes any one of the above-mentioned hard disk fault determination methods applied to periodic rules.

[0108] Optionally, in this embodiment, the above-mentioned computer-readable storage medium may be located in any one of the computer terminals in the computer terminal group in the computer network, or located in any one of the communication devices in the communication device group.

[0109] Optionally, in this embodiment, the computer-readable storage medium is set to store program codes for executing the following steps: determining a feature item with a contribution degree greater than the contribution degree threshold in the operation data of the target hard disk as the target feature item, where the target hard disk is the hard disk that needs to be fault-detected, and the contribution degree refers to the importance of the feature item for fault-detecting the target hard disk; inputting the target feature item into the fault determination model to process the target feature item by using the fault determination model to obtain the fault probability that the target hard disk has a fault, where the fault determination model is a cascaded model obtained by performing periodic training on multiple groups of first training data through machine learning, and the multiple groups of first training data are time series data within a preset time period, and each group of the multiple groups of first training data includes: a sample target feature item, and a first sample fault probability corresponding to the sample target feature item; in the case where the fault probability is greater than the preset probability value, determining that the target hard disk has a fault, or, in the case where the fault probability is not greater than the preset probability value, determining that the target hard disk does not have a fault.

[0110] Optionally, in this embodiment, the computer-readable storage medium is set to store program codes for executing the following steps: obtaining the current operation data of the target hard disk; parsing the operation data to obtain multiple feature items of the operation data and the field name corresponding to each feature item; determining the first linear coefficient of the sample feature item in the initial fault determination model as the single contribution degree of the feature item according to the field name, where the initial fault determination model is obtained by performing training on multiple groups of second training data through machine learning, and each group of the multiple groups of second training data includes: a sample feature item, and a second sample fault probability corresponding to the sample feature item, the sample feature item has the same field name as the feature item, and the single contribution degree refers to the importance of a single feature item for fault-detecting the target hard disk; determining the feature item with a single contribution degree greater than the first contribution degree threshold as the target feature item.

[0111] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining a combination of any two feature items among a plurality of feature items as a feature combination; determining the second linear coefficient of the sample feature combination in the initial fault determination model as the combination contribution degree of the feature combination according to the field names of the two feature items in the feature combination, where the two sample feature items in the sample feature combination have the same field names as the two feature items in the feature combination, and the combination contribution degree refers to the importance of combining the two feature items for fault detection of the target hard disk; determining the feature combination with the combination contribution degree greater than the second contribution degree threshold as the target feature combination, where the second contribution degree threshold is a contribution degree threshold that is the same as or different from the first contribution degree threshold; and determining that the two feature items in the target feature combination are both target feature items.

[0112] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: before determining the first linear coefficient of the sample feature item in the initial fault determination model as the contribution degree of the feature item according to the field name, or determining the combination contribution degree of the sample feature combination in the initial fault determination model according to the field names of the two feature items in the feature combination, obtaining multiple pieces of labeled historical operation data of the target hard disk in a preset historical time period according to the time series, where the labeled historical operation data refers to the historical operation data carrying the historical fault probability label; dividing the multiple pieces of labeled historical operation data to obtain a labeled training data set and a labeled validation data set, where the labeled training data set is used for training to obtain the initial fault determination model, and the labeled validation data set is used for performance verification of the initial fault determination model; iteratively training the machine learning model using the labeled training data set until the iteration termination condition is reached, and determining the currently trained machine learning model as the original fault determination model, where the iteration termination condition includes at least one of the following: the number of iterations reaches the preset number of iterations, and the prediction error is not greater than the error threshold; performing performance verification on the original fault determination model using the labeled validation data set to obtain a verification result; and in the case where the verification result indicates that the performance score of the original fault determination model is higher than the score threshold, determining the original fault determination model as the initial fault determination model.

[0113] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: removing the incomplete tag historical operation data from the multiple tag historical operation data, where the incomplete tag historical operation data refers to the tag historical operation data with missing values; performing one-hot encoding processing on the discrete feature items in the multiple tag historical operation data to convert the data verification of the discrete feature items into a preset target format, where the discrete feature items refer to the feature items with a finite or countable number of discontinuous values; dividing the multiple tag historical operation data according to the sample type to obtain multiple sample data sets, where the sample types include: fault samples, non-fault samples, and each sample data set includes multiple tag historical operation data with the same sample type; using an oversampling algorithm to expand the sample data set with the total number of data less than the preset number to obtain a target sample data set, where the total number of data refers to the number of tag historical operation data in the sample data set; performing normalization processing on the tag historical operation data in the sample data set and the target sample data set to complete the preprocessing operation of the multiple tag historical operation data.

[0114] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: using the XGBoost model to perform feature reconstruction on the multiple tag historical operation data to obtain multiple reconstructed feature items, where the XGBoost model is a model in machine learning models; using the multiple feature reconstruction items to perform iterative training on the FM model until the iterative termination condition is reached, and determining the currently trained FM model as the target FM model, where the FM model is a model different from the XGBoost model in machine learning models; fusing the XGBoost model and the target FM model to obtain an original fault determination model.

[0115] Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for performing the following steps: determining the sample feature items with the first linear coefficient greater than the first contribution threshold in the initial fault determination model as the sample target feature items; determining the sample feature items in the sample feature combination with the second linear coefficient greater than the second contribution threshold in the initial fault determination model as the sample target feature items; performing periodic training on the target FM model using the sample target feature items according to the time series to obtain a fault determination model.

[0116] According to another aspect of the embodiments of the present invention, a processor is further provided, and the processor is used to run a program, where when the program runs, it executes the above-mentioned hard disk fault determination method applied to periodic rules.

[0117] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including computer instructions which, when executed by a processor, perform the method for determining hard disk failures applied to periodic rules as described in any one of the above.

[0118] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0119] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0120] In the several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of the units or modules can be in electrical or other forms.

[0121] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0122] In addition, the functional units in the various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0123] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0124] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for determining hard disk failures applicable to periodic patterns, characterized in that, Including: Determine that a feature item with a contribution degree greater than a contribution degree threshold in the operation data of the target hard disk is a target feature item, where the target hard disk is a hard disk that needs to be subjected to fault detection, and the contribution degree refers to the importance of the feature item for fault detection of the target hard disk; Input the target feature item into a fault determination model to process the target feature item by using the fault determination model to obtain a fault probability that the target hard disk has a fault, where the fault determination model is a cascaded model obtained by performing periodic training on multiple sets of first training data through machine learning, the multiple sets of first training data are time series data within a preset time period, and each set of the multiple sets of first training data includes: a sample target feature item, and a first sample fault probability corresponding to the sample target feature item; In the case where the fault probability is greater than a preset probability value, determine that the target hard disk has the fault, or, in the case where the fault probability is not greater than the preset probability value, determine that the target hard disk does not have the fault.

2. The hard disk failure determination method applied to periodic rules according to claim 1, wherein The contribution degree includes a single contribution degree, and the contribution degree threshold includes a first contribution degree threshold. Determining that a feature item with a contribution degree greater than a contribution degree threshold in the operation data of the target hard disk is a target feature item includes: Obtain the current operation data of the target hard disk; Parse the operation data to obtain multiple feature items of the operation data and the field name corresponding to each feature item; Determine that the first linear coefficient of the sample feature item in the initial fault determination model is the single contribution degree of the feature item according to the field name, where the initial fault determination model is obtained by performing training on multiple sets of second training data through machine learning, and each set of the multiple sets of second training data includes: the sample feature item, and a second sample fault probability corresponding to the sample feature item, the sample feature item has the same field name as the feature item, and the single contribution degree refers to the importance of a single feature item for fault detection of the target hard disk; Determine that the feature item with the single contribution degree greater than the first contribution degree threshold is the target feature item.

3. The hard disk failure determination method applied to periodic rules according to claim 1, characterized in that The contribution degree includes a combined contribution degree, and the contribution degree threshold includes a second contribution degree threshold. Determining that a feature item with a contribution degree greater than a contribution degree threshold in the operation data of the target hard disk is a target feature item includes: Determine that a combination of any two of the multiple feature items is a feature combination; Determine that the second linear coefficient of the sample feature combination in the initial fault determination model is the combined contribution degree of the feature combination according to the field names of the two feature items in the feature combination, where the two sample feature items in the sample feature combination have the same field names as the two feature items in the feature combination, and the combined contribution degree refers to the importance of combining the two feature items for fault detection of the target hard disk; Determine the feature combination whose combined contribution degree is greater than the second contribution degree threshold as the target feature combination, where the second contribution degree threshold is a contribution degree threshold that is the same as or different from the first contribution degree threshold; Determine that both of the two feature items in the target feature combination are the target feature items.

4. The method for determining a hard disk failure applied to a periodic rule according to any one of claims 2 or 3, characterized in that, Further included are: Before determining the first linear coefficient of the sample feature item in the initial fault determination model as the contribution degree of the feature item according to the field name, or determining the combined contribution degree of the sample feature combination in the initial fault determination model according to the field names of the two feature items in the feature combination, obtain multiple labeled historical operation data of the target hard disk within a preset historical time period according to the time series, where the labeled historical operation data refers to the historical operation data carrying historical fault probability labels; Divide the multiple labeled historical operation data to obtain a labeled training data set and a labeled verification data set, where the labeled training data set is used for training to obtain the initial fault determination model, and the labeled verification data set is used for performance verification of the initial fault determination model; Iteratively train the machine learning model using the labeled training data set until the iteration termination condition is reached, and determine the currently trained machine learning model as the original fault determination model, where the iteration termination condition includes at least one of the following: the number of iterations reaches the preset number of iterations, and the prediction error is not greater than the error threshold; Perform performance verification on the original fault determination model using the labeled verification data set to obtain a verification result; In the case where the verification result indicates that the performance score of the original fault determination model is higher than the score threshold, determine the original fault determination model as the initial fault determination model.

5. The method for determining hard disk failures applied to periodic rules according to claim 4, wherein After obtaining multiple labeled historical operation data of the target hard disk within a preset historical time period according to the time series, further included are: Eliminate the incomplete labeled historical operation data in the multiple labeled historical operation data, where the incomplete labeled historical operation data refers to the labeled historical operation data with missing values; Perform one-hot encoding processing on the discrete feature items in the multiple labeled historical operation data to convert the data verification of the discrete feature items into a preset target format, where the discrete feature items refer to the feature items with a finite number or countable discontinuous values; Divide the multiple labeled historical operation data according to the sample type to obtain multiple sample data sets, where the sample type includes: fault samples, non-fault samples, and each sample data set includes multiple labeled historical operation data of the same sample type; Use the oversampling algorithm to expand the sample data set with the total number of data less than the preset quantity to obtain the target sample data set, where the total number of data refers to the number of labeled historical operation data in the sample data set; Perform normalization processing on the labeled historical operation data in the sample data set and the target sample data set to complete the preprocessing operation of the multiple labeled historical operation data.

6. The method for determining hard disk failures applied to periodic rules according to claim 4, wherein Iteratively train a machine learning model using a labeled training dataset until, when the iteration termination condition is reached, determine the machine learning model obtained through the current training as the original fault determination model, including: Use the XGBoost model to perform feature reconstruction on multiple pieces of the labeled historical operation data to obtain multiple reconstructed feature items, where the XGBoost model is one of the models in the machine learning model; Iteratively train the FM model using the multiple feature reconstruction items until, when the iteration termination condition is reached, determine the FM model obtained through the current training as the target FM model, where the FM model is a model different from the XGBoost model in the machine learning model; Fuse the XGBoost model and the target FM model to obtain the original fault determination model.

7. The method for determining a hard disk failure applied to a periodic pattern according to claim 6, wherein After determining the original fault determination model as the initial fault determination model when the verification result indicates that the performance score of the original fault determination model is higher than the score threshold, further include: Determine the sample feature items in the initial fault determination model whose first linear coefficient is greater than the first contribution threshold as the sample target feature items; Determine the sample feature items in the sample feature combination in the initial fault determination model whose second linear coefficient is greater than the second contribution threshold as the sample target feature items; Periodically train the target FM model using the sample target feature items in time series to obtain the fault determination model.

8. A hard disk failure determination device applied to periodic rules, characterized in that, Include: A first determination unit, configured to determine the feature items in the operation data of the target hard disk whose contribution degree is greater than the contribution threshold as the target feature items, where the target hard disk is the hard disk for which fault detection is required, and the contribution degree refers to the importance of the feature items for fault detection of the target hard disk; A first acquisition unit, configured to input the target feature items into the fault determination model to process the target feature items using the fault determination model to obtain the fault probability that the target hard disk has a fault, where the fault determination model is a cascaded model obtained by periodically training using multiple sets of first training data through machine learning, the multiple sets of first training data are time series data within a preset time period, and each set of the multiple sets of first training data includes: sample target feature items and the corresponding first sample fault probability; A second determination unit, configured to determine that the target hard disk has the fault when the fault probability is greater than the preset probability value, or determine that the target hard disk does not have the fault when the fault probability is not greater than the preset probability value.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where the program executes the method for determining hard disk faults applicable to periodic rules according to any one of claims 1 to 7.

10. A computer program product comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, execute the method for determining hard disk faults applicable to periodic rules according to any one of claims 1 to 7.

Citation Information

Cited By

  • A method for failure prediction for a data storage device

    CN122507553A