Fault early-warning method and apparatus for heterogeneous hard disk system
By clustering and grouping the hard disk status attribute data of different types of hard disks in heterogeneous hard disk systems and abnormal detection and processing of multi-tower structure models, hard disk health indicator information is generated to conduct fault warning, and the problem of poor fault warning efficiency and accuracy of heterogeneous hard disk systems in the existing technology is solved, achieving more efficient fault warning and system stability.
Patent Information
- Application Number
- PCT/CN2024/089159
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-04-22
- Publication Date
- 2025-06-26
AI Technical Summary
It is difficult for the prior art to effectively warn of faults in heterogeneous hard disk systems, especially when the hard disk models are different and the data distribution is very different, resulting in poor fault warning efficiency and accuracy.
By obtaining the hard disk status attribute data of different models of hard disks in heterogeneous hard disk systems, clustering and grouping to determine the hard disk class cluster, and input the data of each hard disk class cluster into the hard disk failure prediction model of the multi-tower structure for abnormal detection and processing, and generating hard disk health indicator information for fault warning.
It improves the accuracy and efficiency of fault warning of heterogeneous hard disk system, enhances the stability and security of the system, and reduces operation and maintenance costs.
Smart Images

Figure CN2024089159_26062025_PF_FP_ABST
Abstract
Description
Heterogeneous hard disk system failure warning method and device
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to a Chinese patent application filed with the Patent Office of China on December 18, 2023, with application number 202311736825.5 and entitled “A Method and Device for Fault Warning of Heterogeneous Hard Drive System,” the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of intelligent detection technology, and more particularly to a method and device for early warning of a heterogeneous hard disk system failure. Furthermore, the present application also relates to an electronic device and a processor-readable storage medium. Background Art
[0004] In recent years, with the rapid development of technologies such as big data and cloud computing, data volumes have exploded. Cloud service providers have built massive data centers to provide high-quality services to users, and the stable operation of these data centers has become crucial to user experience. Hard drive failures are the most common in data centers, accounting for approximately 78% of hardware failures. Hard drive failures are physically unique, increasing the probability of failure with age. Hard drive failures can have unpredictable consequences. On the one hand, they can cause tasks or systems running on the drives to crash, leading to service interruptions; on the other hand, they can lead to the loss of large amounts of user-stored data. To improve the reliability and security of data centers, various fault-tolerance mechanisms have been adopted, commonly categorized as passive and active. Passive fault-tolerance is a remedial measure implemented after a hard drive failure occurs. For example, Redundant Arrays of Inexpensive Disks (RAID) technology uses virtualized storage technology to combine multiple hard drives into one or more drive groups to achieve data redundancy and improve performance. Data centers also employ multiple replication strategies to improve storage system reliability. For example, HDFS (Hadoop Distributed File System) addresses storage fault tolerance by performing multiple data backups. While these passive fault-tolerance technologies can ensure data security and reliability, they also pose challenges such as high costs and inefficient storage space utilization. Unlike passive fault-tolerance mechanisms, active fault-tolerance predicts drive failures in advance, enabling timely action to reduce operational costs and improve data center reliability and user experience. Due to its significant advantages, active fault-tolerance has become a hot topic in hard drive fault diagnosis.
[0005] SMART (Self-Monitoring, Analysis, and Reporting Technology) is a typical active fault-tolerance technology that detects and records properties related to drive reliability. In recent years, many studies have constructed machine learning and deep learning hard drive failure prediction models based on SMART information. However, these methods typically assume that training and test data come from the same distribution. However, in real data centers, storage systems consist of thousands or even millions of hard drives, often from different vendors or from the same vendor with different models. These different models are referred to as heterogeneous drive systems. Furthermore, the number of drives in a heterogeneous drive system increases with the occurrence of drive failures. Heterogeneous drives typically have different SMART data distributions. Failure prediction models trained using data from a single drive model are not applicable to other models. Existing transfer learning methods have significant limitations, relying on the number of heterogeneous drive types and failing to consider the impact of the number of different drive models on model parameters. This results in poor fault warning efficiency and accuracy. Therefore, designing a more efficient and user-friendly fault warning solution for heterogeneous drive systems has become an urgent challenge.
[0006] Summary of the Invention
[0007] In a first aspect, the present application provides a heterogeneous hard disk system failure warning method, comprising:
[0008] Obtaining hard disk status attribute data of hard disks of different models in a heterogeneous hard disk system, and clustering the hard disk status attribute data according to the hard disk models and the distribution differences of the hard disk status attribute data to determine the hard disk cluster to which the hard disk status attribute data belongs;
[0009] The hard drive status attribute data corresponding to each hard drive cluster is input into a preset multi-tower hard drive failure prediction model for anomaly detection, thereby obtaining hard drive health indicator information output by the multi-tower hard drive failure prediction model. The hard drive failure prediction model is trained based on sample hard drive status attribute data and the hard drive health label information corresponding to the sample hard drive status attribute data.
[0010] Provide fault warning for heterogeneous hard disk systems based on hard disk health indicator information.
[0011] Furthermore, the hard disk status attribute data corresponding to each hard disk cluster is input into a preset multi-tower hard disk failure prediction model for anomaly detection processing, and hard disk health indicator information output by the multi-tower hard disk failure prediction model is obtained, specifically including:
[0012] According to the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selectively input into the corresponding hard disk individual feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the individual feature parameters corresponding to each hard disk cluster;
[0013] Input the hard disk status attribute data into the hard disk common feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the common feature parameters of different hard disk clusters;
[0014] Based on individual characteristic parameters and common characteristic parameters, the target attribute representation corresponding to different models of hard disks in the heterogeneous hard disk system is determined;
[0015] The target attribute representation is input into the hard disk failure warning function module in the multi-tower structure hard disk failure prediction model for fault reasoning analysis to obtain the hard disk health indicator information output by the hard disk failure warning function module.
[0016] Furthermore, based on the hard disk models and the distribution differences of the hard disk status attribute data, the hard disk status attribute data is clustered and grouped to determine the hard disk cluster to which the hard disk status attribute data belongs, specifically including:
[0017] According to the hard disk models in the heterogeneous hard disk system, hard disk status attribute data belonging to the same hard disk model are allocated to the same hard disk cluster; if the data volume of the hard disk status attribute data of one or more hard disk models is greater than or equal to a preset data volume threshold, the corresponding one or more first hard disk clusters are respectively split into multiple second hard disk clusters; wherein the data volume of the first hard disk cluster is greater than the data volume of the second hard disk cluster; or, if the data volume of the hard disk status attribute data of one or more hard disk models is less than the data volume threshold, the corresponding one or more third hard disk clusters are merged into a fourth hard disk cluster; wherein the data volume of the fourth hard disk cluster is greater than the data volume of the third hard disk cluster;
[0018] The hard disk status attribute data is clustered based on the distribution difference of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, and determine the hard disk cluster to which the hard disk status attribute data belongs.
[0019] Furthermore, the hard disk health index information is a hard disk health score value obtained by performing anomaly detection processing on hard disk status attribute data by a multi-tower hard disk fault prediction model; the hard disk health score value is proportional to the health level of the hard disk;
[0020] Fault warnings are provided for heterogeneous hard disk systems based on hard disk health indicator information. Specifically, the hard disk health score is compared and analyzed with the currently selected scoring threshold. If the hard disk health score is less than the scoring threshold, the heterogeneous hard disk system is determined to have a fault and a corresponding fault warning prompt message is generated.
[0021] Furthermore, after comparing and analyzing the hard disk health score value with the currently selected scoring threshold, the method further includes: if the hard disk health score value is greater than or equal to the scoring threshold, determining that the heterogeneous hard disk system is in a healthy state.
[0022] Furthermore, the heterogeneous hard disk system failure warning method also includes: when the heterogeneous hard disk system is in a healthy state, the hard disk status attribute data with a hard disk health score value greater than or equal to the scoring threshold is used as a new training sample to perform adaptive gradient training on the multi-tower structure hard disk failure prediction model, so as to update the parameters of each module in the multi-tower structure hard disk failure prediction model in real time, and obtain a new multi-tower structure hard disk failure prediction model, so as to use the new multi-tower structure hard disk failure prediction model to perform abnormality detection processing on the subsequently input hard disk status attribute data.
[0023] Furthermore, before obtaining the hard disk status attribute data of different models of hard disks in the heterogeneous hard disk system, the method further includes: performing model training in an offline state to obtain a hard disk failure prediction model of a multi-tower structure;
[0024] Model training is performed offline to obtain a multi-tower hard drive failure prediction model, specifically including:
[0025] Acquire sample hard disk status attribute data of the sample hard disk; wherein the sample hard disk status attribute data includes status attribute data corresponding to a healthy hard disk and status attribute data corresponding to an abnormal hard disk;
[0026] Clustering and grouping the sample hard disk status attribute data according to the hard disk models of the sample hard disks and the distribution differences of the sample hard disk status attribute data, and determining the sample hard disk cluster to which the sample hard disk status attribute data belongs;
[0027] An initial multi-tower structure hard disk failure prediction model is trained based on the sample hard disk status attribute data corresponding to the sample hard disk cluster, and by comparing the impact of multiple hard disk anomaly detection thresholds on the final anomaly detection results, multiple hard disk anomaly detection thresholds are screened to determine the hard disk anomaly detection threshold that meets the preset probability confidence condition, and based on the hard disk anomaly detection threshold that meets the preset probability confidence condition, the scoring threshold of the model parameter is updated to obtain the finally trained multi-tower structure hard disk failure prediction model; wherein, the scoring threshold is the hard disk anomaly detection threshold with a probability confidence greater than or equal to the preset probability confidence condition.
[0028] Furthermore, hard disk status attribute data of different models of hard disks in a heterogeneous hard disk system is obtained, specifically including: obtaining original hard disk status attribute data of different models of hard disks in the heterogeneous hard disk system, filling missing values in the original hard disk status attribute data, and obtaining first hard disk status attribute data; performing feature screening on the first hard disk status attribute data to obtain second hard disk status attribute data; and normalizing the second hard disk status attribute data to obtain hard disk status attribute data.
[0029] Furthermore, based on the individual characteristic parameters and the common characteristic parameters, target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system are determined, specifically including:
[0030] The individual characteristic parameters and the common characteristic parameters are multiplied to obtain the target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system.
[0031] Furthermore, clustering the hard disk status attribute data based on the distribution difference of the hard disk status attribute data to assign the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, and determining the hard disk cluster to which the hard disk status attribute data belongs, specifically includes:
[0032] Determine the metric to use for clustering;
[0033] The hard disk status attribute data is clustered based on the distribution difference and measurement criteria of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster to obtain the hard disk cluster to which the hard disk status attribute data belongs.
[0034] Furthermore, based on the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selectively input into the corresponding hard disk individual feature extraction module in the multi-tower structure hard disk failure prediction model to obtain individual feature parameters corresponding to each hard disk cluster, specifically including:
[0035] Determine identification information of a corresponding hard disk personality feature extraction module according to the hard disk cluster to which the current hard disk status attribute data belongs;
[0036] Based on the identification information, the hard disk status attribute data is selectively input into the corresponding hard disk personality feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the personality feature parameters corresponding to each hard disk cluster.
[0037] Furthermore, the multi-tower hard disk failure prediction model includes multiple hard disk personality feature extraction modules, each hard disk personality feature extraction module processes the hard disk status attribute data of a hard disk cluster, and each hard disk personality feature extraction module corresponds to a target weight parameter.
[0038] In a second aspect, the present application further provides a heterogeneous hard disk system failure warning device, comprising:
[0039] A clustering and grouping unit is used to obtain hard disk status attribute data of different models of hard disks in a heterogeneous hard disk system, and cluster the hard disk status attribute data according to the hard disk models and the distribution differences of the hard disk status attribute data to determine the hard disk cluster to which the hard disk status attribute data belongs;
[0040] a fault analysis unit configured to input the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower hard disk fault prediction model for anomaly detection and processing, thereby obtaining hard disk health indicator information output by the multi-tower hard disk fault prediction model; wherein the hard disk fault prediction model is trained based on sample hard disk status attribute data and hard disk health label information corresponding to the sample hard disk status attribute data;
[0041] The fault warning unit is used to provide fault warning for heterogeneous hard disk systems based on hard disk health indicator information.
[0042] Furthermore, the fault analysis unit is specifically used to:
[0043] According to the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selectively input into the corresponding hard disk individual feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the individual feature parameters corresponding to each hard disk cluster;
[0044] Input the hard disk status attribute data into the hard disk common feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the common feature parameters of different hard disk clusters;
[0045] Based on individual characteristic parameters and common characteristic parameters, the target attribute representation corresponding to different models of hard disks in the heterogeneous hard disk system is determined;
[0046] The target attribute representation is input into the hard disk failure warning function module in the multi-tower structure hard disk failure prediction model for fault reasoning analysis to obtain the hard disk health indicator information output by the hard disk failure warning function module.
[0047] Furthermore, based on the hard disk models and the distribution differences of the hard disk status attribute data, the hard disk status attribute data is clustered and grouped to determine the hard disk cluster to which the hard disk status attribute data belongs, specifically including:
[0048] According to the hard disk models in the heterogeneous hard disk system, hard disk status attribute data belonging to the same hard disk model are allocated to the same hard disk cluster; if the data volume of the hard disk status attribute data of one or more hard disk models is greater than or equal to a preset data volume threshold, the corresponding one or more first hard disk clusters are respectively split into multiple second hard disk clusters; wherein the data volume of the first hard disk cluster is greater than the data volume of the second hard disk cluster; or, if the data volume of the hard disk status attribute data of one or more hard disk models is less than the data volume threshold, the corresponding one or more third hard disk clusters are merged into a fourth hard disk cluster; wherein the data volume of the fourth hard disk cluster is greater than the data volume of the third hard disk cluster;
[0049] The hard disk status attribute data is clustered based on the distribution difference of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, and determine the hard disk cluster to which the hard disk status attribute data belongs.
[0050] Furthermore, the hard disk health index information is a hard disk health score value obtained by performing anomaly detection processing on hard disk status attribute data by a multi-tower hard disk fault prediction model; the hard disk health score value is proportional to the health level of the hard disk;
[0051] The fault warning unit is specifically used to compare and analyze the hard disk health score value with the currently selected scoring threshold. When the hard disk health score value is less than the scoring threshold, it is determined that the heterogeneous hard disk system has a fault and a corresponding fault warning prompt message is generated.
[0052] Furthermore, after comparing and analyzing the hard disk health score value with the currently selected scoring threshold, the fault warning unit is also used to: determine that the heterogeneous hard disk system is in a healthy state when the hard disk health score value is greater than or equal to the scoring threshold.
[0053] Furthermore, the heterogeneous hard disk system failure warning device also includes: a model parameter updating unit, which is used to, when the heterogeneous hard disk system is in a healthy state, use the hard disk status attribute data with a hard disk health score value greater than or equal to the scoring threshold as a new training sample to perform adaptive gradient training on the multi-tower structure hard disk failure prediction model, so as to update the parameters of each module in the multi-tower structure hard disk failure prediction model in real time, and obtain a new multi-tower structure hard disk failure prediction model, so as to use the new multi-tower structure hard disk failure prediction model to perform abnormality detection processing on the subsequently input hard disk status attribute data.
[0054] Furthermore, before obtaining the hard disk status attribute data of different models of hard disks in the heterogeneous hard disk system, the method further includes: a model offline training unit for performing model training in an offline state to obtain a hard disk failure prediction model of a multi-tower structure;
[0055] Model offline training unit, specifically used for:
[0056] Acquire sample hard disk status attribute data of the sample hard disk; wherein the sample hard disk status attribute data includes status attribute data corresponding to a healthy hard disk and status attribute data corresponding to an abnormal hard disk;
[0057] Clustering and grouping the sample hard disk status attribute data according to the hard disk models of the sample hard disks and the distribution differences of the sample hard disk status attribute data, and determining the sample hard disk cluster to which the sample hard disk status attribute data belongs;
[0058] An initial multi-tower structure hard disk failure prediction model is trained based on the sample hard disk status attribute data corresponding to the sample hard disk cluster, and by comparing the impact of multiple hard disk anomaly detection thresholds on the final anomaly detection results, multiple hard disk anomaly detection thresholds are screened to determine the hard disk anomaly detection threshold that meets the preset probability confidence condition, and based on the hard disk anomaly detection threshold that meets the preset probability confidence condition, the scoring threshold of the model parameter is updated to obtain the finally trained multi-tower structure hard disk failure prediction model; wherein, the scoring threshold is the hard disk anomaly detection threshold with a probability confidence greater than or equal to the preset probability confidence condition.
[0059] Furthermore, the clustering grouping unit is specifically used to: obtain original hard disk status attribute data of different models of hard disks in a heterogeneous hard disk system, fill missing values in the original hard disk status attribute data, and obtain first hard disk status attribute data; perform feature screening on the first hard disk status attribute data to obtain second hard disk status attribute data; and perform normalization processing on the second hard disk status attribute data to obtain hard disk status attribute data.
[0060] Furthermore, based on the individual characteristic parameters and the common characteristic parameters, target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system are determined, specifically including:
[0061] The individual characteristic parameters and the common characteristic parameters are multiplied to obtain the target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system.
[0062] Furthermore, clustering the hard disk status attribute data based on the distribution difference of the hard disk status attribute data to assign the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, and determining the hard disk cluster to which the hard disk status attribute data belongs, specifically includes:
[0063] Determine the metric to use for clustering;
[0064] The hard disk status attribute data is clustered based on the distribution difference and measurement criteria of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster to obtain the hard disk cluster to which the hard disk status attribute data belongs.
[0065] Furthermore, based on the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selectively input into the corresponding hard disk individual feature extraction module in the multi-tower structure hard disk failure prediction model to obtain individual feature parameters corresponding to each hard disk cluster, specifically including:
[0066] Determine identification information of a corresponding hard disk personality feature extraction module according to the hard disk cluster to which the current hard disk status attribute data belongs;
[0067] Based on the identification information, the hard disk status attribute data is selectively input into the corresponding hard disk personality feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the personality feature parameters corresponding to each hard disk cluster.
[0068] Furthermore, the multi-tower hard disk failure prediction model includes multiple hard disk personality feature extraction modules, each hard disk personality feature extraction module processes the hard disk status attribute data of a hard disk cluster, and each hard disk personality feature extraction module corresponds to a target weight parameter.
[0069] In a third aspect, the present application also provides an electronic device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of any one of the above heterogeneous hard disk system failure warning methods are implemented.
[0070] In a fourth aspect, the present application also provides a processor-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, the steps of any one of the above heterogeneous hard disk system failure warning methods are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0072] FIG1 is a flow chart of a method for early warning of a heterogeneous hard disk system failure according to one or more embodiments of the present application;
[0073] FIG2 is a schematic diagram of a complete process of a heterogeneous hard disk system failure warning method provided by one or more embodiments of the present application;
[0074] FIG3 is a schematic diagram of a hard disk failure prediction model for a multi-tower structure in a heterogeneous hard disk system failure warning method provided by one or more embodiments of the present application;
[0075] FIG4 is a schematic diagram of the structure of a heterogeneous hard disk system failure warning device provided by one or more embodiments of the present application;
[0076] FIG5 is a schematic diagram of the hardware environment of a heterogeneous hard disk system failure warning method provided by one or more embodiments of the present application;
[0077] FIG6 is a schematic diagram of the physical structure of an electronic device provided by one or more embodiments of the present application. DETAILED DESCRIPTION
[0078] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0079] The following first describes in detail the embodiment of the heterogeneous hard disk system failure warning method of the present application. As shown in Figure 1, it is a flow chart of the heterogeneous hard disk system failure warning method provided by the embodiment of the present application. The specific implementation process includes the following steps:
[0080] Step 101: Obtain hard disk status attribute data of different models of hard disks in a heterogeneous hard disk system, and cluster the hard disk status attribute data according to the hard disk models and distribution differences of the hard disk status attribute data to determine the hard disk cluster to which the hard disk status attribute data belongs.
[0081] In an embodiment of the present application, after obtaining hard disk status attribute data of different hard disk models in a heterogeneous hard disk system, hard disk status attribute data of the same hard disk model can be allocated to the same hard disk cluster based on the hard disk models in the heterogeneous hard disk system; if the data volume of hard disk status attribute data of one or more hard disk models is greater than or equal to a preset data volume threshold, the corresponding one or more first hard disk clusters are respectively split into multiple second hard disk clusters; wherein the data volume of the first hard disk cluster is greater than the data volume of the second hard disk cluster; or, if the data volume of hard disk status attribute data of one or more hard disk models is less than the data volume threshold, the corresponding one or more third hard disk clusters are merged into a fourth hard disk cluster; wherein the data volume of the fourth hard disk cluster is greater than the data volume of the third hard disk cluster. The hard disk status attribute data are clustered based on the distribution differences of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, thereby determining the hard disk cluster to which the hard disk status attribute data belongs. Obtaining hard disk status attribute data for different models of hard disks in a heterogeneous hard disk system involves the following specific implementation steps: obtaining raw hard disk status attribute data for different models of hard disks in the heterogeneous hard disk system, filling missing values in the raw hard disk status attribute data to obtain first hard disk status attribute data; performing feature filtering on the first hard disk status attribute data to obtain second hard disk status attribute data; and normalizing the second hard disk status attribute data to obtain hard disk status attribute data. Clustering the hard disk status attribute data based on the distribution differences of the hard disk status attribute data to assign the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster and determine the hard disk cluster to which the hard disk status attribute data belongs requires first determining a metric for clustering, and then clustering the hard disk status attribute data based on the metric and the distribution differences of the hard disk status attribute data to assign the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster to obtain the hard disk cluster to which the hard disk status attribute data belongs. The hard disk status attribute data includes, but is not limited to, data read error rate, disk startup time, motor start / stop count, number of remapped sectors, seek error rate, seek time, and hard disk power-on time. The data read error rate is the hardware read error rate when reading data from the disk surface. The lower the value, the better. The disk startup time is the average time it takes for the disk to start from a standstill and then accelerate to its rated speed. The lower the value, the better. I won't go into detail here. A drive cluster is a collection of drive status attribute data for drives of the same model or with similar attribute status.
[0082] It should be noted that, as shown in FIG2 , the embodiment of the present application includes an offline training phase, an online prediction phase, and an online model parameter update phase. The offline training phase includes the solid line connection portion in FIG2 , and the online prediction phase and the online model parameter update phase include the dotted line connection portion in FIG2 , that is, the solid line represents the offline training phase, and the dotted line represents the online prediction phase and the online model parameter update phase. That is, before executing this step, it is necessary to perform model training in an offline state in advance to obtain a hard disk failure prediction model for a multi-tower structure. The model training is performed in an offline state to obtain a hard disk failure prediction model for a multi-tower structure. The corresponding specific implementation process includes: obtaining sample hard disk status attribute data of a sample hard disk; wherein the sample hard disk status attribute data includes status attribute data corresponding to a healthy hard disk and status attribute data corresponding to an abnormal hard disk. According to the hard disk model of the sample hard disk and the distribution difference of the sample hard disk status attribute data, the sample hard disk status attribute data is clustered and grouped to determine the sample hard disk cluster to which the sample hard disk status attribute data belongs. An initial multi-tower hard drive failure prediction model is trained based on sample hard drive status attribute data corresponding to sample hard drive clusters. Multiple hard drive anomaly detection thresholds are screened by comparing their impact on the final anomaly detection result, determining a hard drive anomaly detection threshold that satisfies a preset probability confidence condition. The scoring threshold for the model parameters is updated based on the hard drive anomaly detection threshold that satisfies the preset probability confidence condition, thereby obtaining a finally trained multi-tower hard drive failure prediction model. The scoring threshold is a hard drive anomaly detection threshold with a probability confidence greater than or equal to the preset probability confidence condition. That is, when screening the hard drive anomaly detection threshold, the impact of different thresholds on the final prediction result can be compared based on existing normal hard drive data (i.e., hard drive health status data) and faulty hard drive data (i.e., hard drive anomaly status data). The screening result of the anomaly detection threshold (i.e., the scoring threshold) is determined to update the high-confidence scoring threshold for the model parameters.
[0083] Specifically, in the offline training phase, first, based on the hard drive model, hard drive status attribute data (i.e., SMART data), etc., a clustering algorithm is used to divide the hard drive related data of different models of hard drives into multiple hard drive clusters, and the data of these hard drive clusters are used to train the multi-tower anomaly detection model (i.e., the initial multi-tower structured hard drive failure prediction model). The online prediction phase includes obtaining new data through preprocessing, using the multi-tower anomaly detection model completed in the offline training phase (i.e., the trained multi-tower structured hard drive failure prediction model) to predict whether there is a risk of failure at this moment, and issuing an early warning message based on this. In addition, it is determined whether the hard drive health score meets the high confidence condition. If so, the current test sample is used to update the model parameters. The specific execution process of the algorithm is as follows: First, data preprocessing is performed: the SMART data of different models of hard drives are preprocessed, including missing value filling, normalization, feature screening and other preprocessing contents. Then, clustering is performed based on the differences in hard drive models and hard drive status attribute data: the SMART data of the pre-processed hard drives is clustered. To account for the different data volume ratios of hard drives of different models, the following principles are followed when clustering the hard drive data, namely the coarse segmentation stage and the fine segmentation stage: In the coarse segmentation stage, the data of hard drives of the same model can be placed in the same hard drive cluster. If the data volume of one or several hard drive models is too large, the hard drive clusters of the same model are split, that is, a large hard drive cluster (i.e., the first hard drive cluster) is split into multiple small clusters (i.e., the second hard drive cluster); one or several hard drive clusters with relatively small data volumes (i.e., the third hard drive cluster) can be merged to obtain a large hard drive cluster (i.e., the hard drive cluster). In the fine segmentation stage, the above cluster splitting or merging is based on the hard drive model. To further ensure that hard drive data with similar data distribution are aggregated and grouped, a clustering algorithm (such as K-means, Dbscan, hierarchical clustering) can be used to cluster the hard drive status attribute data using a preset metric.
[0084] In training the multi-tower anomaly detection model, the hard disk status attribute data of the multiple hard disk clusters obtained above are used as sample data to effectively train the multi-tower anomaly detection model and obtain a hard disk failure prediction model with a multi-tower structure. The structure of the hard disk failure prediction model with a multi-tower structure is shown in Figure 3. The specific process is as follows: the sample data used for training usually includes SMART feature information, label information of whether the hard disk is abnormal at this moment, and the cluster indicator to which the sample data belongs (Input in Figure 3).
[0085] The hard disk status attribute data of the above-mentioned multiple hard disk clusters are input into the multi-tower anomaly detection model. The Gate Net network structure in the multi-tower anomaly detection model (Gate Net in Figure 3) will select the sample data to be input into the corresponding hard disk-specific weight training module (i.e., Disk-1 Weight, Disk-2 Weight..., Disk-N Weight in Figure 3, where N is a positive integer greater than 1) according to the hard disk cluster to which the current sample data belongs. The hard disk-specific weight training module is also called the hard disk personality feature extraction module or the hard disk-specific parameter module, and outputs the unique parameters (i.e., personality feature parameters) of the hard disk cluster to which the hard disk belongs; in addition, the sample data will be input into the hard disk shared weight training module (i.e., Disk Shared Weight in Figure 3, also called the hard disk common feature extraction module or the hard disk shared parameter module) to obtain the common characteristic parameters (i.e., common feature parameters) of different hard disk clusters, and the output results of the hard disk common feature extraction module and the hard disk personality feature extraction module are multiplied to obtain the final target attribute representation of the hard disk. The final target attribute representation of the hard disk is input into the hard disk fault warning function module (corresponding to the Output in Figure 3) for fault reasoning analysis to obtain the hard disk score (i.e., the hard disk health score value), which is used as the hard disk health indicator information.
[0086] The training target algorithm of the multi-tower anomaly detection model is as follows: the training target is to calculate the predicted value of the output The actual value of the input The maximum similarity between:
[0087] in, Represents the predicted value of the sample data of the i-th hard disk cluster c (i.e., the output hard disk health score value); is the actual value of the corresponding sample data (i.e., the input hard disk health score); N is the number of hard disk clusters divided in the embodiment of the present invention; N c is the total number of sample data in the hard disk cluster c, that is, the total amount of hard disk status attribute data of different models of hard disks.
[0088] Step 102: Input the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower structure hard disk failure prediction model for anomaly detection processing, and obtain the hard disk health indicator information output by the multi-tower structure hard disk failure prediction model; wherein the hard disk failure prediction model is trained based on the sample hard disk status attribute data and the hard disk health label information corresponding to the sample hard disk status attribute data.
[0089] In an embodiment of the present application, the hard disk status attribute data can be selected and input into the corresponding hard disk individual feature extraction module in the hard disk failure prediction model of the multi-tower structure according to the hard disk cluster to which the current hard disk status attribute data belongs, and the individual feature parameters corresponding to each hard disk cluster are obtained. The hard disk status attribute data is input into the hard disk common feature extraction module in the hard disk failure prediction model of the multi-tower structure to obtain the common feature parameters of different hard disk clusters. Based on the individual feature parameters and the common feature parameters, the target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system are determined. The target attribute representations are input into the hard disk failure warning function module in the hard disk failure prediction model of the multi-tower structure for fault reasoning analysis to obtain the hard disk health index information output by the hard disk failure warning function module. Among them, based on the individual feature parameters and the common feature parameters, the target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system are determined, and the corresponding specific implementation process may include: multiplying the individual feature parameters and the common feature parameters to obtain the target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system.
[0090] In the process of selectively inputting the hard disk status attribute data into the corresponding hard disk individual feature extraction module in the multi-tower hard disk failure prediction model based on the hard disk cluster to which the current hard disk status attribute data belongs, and obtaining the individual feature parameters corresponding to each hard disk cluster, it is necessary to first determine the identification information of the corresponding hard disk individual feature extraction module based on the hard disk cluster to which the current hard disk status attribute data belongs, and then, based on the identification information, selectively inputting the hard disk status attribute data into the corresponding hard disk individual feature extraction module in the multi-tower hard disk failure prediction model to obtain the individual feature parameters corresponding to each hard disk cluster. The multi-tower hard disk failure prediction model includes multiple hard disk individual feature extraction modules, each of which processes the hard disk status attribute data of a hard disk cluster, and each hard disk individual feature extraction module corresponds to a target weight parameter.
[0091] As shown in Figure 3, the network structure of the multi-tower hard drive failure prediction model includes multiple sub-networks, including hard drive individual feature extraction module 1, hard drive individual feature extraction module 2, ..., hard drive individual feature extraction module n (i.e., corresponding to Disk-1 Weight, Disk-2 Weight, ..., Disk-N Weight, respectively), a hard drive common feature extraction module (i.e., Disk Shared Weight), a hard drive failure warning function module (i.e., Output), and a weight factor prediction module (i.e., Gate Net network structure). The Gate Net network structure is used to extract the weight factors corresponding to the hard drive status attribute data of different hard drive models in a heterogeneous hard drive system (i.e., the probability value of the data corresponding to the input of each hard drive individual feature extraction module). For example, if N is 3, the weight factor corresponding to the Disk-1 Weight sub-network can be 0.1, the weight factor corresponding to the Disk-2 Weight sub-network can be 0.8, and the weight factor corresponding to the Disk-3 Weight sub-network can be 0.1. The above weight factors are multiplied by the individual features extracted by the three hard drive individual feature extraction modules and then concatenated to obtain the individual feature parameters corresponding to each hard drive cluster. The hard drive status attribute data is input into the hard drive common feature extraction module within the multi-tower hard drive failure prediction model to obtain common feature parameters for different hard drive clusters. By multiplying the individual and common feature parameters, the target attribute representations corresponding to different hard drive models in the heterogeneous hard drive system are determined. The target attribute representations are then input into the hard drive failure warning module within the multi-tower hard drive failure prediction model for fault reasoning analysis, which results in the hard drive health indicator information output by the module.
[0092] In the embodiment of the present application, the algorithm formulas corresponding to the multiple hard disk personality feature extraction modules can be:
[0093] FFN(x)=relu(relu(xW1+b1)W2+b2);
[0094] In this formula, x is the hard disk status attribute data of different models of hard disks; w1 and b1 are the first weight coefficient and the first bias coefficient of the first layer network structure in the module respectively; the relu inside the brackets is the activation function of the first layer; w2 and b2 are the second weight coefficient and the second bias coefficient of the second layer network in the module respectively; the relu outside the brackets is the activation function of the second layer network.
[0095] In the embodiment of the present application, the algorithm formula corresponding to the hard disk common feature extraction module can be:
[0096] FFN(x)=relu(relu(xW1+b1)W2+b2);
[0097] In this formula, x is also the hard disk status attribute data of different models of hard disks; w1 and b1 are the first weight coefficient and the first bias coefficient of the first layer network structure in the module respectively; the activation function of the first layer of relu inside the brackets; w2 and b2 are the second weight coefficient and the second bias coefficient of the second layer network in the module respectively; the activation function of the second layer of relu outside the brackets.
[0098] In the embodiment of the present application, the algorithm formula corresponding to the hard disk failure warning function module can be:
[0099] FFN(x)=sigmoid(xW1+b1);
[0100] In this formula, x is the target attribute representation obtained by multiplying the common representation and unique representation of the hard disk; w1 and b1 are the first weight coefficient and the first bias coefficient of the network structure in this module, respectively; sigmoid is the activation function of the network structure, and the output is 0 to 1, indicating whether the hard disk has failed. A larger score value indicates a greater probability of failure.
[0101] In the embodiment of the present application, the algorithm formula corresponding to the weight coefficient prediction module can be:
[0102] FFN(x)=softmax(relu(xW1+b1)W2+b2);
[0103] In this formula, x is the relevant data of different hard disk models, such as the hard disk model and manufacturer identification; w1 and b1 are the first weight coefficient and the first bias coefficient of the first layer network structure in this module respectively; the relu activation function of the first layer in the brackets, w2 and b2 are the weight coefficient and bias coefficient respectively; softmax is the function of normalizing the network, replacing the corresponding output with a weight factor between 0 and 1, and the sum of all weight factors is equal to 1.
[0104] Step 103: Produce a fault warning for the heterogeneous hard disk system based on the hard disk health indicator information.
[0105] In an embodiment of the present application, the hard disk health index information is a hard disk health score value obtained by performing abnormal detection processing on the hard disk status attribute data by a multi-tower hard disk fault prediction model; the hard disk health score value is proportional to the health level of the hard disk.
[0106] Specifically, the hard drive health score is compared with the currently selected scoring threshold. If the hard drive health score is less than the threshold, the heterogeneous hard drive system is determined to have failed, and a corresponding failure warning is generated. In other words, during the online prediction phase, new data is pre-processed and then fed into the multi-tower hard drive failure prediction model to score the current health level of the hard drives. This then determines whether to issue a warning based on the detection threshold.
[0107] In addition, after comparing and analyzing the hard drive health score value with the currently selected scoring threshold, the heterogeneous hard drive system can be determined to be in a healthy state if the hard drive health score value is greater than or equal to the scoring threshold. If the heterogeneous hard drive system is in a healthy state, the hard drive status attribute data with a hard drive health score value greater than or equal to the scoring threshold can be used as a new training sample to perform adaptive gradient training on the multi-tower hard drive fault prediction model, so as to update the parameters of each module in the multi-tower hard drive fault prediction model in real time and obtain a new multi-tower hard drive fault prediction model. The new multi-tower hard drive fault prediction model can then be used to perform anomaly detection on subsequently input hard drive status attribute data. In other words, during the online update of model parameters, the high-confidence scoring threshold obtained can be used to determine whether the current hard drive score value meets the preset confidence condition. If the condition is met, the current sample is used to update the model parameters; otherwise, the model parameters are not updated.
[0108] In the embodiment of the present application, solutions can be provided for the problems such as the impact of data imbalance of different models of hard disks on the update of model parameters and the adaptive update of model parameters, as follows: 1) Based on information such as the hard disk model and the distribution difference of SMART data, a clustering algorithm is used to cluster and group the SMART data of hard disks of different hard disk models to ensure the consistency of data volume and data distribution difference between hard disk clusters, and the consistency of data distribution within hard disk clusters; 2) Based on the SMART data of different hard disk clusters, a multi-tower fault warning model (i.e., a hard disk fault prediction model with a multi-tower structure) that can adapt to hard disks of different models can be established. The multi-tower fault warning model is mainly divided into a hard disk shared parameter module and a hard disk-specific Parameter module and hard disk fault warning function module, hard disk shared parameters are used to capture the commonalities of hard disks of different models, namely common characteristic parameters; hard disk specific parameters are used to characterize the data distribution properties specific to the hard disk cluster (i.e. individual characteristic parameters), and the two are multiplied together to characterize the final target attribute representation of the current hard disk, the target attribute representation is a feature vector, and the target attribute representation is input into the hard disk fault warning function module to obtain information on whether the hard disk will fail at the current moment; 3) According to the probability confidence of the predicted hard disk failure, the parameters of each module of the multi-tower fault warning model are adaptively gradient updated to realize the real-time update of the parameters of each module in the multi-tower fault warning model, thereby ensuring the prediction accuracy of the multi-tower fault warning model.
[0109] It should be noted that clustering hard drives of different models can solve the problem of data imbalance of hard drive models. The imbalance problem is the application of machine learning in the field of anomaly detection. In the model, the source of parameter updates mostly comes from sample data with a larger proportion (i.e., normal hard drive data in this article), which cannot well represent the characteristics of a small number of sample types; in the prediction results, it is manifested as a high overall accuracy, but a low accuracy on samples with fewer types. Specifically, the proposed solution manifests itself in two aspects: the distribution of data volume of hard drives of different models varies greatly. This solution adopts coarse clustering and fine clustering to achieve a balance in the number of different hard drive models; the data volume of faulty hard drives and normal hard drives differs greatly. Upsampling or adversarial learning can be used to solve the problem for faulty hard drive data, which will not be elaborated here. In addition, the multi-tower structure based on the hard drive failure prediction model can solve the commonalities and differences of hard drive data of different models from the model structure level. That is, the application solution is divided into multiple towers according to the number of hard drive clusters, and the model structure prevents the model parameters from being biased by a large amount of data. In addition, the multi-tower structure remains independent of each other, and is still effective even for hard drive clusters with a small amount of data, further optimizing the sample imbalance problem. This application can filter high-confidence samples to update the parameters of each module in the multi-tower fault warning model. The robustness of hard disk fault warning is reflected in the indicators of increasing the warning rate of faulty hard disks and reducing the false alarm rate of normal hard disks. This application can achieve real-time updating of model parameters in many aspects to ensure that the model can capture the data distribution characteristics of heterogeneous hard disk systems in real time. When updating samples, high-confidence samples are selected to further ensure that the model parameters are not affected by misjudged samples. In addition, due to the time-varying properties of the disk, when constructing the multi-tower fault warning model, the input to the Input module in Figure 3 is sequence data within a time window, and the LSTM model is integrated to better capture this time-varying characteristic. Specifically, for each sample in the time series, in addition to the SMART data of the sample data itself as input, the LSTM structure of the previous sample is also used to hide the traversal input.
[0110] The heterogeneous hard disk system fault warning method of the embodiment of the present application obtains hard disk status attribute data of different models of hard disks in the heterogeneous hard disk system, and clusters and groups the hard disk status attribute data according to the distribution differences of the hard disk models and the hard disk status attribute data, determines the hard disk cluster to which the hard disk status attribute data belongs, inputs the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower structure hard disk fault prediction model for anomaly detection processing, obtains hard disk health indicator information output by the multi-tower structure hard disk fault prediction model, and performs fault warning for the heterogeneous hard disk system based on the hard disk health indicator information, which can effectively improve the accuracy and efficiency of the heterogeneous hard disk system fault warning, thereby improving the stability and security of the operation of the heterogeneous hard disk system.
[0111] Corresponding to the heterogeneous hard disk system failure warning method provided above, the present application also provides a heterogeneous hard disk system failure warning device. Since the embodiment of this device is similar to the above method embodiment, the description is relatively simple. For relevant details, please refer to the description of the above method embodiment. The embodiment of the heterogeneous hard disk system failure warning device described below is only exemplary. Please refer to Figure 4, which is a structural diagram of a heterogeneous hard disk system failure warning device provided in an embodiment of the present application. The heterogeneous hard disk system failure warning device of the present application includes the following parts:
[0112] Clustering and grouping unit 401 is used to obtain hard disk status attribute data of different models of hard disks in a heterogeneous hard disk system, and cluster the hard disk status attribute data according to the hard disk models and the distribution differences of the hard disk status attribute data to determine the hard disk cluster to which the hard disk status attribute data belongs;
[0113] Fault analysis unit 402 is configured to input the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower hard disk fault prediction model for anomaly detection, thereby obtaining hard disk health indicator information output by the multi-tower hard disk fault prediction model; wherein the hard disk fault prediction model is trained based on sample hard disk status attribute data and hard disk health label information corresponding to the sample hard disk status attribute data;
[0114] The fault warning unit 403 is used to provide fault warning for the heterogeneous hard disk system based on the hard disk health indicator information.
[0115] Furthermore, the fault analysis unit is specifically used to:
[0116] According to the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selectively input into the corresponding hard disk individual feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the individual feature parameters corresponding to each hard disk cluster;
[0117] Input the hard disk status attribute data into the hard disk common feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the common feature parameters of different hard disk clusters;
[0118] Based on individual characteristic parameters and common characteristic parameters, the target attribute representation corresponding to different models of hard disks in the heterogeneous hard disk system is determined;
[0119] The target attribute representation is input into the hard disk failure warning function module in the multi-tower structure hard disk failure prediction model for fault reasoning analysis to obtain the hard disk health indicator information output by the hard disk failure warning function module.
[0120] Furthermore, based on the hard disk models and the distribution differences of the hard disk status attribute data, the hard disk status attribute data is clustered and grouped to determine the hard disk cluster to which the hard disk status attribute data belongs, specifically including:
[0121] According to the hard disk models in the heterogeneous hard disk system, hard disk status attribute data belonging to the same hard disk model are allocated to the same hard disk cluster; if the data volume of the hard disk status attribute data of one or more hard disk models is greater than or equal to a preset data volume threshold, the corresponding one or more first hard disk clusters are respectively split into multiple second hard disk clusters; wherein the data volume of the first hard disk cluster is greater than the data volume of the second hard disk cluster; or, if the data volume of the hard disk status attribute data of one or more hard disk models is less than the data volume threshold, the corresponding one or more third hard disk clusters are merged into a fourth hard disk cluster; wherein the data volume of the fourth hard disk cluster is greater than the data volume of the third hard disk cluster;
[0122] The hard disk status attribute data is clustered based on the distribution difference of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, and determine the hard disk cluster to which the hard disk status attribute data belongs.
[0123] Furthermore, the hard disk health index information is a hard disk health score value obtained by performing anomaly detection processing on hard disk status attribute data by a multi-tower hard disk fault prediction model; the hard disk health score value is proportional to the health level of the hard disk;
[0124] The fault warning unit is specifically used to compare and analyze the hard disk health score value with the currently selected scoring threshold. When the hard disk health score value is less than the scoring threshold, it is determined that the heterogeneous hard disk system has a fault and a corresponding fault warning prompt message is generated.
[0125] Furthermore, after comparing and analyzing the hard disk health score value with the currently selected scoring threshold, the fault warning unit is also used to: determine that the heterogeneous hard disk system is in a healthy state when the hard disk health score value is greater than or equal to the scoring threshold.
[0126] Furthermore, the heterogeneous hard disk system failure warning device also includes: a model parameter updating unit, which is used to, when the heterogeneous hard disk system is in a healthy state, use the hard disk status attribute data with a hard disk health score value greater than or equal to the scoring threshold as a new training sample to perform adaptive gradient training on the multi-tower structure hard disk failure prediction model, so as to update the parameters of each module in the multi-tower structure hard disk failure prediction model in real time, and obtain a new multi-tower structure hard disk failure prediction model, so as to use the new multi-tower structure hard disk failure prediction model to perform abnormality detection processing on the subsequently input hard disk status attribute data.
[0127] Furthermore, before obtaining the hard disk status attribute data of different models of hard disks in the heterogeneous hard disk system, the method further includes: a model offline training unit for performing model training in an offline state to obtain a hard disk failure prediction model of a multi-tower structure;
[0128] Model offline training unit, specifically used for:
[0129] Acquire sample hard disk status attribute data of the sample hard disk; wherein the sample hard disk status attribute data includes status attribute data corresponding to a healthy hard disk and status attribute data corresponding to an abnormal hard disk;
[0130] Clustering and grouping the sample hard disk status attribute data according to the hard disk models of the sample hard disks and the distribution differences of the sample hard disk status attribute data, and determining the sample hard disk cluster to which the sample hard disk status attribute data belongs;
[0131] An initial multi-tower structure hard disk failure prediction model is trained based on the sample hard disk status attribute data corresponding to the sample hard disk cluster, and by comparing the impact of multiple hard disk anomaly detection thresholds on the final anomaly detection results, multiple hard disk anomaly detection thresholds are screened to determine the hard disk anomaly detection threshold that meets the preset probability confidence condition, and based on the hard disk anomaly detection threshold that meets the preset probability confidence condition, the scoring threshold of the model parameter is updated to obtain the finally trained multi-tower structure hard disk failure prediction model; wherein, the scoring threshold is the hard disk anomaly detection threshold with a probability confidence greater than or equal to the preset probability confidence condition.
[0132] Furthermore, the clustering grouping unit is specifically used to: obtain original hard disk status attribute data of different models of hard disks in a heterogeneous hard disk system, fill missing values in the original hard disk status attribute data, and obtain first hard disk status attribute data; perform feature screening on the first hard disk status attribute data to obtain second hard disk status attribute data; and perform normalization processing on the second hard disk status attribute data to obtain hard disk status attribute data.
[0133] Furthermore, based on the individual characteristic parameters and the common characteristic parameters, target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system are determined, specifically including:
[0134] The individual characteristic parameters and the common characteristic parameters are multiplied to obtain the target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system.
[0135] Furthermore, clustering the hard disk status attribute data based on the distribution difference of the hard disk status attribute data to assign the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, and determining the hard disk cluster to which the hard disk status attribute data belongs, specifically includes:
[0136] Determine the metric to use for clustering;
[0137] The hard disk status attribute data is clustered based on the distribution difference and measurement criteria of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster to obtain the hard disk cluster to which the hard disk status attribute data belongs.
[0138] Furthermore, based on the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selectively input into the corresponding hard disk individual feature extraction module in the multi-tower structure hard disk failure prediction model to obtain individual feature parameters corresponding to each hard disk cluster, specifically including:
[0139] Determine identification information of a corresponding hard disk personality feature extraction module according to the hard disk cluster to which the current hard disk status attribute data belongs;
[0140] Based on the identification information, the hard disk status attribute data is selectively input into the corresponding hard disk personality feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the personality feature parameters corresponding to each hard disk cluster.
[0141] Furthermore, the multi-tower hard disk failure prediction model includes multiple hard disk personality feature extraction modules, each hard disk personality feature extraction module processes the hard disk status attribute data of a hard disk cluster, and each hard disk personality feature extraction module corresponds to a target weight parameter.
[0142] The heterogeneous hard disk system fault warning device of the embodiment of the present application obtains hard disk status attribute data of different models of hard disks in the heterogeneous hard disk system, and clusters and groups the hard disk status attribute data according to the distribution differences of the hard disk models and the hard disk status attribute data, determines the hard disk cluster to which the hard disk status attribute data belongs, inputs the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower structure hard disk fault prediction model for anomaly detection processing, obtains hard disk health indicator information output by the multi-tower structure hard disk fault prediction model, and performs fault warning for the heterogeneous hard disk system based on the hard disk health indicator information, which can effectively improve the accuracy and efficiency of the heterogeneous hard disk system fault warning, thereby improving the stability and security of the operation of the heterogeneous hard disk system.
[0143] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal, a device terminal or a similar computing device. Taking operation on a computer terminal as an example, FIG5 is a hardware environment diagram of a heterogeneous hard disk system fault warning method according to an embodiment of the present application. As shown in FIG5 , the computer terminal may include one or more (only one is shown in FIG5 ) first processors 502 (the first processor 502 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a first memory 504 for storing data. In an exemplary embodiment, the above-mentioned computer terminal may also include a transmission device 506 and an input and output device 508 for communication functions. It will be understood by those skilled in the art that the structure shown in FIG5 is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include more or fewer components than those shown in FIG5 , or have a different configuration with the same function as shown in FIG5 or more functions than those shown in FIG5 . The first memory 504 can be used to store computer-readable instructions, for example, software programs and modules of application software, such as computer-readable instructions corresponding to the heterogeneous hard disk system fault warning method in the embodiment of the present application. The first processor 502 executes various functional applications and data processing by running the computer-readable instructions stored in the memory 504, that is, implementing the above method. The first memory 504 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the first memory 504 may further include a memory remotely located relative to the first processor 502, and these remote memories can be connected to the computer terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 506 is used to receive or send data via a network. The specific example of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal. In one example, the transmission device 506 includes a network adapter (Network Interface Controller, referred to as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 506 may be a radio frequency (RF) module for communicating with the Internet wirelessly. In this embodiment, a method for early warning of a failure of a heterogeneous hard disk system is provided, which is applied to the above-mentioned computer terminal.
[0144] Alternatively, corresponding to the heterogeneous hard disk system failure warning method provided above, the present application also provides an electronic device. Since the embodiment of the electronic device is similar to the above method embodiment, the description is relatively simple. For relevant details, please refer to the description of the above method embodiment. The electronic device described below is only schematic. As shown in Figure 6, it is a schematic diagram of the physical structure of an electronic device disclosed in an embodiment of the present application. The electronic device may include: a processor 601, a memory 602 and a communication bus 603, wherein the processor 601 and the memory 602 communicate with each other through the communication bus 603, and communicate with the outside through the communication interface 604. The processor 601 can call the logic instructions in the memory 602 to execute a heterogeneous hard disk system fault warning method, which includes: obtaining hard disk status attribute data of different models of hard disks in the heterogeneous hard disk system, and clustering and grouping the hard disk status attribute data according to the hard disk model and the distribution difference of the hard disk status attribute data, and determining the hard disk cluster to which the hard disk status attribute data belongs; inputting the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower structure hard disk failure prediction model for anomaly detection processing, and obtaining hard disk health indicator information output by the multi-tower structure hard disk failure prediction model; wherein the hard disk failure prediction model is trained based on sample hard disk status attribute data and hard disk health label information corresponding to the sample hard disk status attribute data; and performing fault warning for the heterogeneous hard disk system based on the hard disk health indicator information.
[0145] In addition, the logic instructions in the above-mentioned memory 602 can be implemented in the form of a software function module and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a memory chip, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0146] On the other hand, an embodiment of the present application also provides a computer-readable instruction product, which includes computer-readable instructions stored on a processor-readable storage medium, and the computer-readable instructions include program instructions. When the program instructions are executed by a computer, the computer can execute the heterogeneous hard disk system fault warning method provided by the above-mentioned method embodiments. The method includes: obtaining hard disk status attribute data of different models of hard disks in a heterogeneous hard disk system, and clustering and grouping the hard disk status attribute data according to the hard disk model and the distribution difference of the hard disk status attribute data, and determining the hard disk cluster to which the hard disk status attribute data belongs; inputting the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower structure hard disk fault prediction model for anomaly detection processing, and obtaining the hard disk health indicator information output by the multi-tower structure hard disk fault prediction model; wherein the hard disk fault prediction model is obtained by training based on sample hard disk status attribute data and hard disk health label information corresponding to the sample hard disk status attribute data; and performing fault warning for the heterogeneous hard disk system based on the hard disk health indicator information.
[0147] On the other hand, the embodiments of the present application also provide one or more non-volatile computer-readable storage media storing computer-readable instructions, which, when executed by a processor, are implemented to execute the heterogeneous hard disk system fault warning method provided by the above embodiments. The method includes: obtaining hard disk status attribute data of hard disks of different models in the heterogeneous hard disk system, and clustering and grouping the hard disk status attribute data according to the hard disk model and the distribution difference of the hard disk status attribute data, and determining the hard disk cluster to which the hard disk status attribute data belongs; inputting the hard disk status attribute data corresponding to each hard disk cluster into a preset multi-tower structure hard disk fault prediction model for anomaly detection processing, and obtaining hard disk health indicator information output by the multi-tower structure hard disk fault prediction model; wherein the hard disk fault prediction model is trained based on sample hard disk status attribute data and hard disk health label information corresponding to the sample hard disk status attribute data; and performing fault warning for the heterogeneous hard disk system based on the hard disk health indicator information.
[0148] Non-volatile computer-readable storage media can be any available media or data storage devices that can be accessed by the processor, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO)), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSDs)), etc.
[0149] The device embodiments described above are merely illustrative, wherein the modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, i.e., they may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Those skilled in the art can understand and implement the present invention without inventive effort.
[0150] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A heterogeneous hard disk system failure early warning method, characterized in that: include: Obtaining hard disk status attribute data of hard disks of different models in a heterogeneous hard disk system, and clustering and grouping the hard disk status attribute data according to the hard disk models and distribution differences of the hard disk status attribute data, to determine the hard disk cluster to which the hard disk status attribute data belongs; Inputting the hard disk status attribute data corresponding to each of the hard disk clusters into a preset multi-tower hard disk fault prediction model for abnormality detection processing, and obtaining hard disk health indicator information output by the multi-tower hard disk fault prediction model; wherein the hard disk fault prediction model is trained based on sample hard disk status attribute data and hard disk health label information corresponding to the sample hard disk status attribute data; and A fault warning is performed on the heterogeneous hard disk system based on the hard disk health indicator information.
2. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: The hard disk status attribute data corresponding to each hard disk cluster is input into a preset multi-tower hard disk fault prediction model for abnormality detection processing to obtain hard disk health indicator information output by the multi-tower hard disk fault prediction model, specifically including: According to the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selected and input into the corresponding hard disk personality feature extraction module in the hard disk failure prediction model of the multi-tower structure to obtain the personality feature parameters corresponding to each hard disk cluster; Inputting the hard disk status attribute data into a hard disk common feature extraction module in the hard disk failure prediction model of the multi-tower structure to obtain common feature parameters of different hard disk clusters; Determining target attribute representations corresponding to different models of hard disks in the heterogeneous hard disk system based on the individual feature parameters and the common feature parameters; and The target attribute representation is input into the hard disk fault warning function module in the hard disk fault prediction model of the multi-tower structure to perform fault reasoning analysis, and the hard disk health indicator information output by the hard disk fault warning function module is obtained.
3. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: The clustering and grouping of the hard disk status attribute data according to the hard disk model and the distribution difference of the hard disk status attribute data to determine the hard disk cluster to which the hard disk status attribute data belongs specifically includes: According to the hard disk models in the heterogeneous hard disk system, hard disk status attribute data belonging to the same hard disk model are allocated to the same hard disk cluster; If the data volume of the hard disk status attribute data of one or more hard disk models is greater than or equal to a preset data volume threshold, the corresponding one or more first hard disk clusters are respectively split into multiple second hard disk clusters; wherein the data volume of the first hard disk cluster is greater than the data volume of the second hard disk cluster; and The hard disk status attribute data is clustered based on the distribution difference of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster, and determine the hard disk cluster to which the hard disk status attribute data belongs.
4. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: The hard disk health index information is a hard disk health score value obtained by performing anomaly detection processing on the hard disk status attribute data by the hard disk fault prediction model of the multi-tower structure; The hard disk health score value is directly proportional to the hard disk health level; The fault warning for the heterogeneous hard disk system based on the hard disk health indicator information specifically includes: comparing and analyzing the hard disk health score value with the currently selected scoring threshold, and when the hard disk health score value is less than the scoring threshold, determining that the heterogeneous hard disk system has a fault, and generating corresponding fault warning prompt information.
5. The heterogeneous hard disk system failure early warning method according to claim 4, characterized in that: After comparing and analyzing the hard disk health score value with the currently selected scoring threshold, the method further includes: if the hard disk health score value is greater than or equal to the scoring threshold, determining that the heterogeneous hard disk system is in a healthy state.
6. The heterogeneous hard disk system failure early warning method according to claim 5, characterized in that: Also includes: When the heterogeneous hard disk system is in a healthy state, the hard disk status attribute data whose hard disk health score value is greater than or equal to the scoring threshold is used as a new training sample to perform adaptive gradient training on the multi-tower hard disk failure prediction model, so as to update the parameters of each module in the multi-tower hard disk failure prediction model in real time, and obtain a new multi-tower hard disk failure prediction model, so as to use the new multi-tower hard disk failure prediction model to perform abnormality detection processing on the subsequently input hard disk status attribute data.
7. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: Before obtaining the hard disk status attribute data of hard disks of different models in the heterogeneous hard disk system, the method further includes: performing model training in an offline state to obtain a hard disk failure prediction model of the multi-tower structure; The offline model training to obtain the hard disk failure prediction model of the multi-tower structure specifically includes: Acquire sample hard disk status attribute data of the sample hard disk; wherein the sample hard disk status attribute data includes status attribute data corresponding to a healthy hard disk and status attribute data corresponding to an abnormal hard disk; Clustering and grouping the sample hard disk status attribute data according to the hard disk models of the sample hard disks and the distribution differences of the sample hard disk status attribute data, and determining the sample hard disk cluster to which the sample hard disk status attribute data belongs; and An initial multi-tower hard disk fault prediction model is trained based on the sample hard disk status attribute data corresponding to the sample hard disk cluster, and multiple hard disk anomaly detection thresholds are screened by comparing their influence on the final anomaly detection result to determine the hard disk anomaly detection threshold that meets the preset probability confidence condition, and the scoring threshold of the model parameter is updated based on the hard disk anomaly detection threshold that meets the preset probability confidence condition to obtain the finally trained multi-tower hard disk fault prediction model; wherein the scoring threshold is the hard disk anomaly detection threshold whose probability confidence is greater than or equal to the preset probability confidence.
8. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: The obtaining of hard disk status attribute data of hard disks of different models in the heterogeneous hard disk system specifically includes: obtaining original hard disk status attribute data of hard disks of different models in the heterogeneous hard disk system, and performing missing value filling on the original hard disk status attribute data to obtain first hard disk status attribute data; Performing feature screening on the first hard disk status attribute data to obtain second hard disk status attribute data; and The second hard disk status attribute data is normalized to obtain the hard disk status attribute data.
9. The heterogeneous hard disk system failure early warning method according to claim 2, characterized in that: The determining, based on the individual characteristic parameters and the common characteristic parameters, target attribute representations corresponding to hard disks of different models in the heterogeneous hard disk system specifically includes: The individual characteristic parameter and the common characteristic parameter are multiplied to obtain different models of hard disks in the heterogeneous hard disk system. The corresponding target attribute representation.
10. The heterogeneous hard disk system failure early warning method according to claim 3, characterized in that: The clustering of the hard disk status attribute data based on the distribution difference of the hard disk status attribute data to allocate the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster, and determining the hard disk cluster to which the hard disk status attribute data belongs, specifically includes: determining a metric to use for clustering; and The hard disk status attribute data is clustered based on the distribution difference of the hard disk status attribute data and the measurement criterion to assign the hard disk status attribute data to the second hard disk cluster or the fourth hard disk cluster to obtain the hard disk cluster to which the hard disk status attribute data belongs.
11. The heterogeneous hard disk system failure early warning method according to claim 2, characterized in that: According to the hard disk cluster to which the current hard disk status attribute data belongs, the hard disk status attribute data is selected and input into the corresponding hard disk personality feature extraction module in the hard disk failure prediction model of the multi-tower structure to obtain the personality feature parameters corresponding to each hard disk cluster, specifically including: Determine identification information of a corresponding hard disk personality feature extraction module according to the hard disk cluster to which the current hard disk status attribute data belongs; and Based on the identification information, the hard disk status attribute data is selected and input into the corresponding hard disk individual feature extraction module in the multi-tower structure hard disk failure prediction model to obtain the individual feature parameters corresponding to each hard disk cluster.
12. The heterogeneous hard disk system failure early warning method according to claim 2, characterized in that: The multi-tower hard disk fault prediction model includes multiple hard disk personality feature extraction modules, each hard disk personality feature extraction module processes the hard disk status attribute data of a hard disk cluster, and each hard disk personality feature extraction module corresponds to a target weight parameter.
13. A heterogeneous hard disk system failure warning device, characterized in that: include: A clustering and grouping unit, used to obtain hard disk status attribute data of hard disks of different models in a heterogeneous hard disk system, and cluster and group the hard disk status attribute data according to the hard disk models and the distribution differences of the hard disk status attribute data, and determine the hard disk cluster to which the hard disk status attribute data belongs; A fault analysis unit is used to input the hard disk status attribute data corresponding to each of the hard disk clusters into a preset multi-tower hard disk fault prediction model for abnormal detection processing, and obtain hard disk health indicator information output by the multi-tower hard disk fault prediction model; wherein the hard disk fault prediction model is trained based on sample hard disk status attribute data and hard disk health label information corresponding to the sample hard disk status attribute data; as well as A fault warning unit is used to provide fault warning for the heterogeneous hard disk system based on the hard disk health indicator information.
14. An electronic device comprising a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor, characterized in that: When the processor executes the computer-readable instructions, the steps of the heterogeneous hard disk system failure warning method as described in any one of claims 1 to 12 are implemented.
15. One or more non-volatile computer-readable storage media storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the heterogeneous hard disk system failure warning method as described in any one of claims 1 to 12.
16. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: The clustering and grouping of the hard disk status attribute data according to the hard disk model and the distribution difference of the hard disk status attribute data to determine the hard disk cluster to which the hard disk status attribute data belongs specifically includes: According to the hard disk models in the heterogeneous hard disk system, hard disk status attribute data belonging to the same hard disk model are allocated to the same hard disk cluster; If there is one or more hard disk models whose hard disk status attribute data has a data volume smaller than the data volume threshold, merging the corresponding one or more third hard disk clusters into a fourth hard disk cluster; the data volume of the fourth hard disk cluster is larger than the data volume of the third hard disk cluster; and The hard disk status attribute data are clustered based on the distribution difference of the hard disk status attribute data to allocate the hard disk status attribute data to the fourth hard disk cluster, and determine the hard disk cluster to which the hard disk status attribute data belongs.
17. The heterogeneous hard disk system failure early warning method according to claim 12, characterized in that: The algorithm formulas corresponding to the multiple hard disk personality feature extraction modules are: FFN(x)=relu(relu(xW1+b1)W2+b2); Among them, x is the hard disk status attribute data of hard disks of different models; w1 and b1 are respectively the first weight coefficient and the first deviation coefficient of the first layer network structure in the hard disk personality feature extraction module; the relu inside the brackets is the activation function of the first layer; w2 and b2 are respectively the second weight coefficient and the second deviation coefficient of the second layer network in the hard disk personality feature extraction module; the relu outside the brackets is the activation function of the second layer network.
18. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: The algorithm formula corresponding to the hard disk common feature extraction module is: FFN(x)=relu(relu(xW1+b1)W2+b2); Among them, x is the hard disk status attribute data of hard disks of different models; w1 and b1 are respectively the first weight coefficient and the first deviation coefficient of the first layer network structure in the hard disk common feature extraction module; the activation function of the first layer of relu inside the brackets; w2 and b2 are respectively the second weight coefficient and the second deviation coefficient of the second layer network in the hard disk common feature extraction module; the activation function of the second layer of relu network outside the brackets.
19. The heterogeneous hard disk system failure early warning method according to claim 2, characterized in that: The algorithm formula corresponding to the hard disk failure warning function module is: FFN(x)=sigmoid(xW1+b1); Among them, x is the target attribute representation obtained by multiplying the common representation and the unique representation of the hard disk; w1 and b1 are respectively the first weight coefficient and the first deviation coefficient of the network structure in the hard disk fault warning function module; sigmoid is the activation function of the network structure, and the output is 0 to 1, indicating whether the hard disk fails. The larger the score value, the greater the probability of failure.
20. The heterogeneous hard disk system failure early warning method according to claim 1, characterized in that: The hard disk status attribute data includes one or more of data read error rate, disk startup time, motor start and stop count, number of remapped sectors, seek error rate, seek time and hard disk power-on time.
Citation Information
Patent Citations
Method and device of dynamically diagnosing hard disk failure based on S.M.A.R.T (Self-Monitoring Analysis and Reporting Technology) data
CN105260279A
Small sample hard disk fault data generation method, storage medium and computing device
CN112434733A
Hard disk fault prediction method fusing AP clustering and width learning system
CN114116292A
Hard disk fault prediction method and system based on heterogeneous integration
CN114510382A
Hard disk fault prediction method and device, computer readable medium and electronic equipment
CN115705274A
Cited By
Hard disk fault prediction method, electronic equipment and storage medium
CN120849203A
Fault identification method and device based on variance constraint, equipment and medium
CN121350664A
ICT equipment intelligent early warning method and system
CN121486167A
Data center integration model dynamic adaptation method based on transfer learning
CN121919582A