A method for fault judgment based on computer hard disk status indicators

By periodically collecting hard disk status indicators, dimensionality reduction and failure probability calculations, and combining the threshold fine-tuning model to dynamically adjust the failure probability threshold, the limitations of existing hard disk failure prediction methods in data acquisition and feature processing are solved, and more accurate and timely fault judgment and early warning are achieved.

CN119166399BActive Publication Date: 2025-05-23SHENZHEN YIXUN COMPUTER CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411189456.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-28
Publication Date
2025-05-23
Estimated Expiration
2044-08-28

AI Technical Summary

Technical Problem

The existing software-based hard disk failure prediction methods have limitations in the comprehensiveness and accuracy of data acquisition, feature dimensionality reduction and selection, and the selection and optimization of fault classification models, which leads to the inability to respond in a timely manner under high load or abnormal situations, increasing the risk of data loss and system crash.

Method used

By setting a periodic data acquisition mechanism, multi-dimensional characteristic data including disk rotation instability, data transmission rate, and read and write error rate are obtained, and the data is normalized and denoised. Then, data dimensionality reduction is achieved using an autoencoder, key feature vectors are extracted, and fault probability calculation is performed through the support vector machine. Build a threshold fine-tuning model to dynamically adjust the failure probability threshold based on the correlation and change trends of the physical state and operating state collected in real time.

Benefits of technology

It improves the accuracy and timeliness of fault judgment, reduces the chance of false alarms, and ensures the security of data storage and the normal operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119166399B_ABST
    Figure CN119166399B_ABST
Patent Text Reader

Abstract

The present invention provides a method for fault judgment based on computer hard disk status indicators, and relates to the field of computer technology. The present invention sets a periodic data collection mechanism to obtain multi-dimensional feature data including disk rotation instability, data transmission rate, read and write error rate, etc., and performs normalization and denoising on these data; then, an autoencoder is used to achieve data dimension reduction, extract key feature vectors, and then calculate the fault probability through a support vector machine; on this basis, a threshold fine-tuning model is constructed, which can dynamically adjust the fault probability threshold according to the correlation between the physical state and the operating state collected in real time, as well as their changing trends; not only the accuracy and timeliness of fault judgment are improved, but also the probability of false alarms is reduced through the dynamic adjustment mechanism, thereby ensuring the security of data storage and the normal operation of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method for fault judgment based on computer hard disk status indicators. Background Art

[0002] A hard disk is a data storage device consisting of one or more hard circular platters (called "disks" or "platters") covered with magnetic materials that can store data. With the rapid development of information technology, computer hard disks are the core components of data storage, and their reliability and stability are crucial to the performance of the entire computer system. Traditional hard disk fault detection methods mainly rely on hardware-level monitoring and fault code analysis. Although these methods can identify physical failures of hard disks to a certain extent, they have obvious limitations in predicting early failures and potential operating problems of hard disks. In recent years, with the advancement of machine learning and data analysis technology, software-based hard disk fault prediction methods have gradually become a research hotspot. These methods collect hard disk operating status indicators and physical status indicators, and use complex algorithm models to perform data analysis to achieve early warning and prediction of hard disk failures.

[0003] In the prior art, the publication number is CN114758714A, and the name is a hard disk fault prediction method, device, electronic device and storage medium. The method includes: obtaining the working status data of the hard disk at the current moment; using a fault prediction model to process the working status data to obtain the fault prediction result of the hard disk at a preset time in the future; wherein the fault prediction model is trained using a machine learning model based on sample working status data.

[0004] The publication number is CN111611117B, and the name is a hard disk failure prediction method, device, equipment and computer-readable storage medium. When establishing a hard disk failure prediction model for multiple hard disk models, a conversion relationship between various parameters of each hard disk model and the corresponding parameters of a benchmark hard disk model is first established, and then the parameter detection values ​​of the hard disk are converted according to the conversion relationship, thereby eliminating the differences between different hard disk models; the hard disk failure prediction model is trained using the converted parameter detection values ​​and the operating status of the hard disk, thereby establishing a hard disk failure prediction model suitable for multiple hard disk models, which saves time and effort compared to training a hard disk failure prediction model for each hard disk model separately. When using the hard disk failure prediction model to predict hard disk failures, since an association is established between the parameters of each hard disk model and the benchmark hard disk model, a more accurate prediction result can be obtained compared to the prediction model in the prior art that only distinguishes different hard disk failures by model.

[0005] Article No.: 1627-0385 (2005) 02-0035-04 "Discussion on Common Hard Disk Fault Diagnosis, Treatment Steps and Methods" describes the types of computer hard disk failures in the prior art:

[0006] However, existing software-based hard disk failure prediction methods still face some challenges in practical applications. First, the comprehensiveness and accuracy of data collection are key factors affecting the prediction results. Existing methods often only focus on a few indicators and ignore other parameters that may have an important impact on the health status of the hard disk.

[0007] Secondly, the process of feature dimensionality reduction and selection lacks systematicity and pertinence, resulting in the extracted feature vectors not being able to fully reflect the actual status of the hard disk. In addition, the selection and optimization of the fault classification model is also a difficult point. Different models perform quite differently on different data sets, and the generalization ability of the model needs to be improved. Most of the current fault probability calculation models are based on static initial fault probability thresholds, which often rely on empirical values ​​and fail to make dynamic adjustments based on real-time data of the hard disk status. As a result, they are unable to respond in a timely manner under high load or abnormal conditions, thereby increasing the risk of data loss and system crashes.

[0008] The above information disclosed in the above Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to one of ordinary skill in the art. Summary of the invention

[0009] The purpose of the present invention is to provide a method for fault diagnosis based on computer hard disk status indicators to solve the problems raised in the above background technology.

[0010] To achieve the above object, the present invention provides the following technical solutions:

[0011] A method for fault diagnosis based on computer hard disk status indicators, the specific steps include:

[0012] Step S1: Set the collection period of the hard disk to the set {1, 2, ..., n}, where i∈{1, 2, ..., n} represents the index of the i-th data collection in the collection period, and n represents the index of the current n-th data collection, collect the physical state indicators and operation state indicators of the hard disk, where the physical state indicators include the disk rotation instability data and the number of head loading times, and the operation state indicators include the data transmission rate and the read and write error rate, and perform normalization and denoising preprocessing on the collected data to obtain multi-dimensional feature data;

[0013] Step S2: receiving multi-dimensional feature data collected n times, using an autoencoder to reduce the dimension of the multi-dimensional features, and extracting the key feature vector after the dimension reduction;

[0014] Step S3: receiving the key feature vector after dimension reduction, and using the support vector machine to calculate the failure probability of the key feature vector to achieve binary classification of hard disk failures;

[0015] Set the initial failure probability threshold of the hard disk failure, and set the hard disk failure warning trigger condition according to the initial failure probability threshold;

[0016] Step S4: obtaining disk rotation instability data, data transmission rate and read / write error rate, and performing correlation analysis on the disk rotation instability data and the data transmission rate to obtain a first correlation evaluation coefficient, which is used to evaluate the correlation influence degree between the disk rotation instability data and the data transmission rate;

[0017] Performing correlation analysis on the disk rotation instability data and the read / write error rate to obtain a second correlation evaluation coefficient, the second correlation evaluation coefficient is used to evaluate the correlation influence degree between the disk rotation instability data and the read / write error rate;

[0018] Step S5: calculating the change trend data of the physical state indicator in the current collection period, and analyzing and processing the change trend data to generate a first trend evaluation coefficient, which is used to evaluate the change trend of the physical state indicator in the current collection period;

[0019] Calculate the change trend data of the operating status indicator in the current collection cycle, and analyze and process the change trend data to generate a second trend evaluation coefficient, which is used to evaluate the change trend of the operating status indicator in the current collection cycle;

[0020] Step S6: constructing a threshold fine-tuning model by combining the first correlation evaluation coefficient, the second correlation evaluation coefficient, the first trend evaluation coefficient, and the second trend evaluation coefficient, wherein the threshold fine-tuning model is used to provide a fine-tuning strategy for the initial fault probability threshold;

[0021] Step S7: Obtain the fault probability threshold adjusted by the fine-tuning strategy, and adjust the fault warning trigger condition according to the adjusted fault probability threshold, and further calculate the failure probability of the hard disk in the current nth data collection. If the failure probability exceeds the adjusted failure probability threshold, a fault warning is triggered.

[0022] Furthermore, the acquisition of multi-dimensional feature data includes:

[0023] The disk rotation instability data includes the disk rotation speed fluctuation rate and the disk vibration amplitude, and the disk rotation speed fluctuation rate and the disk vibration amplitude are marked as CVb and CZf respectively;

[0024] Combine the disk rotation speed fluctuation rate and disk vibration amplitude, and perform analysis and processing to construct the disk rotation instability value R of the i-th data collection. i , the calculation formula is as follows:

[0025]

[0026] Parameter explanation: R i is the disk rotation instability value of the ith data collection, CVb i is the disk rotation speed fluctuation rate of the ith data collection, CZf i is the disk vibration amplitude of the ith data collection, a1, a2, and a3 are weight coefficients used to adjust the effects of rotation speed fluctuation rate and vibration amplitude on disk rotation instability;

[0027] The physical state index data and the operating state index data are acquired regularly within the set acquisition cycle, and the acquired data are recorded in the database to form a data set D = {(R i ,L i ,T i ,E i )|i∈{1,2,…,n}};

[0028] Among them, R i ,L i ,T i ,E i They represent the disk rotation instability value, head loading times, data transmission rate and read / write error rate of the i-th data acquisition respectively;

[0029] For normalization, the min-max normalization method is used to normalize each index value x to x';

[0030] The normalized data range is (0,1), where

[0031] For denoising, the moving average method is used to remove random noise in the data to smooth the normalized data of each indicator:

[0032] Get multidimensional feature data: After normalization and denoising, the final multidimensional feature data set is expressed as F = {(R′ i ,L′ i ,T′ i ,E′ i )|i∈{k-1,k,…,n}}, where R′ i ,L′ i ,T′ i ,E′ iThey are the physical state index and operation state index after data preprocessing, and k-1 represents the starting point of the acquisition times after denoising.

[0033] Furthermore, the multi-dimensional feature data collected n times is received, and the multi-dimensional features are reduced in dimension using an autoencoder, and the key feature vectors after the dimension reduction are extracted, including:

[0034] The autoencoder is selected as the dimensionality reduction tool; the autoencoder consists of two parts: the encoder and the decoder, where the encoder converts the high-dimensional input data F i Compressed into a low-dimensional feature vector Z i , the decoder then converts Z i Restore to high-dimensional space;

[0035] For each data collection point i, the input multidimensional feature data F i It is expressed as:

[0036] F i = {R′ i ,L′ i ,T′ i ,E′ i}

[0037] The output of the encoder network is a low-dimensional feature vector Z i :

[0038] Z i =f θ (F i )=σ1(W 1 F i +b1)

[0039] Among them, W 1 is the weight matrix of the encoder, b1 is the bias vector, σ1 is the activation function, and θ represents the set of all parameters of the encoder;

[0040] The autoencoder is trained by minimizing the reconstruction error so that the decoder outputs the reconstructed data Approaching the original input data F i ;

[0041] After the autoencoder training is completed, the low-dimensional feature vector Z output by the encoder is directly used i As the key feature vector after dimensionality reduction;

[0042] The feature vector after dimensionality reduction is expressed as:

[0043] Z i ={z i1 ,z i2 ,…,z im}

[0044] Among them, m is the dimension of the feature vector after dimensionality reduction.

[0045] Furthermore, the support vector machine is used to calculate the fault probability of the key feature vector. If the current fault probability exceeds the threshold, a fault warning is triggered, including:

[0046] For the current key feature vector Z i Classify the fault into two categories:

[0047] Using the key eigenvector Z i Construct an SVM classifier; divide the key feature vector into two categories, corresponding to the normal state and fault state of the hard disk, and obtain the known training data set {(Z i ,y i )}, where y i It is a binary classification label. The binary classification label is: 1 for normal state and -1 for fault state;

[0048] After training is completed, the decision function of SVM is defined as:

[0049] f(Z i )=sign(w·Z i +b2)

[0050] Where sign(·) is a sign function. When the input is greater than 0, the output is +1, indicating "normal". When the input is less than or equal to 0, the output is -1, indicating "fault".

[0051] Set the initial failure probability threshold to P fault , the probability is estimated using the following logistic regression model:

[0052]

[0053] Among them, c1 is the parameter used to adjust the probability curve, which is obtained through cross-validation of the model; P fault The value range is (0,1);

[0054] Set and calculate the failure probability of the hard disk in the current nth data collection as P th,n , when P th,n ≥P fault , the hard disk is judged to be in a faulty state, otherwise it is judged to be in a normal state.

[0055] Furthermore, the first correlation evaluation coefficient and the second correlation evaluation coefficient are constructed as follows:

[0056] The first Pearson correlation coefficient between the disk rotation instability data and the data transfer rate is calculated as:

[0057]

[0058] Among them, ρ 1 is the first Pearson correlation coefficient between disk rotation instability and data transfer rate;

[0059] and R′ i and T′ i The mean in the set {1,2,…,n};

[0060] Define the first correlation evaluation coefficient as C RT , the formula is as follows:

[0061] C RT =|ρ 1 |·d1

[0062] Among them, |ρ 1 | is the absolute value of the calculated first Pearson correlation coefficient, indicating the strength of the correlation;

[0063] d1 is the adjustment factor, which is used to adjust the degree of correlation under different hard disk types or workloads;

[0064] First Pearson correlation coefficient ρ 1 The absolute value range of C is between 0 and 1. RT The value range is also between 0 and 1:

[0065] When C RT The closer it is to 1, the stronger the correlation between disk rotation instability and data transfer rate is, which means that disk rotation instability has a greater impact on data transfer rate and is a key factor that leads to a decrease in data transfer efficiency.

[0066] When C RT When it is closer to 0, the correlation between the two is weaker, the impact of disk rotation instability on data transmission rate is smaller, and the probability of failure is smaller;

[0067] Setting C RT The evaluation threshold is C th ; 0.35≤C th ≤0.75, C RT With C th The size judgment between them is used to distinguish between normal state and fault state;

[0068] The second Pearson correlation coefficient between the disk rotation instability data and the read and write error rate is calculated using the formula:

[0069]

[0070] Among them, ρ 2is the second Pearson correlation coefficient between disk rotation instability and read and write error rate;

[0071] and R′ i and E′ i The mean in the set {1,2,…,n};

[0072] Define the second correlation evaluation coefficient as C RE , the second correlation evaluation coefficient C RE The calculation method of is the same as the first correlation evaluation coefficient, and the specific formula is as follows:

[0073] C RE =|ρ 2 |·d2

[0074] Among them, |ρ 2 | is the absolute value of the calculated second Pearson correlation coefficient, indicating the strength of the correlation;

[0075] d2 is the adjustment factor, which is used to adjust the degree of correlation under different hard disk types or workloads;

[0076] C RE The value range is also between 0 and 1;

[0077] When C RE The closer it is to 1, the stronger the correlation between disk rotation instability and read / write error rate is, which means that disk rotation instability has a greater impact on read / write error rate and is the key factor leading to an increase in read / write error rate.

[0078] When C RE The closer it is to 0, the weaker the correlation between the two, the smaller the impact of disk rotation instability on the read and write error rate, and the smaller the probability of failure;

[0079] Setting C RE The evaluation threshold is C Eh ; 0.35≤C Eh ≤0.75, C RE With C Eh The size judgment between them is used to distinguish between normal state and fault state.

[0080] Furthermore, the first trend evaluation coefficient and the second trend evaluation coefficient are constructed as follows:

[0081] Calculate the average trend of disk rotation instability data:

[0082]

[0083] Among them, T RIndicates the average change trend of disk rotation instability; ΔR i,i+1 represents the change in disk rotation instability between the i-th and i+1-th data collections;

[0084] Calculate the average trend of head loading times:

[0085]

[0086] Among them, T L Indicates the average change trend of the number of head loading times; ΔL i,i+1 It represents the change in the number of head loading times between the i-th and i+1-th data acquisitions;

[0087] The following first trend evaluation coefficients are calculated:

[0088]

[0089] Among them, C T is the first trend evaluation coefficient, 0<C T <1, e1 and e2 are the weight coefficients of the corresponding parameters;

[0090] When C T As e1·T approaches 1, R +e2·T L The smaller the output value, the smaller the change trend of the physical status indicator in the current collection cycle;

[0091] When C T As e1·T approaches 0, R +e2·T L The larger the output value, the greater the change trend of the physical status indicator in the current collection cycle;

[0092] Calculate the average change trend of data transmission rate:

[0093]

[0094] Among them, T S Indicates the average change trend of data transmission rate; ΔT i,i+1 It represents the change in data transmission rate between the i-th and i+1-th data collection;

[0095] Calculate the average change trend of read and write error rates:

[0096]

[0097] Among them, T C Indicates the average change trend of the read and write error rate; ΔE i,i+1It represents the change in read / write error rate between the i-th and i+1-th data collection;

[0098] The following second trend evaluation coefficients are calculated:

[0099]

[0100] Among them, C U is the second trend evaluation coefficient, 0<C U <1, e2 and e3 are weight coefficients of corresponding parameters respectively;

[0101] When C U The closer it is to 1, The smaller the output value, the smaller the change trend of the operating status indicator in the current collection cycle;

[0102] When C U The closer it is to 0, The larger the output value, the greater the change trend of the operating status indicator in the current collection cycle.

[0103] Furthermore, a threshold fine-tuning model is constructed. The threshold fine-tuning model is used to provide a fine-tuning strategy for the initial failure probability threshold, specifically including:

[0104] The calculation formula for defining the threshold fine-tuning model is as follows:

[0105]

[0106] Among them, WT1 is a first comprehensive index combining the first correlation evaluation coefficient and the second correlation evaluation coefficient, which overall reflects the correlation degree of the computer hard disk status; WT2 is a second comprehensive index combining the first trend evaluation coefficient and the second trend evaluation coefficient, which overall reflects the trend degree of the computer hard disk status, P fault is the initial failure probability threshold, P 1 ′ is to reduce P fault The fault probability threshold after taking the value; P′ 2 To improve P fault The fault probability threshold mark after taking the value;

[0107] r1, r2, r3, and r4 are the regression coefficients of the corresponding parameters, μ RT , Respectively represent the first correlation evaluation coefficient C RT The mean and standard deviation of are used for normalization; μ RE , Represent the second correlation evaluation coefficient C RE The mean and standard deviation of are used for normalization; η1, η2, η3, and η4 are all positive numbers;

[0108] The division thresholds of the first comprehensive index and the second comprehensive index are respectively set to Q1 and Q2;

[0109] When WT1 ≥ Q1, it means that the correlation between the disk's rotational instability and the data transfer rate is significant; this means that the hard disk is in poor condition, there is a high risk of failure, and the data transfer efficiency is seriously affected;

[0110] When WT1<Q1, it means that the correlation between the rotational instability of the disk and the data transmission rate is weak, the system performs normally, and the failure risk is low;

[0111] When WT2 ≥ Q2, it means that the change trend of the physical status indicator is obvious, indicating that the operation status of the hard disk has fluctuated greatly during the current acquisition cycle. This is caused by changes in the external environment or internal faults of the hard disk.

[0112] When WT2<Q2, it means that the change trend of the physical status indicator is small, which means that the operation status of the hard disk in the current acquisition cycle is relatively stable, the failure risk is low, and the operation can be carried out normally.

[0113] Further, the fine-tuning strategy is as follows:

[0114] When WT1≥Q1 and WT2≥Q2, use P′ 2 Fine-tuning strategy; at this time, the impact of disk rotation instability on data transmission rate and read / write error rate exceeds 75%, and the volatility of physical state and operation state also exceeds 75%; in this case, it indicates that the system failure risk is extremely high, and the initial failure probability threshold needs to be increased, limiting P′ 2 The increase is P fault to within 10% to 20% of the target, ensuring early warning in high-risk situations;

[0115] When WT1≥Q1 and WT2<Q2, use P′ 2 Fine-tuning strategy; at this time, the impact of disk rotation instability on data transmission rate and read-write error rate exceeds 75%, but the volatility of physical state and operating state is less than 25%; although the operating state is relatively stable, the initial failure probability threshold needs to be increased due to the strong correlation between data transmission and read-write error rate; limit P′ 2 The increase is P fault within 10% of the previous value, increasing the system's sensitivity to key indicators;

[0116] When WT1<Q1 and WT2≥Q2, use P′ 2Fine-tuning strategy; at this time, the impact of disk rotation instability on data transfer rate and read / write error rate is less than 25%, but the volatility of physical state and operating state exceeds 75%; in this case, although the data transfer rate shows a low risk of failure, due to the large fluctuations in physical and operating states, the initial failure probability threshold needs to be increased. After increasing the initial failure probability threshold within 15%, the system can better cope with the potential failure risk caused by operating state fluctuations;

[0117] When WT1<Q1 and WT2<Q2, use P 1 ' fine-tuning strategy; at this time, the impact of disk rotation instability on data transmission rate and read / write error rate is less than 25%, and the volatility of physical state and operating state is also less than 25%; in this case, the overall failure risk is low, the initial failure probability threshold can be reduced, the system avoids excessive sensitivity, and the probability of false alarm is reduced;

[0118] According to the failure probability P of the hard disk in the current nth data collection th,n , when P th,n conform to If any of the following conditions are met, the hard disk is judged to be in a faulty state; otherwise, it is judged to be in a normal state.

[0119] Compared with the prior art, the beneficial effects of the present invention are as follows: by setting a periodic data collection mechanism, multi-dimensional feature data including disk rotation instability, data transmission rate, and read / write error rate are obtained, and these data are normalized and denoised; then, an autoencoder is used to achieve data dimensionality reduction, extract key feature vectors, and then the fault probability is calculated through a support vector machine (SVM); on this basis, a threshold fine-tuning model is constructed, which can dynamically adjust the fault probability threshold according to the correlation between the physical state and the operating state collected in real time, as well as their changing trends; not only the accuracy and timeliness of fault judgment are improved, but also the probability of false alarms is reduced through the dynamic adjustment mechanism, thereby ensuring the security of data storage and the normal operation of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0120] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.

[0121] Figure 1 It is a schematic diagram of the overall method flow of the present invention. DETAILED DESCRIPTION

[0122] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0123] Embodiment 1:

[0124] See also Figure 1 , the present invention provides a technical solution:

[0125] A method for fault diagnosis based on computer hard disk status indicators, the specific steps include:

[0126] Step S1: data collection and preprocessing, setting the collection period of the hard disk to the set {1,2,…,n}, where i∈{1,2,…,n} represents the index of the i-th data collection in the collection period, and n represents the index of the current n-th data collection, collecting the physical state indicators and operation state indicators of the hard disk, where the physical state indicators include the disk rotation instability data and the number of head loading times, and the operation state indicators include the data transmission rate and the read-write error rate, and performing normalization and denoising preprocessing on the collected data to obtain multi-dimensional feature data;

[0127] Step S2: feature dimensionality reduction, receiving multi-dimensional feature data collected n times, using an autoencoder to reduce the dimensionality of the multi-dimensional features, and extracting the key feature vectors after dimensionality reduction, so as to reduce the computational complexity and retain important information;

[0128] Step S3: Fault classification and early warning, receiving the key feature vector after dimensionality reduction, using support vector machine to calculate the fault probability of the key feature vector, and realizing the binary classification of hard disk faults;

[0129] According to historical data and experimental data analysis by the expert group, the initial failure probability threshold of the hard disk failure is set, and according to the initial failure probability threshold, the hard disk failure warning trigger condition is set;

[0130] Step S4: constructing a correlation evaluation coefficient, obtaining disk rotation instability data, data transmission rate and read / write error rate, and performing correlation analysis on the disk rotation instability data and the data transmission rate to obtain a first correlation evaluation coefficient, which is used to evaluate the correlation influence degree between the disk rotation instability data and the data transmission rate;

[0131] Performing correlation analysis on the disk rotation instability data and the read / write error rate to obtain a second correlation evaluation coefficient, the second correlation evaluation coefficient is used to evaluate the correlation influence degree between the disk rotation instability data and the read / write error rate;

[0132] Step S5: constructing an evaluation coefficient, calculating the change trend data of the physical state indicator in the current collection period, and analyzing and processing the change trend data to generate a first trend evaluation coefficient, which is used to evaluate the change trend of the physical state indicator in the current collection period;

[0133] Calculate the change trend data of the operating status indicator in the current collection cycle, and analyze and process the change trend data to generate a second trend evaluation coefficient, which is used to evaluate the change trend of the operating status indicator in the current collection cycle;

[0134] Step S6: constructing a comprehensive fine-tuning index, combining the first correlation evaluation coefficient, the second correlation evaluation coefficient, the first trend evaluation coefficient and the second trend evaluation coefficient to construct a threshold fine-tuning model, the threshold fine-tuning model is used to provide a fine-tuning strategy for the initial fault probability threshold;

[0135] Step S7: Obtain the fault probability threshold adjusted by the fine-tuning strategy, and adjust the fault warning trigger condition according to the adjusted fault probability threshold, and further calculate the failure probability of the hard disk in the current nth data collection. If the failure probability exceeds the adjusted failure probability threshold, a fault warning is triggered.

[0136] To further explain, the acquisition of multi-dimensional feature data includes: In data collection, focus on the following two types of indicators:

[0137] Physical status indicators:

[0138] Disk rotation instability data: indicates the stability of disk rotation, obtained through high-precision sensors;

[0139] The disk rotation instability data includes the disk rotation speed fluctuation rate and the disk vibration amplitude, and the disk rotation speed fluctuation rate and the disk vibration amplitude are marked as CVb and CZf respectively;

[0140] Combine the disk rotation speed fluctuation rate and disk vibration amplitude, and perform analysis and processing to construct the disk rotation instability value R of the i-th data collection. i , the calculation formula is as follows:

[0141]

[0142] Parameter explanation: R i is the disk rotation instability value of the ith data collection, CVb i is the disk rotation speed fluctuation rate of the ith data collection, which is obtained by processing the disk rotation speed data through variance or standard deviation, and is used to reflect the fluctuation of the disk rotation speed; CZf iis the disk vibration amplitude of the ith data collection, which is collected by the vibration sensor and represents the vibration intensity of the disk during operation; a1, a2, and a3 are weight coefficients used to adjust the influence of the rotation speed fluctuation rate and the vibration amplitude on the disk rotation instability; the values ​​of a1, a2, and a3 are obtained by fitting historical data, or determined by an expert group through experimental data to ensure appropriate weight distribution;

[0143] With CVb i As the value of increases, the exponential function grows rapidly, which reflects that the rotational fluctuation has a significant amplification effect on the instability. At the same time, the weight coefficients a1 and a2 are used to adjust the weight of the impact of the fluctuation rate on the final result.

[0144] In the form of a fraction, it is ensured that when the vibration amplitude is small, if CZf i ≈0, this term has an effect on R i The influence of is weak; as the vibration amplitude increases, the value of the fraction approaches a3, reflecting the gradual importance of vibration to instability; in addition, the constant 1 in the denominator ensures that the formula will not be singular when the vibration amplitude approaches 0;

[0145] Disk rotation speed fluctuation rate (RPMVariance):

[0146] Definition: Indicates the rate of change of disk rotation speed per unit time, expressed in the form of standard deviation of revolutions per minute (RPM);

[0147] Collection method: Real-time collection through high-precision sensors or hard disk built-in self-monitoring systems (such as SMART);

[0148] The fluctuation of disk rotation speed directly reflects the disk rotation instability and can be quantified as the volatility, which is closely related to the physical state of the disk;

[0149] Disk vibration amplitude (VibrationAmplitude):

[0150] Definition: The amplitude of mechanical vibration generated by the disk during operation, quantified in micrometers (μm) or acceleration (g);

[0151] Collection method: Measure the vibration of the disk during operation through a built-in or external vibration sensor;

[0152] Vibration is one of the direct causes of disk rotation instability, so the vibration amplitude is an important correlation data;

[0153] Head loading times (L): indicates the number of times the head is loaded during the reading and writing process, in times, counted by the hard disk controller;

[0154] Operation status indicators:

[0155] Data transfer rate (T): indicates the amount of data transferred per unit time, in MB / s, and is obtained through the hard disk performance monitoring tool;

[0156] Read / write error rate (E): indicates the number of read / write errors that occur per unit time, in times / hour, obtained through the hard disk self-monitoring system (SMART);

[0157] Data collection: By writing scripts or using hardware monitoring tools, physical status indicator data and operating status indicator data are regularly obtained within the set collection period, and the collected data are recorded in the database to form a data set D = {(R i ,L i ,T i ,E i )|i∈{1,2,…,n}};

[0158] Among them, R i ,L i ,T i ,E i They represent the disk rotation instability value, head loading times, data transmission rate and read / write error rate of the i-th data acquisition respectively;

[0159] Data preprocessing:

[0160] For normalization, in order to eliminate the dimensional influence of different index values, the min-max normalization method is used to normalize each index value x to x':

[0161]

[0162] Among them, x min and x max are the minimum and maximum values ​​of the indicator in the data set respectively; the normalized data range is (0,1), where

[0163] For denoising, the moving average method is used to remove random noise in the data to smooth the normalized data of each indicator:

[0164] Get multidimensional feature data: After normalization and denoising, the final multidimensional feature data set is expressed as F = {(R′ i ,L′ i ,T′ i ,E′ i )|i∈{k-1,k,…,n}}, where R′ i ,L′ i ,T′ i ,E′i They are the physical state index and operation state index after data preprocessing, and k-1 represents the starting point of the acquisition times after denoising.

[0165] Further explanation: receiving multi-dimensional feature data collected n times, using an autoencoder to reduce the dimension of the multi-dimensional features, and extracting the key feature vector after the dimension reduction, including:

[0166] Select Autoencoder as the dimensionality reduction tool; Autoencoder is an unsupervised neural network that can learn a low-dimensional representation of data while retaining as much original information as possible; the specific operations are as follows:

[0167] Construct an autoencoder network: The autoencoder consists of two parts: an encoder and a decoder, where the encoder converts the high-dimensional input data F i Compressed into a low-dimensional feature vector Z i , the decoder then converts Z i Restore to high-dimensional space;

[0168] For each data collection point i, the input multidimensional feature data F i It is expressed as:

[0169] F i = {R′ i ,L′ i ,T′ i ,E′ i}

[0170] The output of the encoder network is a low-dimensional feature vector Z i :

[0171] Z i =f θ (F i )=σ1(W 1 F i +b1)

[0172] Among them, W 1 is the weight matrix of the encoder, b1 is the bias vector, σ1 is the activation function (ReLU or Sigmoid function is used in this embodiment), and θ represents the set of all parameters of the encoder;

[0173] Training the autoencoder: The autoencoder is trained by minimizing the reconstruction error so that the decoder outputs the reconstructed data Approaching the original input data F i ; The reconstruction error is expressed as:

[0174]

[0175] Among them, g φ (Zi ) is the output of the decoder, φ represents the parameter set of the decoder;

[0176] Extract key feature vectors: After the autoencoder training is completed, directly use the low-dimensional feature vector Z output by the encoder part i As the key feature vector after dimensionality reduction;

[0177] At this time, Z i The dimension is much lower than the original F i , but it still retains the main information in the original data and eliminates redundant features;

[0178] The feature vector after dimensionality reduction is expressed as:

[0179] Z i ={z i1 ,z i2 ,…,z im}

[0180] Among them, m is the dimension of the feature vector after dimensionality reduction, m<<4; that is, the dimension after dimensionality reduction is much smaller than the original dimension;

[0181] Determine the validity of the dimensionality reduction results:

[0182] After dimensionality reduction, the extracted key feature vector Z i Evaluate to ensure that it effectively reduces the dimensionality of the data while maintaining the integrity of the information; validated by:

[0183] Reconstruction accuracy verification: Calculate the reconstructed With the original input F i The mean square error (MSE) between them is used to evaluate the effectiveness of dimensionality reduction; if the reconstruction error is small, it means that the eigenvector Z after dimensionality reduction is i Still retains most of the information of the original data;

[0184] Subsequent analysis: Z i The data is input into the subsequent fault judgment model (logistic regression, support vector machine, etc. in this embodiment) and compared with the original data without dimensionality reduction. If the performance of the data after dimensionality reduction in fault judgment is better than or close to the original data, and the computational complexity is significantly reduced, the dimensionality reduction effect is significant.

[0185] Further explanation: Support vector machine is used to calculate the fault probability of key feature vectors. If the current fault probability exceeds the threshold, a fault warning is triggered, including:

[0186] For the current key feature vector Z i Perform the second classification of faults. The specific operations are as follows:

[0187] Building a classification model: Using the key feature vector Z i Construct an SVM classifier; the goal of SVM is to find an optimal hyperplane to divide the key feature vectors into two categories, corresponding to the normal state and the fault state of the hard disk, respectively. The objective function of the model in this embodiment is expressed as:

[0188]

[0189] Where w is the normal vector of the hyperplane, b2 is the bias term, and ξ i is a slack variable used to deal with inseparable data, and C is a penalty parameter used to balance the trade-off between classification interval and classification error;

[0190] Training classification model: Get a known training data set {(Z i ,y i )}, where y i It is a binary classification label. The binary classification label is: 1 for normal state and -1 for fault state;

[0191] Use training data to determine model parameters w and b2; the training process optimizes the model by maximizing the classification interval and minimizing the classification error, so that the classifier can accurately classify the key feature vectors into the correct category;

[0192] Classification decision function: After training, the decision function of SVM is defined as:

[0193] f(Z i )=sign(w·Z i +b2)

[0194] Where sign(·) is a sign function. When the input is greater than 0, the output is +1, indicating "normal". When the input is less than or equal to 0, the output is -1, indicating "fault".

[0195] Set the initial failure probability threshold to P fault , failure probability P fault The calculation of is achieved by mapping the decision values ​​of the SVM to probabilities, and the probability is estimated using the following logistic regression model:

[0196]

[0197] Among them, c1 is the parameter used to adjust the probability curve, which is obtained through cross-validation of the model; the initial failure probability threshold P fault Indicates the possibility of hard disk failure, P fault The value range is (0,1);

[0198] According to the initial failure probability threshold P fault Based on the calculation results, the following hard disk failure judgment is made:

[0199] The early warning triggering condition is: Similarly, according to mapping the decision value of SVM to probability, the failure probability of the hard disk in the current nth data collection is designed and calculated as P th,n , when P th,n ≥P fault When the hard disk is in a fault state, it is judged to be in a normal state otherwise;

[0200] According to the classification results, a fault alarm signal or a normal operation signal is output to prompt the user of the current hard disk status.

[0201] Further explanation: the first correlation evaluation coefficient and the second correlation evaluation coefficient are constructed as follows:

[0202] The first Pearson correlation coefficient between the disk rotation instability data and the data transfer rate is calculated as:

[0203]

[0204] Among them, ρ 1 is the first Pearson correlation coefficient between disk rotation instability and data transfer rate;

[0205] and R′ i and T′ i The mean value in the set {1, 2, ..., n} is calculated in the same way as the conventional mean value calculation method, which will not be described in detail.

[0206] Define the first correlation evaluation coefficient as C RT , the formula is as follows:

[0207] C RT =|ρ 1 |·d1

[0208] Among them, |ρ 1 | is the absolute value of the calculated first Pearson correlation coefficient, indicating the strength of the correlation;

[0209] d1 is an adjustment factor used to adjust the degree of association under different hard disk types or workloads; d1 is determined by an expert group through experimental data and specific application scenarios. In this embodiment, 0.12≤d1≤1;

[0210] First Pearson correlation coefficient ρ 1 The absolute value range of C is between 0 and 1. RT The value range is also between 0 and 1:

[0211] When C RTThe closer it is to 1, the stronger the correlation between disk rotation instability and data transfer rate is, which means that disk rotation instability has a greater impact on data transfer rate and is a key factor that leads to a decrease in data transfer efficiency.

[0212] When C RT When it is closer to 0, the correlation between the two is weaker, the impact of disk rotation instability on data transmission rate is smaller, and the probability of failure is smaller;

[0213] Setting C RT The evaluation threshold is C th ; 0.35≤C th ≤0.75, C th Determined through historical data analysis and practical application experience, C RT With C th The size judgment between them is used to distinguish between normal state and fault state;

[0214] High risk indication: When C RT ≥C th When the disk is not stable, it means that the instability of disk rotation has a significant negative impact on the transfer rate, indicating that the hard disk is in or approaching a failure state; in this case, the hard disk should be inspected in more detail or preventive maintenance measures should be taken directly;

[0215] Low risk indication: When C RT <C th When , it indicates that the impact of disk rotation instability on data transfer rate is within an acceptable range, the hard disk is relatively stable, and the failure risk is within 20%;

[0216] The second Pearson correlation coefficient between the disk rotation instability data and the read and write error rate is calculated using the formula:

[0217]

[0218] Among them, ρ 2 is the second Pearson correlation coefficient between disk rotation instability and read and write error rate;

[0219] and R′ i and E′ i The mean in the set {1,2,…,n};

[0220] Define the second correlation evaluation coefficient as C RE , the second correlation evaluation coefficient C RE The calculation method of is the same as the first correlation evaluation coefficient, and the specific formula is as follows:

[0221] C RE=|ρ 2 |·d2

[0222] Among them, |ρ 2 | is the absolute value of the calculated second Pearson correlation coefficient, indicating the strength of the correlation;

[0223] d2 is an adjustment factor used to adjust the degree of association under different hard disk types or workloads; d2 is determined by an expert group through experimental data and specific application scenarios. In this embodiment, 0.06≤d2≤1;

[0224] C RE The value range is also between 0 and 1;

[0225] When C RE The closer it is to 1, the stronger the correlation between disk rotation instability and read / write error rate is, which means that disk rotation instability has a greater impact on read / write error rate and is the key factor leading to an increase in read / write error rate.

[0226] When C RE The closer it is to 0, the weaker the correlation between the two, the smaller the impact of disk rotation instability on the read and write error rate, and the smaller the probability of failure;

[0227] Setting C RE The evaluation threshold is C Eh ; 0.35≤C Eh ≤0.75, C Eh Determined through historical data analysis and practical application experience, C RE With C Eh The size judgment between them is used to distinguish between normal state and fault state;

[0228] High risk indication: When C RE ≥C Eh When the disk is not stable, it means that the instability of disk rotation has a significant negative impact on the transfer rate, indicating that the hard disk is in or close to failure. In this case, the hard disk should be inspected in more detail or preventive maintenance measures should be taken directly.

[0229] Low risk indication: When C RE <C Eh When , it indicates that the impact of disk rotation instability on the read and write error rate is within an acceptable range, the hard disk status is relatively stable, and the failure risk is within 15%.

[0230] Further explanation: the construction contents of the first trend evaluation coefficient and the second trend evaluation coefficient are as follows:

[0231] Calculate the average trend of disk rotation instability data:

[0232]

[0233] Among them, T R Indicates the average change trend of disk rotation instability; ΔR i,i+1 represents the change in disk rotation instability between the i-th and i+1-th data collections;

[0234] Calculate the average trend of head loading times:

[0235]

[0236] Among them, T L Indicates the average change trend of the number of head loading times; ΔL i,i+1 It represents the change in the number of head loading times between the i-th and i+1-th data acquisitions;

[0237] The following first trend evaluation coefficients are calculated:

[0238]

[0239] Among them, C T is the first trend evaluation coefficient, 0<C T <1, e1 and e2 are the weight coefficients of the corresponding parameters, and The specific values ​​of e1 and e2 are determined by the expert group through experimental data. For example, in high-speed reading and writing scenarios, the number of head loading times has a greater impact on hard disk failure than disk rotation instability, so a higher e2 value needs to be set;

[0240] When C T As e1·T approaches 1, R +e2·T L The smaller the output value, the smaller the change trend of the physical status indicator in the current collection cycle;

[0241] When C T As e1·T approaches 0, R +e2·T L The larger the output value, the greater the change trend of the physical status indicator in the current collection cycle;

[0242] Calculate the average change trend of data transmission rate:

[0243]

[0244] Among them, T S Indicates the average change trend of data transmission rate; ΔT i,i+1 It represents the change in data transmission rate between the i-th and i+1-th data collection;

[0245] Calculate the average change trend of read and write error rates:

[0246]

[0247] Among them, T C Indicates the average change trend of the read and write error rate; ΔE i,i+1 It represents the change in read / write error rate between the i-th and i+1-th data collection;

[0248] The following second trend evaluation coefficients are calculated:

[0249]

[0250] Among them, C U is the second trend evaluation coefficient, 0<C U <1, e2 and e3 are the weight coefficients of the corresponding parameters, and The specific values ​​of e2 and e3 are determined by the expert group through experimental data;

[0251] When C U The closer it is to 1, The smaller the output value, the smaller the change trend of the operating status indicator in the current collection cycle;

[0252] When C U The closer it is to 0, The larger the output value, the greater the change trend of the operating status indicator in the current collection cycle.

[0253] Further explanation: constructing a threshold fine-tuning model, the threshold fine-tuning model is used to provide a fine-tuning strategy for the initial fault probability threshold, specifically including:

[0254] The calculation formula for defining the threshold fine-tuning model is as follows:

[0255]

[0256] Among them, WT1 is a first comprehensive index combining the first correlation evaluation coefficient and the second correlation evaluation coefficient, which overall reflects the correlation degree of the computer hard disk status; WT2 is a second comprehensive index combining the first trend evaluation coefficient and the second trend evaluation coefficient, which overall reflects the trend degree of the computer hard disk status, P fault is the initial failure probability threshold, P 1 ′ is to reduce P fault The fault probability threshold after taking the value; P′ 2 To improve P fault The fault probability threshold mark after taking the value;

[0257] r1, r2, r3, and r4 are regression coefficients of corresponding parameters, which are obtained and determined through historical data training and can reflect the influence of various variables on the failure risk. The values ​​of r1, r2, r3, and r4 are all positive numbers, and r1+r2=1, r3+r4=1. The specific values ​​of r1, r2, r3, and r4 are determined by the expert group through experimental data.

[0258] μ RT , Represent the first correlation evaluation coefficient C RT The mean and standard deviation of are used for normalization; μ RE , Respectively represent the second correlation evaluation coefficient C RE The mean and standard deviation of the data are calculated by conventional methods of existing data processing and will not be described in detail. They are used for normalization processing.

[0259] η1, η2, η3, η4 are all positive constants, and The specific values ​​of η1, η2, η3, and η4 are determined by the expert group through experimental data;

[0260] The division thresholds of the first comprehensive index and the second comprehensive index are respectively set to Q1 and Q2;

[0261] When WT1 ≥ Q1, it means that the correlation between the disk's rotational instability and the data transfer rate is significant; this means that the hard disk is in poor condition, there is a high risk of failure, and the data transfer efficiency is seriously affected; at this time, a detailed hard disk health check is performed to avoid potential data loss or system crash;

[0262] When WT1 < Q1, the correlation between the disk's rotational instability and the data transfer rate is weak, the system is performing normally, and the risk of failure is low. At this point, the hard disk operation can continue, but the status still needs to be monitored regularly to ensure that no potential problems occur.

[0263] When WT2 ≥ Q2, it means that the change trend of the physical status indicator is obvious, indicating that the operating status of the hard disk has fluctuated greatly during the current acquisition cycle. This is caused by changes in the external environment or internal failures of the hard disk. At this time, take measures to check the operating environment and maintenance status of the hard disk.

[0264] When WT2<Q2, it means that the change trend of the physical status indicator is small, indicating that the operation status of the hard disk in the current acquisition cycle is relatively stable, the failure risk is low, and the operation can be carried out normally; however, it is still necessary to observe the long-term trend to prevent the accumulation of potential hidden dangers;

[0265] The fine-tuning strategy is as follows:

[0266] When WT1≥Q1 and WT2≥Q2, use P′ 2 Fine-tuning strategy; at this time, the impact of disk rotation instability on data transmission rate and read / write error rate exceeds 75%, and the volatility of physical state and operation state also exceeds 75%; in this case, it indicates that the system failure risk is extremely high, and the initial failure probability threshold needs to be increased, limiting P′ 2 The increase is P fault to within 10% to 20% of the target, ensuring early warning in high-risk situations;

[0267] When WT1≥Q1 and WT2<Q2, use P′ 2 Fine-tuning strategy; at this time, the impact of disk rotation instability on data transmission rate and read-write error rate exceeds 75%, but the volatility of physical state and operating state is less than 25%; although the operating state is relatively stable, the initial failure probability threshold needs to be increased due to the strong correlation between data transmission and read-write error rate; limit P′ 2 The increase is P fault within 10% of the previous value, increasing the system's sensitivity to key indicators;

[0268] When WT1<Q1 and WT2≥Q2, use P′ 2 Fine-tuning strategy; at this time, the impact of disk rotation instability on data transfer rate and read / write error rate is less than 25%, but the volatility of physical state and operating state exceeds 75%; in this case, although the data transfer rate shows a low risk of failure, due to the large fluctuations in physical and operating states, the initial failure probability threshold needs to be increased. After increasing the initial failure probability threshold within 15%, the system can better cope with the potential failure risk caused by operating state fluctuations;

[0269] When WT1<Q1 and WT2<Q2, use P 1 ' fine-tuning strategy; at this time, the impact of disk rotation instability on data transmission rate and read / write error rate is less than 25%, and the volatility of physical state and operating state is also less than 25%; in this case, the overall failure risk is low, the initial failure probability threshold can be reduced, the system avoids excessive sensitivity, and the probability of false alarm is reduced;

[0270] When WT1 or WT2 changes:

[0271] If WT1 increases by more than 25% and WT2 remains unchanged or changes less than 5%, priority is given to increasing P fault Value to cope with the risk of failure;

[0272] If WT1 decreases by more than 25% and WT2 remains unchanged or changes less than 5%, priority is given to reducing P fault value to avoid over-sensitivity of the system;

[0273] The sample application is as follows:

[0274] Assume P fault =0.5, the following are examples of fault threshold adjustment in various situations:

[0275] High-risk scenario: WT1=80%, WT2=85%;

[0276] Adjusted P′ 2 =0.5+0.2=0.7(70%);

[0277] Medium risk scenario: WT1=80%, WT2=20%;

[0278] Adjusted P′ 2 =0.5+0.1=0.6;

[0279] Low risk scenario: WT1 = 20%, WT2 = 85%;

[0280] Adjusted P′ 2 =0.5+0.15=0.65;

[0281] No-risk scenario: WT1=20%, WT2=20%;

[0282] Adjusted P 1 ′=0.5-0.15=0.35.

[0283] Further explanation: the failure probability of the hard disk in the current nth data collection is further calculated. If the failure probability exceeds the adjusted failure probability threshold, a failure warning is triggered, specifically including:

[0284] According to the failure probability P of the hard disk in the current nth data collection th,n , when P th,n conform to When any of the following conditions are met, the hard disk is judged to be in a faulty state, otherwise it is judged to be in a normal state;

[0285] Once a fault warning is triggered, the system will automatically start the subsequent fault diagnosis procedures, including:

[0286] Real-time monitoring of hard disk status indicators;

[0287] Generate detailed fault reports. The system automatically integrates and analyzes the hard disk status data collected at present and in the past to generate detailed fault reports. The generated fault reports should be saved in standardized formats, including PDF and editable document formats. The reports should be automatically archived in the system log and associated with the specific hard disk serial number for future inquiries.

[0288] Notify the operator to perform necessary troubleshooting or data backup work, as follows:

[0289] Automatic notification mechanism:

[0290] Immediate notification: The system will immediately notify relevant operators through multiple channels (such as email, SMS, real-time notification system), including the summary information of the fault warning, the current hard disk status, and the recommended preliminary treatment measures;

[0291] Priority setting: According to the risk level of the fault (such as high, medium, and low), the system sets the priority of notification. High-priority faults will be sent to the main person in charge and his / her superior management personnel, and medium-priority faults will be notified to general maintenance personnel;

[0292] Troubleshooting Guide:

[0293] Automatically generate troubleshooting suggestions: Based on the analysis results of the fault report, the system will automatically generate detailed troubleshooting suggestions; these suggestions include:

[0294] Reduce the hard disk load and reduce write operations;

[0295] Add heat dissipation equipment to control disk temperature;

[0296] Migrate important data from the hard drive at risk to other storage devices;

[0297] Perform hard disk self-test or boot SMART test;

[0298] Backup operation guide: In the case of potential threats to data security, the system should automatically generate backup guides to help operators quickly back up key data to a safe location. The backup guides should include recommended backup methods (such as mirror backup, incremental backup), backup target devices, and estimated backup time;

[0299] Response confirmation and feedback:

[0300] Confirmation mechanism: After receiving the notification, the operator should confirm in the system that he has received it and start to handle the fault. The system should require the operator to regularly update the processing progress and submit the final processing results;

[0301] Feedback analysis: After the fault handling is completed, the system will analyze the handling effect, record the lessons learned during the handling process, and incorporate them into the reference library for future fault handling.

[0302] Embodiment 2:

[0303] Further explanation based on Example 1, the purpose of this experiment is to verify the effectiveness of the fault warning system based on the threshold fine-tuning model under different hard disk conditions, especially its performance in dynamically adjusting the initial fault probability threshold; the test objects are 5 server hard disks that have been running for more than 2 years, all of which are enterprise-level SATA hard disks, and the average running time of each hard disk in the past year is about 6000 hours; the data collection indicators selected in the experiment include hard disk rotation instability, data transmission rate, read and write error rate, etc.

[0304] During the experiment, the threshold fine-tuning model was actually verified using experimental data. The specific process is as follows:

[0305] 1) Initial state data collection:

[0306] First, based on the SMART (Self-Monitoring Analysis and Reporting Technology) data of the hard disk, the status indicators of each hard disk in the past 48 hours are obtained, and the first correlation evaluation coefficient C is calculated respectively. RT and the second correlation evaluation coefficient C RE , and the related trend evaluation coefficient C T and C U ; The data of these indicators are used to calculate the first comprehensive index WT1 and the second comprehensive index WT2;

[0307] 2) Parameter setting of threshold fine-tuning model:

[0308] According to historical data, the regression coefficients are set to r1 = 0.5, r2 = 0.4, r3 = 0.7, r4 = 0.6, and the partition thresholds are set to Q1 = 0.75, Q2 = 0.75; the initial failure probability threshold P fault =0.5;

[0309] 3) Experimental steps:

[0310] a. Use the threshold fine-tuning model to calculate the WT1 and WT2 of each hard disk. The formula is as follows:

[0311]

[0312] Based on the calculated WT1 and WT2 values, use the following fine-tuning strategy:

[0313] When WT1≥Q1 and WT2≥Q2, use P′ 2 Fine-tune the fault threshold and increase the fault warning sensitivity;

[0314] When WT1<Q1 and WT2<Q2, use P 1 ′, and reduce the sensitivity of fault warning;

[0315] 4) Monitoring and early warning:

[0316] By comparing the real-time data with the fine-tuned fault probability threshold, we can see whether an early warning is triggered. When the threshold is exceeded, the system will trigger an alarm and record the time of the fault and changes in related indicators.

[0317] The experimental data table is as follows:

[0318] Table 1

[0319]

[0320] Data Analysis and Conclusion:

[0321] From the above experimental data, we can see that when WT1 and WT2 are both high (such as hard disks A, C, and E), the system increases the failure probability threshold and triggers an early warning in real-time monitoring, indicating that the hard disk is at risk of failure. For hard disks with low WT1 and WT2 (such as hard disks B and D), the system lowers the threshold and does not trigger an early warning, indicating that the hard disk status is relatively stable.

[0322] Hard disk status distribution and fault warning triggering conditions:

[0323] Hard disks A, C, and E all showed values ​​higher than the set thresholds in the experiment (WT1 and WT2 were both ≥ Q1 and Q2). The rotational instability and trend index of these hard disks were high, 0.76, 0.80, and 0.78, respectively. These high values ​​indicate that these hard disks have large fluctuations during operation, and it is necessary to issue an early warning for intervention.

[0324] Hard disks B and D showed lower WT1 and WT2 values ​​(0.68 and 0.62; 0.62 and 0.60, respectively), so no warning was triggered; this indicates that their status is relatively stable, reducing the risk of false alarms;

[0325] Failure probability threshold after fine-tuning:

[0326] For hard disks A, C, and E, in fine-tuning the failure probability threshold, the system sets P′ 2 Increased to 0.61, 0.64 and 0.63, reflecting an increase in the probability of failure of 22%, 28% and 26% (relative to the initial probability of failure of 0.50); ​​this adjustment ensures that the system is more sensitive in high-risk situations and warns of potential failures in a timely manner;

[0327] For hard disks B and D, the thresholds after fine-tuning were reduced to 0.46 and 0.45, respectively, with reductions of 8% and 10%. This shows that the system effectively avoids false alarms when the hard disks are in normal condition, thus improving the stability and security of the system.

[0328] The relationship between parameters:

[0329] In the formula, the values ​​of WT1 and WT2 directly affect the fine-tuned failure probability threshold. For example, if WT1 increases from 0.68 to 0.76 (such as from hard disk B to hard disk A), the failure probability will increase significantly. This is because in a high-risk state, the system needs to be more sensitive to failures.

[0330] Specifically, if WT1 increases by 10% (from 0.70 to 0.77), assuming that other parameters remain unchanged, it will lead to P′ 2 The increase is 15%, that is, P′ 2 From 0.50 to 0.65; this shows that the increased rotational instability directly affects the adjustment range of the threshold fine-tuning strategy, improving the response speed of the system;

[0331] Through the quantitative fine-tuning mechanism, the failure risk of the hard disk is controlled within an acceptable range; by setting different threshold intervals, the system can adjust the corresponding failure probability threshold under different conditions;

[0332] When WT1≥Q1 and WT2≥Q2, the failure probability threshold is increased by 10%-20%, effectively increasing the risk warning accuracy to 85%-90%;

[0333] When WT1<Q1 and WT2<Q2, the fault probability threshold is reduced by 10%-15%, reducing the risk of false alarm to 5%-10%;

[0334] Through fine-tuning based on actual status, the system significantly reduces the probability of false alarms when the hard disk status is normal. From the table data analysis, it can be seen that hard disks B and D did not trigger alarms, avoiding unnecessary maintenance costs;

[0335] The fine-tuning model of the present invention effectively identifies high-risk hard disks and issues early warnings before failures occur, so that timely measures can be taken to reduce the risk of data loss;

[0336] According to different state changes, the system flexibly adjusts the fault threshold to make the response strategy more targeted; when the risk of failure is high, the threshold is increased to improve vigilance; when the state is stable, the threshold is lowered to reduce interference.

[0337] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.

[0338] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. Those skilled in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software methods depends on the specific application and design constraints of the technical solution.

[0339] The above embodiments are only used to help understand the method and core idea of ​​the present invention. It should be noted that, for those skilled in the art, several improvements and modifications can be made to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.

Claims

1. A method for fault diagnosis based on computer hard disk status indicators, characterized in that: The specific steps include: Step S1: Set the collection period of the hard disk to the set {1,2,…,n}, where i∈{1,2,…,n} represents the index of the i-th data collection in the collection period, and n represents the index of the current n-th data collection, collect the physical state indicators and operation state indicators of the hard disk, where the physical state indicators include disk rotation instability data and head loading times, and the operation state indicators include data transmission rate and read / write error rate, and perform normalization and denoising preprocessing on the collected data to obtain multi-dimensional feature data; the acquisition of multi-dimensional feature data includes: The disk rotation instability data includes the disk rotation speed fluctuation rate and the disk vibration amplitude, and the disk rotation speed fluctuation rate and the disk vibration amplitude are marked as CVb and CZf respectively; Combine the disk rotation speed fluctuation rate and disk vibration amplitude, and perform analysis and processing to construct the disk rotation instability value of the i-th data collection. , the calculation formula is as follows: ; Parameter explanation: is the disk rotation instability value of the ith data collection, is the disk rotation speed fluctuation rate of the ith data collection, is the disk vibration amplitude of the ith data acquisition, are all positive weight coefficients, and ; The physical status indicator data and operation status indicator data are acquired regularly within the set collection cycle, and the collected data are recorded in the database to form a data set ; in, They represent the disk rotation instability value, head loading times, data transmission rate and read / write error rate of the i-th data acquisition respectively; For normalization, the min-max normalization method is used to normalize each index value x to x'; The normalized data range is limited to (0,1), where ; For denoising, the moving average method is used to remove random noise in the data to smooth the normalized data of each indicator: For multidimensional feature data, after normalization and denoising, the final multidimensional feature data set is expressed as ,in and They are the physical state index and the operating state index after data preprocessing, respectively, and k-1 represents the starting point of the number of acquisitions after denoising; Step S2: receiving multi-dimensional feature data collected n times, using an autoencoder to reduce the dimension of the multi-dimensional features, and extracting the key feature vector after the dimension reduction; Step S3: receiving the key feature vector after dimension reduction, and using the support vector machine to calculate the failure probability of the key feature vector to achieve binary classification of hard disk failures; Set the initial failure probability threshold of the hard disk failure, and set the hard disk failure warning trigger condition according to the initial failure probability threshold; Step S4: obtaining disk rotation instability data, data transmission rate and read / write error rate, and performing correlation analysis on the disk rotation instability data and the data transmission rate to obtain a first correlation evaluation coefficient, which is used to evaluate the correlation influence degree between the disk rotation instability data and the data transmission rate; Performing correlation analysis on the disk rotation instability data and the read / write error rate to obtain a second correlation evaluation coefficient, the second correlation evaluation coefficient is used to evaluate the correlation influence degree between the disk rotation instability data and the read / write error rate; Step S5: calculating the change trend data of the physical state indicator in the current collection period, and analyzing and processing the change trend data to generate a first trend evaluation coefficient, which is used to evaluate the change trend of the physical state indicator in the current collection period; Calculate the change trend data of the operation status indicator in the current collection cycle, and analyze and process the change trend data to generate a second trend evaluation coefficient, which is used to evaluate the change trend of the operation status indicator in the current collection cycle; Step S6: constructing a threshold fine-tuning model by combining the first correlation evaluation coefficient, the second correlation evaluation coefficient, the first trend evaluation coefficient, and the second trend evaluation coefficient, wherein the threshold fine-tuning model is used to provide a fine-tuning strategy for the initial fault probability threshold; Step S7: Obtain the fault probability threshold adjusted by the fine-tuning strategy, and adjust the fault warning trigger condition according to the adjusted fault probability threshold, and further calculate the failure probability of the hard disk in the current nth data collection. If the failure probability exceeds the adjusted failure probability threshold, a fault warning is triggered.

2. A method for fault diagnosis based on computer hard disk status indicators according to claim 1, characterized in that: Receive multi-dimensional feature data collected n times, use the autoencoder to reduce the dimension of the multi-dimensional features, and extract the key feature vector after the dimension reduction, including: Choose autoencoder as a dimensionality reduction tool; the autoencoder consists of two parts: encoder and decoder, where the encoder converts high-dimensional input data Compressed into a low-dimensional feature vector , the decoder then Restore to high-dimensional space; For each data collection point i, the input multidimensional feature data It is expressed as: ; The output of the encoder network is a low-dimensional feature vector : ; in, is the encoder weight matrix, is the bias vector, is the activation function, Represents the set of all encoder parameters; The autoencoder is trained by minimizing the reconstruction error so that the decoder outputs the reconstructed data Approximate the original input data ; After the autoencoder training is completed, the output low-dimensional feature vector is directly used As the key feature vector after dimensionality reduction; The feature vector after dimensionality reduction is expressed as: ; Among them, m is the dimension of the feature vector after dimensionality reduction.

3. A method for fault diagnosis based on computer hard disk status indicators according to claim 2, characterized in that: The first correlation evaluation coefficient and the second correlation evaluation coefficient are constructed as follows: The first Pearson correlation coefficient between the disk rotation instability data and the data transfer rate is calculated as: ; in, is the first Pearson correlation coefficient between disk rotation instability and data transfer rate; and They are and The mean in the set {1,2,…,n}; The second Pearson correlation coefficient between the disk rotation instability data and the read and write error rate is calculated using the formula: ; in, is the second Pearson correlation coefficient between disk rotation instability and read and write error rate; and They are and The mean in the set {1,2,…,n}.

4. A method for fault diagnosis based on computer hard disk status indicators according to claim 3, characterized in that: Construct a threshold fine-tuning model. The threshold fine-tuning model is used to provide a fine-tuning strategy for the initial failure probability threshold, including: The calculation formula for defining the threshold fine-tuning model is as follows: ; Among them, WT1 is a first comprehensive index combining the first correlation evaluation coefficient and the second correlation evaluation coefficient, which overall reflects the correlation degree of the computer hard disk status; WT2 is a second comprehensive index combining the first trend evaluation coefficient and the second trend evaluation coefficient. is the first trend evaluation coefficient, is the second trend evaluation coefficient, which reflects the overall trend of the computer hard disk status. is the initial failure probability threshold, To reduce The fault probability threshold mark after taking the value; To improve The fault probability threshold mark after taking the value; r1, r2, r3, and r4 are the regression coefficients of the corresponding parameters. Represent the first correlation evaluation coefficient The mean and standard deviation of are used for normalization; Represent the second correlation evaluation coefficient The mean and standard deviation of are used for normalization; All are normal numbers; The division thresholds of the first comprehensive index and the second comprehensive index are respectively set to Q1 and Q2; when When , it means that the correlation between the disk's rotational instability and the data transfer rate is significant; this means that the hard disk is in a bad state, there is a high risk of failure, and the data transfer efficiency is seriously affected; when When , it means that the correlation between the disk's rotational instability and the data transfer rate is weak, the system performs normally, and the risk of failure is low; when When , it means that the change trend of the physical status indicator is obvious, indicating that the operating status of the hard disk has fluctuated greatly during the current acquisition cycle; when , it means that the change trend of the physical status indicator is small, indicating that the operation status of the hard disk in the current acquisition cycle is relatively stable and the failure risk is low.

Citation Information

Patent Citations

  • Methods, devices, equipment, and computer-readable storage media for predicting hard disk failures.

    CN111611117B

  • Hard disk fault prediction method and device, electronic equipment and storage medium

    CN114758714A

  • SMART threshold value optimizing method orienting magnetic disk fault detection

    CN108228377A