Computer hardware fault analysis method and system

By pre-processing and classifying the historical failure event data of computer hardware, combined with machine learning algorithms and FMEAC algorithms, computer hardware failure analysis methods can more accurately detect and diagnose hardware failures, solving the problem of difficulty in identifying hidden faults and real-time responses in the prior art, and achieving faster and more accurate fault identification and early warning.

CN120011159AInactive Publication Date: 2025-05-16NANTONG YOUZUO NETWORK INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510106764.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art relies on incomplete hardware English alarm history information in computer hardware failure analysis, resulting in the inaccurate failure that occurs instantly and recovers quickly. The real-time fault identification process takes a long time, making it difficult to quickly respond to sudden serious hardware failures in large host systems.

Method used

By obtaining historical fault event data of computer hardware, pre-processing and classification processing, extracting fault characteristic data, and using a fault judgment model based on machine learning algorithms and FMEAC algorithm to calculate risk-first data and comprehensive fault data to determine whether the fault event occurs.

Benefits of technology

It realizes more accurate detection and diagnosis of hardware failures, reduces false alarms and missed reports, and can identify potential abnormal signs before the failure occurs, and issue early warnings to ensure that operation and maintenance personnel have enough time to take preventive measures to avoid or reduce the impact of the failure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011159A_ABST
    Figure CN120011159A_ABST
Patent Text Reader

Abstract

The invention discloses a computer hardware fault analysis method, which comprises the following steps: acquiring historical fault event data of computer hardware, and carrying out preprocessing, classification processing and feature extraction to obtain fault feature data; inputting the fault feature data into a fault judgment model established based on a machine learning algorithm, and outputting to obtain a classified fault judgment value; calculating the distance between the classification fault judgment value and a preset classification fault judgment value threshold value to obtain first fault data; calculating risk priority data corresponding to the fault feature data based on an FMEAC algorithm, and calculating a distance between the risk priority data and a preset risk priority data threshold to obtain second fault data; and calculating the comprehensive fault data of the first fault data and the second fault data, and judging whether a fault event occurs or not by using the comprehensive fault data, so that the fault analysis is more comprehensive and accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer hardware failure, and in particular to a computer hardware failure analysis method and system. Background Art

[0002] By analyzing computer hardware failures and understanding the causes and patterns of hardware failures, it is helpful to implement reasonable maintenance strategies, thereby extending the service life of the hardware. Potential risks and weak links in the hardware system can be identified, and corresponding preventive measures can be taken to effectively reduce the probability of hardware failures, improve the reliability of the entire computer system, and ensure that the system can run continuously and stably.

[0003] At present, the Chinese invention patent with application number CN113537349A discloses a method, device, equipment and storage medium for identifying hardware faults of large mainframes. The method includes: extracting keyword English entities from the hardware English alarm history information of the target host system; quantizing and encoding the keyword English entities according to the frequency of occurrence of each type of letters to obtain a fault feature sequence set; training a hidden Markov model according to the fault feature sequence set to obtain a hardware fault identification model; and using the hardware fault identification model to identify the hardware fault of the target host system. However, the hardware English alarm history information is only a record after the fault occurs, which may miss key information. Some hardware faults that occur instantly and recover quickly may not be fully presented in the alarm information. This leads to inherent defects in the keywords extracted and the trained models based on these incomplete information, and such hidden faults cannot be identified. Extracting data from the alarm history information, training models, and then using them for real-time hardware fault identification takes a long time. For large mainframe systems that require rapid response, once a serious hardware fault occurs, this set of processes is difficult to give identification results in time, missing the best time for emergency repairs, and causing long-term business interruption. Summary of the invention

[0004] The technical problem solved by the present invention is that the hardware English alarm history information is only a record after the fault occurs, and it may miss key information. Some hardware faults that occur instantly and recover quickly may not be fully presented in the alarm information. This leads to inherent defects in the keywords extracted and the models trained based on these incomplete information, and it is impossible to identify such hidden faults. The process of extracting data and training models from the alarm history information and then using them for real-time hardware fault identification is time-consuming. For large host systems that require rapid response, once a serious hardware fault occurs, this process is difficult to give identification results in time, missing the best time for emergency repairs and causing long-term business interruption.

[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: a computer hardware failure analysis method, the specific steps include:

[0006] Step S100, obtaining historical fault event data of computer hardware, preprocessing and classifying the historical fault event data, extracting features of the preprocessed and classified historical fault event data, and obtaining fault feature data;

[0007] Step S200, inputting the fault feature data into a fault judgment model established based on a machine learning algorithm, and outputting a classified fault judgment value, wherein the classified fault judgment value includes a first classified fault judgment value, a second classified fault judgment value, and a third classified fault judgment value;

[0008] Step S300, calculating the distances between the first classification fault judgment value, the second classification fault judgment value and the third classification fault judgment value and a preset classification fault judgment value threshold to obtain first fault data;

[0009] Step S400, calculating the risk priority data corresponding to the fault feature data based on the FMEAC algorithm, calculating the distance between the risk priority data and a preset risk priority data threshold, and obtaining second fault data;

[0010] Step S500: Calculate comprehensive fault data of the first fault data and the second fault data, and use the comprehensive fault data to determine whether a fault event occurs.

[0011] As a preferred solution of the computer hardware failure analysis method described in the present invention, step S100 specifically includes:

[0012] Step S101, querying the computer's fault log to obtain historical fault event data of the computer hardware, the historical fault event data including hardware temperature rise, hardware power short circuit, hardware power disconnection and hardware device running speed reduction, and assigning a fault identifier to the historical fault event data;

[0013] Step S102, preprocessing the historical fault event data, wherein the preprocessing includes removing invalid data, incomplete data and duplicate data of the historical fault event data, and classifying the preprocessed historical fault event data based on temperature category, current category and operating parameter category;

[0014] Step S103, extracting characteristic data of the historical fault event data after the classification processing, obtaining temperature characteristic quantities, current characteristic quantities and operating parameter characteristic quantities under the temperature category, current category and operating parameter category, and normalizing the temperature characteristic quantities, current characteristic quantities and operating parameter characteristic quantities.

[0015] As a preferred solution of the computer hardware failure analysis method described in the present invention, step S200 specifically includes:

[0016] Step S201, using one-hot encoding to encode the normalized temperature feature quantity, current feature quantity and operating parameter feature quantity;

[0017] Step S202, input the encoded temperature characteristic quantity, current characteristic quantity and operating parameter characteristic quantity into the fault judgment model, and output the classified fault judgment value, which includes the first classified fault judgment value, the second classified fault judgment value and the third classified fault judgment value.

[0018] As a preferred solution of the computer hardware failure analysis method described in the present invention, step S300 specifically includes:

[0019] Step S301, the preset classification fault judgment numerical threshold includes a first classification fault judgment numerical threshold, a second classification fault judgment numerical threshold and a third classification fault judgment numerical threshold;

[0020] Step S302, subtracting the first classification fault judgment value threshold from the first classification fault judgment value to obtain first distance data;

[0021] Subtract the second classification fault judgment value from the second classification fault judgment value threshold to obtain second distance data;

[0022] Subtract the third classification fault judgment value from the third classification fault judgment value threshold to obtain third distance data;

[0023] Step S303: Add the first distance data, the second distance data and the third distance data to obtain first fault data.

[0024] As a preferred solution of the computer hardware failure analysis method described in the present invention, step S400 specifically includes:

[0025] Step S401, retrieve the classified historical fault event data, divide the historical fault event data into three impact levels based on the impact degree of the historical fault event data on the performance and function of the computer system, and assign values ​​to obtain impact assignment data;

[0026] Step S402, when the historical fault event data reaches a preset number threshold, the frequency of the historical fault event data of the corresponding temperature category, current category and operating parameter category is counted to obtain frequency data;

[0027] Step S403, retrieving the classified historical fault event data, dividing the historical fault event data into two difficulty levels based on the difficulty of detecting the historical fault event data, and assigning values ​​to obtain difficulty assignment data;

[0028] Step S404, adding the impact value assignment data, frequency data and difficulty value assignment data to obtain the risk priority data, calculating the distance between the risk priority data and a preset risk priority data threshold to obtain second fault data.

[0029] As a preferred solution of the computer hardware failure analysis method of the present invention, step S500 specifically includes:

[0030] Step S501, obtaining first weight data and second weight data corresponding to the historical fault event data;

[0031] Step S502: calculating comprehensive fault data based on the first weight data and the second weight data. The mathematical expression of the comprehensive fault data is:

[0032] Comprehensive fault data = first weight data multiplied by first fault data + second weight data multiplied by second fault data;

[0033] Step S503, calculating the distance between the comprehensive fault data and a preset comprehensive fault data threshold to obtain total fault data, and using the total fault data to determine whether a fault event occurs.

[0034] As a preferred solution of the computer hardware fault analysis method described in the present invention, the method for establishing the fault judgment model specifically includes:

[0035] Obtaining a fault identifier corresponding to the preprocessed historical fault event data, performing binary conversion on the fault identifier, and obtaining a normal identifier corresponding to the historical normal event data;

[0036] The annotated historical fault event data and historical normal event data are divided into training set and test set in a ratio of 7:3;

[0037] The training set is input into the machine learning model for training, and the trained machine learning model is tested using the test set until the accuracy reaches 0.9, thereby obtaining the fault judgment model.

[0038] As a preferred solution of a computer hardware fault analysis method described in the present invention, wherein: the distances between the first classification fault judgment value, the second classification fault judgment value and the third classification fault judgment value and a preset classification fault judgment value threshold are calculated to obtain first fault data, and the mathematical expression of the first fault data is:

[0039] P=Σ(W1-W2)+(R1-R2)+(T1-T2);

[0040] Among them, P represents the first fault data, W1 represents the first category fault judgment value, R1 represents the first category fault judgment value, T1 represents the first category fault judgment value, W2 represents the first category fault judgment value threshold, R2 represents the first category fault judgment value threshold, and T2 represents the first category fault judgment value threshold.

[0041] As a preferred solution of a computer hardware failure analysis method according to the present invention, the impact of the historical failure event data on the performance and function of the computer system is divided into three impact levels and assigned values, which specifically include:

[0042] If the historical fault event data causes the system to be completely inoperable, and the data is lost and cannot be recovered, the impact level is the first level impact level, and the value is 10;

[0043] If the historical fault event data causes the system performance to drop by more than 50%, and some important data is lost but recoverable, the impact level is the second level, and the value is 7;

[0044] If the historical fault event data causes the system performance to drop by less than 20% and there is no data loss, the impact level is the third level of impact severity and is assigned a value of 3;

[0045] The degree of difficulty of the detection based on the historical fault event data is divided into two difficulty levels and the value assignment specifically includes:

[0046] If the historical fault event data can be detected by the detection tool provided by the system, the difficulty level is the first level, and a value of 1 is assigned;

[0047] If the historical fault event data cannot be detected by the detection tools, complex professional equipment and long-term detection are required, the difficulty level is the second level, and the value is 9.

[0048] A computer hardware fault analysis system comprises a data collection module, a fault data calculation module and a fault judgment module;

[0049] The data collection module is used to obtain historical fault event data of computer hardware and perform pre-processing and classification processing;

[0050] The fault data calculation module is used to calculate the first fault data and the second fault data of the historical fault event data based on the fault judgment model and the FMEAC algorithm;

[0051] The fault judgment module is used to calculate comprehensive fault data of the first fault data and the second fault data, and use the comprehensive fault data to judge whether a fault event occurs.

[0052] The beneficial effects of the present invention are as follows: by calculating the risk priority data and the second fault data corresponding to the fault feature data through the FMEAC algorithm, the mode, impact and harmfulness of the hardware failure can be systematically analyzed, and a comprehensive fault knowledge basis is provided for the machine learning algorithm. By outputting the first fault data corresponding to the fault feature data through the machine learning algorithm, the first fault data and the second fault data are combined to analyze the computer hardware failure, which can more accurately detect and diagnose hardware failures, reduce false alarms and missed alarms, identify potential abnormal signs before the failure occurs, and issue early warnings, so that operation and maintenance personnel can have enough time to take preventive measures to avoid the occurrence of failures or reduce the impact of failures. New failure modes can be promptly incorporated into the entire failure analysis system to make the failure analysis more comprehensive and accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A basic flow chart of a computer hardware failure analysis method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.

[0055] Example, see Figure 1 , is an embodiment of the present invention, and provides a computer hardware failure analysis method, the specific steps of which include:

[0056] Step S100, obtaining historical fault event data of computer hardware, preprocessing and classifying the historical fault event data, extracting features of the preprocessed and classified historical fault event data, and obtaining fault feature data;

[0057] Step S200, inputting the fault feature data into a fault judgment model established based on a machine learning algorithm, and outputting a classified fault judgment value, wherein the classified fault judgment value includes a first classified fault judgment value, a second classified fault judgment value, and a third classified fault judgment value;

[0058] Step S300, calculating the distances between the first classification fault judgment value, the second classification fault judgment value and the third classification fault judgment value and a preset classification fault judgment value threshold to obtain first fault data;

[0059] Step S400, calculating the risk priority data corresponding to the fault feature data based on the FMEAC algorithm, calculating the distance between the risk priority data and a preset risk priority data threshold, and obtaining second fault data;

[0060] Step S500: Calculate comprehensive fault data of the first fault data and the second fault data, and use the comprehensive fault data to determine whether a fault event occurs.

[0061] Step S100 specifically includes:

[0062] Step S101, querying the computer's fault log to obtain historical fault event data of the computer hardware, the historical fault event data including hardware temperature rise, hardware power short circuit, hardware power disconnection and hardware device running speed reduction, and assigning a fault identifier to the historical fault event data;

[0063] Step S102, preprocessing the historical fault event data, wherein the preprocessing includes removing invalid data, incomplete data and duplicate data of the historical fault event data, and classifying the preprocessed historical fault event data based on temperature category, current category and operating parameter category;

[0064] Step S103, extracting characteristic data of the historical fault event data after the classification processing, obtaining temperature characteristic quantities, current characteristic quantities and operating parameter characteristic quantities under the temperature category, current category and operating parameter category, and normalizing the temperature characteristic quantities, current characteristic quantities and operating parameter characteristic quantities.

[0065] In this embodiment, eliminating invalid data, incomplete data and duplicate data can remove noise and interference in the data, making the data more accurate, complete and reliable. Incomplete data leads to deviations in the analysis results, and duplicate data increases the burden of data processing and the complexity of analysis. Classifying the pre-processed historical fault event data based on temperature, current and operating parameter categories helps to divide complex fault data according to different physical characteristics, making the data more structured and organized, and facilitating subsequent in-depth analysis and feature extraction of different categories of data, so as to more accurately find out the causes and patterns of various types of faults.

[0066] Step S200 specifically includes:

[0067] Step S201, using one-hot encoding to encode the normalized temperature feature quantity, current feature quantity and operating parameter feature quantity;

[0068] Step S202, input the encoded temperature characteristic quantity, current characteristic quantity and operating parameter characteristic quantity into the fault judgment model, and output the classified fault judgment value, which includes the first classified fault judgment value, the second classified fault judgment value and the third classified fault judgment value.

[0069] In this embodiment, the normalized temperature feature quantity, current feature quantity and operating parameter feature quantity are encoded using one-hot encoding, which can convert the original classification characteristics into a numerical form that is easier for the machine learning model to handle. The temperature feature quantities are divided into "low temperature", "medium temperature" and "high temperature". After one-hot encoding, they will become three binary vectors, eliminating unreasonable implicit relationships such as the numerical size and order of the classification variables, so that the model will not misjudge the erroneous association between categories, thereby improving the accuracy of model learning, avoiding certain features from dominating the model learning due to excessively large values, improving the comprehensiveness of the data input to the fault judgment model, and minimizing the information loss caused by data preprocessing.

[0070] Step S300 specifically includes:

[0071] Step S301, the preset classification fault judgment numerical threshold includes a first classification fault judgment numerical threshold, a second classification fault judgment numerical threshold and a third classification fault judgment numerical threshold;

[0072] Step S302, subtracting the first classification fault judgment value threshold from the first classification fault judgment value to obtain first distance data;

[0073] Subtract the second classification fault judgment value from the second classification fault judgment value threshold to obtain second distance data;

[0074] Subtract the third classification fault judgment value from the third classification fault judgment value threshold to obtain third distance data;

[0075] Step S303: Add the first distance data, the second distance data and the third distance data to obtain first fault data.

[0076] In this embodiment, the numerical threshold value of the first category fault judgment is 1.3, the numerical threshold value of the second category fault judgment is 1.8, and the numerical threshold value of the third category fault judgment is 1.5. The first, second, and third category fault judgment values ​​are divided and each corresponds to a threshold value, which can cover fault conditions of different dimensions. By calculating the distances to the corresponding thresholds respectively and summarizing them into the first fault data, various potential fault risks are integrated into one indicator, which can more comprehensively reflect the overall fault risk of the server and avoid the one-sidedness of a single-dimensional evaluation. The first fault data provides a quantitative and intuitive system health scale for fault judgment. The larger the value, the greater the possibility that the system is far from various faults. Operation and maintenance personnel no longer need to interpret the complex data of different categories separately. Based on this summarized value, they can quickly determine whether the current operating status of the computer hardware system is safe and whether immediate intervention for troubleshooting and maintenance is required.

[0077] Step S400 specifically includes:

[0078] Step S401, retrieve the classified historical fault event data, divide the historical fault event data into three impact levels based on the impact degree of the historical fault event data on the performance and function of the computer system, and assign values ​​to obtain impact assignment data;

[0079] Step S402, when the historical fault event data reaches a preset number threshold, the frequency of the historical fault event data of the corresponding temperature category, current category and operating parameter category is counted to obtain frequency data;

[0080] Step S403, retrieving the classified historical fault event data, dividing the historical fault event data into two difficulty levels based on the difficulty of detecting the historical fault event data, and assigning values ​​to obtain difficulty assignment data;

[0081] Step S404, adding the impact value assignment data, frequency data and difficulty value assignment data to obtain the risk priority data, calculating the distance between the risk priority data and a preset risk priority data threshold to obtain second fault data.

[0082] In this embodiment, the preset number threshold is 100 times. Obtaining the impact assignment data, frequency data and difficulty assignment data corresponding to the historical fault event data can more comprehensively and objectively evaluate the reliability of the system. The severity level classification and assignment provide a quantitative basis for operation and maintenance management and decision-making, and more targeted measures can be taken to improve the reliability of the system, enhance the scientific nature of decision-making, and rationally plan hardware selection, redundant configuration and disaster recovery strategies to ensure that the computer system can operate stably and efficiently in the future.

[0083] Step S500 specifically includes:

[0084] Step S501, obtaining first weight data and second weight data corresponding to the historical fault event data;

[0085] Step S502: calculating comprehensive fault data based on the first weight data and the second weight data. The mathematical expression of the comprehensive fault data is:

[0086] Comprehensive fault data = first weight data multiplied by first fault data + second weight data multiplied by second fault data;

[0087] Step S503, calculating the distance between the comprehensive fault data and a preset comprehensive fault data threshold to obtain total fault data, and using the total fault data to determine whether a fault event occurs.

[0088] In this embodiment, the first weight data is 0.7, the second weight data is 0.3, and the comprehensive fault data threshold is 5.2. The first weight data corresponding to the fault judgment model and the second weight data corresponding to the FMEAC algorithm are obtained, so that fault-related factors of different dimensions and different importance levels can be taken into consideration. After allocating reasonable weights, the comprehensive fault data is calculated, which can avoid a single factor dominating the judgment result, thereby more comprehensively and accurately reflecting the real fault risk status of the hardware. The distance between the comprehensive fault data and the preset comprehensive fault data threshold is calculated to obtain the total fault data, and the possibility of the fault is quantified. The smaller the total fault data, the closer the current hardware state is to the fault, and vice versa, the lower the risk. Compared with simple qualitative judgment, operation and maintenance personnel can quickly understand the current health of the hardware system based on this intuitive numerical value, which minimizes the errors caused by subjective assumptions of operation and maintenance personnel. Different operation and maintenance personnel have different experiences and abilities, and subjective judgments on whether a fault occurs are prone to disagreement. The unified quantitative standard ensures the consistency and accuracy of the judgment process, making the fault judgment process more scientific and rigorous.

[0089] The method for establishing the fault judgment model specifically includes:

[0090] Obtaining a fault identifier corresponding to the preprocessed historical fault event data, performing binary conversion on the fault identifier, and obtaining a normal identifier corresponding to the historical normal event data;

[0091] The annotated historical fault event data and historical normal event data are divided into training set and test set in a ratio of 7:3;

[0092] The training set is input into the machine learning model for training, and the trained machine learning model is tested using the test set until the accuracy reaches 0.9, thereby obtaining the fault judgment model.

[0093] In this embodiment, the fault identifier is obtained and converted into a normal identifier, and the data of both historical faults and normal operation are included in the scope of model construction. The normal identifier after binary conversion forms a sharp contrast with the fault identifier, so that the model can more accurately distinguish between faults and normal states, which helps the model learn the key differences between the two states. The structured data is divided into training sets and test sets in a ratio of 7:3, which can provide sufficient data for the machine learning model to learn the fault mode, the normal operation mode and the complex relationship between the two. The generalization ability of the model is independently evaluated to avoid overfitting the model to the training data, and ensure that the model can also perform well when facing new data. The trained model is tested continuously using the test set, forming an efficient generation optimization process. The result obtained from each test is feedback on the model performance. If the accuracy is not up to standard, the model can be adjusted in a targeted manner, the number of hidden layers can be increased, and the learning rate can be adjusted until the accuracy requirement of 0.9 is met. This ensures that the final fault judgment model has strong stability and reliability, and reduces the waste of operation and maintenance costs and time loss caused by misjudgment.

[0094] The distances between the first classification fault judgment value, the second classification fault judgment value, and the third classification fault judgment value and a preset classification fault judgment value threshold are calculated to obtain first fault data, wherein the mathematical expression of the first fault data is:

[0095] P=Σ(W1-W2)+(R1-R2)+(T1-T2);

[0096] Among them, P represents the first fault data, W1 represents the first category fault judgment value, R1 represents the first category fault judgment value, T1 represents the first category fault judgment value, W2 represents the first category fault judgment value threshold, R2 represents the first category fault judgment value threshold, and T2 represents the first category fault judgment value threshold.

[0097] In this embodiment, by calculating the distance between the fault judgment value and the preset threshold, the priority of different categories of faults can be accurately determined. For example, the first category fault judgment value is closer to the preset threshold, while the second and third categories are relatively far away. This indicates that the first category fault is closer to the state where a fault may occur and should be given priority attention. This helps operation and maintenance personnel to quickly locate the most likely type of fault among many potential faults, take targeted measures in advance, avoid the actual occurrence of faults or reduce the possible impact of faults. By quantifying the risk assessment of hardware failures, operation and maintenance personnel can more accurately understand the health of the system and reasonably allocate resources for fault prevention and repair.

[0098] The impact degree of the historical fault event data on the performance and function of the computer system is divided into three impact levels and assigned values, specifically including:

[0099] If the historical fault event data causes the system to be completely inoperable, and the data is lost and cannot be recovered, the impact level is the first level impact level, and the value is 10;

[0100] If the historical fault event data causes the system performance to drop by more than 50%, and some important data is lost but recoverable, the impact level is the second level, and the value is 7;

[0101] If the historical fault event data causes the system performance to drop by less than 20% and there is no data loss, the impact level is the third level of impact severity and is assigned a value of 3;

[0102] The degree of difficulty of the detection based on the historical fault event data is divided into two difficulty levels and the value assignment specifically includes:

[0103] If the historical fault event data can be detected by the detection tool provided by the system, the difficulty level is the first level, and a value of 1 is assigned;

[0104] If the historical fault event data cannot be detected by the detection tools, complex professional equipment and long-term detection are required, the difficulty level is the second level, and the value is 9.

[0105] In this embodiment, the impact of the historical fault event data on the performance and function of the computer system is divided into three impact levels and assigned values, which can determine the priority of fault handling. When multiple hardware faults occur at the same time, the operation and maintenance personnel can quickly determine which faults need to be handled immediately based on the assigned severity, which helps to formulate maintenance plans and strategies. The difficulty of detecting the historical fault event data is divided into two difficulty levels and assigned values, which can help determine the order of maintenance plans. When facing multiple hardware faults, giving priority to faults with low detection difficulty can restore some system functions more quickly and reduce system downtime.

[0106] A computer hardware fault analysis system comprises a data collection module, a fault data calculation module and a fault judgment module;

[0107] The data collection module is used to obtain historical fault event data of computer hardware and perform pre-processing and classification processing;

[0108] The fault data calculation module is used to calculate the first fault data and the second fault data of the historical fault event data based on the fault judgment model and the FMEAC algorithm;

[0109] The fault judgment module is used to calculate comprehensive fault data of the first fault data and the second fault data, and use the comprehensive fault data to judge whether a fault event occurs.

[0110] By calculating the risk priority data and second fault data corresponding to the fault feature data through the FMEAC algorithm, the mode, impact and harmfulness of hardware failures can be systematically analyzed, providing a comprehensive fault knowledge base for the machine learning algorithm. By outputting the first fault data corresponding to the fault feature data through the machine learning algorithm, the first fault data and the second fault data are combined to analyze computer hardware failures, which can more accurately detect and diagnose hardware failures, reduce false alarms and missed alarms, identify potential abnormal signs before failures occur, and issue early warnings, so that operation and maintenance personnel have enough time to take preventive measures to avoid failures or reduce the impact of failures. New failure modes can be promptly incorporated into the entire failure analysis system, making the failure analysis more comprehensive and accurate.

[0111] It should be understood by those skilled in the art that the embodiments of the present invention can be provided as methods, systems or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program codes. Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic memory, flash memory, disk or optical disk. These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0112] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A computer hardware failure analysis method, characterized in that: include: Step S100, obtaining historical fault event data of computer hardware, preprocessing and classifying the historical fault event data, extracting features of the preprocessed and classified historical fault event data, and obtaining fault feature data; Step S200, inputting the fault feature data into a fault judgment model established based on a machine learning algorithm, and outputting a classified fault judgment value, wherein the classified fault judgment value includes a first classified fault judgment value, a second classified fault judgment value, and a third classified fault judgment value; Step S300, calculating the distances between the first classification fault judgment value, the second classification fault judgment value and the third classification fault judgment value and a preset classification fault judgment value threshold to obtain first fault data; Step S400, calculating the risk priority data corresponding to the fault feature data based on the FMEAC algorithm, calculating the distance between the risk priority data and a preset risk priority data threshold, and obtaining second fault data; Step S500: Calculate comprehensive fault data of the first fault data and the second fault data, and use the comprehensive fault data to determine whether a fault event occurs.

2. A computer hardware failure analysis method as claimed in claim 1, characterized in that: Step S100 specifically includes: Step S101, querying the computer's fault log to obtain historical fault event data of the computer hardware, the historical fault event data including hardware temperature rise, hardware power short circuit, hardware power disconnection and hardware device running speed reduction, and assigning a fault identifier to the historical fault event data; Step S102, preprocessing the historical fault event data, wherein the preprocessing includes removing invalid data, incomplete data and duplicate data of the historical fault event data, and classifying the preprocessed historical fault event data based on temperature category, current category and operating parameter category; Step S103, extracting characteristic data of the historical fault event data after the classification processing, obtaining temperature characteristic quantities, current characteristic quantities and operating parameter characteristic quantities under the temperature category, current category and operating parameter category, and normalizing the temperature characteristic quantities, current characteristic quantities and operating parameter characteristic quantities.

3. A computer hardware failure analysis method as claimed in claim 2, characterized in that: Step S200 specifically includes: Step S201, using one-hot encoding to encode the normalized temperature feature quantity, current feature quantity and operating parameter feature quantity; Step S202, input the encoded temperature characteristic quantity, current characteristic quantity and operating parameter characteristic quantity into the fault judgment model, and output the classified fault judgment value, which includes the first classified fault judgment value, the second classified fault judgment value and the third classified fault judgment value.

4. A computer hardware failure analysis method as claimed in claim 3, characterized in that: Step S300 specifically includes: Step S301, the preset classification fault judgment numerical threshold includes a first classification fault judgment numerical threshold, a second classification fault judgment numerical threshold and a third classification fault judgment numerical threshold; Step S302, subtracting the first classification fault judgment value threshold from the first classification fault judgment value to obtain first distance data; Subtract the second classification fault judgment value from the second classification fault judgment value threshold to obtain second distance data; Subtract the third classification fault judgment value from the third classification fault judgment value threshold to obtain third distance data; Step S303: Add the first distance data, the second distance data and the third distance data to obtain first fault data.

5. A computer hardware failure analysis method as claimed in claim 4, characterized in that: Step S400 specifically includes: Step S401, retrieve the classified historical fault event data, divide the historical fault event data into three impact levels based on the impact degree of the historical fault event data on the performance and function of the computer system, and assign values ​​to obtain impact assignment data; Step S402, when the historical fault event data reaches a preset number threshold, the frequency of the historical fault event data of the corresponding temperature category, current category and operating parameter category is counted to obtain frequency data; Step S403, retrieving the classified historical fault event data, dividing the historical fault event data into two difficulty levels based on the difficulty of detecting the historical fault event data, and assigning values ​​to obtain difficulty assignment data; Step S404, adding the impact value assignment data, frequency data and difficulty value assignment data to obtain the risk priority data, calculating the distance between the risk priority data and a preset risk priority data threshold to obtain second fault data.

6. A computer hardware failure analysis method as claimed in claim 5, characterized in that: Step S500 specifically includes: Step S501, obtaining first weight data and second weight data corresponding to the historical fault event data; Step S502, calculating comprehensive fault data based on the first weight data and the second weight data, the mathematical expression of the comprehensive fault data is: Comprehensive fault data = first weight data multiplied by first fault data + second weight data multiplied by second fault data; Step S503, calculating the distance between the comprehensive fault data and a preset comprehensive fault data threshold to obtain total fault data, and using the total fault data to determine whether a fault event occurs.

7. A computer hardware failure analysis method as claimed in claim 6, characterized in that: The method for establishing the fault judgment model specifically includes: Obtaining a fault identifier corresponding to the preprocessed historical fault event data, performing binary conversion on the fault identifier, and obtaining a normal identifier corresponding to the historical normal event data; The annotated historical fault event data and historical normal event data are divided into training set and test set in a ratio of 7:3; The training set is input into the machine learning model for training, and the trained machine learning model is tested using the test set until the accuracy reaches 0.9, thereby obtaining the fault judgment model.

8. A computer hardware failure analysis method as claimed in claim 7, characterized in that: The distances between the first classification fault judgment value, the second classification fault judgment value, and the third classification fault judgment value and a preset classification fault judgment value threshold are calculated to obtain first fault data, wherein the mathematical expression of the first fault data is: P=Σ(W1-W2)+(R1-R2)+(T1-T2); Among them, P represents the first fault data, W1 represents the first category fault judgment value, R1 represents the first category fault judgment value, T1 represents the first category fault judgment value, W2 represents the first category fault judgment value threshold, R2 represents the first category fault judgment value threshold, and T2 represents the first category fault judgment value threshold.

9. A computer hardware failure analysis method as claimed in claim 8, characterized in that: The impact degree of the historical fault event data on the performance and function of the computer system is divided into three impact levels and assigned values, specifically including: If the historical fault event data causes the system to be completely inoperable, and the data is lost and cannot be recovered, the impact level is the first level impact level, and the value is 10; If the historical fault event data causes the system performance to drop by more than 50%, and some important data is lost but recoverable, the impact level is the second level, and the value is 7; If the historical fault event data causes the system performance to drop by less than 20% and there is no data loss, the impact level is the third level of impact severity and is assigned a value of 3; The degree of difficulty of the detection based on the historical fault event data is divided into two difficulty levels and the value assignment specifically includes: If the historical fault event data can be detected by the detection tool provided by the system, the difficulty level is the first level, and a value of 1 is assigned; If the historical fault event data cannot be detected by the detection tools, complex professional equipment and long-term detection are required, the difficulty level is the second level, and the value is 9.

10. A computer hardware failure analysis system, comprising a computer hardware failure analysis method as claimed in any one of claims 1 to 9, characterized in that: It includes a data collection module, a fault data calculation module and a fault judgment module; The data collection module is used to obtain historical fault event data of computer hardware and perform pre-processing and classification processing; The fault data calculation module is used to calculate the first fault data and the second fault data of the historical fault event data based on the fault judgment model and the FMEAC algorithm; The fault judgment module is used to calculate comprehensive fault data of the first fault data and the second fault data, and use the comprehensive fault data to judge whether a fault event occurs.

Citation Information

Patent Citations

  • Large host hardware fault identification method and device, equipment and storage medium

    CN113537349A