Computer system fault detection method and device based on AI
Through the AI-based computer system fault detection method, computer hardware and network information can be obtained and processed in real time, contributed parameter links, and positioned the root cause of failure by the root cause of the traceability module, the problem of inaccurate fault positioning in the existing technology is solved, and efficient and accurate fault detection and analysis is achieved.
Patent Information
- Application Number
- CN202510280836.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing computer fault detection methods are difficult to accurately determine the root cause of the problem, and require manual analysis, which has problems such as poor interpretation and inaccurate positioning.
The AI-based computer system fault detection method is adopted, and the computer hardware and network information are obtained in real time, and after cleaning and processing, the hardware and network monitoring module are used to perform fault detection, the contribution parameter link is calculated, and the root cause of failure is located through the root cause association traceability module.
It realizes accurate positioning and root cause analysis of computer system failures, improves the accuracy and efficiency of fault detection, reduces manual intervention, and provides clear fault location and repair guidance.
Smart Images

Figure CN120216238A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer fault detection, and specifically to an AI-based computer system fault detection method and device. Background Art
[0002] The stability and reliability of computer systems are crucial to all fields of modern society. Especially in scenarios such as enterprise operation and maintenance, financial transactions, medical systems, industrial control, and cloud computing, which highly rely on computers, faults in computer systems may lead to serious consequences. For example, faults in industrial control systems may cause production line explosions (such as the control failure accident in a US refinery in 2010), and faults in autonomous driving systems may lead to traffic accidents, etc. Therefore, it is very necessary to detect computer faults.
[0003] Most of the existing computer fault detection methods can only detect faults, but cannot accurately locate the root cause of the problem, and rely on manual analysis, resulting in serious limitations. For example, black box models (such as deep neural networks, random forests, etc.) are difficult to explain the root cause of faults, hindering manual intervention for verification. There are cases: (1) A bank's fault detection misjudged a hard disk fault. Due to the lack of interpretability, for safety reasons, the operation and maintenance personnel were forced to replace all hardware; (2) The system detected an abnormal increase in the CPU load of a certain server, but could not determine whether it was caused by process leakage, insufficient memory, or external attacks. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides an AI-based computer system fault detection method and device, which solves the technical problem of difficult fault diagnosis and root cause analysis proposed in the above background art.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions:
[0006] In the first aspect, the present invention provides an AI-based computer system fault detection method, including the following steps:
[0007] Step 1: Communicate with each computer in real-time to obtain the hardware information and network information of each computer, perform cleaning, denoising, missing value filling, and normalization processing on the hardware information and network information, and then send them to the transceiver module in the computer system fault detection device for storage;
[0008] Step 2: Send the hardware information and network information of the computer to Step 3 and Step 4 respectively;
[0009] Step 3: The hardware monitoring module performs fault detection on each hardware of the computer based on the hardware information to obtain the detection values of the computer for each hardware, and determines the risky hardware accordingly; calculates the contribution degree of each monitoring parameter of the risky hardware to output the contribution parameter link of the risky hardware, and sends it to Step 5;
[0010] Step 4: The network monitoring module performs fault detection on each network component of the computer based on the network information to obtain the detection values of the computer for each network component, and determines the risky components accordingly; calculates the contribution degree of each monitoring parameter of the risky components to output the contribution parameter link of the risky components, and sends it to Step 5;
[0011] Step 5: Based on the received contribution parameter link of the risky hardware and the contribution parameter link of the risky components, perform root cause association and tracing to output the hardware fault link and the network fault link.
[0012] Further, the hardware information includes each hardware of the computer and the monitoring data corresponding to each hardware, and marks it as Yj, specifically where n is a positive integer, j = 1, 2, 3... J, J is a positive integer, and j represents the label of any one of the hardware; the network information includes each network component and the monitoring data corresponding to each network component, and marks it as Wi, specifically where n is a positive integer, i = 1, 2, 3... I, I is a positive integer, and i represents the label of any one of the network components.
[0013] Further, the determination method of the risky hardware is as follows:
[0014] Obtain the hardware configuration information of the computer, formulate corresponding standard parameters for each hardware of the computer according to the hardware configuration information, and then perform normalization processing to ensure that the formats of all data sources are consistent, and record the normal values of each monitoring parameter of each hardware as HYj, specifically
[0015]
[0016] Perform a difference calculation one by one between the monitoring data of each hardware and its corresponding normal value of the monitoring parameter to obtain the difference value of each monitoring parameter of each hardware. The specific difference calculation is as follows:
[0017] If the normal value corresponding to a certain parameter in the monitoring parameter is of the highest limit type, subtract the corresponding normal value from the parameter for the difference calculation; if the normal value corresponding to a certain parameter in the monitoring parameter is of the lowest limit type, subtract the parameter from the corresponding normal value for the difference calculation;
[0018] Extract the difference values obtained by performing the difference calculation between each monitoring parameter and its corresponding normal value respectively, and record them as C Y j, where Then, integrate and calculate the difference value C corresponding to each monitoring parameter of the hardware Y j to obtain the deviation value corresponding to the hardware, denoted as PC Y j, and the integration calculation formula is:
[0019]
[0020] where is the rectified linear unit function. Specifically
[0021] Obtain the real-time deviation values of each hardware, construct a discrete graph of the real-time deviation values, and use the standard deviation calculation formula to obtain the variance of the discrete graph, denoted as σYj; select the maximum deviation value in the discrete graph, denoted as αYj, and the minimum deviation value, and calculate the difference between the maximum deviation value and the minimum deviation value to obtain the extreme value of the discrete graph, denoted as γYj; calculate the average value of the deviation values in the discrete graph to obtain the deviation mean of the discrete graph, denoted as βYj; calculate the detection value of the hardware, denoted as SYj, through the formula by using the variance, extreme value, and deviation mean. The formula is: where is the standard deviation threshold; obtain the detection values of each hardware in the computer, and compare them with the set detection thresholds respectively. If the detection value is greater than or equal to the set detection threshold, then mark the hardware as a risky hardware.
[0022] Furthermore, the contribution parameter link output method of the risky hardware is:
[0023] Extract the difference values of each monitoring parameter of the risky hardware at each monitoring moment, and the contribution degree A of each monitoring parameter to the risky hardware Y j, where The specific contribution degree calculation formula is: where t is the monitoring moment, and T is the total time dimension of the discrete graph; compare the contribution degrees of each monitoring parameter with the set contribution threshold. If it is greater than the set contribution threshold, then mark the parameter corresponding to the contribution degree as an effective contribution parameter, and sort the effective contribution parameters in descending order according to their corresponding contribution degrees to form the contribution parameter link of the risky hardware; obtain the contribution parameter links of each risky hardware.
[0024] Furthermore, the judgment method of the risky hardware is:
[0025] Obtain the network component configuration information of the computer, formulate corresponding standard parameters for each network component of the computer according to the network component configuration information, and then perform normalization processing to ensure that the formats of all data sources are consistent, and the normal values HWi of each monitoring parameter of each network component. Specifically
[0026]
[0027] The difference values of each monitoring parameter of each network component are calculated one by one by comparing the monitoring data of each network component with the normal value of its corresponding monitoring parameter. The specific difference calculation is as follows: If the normal value corresponding to a certain parameter in the monitoring parameter is of the highest limit type, then subtract the parameter from the corresponding normal value for differential calculation; if the normal value corresponding to a certain parameter in the monitoring parameter is of the lowest limit type, then subtract the parameter from the corresponding normal value of the parameter for differential calculation; respectively extract the difference values obtained by the differential calculation of each monitoring parameter and its corresponding normal value, and record them as D W i, where Then, the difference values DWi corresponding to each monitoring parameter of the hardware are integrated and calculated to obtain the deviation value corresponding to the hardware, which is marked as PDYj. The integration calculation formula is:
[0028]
[0029] where is a linear rectification function. Specifically even when the difference value is less than or equal to zero, the numerator in the integration calculation formula takes zero, otherwise it remains the original value;
[0030] The real-time deviation values of each network component are obtained, a discrete graph of the real-time deviation values is constructed, and the variance of the discrete graph is obtained using the standard deviation calculation formula and recorded as σWi; the maximum deviation value in the discrete graph is selected and recorded as αWi and the minimum deviation value, and the difference between the maximum deviation value and the minimum deviation value is calculated to obtain the extreme value of the discrete graph, which is recorded as γWi; the deviation values in the discrete graph are averaged to obtain the deviation mean of the discrete graph, which is recorded as βWi; the variance, extreme value, and deviation mean are calculated through a formula to obtain the detection value of the hardware, which is recorded as SWi. The formula is: where is the standard deviation threshold; the detection values of each network component in the computer are obtained and compared with the set detection threshold respectively. If the detection value is greater than or equal to the set detection threshold, the network component is recorded as a risk component.
[0031] Furthermore, the contribution parameter link output mode of the risk hardware is:
[0032] Extract the difference values of each monitoring parameter of the risk component at each monitoring moment and the contribution degree of each monitoring parameter to the risk component; compare the contribution degree of each monitoring parameter with the set contribution threshold. If it is greater than the set contribution threshold, the parameter corresponding to the contribution degree is recorded as an effective contribution parameter, and each effective contribution parameter is sorted in descending order according to its corresponding contribution degree to form the contribution parameter link of the risk component; the contribution parameter links of each risk hardware are obtained.
[0033] Furthermore, the specific way to output the hardware fault link by root cause correlation tracing is:
[0034] S1: Construct a mapping relationship library between hardware monitoring parameters and potential fault causes;
[0035] S2: For each contribution parameter link of the risk hardware, match the root cause item by item, extract the root cause and weight of each effective contribution parameter from the mapping relationship library, multiply the weight of the root cause by the contribution degree of the corresponding effective contribution parameter to obtain the confidence level of the root cause, combine the confidence levels of the same root cause, sort by the total value to generate a root cause link; use Grafana to display the root cause confidence stack chart and real-time mark the main fault path;
[0036] S3: Combine the timing data and hardware dependency relationship to verify the rationality of the root cause chain and generate the final hardware fault link; output the final hardware fault link and the root cause confidence stack chart as the detection result of the risk hardware.
[0037] Furthermore, the specific way of root cause association and traceability to output the network fault link is as follows:
[0038] K1: Establish a mapping relationship library between network monitoring parameters and potential fault causes;
[0039] K2: For each contribution parameter link of the risk network component, match the root cause item by item, extract the root cause and weight of each effective contribution parameter from the mapping relationship library, multiply the weight of the root cause by the contribution degree of the corresponding effective contribution parameter to obtain the confidence level of the root cause, combine the confidence levels of the same root cause, sort by the total value to generate a root cause link; use Grafana to display the root cause confidence stack chart and real-time mark the main fault path;
[0040] K3: Combine the timing data and network topology dependency relationship to verify the rationality of the root cause and generate the final network fault link; output the final fault link and the root cause confidence stack chart as the detection result of the risk component
[0041] In a second aspect, the present invention provides an AI-based computer system fault detection device, including a transceiver module, a hardware monitoring module, a network monitoring module, and a root cause association and traceability module.
[0042] Further, the transceiver module receives and stores the hardware information and network information, and separately sends them to the hardware monitoring module and the network monitoring module; the hardware monitoring module performs fault detection on each hardware of the computer based on the hardware information, obtains the contribution parameter links of each risky hardware, and sends each risky hardware and its corresponding contribution parameter link to the root cause association and traceability module; the network monitoring module performs fault detection on each network component of the computer based on the network information, obtains the contribution parameter links of each risky component, and sends each risky component and its corresponding contribution parameter link to the root cause association and traceability module; the root cause association and traceability module performs root cause association and traceability based on the received contribution parameter links of the risky hardware and the contribution parameter links of the risky components to output the hardware fault link and the network fault link.
[0043] The present invention has the following beneficial effects:
[0044] 1. Based on the hardware configuration information and network component configuration information of the computer, the standard values of each monitoring parameter of the hardware and network components are formulated in the normal network state. By comparing the real-time monitoring data with the standard values, the difference values are calculated respectively according to the "highest limit class" and "lowest limit class", and after being processed by the rectification function, they are integrated into the deviation value of the network component; the standard deviation, extreme value and mean value are calculated by using the scatter plot, and the detection values of the hardware and network components are calculated and analyzed accordingly to judge the risky hardware and risky components; the contribution degrees of each monitoring parameter of the risky hardware and risky components are calculated at each monitoring moment, and after screening out the effective contribution parameters, they are sorted according to the contribution degree to form the contribution parameter links of the risky hardware and risky network components; it can monitor and detect in real time that there are abnormal risks in the risky hardware and network components, and provide quantitative indicators for the early warning of hardware faults and network faults; the contribution parameter links clearly show the key factors leading to hardware faults and network faults, provide data support and decision-making basis for the subsequent root cause tracing, and improve the accuracy of hardware fault and network fault location;
[0045] 2. Through the root cause mapping, confidence aggregation and time sequence verification of the contribution parameter link, the root causes of the hardware faults and network faults of the computer are accurately located. The final output result needs to include operable repair suggestions to directly guide the operation and maintenance response; it can accurately map the abnormal contribution of the risky component to the specific fault root cause, form a fault link sorted by possibility, and provide clear fault location and repair guidance for the operation and maintenance personnel; when actually deployed, it is necessary to dynamically optimize the model by combining the domain knowledge base and real-time monitoring data to improve the accuracy of the fault link.
[0046] Through the three-stage architecture of multi-dimensional detection → dynamic association → probability reasoning, the problem of fuzzy positioning in traditional methods is effectively solved. Description of the Drawings
[0047] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0048] Figure 1 is the flowchart of the method of the present invention;
[0049] Figure 2 is the connection relationship diagram of the device of the present invention. Specific embodiments
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0051] Please refer to Figure 1 , the present invention provides a technical solution: an AI-based computer system fault detection method, including the following steps:
[0052] Step 1: The computer system fault detection device is communicatively connected to each computer to obtain the hardware information and network information of each computer in real time, and perform cleaning, denoising, missing value filling, and normalization processing on the hardware information and network information to ensure that the formats of all data sources are consistent and the time series are synchronized; then send it to the transceiver module in the computer system fault detection device for storage; the specific hardware information includes each hardware and the monitoring data corresponding to each hardware, and mark it as Yj, specifically where n is a positive integer, j = 1, 2, 3... J, J is a positive integer, and j represents the label of any one of the hardware; for example: the main hardware monitored by the computer is CPU (Central Processing Unit), memory, storage device, motherboard, power supply, GPU (Graphics Processing Unit), etc.; specifically: if the label of the CPU is 1, then the monitoring data of the CPU includes CPU temperature, CPU usage rate, frequency, voltage, cache error rate; the network information includes each network component and the monitoring data corresponding to each network component, and mark it as Wi, specifically where n is a positive integer, i = 1, 2, 3... I, I is a positive integer, and i represents the label of any one of the network components; for example: the network components monitored by the computer are mainly network interfaces, switches / routers, transport layer protocols, application layer protocols, wireless networks, network security, etc., specifically: the monitoring parameters of the network interface are broadband utilization rate, packet loss rate, error frame count;
[0053] By storing data in the transceiver module and distributing it to the subsequent hardware monitoring module and network monitoring module as expected; after data preprocessing, noise and outlier interference can be eliminated, ensuring the accuracy and consistency of monitoring metrics and laying a solid data foundation for fault detection.
[0054] Step 2: The computer fault detection device includes a transceiver module, a hardware monitoring module, a network monitoring module, a root cause association module, and a root cause tracing module; the transceiver module stores the hardware information and network information of the computer and sends them to the hardware monitoring module and the network monitoring module respectively;
[0055] Step 3: The hardware monitoring module performs fault detection on each hardware of the computer based on the hardware information. Specifically:
[0056] Obtain the hardware configuration information of the computer. The hardware configuration information includes CPU, motherboard, GPU model, interface type of storage device, motherboard chipset model, GPU video memory capacity, memory capacity, type, etc. Usually, a computer comes with a configuration manual that has detailed descriptions of the computer configuration. Here, no examples will be given one by one; based on the hardware configuration information, corresponding standard parameters are formulated for each hardware of the computer. The specific standard parameters refer to the normal values of each monitoring parameter of each hardware in the normal state (the normal value here refers to the lowest limit or the highest limit in the normal state. For example, the normal value of the CPU temperature is often the highest limit value of the CPU temperature when the computer is normal, and the highest limit is 50 °C; the normal value of the CPU usage rate is the lowest limit value of the CPU usage rate in the normal situation, usually 95%, and less than 95% indicates that there may be a resource bottleneck in the CPU). Then, it is normalized to ensure that the formats of all data sources are consistent, and the normal values of each monitoring parameter of each hardware are recorded as HYj. Specifically
[0057]
[0058] Perform a difference calculation for each monitoring parameter of each hardware one by one with its corresponding normal value of the monitoring parameter to obtain the difference value of each monitoring parameter of each hardware. The specific difference calculation is as follows: If the normal value corresponding to a certain parameter in the monitoring parameter is of the highest limit type, then subtract the corresponding normal value from the parameter for the difference calculation; if the normal value corresponding to a certain parameter in the monitoring parameter is of the lowest limit type, then subtract the parameter from the corresponding normal value of the parameter for the difference calculation; respectively extract the difference values obtained by performing the difference calculation for each monitoring parameter and its corresponding normal value, and record them as C Y j, where Then, integrate and calculate the difference values C Y j corresponding to each monitoring parameter of the hardware to obtain the deviation value corresponding to the hardware, marked as PC Yj, the integration calculation formula is:
[0059]
[0060] where is the rectified linear unit function. Specifically, even when the difference value is less than or equal to zero, the numerator in the integration calculation formula takes zero, otherwise it remains the original value; among the parameters of the computer, if a parameter is less than the highest limit value, for example, the CPU temperature is less than 50 °C, indicating that the CPU temperature is normal, then the difference value corresponding to the CPU temperature is less than zero, and the numerator of the term corresponding to the CPU temperature in the corresponding integration calculation formula is zero, meaning that at this time the CPU temperature term has no contribution to the deviation value;
[0061] Obtain the real-time deviation values of each hardware, construct a discrete graph of the real-time deviation values, and use the standard deviation calculation formula to obtain the variance of the discrete graph, denoted as σYj; select the maximum deviation value in the discrete graph, denoted as αYj and the minimum deviation value, and perform a difference calculation on the maximum deviation value and the minimum deviation value to obtain the extreme value of the discrete graph, denoted as γYj; perform an average calculation on the deviation values in the discrete graph to obtain the deviation average value of the discrete graph, denoted as βYj; calculate the detection value of the hardware through the formula using the variance, extreme value, and deviation average value, denoted as SYj, and the formula is: where is the standard deviation threshold. Specifically, the calculation method of the standard deviation threshold is the mean method or the percentile method. The mean method is: The percentile method is: Obtain the detection values of each hardware in the computer, and compare them with the set detection thresholds respectively. If the detection value is greater than or equal to the set detection threshold, it means that the risk of the hardware being abnormal is greater, and then the hardware is recorded as a risk hardware;
[0062] Extract the difference values of each monitoring parameter of the risk hardware at each monitoring moment, and the contribution degree A Y j of each monitoring parameter to the risk hardware, where The specific contribution degree calculation formula is: where t is the monitoring moment (i.e., a single time dimension), and T is the total time dimension of the discrete graph; compare the contribution degrees of each monitoring parameter with the set contribution threshold. If it is greater than the set contribution threshold, then the parameter corresponding to the contribution degree is recorded as an effective contribution parameter, and each effective contribution parameter is sorted in descending order according to its corresponding contribution degree to form a contribution parameter link of the risk hardware; obtain the contribution parameter links of each risk hardware;
[0063] Send each risk hardware and its corresponding contribution parameter link to the root cause association and traceability module;
[0064] By formulating the monitoring parameter standards of each hardware in the normal state according to the hardware configuration information, it is ensured that the normal values of each hardware accurately reflect the device characteristics; the differences between the collected real-time monitoring data and the standard values are calculated one by one, and the linear rectification function is used to filter the negative deviations in the normal state, and the real-time deviation values of the hardware are integrated; the hardware detection values are calculated through statistical analysis of the scatter plot, and the risk hardware with detection values exceeding the threshold can be found in time; at the same time, the difference values of each parameter of the risk hardware at each monitoring moment are extracted, the contribution degree of each parameter is calculated, and the contribution parameter link is formed by sorting according to the contribution degree; the devices with abnormal risks in the hardware can be accurately detected, and the contribution of each hardware parameter to the fault risk can be quantitatively evaluated. The contribution parameter link provides a detailed basis for subsequent root cause association, supports rapid fault cause location, and effectively assists in operation and maintenance decision-making.
[0065] Step 4: The network monitoring module performs fault detection on each network component of the computer based on network information, specifically:
[0066] Obtain the network component configuration information of the computer. The network component configuration information includes the operating system version, kernel version, system architecture, CPU, motherboard, GPU model, interface type of the storage device, motherboard chipset model, GPU video memory capacity, memory capacity, program driver version number, compatibility report (such as WHQL certification), security patch status, etc. Usually, the computer comes with a configuration manual, which contains detailed descriptions of the computer network security component configuration. Here, no examples will be given one by one; according to the network component configuration information, corresponding standard parameters are formulated for each network component of the computer. The specific standard parameters refer to the normal values of each monitoring parameter of each network component in the normal state of the computer network (the normal value here refers to the lowest limit or the highest limit in the normal state. For example, the normal value of the broadband utilization rate refers to the highest limit value in the normal network situation of the computer, and usually the highest limit value is 70%; the normal value of the signal strength refers to the lowest limit in the normal network situation of the computer, and usually the lowest limit value is -70dBm; when the strength is less than -70dBm, the signal strength is poor, resulting in unstable computer connection, network speed decline, etc.). Then, it is normalized to ensure that the formats of all data sources are consistent, and the normal values of each monitoring parameter of each network component are recorded as HWi. Specifically
[0067] Calculate the difference value of each monitoring parameter of each network component by calculating the difference between the monitoring data of each network component and the normal value of its corresponding monitoring parameter one by one. The specific difference calculation is as follows: If the normal value corresponding to a certain parameter in the monitoring parameter is of the highest limit type, then subtract the corresponding normal value from the parameter for differential calculation; if the normal value corresponding to a certain parameter in the monitoring parameter is of the lowest limit type, then subtract the parameter from the corresponding normal value of the parameter for differential calculation; respectively extract each monitoring parameter and its corresponding normal value for differential calculation to obtain the difference value, and record it as D W i, where Then integrate and calculate the difference values D W i corresponding to each monitoring parameter of the hardware to obtain the deviation value corresponding to the hardware, marked as PD Y j. The integration calculation formula is:
[0068]
[0069] where is the linear rectifier function. Specifically even when the difference value is less than or equal to zero, the numerator in the integration calculation formula takes zero, otherwise it remains the original value;
[0070] Obtain the real-time deviation value of each network component, construct a discrete graph of the real-time deviation value, and use the standard deviation calculation formula to obtain the variance of the discrete graph, denoted as σWi; select the maximum deviation value in the discrete graph, denoted as αWi and the minimum deviation value, and calculate the difference between the maximum deviation value and the minimum deviation value to obtain the extreme value of the discrete graph, denoted as γWi; calculate the average value of the deviation values in the discrete graph to obtain the deviation average value of the discrete graph, denoted as βWi; calculate the detection value of the hardware through the formula by using the variance, extreme value and deviation average value, and the formula is: where is the standard deviation threshold. The specific calculation method of the standard deviation threshold is the mean method or the percentile method. The mean method is: The percentile method is: Obtain the detection values of each network component in the computer and compare them with the set detection threshold respectively. If the detection value is greater than or equal to the set detection threshold, it means that the risk of abnormality of the network component is greater, then mark the network component as a risk component;
[0071] Extract the difference values of each monitoring parameter of the risk component at each monitoring moment, and the contribution degree B W i of each monitoring parameter to the risk component, where The specific contribution degree calculation formula is: where \(t\) is the monitoring time (i.e., a single time dimension), and \(T\) is the total time dimension of the discrete graph; compare the contribution degrees of each monitoring parameter with the set contribution threshold. If it is greater than the set contribution threshold, the parameter corresponding to this contribution degree is recorded as an effective contribution parameter, and each effective contribution parameter is sorted in descending order according to its corresponding contribution degree to form the contribution parameter link of the risk component; obtain the contribution parameter links of each risk component.
[0072] Send each risk component and its corresponding contribution parameter link to the root cause association and tracing module.
[0073] Based on the network component configuration information, the standard values of each monitoring parameter of the network component in the normal network state are formulated. By comparing the real-time monitoring data with the standard values, the difference values are calculated separately according to the "highest limit type" and "lowest limit type", and after being processed by the rectification function, they are integrated into the deviation value of the network component; use the discrete graph to calculate the standard deviation, extreme value and mean value, and calculate and analyze accordingly to obtain the detection value of the network component, and judge the risk component; calculate the contribution degree of each monitoring parameter of the risk component at each monitoring time, screen out the effective contribution parameters and sort them according to the contribution degree to form the contribution parameter link of the risk network component; it can monitor and detect the abnormal risks of network components in real time, provide a quantitative index for network fault warning; the contribution parameter link clearly shows the key factors leading to network faults, provides data support and decision-making basis for subsequent root cause tracing, and improves the accuracy of network fault location.
[0074] Step Five: The root cause association and tracing module performs root cause association and tracing based on the received contribution parameter links of risk hardware and risk components to output the hardware fault link and the network fault link.
[0075] The output method of the hardware fault link is as follows:
[0076] S1: Construct a mapping relationship library between hardware monitoring parameters and potential fault causes. It should be noted that this mapping relationship library is usually formulated by technicians in this field according to the industry internal standards and experience. This mapping relationship library is continuously updated and verified with actual cases to ensure that the mapping relationship reflects the latest device characteristics and fault modes; Example: Root causes associated with CPU usage: Root cause a = abnormal process occupying resources, weight: 0.8; Root cause b = application design defect, weight: 0.4.
[0077] S2: For each contribution parameter link of the risk hardware, match the root cause item by item, extract the root cause (one or more) and weight of each valid contribution parameter from the mapping relationship library, multiply the weight of the root cause by the contribution degree of the valid contribution parameter to which it belongs to obtain the confidence level of the root cause, merge the confidence levels of the same root cause (i.e., sum the confidence levels of the same root cause), sort according to the total value to generate the root cause link; use Grafana to display the root cause confidence stack chart and real-time mark the main fault path;
[0078] S3: Combine the timing data and hardware dependency relationship to verify the rationality of the root cause chain and generate the final hardware fault link; output the final hardware fault link and the root cause confidence stack chart as the detection result of the risk hardware to clarify the fault cause in a timely manner and provide an efficient auxiliary function for the operation and maintenance personnel;
[0079] The output method of the network fault link is as follows:
[0080] K1: Establish a mapping relationship library between network monitoring parameters and potential fault causes. Example: Root cause of "bandwidth utilization": Root cause a = network congestion, weight: 0.8; Root cause b = DDoS attack, weight: 0.7; Root cause c = incorrect switch QoS configuration, weight: 0.6;
[0081] K2: For each contribution parameter link of the risk network component (sorted by contribution degree), match the root cause item by item, extract the root cause (one or more) and weight of each valid contribution parameter from the mapping relationship library, multiply the weight of the root cause by the contribution degree of the valid contribution parameter to which it belongs to obtain the confidence level of the root cause, merge the confidence levels of the same root cause (i.e., sum the confidence levels of the same root cause), sort according to the total value to generate the root cause link; use Grafana to display the root cause confidence stack chart and real-time mark the main fault path;
[0082] K3: Combine the timing data and network topology dependency relationship to verify the rationality of the root cause and generate the final network fault link; output the final fault link and the root cause confidence stack chart as the detection result of the risk component to clarify the fault cause in a timely manner and provide an efficient auxiliary function for the operation and maintenance personnel;
[0083] Through the root cause mapping, confidence aggregation and timing verification of the contribution parameter link, the root causes of the hardware faults and network faults of the computer can be accurately located. The final output result should include operable repair suggestions to directly guide the operation and maintenance response; it can accurately map the abnormal contributions of the risk components to specific fault root causes, form a fault link sorted by possibility, and provide clear fault location and repair guidance for the operation and maintenance personnel; during actual deployment, it is necessary to dynamically optimize the model by combining the domain knowledge base and real-time monitoring data to improve the accuracy of the fault link.
[0084] Example:
[0085] Risk components: network interface; Contribution parameter link: [("Bandwidth utilization rate", 65%), ("Packet loss rate", 25%), ("Error frame count", 10%)]
[0086] Root cause confidence calculation:
[0087] Bandwidth utilization rate (65%): (1) Network congestion: 0.65 * 0.8 = 0.52; (2) DDoS attack: 0.65 * 0.7 = 0.455; (3) QoS configuration error: 0.65 * 0.6 = 0.39
[0088] Packet loss rate (25%): (1) Physical link failure: 0.25 * 0.9 = 0.225; (2) Switch buffer overflow: 0.25 * 0.75 = 0.1875
[0089] Error frame count (10%): (1) Electromagnetic interference: 0.1 * 0.7 = 0.07; (2) Network card failure: 0.1 * 0.85 = 0.085;
[0090] Root cause aggregation and sorting:
[0091] 1. Network congestion (0.52)
[0092] 2. DDoS attack (0.455)
[0093] 3. QoS configuration error (0.39)
[0094] 4. Physical link failure (0.225)
[0095] 5. Switch buffer overflow (0.1875)
[0096] 6. Network card hardware failure (0.085)
[0097] 7. Electromagnetic interference (0.07);
[0098] Combined with time series data and network topology dependencies, verify the rationality of the root cause:
[0099] Time series analysis: If the sudden increase in bandwidth utilization rate is earlier than the increase in packet loss rate → support "network congestion" as the root cause; If the error frame count surges when the factory equipment starts at night → support "electromagnetic interference";
[0100] Dependency verification: Check the switch port statistics. If the number of discarded packets is high and the QoS policy is not enabled → confirm "QoS configuration error"; Packet capture analysis of traffic characteristics: Capture traffic and use Wireshark to analyze DDoS characteristics;
[0101] Output network fault link:
[0102] Network Fault Link Report - Risk Component: Network Interface (eth0);
[0103] Root Cause Link (Sorted by Confidence): 1. Network Congestion (Confidence 52%); Evidence Chain: Bandwidth Utilization Continuously > 85% (Threshold 70%); Traffic Analysis Shows That Peer-to-Peer Applications Occupy 60% of the Bandwidth; Suggested Action: Enable Traffic Shaping (Such as tc qdisc Rate Limiting), Optimize Application Bandwidth Allocation Strategy;
[0104] 2. DDoS Attack (Confidence 45.5%); Evidence Chain: Detected Abnormal UDP Flood Traffic (Source IPs Are Scattered); Firewall Logs Show 10,000+ SYN Requests per Second; Suggested Action: Enable Cloud Cleaning Service (Such as Cloudflare); Configure ACL Rules to Block Malicious IP Segments
[0105] 3. QoS Configuration Error (Confidence 39%); Evidence Chain: Packet Drop Rate of Switch Ports > 15%; QoS Policy Does Not Prioritize Critical Services (Such as VoIP); Suggested Action: Configure DiffServ to Mark Critical Traffic; Enable WRED (Weighted Random Early Detection);
[0106] Minor Factors: Physical Link Failure (Check the Oxidation of Network Cable Connectors); Switch Buffer Overflow (Upgrade Switch Firmware and Adjust Buffer Size).
[0107] Please refer to Figure 2 This invention provides a technical solution: an AI-based computer system fault detection device, including a transceiver module, a hardware monitoring module, a network monitoring module, and a root cause association and tracing module.
[0108] Among them, the transceiver module receives and stores hardware information and network information, and sends them to the hardware monitoring module and the network monitoring module respectively; the hardware monitoring module performs fault detection on each hardware of the computer based on the hardware information, obtains the contribution parameter links of each risk hardware, and sends each risk hardware and its corresponding contribution parameter link to the root cause association and tracing module; the network monitoring module performs fault detection on each network component of the computer based on the network information, obtains the contribution parameter links of each risk component, and sends each risk component and its corresponding contribution parameter link to the root cause association and tracing module; the root cause association and tracing module performs root cause association and tracing based on the received contribution parameter links of risk hardware and risk components to output hardware fault links and network fault links.
[0109] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes, shall be covered by the protection scope of the present invention.
Claims
1. A computer system fault detection method based on AI, characterized in that: The following steps are involved: Step 1: Obtain the hardware information and network information of each computer in real time, clean, denoise, fill in missing values and normalize the hardware information and network information, and then send them to the transceiver module for storage; Step 2: Send the computer's hardware information and network information to step 3 and step 4 respectively; Step 3: The hardware monitoring module performs fault detection on each hardware of the computer based on the hardware information to obtain the detection value of each hardware of the computer, and judges the risky hardware accordingly; Calculate the contribution of each monitoring parameter of the risk hardware to output a contribution parameter link about the risk hardware, and send it to step five; Step 4: The network monitoring module performs fault detection on each network component of the computer based on the network information to obtain the detection value of each network component of the computer, and judges the risk component accordingly; Calculate the contribution of each monitoring parameter of the risk component to output a contribution parameter link of the risk component, and send it to step five; Step 5: Based on the received contribution parameter links of the risky hardware and the contribution parameter links of the risky components, root cause correlation tracing is performed to output hardware fault links and network fault links.
2. The AI-based computer system fault detection method according to claim 1, characterized in that: Hardware information includes the computer hardware and the corresponding monitoring data of each hardware, which is marked as Yj. Where n is a positive integer, j represents the label of any hardware; the network information includes each network component and the monitoring data corresponding to each network component, which is marked as Wi. Wherein n is a positive integer, I is a positive integer, and i represents the number of any network component.
3. The AI-based computer system fault detection method according to claim 2, characterized in that: The risky hardware is judged as follows: Obtain the computer's hardware configuration information, formulate corresponding standard parameters for each hardware of the computer according to the hardware configuration information, and then normalize them to ensure that the formats of each data source are consistent, and record the normal values of each monitoring parameter of each hardware as HYj. The difference between the monitoring data of each hardware and the normal value of its corresponding monitoring parameter is calculated one by one to obtain the difference value of each monitoring parameter of each hardware. The specific difference calculation is: If the normal value corresponding to a parameter among the monitoring parameters is of the highest limit type, the corresponding normal value is subtracted from the parameter for differential calculation; if the normal value corresponding to a parameter among the monitoring parameters is of the lowest limit type, the normal value corresponding to the parameter is subtracted from the parameter for differential calculation; The difference value is obtained by extracting each monitoring parameter and the corresponding normal value and calculating the difference, which is recorded as C Y j, where Then the difference value C corresponding to each monitoring parameter of the hardware Y j is integrated and calculated to obtain the deviation value corresponding to the hardware, which is marked as PC Y j, the integrated calculation formula is: in is a linear rectification function, specifically Get the real-time deviation value of each hardware, build a discrete graph of the real-time deviation value, and use the standard deviation calculation formula to get the variance of the discrete graph, which is recorded as σYj; At The maximum deviation value in the discrete graph is selected as αYj and the minimum deviation value, and the difference between the maximum deviation value and the minimum deviation value is calculated to obtain the extreme value of the discrete graph as γYj; The deviation values in the discrete graph are averaged and the deviation mean of the discrete graph is recorded as βYj; the variance, extreme value and deviation mean are calculated by the formula to obtain the hardware detection value recorded as SYj. The formula is: in is the standard deviation threshold; the detection value of each hardware in the computer is obtained, and it is compared with the set detection threshold respectively. If the detection value is greater than or equal to the set detection threshold, the hardware is recorded as risky hardware.
4. The AI-based computer system fault detection method according to claim 3, characterized in that: The contribution parameter link output mode of risk hardware is: Extract the difference value of each monitoring parameter of the risk hardware at each monitoring time, and calculate the contribution of each monitoring parameter to the risk hardware. Y j, where The contribution calculation formula is: Where t is the monitoring time, and T is the total time dimension of the discrete graph; the contribution of each monitoring parameter is compared with the set contribution threshold. If it is greater than the set contribution threshold, the parameter corresponding to the contribution is recorded as the effective contribution parameter, and each effective contribution parameter is sorted from large to small according to its corresponding contribution to form a contribution parameter link of risk hardware; Get the contribution parameter link of each risk hardware.
5. The AI-based computer system fault detection method according to claim 4, characterized in that: The risky hardware is judged as follows: Obtaining the network component configuration information of the computer, formulating corresponding standard parameters for each network component of the computer according to the network component configuration information, and then normalizing the parameters; The monitoring data of each network component and the normal value of the corresponding monitoring parameter are calculated one by one to obtain the difference value of each monitoring parameter of each network component. The specific difference calculation is as follows: if the normal value corresponding to a parameter in the monitoring parameters is the upper limit value class, then the parameter is subtracted from the corresponding normal value for differential calculation; if the normal value corresponding to a parameter in the monitoring parameters is the lower limit value class, then the normal value corresponding to the parameter is subtracted from the parameter for differential calculation; each monitoring parameter is extracted and the corresponding normal value is calculated for differential calculation to obtain the difference value; then the difference values corresponding to each monitoring parameter of the hardware are integrated and calculated to obtain the deviation value corresponding to the hardware; Obtain the real-time deviation value of each network component, construct a discrete graph of the real-time deviation value, and obtain the discrete graph using the standard deviation calculation formula; Select the maximum deviation value and the minimum deviation value in the discrete graph, and perform difference calculation on the maximum deviation value and the minimum deviation value to obtain the extreme value of the discrete graph; The deviation values in the discrete graph are averaged to obtain the deviation mean of the discrete graph; The variance, extreme value and deviation from the mean are calculated by formula to obtain the hardware detection value; The detection value of each network component in the computer is obtained, and it is compared with the set detection threshold. If the detection value is greater than or equal to the set detection threshold, the network component is recorded as a risk component.
6. The AI-based computer system fault detection method according to claim 5, characterized in that: The contribution parameter link output mode of risk hardware is: Extract the difference value of each monitoring parameter of the risk component at each monitoring time, and the contribution of each monitoring parameter to the risk component; compare the contribution of each monitoring parameter with the set contribution threshold. If it is greater than the set contribution threshold, the parameter corresponding to the contribution is recorded as a valid contribution parameter, and the valid contribution parameters are sorted from large to small according to their corresponding contribution to form a contribution parameter link of the risk component; Get the contribution parameter link of each risk hardware.
7. The AI-based computer system fault detection method according to claim 6, characterized in that: The specific method of root cause correlation tracing to output hardware fault links is as follows: S1: Build a mapping relationship library between hardware monitoring parameters and potential fault causes; S2: For each risk hardware contribution parameter link, match the root cause item by item, extract the root cause and weight of each valid contribution parameter from the mapping relationship library, and multiply the weight of the root cause by the contribution of the valid contribution parameter to which it belongs to obtain the confidence of the root cause. Combine the confidences of the same root causes and generate the root cause link by sorting the total values. Use Grafana to display the root cause confidence stacking chart and mark the main fault paths in real time; S3: Combine the timing data with the hardware dependency to verify the rationality of the root cause chain and generate the final hardware failure link; output the final hardware failure link and the root cause confidence stacking diagram as the detection result of the risky hardware.
8. The AI-based computer system fault detection method according to claim 1, characterized in that: The specific method of root cause correlation tracing to output network fault links is as follows: K1: Establish a mapping relationship library between network monitoring parameters and potential fault causes; K2: For each risk network component's contribution parameter link, match the root cause item by item, extract the root cause and weight of each valid contribution parameter from the mapping relationship library, and multiply the root cause weight by the contribution of the valid contribution parameter to which it belongs to obtain the confidence of the root cause. Combine the confidences of the same root causes and generate the root cause link by sorting the total values. Use Grafana to display the root cause confidence stacking chart and mark the main fault paths in real time; K3: Combine time series data with network topology dependencies to verify the rationality of the root cause and generate the final network fault link; output the final fault link and root cause confidence stacking diagram as the detection result of the risk component.
9. An AI-based computer system fault detection device, characterized in that Applied to the AI-based computer system fault detection method as described in any one of claims 1 to 8, the computer system fault detection device comprises: a transceiver module, a hardware monitoring module, a network monitoring module and a root cause correlation tracing module; The transceiver module receives and stores hardware information and network information, and sends them to the hardware monitoring module and the network monitoring module respectively; the hardware monitoring module performs fault detection on each hardware of the computer based on the hardware information, obtains the contribution parameter link of each risk hardware, and sends each risk hardware and its corresponding contribution parameter link to the root cause association tracing module; the network monitoring module performs fault detection on each network component of the computer based on the network information, obtains the contribution parameter link of each risk component, and sends each risk component and its corresponding contribution parameter link to the root cause association tracing module; the root cause association tracing module performs root cause association tracing based on the contribution parameter link of the risk hardware and the contribution parameter link of the risk component received to output the hardware fault link and the network fault link.