Computer intelligent fault diagnosis and early warning system based on artificial intelligence

By designing an intelligent computer fault diagnosis and early warning system based on artificial intelligence, combining hardware and network monitoring data, and calculating fault risks using multiple algorithm units, we finally realize comprehensive monitoring and early warning of computer hardware and network conditions, solving the problem that existing systems are difficult to consider hardware and network factors at the same time, and improving the accuracy and timeliness of fault warnings.

CN120179513AActive Publication Date: 2025-06-20CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510289084.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-20
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing computer intelligent fault diagnosis and early warning system based on artificial intelligence is difficult to consider the hardware operation health status and network broadband usage while the computer is running, and to focus on monitoring and early warning of computers with high fault diagnosis risks under limited resources.

Method used

An intelligent computer fault diagnosis and early warning system based on artificial intelligence is designed, including a data collection module, a data preprocessing module, a computing processing module and a fault diagnosis and early warning module. The hardware monitoring software monitors the CPU temperature, operating memory usage rate and hard disk memory usage rate in real time, and obtains the network broadband usage rate through network monitoring tools. The final value of the failure risk is calculated by combining three algorithm units (the basic value algorithm unit of failure risk, the adjusted failure risk value algorithm unit and the failure risk final value algorithm unit) and fault diagnosis and early warning are carried out based on this value.

Benefits of technology

It realizes intelligent fault diagnosis and early warning based on the hardware operation health status and network broadband usage when the computer is running, which can more comprehensively reflect the load and health status of the computer system, improves the accuracy and timeliness of fault warning, and focuses on monitoring and early warning of computers with high fault diagnosis risks under limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179513A_ABST
    Figure CN120179513A_ABST
Patent Text Reader

Abstract

The invention discloses a computer intelligent fault diagnosis and early warning system based on artificial intelligence, relates to the technical field of computers, and forms a core architecture of the computer intelligent fault diagnosis and early warning system based on artificial intelligence through mutual cooperation of three algorithm units. According to the computer intelligent fault diagnosis and early warning method, double consideration can be carried out during computer intelligent fault diagnosis and early warning according to the hardware operation condition of a computer and the use condition of a network broadband when the computer runs, so that the load and health condition of a computer system can be reflected more comprehensively; an operator or an automatic system is helped to carry out fault positioning and computer repairing more efficiently, a fault risk final value Ffv obtained through calculation is compared with a warning threshold value and a danger threshold value which are set in a database, and when multiple computers are pre-warned at the same time, under limited resources in the pre-warning system, the pre-warning efficiency is improved. And emphasized diagnosis and early warning are carried out on the computer with relatively high fault diagnosis danger.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and specifically to a computer intelligent fault diagnosis and early warning system based on artificial intelligence. Background Art

[0002] A computer is commonly known as a "computer" and is a modern electronic computing device used for high-speed calculations. It can perform numerical calculations, logical calculations, and also has a storage and memory function. It is a modern intelligent electronic device that can run according to a program and automatically and quickly process a large amount of data.

[0003] Chinese invention CN202410943851.3 discloses a computer intelligent fault diagnosis and early warning system based on artificial intelligence, including an acquisition module for acquiring computer hardware operation state parameters and software operation process parameters; a configuration module for configuring the operation cycle of the acquisition module; this invention comprehensively monitors the safety of the computer operation state by sensing devices to monitor computer hardware operation state parameters and software operation process parameters under the computer operation state, and further, based on the monitored computer parameters, comprehensively analyzes to obtain the computer safety score. Finally, according to the dynamic change of the computer safety score, a fault warning and fault diagnosis are issued to the computer user to ensure that the potential safety hazards during the computer operation process can be timely known to the computer user, and the computer is predictably maintained and serviced.

[0004] However, it is difficult for the above-cited document to consider both the hardware operation health status of the computer itself and the usage of network bandwidth during computer intelligent fault diagnosis and early warning when the computer is running, and when the early warning system monitors and warns multiple computers simultaneously, it is difficult to focus on monitoring and warning the computers with higher fault diagnosis risks under the limited resources in the early warning system.

[0005] Therefore, there is an urgent need for a computer intelligent fault diagnosis and early warning system based on artificial intelligence to solve the above problems. Summary of the Invention

[0006] The purpose of the present invention is to provide a computer intelligent fault diagnosis and early warning system based on artificial intelligence to solve the problems raised in the above background art.

[0007] To achieve the above purpose, the present invention provides the following technical solution: A computer intelligent fault diagnosis and early warning system based on artificial intelligence, including: A data collection module for real-time monitoring and obtaining the current CPU temperature Tcpu, operating memory usage rate Mu, and hard disk memory usage rate Hs of the computer through hardware monitoring software, and for real-time monitoring and obtaining the network bandwidth usage rate through a network monitoring tool, and uploading them to the database together; A data preprocessing module, which is used to decode and preprocess the data information in the database to obtain the parameters participating in the calculation in the calculation processing module; A calculation processing module, which is used to substitute the parameter values obtained after decoding and preprocessing into the fault risk base value algorithm unit to calculate the fault risk base value Fri and upload it to the database; which is used to input the calculated fault risk base value Fri as an input parameter into the adjusted fault risk value algorithm unit to calculate the adjusted fault risk prediction value Al; It is also used to input the calculated fault risk base value Fri and the adjusted fault risk prediction value Al as input parameters into the fault risk final value algorithm unit to calculate the fault risk final value Ffv and upload it to the database; A fault diagnosis and warning module, which is used to perform fault diagnosis and warning according to the fault risk final value Ffv recorded in the database.

[0008] Optionally, the fault diagnosis and warning specifically includes: Setting 1.5 times the fault risk base value Fri as the warning threshold and 1.8 times the fault risk base value Fri as the danger threshold in the database; Comparing the fault risk final value Ffv obtained in one operation of the three algorithm units with the danger threshold and the warning threshold; When the fault risk final value Ffv is lower than the warning threshold, it is determined that the computer is in a good operating state, and the monitoring frequency is maintained at once every ten minutes; When the fault risk final value Ffv is higher than the warning threshold and lower than the danger threshold, the monitoring frequency of the computer is adjusted to once a minute; For a computer with a fault risk final value Ffv higher than the danger threshold, remind the user that the computer will shut down within 5 minutes to protect the hardware in the computer.

[0009] Optionally, the calculation processing module includes a fault risk base value algorithm unit, an adjusted fault risk value algorithm unit, and a fault risk final value algorithm unit.

[0010] Optionally, the fault risk base value algorithm unit is as follows: ;

[0011] Where: Fri represents the fault risk base value; Tcpu represents the current CPU temperature, which is the temperature of the computer's central processing unit CPU; Tmax represents the maximum safe temperature of the CPU, which is the highest temperature that the computer's central processing unit CPU can withstand; Mu represents the running memory utilization rate; Hs represents the hard disk memory utilization rate; α represents the weight coefficient, with a value range from 0.1 to 3, which can be self-adjusted within the basic fault risk value algorithm unit; In the formula calculation: This part reflects the influence degree of the current CPU temperature Tcpu on the calculation of the basic fault risk value Fri through the logarithmic function. As the current CPU temperature Tcpu increases, the calculated value of this part will increase, thus increasing the calculated basic fault risk value Fri. Through the manifestation of the logarithmic function, it represents that as the current CPU temperature Tcpu increases, the increasing amplitude of the basic fault risk value Fri will gradually become smaller. When the temperature is relatively low in the initial stage, the temperature change has a greater impact on the basic fault risk value Fri. As the current CPU temperature Tcpu approaches the maximum safe temperature Tmax of the CPU, the growth rate of the logarithmic function slows down; This part affects the calculation of the basic fault risk value Fri by taking the reciprocal of the negative exponential power of the memory utilization rate Mu plus 1 with the constant e. As the running memory utilization rate Mu increases, the calculated value of this part gets closer to 1, representing that as the running memory utilization rate Mu increases, the system load of the computer is heavier, thus increasing the calculated basic fault risk value Fri; This part reflects the influence of the hard disk memory utilization rate Hs on the basic fault risk value Fri. As the hard disk memory utilization rate Hs increases, the value of this part approaches 1, thus increasing the calculated basic fault risk value Fri, representing that as the hard disk memory utilization rate Hs increases, the system load of the computer increases, thus increasing the calculated basic fault risk value Fri.

[0012] Optionally, the adjusted fault risk value algorithm unit is as follows: ;

[0013] Wherein: Al represents the predicted value of the adjusted fault risk; Fri represents the basic fault risk value; Bm represents the network bandwidth usage value, which is obtained by multiplying the network bandwidth utilization rate monitored in real time by the built-in network monitoring tool of the computer by 100; θ represents the weight coefficient, with a value range from 0 to 1, which can be self-adjusted in the calculation of the adjusted fault risk value algorithm unit; This part reflects the influence degree of the basic fault risk value Fri of temperature Tcpu on the calculation of the adjusted predicted fault risk value Al through a logarithmic function, in order to reflect the non-linear relationship between the basic fault risk value Fri and the adjusted predicted fault risk value Al. The logarithmic transformation can ensure that the adjusted predicted fault risk value Al increases with the increase of the basic fault risk value Fri, but the increasing speed will gradually slow down, avoiding the excessive influence on the adjusted predicted fault risk value Al when the basic fault risk value Fri is too high; This part means that by dividing the network broadband usage value Bm by 50 and taking it as the negative exponential power of the constant e, it affects the calculation of the adjusted predicted fault risk value Al. When the network broadband usage value Bm is lower than 50, with the increase of the network broadband usage value Bm, the increase amplitude of this part is relatively low. When the network broadband usage value Bm is greater than 50, with the increase of the network broadband usage value Bm, the increase amplitude of this part is relatively high, reflecting that the network usage state in the insufficient state has a greater impact on the adjusted predicted fault risk value Al of the calculation result.

[0014] Optionally, the algorithm unit of the final fault risk value is as follows: ;

[0015] Where: Ffv represents the final fault risk value; Al represents the adjusted predicted fault risk value; Sl is the system load value, indicating the current load level of the computer, with a value range from 0 to 100. When taking 0, the load is the smallest, and when taking 100, the load is the largest, which is obtained by real-time monitoring of the system monitoring tool Task Manager; In the formula calculation: This part uses a logarithmic function to smoothly adjust the influence of the adjusted predicted fault risk value Al on the final fault risk value Ffv. With the increase of the adjusted predicted fault risk value Al, the increasing amplitude of the final fault risk value Ffv gradually slows down. Through the logarithmic function, it can avoid the excessive influence on the calculation of the final fault risk value Ffv when the adjusted predicted fault risk value Al is too large; This part reflects the influence of the system load value Sl on the final fault risk value Ffv by dividing the system load value Sl by the maximum load 100 and then adding the constant 1. With the increase of the system load value Sl, the calculated final fault risk value Ffv increases.

[0016] Optionally, the hardware monitoring software includes HWMonitor software and Argus Monitor software.

[0017] Optionally, the decoding preprocessing includes data cleaning and data standardization.

[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: Through the mutual cooperation of three algorithm units, the present invention constitutes the core architecture of a computer intelligent fault diagnosis and early warning system based on artificial intelligence. It can, while the computer is running, take into account both the hardware operating conditions such as CPU temperature, operating memory usage rate, and hard disk memory usage rate, as well as the usage of network bandwidth, when performing intelligent fault diagnosis and early warning on the computer, so as to more comprehensively reflect the load and health status of the computer system during operation, identify the faulty hardware components during computer operation and give an alarm, thereby helping the operator to more efficiently locate faults and repair the computer.

[0019] Through the mutual cooperation of three algorithm units, the present invention calculates the final fault risk value Ffv and compares it with the warning threshold and danger threshold set in the database. When the final fault risk value Ffv is lower than the warning threshold, it is determined that the computer is in a good operating state, and the monitoring frequency of once every ten minutes will be maintained. When the final fault risk value Ffv is higher than the warning threshold and lower than the danger threshold, the monitoring frequency of the computer will be increased. For a computer with a final fault risk value Ffv higher than the danger threshold, the user will be reminded that the computer will shut down within 5 minutes to protect the hardware in the computer, so as to be able to focus on monitoring and early warning of computers with a higher risk of fault diagnosis under the limited resources in the early warning system. Description of the Drawings

[0020] Figure 1 is a flowchart of a computer intelligent fault diagnosis and early warning system based on artificial intelligence; Figure 2 is a schematic diagram of the overall structure of a computer intelligent fault diagnosis and early warning system based on artificial intelligence. Detailed Embodiments

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0022] Embodiment 1, please refer to Figures 1 to 2 , the present invention provides a computer intelligent fault diagnosis and early warning system based on artificial intelligence, including: A data collection module, which is used to monitor and obtain the current CPU temperature Tcpu, the running memory usage rate Mu, and the hard disk memory usage rate Hs of a computer in real time through hardware monitoring software, and is used to monitor and obtain the network broadband usage rate in real time through a network monitoring tool, and upload them to the database together; A data preprocessing module, which is used to decode and preprocess the data information in the database to obtain the parameters participating in the calculation in the calculation processing module; A calculation processing module, which is used to substitute the parameter values obtained after decoding and preprocessing into the fault risk base value algorithm unit to calculate the fault risk base value Fri and upload it to the database; which is used to input the calculated fault risk base value Fri as an input parameter into the adjusted fault risk value algorithm unit to calculate the adjusted fault risk prediction value Al; It is also used to input the calculated fault risk base value Fri and the adjusted fault risk prediction value Al as input parameters into the fault risk final value algorithm unit to calculate the fault risk final value Ffv and upload it to the database; A fault diagnosis and early warning module, which is used to perform fault diagnosis and early warning according to the fault risk final value Ffv recorded in the database, specifically including: Setting 1.5 times the fault risk base value Fri as the warning threshold and 1.8 times the fault risk base value Fri as the danger threshold in the database; Comparing the fault risk final value Ffv obtained in one complete operation of the three algorithm units with the danger threshold and the warning threshold; When the fault risk final value Ffv is lower than the warning threshold, it is determined that the computer is in a good running state, and the monitoring frequency of once every ten minutes is maintained; When the fault risk final value Ffv is higher than the warning threshold and lower than the danger threshold, the monitoring frequency of the computer is adjusted to once a minute; For a computer with a fault risk final value Ffv higher than the danger threshold, the user is reminded that the computer will shut down within 5 minutes to protect the hardware in the computer.

[0023] In this embodiment: Through the mutual cooperation of three algorithm units, the present invention constitutes the core architecture of a computer intelligent fault diagnosis and early warning system based on artificial intelligence. It can take into account both the hardware operation health status of the computer itself and the usage of network bandwidth during computer operation, comprehensively reflecting the load and health status of the computer system, providing a comprehensive view during fault early warning, improving the accuracy and timeliness of early warning, thus detecting potential problems earlier, identifying the hardware components that malfunction during computer operation and giving an alarm, and helping operators or automated systems to more efficiently locate faults and repair the computer when a fault occurs in the computer.

[0024] Moreover, through the mutual cooperation of three algorithm units, the final fault risk value Ffv is calculated and compared with the set warning threshold and danger threshold in the database. When the final fault risk value Ffv is lower than the warning threshold, it is determined that the computer is in a good operating state, and the monitoring frequency is maintained at once every ten minutes. When the final fault risk value Ffv is higher than the warning threshold and lower than the danger threshold, the monitoring frequency of the computer is adjusted to once a minute. For a computer with a final fault risk value Ffv higher than the danger threshold, the user is reminded that the computer will shut down within 5 minutes to protect the hardware in the computer. Therefore, under the limited resources in the early warning system, computers with a higher risk of fault diagnosis can be focused on for monitoring and early warning, which is worthy of popularization and use.

[0025] Please refer to Figures 1 to 2 , the basic fault risk value algorithm unit is as follows: ;

[0026] Where: Fri represents the basic fault risk value; Tcpu represents the current CPU temperature, which is the temperature of the computer's central processing unit CPU. Excessive temperature will cause the performance of the computer to decline; Tmax represents the maximum safe temperature of the CPU, which is the highest temperature that the computer's central processing unit CPU can withstand. Above this temperature, the computer will automatically shut down to protect the hardware in the computer; Mu represents the usage rate of the running memory. As the usage rate of the running memory increases, the computer will run slower and experience lags; Hs represents the usage rate of the hard disk memory. As the usage rate of the hard disk memory increases, the computer will run slower and experience lags; The current CPU temperature Tcpu, the usage rate of the running memory Mu, and the usage rate of the hard disk memory Hs are all detected and obtained in real time through the hardware monitoring software built into the computer. α represents the weight coefficient, with a value range from 0.1 to 3, and it can be self-adjusted within the basic fault risk value algorithm unit; Specifically: The maximum safe temperature Tmax of CPUs with different hardware and system architectures of different computers varies greatly. In high-performance computers or data centers, the maximum safe temperature Tmax of CPUs is higher and high-performance processors are more sensitive to CPU temperature. In the calculation of the basic fault risk value algorithm unit, a higher α value will be selected. For ordinary desktop computers or laptops, the α value will be reduced because the change in their CPU temperature will not have too much impact on system stability.

[0027] In the formula calculation: This part reflects the influence degree of the current CPU temperature Tcpu on the calculation of the basic fault risk value Fri through a logarithmic function. As the current CPU temperature Tcpu increases, the calculated value of this part will increase, thereby increasing the calculated basic fault risk value Fri. Through the manifestation of the logarithmic function, it represents that as the current CPU temperature Tcpu increases, the increase amplitude of the basic fault risk value Fri will gradually become smaller. When the temperature is relatively low at the initial stage, the temperature change has a greater impact on the basic fault risk value Fri. As the current CPU temperature Tcpu approaches the maximum safe temperature Tmax of the CPU, the growth rate of the logarithmic function slows down, making the impact of further temperature increase on the basic fault risk value Fri no longer show a linear increase but gradually tend to be flat, so as to more realistically simulate the relationship between the current CPU temperature Tcpu and the current CPU temperature Tcpu. That is, when the temperature is low, the risk growth is relatively significant, but in the high-temperature range, the system has a certain tolerance, so the growth of the fault risk tends to be flat, avoiding the excessive impact on the calculated basic fault risk value Fri when the current CPU temperature Tcpu is too high; This part affects the calculation of the basic fault risk value Fri by taking the reciprocal of the negative exponential power of the memory usage rate Mu plus 1 with the constant e. As the running memory usage rate Mu increases, the calculated value of this part gets closer to 1, representing that as the running memory usage rate Mu increases, the system load of the computer is heavier, thereby increasing the calculated basic fault risk value Fri; This part reflects the influence of the hard disk memory usage rate Hs on the basic fault risk value Fri. As the hard disk memory usage rate Hs increases, this The value of the part is close to 1, thereby increasing the calculated basic value of the failure risk Fri, which represents that as the hard disk memory utilization rate Hs increases, the system load of the computer increases, thereby increasing the calculated basic value of the failure risk Fri.

[0028] In this embodiment: By comprehensively considering multiple influencing factors such as the current CPU temperature Tcpu, the running memory utilization rate Mu, and the hard disk memory utilization rate Hs, the calculated basic value of the failure risk Fri can more comprehensively reflect the system load and health status. The parameters of the CPU temperature Tcpu, the running memory utilization rate Mu, and the hard disk memory utilization rate Hs can provide a comprehensive view during failure warning, improving the accuracy and timeliness of the warning, so as to detect potential problems earlier. The computer intelligent fault diagnosis and warning system can identify the hardware components most likely to fail through real-time calculation and evaluation of the basic value of the failure risk Fri. For example, if the current CPU temperature Tcpu is too high in the calculation of the basic value of the failure risk Fri, attention needs to be paid to the health status of the cooling system and the CPU itself. If the hard disk memory utilization rate Hs is too high, the read / write situation and storage pressure of the hard disk need to be checked, which can help engineers or automated systems perform fault location and repair more efficiently.

[0029] In practical applications, since the current CPU temperature Tcpu, the running memory utilization rate Mu, and the hard disk memory utilization rate Hs are all real-time monitorable indicators, in the computer intelligent fault diagnosis and warning system based on artificial intelligence, these data can be updated in real time and input into the calculation of the basic value of the failure risk algorithm unit to participate in the calculation of the basic value of the failure risk Fri to continuously track the system health status. When these parameters change, the system will automatically update the basic value of the failure risk Fri to provide the latest warning information. This dynamic monitoring method can make the computer intelligent fault diagnosis and warning system more flexible and help to respond in a timely manner.

[0030] Please refer to Figures 1 to 2 , the adjusted failure risk value algorithm unit is as follows: ;

[0031] Wherein: Al represents the adjusted estimated value of the failure risk; Fri represents the basic value of the failure risk; Bm represents the network broadband usage value, which is obtained by multiplying the network broadband utilization rate obtained by real-time monitoring through the built-in network monitoring tool of the computer by 100; θ represents the weight coefficient, and its value range is from 0 to 1, which can be self-adjusted in the calculation of the adjusted failure risk value algorithm unit; Specifically: In usage scenarios where the computer system is highly sensitive to network status, such as certain real-time data transmission or low-latency applications like online games and financial transactions, in the adjusted failure risk value algorithm unit, θ will be set to a relatively large value (above 0.5) to emphasize the role of the network in failure warning; In formula calculation: This part reflects the influence degree of the failure risk base value Fri temperature Tcpu on the calculation of the adjusted failure risk prediction value Al through the logarithmic function, to reflect the non-linear relationship between the failure risk base value Fri and the adjusted failure risk prediction value Al. The logarithmic transformation can ensure that the adjusted failure risk prediction value Al increases as the failure risk base value Fri increases, but the increasing speed will gradually slow down, avoiding an excessive impact on the adjusted failure risk prediction value Al when the failure risk base value Fri is too high; This part means that by dividing the network broadband usage value Bm by 50 and taking it as the negative exponential power of the constant e, it affects the calculation of the adjusted failure risk prediction value Al. When the network broadband usage value Bm is lower than 50, as the network broadband usage value Bm increases, The increase rate of this part is relatively low, indicating that the network usage status in the abundant state has a relatively small impact on the calculated result of the adjusted failure risk prediction value Al. While when the network broadband usage value Bm is greater than 50, as the network broadband usage value Bm increases, The increase rate of this part is relatively high, representing that when the utilization rate of the network broadband is relatively high, the network usage status in the non-abundant state has a relatively large impact on the calculated result of the adjusted failure risk prediction value Al.

[0032] In this embodiment: By comprehensively considering the network broadband usage value Bm, the usage scenario of the computer, and the basic fault risk value Fri calculated by the basic fault risk value algorithm unit, the adjusted predicted fault risk value Al after computer network status adjustment is calculated. The basic fault risk value Fri provides a basic assessment of the computer fault risk, which combines the usage status of the hardware (such as CPU temperature, memory usage rate, hard disk usage rate, etc.). The value of the basic fault risk value Fri can measure the risk status of the computer system itself and determine the basic fault risk level of the system. Considering the network bandwidth usage value Bw into the adjusted fault risk value algorithm unit represents the communication status between the computer and other devices or systems when the computer is performing tasks. The usage of network bandwidth has a significant impact on the overall performance and reliability of the computer. For example, when the network bandwidth is overloaded, problems such as network congestion, data packet loss, or latency may lead to a decline in system performance, thereby increasing the fault risk. Especially in modern workloads such as remote operation, data transmission, and cloud computing, network fluctuations will have an important impact on the fault risk of the computer.

[0033] The adjusted predicted fault risk value Al calculated by the adjusted fault risk value algorithm unit combines the comprehensive assessment of the hardware health status and network bandwidth, providing a more comprehensive and accurate early warning mechanism. By continuously calculating and analyzing the adjusted predicted fault risk value A, the computer intelligent fault diagnosis and early warning system can identify in advance the fault risks affected by network load. For example, in the case of high hardware load inside the computer, that is, when the value of the basic fault risk value Fri is relatively high, as the network bandwidth usage value Bw increases, the network bandwidth of the computer system is about to approach saturation, and the calculated adjusted predicted fault risk value Al will also increase, indicating that there may be problems such as a decline in system performance or a fault in the computer system. This provides scientific and reliable data information for fault early warning and maintenance personnel, helping maintenance personnel take timely measures to avoid system crashes or serious faults.

[0034] By considering both network bandwidth and hardware health status, the computer intelligent fault diagnosis and early warning system based on artificial intelligence can better simulate the system performance of the computer in a real environment, thereby improving the accuracy of intelligent early warning. For example, in a remote work or cloud computing environment, the network quality and bandwidth usage will fluctuate greatly. At this time, the fault risk of the computer should not only depend on the status of the hardware but also need to comprehensively consider network factors. This can enable the computer intelligent fault diagnosis and early warning system to adapt to different environments and maintain high stability and accuracy under changing network conditions.

[0035] Please refer to Figures 1 to 2 , the final fault risk value algorithm unit is as follows: ;

[0036] Wherein: Ffv represents the final value of the fault risk; Al represents the estimated value of the adjusted fault risk; Sl is the system load value, indicating the current load level of the computer. The value ranges from 0 to 100, with 0 representing the minimum load and 100 representing the maximum load, and it is obtained by real-time monitoring with the system monitoring tool Task Manager; In the formula calculation: This part uses a logarithmic function to smoothly adjust the influence of the estimated value of the adjusted fault risk Al on the final value of the fault risk Ffv. As the estimated value of the adjusted fault risk Al increases, the increase rate of the final value of the fault risk Ffv gradually slows down. The logarithmic function can avoid the excessive influence on the calculation of the final value of the fault risk Ffv when the estimated value of the adjusted fault risk Al is too large; This part reflects the influence of the system load value Sl on the final value of the fault risk Ffv by dividing the system load value Sl by the maximum load 100 and then adding a constant 1. As the system load value Sl increases, the calculated final value of the fault risk Ffv also increases.

[0037] In this embodiment: The final value algorithm unit of the fault risk comprehensively considers the estimated value of the adjusted fault risk Al and the system load Sl after the computer network state is adjusted, and calculates the final value of the fault risk Ffv, which can provide a more accurate and comprehensive fault risk assessment for the computer, avoiding the excessive influence of a single factor. For example, when the computer has multiple situations such as high temperature, high network bandwidth usage, and heavy system load, the system can timely identify potential faults and issue alarms. By continuously monitoring and calculating the final value of the fault risk Ffv of the computer, it can reflect the overall fault risk level of the current computer system in real time. The administrator can effectively monitor the health status of the computer by analyzing the final value of the fault risk Ffv at different time periods of the computer, and perform preventive maintenance according to the warning information. When the calculated final value of the fault risk Ffv is too high, it indicates that the system is in a high-risk state and corresponding measures need to be taken, such as reducing the system load, increasing the network bandwidth, and cleaning the hardware.

[0038] Since the system load value SL is obtained in real time from the task manager, and the network bandwidth usage value Bm in the adjusted estimated failure risk value Al is obtained through real-time network monitoring, the calculated final failure risk value Ffv can provide a dynamic failure risk assessment, which means that the computer intelligent fault diagnosis and early warning system can adjust its risk assessment results in real time to cope with potential problems that may occur under different workloads and network states. For example, when the network bandwidth is about to reach the limit, the final failure risk value algorithm unit will automatically adjust the calculated final failure risk value Ffv without manual intervention. And because the formula includes hardware health, network load, and system load, the administrator can obtain a more specific risk analysis report, so as to be able to make more targeted adjustments and repairs.

[0039] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. Computer intelligent fault diagnosis and early warning system based on artificial intelligence, characterized by: include: The data collection module is used to monitor and obtain the current CPU temperature Tcpu, running memory usage Mu and hard disk memory usage Hs of the computer in real time through the hardware monitoring software, and to monitor and obtain the network bandwidth usage rate in real time through the network monitoring tool, and upload them to the database together; The data preprocessing module is used to decode and preprocess the data information in the database to obtain the parameters involved in the calculation in the calculation processing module; Computational processing module, Used to substitute the parameter value obtained after decoding preprocessing into the fault risk basic value algorithm unit to calculate the fault risk basic value Fri, then input the fault risk basic value Fri as an input parameter into the adjusted fault risk value algorithm unit to calculate the adjusted fault risk estimated value Al, and input the fault risk basic value Fri and the adjusted fault risk estimated value Al as input parameters into the fault risk final value algorithm unit to calculate the fault risk final value Ffv, and upload it to the database; The fault diagnosis and early warning module is used to perform fault diagnosis and early warning according to the final value of the fault risk Ffv recorded in the database.

2. The computer intelligent fault diagnosis and early warning system based on artificial intelligence according to claim 1 is characterized in that: The fault diagnosis and early warning specifically include: In the database, 1.5 times the basic value of the fault risk Fri is set as the warning threshold, and 1.8 times the basic value of the fault risk Fri is set as the danger threshold; Compare the final value Ffv of the fault risk obtained in one operation of the three algorithm units with the danger threshold and the warning threshold; When the final value of the failure risk Ffv is lower than the warning threshold, the computer is deemed to be in a good operating state and the monitoring frequency is maintained every ten minutes; When the final value of the fault risk Ffv is higher than the warning threshold and lower than the danger threshold, the monitoring frequency of the computer is adjusted to once a minute; For computers whose final failure risk value Ffv is higher than the danger threshold, the user is reminded that the computer will be shut down within 5 minutes to protect the hardware inside the computer.

3. The computer intelligent fault diagnosis and early warning system based on artificial intelligence according to claim 2 is characterized in that: The calculation processing module includes a fault risk basic value algorithm unit, an adjusted fault risk value algorithm unit and a fault risk final value algorithm unit.

4. The computer intelligent fault diagnosis and early warning system based on artificial intelligence according to claim 3 is characterized by: The fault risk basic value algorithm unit is as follows: ; in: Fri represents the basic value of failure risk; Tcpu represents the current CPU temperature; Tmax represents the maximum safe temperature of the CPU; Mu represents the running memory usage; Hs represents the hard disk memory usage; α represents the weight coefficient, ranging from 0.1 to 3; In the calculation formula: This part reflects the influence of the current CPU temperature Tcpu on the calculation of the fault risk base value Fri through a logarithmic function. As the current CPU temperature Tcpu increases, The calculated value of this part will increase, thereby increasing the calculated fault risk basic value Fri. The logarithmic function shows that as the current CPU temperature Tcpu increases, the increase in the fault risk basic value Fri will gradually decrease. When the temperature is low in the initial stage, the temperature change has a greater impact on the fault risk basic value Fri. As the current CPU temperature Tcpu approaches the maximum safe temperature Tmax of the CPU, the growth rate of the logarithmic function slows down. This part affects the calculation of the fault risk base value Fri by taking the inverse of the negative exponential power of the constant e plus 1 using the memory usage rate Mu. As the running memory usage rate Mu increases, The closer the calculated value of this part is to 1, the heavier the system load of the computer will be as the running memory usage Mu increases, thus increasing the calculated basic value of the failure risk Fri; This part reflects the impact of the hard disk memory usage rate Hs on the basic value of failure risk Fri. As the hard disk memory usage rate Hs increases, The value of some parts is close to 1, thereby increasing the calculated basic value of failure risk Fri, which means that as the hard disk memory usage rate Hs increases, the system load of the computer increases, thereby increasing the calculated basic value of failure risk Fri.

5. The computer intelligent fault diagnosis and early warning system based on artificial intelligence according to claim 4 is characterized in that: The adjusted fault risk value algorithm unit is as follows: ; in: Al represents the adjusted failure risk estimate; Fri represents the basic value of failure risk; Bm represents the network bandwidth usage value, which is obtained by real-time monitoring of the network bandwidth usage rate obtained through the computer's built-in network monitoring tool and multiplied by 100; θ represents the weight coefficient, ranging from 0 to 1; This part reflects the influence of the fault risk basic value Fri temperature Tcpu on the calculation of the adjusted fault risk estimate value Al through a logarithmic function, so as to reflect the nonlinear relationship of the fault risk basic value Fri on the adjusted fault risk estimate value Al. The logarithmic transformation can ensure that the adjusted fault risk estimate value Al increases with the increase of the fault risk basic value Fri, but the increase rate will gradually slow down, so as to avoid the excessive influence on the adjusted fault risk estimate value Al when the fault risk basic value Fri is too high. This part indicates that the negative exponential power of the constant e after dividing the network bandwidth usage value Bm by 50 affects the calculation of the adjusted fault risk estimate Al. When the network bandwidth usage value Bm is lower than 50, as the network bandwidth usage value Bm increases, The increase in this part is relatively low. When the network bandwidth usage value Bm is greater than 50, as the network bandwidth usage value Bm increases, The increase in this part is relatively high, which is reflected in the fact that the network usage status under insufficient conditions has a greater impact on the adjusted fault risk estimate Al of the calculation results.

6. The computer intelligent fault diagnosis and early warning system based on artificial intelligence according to claim 5 is characterized in that: The fault risk final value algorithm unit is as follows: ; in: Ffv represents the final value of the failure risk; Al represents the adjusted failure risk estimate; Sl system load value, indicating the current load level of the computer, ranging from 0 to 100, with 0 being the smallest load and 100 being the largest load, and is obtained through real-time monitoring by the system monitoring tool Task Manager; In the calculation formula: This part uses a logarithmic function to smoothly adjust the impact of the fault risk estimate Al on the final fault risk value Ffv. As the adjusted fault risk estimate Al increases, the increase in the final fault risk value Ffv gradually slows down. The logarithmic function can avoid excessive impact on the calculation of the final fault risk value Ffv when the adjusted fault risk estimate Al is too large. This part reflects the influence of the system load value S1 on the final value of the fault risk Ffv by dividing the system load value S1 by the maximum load 100 and then adding a constant 1. As the system load value S1 increases, the calculated final value of the fault risk Ffv increases.

7. The computer intelligent fault diagnosis and early warning system based on artificial intelligence according to claim 1 is characterized in that: The hardware monitoring software includes HWMonitor software and Argus Monitor software.

8. The computer intelligent fault diagnosis and early warning system based on artificial intelligence according to claim 1 is characterized in that: The decoding preprocessing includes data cleaning and data standardization.

Citation Information

Patent Citations

  • Computer intelligent fault diagnosis and early warning system based on artificial intelligence

    CN118939458A

  • Distributed system fault judgment and recovery method, cloud operating system applying method and computing platform

    CN118331779A

  • Fault diagnosis prediction method and system based on operation and maintenance scene

    CN119167132A

  • Circuit board fault diagnosis method and system based on artificial intelligence

    CN119203919A

  • Smart APT communication control and status monitoring system

    KR102779346B1