Machine room automation equipment abnormal state monitoring method and device

By collecting and analyzing the sensor and performance data of the computer room automation equipment, calculating the temperature gradient and performance values, the problem of difficulty in comprehensive and accurate monitoring in the existing technology is solved, real-time identification and cause analysis of equipment abnormalities is realized, and the stability of equipment operation is improved.

CN120378329APending Publication Date: 2025-07-25JUXIAN POWER SUPPLY CO STATE GRID SHANDONG ELECTRIC POWER CO

Patent Information

Application Number
CN202510721129.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

It is difficult for the prior art to conduct comprehensive, accurate and real-time monitoring and analysis of the complex operation of computer room automation equipment, especially the slight abnormalities of sensor data and performance data are difficult to detect in a timely manner, which may lead to equipment failures and business interruptions.

Method used

By collecting sensor data and performance data, calculating the average temperature gradient value and comparing it with the gradient threshold, analyzing the causes of abnormalities based on historical data and change characteristics, calculating network performance values and system performance values, and generating specific reason information.

Benefits of technology

It realizes in-depth insight into equipment performance, timely discovers temperature changes abnormalities, comprehensively evaluates the impact of network and system performance abnormalities on equipment, and improves the accuracy and real-time monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120378329A_ABST
    Figure CN120378329A_ABST
Patent Text Reader

Abstract

The invention discloses a method and a device for monitoring an abnormal state of machine room automation equipment, relates to the technical field of equipment monitoring and analysis, and solves the technical problem that complex conditions are difficult to comprehensively and accurately monitor and analyze in real time. Temperature change abnormity can be found in time, slow change, sudden change and local change are further distinguished according to a temperature change rate judgment method, abnormity reasons are matched in combination with historical data and change characteristics, and key indexes such as network performance data, a broadband occupancy value, a packet loss rate and network delay are obtained. The method comprehensively evaluates the influence degree and change trend of network performance abnormity on equipment, considers multiple factors such as CPU, memory and disk I / O for system performance data, accurately judges the influence of the system performance data on the equipment, and realizes deep insight of the equipment performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment monitoring and analysis, and particularly to a method and device for monitoring abnormal states of automated equipment in a computer room. Background Art

[0002] In today's digital age, computer room facilities such as data centers play a crucial role in the normal operation of enterprises and organizations. A large number of automated devices, such as servers and network devices, are deployed in the computer room, and the stable operation of these devices is directly related to business continuity and data security.

[0003] The patent with the publication number CN117992301A discloses a method, device, and system for monitoring equipment anomalies. The method includes: obtaining log data generated during the operation of the equipment; normalizing and storing the log data in an equipment library or an anomaly library, where the equipment library and the anomaly library are pre-established, the equipment library is used to store normal log data, and the anomaly library is used to store abnormal log data; when it is necessary to find target abnormal data, perform a matching search in the anomaly library.

[0004] However, during the operation of the equipment, its sensor data and performance data will continuously change, and the minor anomalies in these data may indicate potential equipment failure risks. For example, an abnormal increase in equipment temperature may be caused by a failure of the cooling system, hardware overload, or other reasons. If not discovered and processed in a timely manner, it may lead to equipment damage or even downtime, and further cause serious consequences such as business interruption and data loss. At the same time, fluctuations in network performance data and changes in system performance data will also have a great impact on the normal operation of the equipment. Traditional monitoring methods are difficult to comprehensively, accurately, and real-timely monitor and analyze these complex situations. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a method and device for monitoring abnormal states of automated equipment in a computer room, which solves the problem of being difficult to comprehensively, accurately, and real-timely monitor and analyze these complex situations.

[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for monitoring abnormal states of automated equipment in a computer room, the method specifically includes the following steps: Step 1: Collect sensor data and performance data of cluster equipment; Step 2: Compare the collected sensor data and performance data with normal values to judge the overall state of the cluster equipment, and at the same time generate equipment normal results or equipment abnormal results; Step 3: Analyze the abnormal situation of sensor data in the device abnormal result. By specifically analyzing the temperature and determining the abnormal situation based on the temperature change, generate specific cause information; Step 4: Analyze the abnormal situation of performance data in the device abnormal result. By separately analyzing the network performance data and system performance data, and judging by calculating the network performance value and system performance value according to the corresponding data information, and at the same time determine the specific cause to generate specific cause information.

[0007] As a further solution of the present invention: The specific method for the step 2 to judge the overall state of the cluster device is: Compare the obtained sensor data and performance data with the corresponding normal values respectively. If both the sensor data and performance data are the same as the normal values, it means that the overall state of the cluster device is normal, and at the same time generate a device normal result. If the sensor data or performance data is different from the normal value, it means that the overall state of the cluster device is abnormal, and at the same time generate a device abnormal result.

[0008] As a further solution of the present invention: The specific way for the step 3 to analyze the abnormal situation of sensor data in the device abnormal result is: Obtain the temperature data corresponding to the sensor data, and at the same time obtain the change situation of the temperature data within the time period t, calculate the temperature gradient corresponding to the temperature data, calculate the gradient mean value corresponding to the temperature gradient, and compare the calculated gradient mean value with the gradient threshold. If the gradient mean value is greater than the gradient threshold, it means that the temperature data changes abnormally, and generate a temperature change abnormal signal. On the contrary, if the temperature gradient mean value is less than the gradient threshold, it means that the temperature data changes normally, and generate a temperature change normal signal, and at the same time analyze the temperature change abnormal signal.

[0009] As a further solution of the present invention: The specific way for the step 3 to analyze the temperature change abnormal signal is: Establish a coordinate system of temperature gradient and time, and at the same time judge the change situation of the temperature gradient, and classify it into slow change, sudden change and local change. Then analyze the abnormal causes corresponding to the slow change, sudden change and local change respectively; Analyze the slow change, obtain the historical data, obtain the similar historical data at the same time, and obtain the abnormal causes corresponding to the similar historical data. Then label the obtained abnormal causes as i, and i = 1, 2,..., j, where j represents the number of abnormal causes. At the same time, obtain the change characteristics corresponding to the slow change, match the change characteristics with the abnormal causes, and screen the abnormal causes in combination with the change mean value corresponding to the slow change to obtain the specific cause, and generate specific cause information; The analysis of sudden changes and local changes is the same as that of slow-changing processes, and the corresponding change characteristics and change means are combined to determine the specific reasons and generate specific reason information.

[0010] As a further solution of the present invention: The specific manner of analyzing the abnormal performance data in the abnormal results of the device in step four is as follows: Analyze the network performance data, obtain the current broadband occupancy value, packet loss rate, and network latency of the cluster device, and compare them with the corresponding normal working values respectively, and at the same time screen out the abnormal network performance data; Obtain the abnormal network performance data, and obtain the corresponding peak value and the mean value of the performance data gradient within time t1, and at the same time calculate the change frequency corresponding to the abnormal network performance data, and substitute the obtained parameters into the formula Calculate the network performance value NPV, where Pmax and Pmin respectively represent the maximum peak value and the minimum peak value corresponding to the abnormal network performance data, F represents the change frequency, G is the mean value of the performance data gradient, and at the same time compare the obtained network performance value NPV with the first preset value, and the specific value of the first preset value is set by the operator.

[0011] As a further solution of the present invention: The specific manner of comparing the obtained network performance value with the first preset value in step four is as follows: If the network performance value NPV is greater than the first preset value, it means that the current abnormal network performance data has an impact on the device and generates an impact information. If the network performance value NPV is less than the first preset value, further analysis and judgment are required. Taking time T as a cycle, analyze the change situation of the network performance value corresponding to the time period T, and at the same time generate the corresponding change information.

[0012] As a further solution of the present invention: The specific manner of analyzing the system performance data in step four is as follows: Obtain the current abnormal CPU usage rate, abnormal memory usage rate, and abnormal disk I / O usage rate of the cluster device, and then substitute the obtained parameters into the formula Calculate the system performance anomaly value SPE, where 、 and are the weight coefficients of the abnormal CPU usage rate, abnormal memory usage rate, and abnormal disk I / O usage rate respectively, and satisfy + + = 1, and the specific values are set by the operator. CPU current and CPU normal respectively represent the current usage rate of the CPU and the normal usage rate of the CPU, MEN current and MEN normalrespectively represent the current memory usage rate and the normal memory usage rate, and IO urrent and IO ormal respectively represent the current disk I / O usage rate and the normal current disk I / O usage rate, and compare the calculated system performance anomaly value SPE with a second preset value.

[0013] As a further solution of the present invention: The specific manner in which the step four compares the calculated system performance anomaly value SPE with the second preset value is: If the system performance anomaly value SPE is greater than the second preset value, it indicates that the current system performance data has an impact on the device, and impact information is generated. On the contrary, if the system performance anomaly value SPE is less than the second preset value, the system performance anomaly value SPE is monitored periodically, and corresponding change information is generated.

[0014] An abnormal state monitoring device for computer room automation equipment includes a device information acquisition module, a device information analysis module, a sensor data analysis module, a performance data analysis module, and a monitoring information output module; The device information acquisition module is used to transmit the sensor data and performance data of the cluster device to the device information analysis module.

[0015] The device information analysis module is used to analyze the overall state of the cluster device based on the acquired sensor data and performance data to generate a device normal result or a device abnormal result, and at the same time transmit the device abnormal result to the sensor data analysis module and the performance data analysis module respectively; The sensor data analysis module is used to analyze the sensor data of the cluster device according to the acquired device abnormal result, specifically analyze the temperature, and determine the abnormal situation based on the temperature change situation to generate specific cause information, and at the same time transmit the specific cause information to the monitoring information output module; The performance data analysis module is used to analyze the performance data of the cluster device according to the acquired device abnormal result, separately analyze the network performance data and the system performance data, calculate the network performance value and the system performance value based on the corresponding data information for judgment, and at the same time determine the specific cause to generate specific cause information, and transmit the specific cause information to the monitoring information output module; The monitoring information output module is used to display the acquired monitoring information to the corresponding operator.

[0016] Beneficial effects The present invention provides a method and device for monitoring the abnormal state of computer room automation equipment. Compared with the prior art, it has the following beneficial effects: By calculating the average temperature gradient and comparing it with the gradient threshold, the present invention can timely detect abnormal temperature changes. Further, according to the temperature change rate judgment method, it distinguishes slow changes, sudden changes and local changes, and combines historical data and change characteristics to match the abnormal reasons. For network performance data, key indicators such as broadband occupancy value, packet loss rate and network latency are obtained, the network performance value NPV is calculated, and by comparing with the first preset value and analyzing the change situation within the time period T, the impact degree and change trend of network performance anomalies on the device are comprehensively evaluated. For system performance data, the system performance anomaly value SPE is calculated, considering multiple factors such as CPU, memory and disk I / O, and compared with the second preset value to accurately judge the impact of system performance data on the device, realizing in-depth insight into the device performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart of the method steps of the present invention; Figure 2 It is a block diagram of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment 1, please refer to Figure 1 , the present application provides a method for monitoring the abnormal state of computer room automation equipment, and the method specifically includes the following steps: Step 1: Collect sensor data and performance data of cluster devices, and the sensor data obtained here includes the temperature of cluster devices, and the performance data includes network performance data and system performance data.

[0020] Step 2: Compare the collected sensor data and performance data with the normal values to judge the overall state of the cluster devices, and at the same time generate a device normal result or a device abnormal result.

[0021] The obtained sensor data and performance data are respectively compared with the corresponding normal values. Here, the normal values are represented as the normal values corresponding to the sensor data, specifically the normal temperature value. For the sensor data, the normal range set by the temperature sensor is 20°C - 25°C. The normal values of the performance data are represented as the normal network performance value and the normal system performance value. For the performance data, in terms of network performance, taking the network bandwidth utilization rate as an example, the normal threshold is set to not exceed 70%. The normal threshold of the CPU usage rate in the system performance is set to not be higher than 80%. If both the sensor data and the performance data are the same as the normal values, it indicates that the overall state of the cluster device is normal, and at the same time, a device normal result is generated. If the sensor data or the performance data is different from the normal value, it indicates that the overall state of the cluster device is abnormal, and at the same time, a device abnormal result is generated. Here, it specifically means that any set of data in the sensor data or the performance data is different from the corresponding normal value. Specifically, for the generated device normal result, continuous monitoring of the cluster device is carried out to ensure the normal operation of the device.

[0022] For example, in the monitoring of a data center cluster device, if it is monitored that the temperature of a certain server remains stable at 23°C, which is within the normal temperature range, and if it is monitored that the network bandwidth utilization rate of the cluster is 50%, and at the same time, the CPU usage rate of a certain key server is actually monitored to be 60%, since both the sensor data and the performance data do not exceed their respective normal threshold ranges, it indicates that the device is operating stably at the current moment and no abnormal conditions occur. On the contrary, if the temperature sensor monitors that the temperature of a certain server suddenly rises to 30°C, exceeding the normal range of 20°C - 25°C, or the network bandwidth utilization rate suddenly soars to 90%, even if other data are normal, as long as such a set of data deviates from the normal threshold, it will be determined that the overall state of the cluster device is abnormal.

[0023] Step 3: Analyze the abnormal situation of the sensor data in the device abnormal result. Through specific analysis of the temperature and based on the temperature change situation, determine the abnormal situation and generate specific cause information.

[0024] Obtain the temperature data corresponding to the sensor data. Here, the temperature sensor data of any group in the cluster device is obtained. Multiple groups of sensors are set at different positions in the cluster device. At the same time, obtain the change of the temperature data within the time period t, and calculate the temperature gradient corresponding to the temperature data. Here, the gradient represents the change amount of temperature per unit time. At the same time, calculate the average gradient corresponding to the temperature gradient, and compare the calculated average gradient with the gradient threshold. The specific value of the gradient threshold is set by the operator. For example, under normal circumstances, the temperature in the server cabinet may not rise more than 2°C per hour. If the average gradient is greater than the gradient threshold, it indicates that the temperature data changes abnormally, and a temperature change abnormal signal is generated. On the contrary, if the average temperature gradient is less than the gradient threshold, it indicates that the temperature data changes normally, and a temperature change normal signal is generated; For example, set the time period to 30 minutes. If within these 30 minutes, the average gradient of the server cabinet is 2.5°C / hour, assuming the gradient threshold is set to 2°C / hour, since 2.5°C / hour is greater than 2°C / hour, the system will generate a temperature change abnormal signal at this time. And if within another time period, the average gradient is 1°C / hour, which is less than 2°C / hour, a temperature change normal signal will be generated.

[0025] Next, analyze the generated temperature change abnormal signal, establish a coordinate system of temperature gradient and time, and at the same time judge the change of the temperature gradient, and classify it into slow change, sudden change and local change. Here, the local change is specifically obtained according to the sensor data at the corresponding position. At the same time, the judgment of the change situation here is based on the temperature change rate judgment method to calculate the temperature change rate (temperature gradient), that is, the change amount of temperature per unit time. Distinguish slow change and sudden change by comparing the temperature change rate with the set threshold. For example, during the operation of a device, record the temperature data. If the temperature rises from 30°C to 30.1°C within 1 minute, the temperature change rate is 0.1°C / minute, which belongs to slow change; but if the temperature rises from 30°C to 30.6°C within 1 minute, the temperature change rate is 0.6°C / minute, which belongs to sudden change. Then analyze the abnormal reasons corresponding to slow change, sudden change and local change respectively; Analyze the slow change, obtain historical data, and at the same time obtain similar historical data, and obtain the abnormal reasons corresponding to the similar historical data. Then label the obtained abnormal reasons as i, and i = 1, 2,..., j, where j represents the number of abnormal reasons. At the same time, obtain the change characteristics corresponding to the slow change. Here, the change characteristics specifically represent the position of the corresponding sensor, and match the change characteristics with the abnormal reasons. At the same time, combine the average change corresponding to the slow change to screen the abnormal reasons to obtain the specific reasons, and generate specific reason information; For example, the abnormal causes corresponding to the slow changes in historical data include problems with the cooling system and increased device load. Among them, the problems with the cooling system include fan failures and blockages or damages to the heat sinks. The increase in device load specifically includes high loads caused by software or services and incompatibilities after hardware configuration upgrades. First, determine the sensor location corresponding to the abnormal temperature data. For example, if it is at the host location, further screening is performed to obtain the corresponding problems with the cooling system, and the specific causes are determined based on the corresponding temperature change values.

[0026] Fan failure: If the temperature of the device is rising slowly and continuously, there may be a problem with the cooling fan.

[0027] Blockage or damage to the heat sink: Excessive accumulation of impurities such as dust and oil on the heat sink will reduce its heat dissipation efficiency.

[0028] High load caused by software or services: When the software or services running on the device occupy too many resources, more heat will be generated.

[0029] Incompatibility after hardware configuration upgrade: If the device has been hardware-upgraded, such as adding memory or replacing the high-performance CPU, etc., it may cause the device temperature to rise due to compatibility issues between the new hardware and the original device or insufficient power supply.

[0030] The analysis of sudden changes and local changes is the same as that of the slow change process, and the specific causes are determined by combining the corresponding change characteristics and change means, and specific cause information is generated.

[0031] For the abnormal temperature that suddenly rises, possible problems include hardware short circuits, hardware component damages, computer room air conditioner failures, and external heat source influences. Among them, a short circuit in the internal circuit of the device is a serious cause for the sudden temperature rise. For example, situations such as the damage of electronic components and the breakage of the insulation layer of the circuit may cause a short circuit. A short circuit will cause the current to increase sharply. According to Joule's law, a large amount of heat will be generated by the device in a short time. Hardware component damage. For example, if there is a fault in the internal circuit of the CPU, its power consumption may increase abnormally, resulting in a sharp rise in temperature.

[0032] Computer room air conditioner failure: If there is a failure in the air conditioning system of the computer room, the ambient temperature will rise rapidly, resulting in an increase in the device temperature accordingly.

[0033] External heat source influence: The sudden appearance of an external heat source near the device, such as the approach of other heating devices or a fire, etc., will also cause the device temperature to rise suddenly.

[0034] Step 4: Analyze the abnormal performance data in the device abnormal results. Analyze the network performance data and system performance data separately, calculate the network performance value and system performance value according to the corresponding data information for judgment, and determine the specific reasons to generate specific reason information.

[0035] Analyze the network performance data, obtain the current broadband occupancy value, packet loss rate, and network latency of the cluster device, and compare them with the corresponding normal working values respectively. Here, the normal working values are the normal broadband occupancy value, normal packet loss rate, and normal network latency respectively. The specific values are set by the operator. Here, the network occupancy value is expressed as the average network speed within time t1. At the same time, screen out the abnormal network performance data. For example, if the obtained packet loss rate is abnormal here, further analyze the packet loss rate. Obtain the abnormal network performance data, and obtain the corresponding peak value and the average gradient of the performance data within time t1. Here, the peak value includes the maximum peak value and the minimum peak value. The gradient average is expressed as the average change difference of the abnormal network performance data within time t1. At the same time, calculate the change frequency corresponding to the abnormal network performance data, and substitute the obtained parameters into the formula Calculate the network performance value NPV, where Pmax and Pmin respectively represent the maximum peak value and the minimum peak value corresponding to the abnormal network performance data, F represents the change frequency, and G is the average gradient of the performance data. At the same time, compare the obtained network performance value NPV with the first preset value, and the specific value of the first preset value is set by the operator. If the network performance value NPV is greater than the first preset value, it means that the current abnormal network performance data has an impact on the device and generate an impact information. If the network performance value NPV is less than the first preset value, further analysis and judgment are required. Taking time T as the cycle, analyze the change of the network performance value corresponding to the time period T, and generate the corresponding change information at the same time. Here, the change information includes stable change and unstable change, and the judgment basis is to identify according to the change difference of the corresponding network performance value.

[0036] Analyze the system performance data, obtain the current abnormal CPU usage rate, abnormal memory usage rate, and abnormal disk I / O usage rate of the cluster device. The abnormal CPU usage rate represents the difference between the current CPU usage ratio and the normal usage ratio. For example, when the normal CPU usage rate is 45% and the current CPU usage rate is 75%, the corresponding abnormal CPU usage rate is 30%. Calculate the abnormal memory usage rate and abnormal disk I / O usage rate in the same way, and then substitute the obtained parameters into the formula Calculate the system performance anomaly value SPE, where 、 and They are the weight coefficients of the abnormal CPU usage rate, the abnormal memory usage rate, and the abnormal disk I / O usage rate respectively, and satisfy + + = 1. The specific values are set by the operator. CPU current and CPU normal represent the current CPU usage rate and the normal CPU usage rate respectively. MEN current and MEN normal represent the current memory usage rate and the normal memory usage rate respectively. IO urrent and IO ormal represent the current disk I / O usage rate and the normal disk I / O usage rate respectively; Compare the calculated system performance anomaly value SPE with the second preset value, and the specific value of the second preset value is set by the operator. If the system performance anomaly value SPE is greater than the second preset value, it means that the current system performance data has an impact on the device and generate impact information. Otherwise, if the system performance anomaly value SPE is less than the second preset value, perform periodic monitoring on the system performance anomaly value SPE and generate corresponding change information. Specifically, the change information here specifically represents the change situation of the system performance anomaly value, such as an increase or decrease change.

[0037] Example 2. Please refer to Figure 2 , this application provides a monitoring device for the abnormal state of computer room automation equipment. The device specifically includes: a device information acquisition module, a device information analysis module, a sensor data analysis module, a performance data analysis module, and a monitoring information output module.

[0038] The device information acquisition module is used to transmit the sensor data and performance data of the cluster device to the device information analysis module.

[0039] The device information analysis module is used to analyze the overall state of the cluster device based on the acquired sensor data and performance data to generate a device normal result or a device abnormal result, and at the same time transmit the device abnormal result to the sensor data analysis module and the performance data analysis module respectively. And the specific processing method here is the same as the processing process in step 2 of Example 1; The sensor data analysis module is used to analyze the sensor data of the cluster device according to the acquired device abnormal result. By specifically analyzing the temperature and determining the abnormal situation based on the temperature change situation, generate specific cause information, and at the same time transmit the specific cause information to the monitoring information output module. And the specific processing method is the same as the processing process in step 3 of Example 1; A performance data analysis module, which is used to analyze the performance data of cluster devices according to the obtained device exception results. By separately analyzing the network performance data and system performance data, and judging by calculating the network performance value and system performance value according to the corresponding data information, while determining the specific reasons to generate specific reason information, and transmitting the specific reason information to the monitoring information output module, and the processing method here is the same as the processing process of step four in Embodiment 1; A monitoring information output module, which is used to display the obtained monitoring information to the corresponding operators.

[0040] At the same time, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.

[0041] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical method of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A method for monitoring the abnormal state of computer room automation equipment, characterized in that, The method specifically includes the following steps: Step 1: Collect the sensor data and performance data of the cluster devices; Step 2: Compare the collected sensor data and performance data with the normal values to judge the overall state of the cluster devices, and at the same time generate a device normal result or a device abnormal result; Step 3: Analyze the abnormal situation of the sensor data in the device abnormal result. Specifically analyze the temperature, and determine the abnormal situation based on the temperature change to generate specific cause information; Step 4: Analyze the abnormal situation of the performance data in the device abnormal result. Specifically analyze the network performance data and system performance data separately, and judge by calculating the network performance value and system performance value according to the corresponding data information, and at the same time determine the specific cause to generate specific cause information.

2. The abnormal state monitoring method of a computer room automation device according to claim 1, characterized in that The specific method for the step 2 to judge the overall state of the cluster devices is: Compare the obtained sensor data and performance data with the corresponding normal values respectively. If both the sensor data and performance data are the same as the normal values, it means that the overall state of the cluster devices is normal, and at the same time generate a device normal result. If the sensor data or performance data is different from the normal value, it means that the overall state of the cluster devices is abnormal, and at the same time generate a device abnormal result.

3. A method for monitoring the abnormal state of computer room automation equipment according to claim 1, characterized in that, The specific way for the step 3 to analyze the abnormal situation of the sensor data in the device abnormal result is: Obtain the temperature data corresponding to the sensor data, and at the same time obtain the change situation of the temperature data within the time period t, calculate the temperature gradient corresponding to the temperature data, and calculate the gradient mean value corresponding to the temperature gradient. Compare the calculated gradient mean value with the gradient threshold. If the gradient mean value is greater than the gradient threshold, it means that the temperature data changes abnormally, and generate a temperature change abnormal signal. On the contrary, if the temperature gradient mean value is less than the gradient threshold, it means that the temperature data changes normally, and generate a temperature change normal signal, and at the same time analyze the temperature change abnormal signal.

4. The abnormal state monitoring method for a computer room automation device according to claim 3, characterized in that, The specific way for the step 3 to analyze the temperature change abnormal signal is: Establish a coordinate system of temperature gradient and time, and at the same time judge the change situation of the temperature gradient, and classify it into slow change, sudden change and local change. Then analyze the abnormal causes corresponding to slow change, sudden change and local change respectively; Analyze the slow change, obtain the historical data, and at the same time obtain the similar historical data, and obtain the abnormal causes corresponding to the similar historical data. Then label the obtained abnormal causes as i, and i = 1, 2, …, j, where j represents the number of abnormal causes. At the same time, obtain the change characteristics corresponding to the slow change, match the change characteristics with the abnormal causes, and screen the abnormal causes in combination with the change mean value corresponding to the slow change to obtain the specific cause, and generate specific cause information; The analysis of sudden change and local change is the same as the process of slow change, and determine the specific cause in combination with the corresponding change characteristics and change mean values, and generate specific cause information.

5. A method for monitoring the abnormal state of computer room automation equipment according to claim 1, characterized in that, The specific way for the step 4 to analyze the abnormal situation of the performance data in the device abnormal result is: Analyze the network performance data to obtain the current broadband occupancy value, packet loss rate, and network latency of the cluster devices, compare them with the corresponding normal working values respectively, and filter out the abnormal network performance data. Obtain abnormal network performance data, obtain the corresponding peak value and the average value of the performance data gradient within time t1, and at the same time calculate the change frequency corresponding to the abnormal network performance data. Substitute the obtained parameters into the formula Calculate the network performance value NPV, where Pmax and Pmin respectively represent the maximum peak value and the minimum peak value corresponding to the abnormal network performance data, F represents the change frequency, G is the average value of the performance data gradient. At the same time, compare the obtained network performance value NPV with the first preset value, and the specific value of the first preset value is set by the operator.

6. The abnormal state monitoring method for a computer room automation device according to claim 5, characterized in that, The specific method for comparing the obtained network performance value with the first preset value in step four is as follows: If the network performance value NPV is greater than the first preset value, it indicates that the current abnormal network performance data has an impact on the device, and an impact message is generated. If the network performance value NPV is less than the first preset value, further analysis and judgment are required. Taking time T as the cycle, analyze the change of the network performance value corresponding to the time cycle T, and generate the corresponding change information at the same time.

7. A method for monitoring the abnormal state of computer room automation equipment according to claim 1, characterized in that The specific method for analyzing the system performance data in step four is as follows: Obtain the current CPU abnormal utilization rate, memory abnormal utilization rate, and disk I / O abnormal utilization rate of the cluster device, and then substitute the obtained parameters into the formula Calculate to obtain the system performance abnormal value SPE, where 、 and are the weight coefficients of the CPU abnormal utilization rate, memory abnormal utilization rate, and disk I / O abnormal utilization rate respectively, and satisfy + + = 1. The specific values are set by the operator. CPU current and CPU normal represent the current CPU utilization rate and the normal CPU utilization rate respectively. MEN current and MEN normal represent the current memory utilization rate and the normal memory utilization rate respectively. IO urrent and IO ormal represent the current disk I / O utilization rate and the normal disk I / O utilization rate respectively. Compare the calculated system performance abnormal value SPE with the second preset value.

8. A method for monitoring the abnormal state of computer room automation equipment according to claim 7, characterized in that, The specific method for comparing the calculated system performance anomaly value SPE with the second preset value in step four is as follows: If the system performance anomaly value SPE is greater than the second preset value, it indicates that the current system performance data has an impact on the device, and an impact message is generated. On the contrary, if the system performance anomaly value SPE is less than the second preset value, the system performance anomaly value SPE is monitored periodically, and the corresponding change information is generated.

9. An abnormal state monitoring device for computer room automation equipment, which is used to execute the abnormal state monitoring method for computer room automation equipment according to any one of claims 1-8, and is characterized in that, It includes a device information acquisition module, a device information analysis module, a sensor data analysis module, a performance data analysis module, and a monitoring information output module; The device information acquisition module is used to transmit the sensor data and performance data of the cluster devices to the device information analysis module; The device information analysis module is used to analyze the overall state of the cluster devices based on the acquired sensor data and performance data to generate a device normal result or a device abnormal result, and at the same time transmit the device abnormal result to the sensor data analysis module and the performance data analysis module respectively; The sensor data analysis module is used to analyze the sensor data of the cluster devices according to the obtained device abnormal result. By specifically analyzing the temperature and determining the abnormal situation based on the temperature change, specific cause information is generated, and at the same time the specific cause information is transmitted to the monitoring information output module; The performance data analysis module is used to analyze the performance data of the cluster devices according to the obtained device abnormal result. By separately analyzing the network performance data and the system performance data, and judging by calculating the network performance value and the system performance value according to the corresponding data information, and at the same time determining the specific cause to generate specific cause information, and transmitting the specific cause information to the monitoring information output module; The monitoring information output module is used to display the acquired monitoring information to the corresponding operator.

Citation Information

Patent Citations

  • Equipment abnormity monitoring method, device and system

    CN117992301A

Cited By

  • Running state optimization method and device of server and medium

    CN120704993A