A Method and System for Monitoring the BMC Operating Status of a Mini Computer Host

By collecting a variety of monitoring data in the microcomputer host and dynamically adjusting the dejitter time window, the problem of low alarm detection accuracy in the prior art is solved, and the accuracy of monitoring and system reliability are improved.

CN119669002BActive Publication Date: 2025-06-17SHENZHEN JIMOKE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510186242.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

In the prior art, a fixed dejitter time window is used to perform alarm detection based on the CPU temperature, which has low accuracy and may lead to system failure or false alarms.

Method used

By collecting the CPU temperature curve, heat sink speed curve and CPU utilization curve monitored by the BMC, combining the degree of temperature abnormality, heat dissipation abnormality and CPU utilization stability, the de-jitter time window is dynamically adjusted to improve the accuracy of alarm detection.

Benefits of technology

By dynamically adjusting the dejitter time window, the accuracy of BMC operating status monitoring of the microcomputer host is improved, the risk of false alarms and faults is reduced, and the reliability of the system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669002B_ABST
    Figure CN119669002B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of anomaly monitoring, and specifically relates to a method and system for monitoring the BMC operating state of a micro computer host. According to the abnormal change of the CPU temperature after the influence of the radiator speed on the CPU temperature and the abnormal rising trend of the CPU temperature, the degree of temperature anomaly is determined; according to the characteristic that the CPU temperature usually decreases after the radiator speed increases, the degree of heat dissipation anomaly is determined; and on the basis of both, the risk of triggering an alarm is determined in combination with the stability of the host operating state; according to the anomaly in the numerical value of the CPU temperature at all times from the occurrence of over-temperature to the alarm, the degree of alarm response requirement is determined; a more accurate dynamic anti-jitter time window is obtained by adaptively determining the window adjustment coefficient in combination with the risk of triggering an alarm and the degree of alarm response requirement, improving the accuracy of alarm detection, and making the effect of monitoring the BMC operating state of the micro computer host according to the dynamic anti-jitter time window better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly monitoring, and particularly to a method and system for monitoring the operating status of a BMC in a microcomputer host. Background Art

[0002] With the continuous development of information technology, microcomputer hosts are increasingly widely used in various information systems. From personal computers to enterprise-level servers, and then to edge computing devices, microcomputer hosts have become an indispensable important device in various computing scenarios due to their small size, strong performance, high energy efficiency, etc. In a microcomputer host, the Baseboard Management Controller (BMC) is a key management component, mainly responsible for functions such as hardware monitoring, diagnosis, fault detection, and remote management of the host. Therefore, real-time monitoring of the operating status of the BMC can help system administrators timely discover potential problems, conduct fault diagnosis and preventive maintenance, and ensure the normal operation of the system.

[0003] Considering that there is a certain error in the temperature detected by the temperature sensor, the existing alarm mechanism in the BMC usually adds a debounce function on the basis of temperature monitoring to avoid false alarms, so as to perform more accurate monitoring of the operating status of the BMC in the microcomputer host. However, the existing technology usually uses a fixed debounce time window for alarm detection based on the CPU temperature. However, the setting of the debounce time window means that there is a time delay in the alarm. If the anomaly of the microcomputer host is relatively serious, then when using a fixed debounce time window for alarm detection, it may lead to system failure or damage due to the lack of timely alarm; on the contrary, if the anomaly of the microcomputer host is relatively minor or there is no anomaly, then when using a fixed debounce time window for alarm detection, false alarms may occur; therefore, the accuracy of the existing technology for alarm detection using a fixed debounce time window based on the CPU temperature is relatively low, resulting in a poor monitoring effect on the operating status of the BMC in the microcomputer host. Summary of the Invention

[0004] In order to solve the problem of relatively low accuracy of the existing technology for alarm detection using a fixed debounce time window based on the CPU temperature, the purpose of this application is to provide a method and system for monitoring the operating status of a BMC in a microcomputer host, and the specific technical solutions adopted are as follows:

[0005] The first aspect of this application provides a method for monitoring the operating status of a BMC in a microcomputer host, including:

[0006] During the operation of the microcomputer host, collect the prior alarm time, CPU temperature curve, radiator rotation speed curve, and CPU utilization rate curve for each monitoring time period monitored by the BMC; obtain the preset debounce time window and the CPU over-temperature start time when the CPU temperature first exceeds the preset temperature threshold in each monitoring time period;

[0007] Determine the degree of temperature anomaly according to the change correlation between the CPU temperature curve and the radiator rotation speed curve and the overall upward trend of the CPU temperature in the CPU temperature curve; determine the degree of heat dissipation anomaly according to the matching situation of the CPU temperature change during the rise of the radiator rotation speed; determine the alarm risk of each monitoring time period according to the overall stability of the CPU utilization rate curve, the degree of temperature anomaly, and the degree of heat dissipation anomaly;

[0008] In each monitoring time period, determine the degree of alarm response requirement for each monitoring time period according to the time length between the prior alarm time and the CPU over-temperature start time, the CPU temperature distribution, and the overall CPU temperature;

[0009] Determine the window adjustment coefficient for each monitoring time period according to the degree of alarm response requirement and the alarm risk; adjust the preset debounce time window according to the window adjustment coefficient to determine the dynamic debounce time window for each monitoring time period; monitor the BMC operating status of the microcomputer host according to the dynamic debounce time window.

[0010] Further, the process of obtaining the degree of temperature anomaly includes:

[0011] On the CPU temperature curve, take the slope value of the line connecting the data point corresponding to each sampling time and the data point corresponding to the previous sampling time as the temperature slope value at each sampling time; determine the temperature rise trend characteristic value according to the mean value of the temperature slope values at all sampling times;

[0012] On the CPU temperature curve, perform a negative correlation mapping on the difference between the mean value of the CPU temperatures at all sampling times and the preset temperature threshold to determine the proximity of the threshold temperature;

[0013] Determine the temperature rotation correlation degree according to the negative correlation mapping of the Pearson correlation coefficient between the CPU temperature curve and the radiator rotation speed curve;

[0014] Determine the temperature anomaly degree for each monitoring time period according to the temperature-rotation speed correlation degree, the threshold temperature proximity degree, and the temperature rising trend eigenvalue; the temperature-rotation speed correlation degree has a negative correlation with the temperature anomaly degree; both the threshold temperature proximity degree and the temperature rising trend eigenvalue have a positive correlation with the temperature anomaly degree.

[0015] Further, the process of determining the temperature anomaly degree for each monitoring time period according to the temperature-rotation speed correlation degree, the threshold temperature proximity degree, and the temperature rising trend eigenvalue includes:

[0016] Normalize the product among the negative correlation mapping value of the temperature-rotation speed correlation degree, the threshold temperature proximity degree, and the temperature rising trend eigenvalue to determine the temperature anomaly degree for each monitoring time period.

[0017] Further, the process of obtaining the heat dissipation anomaly degree includes:

[0018] On the radiator rotation speed curve, take the slope value of the line connecting the data point corresponding to each sampling moment and the data point corresponding to the previous sampling moment as the rotation speed slope value for each sampling moment; take the sampling moments with the rotation speed slope value greater than 0 as the rotation speed rising moments;

[0019] Perform negative correlation mapping on the difference between the rotation speed slope value and the corresponding temperature slope value at each rotation speed rising moment to determine the instantaneous heat dissipation fault degree at each rotation speed rising moment; determine the heat dissipation anomaly degree for each monitoring time period according to the mean value of the instantaneous heat dissipation fault degrees at all rotation speed rising moments.

[0020] Further, the process of obtaining the alarm trigger risk degree includes:

[0021] Perform negative correlation mapping on the standard deviation of the CPU utilization rate at all sampling moments on the CPU utilization rate curve to determine the CPU utilization rate stability degree for each monitoring time period;

[0022] Normalize the product of the negative correlation mapping value of the CPU utilization rate stability degree, the temperature anomaly degree, and the heat dissipation anomaly degree to determine the alarm trigger risk degree for each monitoring time period.

[0023] Further, the process of obtaining the alarm response requirement degree includes:

[0024] Take the monitoring time period in which the CPU over-temperature starting moment is prior to the prior alarm moment in time sequence as the analysis time period; there are a CPU over-temperature starting moment and a prior alarm moment in the analysis time period; set the alarm response requirement degree for other monitoring time periods outside the analysis time period to the preset requirement degree;

[0025] In each analysis time period, all the sampling times between the CPU over-temperature start time and the prior warning time are taken as the delay times; the delay times corresponding to the CPU temperature greater than the preset temperature threshold are taken as the reference abnormal times; the difference between the CPU temperature at each reference abnormal time and the preset temperature threshold is taken as the instantaneous over-temperature degree at each reference abnormal time; the overall over-temperature degree of each analysis time period is determined according to the accumulated value of the instantaneous over-temperature degrees at all the reference abnormal times;

[0026] Normalize the product of the number of delay times and the overall over-temperature degree in each analysis time period to determine the warning response requirement degree of each analysis time period.

[0027] Further, the process of obtaining the window adjustment coefficient includes:

[0028] Perform a negative correlation mapping on the product of the warning response requirement degree and the risk of triggering an alarm to determine the window adjustment coefficient of each monitoring time period.

[0029] Further, the process of obtaining the dynamic debounce time window includes:

[0030] Multiply the window adjustment coefficient by the preset debounce time window to determine the dynamic debounce time window of each monitoring time period.

[0031] Further, the process of monitoring the BMC operating state of the microcomputer host according to the dynamic debounce time window includes:

[0032] Take the dynamic debounce time window of each monitoring time period as the final debounce time window for subsequent monitoring time periods until an alarm occurs and a new dynamic debounce time window is re-determined; monitor the BMC operating state of the microcomputer host according to the final debounce time window of each monitoring time period in combination with the corresponding CPU temperature.

[0033] In a second aspect, the present application provides a BMC operating state monitoring system for a microcomputer host, the system includes:

[0034] A data acquisition and preprocessing module, configured to collect, during the operation of the microcomputer host, the prior warning time, the CPU temperature curve, the radiator rotation speed curve, and the CPU utilization curve of each monitoring time period monitored by the BMC; obtain the preset debounce time window and the CPU over-temperature start time when the CPU temperature first exceeds the preset temperature threshold in each monitoring time period;

[0035] The first determination module is configured to determine the degree of temperature anomaly according to the variation correlation between the CPU temperature curve and the radiator rotation speed curve and the overall upward trend of the CPU temperature in the CPU temperature curve; determine the degree of heat dissipation anomaly according to the matching of the CPU temperature change during the radiator rotation speed increase; and determine the risk of alarm triggering for each monitoring time period according to the overall stability of the CPU utilization curve, the degree of temperature anomaly, and the degree of heat dissipation anomaly.

[0036] The second determination module is configured to determine the degree of alarm response requirement for each monitoring time period according to the time length between the prior alarm time and the temperature anomaly time, the CPU temperature distribution, and the overall magnitude of the CPU temperature in each monitoring time period.

[0037] The operating state monitoring module is configured to determine the window adjustment coefficient for each monitoring time period according to the degree of alarm response requirement and the risk of alarm triggering; adjust the preset debounce time window according to the window adjustment coefficient to determine the dynamic debounce time window for each monitoring time period; and monitor the BMC operating state of the microcomputer host according to the dynamic debounce time window.

[0038] In a third aspect, the present application provides a computer device, including a memory and a processor. The memory is used to store computer program code, and the processor is used to call and run the computer program code from the memory to execute the method according to the first aspect or any embodiment of the first aspect of the present application.

[0039] In a fourth aspect, the present application provides a computer program product, where the computer program product includes computer program code, and when the computer program code is executed, it is used to execute the method according to the first aspect or any embodiment of the first aspect of the present application.

[0040] In a fifth aspect, the present application provides a computer-readable storage medium, where the computer-readable storage medium stores computer program code, and when the computer program code is executed, it is used to execute the method according to the first aspect or any embodiment of the first aspect of the present application.

[0041] The present application has the following beneficial effects:

[0042] First, this application determines the degree of temperature abnormality based on the abnormal change in the CPU temperature and the abnormal rising trend of the CPU temperature after the CPU temperature is affected by the radiator speed. Then, it determines the degree of heat dissipation abnormality based on the characteristic that the CPU temperature usually decreases after the radiator speed increases. Then, based on the degree of temperature abnormality and the degree of heat dissipation abnormality, it combines the CPU utilization curve indicating the stable operation state of the host to determine the risk of triggering an alarm. Then, based on the numerical abnormality of the CPU temperature between the starting moment of CPU overheating and the alarm moment, it determines the degree of alarm response requirement. Thus, it adaptively determines a more accurate window adjustment coefficient for adjusting the debounce time window by combining the risk of triggering an alarm and the degree of alarm response requirement. Finally, it adaptively determines a more accurate dynamic debounce time window based on the window adjustment coefficient, making the accuracy of alarm detection based on the dynamic debounce time window higher, solving the problem of low accuracy of alarm detection using a fixed debounce time window based on the CPU temperature in the prior art, that is, making the monitoring effect of the BMC operation state of the microcomputer host based on the dynamic debounce time window better. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0044] Figure 1 It is a flowchart of a method for monitoring the BMC operation state of a microcomputer host provided by an embodiment of the present invention;

[0045] Figure 2 It is a structural diagram of a system for monitoring the BMC operation state of a microcomputer host provided by an embodiment of the present invention;

[0046] Figure 3 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following specifically describes, in conjunction with the accompanying drawings and preferred embodiments, a method and system for monitoring the BMC operating status of a microcomputer host according to the present invention, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment, and specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form. In addition, the terms "first" and "second" are only used for descriptive purposes and should not be construed as implying or indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs.

[0049] The following specifically describes the specific solution of a method and system for monitoring the BMC operating status of a microcomputer host provided by the present invention in conjunction with the accompanying drawings.

[0050] An embodiment of the present application provides a method for monitoring the BMC operating status of a microcomputer host. Please refer to Figure 1 , which shows a flowchart of a method for monitoring the BMC operating status of a microcomputer host provided by an embodiment of the present invention. The method includes:

[0051] Step S101: During the operation of the microcomputer host, collect the prior warning time, CPU temperature curve, radiator speed curve, and CPU utilization curve for each monitoring time period monitored by the BMC; obtain the preset debounce time window and the CPU over-temperature start time when the CPU temperature first exceeds the preset temperature threshold in each monitoring time period.

[0052] In a specific implementation manner of the embodiment of the present invention, the CPU temperature at each sampling moment is collected by a temperature sensor, and the fan speed and CPU utilization rate of the radiator at each sampling moment are collected by the hardware monitoring system of the microcomputer host; and the CPU temperature, the fan speed of the radiator, and the CPU utilization rate are transmitted to the monitoring module of the BMC to complete the collection; then, all the CPU temperatures in each monitoring time period are arranged in time series and then curve-fitted to obtain a CPU temperature curve; all the fan speeds of the radiators in each monitoring time period are normalized and then arranged in time series and then curve-fitted to obtain a radiator speed curve; all the CPU utilization rates in each monitoring time period are arranged in time series and then curve-fitted to obtain a CPU utilization rate curve; wherein, the time length of each monitoring time period is set to 1 minute, the sampling frequency is set to collect once per second, and the time length of the preset debounce time window is set to 5 seconds, which can be adjusted according to the specific implementation environment.

[0053] In a specific implementation manner of the embodiment of the present invention, the preset temperature threshold is set to 90 degrees Celsius and can be adjusted by itself; the prior warning moment is the moment corresponding to the trigger warning mechanism monitored in each monitoring time period; for each monitoring time period, if the time length of continuously occurring CPU temperature greater than the preset temperature threshold is greater than the corresponding debounce time window, the warning mechanism will be triggered.

[0054] Step S102: Determine the degree of temperature abnormality according to the change correlation between the CPU temperature curve and the radiator speed curve and the overall rising trend of the CPU temperature in the CPU temperature curve; determine the degree of heat dissipation abnormality according to the matching situation of the CPU temperature change during the rising process of the radiator speed; determine the alarm risk of each monitoring time period according to the overall stability of the CPU utilization rate curve, the degree of temperature abnormality, and the degree of heat dissipation abnormality.

[0055] First of all, the problem to be solved is that the accuracy of alarm detection using a fixed debounce time window according to the CPU temperature is relatively low. The reason is that the setting of the debounce time window means that there is a time delay in the alarm. If the abnormality of the microcomputer host is relatively serious, then when using a fixed debounce time window for alarm detection, it may lead to system failure or damage due to the failure to alarm in time; on the contrary, if the abnormality of the microcomputer host is relatively minor or there is no abnormality, then when using a fixed debounce time window for alarm detection, false alarms may occur. Therefore, it is necessary to adaptively adjust the debounce time window in combination with the abnormal state of the microcomputer host.

[0056] When the microcomputer host runs abnormally, the temperature usually changes abnormally. Therefore, in terms of analyzing temperature, first, according to the characteristic that the more obvious the upward trend of the CPU temperature is, the more likely it is to trigger the alarm mechanism, the closer the temperature is to abnormality. And the greater the radiator speed is, the stronger the ability to suppress the CPU temperature. If the speed adjustment of the radiator fan follows the change of the CPU temperature, that is, when the change correlation between the CPU temperature and the radiator speed is relatively high, the CPU temperature usually decreases as the radiator speed increases, resulting in a lower possibility of triggering the warning mechanism. Therefore, first, determine the degree of temperature abnormality based on the change correlation between the CPU temperature curve and the radiator speed curve and the overall upward trend of the CPU temperature in the CPU temperature curve. The greater the degree of temperature abnormality is, the higher the risk of triggering an alarm is.

[0057] Preferably, in some possible implementation manners of the embodiments of the present invention, the process of obtaining the degree of temperature abnormality includes:

[0058] On the CPU temperature curve, the slope value of the line connecting the data point corresponding to each sampling moment and the data point corresponding to the previous sampling moment is used as the temperature slope value of each sampling moment; according to the mean value of the temperature slope values of all sampling moments, determine the temperature upward trend characteristic value; on the CPU temperature curve, perform a negative correlation mapping on the difference between the mean value of the CPU temperatures of all sampling moments and the preset temperature threshold to determine the proximity degree of the threshold temperature; according to the negative correlation mapping of the Pearson correlation coefficient between the CPU temperature curve and the radiator speed curve, determine the temperature-speed correlation degree. In a specific implementation manner of the embodiments of the present invention, the preset temperature threshold is set to 90 degrees and can be adjusted according to the specific implementation environment.

[0059] According to the process of obtaining the temperature slope value, the greater the mean value of the temperature slope values of all sampling moments, that is, the greater the temperature upward trend characteristic value, the more obvious the overall upward trend of the CPU temperature curve is, and the more it has the characteristic of triggering the alarm mechanism. And the greater the proximity degree of the threshold temperature is, the closer the CPU temperature at each sampling moment is to the preset temperature threshold, which means the closer the CPU temperature is to the alarm value, and the more obvious the corresponding alarm trend is, and the higher the possibility of triggering the alarm mechanism is. And when the temperature-speed correlation degree obtained by the negative correlation mapping of the Pearson correlation coefficient of the CPU temperature curve and the radiator speed curve is greater, it means that the CPU temperature and the radiator speed are more negatively correlated, which is more in line with the characteristic that the CPU temperature usually decreases as the radiator speed increases, and the corresponding correlation is greater, and the possibility of triggering the warning mechanism is lower.

[0060] Therefore, further according to the temperature-rotation speed correlation degree, the proximity to the threshold temperature, and the temperature rise trend eigenvalue, the temperature anomaly degree of each monitoring time period is determined; the temperature-rotation speed correlation degree is negatively correlated with the temperature anomaly degree; both the proximity to the threshold temperature and the temperature rise trend eigenvalue are positively correlated with the temperature anomaly degree. Preferably, the process of determining the temperature anomaly degree of each monitoring time period according to the temperature-rotation speed correlation degree, the proximity to the threshold temperature, and the temperature rise trend eigenvalue includes: normalizing the product between the negative correlation mapping value of the temperature-rotation speed correlation degree, the proximity to the threshold temperature, and the temperature rise trend eigenvalue to determine the temperature anomaly degree of each monitoring time period.

[0061] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the temperature anomaly degree is represented by the formula: ; where is the temperature anomaly degree of the th monitoring time period; is the number of sampling moments in the th monitoring time period; is the temperature slope value of the th sampling moment in the th monitoring time period; is the temperature rise trend eigenvalue of the th monitoring time period; is the Pearson correlation coefficient between the CPU temperature curve and the radiator rotation speed curve in the th monitoring time period; is the temperature-rotation speed correlation degree of the th monitoring time period; is the average value of the CPU temperatures of all sampling moments in the th monitoring time period; is the preset temperature threshold; is the proximity to the threshold temperature of the th monitoring time period; is the absolute value symbol. is the negative correlation mapping value of the temperature-rotation speed correlation degree of the th monitoring time period; is the exponential function with the natural constant as the base; is the linear normalization function, which is normalized between 0 and 1 to prevent the temperature anomaly degree from taking negative values and affecting subsequent calculations.

[0062] Then, for the radiator, under normal circumstances, when the radiator speed increases, it will blow out faster air, and the rising trend of temperature will be suppressed or even show a downward trend. If the temperature still shows an upward trend when the radiator speed increases, it indicates that there are abnormalities or faults in the radiator-related components such as heat pipes and thermal grease. The higher the degree of radiator abnormality, the higher the risk of triggering an alarm. Therefore, the degree of heat dissipation abnormality is determined based on the matching situation of the CPU temperature change during the increase in radiator speed. In addition, the stability of the CPU utilization rate can reflect the operating state of the microcomputer host to a certain extent. Therefore, based on the degree of temperature abnormality and the degree of heat dissipation abnormality, the risk of triggering an alarm is determined by combining the overall stability of the CPU utilization rate curve. The greater the risk of triggering an alarm, the more abnormal the microcomputer host, and the smaller the corresponding debounce time window should be.

[0063] Preferably, in a specific implementation manner of the embodiment of the present invention, the process of obtaining the degree of heat dissipation abnormality includes:

[0064] On the radiator speed curve, the slope value of the line connecting the data point corresponding to each sampling moment and the data point corresponding to the previous sampling moment is used as the speed slope value of each sampling moment. The sampling moments with a speed slope value greater than 0 are used as the speed increase moments. The difference between the speed slope value and the corresponding temperature slope value at each speed increase moment is negatively correlated to determine the instantaneous heat dissipation fault degree at each speed increase moment. The degree of heat dissipation abnormality for each monitoring time period is determined based on the average value of the instantaneous heat dissipation fault degrees at all speed increase moments. At each speed increase moment, if the temperature shows a more downward trend, it indicates that the radiator state is more normal. Therefore, when the instantaneous heat dissipation fault degree obtained by negatively correlating the speed slope value and the corresponding temperature slope value is smaller, the radiator state at the corresponding sampling moment is more normal. Further, the degree of heat dissipation abnormality characterizing the radiator is determined by combining all sampling moments.

[0065] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the degree of heat dissipation abnormality is expressed by the formula: ; where is the degree of heat dissipation abnormality of the th monitoring time period; is the number of speed increase moments in the th monitoring time period; is the speed slope value of the rd speed increase moment in the th monitoring time period; is the temperature slope value of the th speed increase moment in the th monitoring time period; is the th in the The instantaneous heat dissipation fault degree at a rotational speed rising moment; is an exponential function with the natural constant as the base.

[0066] Preferably, in a specific implementation manner of the embodiment of the present invention, the process of obtaining the alarm risk includes:

[0067] Perform a negative correlation mapping on the standard deviation of the CPU utilization rate at all sampling moments on the CPU utilization rate curve to determine the stability degree of the CPU utilization rate in each monitoring time period. The larger the standard deviation of the CPU utilization rate, that is, the smaller the stability degree of the CPU utilization rate, the more unstable the CPU utilization rate is, the more abnormal the operating state of the microcomputer host is, the greater the possibility of abnormality, and the higher the alarm risk. Since the greater the degree of radiator abnormality, the greater the degree of temperature abnormality, and the smaller the stability degree of the CPU utilization rate, the higher the alarm risk. Therefore, further normalize the product of the negative correlation mapping value of the CPU utilization rate stability degree, the degree of temperature abnormality, and the degree of heat dissipation abnormality to determine the alarm risk in each monitoring time period, so that the greater the alarm risk, the more abnormal the microcomputer host is, and the smaller the corresponding debounce time window should be.

[0068] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the alarm risk is expressed by the formula: ; where is the alarm risk of the th monitoring time period; is the standard deviation of the CPU utilization rate at all sampling moments on the CPU utilization rate curve in the th monitoring time period; is an exponential function with the natural constant as the base; is the stability degree of the CPU utilization rate in the th monitoring time period; is the degree of heat dissipation abnormality in the th monitoring time period; is the degree of temperature abnormality in the th monitoring time period; is a linear normalization function.

[0069] Step S103: In each monitoring time period, determine the degree of alarm response requirement in each monitoring time period according to the time length between the prior alarm moment and the CPU over-temperature starting moment, the CPU temperature distribution, and the overall CPU temperature size.

[0070] Furthermore, it is necessary to consider that the purpose of setting the debounce time window, that is, the purpose of setting the delay, is to reduce the false alarm rate of the alarm mechanism. However, if the delay is too high, it will lead to insufficient response speed of the alarm mechanism to real anomalies. Therefore, it is necessary to shorten the delay, that is, shorten the debounce time window. Considering that blindly shortening the delay will lead to an increase in the false alarm rate of the alarm, in order to adjust the debounce time window more reasonably and scientifically, it is further necessary to adjust it in combination with the alarm response requirements of each monitoring time period. For a monitoring time period where there are both a prior alarm moment when an alarm has been issued and a CPU over-temperature start moment, when the time length between the prior alarm moment and the temperature anomaly moment is relatively long, it indicates that the time delay between the occurrence of a temperature exceeding the threshold and the occurrence of an alarm behavior is relatively long. To a certain extent, it shows that the time lag of the alarm mechanism in this monitoring time period is relatively high. Then, the corresponding response delay should be reduced. And between the prior alarm moment and the CPU over-temperature start moment, the more the number of moments exceeding the preset temperature threshold and the higher the exceeded temperature, it indicates that the more urgent it is to trigger an alarm, and the corresponding time delay should be smaller. Therefore, considering the time length between the prior alarm moment and the temperature anomaly moment, the CPU temperature distribution, and the overall CPU temperature, the degree of alarm response requirement is determined. When the degree of alarm response requirement is greater, the corresponding debounce time window should be smaller.

[0071] Preferably, in a specific implementation manner of the embodiment of the present invention, the process of obtaining the degree of alarm response requirement includes:

[0072] The monitoring time period in which the CPU over-temperature start moment is before the prior alarm moment in terms of time sequence is used as the analysis time period. There are a CPU over-temperature start moment and a prior alarm moment in the analysis time period. The alarm response requirement degrees of other monitoring time periods outside the analysis time period are set to a preset requirement degree. In a specific implementation manner of the embodiment of the present invention, the preset requirement degree is set to 1, and it can be adjusted according to the specific implementation environment. Since the calculation of the alarm response requirement degree depends on the time delay from the moment of over-temperature to the moment of alarm, other situations that do not meet this condition cannot be applied to subsequent analysis. Therefore, for the integrity of the embodiment, a preset value is set for the alarm response requirement degree in other situations.

[0073] In each analysis time period, all sampling times between the CPU over-temperature start time and the prior warning time are taken as delay times; the delay times corresponding to the CPU temperature being greater than the preset temperature threshold are taken as reference abnormal times; the difference between the CPU temperature at each reference abnormal time and the preset temperature threshold is taken as the instantaneous over-temperature degree at each reference abnormal time; the overall over-temperature degree of each analysis time period is determined according to the cumulative value of the instantaneous over-temperature degrees of all reference abnormal times. First, the more the number of delay times, the higher the delay, and the slower the true abnormal response speed of the corresponding warning mechanism, and the higher the need to shorten the debounce time window; when the number of reference abnormal times is more, it means that the instantaneous over-temperature occurs more frequently or the over-temperature lasts longer, and the accumulated thermal stress will further affect the service life of the hardware, so the need to shorten the debounce time window should also be higher; when the overall over-temperature degree is greater, it means that the overall amplitude exceeding the reasonable temperature of the CPU, that is, the preset temperature threshold, is greater, and a large amplitude exceeding the reasonable temperature of the CPU may cause permanent damage or failure; therefore, when the overall over-temperature degree is greater, the need to shorten the debounce time window should also be higher. Therefore, further normalize the product of the number of delay times and the overall over-temperature degree in each analysis time period to determine the warning response demand degree of each analysis time period.

[0074] In a specific implementation manner of the embodiment of the present invention, the process of obtaining the warning response demand degree is expressed by the formula: ; where is the warning response demand degree of the th analysis time period; is the number of delay times in the th analysis time period; is the number of reference abnormal times in the th analysis time period; is the CPU temperature of the th reference abnormal time in the th analysis time period; is the preset temperature threshold; is the instantaneous over-temperature degree of the th reference abnormal time in the th analysis time period; is the overall over-temperature degree of the th analysis time period; is the linear normalization function.

[0075] Step S104: Determine the window adjustment coefficient for each monitoring time period according to the degree of alarm response requirement and the risk of triggering an alarm; adjust the preset debounce time window according to the window adjustment coefficient to determine the dynamic debounce time window for each monitoring time period; monitor the BMC operating status of the microcomputer host according to the dynamic debounce time window.

[0076] Finally, combine the risk of triggering an alarm and the degree of alarm response requirement to jointly determine the window adjustment coefficient representing the adjustment of the debounce time window, and adaptively determine the dynamic debounce time window according to the window adjustment coefficient, so that the accuracy of alarm detection according to the dynamic debounce time window is higher, that is, the effect of monitoring the BMC operating status of the microcomputer host according to the dynamic debounce time window is improved.

[0077] Preferably, in some possible implementation manners of the embodiments of the present invention, the process of obtaining the window adjustment coefficient includes:

[0078] Since the greater the degree of alarm response requirement and the greater the risk of triggering an alarm, the corresponding debounce time window should be smaller; therefore, further perform a negative correlation mapping on the product between the degree of alarm response requirement and the risk of triggering an alarm to determine the window adjustment coefficient for each monitoring time period. In a specific implementation manner of the embodiments of the present invention, the process of obtaining the window adjustment coefficient is expressed by the formula: ; where is the window adjustment coefficient for the th monitoring time period; is the degree of alarm response requirement for the th monitoring time period; is the risk of triggering an alarm for the th monitoring time period; is a preset negative correlation adjustment constant, and its value is greater than or equal to 1. In the embodiments of the present invention, it is set to 1.3, so that the obtained window adjustment coefficient can have the ability to increase and decrease the debounce time window, and the preset negative correlation adjustment constant can be adjusted by itself.

[0079] Preferably, in some possible implementation manners of the embodiments of the present invention, the process of obtaining the dynamic debounce time window includes: multiplying the window adjustment coefficient by the preset debounce time window to determine the dynamic debounce time window for each monitoring time period; in terms of the formula, it is expressed as: ; where is the dynamic debounce time window for the th monitoring time period; is the preset debounce time window; is the The window adjustment coefficient for each monitoring time period. It should be noted that each debounce time window participating in the calculation by default represents the corresponding time length, because the debounce time window is essentially a time length for debounce detection.

[0080] Preferably, in some possible implementation manners of the embodiments of the present invention, the process of monitoring the BMC operating state of the microcomputer host according to the dynamic debounce time window includes: using the dynamic debounce time window of each monitoring time period as the final debounce time window for subsequent monitoring time periods until a warning occurs and then re-determining a new dynamic debounce time window; monitoring the BMC operating state of the microcomputer host according to the final debounce time window of each monitoring time period in combination with the corresponding CPU temperature. Since the dynamic debounce time window corresponding to each monitoring time period can only be calculated after the end of each monitoring time period, the dynamic debounce time window of each monitoring time period is used as the final debounce time window for subsequent monitoring time periods until a warning occurs and then the dynamic update of the debounce time window is performed. It should be noted that the application of the debounce time window in this application is specifically: only when the duration of the CPU temperature exceeding the preset temperature threshold is greater than the time length of the corresponding debounce time window, a warning message is issued.

[0081] In summary, for a method for monitoring the BMC operating state of a microcomputer host proposed in this application, first, according to the abnormal change of the CPU temperature after the influence of the radiator speed on the CPU temperature and the abnormal rising trend of the CPU temperature, the degree of temperature abnormality is determined; then, according to the characteristic that the CPU temperature usually decreases after the radiator speed increases, the degree of heat dissipation abnormality is determined; then, based on the degree of temperature abnormality and the degree of heat dissipation abnormality, in combination with the CPU utilization curve characterizing the stability of the host operating state, the risk of triggering an alarm is determined; then, according to the abnormality in the numerical value of the CPU temperature between the starting moment of CPU overheating and the warning moment, the degree of warning response requirement is determined; thus, the window adjustment coefficient for more accurately adjusting the debounce time window is adaptively determined by combining the risk of triggering an alarm and the degree of warning response requirement; finally, a more accurate dynamic debounce time window is adaptively determined according to the window adjustment coefficient, so that the accuracy of alarm detection according to the dynamic debounce time window is higher, solving the problem that the accuracy of alarm detection using a fixed debounce time window according to the CPU temperature in the prior art is relatively low, that is, making the effect of monitoring the BMC operating state of the microcomputer host according to the dynamic debounce time window better.

[0082] This application also provides a monitoring system for the BMC operating state of a microcomputer host. Please refer to Figure 2, which shows the structural diagram of a BMC operation status monitoring system for a microcomputer host provided by an embodiment of the present invention. The system includes: a data acquisition and preprocessing module 201, a first determination module 202, a second determination module 203, and an operation status monitoring module 204.

[0083] The data acquisition and preprocessing module 201 is configured to, during the operation of the microcomputer host, collect the prior warning time, CPU temperature curve, radiator rotation speed curve, and CPU utilization rate curve of each monitoring time period monitored by the BMC; obtain a preset debounce time window and the CPU over-temperature start time when the CPU temperature first exceeds the preset temperature threshold in each monitoring time period.

[0084] The first determination module 202 is configured to determine the degree of temperature abnormality according to the change correlation between the CPU temperature curve and the radiator rotation speed curve and the overall upward trend of the CPU temperature in the CPU temperature curve; determine the degree of heat dissipation abnormality according to the matching situation of the CPU temperature change during the rise of the radiator rotation speed; and determine the alarm triggering risk of each monitoring time period according to the overall stability of the CPU utilization rate curve, the degree of temperature abnormality, and the degree of heat dissipation abnormality.

[0085] The second determination module 203 is configured to, in each monitoring time period, determine the degree of alarm response requirement of each monitoring time period according to the time length between the prior warning time and the CPU over-temperature start time, the CPU temperature distribution, and the overall CPU temperature.

[0086] The operation status monitoring module 204 is configured to determine the window adjustment coefficient of each monitoring time period according to the degree of alarm response requirement and the alarm triggering risk; adjust the preset debounce time window according to the window adjustment coefficient to determine the dynamic debounce time window of each monitoring time period; and monitor the BMC operation status of the microcomputer host according to the dynamic debounce time window.

[0087] It should be noted that for the system provided in the above embodiment, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, a BMC operation status monitoring system for a microcomputer host provided in the above embodiment and an embodiment of a BMC operation status monitoring method for a microcomputer host belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0088] An embodiment of the present application also provides a computer device. Please refer to Figure 3, which shows a schematic structural diagram of a computer device provided by an embodiment of the present invention. The computer device includes a memory 301, a processor 302, and a computer program 303 stored in the memory 301 and running on the processor 302. Wherein, when the processor 302 executes the computer program 303, the computer device can execute any one of the BMC operation state monitoring methods of the microcomputer host described above.

[0089] An embodiment of the present application further provides a computer program product. When the computer program product runs on a computer device, the computer device can execute any one of the BMC operation state monitoring methods of the microcomputer host described above.

[0090] An embodiment of the present application further provides a computer-readable storage medium. Computer program code is stored in the computer-readable storage medium. When the computer program code runs on a computer device, the computer device can execute any one of the BMC operation state monitoring methods of the microcomputer host described above.

[0091] In the embodiments provided in the present application, it should be understood that the provided computer device, computer program product, and computer-readable storage medium are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the methods provided above, and will not be elaborated here.

[0092] It should be noted that: the above-mentioned sequence of embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0093] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

Claims

1. A method for monitoring the BMC operation status of a microcomputer host, characterized in that: The method comprises: During the operation of the microcomputer host, the prior alarm time, CPU temperature curve, radiator speed curve and CPU utilization curve of each monitoring time period monitored by the BMC are collected; the preset de-jitter time window and the CPU over-temperature start time when the CPU temperature exceeds the preset temperature threshold for the first time in each monitoring time period are obtained; Determine the degree of temperature anomaly based on the change correlation between the CPU temperature curve and the radiator speed curve and the overall rising trend of the CPU temperature in the CPU temperature curve; determine the degree of heat dissipation anomaly based on the matching of the CPU temperature change during the radiator speed increase process; determine the alarm risk of each monitoring time period based on the overall stability of the CPU utilization curve, the temperature anomaly degree and the heat dissipation anomaly degree; In each monitoring time period, the alarm response requirement level of each monitoring time period is determined according to the time length between the prior alarm moment and the CPU overtemperature start moment, the CPU temperature distribution and the overall CPU temperature; Determine the window adjustment coefficient of each monitoring time period according to the alarm response requirement degree and the alarm risk; adjust the preset de-jitter time window according to the window adjustment coefficient to determine the dynamic de-jitter time window of each monitoring time period; monitor the BMC operation status of the microcomputer host according to the dynamic de-jitter time window.

2. The method for monitoring the BMC operation status of a microcomputer host according to claim 1, characterized in that: The process of obtaining the temperature anomaly degree includes: On the CPU temperature curve, the slope value of the line between the data point corresponding to each sampling moment and the data point corresponding to the previous sampling moment is used as the temperature slope value at each sampling moment; and the temperature rising trend characteristic value is determined according to the average value of the temperature slope values ​​at all sampling moments; On the CPU temperature curve, negative correlation mapping is performed between the mean value of the CPU temperature at all sampling moments and the difference between the preset temperature threshold value to determine the proximity degree of the threshold temperature; Determine the degree of correlation between the temperature and the rotation speed according to a negative correlation mapping of the Pearson correlation coefficient between the CPU temperature curve and the heat sink rotation speed curve; The degree of temperature anomaly in each monitoring time period is determined according to the temperature-speed correlation degree, the threshold temperature proximity degree and the temperature rising trend characteristic value; the temperature-speed correlation degree is negatively correlated with the temperature anomaly degree; the threshold temperature proximity degree and the temperature rising trend characteristic value are both positively correlated with the temperature anomaly degree.

3. The method for monitoring the BMC operation status of a microcomputer host according to claim 2, characterized in that: The process of determining the degree of temperature anomaly in each monitoring time period according to the temperature-rotation speed correlation degree, the threshold temperature proximity degree and the temperature rising trend characteristic value includes: The product of the negative correlation mapping value of the temperature-rotation speed correlation degree, the threshold temperature proximity degree and the temperature rising trend characteristic value is normalized to determine the temperature anomaly degree of each monitoring time period.

4. The method for monitoring the BMC operation status of a microcomputer host according to claim 2, characterized in that: The process of obtaining the abnormal degree of heat dissipation includes: On the radiator speed curve, the slope value of the line between the data point corresponding to each sampling moment and the data point corresponding to the previous sampling moment is used as the speed slope value at each sampling moment; the sampling moment when the speed slope value is greater than 0 is used as the speed rising moment; The difference between the speed slope value and the corresponding temperature slope value at each speed rising moment is negatively correlated to determine the instantaneous heat dissipation fault degree at each speed rising moment; the heat dissipation abnormality degree of each monitoring time period is determined based on the average of the instantaneous heat dissipation fault degrees at all speed rising moments.

5. The method for monitoring the BMC operation status of a microcomputer host according to claim 1, characterized in that: The process of obtaining the alarm risk includes: The CPU utilization standard deviations at all sampling moments on the CPU utilization curve are negatively correlated to determine the stability of the CPU utilization in each monitoring time period; The product of the negative correlation mapping value of the CPU utilization stability, the temperature abnormality and the heat dissipation abnormality is normalized to determine the alarm risk of each monitoring time period.

6. The method for monitoring the BMC operation status of a microcomputer host according to claim 1, characterized in that: The process of obtaining the alarm response requirement degree includes: The monitoring time period before the prior alarm time when the CPU overtemperature starts in the time sequence is used as the analysis time period; the analysis time period contains the CPU overtemperature start time and the prior alarm time; and the alarm response requirement degree of other monitoring time periods outside the analysis time period is set to a preset requirement degree; In each analysis time period, all sampling moments between the CPU over-temperature start moment and the prior alarm moment are taken as delay moments; the delay moment when the corresponding CPU temperature is greater than the preset temperature threshold is taken as the reference abnormal moment; the difference between the CPU temperature at each reference abnormal moment and the preset temperature threshold is taken as the instantaneous over-temperature degree at each reference abnormal moment; the overall over-temperature degree of each analysis time period is determined based on the accumulated value of the instantaneous over-temperature degrees at all reference abnormal moments; The product between the number of delay moments and the overall over-temperature degree in each analysis time period is normalized to determine the alarm response requirement degree in each analysis time period.

7. The method for monitoring the BMC operation status of a microcomputer host according to claim 1, characterized in that: The process of obtaining the window adjustment coefficient includes: A negative correlation mapping is performed on the product of the alarm response requirement degree and the alarm risk to determine the window adjustment coefficient of each monitoring time period.

8. The method for monitoring the BMC operation status of a microcomputer host according to claim 1, characterized in that: The process of acquiring the dynamic de-jittering time window includes: The dynamic de-jitter time window of each monitoring time period is determined by multiplying the product of the window adjustment coefficient and the preset de-jitter time window.

9. The method for monitoring the BMC operation status of a microcomputer host according to claim 1, characterized in that: The process of monitoring the BMC operating status of the microcomputer host according to the dynamic de-jittering time window includes: The dynamic de-jitter time window of each monitoring time period is used as the final de-jitter time window of each subsequent monitoring time period until a new dynamic de-jitter time window is re-determined when an alarm occurs; the BMC operating status of the microcomputer host is monitored based on the final de-jitter time window of each monitoring time period combined with the corresponding CPU temperature.

10. A BMC operation status monitoring system for a microcomputer host, characterized in that: The system comprises: The data collection preprocessing module is used to collect the prior alarm time, CPU temperature curve, radiator speed curve and CPU utilization curve of each monitoring time period monitored by the BMC during the operation of the microcomputer host; obtain the preset de-jitter time window and the CPU over-temperature start time when the CPU temperature exceeds the preset temperature threshold for the first time in each monitoring time period; The first determination module is used to determine the temperature abnormality according to the change correlation between the CPU temperature curve and the radiator speed curve and the overall rising trend of the CPU temperature in the CPU temperature curve; determine the heat dissipation abnormality according to the matching of the CPU temperature change during the radiator speed increase process; determine the alarm risk of each monitoring time period according to the overall stability of the CPU utilization curve, the temperature abnormality and the heat dissipation abnormality; The second determination module is used to determine the alarm response requirement degree of each monitoring time period according to the time length between the priori alarm moment and the CPU overtemperature start moment, the CPU temperature distribution and the overall CPU temperature in each monitoring time period; An operation status monitoring module is used to determine the window adjustment coefficient of each monitoring time period according to the alarm response requirement degree and the alarm risk; adjust the preset de-jitter time window according to the window adjustment coefficient to determine the dynamic de-jitter time window of each monitoring time period; and monitor the BMC operation status of the microcomputer host according to the dynamic de-jitter time window.

Citation Information

Patent Citations

  • Machine room temperature detection control method and device based on BMC, equipment and medium

    CN111273753A

  • Method and equipment for improving heat dissipation efficiency of CPUs

    CN112612349A