Computer host heat dissipation monitoring and alarming system

By collecting multi-source time-series parameters to construct a system collaborative performance model and dynamically fitting the expected temperature rise curve, the problem of the inability to identify collaborative failures of the heat dissipation system in existing technologies is solved. This enables accurate monitoring and early warning of the computer host heat dissipation system, improving hardware stability and operation and maintenance efficiency.

CN121901060APending Publication Date: 2026-04-21SHENZHEN JIUZHOU YUNHAI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing computer host heat dissipation monitoring solutions only focus on the speed and temperature threshold of a single fan, and cannot identify the failure of the heat dissipation system as a whole. This results in hidden faults not being detected in time, which may lead to hardware performance damage or lifespan reduction.

Method used

By collecting multi-source time-series parameters, a system collaborative performance model is constructed, the expected temperature rise curve is dynamically fitted, and the risk of collaborative failure is determined when the actual temperature rise deviates from the expectation, and graded alarms are issued to remind operation and maintenance personnel.

Benefits of technology

It enables accurate assessment and early warning of the collaborative performance of the heat dissipation system, reduces hardware damage caused by collaborative failure, and lowers operation and maintenance costs and time waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901060A_ABST
    Figure CN121901060A_ABST
Patent Text Reader

Abstract

The invention discloses a computer host heat dissipation monitoring and alarming system, and relates to the technical field of computer early warning, and the system comprises a parameter collection module which is used for collecting multi-source time sequence parameters of a computer host heat dissipation system, and the multi-source time sequence parameters at least comprise the temperature of heating hardware, the rotating speed of a heat dissipation fan and the air speed of a key monitoring point in an air duct; the collaborative efficiency analysis module is used for constructing a system collaborative efficiency model based on the multi-source time sequence parameters, and construction and application of the system collaborative efficiency model comprise the steps that the pneumatic conversion efficiency of a current system is calculated according to the corresponding relation between the rotating speed of a cooling fan and the wind speed; according to the method, the collaborative efficiency model is optimized through continuously accumulated time series data, so that the expected temperature rise curve fitting is more accurate, and the risk judgment accuracy is gradually improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer early warning technology, specifically a computer host heat dissipation monitoring and alarm system. Background Technology

[0002] The stable operation of computer mainframes, especially high-performance computing platforms and servers, heavily depends on the efficient operation of their internal cooling systems.

[0003] Currently, monitoring solutions in this field mainly rely on real-time acquisition of the temperature of core heat-generating components (such as CPUs and GPUs) and setting fixed temperature thresholds for alarms and control. When the hardware temperature reaches or exceeds the preset threshold, the system increases the speed of the corresponding cooling fan to enhance heat dissipation. Once the temperature drops back to a safe range, the fan speed is reduced to save energy.

[0004] However, in actual applications, the heat dissipation performance of a computer host is not determined independently by the heat-generating hardware and the cooling fan, but by a collaborative whole consisting of multiple components such as the fan, air duct, heat sink fins, and ventilation holes.

[0005] Existing technologies only focus on the "causal" relationship between terminal temperature and a single fan speed, which has significant limitations: First, it cannot identify hidden faults where "component parameters are normal, but system coordination fails." For example, when the chassis airflow is blocked by dust accumulation, or when the fan and airflow design are mismatched, even if the cooling fan runs at full speed, the cooling airflow it generates cannot effectively deliver cooling to the heat source, resulting in a sharp drop in overall heat dissipation efficiency. At this time, the sensor data of each component may still be within the "normal" range, but the system is already in a state of coordination imbalance. Existing monitoring solutions often only respond when the hardware temperature eventually spikes and triggers a fixed threshold alarm. By then, the hardware may have already suffered irreversible performance damage or lifespan reduction due to prolonged overheating.

[0006] Therefore, there is an urgent need in this field for an innovative solution that can perform online evaluation and early warning of heat dissipation synergy performance from a holistic system perspective. Summary of the Invention

[0007] The purpose of this invention is to provide a computer host heat dissipation monitoring and alarm system to solve the problems mentioned in the background art.

[0008] A computer host heat dissipation monitoring and alarm system, comprising: The parameter acquisition module is used to acquire multi-source timing parameters of the computer host heat dissipation system. The multi-source timing parameters include at least the temperature of the heat-generating hardware, the speed of the cooling fan, and the wind speed at key monitoring points in the air duct. The collaborative performance analysis module is used to construct a system collaborative performance model based on multi-source time-series parameters. The construction and application of the system collaborative performance model includes: Calculate the aerodynamic conversion efficiency of the current system based on the relationship between cooling fan speed and airflow. Based on the time-series data of aerodynamic conversion efficiency and the temperature of the heating hardware, the expected temperature rise curve of the heating hardware is dynamically fitted through a system collaborative efficiency model. The risk assessment module is used to continuously compare the real-time collected hardware temperature data with the expected temperature rise curve, and to determine that the system has a risk of collaborative failure when the actual temperature rise continuously deviates from the expected temperature rise curve and exceeds the dynamic tolerance. The alarm execution module is used to generate a corresponding alarm signal when the risk assessment module determines that there is a risk of collaborative failure.

[0009] Preferably, the key monitoring points in the air duct include at least the air duct inlet, the windward side of the heat-generating hardware, and the leeward side; The multi-source timing parameters also include power consumption data of the heat-generating hardware and ambient temperature data of the host.

[0010] Preferably, when the collaborative performance analysis module calculates the aerodynamic conversion efficiency of the current system: Within a preset sampling window, the cooling fan speed sequence and the corresponding wind speed sequence are collected simultaneously. The collected rotational speed and wind speed sequences were screened for data validity, and abnormal data points caused by transient interference from the sensors were removed. The selected valid data points are divided into multiple intervals according to the rotational speed value, and the statistical characteristic value of the wind speed to rotational speed ratio in each interval is calculated. Based on the temperature change trend of the aforementioned heat-generating hardware, determine whether the system is in a thermal steady state: When the system is in thermal steady state, the mode of the statistical characteristic values ​​within the speed range is selected as the benchmark value of the current aerodynamic conversion efficiency. When the system is in a non-thermal steady state, the weighted average of the statistical characteristic values ​​within the speed range is selected as the benchmark value of the current aerodynamic conversion efficiency, wherein the weighting coefficient is positively correlated with the temperature stability of the corresponding data point; The calculated aerodynamic conversion efficiency benchmark value is compared with the preset efficiency threshold range. If the calculation results exceed the threshold range multiple times in a row, a sensor calibration prompt is triggered and a historical normal value is used as a substitute.

[0011] Preferably, when the collaborative performance analysis module dynamically fits the expected temperature rise curve of the heat-generating hardware: Establish a multivariate regression model based on the aerodynamic conversion efficiency and the temperature of the heating hardware; Before each fitting, the consistency between the changing trend of the aerodynamic conversion efficiency and the changing trend of the power consumption of the heat-generating hardware is analyzed. Evaluate the correspondence between wind speed distribution patterns and aerodynamic conversion efficiency values; When the analysis results are inconsistent, execute the parameter re-acquisition process; A staged fitting method is adopted, and in each stage of the fitting process: Calculate the residual sequence between the current fitted curve and the actual temperature data; Based on the distribution characteristics of the residual sequence, the input parameters of the regression model are screened and optimized by combination. The fitting process is complete when the rate of change of the sum of squared residuals in multiple consecutive fitting stages is less than a set threshold. During model operation, the impact of changes in ambient temperature on the expected temperature rise curve is continuously monitored. When the ambient temperature changes to a predetermined range, the phased fitting process is restarted. The final fitted result is output as the valid expected temperature rise curve.

[0012] Preferably, the dynamic tolerance is dynamically adjusted based on the aerodynamic conversion efficiency; When the aerodynamic conversion efficiency is lower than the first threshold, the dynamic tolerance is narrowed proportionally. When the aerodynamic conversion efficiency is higher than the second threshold, the dynamic tolerance is relaxed proportionally.

[0013] Preferably, it also includes a risk classification module, used when it is determined that there is a risk of collaborative failure in the system: A risk assessment matrix based on a combination of multi-dimensional parameters is constructed. The rows of the matrix represent the degree to which the aerodynamic conversion efficiency deviates from the normal range, and the columns represent the range to which the wind speed distribution uniformity deviates from the benchmark. Analyze the trajectory of the actual temperature rise deviating from the expected temperature rise curve to identify its degree of conformity with typical failure characteristics; Establish a parameter linkage analysis to check the time correspondence between the change in cooling fan speed and the power consumption fluctuation of heat-generating hardware, as well as the spatial distribution relationship between airflow velocity distribution and temperature gradient distribution. Based on the location results, change trajectory characteristics, and parameter linkage analysis conclusions of the risk assessment matrix, the level of collaborative failure risk is determined through a hierarchical decision-making model. The hierarchical decision-making model focuses on examining the situation of multi-parameter collaborative anomalies. When multiple parameters are abnormal at the same time and there are spatiotemporal correlation characteristics, the risk level judgment result is improved. Based on the final determined risk level, the corresponding tiered control strategy will be triggered.

[0014] Preferably, it also includes a fault diagnosis module, used to perform a root cause identification process when it is determined that there is a risk of collaborative failure in the system: A three-dimensional feature space is established based on the numerical range of aerodynamic conversion efficiency, the wind speed distribution uniformity coefficient, and the temperature distribution gradient vector, in which the pre-stored typical fault modes have corresponding regions in the feature space; Calculate the position of the current system state in the feature space and the distance between it and the regions of each typical fault mode; Analyze the time-series variations of aerodynamic conversion efficiency, wind speed distribution, and temperature gradient; Perform influence path analysis between execution parameters to determine the propagation direction of abnormal parameters; Based on the results of the distance, time series variation characteristics, and parameter influence path analysis, a candidate fault set is selected from a predefined root cause library; Sort the candidate fault set by distance from smallest to largest, and check in turn whether each fault mode can explain all the abnormal parameters; The first fault mode that can explain all the abnormal parameters is taken as the root cause of the fault; If no fault mode exists that can explain all the abnormal parameters, then the first two fault modes are selected as the common root cause of the fault. The root cause identification information of the fault is embedded in the alarm signal and output.

[0015] Preferably, the collaborative performance analysis module is further used for: Before dynamically fitting the expected temperature rise curve, monitor the real-time power consumption of the heat-generating hardware and predict its power consumption trend within a future time window. The power consumption change trend is introduced as a priori condition during the dynamic fitting process so that the expected temperature rise curve can reflect the impact of upcoming load changes on temperature.

[0016] Preferably, the dynamic tolerance is adaptively adjusted based on the historical performance of the system's collaborative efficiency model: The accuracy of the predicted temperature rise curve over historical periods in predicting actual temperature rise is statistically analyzed. When the prediction accuracy is high, the dynamic tolerance is narrowed to improve the monitoring sensitivity. When the prediction accuracy is low, the dynamic tolerance is relaxed to reduce the false alarm rate.

[0017] Preferably, when the alarm execution module generates an alarm signal, it performs at least one of the following: Generate graphical or textual warning messages for display in the user interface; Generate the electrical signal that drives the audible and visual alarm. Generate a log file to record the system status.

[0018] Compared with the prior art, the beneficial effects of the present invention are: This invention comprehensively captures multi-dimensional data of the heat dissipation system through a parameter acquisition module, including core parameters such as wind speed at multiple key monitoring points in the air duct, hardware power consumption, and internal ambient temperature. A collaborative performance analysis module constructs a model to fit the expected temperature rise curve, followed by a risk assessment module to accurately identify collaborative failure risks. Finally, a tiered alarm execution module provides early warnings. This solves the problem of traditional single-parameter monitoring's inability to predict collaborative failures in advance, allowing users or maintenance personnel to promptly address potential issues such as air duct blockage, fan failure, and abnormal hardware power consumption. Furthermore, the tiered alarm mechanism adapts to different risk scenarios, reducing interference from invalid alarms. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the system framework structure of the present invention.

[0020] Figure 2 This is a schematic diagram of the core logic flow for heat dissipation monitoring. Detailed Implementation

[0021] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figures 1-2 This application provides a computer host heat dissipation monitoring and alarm system, comprising: The parameter acquisition module is used to collect multi-source timing parameters of the computer host heat dissipation system. The multi-source timing parameters include at least the temperature of the heat-generating hardware, the speed of the cooling fan, the wind speed at key monitoring points in the air duct, the power consumption data of the heat-generating hardware, and the ambient temperature data inside the host. In this embodiment, multi-source timing parameters refer to key data that are continuously collected at fixed time intervals and reflect different dimensions of the heat dissipation system. The collection interval is set to 1 second to ensure real-time capture of state changes.

[0023] The temperature of heat-generating hardware mainly targets core heat-generating components such as the CPU, GPU, and motherboard chipset. It is collected through the built-in digital temperature sensor of the hardware, and the data accuracy can reach ±0.5℃. The cooling fan speed is obtained through the fan's built-in speed signal pin, and the unit is revolutions per minute. Among them, the key monitoring points in the air duct include at least the air inlet of the air duct, the windward side and the leeward side of the heat-generating hardware. The wind speed is collected by a miniature hot wire anemometer in meters per second. The air inlet data reflects the air intake capacity of the fan, and the windward and leeward data can directly reflect the heat exchange effect of the hardware heat sink. The power consumption data of heat-generating hardware is collected through the motherboard power management chip or independent power sensor. The power consumption data of core hardware such as CPU and GPU can be accurate to ±1W. The internal ambient temperature of the host is collected by a temperature sensor placed in a heat-free area in the middle of the chassis to eliminate the interference of ambient temperature on the hardware temperature rise.

[0024] Synchronous acquisition of multi-source data can avoid the one-sidedness of monitoring a single parameter. For example, if the temperature is normal, the potential risks of a sudden drop in wind speed or a sudden increase in power consumption may be ignored. This invention can build a complete operational profile of the heat dissipation system by synchronously acquiring multi-dimensional data. This allows maintenance personnel to not only know whether the current status is normal, but also to trace the root cause of the anomaly, avoiding the waste of time caused by blind troubleshooting. At the same time, it provides a comprehensive and reliable data foundation for subsequent collaborative performance analysis, ensuring the authenticity and effectiveness of the analysis results.

[0025] The collaborative performance analysis module is used to construct a system collaborative performance model based on multi-source time-series parameters. The construction and application of the system collaborative performance model includes the following steps: The first step is to calculate the aerodynamic efficiency of the current system based on the relationship between the cooling fan speed and the airflow. The specific process is as follows: Within a preset sampling window, the cooling fan speed sequence and the corresponding wind speed sequence are collected synchronously. In this embodiment, the sampling window is set to 10 seconds, that is, 10 sets of speed and wind speed data are collected once per second. The synchronous collection is achieved through the clock synchronization mechanism of the parameter acquisition module to ensure that the timestamps of each set of data are completely matched.

[0026] The collected rotational speed and wind speed sequences are screened for data validity, and abnormal data points caused by transient interference from the sensors are removed. Specifically, the screening rule adopts the 3σ criterion, which calculates the mean and standard deviation of the sequence data and removes values ​​that exceed the mean ± 3 times the standard deviation. For example, if the mean of a rotational speed sequence is 2000 revolutions per minute and the standard deviation is 50 revolutions per minute, then values ​​below 1850 revolutions per minute or above 2150 revolutions per minute are judged as abnormal and removed.

[0027] After screening, the valid data points are divided into multiple intervals according to the rotational speed value, with the interval interval set to 200 revolutions per minute, such as 1800-2000 revolutions per minute, 2000-2200 revolutions per minute, etc., and the statistical characteristic values ​​of the wind speed to rotational speed ratio in each interval are calculated, including the mean, median, mode and standard deviation.

[0028] Determine whether the system is in a thermal steady state based on the temperature change trend of the heating hardware. Specifically, the criteria for determining thermal steady state are that the hardware temperature change is ≤0.5℃ for five consecutive acquisition cycles, otherwise it is considered non-thermal steady state.

[0029] When the system is in thermal steady state, the mode of the statistical characteristic values ​​within the current speed range is selected as the benchmark value of aerodynamic conversion efficiency. Specifically, the mode can reflect the efficiency level that most often occurs under steady state and has higher stability. When the system is in a non-thermal steady state, the weighted average of the statistical characteristic values ​​within the speed range is selected as the benchmark value. The weight coefficient is positively correlated with the temperature stability of the corresponding data point. Specifically, the temperature stability is calculated by the temperature fluctuation amplitude 3 seconds before and after the data point acquisition time. The smaller the fluctuation amplitude, the higher the weight. For example, if the temperature fluctuation of a certain data point is 0.2℃, the weight is set to 0.8, and if the fluctuation of another data point is 0.4℃, the weight is set to 0.4.

[0030] The calculated aerodynamic conversion efficiency baseline value is compared with the preset efficiency threshold range, which is calibrated to 60%-95% according to the fan model. If the calculation results exceed this range multiple times (e.g., 3 times), a sensor calibration prompt is triggered and historical normal values ​​are used for replacement (e.g., the average normal efficiency of the same speed range in the past 7 days is used for replacement) to avoid subsequent analysis deviations caused by sensor failure.

[0031] The innovation of this invention lies in the fact that the calculation logic can adapt to thermal change scenarios and is more in line with the actual operating state than fixed ratio calculation. For example, when the host is in a non-thermal steady state after startup, the influence of temperature fluctuation data is weakened by weighted averaging to ensure accurate efficiency calculation; when running stably, the mode is used to improve the reliability of the results.

[0032] This dynamic calculation method, which combines system thermal state, breaks through the limitations of traditional fixed efficiency calculation that cannot adapt to dynamic scenarios such as host start-up and shutdown, load fluctuations, etc. It can reflect the actual coordination capability of fans and air ducts in real time. Even in scenarios where the host operating status changes frequently, it can provide accurate basic parameters for subsequent temperature rise curve fitting and avoid risk misjudgment caused by efficiency calculation deviation.

[0033] The second step in building the collaborative performance model is to dynamically fit the expected temperature rise curve of the heating hardware based on the time-series data of aerodynamic conversion efficiency and the temperature of the heating hardware, combined with the power consumption data of the heating hardware and the internal ambient temperature data of the host, through the system collaborative performance model. The specific process is as follows: First, a multivariate regression model based on aerodynamic conversion efficiency and the temperature of the heating hardware is established. Specifically, the input parameters of this model include aerodynamic conversion efficiency. Current hardware temperature In addition, the power consumption of heat-generating hardware is also included. Internal ambient temperature of the host The core formula is: ,in This represents the expected temperature over the next 5 minutes. These are the weighting coefficients for each variable. This is a constant term.

[0034] Before each fitting, it is necessary to analyze the degree of consistency between the changing trend of aerodynamic conversion efficiency and the changing trend of power consumption of heat-generating hardware. In this embodiment, Pearson correlation coefficient is used for evaluation. A correlation coefficient ≥ 0.6 is considered highly consistent, 0.3~0.6 is moderately consistent, and < 0.3 is inconsistent. For example, if the aerodynamic conversion efficiency increases synchronously or remains stable at a high level when the power consumption increases, it indicates that the heat dissipation system can adapt to the increase in load and the trend is consistent. If the power consumption increases while the aerodynamic conversion efficiency drops sharply, it is considered inconsistent.

[0035] Simultaneously, the correspondence between wind speed distribution patterns and aerodynamic conversion efficiency values ​​is evaluated. Specifically, the wind speed distribution pattern is characterized by the standard deviation of wind speed at the air inlet of the duct, the upwind side, and the downwind side. A standard deviation ≤ 0.2 m / s indicates a uniform distribution, in which case the aerodynamic conversion efficiency should be ≥ 80%. A standard deviation > 0.5 m / s indicates a non-uniform distribution, and the corresponding aerodynamic conversion efficiency is usually ≤ 70%. If it exceeds this corresponding range, it needs to be marked as abnormal.

[0036] When the trend consistency analysis results are inconsistent, or the wind speed distribution does not match the efficiency value, the parameter re-acquisition process is executed. Specifically, the sampling interval is shortened to 0.5 seconds, data is continuously collected for 5 seconds and re-filtered to eliminate anomalies caused by sensor failure or instantaneous load fluctuations.

[0037] The fitting process uses a phased approach, with each phase lasting 1 minute. During each phase of the fitting process: First, calculate the residual sequence between the current fitted curve and the actual temperature data. Specifically, the residual is the difference between the actual temperature and the fitted temperature. Based on the distribution characteristics of the residual sequence, the input parameters of the regression model are screened and combined for optimization. Specifically, if the residual sequence is normally distributed and the mean is close to 0, it indicates that the current parameter combination is reasonable. If the residual is skewed, the weight coefficients are adjusted (such as increasing the weight of the power consumption parameter).

[0038] When the rate of change of the sum of squared residuals in multiple consecutive fitting stages (e.g., 3) is less than a preset threshold (e.g., 5%), the fitting is considered to have converged, and the fitting process is completed.

[0039] During model operation, the impact of ambient temperature changes on the expected temperature rise curve is continuously tracked. When the ambient temperature changes to a predetermined range, for example, when the ambient temperature changes by more than 5°C within 5 minutes, the phased fitting process is restarted to ensure that the curve adapts to environmental fluctuations.

[0040] It should be noted that the initial calibration of model parameters α, β, γ, δ and ε was performed using the gradient descent method. The training set consisted of 1000 sets of historical data covering low, medium and high loads and ambient temperatures of 20℃-40℃. The learning rate was 0.01 and the iterations were 5000 times to bring the mean square error to <0.8℃².

[0041] It should be noted that this fitting logic can accurately adapt to dynamic changes in the system. For example, if a sudden increase in host load leads to an increase in power consumption, and the aerodynamic conversion efficiency increases synchronously and in a consistent trend, the fitted curve will rise smoothly. If the trends are inconsistent, data will be re-collected to avoid fitting bias and provide a reliable benchmark for risk assessment. Compared with traditional fixed curve fitting, this dynamic fitting method is better able to adapt to heat dissipation fluctuations caused by changes in host load, ambient temperature, and other factors. It can predict potential temperature rise risks in advance, giving maintenance personnel sufficient time to intervene and avoid temporary hardware throttling or unexpected shutdowns due to excessively rapid temperature rise, thus ensuring the continuity and stability of host operation.

[0042] The risk assessment module is used to continuously compare the real-time collected hardware temperature data with the expected temperature rise curve, and to determine that the system has a risk of collaborative failure when the actual temperature rise continuously deviates from the expected temperature rise curve and exceeds the dynamic tolerance.

[0043] In this embodiment, the dynamic tolerance refers to the allowable temperature deviation range that is dynamically adjusted according to the hardware type and operating load. The dynamic tolerance of the CPU is set to ±5℃ under high load and ±3℃ under low load; the dynamic tolerance of the GPU is uniformly set to ±4℃ because the heat is more concentrated.

[0044] The criterion for continuous deviation is that the temperature deviation between the actual temperature and the temperature at the corresponding moment of the expected temperature rise curve exceeds the dynamic tolerance within three consecutive acquisition cycles (i.e., 3 seconds).

[0045] For example, if the expected temperature rise curve shows that the CPU temperature will reach 75℃ in 10 seconds, with a dynamic tolerance of ±5℃, but the actual temperature reaches 81℃ after 10 seconds, and the deviation is more than 6℃ for three consecutive seconds, then a risk of coordinated failure is identified. This judgment logic avoids false alarms caused by instantaneous temperature fluctuations. Through continuous comparison and dynamic tolerance settings, it not only eliminates the drawback of traditional single-threshold judgments being easily affected by instantaneous interference and generating false alarms, but also accurately matches the judgment criteria according to hardware type and load status, ensuring the accuracy of risk identification. This advantage is particularly evident in high-load operating scenarios, effectively distinguishing between normal load temperature rise and abnormal coordinated failure temperature rise, reducing the interference of invalid alarms on users, while not overlooking real risks.

[0046] The alarm execution module is used to generate a corresponding alarm signal when the risk assessment module determines that there is a risk of collaborative failure. When generating an alarm signal, at least one of the following operations is performed: generating a graphical or text warning message for display on the user interface, generating an electrical signal to drive the audible and visual alarm, or generating a log file to record the system status. The specific hierarchical execution logic is as follows: When the risk is mild, i.e. the dynamic tolerance has just been exceeded and the duration is short, a graphic and text warning message is generated. A text pop-up window with a red border appears on the operating system desktop, clearly indicating "the heat dissipation efficiency has decreased, and it is recommended to check the cleanliness of the air duct". At the same time, a yellow warning icon and key abnormal parameters such as the current aerodynamic conversion efficiency and wind speed are displayed on the host monitoring software interface. Log files are generated synchronously to record information such as alarm time, risk level, and real-time data of each monitoring point. The log storage path is set to the system default log directory, and the naming format is "heat dissipation alarm_year_month_day_hour_minute_second.log".

[0047] When the risk level is moderate, meaning the deviation exceeds the dynamic tolerance by 50%, an additional electrical signal is generated to drive the audible and visual alarm on top of the risk level operation. This controls the built-in buzzer of the host to sound intermittently at 1-second intervals, while simultaneously triggering the yellow indicator light on the front panel of the host to flash synchronously. The intensity of the audible and visual signals meets the noise and brightness standards of the office environment, avoiding excessive interference. The log file now includes a "moderate risk" label and a brief conclusion on the trend analysis.

[0048] In cases of severe risk, where the deviation exceeds the dynamic tolerance by 100% or the hardware temperature approaches the damage threshold, the graphical and textual warning information is enhanced based on the operation of medium risk. The pop-up window becomes a full-screen top display, the text font is bolded and marked "Emergency: Cooling system failure". At the same time, the abnormal hardware icon is highlighted with a red flashing animation in the monitoring software interface. The system switches the audible and visual alarm to a mode with a continuous buzzer and a constantly lit red indicator light. In addition to basic information, the log file includes parameter change curves for the past minute to facilitate subsequent fault tracing. At the same time, it sends SMS alarms to the pre-bound administrator's mobile phone and automatically activates the hardware frequency reduction strategy to reduce the rate of heat generation.

[0049] Implementing different alarm actions in a tiered manner can avoid excessive alarm for users due to minor risks, while ensuring loss prevention and traceability through multiple warning and recording measures in the event of serious risks, thereby improving the system's usability.

[0050] This tiered alarm mechanism fully considers user needs at different risk levels. Minor risks are addressed solely through interface prompts and log recordings, without affecting normal host operation. Severe risks, on the other hand, utilize multiple audible and visual warnings and proactive frequency reduction strategies to minimize the risk of hardware damage. Simultaneously, detailed log recordings provide comprehensive data support for subsequent troubleshooting, helping maintenance personnel quickly pinpoint the root cause of problems, shorten troubleshooting time, and improve operational efficiency.

[0051] This invention comprehensively captures multi-dimensional data of the heat dissipation system through a parameter acquisition module, including core parameters such as wind speed at multiple key monitoring points in the air duct, hardware power consumption, and internal ambient temperature. The collaborative performance analysis module constructs a model to fit the expected temperature rise curve, the risk assessment module accurately identifies collaborative failure risks, and finally the alarm execution module provides tiered early warnings.

[0052] In this way, not only can the problem of traditional single-parameter monitoring failing to predict collaborative failures in advance be solved, allowing users or maintenance personnel to deal with potential problems such as air duct blockage, fan failure, and abnormal hardware power consumption in a timely manner, but the hierarchical alarm mechanism can also be adapted to different risk scenarios, reducing interference from invalid alarms. Furthermore, by continuously accumulating time-series data, the collaborative efficiency model can be optimized, making the expected temperature rise curve fit more accurately and gradually improving the accuracy of risk assessment.

[0053] As one embodiment of the present invention, the dynamic tolerance is dynamically adjusted based on the aerodynamic conversion efficiency; When the aerodynamic conversion efficiency falls below a first threshold, the dynamic tolerance is narrowed proportionally. In this embodiment, the first threshold is a critical value for aerodynamic conversion efficiency set according to the heat dissipation characteristics of different hardware. For example, the first threshold for CPU is set to 65%, and the first threshold for GPU is set to 60%. This is because the CPU, as the core computing component, has higher requirements for heat dissipation coordination performance, and the sensitivity of risk assessment needs to be increased when the conversion efficiency is slightly low. When the aerodynamic conversion efficiency falls below this threshold, it means that the coordination between the fan and the air duct decreases, and the heat dissipation system cannot efficiently remove heat. At this time, the dynamic tolerance is narrowed by a preset proportion. For example, the normal dynamic tolerance for CPU is ±5℃, which is narrowed to ±2.5℃, thereby enabling earlier detection of abnormal temperature rise trends and avoiding hardware risks caused by heat dissipation failure.

[0054] When the aerodynamic conversion efficiency exceeds the second threshold, the dynamic tolerance is proportionally relaxed. In this embodiment, the second threshold is also set for different hardware: the second threshold for the CPU is set to 85%, and the second threshold for the GPU is set to 80%. When the conversion efficiency exceeds this threshold, it indicates that the heat dissipation system is in good coordination, the fan and airflow are well matched, and the heat dissipation capacity has sufficient redundancy. At this time, the dynamic tolerance is relaxed by a preset ratio. For example, the normal dynamic tolerance for the GPU is ±4℃, which is relaxed to ±6℃. This can effectively avoid false alarms caused by temporary temperature deviations due to instantaneous load fluctuations, reducing unnecessary interference to users.

[0055] This invention allows the dynamic tolerance to be adjusted in real time according to the aerodynamic conversion efficiency, so that the risk judgment criteria can be matched with the actual performance of the heat dissipation system. This improves the timeliness of early warning when the heat dissipation capacity is insufficient and ensures the stability of use when the heat dissipation is good. This avoids the risk of missed judgment and reduces invalid alarms, which helps to improve the practicality and accuracy of the entire heat dissipation monitoring system.

[0056] As one embodiment of the present invention, a risk classification module is also included, used to determine when the system has a risk of collaborative failure: A risk assessment matrix based on multi-dimensional parameter combinations is constructed. The rows of the matrix represent the degree to which the aerodynamic conversion efficiency deviates from the normal range, and the columns represent the range of wind speed distribution uniformity deviating from the benchmark. In this embodiment, the normal range of aerodynamic conversion efficiency is calibrated as 70%-90% according to the hardware model, and the degree of deviation is divided into three levels as "slight deviation (within ±10%)", "moderate deviation (10%-20%)", and "severe deviation (more than 20%)" as the matrix rows. The wind speed distribution uniformity benchmark is a standard deviation of ≤0.2 m / s for wind speeds at the air inlet, upwind side, and downwind side of the duct. Deviations are categorized into three levels as matrix columns: "basically uniform (standard deviation 0.2-0.5 m / s)," "somewhat non-uniform (standard deviation 0.5-0.8 m / s)," and "extremely non-uniform (standard deviation exceeding 0.8 m / s)." This matrix allows for rapid identification of the degree of deviation of key abnormal parameters, providing a fundamental basis for risk classification and avoiding biased judgments caused by a single abnormal parameter.

[0057] The actual temperature rise deviates from the expected temperature rise curve, and the degree of conformity with typical failure characteristics is identified. In this embodiment, typical failure characteristics are pre-entered into the system, including at least the feature library such as "sudden drop in aerodynamic conversion efficiency accompanied by accelerated temperature rise", "uneven wind speed distribution leading to local sudden increase in temperature rise" and "temperature rise fluctuation caused by asynchronous fan speed and power consumption".

[0058] Furthermore, by comparing the overlap and rate of change of the actual temperature rise trajectory with the trajectory in the feature library, the failure type and severity can be preliminarily determined. For example, if the actual temperature rise shows a "step-like sudden increase" and the overlap with the "typical characteristics of air duct blockage" is more than 80%, it is marked as a highly correlated feature.

[0059] This invention supplements risk assessment with a dynamic perspective through trajectory feature analysis, improving the accuracy of predicting potential failures. Compared to traditional static assessments that only focus on current temperature values, trajectory analysis can capture the development trend and evolution of risks, enabling a shift from "post-event alarm" to "pre-event prediction," and allowing maintenance personnel more time to handle faults.

[0060] Establish parameter linkage analysis to check the time correspondence between the change of cooling fan speed and the power consumption fluctuation of heat-generating hardware, as well as the spatial distribution relationship between airflow speed distribution and temperature gradient distribution. In this embodiment, the time correspondence is determined by analyzing the time difference between the change of fan speed and the power consumption fluctuation. Under normal conditions, the fan speed should increase synchronously within 1-2 seconds after the power consumption increases. If the time difference exceeds 5 seconds, it is determined to be an abnormal time correlation. Spatial distribution relationships are determined by comparing the overlap between high wind speed areas and low temperature gradient areas in the wind tunnel. Under normal conditions, the overlap should be ≥70%, and if it is less than 50%, it is considered an abnormal spatial correlation.

[0061] The innovation of this invention lies in its ability to uncover implicit correlations between parameters through parameter linkage analysis. This avoids viewing single parameter anomalies in isolation. For example, anomalies in only fan speed or power consumption may pose a low risk, but if there is a temporal correlation between the two, it could indicate a fan response failure, significantly increasing the risk. This spatiotemporal linkage analysis breaks down the isolation barriers between parameters, enabling the identification of implicit failure risks where "a single parameter is normal but the correlation is abnormal." Such risks are often overlooked in traditional single-parameter monitoring, greatly improving the completeness of risk identification.

[0062] Based on the location results of the risk assessment matrix, the characteristics of the change trajectory, and the conclusions of the linkage analysis between parameters, the level of collaborative failure risk is determined through a hierarchical decision-making model. In this embodiment, the risk level is divided into three levels: "mild risk", "medium risk" and "severe risk". The graded decision-making model adopts a multi-factor weighted scoring mechanism, in which the risk assessment matrix positioning result accounts for 40% of the weight, the consistency of the change trajectory characteristics accounts for 30% of the weight, and the parameter linkage analysis conclusion accounts for 30%.

[0063] Among them, the hierarchical decision-making model focuses on the situation of multi-parameter collaborative anomalies. When multiple parameters are abnormal at the same time and there are spatiotemporal correlation characteristics, the risk level judgment result is improved. It should be understood that this model focuses on examining situations where multiple parameters exhibit synergistic anomalies. When multiple parameters simultaneously show anomalies and exhibit spatiotemporal correlation characteristics, the risk level assessment is upgraded. For example, if the risk assessment matrix is ​​positioned as "moderate deviation" and the trajectory conformity is "moderate correlation," and if spatiotemporal correlation anomalies also exist, the risk level is upgraded from "moderate risk" to "severe risk." This multi-factor fusion decision-making approach overcomes the shortcomings of traditional single-factor decision-making in identifying multi-parameter synergistic failures, ensuring that the risk level assessment more closely reflects actual failure scenarios. By introducing weight allocation and correlation-based upgrade mechanisms, it avoids judgment biases caused by a single factor dominating the assessment and accurately identifies high-risk scenarios with multiple factors overlapping, making the risk level assessment more scientific and reliable.

[0064] Based on the final determined risk level, the corresponding graded control strategy is triggered. In this embodiment, when the risk is mild, an early warning prompt and continuous parameter monitoring strategy are triggered, which only prompts the operation and maintenance personnel through a pop-up window on the monitoring interface, without interfering with the operation of the host. When the risk level is moderate, the active adjustment and early warning enhancement strategies are triggered, controlling the fan speed to increase by 10%-20% and simultaneously activating the audible and visual alarms. In cases of severe risk, an emergency stop-loss and multiple alarm strategies are triggered, immediately reducing hardware frequency by 30%, shutting down non-core processes, sending SMS alarms to the administrator, and recording complete fault data.

[0065] This tiered control strategy achieves a precise match between risk level and intervention intensity, avoiding resource waste and performance loss caused by excessive intervention under mild risks, while also being able to quickly curb the spread of risks through strong measures under severe risks, thus maximizing the balance between system stability and operational efficiency.

[0066] This invention, through the addition of a risk grading module, enables precise grading and differentiated control of collaborative failure risks. Compared to traditional simple grading based solely on temperature thresholds, this module significantly improves the accuracy of risk level determination through multi-dimensional matrix positioning, trajectory feature recognition, and parameter linkage analysis, combined with a decision model focusing on multi-parameter collaborative anomalies. Simultaneously, it triggers differentiated control strategies based on grading results, avoiding excessive intervention in mild risks that could impact host operating efficiency, while quickly mitigating losses in cases of severe risks, significantly enhancing system reliability and usability. More importantly, the module's design is adaptable to host hardware of different brands and configurations; rapid migration can be achieved by adjusting matrix thresholds and feature library parameters, demonstrating strong versatility and scalability, thus addressing the industry pain point of poor adaptability in traditional thermal monitoring systems.

[0067] As one embodiment of the present invention, it also includes a fault diagnosis module, used to perform a fault root cause identification process when it is determined that there is a risk of collaborative failure in the system: A three-dimensional feature space is established based on the aerodynamic conversion efficiency numerical range, the wind speed distribution uniformity coefficient, and the temperature distribution gradient vector. The pre-stored typical fault modes have corresponding regions in this feature space. In this embodiment, the aerodynamic conversion efficiency numerical range is standardized by 0-1 as the X-axis, the wind speed distribution uniformity coefficient is calculated by normalizing the wind speed standard deviation at each monitoring point in the duct as the Y-axis, and the temperature distribution gradient vector is constructed by the temperature difference between the heating hardware and the key points of the duct and converted into a scalar value as the Z-axis. The pre-stored typical fault modes include air duct blockage, fan blade wear, sensor drift, heat sink dust accumulation, and abnormal hardware power consumption. Each mode is marked in a specific area in the three-dimensional feature space using historical fault data. For example, air duct blockage corresponds to the aggregation area of ​​the low-middle range of the X-axis, the high range of the Y-axis, and the high range of the Z-axis.

[0068] The construction of this three-dimensional feature space transforms scattered parameter information into visualized spatial relationships, providing a structured comparison basis for fault mode matching and avoiding the one-sidedness of single-parameter feature matching. Compared with traditional single-parameter or dual-parameter feature matching, three-dimensional space can capture the cooperative abnormal features between parameters, such as the combination of low aerodynamic conversion efficiency, poor wind speed uniformity, and large temperature gradient. It can accurately distinguish faults with similar single-parameter features, such as duct blockage and heat sink dust accumulation, greatly improving the diagnostic accuracy.

[0069] The system calculates the distance between the current system state's position in the feature space and the regions of each typical fault mode. In this embodiment, Euclidean distance is used to calculate the straight-line distance between the current state point and the center of each fault mode region. The smaller the distance value, the higher the similarity between the current state and the fault mode. Distance quantization enables preliminary screening of fault modes, quickly identifying the candidate modes with the highest similarity, narrowing down the scope for subsequent accurate diagnosis, solving the problem of blind trial and error in fault mode testing in traditional manual troubleshooting, and improving the efficiency of root cause identification.

[0070] This study analyzes the time-series variations of aerodynamic conversion efficiency, wind speed distribution, and temperature gradient. In this embodiment, time-series data for these three parameters are extracted over the past 30 seconds to identify trend types (such as sudden drops, gradual drops, and fluctuating increases) and abrupt change points. For example, the time-series characteristic of aerodynamic conversion efficiency corresponding to fan blade wear is a slow decrease accompanied by small fluctuations, while the time-series characteristic of wind speed distribution corresponding to duct blockage is a step-like decrease followed by stabilization at a low level. Analyzing these time-series variation characteristics supplements the limitations of static spatial positioning from a dynamic evolution perspective, avoiding misjudgments of the same static characteristics due to different faults and improving fault mode differentiation. By capturing the process characteristics of parameter changes, gradual faults (such as fan aging) and sudden faults (such as foreign object blockage in the duct) can be accurately identified. Traditional static characteristics can only determine the current state and cannot distinguish the nature of the fault. This dynamic analysis allows maintenance personnel to formulate targeted handling strategies based on the fault evolution pattern.

[0071] The influence path analysis between parameters is performed to determine the transmission direction of abnormal parameters. In this embodiment, a parameter correlation model is constructed based on the physical mechanism of the heat dissipation system. The core influence paths are preset, such as fan speed to wind speed to aerodynamic conversion efficiency to temperature gradient, power consumption to temperature gradient, and finally to wind speed distribution. By analyzing the time sequence and numerical correlation of abnormal parameters, the transmission direction is determined. For example, if the abnormal wind speed distribution occurs first, followed by the decrease in aerodynamic conversion efficiency, and finally the abnormal temperature gradient, the abnormal transmission direction is determined to be wind speed distribution to aerodynamic conversion efficiency and then to temperature gradient.

[0072] This invention's path analysis can pinpoint the initial source of an anomaly, providing a clear tracing direction for root cause localization and preventing derivative anomalies from being misjudged as initial faults. This analysis method aligns with the physical conduction logic of the heat dissipation system, enabling it to penetrate apparent anomalies to find the root cause. For example, it avoids misjudging an increased temperature gradient as the root cause, instead tracing it back to the initial problem of abnormal wind speed distribution, thus resolving the fault at its source.

[0073] Based on the results of distance, time-series variation characteristics, and parameter influence path analysis, a candidate fault set is selected from a predefined root cause database. In this embodiment, the root cause database includes spatial feature intervals, temporal variation rules, and three-dimensional matching conditions for influence path directions corresponding to each typical fault mode. Fault modes that simultaneously meet the following criteria are selected to form the candidate fault set: distance less than a preset threshold, temporal feature matching degree higher than a set proportion, and consistent influence path direction. The candidate set is screened through the superposition of multiple conditions to achieve preliminary accurate filtering, ensuring that fault modes entering the subsequent verification stage all have a high matching basis. This multi-condition superposition screening mechanism avoids the omissions of single-condition screening and significantly improves the quality of the candidate set through three-dimensional matching, reducing the workload of subsequent verification stages and making the diagnostic process more efficient.

[0074] The candidate fault set is sorted by distance from smallest to largest, and each fault mode is checked sequentially to see if it can explain all abnormal parameters. In this embodiment, abnormal parameter coverage is used as the verification standard, that is, whether the typical abnormal manifestations corresponding to the fault mode include all abnormal parameters monitored by the current system. For example, the typical abnormal manifestations of fan blade wear are decreased aerodynamic conversion efficiency, uneven wind speed distribution, and increased temperature gradient. If the current system's abnormal parameters are completely included, then the mode is determined to explain all abnormalities. By verifying each fault mode one by one, the effectiveness of the screening is achieved, ensuring that the root cause can fully cover the abnormal phenomena and avoiding the omission of derivative abnormalities. This full-coverage verification standard can effectively eliminate partially matching interfering fault modes. For example, sensor drift can only explain temperature parameter abnormalities and cannot cover wind speed abnormalities, thus being accurately eliminated, ensuring the uniqueness and accuracy of the root cause.

[0075] The first fault mode that can explain all abnormal parameters is taken as the root cause. In this embodiment, after sorting the candidate fault set by distance, the mode with the highest similarity is verified first. If the typical abnormal behavior of this mode completely matches all current abnormal parameters, it is directly identified as the root cause. This similarity-first + full-parameter explanation judgment rule ensures both the matching accuracy of the root cause and the judgment efficiency, meeting the need for rapid location of high-probability root causes in operation and maintenance scenarios. This rule fully combines the advantages of probability priority and logical rigor, enabling rapid output of highly reliable results in emergency operation and maintenance scenarios, while avoiding sacrificing accuracy for speed, thus balancing diagnostic efficiency and accuracy.

[0076] If no single fault mode can explain all abnormal parameters, the first two fault modes are selected as the common root cause. In this embodiment, if a single fault mode can explain at most 80% of the abnormal parameters, and the combination of the first two fault modes can cover all abnormal parameters, then it is determined to be the common root cause. For example, the combination of fan blade wear and heat sink dust accumulation can explain complex anomalies such as decreased aerodynamic conversion efficiency, uneven wind speed distribution, increased temperature gradient, and sudden increase in local temperature. This combined root cause determination mechanism solves the problem that traditional single root cause determination cannot handle complex faults, improving the diagnostic accuracy in complex scenarios. In actual operation and maintenance, complex faults account for a significant proportion and are extremely difficult to troubleshoot. This mechanism can accurately identify scenarios with multiple causes and one effect, avoiding repeated problems caused by maintenance personnel dealing with only a single fault, and significantly improving the thoroughness of fault resolution.

[0077] The root cause identification information is embedded in the alarm signal output. In this embodiment, the root cause identification information includes the fault mode name, typical feature matching description, and preliminary handling suggestions. For example, the root cause is: air duct blockage; the matching features are: low aerodynamic conversion efficiency range, high wind speed distribution uniformity coefficient, and increased temperature gradient vector; the handling suggestion is: clean debris from the air duct inlet and heat sink area. This information is simultaneously embedded in the background logs and SMS alarms associated with graphic text warnings and audible and visual alarms. This linkage between root cause information and alarm signals allows maintenance personnel to obtain the source of the problem and the direction of handling as soon as they receive the alarm, avoiding blind troubleshooting and significantly shortening the fault resolution time. This design breaks down the information barrier between alarms and diagnosis, upgrading alarm signals from simple risk warnings to solution-oriented instructions, providing direct guidance, especially for inexperienced maintenance personnel, and lowering the maintenance threshold.

[0078] As one embodiment of the present invention, the collaborative performance analysis module is also used for: Before dynamically fitting the expected temperature rise curve, the real-time power consumption of the heat-generating hardware is monitored, and its power consumption trend within a future time window is predicted. In this embodiment, the real-time power consumption of the heat-generating hardware is collected through the motherboard power management chip or an independent power sensor, and the collection frequency is consistent with other parameters to ensure data timestamp synchronization. The duration of the future time window is calibrated according to the thermal response characteristics of the heat dissipation system, and is usually set to 5 minutes. This duration can cover the preheating cycle of most load changes, while avoiding the decrease in accuracy caused by excessively long predictions.

[0079] The power consumption trend prediction employs a time-series prediction model. The input data is a real-time power consumption sequence over the past 60 seconds. The model extracts short-term fluctuation characteristics, periodic patterns, and abrupt change nodes from the power consumption data, outputting the power consumption trend type and confidence level for the next 5 minutes. Trend types include stable trend, slow upward trend, rapid upward trend, slow downward trend, and rapid downward trend. The confidence level characterizes the reliability of the trend prediction. When the confidence level falls below a set threshold, a data supplementation mechanism is triggered to shorten the sampling interval and improve prediction accuracy. For example, if a software startup causes a slight increase in power consumption within 10 seconds, the model can predict that power consumption will enter a slow upward trend within the next 2 minutes, providing advance notice for subsequent temperature rise curve fitting.

[0080] In the dynamic fitting process, the trend of power consumption change is introduced as a prior condition so that the expected temperature rise curve can reflect the impact of upcoming load changes on temperature.

[0081] In this embodiment, the original dynamic fitting process uses a multivariate regression model, with input parameters including aerodynamic conversion efficiency, current hardware temperature, real-time power consumption, and host internal ambient temperature. The newly optimized model imports predicted power consumption trend parameters before fitting, and adjusts the weight allocation strategy for each input parameter according to the trend type.

[0082] The specific weight adjustment logic is as follows: if the prediction is an upward trend, especially a rapid upward trend, increase the weight of the power consumption parameter so that the model takes into account more the heat generation caused by the increased load when fitting the model. If the forecast is a downward trend, the weight of power consumption parameters should be appropriately reduced, while the influence of parameters such as ambient temperature and aerodynamic conversion efficiency should be increased. If the trend is stable, the original weight allocation remains unchanged. Meanwhile, during the phased fitting process, the residual analysis at each stage will additionally introduce a trend fit check, that is, determine whether the deviation between the actual temperature rise and the fitted curve based on the trend prediction is caused by a trend prediction deviation. If the deviation originates from a trend change, the trend prediction process will be restarted and the fitted curve updated.

[0083] The beneficial effect is that it breaks through the limitation of the original temperature rise curve fitting which is based only on current and historical parameters, and realizes advance adaptation to future load changes. When the traditional fitting method suddenly increases the load, the fitted curve will lag, resulting in a delay in risk assessment. However, by introducing the trend of power consumption change, the curve can predict the accelerated temperature rise caused by the increase in load in advance, giving the risk assessment module more time to identify potential collaborative failures.

[0084] As one embodiment of the present invention, the dynamic tolerance is adaptively adjusted based on the historical performance of the system's collaborative efficiency model: The accuracy of the predicted temperature rise curve over historical periods in predicting actual temperature rise is statistically analyzed. In this embodiment, the historical period is set according to the system operation scenario. The daily use scenario of desktop computers is set to 24 hours, and the high load scenario of server room is set to 12 hours, so as to ensure that the statistical data can cover a sufficient number of operating states and respond to changes in system characteristics in a timely manner.

[0085] The expected temperature rise curve is dynamically generated by the collaborative performance analysis module. After each fitting, the predicted temperature sequence and the actual temperature sequence collected during the same period are recorded within a preset time period.

[0086] The prediction accuracy is calculated using a deviation compliance statistical method. First, an allowable deviation value for a single data point is set, determined based on the hardware temperature monitoring accuracy, typically 1℃. Then, the predicted temperature data for each historical period is compared with the actual temperature data. The proportion of data points with deviations within the allowable range is the prediction accuracy. For example, if 86,400 temperature data points are collected over a 24-hour period, and 82,080 data points have a prediction deviation within 1℃, the prediction accuracy is 95%. During the statistical process, data from abnormal scenarios such as sensor malfunctions and sudden power outages are automatically removed to avoid interference from abnormal data in the accuracy calculation.

[0087] When the prediction accuracy is high, the dynamic tolerance is narrowed to improve the monitoring sensitivity. When prediction accuracy is low, the dynamic tolerance can be relaxed to reduce the false alarm rate.

[0088] In this embodiment, two key thresholds are preset to divide three accuracy levels, namely a high accuracy threshold and a low accuracy threshold, wherein the high accuracy threshold is set to 90% and the low accuracy threshold is set to 70%.

[0089] When the statistically obtained prediction accuracy is higher than the high accuracy threshold, it is judged as a high accuracy level; when the prediction accuracy is lower than the low accuracy threshold, it is judged as a low accuracy level; when the prediction accuracy is between the two thresholds, it is judged as a medium accuracy level, and the current dynamic tolerance remains unchanged.

[0090] The accuracy threshold range can be flexibly adjusted according to hardware type and operation and maintenance needs. For example, for server CPUs with extremely high requirements for temperature stability, the high accuracy threshold can be increased to 95% to ensure that the tolerance is narrowed only when the model prediction is reliable enough. For game graphics cards with frequent load fluctuations, the low accuracy threshold can be lowered to 65% to avoid excessively widening the tolerance due to a temporary drop in accuracy caused by short-term drastic load fluctuations.

[0091] Finally, the dynamic tolerance is adaptively adjusted based on the accuracy level. When the prediction accuracy is high, it indicates that the synergistic performance model has high fitting accuracy, and the expected temperature rise curve generated based on this model accurately reflects the actual temperature change pattern. In this case, the dynamic tolerance is narrowed to improve monitoring sensitivity. The adjustment magnitude is positively correlated with the degree to which the accuracy exceeds the high threshold. For example, when the prediction accuracy is 95%, the dynamic tolerance is narrowed by 30% based on the original dynamic tolerance; when the prediction accuracy is 92%, it is narrowed by 15%.

[0092] When prediction accuracy is low, it indicates that the fitting accuracy of the collaborative performance model is insufficient, which may be due to reasons such as hardware aging, long-term environmental changes, or sudden changes in load patterns. In this case, the dynamic tolerance should be relaxed to reduce the false alarm rate. The adjustment range is also positively correlated with the degree to which the accuracy falls below the low threshold. For example, when the prediction accuracy is 60%, the dynamic tolerance should be relaxed by 40%; when the prediction accuracy is 65%, it should be relaxed by 20%. The adjusted dynamic tolerance must also comply with the preset safety boundary range to ensure that excessive adjustment does not lead to missed or false alarms.

[0093] The beneficial effect lies in achieving a deep integration between dynamic tolerance and the actual performance of the collaborative efficiency model. Traditional fixed tolerances or adjustments based solely on real-time parameters cannot adapt to the changing accuracy of the model over time. However, adaptive adjustments based on historical performance ensure that the tolerance range always matches the model's predictive capabilities. This improves monitoring sensitivity when model accuracy is high, enabling earlier detection of potential risks; and reduces false alarms when model accuracy is low, ensuring stable system operation.

[0094] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A computer host heat dissipation monitoring and alarm system, characterized in that, include: The parameter acquisition module is used to acquire multi-source timing parameters of the computer host heat dissipation system. The multi-source timing parameters include at least the temperature of the heat-generating hardware, the speed of the cooling fan, and the wind speed at key monitoring points in the air duct. The collaborative performance analysis module is used to construct a system collaborative performance model based on multi-source time-series parameters. The construction and application of the system collaborative performance model includes: Calculate the aerodynamic conversion efficiency of the current system based on the relationship between cooling fan speed and airflow. Based on the time-series data of aerodynamic conversion efficiency and the temperature of the heating hardware, the expected temperature rise curve of the heating hardware is dynamically fitted through a system collaborative efficiency model. The risk assessment module is used to continuously compare the real-time collected hardware temperature data with the expected temperature rise curve, and to determine that the system has a risk of collaborative failure when the actual temperature rise continuously deviates from the expected temperature rise curve and exceeds the dynamic tolerance. The alarm execution module is used to generate a corresponding alarm signal when the risk assessment module determines that there is a risk of collaborative failure.

2. The computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, The key monitoring points within the air duct include at least the air duct inlet, the windward side of the heat-generating hardware, and the leeward side. The multi-source timing parameters also include power consumption data of the heat-generating hardware and ambient temperature data of the host.

3. The computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, When the collaborative performance analysis module calculates the aerodynamic conversion efficiency of the current system: Within a preset sampling window, the cooling fan speed sequence and the corresponding wind speed sequence are collected simultaneously. The collected rotational speed and wind speed sequences were screened for data validity, and abnormal data points caused by transient interference from the sensors were removed. The selected valid data points are divided into multiple intervals according to the rotational speed value, and the statistical characteristic value of the wind speed to rotational speed ratio in each interval is calculated. Based on the temperature change trend of the aforementioned heat-generating hardware, determine whether the system is in a thermal steady state: When the system is in thermal steady state, the mode of the statistical characteristic values ​​within the speed range is selected as the benchmark value of the current aerodynamic conversion efficiency. When the system is in a non-thermal steady state, the weighted average of the statistical characteristic values ​​within the speed range is selected as the benchmark value of the current aerodynamic conversion efficiency, wherein the weighting coefficient is positively correlated with the temperature stability of the corresponding data point; The calculated aerodynamic conversion efficiency benchmark value is compared with the preset efficiency threshold range. If the calculation results exceed the threshold range multiple times in a row, a sensor calibration prompt is triggered and a historical normal value is used as a substitute.

4. The computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, When the collaborative performance analysis module dynamically fits the expected temperature rise curve of the heat-generating hardware: Establish a multivariate regression model based on the aerodynamic conversion efficiency and the temperature of the heating hardware; Before each fitting, the consistency between the changing trend of the aerodynamic conversion efficiency and the changing trend of the power consumption of the heat-generating hardware is analyzed. Evaluate the correspondence between wind speed distribution patterns and aerodynamic conversion efficiency values; When the analysis results are inconsistent, execute the parameter re-acquisition process; A staged fitting method is adopted, and in each stage of the fitting process: Calculate the residual sequence between the current fitted curve and the actual temperature data; Based on the distribution characteristics of the residual sequence, the input parameters of the regression model are screened and optimized by combination. The fitting process is complete when the rate of change of the sum of squared residuals in multiple consecutive fitting stages is less than a set threshold. During model operation, the impact of changes in ambient temperature on the expected temperature rise curve is continuously monitored. When the ambient temperature changes to a predetermined range, the phased fitting process is restarted. The final fitted result is output as the valid expected temperature rise curve.

5. A computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, The dynamic tolerance is dynamically adjusted based on the aerodynamic conversion efficiency; When the aerodynamic conversion efficiency is lower than the first threshold, the dynamic tolerance is narrowed proportionally. When the aerodynamic conversion efficiency is higher than the second threshold, the dynamic tolerance is relaxed proportionally.

6. A computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, It also includes a risk grading module, used when a system is determined to have a risk of collaborative failure: A risk assessment matrix based on a combination of multi-dimensional parameters is constructed. The rows of the matrix represent the degree to which the aerodynamic conversion efficiency deviates from the normal range, and the columns represent the range to which the wind speed distribution uniformity deviates from the benchmark. Analyze the trajectory of the actual temperature rise deviating from the expected temperature rise curve to identify its degree of conformity with typical failure characteristics; Establish a parameter linkage analysis to check the time correspondence between the change in cooling fan speed and the power consumption fluctuation of heat-generating hardware, as well as the spatial distribution relationship between airflow velocity distribution and temperature gradient distribution. Based on the location results, change trajectory characteristics, and parameter linkage analysis conclusions of the risk assessment matrix, the level of collaborative failure risk is determined through a hierarchical decision-making model. Based on the final determined risk level, the corresponding tiered control strategy will be triggered.

7. A computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, It also includes a fault diagnosis module, used to perform a root cause identification process when the system is determined to have a risk of collaborative failure: A three-dimensional feature space is established based on the numerical range of aerodynamic conversion efficiency, the wind speed distribution uniformity coefficient, and the temperature distribution gradient vector, in which the pre-stored typical fault modes have corresponding regions in the feature space; Calculate the position of the current system state in the feature space and the distance between it and the regions of each typical fault mode; Analyze the time-series variations of aerodynamic conversion efficiency, wind speed distribution, and temperature gradient; Perform influence path analysis between execution parameters to determine the propagation direction of abnormal parameters; Based on the results of the distance, time series variation characteristics, and parameter influence path analysis, a candidate fault set is selected from a predefined root cause library; Sort the candidate fault set by distance from smallest to largest, and check in turn whether each fault mode can explain all the abnormal parameters; The first fault mode that can explain all the abnormal parameters is taken as the root cause of the fault; If no fault mode exists that can explain all the abnormal parameters, then the first two fault modes are selected as the common root cause of the fault. The root cause identification information of the fault is embedded in the alarm signal and output.

8. A computer host heat dissipation monitoring and alarm system according to claim 4, characterized in that, The collaborative performance analysis module is also used for: Before dynamically fitting the expected temperature rise curve, monitor the real-time power consumption of the heat-generating hardware and predict its power consumption trend within a future time window. The power consumption change trend is introduced as a priori condition during the dynamic fitting process so that the expected temperature rise curve can reflect the impact of upcoming load changes on temperature.

9. A computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, The dynamic tolerance is adaptively adjusted based on the historical performance of the system's collaborative efficiency model. The accuracy of the predicted temperature rise curve over historical periods in predicting actual temperature rise is statistically analyzed. When the prediction accuracy is high, the dynamic tolerance is narrowed to improve the monitoring sensitivity. When the prediction accuracy is low, the dynamic tolerance is relaxed to reduce the false alarm rate.

10. A computer host heat dissipation monitoring and alarm system according to claim 1, characterized in that, When the alarm execution module generates an alarm signal, it performs at least one of the following: Generate graphical or textual warning messages for display in the user interface; Generate the electrical signal that drives the audible and visual alarm. Generate log files to record system status.