A method for intelligent temperature monitoring of critical server components

By combining a dynamic threshold model with neural networks and reinforcement learning units, the server temperature threshold is adjusted in real time and an optimized control strategy is generated. This solves the shortcomings of existing systems in terms of real-time performance, dynamic adaptability, and intelligent regulation, and achieves precise control of server temperature and energy consumption optimization.

CN120723583BActive Publication Date: 2025-11-14四川华鲲振宇智能科技有限责任公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511188823.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-14
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing server temperature monitoring systems have shortcomings in terms of real-time performance, dynamic adaptability, comprehensiveness, and intelligent control, and cannot meet the temperature management needs of modern data center servers operating under high density and high load.

Method used

A dynamic threshold model is used to adjust the temperature threshold in real time. Through component correlation analysis and multi-dimensional data coupling analysis, an optimized control strategy is generated. Combined with neural networks and reinforcement learning units, the model parameters are optimized to achieve precise control of server temperature.

Benefits of technology

It enables rapid and precise control of server temperature, reducing the risk of overheating while balancing system stability and energy efficiency. Through spatiotemporal coupling analysis and feedback mechanisms, it continuously optimizes model parameters to adapt to load and environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723583B_ABST
    Figure CN120723583B_ABST
Patent Text Reader

Abstract

This invention relates to an intelligent temperature monitoring method for critical server components. It acquires specified server data and inputs it into a dynamic threshold model for training. Various data are collected in real-time and input into the trained dynamic threshold model. The method adjusts the temperature thresholds of each critical component in real-time based on changes in server load and environment. When the real-time temperature of a component exceeds or equals its corresponding temperature threshold, the temperature thresholds of related components are corrected in real-time. Temperature data from different components at the same time dimension are coupled and analyzed with temperature change curves at different time points. When the real-time temperature of a component exceeds or equals the temperature threshold determined by the dynamic threshold model, an early warning signal is issued. Based on the early warning signal, a temperature control strategy is generated, and corresponding actions are performed on the components. Through dynamic threshold modeling, correlation analysis, and coupling analysis, the server temperature status is comprehensively understood, and optimized control strategies are generated based on the early warning level, achieving rapid and precise control of server temperature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of server heat dissipation technology, and in particular relates to a method for intelligent monitoring of the temperature of important server components. Background Technology

[0002] With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, modern data center servers face the demands of high-density, high-load operation. During prolonged high-load operation, critical components such as the CPU, GPU, memory, and hard drive generate significant heat. Excessive heat in these components directly impacts system performance, leading to decreased processing speed, response delays, and in severe cases, component damage and server downtime, posing a significant risk to the stable operation of the data center. Therefore, accurate and real-time monitoring and control of the temperature of critical server components has become a core element in ensuring data center reliability. Current server temperature monitoring systems have many limitations and struggle to meet the monitoring needs of complex and dynamic environments. First, in terms of data acquisition and threshold setting, traditional systems often rely on timed sampling or fixed threshold alarm mechanisms. Timed sampling suffers from data lag, failing to reflect real-time temperature changes during sudden load increases or environmental shifts, potentially causing the system to fail to respond promptly to temperature anomalies. Fixed thresholds, on the other hand, cannot be dynamically adjusted based on real-time server load or environmental conditions, easily leading to delayed warnings or false alarms, and are ineffective in handling complex and ever-changing workload scenarios.

[0003] Secondly, existing systems often have limitations in terms of monitoring coverage and correlation analysis capabilities. Some systems can only monitor a single type of component, lacking comprehensive coverage of other important components such as GPUs, memory, and hard drives, making it difficult to grasp the overall temperature status of the server. Even if some systems cover multiple components, they fail to achieve correlation analysis between components. When the temperature of a certain component is abnormal, it is impossible to identify its potential impact on related components in a timely manner, resulting in incomplete early warnings and increasing the risk of overall system overheating.

[0004] Furthermore, in terms of early warning and control strategies, traditional systems have relatively simple early warning mechanisms, mostly triggered by a single threshold, and their control strategies lack flexibility and optimization capabilities. After an early warning signal is issued, the cooling system adjustment is often in a fixed mode, failing to dynamically optimize based on component temperature change trends, load conditions, and energy consumption costs. This may lead to excessive energy consumption due to heat dissipation or insufficient adjustment that fails to effectively cool the system. Simultaneously, the regular maintenance and calibration of sensors rely on manual operation, lacking an automatic calibration mechanism. Over long-term use, this can easily lead to data deviations, affecting the accuracy and reliability of monitoring.

[0005] In addition, existing systems have shortcomings in data analysis dimensions, mostly focusing on temperature monitoring at a single time point or for a single component. They fail to couple and analyze temperature data from different components at the same time dimension with temperature change curves at different time points, making it difficult to uncover the inherent patterns and potential risks of temperature changes. This leads to inaccurate judgment of the root causes of temperature anomalies and affects the effectiveness of subsequent control measures.

[0006] In summary, existing server temperature monitoring systems have significant shortcomings in terms of real-time performance, dynamic adaptability, comprehensiveness, and intelligent control, failing to meet the temperature management needs of modern data center servers operating under high density and high load. Therefore, there is an urgent need for an intelligent monitoring method capable of adjusting temperature thresholds in real time, performing component correlation analysis and multi-dimensional data coupling analysis, and generating optimized control strategies based on early warning signals. Summary of the Invention

[0007] The purpose of this invention is to provide a method for intelligent temperature monitoring of important server components, which can adjust temperature thresholds in real time, realize component correlation analysis and multi-dimensional data coupling analysis, and generate optimized control strategies based on early warning signals.

[0008] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:

[0009] A method for intelligent temperature monitoring of critical server components includes the following steps:

[0010] S1: Obtain historical temperature data, load data, and environmental data from the server and input them into the created dynamic threshold model for model training;

[0011] S2: The data acquisition module collects temperature data, load data and environmental data of each component in real time and inputs them into the trained dynamic threshold model. The dynamic threshold model adjusts the temperature threshold of each important component in real time according to the changes in server load and environment.

[0012] S3: When the real-time temperature of a component is greater than or equal to the corresponding temperature threshold, the dynamic threshold model will analyze the temperature change trend of its associated components and correct the temperature threshold of the associated components in real time.

[0013] S4: Real-time temperature data, load data and environmental data of each component are collected and transmitted to the monitoring platform. The monitoring platform couples and analyzes the temperature data of different components at the same time dimension with the temperature change curves at different time points.

[0014] S5: When the real-time temperature of a component is greater than or equal to the temperature threshold determined by the dynamic threshold model, the monitoring system immediately issues an early warning signal.

[0015] S6: Generate a temperature control strategy based on the early warning signal, generate a corresponding temperature control command based on the temperature control strategy and transmit it to the corresponding execution unit, and the execution unit performs the corresponding action.

[0016] Preferably, the specific process of step S1 is as follows:

[0017] S11: Extract historical temperature data, load data, and environmental data from the server database and preprocess the data.

[0018] S12: Extract the preprocessed historical temperature data, load data, and environmental data respectively, and apply the specified features to the data.

[0019] S13: Based on the server's operating status and fault records, determine the model's label. The label is the reasonable temperature threshold range for each component under each specified feature. When the actual temperature exceeds this range, it is considered an abnormal situation.

[0020] S14: Initialize the parameters of the dynamic threshold model, and input the extracted specified features and labels into the dynamic threshold model for model training.

[0021] Preferably, the specific process of inputting the extracted specified features and labels into the dynamic threshold model for model training in step S14 is as follows;

[0022] S141: The dynamic threshold model continuously adjusts the weights and biases of the network based on specified features and labels through neural network units to minimize the error between the predicted value and the label;

[0023] S142: The dynamic threshold model uses a reinforcement learning unit to learn and optimize decision-making strategies based on a preset reward function and interaction with the environment, enabling the model to better adapt to different load scenarios and environmental conditions.

[0024] S143: The neural network unit and the reinforcement learning unit train the model iteratively, and evaluate the specified performance of the model through the test set after reaching the specified number of iterations or the specified conditions. When the evaluation result meets the specified requirements, step S2 is executed.

[0025] Preferably, the specific process of the dynamic threshold model in step S2 adjusting the temperature thresholds of each important component in real time according to changes in server load and environment is as follows:

[0026] S21: The dynamic threshold model extracts specified features from real-time temperature data, load data, and environmental data, maps the extracted specified features to specified intervals, and obtains standardized specified features.

[0027] S22: The standardized specified features are processed in real time using the trained dynamic threshold model, including the neural network unit using the weights and biases optimized during training to perform nonlinear mapping on the input features and make preliminary predictions on the temperature threshold range of each component under the current load and environmental conditions, and the reinforcement learning unit adjusting the prediction results of the neural network in real time according to the reward function based on the current operating status of the server.

[0028] S23: The dynamic threshold model generates real-time temperature warning thresholds for each important component based on real-time calculation results, and formats the threshold data into a format that the monitoring system can recognize, including component identification, current upper threshold value, lower threshold value, and threshold effective time.

[0029] Preferably, the specific process of step S3 is as follows:

[0030] S31: Receive real-time temperature data of each component transmitted by the data acquisition module and compare it with the corresponding component temperature threshold. When the real-time temperature of a component is greater than or equal to its current temperature threshold, immediately trigger the correlation analysis mechanism, mark the component as the triggering component, and record its identifier, temperature value at the time of exceeding the threshold, current load data, and key information of environmental data as the initial basis for correlation analysis.

[0031] S32: Automatically identify all components that are directly or indirectly related to the triggering component based on the preset association relationship map, i.e., associated components;

[0032] S33: Extract temperature data, load data, and environmental data of associated components within a specified time window before and after the triggering component exceeds the threshold.

[0033] S34: Perform trend analysis on the extracted temperature data of related components to obtain trend characteristics including temperature change rate, temperature fluctuation frequency, and correlation between temperature and load.

[0034] S35: Based on the extracted trend features, the temperature correlation strength between the associated component and the triggering component is calculated through mutual information entropy. The correlation strength is represented by a value of 0-1.

[0035] S36: By combining the correlation strength and the difference between the current temperature of the associated component and its own threshold, assess the degree of impact of the overheating of the triggering component on each associated component;

[0036] S37: Dynamically generate threshold correction coefficients based on the degree of influence of related components; for related components with high influence, the correction coefficient value is larger; for related components with low influence, the correction coefficient value is smaller.

[0037] S38: Perform threshold correction calculation: The model uses the formula "corrected threshold = original threshold × correction coefficient - temperature trend compensation value" to calculate the new threshold of the associated components;

[0038] The temperature trend compensation value is determined based on the rate of temperature change of the associated components. If the temperature shows an accelerating upward trend, the compensation value is positive, and the threshold is further reduced; if the temperature tends to stabilize, the compensation value is 0.

[0039] S39: Verify the calculated correction threshold to ensure that it is not lower than the minimum safe temperature of the component and that the correction range is within a reasonable range. After verification, the model will format the correction threshold into structured data that includes component identification, upper and lower limits of the new threshold, correction effective time, and correction basis.

[0040] S40: Update the warning threshold of the associated component and start real-time comparison. If the temperature of the associated component reaches or exceeds the new threshold after the corrected threshold takes effect, the monitoring system will trigger cooling adjustment first according to the warning mechanism and push the associated warning information to the administrator, explaining the potential risk of the component being affected by the triggered component.

[0041] Preferably, the specific process of step S4 is as follows:

[0042] S41: Perform horizontal coupling analysis in the same time dimension: For each time window, the platform extracts the temperature data of all components within that time window, combines the load data and environmental data to perform horizontal coupling analysis, calculates the correlation coefficient of the temperatures of different components using the Pearson coefficient, identifies component clusters with consistent temperature change trends, and analyzes the correlation between load data and temperature data.

[0043] S42: Vertical Coupling Analysis at Different Time Nodes: Extract the temperature change curve of a component or component cluster within a continuous time window, and perform vertical coupling analysis with the temperature curves of the same period in history and similar load scenarios. By comparing the similarity of the curve shapes through the dynamic time warping algorithm, abnormal trends in temperature changes can be identified. At the same time, the derivative changes of the temperature curves are analyzed, and combined with the load fluctuations within the time window, it is determined whether the temperature change is caused by load changes.

[0044] S43: Perform spatiotemporal cross-coupling verification: The platform will cross-verify the results of horizontal and vertical analysis.

[0045] Preferably, step S2, in which the temperature thresholds of each important component are adjusted in real time, also includes a feedback process, as follows:

[0046] Real-time recording of temperature changes, warning triggering status, and cooling system adjustment effects of each component is used to generate feedback data. The feedback data includes whether the component temperature returns to normal after the temperature threshold is adjusted, whether the warning is accurate, and whether the cooling energy consumption is within a reasonable range.

[0047] Feedback data is transmitted to the reinforcement learning unit of the dynamic threshold model, which uses an incremental learning algorithm to fine-tune itself. This includes updating the reward function based on the feedback data when a large deviation in threshold prediction is found under a certain load scenario, and the neural network unit correcting the local weights through backpropagation. This allows the model to continuously adapt to new load patterns and environmental changes that occur during server operation, thereby continuously improving the accuracy and real-time performance of threshold adjustment.

[0048] Preferably, the specific process of step S6 is as follows:

[0049] S61: Perform structured parsing on the warning signal, extract key information including the component identifier that triggered the warning, the difference between the current temperature value and the threshold, the rate of temperature change, the status of related components, the current load data and environmental data, and form a standardized warning event data packet;

[0050] S62: Based on the warning event data packet, classify the warning level according to preset rules, including at least Level 1 warning, Level 2 warning and Level 3 warning. Call the preset basic control strategy library according to the warning level. Level 1 warning matches the powerful heat dissipation strategy; Level 2 warning matches the enhanced heat dissipation strategy; Level 3 warning matches the preventive heat dissipation strategy.

[0051] S63: Call the coupling analysis results to dynamically optimize the temperature control strategy in the basic control strategy library;

[0052] S64: During the strategy generation process, the optimal solution is calculated through a multi-objective optimization algorithm to balance the relationship between heat dissipation efficiency, energy consumption cost and hardware loss.

[0053] The beneficial effects of this invention include:

[0054] This invention provides an intelligent temperature monitoring method for critical server components. It acquires specified server data and inputs it into a dynamic threshold model for training. The method collects various data in real time and inputs it into the trained dynamic threshold model. Based on changes in server load and environment, it adjusts the temperature thresholds of each critical component in real time. When the real-time temperature of a component exceeds or equals its corresponding temperature threshold, it corrects the temperature thresholds of related components in real time. It couples and analyzes temperature data from different components at the same time dimension with temperature change curves at different time points. When the real-time temperature of a component exceeds or equals the temperature threshold determined by the dynamic threshold model, it issues an early warning signal. Based on the early warning signal, it generates a temperature control strategy and executes corresponding actions on the components. By precisely adjusting temperature thresholds through the dynamic threshold model and comprehensively understanding the server temperature status through correlation and coupling analysis, it generates optimized control strategies based on the early warning level, achieving rapid and accurate temperature control of the server, effectively reducing the risk of overheating, and balancing system stability and energy efficiency.

[0055] First, through spatiotemporal coupling analysis, temperature data of different components under the same time dimension are coupled with temperature change curves at different time points for analysis. This not only identifies component clusters with consistent temperature change trends through horizontal analysis, but also compares temperature curves from the same historical period and similar load scenarios through vertical analysis. Combined with tools such as dynamic time warping algorithms, the inherent patterns of temperature changes are uncovered. This overcomes the limitations of traditional systems that monitor a single time point or a single component, enabling more accurate identification of the root cause of temperature anomalies and providing a reliable basis for subsequent regulation.

[0056] Secondly, when generating temperature control strategies based on early warning signals, the system matches the basic strategy according to the early warning level and performs dynamic optimization by combining the results of coupling analysis. At the same time, it balances heat dissipation efficiency, energy consumption cost and hardware loss through multi-objective optimization algorithms.

[0057] Finally, a robust feedback mechanism is established to transmit real-time feedback data, including temperature changes, early warning effectiveness, and cooling energy consumption, to the dynamic threshold model. Incremental learning algorithms continuously optimize model parameters. Neural network units can correct local weights, and reinforcement learning units can update the reward function, enabling the model to continuously adapt to new load patterns and environmental changes, maintaining the accuracy of threshold adjustments over the long term. Attached Figure Description

[0058] Figure 1 This is a flowchart illustrating the intelligent temperature monitoring method for critical server components according to the present invention.

[0059] Figure 2 This is a schematic diagram of the layout of the temperature sensor in the server of the present invention. Detailed Implementation

[0060] The following is in conjunction with the appendix Figures 1-2The present invention will be further described in detail below:

[0061] Example 1

[0062] See appendix Figure 1 As shown, a method for intelligent temperature monitoring of critical server components includes the following steps:

[0063] S1: Obtain historical temperature data, load data, and environmental data from the server and input them into the created dynamic threshold model for model training.

[0064] The specific process of step S1 is as follows:

[0065] S11: Extract historical temperature data, load data, and environmental data from the server database and preprocess the data.

[0066] S12: Extract the preprocessed historical temperature data, load data, and environmental data respectively, and apply the specified features to the data.

[0067] S13: Based on the server's operating status and fault records, determine the model's label. The label is the reasonable temperature threshold range for each component under each specified feature. When the actual temperature exceeds this range, it is considered an abnormal situation.

[0068] S14: Initialize the parameters of the dynamic threshold model, and input the extracted specified features and labels into the dynamic threshold model for model training.

[0069] The specific process of inputting the extracted specified features and labels into the dynamic threshold model for model training in step S14 is as follows;

[0070] S141: The dynamic threshold model continuously adjusts the weights and biases of the network based on specified features and labels through neural network units to minimize the error between the predicted value and the label;

[0071] S142: The dynamic threshold model uses a reinforcement learning unit to learn and optimize decision-making strategies based on a preset reward function and interaction with the environment, enabling the model to better adapt to different load scenarios and environmental conditions.

[0072] S143: The neural network unit and the reinforcement learning unit train the model iteratively, and evaluate the specified performance of the model through the test set after reaching the specified number of iterations or the specified conditions. When the evaluation result meets the specified requirements, step S2 is executed.

[0073] S2: The data acquisition module collects real-time temperature, load, and environmental data from each component and inputs them into the trained dynamic threshold model. The data acquisition module collects temperature data by placing temperature sensors in each component; the layout of the temperature sensors is shown below. Figure 2 As shown, the temperature sensors include IBMC temperature sensors, PICE temperature sensors, memory temperature sensors, CPU temperature sensors, and storage temperature sensors. The dynamic threshold model adjusts the temperature thresholds of each important component in real time according to changes in server load and environment.

[0074] S3: When the real-time temperature of a component is greater than or equal to the corresponding temperature threshold, the dynamic threshold model will analyze the temperature change trend of its associated components and correct the temperature threshold of the associated components in real time.

[0075] S4: Real-time temperature data, load data and environmental data of each component are collected and transmitted to the monitoring platform. The monitoring platform couples and analyzes the temperature data of different components at the same time dimension with the temperature change curves at different time points.

[0076] S5: When the real-time temperature of a component is greater than or equal to the temperature threshold determined by the dynamic threshold model, the monitoring system immediately issues an early warning signal.

[0077] S6: Generate a temperature control strategy based on the early warning signal, generate a corresponding temperature control command based on the temperature control strategy and transmit it to the corresponding execution unit, and the execution unit performs the corresponding action.

[0078] Example 2

[0079] Based on Example 1, the specific process of the dynamic threshold model in step S2 adjusting the temperature thresholds of each important component in real time according to changes in server load and environment is as follows:

[0080] S21: The dynamic threshold model extracts specified features from real-time temperature data, load data, and environmental data, maps the extracted specified features to specified intervals, and obtains standardized specified features.

[0081] S22: The standardized specified features are processed in real time using the trained dynamic threshold model, including the neural network unit using the weights and biases optimized during training to perform nonlinear mapping on the input features and make preliminary predictions on the temperature threshold range of each component under the current load and environmental conditions, and the reinforcement learning unit adjusting the prediction results of the neural network in real time according to the reward function based on the current operating status of the server.

[0082] S23: The dynamic threshold model generates real-time temperature warning thresholds for each important component based on real-time calculation results, and formats the threshold data into a format that the monitoring system can recognize, including component identification, current upper threshold value, lower threshold value, and threshold effective time.

[0083] The specific process of step S3 is as follows:

[0084] S31: The dynamic threshold model receives real-time temperature data of each component transmitted by the data acquisition module and compares it with the corresponding component temperature threshold. When the real-time temperature of a component is greater than or equal to its current temperature threshold, the correlation analysis mechanism is immediately triggered, the component is marked as the trigger component, and its identifier, temperature value at the time of exceeding the threshold, current load data and key information of environmental data are recorded as the initial basis for correlation analysis.

[0085] S32: The dynamic threshold model automatically identifies all components that are directly or indirectly related to the triggering component, i.e., related components, based on a preset correlation graph;

[0086] S33: The dynamic threshold model extracts temperature data, load data, and environmental data of related components within a specified time window before and after the triggering component exceeds the threshold. The time window setting must take into account both the completeness and timeliness of the data, including both the temperature change trend before triggering and the real-time data after triggering, so as to comprehensively analyze the temperature response characteristics of related components. At the same time, the extracted data is subjected to a second validity check to remove invalid data and ensure the reliability of the analysis basis.

[0087] S34: The dynamic threshold model performs trend analysis on the extracted temperature data of related components to obtain trend characteristics, including the rate of temperature change (the magnitude of temperature rise / fall per unit time), the frequency of temperature fluctuations (the number of temperature oscillations in a short period of time), and the correlation between temperature and load. The correlation between temperature and load includes, for example, the degree of synchronous temperature change when the load of related components increases. If the triggering component is the CPU, after it exceeds the threshold, the model will analyze whether the GPU temperature rises faster as the CPU load increases, or whether the memory temperature fluctuates abnormally because the heat dissipation space is occupied by the CPU.

[0088] S35: Based on the extracted trend features, the dynamic threshold model calculates the temperature correlation strength between the associated component and the triggering component through mutual information entropy. The correlation strength is represented by a value of 0-1. The closer the value is to 1, the more synchronized the temperature change trends of the two components are, and the greater the possibility of being affected by the overheating of the triggering component. Because the CPU and GPU share a cooling fan, their temperature correlation strength is usually higher than that between the CPU and the remote hard drive.

[0089] S36: Combining the correlation strength and the difference between the current temperature of the related component and its own threshold, the dynamic threshold model evaluates the impact of the triggering component's overheating on each related component; if the temperature of the related component changes rapidly, the correlation strength is high, and the current temperature is close to its own threshold, it is judged as a high impact; if the temperature changes slowly, the correlation is weak, and the current temperature is far below the threshold, it is judged as a low impact.

[0090] S37: Based on the degree of influence of related components, a threshold correction coefficient is dynamically generated. For related components with a high degree of influence, the correction coefficient is larger, meaning that their threshold needs to be lowered more significantly to provide early warning of potential overheating risks. For related components with a low degree of influence, the correction coefficient is smaller, and the threshold can be fine-tuned. The generation of the correction coefficient also needs to refer to the historical case database. When a certain type of correlation repeatedly shows a situation in historical data where "the triggered component overheats and the related component overheats within a short period of time", the corresponding correction coefficient will be automatically increased to enhance the foresight of the warning.

[0091] S38: Perform threshold correction calculation: The model uses the formula "corrected threshold = original threshold × correction coefficient - temperature trend compensation value" to calculate the new threshold of the associated components;

[0092] The temperature trend compensation value is determined based on the temperature change rate of the associated component. If the temperature shows an accelerating upward trend, the compensation value is positive, further reducing the threshold. If the temperature tends to stabilize, the compensation value is 0. For example, if the original threshold of an associated component is 80℃, the correction coefficient is 0.8, the temperature rise rate is 2℃ / second, and the compensation value is 5℃, then the corrected threshold = 80×0.8-5=59℃. The warning is triggered in advance by lowering the threshold.

[0093] S39: The calculated correction threshold needs to be verified to ensure that it is not lower than the minimum safe temperature of the component, and that it is not lower than the lower limit of the room temperature operation of the electronic components of the component. Furthermore, the correction range must be within a reasonable range, such as a single correction not exceeding 30% of the original threshold, to avoid false alarms due to over-correction. After successful verification, the model will format the correction threshold into structured data that includes the component identifier, the upper and lower limits of the new threshold, the effective time of the correction, and the basis for the correction, such as the strength of the correlation and the degree of impact.

[0094] S40: The corrected threshold is pushed to the server temperature monitoring platform in real time via an encrypted channel. The platform immediately updates the warning thresholds for related components and initiates real-time comparison. If the temperature of a related component reaches or exceeds the new threshold after the corrected threshold takes effect, the monitoring system will prioritize triggering cooling adjustments according to the intelligent warning mechanism. This can involve specifically increasing the cooling fan speed of the related component and pushing related warning information to the administrator, explaining the potential risk of the component being affected by the triggered component.

[0095] Example 3

[0096] Based on Example 1 or Example 2, the specific process of step S4 is as follows:

[0097] S41: Perform horizontal coupling analysis within the same time dimension: For each time window, the platform extracts the temperature data of all components within that time window, and performs horizontal coupling analysis by combining load data and environmental data. It uses the Pearson coefficient to analyze the temperature synchronization of the CPU and GPU, calculates the correlation coefficient of different component temperatures, and identifies component clusters with consistent temperature change trends, including high-load core clusters composed of CPU, GPU, and memory. It analyzes the correlation between load data and temperature data, and uses a multivariate regression model to calculate "the impact of a 10% increase in CPU utilization on its own and related component temperatures." Finally, it combines environmental data to correct the analysis results; for example, when the ambient temperature suddenly rises by 5°C, it adjusts the threshold for judging abnormal component temperatures to eliminate environmental interference factors.

[0098] S42: Vertical Coupling Analysis at Different Time Nodes: The platform extracts the temperature change curve of a component or component cluster within a continuous time window and performs vertical coupling analysis with the temperature curves of the same period in history and under similar load scenarios. Similar load scenarios can be the same CPU utilization range. The similarity of curve shapes is compared using a dynamic time warping algorithm to identify abnormal trends in temperature changes. Simultaneously, the derivative change of the temperature curve (i.e., the rate of temperature change) is analyzed. Combined with load fluctuations within the time window, it is determined whether the temperature change is caused by load changes, because a sudden increase in the temperature rate when the load is stable may be a precursor to hardware failure.

[0099] S43: Perform spatiotemporal cross-coupling verification: The platform cross-verifies the results of horizontal and vertical analysis. When the horizontal analysis reveals that the CPU and memory temperatures rise synchronously within a certain time window, the vertical analysis checks whether this phenomenon has occurred in historical data. If it only occurs under high ambient temperature and sudden load increases, it is judged as a normal load response; if it occurs for the first time under low load and normal conditions, it is marked as abnormal, triggering further investigation. Simultaneously, considering the spatial location of components, the impact of physical distance on temperature coupling is analyzed, because components that are closer together have significantly higher temperature coupling than components that are farther apart, which can rule out misjudgments due to non-physical correlations, such as "synchronous temperature changes between the hard drive and CPU may be due to load linkage rather than heat conduction."

[0100] Step S2, in the process of adjusting the temperature thresholds of each important component in real time, also includes a feedback process, as detailed below:

[0101] Real-time recording of temperature changes, warning triggering status, and cooling system adjustment effects of each component is used to generate feedback data. The feedback data includes whether the component temperature returns to normal after the temperature threshold is adjusted, whether the warning is accurate, and whether the cooling energy consumption is within a reasonable range.

[0102] Feedback data is transmitted to the reinforcement learning unit of the dynamic threshold model. The model uses an incremental learning algorithm to fine-tune itself. When a large deviation in threshold prediction is found in a certain load scenario, the reinforcement learning unit updates the reward function based on the feedback data. The neural network unit corrects the local weights through backpropagation, so that the model can continuously adapt to new load patterns and environmental changes that occur during server operation, and continuously improve the accuracy and real-time performance of threshold adjustment.

[0103] The specific process of step S6 is as follows:

[0104] S61: The monitoring platform performs structured parsing of the early warning signal, extracting key information including the component identifier that triggered the early warning, the difference between the current temperature value and the threshold, the rate of temperature change, the status of related components, current load data, and environmental data, forming a standardized early warning event data packet;

[0105] S62: Based on the warning event data packet, classify the warning level according to preset rules;

[0106] Level 1 warnings apply to scenarios where "a single component's temperature exceeds the threshold by more than 5°C, or the temperature change rate is ≥3°C / second, or multiple components exceed the threshold simultaneously," such as "CPU01 temperature 85°C exceeding the threshold of 78°C, and GPU03 temperature 80°C exceeding the threshold of 75°C." Level 2 warnings apply to scenarios where "a single component's temperature exceeds the threshold by 1-5°C, and the temperature change rate is 1-3°C / second," such as "Hard Disk 05 temperature 62°C exceeding the threshold of 59°C, at a rate of 1.5°C / second." Level 3 warnings apply to scenarios where component temperatures are close to the threshold but the change trend is stable, such as "Memory 02 temperature 70°C, close to the threshold of 71°C, at a rate of 0.3°C / second." Different levels correspond to different response priorities and control strengths.

[0107] S62: Based on the warning level, the preset basic control strategy library is invoked. Level 1 warning is matched with "powerful heat dissipation strategy", which includes a combination of measures such as "fan running at full speed + liquid cooling system at full load" and "temporarily reducing the load of related components"; Level 2 warning is matched with "enhanced heat dissipation strategy", which can "increase the fan speed of the target component to 80% + increase the liquid cooling flow by 20%"; Level 3 warning is matched with "preventive heat dissipation strategy", which can "maintain the fan speed of the target component at 60% and monitor it dynamically".

[0108] S63: Call the coupling analysis results to dynamically optimize the temperature control strategy in the basic control strategy library. For example, when the horizontal coupling analysis shows that "CPU01 and GPU03 are in the same heat dissipation air duct", the speed of all fans in the air duct will be increased simultaneously during the first-level warning to avoid local heat dissipation blind spots.

[0109] When the longitudinal coupling analysis finds that "under similar loads in the same period of history, simply enhancing heat dissipation will cause the CPU temperature to rebound after 30 minutes", a predictive measure of "triggering load balancing scheduling after 15 minutes" will be added to the strategy.

[0110] If the spatial analysis shows that "the target component is located in a heat dissipation dead zone", then the backup heat dissipation device in that area will be activated first.

[0111] S64: During the strategy generation process, the optimal solution is calculated through a multi-objective optimization algorithm to balance the relationship between heat dissipation efficiency, energy consumption cost and hardware loss.

[0112] In summary, the intelligent temperature monitoring method for critical server components provided by this invention acquires specified server data and inputs it into a dynamic threshold model for training. It collects various data in real time and inputs it into the dynamic threshold model, adjusting the temperature thresholds of each critical component in real time based on changes in server load and environment. It also corrects the temperature thresholds of related components in real time, coupling and analyzing temperature data from different components at the same time dimension with temperature change curves at different time points. Based on early warning signals, it generates temperature control strategies and executes corresponding actions on the components. By precisely adjusting temperature thresholds through the dynamic threshold model and comprehensively understanding the server temperature status through correlation and coupling analysis, it generates optimized control strategies based on the early warning level.

[0113] By employing spatiotemporal coupling analysis, temperature data from different components within the same time dimension are coupled with temperature change curves at different time points. This overcomes the limitations of traditional systems that monitor only a single time point or a single component, enabling more accurate identification of the root cause of temperature anomalies and providing a reliable basis for subsequent regulation. When generating temperature control strategies based on early warning signals, the system matches the basic strategy according to the early warning level and dynamically optimizes it based on the coupling analysis results. Simultaneously, a multi-objective optimization algorithm balances heat dissipation efficiency, energy consumption costs, and hardware losses. A robust feedback mechanism is established, transmitting feedback data, including temperature changes, early warning effectiveness, and cooling energy consumption, to the dynamic threshold model in real time. Incremental learning algorithms continuously optimize model parameters. Neural network units can correct local weights, and reinforcement learning units can update the reward function, enabling the model to continuously adapt to new load patterns and environmental changes, maintaining the accuracy of threshold adjustments over the long term.

Claims

1. A method for intelligent temperature monitoring of critical server components, characterized in that, Includes the following steps: S1: Obtain historical temperature data, load data, and environmental data from the server and input them into the created dynamic threshold model for model training; S2: The data acquisition module collects temperature data, load data and environmental data of each component in real time and inputs them into the trained dynamic threshold model. The dynamic threshold model adjusts the temperature threshold of each important component in real time according to the changes in server load and environment. S3: When the real-time temperature of a component is greater than or equal to the corresponding temperature threshold, the dynamic threshold model will analyze the temperature change trend of its associated components and correct the temperature threshold of the associated components in real time. S4: Real-time temperature data, load data and environmental data of each component are collected and transmitted to the monitoring platform. The monitoring platform couples and analyzes the temperature data of different components at the same time dimension with the temperature change curves at different time points. S5: When the real-time temperature of a component is greater than or equal to the temperature threshold determined by the dynamic threshold model, the monitoring system immediately issues an early warning signal. S6: Generate a temperature control strategy based on the early warning signal, generate a corresponding temperature control command based on the temperature control strategy and transmit it to the corresponding execution unit, and the execution unit performs the corresponding action; The specific process of step S3 is as follows: S31: Receive real-time temperature data of each component transmitted by the data acquisition module and compare it with the corresponding component temperature threshold. When the real-time temperature of a component is greater than or equal to its current temperature threshold, immediately trigger the correlation analysis mechanism, mark the component as the triggering component, and record its identifier, temperature value at the time of exceeding the threshold, current load data, and key information of environmental data as the initial basis for correlation analysis. S32: Automatically identify all components that are directly or indirectly related to the triggering component based on the preset association relationship map, i.e., associated components; S33: Extract temperature data, load data, and environmental data of associated components within a specified time window before and after the triggering component exceeds the threshold. S34: Perform trend analysis on the extracted temperature data of related components to obtain trend characteristics including temperature change rate, temperature fluctuation frequency, and correlation between temperature and load. S35: Based on the extracted trend features, the temperature correlation strength between the associated component and the triggering component is calculated through mutual information entropy. The correlation strength is represented by a value of 0-1. S36: By combining the correlation strength and the difference between the current temperature of the associated component and its own threshold, assess the degree of impact of the overheating of the triggering component on each associated component; S37: Dynamically generate threshold correction coefficients based on the degree of influence of related components; for related components with high influence, the correction coefficient value is larger; for related components with low influence, the correction coefficient value is smaller. S38: Perform threshold correction calculation: The model uses the formula "corrected threshold = original threshold × correction coefficient - temperature trend compensation value" to calculate the new threshold of the associated components; The temperature trend compensation value is determined based on the rate of temperature change of the associated components. If the temperature shows an accelerating upward trend, the compensation value is positive, and the threshold is further reduced; if the temperature tends to stabilize, the compensation value is 0. S39: Verify the calculated correction threshold to ensure that it is not lower than the minimum safe temperature of the component and that the correction range is within a reasonable range. After verification, the model will format the correction threshold into structured data that includes component identification, upper and lower limits of the new threshold, correction effective time, and correction basis. S40: Update the warning threshold of the associated component and start real-time comparison. If the temperature of the associated component reaches or exceeds the new threshold after the corrected threshold takes effect, the monitoring system will trigger cooling adjustment first according to the warning mechanism and push the associated warning information to the administrator, explaining the potential risk of the component being affected by the triggered component.

2. The method for intelligent temperature monitoring of critical server components according to claim 1, characterized in that, The specific process of step S1 is as follows: S11: Extract historical temperature data, load data, and environmental data from the server database and preprocess the data. S12: Extract the preprocessed historical temperature data, load data, and environmental data respectively, and apply the specified features to the data. S13: Based on the server's operating status and fault records, determine the model's label. The label is the reasonable temperature threshold range for each component under each specified feature. When the actual temperature exceeds this range, it is considered an abnormal situation. S14: Initialize the parameters of the dynamic threshold model, and input the extracted specified features and labels into the dynamic threshold model for model training.

3. The method for intelligent temperature monitoring of critical server components according to claim 2, characterized in that, The specific process of inputting the extracted specified features and labels into the dynamic threshold model for model training in step S14 is as follows; S141: The dynamic threshold model continuously adjusts the weights and biases of the network based on specified features and labels through neural network units to minimize the error between the predicted value and the label; S142: The dynamic threshold model uses a reinforcement learning unit to learn and optimize decision-making strategies based on a preset reward function and interaction with the environment, enabling the model to better adapt to different load scenarios and environmental conditions. S143: The neural network unit and the reinforcement learning unit train the model iteratively, and evaluate the specified performance of the model through the test set after reaching the specified number of iterations or the specified conditions. When the evaluation result meets the specified requirements, step S2 is executed.

4. The method for intelligent temperature monitoring of critical server components according to claim 1, characterized in that, The specific process of the dynamic threshold model in step S2 adjusting the temperature thresholds of each important component in real time according to changes in server load and environment is as follows: S21: The dynamic threshold model extracts specified features from real-time temperature data, load data, and environmental data, maps the extracted specified features to specified intervals, and obtains standardized specified features. S22: The standardized specified features are processed in real time using the trained dynamic threshold model, including the neural network unit using the weights and biases optimized during training to perform nonlinear mapping on the input features and make preliminary predictions on the temperature threshold range of each component under the current load and environmental conditions, and the reinforcement learning unit adjusting the prediction results of the neural network in real time according to the reward function based on the current operating status of the server. S23: The dynamic threshold model generates real-time temperature warning thresholds for each important component based on real-time calculation results, and formats the threshold data into a format that the monitoring system can recognize, including component identification, current upper threshold value, lower threshold value, and threshold effective time.

5. The method for intelligent temperature monitoring of critical server components according to claim 1, characterized in that, The specific process of step S4 is as follows: S41: Perform horizontal coupling analysis in the same time dimension: For each time window, the platform extracts the temperature data of all components within that time window, combines the load data and environmental data to perform horizontal coupling analysis, calculates the correlation coefficient of the temperatures of different components using the Pearson coefficient, identifies component clusters with consistent temperature change trends, and analyzes the correlation between load data and temperature data. S42: Vertical Coupling Analysis at Different Time Nodes: Extract the temperature change curve of a component or component cluster within a continuous time window, and perform vertical coupling analysis with the temperature curves of the same period in history and similar load scenarios. By comparing the similarity of the curve shapes through the dynamic time warping algorithm, abnormal trends in temperature changes can be identified. At the same time, the derivative changes of the temperature curves are analyzed, and combined with the load fluctuations within the time window, it is determined whether the temperature change is caused by load changes. S43: Perform spatiotemporal cross-coupling verification: The platform will cross-verify the results of horizontal and vertical analysis.

6. The method for intelligent temperature monitoring of critical server components according to claim 1, characterized in that, Step S2, in the process of adjusting the temperature thresholds of each important component in real time, also includes a feedback process, as detailed below: Real-time recording of temperature changes, warning triggering status, and cooling system adjustment effects of each component is used to generate feedback data. The feedback data includes whether the component temperature returns to normal after the temperature threshold is adjusted, whether the warning is accurate, and whether the cooling energy consumption is within a reasonable range. Feedback data is transmitted to the reinforcement learning unit of the dynamic threshold model, which uses an incremental learning algorithm to fine-tune itself. This includes updating the reward function based on the feedback data when a large deviation in threshold prediction is found under a certain load scenario, and the neural network unit correcting the local weights through backpropagation. This allows the model to continuously adapt to new load patterns and environmental changes that occur during server operation, thereby continuously improving the accuracy and real-time performance of threshold adjustment.

7. A method for intelligent temperature monitoring of critical server components according to claim 5, characterized in that, The specific process of step S6 is as follows: S61: Perform structured parsing on the warning signal, extract key information including the component identifier that triggered the warning, the difference between the current temperature value and the threshold, the rate of temperature change, the status of related components, the current load data and environmental data, and form a standardized warning event data packet; S62: Based on the warning event data packet, classify the warning level according to the preset rules, including at least Level 1 warning, Level 2 warning and Level 3 warning, and call the preset basic control strategy library according to the warning level. Level 1 warning is matched with a powerful heat dissipation strategy. Level 2 warning system with enhanced heat dissipation strategy; A three-level early warning system is matched with a preventative heat dissipation strategy; S63: Call the coupling analysis results to dynamically optimize the temperature control strategy in the basic control strategy library; S64: During the strategy generation process, the optimal solution is calculated through a multi-objective optimization algorithm to balance the relationship between heat dissipation efficiency, energy consumption cost and hardware loss.

Citation Information

Patent Citations

  • IDC machine room environment intelligent monitoring method and monitoring system thereof

    CN119201605A

  • Risk early warning method and system and electronic equipment

    CN120162213A