Intelligent monitoring method for temperature of important components of server

By adjusting the temperature threshold in real time through dynamic threshold models and neural network units, and conducting component correlation analysis and multi-dimensional data coupling, the shortcomings of the existing server temperature monitoring system are solved, and rapid and precise control of server temperature and risk reduction are achieved.

CN120723583AActive Publication Date: 2025-09-30四川华鲲振宇智能科技有限责任公司

Patent Information

Application Number
CN202511188823.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-30
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

The existing server temperature monitoring system has defects in real-time performance, dynamic adaptability, comprehensiveness and intelligent regulation, and cannot meet the temperature management needs of modern data center servers under high-density and high-load operation. In particular, there are deficiencies in data collection and threshold setting, monitoring coverage, early warning and control strategies.

Method used

A dynamic threshold model is used to monitor the temperature of important server components. The temperature threshold is adjusted in real time through neural networks and reinforcement learning units. Component correlation analysis and multi-dimensional data coupling analysis are performed to generate an optimized temperature control strategy. The feedback mechanism is combined to optimize the model parameters.

Benefits of technology

It achieves rapid and precise control of server temperature, reduces the risk of overheating, takes into account both system stability and energy efficiency, accurately determines the root cause of temperature anomalies through spatiotemporal coupling analysis, and generates dynamically optimized control strategies to adapt to load and environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723583A_ABST
    Figure CN120723583A_ABST
Patent Text Reader

Abstract

The invention relates to a method for intelligently monitoring the temperature of important parts of a server, which comprises the following steps: acquiring specified data of the server, inputting the specified data into a dynamic threshold model for training, acquiring various items in real time, inputting the trained dynamic threshold model, and adjusting the temperature threshold of each important part in real time according to the load change and environment change of the server; when the real-time temperature of a certain part is greater than or equal to the corresponding temperature threshold value, correcting the temperature threshold value of the associated part in real time, and performing coupling analysis on the temperature data of different parts under the same time dimension and the temperature change curve of different time nodes; when the real-time temperature of a certain part is greater than or equal to the temperature threshold determined by the dynamic threshold model, an early warning signal is sent out; and generating a temperature control strategy based on the early warning signal and executing a corresponding action by an execution component. The server temperature condition is comprehensively mastered through a dynamic threshold value model, correlation analysis and coupling analysis, an optimized control strategy is generated according to the early warning level, and rapid and accurate regulation and control of the server temperature are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of server heat dissipation, and in particular relates to a method for intelligently monitoring the temperature of important server components. Background Art

[0002] With the rapid development of technologies such as cloud computing, big data, and artificial intelligence, servers in modern data centers are facing high-density, high-load operating demands. During prolonged, high-load operation, critical components within servers, such as the CPU, GPU, memory, and hard disk, generate significant heat. Excessive temperature rise can directly impact system performance, leading to reduced computing speeds, response delays, and, in severe cases, component damage and server downtime, posing significant risks to the stable operation of data centers. Therefore, accurate, real-time monitoring and control of the temperatures of critical server components has become crucial for ensuring data center reliability. Current server temperature monitoring systems have numerous limitations and are unable to meet the monitoring needs of complex, dynamic environments. First, regarding data collection and threshold setting, traditional systems often rely on timed sampling or fixed threshold alarm mechanisms. Timed sampling methods suffer from data lag and fail to reflect real-time server temperature changes caused by sudden increases in load or environmental changes, resulting in potential delays in system response to temperature anomalies. Fixed thresholds, on the other hand, cannot dynamically adjust to the server's real-time load or environmental conditions, leading to delayed warnings and false alarms, making them ineffective in addressing complex and changing workload scenarios. Secondly, existing systems often have limitations in monitoring coverage and correlation analysis capabilities. Some systems can only monitor a single component type, lacking comprehensive coverage of other critical components like GPUs, memory, and hard drives, making it difficult to understand the overall temperature status of the server. Even some systems that cover multiple components lack correlation analysis between them. When a component's temperature is abnormal, its potential impact on related components cannot be identified in a timely manner, resulting in incomplete early warnings and increasing the risk of system-wide overheating. Furthermore, in terms of early warning and control strategies, traditional systems have relatively simple early warning mechanisms, often triggering alarms based on a single threshold. Furthermore, these control strategies lack flexibility and optimization capabilities. After the early warning signal is issued, cooling system adjustments often adhere to a fixed pattern, failing to dynamically optimize based on component temperature trends, load conditions, and energy costs. This can lead to excessive energy consumption for heat dissipation or insufficient regulation to effectively reduce temperatures. Furthermore, regular sensor maintenance and calibration rely on manual labor, lacking an automated calibration mechanism. This can lead to data deviations after long-term use, compromising monitoring accuracy and reliability. In addition, the existing system has deficiencies in the data analysis dimension. It mostly stays at the temperature monitoring of a single time point or a single component, and fails to couple the temperature data of different components in the same time dimension with the temperature change curves at different time nodes for analysis. It is difficult to explore the inherent laws and potential risks of temperature changes, resulting in inaccurate judgment of the root causes of temperature anomalies, affecting the effectiveness of subsequent regulatory measures. In general, existing server temperature monitoring systems have significant shortcomings in terms of real-time performance, dynamic adaptability, comprehensiveness, and intelligent control. They are unable to meet the temperature management requirements of modern data center servers operating at high density and high load. Therefore, there is an urgent need for an intelligent monitoring method that can adjust temperature thresholds in real time, implement component correlation analysis and multi-dimensional data coupling analysis, and generate optimized control strategies based on early warning signals. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for intelligent temperature monitoring of important server components, which is used to adjust temperature thresholds in real time, realize component correlation analysis and multi-dimensional data coupling analysis, and generate optimized control strategies based on early warning signals.

[0004] In order to solve the above technical problems, the technical solutions adopted by the present invention are as follows: A method for intelligently monitoring the temperature of important server components includes the following steps: S1: Obtain historical server temperature data, load data, and environmental data and input them into the created dynamic threshold model for model training; S2: The data acquisition module collects temperature data, load data, and environmental data from each component in real time and inputs them into the trained dynamic threshold model. The dynamic threshold model adjusts the temperature thresholds of each important component in real time based on changes in server load and environment. S3: When the real-time temperature of a component is greater than or equal to the corresponding temperature threshold, the dynamic threshold model analyzes the temperature change trend of its associated components through correlation analysis and adjusts the temperature threshold of the associated components in real time; S4: The real-time temperature data, load data, and environmental data of each component are collected and transmitted to the monitoring platform. The monitoring platform couples and analyzes the temperature data of different components in the same time dimension with the temperature change curves at different time nodes. S5: When the real-time temperature of a component is greater than or equal to the temperature threshold determined by the dynamic threshold model, the monitoring system immediately issues an early warning signal; S6: Generate a temperature control strategy based on the early warning signal, generate a corresponding temperature control command based on the temperature control strategy and transmit it to the corresponding execution component, and the execution component performs the corresponding action.

[0005] Preferably, the specific process of step S1 is as follows: S11: extract historical temperature data, load data, and environmental data from the server database and perform data preprocessing on them; S12: extracting the pre-processed historical temperature data, load data and environmental data respectively and performing the specified features on the data; S13: Determine a label for the model based on the server's operating status and fault records. The label is a reasonable temperature threshold range for each component under each specified feature. When the actual temperature exceeds the range, it is considered an abnormal condition. S14: Initializing parameters of the dynamic threshold model, and inputting the extracted specified features and labels into the dynamic threshold model for model training.

[0006] Preferably, the specific process of inputting the extracted specified features and labels into the dynamic threshold model for model training in step S14 is as follows; S141: The dynamic threshold model continuously adjusts the weights and biases of the network based on the specified features and labels through the neural network unit to minimize the error between the predicted value and the label; S142: The dynamic threshold model uses a reinforcement learning unit to learn and optimize decision strategies based on a preset reward function and interaction with the environment, so that the model can better adapt to different load scenarios and environmental conditions; S143: The neural network unit and the reinforcement learning unit perform model training in an iterative manner, and evaluate the specified performance of the model through a test set after reaching a specified number of iterations or specified conditions. When the evaluation result meets the specified requirements, step S2 is executed.

[0007] Preferably, the specific process of the dynamic threshold model in step S2 adjusting the temperature threshold of each important component in real time according to the load change of the server and the environmental change is as follows: S21: The dynamic threshold model extracts specified features from the real-time temperature data, load data, and environmental data, and maps the extracted specified features to a specified interval to obtain standardized specified features; S22: performing real-time calculations on the standardized designated features using the trained dynamic threshold model, including the neural network unit using weights and biases optimized during the training process to perform nonlinear mapping on the input features, preliminarily predicting the temperature threshold range of each component under the current load and environmental conditions, and the reinforcement learning unit combining the current operating status of the server and adjusting the prediction results of the neural network in real time according to the reward function; S23: The dynamic threshold model generates real-time temperature warning thresholds for each important component based on the real-time calculation results, and formats the threshold data into a format recognizable by the monitoring system, including component identification, current threshold upper limit value, threshold lower limit value, and threshold effective time.

[0008] Preferably, the specific process of step S3 is as follows: S31: Receive the real-time temperature data of each component transmitted by the data acquisition module in real time and compare it with the temperature threshold of the corresponding component. When the real-time temperature of a component is greater than or equal to its current temperature threshold, the correlation analysis mechanism is immediately triggered, marking the component as a triggering component and recording its identification, the temperature value at the time of exceeding the threshold, the current load data and the key information of the environmental data as the initial basis for the correlation analysis; S32: Automatically identify all components that are directly or indirectly associated with the trigger component, i.e., associated components, according to a preset association relationship graph; S33: extracting temperature data, load data, and environmental data of associated components within a specified time window before and after the moment when the triggering component exceeds the threshold; S34: performing trend analysis on the extracted temperature data of the associated components to obtain trend characteristics including temperature change rate, temperature fluctuation frequency, and correlation between temperature and load; S35: Based on the extracted trend features, the temperature correlation strength between the associated component and the trigger component is calculated by mutual information entropy, where the correlation strength is represented by a value between 0 and 1; S36: Evaluate the impact of the overtemperature of the triggering component on each associated component based on the correlation strength and the difference between the current temperature of the associated component and its own threshold; S37: Dynamically generate a threshold correction coefficient based on the impact of the associated component; the correction coefficient is larger for the associated component with a high impact, and smaller for the associated component with a low impact; S38: Performing a modified threshold calculation: The model uses the formula "modified threshold = original threshold × correction coefficient - temperature trend compensation value" to calculate the new threshold of the associated component; The temperature trend compensation value is determined according to the temperature change rate of the associated components. If the temperature shows an accelerating upward trend, the compensation value is positive and the threshold is further reduced; if the temperature tends to be stable, the compensation value is 0; S39: Verify the calculated correction threshold to ensure it is not lower than the minimum safe temperature of the component and that the correction range is within a reasonable range. Once verified, the model formats the correction threshold into structured data that includes the component identifier, the new upper and lower threshold limits, the correction effective time, and the correction basis. S40: Update the warning threshold of the associated components and start real-time comparison. If the temperature of the associated components reaches or exceeds the new threshold after the revised threshold takes effect, the monitoring system will trigger cooling adjustment first according to the warning mechanism and push the associated warning information to the administrator, indicating the potential risk of the component being affected by the triggering component.

[0009] Preferably, the specific process of step S4 is as follows: S41: Performing lateral coupling analysis in the same time dimension: For each time window, the platform extracts temperature data for all components within that time window, combines it with load data and environmental data for lateral coupling analysis, calculates the correlation coefficient of different component temperatures using the Pearson coefficient, identifies component clusters with consistent temperature change trends, and analyzes the correlation between load data and temperature data; S42: Longitudinal Coupling Analysis at Different Time Nodes: This involves extracting the temperature variation curve of a component or component cluster within a continuous time window and performing a longitudinal coupling analysis with the temperature curves of the same period and similar load scenarios. Using a dynamic time warping algorithm, the curves are compared for similarity to identify abnormal trends in temperature variation. Furthermore, the derivative of the temperature curve is analyzed, combined with load fluctuations within the time window, to determine whether the temperature variation is caused by load changes. S43: Perform spatiotemporal cross-coupling verification: The platform cross-validates the results of horizontal and vertical analysis.

[0010] Preferably, the process of adjusting the temperature thresholds of the important components in real time in step S2 also includes a feedback process, which is specifically as follows: Real-time recording of temperature changes of each component, warning triggering conditions, and cooling system adjustment effects to generate feedback data, including whether component temperatures return to normal after temperature threshold adjustment, whether warnings are accurate, and whether cooling energy consumption is within a reasonable range; The reinforcement learning unit that transmits feedback data to the dynamic threshold model uses an incremental learning algorithm to fine-tune itself. When it finds that the threshold prediction deviation is large in a certain load scenario, the reinforcement learning unit updates the reward function according to the feedback data, and the neural network unit corrects the local weights through back propagation, so that the model can continuously adapt to new load patterns and environmental changes that occur during server operation, and continuously improve the accuracy and real-time performance of threshold adjustment.

[0011] Preferably, the specific process of step S6 is as follows: S61: Performing structured analysis on the warning signal to extract key information including the component identification that triggered the warning, the difference between the current temperature value and the threshold, the temperature change rate, the status of the associated components, the current load data, and the environmental data to form a standardized warning event data packet; S62: Based on the warning event data packet, the warning level is divided according to the preset rules, including at least level one warning, level two warning, and level three warning. The preset basic control strategy library is called according to the warning level. The level one warning matches the strong heat dissipation strategy; the level two warning matches the enhanced heat dissipation strategy; and the level three warning matches the preventive heat dissipation strategy. S63: Call the coupling analysis results to dynamically optimize the temperature control strategy in the basic control strategy library; S64: During the strategy generation process, the optimal solution is calculated through a multi-objective optimization algorithm to balance the relationship between heat dissipation efficiency, energy consumption cost, and hardware loss.

[0012] The beneficial effects of the present invention include: The present invention provides an intelligent temperature monitoring method for important server components. The method obtains server-specified data and inputs it into a dynamic threshold model for training. It collects various items in real time and inputs them into the trained dynamic threshold model. The temperature thresholds of important components are adjusted in real time according to changes in server load and environment. When the real-time temperature of a component is greater than or equal to the corresponding temperature threshold, the temperature threshold of the associated component is corrected in real time. The temperature data of different components in the same time dimension are coupled and analyzed with the temperature change curves at different time nodes. When the real-time temperature of a component is greater than or equal to the temperature threshold determined by the dynamic threshold model, an early warning signal is issued. A temperature control strategy is generated based on the early warning signal, and the execution component performs the corresponding action. The temperature threshold is accurately adjusted through the dynamic threshold model, the server temperature status is fully understood through correlation analysis and coupling analysis, and an optimized control strategy is generated according to the early warning level. This allows for rapid and precise regulation of the server temperature, effectively reducing the risk of overheating while taking into account both system stability and energy efficiency.

[0013] First, through spatiotemporal coupling analysis, the temperature data of different components in the same time dimension are coupled with the temperature change curves at different time nodes for analysis. Horizontal analysis is used to identify component clusters with consistent temperature change trends, while vertical analysis is used to compare the temperature curves of the same historical period and similar load scenarios. Combined with tools such as dynamic time warping algorithms, the inherent laws of temperature changes are explored, overcoming the limitations of traditional systems that monitor a single time point or a single component. This allows for more accurate determination of the root cause of temperature anomalies, providing a reliable basis for subsequent regulation.

[0014] Secondly, when generating a temperature control strategy based on the warning signal, the system matches the basic strategy according to the warning level, and performs dynamic optimization based on the coupling analysis results. At the same time, it balances heat dissipation efficiency, energy consumption cost and hardware loss through a multi-objective optimization algorithm.

[0015] Finally, a comprehensive feedback mechanism is established to transmit feedback data, including temperature changes, warning effectiveness, and cooling energy consumption, to the dynamic threshold model in real time. This allows for continuous optimization of model parameters through an incremental learning algorithm. Neural network units modify local weights, and reinforcement learning units update the reward function, enabling the model to continuously adapt to new load patterns and environmental changes, maintaining the accuracy of threshold adjustments over the long term. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 The figure is a flow chart of the method for intelligently monitoring the temperature of important server components according to the present invention.

[0017] Figure 2 Schematic diagram of the layout of temperature sensors in the server of the present invention. DETAILED DESCRIPTION

[0018] The following is combined with Figure 1~Figure 2 The present invention is described in further detail: Example 1 See attached Figure 1 As shown, a method for intelligently monitoring the temperature of important server components includes the following steps: S1: Obtain historical server temperature data, load data, and environmental data and input them into the created dynamic threshold model for model training.

[0019] The specific process of step S1 is as follows: S11: extract historical temperature data, load data, and environmental data from the server database and perform data preprocessing on them; S12: extracting the pre-processed historical temperature data, load data and environmental data respectively and performing the specified features on the data; S13: Determine a label for the model based on the server's operating status and fault records. The label is a reasonable temperature threshold range for each component under each specified feature. When the actual temperature exceeds the range, it is considered an abnormal condition. S14: Initializing parameters of the dynamic threshold model, and inputting the extracted specified features and labels into the dynamic threshold model for model training.

[0020] The specific process of inputting the extracted specified features and labels into the dynamic threshold model for model training in step S14 is as follows; S141: The dynamic threshold model continuously adjusts the weights and biases of the network based on the specified features and labels through the neural network unit to minimize the error between the predicted value and the label; S142: The dynamic threshold model uses a reinforcement learning unit to learn and optimize decision strategies based on a preset reward function and interaction with the environment, so that the model can better adapt to different load scenarios and environmental conditions; S143: The neural network unit and the reinforcement learning unit perform model training in an iterative manner, and evaluate the specified performance of the model through a test set after reaching a specified number of iterations or specified conditions. When the evaluation result meets the specified requirements, step S2 is executed.

[0021] S2: The data acquisition module collects the temperature data, load data and environmental data of each component in real time and inputs them into the trained dynamic threshold model. The data acquisition module collects temperature data by setting temperature sensors in each component. The layout of the temperature sensors is shown in Figure 2As shown, the temperature sensors include IBMC temperature sensor, pice temperature sensor, memory temperature sensor, CPU temperature sensor, storage temperature sensor, etc. The dynamic threshold model adjusts the temperature threshold of each important component in real time according to the load changes of the server and the environment changes; S3: When the real-time temperature of a component is greater than or equal to the corresponding temperature threshold, the dynamic threshold model analyzes the temperature change trend of its associated components through correlation analysis and adjusts the temperature threshold of the associated components in real time; S4: The real-time temperature data, load data, and environmental data of each component are collected and transmitted to the monitoring platform. The monitoring platform couples and analyzes the temperature data of different components in the same time dimension with the temperature change curves at different time nodes. S5: When the real-time temperature of a component is greater than or equal to the temperature threshold determined by the dynamic threshold model, the monitoring system immediately issues an early warning signal; S6: Generate a temperature control strategy based on the early warning signal, generate a corresponding temperature control command based on the temperature control strategy and transmit it to the corresponding execution component, and the execution component performs the corresponding action.

[0022] Example 2 Based on Example 1, the specific process of the dynamic threshold model in step S2 adjusting the temperature threshold of each important component in real time according to the load change of the server and the environmental change is as follows: S21: The dynamic threshold model extracts specified features from the real-time temperature data, load data, and environmental data, and maps the extracted specified features to a specified interval to obtain standardized specified features; S22: performing real-time calculations on the standardized designated features using the trained dynamic threshold model, including the neural network unit using weights and biases optimized during the training process to perform nonlinear mapping on the input features, preliminarily predicting the temperature threshold range of each component under the current load and environmental conditions, and the reinforcement learning unit combining the current operating status of the server and adjusting the prediction results of the neural network in real time according to the reward function; S23: The dynamic threshold model generates real-time temperature warning thresholds for each important component based on the real-time calculation results, and formats the threshold data into a format recognizable by the monitoring system, including component identification, current threshold upper limit value, threshold lower limit value, and threshold effective time.

[0023] The specific process of step S3 is as follows: S31: The dynamic threshold model receives the real-time temperature data of each component transmitted by the data acquisition module in real time and compares it with the temperature threshold of the corresponding component. When the real-time temperature of a component is greater than or equal to its current temperature threshold, the correlation analysis mechanism is immediately triggered, marking the component as a triggering component and recording its identification, the temperature value at the time of exceeding the threshold, the current load data, and key environmental data as the initial basis for correlation analysis. S32: The dynamic threshold model automatically identifies all components that are directly or indirectly associated with the trigger component, i.e., associated components, based on a preset association relationship map; S33: The dynamic threshold model extracts temperature data, load data, and environmental data of associated components within a specified time window before and after the triggering component exceeds the threshold. The time window must be set to balance data integrity and timeliness, including both pre-trigger temperature change trends and post-trigger real-time data to comprehensively analyze the temperature response characteristics of associated components. A secondary validity check is performed on the extracted data to eliminate invalid data and ensure the reliability of the analysis basis. S34: The dynamic threshold model performs trend analysis on the extracted temperature data of associated components, obtaining trend characteristics, including the temperature change rate (i.e., the temperature rise / fall amplitude per unit time), the temperature fluctuation frequency (i.e., the number of temperature oscillations within a short period of time), and the correlation between temperature and load. The correlation between temperature and load includes, for example, the degree of synchronous temperature changes when the associated component load increases. If the triggering component is the CPU, after it exceeds the threshold, the model will analyze whether the GPU temperature increases faster as the CPU load increases, or whether the memory temperature fluctuates abnormally due to the CPU occupying the heat dissipation space; S35: Based on the extracted trend features, the dynamic threshold model calculates the temperature correlation strength between the associated component and the triggering component using mutual information entropy. The correlation strength is expressed as a value between 0 and 1. The closer the value is to 1, the more synchronized the temperature change trends between the two components are, and the greater the likelihood of being affected by the overheating of the triggering component. Because the CPU and GPU share a cooling fan, their temperature correlation strength is generally higher than the correlation strength between the CPU and the remote hard drive. S36: The dynamic threshold model evaluates the impact of the triggering component overtemperature on each associated component based on the correlation strength and the difference between the current temperature of the associated component and its own threshold. If the associated component temperature changes rapidly, the correlation strength is high, and the current temperature is close to its own threshold, the impact is determined to be high. If the temperature changes slowly, the correlation is weak, and the current temperature is far below the threshold, the impact is determined to be low. S37: Dynamically generate threshold correction coefficients based on the impact of associated components. For highly impactful components, the correction coefficients are larger, meaning their thresholds need to be lowered more significantly to provide early warning of potential overheating risks. For less impactful components, the correction coefficients are smaller, requiring only minor adjustments to the thresholds. The generation of correction coefficients also relies on a historical case library. When a certain type of association repeatedly shows "overheating of a triggering component followed by overheating of associated components within a short period of time" in historical data, the corresponding correction coefficients are automatically increased to enhance the foresight of the warning. S38: Performing a modified threshold calculation: The model uses the formula "modified threshold = original threshold × correction coefficient - temperature trend compensation value" to calculate the new threshold of the associated component; The temperature trend compensation value is determined based on the temperature change rate of the associated component. If the temperature shows an accelerating upward trend, the compensation value is positive, further lowering the threshold. If the temperature tends to be stable, the compensation value is 0. For example, if the original threshold of a related component is 80°C, the correction coefficient is 0.8, the temperature rise rate is 2°C / second, and the compensation value is 5°C, then the corrected threshold value = 80×0.8-5=59°C, triggering an early warning by lowering the threshold. S39: The calculated correction threshold must be verified to ensure it is not lower than the component's minimum safe temperature. This ensures it is not lower than the lower limit of the component's electronic components' normal operating temperature, and that the correction amplitude is within a reasonable range, such as a single correction not exceeding 30% of the original threshold, to avoid false positives due to excessive corrections. After verification, the model formats the correction threshold into structured data that includes the component identifier, the new upper and lower threshold limits, the effective time of the correction, and the basis for the correction, such as the strength of the correlation and the degree of impact. S40: The revised threshold is pushed to the server temperature monitoring platform in real time via an encrypted channel. The platform immediately updates the warning thresholds for associated components and initiates a real-time comparison. If the temperature of an associated component reaches or exceeds the new threshold after the revised threshold takes effect, the monitoring system will prioritize cooling adjustments according to the intelligent warning mechanism. This can specifically increase the cooling fan speed of the associated component and push a related warning message to the administrator, indicating the potential risk of the component being affected by the triggering component.

[0024] Example 3 Based on Example 1 or Example 2, the specific process of step S4 is as follows: S41: Performs horizontal coupling analysis within the same time dimension: For each time window, the platform extracts temperature data for all components within that time window and combines it with load and environmental data for horizontal coupling analysis. The Pearson coefficient is used to analyze the temperature synchronization between the CPU and GPU, and the correlation coefficient between the temperatures of different components is calculated. This identifies component clusters with consistent temperature trends, including high-load core clusters consisting of the CPU, GPU, and memory. The platform also analyzes the correlation between load and temperature data, and uses a multivariate regression model to calculate the impact of a 10% increase in CPU utilization on the temperature of the CPU itself and associated components. Finally, the analysis results are modified based on environmental data. For example, if the ambient temperature suddenly rises by 5°C, the threshold for judging abnormal component temperatures is adjusted to eliminate environmental interference factors. S42: Vertical Coupling Analysis at Different Time Nodes: The platform extracts the temperature change curve of a component or component cluster within a continuous time window and performs a vertical coupling analysis with the temperature curves of the same period and similar load scenarios. Similar load scenarios can be within the same CPU usage range. A dynamic time warping algorithm is used to compare the similarity of the curve shapes to identify abnormal temperature trends. Simultaneously, the derivative of the temperature curve (i.e., the temperature change rate) is analyzed, combined with load fluctuations within the time window, to determine whether the temperature change is caused by load changes. A sudden increase in the temperature rate when the load is stable may be a precursor to hardware failure. S43: Perform spatiotemporal cross-coupling verification: The platform cross-validates the results of the horizontal and vertical analysis. When the horizontal analysis finds that the CPU and memory temperatures rise synchronously within a certain time window, a vertical analysis is performed to check whether this phenomenon has appeared in historical data. If it only occurs in scenarios with high ambient temperature and sudden load increase, it is judged as a normal load response; if it occurs for the first time under low load and normal environment, it is marked as an anomaly, triggering further investigation. At the same time, combined with the spatial position of the components, the impact of physical distance on temperature coupling is analyzed. Because the temperature coupling of components that are closer is significantly higher than that of components that are farther away, misjudgments of non-physical correlations can be ruled out, such as "the synchronous changes in hard disk and CPU temperatures may be due to load linkage rather than heat conduction."

[0025] The process of adjusting the temperature thresholds of each important component in real time in step S2 also includes a feedback process, which is as follows: Real-time recording of temperature changes of each component, warning triggering conditions, and cooling system adjustment effects to generate feedback data, including whether component temperatures return to normal after temperature threshold adjustment, whether warnings are accurate, and whether cooling energy consumption is within a reasonable range; The feedback data is transmitted to the reinforcement learning unit of the dynamic threshold model. The model uses an incremental learning algorithm to fine-tune itself. When it is found that the threshold prediction deviation is large under a certain load scenario, the reinforcement learning unit updates the reward function according to the feedback data, and the neural network unit corrects the local weights through back propagation, so that the model can continuously adapt to new load patterns and environmental changes that occur during the operation of the server, and continuously improve the accuracy and real-time performance of the threshold adjustment.

[0026] The specific process of step S6 is as follows: S61: The monitoring platform performs a structured analysis on the warning signal, extracting key information including the component identification that triggered the warning, the difference between the current temperature value and the threshold, the temperature change rate, the status of the associated components, the current load data, and the environmental data, to form a standardized warning event data package; S62: Classify the warning level according to the warning event data packet according to the preset rules; Level 1 alerts apply to scenarios where a single component temperature exceeds the threshold by more than 5°C, or the temperature change rate is 3°C / second or more, or multiple components exceed the threshold simultaneously. For example, "CPU01 temperature exceeds the threshold by 85°C, exceeding the threshold by 78°C, and GPU03 temperature exceeds the threshold by 80°C, exceeding the threshold by 75°C." Level 2 alerts apply to scenarios where a single component temperature exceeds the threshold by 1-5°C, or the temperature changes at a rate of 1-3°C / second. For example, "HDD05 temperature exceeds the threshold by 62°C, exceeding the threshold by 59°C, at a rate of 1.5°C / second." Level 3 alerts apply to scenarios where component temperatures approach the threshold but the change trend is stable, for example, "Memory02 temperature approaches the threshold by 70°C, approaching the threshold by 71°C, at a rate of 0.3°C / second." Different levels correspond to different response priorities and control strengths.

[0027] S62: The preset basic control strategy library is called according to the warning level. The first-level warning matches the "strong cooling strategy", which includes the combined measures of "full fan speed + full liquid cooling system load" and "temporary reduction of related component loads". The second-level warning matches the "enhanced cooling strategy", which can "increase the fan speed of the target component to 80% and increase the liquid cooling flow by 20%". The third-level warning matches the "preventive cooling strategy", which can "maintain the fan speed of the target component at 60% and dynamically monitor it". S63: Invokes the coupling analysis results to dynamically optimize the temperature control strategies in the basic control strategy library. This includes, when the horizontal coupling analysis shows that "CPU01 and GPU03 are in the same cooling duct," increasing the speed of all fans in the duct simultaneously during a level 1 warning to avoid local cooling blind spots. If the vertical coupling analysis finds that simply increasing heat dissipation under similar load conditions during the same period in history will cause the CPU temperature to rebound after 30 minutes, a predictive measure to trigger load balancing after 15 minutes will be added to the policy. When the spatial analysis shows that the target component is located in a heat dissipation blind spot, the backup heat dissipation device in that area will be activated first; S64: During the strategy generation process, the optimal solution is calculated through a multi-objective optimization algorithm to balance the relationship between heat dissipation efficiency, energy consumption cost, and hardware loss.

[0028] In summary, the present invention provides a method for intelligent temperature monitoring of important server components. The method obtains server-specified data and inputs it into a dynamic threshold model for training. It collects various items in real time and inputs them into the dynamic threshold model. The temperature thresholds of important components are adjusted in real time according to changes in the server's load and environment. The method corrects the temperature thresholds of related components in real time, and couples and analyzes the temperature data of different components in the same time dimension with the temperature change curves at different time nodes. The method generates a temperature control strategy based on early warning signals and executes the corresponding actions of the components. The temperature threshold is precisely adjusted through the dynamic threshold model, and the server temperature status is fully understood through correlation analysis and coupling analysis. An optimized control strategy is generated based on the early warning level.

[0029] Through spatiotemporal coupling analysis, the temperature data of different components within the same time dimension are coupled with temperature change curves at different time points for analysis. This overcomes the limitations of traditional systems that monitor a single time point or component, enabling more accurate identification of the root cause of temperature anomalies and providing a reliable basis for subsequent control. When generating a temperature control strategy based on the warning signal, the system matches the basic strategy according to the warning level and dynamically optimizes it based on the results of the coupling analysis. A multi-objective optimization algorithm simultaneously balances cooling efficiency, energy costs, and hardware losses. A comprehensive feedback mechanism is implemented to transmit feedback data, including temperature changes, warning effects, and cooling energy consumption, to the dynamic threshold model in real time. Model parameters are continuously optimized using an incremental learning algorithm. Neural network units modify local weights, and reinforcement learning units update the reward function, enabling the model to continuously adapt to new load patterns and environmental changes, maintaining the accuracy of threshold adjustments over the long term.

Claims

1. A method for intelligently monitoring the temperature of important server components, characterized in that: The following steps are involved: S1: Obtain historical server temperature data, load data, and environmental data and input them into the created dynamic threshold model for model training; S2: The data acquisition module collects temperature data, load data, and environmental data from each component in real time and inputs them into the trained dynamic threshold model. The dynamic threshold model adjusts the temperature thresholds of each important component in real time based on changes in server load and environment. S3: When the real-time temperature of a component is greater than or equal to the corresponding temperature threshold, the dynamic threshold model analyzes the temperature change trend of its associated components through correlation analysis and adjusts the temperature threshold of the associated components in real time; S4: The real-time temperature data, load data, and environmental data of each component are collected and transmitted to the monitoring platform. The monitoring platform couples and analyzes the temperature data of different components in the same time dimension with the temperature change curves at different time nodes. S5: When the real-time temperature of a component is greater than or equal to the temperature threshold determined by the dynamic threshold model, the monitoring system immediately issues an early warning signal; S6: Generate a temperature control strategy based on the early warning signal, generate a corresponding temperature control command based on the temperature control strategy and transmit it to the corresponding execution component, and the execution component performs the corresponding action.

2. The method for intelligently monitoring the temperature of important server components according to claim 1, characterized in that: The specific process of step S1 is as follows: S11: extract historical temperature data, load data, and environmental data from the server database and perform data preprocessing on them; S12: extracting the pre-processed historical temperature data, load data and environmental data respectively and performing the specified features on the data; S13: Determine a label for the model based on the server's operating status and fault records. The label is a reasonable temperature threshold range for each component under each specified feature. When the actual temperature exceeds the range, it is considered an abnormal condition. S14: Initializing parameters of the dynamic threshold model, and inputting the extracted specified features and labels into the dynamic threshold model for model training.

3. A method for intelligently monitoring the temperature of important server components according to claim 2, characterized in that: The specific process of inputting the extracted specified features and labels into the dynamic threshold model for model training in step S14 is as follows; S141: The dynamic threshold model continuously adjusts the weights and biases of the network based on the specified features and labels through the neural network unit to minimize the error between the predicted value and the label; S142: The dynamic threshold model uses a reinforcement learning unit to learn and optimize decision strategies based on a preset reward function and interaction with the environment, so that the model can better adapt to different load scenarios and environmental conditions; S143: The neural network unit and the reinforcement learning unit perform model training in an iterative manner, and evaluate the specified performance of the model through a test set after reaching a specified number of iterations or specified conditions. When the evaluation result meets the specified requirements, step S2 is executed.

4. The method for intelligently monitoring the temperature of important server components according to claim 1, wherein: The specific process of the dynamic threshold model in step S2 adjusting the temperature thresholds of important components in real time according to the load changes and environmental changes of the server is as follows: S21: The dynamic threshold model extracts specified features from the real-time temperature data, load data, and environmental data, and maps the extracted specified features to a specified interval to obtain standardized specified features; S22: performing real-time calculations on the standardized designated features using the trained dynamic threshold model, including the neural network unit using weights and biases optimized during the training process to perform nonlinear mapping on the input features, preliminarily predicting the temperature threshold range of each component under the current load and environmental conditions, and the reinforcement learning unit combining the current operating status of the server and adjusting the prediction results of the neural network in real time according to the reward function; S23: The dynamic threshold model generates real-time temperature warning thresholds for each important component based on the real-time calculation results, and formats the threshold data into a format recognizable by the monitoring system, including component identification, current threshold upper limit value, threshold lower limit value, and threshold effective time.

5. The method for intelligently monitoring the temperature of important server components according to claim 1, characterized in that: The specific process of step S3 is as follows: S31: Receive the real-time temperature data of each component transmitted by the data acquisition module in real time and compare it with the temperature threshold of the corresponding component. When the real-time temperature of a component is greater than or equal to its current temperature threshold, the correlation analysis mechanism is immediately triggered, marking the component as a triggering component and recording its identification, the temperature value at the time of exceeding the threshold, the current load data and the key information of the environmental data as the initial basis for the correlation analysis; S32: Automatically identify all components that are directly or indirectly associated with the trigger component, i.e., associated components, according to a preset association relationship graph; S33: extracting temperature data, load data, and environmental data of associated components within a specified time window before and after the moment when the triggering component exceeds the threshold; S34: performing trend analysis on the extracted temperature data of the associated components to obtain trend characteristics including temperature change rate, temperature fluctuation frequency, and correlation between temperature and load; S35: Based on the extracted trend features, the temperature correlation strength between the associated component and the trigger component is calculated by mutual information entropy, where the correlation strength is represented by a value between 0 and 1; S36: Evaluate the impact of the overtemperature of the triggering component on each associated component based on the correlation strength and the difference between the current temperature of the associated component and its own threshold; S37: Dynamically generate a threshold correction coefficient based on the impact of the associated component; the correction coefficient is larger for the associated component with a high impact, and smaller for the associated component with a low impact; S38: Performing a modified threshold calculation: The model uses the formula "modified threshold = original threshold × correction coefficient - temperature trend compensation value" to calculate the new threshold of the associated component; The temperature trend compensation value is determined according to the temperature change rate of the associated components. If the temperature shows an accelerating upward trend, the compensation value is positive and the threshold is further reduced; if the temperature tends to be stable, the compensation value is 0; S39: Verify the calculated correction threshold to ensure it is not lower than the minimum safe temperature of the component and that the correction range is within a reasonable range. Once verified, the model formats the correction threshold into structured data that includes the component identifier, the new upper and lower threshold limits, the correction effective time, and the correction basis. S40: Update the warning threshold of the associated components and start real-time comparison. If the temperature of the associated components reaches or exceeds the new threshold after the revised threshold takes effect, the monitoring system will trigger cooling adjustment first according to the warning mechanism and push the associated warning information to the administrator, indicating the potential risk of the component being affected by the triggering component.

6. The method for intelligently monitoring the temperature of important server components according to claim 1, characterized in that: The specific process of step S4 is as follows: S41: Performing lateral coupling analysis in the same time dimension: For each time window, the platform extracts temperature data for all components within that time window, combines it with load data and environmental data for lateral coupling analysis, calculates the correlation coefficient of different component temperatures using the Pearson coefficient, identifies component clusters with consistent temperature change trends, and analyzes the correlation between load data and temperature data; S42: Longitudinal Coupling Analysis at Different Time Nodes: This involves extracting the temperature variation curve of a component or component cluster within a continuous time window and performing a longitudinal coupling analysis with the temperature curves of the same period and similar load scenarios. Using a dynamic time warping algorithm, the curves are compared for similarity to identify abnormal trends in temperature variation. Furthermore, the derivative of the temperature curve is analyzed, combined with load fluctuations within the time window, to determine whether the temperature variation is caused by load changes. S43: Perform spatiotemporal cross-coupling verification: The platform cross-validates the results of horizontal and vertical analysis.

7. The method for intelligently monitoring the temperature of important server components according to claim 1, characterized in that: The process of adjusting the temperature thresholds of each important component in real time in step S2 also includes a feedback process, which is as follows: Real-time recording of temperature changes of each component, warning triggering conditions, and cooling system adjustment effects to generate feedback data, including whether component temperatures return to normal after temperature threshold adjustment, whether warnings are accurate, and whether cooling energy consumption is within a reasonable range; The reinforcement learning unit that transmits feedback data to the dynamic threshold model uses an incremental learning algorithm to fine-tune itself. When it finds that the threshold prediction deviation is large in a certain load scenario, the reinforcement learning unit updates the reward function according to the feedback data, and the neural network unit corrects the local weights through back propagation, so that the model can continuously adapt to new load patterns and environmental changes that occur during server operation, and continuously improve the accuracy and real-time performance of threshold adjustment.

8. The method for intelligently monitoring the temperature of important server components according to claim 6, characterized in that: The specific process of step S6 is as follows: S61: Performing structured analysis on the warning signal to extract key information including the component identification that triggered the warning, the difference between the current temperature value and the threshold, the temperature change rate, the status of the associated components, the current load data, and the environmental data to form a standardized warning event data packet; S62: Based on the warning event data packet, the warning level is divided according to the preset rules, including at least level 1 warning, level 2 warning, and level 3 warning. The preset basic control strategy library is called according to the warning level, and the level 1 warning is matched with the strong heat dissipation strategy; Secondary warning matching enhanced cooling strategy; Three-level warning matches preventive cooling strategy; S63: Call the coupling analysis results to dynamically optimize the temperature control strategy in the basic control strategy library; S64: During the strategy generation process, the optimal solution is calculated through a multi-objective optimization algorithm to balance the relationship between heat dissipation efficiency, energy consumption cost, and hardware loss.

Citation Information

Patent Citations

  • Information processing method, temperature prediction model training method and device and electronic equipment

    CN114491943A

  • IDC machine room environment intelligent monitoring method and monitoring system thereof

    CN119201605A

  • Temperature collecting, monitoring and alarming method for palm fire fighting

    CN119723856A

  • Risk early warning method and system and electronic equipment

    CN120162213A

  • Intelligent power distribution network line state real-time monitoring, analyzing and evaluating system

    CN120342080A

Cited By

  • Prefabricated solder shell water cooling flow intelligent control method and system and medium

    CN121277260A

  • A prefabricated solder shell water cooling flow intelligent control method, system and medium

    CN121277260B

  • Adjusting device, adjusting method and server board card

    CN121578868A

  • A regulating device, a regulating method and a server board

    CN121578868B

  • Electric control cabinet temperature control method based on simulation

    CN121596937A