Server heat dissipation control method, electronic equipment and storage medium
Through the coordination of deep reinforcement learning-optimized heat dissipation control algorithm and edge computing, the fan speed and temperature threshold are dynamically adjusted, which solves the problem of difficulty in parameter setting and poor adaptability in server heat dissipation control, and improves the efficiency and stability of the heat dissipation system.
Patent Information
- Application Number
- CN202510536003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The server cooling control faces the problems of parameter setting and poor system adaptability, which leads to poor heat dissipation effect and may even cause overheating problems.
The deep reinforcement learning mechanism is used to optimize the heat dissipation control algorithm, combined with edge computing and multi-algorithm coordination, dynamically adjust the fan speed and temperature threshold, and evaluate and optimize based on real-time operation data and historical data.
It improves the efficiency and stability of the cooling system, solves the problems of difficulty in parameter setting and poor adaptability, and ensures the normal operation of the server.
Smart Images

Figure CN120066922A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method for server heat dissipation control, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of information technology, servers, as the core devices for data processing and storage, their performance and stability are crucial for business operations. During the operation of servers, the heat dissipation problem has always been one of the key factors affecting server performance and stability. Most traditional heat dissipation control methods are based on fixed heat dissipation strategies. For example, a fixed fan speed is set or the fan speed is adjusted according to a preset temperature threshold. However, fixed heat dissipation strategies often cannot meet the heat dissipation requirements of servers under different loads and environments, resulting in poor heat dissipation effects and even possible overheating problems, affecting the normal operation of servers.
[0003] To address the problem of poor heat dissipation effect, prior art provides a heat dissipation control method based on the Proportional-Integral-Derivative (PID) control algorithm. Although the heat dissipation control method based on the PID control algorithm can solve the problem, most of the existing heat dissipation control methods based on the PID self-control algorithm rely on traditional machine learning algorithms and the algorithms are relatively single, with low optimization ability under complex working conditions. Therefore, how to solve the problems of difficult parameter tuning and poor system adaptability faced in server heat dissipation control is an urgent problem to be solved currently. Summary of the Invention
[0004] This application provides a method for server heat dissipation control, an electronic device, and a storage medium, to at least solve the problems of difficult parameter tuning and poor system adaptability faced in server heat dissipation control.
[0005] This application provides a method for server heat dissipation control, including: Inputting the real-time operation data of the first server and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server, so as to obtain the target fan speed, where the dynamic temperature threshold is a temperature threshold obtained by analyzing the historical operation data of the first server; Performing fan speed adjustment processing on the first server according to the target fan speed to obtain the adjusted first server and the second operation data of the adjusted first server, and performing data evaluation and analysis processing on the second operation data to obtain the heat dissipation evaluation result corresponding to the adjusted first server; In the case where the heat dissipation evaluation result does not meet the preset heat dissipation conditions, based on the second operation data, the heat dissipation evaluation result, and the historical operation data, the first heat dissipation control algorithm is optimized based on the deep reinforcement learning mechanism to obtain the second heat dissipation control algorithm; Based on the second heat dissipation control algorithm, the second operation data, and the dynamic temperature threshold, heat dissipation control processing is performed on the adjusted first server, and the second heat dissipation control algorithm is transmitted to the preset edge computing device.
[0006] This application also provides a device for server heat dissipation control, including: An input unit, configured to input the real-time operation data of the first server and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server, and obtain the target fan speed, where the dynamic temperature threshold is a temperature threshold obtained by analyzing the historical operation data of the first server; An adjustment unit, configured to perform fan speed adjustment processing on the first server according to the target fan speed to obtain the adjusted first server and the second operation data of the adjusted first server; An analysis unit, configured to perform data evaluation and analysis processing on the second operation data to obtain the heat dissipation evaluation result corresponding to the adjusted first server; An optimization unit, configured to, in the case where the heat dissipation evaluation result does not meet the preset heat dissipation conditions, optimize the first heat dissipation control algorithm based on the second operation data, the heat dissipation evaluation result, and the historical operation data based on the deep reinforcement learning mechanism to obtain the second heat dissipation control algorithm; A control unit, configured to perform heat dissipation control processing on the adjusted first server based on the second heat dissipation control algorithm, the second operation data, and the dynamic temperature threshold; A transmission unit, configured to transmit the second heat dissipation control algorithm to the preset edge computing device.
[0007] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above server heat dissipation control methods when executing the computer program.
[0008] This application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program implements the steps of any of the above server heat dissipation control methods when executed by a processor.
[0009] This application also provides a computer program product, including a computer program, where the computer program implements the steps of any of the above server heat dissipation control methods when executed by a processor.
[0010] The method, electronic device, and storage medium for server heat dissipation control of the present application first perform heat dissipation control processing on the server according to the operating data of the server and the heat dissipation control algorithm. Then, according to the actual heat dissipation requirements of the server, i.e., the preset heat dissipation conditions, the server after heat dissipation control is evaluated. The heat dissipation control algorithm is optimized according to the evaluation results and the operating data of the server after heat dissipation control. At the same time, the optimization of the heat dissipation control algorithm introduces a deep reinforcement learning mechanism, which combines multiple algorithms and combines with edge computing, i.e., the preset edge computing device, to solve the balance problem of difficult parameter tuning and poor system adaptability faced in traditional heat dissipation control. Therefore, the problems of difficult parameter tuning and poor system adaptability faced in server heat dissipation control can be solved, and the technical effects of improving the efficiency and stability of the heat dissipation system can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 It is a flowchart showing the method for server heat dissipation control provided by an embodiment of the present application; Figure 2 It is a flowchart showing the determination process of a dynamic temperature threshold provided by an embodiment of the present application; Figure 3 It is an analysis flowchart of the heat dissipation evaluation results provided by an embodiment of the present application; Figure 4 It is a schematic structural diagram of a server heat dissipation control system provided by an embodiment of the present application; Figure 5 It is a flowchart showing another method for server heat dissipation control provided by an embodiment of the present application; Figure 6 It is a schematic structural diagram of a server heat dissipation control device provided by an embodiment of the present application; Figure 7 It is a schematic structural diagram of another server heat dissipation control device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0014] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0015] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific implementation manners.
[0016] Figure 1 The following is a schematic flowchart of a method for server heat dissipation control provided for an embodiment of this application. In combination with the execution process of the method for server heat dissipation control, the method will be described in detail.
[0017] As Figure 1 shown, the method for server heat dissipation control includes: Step 101, input the real-time operation data and dynamic temperature threshold of the first server into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server, and obtain the target fan speed, where the dynamic temperature threshold is the temperature threshold obtained by analyzing the historical operation data of the first server.
[0018] In the embodiment of this application, the first server is a server that needs heat dissipation control, such as: enterprise-level server, super server, etc. Specifically, the type, category, structure, etc. of the first server are not limited in this application.
[0019] The real-time operation data is the operation data obtained by collecting data from the first server through the Baseboard Management Controller (BMC) of the first server. The real-time operation data includes, but is not limited to: hardware data and status data, etc. The acquisition of the real-time operation data includes, but is not limited to: through high-precision temperature sensors, fan speed sensors, load sensors, and ambient temperature sensors connected to the main board of the first server, real-time collection of temperature data, fan speed data, load data, and ambient temperature data of the server hardware.
[0020] The real-time operation data includes at least: temperature data, fan speed data, load data, and ambient temperature data. The temperature data includes, but is not limited to, the temperatures of key hardware such as the Central Processing Unit (CPU), memory, and hard disk. The fan speed data includes, but is not limited to, the speeds of each fan inside the first server. The temperature data accurately reflects the real-time temperature status of the key hardware of the first server. The fan speed data monitors the operating status of the cooling fans of the first server. The load data reflects the working intensity of the first server, and the ambient temperature data provides external environment information.
[0021] Among them, the dynamic temperature threshold is not a fixed value, but an adaptive threshold generated by analyzing the historical operation data of the first server (such as: the temperature fluctuation pattern in the past 24 hours, the peak load period, the ambient temperature change trend). For example: in a scenario where the load of the first server shows periodic fluctuations, the temperature change trend of the first server in the future period is predicted through a machine learning algorithm (such as: the Long Short-Term Memory (LSTM) time series model), and then the temperature threshold is dynamically adjusted to appropriately relax the threshold (increase the temperature threshold) in the low-load stage to reduce fan energy consumption, and tighten the threshold (decrease the temperature threshold) in the high-load stage to prevent overheating.
[0022] The first cooling control algorithm is a self-defined cooling control algorithm, such as: an algorithm based on the PID control framework, etc. For the convenience of understanding, the first cooling control algorithm will be described below by taking the algorithm based on the PID control framework as an example.
[0023] The parameters of the first cooling control algorithm (such as: proportional coefficient P, integral coefficient I, derivative coefficient D, etc.) are set by experience or obtained through pre-training with historical data. When calculating the target fan speed, the algorithm inputs the deviation between the temperature data in the real-time operation data and the dynamic temperature threshold, the deviation change rate (differential term), and the cumulative historical deviation (integral term) into the PID formula, and combines the influence weights of the load and ambient temperature to comprehensively calculate the optimal fan speed that meets the current cooling requirements. In this process, the first cooling control algorithm not only considers immediate temperature adjustment, but also introduces a prediction of future working conditions through the dynamic threshold to avoid the lag response caused by traditional fixed thresholds.
[0024] Input the collected multi-source data (real-time operation data) into the PID control algorithm based on deep reinforcement learning (the first heat dissipation control algorithm). At the same time, the first heat dissipation control algorithm includes a deep neural network with multiple hidden layers as the policy network, which takes the state space (i.e., multi-source data) as the input and outputs the optimal PID parameter adjustment value (the adjustment value of the first heat dissipation control algorithm) in the current state. According to the adjusted PID parameters, combined with the deviation, deviation change rate, and integral value between the current temperature (temperature data in real-time operation data) and the target temperature (dynamic temperature threshold), calculate the target fan speed. During the training process of the deep reinforcement learning algorithm, the policy network is continuously optimized according to the feedback of the reward function, and gradually learns the optimal control strategy.
[0025] Furthermore, it should be noted that regardless of whether the first server is overheated, the target fan speed will be calculated through the first heat dissipation control algorithm to determine the optimal fan speed for the first server to meet the current heat dissipation requirements.
[0026] Step 102: Adjust the fan speed of the first server according to the target fan speed to obtain the adjusted first server and the second operation data of the adjusted first server, and perform data evaluation and analysis on the second operation data to obtain the heat dissipation evaluation result corresponding to the adjusted first server.
[0027] In the embodiments of the present application, when performing the fan speed adjustment process, it can be but is not limited to: sending the target fan speed to the fan controller of the first server to adjust the fan speed, etc.
[0028] After adjusting the fan speed of the first server according to the target fan speed, the second operation data includes but is not limited to: the adjusted temperature change curve, the actual fan speed, the instantaneous energy consumption, and the load status, etc. The data evaluation and analysis of the second operation data can be completed through but is not limited to: a preset evaluation model. Specifically, the present application does not limit the preset evaluation model.
[0029] When performing data evaluation and analysis on the second operation data through the preset evaluation model, the heat dissipation evaluation result can be generated from but is not limited to the following dimensions: Temperature stability: Whether the temperature of key components quickly converges within the dynamic threshold range and whether the fluctuation amplitude is lower than the allowable upper limit; Energy consumption efficiency: Whether the fan speed adjustment minimizes power consumption while meeting the heat dissipation requirements, avoiding frequent start-stop or continuous high-speed operation; Response timeliness: Whether the time delay from temperature overrun to the effective fan speed adjustment is within the tolerance range of the system; Environmental adaptability: Whether the heat dissipation strategy can effectively offset the impact brought by sudden changes in environmental temperature (such as: failure of the computer room air conditioner).
[0030] The heat dissipation evaluation result can be presented in the form of a quantitative score or a classification label (such as excellent, medium, poor). Specifically, the present application does not limit the presentation form of the heat dissipation evaluation result.
[0031] Furthermore, during the heat dissipation control process, the BMC of the first server continuously collects the temperature data, fan speed data, load data, and ambient temperature data after heat dissipation control (the adjusted first server), combines these data to obtain the second operation data, and stores it in the local training data set. As time goes by, the training data set is continuously updated, containing more data under different working conditions, providing rich data support for the continuous optimization of the deep reinforcement learning algorithm.
[0032] Step 103, when the heat dissipation evaluation result does not meet the preset heat dissipation conditions, based on the second operation data, the heat dissipation evaluation result, and the historical operation data, the first heat dissipation control algorithm is optimized based on the deep reinforcement learning mechanism to obtain the second heat dissipation control algorithm.
[0033] In the embodiment of the present application, the preset heat dissipation conditions are custom-set heat dissipation conditions, such as: temperature safety range, maximum allowable energy consumption, response time threshold, etc. Specifically, the present application does not limit the preset heat dissipation conditions.
[0034] The heat dissipation evaluation result can be compared with the preset heat dissipation conditions (such as: temperature safety range, maximum allowable energy consumption, response time threshold). For example: if the CPU temperature of the adjusted first server still continuously exceeds the dynamic temperature threshold, or the fan speed of the first server fluctuates frequently and significantly, resulting in the energy consumption exceeding the maximum allowable energy consumption, the heat dissipation evaluation result will be marked as not meeting the preset heat dissipation conditions.
[0035] When the heat dissipation evaluation result is marked as not meeting the preset heat dissipation conditions, it is necessary to optimize the PID parameters (parameters of the first heat dissipation control algorithm) through deep reinforcement learning. The BMC regularly trains the deep reinforcement learning algorithm with the latest training samples. During the training process, the algorithm continuously adjusts the parameters of the policy network through interaction with the environment (i.e., the heat dissipation control process) to maximize the cumulative reward. When the algorithm converges or reaches the preset number of training times, the optimized first heat dissipation control algorithm is obtained for subsequent heat dissipation control.
[0036] Furthermore, when the heat dissipation assessment does not meet the standard, the algorithm optimization process based on deep reinforcement learning (DRL) will be triggered. At this time, the second operation data and historical data together constitute the training sample set, which is input into the DRL algorithm for policy iteration. The reward function of the DRL algorithm is designed as a multi-objective optimization form: positive rewards: temperature returns to the safe range, energy consumption is reduced, and response speed is improved; negative penalties: temperature continues to exceed the standard, fan speed oscillates, and energy consumption exceeds the limit.
[0037] The DRL algorithm uses the exploration-utilization mechanism to simulate the effects of different combinations of parameters of the first heat dissipation control algorithm on the heat dissipation effect in a simulation environment, and gradually learns how to dynamically adjust the parameters of the first heat dissipation control algorithm according to real-time conditions (such as sudden increase in load and sudden rise in ambient temperature). For example, in scenarios where the load fluctuates drastically, it tends to increase the differential coefficient to quickly suppress temperature changes, while reducing the integral coefficient to prevent overshoot. The optimized second heat dissipation control algorithm not only corrects the defects of the original algorithm, but may also introduce new control logic (such as load prediction feedforward compensation), thereby improving robustness under complex working conditions.
[0038] Step 104: Perform heat dissipation control processing on the adjusted first server based on the second heat dissipation control algorithm, the second operating data, and the dynamic temperature threshold, and transmit the second heat dissipation control algorithm to the preset edge computing device.
[0039] In the embodiment of the present application, the optimized algorithm, i.e., the second heat dissipation control algorithm, is first applied to the adjusted first server to verify the effectiveness of the second heat dissipation control algorithm in the actual environment. At the same time, edge collaboration and algorithm deployment are performed, specifically, but not limited to the following methods: the second heat dissipation control algorithm is transmitted to a preset edge computing device through an encrypted channel, and the preset edge computing device is a custom storage device, such as an edge server deployed locally in the computer room, a BMC built-in computing module, an external cloud server device, etc.
[0040] The preset edge computing device stores and updates the second heat dissipation control algorithm, so that it can quickly respond to similar heat dissipation problems. For cross-server collaborative heat dissipation requirements (such as multiple servers sharing a heat dissipation duct), the preset edge computing device can exchange heat dissipation strategies and real-time status with other nodes based on blockchain technology to achieve distributed decision-making. For example, when the preset edge computing device detects a surge in the load of an adjacent server, it can fine-tune the fan speed of this server in advance to avoid local overheating caused by hot air reflux.
[0041] The method, electronic device, and storage medium for server heat dissipation control in this application first perform heat dissipation control processing on the server according to the operating data of the server and the heat dissipation control algorithm. Then, according to the actual heat dissipation requirements of the server, that is, the preset heat dissipation conditions, the server after heat dissipation control is evaluated. The heat dissipation control algorithm is optimized according to the evaluation results and the operating data of the server after heat dissipation control. At the same time, the optimization of the heat dissipation control algorithm introduces a deep reinforcement learning mechanism, which combines multiple algorithms and combines with edge computing, that is, the preset edge computing device, to solve the balance problem of difficult parameter tuning and poor system adaptability faced in traditional heat dissipation control. Therefore, the problems of difficult parameter tuning and poor system adaptability faced in server heat dissipation control can be solved, and the technical effects of improving the efficiency and stability of the heat dissipation system can be achieved.
[0042] In an implementable manner of the embodiment of this application, before performing the fan speed calculation process to obtain the target fan speed, it is necessary to first determine the size of the dynamic temperature threshold to facilitate the calculation of the target fan speed. Specifically, regarding the determination of the dynamic temperature threshold, the embodiment of this application provides a schematic diagram of the determination process of the dynamic temperature threshold, as Figure 2 shown, including: Step 201, obtain the historical operating data of the first server, and perform data analysis processing on the historical operating data to obtain the temperature change cycle data, temperature change trend data, load change information, and environmental temperature information corresponding to the first server, where the historical operating data is the operating data of the first server before the real-time operating data.
[0043] In the embodiment of this application, the historical operating data refers to the operating records accumulated by the first server before the real-time operating data acquisition time point, including but not limited to: cycle data from several days to several months, for example: historical temperature curves (such as: minute-level temperature sampling of CPU and memory), load change logs (such as: time-series records of CPU utilization rate and memory occupancy rate), fan speed history, environmental temperature records (such as: historical data of the temperature and humidity sensor in the computer room), and the task scheduling information of the first server (such as: batch processing job period, high-load task trigger event), etc.
[0044] The data analysis processing of the historical operating data can adopt but is not limited to: multi-modal analysis methods. Specifically, regarding the analysis of the temperature change cycle data, it can adopt but is not limited to the following methods: identifying the periodic fluctuation law of the temperature of the first server through Fourier transform or periodic detection algorithms. For example: The night backup task in the data center may cause a load peak at dawn every day, and the corresponding CPU temperature shows a 24-hour periodic rise.
[0045] Regarding the analysis of temperature change trend data, the following methods can be adopted, but are not limited to: using linear regression or trend decomposition algorithms (such as Seasonal and Trend decomposition using Loess, STL) to extract the long-term trend of temperature. For example, due to the slow decline in heat dissipation efficiency caused by hardware aging, the temperature baseline under the same load increases year by year.
[0046] Regarding the analysis of load change information, the following methods can be adopted, but are not limited to: analyzing the correlation between the load of the first server and the temperature, and establishing a load-temperature mapping model. For example, GPU-intensive tasks may cause a sharp increase in temperature in a specific hardware area, while Input / Output (I / O)-intensive tasks have less impact on temperature.
[0047] Regarding the analysis of ambient temperature information, the following methods can be adopted, but are not limited to: statistically analyzing the daily changes, seasonal fluctuations, and abnormal events (such as a sudden increase in temperature caused by air conditioner failure) of the ambient temperature in the computer room where the first server is located, and evaluating the influence weight of the external environment on the server's heat dissipation.
[0048] Step 202: Perform temperature prediction processing on the first server according to the load change information, ambient temperature information, temperature change cycle data, and temperature change trend data to obtain target temperature change data.
[0049] In the embodiments of the present application, when performing temperature prediction processing, the following methods can be adopted, but are not limited to: using a hybrid model that combines time series prediction and causal reasoning for processing. Specifically, a time series prediction model (such as LSTM, Transformer) can be used: taking the historical temperature sequence and load cycle data as inputs to predict the temperature change trajectory in the next period of time (such as within the next hour).
[0050] Causal reasoning model: Analyze the causal relationship between ambient temperature, load mutation events (such as virtual machine migration, sudden computing tasks), and temperature fluctuations, and quantify the contribution of external factors to temperature. For example, if the CPU temperature rises by 0.5℃ for every 1℃ increase in ambient temperature, the impact of the ambient temperature rise will be superimposed in the prediction.
[0051] Uncertainty modeling: Evaluate the confidence interval of the prediction result through Monte Carlo simulation or Bayesian neural network to avoid misjudgment caused by a single predicted value.
[0052] The finally generated target temperature change data includes, but is not limited to: predicted temperature curve, confidence interval range, and key inflection points (such as the time and amplitude of the predicted temperature peak). For example, the content of the target temperature change data can be interpreted as predicting that the CPU temperature will rise by 8℃ due to a newly started training task within the next 30 minutes and will return to the baseline 45 minutes later.
[0053] Step 203: When it is determined according to the target temperature change data that the temperature of the first server is rising, set the dynamic temperature threshold to the first threshold.
[0054] In the embodiments of the present application, the first threshold is the threshold set when the first server is in a temperature rising scenario. For example, when the target temperature change data indicates that the temperature of the first server will rise significantly (for example, the predicted temperature rise exceeds the safety baseline by more than 3°C), the dynamic temperature threshold is set to the lower first threshold. The first threshold is usually close to the buffer value of the hardware safety upper limit (for example, set to 85°C instead of the maximum allowable 90°C), and by triggering the fan speed increase or load migration in advance, potential overheating risks are prevented. For example, when it is predicted that the GPU will heat up due to rendering tasks, the threshold is lowered from 88°C to 83°C in advance, so that the heat dissipation system can intervene in regulation before the temperature approaches the critical value.
[0055] Step 204: When it is determined according to the target temperature change data that the temperature of the first server is decreasing, set the dynamic temperature threshold to the second threshold, where the second threshold is greater than the first threshold.
[0056] In the embodiments of the present application, the second threshold is the threshold set when the first server is in a temperature decreasing scenario. For example, when the target temperature change data shows that the temperature of the first server will drop (for example, during the low-load period at night or when the environmental air conditioner's refrigeration is enhanced), the dynamic temperature threshold is adjusted up to the higher second threshold (for example, from 83°C to 88°C). The higher threshold allows the fan to maintain a lower speed within a safe range, reducing unnecessary energy consumption. For example, if it is predicted that the load will drop below 10% after 2 am and the environmental temperature drops by 5°C, the system can relax the threshold to 87°C, enabling the fan to operate at 40% speed instead of the default 60%, thus achieving quiet operation and energy conservation.
[0057] Use the time series analysis algorithm to analyze the historical temperature data of the key components of the server, and extract the periodic and trend characteristics of the temperature change. Combine the load change pattern of the server and the environmental temperature change trend, and predict the temperature change in the next period of time through the machine learning algorithm. According to the prediction results, dynamically adjust the temperature threshold of the key components. For example, when it is predicted that the server load will increase significantly, lower the temperature threshold in advance so as to start the fan speed regulation in time; when it is predicted that the environmental temperature will drop, appropriately increase the temperature threshold to reduce the unnecessary operation of the fan.
[0058] Through the dynamic adjustment of the threshold through historical data analysis and prediction, a closed-loop control logic of prediction-prevention-regulation is constructed. Before the actual temperature limit is exceeded, the cooling strategy is adjusted in advance based on the prediction results to avoid the temperature transient overshoot or frequent start-stop of the fan caused by the lag response. Through the dynamic relaxation of the threshold (the second threshold), the low-power operation time is maximized on the premise of safety, and the overall power usage effectiveness (PUE) of the data center is reduced. In response to sudden loads or environmental anomalies (such as air conditioner failures), through confidence interval evaluation and rapid threshold tightening (the first threshold), the fault tolerance of the system to uncertainties is improved.
[0059] In an implementable manner of the embodiment of the present application, when calculating the fan speed to obtain the target fan speed, the following methods can also be used but are not limited to: comparing the server temperature in the real-time operation data with the dynamic temperature threshold to obtain a first comparison result, where the server temperature is the real-time temperature of the first server; in the case where it is determined according to the first comparison result that the server temperature is greater than the dynamic temperature threshold, input the real-time operation data and the dynamic temperature threshold into the first cooling control algorithm to calculate the fan speed of the first server to obtain a first fan speed, where the first cooling control algorithm is a cooling control algorithm based on deep reinforcement learning extracted from a preset edge computing device; in the case where it is determined according to the first comparison result that the server temperature is less than or equal to the dynamic temperature threshold, input the real-time operation data and the dynamic temperature threshold into the first cooling control algorithm to calculate the fan speed of the first server to obtain a second fan speed; where the target fan speed includes the first fan speed and the second fan speed, and the second fan speed is less than the first fan speed.
[0060] In the embodiment of the present application, the server temperature in the real-time operation data refers to the current temperature values of the key components (such as CPU, GPU, memory) of the first server collected in real time by the baseboard management controller (BMC), which is usually updated at a frequency of seconds or milliseconds.
[0061] The comparison between the real-time operation data and the dynamic temperature threshold includes, but is not limited to: numerically comparing the server temperature in the real-time operation data with the dynamic temperature threshold and outputting a first comparison result, and the first comparison result includes two states: The over-limit state, that is, the state where the server temperature is greater than the dynamic temperature threshold. Among them, when the temperature of any key component of the first server is greater than the dynamic temperature threshold, it is determined that the first server is in the over-limit state. For example: if the dynamic temperature threshold is set to 85 °C and the real-time temperature of the CPU is 87 °C, the over-limit state is triggered.
[0062] The safe state is the state where the server temperature is less than or equal to the dynamic temperature threshold. Among them, when the temperatures of all critical components of the first server are less than or equal to the dynamic temperature threshold, the first server is determined to be in a safe state. For example: the memory temperature is 83°C (dynamic temperature threshold 85°C), and the GPU temperature is 80°C (dynamic temperature threshold 85°C), then it is in a safe state.
[0063] The differential calculation of the first fan speed and the second fan speed refers to generating the target fan speed using different control strategies according to different first comparison results. Among them, the calculation of the first fan speed in the over-limit state is as follows: when it is determined to be over-limit, that is, when the server temperature is greater than the dynamic temperature threshold, enter the high-response heat dissipation mode. At this time, the real-time operation data (including: over-limit temperature value, load intensity, ambient temperature) and the dynamic temperature threshold are jointly input into the first heat dissipation control algorithm. The first heat dissipation control algorithm is a PID algorithm based on deep reinforcement learning extracted from a preset edge computing device. Based on historical training data and real-time operation data, a fan speed command (i.e., the first fan speed) suitable for the over-limit scenario is output. For example: when the CPU temperature is over-limit and the load remains high, the algorithm may calculate a relatively high speed (such as: 5000 revolutions per minute (RPM)), and superimpose a dynamic adjustment factor (such as: an additional 200 RPM increase according to the rising ambient temperature) to quickly suppress the rising trend of the temperature.
[0064] The calculation of the second fan speed in the safe state is as follows: when the temperature is within the safe range, that is, when the server temperature is less than or equal to the dynamic temperature threshold, switch to the energy efficiency priority mode. At this time, the same real-time operation data and the dynamic temperature threshold are input into the first heat dissipation control algorithm, but the algorithm will output a lower second fan speed according to the characteristics of the safe state (such as: the difference between the temperature and the threshold, the decreasing trend of the load). For example: if the CPU temperature is 82°C (threshold 85°C) and the load is stable at 30%, the algorithm may set the speed to 3000 RPM, which is only 60% of the first speed. In this mode, the algorithm actively reduces the speed to reduce power consumption through the energy consumption optimization target in the reinforcement learning reward mechanism, while ensuring that the temperature always remains within the safe range.
[0065] Among them, the first heat dissipation control algorithm is deployed on a preset edge computing device. Through model lightweighting, pruning and quantization can be performed on the DRL network of the first heat dissipation control algorithm to reduce the computational complexity while retaining the core decision-making ability, enabling it to run in real time on embedded hardware (such as: BMC).
[0066] Through local incremental training, the first heat dissipation control algorithm is fine-tuned regularly using newly collected local data to gradually adapt to long-term changing factors such as server hardware aging and radiator dust accumulation, avoiding performance degradation of the first heat dissipation control algorithm.
[0067] Provide an abnormal fusing mechanism. When the algorithm calculation is abnormal (e.g., the output rotation speed exceeds the physical limit of the fan), the preset edge computing device automatically switches to the preset conservative control strategy (e.g., fixed PID parameter mode) to ensure system safety.
[0068] Implement dynamic hierarchical control of the heat dissipation strategy through temperature threshold comparison: prioritize heat dissipation efficiency in the over-limit state, and quickly calculate high-speed rotation commands through a heat dissipation control algorithm based on deep reinforcement learning to suppress temperature rise; focus on energy efficiency optimization in the safe state, and use the multi-objective optimization ability of the same algorithm to reduce the rotation speed. The differential control mechanism overcomes the rigidity of the traditional single rotation speed strategy, avoiding both energy waste caused by excessive heat dissipation and quickly responding when temperature risks occur. At the same time, the algorithm deployment mode based on edge computing reduces response latency, ensuring the real-time and reliability of critical heat dissipation commands, and providing a solid guarantee for the stable operation of high-density servers.
[0069] In an implementable manner of the embodiment of the present application, when calculating the first fan rotation speed, the following method can also be used but is not limited to: calculate the first temperature difference between the server temperature and the dynamic temperature threshold; input the first temperature difference, real-time operation data, and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan rotation speed calculation processing on the first server to obtain the first fan rotation speed, where the magnitude of the first temperature difference is positively correlated with the magnitude of the first fan rotation speed.
[0070] In the embodiment of the present application, the first temperature difference refers to the numerical difference between the real-time temperature (i.e., the server temperature) of the key components (such as CPU, etc.) of the first server and the dynamic temperature threshold. The first temperature difference obtains the current temperature value through real-time temperature acquisition and performs a subtraction operation with the dynamic temperature threshold. For example: if the real-time CPU temperature is 88 °C and the dynamic temperature threshold is set to 85 °C, then the first temperature difference is +3 °C; if the real-time temperature is 82 °C and the threshold is 85 °C, then the difference is -3 °C. This difference not only reflects the deviation degree of the current temperature from the safety boundary but also implies the urgency of the heat dissipation requirement, providing a quantitative basis for subsequent fan rotation speed calculation.
[0071] When generating the first fan rotation speed according to the first heat dissipation control algorithm, it can be achieved through but not limited to the following methods: State space construction: Use the first temperature difference as the core state variable, and jointly form a multi-dimensional input vector with parameters such as load and ambient temperature. For example: the input vector can be expressed as [temperature difference ΔT = +3 °C, CPU load = 90%, ambient temperature = 25 °C, current fan rotation speed = 4000 RPM].
[0072] Differential weight dynamic allocation: Automatically learn the influence weight of the temperature difference on the fan speed in different scenarios through a neural network. For example, in a high-load scenario (such as when the CPU utilization rate exceeds 80%), a higher weight may be assigned to the positive temperature difference to quickly respond to potential overheating risks; while in a low-load scenario, the model may reduce the weight and prioritize energy consumption optimization.
[0073] Positive correlation control logic implementation: Strengthen the control strategy of the positive correlation between the first temperature difference and the fan speed through the reward mechanism in the training data. When the first temperature difference increases (such as when ΔT increases from +2°C to +5°C), a higher fan speed is inclined to be output (such as increasing from 4500 RPM to 6000 RPM) to accelerate heat dissipation; conversely, when the first temperature difference decreases, the speed is gradually reduced to reduce energy consumption. This step can be achieved through the design of the reward function of DRL. For example, conduct multi-objective trade-offs on the energy consumption cost of the reduction amplitude of the first temperature difference and the increase amplitude of the speed per unit time to ensure that the positive correlation meets the overall optimization goal of the system.
[0074] Directly reflect the urgency of the heat dissipation requirement through the magnitude of the first temperature difference, and dynamically adjust the speed in combination with multi-dimensional data such as load and environment to avoid overheating or insufficient heat dissipation caused by a one-size-fits-all strategy. Ensure that the fan speed increases with the increase of the first temperature difference through the positive correlation logic, but at the same time suppress unnecessary excessive speed through the multi-objective optimization mechanism of the first heat dissipation control algorithm, and achieve a balance between rapid cooling and energy conservation and consumption reduction.
[0075] In an implementable manner of the embodiment of the present application, when calculating the second fan speed, the following method can also be adopted but is not limited to: Calculate the second temperature difference between the server temperature and the dynamic temperature threshold; input the second temperature difference, real-time operation data, and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain the second fan speed, where the magnitude of the second temperature difference is negatively correlated with the magnitude of the second fan speed.
[0076] In the embodiment of the present application, the second temperature difference also refers to the numerical difference between the real-time temperature (i.e., the server temperature) of the key components (such as the CPU, etc.) of the first server and the dynamic temperature threshold, but its calculation scenario is limited to the case where the real-time temperature is less than or equal to the dynamic temperature threshold. At this time, the difference is zero or negative (such as the difference between the real-time temperature of 82°C and the threshold of 85°C is -3°C), reflecting the distance between the current temperature and the safety upper limit. Different from the positive difference in the over-limit state, the negative difference characterizes the redundancy ability of the heat dissipation system and provides a quantitative basis for energy efficiency optimization. For example, when the difference reaches -5°C, it indicates that the server temperature is much lower than the threshold, and there is room to reduce the fan speed to save energy.
[0077] The magnitude of the second temperature difference is negatively correlated with the magnitude of the second fan speed, which means that when the temperature difference increases negatively (i.e., the real-time temperature is further lower than the threshold), the target fan speed decreases accordingly. When generating the second fan speed according to the first heat dissipation control algorithm, it can be achieved by, but not limited to, the following methods: Difference-sign sensitive state encoding: The first heat dissipation control algorithm separates the positive and negative of the temperature difference at the input layer. For negative differences, i.e., the second temperature difference, the energy efficiency optimization sub-network is activated to learn how to minimize the fan speed while maintaining temperature safety through historical data. For example, the input vector [ΔT = -3°C, load = 20%, ambient temperature = 22°C] may trigger the low-speed decision branch.
[0078] Multi-objective trade-off of the reward function: In a safe state, the reward function of the first heat dissipation control algorithm focuses on energy consumption and noise suppression. For example, when the fan speed decreases and the temperature remains stable, a positive reward is given; if the speed is too low and the temperature approaches the threshold (e.g., ΔT rises from -5°C to -1°C), a slight penalty is imposed to prompt the risk. Through repeated training, the algorithm autonomously establishes the association rule that the larger the second temperature difference → the lower the speed → the higher the reward, forming a negative correlation control strategy.
[0079] Dynamic relaxation constraint: Dynamically adjust the allowable range of the second temperature difference according to the ambient temperature and load trend. For example, when the ambient temperature is low (e.g., 18°C) and the load is stable, the second temperature difference is allowed to reach -8°C before reducing the speed; while when the ambient temperature is high (e.g., 28°C) or the load has an upward trend, the second temperature difference is limited within -4°C, reserving heat dissipation margin to cope with potential temperature rise.
[0080] The calculation process of the second fan speed includes but is not limited to the following steps: Difference interval grading: Divide the second temperature difference into multiple intervals (e.g., -1°C to -3°C is mild redundancy, -3°C to -6°C is high redundancy, less than -6°C is excessive redundancy), and each interval corresponds to a different speed adjustment strategy. For example, in the high redundancy interval, the algorithm can reduce the speed through a linear relationship (for every -1°C increase in ΔT, the speed decreases by 200 RPM); while in the excessive redundancy interval, a non-linear attenuation strategy is enabled (e.g., when ΔT drops from -6°C to -8°C, the speed only decreases by 50 RPM) to avoid local hot spots caused by fan stoppage.
[0081] Environmental coupling compensation: When calculating the speed, compensate for the impact of the ambient temperature on the heat dissipation efficiency. For example, when the ambient temperature rises (e.g., from 20°C to 25°C), even if ΔT remains -4°C, the algorithm will slightly increase the speed (e.g., increase by 100 RPM) to offset the potential impact of the deterioration of the external thermal environment on the heat dissipation effect.
[0082] Load change prediction: By analyzing real-time load data (such as: process queue length, I / O throughput), predict the future short-term load change trend. If the load may increase (such as: detecting a batch task start signal), the algorithm will reserve a small amount of redundancy on the current rotation speed (such as: maintaining an additional 200 RPM) to avoid the rapid approach of the temperature to the threshold due to a sudden increase in load.
[0083] Through the negative correlation control of the second temperature difference value and the fan rotation speed, when there is sufficient temperature redundancy, actively reduce the fan rotation speed to reduce power consumption and noise pollution. Through the dynamic relaxation constraint and load prediction mechanism, avoid the rapid rise of the temperature back to near the threshold due to excessive speed reduction, and maintain the smoothness of the temperature curve.
[0084] In an implementable manner of the embodiment of the present application, when performing fan rotation speed adjustment processing on the first server according to the target fan rotation speed, the following methods can also be used but are not limited to: Transmit the target fan rotation speed to the fan controller of the first server, and control the fan rotation speed of the first server to the target fan rotation speed through the fan controller to obtain the adjusted first server; when it is determined that the fan rotation speed of the first server is the target fan rotation speed, obtain the operation data of the first server to obtain the second operation data.
[0085] In the embodiment of the present application, the target fan rotation speed is an ideal rotation speed value calculated by the first heat dissipation control algorithm, including the first fan rotation speed (high rotation speed in the over-limit state) or the second fan rotation speed (low rotation speed in the safe state). The target fan rotation speed can be transmitted to the fan controller of the first server through the communication interface of the baseboard management controller (BMC). The fan controller is a hardware module dedicated to managing the fan operation inside the first server, usually integrated on the motherboard or an independent control card, and is responsible for converting the digital rotation speed instruction into a physical drive signal.
[0086] After confirming that the fan rotation speed has stabilized to the target fan rotation speed, collect the second operation data, and the second operation data includes but is not limited to: Temperature change response data, at least including: the temperature change curve, cooling rate, steady-state temperature value, and fluctuation amplitude of key components (such as: CPU) after the rotation speed adjustment. Energy consumption data, at least including: the real-time power consumption, cumulative energy consumption, and power factor of the fan at the target rotation speed. Noise and vibration data, at least including: the noise decibel value and chassis vibration amplitude collected by a microphone or an acceleration sensor when the fan is running at high speed, used to evaluate the impact of the heat dissipation strategy on the device working environment. Load-related data, at least including: record the load status of the server during the rotation speed adjustment (such as: CPU utilization rate, memory bandwidth occupancy rate), and analyze the coupling relationship between the heat dissipation effect and the load fluctuation.
[0087] Through the gradient adjustment and feedback verification mechanism of the fan controller, ensure that the rotation speed command is accurately executed, and avoid control deviation caused by mechanical inertia or electrical noise. Provide real working condition feedback for the deep reinforcement learning algorithm through the second operating data, which is used to correct the prediction error of the first heat dissipation control algorithm (for example, when the actual temperature reduction speed is lower than expected, adjust the mapping relationship between the difference value and the rotation speed) or optimize the reward function weight (for example, reduce the reward score for the high rotation speed strategy in a high-noise scenario). By monitoring the adjusted temperature and energy consumption in real time, abnormal states (such as non-compliance with the rotation speed caused by fan failure) can be quickly identified, and a redundant heat dissipation solution (such as enabling a standby fan or migrating the computing load) can be triggered to prevent risks caused by local overheating.
[0088] In an implementable manner of the embodiment of the present application, when performing data evaluation and analysis processing on the second operating data, the following methods can also be used but are not limited to: perform data analysis processing on the second operating data according to historical operating data and preset evaluation indicators to obtain a heat dissipation evaluation result, where the heat dissipation evaluation result at least includes a temperature evaluation result, a rotation speed evaluation result, and an energy consumption evaluation result.
[0089] In the embodiment of the present application, the historical operating data is used as an evaluation benchmark, including but not limited to: heat dissipation performance records of the first server under various typical working conditions, for example: Temperature baseline library: historical average temperature and reasonable fluctuation range corresponding to different load levels (such as 30%, 60%, 90% CPU utilization); rotation speed efficiency model: typical heat dissipation capacity (such as a temperature reduction of 2°C per minute) at a specific rotation speed (such as 4000 RPM) and the historical best rotation speed-temperature response relationship; energy consumption mode library: fan energy consumption baseline and energy efficiency optimization threshold under different ambient temperatures (such as 20°C, 25°C, 30°C).
[0090] By performing time series alignment and scenario matching on the second operating data and the historical operating data (such as screening records with the same load interval and similar environmental conditions), a comparable evaluation framework is established to eliminate the interference of external variables. For example: when the CPU load in the second operating data is 75%, extract all records with a load between 70% - 80% from the historical library as the comparison benchmark.
[0091] The preset evaluation indicators are multi-dimensional analysis indicators set in advance, including but not limited to: analysis indicators covering three dimensions of heat dissipation performance, control efficiency, and energy efficiency. Through the analysis of the second operating data using the preset evaluation indicators and the historical operating data, a heat dissipation evaluation result can be obtained that at least includes a temperature evaluation result, a rotation speed evaluation result, and an energy consumption evaluation result. Specifically, the evaluation process of the temperature evaluation result includes but not limited to: Stability evaluation: Calculate the standard deviation and the maximum fluctuation amplitude of the adjusted temperature curve to determine whether the temperature converges stably within the dynamic threshold range. For example, if the temperature fluctuation is less than ±1°C within 10 minutes after adjustment, it is marked as high stability. Overlimit risk evaluation: Statistically analyze the duration and frequency of the temperature approaching the threshold (e.g., reaching 95% of the threshold) to evaluate the safety margin of the heat dissipation strategy. For example, if the temperature touches 84°C (dynamic temperature threshold of 85°C) multiple times after adjustment, it indicates a potential overlimit risk. Heat dissipation efficiency evaluation: Quantify the cooling effect brought by the increase in unit rotational speed based on the ratio of the temperature change rate to the rotational speed adjustment amount. For example, if the temperature drops at a rate of 3°C per minute after the rotational speed increases by 1000 RPM, the heat dissipation efficiency is rated as excellent.
[0092] The evaluation process of the rotational speed assessment results includes, but is not limited to: Response timeliness evaluation: Measure the total delay from temperature overlimit detection to the rotational speed stabilizing to the target fan rotational speed to judge the real-time performance of the control link. For example, if the delay exceeds 5 seconds, it is regarded as a response lag. Rotational speed smoothness evaluation: Analyze the deviation rate and fluctuation frequency between the actual rotational speed and the target fan rotational speed to evaluate the execution accuracy of the control instruction. For example, if the rotational speed fluctuates within the range of ±2% without periodic oscillation, it is rated as high smoothness. Strategy rationality evaluation: Compare the difference between the actual rotational speed and the historical optimal rotational speed strategy to identify whether there is over-adjustment or conservative control. For example, if the current rotational speed is 10% higher than the best value under the same historical working conditions, it indicates that the strategy may be redundant.
[0093] The evaluation process of the energy consumption assessment results includes, but is not limited to: Absolute energy consumption evaluation: Calculate the total energy consumption of the heat dissipation system during the adjustment period and compare it with the historical average value of the same scenario to judge the improvement or degradation of energy efficiency. For example, if the energy consumption is reduced by 15% and the temperature is stable, it is marked as an energy efficiency improvement. Energy Temperature Ratio (ETR): Defined as the energy consumed per unit temperature drop (e.g., 0.5 kWh of electricity is consumed per 1°C drop), which is used to horizontally compare the energy efficiency levels of different heat dissipation strategies. Load-related energy efficiency evaluation: Analyze the sensitivity of energy consumption to load changes to identify the energy efficiency bottleneck under high load. For example, if the energy consumption surges by 200% when the load increases from 50% to 80%, it indicates that there is room for optimization of the heat dissipation system under high load scenarios.
[0094] The heat dissipation assessment results can be generated through, but are not limited to, the following processes: Data alignment and normalization: Align the second run data with the historical data in a time window, unify the units, and remove outliers to ensure comparability.
[0095] Indicator weighted calculation: According to the first server type and the operation and maintenance strategy, dynamic weights are assigned to each evaluation indicator. For example, in a high-performance computing server, the weight of temperature stability may be as high as 60%, while the weight of energy consumption only accounts for 20%; in an edge low-power server, the weight of energy consumption can be increased to 50%.
[0096] Comprehensive evaluation output: Using fuzzy logic or a weighted scoring model, the scores of each dimension are fused into an overall evaluation result (such as a percentage score or a grade of excellent / good / fair / poor), and diagnostic suggestions are generated (such as suggesting reducing the rotation speed by 200 RPM to optimize energy efficiency).
[0097] Through in-depth evaluation of multi-dimensional and multi-source data, and directly feedback the heat dissipation evaluation results to the reward function of the deep reinforcement learning algorithm to guide the priority learning of efficient, low-power, and stable control strategies, accelerating the convergence of the algorithm. The breakdown evaluation results of temperature, rotation speed, and energy consumption help the operation and maintenance personnel quickly locate the bottleneck of the heat dissipation system (such as the performance decay of a specific fan or the design defect of the heat dissipation air duct), and formulate a targeted maintenance plan. Combining the comparative analysis of historical data, the system can identify long-term trends (such as the energy efficiency decline caused by dust accumulation in the radiator), automatically trigger model retraining or hardware maintenance alarms, and improve the reliability within the system life cycle.
[0098] In an implementable manner of the embodiment of the present application, after obtaining the heat dissipation evaluation result, it is necessary to determine whether the heat dissipation evaluation result meets the preset heat dissipation conditions. Specifically, regarding the analysis method of the heat dissipation evaluation result, the embodiment of the present application provides an analysis flowchart of the heat dissipation evaluation result, as Figure 3 shown, including: Step 301, compare the temperature evaluation result with the temperature evaluation indicator in the preset evaluation indicators to obtain a temperature evaluation deviation value, where the preset evaluation indicators at least include a rotation speed evaluation indicator, an energy consumption evaluation indicator, and a temperature evaluation indicator.
[0099] In the embodiment of the present application, the temperature evaluation indicator is a preset temperature-related performance standard, including but not limited to: a temperature stability threshold (such as an allowable fluctuation range of ±2°C), an over-limit risk tolerance (such as the proportion of the time when the temperature is close to the threshold ≤10%), and a heat dissipation efficiency baseline (such as a temperature reduction rate of ≥1°C per minute). The temperature evaluation result includes but is not limited to: the temperature stability value obtained by actual measurement or calculation, the over-limit risk level, and the heat dissipation efficiency score.
[0100] The comparison process can be achieved by, but not limited to, deviation calculation. The difference or ratio operation is performed on each sub-item in the temperature evaluation result (such as stability, over-limit risk, efficiency) and the corresponding temperature evaluation index to generate a temperature evaluation deviation value. For example, if the preset stability requirement is that the fluctuation ≤ ±1°C, and the actual fluctuation is ±1.5°C, then the deviation value is +0.5°C; if the preset heat dissipation efficiency is to cool down 2°C per minute and the actual value is 2.3°C, then the deviation value is -0.3°C (a negative value indicates better than expected). The absolute value of the deviation value reflects the degree of deviation, and the sign indicates the direction (a positive value means not meeting the standard, and a negative value means exceeding the target).
[0101] Step 302, perform a comparison process based on the rotation speed evaluation result and the rotation speed evaluation index to obtain a rotation speed evaluation deviation value, and perform a comparison process based on the energy consumption evaluation result and the energy consumption evaluation index to obtain an energy consumption evaluation deviation value.
[0102] In the embodiments of the present application, the rotation speed evaluation index is a preset rotation speed-related performance standard, including but not limited to: the upper limit of response time (such as: from the instruction issuance to the rotation speed stabilization ≤ 3 seconds), the rotation speed control accuracy (such as: the deviation between the actual rotation speed and the target value ≤ ±3%), and the strategy rationality threshold (such as: the rotation speed adjustment range does not exceed ±15% of the historical optimal value). The rotation speed evaluation result includes but not limited to: the measured response time, the rotation speed deviation rate, and the strategy difference degree. In the comparison process, for example: the actual response time and the preset upper limit (such as: 3 seconds), calculate the rotation speed evaluation deviation value (such as: if the actual is 4 seconds, then the deviation is +1 second); similarly, the rotation speed deviation rate (such as: the actual deviation is 5% and the preset is 3%) generates a deviation value of +2%.
[0103] The energy consumption evaluation index is a preset energy consumption-related performance standard, including but not limited to: the maximum allowable power consumption (such as: a single fan ≤ 20W), the energy consumption-temperature ratio (such as: ETR ≤ 0.6 kWh / °C), and the load sensitivity limit (such as: when the load increases by 30%, the energy consumption increase rate ≤ 50%). The energy consumption evaluation result includes but not limited to: the measured power consumption, the ETR value, and the slope of the load-energy consumption curve. In the comparison process, for example: if the actual ETR is 0.7 kWh / °C and the preset index is 0.6, then the deviation value is +0.1; if the energy consumption only increases by 40% when the load increases by 30%, then the deviation value is -10%.
[0104] Step 303, determine whether the temperature evaluation result meets the preset heat dissipation condition according to the temperature evaluation deviation value, determine whether the rotation speed evaluation result meets the preset heat dissipation condition according to the rotation speed evaluation deviation value, and determine whether the energy consumption evaluation result meets the preset heat dissipation condition according to the energy consumption evaluation deviation value.
[0105] In the embodiments of the present application, the preset heat dissipation conditions at least include a set of allowable range of deviation values for each evaluation dimension. For example: Temperature dimension: stability deviation ≤ +0.5°C, overlimit risk deviation ≤ +5%, efficiency deviation ≥ -0.2°C / minute; Rotation speed dimension: response time deviation ≤ +1 second, control accuracy deviation ≤ +2%, strategy difference degree ≤ ±10%; Energy consumption dimension: power consumption deviation ≤ +10%, ETR deviation ≤ +0.1, load sensitivity deviation ≤ +15%. Specifically, the content of the preset heat dissipation conditions is not limited in the present application.
[0106] In the process of determining whether it meets the preset heat dissipation conditions, each deviation value is checked one by one to see if it exceeds the allowable range. For example: If the temperature stability deviation is +0.6°C (exceeding the upper limit of +0.5°C), then mark that the temperature evaluation does not meet the requirements; If the rotation speed control accuracy deviation is +1.5% (within the tolerance of +2%), then mark it as meeting the requirements. This process can be implemented through a rule engine or fuzzy logic, supporting flexible processing of boundary conditions (e.g., deviation values close to the threshold can trigger a warning instead of directly determining failure).
[0107] Step 304, in the temperature evaluation result, rotation speed evaluation result, and energy consumption evaluation result, if any one of the evaluation results does not meet the preset heat dissipation conditions, it is determined that the heat dissipation evaluation result does not meet the preset heat dissipation conditions.
[0108] In the embodiments of the present application, when the evaluation result of any one dimension of temperature, rotation speed, and energy consumption does not meet the preset heat dissipation conditions, it is determined that the overall heat dissipation evaluation result does not meet the preset heat dissipation conditions. For example: Even if the rotation speed and energy consumption evaluations are both up to standard, but the temperature stability exceeds the limit, the optimization process is still triggered. This design reflects the strict priority setting of the system for heat dissipation safety, while taking into account the balance of energy efficiency and control performance.
[0109] Through multi-dimensional deviation calculation, the weak links of the heat dissipation strategy can be accurately located (such as too long response delay or low energy efficiency), avoiding the one-sidedness of traditional single-index determination. The preset heat dissipation conditions can be dynamically adjusted according to the server life cycle stage and environmental characteristics. For example: The energy consumption deviation tolerance can be relaxed for old servers, while the temperature stability requirements can be tightened for newly built high-performance servers. When any core indicator (such as temperature safety) does not meet the standard, emergency optimization is triggered first to ensure system reliability; For secondary indicators (such as slight energy efficiency deviation), they can be included in the periodic optimization tasks.
[0110] In an implementable manner of the embodiment of the present application, when determining whether the temperature evaluation result, the rotation speed evaluation result, and the energy consumption evaluation result meet the preset heat dissipation conditions, the following methods can be adopted but are not limited to: when the temperature evaluation deviation value is greater than the first preset evaluation deviation threshold, it is determined that the temperature evaluation result does not meet the preset heat dissipation conditions; when the rotation speed evaluation deviation value is greater than the second preset evaluation deviation threshold, it is determined that the rotation speed evaluation result does not meet the preset heat dissipation conditions; when the energy consumption evaluation deviation value is greater than the third preset evaluation deviation threshold, it is determined that the energy consumption evaluation result does not meet the preset heat dissipation conditions.
[0111] In the embodiment of the present application, through multi-dimensional independent determination, the defective links of the heat dissipation strategy can be quickly located. For example: if only the energy consumption exceeds the standard, the optimization direction focuses on reducing the rotation speed or introducing load migration; if both the temperature and the rotation speed do not meet the standard, it is necessary to comprehensively adjust the PID parameters and the fan response logic. It provides core support for the continuous optimization and refined operation and maintenance of the server heat dissipation system.
[0112] In an implementable manner of the embodiment of the present application, when performing algorithm optimization processing on the first heat dissipation control algorithm, the following methods can be adopted but are not limited to: performing algorithm optimization processing on the first heat dissipation control algorithm through a preset deep reinforcement learning algorithm according to the second operation data, the heat dissipation evaluation result, and the historical operation data; until the preset optimization times or the preset convergence condition are reached, the second heat dissipation control algorithm is obtained.
[0113] In the embodiment of the present application, the BMC regularly trains the first heat dissipation control algorithm based on deep reinforcement learning using the latest training samples. During the training process, the algorithm continuously adjusts the parameters of the policy network through interaction with the environment (i.e., the heat dissipation control process) to maximize the cumulative reward. When the algorithm converges or reaches the preset number of training times, the optimized parameters of the first heat dissipation control algorithm are obtained for subsequent heat dissipation control.
[0114] The second operation data provides the actual working condition feedback after the fan rotation speed is adjusted, including but not limited to: the temperature response curve, the energy consumption change, and the control stability index; the heat dissipation evaluation result quantifies the performance defects of the current policy in terms of temperature safety, control accuracy, and energy efficiency (such as: the temperature fluctuation exceeds the standard or the energy consumption is too high); the historical operation data covers diverse working condition records accumulated during the long-term operation of the server, providing a wide range of training samples and a pattern recognition basis for algorithm optimization. The three together constitute a multi-dimensional input space based on the deep reinforcement learning mechanism, ensuring that the optimization process considers both the immediate control effect and the historical experience and the system-level performance goals.
[0115] The preset number of optimization times is the number of optimization times set by the user, which is a safeguard mechanism to avoid infinite iteration. For example, the maximum number of iterations for local training is 1000 times, etc. Specifically, the embodiments of the present application do not limit the preset number of optimization times.
[0116] The algorithm can autonomously identify long-term factors such as server hardware aging and changes in the efficiency of the heat dissipation air duct, and dynamically adjust the parameters of the first heat dissipation control algorithm to maintain the best heat dissipation efficiency; it can achieve dynamic trade-offs among conflicting goals such as temperature safety, energy consumption economy, and control stability. For example, in a slightly over-limit scenario, it selects to moderately increase the rotation speed instead of the maximum rotation speed, taking into account both the cooling speed and the energy consumption cost; the second heat dissipation control algorithm has stronger anti-interference ability. When there is a sudden load impact or a drastic change in the ambient temperature, it can quickly respond and suppress temperature fluctuations to avoid cascading failures.
[0117] In an implementable manner of the embodiments of the present application, after transmitting the second heat dissipation control algorithm to the preset edge computing device, the following methods can be adopted but are not limited to: in response to a control instruction for heat dissipation control of the second server, extract the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server.
[0118] In the embodiments of the present application, the preset edge computing device receives the collected data, that is, the third operation data, in real time and analyzes the data in real time. For simple changes in heat dissipation requirements, such as small temperature fluctuations or slight load changes, the preset edge computing device quickly adjusts the parameters of the second heat dissipation control algorithm and the fan rotation speed according to the locally preset rules and lightweight models. When encountering complex heat dissipation scenarios, such as a sudden large increase in server load or a sharp change in ambient temperature, the preset edge computing device uploads the key data to the cloud computing platform. The cloud computing platform uses big data analysis and complex machine learning models to deeply process the data, generates an optimized control strategy, and feeds the strategy back to the preset edge computing device, which then performs the corresponding adjustment operations.
[0119] Among them, the second server can be the same as or different from the first server. The second server refers to other server nodes different from the first server. The hardware configuration, load characteristics, and heat dissipation requirements of the second server may be different from those of the first server (for example, using different models of CPUs or being deployed in different rack positions in the computer room). When the system detects that the second server needs heat dissipation control (for example, detecting that its temperature approaches the dynamic threshold or the load increases significantly), a control instruction is generated and sent to the preset edge computing device. The control instruction includes but is not limited to: the unique identifier of the second server, the current operation status summary (such as temperature, load, ambient temperature), and the heat dissipation control target (such as priority energy efficiency or forced cooling).
[0120] After receiving the instruction, the preset edge computing device performs heat dissipation control processing on the second server through the second heat dissipation control algorithm. In a multi-server cluster environment, each server shares heat dissipation-related data through the blockchain network. At regular time intervals, local temperature data, fan speed data, load information, heat dissipation control algorithm parameters, etc. are encapsulated into blocks and transmitted to the preset edge computing device, and then broadcast and verified in the blockchain network through the consensus mechanism. Based on the shared data on the blockchain, the distributed heat dissipation optimization algorithm runs regularly, and calculates the optimal fan speed and heat dissipation control algorithm parameter adjustment scheme for each server according to the real-time status of each server. Each server adjusts the local heat dissipation control strategy according to the algorithm results to achieve coordinated heat dissipation of multiple servers.
[0121] Through the optimized heat dissipation control algorithm, that is, the second heat dissipation control algorithm, it is quickly promoted within the cluster, avoiding the waste of resources for each server to train independently, and significantly shortening the strategy convergence time of new nodes. For example, when a new server is deployed, the matching optimization model in the existing algorithm library can be directly called without starting from scratch.
[0122] In an implementable manner of the embodiment of the present application, when extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server, the following methods can also be used but are not limited to: obtaining the third running data of the second server, and performing data comparison processing according to the third running data and the real-time running data to obtain the running difference data between the first server and the second server; when the running difference data meets the preset first running difference condition, extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server; when the running difference data does not meet the preset first running difference condition and meets the second running difference condition, adjusting the second heat dissipation control algorithm according to the preset algorithm update rule in the preset edge computing device to obtain the adjusted second heat dissipation control algorithm, and performing heat dissipation control processing on the second server according to the adjusted second heat dissipation control algorithm; when the running difference data does not meet the second running difference condition, obtaining the second historical data of the second server, and performing algorithm optimization processing on the second heat dissipation control algorithm according to the second historical data and the third running data to obtain the optimized second heat dissipation control algorithm; performing heat dissipation control processing on the second server according to the optimized second heat dissipation control algorithm.
[0123] In the embodiment of the present application, the third running data refers to the real-time running status information of the second server, including but not limited to: hardware parameters: CPU model, radiator type, fan specifications, etc.; dynamic indicators: current temperature, load intensity (such as: CPU utilization rate, memory bandwidth), ambient temperature, fan speed; historical characteristics: typical temperature rise mode, heat dissipation response delay, energy consumption baseline.
[0124] Data comparison processing can generate operation difference data through the following steps, but not limited to: Feature extraction: Extract key features from the real-time operation data of the first server and the second server, such as: the slope of the load-temperature curve, the heat dissipation efficiency coefficient (temperature reduction ability per unit speed), and the environmental temperature sensitivity.
[0125] Similarity calculation: Use cosine similarity or Euclidean distance to measure the temperature response difference, energy consumption difference, and control stability difference between the two servers in the same load range (such as: 60%-80% CPU utilization).
[0126] Difference quantification: Convert the similarity result into an interpretable difference score. For example: If the temperature difference between the two servers under the same load is +5°C, the difference score is marked as high; if the energy consumption difference is less than 5%, it is marked as low.
[0127] Among them, the preset first operation difference condition is the threshold for allowing direct reuse of the second heat dissipation control algorithm. For example: The heat dissipation architectures of the second server and the first server are similar (such as: both use a liquid cooling + air cooling hybrid design); the load-temperature response difference ≤ 10%, the energy consumption difference ≤ 8%; the duct efficiency difference at the rack positions where they are located ≤ 15%.
[0128] When the operation difference data meets the first operation difference condition, the preset edge computing device directly invokes the second heat dissipation control algorithm and deploys it to the BMC of the second server. For example: If the two servers are of the same model and are in adjacent racks, the optimized PID parameters and speed strategy can be directly applied.
[0129] The preset second operation difference condition is the threshold for allowing adaptation through lightweight adjustment of the algorithm. For example: The radiator types are the same but the fan specifications are different (such as the speed range difference ≤ 20%); the load-temperature response difference ≤ 25%, the energy consumption difference ≤ 20%; the duct efficiency difference ≤ 30%.
[0130] In this case, trigger the preset algorithm update rule to adjust the second heat dissipation control algorithm. Regarding the process of adjusting the second heat dissipation control algorithm, it can be implemented in the following ways, but not limited to: Parameter scaling: Adjust the parameters of the second heat dissipation control algorithm according to hardware differences. For example, if the maximum rotation speed of the second server's fan is 80% of that of the first server, then scale the proportionality coefficient P by 1.25 times to compensate for the difference in rotation speed limits. Model fine-tuning: Use a small amount of real-time data from the second server (such as the running records in the last 1 hour) to perform transfer learning on the DRL model on a preset edge device and update the output layer weights of the policy network. Rule injection: Add rules for compensating environmental differences. For example, if the environmental temperature of the second server is relatively high, then add a fixed offset (such as +200 RPM) to the target rotation speed calculation. While retaining the core control logic, the adjusted second heat dissipation control algorithm adapts to the local characteristics of the second server. For example, it increases the heat dissipation margin for servers in high environmental temperature areas.
[0131] When the running difference data does not meet the second running difference condition (such as: significantly different hardware architectures, performance deviation > 30%), enter the deep optimization process for the second heat dissipation control algorithm: Integration of second historical data: Extract the running data accumulated by the second server over a long period, including extreme operating conditions (such as full-load stress tests), historical heat dissipation events (such as overheat alarm records), and maintenance logs (such as the efficiency changes after radiator cleaning). Joint training: Mix the second historical data with the third running data and retrain the DRL model: State space expansion: Add features unique to the second server (such as video memory temperature, liquid cooling pump speed); Reward function customization: Adjust the reward weights according to the operation and maintenance objectives of the second server (such as giving priority to ensuring heat dissipation of key components); Cross-model knowledge distillation: Extract general strategies (such as temperature-rotation speed mapping rules) from the optimized model of the first server to accelerate the convergence of the new model. Edge-cloud collaborative optimization: If the local computing power is insufficient, encrypt the data and upload it to the cloud, use distributed training to generate the optimized second heat dissipation control algorithm, and then send it back to the edge device for verification and deployment.
[0132] Through hierarchical condition judgment and policy branches, elastic control from direct reuse to deep customization is achieved, taking into account both efficiency and accuracy. Mild adjustments avoid the computational overhead of full-scale retraining, while deep optimization ensures precise control of heterogeneous hardware and optimizes edge computing power allocation. After the optimization results of the second server are stored on the blockchain, they can be called by other similar nodes in the cluster to form an evolving heat dissipation strategy library.
[0133] To facilitate the understanding of the implementation process of this application, the embodiments of this application also provide a schematic diagram of the system structure for server heat dissipation control, as Figure 4As shown in the figure, the component monitoring module is a module in the first server, such as BMC in the first server, etc. It includes various components of the first server and is used to monitor the real-time operation data of the first server. The BMC data acquisition module is used to obtain the real-time operation data for heat dissipation control processing. The PID control calculation module is used to calculate the target fan speed based on the first heat dissipation control algorithm. The fan controller is used to control the fan speed of the first server to the target fan speed. The PID parameter learning module is used to optimize the first heat dissipation control algorithm to obtain the second heat dissipation control algorithm. The performance evaluation module is used to perform data evaluation and analysis on the second operation data to obtain the heat dissipation evaluation result.
[0134] Furthermore, for the system based on server heat dissipation control, an embodiment of the present application also provides a schematic flow diagram of another server heat dissipation control method, as Figure 5 shown in the figure, which includes: S1: Import an initial heat dissipation strategy for the first server and preliminarily set PID regulation parameters (the first heat dissipation control algorithm parameters) for each key component, including: the temperature threshold of the key component, the temperature regulation point, etc. It should be noted that when controlling the heat dissipation of the first server, the initial parameters can be provided in advance, or the first heat dissipation control algorithm parameters can be obtained from the preset edge computing device.
[0135] S2: Real-time collect multi-source data of the first server hardware.
[0136] The BMC real-time collects the temperature data, fan speed data, load data, and ambient temperature data of the server hardware through high-precision temperature sensors, fan speed sensors, load sensors, and ambient temperature sensors connected to the server motherboard. The temperature data includes the temperatures of key hardware such as the CPU, memory, and hard disk. The fan speed data includes the speeds of each fan inside the server. The temperature data accurately reflects the real-time temperature status of the key hardware, the fan speed data monitors the operating state of the heat dissipation fans, the load data reflects the working intensity of the server, and the ambient temperature data provides external environmental information.
[0137] S3: Calculate the target speed based on deep reinforcement learning.
[0138] Input the collected multi-source data into the PID control algorithm based on deep reinforcement learning (the first heat dissipation control algorithm). Among them, a deep neural network with multiple hidden layers serves as the policy network, taking the state space (i.e., real-time operation data) as the input and outputting the optimal PID parameters (optimal parameters of the first heat dissipation control algorithm) in the current state. According to the optimal PID parameters, combined with the deviation between the current temperature and the target temperature, the rate of change of the deviation, and the integral value (dynamic temperature threshold), calculate the target fan speed. During the training process of the deep reinforcement learning algorithm, continuously optimize the policy network according to the feedback of the reward function, and gradually learn the optimal control strategy.
[0139] S4: Send the fan speed adjustment amount to the fan controller to adjust the fan speed.
[0140] The BMC sends the calculated fan speed adjustment amount to the fan controller, and the fan controller adjusts the fan speed according to the received adjustment amount to achieve heat dissipation control.
[0141] S5: Training sample collection and update.
[0142] During the heat dissipation control process, the BMC continuously collects the temperature data, fan speed data, load data, and ambient temperature data after heat dissipation control, forms these data into training samples, and stores them in the local training dataset. As time goes by, the training dataset is continuously updated, containing more data under different working conditions, providing rich data support for the continuous optimization of the deep reinforcement learning algorithm.
[0143] S6: Optimize the PID parameters by deep reinforcement learning.
[0144] The BMC regularly trains the first heat dissipation control algorithm based on the deep reinforcement learning mechanism using the latest training samples. During the training process, the algorithm continuously adjusts the parameters of the policy network through interaction with the environment (i.e., the heat dissipation control process) to maximize the cumulative reward. When the algorithm converges or reaches the preset number of training times, the optimized PID parameters (parameters of the second heat dissipation control algorithm) are obtained for subsequent heat dissipation control.
[0145] S7: Collaborative processing of edge computing and cloud computing.
[0146] The edge computing device receives the data collected by the BMC in real time and analyzes the data in real time. For simple changes in heat dissipation requirements, such as small temperature fluctuations or slight load changes, the edge computing device quickly adjusts the parameters of the second heat dissipation control algorithm and the fan speed according to the locally preset rules and lightweight models. When encountering complex heat dissipation scenarios, such as a sudden significant increase in server load or a sharp change in ambient temperature, the edge computing device uploads the key data to the cloud computing platform. The cloud computing platform uses big data analysis and complex machine learning models to deeply process the data, generates optimized control strategies, and feeds back the strategies to the edge computing device, which executes the corresponding adjustment operations.
[0147] S8: Blockchain distributed heat dissipation management.
[0148] In a multi-server cluster environment, the BMCs of each server share heat dissipation-related data through the blockchain network. At regular time intervals, each BMC encapsulates the local temperature data, fan speed data, load information, heat dissipation control algorithm parameters, etc. into blocks and broadcasts and verifies them in the blockchain network through a consensus mechanism. Based on the shared data on the blockchain, the distributed heat dissipation optimization algorithm runs regularly, calculates the optimal fan speed and the adjustment plan of the heat dissipation control algorithm parameters for each server according to the real-time status of each server. The BMCs of each server adjust the local heat dissipation control strategy according to the algorithm results to achieve coordinated heat dissipation of multiple servers.
[0149] S9: Adaptive dynamic temperature threshold adjustment.
[0150] The BMC uses time series analysis algorithms to analyze the historical temperature data of the key components of the server, and extracts the periodic and trend characteristics of the temperature change. Combining the load change pattern of the server and the ambient temperature change trend, it predicts the temperature change in the next period of time through machine learning algorithms. According to the prediction results, the temperature threshold of the key components is dynamically adjusted. For example, when it is predicted that the server load will increase significantly, the temperature threshold is lowered in advance to start the fan speed regulation in time; when it is predicted that the ambient temperature will drop, the temperature threshold is appropriately increased to reduce the unnecessary operation of the fan.
[0151] S10: Performance evaluation and strategy optimization.
[0152] The BMC regularly evaluates the heat dissipation control effect according to the preset evaluation indicators, such as temperature stability, temperature fluctuation range, fan speed adjustment frequency, energy consumption, etc. By comparing and analyzing with historical data and ideal indicators, it judges the advantages and disadvantages of the current heat dissipation control strategy. According to the evaluation results, adjust the hyperparameters of the deep reinforcement learning algorithm, optimize the cooperation strategy between edge computing and cloud computing, improve the distributed heat dissipation management algorithm or adjust the calculation method of the adaptive temperature threshold, continuously optimize the heat dissipation control strategy, and improve the performance of the heat dissipation system.
[0153] In summary, the embodiments of the present application can achieve the following technical effects: 1. First, perform heat dissipation control processing on the server according to the server's operating data and the heat dissipation control algorithm. Then, evaluate the server after heat dissipation control according to the actual heat dissipation requirements of the server, i.e., the preset heat dissipation conditions. Optimize the heat dissipation control algorithm based on the evaluation results and the operating data of the server after heat dissipation control. At the same time, the optimization of the heat dissipation control algorithm introduces a deep reinforcement learning mechanism. By collaborating multiple algorithms and combining with edge computing, i.e., the preset edge computing device, the balance problem of difficult parameter tuning and poor system adaptability faced in traditional heat dissipation control is solved. Therefore, the problems of difficult parameter tuning and poor system adaptability in server heat dissipation control can be solved, achieving the technical effect of improving the efficiency and stability of the heat dissipation system.
[0154] 2. The introduction of deep reinforcement learning in the present application enables the PID controller to have high intelligence and adaptability, and can automatically optimize the PID parameters under complex and changeable working conditions, realizing precise heat dissipation control and significantly improving the control efficiency and stability of the system.
[0155] 3. Through the collaborative decision-making mode of edge computing and cloud computing in the present application, the real-time performance and accuracy of heat dissipation control are ensured. The edge computing device quickly processes local data, reducing the response delay; the cloud computing platform provides powerful computing and analysis capabilities to handle complex heat dissipation scenarios. The combination of the two improves the overall performance of heat dissipation control.
[0156] 4. The present application realizes the collaborative work between multiple servers through the blockchain-based distributed heat dissipation management system, effectively improving the heat dissipation efficiency of the server cluster and reducing energy consumption. Blockchain technology ensures the secure sharing and consistency of data, enhancing the reliability and scalability of the system.
[0157] 5. Through the adaptive dynamic temperature threshold adjustment strategy in the present application, the temperature threshold is dynamically adjusted according to the actual operating conditions of the server and environmental changes, realizing more precise heat dissipation control, improving the energy utilization rate of the server, reducing unnecessary fan speed regulation, and reducing equipment loss.
[0158] 6. The present application realizes the automation and intelligence of heat dissipation control, reducing the cost and risk of manual intervention. Efficient heat dissipation control ensures the stable operation of the server, improves the performance and service life of the server, and brings significant economic benefits to the operation of the data center.
[0159] 7. This application integrates cutting-edge technologies such as deep reinforcement learning, edge computing, cloud computing, and blockchain, bringing innovative solutions to the field of server heat dissipation control, promoting the application and development of related technologies in this field, and enhancing the technical level of the entire industry.
[0160] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0161] The embodiment of this application also provides a device for server heat dissipation control. Figure 6 As shown in the structural schematic diagram of a device for server heat dissipation control provided by this application, Figure 6 it includes: An input unit 61, configured to input the real-time operation data of the first server and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server, so as to obtain the target fan speed, where the dynamic temperature threshold is a temperature threshold obtained by analyzing the historical operation data of the first server; An adjustment unit 62, configured to perform fan speed adjustment processing on the first server according to the target fan speed to obtain the adjusted first server and the second operation data of the adjusted first server; An analysis unit 63, configured to perform data evaluation and analysis processing on the second operation data to obtain the heat dissipation evaluation result corresponding to the adjusted first server; An optimization unit 64, configured to, when the heat dissipation evaluation result does not meet the preset heat dissipation condition, perform algorithm optimization processing on the first heat dissipation control algorithm based on the second operation data, the heat dissipation evaluation result, and the historical operation data based on the deep reinforcement learning mechanism to obtain the second heat dissipation control algorithm; A control unit 65, configured to perform heat dissipation control processing on the adjusted first server based on the second heat dissipation control algorithm, the second operation data, and the dynamic temperature threshold; A transmission unit 66, configured to transmit the second heat dissipation control algorithm to a preset edge computing device.
[0162] In an embodiment of this application, as Figure 7 shown, the device for server heat dissipation control further includes: A first processing unit 67, configured to obtain the historical operation data of the first server and perform data analysis processing on the historical operation data to obtain the temperature change cycle data, temperature change trend data, load change information, and ambient temperature information corresponding to the first server, where the historical operation data is the operation data of the first server before the real-time operation data; The first processing unit 67 is further configured to perform temperature prediction processing on the first server according to the load change information, the ambient temperature information, the temperature change cycle data, and the temperature change trend data to obtain target temperature change data; The first processing unit 67 is further configured to, when it is determined according to the target temperature change data that the temperature of the first server rises, set the dynamic temperature threshold to the first threshold; The first processing unit 67 is further configured to, when it is determined according to the target temperature change data that the temperature of the first server drops, set the dynamic temperature threshold to the second threshold, where the second threshold is greater than the first threshold.
[0163] In an embodiment of the present application, the input unit 61 is further configured to: Perform data comparison processing on the server temperature in the real-time operation data and the dynamic temperature threshold to obtain a first comparison result, where the server temperature is the real-time temperature of the first server; When it is determined according to the first comparison result that the server temperature is greater than the dynamic temperature threshold, input the real-time operation data and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a first fan speed, where the first heat dissipation control algorithm is a heat dissipation control algorithm based on deep reinforcement learning extracted from a preset edge computing device, When it is determined according to the first comparison result that the server temperature is less than or equal to the dynamic temperature threshold, input the real-time operation data and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a second fan speed; Wherein, the target fan speed includes the first fan speed and the second fan speed, and the second fan speed is less than the first fan speed.
[0164] In an embodiment of the present application, the input unit 61 is further configured to: Calculate a first temperature difference between the server temperature and the dynamic temperature threshold; Input the first temperature difference, the real-time operation data, and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a first fan speed, where the magnitude of the first temperature difference is positively correlated with the magnitude of the first fan speed.
[0165] In an embodiment of the present application, the input unit 61 is further configured to: Calculate a second temperature difference between the server temperature and the dynamic temperature threshold; Input the second temperature difference, the real-time operation data, and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a second fan speed, where the magnitude of the second temperature difference is negatively correlated with the magnitude of the second fan speed.
[0166] In an embodiment of the present application, the adjustment unit 62 is further configured to: Transmit the target fan speed to the fan controller of the first server, and control the fan speed of the first server to the target fan speed through the fan controller, so as to obtain the adjusted first server; When it is determined that the fan speed of the first server is the target fan speed, obtain the operation data of the first server to obtain the second operation data.
[0167] In an embodiment of the present application, the analysis unit 63 is further configured to perform data analysis processing on the second operation data according to the historical operation data and the preset evaluation indicators to obtain a heat dissipation evaluation result, where the heat dissipation evaluation result at least includes a temperature evaluation result, a speed evaluation result, and an energy consumption evaluation result.
[0168] In an embodiment of the present application, as Figure 7 shown, the server heat dissipation control device further includes: A second processing unit 68, configured to compare the temperature evaluation result with the temperature evaluation indicator in the preset evaluation indicators to obtain a temperature evaluation deviation value, where the preset evaluation indicators at least include a speed evaluation indicator, an energy consumption evaluation indicator, and a temperature evaluation indicator; The second processing unit 68 is further configured to compare the speed evaluation result with the speed evaluation indicator to obtain a speed evaluation deviation value, and compare the energy consumption evaluation result with the energy consumption evaluation indicator to obtain an energy consumption evaluation deviation value; The second processing unit 68 is further configured to determine whether the temperature evaluation result meets the preset heat dissipation conditions according to the temperature evaluation deviation value, determine whether the speed evaluation result meets the preset heat dissipation conditions according to the speed evaluation deviation value, and determine whether the energy consumption evaluation result meets the preset heat dissipation conditions according to the energy consumption evaluation deviation value; The second processing unit 68 is further configured to determine that the heat dissipation evaluation result does not meet the preset heat dissipation conditions when any one of the temperature evaluation result, the speed evaluation result, and the energy consumption evaluation result does not meet the preset heat dissipation conditions.
[0169] In an embodiment of the present application, the second processing unit 68 is further configured to: When the temperature evaluation deviation value is greater than the first preset evaluation deviation threshold, determine that the temperature evaluation result does not meet the preset heat dissipation conditions; When the speed evaluation deviation value is greater than the second preset evaluation deviation threshold, determine that the speed evaluation result does not meet the preset heat dissipation conditions; When the energy consumption evaluation deviation value is greater than the third preset evaluation deviation threshold, determine that the energy consumption evaluation result does not meet the preset heat dissipation conditions.
[0170] In an embodiment of the present application, the optimization unit 64 is further configured to: Perform algorithm optimization processing on the first heat dissipation control algorithm according to the second operation data, the heat dissipation evaluation result, and the historical operation data through a preset deep reinforcement learning algorithm; Until the preset optimization times or the preset convergence condition is reached, the second heat dissipation control algorithm is obtained.
[0171] In an embodiment of the present application, as Figure 7 shown, the device for server heat dissipation control further includes: A third processing unit 69, configured to, in response to a control instruction for heat dissipation control of a second server, extract the second heat dissipation control algorithm from a preset edge computing device to perform heat dissipation control processing on the second server.
[0172] In an embodiment of the present application, the third processing unit 69 is further configured to: Obtain the third operation data of the second server, and perform data comparison processing according to the third operation data and the real-time operation data to obtain the operation difference data between the first server and the second server; When the operation difference data meets the preset first operation difference condition, extract the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server; When the operation difference data does not meet the preset first operation difference condition and meets the second operation difference condition, adjust the second heat dissipation control algorithm according to the preset algorithm update rule in the preset edge computing device to obtain the adjusted second heat dissipation control algorithm, and perform heat dissipation control processing on the second server according to the adjusted second heat dissipation control algorithm; When the operation difference data does not meet the second operation difference condition, obtain the second historical data of the second server, and perform algorithm optimization processing on the second heat dissipation control algorithm according to the second historical data and the third operation data to obtain the optimized second heat dissipation control algorithm; Perform heat dissipation control processing on the second server according to the optimized second heat dissipation control algorithm.
[0173] For the description of the features in the corresponding embodiment of the device for server heat dissipation control, reference can be made to the relevant description in the corresponding embodiment of the method for server heat dissipation control, which will not be elaborated here one by one.
[0174] An embodiment of the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments for server heat dissipation control.
[0175] Embodiments of the present application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the method embodiments of the above server heat dissipation control when running.
[0176] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs that can store computer programs.
[0177] Embodiments of the present application also provide a computer program product, where the computer program product includes a computer program, and the steps in any of the method embodiments of the above server heat dissipation control are implemented when the computer program is executed by a processor.
[0178] Embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and the steps in any of the method embodiments of the above server heat dissipation control are implemented when the computer program is executed by a processor.
[0179] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0180] The above has introduced in detail a method, an electronic device, and a storage medium for server heat dissipation control provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for controlling heat dissipation of a server, characterized in that: include: Inputting the real-time operation data and the dynamic temperature threshold of the first server into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain a target fan speed, wherein the dynamic temperature threshold is a temperature threshold obtained by analyzing the historical operation data of the first server; Performing fan speed adjustment processing on the first server according to the target fan speed to obtain an adjusted first server and second operating data of the adjusted first server, and performing data evaluation and analysis processing on the second operating data to obtain a heat dissipation evaluation result corresponding to the adjusted first server; When the heat dissipation evaluation result does not meet the preset heat dissipation condition, the first heat dissipation control algorithm is optimized based on a deep reinforcement learning mechanism according to the second operation data, the heat dissipation evaluation result and the historical operation data to obtain a second heat dissipation control algorithm; Based on the second heat dissipation control algorithm, the second operating data and the dynamic temperature threshold, heat dissipation control processing is performed on the adjusted first server, and the second heat dissipation control algorithm is transmitted to a preset edge computing device.
2. The method for controlling heat dissipation of a server according to claim 1, characterized in that: Before inputting the real-time operation data and the dynamic temperature threshold of the first server into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain the target fan speed, the method further includes: Acquire the historical operation data of the first server, and perform data analysis and processing on the historical operation data to obtain temperature change cycle data, temperature change trend data, load change information, and ambient temperature information corresponding to the first server, wherein the historical operation data is the operation data of the first server before the real-time operation data; Performing temperature prediction processing on the first server according to the load change information, the ambient temperature information, the temperature change cycle data, and the temperature change trend data to obtain target temperature change data; In a case where it is determined according to the target temperature change data that the temperature of the first server increases, setting the dynamic temperature threshold to a first threshold; In a case where it is determined according to the target temperature change data that the temperature of the first server decreases, the dynamic temperature threshold is set to a second threshold, wherein the second threshold is greater than the first threshold.
3. The method for controlling heat dissipation of a server according to claim 1, characterized in that: The step of inputting the real-time operation data and the dynamic temperature threshold of the first server into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the target fan speed includes: Performing data comparison processing on the server temperature in the real-time operation data and the dynamic temperature threshold to obtain a first comparison result, wherein the server temperature is the real-time temperature of the first server; In the case where it is determined according to the first comparison result that the temperature of the server is greater than the dynamic temperature threshold, the real-time operation data and the dynamic temperature threshold are input into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a first fan speed, wherein the first heat dissipation control algorithm is a heat dissipation control algorithm based on deep reinforcement learning extracted from a preset edge computing device; In the case where it is determined according to the first comparison result that the server temperature is less than or equal to the dynamic temperature threshold, the real-time operation data and the dynamic temperature threshold are input into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a second fan speed; The target fan speed includes the first fan speed and the second fan speed, and the second fan speed is less than the first fan speed.
4. The method for controlling heat dissipation of a server according to claim 3, characterized in that: Inputting the real-time operation data and the dynamic temperature threshold into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the first fan speed includes: Calculating a first temperature difference between the server temperature and the dynamic temperature threshold; The first temperature difference, the real-time operating data, and the dynamic temperature threshold are input into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the first fan speed, wherein the size of the first temperature difference is positively correlated with the size of the first fan speed.
5. The method for controlling heat dissipation of a server according to claim 3, characterized in that: Inputting the real-time operation data and the dynamic temperature threshold into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the second fan speed includes: calculating a second temperature difference between the server temperature and the dynamic temperature threshold; The second temperature difference, the real-time operating data, and the dynamic temperature threshold are input into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the second fan speed, wherein the size of the second temperature difference is negatively correlated with the size of the second fan speed.
6. The method for controlling heat dissipation of a server according to claim 1, characterized in that: The fan speed adjustment process is performed on the first server according to the target fan speed to obtain the adjusted first server and the second operation data of the adjusted first server, including: Transmitting the target fan speed to a fan controller of the first server, and controlling the fan speed of the first server to the target fan speed by the fan controller, to obtain the adjusted first server; When it is determined that the fan speed of the first server is the target fan speed, the operation data of the first server is acquired to obtain the second operation data.
7. The method for controlling heat dissipation of a server according to claim 1, characterized in that: The performing data evaluation and analysis on the second operation data to obtain the heat dissipation evaluation result corresponding to the adjusted first server includes: The second operating data is analyzed and processed according to the historical operating data and the preset evaluation index to obtain the heat dissipation evaluation result, wherein the heat dissipation evaluation result at least includes a temperature evaluation result, a rotation speed evaluation result and an energy consumption evaluation result.
8. The method for controlling heat dissipation of a server according to claim 7, characterized in that: After performing data evaluation and analysis on the second operating data to obtain a heat dissipation evaluation result corresponding to the adjusted first server, the method further includes: Comparing the temperature evaluation result with the temperature evaluation index in the preset evaluation index to obtain a temperature evaluation deviation value, wherein the preset evaluation index at least includes a rotation speed evaluation index, an energy consumption evaluation index and the temperature evaluation index; Comparing the speed evaluation result with the speed evaluation index to obtain a speed evaluation deviation value, and comparing the energy consumption evaluation result with the energy consumption evaluation index to obtain an energy consumption evaluation deviation value; Determine whether the temperature assessment result meets the preset heat dissipation condition according to the temperature assessment deviation value, determine whether the speed assessment result meets the preset heat dissipation condition according to the speed assessment deviation value, and determine whether the energy consumption assessment result meets the preset heat dissipation condition according to the energy consumption assessment deviation value; When any one of the temperature evaluation result, the rotation speed evaluation result, and the energy consumption evaluation result does not meet the preset heat dissipation condition, it is determined that the heat dissipation evaluation result does not meet the preset heat dissipation condition.
9. The method for controlling heat dissipation of a server according to claim 8, characterized in that: Determining whether the temperature evaluation result meets the preset heat dissipation condition according to the temperature evaluation deviation value, determining whether the speed evaluation result meets the preset heat dissipation condition according to the speed evaluation deviation value, and determining whether the energy consumption evaluation result meets the preset heat dissipation condition according to the energy consumption evaluation deviation value includes: In a case where the temperature evaluation deviation value is greater than a first preset evaluation deviation threshold, determining that the temperature evaluation result does not meet the preset heat dissipation condition; When the rotation speed evaluation deviation value is greater than a second preset evaluation deviation threshold, determining that the rotation speed evaluation result does not meet the preset heat dissipation condition; When the energy consumption evaluation deviation value is greater than a third preset evaluation deviation threshold, it is determined that the energy consumption evaluation result does not meet the preset heat dissipation condition.
10. The method for controlling heat dissipation of a server according to claim 1, characterized in that: The performing of algorithm optimization processing on the first heat dissipation control algorithm based on a deep reinforcement learning mechanism according to the second operation data, the heat dissipation evaluation result and the historical operation data to obtain a second heat dissipation control algorithm includes: Performing algorithm optimization processing on the first heat dissipation control algorithm by using a preset deep reinforcement learning algorithm according to the second operation data, the heat dissipation evaluation result and the historical operation data; Until a preset number of optimizations or a preset convergence condition is reached, the second heat dissipation control algorithm is obtained.
11. The method for controlling heat dissipation of a server according to claim 1, characterized in that: After transmitting the second heat dissipation control algorithm to the preset edge computing device, the method further includes: In response to a control instruction to perform heat dissipation control on the second server, the second heat dissipation control algorithm is extracted from the preset edge computing device to perform heat dissipation control processing on the second server.
12. The method for controlling heat dissipation of a server according to claim 11, characterized in that: Extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server includes: Acquire the third operation data of the second server, and perform data comparison processing according to the third operation data and the real-time operation data to obtain operation difference data between the first server and the second server; When the operation difference data satisfies a preset first operation difference condition, extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server; When the operation difference data does not satisfy the preset first operation difference condition but satisfies the second operation difference condition, adjusting the second heat dissipation control algorithm according to the preset algorithm update rule in the preset edge computing device to obtain an adjusted second heat dissipation control algorithm, and performing heat dissipation control processing on the second server according to the adjusted second heat dissipation control algorithm; When the operation difference data does not satisfy the second operation difference condition, obtaining second historical data of the second server, and performing algorithm optimization processing on the second heat dissipation control algorithm according to the second historical data and the third operation data to obtain an optimized second heat dissipation control algorithm; Perform heat dissipation control processing on the second server according to the optimized second heat dissipation control algorithm.
13. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of the method for controlling heat dissipation of a server as claimed in any one of claims 1 to 12 when executing the computer program.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for controlling heat dissipation of a server as claimed in any one of claims 1 to 12.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for controlling heat dissipation of a server as claimed in any one of claims 1 to 12 are implemented.
Citation Information
Patent Citations
Pure electric vehicle cooling system control method based on deep reinforcement learning
CN109193075A
Heat dissipation control method and device of server, storage medium and electronic device
CN116027868A
Server thermal management method based on artificial intelligence
CN117234301A
Heat dissipation control method and device of server, storage medium and electronic equipment
CN118567454A
Self-adaptive setting method and device for fan control parameters, storage medium and electronic equipment
CN118672368A
Cited By
Control method and device of heat dissipation equipment, storage medium and electronic equipment
CN120335582A
Control methods, devices, storage media and electronic equipment for heat dissipation equipment
CN120335582B
Cooling performance evaluation method and device of cold plate, storage medium and electronic equipment
CN120336146A
Cooling equipment control method and device, equipment and storage medium
CN120353134A
Heat dissipation device control method, device, equipment and storage medium
CN120353134B