Server heat dissipation control method, electronic device and storage medium
Through the thermal dissipation control algorithm optimized by deep reinforcement learning and edge computing, the problem of difficulty in parameter setting and poor adaptability in traditional server thermal control is solved, and more efficient and stable thermal dissipation effect is achieved, improving the performance and stability of the server.
Patent Information
- Application Number
- CN202510536003.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-27
AI Technical Summary
There are problems in traditional server cooling control methods such as difficulty in parameter setting and poor system adaptability, resulting in poor heat dissipation effect and affecting the performance and stability of the server.
The heat dissipation control algorithm based on deep reinforcement learning is adopted, combined with edge computing, and fan speed is optimized through real-time running data and dynamic temperature thresholds, multiple algorithms are used to coordinate heat dissipation control, and algorithm optimization is carried out when preset heat dissipation conditions are not met.
It improves the efficiency and stability of the server cooling system, solves the problems of difficulty in parameter setting and poor system adaptability, and achieves better heat dissipation effect and energy efficiency balance.
Smart Images

Figure CN120066922B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, electronic device, and storage medium for controlling heat dissipation in a server. Background Art
[0002] With the rapid development of information technology, servers, as core devices for data processing and storage, are crucial to business operations in terms of performance and stability. Heat dissipation has always been one of the key factors affecting server performance and stability during server operation. Traditional heat dissipation control methods are mostly based on fixed heat dissipation strategies, such as setting a fixed fan speed or adjusting the fan speed based on a preset temperature threshold. However, fixed heat dissipation strategies often fail to adapt to the heat dissipation requirements of servers under different loads and environments, resulting in poor heat dissipation and may even cause overheating, affecting the normal operation of the server.
[0003] To address the issue of poor heat dissipation, existing technologies offer heat dissipation control methods based on the Proportional-Integral-Derivative (PID) control algorithm. While these PID-based heat dissipation control methods can address the issue, most existing PID-based, custom-built heat dissipation control methods rely on traditional machine learning algorithms, which are relatively simple and have low optimization capabilities under complex operating conditions. Therefore, addressing the difficulties in parameter tuning and poor system adaptability faced by server heat dissipation control is a pressing issue. Summary of the Invention
[0004] The present application provides a method, electronic device and storage medium for controlling heat dissipation of a server, so as to at least solve the problems of difficult parameter setting and poor system adaptability faced in controlling heat dissipation of the server.
[0005] This application provides a method for controlling heat dissipation in a server, including:
[0006] Inputting the real-time operating data and the dynamic temperature threshold of the first server into a first heat dissipation control algorithm to calculate the fan speed of the first server to obtain a target fan speed, wherein the dynamic temperature threshold is a temperature threshold obtained by analyzing the historical operating data of the first server;
[0007] Adjusting the fan speed of the first server according to the target fan speed to obtain the adjusted first server and second operating data of the adjusted first server, and performing data evaluation and analysis on the second operating data to obtain a heat dissipation evaluation result corresponding to the adjusted first server;
[0008] When the heat dissipation evaluation result does not meet the preset heat dissipation condition, the first heat dissipation control algorithm is optimized based on the deep reinforcement learning mechanism according to the second operating data, the heat dissipation evaluation result and the historical operating data to obtain a second heat dissipation control algorithm;
[0009] The adjusted first server is subjected to heat dissipation control processing based on the second heat dissipation control algorithm, the second operating data, and the dynamic temperature threshold, and the second heat dissipation control algorithm is transmitted to the preset edge computing device.
[0010] The present application also provides a device for controlling heat dissipation of a server, comprising:
[0011] an input unit, configured to input the real-time operating data of the first server and a dynamic temperature threshold into a first heat dissipation control algorithm to calculate a fan speed of the first server to obtain a target fan speed, wherein the dynamic temperature threshold is a temperature threshold obtained by analyzing historical operating data of the first server;
[0012] an adjusting unit, configured to adjust the fan speed of the first server according to the target fan speed, to obtain the adjusted first server and second operating data of the adjusted first server;
[0013] an analyzing unit, configured to perform data evaluation and analysis on the second operating data to obtain an adjusted heat dissipation evaluation result corresponding to the first server;
[0014] an optimization unit, configured to optimize the first heat dissipation control algorithm based on a deep reinforcement learning mechanism according to the second operating data, the heat dissipation evaluation result, and the historical operating data to obtain a second heat dissipation control algorithm when the heat dissipation evaluation result does not meet the preset heat dissipation condition;
[0015] a control unit, configured to perform heat dissipation control processing on the adjusted first server based on the second heat dissipation control algorithm, the second operating data, and the dynamic temperature threshold;
[0016] A transmission unit is used to transmit the second heat dissipation control algorithm to a preset edge computing device.
[0017] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server heat dissipation control methods when executing the computer program.
[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server heat dissipation control methods are implemented.
[0019] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned server heat dissipation control methods when executed by a processor.
[0020] The server heat dissipation control method, electronic device and storage medium of the present application first performs heat dissipation control on the server according to the server's operating data and the heat dissipation control algorithm, and then evaluates the server after heat dissipation control according to the server's actual heat dissipation requirements, i.e., preset heat dissipation conditions, and optimizes the heat dissipation control algorithm based on the evaluation results and the operating data of the server after heat dissipation control. At the same time, the optimization of the heat dissipation control algorithm introduces a deep reinforcement learning mechanism, which solves the balance problem of parameter setting difficulties and poor system adaptability faced in traditional heat dissipation control through the collaboration of multiple algorithms and combined with edge computing, i.e., preset edge computing devices. Therefore, it can solve the problems of parameter setting difficulties and poor system adaptability faced in server heat dissipation control, and achieve the technical effect of improving the efficiency and stability of the heat dissipation system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 A flow chart of a method for controlling heat dissipation of a server provided in an embodiment of the present application;
[0023] Figure 2 A schematic diagram of a flow chart for determining a dynamic temperature threshold provided in an embodiment of the present application;
[0024] Figure 3 A flow chart for analyzing heat dissipation evaluation results provided in an embodiment of the present application;
[0025] Figure 4 A schematic diagram of a server heat dissipation control system provided in an embodiment of the present application;
[0026] Figure 5 A flow chart of another method for controlling heat dissipation in a server provided in an embodiment of the present application;
[0027] Figure 6 A schematic diagram of the structure of a server heat dissipation control device provided in an embodiment of the present application;
[0028] Figure 7 A schematic structural diagram of another server heat dissipation control device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0030] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0031] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0032] Figure 1 A flow chart of a method for controlling heat dissipation of a server provided in an embodiment of the present application is provided, and the method is described in detail in conjunction with the execution flow of the method for controlling heat dissipation of the server.
[0033] like Figure 1 As shown, the server heat dissipation control method includes:
[0034] In step 101, the real-time operating data and the dynamic temperature threshold of the first server are input into a first heat dissipation control algorithm to calculate the fan speed of the first server to obtain a target fan speed, wherein the dynamic temperature threshold is a temperature threshold obtained by analyzing the historical operating data of the first server.
[0035] In an embodiment of the present application, the first server is a server that requires heat dissipation control, such as an enterprise-level server, a super server, etc. Specifically, the present application does not limit the type, category, and structure of the first server.
[0036] Real-time operating data is operating data obtained by collecting data from the first server through the baseboard management controller (BMC) of the first server. Real-time operating data includes, but is not limited to, hardware data and status data. Acquisition of real-time operating data includes, but is not limited to, real-time collection of server hardware temperature data, fan speed data, load data, and ambient temperature data through high-precision temperature sensors, fan speed sensors, load sensors, and ambient temperature sensors connected to the motherboard of the first server.
[0037] Real-time operating data includes at least temperature data, fan speed data, load data, and ambient temperature data. Temperature data includes, but is not limited to, the temperatures of key hardware such as the central processing unit (CPU), memory, and hard disk. Fan speed data includes, but is not limited to, the speed of each fan within the first server. Temperature data accurately reflects the real-time temperature status of the first server's key hardware. Fan speed data monitors the operating status of the first server's cooling fans. Load data reflects the first server's workload, and ambient temperature data provides information about the external environment.
[0038] The dynamic temperature threshold isn't a fixed value; it's an adaptive threshold generated by analyzing the first server's historical operating data (e.g., temperature fluctuation patterns, peak load periods, and ambient temperature trends over the past 24 hours). For example, if the first server's load exhibits periodic fluctuations, machine learning algorithms (e.g., Long Short-Term Memory (LSTM) time series models) can be used to predict the first server's temperature trends over future periods. This allows the temperature threshold to be dynamically adjusted, allowing it to be relaxed (increased) during low-load periods to reduce fan power consumption, and tightened (decreased) during high-load periods to prevent overheating.
[0039] The first heat dissipation control algorithm is a custom selected heat dissipation control algorithm, for example, an algorithm based on a PID control framework. For ease of understanding, the first heat dissipation control algorithm will be described below using the algorithm based on the PID control framework as an example.
[0040] The parameters of the first cooling control algorithm (such as the proportional coefficient P, integral coefficient I, and differential coefficient D) are set empirically or obtained through pre-training with historical data. When calculating the target fan speed, the algorithm inputs the deviation between the temperature data in the real-time operating data and the dynamic temperature threshold, the rate of change of the deviation (the differential term), and the accumulated historical deviation (the integral term) into the PID formula. Combined with the weighted influence of load and ambient temperature, the algorithm comprehensively calculates the optimal fan speed to meet the current cooling requirements. During this process, the first cooling control algorithm not only considers immediate temperature adjustment but also introduces a prediction of future operating conditions through dynamic thresholds, avoiding the delayed response caused by traditional fixed thresholds.
[0041] The collected multi-source data (real-time operating data) is input into a PID control algorithm (the first cooling control algorithm) based on deep reinforcement learning. This algorithm, which includes a deep neural network with multiple hidden layers as its policy network, takes the state space (i.e., the multi-source data) as input and outputs the optimal PID parameter adjustment values (the first cooling control algorithm adjustment values) for the current state. The target fan speed is calculated based on the adjusted PID parameters, the deviation between the current temperature (the temperature data from the real-time operating data) and the target temperature (the dynamic temperature threshold), the rate of change of the deviation, and the integral value. During training, the deep reinforcement learning algorithm continuously optimizes the policy network based on feedback from the reward function, gradually learning the optimal control strategy.
[0042] Furthermore, it should be noted that, regardless of whether the temperature of the first server is too high, the target fan speed will be calculated using the first heat dissipation control algorithm to determine the optimal fan speed of the first server that meets the current heat dissipation requirement.
[0043] Step 102 , adjust the fan speed of the first server according to the target fan speed to obtain the adjusted first server and second operating data of the adjusted first server, and perform data evaluation and analysis on the second operating data to obtain a heat dissipation evaluation result corresponding to the adjusted first server.
[0044] In the embodiment of the present application, when performing the fan speed adjustment process, it is possible to adopt, but not limited to: sending the target fan speed to the fan controller of the first server, adjusting the fan speed, etc.
[0045] After adjusting the fan speed of the first server according to the target fan speed, the second operating data includes, but is not limited to, the adjusted temperature change curve, the actual fan speed, instantaneous energy consumption, and load status. Data evaluation and analysis of the second operating data can be performed using, but is not limited to, a preset evaluation model. Specifically, this application does not limit the preset evaluation model.
[0046] When the second operating data is evaluated and analyzed using a preset evaluation model, heat dissipation evaluation results may be generated from, but are not limited to, the following dimensions:
[0047] Temperature stability: Whether the temperature of key components converges quickly within the dynamic threshold range and whether the fluctuation amplitude is lower than the allowable upper limit; Energy efficiency: Whether fan speed adjustment minimizes power consumption while meeting heat dissipation requirements, avoiding frequent starts and stops or continuous high-speed operation; Response timeliness: Whether the time delay from temperature exceeding the limit to fan speed adjustment taking effect is within the system tolerance range; Environmental adaptability: Whether the heat dissipation strategy effectively offsets the impact of sudden changes in ambient temperature (such as: air conditioning failure in the computer room).
[0048] The heat dissipation evaluation results can be presented in the form of quantitative scores or classification labels (such as excellent, medium, and poor). Specifically, this application does not limit the presentation form of the heat dissipation evaluation results.
[0049] Furthermore, during the cooling control process, the first server's BMC continuously collects temperature data, fan speed data, load data, and ambient temperature data from the adjusted first server. This data is combined to generate secondary operating data and stored in a local training dataset. Over time, the training dataset is continuously updated to include data from more diverse operating conditions, providing rich data support for the continuous optimization of deep reinforcement learning algorithms.
[0050] Step 103: When the heat dissipation evaluation result does not meet the preset heat dissipation conditions, the first heat dissipation control algorithm is optimized based on a deep reinforcement learning mechanism according to the second operating data, the heat dissipation evaluation result, and the historical operating data to obtain a second heat dissipation control algorithm.
[0051] In an embodiment of the present application, the preset heat dissipation conditions are custom-set heat dissipation conditions, such as: temperature safety range, maximum allowable energy consumption, response time threshold, etc. Specifically, this application does not impose any restrictions on the preset heat dissipation conditions.
[0052] The thermal assessment results can be compared against pre-set thermal conditions (e.g., temperature safety range, maximum allowable power consumption, and response time threshold). For example, if the CPU temperature of server 1 remains above the dynamic temperature threshold after adjustment, or if the fan speed of server 1 fluctuates significantly and frequently, causing power consumption to exceed the maximum allowable power consumption, the thermal assessment result will be marked as not meeting the pre-set thermal conditions.
[0053] If the heat dissipation assessment result indicates that the system does not meet the preset heat dissipation conditions, deep reinforcement learning is used to optimize the PID parameters (parameters of the first heat dissipation control algorithm). The BMC regularly trains the deep reinforcement learning algorithm using the latest training samples. During training, the algorithm continuously adjusts the policy network parameters through interaction with the environment (i.e., the heat dissipation control process) to maximize the cumulative reward. When the algorithm converges or reaches the preset number of training cycles, the optimized first heat dissipation control algorithm is obtained and used for subsequent heat dissipation control.
[0054] Furthermore, if the heat dissipation assessment fails to meet the standards, a Deep Reinforcement Learning (DRL)-based algorithm optimization process is triggered. The second-run data and historical data together form a training sample set, which is fed into the DRL algorithm for policy iteration. The DRL algorithm's reward function is designed as a multi-objective optimization method: positive rewards include: temperature returning to a safe range, energy consumption reduction, and improved response speed; negative penalties include: temperature continuously exceeding the standard, fan speed fluctuations, and energy consumption exceeding the limit.
[0055] The DRL algorithm uses an exploration-exploitation mechanism to simulate the effects of different parameter combinations of the primary cooling control algorithm on cooling performance in a simulation environment. It gradually learns how to dynamically adjust the primary cooling control algorithm's parameters based on real-time conditions (such as sudden load increases or ambient temperature rises). For example, in scenarios with severe load fluctuations, it tends to increase the differential coefficient to quickly suppress temperature fluctuations, while reducing the integral coefficient to prevent overshoot. The optimized secondary cooling control algorithm not only corrects the flaws of the original algorithm but may also introduce new control logic (such as load prediction feedforward compensation) to improve robustness under complex operating conditions.
[0056] Step 104: Perform heat dissipation control processing on the adjusted first server based on the second heat dissipation control algorithm, the second operating data, and the dynamic temperature threshold, and transmit the second heat dissipation control algorithm to the preset edge computing device.
[0057] In an embodiment of the present application, the optimized algorithm, i.e., the second heat dissipation control algorithm, is first applied to the adjusted first server to verify its effectiveness in a real-world environment. Simultaneously, edge collaboration and algorithm deployment are performed. Specifically, this can be achieved through, but not limited to, the following methods: the second heat dissipation control algorithm is transmitted via an encrypted channel to a preset edge computing device, where the preset edge computing device is a custom storage device, such as an edge server deployed locally in a computer room, a built-in computing module in a BMC, or an external cloud server device.
[0058] The pre-configured edge computing device stores and updates a second cooling control algorithm, enabling it to quickly respond to similar cooling issues. For cross-server collaborative cooling needs (e.g., multiple servers sharing a cooling duct), the pre-configured edge computing device can leverage blockchain technology to exchange cooling strategies and real-time status with other nodes, enabling distributed decision-making. For example, if the pre-configured edge computing device detects a surge in load on a neighboring server, it can proactively fine-tune the fan speed of that server to prevent local overheating caused by hot air backflow.
[0059] The server heat dissipation control method, electronic device and storage medium of the present application first performs heat dissipation control on the server according to the server's operating data and the heat dissipation control algorithm, and then evaluates the server after heat dissipation control according to the server's actual heat dissipation requirements, i.e., preset heat dissipation conditions, and optimizes the heat dissipation control algorithm based on the evaluation results and the operating data of the server after heat dissipation control. At the same time, the optimization of the heat dissipation control algorithm introduces a deep reinforcement learning mechanism, which solves the balance problem of parameter setting difficulties and poor system adaptability faced in traditional heat dissipation control through the collaboration of multiple algorithms and combined with edge computing, i.e., preset edge computing devices. Therefore, it can solve the problems of parameter setting difficulties and poor system adaptability faced in server heat dissipation control, and achieve the technical effect of improving the efficiency and stability of the heat dissipation system.
[0060] In one possible implementation of the embodiment of the present application, before performing fan speed calculation processing to obtain the target fan speed, it is necessary to first determine the size of the dynamic temperature threshold in order to provide conditions for calculating the target fan speed. Specifically, regarding the determination of the dynamic temperature threshold, the embodiment of the present application provides a schematic flow chart of determining the dynamic temperature threshold, as shown in FIG. Figure 2 As shown, including:
[0061] Step 201, obtain the historical operating data of the first server, and perform data analysis and processing on the historical operating data to obtain the temperature change cycle data, temperature change trend data, load change information and ambient temperature information corresponding to the first server, wherein the historical operating data is the operating data of the first server before the real-time operating data.
[0062] In an embodiment of the present application, historical operation data refers to the operation records accumulated by the first server before the real-time operation data collection time point, including but not limited to: periodic data of several days to several months, such as: historical temperature curves (such as: minute-level temperature sampling of CPU and memory), load change logs (such as: time series records of CPU utilization and memory occupancy), fan speed history, ambient temperature records (such as: historical data of temperature and humidity sensors in the computer room) and task scheduling information of the first server (such as: batch job time period, high-load task triggering events), etc.
[0063] Data analysis and processing of historical operating data may employ, but is not limited to, multimodal analysis methods. Specifically, analysis of temperature variation periodicity data may employ, but is not limited to, the following methods: identifying the periodic fluctuation pattern of the first server's temperature through Fourier transform or a periodicity detection algorithm. For example, a data center's nighttime backup tasks may result in a peak load in the early morning hours, resulting in a 24-hour periodic increase in CPU temperature.
[0064] Temperature trend analysis can be done using, but is not limited to, linear regression or trend decomposition algorithms (such as seasonal and trend decomposition using Loess (STL)) to extract long-term temperature trends. For example, a gradual decrease in heat dissipation efficiency due to hardware aging may cause the temperature baseline to increase annually under the same load.
[0065] Load change information analysis can be performed using, but is not limited to, the following methods: analyzing the correlation between the load and temperature of the first server and establishing a load-temperature mapping model. For example, GPU-intensive tasks may cause a sharp increase in the temperature of a specific hardware area, while input / output (I / O)-intensive tasks have less impact on the temperature.
[0066] The analysis of ambient temperature information can be carried out in the following ways, but is not limited to: counting the average daily changes, seasonal fluctuations and abnormal events (such as a sudden temperature rise caused by an air conditioning failure) of the ambient temperature in the computer room where the first server is located, and evaluating the weight of the impact of the external environment on the server's heat dissipation.
[0067] Step 202 : Perform temperature prediction processing on the first server according to the load change information, the ambient temperature information, the temperature change cycle data, and the temperature change trend data to obtain target temperature change data.
[0068] In an embodiment of the present application, when performing temperature prediction processing, a hybrid model that integrates time series prediction and causal reasoning can be used for processing, but is not limited to: a hybrid model that integrates time series prediction and causal reasoning can be used for processing. Specifically, a time series prediction model (such as LSTM, Transformer) can be used: using historical temperature series and load cycle data as input, predict the temperature change trajectory in the future (within the next hour).
[0069] Causal inference model: Analyzes the causal relationship between ambient temperature, sudden load events (such as virtual machine migration and burst computing tasks), and temperature fluctuations, quantifying the contribution of external factors to temperature. For example, if a 1°C increase in ambient temperature causes a 0.5°C increase in CPU temperature, the impact of the ambient temperature increase is factored into the prediction.
[0070] Uncertainty modeling: Evaluate the confidence interval of the prediction results through Monte Carlo simulation or Bayesian neural network to avoid misjudgment caused by a single prediction value.
[0071] The resulting target temperature change data includes, but is not limited to, the predicted temperature curve, confidence interval, and key inflection points (e.g., the expected time and magnitude of the temperature peak). For example, the target temperature change data can be interpreted as predicting that the CPU temperature will rise by 8°C in the next 30 minutes due to a newly started training task, and will return to the baseline after 45 minutes.
[0072] Step 203 : When it is determined according to the target temperature change data that the temperature of the first server has increased, the dynamic temperature threshold is set to a first threshold.
[0073] In an embodiment of the present application, the first threshold is a threshold set when the first server is in a temperature-rising scenario. For example, when the target temperature change data indicates that the temperature of the first server will rise significantly (e.g., the predicted temperature rise exceeds the safety baseline by more than 3°C), the dynamic temperature threshold is set to a lower first threshold. The first threshold is usually close to the buffer value of the hardware safety upper limit (e.g., set to 85°C instead of the maximum allowable 90°C), and prevents potential overheating risks by triggering fan speed increase or load migration in advance. For example, when it is predicted that the GPU is about to heat up due to a rendering task, the threshold is lowered from 88°C to 83°C in advance, allowing the cooling system to intervene and regulate before the temperature approaches the critical value.
[0074] Step 204 : When it is determined according to the target temperature change data that the temperature of the first server decreases, the dynamic temperature threshold is set to a second threshold, wherein the second threshold is greater than the first threshold.
[0075] In an embodiment of the present application, the second threshold is a threshold set when the first server is in a temperature reduction scenario. For example, when the target temperature change data shows that the temperature of the first server will drop (such as during a low-load period at night or when the ambient air conditioning cooling is enhanced), the dynamic temperature threshold is raised to a higher second threshold (such as from 83°C to 88°C). A higher threshold allows the fan to maintain a lower speed within a safe range, reducing unnecessary energy consumption. For example, if the load is predicted to drop below 10% after 2 a.m. and the ambient temperature drops by 5°C, the system can relax the threshold to 87°C, causing the fan to run at 40% speed instead of the default 60%, thereby achieving quietness and energy saving.
[0076] Using a time series analysis algorithm, we analyze historical temperature data from key server components to identify periodic and trending characteristics of temperature changes. Combining server load patterns and ambient temperature trends, we use a machine learning algorithm to predict future temperature changes. Based on these predictions, we dynamically adjust the temperature thresholds for key components. For example, if a significant increase in server load is predicted, we lower the temperature threshold in advance to enable fan speed control. If a drop in ambient temperature is predicted, we raise the temperature threshold appropriately to reduce unnecessary fan operation.
[0077] Through dynamic threshold adjustment based on historical data analysis and prediction, a closed-loop control logic of prediction, prevention, and regulation is established. By adjusting the cooling strategy in advance based on prediction results before the temperature actually exceeds the limit, transient temperature overshoot or frequent fan startup and shutdown caused by delayed response is avoided. Through dynamic threshold relaxation (second threshold), low-power operation time is maximized while ensuring safety, reducing the overall power usage effectiveness (PUE) of the data center. In response to sudden loads or environmental anomalies (such as air conditioning failures), confidence interval assessment and rapid threshold tightening (first threshold) are used to improve the system's tolerance to uncertainty.
[0078] In one implementable method of an embodiment of the present application, when performing fan speed calculation processing to obtain a target fan speed, it can also be implemented by, but not limited to, the following method: performing data comparison processing based on the server temperature in the real-time operation data and the dynamic temperature threshold to obtain a first comparison result, wherein the server temperature is the real-time temperature of the first server; when it is determined according to the first comparison result that the server temperature is greater than the dynamic temperature threshold, the real-time operation data and the dynamic temperature threshold are input into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain the first fan speed, wherein the first heat dissipation control algorithm is a heat dissipation control algorithm based on deep reinforcement learning extracted from a preset edge computing device, and when it is determined according to the first comparison result that the server temperature is less than or equal to the dynamic temperature threshold, the real-time operation data and the dynamic temperature threshold are input into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain the second fan speed; wherein the target fan speed includes the first fan speed and the second fan speed, and the second fan speed is less than the first fan speed.
[0079] In an embodiment of the present application, the server temperature in the real-time operation data refers to the current temperature values of the key components (such as CPU, GPU, memory) of the first server collected in real time by the baseboard management controller (BMC), which is usually updated at a frequency of seconds or milliseconds.
[0080] The comparison of the real-time operation data with the dynamic temperature threshold includes, but is not limited to: comparing the server temperature in the real-time operation data with the dynamic temperature threshold, and outputting a first comparison result. The first comparison result includes two states:
[0081] An over-limit state occurs when the server temperature exceeds the dynamic temperature threshold. When the temperature of any key component of a server exceeds the dynamic temperature threshold, the server is considered to be in an over-limit state. For example, if the dynamic temperature threshold is set to 85°C and the real-time CPU temperature is 87°C, an over-limit state is triggered.
[0082] A safe state refers to a state where the server temperature is less than or equal to the dynamic temperature threshold. A server is considered safe when all key component temperatures are less than or equal to the dynamic temperature threshold. For example, a server is considered safe if the memory temperature is 83°C (dynamic temperature threshold 85°C) and the GPU temperature is 80°C (dynamic temperature threshold 85°C).
[0083] Differentiated calculation of the first and second fan speeds involves generating target fan speeds using different control strategies based on the first comparison results. For the first fan speed in an over-limit state, when the server temperature exceeds the dynamic temperature threshold, the system enters high-response cooling mode. Real-time operating data (including the over-limit temperature, load intensity, and ambient temperature) and the dynamic temperature threshold are input into the first cooling control algorithm. The first cooling control algorithm is a deep reinforcement learning-based PID algorithm extracted from a pre-defined edge computing device. Based on historical training data and real-time operating data, it outputs a fan speed command (i.e., the first fan speed) tailored to the over-limit scenario. For example, if the CPU temperature exceeds the limit and the load remains high, the algorithm may calculate a higher speed (e.g., 5000 revolutions per minute (RPM)) and apply a dynamic adjustment factor (e.g., an additional 200 RPM based on rising ambient temperature) to quickly suppress the temperature increase.
[0084] The second fan speed in the safe state is calculated as follows: when the temperature is within the safe range, that is, the server temperature is less than or equal to the dynamic temperature threshold, it switches to energy efficiency priority mode. At this time, the same real-time operating data and dynamic temperature threshold are input into the first heat dissipation control algorithm, but the algorithm will output a lower second fan speed based on the characteristics of the safe state (such as the difference between the temperature and the threshold, and the load reduction trend). For example: If the CPU temperature is 82°C (threshold 85°C) and the load is stable at 30%, the algorithm may set the speed to 3000 RPM, which is only 60% of the first speed. In this mode, the algorithm uses the energy consumption optimization goal in the reinforcement learning reward mechanism to actively reduce the speed to reduce power consumption, while ensuring that the temperature always remains in the safe range.
[0085] Among them, the first heat dissipation control algorithm is deployed on the preset edge computing device. The DRL network of the first heat dissipation control algorithm can be pruned and quantized through model lightweighting, thereby reducing the computational complexity while retaining the core decision-making capabilities, so that it can run in real time on embedded hardware (such as BMC).
[0086] Through localized incremental training, the first heat dissipation control algorithm is regularly fine-tuned using new data collected locally, gradually adapting to long-term changing factors such as server hardware aging and radiator dust accumulation, thereby avoiding performance degradation of the first heat dissipation control algorithm.
[0087] An abnormality fuse mechanism is provided. When the algorithm calculation is abnormal (for example, the output speed exceeds the physical limit of the fan), the preset edge computing device automatically switches to the preset conservative control strategy (for example, fixed PID parameter mode) to ensure system safety.
[0088] Dynamic hierarchical control of cooling strategies is achieved through temperature threshold comparison: In over-limit conditions, cooling efficiency is prioritized, and a deep reinforcement learning-based cooling control algorithm is used to rapidly calculate high-speed commands to suppress temperature rise. In safe conditions, the focus is on optimizing energy efficiency, leveraging the multi-objective optimization capabilities of the same algorithm to reduce the speed. This differentiated control mechanism overcomes the rigidity of traditional single-speed strategies, preventing energy waste from excessive cooling while enabling rapid response when temperature risks arise. Furthermore, the edge computing-based algorithm deployment model reduces response latency, ensuring the real-time and reliability of critical cooling commands, and providing a solid foundation for the stable operation of high-density servers.
[0089] In one possible implementation of the embodiment of the present application, when calculating the first fan speed, the following method can also be used but is not limited to: calculating the first temperature difference between the server temperature and the dynamic temperature threshold; inputting the first temperature difference, real-time operating data and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain the first fan speed, wherein the size of the first temperature difference is positively correlated with the size of the first fan speed.
[0090] In an embodiment of the present application, the first temperature difference refers to the numerical difference between the real-time temperature (i.e., server temperature) of a key component of the first server (e.g., the CPU) and the dynamic temperature threshold. The first temperature difference is obtained by acquiring the current temperature value through real-time temperature acquisition and subtracting it from the dynamic temperature threshold. For example, if the real-time CPU temperature is 88°C and the dynamic temperature threshold is set to 85°C, the first temperature difference is +3°C; if the real-time temperature is 82°C and the threshold is 85°C, the difference is -3°C. This difference not only reflects the degree of deviation between the current temperature and the safety boundary but also implies the urgency of the cooling requirement, providing a quantitative basis for subsequent fan speed calculations.
[0091] Generating the first fan speed according to the first heat dissipation control algorithm may be achieved through, but not limited to, the following methods:
[0092] State space construction: The first temperature difference is used as the core state variable, along with parameters such as load and ambient temperature, to form a multidimensional input vector. For example, the input vector could be represented as [temperature difference ΔT = +3°C, CPU load = 90%, ambient temperature = 25°C, current fan speed = 4000 RPM].
[0093] Dynamic weighting of temperature differences: A neural network automatically learns the weights of how temperature differences affect fan speed in different scenarios. For example, in high-load scenarios (e.g., CPU utilization exceeding 80%), positive temperature differences may be given a higher weight to quickly respond to potential overheating risks; in low-load scenarios, the model may reduce the weight to prioritize energy efficiency.
[0094] Positive correlation control logic implementation: The control strategy of positively correlating the first temperature difference with fan speed is reinforced through a reward mechanism in the training data. When the first temperature difference increases (e.g., ΔT rises from +2°C to +5°C), the fan speed is increased (e.g., from 4500 RPM to 6000 RPM) to accelerate heat dissipation. Conversely, as the first temperature difference decreases, the speed is gradually reduced to reduce energy consumption. This step can be implemented through the design of a DRL reward function. For example, a multi-objective trade-off is performed between the reduction in the first temperature difference per unit time and the energy cost of the speed increase to ensure a positive correlation that meets the overall system optimization goals.
[0095] The magnitude of the first temperature difference directly reflects the urgency of cooling needs, and dynamically adjusts the speed based on multiple factors such as load and environmental data, avoiding over- or under-cooling caused by a one-size-fits-all approach. Positive correlation logic ensures that fan speed increases as the first temperature difference increases, while simultaneously suppressing unnecessary speed increases through the multi-objective optimization mechanism of the first cooling control algorithm, achieving a balance between rapid cooling and energy conservation and consumption reduction.
[0096] In one possible implementation of the embodiment of the present application, when calculating the second fan speed, the following method can also be used but is not limited to: calculating the second temperature difference between the server temperature and the dynamic temperature threshold; inputting the second temperature difference, real-time operating data and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain the second fan speed, wherein the size of the second temperature difference is negatively correlated with the size of the second fan speed.
[0097] In the embodiments of the present application, the second temperature difference also refers to the numerical difference between the real-time temperature (i.e., server temperature) of the key components of the first server (such as the CPU, etc.) and the dynamic temperature threshold, but its calculation scenario is limited to the case where the real-time temperature is less than or equal to the dynamic temperature threshold. In this case, the difference is zero or a negative value (such as the difference between the real-time temperature of 82°C and the threshold of 85°C is -3°C), reflecting the distance between the current temperature and the upper safety limit. Unlike the positive difference in the over-limit state, the negative difference characterizes the redundancy capacity of the cooling system and provides a quantitative basis for energy efficiency optimization. For example, when the difference reaches -5°C, it indicates that the server temperature is far below the threshold, and there is room to reduce the fan speed to save energy.
[0098] The second temperature difference is negatively correlated with the second fan speed, meaning that when the temperature difference increases negatively (i.e., the real-time temperature further falls below the threshold), the target fan speed decreases accordingly. Generating the second fan speed according to the first heat dissipation control algorithm can be achieved through, but not limited to, the following methods:
[0099] Sign-sensitive state encoding: The first cooling control algorithm separates positive and negative temperature differences at the input layer. For negative differences (i.e., the second temperature difference), the energy efficiency optimization subnetwork is activated. Using historical data, it learns how to minimize fan speed while maintaining safe temperatures. For example, the input vector [ΔT = -3°C, load = 20%, ambient temperature = 22°C] might trigger a low-speed decision branch.
[0100] Multi-objective trade-offs in the reward function: In a safe state, the reward function of the first cooling control algorithm prioritizes energy consumption and noise suppression. For example, when the fan speed decreases while the temperature remains stable, a positive reward is given. However, if the speed drops too low, causing the temperature to approach a threshold (e.g., ΔT rises from -5°C to -1°C), a slight penalty is applied to indicate risk. Through repeated training, the algorithm autonomously establishes the association rule that a larger second temperature difference → lower speed → higher reward, forming a negative correlation control strategy.
[0101] Dynamically relax constraints: Dynamically adjust the allowable range of the second temperature difference based on ambient temperature and load trends. For example, when the ambient temperature is low (e.g., 18°C) and the load is stable, the second temperature difference can be reduced to -8°C before the speed is reduced. When the ambient temperature is high (e.g., 28°C) or the load is increasing, the second temperature difference is limited to -4°C to reserve heat dissipation margin to cope with potential temperature rise.
[0102] The calculation process of the second fan speed includes but is not limited to the following steps:
[0103] Differential range grading: The second temperature differential is divided into multiple ranges (e.g., -1°C to -3°C for mild redundancy, -3°C to -6°C for high redundancy, and less than -6°C for excessive redundancy). Each range corresponds to a different speed adjustment strategy. For example, in the high redundancy range, the algorithm can reduce the speed linearly (200 RPM for every -1°C increase in ΔT). In the excessive redundancy range, a nonlinear attenuation strategy is used (e.g., when ΔT drops from -6°C to -8°C, the speed is reduced by only 50 RPM), preventing fan stalls and localized hotspots.
[0104] Environmental Coupling Compensation: When calculating speed, the algorithm takes into account the effect of ambient temperature on heat dissipation efficiency to compensate. For example, when the ambient temperature rises (e.g., from 20°C to 25°C), even if ΔT remains at -4°C, the algorithm will slightly increase the speed (e.g., by 100 RPM) to offset the potential impact of a deteriorating external thermal environment on heat dissipation.
[0105] Load change prediction: By analyzing real-time load data (such as process queue length and I / O throughput), the algorithm predicts short-term load fluctuations. If the load is likely to increase (for example, when a batch task start signal is detected), the algorithm will maintain a small margin (for example, an additional 200 RPM) above the current speed to prevent a sudden increase in load from causing the temperature to quickly approach the threshold.
[0106] By controlling the negative correlation between the second temperature difference and fan speed, the system proactively reduces fan speed when sufficient temperature margin is met, reducing power consumption and noise pollution. Dynamically relaxing constraints and pre-determining loads prevent excessive speed reduction from causing a rapid temperature rise near the threshold, maintaining a smooth temperature curve.
[0107] In one possible implementation of the embodiment of the present application, when the fan speed of the first server is adjusted according to the target fan speed, it can also be implemented in but not limited to the following manner: the target fan speed is transmitted to the fan controller of the first server, and the fan speed of the first server is controlled to the target fan speed through the fan controller to obtain the adjusted first server; when it is determined that the fan speed of the first server is the target fan speed, the operating data of the first server is obtained to obtain the second operating data.
[0108] In an embodiment of the present application, the target fan speed is an ideal speed value calculated by the first heat dissipation control algorithm, including a first fan speed (a high speed in an over-limit state) or a second fan speed (a low speed in a safe state). The target fan speed can be transmitted to the fan controller of the first server via a communication interface of the baseboard management controller (BMC). The fan controller is a hardware module within the first server that specifically manages fan operation. It is typically integrated on the motherboard or a separate control card and is responsible for converting digital speed commands into physical drive signals.
[0109] After confirming that the fan speed has stabilized to the target fan speed, collect second operating data, which includes but is not limited to:
[0110] Temperature change response data, including at least: the temperature change curve, cooling rate, steady-state temperature value, and fluctuation amplitude of key components (such as the CPU) after speed adjustment. Energy consumption data, including at least: the real-time power consumption, cumulative energy consumption, and power factor of the fan at the target speed. Noise and vibration data, including at least: the noise decibel value and chassis vibration amplitude collected by microphones or accelerometers when the fan is running at high speed, used to evaluate the impact of the cooling strategy on the equipment working environment. Load-related data, including at least: recording the server load status during speed adjustment (such as CPU utilization and memory bandwidth occupancy), and analyzing the coupling relationship between cooling effect and load fluctuations.
[0111] The fan controller's gradient adjustment and feedback verification mechanisms ensure accurate execution of speed commands, avoiding control deviations caused by mechanical inertia or electrical noise. Secondary operating data provides real-world operating condition feedback to the deep reinforcement learning algorithm, which can be used to correct the prediction error of the primary cooling control algorithm (e.g., adjusting the difference-speed mapping when the actual cooling rate is lower than expected) or optimize the reward function weight (e.g., reducing the reward score for high-speed strategies in high-noise scenarios). Real-time monitoring of adjusted temperature and energy consumption allows for rapid identification of abnormal conditions (e.g., substandard speed due to fan failure) and triggering redundant cooling solutions (e.g., enabling backup fans or migrating computing loads) to prevent risks caused by local overheating.
[0112] In one possible implementation method of the embodiment of the present application, when performing data evaluation and analysis on the second operating data, it can also be implemented by, but not limited to, the following method: performing data analysis and processing on the second operating data based on historical operating data and preset evaluation indicators to obtain a heat dissipation evaluation result, wherein the heat dissipation evaluation result includes at least a temperature evaluation result, a speed evaluation result, and an energy consumption evaluation result.
[0113] In the embodiment of the present application, the historical operating data used as the evaluation benchmark includes, but is not limited to: heat dissipation performance records of the first server under various typical working conditions, for example:
[0114] Temperature baseline library: historical average temperatures and reasonable fluctuation ranges corresponding to different load levels (such as 30%, 60%, and 90% CPU utilization); speed efficiency model: typical heat dissipation capacity (such as 2°C cooling per minute) and historical optimal speed-temperature response relationship at a specific speed (such as 4000RPM); energy consumption mode library: fan energy consumption baselines and energy efficiency optimization thresholds under different ambient temperatures (such as 20°C, 25°C, and 30°C).
[0115] By aligning the second-run data with the historical data and matching them with the scenarios (e.g., selecting records with the same load range and similar environmental conditions), a comparable evaluation framework is established to eliminate external variable interference. For example, when the CPU load in the second-run data is 75%, all records with loads between 70% and 80% are extracted from the historical database as a comparison baseline.
[0116] The preset evaluation indicators are pre-set multi-dimensional analysis indicators, including but not limited to: analysis indicators covering three dimensions of heat dissipation performance, control efficiency and energy efficiency. By analyzing the second operation data with the preset evaluation indicators and historical operation data, a heat dissipation evaluation result including at least a temperature evaluation result, a speed evaluation result and an energy consumption evaluation result can be obtained.
[0117] Specifically, the evaluation process of temperature assessment results includes but is not limited to:
[0118] Stability evaluation: Calculates the standard deviation and maximum fluctuation of the adjusted temperature curve to determine whether the temperature is stable and converging within the dynamic threshold range. For example, if the temperature fluctuates less than ±1°C within 10 minutes after adjustment, it is marked as highly stable. Over-limit risk assessment: Counts the duration and frequency of temperature close to the threshold (e.g., reaching 95% of the threshold) to evaluate the safety margin of the cooling strategy. For example, if the adjusted temperature repeatedly reaches 84°C (dynamic temperature threshold 85°C), it indicates a potential over-limit risk. Heat dissipation efficiency evaluation: Quantifies the cooling effect brought about by a unit speed increase based on the ratio of the temperature change rate to the speed adjustment amount. For example, if the temperature drops by 3°C per minute after increasing the speed by 1000 RPM, the heat dissipation efficiency is rated as excellent.
[0119] The evaluation process of the speed assessment results includes but is not limited to:
[0120] Response timeliness evaluation: Measure the total delay from temperature over-limit detection to speed stabilization to target fan speed to determine the real-time performance of the control link. For example, if the delay exceeds 5 seconds, it is considered a response lag. Speed stability evaluation: Analyze the deviation rate and fluctuation frequency between the actual speed and the target fan speed to evaluate the execution accuracy of the control instructions. For example, if the speed fluctuates within the range of ±2% and there is no periodic oscillation, it is rated as high stability. Strategy rationality evaluation: Compare the actual speed with the historical optimal speed strategy to identify whether there is over-adjustment or conservative control. For example, if the current speed is 10% higher than the historical optimal value under the same working conditions, it indicates that the strategy may be redundant.
[0121] The evaluation process of energy consumption assessment results includes but is not limited to:
[0122] Absolute energy consumption evaluation: Calculates the total energy consumption of the cooling system during the adjustment period and compares it with the historical average value of the same scenario to determine whether energy efficiency has improved or degraded. For example, if energy consumption decreases by 15% and the temperature is stable, it is marked as energy efficiency improvement. Energy Temperature Ratio (ETR): Defined as the energy consumed per unit temperature drop (for example, 0.5 kWh per 1°C decrease), it is used to compare the energy efficiency levels of different cooling strategies. Load-related energy efficiency evaluation: Analyzes the sensitivity of energy consumption to load changes and identifies energy efficiency bottlenecks under high load. For example, if energy consumption surges by 200% when the load increases from 50% to 80%, it indicates that there is room for optimization of the cooling system in high-load scenarios.
[0123] Thermal assessment results can be generated through, but not limited to, the following processes:
[0124] Data alignment and normalization: Align the time windows of the second run data with the historical data, unify the units, and remove outliers to ensure comparability.
[0125] Weighted metric calculation: Dynamic weights are assigned to each evaluation metric based on the server type and O&M strategy. For example, in a high-performance computing server, temperature stability might be weighted as high as 60%, while energy consumption might only be weighted 20%. In an edge low-power server, energy consumption might be weighted as high as 50%.
[0126] Comprehensive evaluation output: Using fuzzy logic or a weighted scoring model, the scores of each dimension are integrated into an overall evaluation result (e.g., a 100-point score or excellent / good / fair / poor grades), and diagnostic recommendations are generated (e.g., a recommendation to reduce the speed by 200 RPM to optimize energy efficiency).
[0127] Through in-depth evaluation of multi-dimensional, multi-source data and direct feedback of heat dissipation assessment results into the reward function of the deep reinforcement learning algorithm, this system prioritizes learning of efficient, low-consumption, and stable control strategies, accelerating algorithm convergence. Detailed evaluation results based on temperature, speed, and energy consumption help operations and maintenance personnel quickly identify cooling system bottlenecks (such as performance degradation of specific fans or design flaws in cooling ducts) and develop targeted maintenance plans. Combined with comparative analysis of historical data, the system can identify long-term trends (such as energy efficiency degradation caused by dust accumulation on radiators), automatically triggering model retraining or hardware maintenance alerts, and improving system reliability throughout its lifecycle.
[0128] In one possible implementation of the embodiment of the present application, after obtaining the heat dissipation evaluation result, it is necessary to determine whether the heat dissipation evaluation result meets the preset heat dissipation conditions. Specifically, regarding the analysis method of the heat dissipation evaluation result, the embodiment of the present application provides an analysis flow chart of the heat dissipation evaluation result, such as Figure 3 As shown, including:
[0129] Step 301 : Compare the temperature evaluation result with the temperature evaluation indicators in the preset evaluation indicators to obtain a temperature evaluation deviation value, wherein the preset evaluation indicators at least include a speed evaluation indicator, an energy consumption evaluation indicator, and a temperature evaluation indicator.
[0130] In the embodiments of this application, temperature assessment indicators are preset temperature-related performance standards, including but not limited to: a temperature stability threshold (e.g., an allowable fluctuation range of ±2°C), an over-limit risk tolerance (e.g., the percentage of time the temperature is close to the threshold ≤ 10%), and a heat dissipation efficiency baseline (e.g., a cooling rate of ≥ 1°C per minute). Temperature assessment results include but are not limited to: a measured or calculated temperature stability value, an over-limit risk level, and a heat dissipation efficiency score.
[0131] Comparison processing can be achieved through, but is not limited to, deviation calculations. This involves performing a difference or ratio calculation on each sub-item in the temperature assessment results (e.g., stability, risk of exceeding limits, and efficiency) with the corresponding temperature assessment indicator to generate a temperature assessment deviation. For example, if the preset stability requirement is a fluctuation of ≤±1°C and the actual fluctuation is ±1.5°C, the deviation value is +0.5°C. If the preset cooling efficiency is a cooling of 2°C per minute and the actual cooling efficiency is 2.3°C, the deviation value is -0.3°C (negative values indicate better than expected performance). The absolute value of the deviation reflects the degree of deviation, and the sign indicates the direction (positive values indicate failure to meet the target, negative values indicate overachievement).
[0132] Step 302 : Compare the rotational speed evaluation result with the rotational speed evaluation index to obtain a rotational speed evaluation deviation value, and compare the energy consumption evaluation result with the energy consumption evaluation index to obtain an energy consumption evaluation deviation value.
[0133] In the embodiments of the present application, the speed evaluation index is a preset speed-related performance standard, including but not limited to: an upper limit on the response time (e.g., from the issuance of the instruction to the stabilization of the speed ≤ 3 seconds), speed control accuracy (e.g., the deviation between the actual speed and the target value ≤ ±3%), and a strategy rationality threshold (e.g., the speed adjustment range does not exceed ±15% of the historical optimal value). The speed evaluation results include but are not limited to: the measured response time, the speed deviation rate, and the strategy difference. In the comparison process, for example: the actual response time and the preset upper limit (e.g., 3 seconds) are compared to calculate the speed evaluation deviation value (e.g., the actual 4 seconds is a +1 second deviation); similarly, the speed deviation rate (e.g., the actual deviation is 5% and the preset 3%) generates a deviation value of +2%.
[0134] Energy consumption assessment indicators are pre-set energy-related performance standards, including but not limited to: maximum allowable power consumption (e.g., single fan ≤ 20W), energy consumption to temperature ratio (e.g., ETR ≤ 0.6kWh / °C), and load sensitivity limits (e.g., energy consumption increase ≤ 50% when load increases by 30%). Energy consumption assessment results include, but are not limited to, measured power consumption, ETR value, and the slope of the load-energy consumption curve. For example, if the actual ETR is 0.7kWh / °C and the pre-set indicator is 0.6, the deviation is +0.1; if energy consumption only increases by 40% when the load increases by 30%, the deviation is -10%.
[0135] Step 303: determine whether the temperature evaluation result meets the preset heat dissipation conditions according to the temperature evaluation deviation value, determine whether the speed evaluation result meets the preset heat dissipation conditions according to the speed evaluation deviation value, and determine whether the energy consumption evaluation result meets the preset heat dissipation conditions according to the energy consumption evaluation deviation value.
[0136] In an embodiment of the present application, the preset heat dissipation conditions include at least a set of allowable ranges for the deviation values of each evaluation dimension, for example: for the temperature dimension: stability deviation ≤ +0.5°C, over-limit risk deviation ≤ +5%, efficiency deviation ≥ -0.2°C / minute; for the speed dimension: response time deviation ≤ +1 second, control accuracy deviation ≤ +2%, strategy difference ≤ ±10%; for the energy consumption dimension: power consumption deviation ≤ +10%, ETR deviation ≤ +0.1, load sensitivity deviation ≤ +15%. Specifically, this application does not limit the content of the preset heat dissipation conditions.
[0137] When determining whether the preset cooling conditions are met, each deviation is individually checked to see if it exceeds the allowable range. For example, if the temperature stability deviation is +0.6°C (exceeding the upper limit of +0.5°C), the temperature assessment is marked as non-compliant; if the speed control accuracy deviation is +1.5% (within the +2% tolerance), it is marked as compliant. This process can be implemented using a rules engine or fuzzy logic, allowing for flexible handling of boundary conditions (for example, deviations approaching a threshold can trigger a warning rather than an outright failure).
[0138] Step 304 : If any one of the temperature evaluation result, the rotation speed evaluation result, and the energy consumption evaluation result does not meet the preset heat dissipation condition, determine that the heat dissipation evaluation result does not meet the preset heat dissipation condition.
[0139] In the embodiments of this application, if the evaluation results for any one of temperature, speed, or energy consumption fail to meet the preset cooling conditions, the overall cooling evaluation result is determined to fail the preset cooling conditions. For example, even if the speed and energy consumption evaluations meet the standards, but the temperature stability exceeds the limit, the optimization process is still triggered. This design reflects the system's strict prioritization of cooling safety while balancing energy efficiency and control performance.
[0140] Through multi-dimensional deviation calculation, we can precisely pinpoint weaknesses in cooling strategies (such as excessive response latency or low energy efficiency), avoiding the biased nature of traditional single-metric assessments. Preset cooling conditions can be dynamically adjusted based on the server lifecycle stage and environmental characteristics. For example, energy consumption deviation tolerances can be relaxed for older servers, while temperature stability requirements can be tightened for new, high-performance servers. When any core metric (such as temperature safety) falls below standard, emergency optimization is prioritized to ensure system reliability. Minor metrics (such as minor energy efficiency deviations) can be incorporated into periodic optimization tasks.
[0141] In one possible implementation of the embodiment of the present application, when judging whether the temperature evaluation results, the speed evaluation results and the energy consumption evaluation results meet the preset heat dissipation conditions, it can be implemented in but not limited to the following manner: when the temperature evaluation deviation value is greater than the first preset evaluation deviation threshold, it is determined that the temperature evaluation result does not meet the preset heat dissipation conditions; when the speed evaluation deviation value is greater than the second preset evaluation deviation threshold, it is determined that the speed evaluation result does not meet the preset heat dissipation conditions; when the energy consumption evaluation deviation value is greater than the third preset evaluation deviation threshold, it is determined that the energy consumption evaluation result does not meet the preset heat dissipation conditions.
[0142] In the embodiments of this application, independent, multi-dimensional assessments are used to quickly identify deficiencies in cooling strategies. For example, if only energy consumption exceeds standards, optimization may focus on reducing speed or introducing load migration. If both temperature and speed fall short, comprehensive adjustments to PID parameters and fan response logic are necessary. This provides core support for the continuous optimization and refined operation and maintenance of server cooling systems.
[0143] In one possible implementation method of the embodiment of the present application, when performing algorithm optimization processing on the first heat dissipation control algorithm, it can also be implemented in but not limited to the following method: performing algorithm optimization processing on the first heat dissipation control algorithm through a preset deep reinforcement learning algorithm based on the second operating data, heat dissipation evaluation results and historical operating data; until a preset number of optimizations or a preset convergence condition is reached, and a second heat dissipation control algorithm is obtained.
[0144] In the embodiments of this application, the BMC regularly trains a first heat dissipation control algorithm based on deep reinforcement learning using the latest training samples. During training, the algorithm continuously adjusts the policy network parameters through interaction with the environment (i.e., the heat dissipation control process) to maximize the cumulative reward. When the algorithm converges or reaches a preset number of training cycles, the optimized first heat dissipation control algorithm parameters are obtained and used for subsequent heat dissipation control.
[0145] Secondary operational data provides feedback on actual operating conditions after fan speed adjustments, including but not limited to temperature response curves, energy consumption variations, and control stability indicators. Thermal assessment results quantify performance deficiencies of the current strategy in terms of temperature safety, control accuracy, and energy efficiency (e.g., excessive temperature fluctuations or excessive energy consumption). Historical operational data encompasses diverse operating conditions accumulated over the server's long-term operation, providing a broad training sample and pattern recognition foundation for algorithm optimization. Together, these three data elements form a multi-dimensional input space based on deep reinforcement learning mechanisms, ensuring that the optimization process considers both immediate control effects and historical experience and system-level performance goals.
[0146] The preset number of optimizations is a custom-set number of optimizations, which is used as a safeguard mechanism to avoid infinite iterations, for example: the maximum number of iterations for local training is 1,000, etc. Specifically, regarding the preset number of optimizations, the embodiments of this application do not impose any restrictions.
[0147] The algorithm can autonomously identify long-term factors such as server hardware aging and changes in cooling duct efficiency, and dynamically adjust the parameters of the first cooling control algorithm to maintain optimal cooling performance; it achieves dynamic trade-offs between conflicting goals such as temperature safety, energy economy, and control stability. For example, in a mild over-limit scenario, it chooses to moderately increase the speed rather than the maximum speed, taking into account both the cooling speed and energy cost; the second cooling control algorithm has stronger anti-interference capabilities, and can quickly respond to and suppress temperature fluctuations in the event of sudden load shocks or drastic changes in ambient temperature, thereby avoiding cascading failures.
[0148] In one possible implementation of an embodiment of the present application, after transmitting the second heat dissipation control algorithm to a preset edge computing device, the following method may also be used but is not limited to: in response to a control instruction to perform heat dissipation control on the second server, extract the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server.
[0149] In an embodiment of the present application, a preset edge computing device receives the collected data, i.e., the third operating data, in real time and analyzes the data in real time. For simple changes in heat dissipation requirements, such as small temperature fluctuations or slight changes in load, the preset edge computing device quickly adjusts the second heat dissipation control algorithm parameters and fan speed according to local preset rules and lightweight models. When encountering complex heat dissipation scenarios, such as a sudden and substantial increase in server load or a sharp change in ambient temperature, the preset edge computing device uploads key data to the cloud computing platform. The cloud computing platform uses big data analysis and complex machine learning models to deeply process the data, generate optimized control strategies, and feed back the strategies to the preset edge computing device, which then performs corresponding adjustment operations.
[0150] The second server can be the same as or different from the first server. The second server refers to a server node different from the first server. The hardware configuration, load characteristics, and heat dissipation requirements of the second server may differ from those of the first server (e.g., using a different CPU model or being deployed in a different rack location in the computer room). When the system detects that the second server requires heat dissipation control (e.g., detecting that its temperature is approaching a dynamic threshold or that its load has increased significantly), a control instruction is generated and sent to the preset edge computing device. The control instruction includes, but is not limited to: a unique identifier for the second server, a summary of its current operating status (e.g., temperature, load, ambient temperature), and a heat dissipation control target (e.g., prioritizing energy efficiency or forced cooling).
[0151] After receiving the instruction, the preset edge computing device performs heat dissipation control on the second server using the second heat dissipation control algorithm. In a multi-server cluster environment, each server shares heat dissipation-related data via the blockchain network. At regular intervals, local temperature data, fan speed data, load information, and heat dissipation control algorithm parameters are packaged into blocks and transmitted to the preset edge computing device. These blocks are then broadcast and verified on the blockchain network through a consensus mechanism. Based on the shared data on the blockchain, the distributed heat dissipation optimization algorithm runs periodically, calculating the optimal fan speed and heat dissipation control algorithm parameter adjustment plan for each server based on the real-time status of each server. Based on the algorithm results, each server adjusts its local heat dissipation control strategy to achieve collaborative heat dissipation for multiple servers.
[0152] By rapidly promoting the optimized heat dissipation control algorithm, the second heat dissipation control algorithm, across the cluster, we avoid wasting resources training each server independently and significantly shorten the policy convergence time for new nodes. For example, when deploying a new server, we can directly call the matching optimization model from the existing algorithm library without having to learn from scratch.
[0153] In one implementation of the embodiment of the present application, when extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server, it can also be implemented in but not limited to the following manner: obtaining the third operating data of the second server, and performing data comparison processing based on the third operating data and the real-time operating data to obtain operating difference data between the first server and the second server; when the operating difference data meets the preset first operating difference condition, extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server; when the operating difference data does not meet the preset first operating difference condition but meets the second operating difference condition, adjusting the second heat dissipation control algorithm according to the preset algorithm update rule in the preset edge computing device to obtain an adjusted second heat dissipation control algorithm, and performing heat dissipation control processing on the second server according to the adjusted second heat dissipation control algorithm; when the operating difference data does not meet the second operating difference condition, obtaining the second historical data of the second server, and performing algorithm optimization processing on the second heat dissipation control algorithm based on the second historical data and the third operating data to obtain an optimized second heat dissipation control algorithm; and performing heat dissipation control processing on the second server according to the optimized second heat dissipation control algorithm.
[0154] In an embodiment of the present application, the third operating data refers to the real-time operating status information of the second server, including but not limited to: hardware parameters: CPU model, radiator type, fan specifications, etc.; dynamic indicators: current temperature, load intensity (such as: CPU utilization, memory bandwidth), ambient temperature, fan speed; historical characteristics: typical temperature rise pattern, heat dissipation response delay, energy consumption baseline.
[0155] The data comparison process can generate running difference data through but not limited to the following steps:
[0156] Feature extraction: Extract key features from the real-time operating data of the first and second servers, such as the slope of the load-temperature curve, the heat dissipation efficiency coefficient (cooling capacity per unit speed), and the ambient temperature sensitivity.
[0157] Similarity calculation: Use cosine similarity or Euclidean distance to measure the temperature response differences, energy consumption differences, and control stability differences between two servers in the same load range (e.g., 60%-80% CPU utilization).
[0158] Difference quantification: This converts similarity results into interpretable difference scores. For example, if the temperature difference between two servers under the same load is +5°C, the difference score is marked as high; if the energy consumption difference is less than 5%, it is marked as low.
[0159] Among them, the preset first operating difference condition is the threshold that allows direct reuse of the second heat dissipation control algorithm, for example: the heat dissipation architecture of the second server is similar to that of the first server (for example, both adopt a liquid cooling + air cooling hybrid design); the load-temperature response difference is ≤10%, the energy consumption difference is ≤8%; the air duct efficiency difference of the rack position is ≤15%.
[0160] When the operational difference data meets the first operational difference condition, the preset edge computing device directly invokes the second thermal control algorithm and deploys it to the BMC of the second server. For example, if the two servers are of the same model and located in adjacent racks, the optimized PID parameters and speed strategy can be directly applied.
[0161] The preset second operating difference condition is a threshold that allows lightweight adjustment of the adaptation algorithm, for example: the radiator type is the same but the fan specifications are different (such as the speed range difference ≤ 20%); the load-temperature response difference is ≤ 25%, the energy consumption difference is ≤ 20%; the air duct efficiency difference is ≤ 30%.
[0162] In this case, the preset algorithm update rule is triggered to adjust the second heat dissipation control algorithm. The process of adjusting the second heat dissipation control algorithm can be implemented in, but not limited to, the following ways:
[0163] Parameter scaling: Adjust the parameters of the second heat dissipation control algorithm according to hardware differences. For example, if the maximum fan speed of the second server is 80% of that of the first server, the proportional coefficient P is amplified by 1.25 times to compensate for the difference in the upper limit of the speed. Model fine-tuning: Utilize a small amount of real-time data from the second server (such as the operation records of the last hour) to perform transfer learning on the DRL model on the preset edge device and update the output layer weights of the policy network. Rule injection: Add environmental difference compensation rules. For example, if the ambient temperature of the second server is high, add a fixed offset (such as +200RPM) to the target speed calculation. The adjusted second heat dissipation control algorithm adapts to the local characteristics of the second server while retaining the core control logic, for example, increasing the heat dissipation margin for servers in high ambient temperature areas.
[0164] When the operating difference data does not meet the second operating difference condition (for example, the hardware architecture is very different, and the performance deviation is greater than 30%), the second heat dissipation control algorithm is deeply optimized:
[0165] Secondary historical data integration: Extracts long-term operational data from the secondary server, including extreme operating conditions (e.g., full-load stress testing), historical cooling events (e.g., overheating alarms), and maintenance logs (e.g., performance changes after radiator cleaning). Joint training: Blends secondary historical data with tertiary operational data to retrain the DRL model. State space expansion: Incorporates unique features of the secondary server (e.g., memory temperature, liquid cooling pump speed); reward function customization: Adjusts reward weights based on the secondary server's operational and maintenance objectives (e.g., prioritizing cooling of key components); Cross-model knowledge distillation: Extracts common policies (e.g., temperature-speed mapping rules) from the primary server's optimization model to accelerate convergence of the new model. Edge-cloud collaborative optimization: If local computing power is insufficient, data is encrypted and uploaded to the cloud. Distributed training is used to generate an optimized secondary cooling control algorithm, which is then transmitted back to the edge device for verification and deployment.
[0166] Through hierarchical conditional judgment and policy branching, flexible control is achieved, from direct reuse to deep customization, balancing efficiency and accuracy. Light adjustments avoid the computational overhead of full retraining, while deep optimization ensures precise control of heterogeneous hardware and optimizes the allocation of edge computing power. The optimization results of the second server are stored on the blockchain and can be used by other similar nodes in the cluster, forming a continuously evolving library of cooling strategies.
[0167] In order to facilitate understanding of the implementation process of this application, the embodiment of this application also provides a system structure diagram of server heat dissipation control, such as Figure 4 As shown, the component monitoring module is a module in the first server, for example: the BMC in the first server, etc., which includes various components of the first server and is used to monitor the real-time operation data of the first server; the BMC data acquisition module is used to obtain real-time operation data for heat dissipation control processing; the PID control calculation module is used to calculate the target fan speed based on the first heat dissipation control algorithm; the fan controller is used to control the fan speed of the first server to the target fan speed; the PID parameter learning module is used to optimize the first heat dissipation control algorithm to obtain a second heat dissipation control algorithm; the performance evaluation module is used to perform data evaluation and analysis on the second operation data to obtain a heat dissipation evaluation result.
[0168] Furthermore, based on the server heat dissipation control system, the embodiment of the present application also provides a flow chart of another server heat dissipation control method, such as Figure 5 As shown, including:
[0169] S1: Import an initial version of the cooling strategy into the first server, and preliminarily set the PID control parameters (first cooling control algorithm parameters) for each key component, including: key component temperature threshold, temperature control point, etc. It should be noted that when performing cooling control on the first server, initial parameters can be provided in advance, or the first cooling control algorithm parameters can be obtained from a preset edge computing device.
[0170] S2: Real-time collection of multi-source data of the first server hardware.
[0171] The BMC uses high-precision temperature sensors, fan speed sensors, load sensors, and ambient temperature sensors connected to the server motherboard to collect real-time data on server hardware temperature, fan speed, load, and ambient temperature. Temperature data includes the temperatures of key hardware such as the CPU, memory, and hard disk, while fan speed data includes the speed of each fan within the server. Temperature data accurately reflects the real-time temperature status of key hardware, fan speed data monitors the operating status of cooling fans, load data reflects the server's workload, and ambient temperature data provides information about the external environment.
[0172] S3: Target speed calculation based on deep reinforcement learning.
[0173] The collected multi-source data is input into a PID control algorithm (the first cooling control algorithm) based on deep reinforcement learning. A deep neural network with multiple hidden layers serves as the policy network. It takes the state space (i.e., real-time operating data) as input and outputs the optimal PID parameters for the current state (the optimal first cooling control algorithm parameters). Based on the optimal PID parameters, the target fan speed is calculated by combining the deviation between the current and target temperatures, the rate of change of the deviation, and the integral value (dynamic temperature threshold). During training, the deep reinforcement learning algorithm continuously optimizes the policy network based on feedback from the reward function, gradually learning the optimal control strategy.
[0174] S4: Send the fan speed adjustment amount to the fan controller to adjust the fan speed.
[0175] The BMC sends the calculated fan speed adjustment value to the fan controller, and the fan controller adjusts the fan speed according to the received adjustment value to achieve heat dissipation control.
[0176] S5: Training sample collection and update.
[0177] During the cooling control process, the BMC continuously collects temperature data, fan speed data, load data, and ambient temperature data after cooling control. This data is combined into training samples and stored in a local training dataset. Over time, the training dataset is continuously updated to include data from more diverse operating conditions, providing rich data support for the continuous optimization of deep reinforcement learning algorithms.
[0178] S6: Deep reinforcement learning to optimize PID parameters.
[0179] The BMC regularly trains the first thermal control algorithm using the latest training samples based on a deep reinforcement learning mechanism. During training, the algorithm continuously adjusts the policy network parameters through interaction with the environment (i.e., the thermal control process) to maximize the cumulative reward. When the algorithm converges or reaches the preset number of training cycles, the optimized PID parameters (parameters of the second thermal control algorithm) are obtained and used for subsequent thermal control.
[0180] S7: Collaborative processing of edge computing and cloud computing.
[0181] Edge computing devices receive and analyze data collected by the BMC in real time. For simple changes in cooling requirements, such as small temperature fluctuations or slight changes in load, edge computing devices quickly adjust the parameters of the secondary cooling control algorithm and fan speed based on locally pre-set rules and lightweight models. When encountering complex cooling scenarios, such as a sudden and significant increase in server load or a sharp change in ambient temperature, edge computing devices upload key data to the cloud computing platform. The cloud computing platform uses big data analysis and complex machine learning models to deeply process the data, generate optimized control strategies, and feed these strategies back to the edge computing devices, which then perform the corresponding adjustments.
[0182] S8: Blockchain distributed heat dissipation management.
[0183] In a multi-server cluster environment, each server's BMC shares cooling-related data via the blockchain network. At regular intervals, each BMC encapsulates local temperature data, fan speed data, load information, and cooling control algorithm parameters into blocks, which are broadcast and verified on the blockchain network through a consensus mechanism. Based on the shared data on the blockchain, a distributed cooling optimization algorithm runs periodically, calculating the optimal fan speed and cooling control algorithm parameter adjustment plan for each server based on the real-time status of each server. Based on the algorithm's results, each server's BMC adjusts its local cooling control strategy, achieving coordinated cooling across multiple servers.
[0184] S9: Adaptive dynamic temperature threshold adjustment.
[0185] The BMC uses a time series analysis algorithm to analyze historical temperature data of key server components, extracting periodic and trend characteristics of temperature changes. Combining server load patterns and ambient temperature trends, it uses a machine learning algorithm to predict future temperature changes. Based on these predictions, it dynamically adjusts the temperature thresholds of key components. For example, if a significant increase in server load is predicted, the temperature threshold is lowered in advance to enable fan speed control. If a drop in ambient temperature is predicted, the temperature threshold is appropriately raised to reduce unnecessary fan operation.
[0186] S10: Performance evaluation and strategy optimization.
[0187] The BMC regularly evaluates the effectiveness of thermal control based on preset metrics, such as temperature stability, temperature fluctuation range, fan speed adjustment frequency, and energy consumption. By comparing and analyzing historical data and ideal indicators, the BMC determines the effectiveness of the current thermal control strategy. Based on the evaluation results, the BMC adjusts hyperparameters of the deep reinforcement learning algorithm, optimizes the collaborative strategy between edge computing and cloud computing, improves the distributed thermal management algorithm, or adjusts the adaptive temperature threshold calculation method to continuously optimize the thermal control strategy and enhance the performance of the cooling system.
[0188] In summary, the embodiments of the present application can achieve the following technical effects:
[0189] 1. The server is first subjected to heat dissipation control based on its operating data and the heat dissipation control algorithm. The server after heat dissipation control is then evaluated based on the server's actual heat dissipation requirements, i.e., the preset heat dissipation conditions. The heat dissipation control algorithm is then optimized based on the evaluation results and the server's operating data after heat dissipation control. The optimization of the heat dissipation control algorithm also incorporates a deep reinforcement learning mechanism. By combining multiple algorithms with edge computing, i.e., preset edge computing devices, this algorithm addresses the balancing issues of difficult parameter setting and poor system adaptability faced in traditional heat dissipation control. This approach can address the difficulties of parameter setting and poor system adaptability faced in server heat dissipation control, achieving the technical effect of improving the efficiency and stability of the heat dissipation system.
[0190] 2. This application introduces deep reinforcement learning to make the PID controller highly intelligent and adaptive, and can automatically optimize PID parameters under complex and changeable working conditions, achieve precise heat dissipation control, and significantly improve the control efficiency and stability of the system.
[0191] 3. This application ensures real-time and accurate cooling control through a collaborative decision-making model between edge computing and cloud computing. Edge computing devices rapidly process local data, reducing response latency; the cloud computing platform provides powerful computing and analysis capabilities to handle complex cooling scenarios. The combination of the two improves the overall performance of cooling control.
[0192] 4. This application achieves multi-server collaboration through a blockchain-based distributed heat dissipation management system, effectively improving the cooling efficiency of the server cluster and reducing energy consumption. Blockchain technology ensures secure data sharing and consistency, enhancing system reliability and scalability.
[0193] 5. This application uses an adaptive dynamic temperature threshold adjustment strategy to dynamically adjust the temperature threshold according to the actual operating conditions and environmental changes of the server, thereby achieving more precise heat dissipation control, improving the energy utilization of the server, reducing unnecessary fan speed adjustments, and reducing equipment losses.
[0194] 6. This application achieves automated and intelligent heat dissipation control, reducing the cost and risk of manual intervention. Efficient heat dissipation control ensures stable server operation, improves server performance and service life, and brings significant economic benefits to data center operations.
[0195] 7. This application integrates cutting-edge technologies such as deep reinforcement learning, edge computing, cloud computing, and blockchain, bringing innovative solutions to the field of server heat dissipation control, promoting the application and development of related technologies in this field, and improving the technical level of the entire industry.
[0196] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0197] The embodiment of the present application also provides a device for controlling heat dissipation of a server. Figure 6 This is a schematic diagram of the structure of a server heat dissipation control device provided by this application, such as Figure 6 As shown, including:
[0198] An input unit 61 is configured to input the real-time operating data of the first server and a dynamic temperature threshold into a first heat dissipation control algorithm to calculate a fan speed for the first server to obtain a target fan speed, wherein the dynamic temperature threshold is a temperature threshold obtained by analyzing historical operating data of the first server;
[0199] An adjusting unit 62 is configured to adjust the fan speed of the first server according to the target fan speed to obtain the adjusted first server and second operating data of the adjusted first server;
[0200] An analysis unit 63 is configured to perform data evaluation and analysis on the second operating data to obtain an adjusted heat dissipation evaluation result corresponding to the first server;
[0201] An optimization unit 64 is configured to optimize the first heat dissipation control algorithm based on a deep reinforcement learning mechanism according to the second operating data, the heat dissipation evaluation result, and the historical operating data to obtain a second heat dissipation control algorithm when the heat dissipation evaluation result does not meet the preset heat dissipation condition;
[0202] a control unit 65 configured to perform heat dissipation control processing on the adjusted first server based on the second heat dissipation control algorithm, the second operating data, and the dynamic temperature threshold;
[0203] The transmission unit 66 is used to transmit the second heat dissipation control algorithm to a preset edge computing device.
[0204] In one embodiment of the present application, Figure 7 As shown, the server heat dissipation control device also includes:
[0205] A first processing unit 67 is configured to obtain historical operating data of the first server and perform data analysis and processing on the historical operating data to obtain temperature change cycle data, temperature change trend data, load change information, and ambient temperature information corresponding to the first server, wherein the historical operating data is operating data of the first server before the real-time operating data;
[0206] The first processing unit 67 is further configured to perform temperature prediction processing on the first server based on the load change information, the ambient temperature information, the temperature change cycle data, and the temperature change trend data to obtain target temperature change data;
[0207] The first processing unit 67 is further configured to, when it is determined based on the target temperature change data that the temperature of the first server has increased, set the dynamic temperature threshold to a first threshold;
[0208] The first processing unit 67 is further configured to, when it is determined according to the target temperature change data that the temperature of the first server decreases, set the dynamic temperature threshold to a second threshold, wherein the second threshold is greater than the first threshold.
[0209] In one embodiment of the present application, the input unit 61 is further configured to:
[0210] Performing data comparison processing based on the server temperature in the real-time operation data and the dynamic temperature threshold to obtain a first comparison result, wherein the server temperature is the real-time temperature of the first server;
[0211] When it is determined according to the first comparison result that the server temperature is greater than the dynamic temperature threshold, the real-time operating data and the dynamic temperature threshold are input into a first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a first fan speed, wherein the first heat dissipation control algorithm is a heat dissipation control algorithm based on deep reinforcement learning extracted from a preset edge computing device.
[0212] When it is determined according to the first comparison result that the server temperature is less than or equal to the dynamic temperature threshold, the real-time operating data and the dynamic temperature threshold are input into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain a second fan speed;
[0213] The target fan speed includes a first fan speed and a second fan speed, and the second fan speed is lower than the first fan speed.
[0214] In one embodiment of the present application, the input unit 61 is further configured to:
[0215] Calculating a first temperature difference between the server temperature and a dynamic temperature threshold;
[0216] The first temperature difference, real-time operating data, and dynamic temperature threshold are input into a first heat dissipation control algorithm to calculate the fan speed of the first server to obtain a first fan speed, wherein the magnitude of the first temperature difference is positively correlated with the magnitude of the first fan speed.
[0217] In one embodiment of the present application, the input unit 61 is further configured to:
[0218] calculating a second temperature difference between the server temperature and a dynamic temperature threshold;
[0219] The second temperature difference, real-time operating data, and dynamic temperature threshold are input into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the second fan speed, wherein the magnitude of the second temperature difference is negatively correlated with the magnitude of the second fan speed.
[0220] In one embodiment of the present application, the adjustment unit 62 is further configured to:
[0221] Transmitting the target fan speed to a fan controller of the first server, and controlling the fan speed of the first server to the target fan speed by the fan controller to obtain an adjusted first server;
[0222] When it is determined that the fan speed of the first server is the target fan speed, the operating data of the first server is acquired to obtain second operating data.
[0223] In one embodiment of the present application, the analysis unit 63 is also used to perform data analysis and processing on the second operating data based on historical operating data and preset evaluation indicators to obtain a heat dissipation evaluation result, wherein the heat dissipation evaluation result at least includes a temperature evaluation result, a speed evaluation result and an energy consumption evaluation result.
[0224] In one embodiment of the present application, Figure 7 As shown, the server heat dissipation control device also includes:
[0225] a second processing unit 68 for comparing the temperature evaluation result with a temperature evaluation index in a preset evaluation index to obtain a temperature evaluation deviation value, wherein the preset evaluation index includes at least a speed evaluation index, an energy consumption evaluation index, and a temperature evaluation index;
[0226] The second processing unit 68 is further configured to compare the speed evaluation result with the speed evaluation index to obtain a speed evaluation deviation value, and to compare the energy consumption evaluation result with the energy consumption evaluation index to obtain an energy consumption evaluation deviation value;
[0227] The second processing unit 68 is further configured to determine whether the temperature evaluation result meets the preset heat dissipation condition according to the temperature evaluation deviation value, determine whether the speed evaluation result meets the preset heat dissipation condition according to the speed evaluation deviation value, and determine whether the energy consumption evaluation result meets the preset heat dissipation condition according to the energy consumption evaluation deviation value;
[0228] The second processing unit 68 is further configured to determine that the heat dissipation evaluation result does not meet the preset heat dissipation condition if any one of the temperature evaluation result, the rotation speed evaluation result, and the energy consumption evaluation result does not meet the preset heat dissipation condition.
[0229] In one embodiment of the present application, the second processing unit 68 is further configured to:
[0230] When the temperature evaluation deviation value is greater than a first preset evaluation deviation threshold, it is determined that the temperature evaluation result does not meet the preset heat dissipation condition;
[0231] When the rotation speed evaluation deviation value is greater than a second preset evaluation deviation threshold, it is determined that the rotation speed evaluation result does not meet the preset heat dissipation condition;
[0232] When the energy consumption evaluation deviation value is greater than a third preset evaluation deviation threshold, it is determined that the energy consumption evaluation result does not meet the preset heat dissipation condition.
[0233] In one embodiment of the present application, the optimization unit 64 is further configured to:
[0234] Performing algorithm optimization processing on the first heat dissipation control algorithm by using a preset deep reinforcement learning algorithm according to the second operating data, the heat dissipation evaluation result, and the historical operating data;
[0235] Until a preset number of optimizations or a preset convergence condition is reached, a second heat dissipation control algorithm is obtained.
[0236] In one embodiment of the present application, Figure 7 As shown, the server heat dissipation control device also includes:
[0237] The third processing unit 69 is configured to extract a second heat dissipation control algorithm from a preset edge computing device to perform heat dissipation control processing on the second server in response to a control instruction for performing heat dissipation control on the second server.
[0238] In one embodiment of the present application, the third processing unit 69 is further configured to:
[0239] Obtaining the third operating data of the second server, and performing data comparison processing based on the third operating data and the real-time operating data to obtain operating difference data between the first server and the second server;
[0240] When the operation difference data satisfies a preset first operation difference condition, extracting a second heat dissipation control algorithm from a preset edge computing device to perform heat dissipation control processing on the second server;
[0241] When the operation difference data does not satisfy the preset first operation difference condition but satisfies the second operation difference condition, adjusting the second heat dissipation control algorithm according to a preset algorithm update rule in the preset edge computing device to obtain an adjusted second heat dissipation control algorithm, and performing heat dissipation control processing on the second server according to the adjusted second heat dissipation control algorithm;
[0242] When the operation difference data does not satisfy the second operation difference condition, obtaining second historical data of the second server, and performing algorithm optimization processing on the second heat dissipation control algorithm according to the second historical data and the third operation data to obtain an optimized second heat dissipation control algorithm;
[0243] Heat dissipation control processing is performed on the second server according to the optimized second heat dissipation control algorithm.
[0244] For the description of the features in the embodiment corresponding to the device for controlling heat dissipation of a server, reference can be made to the relevant description of the embodiment corresponding to the method for controlling heat dissipation of a server, which will not be described in detail here.
[0245] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps of any of the above-mentioned server heat dissipation control method embodiments.
[0246] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned server heat dissipation control method embodiments when running.
[0247] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0248] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server heat dissipation control method embodiments are implemented.
[0249] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned server heat dissipation control method embodiments.
[0250] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0251] The above is a detailed introduction to a method, electronic device, and storage medium for controlling heat dissipation in a server provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that, for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for controlling heat dissipation of a server, characterized in that: include: Inputting the real-time operating data and a dynamic temperature threshold of the first server into a first heat dissipation control algorithm to calculate the fan speed of the first server to obtain a target fan speed, wherein the dynamic temperature threshold is a temperature threshold obtained by analyzing the historical operating data of the first server; Adjusting the fan speed of the first server according to the target fan speed to obtain an adjusted first server and second operating data of the adjusted first server, and performing data evaluation and analysis on the second operating data to obtain a heat dissipation evaluation result corresponding to the adjusted first server; When the heat dissipation evaluation result does not meet the preset heat dissipation condition, performing algorithm optimization processing on the first heat dissipation control algorithm based on a deep reinforcement learning mechanism according to the second operating data, the heat dissipation evaluation result, and the historical operating data to obtain a second heat dissipation control algorithm; performing heat dissipation control processing on the adjusted first server based on the second heat dissipation control algorithm, the second operating data, and the dynamic temperature threshold, and transmitting the second heat dissipation control algorithm to a preset edge computing device; The determination of the dynamic temperature threshold comprises: Obtaining the historical operating data of the first server, and performing data analysis and processing on the historical operating data to obtain temperature change cycle data, temperature change trend data, load change information, and ambient temperature information corresponding to the first server, wherein the historical operating data is operating data of the first server before the real-time operating data; performing temperature prediction processing on the first server according to the load change information, the ambient temperature information, the temperature change cycle data, and the temperature change trend data to obtain target temperature change data; When it is determined according to the target temperature change data that the temperature of the first server has increased, setting the dynamic temperature threshold to a first threshold, wherein the first threshold is a threshold obtained by performing a threshold lowering process on the dynamic temperature threshold; When it is determined that the temperature of the first server has dropped according to the target temperature change data, the dynamic temperature threshold is set to a second threshold, wherein the second threshold is a threshold obtained by performing a threshold increase processing on the dynamic temperature threshold, and the second threshold is greater than the first threshold.
2. The method for controlling heat dissipation of a server according to claim 1, wherein: Inputting the real-time operating data and the dynamic temperature threshold of the first server into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the target fan speed includes: performing data comparison processing on the server temperature in the real-time operation data and the dynamic temperature threshold to obtain a first comparison result, wherein the server temperature is the real-time temperature of the first server; If it is determined according to the first comparison result that the server temperature is greater than the dynamic temperature threshold, inputting the real-time operating data and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a first fan speed, wherein the first heat dissipation control algorithm is a heat dissipation control algorithm based on deep reinforcement learning extracted from a preset edge computing device; If it is determined according to the first comparison result that the server temperature is less than or equal to the dynamic temperature threshold, inputting the real-time operating data and the dynamic temperature threshold into the first heat dissipation control algorithm to perform fan speed calculation processing on the first server to obtain a second fan speed; The target fan speed includes the first fan speed and the second fan speed, and the second fan speed is less than the first fan speed.
3. The method for controlling heat dissipation of a server according to claim 2, wherein: Inputting the real-time operating data and the dynamic temperature threshold into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the first fan speed includes: Calculating a first temperature difference between the server temperature and the dynamic temperature threshold; The first temperature difference, the real-time operating data, and the dynamic temperature threshold are input into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the first fan speed, wherein the magnitude of the first temperature difference is positively correlated with the magnitude of the first fan speed.
4. The method for controlling heat dissipation of a server according to claim 2, wherein: Inputting the real-time operating data and the dynamic temperature threshold into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the second fan speed includes: calculating a second temperature difference between the server temperature and the dynamic temperature threshold; The second temperature difference, the real-time operating data, and the dynamic temperature threshold are input into the first heat dissipation control algorithm to calculate the fan speed of the first server to obtain the second fan speed, wherein the magnitude of the second temperature difference is negatively correlated with the magnitude of the second fan speed.
5. The method for controlling heat dissipation of a server according to claim 1, wherein: The adjusting the fan speed of the first server according to the target fan speed to obtain the adjusted first server and the second operating data of the adjusted first server includes: transmitting the target fan speed to a fan controller of the first server, and controlling the fan speed of the first server to the target fan speed by the fan controller, thereby obtaining the adjusted first server; When it is determined that the fan speed of the first server is the target fan speed, the operating data of the first server is acquired to obtain the second operating data.
6. The method for controlling heat dissipation of a server according to claim 1, wherein: The performing data evaluation and analysis on the second operating data to obtain the adjusted heat dissipation evaluation result corresponding to the first server includes: The second operating data is analyzed and processed according to the historical operating data and the preset evaluation index to obtain the heat dissipation evaluation result, wherein the heat dissipation evaluation result at least includes a temperature evaluation result, a rotation speed evaluation result and an energy consumption evaluation result.
7. The method for controlling heat dissipation of a server according to claim 6, wherein: After performing data evaluation and analysis on the second operating data to obtain the adjusted heat dissipation evaluation result corresponding to the first server, the method further includes: Comparing the temperature evaluation result with the temperature evaluation indicators in the preset evaluation indicators to obtain a temperature evaluation deviation value, wherein the preset evaluation indicators include at least a speed evaluation indicator, an energy consumption evaluation indicator, and the temperature evaluation indicator; Comparing the speed evaluation result with the speed evaluation index to obtain a speed evaluation deviation value, and comparing the energy consumption evaluation result with the energy consumption evaluation index to obtain an energy consumption evaluation deviation value; Determine whether the temperature evaluation result meets the preset heat dissipation condition according to the temperature evaluation deviation value, determine whether the speed evaluation result meets the preset heat dissipation condition according to the speed evaluation deviation value, and determine whether the energy consumption evaluation result meets the preset heat dissipation condition according to the energy consumption evaluation deviation value; If any one of the temperature evaluation result, the rotation speed evaluation result, and the energy consumption evaluation result does not meet the preset heat dissipation condition, it is determined that the heat dissipation evaluation result does not meet the preset heat dissipation condition.
8. The method for controlling heat dissipation of a server according to claim 7, wherein: Determining whether the temperature evaluation result meets the preset heat dissipation condition according to the temperature evaluation deviation value, determining whether the speed evaluation result meets the preset heat dissipation condition according to the speed evaluation deviation value, and determining whether the energy consumption evaluation result meets the preset heat dissipation condition according to the energy consumption evaluation deviation value includes: When the temperature evaluation deviation value is greater than a first preset evaluation deviation threshold, determining that the temperature evaluation result does not meet the preset heat dissipation condition; When the rotation speed evaluation deviation value is greater than a second preset evaluation deviation threshold, determining that the rotation speed evaluation result does not meet the preset heat dissipation condition; When the energy consumption evaluation deviation value is greater than a third preset evaluation deviation threshold, it is determined that the energy consumption evaluation result does not meet the preset heat dissipation condition.
9. The method for controlling heat dissipation of a server according to claim 1, wherein: The performing algorithm optimization processing on the first heat dissipation control algorithm based on a deep reinforcement learning mechanism according to the second operating data, the heat dissipation evaluation result, and the historical operating data to obtain a second heat dissipation control algorithm includes: performing algorithm optimization processing on the first heat dissipation control algorithm by using a preset deep reinforcement learning algorithm according to the second operating data, the heat dissipation evaluation result, and the historical operating data; Until a preset number of optimizations or a preset convergence condition is reached, the second heat dissipation control algorithm is obtained.
10. The method for controlling heat dissipation of a server according to claim 1, wherein: After transmitting the second heat dissipation control algorithm to the preset edge computing device, the method further includes: In response to a control instruction for performing heat dissipation control on the second server, the second heat dissipation control algorithm is extracted from the preset edge computing device to perform heat dissipation control processing on the second server.
11. The method for controlling heat dissipation of a server according to claim 10, wherein: Extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server includes: Acquire the third operating data of the second server, and perform data comparison processing based on the third operating data and the real-time operating data to obtain operating difference data between the first server and the second server; When the operation difference data satisfies a preset first operation difference condition, extracting the second heat dissipation control algorithm from the preset edge computing device to perform heat dissipation control processing on the second server; When the operation difference data does not satisfy the preset first operation difference condition but satisfies the second operation difference condition, adjusting the second heat dissipation control algorithm according to a preset algorithm update rule in the preset edge computing device to obtain an adjusted second heat dissipation control algorithm, and performing heat dissipation control processing on the second server according to the adjusted second heat dissipation control algorithm; If the operation difference data does not satisfy the second operation difference condition, obtaining second historical data of the second server, and performing algorithm optimization processing on the second heat dissipation control algorithm based on the second historical data and the third operation data to obtain an optimized second heat dissipation control algorithm; Perform heat dissipation control processing on the second server according to the optimized second heat dissipation control algorithm.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor is configured to implement the steps of the method for controlling heat dissipation of a server as claimed in any one of claims 1 to 11 when executing the computer program.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method for controlling heat dissipation of a server according to any one of claims 1 to 11 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for controlling heat dissipation of a server as claimed in any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Pure electric vehicle cooling system control method based on deep reinforcement learning
CN109193075A
Self-adaptive setting method and device for fan control parameters, storage medium and electronic equipment
CN118672368A