Control methods, devices, storage media and electronic equipment for heat dissipation equipment
By using a reinforcement learning network model to adjust the control actions of the heat dissipation equipment in real time, the problem of inaccurate heat dissipation under fixed threshold control is solved, and precise heat dissipation and energy optimization of the equipment are achieved.
Patent Information
- Application Number
- CN202510815951.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-06-18
AI Technical Summary
In existing technologies, fixed threshold control heat dissipation devices cannot accurately adapt to changes in equipment load and environment, resulting in inaccurate heat dissipation.
A reinforcement learning network model is used to obtain the state parameters of the target device, determine the dynamic state characteristics, and optimize the control actions of the heat dissipation device based on these characteristics. This includes real-time adjustments using a time loop network and a reinforcement learning network model.
It enables precise control of the heat dissipation equipment, adapts to changes in equipment load and environment, improves the accuracy and efficiency of heat dissipation, reduces fan power consumption, and optimizes energy use.
Smart Images

Figure CN120335582B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a control method, apparatus, storage medium and electronic device for a heat dissipation device. Background Technology
[0002] In related technologies, a common method for cooling equipment is to control its activation and deactivation when the equipment temperature reaches a fixed threshold. However, when the equipment load or environment changes continuously, this method of controlling the activation and deactivation of the cooling equipment using a fixed threshold cannot provide precise cooling.
[0003] This indicates that the relevant technologies have a problem with accurately dissipating heat from the equipment.
[0004] There is currently no effective solution to the aforementioned problems in the relevant technologies. Summary of the Invention
[0005] This application provides a control method, apparatus, storage medium, and electronic device for a heat dissipation device, to at least solve the problem of inaccurate heat dissipation for devices in the related art.
[0006] This application provides a control method for a heat dissipation device, comprising: acquiring first state parameters of a target device within a predetermined time period; determining first dynamic state features of the first state parameters, wherein the first dynamic state features are used to represent the dependency relationship between each parameter included in the first state parameters; inputting the first dynamic state features into a reinforcement learning network model to determine a target control action, wherein the reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device before and after the target device operates according to the control action, and the control action is an action determined by the reinforcement learning network model before determining the target control action; and controlling the heat dissipation device to operate according to the target control action.
[0007] In an exemplary embodiment, determining a first dynamic state feature of a first state parameter includes: determining a time difference feature between each parameter included in the first state parameter; and inputting the time difference feature and the first state parameter into a time loop network to obtain the first dynamic state feature.
[0008] In an exemplary embodiment, after inputting the first dynamic state features into the reinforcement learning network model to determine the target control action, the method further includes: determining second state parameters of the target device; determining a reward function based on the first state parameters and the second state parameters; and updating the model parameters of the reinforcement learning network model based on the reward function.
[0009] In one exemplary embodiment, determining a reward function based on a first state parameter and a second state parameter includes: determining a maximum temperature, an average temperature, and a first rotational speed of the heat dissipation device included in the second state parameter; determining the power consumption of the heat dissipation device and a second rotational speed included in the first state parameter; and determining a reward function based on the maximum temperature, average temperature, first rotational speed, power consumption, and second rotational speed.
[0010] In one exemplary embodiment, determining a reward function based on a maximum temperature, an average temperature, a first rotational speed, power consumption, and a second rotational speed includes: determining a first difference between the maximum temperature and a predetermined safe temperature; determining a target distance between the average temperature and a predetermined target temperature; determining the absolute value of a second difference between the first rotational speed and the second rotational speed; determining a first negative of the product of the first difference and a first weight; determining a second negative of the product of the target distance and the second weight; determining a third negative of the product of power consumption and a third weight; and determining a fourth negative of the product of the absolute value and a fourth weight; and determining the sum of the first negative, second negative, third negative, and fourth negative as the reward function.
[0011] In an exemplary embodiment, updating the model parameters of a reinforcement learning network model based on a reward function includes: inputting a first dynamic state feature and a target control action into a value network included in the reinforcement learning network model to obtain a first value evaluation value of the target control action; updating the first model parameters of a policy network included in the reinforcement learning network model based on the reward function and the first value evaluation value, wherein the policy network is used to determine the target control action; and updating the second model parameters of the value network based on the reward function and the first value evaluation value.
[0012] In one exemplary embodiment, updating the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the value assessment value includes: determining the policy eligibility trace based on the reward function; and updating the first model parameters based on the policy eligibility trace.
[0013] In one exemplary embodiment, updating the second model parameters of a value network based on a reward function and a first value evaluation value includes: determining a second dynamic state feature of the second state parameters; determining a second value evaluation value corresponding to the second dynamic state feature; and updating the second model parameters based on the reward function, the first value evaluation value, and the second value evaluation value.
[0014] In an exemplary embodiment, obtaining a first state parameter of a target device within a predetermined time period includes: synchronously acquiring second state parameters collected by multiple parameter acquisition devices at a target frequency, wherein the parameter acquisition devices are devices deployed in the target device for acquiring the first state parameters; determining a sliding time window, wherein the duration of the sliding time window is the same as the duration of the predetermined time period; and determining the first state parameter from the second state parameters according to the sliding time window.
[0015] In one exemplary embodiment, controlling the heat dissipation device to operate according to a target control action includes: adding pre-set noise to the target control action to obtain a processing control action; and controlling the heat dissipation device to operate according to the processing control action.
[0016] In one exemplary embodiment, before inputting the first dynamic state feature into the reinforcement learning network model, the method further includes: determining a pre-defined sparsity; and initializing the network parameters of the reinforcement learning network model based on the sparsity.
[0017] This application also provides a control device for a heat dissipation device, comprising: an acquisition module for acquiring first state parameters of a target device within a predetermined time period; a first determination module for determining first dynamic state features of the first state parameters, wherein the first dynamic state features represent the dependencies between each parameter included in the first state parameters; a second determination module for inputting the first dynamic state features into a reinforcement learning network model to determine a target control action, wherein the reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device before and after the target device operates according to the control action, and the control action is an action determined by the reinforcement learning network model before determining the target control action; and a control module for controlling the heat dissipation device to operate according to the target control action.
[0018] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the control method of any of the above-described heat dissipation devices.
[0019] This application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the control method for any of the above-described heat dissipation devices.
[0020] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the control method for any of the above-described heat dissipation devices.
[0021] Through this application, since the reinforcement learning network model is iteratively updated based on the state parameters of the target device's heat dissipation equipment before and after the control action, the reinforcement learning network model can continuously evolve and optimize with continuous interaction with the environment. This allows the reinforcement learning network model to automatically adapt to changes in device load and environment. Furthermore, the model input of the reinforcement learning network model is the first dynamic state feature, which characterizes the dependencies between parameters of the target device. Therefore, by using the key feature vectors of the target device's current dynamic state and future trends, the reinforcement learning network model can accurately determine the target control action. Through accurate target control actions, it can then precisely control the heat dissipation of the heat dissipation equipment. Therefore, it can solve the technical problem in related technologies that cannot accurately dissipate heat from the device, achieving the technical effect of accurately dissipating heat from the device. Attached Figure Description
[0022] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a control method for a heat dissipation device according to an embodiment of this application;
[0024] Figure 2 This is a system schematic diagram of a control method for an execution device according to an embodiment of this application;
[0025] Figure 3 This is a flowchart of a control method for a heat dissipation device according to a specific embodiment of this application;
[0026] Figure 4 This is a schematic diagram of the control device structure of a heat dissipation device according to an embodiment of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.
[0028] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0029] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] The specific application environment architecture or specific hardware architecture on which the control method for the heat dissipation device depends is described here.
[0031] The control methods for heat dissipation equipment can be applied to devices equipped with heat dissipation devices, such as servers, computers, laptops, and desktop computer cases. These heat dissipation devices can include fans, heat sinks, etc.
[0032] The embodiments of this application provide a control method for a heat dissipation device. The method is described in detail below in conjunction with the execution flow of the control method for the heat dissipation device.
[0033] Figure 1 This is a flowchart of a control method for a heat dissipation device according to an embodiment of this application, such as... Figure 1 As shown, the process includes:
[0034] Step S102: Obtain the first status parameters of the target device within a predetermined time period;
[0035] In this embodiment, the target device can be a server, computer, laptop, desktop computer chassis, etc. The first state parameter can include the temperature of various components of the target device, such as the temperature of the central processing unit (CPU), the temperature of the memory, the temperature of the power supply, and the temperature of the air inlet and outlet. It can also include the power consumption of various components in the target device, the power consumption of the cooling device, the rotation speed, etc.
[0036] In this embodiment, high-frequency, high-precision sensors can be deployed in various components of the target device. These sensors may include temperature sensors, power consumption sensors, and speed sensors. The execution system of the heat dissipation device control method can continuously collect data (temperature, power consumption, fan speed, etc.) from high-frequency, high-precision sensors deployed in key locations inside the server (such as CPU, GPU, memory, power supply, air inlet / outlet vents, etc.) to form a multi-dimensional real-time data stream. The first state parameter is obtained based on the real-time data stream.
[0037] Specifically, when the target device is a server, a high-density sensor network can be deployed within the server, including temperature sensors (thermocouples or digital temperature sensors with an accuracy requirement of ±0.5°C), power consumption sensors (monitoring the power consumption of components such as the CPU, GPU (graphics processor), and fans), and fan speed sensors (providing closed-loop feedback). The data collected by the high-density sensor network is determined as the first state parameter.
[0038] Step S104: Determine the first dynamic state feature of the first state parameter, wherein the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameter;
[0039] In this embodiment, the original sensor data stream, i.e., the first state parameters, can be fed into a streaming feature extraction module. The streaming feature extraction module can use time series analysis techniques to extract key feature vectors, i.e., the first dynamic state features, from the data stream in real time, reflecting the current dynamic thermal state and future trends of the system. The streaming feature extraction module can include convolutional networks, recurrent temporal networks, etc. The recurrent temporal network can be a Long Short-Term Memory (LSTM) network.
[0040] Step S106: Input the first dynamic state feature into the reinforcement learning network model to determine the target control action. The reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device of the target device before and after the control action is executed. The control action is the action determined by the reinforcement learning network model before the target control action is determined.
[0041] In this embodiment, the target control action can be an action performed by controlling the heat dissipation device. When the heat dissipation device is a fan, the target control action can include speed, time, etc. It can also include the speed corresponding to each time point, that is, the fan can rotate at the same speed or at different speeds. For example, it can first rotate at high speed and then at low speed. The high speed rotation first is to quickly reduce the temperature of the target device, and then the low speed rotation is to achieve the purpose of precise temperature control, so as to achieve a balance between heat dissipation effect (maintaining a safe temperature and close to the target temperature) and energy efficiency (minimizing fan power consumption).
[0042] In this embodiment, the reinforcement learning network model can be a streaming reinforcement learning network model, represented by the Stream-X algorithm, which focuses on processing continuous data streams and performing real-time, efficient online incremental learning. Its core characteristics include: instant parameter updates based on single or mini-batch samples, no need to store massive amounts of historical data, and effective handling of delayed rewards and nonlinear relationships using mechanisms such as Eligibility Traces. These features naturally align with the real-time, dynamic, and resource-constrained requirements of target device heat dissipation control. A target device heat dissipation control system capable of real-time, rapid perception, intelligent decision-making, and continuous optimization can be constructed based on the streaming reinforcement learning network model to dissipate heat from the target device.
[0043] Step S108: Control the heat dissipation equipment to operate according to the target control action.
[0044] In this embodiment, the target control action can be converted into pulse width modulation (PWM) and the PWM can be sent to the heat dissipation device so that the heat dissipation device can operate according to the PWM to reduce the temperature of the target device.
[0045] Optionally, the entity performing the above steps can be the target device, processor, or a device with integrated data processing capabilities. When the target device is a server and the cooling device is a fan, a system diagram illustrating the control method for executing the device can be found in the appendix. Figure 2 ,like Figure 2 As shown, the first state parameters, such as temperature and power consumption data, can be obtained by the sensor and sent to the reinforcement learning agent (RL agent). The RL agent determines the target control action and sends the target control action to the board management controller (BMC). The BMC determines the PWM based on the target control action and sends the PWM to the fan to control the fan rotation.
[0046] Through this application, since the reinforcement learning network model is iteratively updated based on the state parameters of the target device's heat dissipation equipment before and after the control action, the reinforcement learning network model can continuously evolve and optimize with continuous interaction with the environment. This allows the reinforcement learning network model to automatically adapt to changes in device load and environment. Furthermore, the model input of the reinforcement learning network model is the first dynamic state feature, which characterizes the dependencies between parameters of the target device. Therefore, by using the key feature vectors of the target device's current dynamic state and future trends, the reinforcement learning network model can accurately determine the target control action. Through accurate target control actions, it can then precisely control the heat dissipation of the heat dissipation equipment. Therefore, it can solve the technical problem in related technologies that cannot accurately dissipate heat from the device, achieving the technical effect of accurately dissipating heat from the device.
[0047] In an exemplary embodiment, determining a first dynamic state feature of a first state parameter includes: determining a time difference feature between each parameter included in the first state parameter; inputting the time difference feature and the first state parameter into a time loop network to obtain the first dynamic state feature. In this embodiment, the first state parameter can be processed to capture dynamic information, i.e., the first dynamic state feature. For example, one-dimensional causal convolution can be used to calculate the time difference between data points to obtain a time difference feature, which characterizes the rate and trend of temperature / power consumption change. Causal convolution ensures that the calculation depends only on current and past data, meeting real-time requirements. Its mathematical form can be expressed as: DiffFeat t =Conv1D causal (Windoes(X t )), where Conv1D causal It is a one-dimensional causal convolution, Windoees(X) t ) represents the data from the most recent W time steps, i.e., the first state parameter.
[0048] In this embodiment, the time-recurrent network can be an LSTM, which inputs the first state parameters and extracted temporal difference features into an online-updating Long Short-Term Memory (LSTM) network. The recurrent structure of the LSTM enables it to effectively capture long-term dependencies in time-series data, integrate historical information, and output a hidden state that encodes the complex dynamics of the current system. This process can be represented as: h t =LSTM online (InputFeat t h (t-1) ), where InputFeat t It is [X] (t-w+1) …, X t [with DiffFeat] tThe pipeline combines these elements. Ultimately, the pipeline outputs a fixed-dimensional feature vector S. t , that is, the first dynamic state feature, which is the state representation of the reinforcement learning agent at time step t. Where h t It is S t The intermediate representation, in the continuous calculation process, is the last h. t That is, S t .
[0049] In the above embodiments, temporal difference features, such as the rate of change of server temperature over time, help the model capture the trend of parameter changes over time. Temporal recurrent networks (such as LSTM or GRU) can process sequential data, enabling the model to understand the dependencies between parameters, including but not limited to the interaction between temperature and fan speed. Through this technical feature, the model can more accurately predict the control effect of the cooling equipment, thereby making more optimized decisions and achieving precise control of server temperature.
[0050] In an exemplary embodiment, after inputting the first dynamic state feature into the reinforcement learning network model to determine the target control action, the method further includes: determining a second state parameter of the target device; determining a reward function based on the first state parameter and the second state parameter; and updating the model parameters of the reinforcement learning network model based on the reward function. In this embodiment, the current first state parameter of the target device can be determined as S. t The target control action is defined as At. After the heat dissipation device executes At, the state of the target device transitions to a new second state parameter S. t+1 The reward function can be determined based on the first and second state parameters, and the model parameters of the reinforcement learning network model can be updated according to the reward function.
[0051] In the above embodiments, the reward function can be used to quantify the quality of the target control action, namely, the heat dissipation effect and the power consumption cost. By quantifying the heat dissipation effect and power consumption cost, the model is guided to learn a better control strategy. For example, if the second state parameters show that the server temperature decreases and power consumption decreases, the reward function value is higher, and vice versa. By updating the model parameters, the model can continuously optimize its decision-making process, including but not limited to adjusting the control strategy of the heat dissipation equipment to achieve better heat dissipation and lower power consumption.
[0052] In one exemplary embodiment, determining a reward function based on first state parameters and second state parameters includes: determining a maximum temperature, an average temperature, and a first rotational speed of the heat dissipation device included in the second state parameters; determining the power consumption of the heat dissipation device and a second rotational speed included in the first state parameters; and determining a reward function based on the maximum temperature, average temperature, first rotational speed, power consumption, and second rotational speed. In this embodiment, the second state parameters may include the temperatures of various components, and the maximum temperature and average temperature may be determined. The second state parameters may also include the first rotational speed of the heat dissipation device. Furthermore, the power consumption of the heat dissipation device and the second rotational speed included in the first state parameters may be determined, and the reward function is determined based on the maximum temperature, average temperature, first rotational speed, power consumption, and second rotational speed.
[0053] In this embodiment, the maximum temperature can be compared with a preset safety threshold. If the maximum temperature is lower than the safety threshold, the reward value is increased. The deviation between the average temperature and the target operating temperature is calculated, and the reward value is adjusted according to the magnitude of the deviation; the smaller the deviation, the higher the reward value. Based on the difference between the first and second rotational speeds, a smoothness penalty is introduced. If the change in the first and second rotational speeds is large, the reward value is reduced to promote the smoothness of the control action. Based on the power consumption of the heat dissipation device, an energy consumption cost penalty is introduced; the higher the power consumption, the greater the penalty, thereby reducing the reward value. Taking into account the maximum temperature, average temperature, the difference between the first and second rotational speeds, and power consumption, the final value of the reward function is determined by a weighted summation. The weights of each parameter are flexibly adjusted according to the priority of heat dissipation control and energy efficiency requirements to balance the relationship between heat dissipation effect and energy consumption.
[0054] Among these features, the safety threshold comparison emphasizes the importance of incorporating safe temperature considerations into the reward function, ensuring that the agent's decisions do not exceed the safe temperature range. Temperature deviation adjustment focuses on maintaining the temperature within the target operating range to improve the overall efficiency and stability of the system. Smoothness penalty addresses the rate of change of control actions, aiming to avoid unnecessary frequent adjustments and drastic fluctuations, which helps extend hardware lifespan and reduce noise. Energy cost penalty introduces a direct consideration of power consumption, aiming to encourage the agent to adopt more energy-efficient strategies, aligning with the development trend of green computing. Weighted summation proposes a comprehensive evaluation mechanism; by adjusting parameter weights, the agent can flexibly respond to the heat dissipation and energy consumption balance requirements under different scenarios, achieving more comprehensive optimization.
[0055] In this embodiment, the determination of the reward function takes into account the balance between temperature control and power consumption costs. For example, a decrease in maximum and average temperature will increase the value of the reward function, while an increase in power consumption and drastic changes in rotation speed will decrease the value of the reward function. This design allows the model to automatically balance heat dissipation and power consumption costs during the learning process, achieving more efficient and stable server temperature control.
[0056] In one exemplary embodiment, determining a reward function based on a maximum temperature, an average temperature, a first rotational speed, power consumption, and a second rotational speed includes: determining a first difference between the maximum temperature and a predetermined safe temperature; determining a target distance between the average temperature and a predetermined target temperature; determining the absolute value of a second difference between the first rotational speed and the second rotational speed; determining a first negative of the product of the first difference and a first weight; determining a second negative of the product of the target distance and a second weight; determining a third negative of the product of power consumption and a third weight; and determining a fourth negative of the product of the absolute value and a fourth weight; and determining the sum of the first negative, second negative, third negative, and fourth negative values as the reward function. In this embodiment, action A is performed. t Afterwards, the system transitions to a new state S. (t+1) and received a reward signal R t Reward function R(S) t A t ,S (t+1) It is designed to quantify the quality of an action, and its form is usually as follows: .in, Indicates the first weight. Indicates the second weight. Indicates the third weight. This indicates the fourth weight. The maximum temperature is... The safe temperature is , Indicates average temperature. Indicates the target temperature. Indicates power consumption, |△a t | represents the absolute value of the second difference between the first speed and the second speed, i.e., the change in control action. - These are weighting coefficients used to balance different objectives.
[0057] In this embodiment, by introducing weights, the model's emphasis on different parameters can be flexibly adjusted. For example, if maintaining a safe temperature is more important, the first weight can be increased, causing the model to focus more on temperature control and less on power consumption when making decisions. This design allows the model to adapt to different application scenarios, including but not limited to thermal management in data centers and temperature control in high-performance computing servers, enabling wider application.
[0058] In an exemplary embodiment, updating the model parameters of a reinforcement learning network model based on a reward function includes: inputting a first dynamic state feature and a target control action into a value network included in the reinforcement learning network model to obtain a first value evaluation value of the target control action; updating the first model parameters of a policy network included in the reinforcement learning network model based on the reward function and the first value evaluation value, wherein the policy network is used to determine the target control action; and updating the second model parameters of the value network based on the reward function and the first value evaluation value. In this embodiment, the reinforcement learning network model may include a policy network and a value network. The policy network is in state S. t Given input, output an action A. t Action A t It can be a one-dimensional vector normalized to the interval [0, 1], representing the PWM duty cycle of the corresponding fan. The value network uses state S. t And Action A t As input, the Agent outputs the current state and the value assessment of the action. At each control decision point t, the Agent, based on the current state S,... t and strategy π θ Select action A t To balance exploration and exploitation, noise can be added to the output when selecting actions.
[0059] In the above embodiments, the obtained empirical tuples (S) can be utilized. t A t R t S (t+1) This is used to update the policy network parameters θ, i.e., the first model parameters, and the value network parameters φ, i.e., the second model parameters.
[0060] In the above embodiments, the updating of the value evaluation model and the policy network is the core of the reinforcement learning model's learning. Guided by the reward function, the model can continuously optimize its policy network to determine better control actions, while simultaneously updating the value evaluation model, enabling the model to more accurately evaluate the value of different control actions. This mechanism ensures that the model can continuously learn and optimize, including but not limited to automatically adjusting the heat dissipation strategy when the server load changes, achieving more efficient and intelligent temperature control.
[0061] In an exemplary embodiment, updating the first model parameters of the policy network included in the reinforcement learning network model based on a reward function and a value evaluation value includes: determining a policy eligibility trace based on the reward function; and updating the first model parameters based on the policy eligibility trace. In this embodiment, the policy network is used to generate actions, and the value network is used to evaluate the quality of the actions. A policy eligibility trace z can be introduced when updating the policy network. θ , The parameters have the following meanings:
[0062] Discount factor: Values range from 0 to 1 and are used to weigh the importance of future rewards. The closer the value is to 1, the more the model emphasizes long-term rewards.
[0063] The trace decay parameter, ranging from 0 to 1, controls the weight of historical information. It is similar to the decay of qualification traces in reinforcement learning. The smaller the value, the faster the historical impact diminishes.
[0064] Policy gradient: Represents the gradient of the logarithmic probability of the policy with respect to the parameter θ, guiding how to adjust θ to increase the probability of action A in state S.
[0065] The entropy hyperparameter controls the weight of the entropy term and affects the intensity of the exploration.
[0066] sign( The sign of the TD error is indicated by: = TD error represents the difference between the actual reward of an action and its estimated value. The value can be +1 (positive) or -1 (negative), reflecting the quality of the action.
[0067] Path entropy: used to measure the strategy The randomness of the strategy. The greater the entropy, the more random the strategy; the smaller the entropy, the more deterministic the strategy.
[0068] Entropy term Based on the direction adjustment of TD error: if > 0 (actual return exceeded expectations), sign( = +1, the entropy term encourages more exploration and maintains the randomness of the strategy.
[0069] if < 0 (actual return is lower than expected), sign( = -1, the entropy term reduces exploration, prompting the strategy to favor known good moves.
[0070] In this embodiment, to ensure the stability of online learning, the policy network introduces a pruning mechanism, KL divergence, and ObGD constraints to limit the step size of each policy network update: The parameters have the following meanings:
[0071] Probability ratio: Action A under the new strategy and the old strategy tThe ratio of probabilities reflects the magnitude of policy updates.
[0072] Advantage function: Estimates action A in state S t The relative quality of something is usually determined by... calculate.
[0073] It is a clipping function: [The function will...] Limited to | Within the range, It is the pruning factor (usually 0.1 to 0.2) to prevent the strategy update from being too large.
[0074] It is the KL divergence penalty coefficient: it controls the weight of the KL divergence and limits the difference between the old and new strategies.
[0075] KL divergence: a measure of a new strategy and old strategies The larger the value, the greater the strategy change.
[0076] ObGD constraint: is a kind of policy-based qualification trace A gradient descent variant.
[0077] By using the min function and clip operation, it is ensured that the policy update does not deviate too far from the old policy, thus avoiding the instability of traditional policy gradient methods. KL divergence: as an additional regularization term, it limits the difference between the old and new policies, further enhancing stability.
[0078] In the above embodiments, the policy qualification trace is the trajectory of the policy network during the decision-making process. Guided by the reward function, the model can identify which decision paths are effective and which are ineffective, thereby optimizing the parameters of the policy network. This mechanism enables the model to continuously learn and improve its decision-making strategy, including but not limited to automatically adjusting the control parameters of the heat dissipation equipment under different environmental conditions to achieve more precise temperature control.
[0079] In an exemplary embodiment, updating the second model parameters of a value network based on a reward function and a first value evaluation value includes: determining a second dynamic state feature of the second state parameters; determining a second value evaluation value corresponding to the second dynamic state feature; and updating the second model parameters based on the reward function, the first value evaluation value, and the second value evaluation value. In this embodiment, the second value evaluation value can be represented as... The first value evaluation value can be expressed as The value network can be updated by minimizing the square of the TD error. The loss function can be expressed as... These updates are made in real time based on current experience. Among them, It is the target value, which represents the value of the next state after the current reward is added to the discount. It reflects the value network's prediction of the current state. This indicates the difference between the target value and the estimated value.
[0080] In the above embodiments, the value evaluation model is updated based on current and future state characteristics, as well as the reward function, ensuring that the model can accurately assess the impact of the current control action on the future state. This mechanism enables the model to more intelligently predict heat dissipation effects, including but not limited to adjusting heat dissipation strategies in advance when server load changes, achieving more efficient and stable temperature control.
[0081] In an exemplary embodiment, obtaining a first state parameter of a target device within a predetermined time period includes: synchronously acquiring second state parameters collected by multiple parameter acquisition devices at a target frequency, wherein the parameter acquisition devices are devices deployed in the target device for acquiring the first state parameters; determining a sliding time window, wherein the duration of the sliding time window is the same as the duration of the predetermined time period; and determining the first state parameter from the second state parameters according to the sliding time window. In this embodiment, a real-time feature extraction pipeline can be established. This pipeline receives a continuous data stream X. t This refers to the second state parameter. A sliding window mechanism is applied to extract data from the most recent W time steps. This yields the first state parameters. The data acquisition unit can be configured to use the low-latency synchronous PTP protocol to ensure high-frequency synchronous acquisition of all sensor data, with a sampling period down to the millisecond level, forming a vectorized time-series data stream. Where t represents the discrete time step. This represents the CPU and GPU temperature and power consumption at time t, as well as the memory temperature and fan speed. The specific value of N varies depending on the server configuration.
[0082] In the above embodiments, the selection of the target frequency, for example, collecting data every 10 seconds, ensures that the model obtains real-time state parameters, thereby making more timely control decisions. This design allows the model to adapt to different server environments and achieve wider application. A sliding time window, such as temperature data from the last 5 minutes, helps the model understand the short-term trends of current state parameters. This design enables the model to more accurately capture parameter change trends, including but not limited to responding quickly when server load suddenly increases, achieving more precise temperature control.
[0083] In one exemplary embodiment, controlling a heat dissipation device to operate according to a target control action includes: adding pre-set noise to the target control action to obtain a processed control action; and controlling the heat dissipation device to operate according to the processed control action. In this embodiment, the noise can be pre-set noise, such as Gaussian white noise. Adding noise can increase the likelihood of the agent exploring the environment. In the early stages of reinforcement learning, the agent needs to try different actions to understand the environment's response, thereby learning the optimal strategy. By adding noise to the target control action, the agent can deviate from the original strategy to some extent, trying action paths that have not been explored or have been explored less, which helps to discover potentially better behavioral strategies.
[0084] Adding noise during training helps prevent agent policies from overfitting to specific environmental states. Overfitting refers to the phenomenon where a model performs well on training data but generalizes poorly on unseen data. By adding noise to control actions, the agent encounters more variations during learning, which encourages the policy network to learn more generalized features rather than simply memorizing patterns from the training data.
[0085] When facing rapidly changing environments, introducing noise can improve the real-time response capability of control systems. Agents can not only make decisions based on the current environmental state, but also, through the introduction of noise, their strategies can anticipate and adapt to potential future state changes to a certain extent.
[0086] In server heat dissipation control scenarios, the aforementioned benefits mean that the agent can adjust heat dissipation equipment (such as fans) more flexibly and intelligently, maximizing energy efficiency and reducing hardware wear while meeting heat dissipation requirements. It can also quickly adapt to changes in server load, ambient temperature, and other conditions, maintaining stable and efficient heat dissipation performance.
[0087] In an exemplary embodiment, before inputting the first dynamic state feature into the reinforcement learning network model, the method further includes: determining a pre-defined sparsity; and initializing the network parameters of the reinforcement learning network model based on the sparsity. In this embodiment, a reinforcement learning agent may be initialized, which includes a policy network π. θ (A|S) and Value Network V ∅ (S). Both networks use layer normalization to enhance stability. During initialization, the weights of the policy network and value network can be initialized using the SparseInit technique, where the weights w are initialized according to sparsity, and all biases are set to 0.
[0088] In the above embodiments, the sparsity setting, such as the proportion of zero values in the network parameters, can reduce the number of parameters in the model and improve the training efficiency and generalization ability of the model. This design enables the model to achieve more efficient learning with limited computing resources, including but not limited to deploying reinforcement learning models on edge servers to achieve more intelligent temperature control.
[0089] The control method for the heat dissipation device is described below with reference to specific implementation methods:
[0090] Figure 3 This is a flowchart of a control method for a heat dissipation device according to a specific embodiment of this application, such as... Figure 3 As shown, the process includes:
[0091] Step S302, High-density sensor network. This involves system initialization and high-frequency data stream construction.
[0092] Deploy a high-density sensor network within the server, including temperature sensors (thermocouples or digital temperature sensors with an accuracy requirement of ±0.5°C), power consumption sensors (monitoring the power consumption of components such as CPU, GPU, and fans), and fan speed sensors (providing closed-loop feedback).
[0093] The data acquisition unit is configured with a low-latency synchronous PTP protocol to ensure high-frequency synchronous acquisition of all sensor data, with a sampling period down to the millisecond level, forming a vectorized time-series data stream. Where t represents the discrete time step. This represents the CPU and GPU temperatures and power consumption at time t, as well as the memory temperature and fan speed. The specific value of N varies depending on the server configuration.
[0094] Step S304, Streaming Feature Extraction Engineering Layer. This involves streaming feature extraction and state construction.
[0095] Establish a real-time feature extraction pipeline. This pipeline receives a continuous data stream. (Corresponding to the second state parameter mentioned above).
[0096] First, a sliding window mechanism is applied to extract data from the most recent W time steps. (Corresponding to the first state parameter mentioned above).
[0097] Secondly, the data within the window is processed to capture dynamic information. One-dimensional causal convolution is used to calculate the temporal difference between data points (corresponding to the aforementioned temporal difference feature) to characterize the rate and trend of temperature / power consumption change. Causal convolution ensures that the calculation depends only on current and past data, meeting real-time requirements. Its mathematical form can be expressed as: ,in It is a one-dimensional causal convolution. This represents the data from the most recent W time steps.
[0098] Subsequently, the original data and extracted differential features are input into an online-updated Long Short-Term Memory (LSTM) network. The recurrent structure of the LSTM enables it to effectively capture long-term dependencies in time-series data, integrate historical information, and output a hidden state that encodes the complex dynamics of the current system. This process can be represented as: ,in yes and The combination of.
[0099] Ultimately, the pipeline outputs a fixed-dimensional feature vector. (Corresponding to the first dynamic state feature mentioned above), this vector is the state representation of the reinforcement learning agent at time step t.
[0100] Step S306, Policy Network Layer.
[0101] Initialize a reinforcement learning agent that contains a policy network. With value network Both networks employ layer normalization to enhance stability. During initialization, the weights of the policy and value networks are initialized using the SpareseInit technique, with weights initialized based on sparsity. All biases are set to 0.
[0102] Policy networks are stateful Given input, output an action. ,action It is a one-dimensional vector normalized to the [0,1] interval, representing the PWM duty cycle of the corresponding fan. The value network is based on states. and actions As input, output the value assessment of the current state and action.
[0103] At each control decision point t, the Agent determines the current state. and strategy Select Action To balance exploration and exploitation, noise is added to the output when selecting actions.
[0104] Execute action Afterwards, the system transitions to a new state. and received a reward signal. Reward function Designed to quantify the quality of an action, it typically takes the following form: in This is the highest temperature of the key components. It is the target operating temperature. Δa is the fan power consumption, and Δa is the change in control action. These are weighting coefficients used to balance different objectives.
[0105] The agent utilizes the acquired experience tuples To update its policy network parameters and value network parameters Here, an online PPO variant based on the Actor-Critic framework is used, where a policy network generates actions and a value network evaluates the merits of those actions. Policy qualification traces are introduced when updating the policy network.
[0106]
[0107] Among them, the entropy term The exploration is based on adjusting the direction of the TD error. To ensure the stability of online learning, the policy network introduces a pruning mechanism, KL divergence, and ObGD constraints to limit the step size of each policy network update. .in, These are the parameters before the update. It is the cutting factor. It is the KL divergence penalty coefficient. It is the probability ratio of the new strategy to the old strategy.
[0108] The value network is updated by minimizing the square of the TD error:
[0109] These updates are made in real time based on current experience.
[0110] Step S308, action execution.
[0111] Normalized action A output by the Agent t The PWM signal, which is mapped to a signal acceptable to the fan controller, is sent to the fan hardware via the BMC to precisely control its speed.
[0112] Step S310, Value Network Evaluation Layer.
[0113] The system forms a fast real-time control loop between steps S304 and S308.
[0114] Meanwhile, the online learning mechanism in step S306 enables the Agent's policy π θIt evolves and optimizes through continuous interaction with the environment. This allows the system to automatically adapt to dynamic changes caused by varying server workloads, ambient temperature fluctuations, and even fan aging, maintaining efficient and stable heat dissipation performance over the long term.
[0115] Preliminary simulation verification shows that this application can significantly reduce control latency, shortening the control loop response time from hundreds of milliseconds in traditional methods to the millisecond level, and significantly reduce the temperature fluctuation of key components to effectively suppress peak temperatures. Through precise on-demand heat dissipation, it significantly reduces the average power consumption of fans (by 37%), improves the data center PUE (from 1.25 to 1.18), and enables the system to reach stable control from startup within tens of seconds and quickly respond to load changes. The stable thermal environment and optimized fan usage help extend hardware lifespan.
[0116] In the aforementioned embodiments, the Agent receives real-time status. And based on its internal strategy of maintaining and continuously optimizing online. Output optimal fan control action (i.e., the fan's precise PWM duty cycle). Crucially, the agent's learning process is online and incremental: it leverages recent interaction experiences (states) ,action Rewards received and the new state ), combined with KL divergence and The method adjusts its strategy parameters in real time. This online learning capability allows the system to continuously optimize performance without service interruption and quickly adapt to environmental non-stationarity. The agent outputs decision-making actions. The signal is converted into a specific physical control signal (standard PWM signal) and sent to the server's fan controller to precisely adjust the fan's operating speed. The system calculates a scalar reward value based on observations after the action is executed (new sensor readings, power consumption, etc.). The reward function is carefully designed to quantify the effectiveness of the current control strategy, aiming to guide the agent to learn a strategy that balances heat dissipation (maintaining a safe temperature, close to the target temperature) and energy efficiency (minimizing fan power consumption). This reward signal is key to driving the agent's online learning. This is achieved by extracting real-time states. It can accurately capture the time dynamic characteristics of complex thermal flow fields.
[0117] In the aforementioned embodiments, sensor data streams are received in real time; dynamic state features are extracted through streaming processing; the reinforcement learning agent selects fan control actions based on an online-updated strategy; and the strategy is incrementally updated based on a reward signal designed to balance heat dissipation and energy efficiency. Through the design and balancing of a multi-objective reward function, the relationship between heat dissipation safety, performance objectives, energy consumption, and control smoothness is accurately quantified and balanced. A stable and efficient online RL algorithm ensures the convergence and stability of the reinforcement learning process under resource-constrained and non-stationary environments, particularly through the engineering implementation and optimization of the Stream-X paradigm algorithm.
[0118] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0119] Embodiments of this application also provide a control device for a heat dissipation device. Figure 4 This is a schematic diagram of the control device structure of the heat dissipation device according to an embodiment of this application, as shown below. Figure 4 As shown, the device includes:
[0120] The acquisition module 42 is used to acquire the first status parameters of the target device within a predetermined time period;
[0121] The first determining module 44 is used to determine the first dynamic state feature of the first state parameter, wherein the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameter;
[0122] The second determining module 46 is used to input the first dynamic state features into the reinforcement learning network model to determine the target control action. The reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device of the target device before and after the control action is executed. The control action is the action determined by the reinforcement learning network model before the target control action is determined.
[0123] The control module 48 is used to control the heat dissipation equipment to operate according to the target control action.
[0124] In an exemplary embodiment, the first determining module 44 can determine the first dynamic state feature of the first state parameter by: determining the time difference feature between each parameter included in the first state parameter; inputting the time difference feature and the first state parameter into a time loop network to obtain the first dynamic state feature.
[0125] In an exemplary embodiment, the control device of the heat dissipation device can be used to determine a second state parameter of the target device after inputting a first dynamic state feature into a reinforcement learning network model to determine a target control action; determine a reward function based on the first state parameter and the second state parameter; and update the model parameters of the reinforcement learning network model based on the reward function.
[0126] In an exemplary embodiment, the control device of the heat dissipation device can determine the reward function based on the first state parameters and the second state parameters in the following manner: determining the maximum temperature, average temperature and the first rotation speed of the heat dissipation device included in the second state parameters; determining the power consumption and the second rotation speed of the heat dissipation device included in the first state parameters; and determining the reward function based on the maximum temperature, average temperature, first rotation speed, power consumption and the second rotation speed.
[0127] In an exemplary embodiment, the control device of the heat dissipation device can determine a reward function based on a maximum temperature, an average temperature, a first rotational speed, power consumption, and a second rotational speed in the following manner: determining a first difference between the maximum temperature and a predetermined safe temperature; determining a target distance between the average temperature and a predetermined target temperature; determining the absolute value of a second difference between the first rotational speed and the second rotational speed; determining a first negative of the product of the first difference and a first weight; determining a second negative of the product of the target distance and the second weight; determining a third negative of the product of power consumption and a third weight; and determining a fourth negative of the product of the absolute value and a fourth weight; and determining the sum of the first negative, second negative, third negative, and fourth negative as the reward function.
[0128] In an exemplary embodiment, the control device of the heat dissipation device can update the model parameters of the reinforcement learning network model based on the reward function in the following manner: inputting the first dynamic state features and the target control action into the value network included in the reinforcement learning network model to obtain a first value evaluation value of the target control action; updating the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the first value evaluation value, wherein the policy network is used to determine the target control action; and updating the second model parameters of the value network based on the reward function and the first value evaluation value.
[0129] In an exemplary embodiment, the control device of the heat dissipation device can update the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the value evaluation value in the following manner: determining the policy qualification trace based on the reward function; and updating the first model parameters based on the policy qualification trace.
[0130] In an exemplary embodiment, the control device of the heat dissipation device can update the second model parameters of the value network based on the reward function and the first value evaluation value in the following manner: determining the second dynamic state feature of the second state parameter; determining the second value evaluation value corresponding to the second dynamic state feature; and updating the second model parameters based on the reward function, the first value evaluation value, and the second value evaluation value.
[0131] In an exemplary embodiment, the acquisition module 42 can acquire the first state parameter of the target device within a predetermined time period by: synchronously acquiring the second state parameter collected by multiple parameter acquisition devices at a target frequency, wherein the parameter acquisition devices are devices deployed in the target device for collecting the first state parameter; determining a sliding time window, wherein the duration of the sliding time window is the same as the duration of the predetermined time period; and determining the first state parameter from the second state parameter according to the sliding time window.
[0132] In an exemplary embodiment, the control module 48 can control the heat dissipation device to operate according to the target control action by adding pre-set noise to the target control action to obtain the processing control action; and controlling the heat dissipation device to operate according to the processing control action.
[0133] In one exemplary embodiment, the control device of the heat dissipation device can also be used to determine a pre-set sparsity before inputting the first dynamic state feature into the reinforcement learning network model; and initialize the network parameters of the reinforcement learning network model based on the sparsity.
[0134] For a description of the features of the control device of the heat dissipation device in the corresponding embodiment, please refer to the relevant description of the control method of the heat dissipation device in the corresponding embodiment, which will not be repeated here.
[0135] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described embodiments of the control method for a heat dissipation device.
[0136] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described embodiments of the control method for a heat dissipation device when it is run.
[0137] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0138] The embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described control method embodiments for a heat dissipation device.
[0139] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described control method embodiments for a heat dissipation device.
[0140] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0141] The control method, apparatus, storage medium, and electronic device for a heat dissipation device provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A control method for a heat dissipation device, characterized in that, include: Obtain the first status parameters of the target device within a predetermined time period; Determine a first dynamic state feature of the first state parameter, wherein the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameter; The first dynamic state feature is input into the reinforcement learning network model to determine the target control action. The reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device of the target device before and after the control action is executed. The control action is the action determined by the reinforcement learning network model before the target control action is determined. Control the heat dissipation device to operate according to the target control action; After inputting the first dynamic state features into a reinforcement learning network model to determine the target control action, the method further includes: determining a second state parameter of the target device; and determining a reward function based on the first state parameter and the second state parameter. The step of determining the reward function based on the first state parameter and the second state parameter includes: determining the maximum temperature, average temperature, and first rotation speed of the heat dissipation device included in the second state parameter; determining the power consumption and second rotation speed of the heat dissipation device included in the first state parameter; and determining the reward function based on the maximum temperature, the average temperature, the first rotation speed, the power consumption, and the second rotation speed. The step of determining the reward function based on the maximum temperature, the average temperature, the first rotational speed, the power consumption, and the second rotational speed includes: determining a first difference between the maximum temperature and a predetermined safe temperature; determining a target distance between the average temperature and a predetermined target temperature; determining the absolute value of a second difference between the first rotational speed and the second rotational speed; determining a first negative of the product of the first difference and a first weight; determining a second negative of the product of the target distance and a second weight; determining a third negative of the product of the power consumption and a third weight; and determining a fourth negative of the product of the absolute value and a fourth weight; and determining the sum of the first negative, the second negative, the third negative, and the fourth negative as the reward function.
2. The control method for the heat dissipation device according to claim 1, characterized in that, The first dynamic state feature for determining the first state parameter includes: Determine the temporal difference characteristics between each parameter included in the first state parameters; The time difference feature and the first state parameter are input into the time cyclic network to obtain the first dynamic state feature.
3. The control method for the heat dissipation device according to claim 1, characterized in that, After inputting the first dynamic state features into a reinforcement learning network model to determine the target control action, the method further includes: The model parameters of the reinforcement learning network model are updated based on the reward function.
4. The control method for the heat dissipation device according to claim 3, characterized in that, The step of updating the model parameters of the reinforcement learning network model based on the reward function includes: The first dynamic state feature and the target control action are input into the value network included in the reinforcement learning network model to obtain the first value evaluation value of the target control action; The first model parameters of the policy network included in the reinforcement learning network model are updated based on the reward function and the first value evaluation value, wherein the policy network is used to determine the target control action; The second model parameters of the value network are updated based on the reward function and the first value evaluation value.
5. The control method for the heat dissipation device according to claim 4, characterized in that, The step of updating the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the value evaluation value includes: The policy eligibility trace is determined based on the reward function; The parameters of the first model are updated based on the qualification trace of the strategy.
6. The control method for the heat dissipation device according to claim 4, characterized in that, The step of updating the second model parameters of the value network based on the reward function and the first value evaluation value includes: Determine the second dynamic state characteristic of the second state parameter; Determine the second value evaluation value corresponding to the second dynamic state characteristic; The second model parameters are updated based on the reward function, the first value evaluation value, and the second value evaluation value.
7. The control method for the heat dissipation device according to claim 1, characterized in that, The acquisition of the first state parameters of the target device within a predetermined time period includes: The second state parameters are acquired synchronously from multiple parameter acquisition devices at a target frequency, wherein the parameter acquisition devices are devices deployed in the target device and are used to acquire the first state parameters; Determine a sliding time window, wherein the duration of the sliding time window is the same as the duration of the predetermined time period; The first state parameter is determined from the second state parameter according to the sliding time window.
8. The control method for the heat dissipation device according to claim 1, characterized in that, Controlling the heat dissipation device to operate according to the target control action includes: Pre-set noise is added to the target control action to obtain the processing control action; Control the heat dissipation device to operate according to the processing control action.
9. The control method for the heat dissipation device according to claim 1, characterized in that, Before inputting the first dynamic state feature into the reinforcement learning network model, the method further includes: Determine the pre-defined sparsity; The network parameters of the reinforcement learning network model are initialized based on the sparsity.
10. A control device for a heat dissipation equipment, characterized in that, include: The acquisition module is used to acquire the first status parameters of the target device within a predetermined time period; The first determining module is used to determine a first dynamic state feature of the first state parameter, wherein the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameter; The second determining module is used to input the first dynamic state features into the reinforcement learning network model to determine the target control action. The reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device of the target device before and after the control action is executed. The control action is the action determined by the reinforcement learning network model before the target control action is determined. The control module is used to control the heat dissipation device to operate according to the target control action; The device is further configured to, after inputting the first dynamic state features into a reinforcement learning network model to determine the target control action,: determine the second state parameters of the target device; and determine a reward function based on the first state parameters and the second state parameters; The device implements the determination of the reward function based on the first state parameter and the second state parameter in the following manner: determining the maximum temperature, average temperature and the first rotation speed of the heat dissipation device included in the second state parameter; determining the power consumption and the second rotation speed of the heat dissipation device included in the first state parameter; and determining the reward function based on the maximum temperature, the average temperature, the first rotation speed, the power consumption and the second rotation speed. The device determines the reward function based on the maximum temperature, the average temperature, the first rotational speed, the power consumption, and the second rotational speed in the following manner: determining a first difference between the maximum temperature and a predetermined safe temperature; determining a target distance between the average temperature and a predetermined target temperature; determining the absolute value of a second difference between the first rotational speed and the second rotational speed; determining a first negative of the product of the first difference and a first weight; determining a second negative of the product of the target distance and a second weight; determining a third negative of the product of the power consumption and a third weight; and determining a fourth negative of the product of the absolute value and a fourth weight; and determining the sum of the first negative, the second negative, the third negative, and the fourth negative as the reward function.
11. An electronic device, characterized in that, include: memory for storing computer programs; A processor, configured to implement the steps of the control method for the heat dissipation device as described in any one of claims 1 to 9 when executing the computer program.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the control method for the heat dissipation device as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the control method for the heat dissipation device as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Electricity consumption prediction method and device, and storage medium
CN117709521A
Server heat dissipation control method, electronic equipment and storage medium
CN120066922A
Cooling equipment control method and device, equipment and storage medium
CN120353134A