Control method and device of heat dissipation equipment, storage medium and electronic equipment

By obtaining the device status parameters and using the reinforcement learning network model to optimize the control actions of the heat dissipation equipment, the problem of inaccurate heat dissipation caused by changes in the equipment load and environment is solved, and the equipment's precise heat dissipation and energy efficiency balance is achieved.

CN120335582AActive Publication Date: 2025-07-18JINAN INSPUR DATA TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510815951.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-07-18
Estimated Expiration
2045-06-18

AI Technical Summary

Technical Problem

In the prior art, when the equipment load or environment changes, the heat dissipation device is controlled to start and stop and cannot accurately dissipate heat.

Method used

By obtaining the status parameters of the target device, determining its dynamic state characteristics, and using the reinforcement learning network model iterative update, optimizing the control actions of the heat dissipation device, including the application of time difference characteristics and time loop network, combining the update of the reward function and policy network, to achieve accurate heat dissipation.

Benefits of technology

The reinforcement learning network model can adapt to equipment load and environment changes, accurately determine control actions, achieve accurate heat dissipation, reduce average fan power consumption, and improve system stability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120335582A_ABST
    Figure CN120335582A_ABST
Patent Text Reader

Abstract

The invention discloses a control method and device of a heat dissipation device, a storage medium and an electronic device, and relates to the technical field of computers, and a reinforcement learning network model is iteratively updated according to state parameters before and after the heat dissipation device of a target device operates according to a control action. Therefore, the reinforcement learning network model can be evolved and optimized continuously along with continuous interaction with the environment. This enables the reinforcement learning network model to automatically adapt to load changes and environmental changes of the device. Moreover, the model input of the reinforcement learning network model is the first dynamic state feature, and the first dynamic state feature represents the dependency relationship between the parameters of the target equipment, so that the key feature vectors of the current dynamic state and the future trend of the target equipment can be obtained, the reinforcement learning network model can accurately determine the target control action, and the control accuracy of the target equipment is improved. Therefore, the technical problem that heat dissipation cannot be accurately carried out on the equipment in the prior art can be solved, and the technical effect of accurately carrying out heat dissipation on the equipment is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technologies, and in particular, to a control method, device, storage medium, and electronic device for a heat dissipation device. Background Art

[0002] In the related art, generally, when the temperature of a device reaches a fixed threshold, the heat dissipation device is controlled to start and stop to dissipate heat from the device. However, when the load or environment of the device changes continuously, the method of controlling the start and stop of the heat dissipation device using a fixed threshold cannot dissipate heat accurately.

[0003] It can be seen that there is a problem in the related art that the device cannot be cooled accurately.

[0004] For the above problems existing in the related art, no effective solution has been proposed yet. Summary of the Invention

[0005] The present application provides a control method, device, storage medium, and electronic device for a heat dissipation device, so as to at least solve the problem in the related art that the device cannot be cooled accurately.

[0006] The present application provides a control method for a heat dissipation device, including: obtaining first state parameters of a target device within a predetermined time period; determining a first dynamic state feature of the first state parameters, where the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameters; inputting the first dynamic state feature into a reinforcement learning network model to determine a target control action, where the reinforcement learning network model is iteratively updated based on the state parameters before and after the heat dissipation device of the target device operates according to the control action, and the control action is the action determined by the reinforcement learning network model before determining the target control action; controlling the heat dissipation device to operate according to the target control action.

[0007] In an exemplary embodiment, determining the first dynamic state feature of the first state parameters includes: determining a time difference feature between each parameter included in the first state parameters; inputting the time difference feature and the first state parameters into a time recurrent network to obtain the first dynamic state feature.

[0008] In an exemplary embodiment, after inputting the first dynamic state feature into the reinforcement learning network model to determine the target control action, the method further includes: determining second state parameters of the target device; determining a reward function based on the first state parameters and the second state parameters; updating the model parameters of the reinforcement learning network model based on the reward function.

[0009] In an exemplary embodiment, determining a reward function based on a first state parameter and a second state parameter includes: determining a maximum temperature, an average temperature, and a first rotation speed of a heat dissipation device included in the second state parameter; determining a power consumption of the heat dissipation device and a second rotation speed included in the first state parameter; and determining the reward function based on the maximum temperature, the average temperature, the first rotation speed, the power consumption, and the second rotation speed.

[0010] In an exemplary embodiment, determining the reward function based on the maximum temperature, the average temperature, the first rotation speed, the power consumption, and the second rotation speed includes: determining a first difference between the maximum temperature and a pre-determined safety temperature; determining a target distance between the average temperature and a pre-determined target temperature; determining an absolute value of a second difference between the first rotation speed and the second rotation speed; determining a first opposite number of a product of the first difference and a first weight, determining a second opposite number of a product of the target distance and a second weight, determining a third opposite number of a product of the power consumption and a third weight, and determining a fourth opposite number of a product of the absolute value and a fourth weight; and determining a sum value of the first opposite number, the second opposite number, the third opposite number, and the fourth opposite number as the reward function.

[0011] In an exemplary embodiment, updating model parameters of a reinforcement learning network model based on the reward function includes: inputting a first dynamic state feature and a target control action into a value network included in the reinforcement learning network model to obtain a first value evaluation value of the target control action; updating a first model parameter of a policy network included in the reinforcement learning network model based on the reward function and the first value evaluation value, where the policy network is used to determine the target control action; and updating a second model parameter of the value network based on the reward function and the first value evaluation value.

[0012] In an exemplary embodiment, updating the first model parameter of the policy network included in the reinforcement learning network model based on the reward function and the value evaluation value includes: determining a policy eligibility trace based on the reward function; and updating the first model parameter based on the policy eligibility trace.

[0013] In an exemplary embodiment, updating the second model parameter of the value network based on the reward function and the first value evaluation value includes: determining a second dynamic state feature of the second state parameter; determining a second value evaluation value corresponding to the second dynamic state feature; and updating the second model parameter based on the reward function, the first value evaluation value, and the second value evaluation value.

[0014] In an exemplary embodiment, obtaining first state parameters of a target device within a predetermined time period includes: synchronously obtaining second state parameters collected by a plurality of parameter collection devices at a target frequency, where the parameter collection devices are devices deployed in the target device and are used to collect the first state parameters; determining a sliding time window, where the duration of the sliding time window is the same as the duration of the predetermined time period; and determining the first state parameters from the second state parameters according to the sliding time window.

[0015] In an exemplary embodiment, controlling a heat dissipation device to operate according to a target control action includes: adding preset noise to the target control action to obtain a processed control action; and controlling the heat dissipation device to operate according to the processed control action.

[0016] In an exemplary embodiment, before inputting a first dynamic state feature into a reinforcement learning network model, the method further includes: determining a preset sparsity; and initializing network parameters of the reinforcement learning network model based on the sparsity.

[0017] This application also provides a control device for a heat dissipation device, including: an acquisition module, configured to acquire first state parameters of a target device within a predetermined time period; a first determination module, configured to determine a first dynamic state feature of the first state parameters, where the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameters; a second determination module, configured to input the first dynamic state feature into a reinforcement learning network model to determine a target control action, where the reinforcement learning network model is iteratively updated based on state parameters before and after the heat dissipation device of the target device operates according to a control action, and the control action is an action determined by the reinforcement learning network model before determining the target control action; and a control module, configured to control the heat dissipation device to operate according to the target control action.

[0018] This application also provides an electronic device, including: a memory, configured to store a computer program; and a processor, configured to implement the steps of any one of the above heat dissipation device control methods when executing the computer program.

[0019] This application also provides a computer-readable storage medium, in which a computer program is stored, where the computer program implements the steps of any one of the above heat dissipation device control methods when executed by a processor.

[0020] This application also provides a computer program product, including a computer program, where the computer program implements the steps of any one of the above heat dissipation device control methods when executed by a processor.

[0021] With this application, since the reinforcement learning network model is iteratively updated according to the state parameters of the heat dissipation device of the target device before and after the control action is executed, the reinforcement learning network model can continuously evolve and optimize as it continues to interact with the environment. This enables the reinforcement learning network model to automatically adapt to changes in the device load and environmental changes. In addition, the model input of the reinforcement learning network model is the first dynamic state feature, and the first dynamic state feature characterizes the dependency relationship between the parameters of the target device. Therefore, it is possible to obtain the key feature vector of the current dynamic state and future trend of the target device, enabling the reinforcement learning network model to accurately determine the target control action. Through the accurate target control action, it is then possible to control the precise heat dissipation of the heat dissipation device. Therefore, the technical problem of being unable to accurately dissipate heat for the device in the related art can be solved, achieving the technical effect of accurately dissipating heat for the device. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0023] Figure 1 FIG. is a flowchart of a control method for a heat dissipation device according to an embodiment of the present application;

[0024] Figure 2 FIG. is a schematic diagram of a system for a control method of an execution device according to an embodiment of the present application;

[0025] Figure 3 FIG. is a flowchart of a control method for a heat dissipation device according to a specific embodiment of the present application;

[0026] Figure 4 FIG. is a schematic structural diagram of a control device for a heat dissipation device according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0028] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0029] In order to enable those skilled in the art of this technical field to better understand the solution of this application, the following further describes this application in detail with reference to the accompanying drawings and specific embodiments.

[0030] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the control method of the heat dissipation device depends, the specific application environment architecture or specific hardware architecture is described herein.

[0031] The control method of the heat dissipation device can be applied to devices equipped with heat dissipation devices such as servers, computers, laptops, desktop computer cases, etc. Among them, the heat dissipation device can be a fan, heat dissipation fins, etc.

[0032] The embodiments of this application provide a control method for a heat dissipation device. The method is described in detail in combination with the execution process of the control method of the heat dissipation device.

[0033] Figure 1 As shown in the flowchart of the control method for the heat dissipation device according to the embodiments of this application, as Figure 1 shown, the process includes:

[0034] Step S102, obtaining the first state parameter of the target device within a predetermined time period;

[0035] In this embodiment, the target device can be a server, a computer, a laptop, a desktop computer case, etc. The first state parameter may include the temperatures of various components of the target device, such as the temperature of the central processing unit (CPU), the temperature of the memory, the temperature of the power supply, the temperatures of the air inlet and outlet, and may also include the power consumption of various components in the target device, the heat dissipation power consumption of the heat dissipation device, the rotation speed, etc.

[0036] In this embodiment, high-frequency and high-precision sensors can be deployed in various components of the target device. The sensors can include temperature sensors, power consumption sensors, rotation speed sensors, etc. The execution system of the control method of the heat dissipation device can continuously collect data from high-frequency and high-precision sensors (temperature, power consumption, fan rotation speed, etc.) deployed at key positions inside the server (such as CPU, GPU, memory, power supply, air inlet and outlet, etc.) to form a multi-dimensional real-time data stream. The first state parameter is obtained based on the real-time data stream.

[0037] Specifically, when the target device is a server, a high-density sensor network can be deployed inside the server, including temperature sensors (thermocouples or digital temperature sensors with an accuracy requirement of ±0.5°C), power consumption sensors (monitoring the power consumption of components such as CPUs, GPUs (image processors), and fans), and fan speed sensors (providing closed-loop feedback), etc. The data collected by the high-density sensor network is determined as the first state parameter.

[0038] Step S104, determining the first dynamic state feature of the first state parameter, where the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameter;

[0039] In this embodiment, the original sensor data stream, that is, the first state parameter, can be fed into a streaming feature extraction module. The streaming feature extraction module can use time series analysis techniques to extract key feature vectors that can reflect the current dynamic thermal state and future trend of the system in real time from the data stream, that is, the first dynamic state feature. Among them, the streaming feature extraction module can include convolutional networks, time recurrent networks, etc. The time recurrent network can be a long short-term memory network (LSTM).

[0040] Step S106, inputting the first dynamic state feature into the reinforcement learning network model to determine the target control action, where the reinforcement learning network model is iteratively updated based on the state parameters before and after the heat dissipation device of the target device operates according to the control action, and the control action is the action determined by the reinforcement learning network model before determining the target control action;

[0041] In this embodiment, the target control action can be the action executed by the heat dissipation device. When the heat dissipation device is a fan, the target control action can include speed, time, etc. It can also include the speed corresponding to each time point, that is, the fan can rotate at the same speed or at different speeds. For example, it can rotate at a high speed first and then at a low speed. Rotating at a high speed first can quickly reduce the temperature of the target device, and then rotating at a low speed can achieve the purpose of precise temperature control, realizing the effects of balancing the heat dissipation effect (maintaining a safe temperature, approaching the target temperature) and energy efficiency (minimizing the power consumption of the fan).

[0042] In this embodiment, the reinforcement learning network model can be a Streaming Reinforcement Learning network model, represented by the paradigm represented by the Stream-X algorithm, which focuses on processing continuous data streams and performing real-time and efficient online incremental learning. Its core characteristics include: immediate parameter updates based on single samples or small batches of samples, no need to store a large amount of historical data, and the use of mechanisms such as Eligibility Traces to effectively handle delayed rewards and non-linear relationships. These characteristics naturally meet the real-time, dynamic, and resource-constrained requirements of the target device's heat dissipation control. A target device heat dissipation control system that can perceive, make intelligent decisions, and continuously optimize in real time can be constructed according to the streaming reinforcement learning network model to dissipate heat from the target device.

[0043] Step S108, control the heat dissipation device to operate according to the target control action.

[0044] In this embodiment, the target control action can be converted into Pulse Width Modulation (PWM), and the PWM is sent to the heat dissipation device so that the heat dissipation device operates according to the PWM to reduce the temperature of the target device.

[0045] Optionally, the execution subject of the above steps can be the target device, the processor, etc., or a device integrated with data processing capabilities. When the target device is a server and the heat dissipation device is a fan, the system schematic diagram of the control method of the execution device can be seen in the appendix Figure 2 , as Figure 2 shown, the first state parameters, such as temperature and power consumption data, can be obtained by the sensor, the first state parameters are sent to the Reinforcement Learning Agent (RL Agent), the target control action is determined by the RL Agent, and the target control action is sent to the Baseboard Management Controller BMC. The BMC determines the PWM according to the target control action and sends the PWM to the fan to control the rotation of the fan.

[0046] With this application, since the reinforcement learning network model is iteratively updated according to the state parameters of the heat dissipation device of the target device before and after the control action is executed, the reinforcement learning network model can continuously evolve and optimize as it continues to interact with the environment. This enables the reinforcement learning network model to automatically adapt to changes in the device load and the environment. In addition, the model input of the reinforcement learning network model is the first dynamic state feature, and the first dynamic state feature characterizes the dependency relationship between the parameters of the target device. Therefore, it is possible to obtain the key feature vector of the current dynamic state and future trend of the target device, enabling the reinforcement learning network model to accurately determine the target control action. Through the accurate target control action, the precise heat dissipation of the heat dissipation device can be controlled, thus solving the technical problem of being unable to accurately dissipate heat from the device in the related art and achieving the technical effect of accurately dissipating heat from the device.

[0047] In an exemplary embodiment, determining the first dynamic state feature of the first state parameter includes: determining the time difference feature between each parameter included in the first state parameter; inputting the time difference feature and the first state parameter into a temporal recurrent network to obtain the first dynamic state feature. In this embodiment, the first state parameter can be processed to capture dynamic information, that is, the first dynamic state feature. For example, one-dimensional causal convolution can be used to calculate the time difference between data points to obtain the time difference feature, which is used to characterize the change rate and trend of temperature / power consumption. Causal convolution ensures that the calculation only depends on the current and past data, meeting the real-time requirement. Its mathematical form can be expressed as: DiffFeat t =Conv1D causal (Windoes(X t ))), where Conv1D causal is one-dimensional causal convolution, and Windoes(X t ) is the data of the last W time steps, that is, the first state parameter.

[0048] In this embodiment, the temporal recurrent network can be an LSTM. The first state parameter and the extracted time difference feature are input into an online-updated long short-term memory network (LSTM). The recurrent structure of the LSTM enables it to effectively capture the long-term dependencies in the time series data, integrate historical information, and output a hidden state encoding the complex dynamics of the current system. This process can be expressed as: h t =LSTM online (InputFeat t , h (t-1) ), where InputFeat t is [X (t-w+1) …, X t and DiffFeat tCombination. Finally, the pipeline outputs a feature vector S with a fixed dimension t , which is the first dynamic state feature, and this vector is the state representation of the reinforcement learning agent at time step t. Among them, h t is the intermediate representation of S t . During continuous calculation, the last h t is S t .

[0049] In the above embodiment, the temporal difference feature, for example, the rate of change of the server temperature over time, can help the model capture the trend of parameter change over time. Temporal recurrent networks (such as LSTM or GRU) can process sequential data, enabling the model to understand the dependencies between parameters, including but not limited to the mutual influence between temperature and fan speed. Through this technical feature, the model can more accurately predict the control effect of the heat dissipation device, thereby making more optimized decisions to achieve precise control of the server temperature.

[0050] In an exemplary embodiment, after inputting the first dynamic state feature into the reinforcement learning network model to determine the target control action, the method further includes: determining the second state parameter of the target device; determining the reward function based on the first state parameter and the second state parameter; updating the model parameters of the reinforcement learning network model based on the reward function. In this embodiment, the current first state parameter of the target device can be determined as S t , determining the target control action as At. After the heat dissipation device executes At, the state of the target device transitions to the new second state parameter S t+1 . The reward function can be determined according to the first state parameter and the second state parameter, and the model parameters of the reinforcement learning network model can be updated according to the reward function.

[0051] In the above embodiment, the reward function can be used to quantify the quality of the target control action, that is, the heat dissipation effect and the quantified power consumption cost. By quantifying the heat dissipation effect and the power consumption cost, it guides the model to learn better control strategies. For example, if the second state parameter shows that the server temperature decreases and the power consumption decreases, the reward function value is higher, otherwise it is lower. By updating the model parameters, the model can continuously optimize its decision-making process, including but not limited to adjusting the control strategy of the heat dissipation device to achieve better heat dissipation effect and lower power consumption.

[0052] In an exemplary embodiment, determining a reward function based on a first state parameter and a second state parameter includes: determining a maximum temperature, an average temperature included in the second state parameter, and a first rotation speed of a heat dissipation device; determining a power consumption of the heat dissipation device and a second rotation speed included in the first state parameter; and determining the reward function based on the maximum temperature, the average temperature, the first rotation speed, the power consumption, and the second rotation speed. In this embodiment, the second state parameter may include the temperatures of various components, and the maximum temperature and the average temperature thereof can be determined. The second state parameter may further include the first rotation speed of the heat dissipation device. The power consumption of the heat dissipation device and the second rotation speed included in the first state parameter can also be determined, and the reward function is determined according to the maximum temperature, the average temperature, the first rotation speed, the power consumption, and the second rotation speed.

[0053] In this embodiment, the maximum temperature can be compared with a preset safety threshold. If the maximum temperature is lower than the safety threshold, the reward value is increased; the deviation between the average temperature and the target operating temperature is calculated, and the reward value is adjusted according to the magnitude of the deviation. The smaller the deviation, the higher the reward value; according to the difference between the first rotation speed and the second rotation speed, an action smoothness penalty is introduced. If the change amplitude between the first rotation speed and the second rotation speed is large, the reward value is reduced to promote the smoothness of the control action; according to the power consumption of the heat dissipation device, an energy consumption cost penalty term is introduced. The higher the power consumption, the greater the penalty, thereby reducing the reward value; considering the maximum temperature, the average temperature, the difference between the first rotation speed and the second rotation speed, and the power consumption comprehensively, the final value of the reward function is determined by weighted summation, where the weights of each parameter are flexibly adjusted according to the heat dissipation control priority and the energy efficiency requirement to balance the relationship between the heat dissipation effect and the energy consumption.

[0054] Among them, the safety threshold comparison emphasizes the importance of incorporating the consideration of the safety temperature into the reward function, ensuring that the decisions of the agent do not allow the temperature to exceed the safe range. The temperature deviation adjustment focuses on the agent maintaining the temperature within the target operating range to improve the overall efficiency and stability of the system. The smoothness penalty focuses on the change rate of the control action, aiming to avoid unnecessary frequent adjustments and violent fluctuations, which is beneficial to extending the hardware life and reducing the noise. The energy consumption cost penalty introduces a direct consideration of the power consumption, aiming to encourage the agent to adopt more energy-saving strategies, in line with the development trend of green computing. The weighted summation proposes a comprehensive evaluation mechanism. By adjusting the parameter weights, the agent can flexibly meet the heat dissipation and energy consumption balance requirements in different scenarios and achieve more comprehensive optimization.

[0055] In this embodiment, the determination of the reward function takes into account the balance between temperature control and power consumption cost. For example, the reduction of the maximum temperature and the average temperature will increase the value of the reward function, while the increase of the power consumption and the drastic change of the rotation speed will reduce the value of the reward function. This design enables the model to automatically balance the heat dissipation effect and the power consumption cost during the learning process, achieving more efficient and stable server temperature control.

[0056] In an exemplary embodiment, a reward function is determined based on a maximum temperature, an average temperature, a first rotational speed, a power consumption, and a second rotational speed, including: determining a first difference between the maximum temperature and a pre-determined safe temperature; determining a target distance between the average temperature and a pre-determined target temperature; determining an absolute value of a second difference between the first rotational speed and the second rotational speed; determining a first negative opposite of a product of the first difference and a first weight, determining a second negative opposite of a product of the target distance and a second weight, determining a third negative opposite of a product of the power consumption and a third weight, and determining a fourth negative opposite of a product of the absolute value and a fourth weight; and determining a sum value of the first negative opposite, the second negative opposite, the third negative opposite, and the fourth negative opposite as the reward function. In this embodiment, after performing action A t , the system transitions to a new state S (t+1) and receives a reward signal R t . The reward function R(S t , A t , S (t+1) ) is designed to quantify the quality of an action, and its form is usually . Among them, represents the first weight, represents the second weight, represents the third weight, represents the fourth weight. The maximum temperature is , the safe temperature is , represents the average temperature, represents the target temperature. represents the power consumption, and |△a t | represents the absolute value of the second difference between the first rotational speed and the second rotational speed, that is, the control action change amount. -[[-END]] are weight coefficients used to balance different objectives.

[0057] In this embodiment, by introducing weights, the model can flexibly adjust the degree of emphasis on different parameters. For example, if maintaining the safe temperature is more important, the first weight can be increased, so that the model pays more attention to temperature control during decision-making and reduces the emphasis on power consumption. This design enables the model to adapt to different application scenarios, including but not limited to the thermal management of data centers and the temperature control of high-performance computing servers, and realizes a wider range of applications.

[0058] In an exemplary embodiment, updating the model parameters of the reinforcement learning network model based on a reward function includes: inputting the first dynamic state feature and the target control action into the value network included in the reinforcement learning network model to obtain a first value evaluation of the target control action; updating the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the first value evaluation, where the policy network is used to determine the target control action; updating the second model parameters of the value network based on the reward function and the first value evaluation. In this embodiment, the reinforcement learning network model may include a policy network and a value network. The policy network takes the state S t as input and outputs an action A t , and the action A t can be a one-dimensional vector normalized to the interval [0, 1], representing the PWM duty cycle of the corresponding fan. The value network takes the state S t and the action A t as input and outputs the value evaluation of the current state and action. At each control decision point t, the Agent selects the action A t according to the current state S θ and the policy π t . To balance exploration and exploitation, noise can be added to the output result when selecting an action.

[0059] In the above embodiment, the obtained experience tuple (S t , A t , R t , S (t+1) ) can be used to update its policy network parameters θ, that is, the first model parameters, and the value network parameters φ, that is, the second model parameters.

[0060] In the above embodiment, the update of the value evaluation model and the policy network is the core of the reinforcement learning model learning. Under the guidance of the reward function, the model can continuously optimize its policy network to determine better control actions, and at the same time update the value evaluation model, so that the model can more accurately evaluate the values of different control actions. This mechanism ensures that the model can continuously learn and optimize, including but not limited to automatically adjusting the heat dissipation strategy when the server load changes, achieving more efficient and intelligent temperature control.

[0061] In an exemplary embodiment, updating the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the value evaluation includes: determining a policy eligibility trace based on the reward function; updating the first model parameters based on the policy eligibility trace. In this embodiment, the policy network is used to generate actions, and the value network is used to evaluate the pros and cons of actions. When updating the policy network, the policy eligibility trace z θ , can be introduced. The meanings of the parameters are as follows:

[0062] Represents the discount factor: The value range is between 0 and 1, and it is used to weigh the importance of future rewards. The closer it is to 1, the more the model values long-term rewards.

[0063] Represents the trace decay parameter: The value range is between 0 and 1, and it controls the weight of historical information. Similar to the decay of eligibility traces in reinforcement learning, The smaller it is, the faster the historical influence decays.

[0064] Represents the policy gradient: It represents the gradient of the log probability of the policy with respect to the parameter θ, guiding how to adjust θ to increase the probability of action A in state S.

[0065] Represents the entropy term hyperparameter: It controls the weight of the entropy term and affects the intensity of exploration.

[0066] sign( ) represents the sign of the TD error: where = is the TD error, which represents the difference between the actual return of the action and the estimated value. sign( ) takes values of +1 (positive) or -1 (negative), reflecting the goodness or badness of the action.

[0067] is the policy entropy: It is used to measure the randomness of the policy. The larger the entropy, the more random the policy; the smaller the entropy, the more deterministic the policy.

[0068] Entropy term adjusts exploration according to the direction of the TD error: If > 0 (the actual return exceeds the expectation), sign( ) = +1, and the entropy term encourages more exploration and maintains the randomness of the policy.

[0069] If < 0 (the actual return is lower than the expectation), sign( ) = -1, and the entropy term reduces exploration and prompts the policy to be more inclined to known good actions.

[0070] In this embodiment, to ensure the stability of online learning, the policy network introduces a clipping mechanism, KL divergence, and ObGD constraints to limit the step size of each update of the policy network: . The parameter meanings are as follows:

[0071] is the probability ratio: The actions A under the new policy and the old policy tThe ratio of probabilities, reflecting the magnitude of policy update.

[0072] The advantage function: Estimates the relative goodness or badness of action A in state S t under which, usually through calculation.

[0073] is the clipping function: Constrains within the range of | |, where is the clipping coefficient (usually from 0.1 to 0.2), preventing the policy update from being too large.

[0074] is the KL divergence penalty coefficient: Controls the weight of the KL divergence, restricting the difference between the new and old policies.

[0075] is the KL divergence: Measures the difference between the new policy and the old policy The larger the value, the greater the policy change.

[0076] is the ObGD constraint: A variant of gradient descent based on policy eligibility traces of a certain kind.

[0077] Through the min function and clip operation, it ensures that the policy update does not deviate too far from the old policy. Avoids the instability of traditional policy gradient methods. KL divergence: As an additional regularization term, it restricts the difference between the new and old policies, further enhancing stability.

[0078] In the above embodiments, the policy eligibility trace is the trajectory of the policy network during the decision-making process. Guided by the reward function, the model can identify which decision paths are effective and which are ineffective, thereby optimizing the parameters of the policy network. This mechanism enables the model to continuously learn and improve its decision-making strategy, including but not limited to automatically adjusting the control parameters of the heat dissipation device under different environmental conditions to achieve more accurate temperature control.

[0079] In an exemplary embodiment, updating the second model parameters of the value network based on the reward function and the first value evaluation value includes: determining the second dynamic state feature of the second state parameter; determining the second value evaluation value corresponding to the second dynamic state feature; updating the second model parameters based on the reward function, the first value evaluation value, and the second value evaluation value. In this embodiment, the second value evaluation value can be expressed as , and the first value evaluation value can be expressed as . The value network can be updated by minimizing the square of the TD error. The loss function can be expressed as 。These updates are performed in real time based on current experience. Among them, is the target value, which represents the current reward plus the value of the next state after discounting. reflects the prediction of the value network for the current state. represents the difference between the target value and the estimated value.

[0080] In the above embodiments, the update of the value evaluation model is based on current and future state features, as well as the reward function, ensuring that the model can accurately evaluate the impact of the current control action on future states. This mechanism enables the model to predict the heat dissipation effect more intelligently, including but not limited to adjusting the heat dissipation strategy in advance when the server load changes, achieving more efficient and stable temperature control.

[0081] In an exemplary embodiment, obtaining the first state parameters of the target device within a predetermined time period includes: synchronously obtaining the second state parameters collected by multiple parameter collection devices at a target frequency, where the parameter collection devices are devices deployed in the target device and are used to collect the first state parameters; determining a sliding time window, where the duration of the sliding time window is the same as the duration of the predetermined time period; and determining the first state parameters from the second state parameters according to the sliding time window. In this embodiment, a real-time feature extraction pipeline can be established. This pipeline receives a continuous data stream X t , that is, the second state parameters. Applying the sliding window mechanism, the data of the most recent W time steps is intercepted , that is, the first state parameters are obtained. The data acquisition unit can be configured to use the PTP protocol with low-latency synchronization to ensure high-frequency synchronous acquisition of all sensor data, and the sampling period can reach the millisecond level, forming a vectorized time series data stream where t represents discrete time steps. represents the temperature, power consumption of the CPU and GPU at time t, the temperature of the memory, and the speed feedback of the fan. The specific value of N varies according to different server configurations.

[0082] In the above embodiments, the selection of the target frequency, for example, collecting once every 10 seconds, can ensure that the model obtains real-time state parameters, thereby making more timely control decisions. This design enables the model to adapt to different server environments and achieve a wider range of applications. The sliding time window, for example, the temperature data within the most recent 5 minutes, can help the model understand the short-term trend of the current state parameters. This design enables the model to capture the parameter change trend more accurately, including but not limited to quickly responding when the server load suddenly increases, achieving more precise temperature control.

[0083] In an exemplary embodiment, controlling a heat dissipation device to operate according to a target control action includes: adding a preset noise to the target control action to obtain a processed control action; and controlling the heat dissipation device to operate according to the processed control action. In this embodiment, the noise can be a preset noise, such as Gaussian white noise. Adding noise can increase the possibility of the agent exploring the environment. In the early stage of reinforcement learning, the agent needs to try different actions to understand the reaction of the environment, so as to learn the optimal strategy. By adding noise to the target control action, the agent can deviate from the original strategy to a certain extent and try action paths that have not been explored or less explored, which helps to discover potentially better behavioral strategies.

[0084] During the training process, adding noise helps prevent the agent's strategy from overfitting to specific environmental states. Overfitting refers to the phenomenon that the model performs well on the training data but has poor generalization ability on unseen data. By adding noise to the control action, the agent will encounter more variations during learning, which can prompt the policy network to learn more general features rather than just memorizing the patterns in the training data.

[0085] When facing a rapidly changing environment, the action with added noise can improve the real-time response ability of the control system. The agent can not only make decisions based on the current environmental state, but also through the introduction of noise, the agent's strategy can anticipate and adapt to possible future state changes to a certain extent.

[0086] In the server heat dissipation control scenario, the above beneficial effects mean that the agent can adjust the heat dissipation device (such as a fan) more flexibly and intelligently, maximize energy efficiency while meeting the heat dissipation requirements, reduce hardware losses, and can quickly adapt when conditions such as server load and environmental temperature change, maintaining stable and efficient heat dissipation performance.

[0087] In an exemplary embodiment, before inputting the first dynamic state feature into the reinforcement learning network model, the method further includes: determining a preset sparsity; and initializing the network parameters of the reinforcement learning network model based on the sparsity. In this embodiment, a reinforcement learning Agent can be initialized, and the Agent includes a policy network π θ (A|S) and a value network V ∅ (S). Both networks use layer normalization to enhance stability. During initialization, the SpareseInit technique can be used to initialize the weights of the policy network and the value network. When initializing, the weights w are initialized according to the sparsity, and all biases are set to 0.

[0088] In the above embodiments, the setting of the sparsity, for example, the proportion of zero values in the network parameters, can reduce the number of model parameters, improve the training efficiency and generalization ability of the model. This design enables the model to achieve more efficient learning with limited computing resources, including but not limited to deploying a reinforcement learning model on an edge server to achieve more intelligent temperature control.

[0089] The control method of the heat dissipation device will be described below in conjunction with specific embodiments:

[0090] Figure 3 It is a flowchart of the control method of the heat dissipation device according to a specific embodiment of the present application, as Figure 3 shown, and the process includes:

[0091] Step S302, high-density sensor network. That is, system initialization and high-frequency data stream construction.

[0092] Deploy a high-density sensor network inside the server, including temperature sensors (thermocouples or digital temperature sensors with an accuracy requirement of ±0.5°C), power consumption sensors (monitoring the power consumption of components such as CPUs, GPUs, and fans), and fan speed sensors (providing closed-loop feedback).

[0093] Configure the data acquisition unit, and use the low-latency synchronous PTP protocol to ensure that all sensor data is acquired synchronously at a high frequency. The sampling period can reach the millisecond level, forming a vectorized time-series data stream where t represents discrete time steps. represents the temperature and power consumption of the CPU and GPU, the temperature of the memory, and the fan speed feedback at time t. The specific value of N varies according to different server configurations.

[0094] Step S304, streaming feature extraction engineering layer. That is, streaming feature extraction and state construction.

[0095] Establish a real-time feature extraction pipeline. This pipeline receives continuous data streams (corresponding to the above second state parameter).

[0096] First, apply the sliding window mechanism to intercept the data of the last W time steps (corresponding to the above first state parameter).

[0097] Secondly, process the data within the window to capture dynamic information. Use one-dimensional causal convolution to calculate the time difference between data points (corresponding to the above time difference feature) to characterize the change rate and trend of temperature / power consumption. Causal convolution ensures that the calculation only depends on the current and past data, meeting the real-time requirements. Its mathematical form can be expressed as: , where is a one-dimensional causal convolution, and is the data for the most recent W time steps.

[0098] Subsequently, the original data and the extracted differential features are input into an online-updated long short-term memory network (LSTM). The recurrent structure of the LSTM enables it to effectively capture long-term dependencies in time series data, integrate historical information, and output a hidden state encoding the complex dynamics of the current system. This process can be expressed as: , where is combined with .

[0099] Finally, the pipeline outputs a feature vector of a fixed dimension (corresponding to the above first dynamic state feature), and this vector is the state representation of the reinforcement learning Agent at time step t.

[0100] Step S306, policy network layer.

[0101] Initialize a reinforcement learning Agent that includes a policy network and a value network . Both networks use layer normalization to enhance stability. At initialization, the weights of the policy network and the value network are initialized using the SpareseInit technique, and the weights are initialized according to the sparsity , and all biases are set to 0.

[0102] The policy network takes the state as input and outputs an action , and the action is a one-dimensional vector normalized to the interval [0,1], representing the PWM duty cycle of the corresponding fan. The value network takes the state and the action as input and outputs the value evaluation of the current state and action.

[0103] At each control decision point t, the Agent selects an action according to the current state and the policy . To balance exploration and exploitation, noise is added to the output result when selecting an action.

[0104] After executing the action , the system transitions to a new state , and receives a reward signal . The reward function is designed to quantify the goodness or badness of an action, and its form is usually: where is the highest temperature of the key component is the target operating temperature is the fan power consumption, and Δa is the change in control action is the weight coefficient, which is used to balance different objectives

[0105] The Agent uses the obtained experience tuple to update its policy network parameters and value network parameters . Here, an online PPO variant based on the Actor-Critic framework is used. The policy network is used to generate actions, and the value network is used to evaluate the quality of actions. When updating the policy network, policy eligibility traces are introduced

[0106]

[0107] Among them, the entropy term adjusts exploration according to the direction of the TD error. To ensure the stability of online learning, the policy network introduces a clipping mechanism, KL divergence, and ObGD constraints to limit the step size of each policy network update: . Among them, are the parameters before update is the clipping coefficient is the KL divergence penalty coefficient is the probability ratio of the new and old policies

[0108] The value network is updated by minimizing the square of the TD error:

[0109] . These updates are performed in real time based on the current experience

[0110] Step S308, action execution

[0111] The normalized action A output by the Agent t is mapped to the PWM signal acceptable to the fan controller and sent to the fan hardware through BMC to precisely control its rotation speed

[0112] Step S310, value network evaluation layer

[0113] The system forms a fast real-time control loop between step S304 and step S308

[0114] At the same time, the online learning mechanism in step S306 enables the policy π of the Agent θIt continuously evolves and optimizes with the continuous interaction with the environment. This enables the system to automatically adapt to the dynamic changes of the system caused by changing server workloads, environmental temperature fluctuations, and even fan aging, and maintain efficient and stable heat dissipation performance in the long term.

[0115] Through preliminary simulation verification, this application can significantly reduce control latency. The control loop response time is shortened from hundreds of milliseconds in traditional methods to the millisecond level, and the temperature fluctuation amplitude of key components is significantly reduced, effectively suppressing the peak temperature. Through precise on-demand heat dissipation, the average power consumption of the fan is significantly reduced (by 37%), the PUE of the data center is improved (from 1.25 to 1.18), the system can reach stable control within dozens of seconds from startup, and quickly respond to load mutations. The stable thermal environment and optimized fan usage contribute to extending the hardware life.

[0116] In the foregoing embodiment, the Agent receives the real-time state , and based on the strategy maintained and continuously optimized online within it , outputs the optimal fan control action (i.e., the precise PWM duty cycle of the fan). Importantly, the learning process of this Agent is online and incremental: it utilizes the just-past interaction experience (state , action , the obtained reward , and the new state ), combines the KL divergence and method, and adjusts its policy parameters in real time. This online learning ability enables the system to continuously optimize performance without interrupting services and can quickly adapt to the non-stationarity of the environment. The decision action output by the Agent is converted into a specific physical control signal (standard PWM signal) and sent to the fan controller of the server, thereby precisely adjusting the running speed of the fan. The system calculates a scalar reward value based on the observed results (new sensor readings, power consumption, etc.) after executing the action. This reward function is carefully designed to quantify the quality of the current control strategy, and its goal is to guide the Agent to learn a strategy that can balance the heat dissipation effect (maintaining a safe temperature, approaching the target temperature) and energy efficiency (minimizing the fan power consumption). This reward signal is the key to driving the Agent's online learning. By extracting the real-time state , the time dynamic characteristics of the complex heat flow field can be accurately captured.

[0117] In the foregoing embodiments, a sensor data stream is received in real time; dynamic state features are extracted through stream processing; a reinforcement learning agent selects a fan control action according to an online updated policy; and the policy is incrementally updated based on a reward signal aimed at balancing heat dissipation and energy efficiency. Through the design and balance of a multi-objective reward function, the relationship among heat dissipation safety, performance objectives, energy consumption, and control smoothness is accurately quantified and balanced. Through a stable and efficient online RL algorithm, the reinforcement learning process in a resource-constrained and non-stationary environment is ensured to converge and be stable, especially the engineering implementation and tuning of the Stream-X paradigm algorithm.

[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.

[0119] The embodiments of the present application further provide a control device for a heat dissipation device. Figure 4 As a schematic structural diagram of a control device for a heat dissipation device according to an embodiment of the present application, as Figure 4 shown, the device includes:

[0120] An acquisition module 42, configured to acquire first state parameters of a target device within a predetermined time period;

[0121] A first determination module 44, configured to determine a first dynamic state feature of the first state parameters, where the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameters;

[0122] A second determination module 46, configured to input the first dynamic state feature into a reinforcement learning network model to determine a target control action, where the reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device of the target device before and after running according to the control action, and the control action is an action determined by the reinforcement learning network model before determining the target control action;

[0123] A control module 48, configured to control the heat dissipation device to run according to the target control action.

[0124] In an exemplary embodiment, the first determination module 44 may implement the determination of the first dynamic state feature of the first state parameters in the following manner: determining a time difference feature between each parameter included in the first state parameters; and inputting the time difference feature and the first state parameters into a time recurrent network to obtain the first dynamic state feature.

[0125] In an exemplary embodiment, the control device of the heat dissipation device can be used to determine the second state parameter of the target device after inputting the first dynamic state feature into the reinforcement learning network model to determine the target control action; determine the reward function based on the first state parameter and the second state parameter; and update the model parameters of the reinforcement learning network model based on the reward function.

[0126] In an exemplary embodiment, the control device of the heat dissipation device can determine the reward function based on the first state parameter and the second state parameter in the following manner: determine the maximum temperature, average temperature included in the second state parameter, and the first rotation speed of the heat dissipation device; determine the power consumption of the heat dissipation device and the second rotation speed included in the first state parameter; and determine the reward function based on the maximum temperature, average temperature, first rotation speed, power consumption, and second rotation speed.

[0127] In an exemplary embodiment, the control device of the heat dissipation device can determine the reward function based on the maximum temperature, average temperature, first rotation speed, power consumption, and second rotation speed in the following manner: determine the first difference between the maximum temperature and the pre-determined safe temperature; determine the target distance between the average temperature and the pre-determined target temperature; determine the absolute value of the second difference between the first rotation speed and the second rotation speed; determine the first negative opposite of the product of the first difference and the first weight, determine the second negative opposite of the product of the target distance and the second weight, determine the third negative opposite of the product of the power consumption and the third weight, and determine the fourth negative opposite of the product of the absolute value and the fourth weight; and determine the sum value of the first negative opposite, the second negative opposite, the third negative opposite, and the fourth negative opposite as the reward function.

[0128] In an exemplary embodiment, the control device of the heat dissipation device can update the model parameters of the reinforcement learning network model based on the reward function in the following manner: input the first dynamic state feature and the target control action into the value network included in the reinforcement learning network model to obtain the first value evaluation value of the target control action; update the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the first value evaluation value, where the policy network is used to determine the target control action; and update the second model parameters of the value network based on the reward function and the first value evaluation value.

[0129] In an exemplary embodiment, the control device of the heat dissipation device can update the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the value evaluation value in the following manner: determine the policy eligibility trace based on the reward function; and update the first model parameters based on the policy eligibility trace.

[0130] In an exemplary embodiment, the control device of the heat dissipation device may update the second model parameters of the value network based on a reward function and a first value evaluation value in the following manner: determining a second dynamic state feature of the second state parameter; determining a second value evaluation value corresponding to the second dynamic state feature; and updating the second model parameters based on the reward function, the first value evaluation value, and the second value evaluation value.

[0131] In an exemplary embodiment, the acquisition module 42 may acquire the first state parameter of the target device within a predetermined time period in the following manner: synchronously acquiring the second state parameters collected by a plurality of parameter acquisition devices at a target frequency, where the parameter acquisition devices are devices deployed in the target device and are used to acquire the first state parameter; determining a sliding time window, where the duration of the sliding time window is the same as the duration of the predetermined time period; and determining the first state parameter from the second state parameters according to the sliding time window.

[0132] In an exemplary embodiment, the control module 48 may control the heat dissipation device to operate according to a target control action in the following manner: adding a preset noise to the target control action to obtain a processed control action; and controlling the heat dissipation device to operate according to the processed control action.

[0133] In an exemplary embodiment, before inputting the first dynamic state feature into the reinforcement learning network model, the control device of the heat dissipation device may also be used to determine a preset sparsity; and initialize the network parameters of the reinforcement learning network model based on the sparsity.

[0134] For the description of the features in the corresponding embodiments of the control device of the heat dissipation device, reference may be made to the relevant descriptions in the corresponding embodiments of the control method of the heat dissipation device, which will not be elaborated here one by one.

[0135] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the control method of the heat dissipation device.

[0136] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the control method of the heat dissipation device when running.

[0137] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk, or optical disc, and other various media that can store computer programs.

[0138] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the control method for the heat dissipation device.

[0139] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the control method for the heat dissipation device.

[0140] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0141] The above has introduced in detail a control method, device, storage medium, and electronic device for a heat dissipation device provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A control method for a heat dissipation device, characterized in that, Including: Obtain the first state parameters of the target device within a predetermined time period; Determine the first dynamic state feature of the first state parameters, where the first dynamic state feature is used to represent the dependency relationship between each parameter included in the first state parameters; Input the first dynamic state feature into a reinforcement learning network model to determine a target control action, where the reinforcement learning network model is iteratively updated based on the state parameters of the heat dissipation device of the target device before and after running according to the control action, and the control action is the action determined by the reinforcement learning network model before determining the target control action; Control the heat dissipation device to run according to the target control action.

2. The control method of the heat dissipation device according to claim 1, characterized in that The determining the first dynamic state feature of the first state parameters includes: Determine the time difference feature between each parameter included in the first state parameters; Input the time difference feature and the first state parameters into a time recurrent network to obtain the first dynamic state feature.

3. The control method of the heat dissipation device according to claim 1, characterized in that After inputting the first dynamic state feature into the reinforcement learning network model to determine the target control action, the method further includes: Determine the second state parameters of the target device; Determine a reward function based on the first state parameters and the second state parameters; Update the model parameters of the reinforcement learning network model based on the reward function.

4. The control method of the heat dissipation device according to claim 3, characterized in that The determining the reward function based on the first state parameters and the second state parameters includes: Determine the maximum temperature, the average temperature included in the second state parameters, and the first rotation speed of the heat dissipation device; Determine the power consumption of the heat dissipation device and the second rotation speed included in the first state parameters; Determine the reward function based on the maximum temperature, the average temperature, the first rotation speed, the power consumption, and the second rotation speed.

5. The control method of the heat dissipation device according to claim 4, characterized in that The determining the reward function based on the maximum temperature, the average temperature, the first rotation speed, the power consumption, and the second rotation speed includes: Determine the first difference between the maximum temperature and a pre-determined safe temperature; Determine the target distance between the average temperature and a pre-determined target temperature; Determine the absolute value of the second difference between the first rotation speed and the second rotation speed; Determine the first opposite number of the product of the first difference and the first weight, determine the second opposite number of the product of the target distance and the second weight, determine the third opposite number of the product of the power consumption and the third weight, and determine the fourth opposite number of the product of the absolute value and the fourth weight; Determine the sum value of the first opposite number, the second opposite number, the third opposite number, and the fourth opposite number as the reward function.

6. The control method of the heat dissipation device according to claim 3, characterized in that Updating the model parameters of the reinforcement learning network model based on the reward function includes: Inputting the first dynamic state feature and the target control action into a value network included in the reinforcement learning network model to obtain a first value evaluation of the target control action; Updating first model parameters of a policy network included in the reinforcement learning network model based on the reward function and the first value evaluation, wherein the policy network is used to determine the target control action; Updating second model parameters of the value network based on the reward function and the first value evaluation.

7. The control method of a heat dissipation device according to claim 6, characterized in that The updating the first model parameters of the policy network included in the reinforcement learning network model based on the reward function and the value evaluation includes: Determining a policy eligibility trace based on the reward function; Updating the first model parameters based on the policy eligibility trace.

8. The control method of a heat dissipation device according to claim 6, characterized in that The updating the second model parameters of the value network based on the reward function and the first value evaluation includes: Determining a second dynamic state feature of the second state parameter; Determining a second value evaluation corresponding to the second dynamic state feature; Updating the second model parameters based on the reward function, the first value evaluation, and the second value evaluation.

9. The control method of a heat dissipation device according to claim 1, characterized in that The obtaining the first state parameter of the target device within a predetermined time period includes: Synchronously obtaining second state parameters collected by a plurality of parameter collection devices at a target frequency, wherein the parameter collection devices are devices deployed in the target device and are used to collect the first state parameter; Determining a sliding time window, wherein the duration of the sliding time window is the same as the duration of the predetermined time period; Determining the first state parameter from the second state parameters according to the sliding time window.

10. The control method of a heat dissipation device according to claim 1, characterized in that The controlling the heat dissipation device to operate according to the target control action includes: Adding preset noise to the target control action to obtain a processed control action; Controlling the heat dissipation device to operate according to the processed control action.

11. The control method of a heat dissipation device according to claim 1, characterized in that Before inputting the first dynamic state feature into the reinforcement learning network model, the method further includes: Determining a preset sparsity; Initializing network parameters of the reinforcement learning network model based on the sparsity.

12. A control device for a heat dissipation device, characterized in that, including: An obtaining module, configured to obtain a first state parameter of a target device within a predetermined time period; A first determining module, configured to determine a first dynamic state feature of the first state parameter, wherein the first dynamic state feature is used to represent a dependency relationship between each parameter included in the first state parameter; A second determination module, configured to input the first dynamic state feature into a reinforcement learning network model to determine a target control action, where the reinforcement learning network model is iteratively updated based on state parameters before and after the heat dissipation device of the target device operates according to the control action, and the control action is an action determined by the reinforcement learning network model before determining the target control action; A control module, configured to control the heat dissipation device to operate according to the target control action.

13. An electronic device, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to implement the steps of the control method of the heat dissipation device according to any one of claims 1 to 11 when executing the computer program.

14. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the control method of the heat dissipation device according to any one of claims 1 to 11 when executed by a processor.

15. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the control method of the heat dissipation device according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Data center energy consumption optimization control method based on deep reinforcement learning

    CN114511208A

  • Electricity consumption prediction method and device, and storage medium

    CN117709521A

  • Cooling control method and device of data center based on safe deep reinforcement learning

    CN119172985A

  • Inner fan control method, device and equipment, medium and computer program product

    CN119288902A

  • Server heat dissipation control method, electronic equipment and storage medium

    CN120066922A