Data center IT load and cooling system cooperative control method based on TD3 algorithm
Through a deep reinforcement learning model based on the TD3 algorithm, the control lag and server load imbalance problems of the data center cooling system were solved, and coordinated optimization of the server and cooling system was achieved, reducing energy consumption and improving system efficiency.
Patent Information
- Application Number
- CN202511149136.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies for optimizing data center cooling systems suffer from control lags, unbalanced server loads, and local hot spots or cooling waste caused by large temperature differences, and fail to effectively coordinate the management of IT load and cooling system energy consumption.
A deep reinforcement learning model based on the TD3 algorithm is adopted. By obtaining data sets of server status, environment and task characteristics, a policy network Actor and a value network Critic are constructed, the state space, action space and reward function are defined, and model training is performed to output work task allocation and cooling system control strategies to form a dynamic closed-loop optimization.
It achieves the coordinated optimization of server IT load and cooling system, reduces energy consumption, minimizes temperature differences, avoids local hot spots, ensures stable operation of servers and improves the energy efficiency of the cooling system.
Smart Images

Figure CN120743544A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep reinforcement learning, and specifically to a collaborative control method for IT load and cooling system of a data center based on the TD3 algorithm. Background Art
[0002] A data center is a physical facility or complex of buildings used for the centralized management, storage, processing, and transmission of large amounts of data. It typically consists of servers, network equipment, storage systems, power supplies, cooling systems, and more. Data centers house multiple servers, and tasks are assigned to different servers for processing. The IT load on each server reflects the performance pressure on that server. Large variations in IT load can also lead to significant temperature differences between servers. For data centers that use rooms or racks as cooling units, the cooling system must increase its cooling capacity to ensure the safe operation of all servers. This can easily lead to overcooling in some areas, affecting server efficiency, and waste energy.
[0003] In the related art, there are two main types of optimization methods for data center cooling systems. The first type of method is to train the cooling system's historical operating data to find the operating parameter combination that minimizes energy consumption. For example, the total energy consumption value under different operating parameter combinations is calculated according to a preset optimization algorithm. Ultimately, the operating parameter combination with the lowest total energy consumption and not in an unstable operating state is used as the optimal operating parameter combination to guide the regulation of the cooling system to achieve energy saving effects. This type of method can maximize the energy saving effect of the cooling system while ensuring the safety of the data center system. The second type of method can regulate the cooling system based on real-time load, adjusting the cooling supply according to the actual load of the computer room to meet the immediate cooling needs of the data center, realizing a one-way optimization process of "load-cooling capacity".
[0004] However, the first type of method mentioned above is limited by historical operating data and cannot dynamically adapt to load changes. It also has obvious control lag problems. In addition, the first type of method has a single optimization goal and does not consider server load-related issues. It is easy for some servers to be overloaded due to unbalanced server loads, and local hot spots or cooling waste due to large temperature differences between servers. The second type of method mentioned above can regulate the cooling system in advance according to the load distribution, which can solve the problems of control lag and load imbalance to a certain extent. However, this type of method only considers the one-way optimization process from "load-cooling capacity" during load scheduling, and does not take into account the real-time temperature differences of servers. It is easy for local hot spots or cooling waste due to large temperature differences between servers. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, this application provides a collaborative control method for data center IT load and cooling system based on the TD3 algorithm, which solves the current problem of large limitations in the regulation and management of various equipment in the data center and the unsatisfactory balance between server IT load and cooling system energy consumption.
[0006] To achieve the above objectives, this application is implemented through the following technical solutions: In the first aspect, an embodiment of the present application provides a method for collaborative control of IT load and cooling system of a data center based on the TD3 algorithm, and the method includes: obtaining an initial data set that characterizes the server status, environment and task characteristics of the data center; the data center includes a cooling system and multiple servers; performing outlier detection, missing value filling and normalization on the initial data set to obtain a target data set; building a deep reinforcement learning model based on the TD3 algorithm and including a policy network Actor and a value network Critic, determining the state space, action space and reward function of the model; the state variables in the state space are obtained by covering the individual states of the servers , task characteristics and environmental characteristics to expand the dimension and finely characterize the dynamic changes of the data center; the action variables in the action space are used to construct the task allocation strategy and power control strategy; the reward function corresponds to the energy consumption of the cooling system, load balancing, server temperature difference and temperature compliance; the deep reinforcement learning model is trained based on the target data set to obtain the target model; the work task allocation strategy and the cooling system control strategy are output through the interaction between the target model and the environment to allocate servers to the tasks to be assigned, and the cooling system is regulated to build a dynamic closed loop of load distribution, cooling state feedback and load redistribution; among them, the work task allocation strategy and the cooling system control strategy are determined by the action variable representation of the action space.
[0007] According to the first aspect of an embodiment of the present application, the initial data set includes a first data set, a second data set, and a third data set, which respectively characterize the server status, environment, and task characteristics; the first data set includes the server temperature, CPU load, and memory usage; the second data set includes the ambient temperature and ambient humidity of the data center; the third data set includes the computational amount, memory requirements, and priority label of the task to be assigned.
[0008] According to the first aspect of the embodiment of the present application, the state space includes multiple state variables, which characterize the specific state of the data center at a certain moment and are used to provide the basic information required for decision-making for the deep reinforcement learning model; the individual state of the server corresponds to the server temperature, CPU load and memory usage; the task characteristics correspond to the computational amount of the task to be assigned and the memory requirement of the task to be assigned; the environmental characteristics correspond to the ambient temperature and ambient humidity of the data center.
[0009] According to a first aspect of the embodiment of the present application, the state space satisfies the expression: Where, represents the state space, represents the state variable at time 1, represents the state variables at time 2, represents the state variables at time 3, represents the state variable at time t; represents the temperature of the i-th server at time t, represents the CPU load of the i-th server at time t, represents the memory usage of the i-th server at time t; In matrix form, it represents the computational and memory requirements of the queue to be allocated at time t. represents the ambient temperature of the data center at time t, represents the ambient humidity of the data center at time t, and N is the total number of servers in the data center.
[0010] According to a first aspect of an embodiment of the present application, the action space includes a plurality of action variables, and the action variables satisfy the expression: in, represents the action variable, corresponds to the task allocation strategy and is in binary form and represents the decision of assigning task j to server i at time t. It corresponds to the power control strategy and represents the power of the k-th refrigeration equipment at time t.
[0011] According to the first aspect of the embodiment of the present application, the reward function is an immediate feedback signal after the agent performs each action to evaluate each decision of the model; the reward function satisfies the expression: Where, is the reward function, The baseline energy consumption of the cooling system is set as is the energy consumption of the cooling system at time t, Indicates the weight of the cooling system energy consumption reward in the overall reward; is the standard deviation of CPU load at time t, is the preset CPU load target standard deviation, Indicates the weight of load balancing reward in the overall reward; is the standard deviation of server temperature at time t, is the preset server temperature target standard deviation, Indicates the weight of the server temperature difference reward in the overall reward; represents the temperature of the i-th server at time t; Satisfies the expression: Where M is an arbitrarily large positive number, a and b are the set server temperature boundary values, Indicates the weight of the server temperature compliance reward in the overall reward.
[0012] According to the first aspect of the embodiment of the present application, the value network Critic includes two independent Critic networks, the parameter update processes of the two Critic networks are independent of each other, and each Critic network includes a main Critic network and a target Critic network; the policy network Actor includes a main Actor network and a target Actor network, the main Actor network is used to generate the policy of the current state, and the target Actor network is used to generate the next action of the next state in the target Critic network; wherein, the two main Critic networks respectively evaluate the state action under the current policy, output two Q function values to calculate the time difference error, update their own parameters, and provide an update basis to the policy network Actor; the two target Critic networks evaluate the next state action pair and output two target Q function values.
[0013] According to the first aspect of the embodiment of the present application, the aforementioned training of the deep reinforcement learning model based on the target data set to obtain the target model can specifically include the following steps: outputting action variables based on the current state through the policy network Actor; wherein the action variables are used to control task allocation and regulate the cooling system; respectively evaluating the action variables taken in the current state through two main Critic networks, and continuing to execute the cumulative reward expectations corresponding to subsequent actions to obtain two Q function values; determining the minimum of the two target Q function values generated by the two target Critic networks as the target Q value; based on the target Q value, reversely calculating the actual Q value of the previous step, calculating the mean square error between the actual Q value and the two Q function values generated by the two main Critic networks, and taking the sum to obtain the total error; using the optimizer to update the parameters of the two main Critic networks according to the total error; using the policy delay mechanism to update the policy network Actor, and obtaining the target model through iterative interactive training of the policy network Actor and the value network Critic.
[0014] According to the first aspect of the embodiment of the present application, the aforementioned outlier detection, missing value filling and normalization processing of the initial data set to obtain the target data set can specifically include the following steps: calculating the first median of the data in the initial data set through the Hampel identifier, and calculating the distance from each data point in the initial data set to the first median to obtain multiple distances; determining the second median of the multiple distances as the median absolute deviation MAD, and correcting the median absolute deviation MAD through a preset adjustment parameter to obtain a target threshold ; Determine the distance greater than The data points corresponding to the target distance are outliers and are removed from the initial data set to obtain the first array after removal; the missing values of the first array are filled by linear interpolation to obtain the second array; the data in the second array are normalized by the maximum and minimum normalization method to obtain the target data set.
[0015] In the second aspect, an embodiment of the present application provides a data center IT load and cooling system collaborative control system based on the TD3 algorithm. The data center IT load and cooling system collaborative control system based on the TD3 algorithm includes: a data acquisition module, a data processing module, a model building module, a model training module and a strategy output module.
[0016] Specifically, the data acquisition module is used to obtain an initial data set that characterizes the server status, environment, and task characteristics of the data center; the data center includes a cooling system and multiple servers; the data processing module is used to detect outliers, fill in missing values, and normalize the initial data set to obtain a target data set; the model building module is used to build a deep reinforcement learning model based on the TD3 algorithm and including a policy network Actor and a value network Critic, and determine the model's state space, action space, and reward function; the state variables in the state space expand the dimension and finely characterize the dynamic changes of the data center by covering the individual status of the server, task characteristics, and environmental characteristics; the action variables in the action space are used to construct task allocation strategies and power control strategies; the reward function corresponds to cooling system energy consumption, load balancing, server temperature difference, and temperature compliance; the model training module is used to train the deep reinforcement learning model based on the target data set to obtain a target model; the policy output module is used to output a work task allocation strategy and a cooling system control strategy through the interaction between the target model and the environment, so as to allocate servers to the tasks to be assigned and regulate the cooling system to build a dynamic closed loop for load distribution, cooling state feedback, and load redistribution; wherein the work task allocation strategy and cooling system control strategy are determined by the action variable representation of the action space.
[0017] This application provides a method for collaboratively controlling IT loads and cooling systems in data centers based on the TD3 algorithm. Compared with existing technologies, it has the following advantages: In order to collaboratively optimize the IT load and cooling system of a data center, this application constructs a deep reinforcement learning model based on the TD3 algorithm, and determines the state space, action space, and reward function of the deep reinforcement learning model; the state space can define the perception range of the agent, the action space determines the control ability of the agent, and the reward function guides the agent to select the optimal action in a given state space. The data required for model training corresponds to information covering various aspects such as the server status, environment, and task characteristics of the data center, and considers the collaborative management of multiple goals; the action variable can represent the process of the target model making work task allocation decisions and cooling system regulation decisions based on the current environmental status. After continuous exploration and learning, the target model can interact with the current environment to output a comprehensively balanced work task allocation strategy and cooling system control strategy, thereby guiding the work task allocation and controlling the operation of the cooling system to achieve a balance between the server IT load and the cooling system, thereby ensuring the safe and stable operation of the server and reducing the energy consumption of the cooling system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 This is a flow chart of a method for collaboratively controlling IT load and cooling system in a data center based on the TD3 algorithm provided in an embodiment of the present application; Figure 2 yes Figure 1 An exemplary flow chart of S140; Figure 3 This is a structural diagram of a data center IT load and cooling system collaborative control system based on the TD3 algorithm provided in an embodiment of the present application. DETAILED DESCRIPTION To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0020] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0021] The embodiments of the present application provide a method for collaboratively controlling the IT load and cooling system of a data center based on the TD3 algorithm, thereby solving the problem that the current regulation and management of various equipment in a data center is quite limited, and the trade-off between the server IT load and the energy consumption of the cooling system is not ideal.
[0022] In order to better understand the above technical solution, the above technical solution will be described in detail below with reference to the accompanying drawings and specific implementation methods.
[0023] The following first introduces a method for collaborative control of data center IT load and cooling system based on the TD3 algorithm provided in an embodiment of the present application.
[0024] The embodiment of the present application provides a flow chart of a method for collaboratively controlling IT load and cooling system of a data center based on the TD3 algorithm, as shown in FIG. Figure 1 As shown, the data center IT load and cooling system coordinated control method based on the TD3 algorithm may include the following steps S110-S150.
[0025] S110 , obtaining an initial data set representing server status, environment, and task characteristics of a data center; the data center includes a cooling system and multiple servers.
[0026] S120 , performing outlier detection, missing value filling, and normalization processing on the initial data set to obtain a target data set.
[0027] S130. Build a deep reinforcement learning model based on the TD3 algorithm and including the policy network Actor and the value network Critic, and determine the model's state space, action space, and reward function; the state variables in the state space expand the dimension and finely characterize the dynamic changes of the data center by covering the individual state of the server, task characteristics, and environmental characteristics; the action variables in the action space are used to construct task allocation strategies and power control strategies; the reward function corresponds to the cooling system energy consumption, load balancing, server temperature difference, and temperature compliance.
[0028] S140. Train the deep reinforcement learning model based on the target data set to obtain a target model.
[0029] S150. Output the work task allocation strategy and the cooling system control strategy through the interaction between the target model and the environment, so as to allocate servers to the tasks to be assigned, and regulate the cooling system to build a dynamic closed loop of load distribution, cooling state feedback and load redistribution; wherein, the work task allocation strategy and the cooling system control strategy are determined by the action variable representation of the action space.
[0030] The above is a specific implementation method of a collaborative control method of data center IT load and cooling system based on the TD3 algorithm provided in an embodiment of the present application. It can be understood that the Twin Delayed Deep Deterministic Policy Gradient algorithm (TD3 algorithm) is a deep reinforcement learning DRL algorithm based on the Actor-critic framework. The TD3 algorithm constructs a policy network Actor and a value network Critic. By training the deep reinforcement learning model, it can solve the over-estimation problem of the traditional DDPG (Deep Deterministic Policy Gradient) algorithm, thereby ensuring the reliability of the decision.
[0031] It should be noted that in the TD3 algorithm, the state space defines the agent's perceptual range, the action variables determine the agent's control capabilities, the state space must contain the key information required for reward calculation, and the reward function must guide the agent to select the optimal action within a given state space; this collaborative relationship is the foundation of the TD3 algorithm's efficient and continuous control. This application, for collaborative optimization of data center IT loads and cooling systems, builds a deep reinforcement learning model based on the TD3 algorithm and determines the deep reinforcement learning model's state space, action variables, and reward function.
[0032] Furthermore, the data required for model training corresponds to information covering various aspects of the data center, such as the server status, environmental data, and task characteristics. This application considers the collaborative management of multiple objectives. The action variable can characterize the process by which the target model makes work task allocation decisions and cooling system regulation decisions based on the current environmental status. After continuous exploration and learning, the target model can interact with the current environment to output a comprehensive and balanced work task allocation strategy and cooling system control strategy, thereby guiding task allocation, adjusting the power of the cooling system, and achieving a balance between the server IT load and the cooling system cooling, thereby ensuring the safe and stable operation of the server and reducing the energy consumption of the cooling system.
[0033] In some embodiments, the initial data set includes a first data set, a second data set, and a third data set, respectively characterizing server status, environment, and task characteristics; The first data set includes server temperature, CPU load, and memory usage; the second data set includes the ambient temperature and humidity of the data center; and the third data set includes the computational load, memory requirements, and priority tags of the tasks to be assigned.
[0034] In the embodiments of the present application, it can be understood that the present application uses a deep reinforcement learning algorithm to build a model, taking into account multiple data such as server temperature, CPU load, memory usage, and cooling system energy consumption, thereby balancing the server IT load rate, server temperature, and cooling system energy consumption, and ultimately achieving an optimization effect. In some embodiments, the aforementioned outlier detection, missing value filling, and normalization processing are performed on the initial data set to obtain the target data set, that is, the aforementioned S120 may specifically include the following steps: S210, calculating a first median of the data in the initial data set using a Hampel identifier, and calculating a distance from each data point in the initial data set to the first median to obtain a plurality of distances; S220: Determine the second median of the multiple distances as the median absolute deviation (MAD), and modify the median absolute deviation (MAD) using a preset adjustment parameter to obtain a target threshold. ; S230, determine the distance greater than The data points corresponding to the target distance are outliers and are removed from the initial data set to obtain a first array after removal; S240, using linear interpolation to fill missing values in the first array to obtain a second array; S250 , normalizing the data in the second array using a maximum and minimum value normalization method to obtain a target data set.
[0035] In the embodiments of the present application, it is understood that after obtaining the initial data set, the target data set is obtained by performing preprocessing operations such as outlier detection, missing value filling, and normalization, which can avoid the problem of excessive data deviation caused by extreme cases and help improve data accuracy. The Hampel identifier relies on the median absolute deviation (MAD) and uses a rolling window to identify outliers. MAD is a robust measure of data dispersion. In the presence of a small number of extreme values, since the median is relatively stable, the Hampel identifier has high robustness and good tolerance to noise and non-normal distribution in the data.
[0036] In some embodiments, the state space includes multiple state variables, which represent the specific state of the data center at a certain moment and are used to provide the basic information required for decision-making for the deep reinforcement learning model; the individual state of the server corresponds to the server temperature, CPU load and memory usage; the task characteristics correspond to the computational amount of the task to be assigned and the memory requirement of the task to be assigned; the environmental characteristics correspond to the ambient temperature and ambient humidity of the data center.
[0037] In the embodiments of the present application, it is understood that the present application considers collaborative control analysis in multiple dimensions, focusing on the balance of cooling system energy consumption, IT load, and server temperature; reducing the temperature difference between multiple servers, avoiding local hot spots, and reducing cooling waste. The present application constructs a deep collaborative mechanism. The target model can output the work task allocation strategy and cooling system control strategy based on the current state information, forming a complete "load distribution-cooling state feedback-load redistribution" closed loop, solving the local hot spot problem.
[0038] In one example, the state space satisfies the expression: Where, represents the state space, represents the state variable at time 1, represents the state variables at time 2, represents the state variables at time 3, represents the state variable at time t; represents the temperature of the i-th server at time t, represents the CPU load of the i-th server at time t, represents the memory usage of the i-th server at time t; In matrix form, it represents the computational and memory requirements of the queue to be allocated at time t. represents the ambient temperature of the data center at time t, represents the ambient humidity of the data center at time t, and N is the total number of servers in the data center.
[0039] In one example, the action space includes multiple action variables, and the action variables satisfy the expression: in, represents the action variable, corresponds to the task allocation strategy and is in binary form and represents the decision of assigning task j to server i at time t. It corresponds to the power control strategy and represents the power of the k-th refrigeration equipment at time t.
[0040] It can be understood that since the process of the aforementioned target model making work task allocation decisions and cooling system control decisions based on the current environmental status is determined by the action variable representation; the work task allocation strategy output by the interaction between the target model and the environment is used to allocate the tasks to be assigned to the appropriate servers in the data center, and the cooling system control strategy is used to adjust the power of each refrigeration equipment.
[0041] Based on this, this application considers task allocation and the power of the cooling system when constructing action variables; takes the power of the refrigeration equipment as part of the reward function in the optimization model, and uses deep reinforcement learning algorithms to continuously explore and learn. By incorporating the control parameters of the cooling system into the action variables for training together, a cooling system control strategy consistent with the overall operation of the data center is generated, and the optimal cooling system control strategy is obtained, which realizes precise cooling and solves the control lag problem, thereby realizing intelligent cooling system control.
[0042] In some embodiments, the reward function is the immediate feedback signal after each action performed by the agent to evaluate each decision of the model. This application considers several factors when constructing the reward function, such as cooling system energy consumption, CPU load, server temperature difference, and server temperature compliance. The deep reinforcement learning model of this application incorporates temperature, energy consumption, and load rate into the reward function, enabling collaborative control analysis.
[0043] Specifically, the lower the cooling system's energy consumption, the higher the reward; the more balanced the CPU load, the higher the reward; the smaller the server temperature difference, the higher the reward; and positive rewards are given when the server temperature is within the set optimal temperature range [a, b], and negative rewards otherwise. This application uses a reward function to balance cooling system energy consumption, server load rate, server temperature difference, and server temperature range, solving the independent optimization or one-way linkage problems of traditional methods.
[0044] In one example, the reward function satisfies the expression: Where, is the reward function, The baseline energy consumption of the cooling system is set as is the energy consumption of the cooling system at time t, Indicates the weight of the cooling system energy consumption reward in the overall reward; is the standard deviation of CPU load at time t, is the preset CPU load target standard deviation, Indicates the weight of load balancing reward in the overall reward; is the standard deviation of server temperature at time t, is the preset server temperature target standard deviation, Indicates the weight of the server temperature difference reward in the overall reward; represents the temperature of the i-th server at time t; Satisfies the expression: Where M is an arbitrarily large positive number, a and b are the set server temperature boundary values, Indicates the weight of the server temperature compliance reward in the overall reward.
[0045] It is understandable that this application adds a temperature constraint to the reward function, balancing the different server temperatures through the reward function component of the server temperature standard deviation. By setting an appropriate temperature range and M value, actions that cause temperature non-compliance are penalized to ensure safe and stable server operation. In addition, this application uses a deep reinforcement learning algorithm to dynamically allocate tasks based on the real-time status of the server's temperature, load rate, and memory, so that the server load rate standard deviation remains within a certain threshold, thereby balancing the load rates of different servers and achieving balanced IT load scheduling.
[0046] In some embodiments, the value network Critic includes two independent Critic networks, the parameter update processes of the two Critic networks are independent of each other, and each Critic network includes a main Critic network and a target Critic network; the strategy network Actor includes a main Actor network and a target Actor network, the main Actor network is used to generate the strategy for the current state, and the target Actor network is used to generate the next action in the target Critic network for the next state.
[0047] Among them, the two main critic networks respectively evaluate the state action under the current strategy, output two Q function values to calculate the time difference error, update their own parameters, and provide update basis to the policy network Actor; the two target critic networks evaluate the next state action pair and output two target Q function values.
[0048] In the embodiment of the present application, it can be understood that the value network Critic is mainly used to evaluate the quality of the actions taken by the policy network Actor. The value network Critic learns the state value function, which estimates the expected value of the future cumulative rewards that may be obtained by following the current strategy under a specific state; when the policy network Actor takes an action, the environment will feedback a new state and immediate reward, which can be determined according to the reward function; the value network Critic will evaluate the long-term value brought by the action based on this new state and previous experience.
[0049] In some embodiments, as Figure 2 As shown, the aforementioned deep reinforcement learning model is trained based on the target data set to obtain the target model, that is, the aforementioned S140 may specifically include the following steps: S310, outputting action variables based on the current state through the policy network Actor; wherein the action variables are used to control task allocation and regulate the cooling system; S320, using two main critic networks to evaluate the action variables taken in the current state and continue to execute the cumulative reward expectations corresponding to subsequent actions to obtain two Q function values; S330, determining the minimum of the two target Q function values generated by the two target Critic networks as the target Q value; S340, reversely calculate the actual Q value of the previous step based on the target Q value, calculate the mean square error between the actual Q value and the two Q function values generated by the two main critic networks, and sum them to obtain the total error; S350, using the optimizer to update the parameters of the two main critic networks according to the total error; S360, use the policy delay mechanism to update the policy network Actor, and obtain the target model through iterative interactive training of the policy network Actor and the value network Critic.
[0050] In the examples of this application, it can be understood that the TD3 algorithm's workflow primarily includes: feeding the current state and action into the network, which predicts the expected reward for the next state; calculating the gradient between the two evaluators; updating the parameters between the two networks; and repeating these steps until the network reaches a stable state. The policy network (Actor) outputs actions based on the policy, and the value network (Critic) evaluates the actions. Through the interaction and repeated training of the two, the entire system can learn the optimal policy more quickly and stably, thereby making more reasonable decisions in complex environments.
[0051] During the model training process, the minimum of the two target Q function values generated by the two target Critic networks is selected as the target Q value, thereby reducing the estimated Q value variance, avoiding overestimation, and making the result more stable. Specifically, the main Critic network and the target Critic network are updated at each time step, and the main Actor network is updated after every d steps, where d is a positive integer. The main Critic network calculates the mean square error in the aforementioned step S340 at each step, and updates the parameters of the main Critic network in reverse based on the size of the mean square error, in order to minimize the mean square error; the target Critic network is soft-updated based on the parameters of the main Critic network, that is, the parameters of the target Critic network are not directly equal to the parameters of the main Critic network, but are assigned a certain proportion based on the parameters of the main Critic network.
[0052] Furthermore, during the update process of the main actor network, the gradient is calculated based on the main critic network. The parameters of the main actor network that maximizes the main critic network's evaluation of the current policy are found and updated. The update frequency is set by parameters. The target actor network is soft-updated based on the parameters of the main actor network.
[0053] It should be understood that since the main actor network generates the policy, if the main critic network is updated every time it is updated, the generated policy will be particularly volatile. Therefore, this application delays the update of the main actor network to allow it to initially converge and stabilize before updating it under more stable conditions. The delayed update mechanism may include: when the number of target Q values reaches a certain index, the main actor network is updated to give the model sufficient time to learn accurate Q value estimates.
[0054] In some embodiments, the present application provides a data center IT load and cooling system collaborative control system 400 based on the TD3 algorithm, such as Figure 3 As shown, the data center IT load and cooling system coordinated control system 400 based on the TD3 algorithm may include the following modules: A data acquisition module 410 is configured to acquire an initial data set representing server status, environment, and task characteristics of a data center, wherein the data center includes a cooling system and a plurality of servers; The data processing module 420 is used to perform outlier detection, missing value filling and normalization on the initial data set to obtain a target data set; Model building module 430 is used to build a deep reinforcement learning model based on the TD3 algorithm and including a policy network Actor and a value network Critic, and determine the model's state space, action space, and reward function. The state variables in the state space expand the dimension and finely characterize the dynamic changes of the data center by covering individual server states, task characteristics, and environmental characteristics. The action variables in the action space are used to construct task allocation strategies and power control strategies. The reward function corresponds to cooling system energy consumption, load balancing, server temperature difference, and temperature compliance. Model training module 440 is used to train the deep reinforcement learning model based on the target data set to obtain a target model. The strategy output module 450 is used to output the work task allocation strategy and the cooling system control strategy through the interaction between the target model and the environment, so as to allocate servers to the tasks to be assigned and regulate the cooling system to build a dynamic closed loop of load distribution, cooling state feedback and load redistribution; wherein the work task allocation strategy and the cooling system control strategy are determined by the action variable representation of the action space.
[0055] According to an embodiment of the present application, any multiple modules among the data acquisition module 410, the data processing module 420, the model building module 430, the model training module 440, and the strategy output module 450 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module.
[0056] In some embodiments, the data processing module 420 may be specifically configured to: Calculate the first median of the data in the initial data set through the Hampel identifier, and calculate the distance from each data point in the initial data set to the first median to obtain multiple distances; Determine the second median of multiple distances as the median absolute deviation MAD, and correct the median absolute deviation MAD through the preset adjustment parameters to obtain the target threshold ; Determine the distance greater than The data points corresponding to the target distance are outliers and are removed from the initial data set to obtain a first array after removal; Use linear interpolation to fill missing values in the first array to obtain the second array; The data in the second array is normalized using the maximum and minimum value normalization method to obtain the target data set.
[0057] In some embodiments, the model training module 440 may be specifically used to: The policy network Actor outputs action variables based on the current state; the action variables are used to control task allocation and regulate the cooling system; The two main critic networks evaluate the action variables taken in the current state and the cumulative reward expectations corresponding to the subsequent actions to obtain two Q function values; Determine the minimum of the two target Q function values generated by the two target Critic networks as the target Q value; Based on the target Q value, reversely calculate the actual Q value of the previous step, calculate the mean square error between the actual Q value and the two Q function values generated by the two main critic networks, and take the sum to get the total error; Use the optimizer to update the parameters of the two main critic networks based on the total error; The policy network Actor is updated using the policy delay mechanism, and the target model is obtained through iterative interactive training of the policy network Actor and the value network Critic.
[0058] Figure 3 Each module in the system shown has the function of implementing each step in the aforementioned data center IT load and cooling system coordinated control method based on the TD3 algorithm, and can achieve its corresponding technical effects. For the sake of brevity, it will not be repeated here.
[0059] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A collaborative control method for IT load and cooling system of a data center based on TD3 algorithm, characterized in that: include: Acquire an initial dataset characterizing server status, environment, and task characteristics of a data center; the data center includes a cooling system and a plurality of servers; Performing outlier detection, missing value filling, and normalization on the initial data set to obtain a target data set; A deep reinforcement learning model based on the TD3 algorithm, including a policy network (Actor) and a value network (Critic), was constructed to determine the model's state space, action space, and reward function. The state variables in the state space expanded and refined the dynamic changes of the data center by covering individual server states, task characteristics, and environmental characteristics. The action variables in the action space were used to construct task allocation and power control strategies. The reward function corresponded to cooling system energy consumption, load balancing, server temperature differences, and temperature compliance. Training the deep reinforcement learning model based on the target data set to obtain a target model; The work task allocation strategy and the cooling system control strategy are output through the interaction between the target model and the environment to allocate servers to the tasks to be assigned, and the cooling system is regulated to build a dynamic closed loop of load distribution, cooling state feedback and load redistribution; wherein the work task allocation strategy and the cooling system control strategy are determined by the action variable representation of the action space.
2. The collaborative control method for IT load and cooling system of a data center based on the TD3 algorithm according to claim 1, characterized in that: The initial data set includes a first data set, a second data set and a third data set respectively characterizing the server state, the environment and the task characteristics; The first data set includes server temperature, CPU load and memory usage; the second data set includes the ambient temperature and humidity of the data center; and the third data set includes the computational load, memory requirements and priority tag of the task to be assigned.
3. The collaborative control method for IT load and cooling system of a data center based on the TD3 algorithm according to claim 2, characterized in that: The state space includes multiple state variables, which represent the specific state of the data center at a certain moment and are used to provide basic information required for decision-making for the deep reinforcement learning model; the individual server state corresponds to the server temperature, the CPU load and the memory usage; the task characteristics correspond to the computational amount of the task to be assigned and the memory requirement of the task to be assigned; the environmental characteristics correspond to the ambient temperature and ambient humidity of the data center.
4. The data center IT load and cooling system coordinated control method based on the TD3 algorithm according to claim 3 is characterized in that: The state space satisfies the expression: Where, represents the state space, represents the state variable at time 1, represents the state variables at time 2, represents the state variables at time 3, represents the state variable at time t; represents the temperature of the i-th server at time t, represents the CPU load of the i-th server at time t, represents the memory usage of the i-th server at time t; In matrix form, it represents the computational and memory requirements of the queue to be allocated at time t. represents the ambient temperature of the data center at time t, represents the ambient humidity of the data center at time t, and N is the total number of servers in the data center.
5. The data center IT load and cooling system coordinated control method based on the TD3 algorithm according to claim 1, characterized in that: The action space includes a plurality of action variables, and the action variables satisfy the expression: in, represents the action variable, corresponds to the task allocation strategy, and is in binary form and represents the decision of assigning task j to server i at time t. corresponds to the power control strategy and represents the power of the k-th refrigeration device at time t.
6. The data center IT load and cooling system coordinated control method based on the TD3 algorithm according to claim 5, characterized in that: The reward function is the immediate feedback signal after the agent performs each action to evaluate each decision of the model; the reward function satisfies the expression: Where, is the reward function, The baseline energy consumption of the cooling system is set as is the energy consumption of the cooling system at time t, Indicates the weight of the cooling system energy consumption reward in the overall reward; is the standard deviation of CPU load at time t, is the preset CPU load target standard deviation, Indicates the weight of load balancing reward in the overall reward; is the standard deviation of server temperature at time t, is the preset server temperature target standard deviation, Indicates the weight of the server temperature difference reward in the overall reward; represents the temperature of the i-th server at time t; Satisfies the expression: Where M is an arbitrarily large positive number, a and b are the set server temperature boundary values, Indicates the weight of the server temperature compliance reward in the overall reward.
7. The method for collaboratively controlling IT load and cooling system of a data center based on the TD3 algorithm according to claim 1, characterized in that: The value network Critic includes two independent Critic networks, the parameter update processes of the two Critic networks are independent of each other, and each Critic network includes a main Critic network and a target Critic network; The policy network Actor includes a main Actor network and a target Actor network. The main Actor network is used to generate a policy for the current state, and the target Actor network is used to generate the next action for the next state in the target Critic network. Among them, the two main critic networks respectively evaluate the state action under the current strategy, output two Q function values to calculate the time difference error, update their own parameters, and provide update basis to the strategy network Actor; the two target critic networks evaluate the next state action pair and output two target Q function values.
8. The data center IT load and cooling system coordinated control method based on the TD3 algorithm according to claim 7, characterized in that: The training of the deep reinforcement learning model based on the target data set to obtain a target model includes: Outputting action variables based on the current state through the policy network Actor; wherein the action variables are used to control task allocation and regulate the cooling system; The two main critic networks respectively evaluate the cumulative reward expectations corresponding to taking the action variable in the current state and continuing to perform subsequent actions to obtain two Q function values; Determine the minimum of the two target Q function values generated by the two target Critic networks as the target Q value; Based on the target Q value, reversely calculate the actual Q value of the previous step, calculate the mean square error between the actual Q value and the two Q function values generated by the two main critic networks, and take the sum to obtain the total error; Using an optimizer to update the parameters of the two main critic networks according to the total error; The policy network Actor is updated using a policy delay mechanism, and the target model is obtained through iterative interactive training of the policy network Actor and the value network Critic.
9. The data center IT load and cooling system coordinated control method based on the TD3 algorithm according to claim 1, characterized in that: The performing outlier detection, missing value filling and normalization processing on the initial data set to obtain a target data set includes: Calculating a first median of the data in the initial data set using a Hampel identifier, and calculating a distance from each data point in the initial data set to the first median to obtain a plurality of distances; Determine the second median of the multiple distances as the median absolute deviation MAD, and correct the median absolute deviation MAD by a preset adjustment parameter to obtain the target threshold ; Determine the distances greater than The data points corresponding to the target distance are outliers, and are removed from the initial data set to obtain a first array after removal; Filling missing values in the first array using linear interpolation to obtain a second array; The data in the second array is normalized using a maximum and minimum value normalization method to obtain a target data set.
10. A data center IT load and cooling system collaborative control system based on the TD3 algorithm, characterized in that: include: A data acquisition module, used to obtain an initial data set that characterizes the server status, environment, and task characteristics of the data center; The data center includes a cooling system and a plurality of servers; A data processing module is used to perform outlier detection, missing value filling and normalization on the initial data set to obtain a target data set; A model building module is used to build a deep reinforcement learning model based on the TD3 algorithm and including a policy network actor and a value network critic, and to determine the model's state space, action space, and reward function. The state variables in the state space expand the dimension and finely characterize the dynamic changes of the data center by covering individual server states, task characteristics, and environmental characteristics. The action variables in the action space are used to construct task allocation strategies and power control strategies. The reward function corresponds to cooling system energy consumption, load balancing, server temperature difference, and temperature compliance. A model training module, configured to train the deep reinforcement learning model based on the target dataset to obtain a target model; A strategy output module is used to output a work task allocation strategy and a cooling system control strategy through the interaction between the target model and the environment, so as to allocate servers to the tasks to be assigned and regulate the cooling system to build a dynamic closed loop of load distribution, cooling state feedback and load redistribution; wherein the work task allocation strategy and the cooling system control strategy are determined by the action variable representation of the action space.
Citation Information
Patent Citations
Multi-data center collaborative energy-saving method based on multi-agent reinforcement learning
CN113064480A
Data center energy consumption optimization method and device and storage medium
CN116795198A
Energy-saving predictive control method suitable for data center water cooling system
CN117408170A
Method for optimizing joint operation of task scheduling and cooling control of data center
CN117891578A
Task execution method, electronic equipment, storage medium and program product
CN120085991A
Cited By
Cold storage PID control parameter optimization method, device and equipment and storage medium
CN121386355A
Deep reinforcement learning adaptive regulation and control method and system of air-liquid mixed cooling system
CN121604369A
Intelligent control method and system for immersion liquid cooling system based on AI
CN121879541A
Dynamic control method for wind-liquid collaborative cooling tail end of data center under load fluctuation
CN121920243A