Joint optimization method and device based on task scheduling and thermal management and medium

By building a dual-time scale control framework, combining task scheduling and thermal management, using deep reinforcement learning and multi-agent algorithms to optimize task allocation and cooling, the problem of high energy consumption and thermal management mismatch of computing-intensive tasks is solved, and the energy efficiency and stability of the data center are improved.

CN120295780APending Publication Date: 2025-07-11CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510366365.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Computation-intensive tasks have high energy consumption and high thermal load, resulting in mismatch between thermal management and task volume, resulting in waste of resources and waste of cooling capacity or the generation of local hot spots.

Method used

A dual-time-scale control framework is built, combining task scheduling and thermal management, and a second-level task scheduling decision is made on a fast time scale through deep reinforcement learning algorithms, and a multi-agent reinforcement learning algorithm performs minute-level collaborative thermal management on a slow time scale to optimize task allocation and cooling requirements.

Benefits of technology

The energy efficiency improvement of the data center, the operation cost reduction and the task stay time are shortened, ensuring the thermal stability and resource utilization of the data center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295780A_ABST
    Figure CN120295780A_ABST
Patent Text Reader

Abstract

The invention discloses a joint optimization method and equipment based on task scheduling and thermal management and a medium, which can dynamically adjust the matching of refrigeration demand and supply by combining computational intensive task scheduling and thermal management, and can improve the matching efficiency for the problem of mismatching of time constants between an IT subsystem and a refrigeration subsystem. A dual-time scale control framework is introduced, and second-level or millisecond-level task scheduling decision is carried out by using a deep reinforcement learning algorithm on a fast time scale; on a slow time scale, a multi-agent reinforcement learning algorithm is adopted to realize minute-level collaborative thermal management across multiple data centers. According to the method, the operation cost, the task residence time and the energy use efficiency can be well balanced, the joint target of task scheduling and heat management is optimized, the multi-target optimization challenge faced in the geographic distribution data center is effectively dealt with, and the energy efficiency and the stability of the system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of thermal load management, and particularly relates to a joint optimization method, device and medium based on task scheduling and thermal management. Background Art

[0002] With the development of artificial intelligence and large language models, artificial intelligence has generated a large number of computationally intensive tasks. These computationally intensive tasks not only consume high energy but also have a large thermal load, and their execution may disrupt the thermal balance of the data center. The main sources of energy consumption in the data center are the IT subsystem and the cooling subsystem. The time constants of these two subsystems do not match. The task scheduling of the IT subsystem is usually completed at the second or millisecond level, while the adjustment of the cooling subsystem is limited by thermal inertia and usually takes several minutes (such as 15 minutes). This mismatch in time constants may lead to problems such as hot spots, unnecessary cooling fluctuations, and over-provisioning of cooling capacity. Moreover, large-scale tasks (such as training large language models) require a large amount of computing resources, usually exceeding the capabilities of a single data center. Distributed data centers provide solutions by efficiently managing these large-scale workloads. However, their distributed nature introduces high-dimensional and highly dynamic optimization challenges. These challenges stem from the diversity of devices and facilities, as well as the changes in workload, environmental conditions, and system state over time, further increasing the complexity of achieving energy-efficient operation.

[0003] Many studies have been devoted to improving the energy efficiency of data centers, mainly divided into three categories: task scheduling, thermal management, and joint optimization of the two. Existing task scheduling methods improve resource utilization and reduce costs by allocating tasks to appropriate data center servers, including heuristic-based, model-based, and learning-based methods. However, only focusing on task scheduling often ignores the impact of task computing load on the thermal balance of the data center. Existing thermal management methods usually control the cooling of the data center by adjusting the set temperature or air flow rate of the air cooling unit (ACU), but these methods usually ignore the dynamics of the IT subsystem and only consider the rack temperature or indoor air temperature for cooling subsystem control. Due to the lack of coordination between IT and cooling and the pursuit of safe operation, these isolated methods often lead to waste of cooling capacity or the generation of local hot spots. Summary of the Invention

[0004] The technical problem to be solved by the present invention is that computing-intensive tasks have high energy consumption and large heat loads, resulting in a mismatch between thermal management and the amount of tasks, causing waste of resources. The purpose is to provide a joint optimization method, device, and medium based on task scheduling and thermal management. By constructing a dual-time-scale control framework and performing joint regulation of task scheduling and thermal management based on the dual-time-scale control framework, the energy efficiency of the data center can be improved, and the operating expenses and task sojourn time can be reduced. By comprehensively optimizing task scheduling and thermal management and balancing the computing load and cooling demand, the thermal stability of the data center can be ensured.

[0005] The present invention is realized through the following technical solutions:

[0006] In the first aspect of the present invention, a joint optimization method based on task scheduling and thermal management is provided, including the following specific steps:

[0007] Construct a dual-time-scale control framework, where the dual-time-scale includes a slow time scale and a fast time scale;

[0008] Construct an operating expense model to obtain the operating cost of the geographically distributed data center for executing tasks;

[0009] Construct an energy utilization efficiency model to obtain the average energy utilization efficiency of the geographically distributed data center;

[0010] Obtain the task waiting time, and construct an optimization model with the goal of minimizing the operating cost of the geographically distributed data center for executing tasks, the average energy utilization efficiency of the geographically distributed data center, and the task waiting time;

[0011] Based on the optimization model, generate optimal task scheduling decisions and multi-data-center collaborative thermal management decisions, including:

[0012] On the fast time scale, construct a task scheduling agent and generate task scheduling decisions;

[0013] On the slow time scale, construct a thermal management agent and generate multi-data-center collaborative thermal management decisions;

[0014] Based on the shared state space, co-train the task scheduling agent and the thermal management agent;

[0015] Based on the trained task scheduling agent and thermal management agent, generate optimal task scheduling decisions and multi-data-center collaborative thermal management decisions.

[0016] Further, the construction of the dual-time-scale control framework specifically includes the steps of establishing a slow time scale and a fast time scale:

[0017] Divide each working day into N slow time steps n, where n ∈ {1, 2,..., N};

[0018] Each slow time step is divided into M fast time steps m, where m ∈ {1, 2,..., M}.

[0019] Furthermore, the construction of the operating expenditure model to obtain the operating cost of a geographically distributed data center for executing tasks specifically includes:

[0020] Obtain the electricity prices of the power grid, solar energy, and wind energy, obtain the consumption ratios of the power grid, solar energy, and wind energy, and construct the operating expenditure model;

[0021] Based on the operating expenditure model, obtain the operating costs of the IT subsystem and the cooling subsystem of the data center at time τ.

[0022] Furthermore, the construction of the energy usage efficiency model to obtain the average energy usage efficiency of a geographically distributed data center specifically includes:

[0023] Obtain the cooling energy consumption of the data center at time τ and the total energy consumption of the center at time τ;

[0024] Based on the double time scale, construct the energy usage efficiency model;

[0025] Based on the energy usage efficiency model, obtain the average energy usage efficiency of the geographically distributed data center.

[0026] Furthermore, the construction of the optimization model specifically includes:

[0027] Dynamically obtain the air flow rate of the computer room air conditioner unit of the data center, the operating cost of the IT subsystem of the data center at time τ, the operating cost of the cooling system of the data center at time τ, the average energy usage efficiency of the geographically distributed data center, the operating cost of the data center, and the penalty term for data center overheating;

[0028] Taking the minimum of the operating cost of the geographically distributed data center for executing tasks, the average energy usage efficiency of the geographically distributed data center, and the task waiting time as the goal, construct the target optimization function;

[0029] Solve the target optimization function to obtain the optimization model.

[0030] Furthermore, on the fast time scale, construct a task scheduling agent and generate a task scheduling decision, specifically including:

[0031] Construct a task scheduling agent based on Markov decision-making;

[0032] Define the state space as the state vector of the task scheduling agent of the geographically distributed data center;

[0033] Define the action space as the selection of allocating candidate tasks to the data center;

[0034] Define a reward function based on the operating cost of the IT subsystem in the data center at time τ, the operating cost of the cooling system in the data center at time τ, and the task waiting time.

[0035] According to the state space, action space, and reward function, train a task scheduling agent through a reinforcement learning algorithm to generate task scheduling decisions.

[0036] Furthermore, on the slow time scale, construct a thermal management agent and generate a multi-data center collaborative thermal management decision, specifically including:

[0037] Construct a thermal management agent based on Markov decision-making.

[0038] Define the state space as the union of all states of geographically distributed data centers.

[0039] Define the action space as continuous control actions for adjusting the air flow rate of computer room air conditioning units.

[0040] Define a reward function based on the average temperature, temperature control, and energy use efficiency of the data center at time τ.

[0041] According to the state space, action space, and reward function, train the thermal management agent through a reinforcement learning algorithm to generate a multi-data center collaborative thermal management decision.

[0042] Furthermore, the collaborative training of the task scheduling agent and the thermal management agent based on state space sharing specifically includes:

[0043] At the beginning of each slow time step, initialize the parameters of the task scheduling agent and the thermal management agent.

[0044] The task scheduling agent selects an action based on the observed environmental state, obtains a reward and then transfers to the next state, while storing the experience in the experience replay buffer of the task scheduling agent.

[0045] At the end of each slow time step, the thermal management agent obtains local observations, selects an action to explore and obtains a reward, then transfers to the next state, while storing the experience in the experience replay buffer of the thermal management agent.

[0046] Alternately sample the experience replay buffers of the task scheduling agent and the thermal management agent to update the network parameters of the task scheduling agent and the thermal management agent.

[0047] The second aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a joint optimization method based on task scheduling and thermal management.

[0048] In the third aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, a joint optimization method based on task scheduling and thermal management is implemented.

[0049] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0050] By combining compute-intensive task scheduling and thermal management, the present invention can dynamically adjust the matching between cooling demand and supply. Aiming at the problem of mismatch in time constants between the IT subsystem and the cooling subsystem, a dual-time-scale control framework is introduced. On the fast time scale, the deep reinforcement learning algorithm DQN (Deep Q-Network) is used to make task scheduling decisions at the second or millisecond level; on the slow time scale, the multi-agent reinforcement learning algorithm QMIX (Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning) is used to achieve minute-level collaborative thermal management across multiple data centers. It can achieve a good balance among the operating cost OPEX, task residence time, and power usage effectiveness PUE, optimize the joint objectives of task scheduling and thermal management, effectively address the multi-objective optimization challenges faced in geographically distributed data centers, and significantly improve the energy efficiency and stability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0052] Figure 1 is the optimization process in the embodiments of the present invention;

[0053] Figure 2 is the dual-time-scale control framework in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the embodiments and the drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and do not limit the present invention.

[0055] As a possible implementation, as Figure 1As shown in the figure, this embodiment provides a joint optimization method based on task scheduling and thermal management, including the following specific steps:

[0056] Construct a dual-time-scale control framework, where the dual-time scale includes a slow time scale and a fast time scale;

[0057] Construct an operating expenditure model to obtain the operating cost of a geographically distributed data center for executing tasks;

[0058] Construct an energy usage efficiency model to obtain the average energy usage efficiency of a geographically distributed data center;

[0059] Obtain the task waiting time, and construct an optimization model with the goal of minimizing the operating cost of a geographically distributed data center for executing tasks, the average energy usage efficiency of a geographically distributed data center, and the task waiting time;

[0060] Based on the optimization model, generate optimal task scheduling decisions and multi-data center collaborative thermal management decisions, including:

[0061] On the fast time scale, construct a task scheduling agent and generate task scheduling decisions;

[0062] On the slow time scale, construct a thermal management agent and generate multi-data center collaborative thermal management decisions;

[0063] Based on state space sharing, co-train the task scheduling agent and the thermal management agent;

[0064] Based on the trained task scheduling agent and thermal management agent, generate optimal task scheduling decisions and multi-data center collaborative thermal management decisions.

[0065] In this embodiment, by combining compute-intensive task scheduling and thermal management, the matching of cooling demand and supply can be dynamically adjusted. For the problem of time constant mismatch between the IT subsystem and the cooling subsystem, a dual-time-scale control framework is introduced. On the fast time scale, a deep reinforcement learning algorithm is used to make task scheduling decisions at the second or millisecond level; on the slow time scale, a multi-agent reinforcement learning algorithm is used to achieve minute-level collaborative thermal management across multiple data centers. It can achieve a good balance among operating cost, task residence time, and energy usage efficiency, optimize the joint goal of task scheduling and thermal management, effectively address the multi-objective optimization challenges faced in geographically distributed data centers, and significantly improve the energy efficiency and stability of the system.

[0066] In some possible embodiments, as Figure 2 shown, the specific steps for constructing a dual-time-scale control framework include establishing a slow time scale and a fast time scale:

[0067] Each working day is divided into N slow time steps n, where n ∈ {1, 2,..., N}, and the duration of each slow time slot is l;

[0068] Each slow time step is divided into M fast time steps m, where m ∈ {1, 2,..., M}, and the duration of each fast time slot is d.

[0069] In some possible embodiments, an operating expenditure model is constructed to obtain the operating costs of a geographically distributed data center performing tasks, specifically including:

[0070] Since there are differences in grid electricity prices and renewable energy prices in different regions, we define a comprehensive energy price to evaluate the operating costs (OPEX) of performing tasks in geographically dispersed data centers. Here, the comprehensive energy price of data center c at time slot τ is defined as follows:

[0071]

[0072] Among them, and represent the electricity prices of the grid, solar energy, and wind energy respectively. represent the consumption proportions of the grid, solar energy, and wind energy respectively, and their definitions are as follows:

[0073]

[0074]

[0075]

[0076] Among them, represents the total energy consumption of the IT subsystem of data center c at time τ, represents the total energy consumption of the cooling subsystem of data center c at time τ;

[0077] and represent the renewable energy available to data center c in the fast time slot τ:

[0078]

[0079]

[0080] Among them, κ c represents the efficiency of converting solar energy into electrical energy, A c [m 2 represents the effective irradiation area of the solar panels, I c [W·m -2 represents the solar irradiance, η cDenotes the efficiency of converting wind energy into electrical energy, ζ c [m 2 represents the rotor area of the wind turbine, ρ air [kg·m -3 . and q c [m·s -1 represent the air density and wind speed respectively, and d represents the duration of the time slot τ.

[0081] Therefore, the operating cost of the IT subsystem and the cooling subsystem of the data center c at time τ is:

[0082]

[0083]

[0084] In some possible embodiments, an energy usage efficiency model is constructed to obtain the average energy usage efficiency of a geographically distributed data center, which specifically includes: obtaining the cooling energy consumption of the data center at time τ and the total energy consumption of the center at time τ; constructing an energy usage efficiency model based on a double time scale; and obtaining the average energy usage efficiency of the geographically distributed data center based on the energy usage efficiency model. The specific calculation steps include:

[0085]

[0086] The average PUE of C geographically distributed data centers can be calculated as:

[0087]

[0088] where PUE represents the energy utilization efficiency, M represents the number of fast time steps, N represents the number of slow time steps, and C represents the number of data centers.

[0089] In some possible embodiments, constructing an optimization model specifically includes:

[0090] Dynamically obtaining the air flow rate of the computer room air conditioning unit of the data center, the operating cost of the IT subsystem of the data center at time τ, the operating cost of the cooling system of the data center at time τ, the average energy usage efficiency of the geographically distributed data center, the operating cost of the data center, and the penalty term for data center overheating;

[0091] Taking the minimum operating cost of the geographically distributed data center for executing tasks, the average energy usage efficiency of the geographically distributed data center, and the task waiting time as the goals, constructing an objective optimization function;

[0092] Solving the objective optimization function to obtain the optimization model.

[0093] In some possible embodiments, the objective optimization function is:

[0094]

[0095] Among them, μ1 - μ5 represent weight factors, It is represented that Γ IT and Γ cooling represent the operating costs of the IT subsystem and the cooling subsystem of data center c at time τ, and W s represents the task waiting time. W g = eg - bg represents the sojourn time of task g, which is obtained from the submission time of task g and the time when task g starts to execute. PUE represents the energy utilization efficiency, and Ξ represents the penalty term for overheating of the data center, and its definition is as follows: Among them, T c (τ) represents the temperature of data center c at time τ, and ψ T represents the safety temperature threshold of the data center.

[0096] In some possible embodiments, on a fast time scale, a task scheduling agent is constructed and a task scheduling decision is generated, specifically including:

[0097] Construct a task scheduling agent based on Markov decision-making;

[0098] Define the state space as the state vector of the task scheduling agent of the geographically distributed data center;

[0099] Define the action space as the selection of candidate tasks assigned to the data center;

[0100] Define a reward function based on the operating cost of the IT subsystem of the data center at time τ, the operating cost of the cooling system of the data center at time τ, and the task waiting time;

[0101] According to the state space, action space, and reward function, train the task scheduling agent through a reinforcement learning algorithm to generate a task scheduling decision.

[0102] In some possible embodiments, on a slow time scale, a thermal management agent is constructed and a multi - data - center collaborative thermal management decision is generated, specifically including:

[0103] Construct a thermal management agent based on Markov decision-making;

[0104] Define the state space as the union of all states of the geographically distributed data center;

[0105] Define the action space as the continuous control action of adjusting the air flow rate of the computer room air conditioning unit;

[0106] Define a reward function based on the average temperature of the data center at time τ, temperature control, and energy use efficiency;

[0107] Based on the state space, action space, and reward function, a thermal management agent is trained through a reinforcement learning algorithm to generate multi-data center collaborative thermal management decisions.

[0108] In some possible embodiments, collaborative training of the task scheduling agent and the thermal management agent is performed based on state space sharing, specifically including:

[0109] Simultaneously and centrally train the task scheduling agent and the thermal management agent. The training process is carried out for T epochs. Each epoch contains N slow time steps, and each slow time step is divided into M fast time steps;

[0110] At the beginning of each slow time step, initialize the parameters of the task scheduling agent and the thermal management agent, including: the parameters of the task scheduling agent, and the parameters of the target network are copied from the online network; the parameters of the thermal management agent, policy network, and hybrid network, and the parameters of the target network are also copied from the online network; the exploration noise of the task scheduling agent and the exploration noise of the thermal management agent;

[0111] The task scheduling agent selects an action based on the observed environmental state, obtains a reward and then transfers to the next state, and at the same time stores the experience in the experience replay buffer of the task scheduling agent;

[0112] At the end of each slow time step, the thermal management agent obtains local observations, selects an action to explore and obtains a reward, then transfers to the next state, and at the same time stores the experience in the experience replay buffer of the thermal management agent;

[0113] Alternately sample the experience replay buffers of the task scheduling agent and the thermal management agent to update the network parameters of the task scheduling agent and the thermal management agent.

[0114] In some possible embodiments, the specific calculation process for constructing the task scheduling agent and the thermal management agent includes:

[0115] Step 1: Fast time scale control based on the DQN algorithm:

[0116] Construct a task scheduling agent based on Markov decision-making;

[0117] Define the state vector of the geographically distributed data center as

[0118] Action space: The task scheduler needs to assign a candidate task to an appropriate data center k. Therefore, the corresponding action can be expressed as k ∈ {1, 2,..., |C|};

[0119] Reward function: The reward function under the fast time scale is defined as: R fast= ξ1 - μ1Γ IT - μ2Γ cooling - μ3W s ;

[0120] where s q = <u g , q len > represents the state of the arrival queue, u g represents the number of processors required for the current candidate task, and q len is the length of the current arrival queue; represents the state of data center c, where } represents the utilization rate of the servers in data center c, and u c represents the total number of available processors in data center c, represents the operating cost of data center c, f c represents the set of air flow rate settings of all CRAC units in data center c, T c represents the temperature of c; ξ1 is a positive value to ensure that the reward is positive. μ1, μ2, and μ3 are weight factors.

[0121] Step 2: Slow Time Scale Control Based on the QMIX Algorithm

[0122] This section models the thermal management problem of geographically distributed data centers as a Markov game Each data center is regarded as an agent, denoted as ε = {e1, e2,..., e |C|}. is the global state space, which is the same as the state space in Step 1. The observation profile is the joint observation of all agents All possible constitute The action profile is the joint action of all agents All possible constitute The transition to the new state S′ is determined by the transition probability p: All agents share a unified reward function R slow : This function is used to optimize the overall performance and encourage cooperative behavior. γ is the discount factor.

[0123] By defining its components as follows, the optimization problem is transformed into a distributed partially observable Markov decision process (Dec-POSMDP):

[0124] Observation space: The observation of agent e c (c ∈ C) is defined as where q len is the length of the arrival queue, and φ crepresents the utilization rate of servers in data center c, T c , is the average temperature inside data center c, is the set of air flow rate settings of all computer room air conditioning (CRAC) units in data center c during the previous cooling control decision cycle, represents the observation space of data center c. Then, a global observation vector O is formed by aggregating the observations of all agents in the system.

[0125] Action space: Each agent e c will take an action to adjust the air flow rate of the computer room air conditioning (CRAC) units in data center c.

[0126] Reward function: The global reward R slow is defined as:

[0127] where ξ2 is a positive constant to ensure that R slow remains positive. μ4 and μ5 are weight factors, and T c (τ) represents the average temperature of data center c during τ. ψ T represents the safety temperature threshold of all data centers.

[0128] As a possible implementation manner, this embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements a joint optimization method based on task scheduling and thermal management.

[0129] As a possible implementation manner, this embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a joint optimization method based on task scheduling and thermal management.

[0130] The specific implementation manners described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are only specific implementation manners of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A joint optimization method based on task scheduling and thermal management, characterized in that, It includes the following specific steps: Construct a dual-time-scale control framework, where the dual-time scale includes a slow time scale and a fast time scale; Construct an operating expenditure model to obtain the operating cost of a geographically distributed data center for executing tasks; Construct an energy usage efficiency model to obtain the average energy usage efficiency of a geographically distributed data center; Obtain the task waiting time, and construct an optimization model with the goal of minimizing the operating cost of a geographically distributed data center for executing tasks, the average energy usage efficiency of a geographically distributed data center, and the task waiting time; Based on the optimization model, generate optimal task scheduling decisions and multi-data center collaborative thermal management decisions, including: On the fast time scale, construct a task scheduling agent and generate task scheduling decisions; On the slow time scale, construct a thermal management agent and generate multi-data center collaborative thermal management decisions; Based on state space sharing, co-train the task scheduling agent and the thermal management agent; Based on the trained task scheduling agent and thermal management agent, generate optimal task scheduling decisions and multi-data center collaborative thermal management decisions.

2. The joint optimization method based on task scheduling and thermal management according to claim 1, wherein The specific steps of constructing the dual-time-scale control framework include establishing a slow time scale and a fast time scale: Divide each working day into N slow time steps n, where n ∈ {1, 2,..., N}; Divide each slow time step into M fast time steps m, where m ∈ {1, 2,..., M}.

3. The joint optimization method based on task scheduling and thermal management according to claim 1, characterized in that The specific steps of constructing the operating expenditure model to obtain the operating cost of a geographically distributed data center for executing tasks include: Obtain the electricity prices of the power grid, solar energy, and wind energy, obtain the consumption ratios of the power grid, solar energy, and wind energy, and construct an operating expenditure model; Based on the operating expenditure model, obtain the operating costs of the IT subsystem and the cooling subsystem of the data center at time τ.

4. The joint optimization method based on task scheduling and thermal management according to claim 1, wherein The specific steps of constructing the energy usage efficiency model to obtain the average energy usage efficiency of a geographically distributed data center include: Obtain the cooling energy consumption of the data center at time τ and the total energy consumption of the center at time τ; Based on the dual-time scale, construct an energy usage efficiency model; Based on the energy usage efficiency model, obtain the average energy usage efficiency of a geographically distributed data center.

5. The joint optimization method based on task scheduling and thermal management according to claim 1, characterized in that The specific steps of constructing the optimization model include: Dynamically obtain the air flow rate of the computer room air conditioner unit of the data center, the operating cost of the IT subsystem of the data center at time τ, the operating cost of the cooling system of the data center at time τ, the average energy usage efficiency of the geographically distributed data center, the operating cost of the data center, and the penalty term for data center overheating; With the goal of minimizing the operating cost of a geographically distributed data center for executing tasks, the average energy usage efficiency of a geographically distributed data center, and the task waiting time, construct an objective optimization function; Solve the objective optimization function to obtain the optimization model.

6. The joint optimization method based on task scheduling and thermal management according to claim 1, wherein The specific steps of constructing a task scheduling agent and generating task scheduling decisions on the fast time scale include: Construct a task scheduling agent based on Markov decision-making; Define the state space as the state vector of the task scheduling agent of the geographically distributed data center; Define the action space as the selection of assigning candidate tasks to data centers; Define a reward function based on the operating cost of the IT subsystem in the data center at time τ, the operating cost of the cooling system in the data center at time τ, and the task waiting time. According to the state space, action space, and reward function, train a task scheduling agent through a reinforcement learning algorithm to generate task scheduling decisions.

7. The joint optimization method based on task scheduling and thermal management according to claim 1, characterized in that On the slow time scale, construct a thermal management agent and generate a multi-data center collaborative thermal management decision, specifically including: Construct a thermal management agent based on Markov decision-making. Define the state space as the union of all states of the geographically distributed data centers. Define the action space as the continuous control action of adjusting the air flow rate of the computer room air conditioning unit. Define a reward function based on the average temperature, temperature control, and energy use efficiency of the data center at time τ. According to the state space, action space, and reward function, train the thermal management agent through a reinforcement learning algorithm to generate a multi-data center collaborative thermal management decision.

8. The joint optimization method based on task scheduling and thermal management according to claim 1, wherein The collaborative training of the task scheduling agent and the thermal management agent based on state space sharing specifically includes: At the beginning of each slow time step, initialize the parameters of the task scheduling agent and the thermal management agent. The task scheduling agent selects an action based on the observed environmental state, obtains a reward and then transfers to the next state, while storing the experience in the experience replay buffer of the task scheduling agent. At the end of each slow time step, the thermal management agent obtains local observations, selects an action to explore and obtains a reward, then transfers to the next state, while storing the experience in the experience replay buffer of the thermal management agent. Alternately sample the experience replay buffers of the task scheduling agent and the thermal management agent to update the network parameters of the task scheduling agent and the thermal management agent.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the joint optimization method based on task scheduling and thermal management according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the joint optimization method based on task scheduling and thermal management according to any one of claims 1 to 8.