A data center micro-grid electric calculation cooperative double-layer deep reinforcement learning scheduling method

By employing a two-layer deep reinforcement learning scheduling method for data center microgrids, the problem of coordinated scheduling of power and computing power in geographically distributed data center microgrids is solved. This method achieves safe, economical, and stable system operation and energy optimization, reduces the cost of non-renewable energy sources, and improves the system's convergence speed.

CN120147060BActive Publication Date: 2025-11-25HARBIN INST OF TECH

Patent Information

Application Number
CN202510208164.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-11-25
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

Geographically distributed data center microgrids struggle to achieve secure, economical, and stable coordinated scheduling of power and computing power when faced with complex computing tasks and the high uncertainty of renewable energy.

Method used

A two-layer deep reinforcement learning scheduling method for data center microgrids is adopted to construct a power computing power collaborative scheduling model and a two-layer multi-agent deep reinforcement learning framework. Combined with task allocation and reward mechanisms at the global and local layers, energy management and task scheduling are optimized.

Benefits of technology

It realizes the coordinated scheduling of power and computing power in the microgrid of geographically distributed data centers, improves the system's security, reliability and economy, reduces the cost of non-renewable energy by 15%-20%, and increases the convergence speed by 5 times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147060B_ABST
    Figure CN120147060B_ABST
Patent Text Reader

Abstract

The application discloses a data center micro-grid electric calculation cooperative double-layer deep reinforcement learning scheduling method and belongs to the technical field of power systems. Firstly, aiming at the high renewable energy penetration scene of the geographical distributed data center micro-grid, a double-layer multi-agent deep reinforcement learning framework is proposed. The double-layer multi-agent deep reinforcement learning framework containing a global layer and a local layer is constructed. The global layer is distributed by centralized deep reinforcement learning training to calculate the task, and the local layer is optimized by decentralized deep reinforcement learning to internally dispatch energy. The layered multi-agent double-delay deep deterministic policy gradient algorithm is combined to realize cross-layer interaction and rapid convergence. Through space-time task adjustment and energy storage cooperation, the application reduces the cost of non-renewable energy by 15%-20%, and the convergence speed is improved by 5 times compared with the prior art. Moreover, the application supports dynamic data backup, significantly improves the system reliability and economy while maintaining the service quality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of power systems, and particularly relates to a data center micro-grid electricity and computing collaborative double-layer deep reinforcement learning scheduling method. BACKGROUND

[0002] The complexity of the geographical distributed data center micro-grid, together with the complex nature of computing tasks and high uncertainty related to renewable energy generation, requires advanced real-time and effective power and computing scheduling methods. Ensuring the safety, economy and stable operation of the system while maintaining the quality of service has become a key challenge in the power and computing management of the geographical distributed data center micro-grid. Therefore, it is of great significance to develop a power and computing collaborative scheduling method with time and space regulation capability for the data center micro-grid. SUMMARY

[0003] In view of the above problems, the application provides a data center micro-grid electricity and computing collaborative double-layer deep reinforcement learning scheduling method, which is suitable for real-time scheduling and cost optimization in a high renewable energy penetration scenario, realizes the power and computing collaborative scheduling of the geographical distributed data center micro-grid, and ensures the safety, economy and stable operation of the system while maintaining the quality of service.

[0004] The technical scheme adopted by the application is as follows:

[0005] A data center micro-grid electricity and computing collaborative double-layer deep reinforcement learning scheduling method comprises the following steps:

[0006] S1: constructing a power and computing collaborative scheduling model of the geographical distributed data center micro-grid, task flexibility requirements and an integrated data protection mechanism,

[0007] The power and computing collaborative scheduling model comprises a data center energy consumption model, an energy storage system model, a computing task model and a data center backup data energy consumption model, wherein the constraints of the energy storage system model include upper and lower constraints of charge and discharge power and upper and lower constraints of capacity;

[0008] The task flexibility requirement is that the real-time processing of computing tasks by each data center cannot exceed the maximum computing task amount that can be processed by all servers of the data center;

[0009] The data protection mechanism is that the data center micro-grid saves data locally to prevent data transmission errors;

[0010] S2: constructing a double-layer multi-agent deep reinforcement learning framework comprising a global layer and a local layer:

[0011] The global layer adopts a centralized deep reinforcement learning algorithm to calculate tasks based on renewable energy prediction and real-time task demand distribution, the centralized deep reinforcement learning algorithm is a deep reinforcement learning algorithm for the global layer to adopt an agent to distribute computing tasks arriving at the global layer in real time according to the new energy prediction output of the geographic distributed data center microgrid in the next two hours, the number of active servers and real-time residual computing task quantity information;

[0012] The local layer adopts a decentralized deep reinforcement learning algorithm to optimize energy management and task scheduling within each data center microgrid while meeting quality of service requirements, the decentralized reinforcement learning algorithm is an algorithm for each geographic distributed data center microgrid to adopt an agent to schedule real-time arriving computing tasks according to new energy output and the number of active servers in the next two hours;

[0013] S3: interact between the global layer and the local layer, the global layer distributes computing tasks to each local layer, and the local layer feeds back the unfinished computing tasks to the global layer to affect the global layer distribution;

[0014] S4: guide the global layer to distribute computing tasks to the geographic distributed data center microgrid through a reward function, ensure that the local layer focuses on maximizing the use of local renewable energy within a day, and the local layer minimizes non-renewable energy cost, thereby realizing power and computing power collaborative optimization scheduling of the geographic distributed data center microgrid.

[0015] Further, in step S1, the data center energy consumption model is shown in formulas (1)-(6), the energy storage system model is shown in formulas (7)-(11), and the computing task model is shown in formulas (12)-(20), including batch processing computing task model and interactive computing task model, wherein formulas (12)-(16) are batch processing computing task model, wherein formulas (15) and (16) are batch processing computing task model constraints, indicating that batch processing computing tasks are completed within the maximum tolerance time and satisfy the non-negative constraint at any time; formulas (17)-(21) are interactive computing task model, wherein formulas (18) and (19) are interactive processing computing task model constraints, indicating that interactive computing tasks are completed within the maximum tolerance time and satisfy the non-negative constraint at any time; the data center backup data energy consumption model is shown in formula (21),

[0016] P i,DC (t)=x i (t)[P i,idle +(PUE i -1)P i,peak ]+x i (t)(P i,peak -P i,idle )γi (t)+ξ i (1)

[0017] ξ i ≤P i,DC (t)≤M i PUE i ·P i,peak +ξ i (2)

[0018]

[0019] 0≤x i (t)≤M i (6)

[0020] B i (t+1)=B i (t)+(u i,ch P i,ch (t)η i,ch -u i,dis P i,dis (t) / η i,dis )Δt (7)

[0021] u i,ch +u i,dis =1,u i,ch ≥0,u i,dis ≥0 (8)

[0022]

[0023] 0≤P i,dis (t)≤ p i,dis (10)

[0024]

[0025] In the formula: P i,DC (t) represents the energy consumption of data center i at time t; x i (t) represents the number of active servers in data center i at time t; PUE i Energy efficiency of data center i; P i,idle P i,peak These represent the energy consumption of server i when idle and when fully loaded, respectively; γ i (t) represents the average utilization rate of servers in data center i at time t; ξ i Indicates the energy consumption of data center infrastructure; M i ω represents the maximum number of servers in data center i; i (t) represents the amount of batch processing computation tasks that data center i needs to process at time t; is the number of batch computing tasks arriving at data center i at time t; is the number of batch computing tasks arriving at data center i at time t; is the maximum tolerable time of batch computing tasks arriving at data center i at time t; is the maximum tolerable time of batch computing tasks arriving at data center i at time t; i (t) is the capacity of the energy storage system included in data center microgrid i at time t; is the capacity of the energy storage system included in data center microgrid i at time t; B i is the capacity of the energy storage system included in data center microgrid i at time t; i,ch (t) is the capacity of the energy storage system included in data center microgrid i at time t; i,dis (t) is the charge and discharge power of the energy storage system included in data center microgrid i at time t; is the charge and discharge power of the energy storage system included in data center microgrid i at time t; p i,dis is the charge and discharge power of the energy storage system included in data center microgrid i at time t; i,ch is the charge and discharge power of the energy storage system included in data center microgrid i at time t; i,dis is the charge and discharge power of the energy storage system included in data center microgrid i at time t; i,ch is the charge and discharge power of the energy storage system included in data center microgrid i at time t; i,dis is the charge and discharge power of the energy storage system included in data center microgrid i at time t; j is the arrival time set of the jth type of batch computing task type; is the number of batch computing tasks arriving at data center i at time t; jm is the number of batch computing tasks arriving at data center i at time t; is the number of initial batch computing tasks arriving at data center i at time t0; is the number of initial interactive computing tasks arriving at data center i at time t; i (t) is the number of interactive computing tasks processed by data center i at time t; is the number of batch computing tasks arriving at data center i at time t; is the adjustable amount of the number of batch computing tasks arriving at data center i at time t; i (t) is the adjustable amount of the number of interactive computing tasks processed by data center i at time t; is the energy consumption of data center i backing up data at time t; is the number of data backed up by data center i; i is the energy consumption coefficient of data backed up by data center i; i (t u0 | t u0-1 ) is the number of batch computing tasks arriving at data center microgrid i at time t; u0-1 is the number of batch computing tasks arriving at data center microgrid i at time t; u0 is the number of batch computing tasks arriving at data center microgrid i at time t.

[0026] Further, in step S2, a two-layer multi-agent deep reinforcement learning framework containing a global layer and a local layer is constructed, as shown in equations (22)-(25),

[0027]

[0028] wherein s global represents the global layer state space; i=1,...,S represents the state space of the data center microgrid i; P i,green (t), i=1,...,S is the renewable energy output of the data center microgrid i at time t;

[0029] R(ω i (t u0 |t u0-1 )), i=1,...,S is the batch processing computing task arriving at the data center microgrid i at time t u0-1 remaining at time t u0 for the data center microgrid i; u global represents the action space of the global layer; i=1,...,S represents the action space of the data center microgrid i; i=1,...,S is the batch processing computing task and interactive computing task allocated by the global layer to the data center microgrid i at time t; u0 i=1,...,S is the N i th non-renewable energy unit output of the data center microgrid i.

[0030] Further, in step S3, the global layer allocates computing tasks to each local layer, and the local layer feeds back the unfinished computing tasks to the global layer to affect the global layer allocation, as shown in equations (26)-(28),

[0031]

[0032] R(ω i (t1|t0))=0 (28)

[0033] wherein T l represents the scheduling time step of the local layer in one day; is an indicator function.

[0034] Further, in step S4, the global layer reward function and the local layer reward function are as shown in equations (29)-(37),

[0035]

[0036] P i,green (t) = P i,WT (t) + P​i,PV (t) (31)

[0037]

[0038]

[0039] wherein: r global and i = 1, …, S represent the reward functions of global layer and local layer respectively; P i,green (t) represents the renewable energy output of the geographic location of the data center microgrid i at time t; P i,WT (t) and P i,PV (t) represent the wind power output and photovoltaic output of the geographic location of the data center microgrid i at time t, respectively; and P i,WT represent the upper and lower limit constraints of the wind power output of the geographic location of the data center microgrid i, respectively; and P i,PV represent the upper and lower limit constraints of the photovoltaic output of the geographic location of the data center microgrid i, respectively; N i represents the non-renewable energy unit included in the data center microgrid i; represents the output of the l i th non-renewable energy unit of the data center microgrid i at time t; and represent the upper and lower limit constraints of the output of the l i th non-renewable energy unit of the data center microgrid i, respectively; and represent the increased or decreased output of the fuel unit per unit time, respectively;

[0040] wherein, formula (37) is the ramping constraint of the l i th non-renewable energy unit of the data center microgrid i.

[0041] Another object of the present application is achieved by a data center microgrid electric calculation collaborative double-layer deep reinforcement learning scheduling system, comprising: a memory, a processor and a computer program stored on the memory and executable on the processor, when the computer program is executed by the processor, a data center microgrid electric calculation collaborative double-layer deep reinforcement learning scheduling method as described above is realized.

[0042] The application has the advantages and beneficial effects: the application realizes power and computing power collaborative scheduling of geographical distributed data center microgrid, improves the security, reliability, economy and stability of the system while maintaining the quality of service, reduces the cost of non-renewable energy by 15%-20% through space-time task adjustment and energy storage cooperation, improves the convergence speed by 5 times compared with the prior art, and supports dynamic data backup. BRIEF DESCRIPTION OF DRAWINGS

[0043] Figure 1 is a flowchart of the application;

[0044] Figure 2 is a convergence comparison chart; (a) is a convergence chart trained using the existing scheduling method, and (b) is a convergence chart trained using the scheduling method of the application;

[0045] Figure 3 is a cost comparison chart of power and computing power scheduling of data center microgrid. DETAILED DESCRIPTION

[0046] In order to make the purpose, technical scheme and advantages of the application more clear and obvious, the application will be further described in detail below in combination with embodiments, and it should be understood that the specific embodiments described here are only used to explain the application, and are not used to limit the application.

[0047] Embodiment 1:

[0048] As shown in Figure 1 , a data center microgrid electric and computing power collaborative double-layer deep reinforcement learning scheduling method comprises the following steps:

[0049] S1: constructing an electric power and computing power collaborative scheduling model of a geographical distributed data center microgrid, task flexibility requirements and an integrated data protection mechanism,

[0050] The electric power and computing power collaborative scheduling model comprises a data center energy consumption model, an energy storage system model, a computing task model and a data center backup data energy consumption model, wherein the constraints of the energy storage system model include upper and lower constraints of charge and discharge power and upper and lower constraints of its capacity;

[0051] The task flexibility requirement is that the real-time processing of computing tasks by each data center cannot exceed the maximum computing task amount that can be processed by all servers of the data center;

[0052] The data protection mechanism is that the data center microgrid saves data locally to prevent data transmission errors;

[0053] The data center energy consumption model is shown as formulas (1)-(6), the energy storage system model is shown as formulas (7)-(11), and the computing task model is shown as formulas (12)-(20), including a batch computing task model and an interactive computing task model, wherein formulas (12)-(16) are the batch computing task model, wherein formulas (15) and (16) are constraints of the batch computing task model, indicating that the batch computing task is completed within a maximum tolerance time and satisfies a non-negative constraint at any time; formulas (17)-(21) are the interactive computing task model, wherein formulas (18) and (19) are constraints of the interactive computing task model, indicating that the interactive computing task is completed within a maximum tolerance time and satisfies a non-negative constraint at any time; the data center backup data energy consumption model is shown as formula (21),

[0054] P i,DC (t)=x i (t)[P i,idle +(PUE i -1)P i,peak ]+x i (t)(P i,peak -P i,idle )γ i (t)+ξ i (1)

[0055] ξ i ≤P i,DC (t)≤M i ·PUE i ·P i,peak +ξ i (2)

[0056]

[0057] 0≤x i (t)≤M i (6)

[0058] B i (t+1)=B i (t)+(u i,ch P i,ch (t)η i,ch -u i,dis P i,dis (t) / η i,dis )Δt (7)

[0059] u i,ch +u i,dis =1,u i,ch ≥0,u i,dis ≥0 (8)

[0060]

[0061] 0≤P i,dis (t)≤ p i,dis (10)

[0062]

[0063] In the formula: P i,DC (t) represents the energy consumption of data center i at time t; x i (t) represents the number of active servers in data center i at time t; PUE i Energy efficiency of data center i; P i,idle P i,peak These represent the energy consumption of server i when idle and when fully loaded, respectively; γ i (t) represents the average utilization rate of servers in data center i at time t; ξ i Indicates the energy consumption of data center infrastructure; M i ω represents the maximum number of servers in data center i; i (t) represents the amount of batch processing computation tasks that data center i needs to process at time t; For data center i to process at time t Batch processing tasks that arrive at a specific time; for The maximum tolerable time for batch processing tasks arriving at a given time; B i (t) represents the capacity of the energy storage system included in the data center microgrid i at time t; as well as B i P represents the upper and lower capacity limits of the energy storage system included in the data center microgrid i; i,ch (t) and P i,dis (t) represents the charging and discharging power of the energy storage system included in the data center microgrid i at time t; as well as p i,dis This represents the upper and lower limits of the charging and discharging power of the energy storage system included in the data center microgrid i; η i,ch and η i,dis This represents the charging and discharging efficiency of the energy storage system included in the data center microgrid i; u i,ch and u i,dis Δt represents the charging and discharging binary variables of the energy storage system included in the data center microgrid i; Ω represents the time interval; J represents the batch processing task set; τ represents the batch processing task type; Δt represents the time interval; Ω represents the set of batch processing tasks; J represents the batch processing task type; τ represents the time interval; Δt ... j For the j-th batch processing task type, calculate the arrival time set; Data center i processing Batch processing tasks that arrive at a specific time; S represents the set of data center microgrids; is the initial batch processing computing task quantity arriving at the data center i at time t0; is the initial interactive computing task quantity arriving at the data center i at time t; λ i (t) is the interactive computing task quantity processed by the data center i at time t; is the initial batch processing computing task quantity arriving at the data center i at time t; is the adjustable quantity of the batch processing computing task quantity arriving at the data center i at time t; Δλ i (t) is the adjustable quantity of the interactive computing task processed by the data center i at time t; is the energy consumption of the data center i for backing up data at time t; is the data quantity of the data center i for backing up data; α i is the energy consumption coefficient of the data center i for backing up data; R(ω i (t u0 is the energy consumption of the data center i for backing up data at time t; u0-1 is the energy consumption of the data center i for backing up data at time t; u0-1 is the batch processing computing task quantity arriving at the data center microgrid i at time t; u0 is the batch processing computing task quantity remaining in the data center microgrid i at time t.

[0064] S2: constructing a two-layer multi-agent deep reinforcement learning framework containing a global layer and a local layer:

[0065] The global layer adopts a centralized deep reinforcement learning algorithm to allocate computing tasks based on renewable energy prediction and real-time task demand. The centralized deep reinforcement learning algorithm is a deep reinforcement learning algorithm in which one agent allocates computing tasks arriving at the global layer in real time according to the new energy prediction output, the number of active servers, and the real-time remaining computing task quantity information of the geographical distributed data center microgrids in the next two hours;

[0066] The local layer adopts a decentralized deep reinforcement learning algorithm to optimize the internal energy management and task scheduling of each data center microgrid under the requirement of service quality. The decentralized reinforcement learning algorithm is an algorithm in which each geographical distributed data center microgrid adopts one agent to schedule computing tasks arriving in real time according to the new energy output and the number of active servers;

[0067] wherein the two-layer multi-agent deep reinforcement learning framework containing the global layer and the local layer is constructed as shown in formulas (22)-(25),

[0068]

[0069] In the formula, s global represents the global layer state space; i=1,...,S represents the state space of the data center microgrid i; P i,green(t), i = 1, ..., S represents the renewable energy output of the data center microgrid i at time t;

[0070] R(ω i (t u0 |t u0-1 ), i = 1, ..., S is t u0-1 Time to arrive at the data center microgrid i in t u0 The remaining batch processing tasks at that time; u global Represents the action space of the global layer; i = 1, ..., S represents the action space of data center microgrid i; i = 1, ..., S represents the global layer at time t u0 Time is allocated to batch computing tasks and interactive computing tasks of data center microgrid i; i = 1, ..., S is the Nth node of the data center microgrid i. i The power output of Taiwan's non-renewable energy units.

[0071] S3: The global layer interacts with the local layers. The global layer allocates computational tasks to each local layer, and the local layers feed back unfinished computational tasks to the global layer to affect the global layer's allocation. As shown in formulas (26)-(28), the global layer allocates computational tasks to each local layer, and the local layers feed back unfinished computational tasks to the global layer to affect the global layer's allocation.

[0072]

[0073] R(ω i (t1∣t0))=0 (28)

[0074] In the formula: T l This indicates the scheduling time step for a local layer within one day. This is an indicator function.

[0075] S4: The reward function guides the global layer to allocate computing tasks to the geographic distributed data center microgrid, ensuring that the local layer focuses on maximizing the use of local renewable energy within a day and minimizing the cost of non-renewable energy, thereby achieving the collaborative optimization scheduling of power computing power in the geographic distributed data center microgrid; where the reward function of the global layer and the reward function of the local layer are shown in formulas (29)-(37).

[0076]

[0077] P i,green (t)=P i,WT (t)+P i,PV (t) (31)

[0078]

[0079] In the formula: r global And i = 1, …, S respectively represent the reward function of the global layer and the local layer; P i,green (t) represents the renewable energy output of the geographic location t of the data center microgrid i at the moment; P i,WT (t) and P i,PV (t) respectively represent the wind power output and photovoltaic output of the geographic location t of the data center microgrid i at the moment; And P i,WT respectively represent the upper and lower limit constraints of the wind power output of the geographic location of the data center microgrid i; And P i,PV respectively represent the upper and lower limit constraints of the photovoltaic output of the geographic location of the data center microgrid i; N i represent the non-renewable energy units contained in the data center microgrid i; represent the l i th non-renewable energy unit output of the data center microgrid i at t And respectively represent the upper and lower limit constraints of the l i th non-renewable energy unit output of the data center microgrid i; And respectively represent the increased or decreased output of the fuel unit per unit time;

[0080] Wherein, formula (37) is the ramping constraint of the l i th non-renewable energy unit of the data center microgrid i; After the above steps, the power and computing power collaborative optimization scheduling of the geographic distributed data center microgrid can be realized.

[0081] Embodiment 2

[0082] As Figure 2 shown, Global level represents the global layer, and Local level represents the local layer; DC1, DC2, and DC3 represent three data center microgrids. As can be seen, Figure 2 (a) At least 2000 iterations are required to reach convergence using the existing algorithm; Figure 2 (b) The algorithm used by the present application can reach convergence in only 500 iterations. As can be seen, the method used by the present application is 5 times faster than the prior art.

[0083] As Figure 3As shown, Random represents that no optimization algorithm is used for the geographic distributed data center micro-grid power and computing power collaborative scheduling, Bi-TD3 represents that the prior art is used for the geographic distributed data center micro-grid power and computing power collaborative scheduling algorithm; Proposed represents that the application is used for the geographic distributed data center micro-grid power and computing power collaborative scheduling algorithm; the cost comparison of the power and computing power scheduling of the data center micro-grid by using the optimization scheduling method, the prior method and the method proposed in the application can show that the method used in the application significantly reduces the non-renewable energy cost by 15%-20% compared with the method without using the optimization scheduling method, and has a certain degree of superiority compared with the prior method.

Claims

1. A data center microgrid electric algorithm collaborative double-layer deep reinforcement learning scheduling method, characterized in that, The method comprises the following steps: S1: constructing a power-computing collaborative scheduling model of a geographic distributed data center microgrid, task flexibility requirements and an integrated data protection mechanism, The power-computing collaborative scheduling model comprises a data center energy consumption model, an energy storage system model, a computing task model and a data center backup data energy consumption model, wherein the constraints of the energy storage system model comprise upper and lower constraints of charging and discharging power and upper and lower constraints of capacity; The task flexibility requirement is that the real-time processing of computing tasks by each data center cannot exceed the maximum computing task amount that can be processed by all servers of the data center; The data protection mechanism is that the data center microgrid locally saves data to prevent data transmission errors; S2: constructing a two-layer multi-agent deep reinforcement learning framework comprising a global layer and a local layer: The global layer adopts a centralized deep reinforcement learning algorithm, and the computing tasks are distributed based on renewable energy prediction and real-time task demand, wherein the centralized deep reinforcement learning algorithm is a deep reinforcement learning algorithm in which one agent of the global layer distributes the computing tasks that arrive at the global layer in real time according to the new energy output of the geographic distributed data center microgrid in the next two hours, the number of active servers and the real-time residual computing task amount information; The local layer adopts a decentralized deep reinforcement learning algorithm, and optimizes the internal energy management and task scheduling of each data center microgrid under the requirement of service quality, wherein the decentralized reinforcement learning algorithm is an algorithm in which one agent of each geographic distributed data center microgrid schedules the computing tasks that arrive in real time according to the new energy output and the number of active servers in the next two hours; S3: interacting between the global layer and the local layer, the global layer distributing computing tasks to each local layer, and the local layer feeding back the unfinished computing tasks to the global layer to affect the distribution of the global layer; S4: guiding the global layer to distribute computing tasks to the geographic distributed data center microgrid through a reward function, ensuring that the local layer maximizes the use of local renewable energy within a day, and the local layer minimizes the cost of non-renewable energy, thereby realizing the collaborative optimization and scheduling of the power and computing power of the geographic distributed data center microgrid.

2. The data center microgrid electric calculation cooperative double-layer deep reinforcement learning scheduling method according to claim 1, characterized in that, In step S1, the data center energy consumption model is shown in formulas (1)-(6), the energy storage system model is shown in formulas (7)-(11), and the computing task model is shown in formulas (12)-(20), comprising a batch processing computing task model and an interactive computing task model, wherein formulas (12)-(16) are the batch processing computing task model, formulas (15) and (16) are constraints of the batch processing computing task model, indicating that the batch processing computing task is completed within the maximum tolerance time and satisfies the non-negative constraint at any time; formulas (17)-(21) are the interactive computing task model, wherein formulas (18) and (19) are constraints of the interactive processing computing task model, indicating that the interactive computing task is completed within the maximum tolerance time and satisfies the non-negative constraint at any time; and the data center backup data energy consumption model is shown in formula (21). P i,DC (t) = x i (t) [P i,idle + (PUE i - 1) P i,peak ] + x i (t) (P i,peak - P i,idle ) γ i (t) + ξ i (1) ξ i ≤ P i,DC (t) ≤ M i · PUE i · P i,peak + ξ i (2) 0 < x i (t) < M i (6) B i (t+1) = B i (t) + u i,ch P i,ch (t) η i,ch -u i,dis P i,dis (t) / η i,dis ) Δt (7) u i,ch +u i,dis =1,u i,ch ≥0,u i,dis ≥0 (8) 0 < P i,dis (t) ≤ p i,dis (10) In the formula: P i,DC (t) represents the energy consumption of data center i at time t; x i (t) represents the number of active servers in data center i at time t; PUE i Energy efficiency of data center i; P i,idle P i,peak These represent the energy consumption of server i when idle and when fully loaded, respectively; γ i (t) represents the average utilization rate of servers in data center i at time t; ξ i Indicates the energy consumption of data center infrastructure; M i ω represents the maximum number of servers in data center i; i (t) represents the amount of batch processing computation tasks that data center i needs to process at time t; For data center i to process at time t Batch processing tasks that arrive at a specific time; for The maximum tolerable time for batch processing tasks arriving at a given time; B i (t) represents the capacity of the energy storage system included in the data center microgrid i at time t; as well as B i P represents the upper and lower capacity limits of the energy storage system included in the data center microgrid i; i,ch (t) and P i,dis (t) represents the charging and discharging power of the energy storage system included in the data center microgrid i at time t; as well as p i,dis This represents the upper and lower limits of the charging and discharging power of the energy storage system included in the data center microgrid i; η i,ch and η i,dis This represents the charging and discharging efficiency of the energy storage system included in the data center microgrid i; u i,ch and u i,dis Δt represents the charging and discharging binary variables of the energy storage system included in the data center microgrid i; Ω represents the time interval; J represents the batch processing task set; τ represents the batch processing task type; Δt represents the time interval; Ω represents the set of batch processing tasks; J represents the batch processing task type; τ represents the time interval; Δt ... j For the j-th batch processing task type, calculate the arrival time set; Data center i processing Batch processing tasks that arrive at a specific time; S represents the set of data center microgrids; The initial batch processing computation tasks arriving at data center i at time t0; Let λ be the initial number of interactive computing tasks arriving at data center i at time t; i (t) represents the amount of interactive computing tasks processed by data center i at time t; For data center i to process at time t The adjustable quantity of the number of batch processing tasks arriving at a given time; Δλ i (t) represents the adjustable amount of interactive computing tasks processed by data center i at time t; Energy consumption for data center i to back up data at time t; The number of data backups for data center i; α i Energy consumption coefficient for backing up data in data center i; R(ω) i (t u0 |t u0-1 )) is t u0-1 Time to arrive at the data center microgrid i in t u0 The remaining batch processing tasks at any given time.

3. The data center microgrid electric calculation cooperative double-layer deep reinforcement learning scheduling method according to claim 2, characterized in that, In step S2, a two-layer multi-agent deep reinforcement learning framework containing a global layer and a local layer is constructed, as shown in formulas (22)-(25), where s global represents the global layer state space; represents the state space of data center microgrid i; P i,green (t), i = 1,..., S is the renewable energy output of data center microgrid i at time t; R(ω i (t u0 |t u0-1 ), i = 1, ..., S is t u0-1 Time to arrive at the data center microgrid i in t u0 The remaining batch processing tasks at that time; u global Represents the action space of the global layer; Represents the action space of data center microgrid i; For the global layer at t u0 Time is allocated to batch computing tasks and interactive computing tasks of data center microgrid i; For the Nth data center microgrid i i The power output of Taiwan's non-renewable energy units.

4. The data center microgrid electric calculation cooperative double-layer deep reinforcement learning scheduling method according to claim 3, characterized in that, In step S3, the global layer allocates computing tasks to each local layer, and the local layer feeds back the unfinished computing tasks to the global layer to affect the allocation of the global layer, as shown in formulas (26)-(28), R(ω i (t1 | t0) = 0 (28) where: T l denotes the local layer one day scheduling time step; is an indicator function.

5. The data center microgrid electric algorithm collaborative double-deep reinforcement learning scheduling method according to claim 4, characterized in that, In step S4, the global layer reward function and the local layer reward function are as shown in formulas (29)-(37), P i,green (t) = P i,WT (t) + P i,PV (t) (31) where r global and r represent the reward functions of global and local layers, respectively; P i,green (t) represents the renewable energy output of data center microgrid i at geographical location t; P i,WT (t) and P i,PV (t) represent the wind power output and photovoltaic power output of data center microgrid i at geographical location t, respectively; and r P i,WT represent the upper and lower limits of wind power output of data center microgrid i at geographical location, respectively; and r P i,PV represent the upper and lower limits of photovoltaic power output of data center microgrid i at geographical location, respectively; N i represents the non-renewable energy units contained in data center microgrid i; represents the l i th non-renewable energy unit output of data center microgrid i at t; and r represent the upper and lower limits of the l i th non-renewable energy unit output of data center microgrid i, respectively; and r represent the increased or decreased output of fuel units per unit time, respectively; wherein formula (37) is the lth objective function of the data center microgrid i i Non-renewable energy unit climbing constraint.

6. A data center microgrid electric algorithm collaborative double-deep reinforcement learning scheduling system, comprising: The memory, the processor and the computer program stored on the memory and executable on the processor, wherein the computer program is executed by the processor to implement the data center micro-grid electric computer collaborative two-layer deep reinforcement learning scheduling method of any one of claims 1-5.

Citation Information

Patent Citations

  • Active power distribution network distributed power supply collaborative optimization operation method and system

    CN116896112A

  • Multi-mobile emergency power supply toughness optimization scheduling method based on data driving

    CN118449131A

Cited By

  • Power computing reinforcement learning scheduling method considering energy storage life

    CN122660106A