Multi-region energy internet optimization scheduling method and device

By building a multi-agent reinforcement learning framework for multi-energy park system, and combining multi-agent deep deterministic strategy gradient algorithm and Dropout method, the problem of energy scheduling overfitting is solved, and more efficient and stable energy Internet optimization scheduling is achieved.

CN120031272APending Publication Date: 2025-05-23STATE GRID HEBEI ELECTRIC POWER CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411841791.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When the prior art uses multi-agent deep reinforcement learning algorithm for energy Internet optimization scheduling, the problem of energy scheduling overfitting is prone to occur.

Method used

Build a multi-agent reinforcement learning framework for multi-energy park system, combine multi-agent depth deterministic strategy gradient algorithm and Dropout method to build an energy scheduling model, and train the model through input and output of state space and action space.

Benefits of technology

It enhances the adaptability and robustness of the system, improves the generalization ability of the model, and ensures that the energy scheduling model can maintain high scheduling accuracy and stability when facing new or unknown scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031272A_ABST
    Figure CN120031272A_ABST
Patent Text Reader

Abstract

The invention provides a multi-region energy internet optimization scheduling method and device, and relates to the technical field of energy internet optimization scheduling. The method comprises the steps that a multi-agent reinforcement learning framework of the multi-energy park system is constructed, the multi-agent reinforcement learning framework comprises a state space and an action space, the state space comprises the state of each agent at each time point, the action space comprises the action of each agent at each time point, and the agent is a control center of the multi-energy park; building an energy scheduling model based on a multi-agent depth deterministic strategy gradient algorithm and a Dropout method; training an energy scheduling model by taking the state space as input and the action space as output; and determining an energy scheduling scheme of the multi-energy park system based on the trained energy scheduling model. According to the invention, the adaptability and robustness of the system can be enhanced, and the scheduling efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of energy internet optimization and scheduling, and in particular to a multi-region energy internet optimization and scheduling method and device. Background Art

[0002] With the global emphasis on environmental protection and sustainable development, the use of renewable energy (such as solar energy, wind energy, etc.) has become an important direction of energy transformation. Regional energy internet promotes the green transformation of energy structure by integrating and optimizing renewable energy resources and realizing their efficient and stable access and utilization. In particular, energy synergy between different regions and the interactive use of energy resources are crucial. However, issues such as energy supply and demand balance, energy cost minimization, and the stability of multi-agent energy systems have posed many challenges to the research in the field of multi-regional energy scheduling.

[0003] In recent years, some studies have adopted model-based methods to solve the multi-region energy scheduling problem, including meta-heuristic optimization algorithms such as random mixed integer linear programming, genetic algorithms, and solutions such as the alternating direction multiplier method. Although these methods can achieve optimal decisions, they rely heavily on accurate modeling or prediction. With the development of artificial intelligence, many scholars have carried out research based on model-free methods. Deep reinforcement learning (DRL) combines the perception ability of deep learning with the decision-making ability of reinforcement learning. It is widely used in the energy field. However, most previous studies only use single-agent deep reinforcement learning algorithms to solve the regional energy Internet scheduling problem, which is not suitable for multi-region coordinated optimization. The multi-agent deep reinforcement learning (MADRL) framework, as an extension of DRL, is used to solve the energy scheduling problem. The multi-agent deep reinforcement learning algorithm considers the interaction and learning of multiple agents in the same environment, and has stronger adaptability and robustness. At present, the application of multi-agent deep reinforcement learning algorithms in the optimal scheduling of energy Internet is in its infancy, and there are still many challenges and problems, such as the overfitting problem faced by existing multi-agent reinforcement learning due to environmental non-stationarity and partial observation problems, as well as the privacy issue of agent information interaction. Summary of the invention

[0004] The present application provides a multi-region energy Internet optimization scheduling method and device to solve the problem of overfitting of energy scheduling caused by the complexity of the comprehensive energy environment in the prior art when using a multi-agent deep reinforcement learning algorithm to optimize the scheduling of the energy Internet.

[0005] In a first aspect, the present application provides a multi-region energy internet optimization scheduling method, comprising:

[0006] Construct a multi-agent reinforcement learning framework for a multi-energy park system, wherein the multi-agent reinforcement learning framework includes a state space and an action space, wherein the state space includes the state of each agent at each time point, and the action space includes the action of each agent at each time point, and the agent is the control center of the multi-energy park;

[0007] Build an energy scheduling model based on multi-agent deep deterministic policy gradient algorithm and Dropout method;

[0008] Training the energy scheduling model using the state space as input and the action space as output;

[0009] Based on the trained energy scheduling model, the energy scheduling plan of the multi-energy park system is determined.

[0010] In a second aspect, the present application provides a multi-region energy internet optimization scheduling device, comprising:

[0011] A framework construction module, used to construct a multi-agent reinforcement learning framework for a multi-energy park system, wherein the multi-agent reinforcement learning framework includes a state space and an action space, wherein the state space includes the state of each agent at each time point, and the action space includes the action of each agent at each time point, and the agent is the control center of the multi-energy park;

[0012] Model building module, used to build energy scheduling model based on multi-agent deep deterministic policy gradient algorithm and Dropout method;

[0013] A model training module, used for training the energy scheduling model by taking the state space as input and the action space as output;

[0014] The scheme determination module is used to determine the energy scheduling scheme of the multi-energy park system based on the trained energy scheduling model.

[0015] The present application provides a multi-region energy Internet optimization scheduling method and device, by constructing a multi-agent reinforcement learning framework of a multi-energy park system, the multi-agent reinforcement learning framework includes a state space and an action space, the state space includes the state of each agent at each time point, the action space includes the action of each agent at each time point, and the agent is the control center of the multi-energy park; an energy scheduling model is built based on a multi-agent deep deterministic policy gradient algorithm and a Dropout method; the state space is used as input and the action space is used as output to train the energy scheduling model; based on the trained energy scheduling model, the energy scheduling scheme of the multi-energy park system is determined. The present application utilizes a multi-agent reinforcement learning framework to learn and adapt to the dynamic changes of various energy sources in the multi-energy park system, so that the system can make timely and effective adjustments when facing uncertain factors, enhance the system's adaptability and robustness, and introduce the Dropout method to improve the generalization ability of the model, so that the energy scheduling model can still maintain a high scheduling accuracy and stability when facing new or unknown energy scheduling scenarios. At the same time, by building a multi-agent reinforcement learning framework, we can fully consider the complementarity and conversion relationship between various energy sources, realize the coordinated dispatch of energy, help optimize energy configuration, reduce energy waste, and improve energy utilization efficiency. By building an intelligent energy dispatch model, we can realize the intelligent management of the multi-energy park system, which helps to reduce the burden of manual dispatch, improve dispatch efficiency, reduce the risk of human error, and help achieve the sustainable development of the multi-energy park system. By optimizing the energy dispatch plan, we can reduce energy consumption and emissions and promote green and low-carbon energy utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0017] Figure 1 It is a flowchart for implementing the multi-region energy Internet optimization scheduling method provided in the embodiment of the present application;

[0018] Figure 2 It is a diagram of the multi-energy regional energy interconnection structure provided in the embodiment of the present application;

[0019] Figure 3 is a training flow chart of the energy scheduling model provided in the embodiment of the present application;

[0020] Figure 4 It is a structural diagram of a multi-region energy Internet optimization scheduling device provided in an embodiment of the present application. Detailed implementation manners

[0021] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0022] To make the objectives, technical solutions, and advantages of the present application clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.

[0023] Figure 1 The implementation flowchart of the multi-region energy Internet optimal scheduling method provided for the embodiments of the present application is described in detail as follows:

[0024] In step 101, a multi-agent reinforcement learning framework for the multi-energy park system is constructed. The multi-agent reinforcement learning framework includes a state space and an action space. The state space includes the states of each agent at each time point, and the action space includes the actions of each agent at each time point. The agent is the control center of the multi-energy park.

[0025] In the embodiments of the present application, based on the multi-energy park system, a multi-agent reinforcement learning framework for the energy Internet optimal scheduling problem is constructed. Among them, the multi-agent reinforcement learning framework includes a state space and an action space. Among them, the state space includes the states of each agent at each time point, and the action space includes the actions of each agent at each time point. The agent is the control center of the multi-energy park. Refer to Figure 2 , the agent is the control center of each park,

[0026] Since the same energy conversion equipment is used in each park, the same state space, action space, and reward function are used for each agent.

[0027] In a possible implementation manner, before constructing the multi-agent reinforcement learning framework of the multi-energy park system, the method may further include:

[0028] Construct an energy Internet system model for the collaborative optimization operation of the multi-energy park. The energy Internet system model takes the minimum daily system operation cost of a single multi-energy area as the optimization goal;

[0029] Among them, the objective function of the energy Internet system model is:

[0030] minC i =min(C i,gas +(C i,b -Ci,s )+(C i,BES +C i,HES ))

[0031] Among them, C i is the system operation cost of the multi-energy park on day i, C i,gas The cost of purchasing natural gas for multi-energy park i, C i,b is the electricity purchase cost of multi-energy park i, C i,s is the electricity sales revenue of multi-energy park i, C i,BES is the charging and discharging depreciation cost of the electric energy storage in the multi-energy park i, C i,HES is the charging and discharging depreciation cost of the thermal energy storage in multi-energy park i.

[0032] Among them, refer to Figure 2 A single park includes common units such as electric boiler (EB), gas boiler (GB), photovoltaic array (PV), battery energy storage (BES) and heat energy storage (HES), as well as user electric-heat load. The parks can transfer electric energy in both directions through the energy bus, and transmit information to the dispatch center. No information is transmitted between parks to protect privacy.

[0033] The various parks exchange electricity through the energy bus and settle accounts through the internal market. In the system, electric boilers and gas boilers consume electricity and natural gas to generate heat respectively, and the corresponding model is as follows:

[0034] h EB (t) = p EB (t)η EB (1)

[0035] h GB (t) = M GB (t)η GB (2)

[0036] Among them, h EB (t) is the thermal power output of the electric boiler at time t, h GB (t) is the thermal power output of the gas boiler at time t, p EB (t) is the electric power consumed by the electric boiler at time t, M GB (t) is the natural gas consumption in the gas boiler at time t, η EB is the conversion efficiency of the electric boiler, η GB It is the conversion efficiency of gas boiler.

[0037] In addition, the calculation formula of the state of charge (SOC) of the energy storage at time t is:

[0038]

[0039] Among them, c SOC (t) is the state of charge of the energy storage at time t, p BES (t) is the charging / discharging power of the energy storage at time t, p BES (t)>0 indicates discharge, p BES (t) < 0 means charging; Q BES is the capacity of electrical energy storage, is the charge state of the energy storage at the initial moment, Δt is the scheduling time scale, η BES is the charge / discharge coefficient of the electric energy storage. The operation mode of the electric energy storage is similar to that of the thermal energy storage, and will not be described in detail in this embodiment.

[0040] Optionally, before constructing the multi-agent reinforcement learning framework of the multi-energy park system, the embodiment of the present application also needs to establish an energy Internet system model for the coordinated optimization operation of the multi-energy park. Among them, the energy Internet system model takes the minimum system operation cost of a single multi-energy park as the optimization goal, and the objective function of the energy Internet system model is:

[0041] minC i =min(C i,gas +(C i,b -C i,s )+(C i,BES +C i,HES ))

[0042] Among them, C i is the system operation cost of the multi-energy park on day i, C i,gas The cost of purchasing natural gas for multi-energy park i, C i,b is the electricity purchase cost of multi-energy park i, C i,s is the electricity sales revenue of multi-energy park i, C i,BES is the charging and discharging depreciation cost of the electric energy storage in the multi-energy park i, C i,HES is the charging and discharging depreciation cost of the thermal energy storage in multi-energy park i.

[0043] The cost calculation formula for purchasing natural gas is:

[0044]

[0045] Among them, ε gas is the unit calorific value price of natural gas, h i,GB (t) is the thermal power output of the gas boiler in the multi-energy park i in time period t, T is the total time period of the system scheduling, and Δt is the time slot length.

[0046] The internal market transaction cost is the electricity purchase cost minus the electricity sales revenue, that is, C i,b -C i,s , C i,b and C i,s are the electricity purchase cost and electricity sales revenue of multi-energy park i respectively, and the calculation formula is as follows:

[0047]

[0048] Among them, p i,b (t) is the amount of electricity purchased by multi-energy park i in period t, p i,s (t) is the electricity sales of multi-energy park i in time period ε, ε b is the electricity purchase price in the internal market at time period t, ε s is the electricity price in the internal market at time period t.

[0049] The depreciation cost of charging and discharging of electric energy storage C i,BES The calculation formula is:

[0050]

[0051] Among them, p i,BES (t) is the charging / discharging power of the electric energy storage in multi-energy park i in time period t, p i,BES (t)>0 means that the energy storage is in the discharge state, p i,BES (t)<0 means it is in charging state, ρ BES is the depreciation cost coefficient of electric energy storage.

[0052] Thermal storage charging and discharging depreciation cost C i,HES The calculation formula is:

[0053]

[0054] Among them, p i,HES (t) is the charging / discharging power of the electric energy storage in multi-energy park i in time period t, p i,HES (t)>0 means that the thermal energy storage is in the state of heat release, p i,HES (t)<0 indicates heat storage state, ρ HES is the depreciation cost coefficient of thermal energy storage.

[0055] Correspondingly, the constraints of the energy Internet system model include power balance constraints, equipment operation constraints and clearing mechanism constraints.

[0056] The power balance constraint is:

[0057]

[0058] Among them, p i,b(t) is the electricity purchase amount of multi-energy park i, p i,s (t) is the electricity sales of multi-energy park i, p i,OV (t) is the output power of photovoltaic power in multi-energy park i in time period t, p i,BES (t) is the charging / discharging power of the multi-energy park i electric energy storage, p i,EB (t) is the input power of electric boiler i in the multi-energy park, p i,load (t) is the electric load of multi-energy park i in time period t, h i,GB (t) is the thermal power output of the gas boiler in multi-energy park i in time period t, h i,EB (t) is the thermal power output of the electric boiler in multi-energy park i in time period t, h i,HES (t) is the charging / discharging power of the thermal energy storage in multi-energy park i during time period t, h i,load (t) is the heat load of multi-energy park i in time period t.

[0059] The equipment operation constraints are:

[0060]

[0061] Among them, h GB (t) is the thermal power output of the gas boiler, is the lower limit of the thermal power output of the gas boiler, is the upper limit of the thermal power output of the gas boiler, h EB (t) is the thermal power output of the electric boiler, is the lower limit of the thermal power output of the electric boiler, is the upper limit of the thermal power output of the electric boiler, p BES (t) is the charging and discharging power of the energy storage, is the lower limit of the energy storage charging / discharging power, is the upper limit of the energy storage charge / discharge power, h HES (t) is the thermal energy storage charging and discharging power, is the lower limit of the thermal energy storage charging / discharging power, is the upper limit of the thermal energy storage charging / discharging power, C SOC (t) is the state of charge of the energy storage device in time period t, is the lower limit of the state of charge of the energy storage device, It is the upper limit of the state of charge of the energy storage device.

[0062] The clearing mechanism constraints are:

[0063]

[0064] Among them, ε s (t) is the electricity price in the internal market, ε b (t) is the electricity purchase price in the internal market, εm (t) is the internal clearing price, p b (t) is the total electricity purchase of all multi-energy parks, p s (t) is the total electricity sales of all multi-energy parks, ε grid,s (t) is the price of electricity sold to the grid, ε grid,b (t) is the price of electricity purchased from the power grid.

[0065] Among them, ε m (t) satisfies ε grid,s (t)≤ε m (t)≤ε grid,b (t), where ε is set in this embodiment. m (t) = (ε grid,b (t)+ε grid,s (t)) / 2.

[0066] The above settings can ensure that ε s (t) and ε b (t) respectively satisfy ε grid,s (t)≤ε s (t)≤ε m (t) and ε m (t)≤ε b (t)≤ε grid,b (t).

[0067] In one possible implementation, a multi-agent reinforcement learning framework for building a multi-functional campus system may include:

[0068] Obtain the user's electrical load and thermal load demand, photovoltaic power generation, charge state of electric energy storage, the price of selling and purchasing electricity from the grid, and the dispatch period, and construct the state space;

[0069] Obtain the processing of each device in the multi-energy park system as well as the power purchase and sales in the internal market to build an action space.

[0070] Optionally, for each intelligent agent, the state space includes the user's electrical load and thermal load demand, photovoltaic power generation power, the state of charge of the electric energy storage, the price of selling and purchasing electricity from the power grid, and the scheduling period. In this embodiment, the state space is composed of the user's electrical load and thermal load demand, photovoltaic power generation power, the state of charge of the electric energy storage, the price of selling and purchasing electricity from the power grid, and the scheduling period. The state space is:

[0071] s i,t ={p i,load (t),h i,load (t),p i,PV (t),c i,SOC (t),ε grid,s (t),εgrid,b (t),t} (10)

[0072] Among them, s i,t is the state of multi-energy park i at time period t, p i,load (t) is the electric load of multi-energy park i in time period t, h i,load (t) is the heat load of multi-energy park i in period t, p i,PV (t) is the output power of photovoltaic power in multi-energy park i in time period t, c i,SOC (t) is the state of charge of the energy storage device in multi-energy park i at time period t, ε grid,s (t) is the price of electricity sold to the grid, ε grid,b (t) is the price of electricity purchased from the grid, and t is the time period.

[0073] For each agent, the action space can be composed of the processing of each device in the multi-energy park system and the purchase and sale of electricity in the internal market.

[0074] From formula (1), we can see that h i,EB (t) After determination, p i,EB (t) can be determined by calculation; by the formula h i,GB (t)+h i,EB (t)+h i,HES (t) = h i,load (t) It can be seen that h i,GB (t) After determination, h i,HES (t) can be determined by calculation; further by formula p i,b (t)+p i,s (t)+p i,PV (t)+p i,BES (t)-p i,EB (t) = p i,load (t) It can be seen that, determine p i,BES (t) after, p i,b (t) and p i,s (t) can also be determined by calculation. Therefore, the action space can be expressed as:

[0075] a i,t ={h i,GB (t),h i,EB (t),p i,BES (t)} (11)

[0076] Among them, a i,t is the action of multi-energy park i in time period t, h i,GB (t) is the thermal power output of the gas boiler in multi-energy park i in time period t, h i,EB (t) is the thermal power output of the electric boiler in multi-energy park i in time period t, p i,BES(t) is the charging / discharging power of the multi-energy park i-electric energy storage.

[0077] In addition, the multi-agent reinforcement learning framework of this embodiment also includes a reward function, which is calculated from the state space and the action space. The reward function can be divided into two parts: system operation cost penalty and constraint violation penalty.

[0078] r t (s t ,a t )=-(β 1 C i +β 2 F t ) (12)

[0079]

[0080] Among them, r t (s t , a t ) is the reward function, s t is the state space, a t is the action space, β 1 is the system operation cost penalty coefficient, β 2 is the penalty coefficient for system violation of constraints, H GB is the operating limit of the gas boiler, H EB is the operating limit of the electric boiler, P BES is the operating limit of electric energy storage, H HES is the operating limit of thermal energy storage.

[0081] Among them, H GB The value of is shown in formula (14), H EB , P BES and H HES The calculation method of H GB similar.

[0082]

[0083] in, is the lower limit of the thermal power output of the electric boiler, It is the upper limit of the thermal power output of the electric boiler.

[0084] The embodiment of the present application can fully consider the complementarity and conversion relationship between various energy sources (such as electricity, heat, natural gas, etc.) by constructing a multi-agent reinforcement learning framework for a multi-energy park system, and realize the coordinated scheduling of energy, which helps to optimize energy allocation, reduce energy waste, and improve energy utilization efficiency. At the same time, the multi-agent reinforcement learning framework can learn and adapt to the dynamic changes of various energy sources in the multi-energy park system, such as load fluctuations, uncertainty in renewable energy power generation, etc., so that the system can make timely and effective adjustments in the face of uncertain factors, and enhance the system's adaptability and robustness.

[0085] In step 102, an energy scheduling model is built based on a multi-agent deep deterministic policy gradient algorithm and a Dropout method.

[0086] In the embodiment of the present application, the energy scheduling model is constructed by integrating the multi-agent deep deterministic policy gradient algorithm and the Dropout method. In order to effectively prevent overfitting, two policy functions are used to reduce the adverse effects of overestimation errors, increase the robustness of the model, and improve training efficiency. In this embodiment, the Dropout method is used to randomly ignore a part of neurons during the training process of each agent network, reduce the dependency between neurons, and thus reduce the risk of overfitting of the model for specific inputs.

[0087] The energy scheduling model constructed by the embodiment of the present application using the multi-agent deep deterministic policy gradient algorithm and the Dropout method can handle complex energy scheduling problems and provide accurate scheduling solutions, thereby ensuring the stable operation of the multi-energy park system. At the same time, the introduction of the Dropout method further improves the generalization ability of the model, so that the energy scheduling model can still maintain high scheduling accuracy and stability when facing new or unknown energy scheduling scenarios.

[0088] In step 103, the energy scheduling model is trained using the state space as input and the action space as output.

[0089] In the embodiment of the present application, the state space in step 101 is used as input, and the action space in step 101 is used as output to train the energy scheduling model constructed in step 102.

[0090] In a possible implementation, the energy scheduling model may include a first strategy network, a second strategy network, and a value network.

[0091] Optionally, the embodiment of the present application constructs two strategy networks (Actor1 network and Actor2 network) and a value network Critic network for each agent i.

[0092] The embodiment of the present application uses two policy networks for training to alleviate the overfitting phenomenon of the multi-agent deep deterministic policy gradient algorithm, and reduces the impact of the algorithm's overestimation error by selecting the smaller value of the reward value calculated by the action generated by the two policy networks as the final reward.

[0093] The embodiment of the present application uses the value network Critic network to verify whether the training of the two strategy networks is complete.

[0094] The training idea is: in each round of the training process, each agent calculates the actions generated by the policy network in the current state and determines the actual action to be performed. After all agents perform the action, they calculate the reward value and observe the new state, record the data, and store it in the experience buffer pool. Determine whether the experience buffer pool has reached its maximum capacity at this time. If not, s t+1 As the initial state of the next set of data, repeat the above operation. If the experience buffer pool reaches the maximum capacity, iterate the initial data, because the initial data is difficult to be used as effective training data due to the randomness of the network parameters. Then, randomly extract a batch of data in the experience buffer pool to update the value network and the policy network. Repeat the above steps until the training is completed. When conducting multi-energy regional energy Internet scheduling tests, you can use the current system state s t , use the trained strategy network to select the scheduling action. Then, execute the action and enter the next state, so as to realize the real-time optimization scheduling of the energy Internet.

[0095] In one possible implementation, training the energy scheduling model with the state space as input and the action space as output may include:

[0096] Initialize the first network parameters of the first strategy network, the second network parameters of the second strategy network, and the third network parameters of the value network, and use the initialized first network parameters as the current first network parameters of the current first strategy network, the initialized second network parameters as the current second network parameters of the current second strategy network, and the initialized third network parameters as the current third network parameters of the current value network;

[0097] Assigning the current first network parameter to the current first target network of the current first strategy network, assigning the current second network parameter to the current second target network of the current second strategy network, and assigning the current third network parameter to the current third target network of the current value network;

[0098] Obtain the current state, input the current state into the current first strategy network, output the current first action, and input the current state into the current second strategy network, output the current second action;

[0099] Input the current state and the current first action into the current value network to obtain a first output value, input the current state and the current second action into the current value network to obtain a second output value, select the minimum output value from the first output value and the second output value as the target output value, and use the policy network corresponding to the target output value as the target policy network, and use the action corresponding to the target output value as the target action;

[0100] After the target policy network executes the target action, obtain the reward value and the new state, and use the current state, the target action, the new state, the reward value, the network parameters of the target network corresponding to the target policy network, and the network parameters of the current third target network to update the network parameters of the third network of the current value network to obtain a new value network, and use the updated network parameters of the third network to update the network parameters of the current third target network to obtain a new third target network;

[0101] Determine whether the current iteration number is an odd iteration number;

[0102] If the current iteration number is an odd iteration number, calculate new first network parameters using the current first action to obtain a new first policy network, and use the new first network parameters to update the network parameters of the current first target network to obtain a new first target network;

[0103] Determine whether the current iteration number reaches the maximum iteration number;

[0104] If the current iteration number reaches the maximum iteration number, determine the new first policy network as the trained energy scheduling model;

[0105] If the current iteration number does not reach the maximum iteration number, update the current state using the new state, update the current first policy network using the new first policy network, update the current first target network using the new first target network, update the current value network using the new value network, update the current third target network using the new third target network, and increment the current iteration number by one, and return to the step of obtaining the current state, inputting the current state into the current first policy network to output the current first action, and inputting the current state into the current second policy network to output the current second action to continue execution.

[0106] Among them, the Actor1 network π in the embodiments of the present application i1 、Actor2 network π i2 and Critic network Q i all have their respective target networks, the first target network π i ′ 1 、the second target network π i ′ 2and the third target network Q i ′ First, use n groups (assuming there are n parks, i.e. n multi-functional parks) of random parameters to initialize the network parameters of the Actor1 network of each agent i and the network parameters of the Actor2 network And the network parameters of the Critic network And copy the parameters to each network's respective target network, that is, and For each agent i, according to the current policy network π i1 and π i2 Select action a ik,t ~π i,k (s i,t )+∈,∈ is the strategy noise, which obeys the normal distribution k∈{1,2} represents the strategy π i1 and π i2 The two actions are input into the Critic network Q together with the state i , select the action with the smaller reward value as the action to be performed by agent i. After executing all actions a k,t =(a 1k,t ,…,a nk,t ), calculate the reward value r k,t , and get the new state s k,t+1 =(s 1k,t+1 ,…,s nk,t+1 ), the data tuple (s t ,a k,t ,r k,t ,a k,t ) is saved in the buffer experience pool.

[0107] Reference Figure 3 , the specific training process is as follows:

[0108] Step 3.1, randomly initialize the first policy network π of each agent i1 The first network parameter Second Strategy Network π i2 The second network parameters and value network Q i The third network parameter And initialize the first network parameters As the current first network parameter The second network parameters to be initialized As the current second network parameter And the third network parameters to be initialized As the current third network parameter

[0109] Step 3.2, copy the network parameters of each initialized network to their respective target networks, that is, assign the current first network parameter to the current first policy network π i1 of the current first target network π i ′ 1 , assign the current second network parameter to the current second target network π i2 of the current second policy network π i ′ 2 , and assign the current third network parameter to the current third target network Q i of the current value network Q i ′ .

[0110] Step 3.3, obtain the current state s i,t , and input the current state s i,t into the current first policy network π i1 to obtain the current first action a i1 , and input the current state s i,t into the current second policy network π i2 to obtain the current second action a i2 .

[0111] Step 3.4, input the current state s i,t and the current first action a i1 into the current value network Q i to obtain the first output value. Input the current state s i,t and the current second action a i2 into the current value network Q i to obtain the second output value. Select the minimum output value from the first output value and the second output value as the target output value, and use the policy network corresponding to the target output value as the target policy network, and use the action corresponding to the target output value as the target action.

[0112] Step 3.5, after the target policy network executes the target action, obtain the current reward value r i,t and the new state s i,t+1 , and store the reward value r i,t and the new state s i,t+1 in the experience pool.

[0113] Step 3.6, use the current state s i,t , the target action, the new state s i,t+1 , the reward value ri,t , the network parameters of the target network corresponding to the target strategy network and the current third target network Q i ′ The network parameters are updated to update the current value network Q i The third network parameter Get the new value network Q i ; and using the updated third network parameters Update the current third target network Q i ′ The network parameters of the new third target network Q i ′ .

[0114] Step 3.7, determine whether the current number of iterations is an odd number of iterations. If so, go to step 3.8.

[0115] Step 3.8, using the current first action a i1 Calculate the new first network parameters Using the new first network parameters Update to get the new first strategy network π i1 , and use the new first network parameters Update the network parameters of the current first target network to obtain a new first target network π i ′ 1 .

[0116] Step 3.9, determine whether the current number of iterations reaches the maximum number of iterations. If so, go to step 3.10; if not, go to step 3.11.

[0117] Step 3.10, the new first strategy network π i1 The energy scheduling model is determined to be trained.

[0118] Step 3.11, using the new state s i,t+1 Update current status i,t , update the current first strategy network using the new first strategy network, update the current first target network using the new first target network, update the current value network using the new value network, update the current third target network using the new third target network, and increase the current number of iterations by one, and return to step 3.3 to continue execution.

[0119] In a possible implementation, after determining whether the current number of iterations is an odd number of iterations, the method may further include:

[0120] If the current number of iterations is not an odd number of iterations, the new second network parameters are calculated using the current second action to obtain a new second strategy network, and the network parameters of the current second target network are updated using the new second strategy network parameters to obtain a new second target network;

[0121] Accordingly, judging whether the current number of iterations reaches the maximum number of iterations includes:

[0122] If the current number of iterations reaches the maximum number of iterations, the new second strategy network is determined as the trained energy scheduling model;

[0123] If the current number of iterations does not reach the maximum number of iterations, the current state is updated with the new state, the current second policy network is updated with the new second policy network, the current second target network is updated with the new second target network, the current value network is updated with the new value network, the current third target network is updated with the new third target network, and the current number of iterations is increased by one, and the current state is returned to obtain the current state, the current state is input into the current first policy network, the current first action is output, and the current state is input into the current second policy network, the current second action step is output to continue execution.

[0124] Optionally, after step 3.7, if the number of iterations is not an odd number of iterations, go to step 3.12.

[0125] Step 3.12, using the current second action a i2 Calculate new second network parameters Using the new second network parameters Update to get the new second strategy network π i2 , and use the new second network parameters Update the network parameters of the current second target network to obtain a new second target network π i ′ 2 .

[0126] Accordingly, after step 3.9, if it is reached, go to step 3.13; if it is not reached, go to step 3.14.

[0127] Step 3.13, the new second strategy network π i2 The energy scheduling model is determined to be trained.

[0128] Step 3.14, using the new state s i,t+1 Update current status i,t, update the current second policy network using the new second policy network, update the current second target network using the new second target network, update the current value network using the new value network, update the current third target network using the new third target network, increment the current iteration count by one, and return to step 3.3 to continue execution.

[0129] In a possible implementation, using the current state, target action, new state, reward value, network parameters of the target network corresponding to the target policy network, and network parameters of the current third target network, update the third network parameters of the current value network to obtain a new value network, and use the updated third network parameters to update the network parameters of the current third target network to obtain a new third target network, which may include:

[0130] Input the current state, target action, new state, reward value, network parameters of the target network corresponding to the target policy network, and network parameters of the current third target network into the minimum loss function;

[0131] Input the minimum loss function into the first gradient formula to obtain new third network parameters, and use the new third network parameters to update the third network parameters of the current value network to obtain a new value network;

[0132] Input the updated third network parameters into the first formula to obtain the network parameters of the new third target network, and use the network parameters of the new third target network to obtain a new third target network;

[0133] Among them, the minimum loss function is:

[0134]

[0135] Among them, is the minimum loss function, y i,t is the output value, r i,t is the reward value, Q i ′ is the current third target network, s i,t+1 is the state of the multi - energy park i at time t + 1, π i ′ k (s i,t+1 ) is the target network corresponding to the target policy network, is the third network parameter of the current value network, Q i is the current value network, s i,t is the state of the multi - energy park i at time t, a i,t is the action of the multi - energy park i at time t, k is the serial number of the policy network, and γ is the discount factor;

[0136] The first gradient formula is:

[0137]

[0138] in, For Perform gradient calculation, E(·) is the expected function, μ i is the value network learning rate of multi-energy park i, min(·) is the minimum function;

[0139] The first formula is:

[0140]

[0141] in, are the network parameters of the new third target network, and α is the soft update coefficient.

[0142] Optionally, step 3.6 is specifically as follows: input the current state, target action, new state, reward value, network parameters of the target network corresponding to the target policy network, and network parameters of the current third target network into the minimum loss function And input the minimum loss function into the first gradient formula The new third network parameters are obtained And using the new third network parameters Update the third network parameters of the current value network to obtain a new value network. Then set the network parameters of the updated third network Enter the first formula The network parameters of the new third target network are obtained And use the network parameters of the new third target network A new third target network is obtained.

[0143] In a possible implementation, calculating a new first network parameter by using the current first action to obtain a new first policy network, and updating a network parameter of the current first target network by using the new first network parameter to obtain a new first target network may include:

[0144] Inputting the current first action into the second gradient formula to obtain new first network parameters, and using the new first network parameters to obtain a new first strategy network;

[0145] Inputting the new first network parameters into the second formula to obtain the new network parameters of the first target network, and using the new network parameters of the first target network to obtain the new first target network;

[0146] Among them, the second gradient formula is:

[0147]

[0148] in, For Perform gradient calculation, J(π i1 ) is the action value function under the first policy network, π i1 For the first strategy network, For a i1,t Perform gradient calculation, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, is the first network parameter of the first strategy network, τ i is the strategy network learning rate of multi-functional park i, a i1,t is the action generated when the first strategy network is adopted;

[0149] The second formula is:

[0150]

[0151] in, is the network parameter of the new first target network, and α is the soft update coefficient.

[0152] Optionally, step 3.8 is specifically: input the current first action into the second gradient formula The new first network parameters are obtained And use the new first network parameters Get the new first strategy network. Then set the new first network parameters Enter the second formula The network parameters of the new first target network are obtained And use the network parameters of the new first target network Get the new first target network.

[0153] In a possible implementation, calculating a new second network parameter using the current second action to obtain a new second policy network, and updating a network parameter of the current second target network using the new second policy network parameter to obtain a new second target network may include:

[0154] Input the current second action into the third gradient formula to obtain new second network parameters, and use the new second network parameters to obtain a new second strategy network;

[0155] Inputting the new second network parameters into the third formula to obtain the new network parameters of the second target network, and using the new network parameters of the second target network to obtain the new second target network;

[0156] Among them, the third gradient formula is:

[0157]

[0158] in, For Perform gradient calculation, J(π i2 ) is the action value function under the second policy network, π i2 is the second strategy network, For a i2,t Perform gradient calculation, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, is the second network parameter of the second strategy network, τ i is the strategy network learning rate of multi-functional park i, a i2,t is the action generated when the second strategy network is adopted;

[0159] The third formula is:

[0160]

[0161] in, is the network parameter of the new second target network, and α is the soft update coefficient.

[0162] Optionally, step 3.12 is specifically: input the current second action into the third gradient formula The new second network parameters are obtained Using the new second network parameters Get a new second strategy network. Then, set the new second network parameters Enter the third formula In the process, network parameters of a new second target network are obtained, and a new second target network is obtained by using the network parameters of the new second target network.

[0163] In step 104, an energy scheduling plan for the multi-energy park system is determined based on the trained energy scheduling model.

[0164] In an embodiment of the present application, the current state is input into the energy scheduling model trained in step 103, and the action of the multi-energy park system in the current state can be output, that is, the current state and the corresponding action are used as the energy scheduling plan of the multi-energy park system.

[0165] The present application provides a multi-region energy Internet optimization scheduling method, by constructing a multi-agent reinforcement learning framework of a multi-energy park system, the multi-agent reinforcement learning framework includes a state space and an action space, the state space includes the state of each agent at each time point, the action space includes the action of each agent at each time point, and the agent is the control center of the multi-energy park; an energy scheduling model is built based on a multi-agent deep deterministic policy gradient algorithm and a Dropout method; the state space is used as input and the action space is used as output to train the energy scheduling model; based on the trained energy scheduling model, the energy scheduling scheme of the multi-energy park system is determined. The present application utilizes a multi-agent reinforcement learning framework to learn and adapt to the dynamic changes of various energy sources in a multi-energy park system, so that the system can make timely and effective adjustments when facing uncertain factors, enhance the system's adaptability and robustness, and introduce the Dropout method to improve the generalization ability of the model, so that the energy scheduling model can still maintain a high scheduling accuracy and stability when facing new or unknown energy scheduling scenarios. At the same time, by building a multi-agent reinforcement learning framework, we can fully consider the complementarity and conversion relationship between various energy sources, realize the coordinated dispatch of energy, help optimize energy configuration, reduce energy waste, and improve energy utilization efficiency. By building an intelligent energy dispatch model, we can realize the intelligent management of the multi-energy park system, which helps to reduce the burden of manual dispatch, improve dispatch efficiency, reduce the risk of human error, and help achieve the sustainable development of the multi-energy park system. By optimizing the energy dispatch plan, we can reduce energy consumption and emissions and promote green and low-carbon energy utilization.

[0166] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0167] The following is an embodiment of the device of the present application. For details not described in detail, please refer to the corresponding method embodiment described above.

[0168] Figure 4 The schematic diagram of the structure of the multi-region energy internet optimization scheduling device provided in the embodiment of the present application is shown. For the convenience of explanation, only the part related to the embodiment of the present application is shown, which is described in detail as follows:

[0169] like Figure 4 As shown, the multi-region energy Internet optimization scheduling device 4 includes:

[0170] A framework construction module 41 is used to construct a multi-agent reinforcement learning framework for a multi-energy park system. The multi-agent reinforcement learning framework includes a state space and an action space. The state space includes the state of each agent at each time point, and the action space includes the action of each agent at each time point. The agent is the control center of the multi-energy park.

[0171] A model building module 42, used to build an energy scheduling model based on a multi-agent deep deterministic policy gradient algorithm and a Dropout method;

[0172] A model training module 43, for training an energy scheduling model using the state space as input and the action space as output;

[0173] The scheme determination module 44 is used to determine the energy scheduling scheme of the multi-energy park system based on the trained energy scheduling model.

[0174] The present application provides a multi-region energy Internet optimization scheduling device, which constructs a multi-agent reinforcement learning framework of a multi-energy park system. The multi-agent reinforcement learning framework includes a state space and an action space. The state space includes the state of each agent at each time point, and the action space includes the action of each agent at each time point. The agent is the control center of the multi-energy park; an energy scheduling model is built based on a multi-agent deep deterministic policy gradient algorithm and a Dropout method; the state space is used as input and the action space is used as output to train the energy scheduling model; based on the trained energy scheduling model, the energy scheduling scheme of the multi-energy park system is determined. The present application utilizes a multi-agent reinforcement learning framework to learn and adapt to the dynamic changes of various energy sources in the multi-energy park system, so that the system can make timely and effective adjustments when facing uncertain factors, enhance the system's adaptability and robustness, and introduce the Dropout method to improve the generalization ability of the model, so that the energy scheduling model can still maintain a high scheduling accuracy and stability when facing new or unknown energy scheduling scenarios. At the same time, by building a multi-agent reinforcement learning framework, we can fully consider the complementarity and conversion relationship between various energy sources, realize the coordinated dispatch of energy, help optimize energy configuration, reduce energy waste, and improve energy utilization efficiency. By building an intelligent energy dispatch model, we can realize the intelligent management of the multi-energy park system, which helps to reduce the burden of manual dispatch, improve dispatch efficiency, reduce the risk of human error, and help achieve the sustainable development of the multi-energy park system. By optimizing the energy dispatch plan, we can reduce energy consumption and emissions and promote green and low-carbon energy utilization.

[0175] In a possible implementation, the energy scheduling model may include a first strategy network, a second strategy network, and a value network, and the model training module may be used to:

[0176] Initialize the first network parameters of the first strategy network, the second network parameters of the second strategy network, and the third network parameters of the value network, and use the initialized first network parameters as the current first network parameters of the current first strategy network, the initialized second network parameters as the current second network parameters of the current second strategy network, and the initialized third network parameters as the current third network parameters of the current value network;

[0177] Assigning the current first network parameter to the current first target network of the current first strategy network, assigning the current second network parameter to the current second target network of the current second strategy network, and assigning the current third network parameter to the current third target network of the current value network;

[0178] Obtain the current state, input the current state into the current first strategy network, output the current first action, and input the current state into the current second strategy network, output the current second action;

[0179] Input the current state and the current first action into the current value network to obtain a first output value, input the current state and the current second action into the current value network to obtain a second output value, and select the minimum output value from the first output value and the second output value as the target output value, and use the policy network corresponding to the target output value as the target policy network, and use the action corresponding to the target output value as the target action;

[0180] After the target policy network executes the target action, the reward value and the new state are obtained, and the third network parameters of the current value network are updated using the current state, the target action, the new state, the reward value, the network parameters of the target network corresponding to the target policy network, and the network parameters of the current third target network to obtain a new value network, and the network parameters of the current third target network are updated using the updated third network parameters to obtain a new third target network;

[0181] Determine whether the current number of iterations is an odd number of iterations;

[0182] If the current number of iterations is an odd number of iterations, the new first network parameters are calculated using the current first action to obtain a new first strategy network, and the network parameters of the current first target network are updated using the new first network parameters to obtain a new first target network;

[0183] Determine whether the current number of iterations has reached the maximum number of iterations;

[0184] If the current number of iterations reaches the maximum number of iterations, the new first strategy network is determined as the trained energy scheduling model;

[0185] If the current number of iterations does not reach the maximum number of iterations, the current state is updated with the new state, the current first policy network is updated with the new first policy network, the current first target network is updated with the new first target network, the current value network is updated with the new value network, the current third target network is updated with the new third target network, and the current number of iterations is increased by one, and the current state is returned to obtain the current state, the current state is input into the current first policy network, the current first action is output, and the current state is input into the current second policy network, and the current second action step is output to continue execution.

[0186] In a possible implementation, the model training module can also be used to:

[0187] If the current number of iterations is not an odd number of iterations, the new second network parameters are calculated using the current second action to obtain a new second strategy network, and the network parameters of the current second target network are updated using the new second strategy network parameters to obtain a new second target network;

[0188] Accordingly, it is determined whether the current number of iterations reaches the maximum number of iterations;

[0189] If the current number of iterations reaches the maximum number of iterations, the new second strategy network is determined as the trained energy scheduling model;

[0190] If the current number of iterations does not reach the maximum number of iterations, the current state is updated with the new state, the current second policy network is updated with the new second policy network, the current second target network is updated with the new second target network, the current value network is updated with the new value network, the current third target network is updated with the new third target network, and the current number of iterations is increased by one, and the current state is returned to obtain the current state, the current state is input into the current first policy network, the current first action is output, and the current state is input into the current second policy network, the current second action step is output to continue execution.

[0191] In a possible implementation, the model training module can also be used to:

[0192] Input the current state, target action, new state, reward value, network parameters of the target network corresponding to the target policy network, and network parameters of the current third target network into the minimum loss function;

[0193] Input the minimum loss function into the first gradient formula to obtain new third network parameters, and use the new third network parameters to update the third network parameters of the current value network to obtain a new value network;

[0194] Inputting the updated third network parameters into the first formula to obtain new network parameters of the third target network, and obtaining a new third target network using the new network parameters of the third target network;

[0195] Among them, the minimum loss function is:

[0196]

[0197] in, is the minimum loss function, y i,t is the output value, r i,t is the reward value, Q i ′ is the current third target network, s i,t+i is the state of multi-energy park i at time period t+1, π i ′ k (s i,t+i ) is the target network corresponding to the target strategy network, is the third network parameter of the current value network, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, a i,t is the action of multi-energy park i in time period t, k is the sequence number of the strategy network, and γ is the discount coefficient;

[0198] The first gradient formula is:

[0199]

[0200] in, For Perform gradient calculation, E(·) is the expected function, μ i is the value network learning rate of multi-energy park i, min(·) is the minimum function;

[0201] The first formula is:

[0202]

[0203] in, are the network parameters of the new third target network, and α is the soft update coefficient.

[0204] In a possible implementation, the model training module can also be used to:

[0205] Inputting the current first action into the second gradient formula to obtain new first network parameters, and using the new first network parameters to obtain a new first strategy network;

[0206] Inputting the new first network parameters into the second formula to obtain the new network parameters of the first target network, and using the new network parameters of the first target network to obtain the new first target network;

[0207] Among them, the second gradient formula is:

[0208]

[0209] in, For Perform gradient calculation, J(π i1 ) is the action value function under the first policy network, π i1 For the first strategy network, For a i1,t Perform gradient calculation, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, is the first network parameter of the first strategy network, τ i is the strategy network learning rate of multi-functional park i, a ii,t is the action generated when the first strategy network is adopted;

[0210] The second formula is:

[0211]

[0212] in, is the network parameter of the new first target network, and α is the soft update coefficient.

[0213] In a possible implementation, the model training module can also be used to:

[0214] Input the current second action into the third gradient formula to obtain new second network parameters, and use the new second network parameters to obtain a new second strategy network;

[0215] Inputting the new second network parameters into the third formula to obtain the new network parameters of the second target network, and using the new network parameters of the second target network to obtain the new second target network;

[0216] Among them, the third gradient formula is:

[0217]

[0218] in, For Perform gradient calculation, J(π i2 ) is the action value function under the second policy network, π i2 is the second strategy network, For a i2,tPerform gradient calculation, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, is the second network parameter of the second strategy network, τ i is the strategy network learning rate of Multi-Energy Park I, a i2,t is the action generated when the second strategy network is adopted;

[0219] The third formula is:

[0220]

[0221] in, is the network parameter of the new second target network, and α is the soft update coefficient.

[0222] In one possible implementation, the framework building blocks may be used to:

[0223] Obtain the user's electrical load and thermal load demand, photovoltaic power generation, charge state of electric energy storage, the price of selling and purchasing electricity from the grid, and the dispatch period, and construct the state space;

[0224] Obtain the processing of each device in the multi-energy park system as well as the power purchase and sales in the internal market to build an action space.

[0225] In a possible implementation, the device may further include an energy model building module, and the energy model building module may be used to:

[0226] Construct an energy internet system model for the coordinated optimization of multi-energy parks. The energy internet system model takes minimizing the daily system operation cost of a single multi-energy zone as its optimization goal.

[0227] Among them, the objective function of the energy Internet system model is:

[0228] minC i =min(C i,gas +(C i,b -C i,s )+(C i,BES +C i,HES ))

[0229] Among them, C i is the system operation cost of the multi-energy park on day i, C i,gas The cost of purchasing natural gas for multi-energy park i, C i,b is the electricity purchase cost of multi-energy park i, C i,s is the electricity sales revenue of multi-energy park i, C i,BES is the charging and discharging depreciation cost of the electric energy storage in the multi-energy park i, C i,HESis the charging and discharging depreciation cost of the thermal energy storage in multi-energy park i.

[0230] In one possible implementation, the constraints of the energy internet system model may include power balance constraints, equipment operation constraints, and clearing mechanism constraints;

[0231] The power balance constraint is:

[0232]

[0233] Among them, p i,b (t) is the electricity purchase amount of multi-energy park i, p i,s (t) is the electricity sales of multi-energy park i, p i,PV (t) is the output power of photovoltaic power in multi-energy park i in time period t, p i,BES (t) is the charging / discharging power of the multi-energy park i electric energy storage, p i,EB (t) is the input power of electric boiler i in the multi-energy park, p i,load (t) is the electric load of multi-energy park i in time period t, h i,GB (t) is the thermal power output of the gas boiler in multi-energy park i in time period t, h i,EB (t) is the thermal power output of the electric boiler in multi-energy park i in time period t, h i,HES (t) is the charging / discharging power of the thermal energy storage in multi-energy park i during time period t, h i,load (t) is the heat load of multi-energy park i in time period t;

[0234] The equipment operation constraints are:

[0235]

[0236] Among them, h GB (t) is the thermal power output of the gas boiler, is the lower limit of the thermal power output of the gas boiler. is the upper limit of the thermal power output of the gas boiler, h EB (t) is the thermal power output of the electric boiler, is the lower limit of the thermal power output of the electric boiler, is the upper limit of the thermal power output of the electric boiler, p BES (t) is the charging and discharging power of the energy storage, is the lower limit of the energy storage charging / discharging power, is the upper limit of the energy storage charge / discharge power, h HES (t) is the thermal energy storage charging and discharging power, is the lower limit of the thermal energy storage charging / discharging power, is the upper limit of the thermal energy storage charging / discharging power, C SOC (t) is the state of charge of the energy storage device in time period t, is the lower limit of the state of charge of the energy storage device, is the upper limit of the state of charge of the energy storage device;

[0237] The clearing mechanism constraints are:

[0238]

[0239] Among them, ε s (t) is the electricity price in the internal market, ε b (t) is the electricity purchase price in the internal market, ε m (t) is the internal clearing price, p b (t) is the total electricity purchase of all multi-energy parks, p s (t) is the total electricity sales of all multi-energy parks, ε grid,s (t) is the price of electricity sold to the grid, ε grid,b (t) is the price of electricity purchased from the power grid.

[0240] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0241] Those of ordinary skill in the art will appreciate that the templates, units, and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0242] If the module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, it can implement the steps of the above-mentioned embodiments of the multi-regional energy Internet optimization scheduling method. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device that can carry the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal and software distribution medium.

[0243] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A multi-region energy internet optimization scheduling method, characterized in that: include: Construct a multi-agent reinforcement learning framework for a multi-energy park system, wherein the multi-agent reinforcement learning framework includes a state space and an action space, wherein the state space includes the state of each agent at each time point, and the action space includes the action of each agent at each time point, and the agent is the control center of the multi-energy park; Build an energy scheduling model based on multi-agent deep deterministic policy gradient algorithm and Dropout method; Training the energy scheduling model using the state space as input and the action space as output; Based on the trained energy scheduling model, the energy scheduling plan of the multi-energy park system is determined.

2. The multi-regional energy internet optimization scheduling method according to claim 1 is characterized in that: The energy scheduling model includes a first strategy network, a second strategy network and a value network. The state space is used as input and the action space is used as output. Training the energy scheduling model includes: Initialize the first network parameters of the first policy network, the second network parameters of the second policy network, and the third network parameters of the value network, and use the initialized first network parameters as the current first network parameters of the current first policy network, the initialized second network parameters as the current second network parameters of the current second policy network, and the initialized third network parameters as the current third network parameters of the current value network; Assigning the current first network parameter to the current first target network of the current first strategy network, assigning the current second network parameter to the current second target network of the current second strategy network, and assigning the current third network parameter to the current third target network of the current value network; Obtain the current state, input the current state into the current first strategy network, output the current first action, and input the current state into the current second strategy network, output the current second action; Input the current state and the current first action into the current value network to obtain a first output value, input the current state and the current second action into the current value network to obtain a second output value, and select the minimum output value from the first output value and the second output value as the target output value, and use the policy network corresponding to the target output value as the target policy network, and use the action corresponding to the target output value as the target action; After the target policy network executes the target action, a reward value and a new state are obtained, and the third network parameters of the current value network are updated using the current state, the target action, the new state, the reward value, the network parameters of the target network corresponding to the target policy network, and the network parameters of the current third target network to obtain a new value network, and the network parameters of the current third target network are updated using the updated third network parameters to obtain a new third target network; Determine whether the current number of iterations is an odd number of iterations; If the current number of iterations is an odd number of iterations, a new first network parameter is calculated using the current first action to obtain a new first strategy network, and the network parameter of the current first target network is updated using the new first network parameter to obtain a new first target network; Determine whether the current number of iterations has reached the maximum number of iterations; If the current number of iterations reaches the maximum number of iterations, determining the new first strategy network as the trained energy scheduling model; If the current number of iterations does not reach the maximum number of iterations, the current state is updated using the new state, the current first policy network is updated using the new first policy network, the current first target network is updated using the new first target network, the current value network is updated using the new value network, the current third target network is updated using the new third target network, and the current number of iterations is increased by one, and the current state is obtained and returned, the current state is input into the current first policy network, the current first action is output, and the current state is input into the current second policy network, and the current second action step is output to continue execution.

3. The multi-regional energy internet optimization scheduling method according to claim 2 is characterized in that: After determining whether the current number of iterations is an odd number of iterations, the method further includes: If the current number of iterations is not an odd number of iterations, new second network parameters are calculated using the current second action to obtain a new second strategy network, and network parameters of the current second target network are updated using the new second strategy network parameters to obtain a new second target network; Correspondingly, the determining whether the current number of iterations reaches the maximum number of iterations includes: If the current number of iterations reaches the maximum number of iterations, determining the new second strategy network as the trained energy scheduling model; If the current number of iterations does not reach the maximum number of iterations, the current state is updated using the new state, the current second policy network is updated using the new second policy network, the current second target network is updated using the new second target network, the current value network is updated using the new value network, the current third target network is updated using the new third target network, and the current number of iterations is increased by one, and the current state is obtained and returned, the current state is input into the current first policy network, the current first action is output, and the current state is input into the current second policy network, the current second action step is output and continued.

4. The multi-regional energy internet optimization scheduling method according to claim 2 is characterized in that: The method uses the current state, the target action, the new state, the reward value, the network parameters of the target network corresponding to the target policy network, and the network parameters of the current third target network to update the third network parameters of the current value network to obtain a new value network, and uses the updated third network parameters to update the network parameters of the current third target network to obtain a new third target network, including: Inputting the current state, the target action, the new state, the reward value, the network parameters of the target network corresponding to the target policy network, and the network parameters of the current third target network into a minimum loss function; Inputting the minimum loss function into the first gradient formula to obtain new third network parameters, and using the new third network parameters to update the third network parameters of the current value network to obtain the new value network; Inputting the updated third network parameters into the first formula to obtain new network parameters of the third target network, and obtaining the new third target network using the new network parameters of the third target network; Wherein, the minimum loss function is: in, is the minimum loss function, y i,t is the output value, r i,t is the reward value, Q i ′ is the current third target network, s i,t+1 is the state of multi-energy park i at time period t+1, π i ′ k (s i,t+1 ) is the target network corresponding to the target policy network, is the third network parameter of the current value network, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, a i,t is the action of multi-energy park i in time period t, k is the sequence number of the strategy network, and γ is the discount coefficient; The first gradient formula is: in, For Perform gradient calculation, E(·) is the expected function, μ i is the value network learning rate of multi-energy park i, min(·) is the minimum function; The first formula is: in, are the network parameters of the new third target network, and α is the soft update coefficient.

5. The multi-regional energy internet optimization scheduling method according to claim 2 is characterized in that: The method of calculating new first network parameters by using the current first action to obtain a new first strategy network, and updating network parameters of the current first target network by using the new first network parameters to obtain a new first target network includes: Inputting the current first action into the second gradient formula to obtain the new first network parameters, and using the new first network parameters to obtain the new first strategy network; Inputting the new first network parameters into the second formula to obtain the network parameters of the new first target network, and using the network parameters of the new first target network to obtain the new first target network; Among them, the second gradient formula is: in, For Perform gradient calculation, J(π i1 ) is the action value function under the first policy network, π i1 For the first strategy network, For a i1,t Perform gradient calculation, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, is the first network parameter of the first strategy network, τ i is the strategy network learning rate of multi-functional park i, a i1,t is the action generated when the first strategy network is adopted; The second formula is: in, are the network parameters of the new first target network, and α is the soft update coefficient.

6. The multi-regional energy internet optimization scheduling method according to claim 3 is characterized in that: The method of calculating new second network parameters by using the current second action to obtain a new second strategy network, and updating network parameters of the current second target network by using the new second strategy network parameters to obtain a new second target network includes: Inputting the current second action into the third gradient formula to obtain the new second network parameters, and using the new second network parameters to obtain the new second strategy network; Inputting the new second network parameters into a third formula to obtain new network parameters of the second target network, and using the new network parameters of the second target network to obtain the new second target network; Wherein, the third gradient formula is: in, For Perform gradient calculation, J(π i2 ) is the action value function under the second policy network, π i2 is the second strategy network, For a i2,t Perform gradient calculation, Q i is the current value network, s i,t is the state of multi-energy park i in time period t, is the second network parameter of the second strategy network, τ i is the strategy network learning rate of multi-functional park i, a i2,t is the action generated when the second strategy network is adopted; The third formula is: in, are the network parameters of the new second target network, and α is the soft update coefficient.

7. The multi-regional energy internet optimization scheduling method according to claim 1 is characterized in that: The multi-agent reinforcement learning framework for constructing a multi-functional park system includes: Obtain the user's electric load and heat load demand, photovoltaic power generation power, charge state of electric energy storage, the price of selling and purchasing electricity from the power grid, and the dispatching period, and construct the state space; The processing of each device in the multi-energy park system and the power purchase and sales in the internal market are obtained to construct the action space.

8. The multi-regional energy internet optimization scheduling method according to claim 1 is characterized in that: Before constructing the multi-agent reinforcement learning framework of the multi-energy park system, the method further includes: Constructing an energy internet system model for the coordinated optimization operation of multiple energy parks, wherein the energy internet system model takes minimizing the system operation cost within a single multi-energy zone as the optimization goal; Among them, the objective function of the energy Internet system model is: minC i =min(C i,gas +(C i,b -C i,s )+(C i,BES +C i,HEs )) Among them, C i is the system operation cost of the multi-energy park on day i, C i,gas The cost of purchasing natural gas for multi-energy park i, C i,b is the electricity purchase cost of multi-energy park i, C i,s is the electricity sales revenue of multi-energy park i, C i,BES is the charging and discharging depreciation cost of the electric energy storage in the multi-energy park i, C i,HES is the charging and discharging depreciation cost of the thermal energy storage in multi-energy park i.

9. The multi-regional energy internet optimization scheduling method according to claim 8 is characterized in that: The constraints of the energy internet system model include power balance constraints, equipment operation constraints and clearing mechanism constraints; The power balance constraint is: Among them, p i,b (t) is the electricity purchase amount of multi-energy park i, p i,s (t) is the electricity sales of multi-energy park i, p i,PV (t) is the output power of photovoltaic power in multi-energy park i in time period t, p i,BeS (t) is the charging / discharging power of the multi-energy park i electric energy storage, p i,EB (t) is the input power of electric boiler i in the multi-energy park, p i,load (t) is the electric load of multi-energy park i in time period t, h i,GB (t) is the thermal power output of the gas boiler in multi-energy park i in time period t, h i,EB (t) is the thermal power output of the electric boiler in multi-energy park i in time period t, h i,HES (t) is the charging / discharging power of the thermal energy storage in multi-energy park i during time period t, h i,load (t) is the heat load of multi-energy park i in time period t; The equipment operation constraints are: Among them, h GB (t) is the thermal power output of the gas boiler, is the lower limit of the thermal power output of the gas boiler, is the upper limit of the thermal power output of the gas boiler, h EB (t) is the thermal power output of the electric boiler, is the lower limit of the thermal power output of the electric boiler, is the upper limit of the thermal power output of the electric boiler, p BES (t) is the charging and discharging power of the energy storage, is the lower limit of the energy storage charging / discharging power, is the upper limit of the energy storage charge / discharge power, h HES (t) is the thermal energy storage charging and discharging power, is the lower limit of the thermal energy storage charging / discharging power, is the upper limit of the thermal energy storage charging / discharging power, C SOC (t) is the state of charge of the energy storage device in time period t, is the lower limit of the state of charge of the energy storage device, is the upper limit of the state of charge of the energy storage device; The clearing mechanism constraints are: Among them, ε s (t) is the electricity price in the internal market, ε b (t) is the electricity purchase price in the internal market, ε m (t) is the internal clearing price, p b (t) is the total electricity purchase of all multi-energy parks, p s (t) is the total electricity sales of all multi-energy parks, ε grid,s (t) is the price of electricity sold to the grid, ε grid,b (t) is the price of electricity purchased from the power grid.

10. A multi-region energy internet optimization and dispatching device, characterized in that: include: A framework construction module, used to construct a multi-agent reinforcement learning framework for a multi-energy park system, wherein the multi-agent reinforcement learning framework includes a state space and an action space, wherein the state space includes the state of each agent at each time point, and the action space includes the action of each agent at each time point, and the agent is the control center of the multi-energy park; Model building module, used to build energy scheduling model based on multi-agent deep deterministic policy gradient algorithm and Dropout method; A model training module, used for training the energy scheduling model by taking the state space as input and the action space as output; The scheme determination module is used to determine the energy scheduling scheme of the multi-energy park system based on the trained energy scheduling model.