Micro-grid energy management double-layer optimization method and system based on deep reinforcement learning

By applying a two-layer optimization method of deep reinforcement learning in microgrids, the problems of long calculation time and difficulty in cost control caused by uncertainty in microgrid energy management are solved, and a fast and economical energy management strategy is achieved.

CN120165360APending Publication Date: 2025-06-17STATE GRID HUBEI MARKETING SERVICE CENT (MEASUREMENT CENT)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510149514.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The energy management of microgrids faces uncertainty about renewable energy and load requirements, which leads to traditional methods requiring a lot of computing time and difficulty in solving grid scheduling strategies in real time, and difficult to achieve cost-minimization control.

Method used

The two-layer optimization method of microgrid energy management based on deep reinforcement learning is adopted. By constructing Markov decision-making process formulas and deep reinforcement learning agents, combining the upper-level optimization model and the lower-level optimization model, the optimal strategy solution for the microgrid system is achieved.

Benefits of technology

It realizes minimizing operating costs within T periods, simplifies the calculation process, can quickly give optimization strategies, adapt to the real-time state of the microgrid system, and improves the efficiency and economicality of energy management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165360A_ABST
    Figure CN120165360A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-grid energy management double-layer optimization method and system based on deep reinforcement learning. The method comprises the following steps: step 1, constructing a Markov decision process formula of a micro-grid system in T time periods; step 2, constructing a deep reinforcement learning agent as an upper layer optimization model, and solving a scheduling strategy taking maximization of accumulated rewards of the micro-grid system in T time periods as a target; 3, constructing an optimization solver as a lower-layer optimization model, and solving the action of the fuel generator set with the goal of maximizing the tth time period reward rt (st, at); 4, a game model between the upper layer optimization model and the lower layer optimization model is constructed, and the game model is configured in the mode that when the scheduling strategy converges to a target strategy, the target strategy serves as the optimization strategy of the micro-grid system. According to the technical scheme of the invention, rapid, economical and effective scheduling of the micro-grid system under different conditions can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microgrid energy management, and in particular, to a two-layer optimization method and system for microgrid energy management based on deep reinforcement learning. Background Art

[0002] In order to alleviate the shortage of fossil fuels and the environmental pollution caused by the overuse of fossil fuels, renewable energy has received extensive attention and has developed rapidly. Microgrids provide an effective way to integrate renewable energy into the power grid.

[0003] A microgrid is a small power generation and distribution system composed of renewable energy power generation devices, energy storage devices, energy conversion devices, loads, control devices, etc. Relying on technologies such as energy management and operation control, the microgrid can operate economically and stably.

[0004] However, the uncertainty of renewable energy and load demand brings great challenges to the energy management of microgrids. Many different methods have been applied to solve the energy management problem of microgrids. These methods include mathematical programming, robust optimization, heuristic algorithms, and model predictive control algorithms. However, these traditional methods require the establishment of an accurate prediction model for renewable energy and load demand, and another challenge of these methods is the large amount of computing time required. For this reason, there is an urgent need for a new energy management method that can instantaneously solve the grid scheduling strategy and at the same time achieve cost minimization control. Summary of the Invention

[0005] The present invention aims to solve at least one of the problems in the related art to some extent. Embodiments of the present invention provide a two-layer optimization method and system for microgrid energy management based on deep reinforcement learning to solve the optimal strategy of the microgrid system and then optimize the energy management.

[0006] In a first aspect, the present invention provides a two-layer optimization method for microgrid energy management based on deep reinforcement learning, including:

[0007] Step 1, construct a Markov decision process formula for the microgrid system within T time periods: (s, a, p, r); s is the environmental state; a is the action; p is the transition probability, which represents the mapping between the states at two adjacent time steps; r is the reward; where, in the t-th time period, the Markov decision process formula is expressed as: (s t , a t , p t , r t (s t , a t )); s t is the environmental state in the t-th time period; a tis the action of the energy storage system in the t-th period; p t is the environmental state s in the t-th period t transitioning to the environmental state s t+1 is the transition probability; r t (s t , a t ) is the reward in the t-th period;

[0008] Step 2, construct a deep reinforcement learning agent as the upper-layer optimization model. The upper-layer optimization model is based on the SAC algorithm in deep reinforcement learning, solves the scheduling strategy aiming to maximize the cumulative reward of the microgrid system within T periods, and sends the environmental state s t in the t-th period under the scheduling strategy and the action a t of the energy storage system to the lower-layer optimization model;

[0009] Step 3, construct an optimization solver as the lower-layer optimization model. The lower-layer optimization model receives the environmental state s t in the t-th period solved by the upper-layer optimization model and the action a t of the energy storage system, and solves the action of the fuel generator set aiming to maximize the reward r t (s t , a t ) in the t-th period, and feeds it back to the upper-layer optimization model;

[0010] Step 4, construct a game model between the upper-layer optimization model and the lower-layer optimization model. The game model is configured to: when the scheduling strategy converges to the target strategy, take the target strategy as the optimization strategy of the microgrid system.

[0011] Furthermore, the microgrid system includes: a fuel generator set, a photovoltaic system, an energy storage system, and a local load;

[0012] In the t-th period, in the Markov decision process formula:

[0013] The environmental state s t is expressed as:

[0014]

[0015] where, is the output power of the photovoltaic system in the t-th period, is the power demand of the local load in the t-th period, represents the real-time electricity price in the t-th period, S t is the charging state of the energy storage system in the t-th period;

[0016] The action a t of the energy storage system is expressed as:

[0017]

[0018] Among them, is the charge and discharge action of the energy storage system in the t-th period. When is a charging action, When is a discharging action, and represent the charging power and discharging power of the energy storage system in the t-th period. S t is the charging state of the energy storage system in the t-th period. S max and S min are the upper and lower limits allowed for the charging state respectively. is the capacity of the energy storage system. θ ch and θ dis are the charging efficiency and discharging efficiency of the energy storage system. Δt represents the change in time;

[0019] p t represents the probability that the environmental state s t and action a t of the microgrid system in the (t + 1)-th period is s t+1 ' after:

[0020] p t = Probability{s t+1 = s'|s t = s, a t = a}

[0021] Among them, s' is the changed environmental state;

[0022] The reward r t (s t , a t ) in the t-th period is expressed as:

[0023]

[0024] Among them, r t (s t , a t ) is the reward function, which is defined by the negative value of the operating cost of the microgrid system in the t-th period. is the cost of the microgrid system exchanging power with the main grid in the t-th period. is the power generation cost of the i-th unit among N fuel generating units. N is the total number of fuel generating units.

[0025] Furthermore, The calculation formula of

[0026]

[0027] Among them, represents the cost of exchanging power with the main grid in the t-th time period, is the power of the microgrid system exchanging power with the main grid in the t-th time period, indicates purchasing power from the main grid when indicates selling power to the main grid when represents the real-time electricity price in the t-th time period, η represents the discount of the selling price, and Δt represents the change in time.

[0028] Furthermore, step 2 specifically includes:

[0029] Construct the policy function and value function Q(s t , a t ) of the microgrid system according to the cumulative reward within T time periods;

[0030] Using the SAC algorithm in deep reinforcement learning, select any policy from the policy function as the current policy, perform policy evaluation and policy improvement on the current policy, repeat the process of policy evaluation and policy improvement until the policy converges, and use the converged policy as the scheduling policy;

[0031] The process of the policy evaluation is: evaluate the current policy by calculating the value function Q(s t , a t ); the process of the policy improvement is: guide the policy improvement through the calculation result of the value function Q(s t , a t ).

[0032] Furthermore, the solution process of the scheduling policy specifically includes: based on the AC architecture in the SAC algorithm, construct the A network and the C network;

[0033] Use the A network Q ω (s t , a t ) with parameters ω to estimate the policy function and perform policy evaluation;

[0034] Use the C network with parameters to estimate the value function and perform policy improvement; Based on the time difference theory, update the parameters of the C network by minimizing the loss function

[0035] Update the parameters ω of the A network using the minimum loss function based on the KL divergence.

[0036] ​Further, in the game model in Step 4, the upper-layer optimization model is used to send the environmental state s at the t-th time period to the lower-layer optimization model t and the action a of the energy storage system t , and update the optimal scheduling strategy according to the optimal action of the fuel power generation unit fed back by the lower-layer optimization model; the lower-layer optimization model is used to receive the environmental state s at the t-th time period sent by the upper-layer optimization model t and the action a of the energy storage system t , and select the optimal action of the fuel power generation unit and feed it back to the upper-layer optimization model; the target strategy is the optimal scheduling strategy when the upper-layer optimization model and the lower-layer optimization model reach equilibrium.

[0037] Further, during the operation of the microgrid system, the constraint conditions include: fuel power generation unit constraints, energy storage system constraints, main grid constraints, and power balance constraints;

[0038] The fuel power generation unit constraints are:

[0039]

[0040] where, is the active power output of the i-th fuel power generation unit at the t-th time period, are respectively the minimum and maximum active power outputs of the i-th fuel power generation unit; is the maximum DG ramp limit of the i-th fuel power generation unit;

[0041] The energy storage system constraints are:

[0042]

[0043] S min ≤S t ≤S max

[0044] where, and represent the charging power and discharging power of the energy storage system at the t-th time period, P ES is the maximum charging / discharging power of the energy storage system, means that the energy storage system cannot charge and discharge simultaneously, S t represents the charging state of the energy storage system at the t-th time period, S t-1 represents the charging state of the energy storage system at the (t - 1)-th time period, represents the capacity of the energy storage system, θ ch and θ dis represent the charging efficiency and discharging efficiency of the energy storage system, Δt represents the change in time, S max and S minThey are the upper and lower limits allowed for the charging state respectively;

[0045] The main grid constraint is:

[0046]

[0047] Wherein, is the power for the microgrid system to exchange power with the main grid, is the maximum power to exchange power with the main grid at the point of common coupling;

[0048] The power balance constraint is:

[0049]

[0050] Wherein, is the active power output of the i-th fuel power generation unit in the t-th time period, and represent the charging and discharging power of the energy storage system in the t-th time period, is the output power of the photovoltaic system in the t-th time period, is the power for the microgrid system to exchange power with the main grid in the t-th time period, is the power demand of the local load in the t-th time period.

[0051] In a second aspect, embodiments of the present disclosure provide a two-layer optimization system for microgrid energy management based on deep reinforcement learning. The system can implement the method provided in the first aspect. The system includes:

[0052] A first construction unit for constructing a Markov decision process formula (s, a, p, r) of the microgrid system within T time periods; s is the environmental state; a is the action; p is the transition probability, indicating a mapping between the states at two adjacent time steps; r is the reward; wherein, in the t-th time period, the Markov decision process formula is expressed as: (s t , a t , p t , r t (s t , a t )); s t is the environmental state in the t-th time period; a t is the action of the energy storage system in the t-th time period; p t is the transition probability from the environmental state s t to the environmental state s t+1 in the t-th time period; r t (s t , a t ) is the reward in the t-th time period;

[0053] The second construction unit is used to construct a deep reinforcement learning agent as the upper-layer optimization model. The upper-layer optimization model is based on the SAC algorithm in deep reinforcement learning, solves the scheduling strategy aiming to maximize the cumulative reward of the microgrid system within T time periods, and sends the environmental state s t at the t-th time period under the scheduling strategy t and the action a of the energy storage system to the lower-layer optimization model;

[0054] The third construction unit is used to construct an optimization solver as the lower-layer optimization model. The lower-layer optimization model receives the environmental state s t at the t-th time period solved by the upper-layer optimization model t and the action a of the energy storage system, and solves the action of the fuel power generation unit aiming to maximize the reward r t (s t , a t ) at the t-th time period, and feeds it back to the upper-layer optimization model;

[0055] The fourth construction unit is used to construct a game model between the upper-layer optimization model and the lower-layer optimization model. The game model is configured to: when the scheduling strategy converges to the target strategy, use the target strategy as the optimization strategy of the microgrid system.

[0056] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:

[0057] One or more processors;

[0058] A memory for storing one or more programs;

[0059] When the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the first aspect.

[0060] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method provided in the first aspect.

[0061] Compared with the prior art, the beneficial effects of the present invention are:

[0062] An embodiment of the present invention provides a two - layer optimization method and system for micro - grid energy management based on deep reinforcement learning. The upper - layer deep reinforcement learning agent is responsible for arranging appropriate actions of the energy storage to maximize the cumulative reward. Based on the output of the upper layer, the lower - layer non - linear optimization solver aims to select the best actions of the fuel - generating units to maximize the immediate reward and return the result to the upper layer. After the game model is trained, it can quickly give the final optimization strategy according to the immediate state of the micro - grid system to minimize the operating cost within T time periods.

[0063] Other features and advantages of the present invention will be described in the following specification, and will, in part, be obvious from the specification, or can be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only the preferred embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0065] Figure 1 is a flowchart of a two - layer optimization method for micro - grid energy management based on deep reinforcement learning provided by an embodiment of the present invention;

[0066] Figure 2 is a schematic diagram of the model training and solving mechanism of the two - layer optimization model provided by an embodiment of the present invention;

[0067] Figure 3 is a daily reward convergence curve graph of the micro - grid system provided by an embodiment of the present invention;

[0068] Figure 4 is a schematic diagram of the typical daily electricity price, photovoltaic power generation, load demand, and output provided by an embodiment of the present invention;

[0069] Figure 5 is a schematic diagram of the 7 - day cumulative operating cost provided by an embodiment of the present invention;

[0070] Figure 6 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0071] The following describes the principles and features of the present invention with reference to the drawings. The listed embodiments are only used to explain the present invention and are not used to limit the scope of the present invention.

[0072] Unless otherwise defined, the technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar words used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, words such as "a", "an" or "the" do not denote a quantity limitation, but mean that there is at least one. Words such as "comprising" or "including" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.

[0073] In the respective drawings, like elements are denoted by like reference numerals. For the sake of clarity, not all parts in the drawings are drawn to scale. In addition, some well-known parts may not be shown in the figures.

[0074] Many specific details of this disclosure are described below in order to understand this disclosure more clearly. However, as those skilled in the art can understand, this disclosure can be implemented without these specific details.

[0075] Embodiment 1

[0076] Figure 1 The flowchart of a two-layer optimization method for microgrid energy management based on deep reinforcement learning provided for Embodiment 1 of the present invention is as Figure 1 shown, and this method includes:

[0077] Step 1, construct the Markov decision process formula for the microgrid system within T time periods: (s, a, p, r); s is the environmental state; a is the action; p is the transition probability, representing the mapping between the states at two adjacent time steps; r is the reward.

[0078] For the sake of simplifying the solution, in Embodiment of the present invention, for the t-th time period of the operation of the microgrid system, the Markov decision process formula is transformed, and it is expressed as: (s t , a t , p t , r t (s t , a t ))).

[0079] Among them, s t is the environmental state of the t-th time period, a t is the action of the energy storage system in the t-th time period, p t is the transition probability from the environmental state s t to the environmental state s t+1 in the t-th time period, r t (s t , at ) is the reward for the t-th time period.

[0080] Before proceeding to the next step, it should be noted that: The microgrid system includes: a fuel power generation unit, a photovoltaic system, an energy storage system, and local loads. For example: The microgrid system can consist of several fuel generators, a photovoltaic system, an energy storage system, and some local loads, operating in grid-connected mode.

[0081] Table 1: Microgrid operating parameters

[0082]

[0083] In the t-th time period, in the Markov decision process formula:

[0084] Environmental state s t Is expressed as:

[0085]

[0086] Where, Is the output power of the photovoltaic system in the t-th time period, Is the power demand of the local load in the t-th time period, Represents the real-time electricity price in the t-th time period, S t Is the charge state of the energy storage system in the t-th time period.

[0087] Action a of the energy storage system t Is expressed as:

[0088]

[0089] Where, Is the charge and discharge action of the energy storage system in the t-th time period. When Is a charging action, When Is a discharging action, And Represent the charging power and discharging power of the energy storage system in the t-th time period, S t Is the charge state of the energy storage system in the t-th time period, S max And S min Are respectively the upper and lower limits allowed for the charge state, Is the capacity of the energy storage system, θ ch And θ dis Are the charging efficiency and discharging efficiency of the energy storage system, and Δt represents the change in time.

[0090] p t Represents given the environmental state s in the t-th time period t And action a tAfter that, the probability that the environmental state s of the microgrid system in the (t + 1)-th period is s': t+1 is:

[0091] p t = Probability{s t+1 = s'|s t = s,a t = a}

[0092] where s' is the changed environmental state.

[0093] The reward r in the t-th period t (s t ,a t ) is expressed as:

[0094]

[0095] where r t (s t ,a t ) is the reward function, defined by the negative value of the operating cost of the microgrid system in the t-th period, is the cost of the microgrid system exchanging power with the main grid in the t-th period, is the power generation cost of the i-th unit among N fuel power generation units, and N is the total number of fuel power generation units.

[0096] It should be noted that the goal of microgrid energy management is to minimize the operating cost within T periods, including the cost of exchanging power with the main grid and the power generation cost of fuel power generation units. Minimizing the operating cost within T periods is expressed as:

[0097]

[0098] where, represents the cost of the microgrid system exchanging power with the main grid in the t-th period, represents the power generation cost of the i-th unit among N fuel power generation units.

[0099] The calculation formula of

[0100]

[0101] where, represents the cost of the microgrid system exchanging power with the main grid in the t-th period, is the power of the microgrid system exchanging power with the main grid in the t-th period, means purchasing power from the main grid when means selling power to the main grid when Let \(p_t\) denote the real-time electricity price in the \(t\)-th time period, \(\eta\) denote the discount of the selling price, and \(\Delta t\) denote the change in time.

[0102] The calculation formula of \(p_t\) is:

[0103]

[0104] Wherein, \(P_{i,t}\) is the active power output of the \(i\)-th fuel power generation unit in the \(t\)-th time period, \(a_i\), \(b_i\), \(c_i\), \(d_i\), \(e_i\), \(f_i\) are all constants.

[0105] In addition, during the operation of the microgrid system, certain constraint conditions need to be satisfied.

[0106] The constraint conditions include: fuel power generation unit constraints, energy storage system constraints, main grid constraints, and power balance constraints.

[0107] The fuel power generation unit constraints are:

[0108]

[0109] Wherein, \(P_{i,t}\) is the active power output of the \(i\)-th fuel power generation unit in the \(t\)-th time period, \(P_{i,\min}\) and \(P_{i,\max}\) are the minimum and maximum active power outputs of the \(i\)-th fuel power generation unit respectively; \(R_i\) is the maximum DG ramp rate limit of the \(i\)-th fuel power generation unit.

[0110] The energy storage system constraints are:

[0111]

[0112] \(S_{t}^{\text{ch}}\) min \(\leq S_{t}^{\text{ch}}\) t \(\leq S_{t}^{\text{ch}}\) max

[0113] Wherein, \(S_{t}^{\text{ch}}\) and \(S_{t}^{\text{dch}}\) represent the charging power and discharging power of the energy storage system in the \(t\)-th time period, \(P_{\text{max}}\) ES is the maximum charging / discharging power of the energy storage system, \(S_{t}^{\text{ch}}\cdot S_{t}^{\text{dch}} = 0\) means that the energy storage system cannot charge and discharge simultaneously, \(S_t\) t represents the charging state of the energy storage system in the \(t\)-th time period, \(S_{t - 1}\) t-1 represents the charging state of the energy storage system in the \((t - 1)\)-th time period, \(S_{\text{max}}\) represents the capacity of the energy storage system, \(\theta_{\text{ch}}\) ch and \(\theta_{\text{dch}}\) dis represent the charging efficiency and discharging efficiency of the energy storage system, \(\Delta t\) represents the change in time, \(S_{\text{max}}\) max and \(S_{\text{min}}\) min are the allowable upper and lower limits of the charging state respectively.

[0114] The main grid constraint is:

[0115]

[0116] Among them, is the power for the microgrid system to exchange electricity with the main grid, is the maximum power to exchange electricity with the main grid at the point of common coupling.

[0117] The power balance constraint is:

[0118]

[0119] Among them, is the active power output of the i-th fuel power generation unit in the t-th time period, and represent the charging and discharging power of the energy storage system in the t-th time period, is the output power of the photovoltaic system in the t-th time period, is the power for the microgrid system to exchange electricity with the main grid in the t-th time period, is the power demand of the local load in the t-th time period.

[0120] Step 2: Construct a deep reinforcement learning agent as the upper-layer optimization model. The upper-layer optimization model is based on the SAC algorithm in deep reinforcement learning, solves the scheduling strategy aiming to maximize the cumulative reward of the microgrid system within T time periods, and sends the environmental state s t at the t-th time period under the scheduling strategy and the action a t of the energy storage system to the lower-layer optimization model.

[0121] Step 2 specifically includes:

[0122] Step 201: Construct the policy function and value function Q(s t ,a t ) of the microgrid system according to the cumulative reward within T time periods.

[0123] A: For the policy function, this embodiment of the present invention provides multiple different optional formulas. As follows:

[0124] Formula A1:

[0125]

[0126] Among them, argmax represents that under the current policy, the cumulative reward reaches the maximum value; ρ policy is in the state s t - action a tThe policy distribution for time synchronization, where E represents the mathematical expectation.

[0127] Formula A2:

[0128]

[0129] Formula A2 is an optimization of Formula A1, replacing the immediate reward r at the t-th time period t (s t ,a t ) with r t (s t ,a t ) + βH(policy(·|s t ))).

[0130] Among them, β is the temperature coefficient that determines the importance of entropy; H(policy(·|s t ))) is the entropy of the policy policy. By adding an entropy term to the reward function, it is possible to maximize the cumulative reward while optimizing the policy. In addition, this formula can solve the Markov decision process under continuous actions in the present invention.

[0131] Formula A3;

[0132]

[0133] To adapt to more complex tasks, in the SAC algorithm, an energy-based model can also be used to represent the policy function, β is the temperature coefficient that determines the importance of entropy, and Q(s t ,a t ) is the value function.

[0134] Formula A4;

[0135]

[0136] Formula A4 is an optimization of Formula A3. The ideal policy of SAC is the energy-based policy in Formula A3, but it cannot be directly sampled. Therefore, a normal distribution with mean μ and standard deviation σ is applied to replace the energy-based policy, and the KL divergence is used to narrow the gap between the normal distribution and the energy-based policy.

[0137] Among them, Ω is the set of optional policies, which is a set of normal distributions with parameters μ and σ; Z is a log partition function.

[0138] B: For the value function, the embodiments of the present invention provide a variety of different optional formulas. As follows:

[0139] Formula B1:

[0140]

[0141] Among them, formula B1 is used in conjunction with formula A2. β is the temperature coefficient that determines the importance of entropy; H(policy(·|s t ))] is the entropy of the policy policy; λ t is a discount factor used to determine future rewards, greater than or equal to 0 and less than or equal to 1, s0 represents the initial environmental state, and a0 represents the initial action.

[0142] Formula B2:

[0143]

[0144] Formula B2 updates the Q-function value using the Bellman equation to perform Q-function value iteration and can be used in conjunction with formula A3 or A4.

[0145] Step 202: Use the SAC algorithm in deep reinforcement learning to select any policy from the policy function as the current policy, perform policy evaluation and policy improvement on the current policy, repeat the process of policy evaluation and policy improvement until the policy converges, and use the converged policy as the scheduling policy.

[0146] Among them, the process of policy evaluation is: evaluate the current policy by calculating the value function Q(s t ,a t ); the process of policy improvement is: guide the policy improvement through the calculation result of the value function Q(s t ,a t ).

[0147] Preferably, in the process of solving the scheduling policy, step 202 specifically includes:

[0148] Step 2021: Based on the AC architecture in the SAC algorithm, construct an A network and a C network.

[0149] Table 2: Hyperparameters of A and C deep neural networks

[0150] Description Concept Value Duration T 24 Maximum number of times E 800 Initial network update times M 16 Interaction times with the environment L 50 Discount factor λ 0.99 Learning rate of network A <![CDATA[β policy > 0.0001 Learning rate of network C <![CDATA[β Q > 0.0001 Number of hidden layers - 3 Number of neurons in each hidden layer - 256 Mini-batch size - 1024 Size of replay buffer - 1000000

[0151] Both the A and C deep neural networks have three fully connected hidden layers, with 256 neurons in each layer. The activation function of the hidden layer is ReLU. The output layers of the A network use the tanh and Softplus activation functions as the mean and standard respectively, and the output layer of the C network uses the Softplus activation function.

[0152] Step 2022: Use the A network Q with parameter ω ω (s t ,a t ) to estimate the policy function and perform policy evaluation.

[0153] Step 2023: Use the parameter C Network To estimate the value function and make strategy improvements.

[0154] Step 2024: Based on the time difference theory, update the parameters of the C network by minimizing the loss function. The parameters ω of the A network are updated using the minimum loss function based on KL divergence.

[0155] Based on the time difference theory, the C network is updated by minimizing the loss function. According to formula B2, the loss function of the C network in training can be defined as:

[0156]

[0157] Among them, B is the experience replay buffer, which saves the state, action, and reward obtained from each interaction with the environment; and They are two C networks with parameters ω1 and ω2 respectively; and There are two parameters and target network.

[0158] The input to the actor network is the state s t , the output of the network is the mean μ and standard deviation σ of the normal distribution representing the action distribution. Since it is not possible to sample directly from the normal distribution, a reparameterization technique is used to sample the action:

[0159]

[0160] Among them, the function represents the actor network with output parameters μ and σ; ε t is the noise sampled from a standard normal distribution.

[0161] The loss function of the A network during training can be defined as:

[0162]

[0163] In addition, the hyperparameter β is automatically adjusted and the loss function is defined as:

[0164]

[0165] where ψ is a predefined threshold for the minimum policy entropy.

[0166] Step 3: Construct an optimization solver as the lower optimization model. The lower optimization model receives the environmental state s of the tth period solved by the upper optimization model. tWith the operation a of the energy storage system t , and solve to maximize the reward r in the t-th period t (s t , a t ) as the operation of the fuel power generation unit, and feedback it to the upper-layer optimization model.

[0167] The goal of the lower-layer optimization model is to optimize the operation of the fuel power generation unit to minimize the operating cost in the t-th period:

[0168]

[0169] The lower-layer optimization model is used to solve the non-linear programming problem of the fuel power generation unit and return the result to the upper-layer optimization model to calculate the reward in the t-th period. If the reward calculated by the lower-layer optimization model cannot meet the reward in the t-th period of the scheduling strategy provided by the upper-layer optimization model, the upper-layer optimization model needs to update the scheduling strategy.

[0170] Step 4: Construct a game model between the upper-layer optimization model and the lower-layer optimization model. The game model is configured to: when the scheduling strategy converges to the target strategy, use the target strategy as the optimization strategy of the microgrid system.

[0171] In the game model in Step 4:

[0172] The upper-layer optimization model is used to send the environmental state s in the t-th period t and the operation a of the energy storage system t to the lower-layer optimization model, and update the scheduling strategy according to the optimal operation of the fuel power generation unit feedback by the lower-layer optimization model.

[0173] The lower-layer optimization model is used to receive the environmental state s in the t-th period sent by the upper-layer optimization model t and the operation a of the energy storage system t , and select the optimal operation of the fuel power generation unit and feedback it to the upper-layer optimization model.

[0174] The target strategy is the optimal scheduling strategy when the upper-layer optimization model and the lower-layer optimization model reach equilibrium.

[0175] Figure 2 is a schematic diagram of the model training and solving mechanism of the two-layer optimization model provided by the embodiments of the present invention. As Figure 2 shown, the upper-layer optimization model is responsible for arranging the appropriate operation of the energy storage to maximize the cumulative reward. Based on the output of the upper layer, the lower-layer non-linear programming solver aims to select the operation of the fuel power generation unit to maximize the immediate reward and return the result to the upper layer.

[0176] During the offline training process of the two-layer optimization model, the actions of the energy storage can be dynamically selected according to historical experience, and the actions of the fuel generator sets under different energy storage actions can be optimized, the optimization results can be compared, the proxy topology selection can be continuously modified, and the final optimization strategy can be converged to.

[0177] During the online solution process, the trained intelligent agent, that is, the two-layer optimization model, can give the final optimization result according to the immediate state.

[0178] (1) Offline training results

[0179] In each training scenario, one day is randomly selected from the test set to evaluate the performance of the participant network, and the SAC is used to solve the energy management problem of the microgrid and compared with the method in the present invention.

[0180] Figure 3 is the daily reward convergence curve graph of the microgrid system provided by the embodiment of the present invention. Combining Figure 3 , the convergence curve obtained by using the method of the present invention is obviously above the convergence curve obtained only by using the SAC method, which means that the method of the present invention can obtain a higher reward value.

[0181] In the exploration and learning stage, since the generator set behavior obtained by the method in the present invention is optimal, while the generator set behavior obtained only by using the SAC method is still in the exploration stage, the agent of this method obtains a higher daily reward than the agent only using SAC. Therefore, a large number of ineffective searches are avoided and the convergence speed of the algorithm is improved. Combining Figure 3 , although after 100 and 200 operations, the method proposed in the present application and the method only using SAC converge within a certain range, obviously, the method proposed in the present application is more stable. During most of the training time, the method of the present application is better than the method only using SAC.

[0182] (2) Online solution results

[0183] The well-trained intelligent agent is applied to solve the microgrid energy management problem, and a typical day is selected to test the performance of this method.

[0184] Figure 4 is the schematic diagram of the typical day electricity price, photovoltaic power generation, load demand and output provided by the embodiment of the present invention. As Figure 4 shown, seven different conditions are selected, and the cumulative daily operating costs of the method proposed in the present application and the other two methods are as Figure 5 shown.

[0185] Figure 5It is a schematic diagram of the 7-day cumulative operating cost provided by the embodiments of the present invention. The method proposed in this application is superior to the method that only uses SAC. Compared with the calculation results of the mathematical programming method, the method proposed in this application only reduces by 14.4%. However, the mathematical programming method requires accurate prediction of the next 24 hours, and its result is only an ideal optimum. While the method proposed in this application can obtain similar results based on only immediate information.

[0186] In view of the uncertainty existing in the microgrid energy management problem, the embodiments of the present invention propose a two-layer optimization method based on deep reinforcement learning. The upper layer uses the SAC algorithm to optimize the charging and discharging actions of the microgrid, and the lower layer uses a nonlinear programming solver to solve the role of the fuel generator set. Through testing on the microgrid, the training results show that the technical solution in the embodiments of the present invention has a fast convergence speed and stable performance. By introducing the lower-layer nonlinear programming, a large number of ineffective searches are avoided, the design of the deep reinforcement learning reward function is simplified, and the search speed and quality are improved. According to the test result analysis, it shows that the present invention can flexibly select the actions of the energy storage and the fuel generator set according to the electricity price, load demand, and photovoltaic power generation output, and perform economic and effective scheduling under different conditions of the microgrid.

[0187] Embodiment 2

[0188] The embodiments of the present invention also provide a two-layer optimization system for microgrid energy management based on deep reinforcement learning. The system includes:

[0189] A first construction unit, configured to construct a Markov decision process formula for the microgrid system within T time periods: (s, a, p, r); s is the environmental state; a is the action; p is the transition probability, indicating the mapping between the states at two adjacent time steps; r is the reward; wherein, at the t-th time period, the Markov decision process formula is expressed as: (s t , a t , p t , r t (s t , a t )); s t is the environmental state at the t-th time period; a t is the action of the energy storage system at the t-th time period; p t is the transition probability from the environmental state s t to the environmental state s t+1 at the t-th time period; r t (s t , a t ) is the reward at the t-th time period;

[0190] The second construction unit is used to construct a deep reinforcement learning agent as the upper-layer optimization model. The upper-layer optimization model is based on the SAC algorithm in deep reinforcement learning, solves the scheduling strategy aiming to maximize the cumulative reward of the microgrid system within T time periods, and sends the environmental state s t at the t-th time period under the scheduling strategy and the action a t of the energy storage system to the lower-layer optimization model;

[0191] The third construction unit is used to construct an optimization solver as the lower-layer optimization model. The lower-layer optimization model receives the environmental state s t at the t-th time period solved by the upper-layer optimization model and the action a t of the energy storage system, and solves the action of the fuel power generation unit aiming to maximize the reward r t (s t , a t ), and feeds it back to the upper-layer optimization model;

[0192] The fourth construction unit is used to construct a game model between the upper-layer optimization model and the lower-layer optimization model. The game model is configured to: when the scheduling strategy converges to the target strategy, use the target strategy as the optimization strategy of the microgrid system.

[0193] For the specific descriptions of the above modules, please refer to the content in the previous embodiments and will not be elaborated here.

[0194] Embodiment 3

[0195] Based on the same inventive concept, an embodiment of the present invention also provides an electronic device. Figure 6 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. As Figure 6 shown, an electronic device provided by an embodiment of the present invention includes: one or more processors 101, a memory 102, and one or more N / O interfaces 103. One or more programs are stored on the memory 102. When the one or more programs are executed by the one or more processors, the one or more processors implement the electric vehicle aggregator charging service pricing optimization method as described in any of the above embodiments; one or more N / O interfaces 103 are connected between the processor and the memory and are configured to implement information interaction between the processor and the memory.

[0196] Among them, the processor 101 is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory 102 is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the N / O interface (read / write interface) 103 is connected between the processor 101 and the memory 102 and can realize the information interaction between the processor 101 and the memory 102, including but not limited to a data bus (Bus), etc.

[0197] In some embodiments, the processor 101, the memory 102, and the N / O interface 103 are interconnected through a bus 104 and are further connected to other components of the computing device.

[0198] In some embodiments, the one or more processors 101 include a field programmable gate array.

[0199] According to an embodiment of the present invention, there is also provided a computer-readable medium. A computer program is stored on the computer-readable medium, where, when the program is executed by a processor, it implements the steps in any one of the above-mentioned embodiments of the charging service pricing optimization method for an electric vehicle aggregator.

[0200] Specifically, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, including a computer program carried on a machine-readable medium, where the computer program includes program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it executes the above-mentioned functions defined in the system of the present invention.

[0201] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0202] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the foregoing module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0203] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A two-layer optimization method for microgrid energy management based on deep reinforcement learning, characterized in that: The method comprises: Step 1: Construct the Markov decision process formula of the microgrid system in T time periods: (s, a, p, r); s is the environment state; a is the action; p is the transition probability, which means mapping between the states of two adjacent time steps; r is the reward; Among them, in the tth time period, the Markov decision process formula is expressed as: (s t ,a t ,p t ,r t (s t ,a t ));s t is the environmental state at the tth period; a t is the action of the energy storage system in the tth period; p t is the environmental state s in the tth period t Turn to environment state s t+1 The transition probability of t (s t ,a t ) is the reward for the tth period; Step 2: Construct a deep reinforcement learning agent as the upper optimization model. The upper optimization model is based on the SAC algorithm in deep reinforcement learning to solve the scheduling strategy with the goal of maximizing the cumulative reward of the microgrid system in T time periods, and calculate the environmental state s in the tth time period under the scheduling strategy. t Action of the energy storage system t Send to the lower optimization model; Step 3: construct an optimization solver as a lower-level optimization model, which receives the environmental state s of the tth period solved by the upper-level optimization model. t Action of the energy storage system t , and solve to maximize the reward r in the tth period t (s t ,a t ) is the action of the target fuel generator set, and is fed back to the upper optimization model; Step 4: construct a game model between the upper-layer optimization model and the lower-layer optimization model, wherein the game model is configured such that when the dispatching strategy converges to a target strategy, the target strategy is used as the optimization strategy for the microgrid system.

2. The method according to claim 1, characterized in that The microgrid system includes: a fuel generator set, a photovoltaic system, an energy storage system and a local load; In the tth period, in the Markov decision process formula: Environmental status t It is expressed as: Among them, P t PV is the output power of the photovoltaic system in the tth period, P t 2 t is the power demand of the local load in the tth period, represents the real-time electricity price in the tth period, S t is the charging state of the energy storage system in the tth period; Energy storage system action t It is expressed as: Among them, P t Es It is the charging and discharging action of the energy storage system in the tth period. t Es When charging, P t Es >0,P t ch =P t Es , P t dis =0; when P t Es During discharge operation, P t Es <0, P t dis =-P t Es , P t ch =0;P t ch and P t dis represents the charging power and discharging power of the energy storage system in the tth period, S t is the charging state of the energy storage system in the tth period, S max and S min are the upper and lower limits of the charging state, is the capacity of the energy storage system, θ ch and θ dis are the charging efficiency and discharging efficiency of the energy storage system, and Δt represents the change in time; p t Represents the environmental state s at a given time period t t and action a t After that, the environmental state s of the microgrid system in the t+1th period t+1 The probability of being s' is: p t =Probability{s t+1 =s'|s t =s,a t =a} Among them, s' is the environmental state after the change; The reward r in the tth period t (s t ,a t ) is expressed as: Among them, r t (s t ,a t ) is the reward function, defined by the negative value of the microgrid system operating cost in the tth period, is the cost of exchanging electricity between the microgrid system and the main grid during the tth period, is the power generation cost of the i-th unit among N fuel power generation units, and N is the total number of fuel power generation units.

3. The method according to claim 2, characterized in that The calculation formula is: in, It is expressed as the cost of exchanging electricity with the main grid in the tth period, P t 1 is the power exchanged between the microgrid system and the main grid in the tth period, P t 1 >0 means purchasing electricity from the main grid, P t 1 When ≤0, it means selling electricity to the main grid. It represents the real-time electricity price in the tth period, η represents the discount of the selling price, and Δt represents the change in time.

4. The method according to claim 1, characterized in that Step 2 specifically includes: According to the cumulative rewards in T time periods, the strategy function and value function Q(s) of the microgrid system are constructed. t ,a t ); Using the SAC algorithm based on deep reinforcement learning, any strategy is selected from the strategy function as the current strategy, strategy evaluation and strategy improvement are performed on the current strategy, and the strategy evaluation and strategy improvement process is repeated until the strategy converges, and the strategy at the time of convergence is used as the scheduling strategy; The process of strategy evaluation is as follows: by calculating the value function Q(s t ,a t ) evaluates the current strategy; the process of strategy improvement is: through the value function Q(s t ,a t ) calculation results guide strategy improvement.

5. The method according to claim 4, characterized in that The solution process of the scheduling strategy specifically includes: constructing the A network and the C network based on the AC architecture in the SAC algorithm; Using A network Q with parameters ω ω (s t ,a t ) to estimate the policy function and perform policy evaluation; Use with parameters C Network To estimate the value function and improve the strategy; Based on the time difference theory, the parameters of the C network are updated by minimizing the loss function The parameters ω of the A network are updated using the minimum loss function based on KL divergence.

6. The method according to claim 1, characterized in that In the game model in step 4, the upper optimization model is used to send the environment state s of the tth period to the lower optimization model. t Action of the energy storage system t , and, updating the scheduling strategy according to the action of the optimal fuel generator set fed back by the lower optimization model; the lower optimization model is used to receive the environmental state s of the tth period sent by the upper optimization model t Action of the energy storage system t , and, selecting the optimal action of the fuel generator set and feeding it back to the upper optimization model; the target strategy is the optimal scheduling strategy when the upper optimization model and the lower optimization model reach equilibrium.

7. The method according to claim 1, characterized in that During the operation of the microgrid system, the constraints include: fuel generator constraints, energy storage system constraints, main grid constraints, and power balance constraints; The fuel generator set constraints are: in, is the active power output of the i-th fuel generator set in the t-th period, P i DGMIN , P i DGMAX are the minimum and maximum active output of the i-th fuel generator set; P i DGR is the maximum DG ramp limit of the ith fuel generator set; The energy storage system constraints are: 0≤P t ch ,P t dis ≤P ES P t ch P t dis =0 S min ≤S t ≤S max Among them, P t ch and P t dis represents the charging power and discharging power of the energy storage system in the tth period, P ES is the maximum charging / discharging power of the energy storage system, P t ch P t dis =0 means the energy storage system cannot charge and discharge at the same time, S t It is represented as the charging state of the energy storage system in the tth period, S t-1 It is represented as the charging state of the energy storage system in the t-1th period, represents the capacity of the energy storage system, θ ch and θ dis represents the charging efficiency and discharging efficiency of the energy storage system, Δt represents the change in time, S max and S min are the upper and lower limits of the allowed state of charge, respectively; The main grid constraints are: Among them, P t 1 It is the power of the microgrid system exchanging electricity with the main grid. It is the maximum power that exchanges electricity with the main grid at the point of common coupling; The power balance constraint is: in, is the active power output of the i-th fuel generator set in the t-th period, P t ch and P t dis represents the charging and discharging power of the energy storage system in the tth period, P t PV is the output power of the photovoltaic system in the tth period, P t 1 is the power exchanged between the microgrid system and the main grid in the tth period, P t 2 is the power demand of the local load in the tth period.

8. A two-layer optimization system for microgrid energy management based on deep reinforcement learning, characterized in that: The system can implement the method according to any one of claims 1 to 7, and the system comprises: The first construction unit is used to construct the Markov decision process formula of the microgrid system in T time periods: (s, a, p, r); s is the environment state; a is the action; p is the transition probability, which means mapping between the states of two adjacent time steps; r is the reward; Among them, in the tth time period, the Markov decision process formula is expressed as: (s t ,a t ,p t ,r t (s t ,a t ));s t is the environmental state at the tth period; a t is the action of the energy storage system in the tth period; p t is the environmental state s in the tth period t Turn to environment state s t+1 The transition probability of t (s t ,a t ) is the reward for the tth period; The second construction unit is used to construct a deep reinforcement learning agent as an upper optimization model. The upper optimization model is based on the SAC algorithm in deep reinforcement learning to solve the scheduling strategy with the goal of maximizing the cumulative reward of the microgrid system in T time periods, and the environmental state s in the tth time period under the scheduling strategy t Action of the energy storage system t Send to the lower optimization model; The third construction unit is used to construct an optimization solver as a lower-level optimization model, wherein the lower-level optimization model receives the environmental state s of the tth period solved by the upper-level optimization model. t Action of the energy storage system t , and solve to maximize the reward r in the tth period t (s t ,a t ) is the action of the target fuel generator set, and is fed back to the upper optimization model; The fourth construction unit is used to construct a game model between the upper optimization model and the lower optimization model, and the game model is configured to: when the scheduling strategy converges to the target strategy, use the target strategy as the optimization strategy of the microgrid system.

9. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as claimed in any one of claims 1 to 7.

10. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Multi-microgrid system layered reinforcement learning optimization method and system, and storage medium

    CN115115211A

  • Full-electric ship power generation and navigation scheduling joint optimization method based on deep reinforcement learning

    CN115841075A

  • System and method for real-time distributed micro-grid optimization using price signals

    US20230024900A1