An optimal strategy method for distribution network-microgrid collaboration based on multi-agent deep reinforcement learning robust reward function

By introducing a multi-agent deep reinforcement learning model in the distribution grid-microgrid system, the charging and discharging strategies of the energy storage system are trained, and the problem of traditional optimization methods ignoring the impact of energy storage and distributed power supplies is solved, achieving a more efficient and safe collaborative optimization effect.

CN118917678BActive Publication Date: 2025-05-23HOHAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410823737.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2025-05-23
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

Traditional optimization methods can only perform robust calculations of single-time sections, ignoring the impact of energy storage and distributed power supply on subsequent states, resulting in a decrease in microgrid operating income and an increase in distribution network operating costs.

Method used

The robust reward function based on multi-agent deep reinforcement learning is adopted, and the charging and discharging power of the energy storage system in the microgrid is trained through the multi-agent deep reinforcement learning model, and combined with the microgrid operation constraints and marginal electricity price of the distribution network nodes, the optimal strategy of synergy between the distribution network-microgrid is realized.

Benefits of technology

It improves the robustness and economy of multi-time sectional optimization of distribution grid-microgrids, ensures the safety of coordinated operation, and improves the comprehensive benefits of microgrids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118917678B_ABST
    Figure CN118917678B_ABST
Patent Text Reader

Abstract

The present invention discloses a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function. In the microgrid, the non-cooperative game strategy of the microgrid is mined through multi-agent deep reinforcement learning, and the strategy that minimizes the distribution network operation cost and maximizes the comprehensive benefit of the microgrid is accurately given on the basis of considering future benefits, thereby realizing the distribution network-microgrid collaborative optimal strategy. A robust reward function under the worst case is constructed by using an uncertain Markov process; the strategy difference is converted into a reward difference by the total variation distance. Taking into account the constraints and objective functions of the distribution network and the microgrid, a distribution network-microgrid collaborative optimal strategy model based on a multi-agent deep reinforcement learning robust reward function is constructed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distribution network-microgrid collaborative optimization, and specifically relates to a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function. Background Art

[0002] With the widespread application of distributed energy and microgrids, the operation and planning of modern distribution networks have become increasingly complex. High penetration of renewable energy and energy storage systems enable microgrids to achieve energy self-sufficiency in local areas for some time periods. However, the uncertainty of renewable energy and load can cause a "dangerous" state in microgrids and distribution networks. This state refers to the situation where the changes in renewable energy and load within the grid strategy instruction interval cause the system to exceed safe operation constraints. Therefore, considering the worst case within the instruction interval and giving accurate robust instructions, especially when the state is close to the operating constraints, becomes crucial for the optimization strategy of the power system and grid security.

[0003] However, traditional optimization methods can only perform robustness calculations on a single time section, thereby ignoring the impact of time-coupled devices such as energy storage and distributed power sources on subsequent states, which can easily lead to reduced microgrid operating benefits and increased distribution network operating costs. Therefore, it is particularly important to study robust optimization methods for distribution network-microgrid collaborative optimization problems. Therefore, it is urgent to develop a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function to improve the robustness and economy of distribution network-microgrid collaborative optimization under multiple time sections. Summary of the invention

[0004] Purpose of the invention: The technical problem to be solved by the present invention is to provide a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function in view of the deficiencies in the prior art. The present invention takes into account the collaboration of the distribution network and the microgrid, introduces a multi-agent deep reinforcement learning model to train the charging and discharging power of the energy storage system in the microgrid, and converts the charging and discharging power of the energy storage system into the comprehensive benefit of the microgrid through the microgrid optimal power flow model that considers the microgrid operation constraints and the marginal electricity price of the distribution network nodes. The distribution network-related constraints are introduced to implement a collaborative strategy for the distribution network-microgrid. The present invention can not only realize the accurate interactive power of the microgrid, but also consider the worst case within the instruction interval on this basis, thereby ensuring the safety of collaborative operation.

[0005] Technical solution: In order to solve the above technical problems, the present invention provides a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function, the method comprising the following steps:

[0006] Step 1: Obtain the network parameters and operating parameters of the distribution network; the network parameters and operating parameters of the microgrid; initialize the node marginal electricity price and the microgrid interactive power; initialize the multi-agent reinforcement learning strategy network parameters and value network parameters; and initialize the time;

[0007] Step 2: Obtain the active input active power and reactive power data of the distribution network and microgrid nodes at time t; and the power generation data of the microgrid photovoltaic array and wind turbine at time t;

[0008] Step 3: Based on the interactive power of the microgrid at time t, combined with the network parameters and operating parameters of the distribution network in step 1 and the active power and reactive power data actively input by the distribution network nodes in step 2, with the flow constraint and other constraints as constraints, and the minimum distribution network operation cost as the objective function; solve to obtain the minimum value of the distribution network operation cost at time t, and calculate the marginal electricity price of the node at time t by calculating the partial derivative of the active power transmitted by the distribution network line with the node as the terminal node with respect to the distribution network operation cost;

[0009] Step 4: Based on the marginal electricity price of the node at time t, combined with the active power and reactive power data, photovoltaic array and wind turbine generation data actively input by the microgrid node at time t in step 2, the strategy network of multi-agent deep reinforcement learning is input to obtain the strategy at time t, which is the sampling probability of the energy storage system / inverter power ratio;

[0010] Based on the marginal electricity price at the node at time t, combined with the active power and reactive power data of the microgrid node at time t in step 2, the power generation data of the photovoltaic array and the wind turbine, and the power ratio of the energy storage system / inverter at time t obtained by strategy sampling, combined with the microgrid network parameters and operation parameters in step 1, with the flow constraint, energy storage constraint, and inverter constraint as the constraint conditions, and the maximum comprehensive benefit of the microgrid as the objective function; solve to obtain the maximum comprehensive benefit of the microgrid at time t and the interactive power of the microgrid at time t;

[0011] Step 5: Repeat steps 3 and 4 to update the node marginal electricity price and microgrid interaction power at time t until the node marginal electricity price difference of each node before and after the update is less than 0.01 yuan / kWh, and output the node marginal electricity price at time t, the maximum value of the microgrid comprehensive benefit at time t, the microgrid interaction power at time t, and the energy storage system / inverter power ratio at time t obtained by strategy sampling;

[0012] Step 6: Based on the maximum value of the comprehensive benefit of the microgrid at time t output in step 5, assign it to the multi-agent deep reinforcement learning as the actual non-robust reward, and obtain the actual robust reward of the multi-agent deep reinforcement learning based on the actual non-robust reward and the robust reward function of the multi-agent deep reinforcement learning;

[0013] Based on the active input of active power and reactive power data, photovoltaic array and wind turbine generation data by the microgrid node at time t in step 2, step 5 outputs the marginal electricity price of the node at time t, the maximum value of the microgrid comprehensive benefit at time t and the microgrid interactive power at time t, the energy storage system / inverter power ratio at time t obtained by strategy sampling, and inputs the multi-agent deep reinforcement learning value network to obtain the network robust reward; calculate the value partial derivative of the value network parameter for the square of the difference between the actual robust reward and the network robust reward, and update the value network parameter by gradient descent based on the value partial derivative; calculate the policy partial derivative of the policy network parameter for the network robust reward, and update the policy network parameter by gradient ascent based on the policy partial derivative;

[0014] Step 7: Repeat steps 2 to 6 until the calculation of the distribution network and microgrid load data, the microgrid photovoltaic array and the wind turbine generation data is completed.

[0015] Furthermore, in step 3, with power flow constraints and other constraints as constraints, the objective function is to minimize the distribution network operation cost:

[0016] Objective function:

[0017]

[0018] Constraints:

[0019] 1) Power flow constraints:

[0020]

[0021] 2) Other constraints:

[0022]

[0023] Where pd represents the distribution network, F t pd represents the distribution network operation cost at time t, G represents the distributed generation set, g represents the distributed generation, a g , b g and c g They represent the quadratic coefficient, the primary coefficient and the constant term of the power generation cost of the distributed generation g, respectively. represents the power generated by distributed generation g at time t / t-1, c ma P represents the price per unit of power purchased by the distribution network from the main grid. t ma represents the power purchased by the distribution network from the main grid at time t, B represents the set of distribution network nodes, i and j represent the distribution network node numbers, c sh It represents the unit power price of load shedding in the distribution network. represents the load shedding power of node i at time t, M represents the microgrid set, m represents the microgrid number, P t m represents the interactive power of microgrid m at time t, represents the marginal electricity price of microgrid node m at time t, P i,t Indicates that node i actively inputs active power at time t, F i represents the end node of the line starting from node i, T i represents the starting node of the line with node i as the terminal node, P ij,t / P ji,t represents the active power transmitted from node i to node j / from node j to node i at time t, ij represents the line from node j to node i, I ij,t / I ji,t represents the square of the line current transmitted from node i to node j / from node j to node i at time t, R ij represents the line ij resistance, Q i,t Indicates that node i actively inputs reactive power at time t, Q ij,t / Q ji,t represents the reactive power of the line transmission from node i to node j / from node j to node i at time t, X ij Represents the line ij reactance, V i,t / V j,t represents the square of the voltage of node i / j at time t, L represents the line set, represents the square of the maximum line transmission current of line ij, Indicates the maximum / minimum active power of node i, represents the maximum / minimum reactive power of node i, represents the square of the maximum / minimum voltage at node i, RU g / RD g Indicates the maximum up / down slope value of the distributed generation g in adjacent time intervals.

[0024] Furthermore, in step 4, with power flow constraint, energy storage constraint and inverter constraint as constraint conditions, the objective function is to maximize the comprehensive benefit of the microgrid:

[0025] Objective function:

[0026]

[0027] Constraints:

[0028] 1) Flow constraint: Same as (A-2)-(A-7)

[0029] 2) Energy storage constraints:

[0030]

[0031]

[0032] 3) Inverter constraints:

[0033]

[0034] In the formula, F t m represents the comprehensive benefit of microgrid m at time t, E / P / W represents the energy storage system / photovoltaic array / wind turbine set, e / p / w represents the number of energy storage system / photovoltaic array / wind turbine, c e represents the unit power cost of charging or discharging the energy storage system, c p / c w P represents the unit power generation cost of the photovoltaic array / wind turbine, t m,e / P t m,p / P t m,w represents the active power output of the microgrid m energy storage system / photovoltaic array / wind turbine at time t, It represents the maximum charging / minimum discharging power of microgrid m energy storage system e at time t, Indicates the maximum / minimum charge state of the energy storage system, ζ indicates the self-discharge coefficient of the energy storage system, Δt indicates the five-minute time interval, represents the charge state of the energy storage system e in the microgrid m at time t / t+1, C m,e represents the rated capacity of the microgrid m energy storage system e, represents the charging / discharging state of the energy storage system e in the microgrid m at time t, η ch / η dis Indicates the charging / discharging efficiency of the energy storage system, Indicates that the microgrid m energy storage system e sets the maximum charging / minimum discharging power, max / min indicates the maximum / minimum value, represents the power ratio of microgrid m to energy storage system e, It represents the power ratio of the inverter of the microgrid m photovoltaic array / wind turbine group at time t, Indicates the maximum / minimum output reactive power of the inverter of the microgrid m photovoltaic array p at time t, It represents the maximum / minimum output reactive power of the inverter of wind turbine w in microgrid m at time t, P represents the maximum apparent output power of the inverter of the microgrid m photovoltaic array p / wind turbine w, t m,p / P t m,w It represents the active power output of the photovoltaic array p / wind turbine w where the inverter of the microgrid m is located at time t, Indicates the maximum output reactive power set by the inverter of the microgrid m photovoltaic array p / wind turbine w, It represents the reactive power output by the inverter of the photovoltaic array p of the microgrid m / wind turbine w at time t.

[0035] Furthermore, in step 4, the active power and reactive power data of the microgrid node at time t in step 2, the power generation data of the photovoltaic array and the wind turbine, and the marginal electricity price of the node at time t in step 3 are actively input into the multi-agent deep reinforcement learning to obtain the strategy at time t. The strategy is the sampling probability of the energy storage system / inverter power ratio:

[0036]

[0037] In the formula, represents the state / action of the deep reinforcement learning of the multi-agent microgrid m at time t, lo represents the number of the load node, It indicates that the load node lo of microgrid m actively inputs active power / reactive power at time t, Indicates the number of the input layer / intermediate layer / output layer of the policy network. represents the value of the vth element in the middle layer of the microgrid m strategy network at time t, represents the network activation function, represents the weight from the uth element in the input layer to the vth element in the middle layer of the microgrid m strategy network at time t, represents the u-th element of the deep reinforcement learning state of the multi-agent microgrid m at time t, represents the output layer of the microgrid m strategy network at time t The value of the element, It represents the time t from the vth element in the middle layer to the output layer in the strategy network of microgrid m The weight of the element, represents the mean action value of the policy network, CL represents the policy function of the policy network, Represents policy network actions Selection probability, π represents pi, Σ represents the 3*3 covariance matrix of the policy network, The natural constant e of the policy network is represented by Second power, Represents the policy network matrix The transpose of .

[0038] Furthermore, in step 6, the actual non-robust reward, actual robust reward and network robust reward of multi-agent deep reinforcement learning are:

[0039] Actual non-robust reward:

[0040]

[0041] Actual robust reward:

[0042]

[0043] Network robustness reward:

[0044]

[0045] In the formula, r t m represents the reward of multi-agent deep reinforcement learning of microgrid m at time t, represents the non-robust strategy of deep reinforcement learning for multi-agent microgrid m, Represents a non-robust strategy Down state Value function, A / S represents the set of actions / states, γ represents the discount factor, Indicates status Adopt an action Transfer to state The probability of represents the state of deep reinforcement learning of multi-agent microgrid m at time t+1, Represents a non-robust strategy Under non-robust rewards, Represents a non-robust strategy Down state The distribution probability of Represents a robust strategy Under robust rewards, Representation strategy and Down state The total variation distance, x / y represents the value network input layer / intermediate layer number, represents the value of the yth element in the middle layer of the value network of microgrid m at time t, represents the weight of the xth element in the input layer to the yth element in the middle layer of the value network of microgrid m at time t, Represents the deep reinforcement learning state of multi-agent microgrid m at time t -action The xth element of represents the value network output of microgrid m at time t, Represents the weight of the yth element in the middle layer of the value network of microgrid m to the output layer at time t.

[0046] Furthermore, in step 6, the value partial derivative, gradient descent, policy partial derivative, and gradient ascent are as follows;

[0047] Partial derivative of value:

[0048]

[0049] Gradient Descent:

[0050]

[0051] Strategy partial derivatives:

[0052]

[0053] Gradient Ascent:

[0054]

[0055] In the formula, α / β represents the learning rate of the value network / strategy network, represents the weight from the xth element of the input layer to the yth element of the middle layer of the value network of microgrid m at time t+1, represents the weight from the yth element in the middle layer of the value network of microgrid m to the output at time t+1, represents the weight from the uth element in the input layer to the vth element in the middle layer of the microgrid m-strategy network at time t+1, It represents the time t+1 from the vth element in the middle layer to the output layer in the strategy network of the microgrid m The weight of the element, It represents the partial derivative of the value of the weight of the yth element in the middle layer of the value network of microgrid m at time t to the output layer with respect to the square of the difference between the actual robust cumulative reward and the network robust cumulative reward, It represents the value partial derivative of the weight of the x-th element in the input layer to the y-th element in the middle layer of the value network of microgrid m at time t with respect to the square of the difference between the actual robust cumulative reward and the network robust cumulative reward, represents the symbol of partial derivative, It represents the time t from the vth element in the middle layer to the output layer in the strategy network of microgrid m The partial derivative of the element weight with respect to the network’s robust cumulative reward strategy, Represents the policy partial derivative of the weight from the u-th element in the input layer to the v-th element in the intermediate layer of the microgrid m strategy network at time t with respect to the network’s robust cumulative reward.

[0056] Beneficial effects: Compared with the prior art, the technical solution of the present invention has the following beneficial effects:

[0057] Compared with the traditional model collaborative optimization scheme, the technical scheme of the present invention reduces the calculation time during execution by introducing deep learning training, generates training data through multi-agent deep reinforcement learning and performs weighted summation of multi-time section benefits based on discount factors, and then accelerates the training convergence of multiple microgrids through a multi-agent framework. A robust reward function under the worst case is constructed by using an uncertain Markov process; considering the data pairs in the buffer, the state difference is converted into a strategy difference through the dual form of the Hall inequality, and the strategy difference is converted into a reward difference through the total variation distance. The results of the example test show that the method proposed in the present invention can improve the comprehensive benefits of the microgrid compared with the existing method, considers the worst case of the instruction interval, and improves the stability of the microgrid. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 is a single round training flow chart of the method of the present invention;

[0059] Figure 2 This is a calculation example diagram of the power grid model of distribution network-microgrid collaboration;

[0060] Figure 3 It is a comparison chart of microgrid interaction benefits and calculation time under different methods. DETAILED DESCRIPTION

[0061] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, various equivalent modifications of the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0062] like Figure 1 As shown, the present invention proposes a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function, and the method includes the following steps:

[0063] Step 1: Obtain the network parameters and operating parameters of the distribution network; the network parameters and operating parameters of the microgrid; initialize the node marginal electricity price and the microgrid interactive power; initialize the multi-agent reinforcement learning strategy network parameters and value network parameters; and initialize the time;

[0064] Step 2: Obtain the active input active power and reactive power data of the distribution network and microgrid nodes at time t; and the power generation data of the microgrid photovoltaic array and wind turbine at time t;

[0065] Step 3: Based on the interactive power of the microgrid at time t, combined with the network parameters and operating parameters of the distribution network in step 1 and the active power and reactive power data actively input by the distribution network nodes in step 2, with the flow constraint and other constraints as constraints, and the minimum distribution network operation cost as the objective function; solve to obtain the minimum value of the distribution network operation cost at time t, and calculate the marginal electricity price of the node at time t by calculating the partial derivative of the active power transmitted by the distribution network line with the node as the terminal node with respect to the distribution network operation cost;

[0066] Step 4: Based on the marginal electricity price of the node at time t, combined with the active power and reactive power data, photovoltaic array and wind turbine generation data actively input by the microgrid node at time t in step 2, the strategy network of multi-agent deep reinforcement learning is input to obtain the strategy at time t, which is the sampling probability of the energy storage system / inverter power ratio;

[0067] Based on the marginal electricity price at the node at time t, combined with the active power and reactive power data of the microgrid node at time t in step 2, the power generation data of the photovoltaic array and the wind turbine, and the power ratio of the energy storage system / inverter at time t obtained by strategy sampling, combined with the microgrid network parameters and operation parameters in step 1, with the flow constraint, energy storage constraint, and inverter constraint as the constraint conditions, and the maximum comprehensive benefit of the microgrid as the objective function; solve to obtain the maximum comprehensive benefit of the microgrid at time t and the interactive power of the microgrid at time t;

[0068] Step 5: Repeat steps 3 and 4 to update the node marginal electricity price and microgrid interaction power at time t until the node marginal electricity price difference of each node before and after the update is less than 0.01 yuan / kWh, and output the node marginal electricity price at time t, the maximum value of the microgrid comprehensive benefit at time t, the microgrid interaction power at time t, and the energy storage system / inverter power ratio at time t obtained by strategy sampling;

[0069] Step 6: Based on the maximum value of the comprehensive benefit of the microgrid at time t output in step 5, assign it to the multi-agent deep reinforcement learning as the actual non-robust reward, and obtain the actual robust reward of the multi-agent deep reinforcement learning based on the actual non-robust reward and the robust reward function of the multi-agent deep reinforcement learning;

[0070] Based on the active input of active power and reactive power data, photovoltaic array and wind turbine generation data by the microgrid node at time t in step 2, step 5 outputs the marginal electricity price of the node at time t, the maximum value of the microgrid comprehensive benefit at time t and the microgrid interactive power at time t, the energy storage system / inverter power ratio at time t obtained by strategy sampling, and inputs the multi-agent deep reinforcement learning value network to obtain the network robust reward; calculate the value partial derivative of the value network parameter for the square of the difference between the actual robust reward and the network robust reward, and update the value network parameter by gradient descent based on the value partial derivative; calculate the policy partial derivative of the policy network parameter for the network robust reward, and update the policy network parameter by gradient ascent based on the policy partial derivative;

[0071] Step 7: Repeat steps 2 to 6 until the calculation of the distribution network and microgrid load data, the microgrid photovoltaic array and the wind turbine generation data is completed.

[0072] Furthermore, in step 3, with power flow constraints and other constraints as constraints, the objective function is to minimize the distribution network operation cost:

[0073] Objective function:

[0074]

[0075] Constraints:

[0076] 1) Power flow constraints:

[0077]

[0078] 2) Other constraints:

[0079]

[0080] Where pd represents the distribution network, F t pd represents the distribution network operation cost at time t, G represents the distributed generation set, g represents the distributed generation, a g , b g and c g They represent the quadratic coefficient, the primary coefficient and the constant term of the power generation cost of the distributed generation g, respectively. represents the power generated by distributed generation g at time t / t-1, c ma P represents the price per unit of power purchased by the distribution network from the main grid. t ma represents the power purchased by the distribution network from the main grid at time t, B represents the set of distribution network nodes, i and j represent the distribution network node numbers, c sh It represents the unit power price of load shedding in the distribution network. Denote the load shedding power of node \(i\) at time \(t\), \(M\) represents the set of microgrids, \(m\) represents the microgrid number, \(P\) t m Denote the interactive power of microgrid \(m\) at time \(t\), Denote the nodal marginal price of microgrid \(m\) at time \(t\), \(P\) i,t Denote the active input power of node \(i\) at time \(t\), \(F\) i Denote the end node of the line starting from node \(i\), \(T\) i Denote the start node of the line with node \(i\) as the end node, \(P\) ij,t / P ji,t Denote the active power transmitted through the line from node \(i\) to node \(j\) / from node \(j\) to node \(i\) at time \(t\), \(ij\) represents the line from node \(j\) to node \(i\), \(I\) ij,t / I ji,t Denote the square of the current transmitted through the line from node \(i\) to node \(j\) / from node \(j\) to node \(i\) at time \(t\), \(R\) ij Denote the resistance of line \(ij\), \(Q\) i,t Denote the reactive power input of node \(i\) at time \(t\), \(Q\) ij,t / Q ji,t Denote the reactive power transmitted through the line from node \(i\) to node \(j\) / from node \(j\) to node \(i\) at time \(t\), \(X\) ij Denote the reactance of line \(ij\), \(V\) i,t / V j,t Denote the square of the voltage of node \(i\) / j at time \(t\), \(L\) represents the set of lines, Denote the square of the maximum line current transmitted through line \(ij\), Denote the maximum / minimum active power of node \(i\), Denote the maximum / minimum reactive power of node \(i\), Denote the square of the maximum / minimum voltage of node \(i\), \(RU\) g / RD g Denote the maximum up / down ramp of the distributed generator \(g\) over adjacent time intervals.

[0081] Furthermore, in step 4, with power flow constraints, energy storage constraints, and inverter constraints as the constraint conditions, and the maximum comprehensive benefit of the microgrid as the objective function:

[0082] Objective function:

[0083]

[0084] Constraint conditions:

[0085] 1) Power flow constraints: same as equations (A-2)-(A-7)

[0086] 2) Energy storage constraints:

[0087]

[0088]

[0089] 3) Inverter constraints:

[0090]

[0091] In the formula, F t m represents the comprehensive benefit of microgrid m at time t, E / P / W represents the energy storage system / photovoltaic array / wind turbine set, e / p / w represents the number of energy storage system / photovoltaic array / wind turbine, c e represents the unit power cost of charging or discharging the energy storage system, c p / c w P represents the unit power generation cost of the photovoltaic array / wind turbine, t m,e / P t m,p / P t m,w represents the active power output of the microgrid m energy storage system / photovoltaic array / wind turbine at time t, It represents the maximum charging / minimum discharging power of microgrid m energy storage system e at time t, Indicates the maximum / minimum charge state of the energy storage system, ζ indicates the self-discharge coefficient of the energy storage system, Δt indicates the five-minute time interval, represents the charge state of the energy storage system e in the microgrid m at time t / t+1, C m,e represents the rated capacity of the microgrid m energy storage system e, represents the charging / discharging state of the energy storage system e in the microgrid m at time t, η ch / η dis Indicates the charging / discharging efficiency of the energy storage system, Indicates that the microgrid m energy storage system e sets the maximum charging / minimum discharging power, max / min indicates the maximum / minimum value, represents the power ratio of microgrid m to energy storage system e, It represents the power ratio of the inverter of the microgrid m photovoltaic array / wind turbine group at time t, Indicates the maximum / minimum output reactive power of the inverter of the microgrid m photovoltaic array p at time t, It represents the maximum / minimum output reactive power of the inverter of wind turbine w in microgrid m at time t, P represents the maximum apparent output power of the inverter of the microgrid m photovoltaic array p / wind turbine w, t m,p / P t m,w It represents the active power output of the photovoltaic array p / wind turbine w where the inverter of the microgrid m is located at time t, Indicates the maximum output reactive power set by the inverter of the microgrid m photovoltaic array p / wind turbine w, It represents the reactive power output by the inverter of the photovoltaic array p of the microgrid m / wind turbine w at time t.

[0092] Furthermore, in step 4, the active power and reactive power data of the microgrid node at time t in step 2, the power generation data of the photovoltaic array and the wind turbine, and the marginal electricity price of the node at time t in step 3 are actively input into the multi-agent deep reinforcement learning to obtain the strategy at time t. The strategy is the sampling probability of the energy storage system / inverter power ratio:

[0093]

[0094] In the formula, represents the state / action of the deep reinforcement learning of the multi-agent microgrid m at time t, lo represents the number of the load node, It indicates that the load node lo of microgrid m actively inputs active power / reactive power at time t, Indicates the number of the input layer / intermediate layer / output layer of the policy network. represents the value of the vth element in the middle layer of the microgrid m strategy network at time t, represents the network activation function, represents the weight from the uth element in the input layer to the vth element in the middle layer of the microgrid m strategy network at time t, represents the u-th element of the deep reinforcement learning state of the multi-agent microgrid m at time t, represents the output layer of the microgrid m strategy network at time t the value of the element, It represents the time t from the vth element in the middle layer to the output layer in the strategy network of microgrid m The weight of the element, represents the mean action value of the policy network, CL represents the policy function of the policy network, Represents policy network actions Selection probability, π represents pi, Σ represents the 3*3 covariance matrix of the policy network, The natural constant e of the policy network is represented by Second power, Represents the policy network matrix The transpose of .

[0095] Furthermore, in step 6, the actual non-robust reward, actual robust reward and network robust reward of multi-agent deep reinforcement learning are:

[0096] Actual non-robust reward:

[0097]

[0098] Actual robust reward:

[0099]

[0100] Network robustness reward:

[0101]

[0102] In the formula, r t m represents the reward of multi-agent deep reinforcement learning of microgrid m at time t, represents the non-robust strategy of deep reinforcement learning for multi-agent microgrid m, Represents a non-robust strategy Down state Value function, A / S represents the set of actions / states, γ represents the discount factor, Indicates status Adopt an action Transfer to state The probability of represents the state of deep reinforcement learning of multi-agent microgrid m at time t+1, Represents a non-robust strategy Under non-robust rewards, Represents a non-robust strategy Down state The distribution probability of Represents a robust strategy Under robust rewards, Representation strategy and Down state The total variation distance, x / y represents the value network input layer / intermediate layer number, represents the value of the yth element in the middle layer of the value network of microgrid m at time t, represents the weight of the xth element in the input layer to the yth element in the middle layer of the value network of microgrid m at time t, Represents the deep reinforcement learning state of multi-agent microgrid m at time t -action The xth element of represents the value network output of microgrid m at time t, Represents the weight of the yth element in the middle layer of the value network of microgrid m to the output layer at time t.

[0103] Furthermore, in step 6, the value partial derivative, gradient descent, policy partial derivative, and gradient ascent are as follows;

[0104] Partial derivative of value:

[0105]

[0106] Gradient Descent:

[0107]

[0108] Strategy partial derivatives:

[0109]

[0110] Gradient Ascent:

[0111]

[0112] In the formula, α / β represents the learning rate of the value network / strategy network, represents the weight from the xth element of the input layer to the yth element of the middle layer of the value network of microgrid m at time t+1, represents the weight from the yth element in the middle layer of the value network of microgrid m to the output at time t+1, represents the weight from the uth element in the input layer to the vth element in the middle layer of the microgrid m-strategy network at time t+1, It represents the time t+1 from the vth element in the middle layer to the output layer in the strategy network of the microgrid m The weight of the element, It represents the partial derivative of the value of the weight of the yth element in the middle layer of the value network of microgrid m at time t to the output layer with respect to the square of the difference between the actual robust cumulative reward and the network robust cumulative reward, It represents the value partial derivative of the weight of the x-th element in the input layer to the y-th element in the middle layer of the value network of microgrid m at time t with respect to the square of the difference between the actual robust cumulative reward and the network robust cumulative reward, represents the symbol of partial derivative, It represents the time t from the vth element in the middle layer to the output layer in the strategy network of microgrid m The partial derivative of the element weight with respect to the network’s robust cumulative reward strategy, Represents the policy partial derivative of the weight from the u-th element in the input layer to the v-th element in the intermediate layer of the microgrid m strategy network at time t with respect to the network’s robust cumulative reward.

[0113] Case Analysis

[0114] The following example illustrates the superiority of the optimal strategy for the coordination of distribution network and microgrid based on the multi-agent deep reinforcement learning robust reward function. Figure 2The improved IEEE 33-node power system shown in FIG. 1 is a schematic diagram of an improved IEEE 33-node power system. In order to compare the superiority of the method proposed in the present invention, the optimal strategy method for distribution network-microgrid based on the mechanism model and the optimal strategy method for distribution network-microgrid based on the robust reward function of multi-agent deep reinforcement learning proposed in the present invention are respectively used to perform the distribution network-microgrid coordination strategy. The present invention is implemented through the python platform and the Gurobi solver is used to solve the optimization problem.

[0115] Based on this example, the interactive benefits of the optimal strategy method for distribution network-microgrid based on the mechanism model and the multi-agent deep reinforcement learning robust reward function proposed in this invention are compared (see the results for details). Figure 3 ) and the comparison of the interaction benefits and computing time of the two methods (see Table 1 for the results). This shows that the multi-agent deep reinforcement learning considers the benefits of multiple time sections through the discount factor, and the interactive power generated by the strategic energy storage system is better than the mechanism model method that only considers a single time section (Table 1 and Figure 3 ). Compared with the mechanism model calculation, the present invention reduces the calculation time by adding a deep learning neural network (Table 1).

[0116] Table 1 Comparison of interaction benefits and calculation time of different methods

[0117]

[0118] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A distribution network-microgrid collaboration optimal strategy method based on multi-agent deep reinforcement learning robust reward function, characterized in that: The method comprises the following steps: Step 1: Obtain the network parameters and operating parameters of the distribution network, the network parameters and operating parameters of the microgrid; initialize the node marginal electricity price and the microgrid interactive power, the multi-agent reinforcement learning strategy network parameters and the value network parameters, and initialize the time; Step 2: Obtain the active input active power and reactive power data of the distribution network and microgrid nodes at time t; and the power generation data of the microgrid photovoltaic array and wind turbine at time t; Step 3: Based on the interactive power of the microgrid at time t, combined with the network parameters and operating parameters of the distribution network in step 1 and the active power and reactive power data actively input by the distribution network nodes in step 2, with the flow constraint and other constraints as constraints, and the minimum distribution network operation cost as the objective function; solve to obtain the minimum value of the distribution network operation cost at time t, and calculate the marginal electricity price of the node at time t by calculating the partial derivative of the active power transmitted by the distribution network line with the node as the terminal node with respect to the distribution network operation cost; Step 4: Based on the marginal electricity price of the node at time t, combined with the active power and reactive power data, photovoltaic array and wind turbine generation data actively input by the microgrid node at time t in step 2, the strategy network of multi-agent deep reinforcement learning is input to obtain the strategy at time t, which is the sampling probability of the energy storage system / inverter power ratio; Based on the marginal electricity price at the node at time t, combined with the active power and reactive power data of the microgrid node at time t in step 2, the power generation data of the photovoltaic array and the wind turbine, and the power ratio of the energy storage system / inverter at time t obtained by strategy sampling, combined with the microgrid network parameters and operation parameters in step 1, with the flow constraint, energy storage constraint, and inverter constraint as the constraint conditions, and the maximum comprehensive benefit of the microgrid as the objective function; solve to obtain the maximum comprehensive benefit of the microgrid at time t and the interactive power of the microgrid at time t; Step 5: Repeat steps 3 and 4 to update the node marginal electricity price and microgrid interaction power at time t until the node marginal electricity price difference of each node before and after the update is less than 0.01 yuan / kWh, and output the node marginal electricity price at time t, the maximum value of the microgrid comprehensive benefit at time t, the microgrid interaction power at time t, and the energy storage system / inverter power ratio at time t obtained by strategy sampling; Step 6: Based on the maximum value of the comprehensive benefit of the microgrid at time t output in step 5, assign it to the multi-agent deep reinforcement learning as the actual non-robust reward, and obtain the actual robust reward of the multi-agent deep reinforcement learning based on the actual non-robust reward and the robust reward function of the multi-agent deep reinforcement learning; Based on the active input of active power and reactive power data, photovoltaic array and wind turbine generation data by the microgrid node at time t in step 2, step 5 outputs the marginal electricity price of the node at time t, the maximum value of the microgrid comprehensive benefit at time t and the microgrid interactive power at time t, the energy storage system / inverter power ratio at time t obtained by strategy sampling, and inputs the multi-agent deep reinforcement learning value network to obtain the network robust reward; calculate the value partial derivative of the value network parameter for the square of the difference between the actual robust reward and the network robust reward, and update the value network parameter by gradient descent based on the value partial derivative; calculate the policy partial derivative of the policy network parameter for the network robust reward, and update the policy network parameter by gradient ascent based on the policy partial derivative; Step 7: Repeat steps 2 to 6 until the calculation of the distribution network and microgrid load data, the microgrid photovoltaic array and the wind turbine generation data is completed.

2. According to claim 1, a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function is characterized in that: In step 3, with power flow constraints and other constraints as constraints, the objective function is to minimize the distribution network operation cost: Objective function: Constraints: 1) Power flow constraints: 2) Other constraints: Where pd represents the distribution network, F t pd represents the distribution network operation cost at time t, G represents the distributed generation set, g represents the distributed generation, a g , b g and c g They represent the quadratic coefficient, the primary coefficient and the constant term of the power generation cost of the distributed generation g, respectively. represents the power generated by distributed generation g at time t / t-1, c ma P represents the price per unit of power purchased by the distribution network from the main grid. t ma represents the power purchased by the distribution network from the main grid at time t, B represents the set of distribution network nodes, i and j represent the distribution network node numbers, c sh It represents the unit power price of load shedding in the distribution network. represents the load shedding power of node i at time t, M represents the microgrid set, m represents the microgrid number, P t m represents the interactive power of microgrid m at time t, represents the marginal electricity price of microgrid node m at time t, P i,t Indicates that node i actively inputs active power at time t, F i represents the end node of the line starting from node i, T i represents the starting node of the line with node i as the terminal node, P ij,t / P ji,t represents the active power transmitted from node i to node j / from node j to node i at time t, ij represents the line from node j to node i, I ij,t / I ji,t represents the square of the line current transmitted from node i to node j / from node j to node i at time t, R ij represents the line ij resistance, Q i,t Indicates that node i actively inputs reactive power at time t, Q ij,t / Q ji,t represents the reactive power of the line transmission from node i to node j / from node j to node i at time t, X ij Represents the line ij reactance, V i,t / V j,t represents the square of the voltage of node i / j at time t, L represents the line set, represents the square of the maximum line transmission current of line ij, Indicates the maximum / minimum active power of node i, represents the maximum / minimum reactive power of node i, represents the square of the maximum / minimum voltage at node i, RU g / RD g Indicates the maximum up / down slope value of the distributed generation g in adjacent time intervals.

3. According to claim 2, a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function is characterized in that: In step 4, with power flow constraint, energy storage constraint and inverter constraint as constraint conditions, the objective function is to maximize the comprehensive benefit of the microgrid: Objective function: Constraints: 1) Flow constraint: Same as (A-2)-(A-7) 2) Energy storage constraints: 3) Inverter constraints: In the formula, F t m represents the comprehensive benefit of microgrid m at time t, E / P / W represents the energy storage system / photovoltaic array / wind turbine set, e / p / w represents the number of energy storage system / photovoltaic array / wind turbine, c e represents the unit power cost of charging or discharging the energy storage system, c p / c w P represents the unit power generation cost of the photovoltaic array / wind turbine, t m,e / P t m,p / P t m,w represents the active power output of the microgrid m energy storage system / photovoltaic array / wind turbine at time t, It represents the maximum charging / minimum discharging power of microgrid m energy storage system e at time t, Indicates the maximum / minimum charge state of the energy storage system, ζ indicates the self-discharge coefficient of the energy storage system, Δt indicates the five-minute time interval, represents the charge state of the energy storage system e in the microgrid m at time t / t+1, C m,e represents the rated capacity of the microgrid m energy storage system e, represents the charging / discharging state of the energy storage system e in the microgrid m at time t, η ch / η dis Indicates the charging / discharging efficiency of the energy storage system, Indicates that the microgrid m energy storage system e sets the maximum charging / minimum discharging power, max / min indicates the maximum / minimum value, represents the power ratio of microgrid m to energy storage system e, It represents the power ratio of the inverter of the microgrid m photovoltaic array / wind turbine group at time t, Indicates the maximum / minimum output reactive power of the inverter of the microgrid m photovoltaic array p at time t, It represents the maximum / minimum output reactive power of the inverter of wind turbine w in microgrid m at time t, P represents the maximum apparent output power of the inverter of the microgrid m photovoltaic array p / wind turbine w, t m,p / P t m,w It represents the active power output of the photovoltaic array p / wind turbine w where the inverter of the microgrid m is located at time t, Indicates the maximum output reactive power set by the inverter of the microgrid m photovoltaic array p / wind turbine w, It represents the reactive power output by the inverter of the photovoltaic array p of the microgrid m / wind turbine w at time t.

4. According to claim 1, a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function is characterized in that: In step 4, the active power and reactive power data of the microgrid node at time t in step 2, the power generation data of the photovoltaic array and the wind turbine, and the marginal electricity price of the node at time t in step 3 are actively input into the multi-agent deep reinforcement learning to obtain the strategy at time t. The strategy is the sampling probability of the energy storage system / inverter power ratio: In the formula, represents the state / action of the deep reinforcement learning of the multi-agent microgrid m at time t, lo represents the number of the load node, It indicates that the load node lo of microgrid m actively inputs active power / reactive power at time t, Indicates the number of the input layer / intermediate layer / output layer of the policy network. represents the value of the vth element in the middle layer of the microgrid m strategy network at time t, represents the network activation function, represents the weight from the uth element in the input layer to the vth element in the middle layer of the microgrid m strategy network at time t, represents the u-th element of the deep reinforcement learning state of the multi-agent microgrid m at time t, represents the output layer of the microgrid m strategy network at time t the value of the element, It represents the time t from the vth element in the middle layer to the output layer in the strategy network of microgrid m The weight of the element, represents the mean action value of the policy network, CL represents the policy function of the policy network, Represents policy network actions Selection probability, π represents pi, Σ represents the 3*3 covariance matrix of the policy network, The natural constant e of the policy network is represented by Second power, Represents the policy network matrix The transpose of .

5. According to claim 4, a distribution network-microgrid collaborative optimal strategy method based on multi-agent deep reinforcement learning robust reward function is characterized in that: In step 6, the actual non-robust reward, actual robust reward, and network robust reward of multi-agent deep reinforcement learning are: Actual non-robust reward: Actual robust reward: Network robustness reward: In the formula, represents the reward of multi-agent deep reinforcement learning of microgrid m at time t, represents the non-robust strategy of deep reinforcement learning for multi-agent microgrid m, Represents a non-robust strategy Down state Value function, A / S represents the set of actions / states, γ represents the discount factor, Indicates status Adopt an action Transfer to state The probability of represents the state of deep reinforcement learning of multi-agent microgrid m at time t+1, Represents a non-robust strategy Under non-robust rewards, Represents a non-robust strategy Down state The distribution probability of Represents a robust strategy Under robust rewards, Representation strategy and Down state The total variation distance, x / y represents the value network input layer / intermediate layer number, represents the value of the yth element in the middle layer of the value network of microgrid m at time t, represents the weight of the xth element in the input layer to the yth element in the middle layer of the value network of microgrid m at time t, Represents the deep reinforcement learning state of multi-agent microgrid m at time t -action The xth element of represents the value network output of microgrid m at time t, Represents the weight of the yth element in the middle layer of the value network of microgrid m to the output layer at time t.

6. The optimal strategy method for distribution network-microgrid collaboration based on multi-agent deep reinforcement learning robust reward function according to claim 5 is characterized in that: In step 6, the value partial derivative, gradient descent, policy partial derivative, and gradient ascent are as follows; Partial derivative of value: Gradient Descent: Strategy partial derivatives: Gradient Ascent: In the formula, α / β represents the learning rate of the value network / strategy network, represents the weight from the xth element in the input layer to the yth element in the middle layer of the value network of microgrid m at time t+1, represents the weight from the yth element in the middle layer of the value network of microgrid m to the output at time t+1, represents the weight from the uth element in the input layer to the vth element in the middle layer of the microgrid m-strategy network at time t+1, It represents the time t+1 from the vth element in the middle layer to the output layer in the strategy network of the microgrid m The weight of the element, It represents the partial derivative of the value of the weight of the yth element in the middle layer of the value network of microgrid m at time t to the output layer with respect to the square of the difference between the actual robust cumulative reward and the network robust cumulative reward, It represents the value partial derivative of the weight of the x-th element in the input layer to the y-th element in the middle layer of the value network of microgrid m at time t with respect to the square of the difference between the actual robust cumulative reward and the network robust cumulative reward, represents the symbol of partial derivative, It represents the time t from the vth element in the middle layer to the output layer in the strategy network of microgrid m The partial derivative of the element weight with respect to the network’s robust cumulative reward strategy, Represents the policy partial derivative of the weight from the u-th element in the input layer to the v-th element in the intermediate layer of the microgrid m strategy network at time t with respect to the network’s robust cumulative reward.

Citation Information

Patent Citations

  • Power distribution network optimization method based on multi-agent deep reinforcement learning

    CN114725936A

  • Micro-grid energy storage optimization scheduling method based on deep reinforcement learning

    CN117833285A