A dual-layer coordinated scheduling method and system for microgrid groups

By constructing a two-layer optimization scheduling model for microgrid clusters, using variational autoencoders to improve the SAC algorithm and introducing a time attenuation factor, and combining the residual balance method to improve the ADMM algorithm, the problems of slow convergence and inaccurate optimization results in the coordinated scheduling of microgrid clusters are solved, and more efficient coordinated scheduling of microgrid clusters is achieved.

CN119944649BActive Publication Date: 2025-09-26STATE GRID HUBEI ELECTRIC POWER CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510093308.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-09-26
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

In the existing technology, the research scope of deep reinforcement learning algorithms is limited to a single microgrid, and fails to fully consider the collaborative scheduling problem after multiple microgrids are interconnected to form a microgrid cluster, resulting in slow convergence speed and inaccurate optimization results.

Method used

A two-layer collaborative scheduling method for microgrid groups is proposed. By constructing a two-layer optimization scheduling model for microgrid groups, the SAC algorithm is improved by using variational autoencoders and the time decay factor is introduced. The empirical priority is dynamically adjusted, and the ADMM algorithm is improved based on the residual balance method. The penalty parameters are adaptively adjusted to optimize the scheduling strategy.

Benefits of technology

It significantly improves sample utilization efficiency and training convergence speed, enhances the efficiency and robustness of microgrid collaborative scheduling, simplifies the parameter adjustment process, and enhances adaptability and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119944649B_ABST
    Figure CN119944649B_ABST
Patent Text Reader

Abstract

The present invention discloses a microgrid cluster dual-layer collaborative scheduling method and system, belonging to the field of microgrid cluster collaborative scheduling technology. The method constructs a microgrid cluster dual-layer optimization scheduling model, including an upper-layer model and a lower-layer model. The upper-layer model is established with the goal of minimizing operating costs, while the lower-layer model is established with the goal of minimizing operating costs of sub-microgrids. A VAE is used to improve the SAC algorithm, and a time decay factor is introduced in the priority experience playback to dynamically adjust the priority of experiences, balance the sampling frequency of experiences entering the experience pool at different times, and solve the upper-layer model. The ADMM algorithm is improved based on the residual balance method, and the penalty parameter is adaptively adjusted to solve the lower-layer model. The optimized scheduling results after the adjustment of the upper-layer model and the optimized scheduling results of the lower-layer model are integrated to output the optimal microgrid cluster dual-layer optimization scheduling strategy. This method improves the convergence speed and optimization accuracy of the microgrid cluster dual-layer collaborative scheduling method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of microgrid cluster collaborative scheduling, and more specifically to a microgrid cluster dual-layer collaborative scheduling method and system. Background Art

[0002] With the widespread adoption of distributed clean energy, microgrids have gained widespread application as an effective means of absorbing renewable energy. However, the scheduling of a single microgrid is unable to meet the increasingly complex energy demands. By interconnecting geographically adjacent microgrids to form microgrid clusters, diverse distributed energy resources can be more effectively integrated and utilized. Microgrid clusters enable energy sharing and coordinated scheduling among microgrids, balancing regional power supply and demand. Therefore, research on coordinated scheduling methods for microgrid clusters can not only improve energy efficiency and reduce operating costs, but also optimize resource allocation and enhance system robustness.

[0003] Currently, existing approaches to optimizing microgrid cluster scheduling primarily include heuristic algorithms and reinforcement deep learning network algorithms. Traditional heuristic algorithms struggle to effectively handle the uncertainty of distributed generation, resulting in insufficient robustness in scheduling solutions. A microgrid coordinated optimization scheduling model based on a reinforcement deep learning network algorithm utilizes an experience replay mechanism and a frozen network parameter mechanism to enhance the performance of the deep reinforcement learning algorithm. This model also achieves energy management and optimization of micro-energy networks with economic efficiency as its primary objective. Simulation results demonstrate that, compared to heuristic algorithms, the deep reinforcement learning algorithm is more efficient at finding optimal solutions and continuously optimizing scheduling solutions.

[0004] However, the research scope of the deep reinforcement learning algorithm adopted by the above-mentioned existing technologies is limited to a single microgrid, and fails to fully consider the collaborative scheduling problem after multiple microgrids are interconnected to form a microgrid cluster, resulting in slow convergence speed and inaccurate optimization results. Summary of the Invention

[0005] In response to the problems existing in the above-mentioned fields, the present invention proposes a two-layer collaborative scheduling method and system for microgrid clusters, which can solve the technical problems that the research scope of the deep reinforcement learning algorithm adopted in the existing technology is limited to a single microgrid, and fails to fully consider the collaborative scheduling problem after multiple microgrids are interconnected to form a microgrid cluster, resulting in slow convergence speed and inaccurate optimization results.

[0006] To solve the above technical problems, the present invention discloses a two-layer coordinated scheduling method for a microgrid group, comprising the following steps:

[0007] Collect historical operating data on the power generation demand and energy storage status of the microgrid group;

[0008] A two-layer optimization scheduling model for a microgrid group is constructed; the two-layer optimization scheduling model for a microgrid group includes an upper-layer model and a lower-layer model; the upper-layer model is established with the goal of minimizing operating costs; the lower-layer model is established with the goal of minimizing operating costs of sub-microgrids, and the balance of power supply and demand is used as an indicator to measure overall scheduling decisions and fed back to the upper-layer model;

[0009] The low-dimensional latent variables of the variational autoencoder are used to replace the original state variables to improve the policy network and value network of the SAC algorithm. A time decay factor is introduced into the priority experience replay mechanism to dynamically adjust the priority of experience and balance the sampling frequency of experience entering the experience pool at different times. This results in an improved SAC algorithm for solving the upper-level model. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain the improved ADMM algorithm for solving the lower-level model.

[0010] The collected historical operating data of the microgrid group's power generation demand and energy storage status are input into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after adjustment of the upper model and the optimized scheduling results of the lower model;

[0011] The optimized scheduling results after adjustment of the upper-layer model and the optimized scheduling results of the lower-layer model are integrated to output the optimal two-layer optimized scheduling strategy for the microgrid group.

[0012] Preferably, the upper layer model is established with the goal of minimizing operating costs, comprising the following steps:

[0013] The objective function of the upper model is:

[0014]

[0015]

[0016] Where Z MMG is the total operating cost of the microgrid; C g is the operating cost of all power generation units in the microgrid; C ess is the operating cost of all energy storage systems in the microgrid; C dn The transaction cost between the microgrid and the upper-level power grid; Microgrid group and sub-microgrid MG n Transaction costs; N T 、N g 、N ess and N MG are the time period, the number of power generation units, the number of energy storage devices and the number of sub-microgrids respectively; ρ fuel and ρ maintare the fuel cost and maintenance cost per unit of electricity generated; and Respectively, unit g i startup and shutdown costs; For unit g i Output during period t; and They are start and stop states respectively. Indicates startup; Indicates shutdown; and They are the charging cost, discharging cost and maintenance cost per unit of electricity; and Energy storage system j The charging and discharging power; and are the purchase price and sales price of electricity when the microgrid group trades with the upper power grid during period t; and are the purchased power and sold power of the microgrid group and the upper power grid during the transaction in period t; and They are microgrid group and sub-microgrid MG n The purchase price and sale price of electricity during the transaction in period t; and They are microgrid group and sub-microgrid MG n The purchase and sale of electricity during the t period of trading;

[0017] The constraints of the objective function of the upper model are as follows:

[0018] Microgrid power balance constraints:

[0019]

[0020] Energy storage constraints:

[0021]

[0022] Interactive power constraints:

[0023]

[0024] Power supply and demand balance constraints:

[0025]

[0026] Where, is the load demand of the microgrid group in period t; and are the maximum charging and discharging powers of energy storage system j, respectively; is the energy state of energy storage system j in period t; η essj,c and η essj,d are the charging and discharging efficiencies of energy storage system j, respectively; is the maximum energy state of energy storage system j; and are the minimum and maximum values ​​of the state of charge, respectively; and are the minimum and maximum values ​​of the interaction power between the microgrid group and the upper power grid respectively; and They are microgrid group and sub-microgrid MG n The minimum and maximum values ​​of the interaction power; N represents the number of microgrids. and are the actual power generation and actual storage capacity of microgrid i; P exchange,i is the total interactive power of microgrid i; L i is the load demand of microgrid i; σ is the power balance error.

[0027] Preferably, the lower layer model is established with the goal of minimizing the operating cost of the sub-microgrid, and includes the following steps:

[0028] The objective function of the lower model is:

[0029]

[0030] Where, MG n Operating costs; C ld is the operating cost of all loads; C MMG is the transaction cost between the sub-microgrid and the microgrid group; is the demand cost coefficient of load k; is the required power of load k; and The distribution is the purchase price and sales price of electricity when the sub-microgrid and the microgrid group trade in period t; and are the purchased power and sold power of the sub-microgrid and microgrid group during the transaction in period t, respectively;

[0031] The constraints of the objective function of the lower model include:

[0032] Sub-microgrid power balance constraints:

[0033]

[0034] Sub-microgrid switching power constraints:

[0035]

[0036] Unit output climbing constraint:

[0037]

[0038] Where, is the interaction power between the sub-microgrid and the microgrid group; is the total load demand during period t; is the interaction power between sub-microgrids; and are the minimum and maximum values ​​of the interaction power between the sub-microgrid and the microgrid group, respectively; and are the minimum and maximum values ​​of the interaction power between sub-microgrids, respectively; and are the minimum and maximum output power of the fuel cell, respectively; and are the maximum values ​​of the climbing output of the fuel cell and diesel generator respectively; and are the minimum and maximum output power of the diesel generator respectively.

[0039] Preferably, the improvement of the strategy network and value network of the SAC algorithm specifically includes:

[0040] The original state variable s obtained from the environment is input into the trained VAE model to obtain a low-dimensional latent variable z. The low-dimensional latent variable z is used in the SAC algorithm to replace the original state variable s for strategy optimization and value evaluation.

[0041] The expression of the policy network function of the improved SAC algorithm is:

[0042]

[0043] In the formula, q(z t |s t ) is a given state s t The latent variable z under t The probability distribution of π θ (a t |z t ) is the policy network in a given state z t Output action a t The probability of; α is the temperature parameter, which controls the weight of the entropy term in the objective function; Q φ (z t ,a t ) is the output value of the value network; is the latent variable z generated by the conditional probability distribution t The mathematical expectation of To select action a according to the current strategy tThe mathematical expectation of

[0044] The expression of the value network function of the improved SAC algorithm is:

[0045]

[0046] Where D is the experience replay pool; a t+1 and z t+1 are the state and action of the agent at time t+1 respectively; r(z t ,a t ) is the agent in a given state z t Output action a t The reward after Q φ (z t ,a t ) is the output value of the target value function.

[0047] Preferably, balancing the sampling frequencies of experiences entering the experience pool at different times comprises the following steps:

[0048] The empirical priority formula after introducing the time decay factor is:

[0049] δ i =r+γQ(z t+1 ,a t+1 )-Q(z t ,a t )

[0050]

[0051] Where β is the time attenuation factor, and its value range is (0,1); δ i is the TD error value of state transition i; ε is a constant to prevent the priority from being zero; t i is the experience τ i The number of time steps into the experience pool; r is the action a taken by the agent at the current time step t The reward obtained later; γ is the discount factor, which measures the importance of future rewards and current rewards, and its value range is [0,1]; Q(z t+1 ,a t+1 ) is the next state z t+1 and action a t+1 Q value; Q(z t ,a t ) is the current state z t and action a t Q value; p(τ i ) is the empirical τ i Adjusted priorities;

[0052] Calculate the probability of each experience being sampled based on the adjusted priority:

[0053]

[0054] Where N is the total number of experiences in the experience pool, is the total priority of all experiences in the experience pool, P(τ i ) is the empirical τ i The probability of being adopted.

[0055] Preferably, solving the upper model comprises the following steps:

[0056] Step 1: Based on the upper model, define the Markov decision process MDP;

[0057] Define the MDP as a five-tuple<S,A,P,r,γ> , where S is the set of state spaces, A is the set of action spaces, P is the state transition probability, r is the reward function, and γ is the discount factor; the state transition probability is learned by the agent through its interaction with the environment, and the discount factor weighs the short-term and long-term effects of the current action;

[0058] In the state space, the state variables of the upper model are defined as the amount of electricity purchased, the power generation of each sub-microgrid, the energy storage status of each sub-microgrid, the load demand of each sub-microgrid, and the balance of power supply and demand of each sub-microgrid. The state space vector is expressed as:

[0059]

[0060] Where, P buy The amount of electricity purchased indicates the amount of electricity purchased by the microgrid group; is the power generation, which represents the amount of electricity generated by all power generation equipment in sub-microgrid i; is the energy storage state, which indicates the amount of electric energy stored in the energy storage system of sub-microgrid i; is the load demand, which represents the power demand of sub-microgrid i; is the power supply and demand balance, which represents the power supply and demand status of sub-microgrid i;

[0061] Action space, the action variables of the upper model are defined as the adjustment of each sub-microgrid's power purchase, power generation, energy storage, and power exchange between sub-microgrids. The action space vector is expressed as:

[0062]

[0063] Where, is the electricity purchase adjustment of microgrid i; The power generation adjustment of microgrid i; is the energy storage adjustment of microgrid i; is the power exchange amount adjustment of microgrid i;

[0064] The designed reward function is:

[0065]

[0066] ξ(k)=ξ max (1-e -μk )

[0067] Where c buy 、c gen and c storage are the cost coefficients of unit electricity purchase cost, unit power generation and unit energy storage respectively; n is the number of sub-microgrids included in the microgrid group; is the amount of electricity purchased by the microgrid group at the current moment t; is the power generation power of microgrid i; is the energy storage capacity of microgrid i; is the power supply and demand balance of microgrid i; ξ(k) is the penalty coefficient, where ξ max is the maximum penalty coefficient, μ is the adjustment speed parameter, which controls the growth rate of the penalty coefficient, k is the number of iterations, and the initial value is 0;

[0068] Step 1: Define the Markov decision process;

[0069] Step 2: Initialize the parameters of the policy network, value network, and target value network, and set the experience replay pool for the agent;

[0070] Step 3: According to the current state s t and the strategy selects action a t And execute, get the new state s t+1 and reward r t , transfer samples (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool;

[0071] Step 4: Randomly extract historical running data collected from the experience replay pool as sample data, calculate the target Q value and policy loss, and update the value network parameters and policy network parameters by minimizing the loss function;

[0072] Step 5: Repeat the process of data collection and policy optimization until the policy network converges and outputs the optimized scheduling results of the upper-level model.

[0073] Preferably, the improved ADMM algorithm comprises the following steps:

[0074] By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted;

[0075] The calculation formulas for the relative original residual and the relative dual residual include:

[0076] r (k) =Ax (k) +Bz (k) -c

[0077]

[0078] Where x∈R n and z∈R m are all variables in the optimization problem; c∈R p is a constant vector of linear constraints, representing the offset of the constraints; A∈R p×n and B∈R p×m are coefficient matrices of linear constraints; x and y are original and dual variables respectively; r (k) is the original residual, is the relative original residual; s (k) is the dual residual, is the relative dual residual; is the two-norm relative to the original residual, is the binorm of the relative dual residual; τ and γ are constants, is the penalty parameter;

[0079] Adaptively adjust the penalty parameter by calculating the relative original residual and the relative dual residual

[0080] Preferably, the output optimal microgrid group dual-layer optimization scheduling strategy includes the following steps:

[0081] Step 1: Collect historical operating data of the microgrid group, including power demand and energy storage status, and clean and standardize the collected historical operating data of the microgrid group;

[0082] Step 2: Initialize the environment, establish a two-layer collaborative scheduling model for the microgrid cluster, and define the Markov decision process;

[0083] Step 3: Initialize the parameters of the improved SAC algorithm and the improved ADMM algorithm, and set the experience replay pool to store the experience samples collected from the environment;

[0084] Step 4: Based on the current state, the agent selects an action based on the output of the policy network and executes it, and stores the current state, action, reward, and new state into the experience replay pool;

[0085] Step 5: Randomly select a batch of collected historical operation data of the microgrid group from the experience replay pool as sample data, input it into the value network and policy network, calculate the target Q value and loss function, and update the network parameters;

[0086] Step 6: Determine whether the loss function converges. If so, output the obtained power exchange amount. Otherwise, return to step 4 and execute the strategy optimization process.

[0087] Step 7: Iteratively update the primal residual and the dual residual according to the calculation formula of the relative primal residual and the relative dual residual; when the convergence condition is met, the optimal power generation and energy storage state of each sub-microgrid are obtained;

[0088] Step 8: The power exchange quantity solved by the upper model is passed to the lower model, and the power supply and demand balance of each sub-microgrid is calculated and fed back to the upper model;

[0089] Step 9: If the power supply and demand balance satisfies the constraints, the optimal scheduling result of the lower model is output; if the constraints are not met, the number of iterations k = k + 1 is adjusted, the state space is updated, the reward function is adjusted, and the process returns to step 4 to re-execute the policy optimization process until the power supply and demand balance satisfies the constraints. The optimized scheduling result after the adjustment of the upper model is obtained;

[0090] Step 10: Combine the optimized scheduling results after adjustment of the upper-layer model and the optimal scheduling results of the lower-layer model to output the optimal microgrid group two-layer optimization scheduling strategy.

[0091] Preferably, a microgrid group dual-layer collaborative scheduling system is also included, including:

[0092] Data collection module, used to collect historical operating data of the microgrid group's power generation demand and energy storage status;

[0093] A microgrid cluster two-layer optimization scheduling model construction module is used to construct a microgrid cluster two-layer optimization scheduling model; the microgrid cluster two-layer optimization scheduling model includes an upper-layer model and a lower-layer model; the upper-layer model is established with the goal of minimizing operating costs; the lower-layer model is established with the goal of minimizing operating costs of sub-microgrids, and the balance of power supply and demand is used as an indicator to measure overall scheduling decisions and fed back to the upper-layer model;

[0094] The solution algorithm improvement module is used to replace the original state variables with low-dimensional latent variables of the variational autoencoder, improve the policy network and value network of the SAC algorithm, and introduce a time decay factor in the priority experience replay mechanism to dynamically adjust the priority of experience and balance the sampling frequency of experience entering the experience pool at different times. This results in an improved SAC algorithm for solving the upper-level model. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain an improved ADMM algorithm for solving the lower-level model.

[0095] The data output module is used to input the collected historical operating data of the microgrid group's power generation demand and energy storage status into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after adjustment of the upper-level model and the optimized scheduling results of the lower-level model; the optimized scheduling results after adjustment of the upper-level model and the optimized scheduling results of the lower-level model are integrated to output the optimal two-level optimized scheduling strategy for the microgrid group.

[0096] Compared with the prior art, the present invention has the following beneficial effects:

[0097] The present invention proposes a two-layer collaborative scheduling method for a microgrid group. This method constructs a two-layer optimization scheduling model for a microgrid group. The SAC algorithm is improved using VAE, and low-dimensional latent variables are used instead of the original state variables. The policy network and value network of the SAC algorithm are improved, which can significantly improve the sample utilization efficiency and the training convergence speed, thereby showing better results in complex tasks. A time decay factor is introduced into the priority experience replay of the SAC algorithm to reduce the priority of experiences entering the experience pool early and increase the chance of experiences entering the experience pool later. The priority of experiences can be dynamically adjusted, effectively balancing the sampling frequency of experiences entering the experience pool at different times, thereby improving the sample adoption efficiency and the learning efficiency of the algorithm, and solving the upper-layer model. The ADMM algorithm is improved based on the residual balance method, and the penalty parameters are adaptively adjusted. By balancing the original residual and the dual residual, and by adaptively adjusting the penalty parameters, the improved ADMM algorithm converges more quickly and stably, and can effectively handle the optimization operation and resource allocation problems within the sub-microgrid. At the same time, no matter how the scale and numerical range of the problem change, this strategy can maintain the effectiveness and convergence of the algorithm, simplify the parameter adjustment process, and improve the adaptability and robustness of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] Figure 1 This is a flow chart of the dual-layer collaborative scheduling method for microgrid groups proposed by the present invention;

[0099] Figure 2 This is the optimized scheduling result of the microgrid group using the microgrid group double-layer collaborative scheduling method of the present invention. DETAILED DESCRIPTION

[0100] The following is a combination of the embodiments of the present invention Figure 1-Figure 2 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms used in the present invention are only used to describe specific implementation methods and are not intended to limit the present invention.

[0101] like Figure 1 As shown, the present invention proposes a microgrid dual-layer coordinated scheduling method, comprising the following steps:

[0102] The present invention collects historical operating data of the power generation demand and energy storage status of the microgrid group;

[0103] Taking into account the operation characteristics and control requirements of the microgrid group, a two-layer optimization scheduling model of the microgrid group is constructed, including: an upper model and a lower model.

[0104] The upper-level model is built with the goal of minimizing operating costs, determining the optimal energy scheduling solution and passing it to the lower-level model. The lower-level model is built with the goal of minimizing operating costs for each sub-microgrid, determining the specific operating strategy for each sub-microgrid, and executing the current strategy to calculate the power supply and demand balance for each sub-microgrid. This is then fed back to the upper-level model to help it optimize the scheduling strategy.

[0105] The low-dimensional latent variables of the variational autoencoder (VAE) are used to replace the original state variables, and the policy network and value network of the SAC algorithm are improved. A time decay factor is introduced into the priority experience replay mechanism of the SAC algorithm to dynamically adjust the priority of the experience and balance the sampling frequency of the experience entering the experience pool at different times. The improved SAC algorithm is used to solve the upper-level model to realize the overall energy scheduling decision of the microgrid group, thereby improving the system efficiency and generalization ability.

[0106] The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted and the lower-level model is solved. The improved ADMM algorithm can effectively handle the optimization operation and resource allocation problems within the sub-microgrid.

[0107] According to the collected historical operation data of the microgrid group's power generation demand and energy storage status, a batch of sample data is selected from the historical operation data and input into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results of the upper model and the lower model.

[0108] The optimized scheduling results of the upper-level model are input into the lower-level model to obtain the optimized scheduling results after adjustment of the upper-level model; the optimized scheduling results after adjustment of the upper-level model and the optimized scheduling results of the lower-level model are optimized and integrated to obtain the optimal collaborative scheduling optimization strategy.

[0109] Specifically, the objective function of the upper model is:

[0110]

[0111] Where Z MMG is the total operating cost of the microgrid; C g is the operating cost of all power generation units in the microgrid; C ess is the operating cost of all energy storage systems in the microgrid; C dn The transaction cost between the microgrid and the upper-level power grid; Microgrid group and sub-microgrid MG n Transaction costs; N T 、N g 、N ess and N MG are the time period, the number of power generation units, the number of energy storage devices and the number of sub-microgrids respectively; ρ fuel and ρ maint are the fuel cost and maintenance cost per unit of electricity generated; and Respectively, unit g i startup and shutdown costs; For unit g i Output during period t; and They are start and stop states respectively. Indicates startup; Indicates shutdown; and They are the charging cost, discharging cost and maintenance cost per unit of electricity; and Energy storage system j The charging and discharging power; and are the purchase price and sales price of electricity when the microgrid group trades with the upper power grid during period t; and are the purchased power and sold power of the microgrid group and the upper power grid during the transaction in period t; and They are microgrid group and sub-microgrid MG n The purchase price and sale price of electricity during the transaction in period t; and They are microgrid group and sub-microgrid MG nThe purchase and sale of electricity during the t period of trading.

[0112] The constraints of the objective function of the upper model are as follows:

[0113] Microgrid power balance constraints:

[0114]

[0115] Energy storage constraints:

[0116]

[0117] Interactive power constraints:

[0118]

[0119] Power supply and demand balance constraints:

[0120]

[0121] Where, is the load demand of the microgrid group in period t; and are the maximum charging and discharging powers of energy storage system j, respectively; is the energy state of energy storage system j in period t; η essj,c and η essj,d are the charging and discharging efficiencies of energy storage system j, respectively; is the maximum energy state of energy storage system j; and are the minimum and maximum values ​​of the state of charge, respectively; and are the minimum and maximum values ​​of the interaction power between the microgrid group and the upper power grid respectively; and They are microgrid group and sub-microgrid MG n The minimum and maximum values ​​of the interaction power; N represents the number of microgrids. and are the actual power generation and actual storage capacity of microgrid i; P exchange,i is the total interactive power of microgrid i; L i is the load demand of microgrid i; σ is the power balance error.

[0122] The objective function of the lower model is:

[0123]

[0124] Where, MG n Operating costs; C ldis the operating cost of all loads; C MMG is the transaction cost between the sub-microgrid and the microgrid group; is the demand cost coefficient of load k; is the required power of load k; and The distribution is the purchase price and sales price of electricity when the sub-microgrid and the microgrid group trade in period t; and They are respectively the purchased power and sold power of the sub-microgrid and the microgrid group during transactions in period t.

[0125] The constraints of the objective function of the lower model include:

[0126] Sub-microgrid power balance constraints:

[0127]

[0128] Sub-microgrid switching power constraints:

[0129]

[0130] Unit output climbing constraint:

[0131]

[0132] Where, is the interaction power between the sub-microgrid and the microgrid group; is the total load demand during period t; is the interaction power between sub-microgrids; and are the minimum and maximum values ​​of the interaction power between the sub-microgrid and the microgrid group, respectively; and are the minimum and maximum values ​​of the interaction power between sub-microgrids, respectively; and are the minimum and maximum output power of the fuel cell, respectively; and are the maximum values ​​of the climbing output of the fuel cell and diesel generator respectively; and are the minimum and maximum output power of the diesel generator respectively.

[0133] The SAC algorithm is improved based on the variational autoencoder VAE. At the same time, the time decay factor is introduced to improve the priority experience replay mechanism of the SAC algorithm to solve the upper-level model.

[0134] The SAC algorithm is an advanced reinforcement learning algorithm that aims to improve the exploration of policies by maximizing their entropy. However, the SAC algorithm still has shortcomings such as low sample utilization when dealing with high-dimensional states.

[0135] Therefore, the present invention improves the SAC algorithm based on the variational autoencoder VAE, improves the utilization efficiency of samples, and further improves the convergence speed of the algorithm.

[0136] (1)VAE

[0137] VAE is a generative model that captures the underlying structure of data by learning latent representations of states. It collects variables from the environment, including the current state, action, reward, and next state. It designs the VAE network structure, including an encoder and a decoder. The encoder compresses high-dimensional data into a low-dimensional latent space, while the decoder reconstructs the data in the low-dimensional latent space into high-dimensional data. It also defines its loss function, which optimizes the VAE parameters by minimizing the reconstruction error and KL divergence, so that the compressed low-dimensional latent space retains as much information as possible from the original state. The specific expression of VAE is:

[0138]

[0139] Where s is the original data; is the reconstructed data obtained by the decoder; β is the parameter that controls the trade-off between reconstruction error and KL divergence; D KL is the KL divergence between the approximate posterior distribution q(z|s) and the prior distribution p(z) of the latent variable z.

[0140] The VAE is trained using the collected historical running data, and the optimal encoder and decoder parameters are obtained by minimizing the loss function. The trained VAE can retain the main features of high-dimensional data and provide effective low-dimensional representation for subsequent tasks.

[0141] The original state variable s obtained from the environment is input into the trained VAE model to obtain a low-dimensional latent variable z. The low-dimensional latent variable z is used in the SAC algorithm to replace the original state variable s for strategy optimization and value evaluation.

[0142] The expression of the policy network function of the improved SAC algorithm is:

[0143]

[0144] In the formula, q(z t |s t ) is a given state s t The latent variable z under t The probability distribution of π θ (a t |z t ) is the policy network in a given state z t Output action a tThe probability of; α is the temperature parameter, which controls the weight of the entropy term in the objective function; Q φ (z t ,a t ) is the output value of the value network; is the latent variable z generated by the conditional probability distribution t The mathematical expectation of To select action a according to the current strategy t The mathematical expectation of .

[0145] The expression of the value network function of the improved SAC algorithm is:

[0146]

[0147]

[0148] Where D is the experience replay pool; a t+1 and z t+1 are the state and action of the agent at time t+1 respectively; r(z t ,a t ) is the agent in a given state z t Output action a t The reward after Q φ (z t ,a t ) is the output value of the target value function.

[0149] By combining VAE and SAC algorithms and using low-dimensional latent variables to improve the policy network and value network of the SAC algorithm, the sample utilization efficiency and training convergence speed can be significantly improved, thereby achieving better results in complex tasks.

[0150] (2) Improve the priority experience replay mechanism of the SAC algorithm

[0151] Current preferential experience replay measures focus on acquiring high-value experiences from the experience pool, but fail to consider the order in which experiences enter the pool, potentially leading to an imbalance in the frequency of adoption. Specifically, experiences that enter the pool earlier are more likely to be sampled frequently, while experiences that enter later may be ignored.

[0152] However, the experience that enters later may be more valuable than the experience that enters the experience pool earlier and is frequently sampled. Therefore, based on the priority experience playback, the present invention introduces a time decay factor to reduce the priority of the experience that enters the experience pool earlier and increase the chance of the experience that enters the experience pool later.

[0153] Set the time decay factor to β, with a value range of (0,1). The empirical priority formula after introducing the time decay factor is:

[0154] δ i =r+γQ(z t+1 ,a t+1 )-Q(z t ,a t )

[0155]

[0156] Where, δ i is the TD error value of state transition i; ε is a constant to prevent the priority from being zero; t i is the experience τ i The number of time steps into the experience pool; r is the action a taken by the agent at the current time step t The reward obtained later; γ is the discount factor, which measures the importance of future rewards and current rewards, and its value range is [0,1]; Q(z t+1 ,a t+1 ) is the next state z t+1 and action a t+1 Q value; Q(z t ,a t ) is the current state z t and action a t Q value; p(τ i ) is the empirical τ i The adjusted priority.

[0157] Calculate the probability of each experience being sampled based on the adjusted priority:

[0158]

[0159] Where N is the total number of experiences in the experience pool, is the total priority of all experiences in the experience pool, P(τ i ) is the empirical τ i The probability of being adopted.

[0160] By introducing the time decay factor, the priority of experience can be dynamically adjusted, effectively balancing the sampling frequency of experiences entering the experience pool at different times, thereby improving the sample adoption efficiency and the learning efficiency of the algorithm.

[0161] The SAC algorithm is improved based on VAE, and the time decay factor is introduced to improve the priority experience replay mechanism to achieve the optimization solution of the upper model. The specific steps include:

[0162] Step 1: Based on the upper model, define the Markov decision process MDP. Generally, the MDP is defined as a five-tuple<S,A,P,r,γ> , where S is the set of state spaces, A is the set of action spaces, P is the state transition probability, r is the reward function, and γ is the discount factor, where:

[0163] The state transition probabilities are usually learned by the agent itself through interactions with the environment, while the discount factor weighs the short-term and long-term effects of the current action.

[0164] State space, which is usually the variable that affects the operating state of the microgrid group. This paper defines the state variables of the upper model as the amount of electricity purchased, the power generation of each sub-microgrid, the energy storage status of each sub-microgrid, the load demand of each sub-microgrid, and the balance of power supply and demand of each sub-microgrid. The state space vector is expressed as:

[0165]

[0166] Where, P buy The amount of electricity purchased indicates the amount of electricity purchased by the microgrid group; is the power generation, which represents the amount of electricity generated by all power generation equipment in sub-microgrid i; is the energy storage state, which indicates the amount of electric energy stored in the energy storage system of sub-microgrid i; is the load demand, which represents the power demand of sub-microgrid i; is the power supply and demand balance, which represents the power supply and demand status of sub-microgrid i.

[0167] Action space. This paper defines the action variables of the upper model as the adjustment of each sub-microgrid’s power purchase, power generation, energy storage, and power exchange between sub-microgrids. The action space vector is expressed as:

[0168]

[0169] Where, is the electricity purchase adjustment of microgrid i; The power generation adjustment of microgrid i; is the energy storage adjustment of microgrid i; is the power exchange adjustment of microgrid i.

[0170] The optimization goal of the upper-level model established in this invention is to minimize the operating cost. Based on this optimization goal, the reward function designed is:

[0171]

[0172] ξ(k)=ξ max (1-e -μk )

[0173] Where c buy 、c gen and c storage are the cost coefficients of unit electricity purchase cost, unit power generation and unit energy storage respectively; n is the number of sub-microgrids included in the microgrid group; is the amount of electricity purchased by the microgrid group at the current moment t; is the power generation power of microgrid i; is the energy storage capacity of microgrid i; is the power supply and demand balance of microgrid i; ξ(k) is the penalty coefficient, where ξ max is the maximum penalty coefficient, μ is the adjustment speed parameter, which controls the growth rate of the penalty coefficient, k is the number of iterations, and the initial value is 0.

[0174] Using the improved SAC algorithm, the basic process of solving the upper model is as follows:

[0175] Step 1: Define the Markov decision process;

[0176] Step 2: Initialize the parameters of the policy network, value network, and target value network, and set the experience replay pool for the agent;

[0177] Step 3: According to the current state s t and the strategy selects action a t And execute, get the new state s t+1 and reward r t , transfer samples (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool;

[0178] Step 4: Randomly extract historical running data collected from the experience replay pool as sample data, calculate the target Q value and policy loss, and update the value network parameters and policy network parameters by minimizing the loss function;

[0179] Step 5: Repeat the process of data collection and strategy optimization until the strategy network converges and outputs the optimal solution of the upper-level model.

[0180] The present invention also improves the ADMM algorithm based on the residual balance method to obtain an improved ADMM algorithm for solving the lower layer model.

[0181] The ADMM algorithm is a commonly used optimization algorithm, but its convergence speed is highly dependent on the choice of penalty parameters. Inappropriate penalty parameters may lead to slow convergence or even no convergence.

[0182] Therefore, the present invention proposes an adaptive penalty parameter strategy based on residual balance, which adaptively adjusts the penalty parameter by balancing the original residual and the dual residual, thereby improving the convergence speed and performance of the ADMM algorithm.

[0183] Specifically, the present invention adaptively adjusts the penalty parameter by comparing the relative original residual and the relative dual residual. Using relative residuals can improve the robustness of the algorithm and automatically adjust the penalty parameter when the optimization problem formula changes. The calculation formulas for the relative original residual and the relative dual residual include:

[0184] r (k) =Ax (k) +Bz (k) -c

[0185]

[0186] Where x∈R n and z∈R m are all variables in the optimization problem; c∈R p is a constant vector of linear constraints, representing the offset of the constraints; A∈R p×n and B∈R p×m are coefficient matrices of linear constraints; x and y are original and dual variables respectively; r (k) is the original residual, is the relative original residual; s (k) is the dual residual, is the relative dual residual; is the two-norm relative to the original residual, is the binorm of the relative dual residual; τ and γ are constants, is the penalty parameter.

[0187] Adaptively adjust the penalty parameter by calculating the relative original residual and the relative dual residual This allows the improved ADMM algorithm to converge more quickly and stably. Furthermore, regardless of the scale and numerical range of the problem, this strategy maintains the algorithm's effectiveness and convergence, simplifies the parameter adjustment process, and improves the algorithm's adaptability and robustness.

[0188] By optimizing and integrating the optimized dispatch results after adjustment of the upper-layer model and the optimized dispatch results of the lower-layer model, the final optimal microgrid coordinated dispatch strategy is obtained, which includes the following steps:

[0189] Step 1: Collect historical operating data of the microgrid group, including power demand and energy storage status, and clean and standardize the collected historical operating data of the microgrid group;

[0190] Step 2: Initialize the environment, establish a two-layer collaborative scheduling model for the microgrid cluster, and define the Markov decision process;

[0191] Step 3: Initialize the parameters of the improved SAC algorithm and ADMM algorithm, and set the experience replay pool to store the experience samples collected from the environment;

[0192] Step 4: Based on the current state, the agent selects an action based on the output of the policy network and executes it, and stores the current state, action, reward, and new state into the experience replay pool;

[0193] Step 5: Randomly select a batch of collected historical operation data of the microgrid group from the experience replay pool as sample data, input it into the value network and policy network, calculate the target Q value and loss function, and update the network parameters;

[0194] Step 6: Determine whether the loss function converges. If so, output the obtained power exchange amount. Otherwise, return to step 4 and execute the strategy optimization process.

[0195] Step 7: Iteratively update the primal residual and the dual residual according to the calculation formula of the relative primal residual and the relative dual residual; when the convergence condition is met, the optimal power generation and energy storage state of each sub-microgrid are obtained;

[0196] Step 8: The power exchange quantity solved by the upper model is passed to the lower model, and the power supply and demand balance of each sub-microgrid is calculated and fed back to the upper model;

[0197] Step 9: If the power supply and demand balance satisfies the constraints, the optimal scheduling result of the lower model is output; if the constraints are not met, the number of iterations k=k+1 is adjusted, the state space is updated, the reward function is adjusted, and the process returns to step 4 to re-execute the policy optimization process until the power supply and demand balance satisfies the constraints. The optimized scheduling result after the adjustment of the upper model is obtained;

[0198] Step 10: Combine the optimized scheduling results after adjustment of the upper-layer model and the optimal scheduling results of the lower-layer model to output the optimal microgrid group two-layer optimization scheduling strategy.

[0199] The present invention proposes a microgrid group two-layer coordinated scheduling system, comprising:

[0200] Data collection module, used to collect historical operating data of the microgrid group's power generation demand and energy storage status;

[0201] A microgrid cluster two-layer optimization scheduling model construction module is used to construct a microgrid cluster two-layer optimization scheduling model; the microgrid cluster two-layer optimization scheduling model includes an upper-layer model and a lower-layer model; the upper-layer model is established with the goal of minimizing operating costs; the lower-layer model is established with the goal of minimizing operating costs of sub-microgrids, and the balance of power supply and demand is used as an indicator to measure overall scheduling decisions and fed back to the upper-layer model;

[0202] The solution algorithm improvement module is used to replace the original state variables with low-dimensional latent variables of the variational autoencoder, improve the policy network and value network of the SAC algorithm, and introduce a time decay factor in the priority experience replay mechanism to dynamically adjust the priority of experience and balance the sampling frequency of experience entering the experience pool at different times. This results in an improved SAC algorithm for solving the upper-level model. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain an improved ADMM algorithm for solving the lower-level model.

[0203] The data output module is used to input the collected historical operating data of the microgrid group's power generation demand and energy storage status into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after adjustment of the upper-level model and the optimized scheduling results of the lower-level model; the optimized scheduling results after adjustment of the upper-level model and the optimized scheduling results of the lower-level model are integrated to output the optimal two-level optimized scheduling strategy for the microgrid group.

[0204] The proposed dual-layer coordinated scheduling method for microgrid clusters constructs a dual-layer optimized scheduling model for microgrid clusters based on their operational characteristics and control requirements. The upper-layer model is established with the goal of minimizing operating costs, while the lower-layer model is established with the goal of minimizing operating costs for sub-microgrids. The power supply and demand balance is fed back to the upper-layer model as an indicator for measuring overall scheduling decisions, thereby rationally adjusting the upper-layer model's decisions. The SAC algorithm is improved based on VAE, and a time decay factor is introduced based on the prioritized experience replay mechanism to improve the SAC algorithm's experience replay mechanism. The improved VAE-SAC algorithm is used to solve the upper-layer model, achieving overall energy scheduling decisions for the microgrid cluster and improving system efficiency and generalization capabilities. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to solve the lower-layer model. The improved ADMM algorithm can effectively address optimal operation and resource allocation issues within sub-microgrids. By inputting the optimization results of the upper-level model into the lower-level model, the upper and lower-level scheduling results are optimized and integrated to obtain the final microgrid group collaborative scheduling plan.

[0205] Example

[0206] like Figure 1The figure shows a flow chart of the proposed two-tier coordinated scheduling method for microgrids. The method includes: constructing a two-tier optimized scheduling model for a microgrid cluster based on the operating characteristics and control requirements of the microgrid cluster. The upper-tier model is established with the goal of minimizing operating costs, while the lower-tier model is established with the goal of minimizing operating costs for sub-microgrids. The balance of power supply and demand is fed back to the upper-tier model as an indicator for measuring overall scheduling decisions, thereby rationally adjusting the upper-tier decisions. The SAC algorithm is improved based on VAE, and a time decay factor is introduced based on the prioritized experience replay mechanism to improve the SAC algorithm's experience replay mechanism. The improved SAC algorithm is used to solve the upper-tier model, achieving overall energy scheduling decisions for the microgrid cluster and improving system efficiency and generalization capabilities. The ADMM algorithm is improved based on the residual balance method to solve the lower-tier model. The improved ADMM algorithm can effectively address optimal operation and resource allocation within the sub-microgrids. The optimization results of the upper-tier model are input into the lower-tier model, optimizing and integrating the upper and lower-tier scheduling results to obtain the final coordinated scheduling solution.

[0207] like Figure 2 The figure shows the optimized scheduling results of each sub-microgrid in the microgrid group according to the method proposed in the present invention. By analyzing the data in the figure, it can be seen that Figure (a) shows that sub-microgrid 1 mainly acts as the power demander, while Figures (b) and (c) show that sub-microgrid 2 and sub-microgrid 3 respectively have the dual roles of power supply and demand. During the period of low photovoltaic power generation and high load demand, each sub-microgrid can release electricity through the energy storage system; during the peak period of photovoltaic power generation, the energy storage system can be charged, thereby improving energy utilization. In addition, during the period of high photovoltaic power generation, sub-microgrid 2 and sub-microgrid 3 can sell excess electricity to sub-microgrid 1. The rational distribution of electricity between different sub-microgrids through power sales and purchases further improves the stability and reliability of the overall power grid. Therefore, it can be concluded that the method proposed in the present invention can reasonably optimize the operation of the microgrid group and improve the operating efficiency and economic benefits of the microgrid group.

[0208] As shown in Table 1, the operating indicators of different microgrid cluster collaborative scheduling methods are compared. The present invention sets up four different schemes to verify the efficiency of the microgrid cluster two-layer collaborative scheduling method proposed in the present invention.

[0209] Scheme 1 is a microgrid cluster optimization scheduling method based on the PSO algorithm, Scheme 2 is a microgrid cluster optimization scheduling method based on the ADMM algorithm, Scheme 3 is a microgrid cluster optimization scheduling method based on the DQN algorithm, and Scheme 4 is a microgrid cluster optimization scheduling method proposed in this invention.

[0210] Table 1 Comparison of operating indicators of different microgrid collaborative scheduling methods

[0211] plan Operating cost / yuan Energy utilization rate / % Load balance rate / % Scheduling time / s PSO Algorithm 35899.56 85.4 94.6 20.3 ADMM algorithm 34325.22 88.6 96.3 18.6 DQN Algorithm 32753.12 92.2 97.5 15.4 The method proposed by the present invention 28016.36 97.2 98.9 13.7

[0212] By analyzing the data in the table, it can be seen that the operating cost of Scheme 4, that is, the microgrid cluster optimization scheduling method proposed by the present invention, is 28,016.36 yuan, which is reduced by 21.94%, 18.38%, and 14.45% compared with Schemes 1, 2, and 3, respectively; the energy utilization rate of Scheme 4 is 97.2%, which is increased by 13.85%, 9.72%, and 5.43% compared with Schemes 1, 2, and 3, respectively; the load balancing rate of Scheme 4 is 98.9%, which is increased by 4.55%, 2.70%, and 1.44% compared with Schemes 1, 2, and 3, respectively; the scheduling time of Scheme 4 is 13.7 seconds, which is reduced by 32.51%, 26.34%, and 11.04% compared with Schemes 1, 2, and 3, respectively.

[0213] It can be seen that the method proposed in the present invention shows significant advantages in terms of operating cost, energy utilization, load balancing rate and scheduling time.

[0214] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

[0215] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.

Claims

1. A two-layer coordinated scheduling method for a microgrid group, characterized in that: The following steps are involved: Collect historical operating data on the power generation demand and energy storage status of the microgrid group; A two-layer optimization scheduling model for a microgrid group is constructed; the two-layer optimization scheduling model for a microgrid group includes an upper-layer model and a lower-layer model; the upper-layer model is established with the goal of minimizing operating costs; the lower-layer model is established with the goal of minimizing operating costs of sub-microgrids, and the balance of power supply and demand is used as an indicator to measure overall scheduling decisions and fed back to the upper-layer model; The low-dimensional latent variables of the variational autoencoder are used to replace the original state variables. The policy network and value network of the SAC algorithm are improved. A time decay factor is introduced into the priority experience replay mechanism to dynamically adjust the priority of experience and balance the sampling frequency of experience entering the experience pool at different times. This improves the SAC algorithm and solves the upper model. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain the improved ADMM algorithm and solve the underlying model. The collected historical operating data of the microgrid group's power generation demand and energy storage status are input into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after adjustment of the upper model and the optimized scheduling results of the lower model; The optimized scheduling results after adjustment of the upper-layer model and the optimized scheduling results of the lower-layer model are integrated to output the optimal two-layer optimized scheduling strategy for the microgrid group.

2. The microgrid dual-layer coordinated scheduling method according to claim 1, characterized in that: The upper layer model is established with the goal of minimizing operating costs, and includes the following steps: The objective function of the upper model is: Where Z MMG is the total operating cost of the microgrid; C g is the operating cost of all power generation units in the microgrid; C ess is the operating cost of all energy storage systems in the microgrid; C dn The transaction cost between the microgrid and the upper-level power grid; Microgrid group and sub-microgrid MG n Transaction costs; N T 、N g 、N ess and N MG are the time period, the number of power generation units, the number of energy storage devices and the number of sub-microgrids respectively; ρ fuel and ρ maint are the fuel cost and maintenance cost per unit of electricity generated; and Respectively, unit g i startup and shutdown costs; For unit g i Output during period t; and They are start and stop states respectively. Indicates startup; Indicates shutdown; and They are the charging cost, discharging cost and maintenance cost per unit of electricity; and Energy storage system j The charging and discharging power; and are the purchase price and sales price of electricity when the microgrid group trades with the upper power grid during period t; and are the purchased power and sold power of the microgrid group and the upper power grid during the transaction in period t; and They are microgrid group and sub-microgrid MG n The purchase price and sale price of electricity during the transaction in period t; and They are microgrid group and sub-microgrid MG n The purchase and sale of electricity during the t period of trading; The constraints of the objective function of the upper model are as follows: Microgrid power balance constraints: Energy storage constraints: Interactive power constraints: Power supply and demand balance constraints: Where, is the load demand of the microgrid group in period t; and are the maximum charging and discharging powers of energy storage system j, respectively; is the energy state of energy storage system j in period t; η essj,c and η essj,d are the charging and discharging efficiencies of energy storage system j, respectively; is the maximum energy state of energy storage system j; and are the minimum and maximum values ​​of the state of charge, respectively; and are the minimum and maximum values ​​of the interaction power between the microgrid group and the upper power grid respectively; and They are microgrid group and sub-microgrid MG n The minimum and maximum values ​​of the interaction power; N represents the number of microgrids. and are the actual power generation and actual storage capacity of microgrid i; P exchange,i is the total interactive power of microgrid i; L i is the load demand of microgrid i; σ is the power balance error.

3. The microgrid dual-layer coordinated scheduling method according to claim 2, characterized in that: The lower layer model is established with the goal of minimizing the operating cost of the sub-microgrid, and includes the following steps: The objective function of the lower model is: Where, MG n Operating costs; C ld is the operating cost of all loads; C MMG is the transaction cost between the sub-microgrid and the microgrid group; is the demand cost coefficient of load k; is the required power of load k; and The distribution is the purchase price and sales price of electricity when the sub-microgrid and the microgrid group trade in period t; and are the purchased power and sold power of the sub-microgrid and microgrid group during transactions in period t, respectively; The constraints of the objective function of the lower model include: Sub-microgrid power balance constraints: Sub-microgrid switching power constraints: Unit output climbing constraint: Where, is the interaction power between the sub-microgrid and the microgrid group; is the total load demand during period t; is the interaction power between sub-microgrids; and are the minimum and maximum values ​​of the interaction power between the sub-microgrid and the microgrid group, respectively; and are the minimum and maximum values ​​of the interaction power between sub-microgrids, respectively; and are the minimum and maximum output power of the fuel cell, respectively; and are the maximum values ​​of the climbing output of the fuel cell and diesel generator respectively; and are the minimum and maximum output power of the diesel generator respectively.

4. The microgrid dual-layer coordinated scheduling method according to claim 3, characterized in that: The improvements to the strategy network and value network of the SAC algorithm specifically include: The original state variable s obtained from the environment is input into the trained VAE model to obtain a low-dimensional latent variable z. The low-dimensional latent variable z is used in the SAC algorithm to replace the original state variable s for strategy optimization and value evaluation. The expression of the policy network function of the improved SAC algorithm is: In the formula, q(z t |s t ) is a given state s t The latent variable z under t The probability distribution of π θ (a t |z t ) is the policy network in a given state z t Output action a t The probability of; α is the temperature parameter, which controls the weight of the entropy term in the objective function; Q φ (z t ,a t ) is the output value of the value network; is the latent variable z generated by the conditional probability distribution t The mathematical expectation of To select action a according to the current strategy t The mathematical expectation of The expression of the value network function of the improved SAC algorithm is: Where D is the experience replay pool; a t+1 and z t+1 are the state and action of the agent at time t+1 respectively; r(z t ,a t ) is the agent in a given state z t Output action a t The reward after Q φ (z t ,a t ) is the output value of the target value function.

5. The microgrid dual-layer coordinated scheduling method according to claim 4, characterized in that: The method of balancing the sampling frequencies of experiences entering the experience pool at different times comprises the following steps: The empirical priority formula after introducing the time decay factor is: δ i =r+γQ(z t+1 ,a t+1 )-Q(z t ,a t ) Where β is the time attenuation factor, and its value range is (0,1); δ i is the TD error value of state transition i; ε is a constant to prevent the priority from being zero; t i is the experience τ i The number of time steps into the experience pool; r is the action a taken by the agent at the current time step t The reward obtained later; γ is the discount factor, which measures the importance of future rewards and current rewards, and its value range is [0,1]; Q(z t+1 ,a t+1 ) is the next state z t+1 and action a t+1 Q value; Q(z t ,a t ) is the current state z t and action a t Q value; p(τ i ) is the empirical τ i Adjusted priorities; Calculate the probability of each experience being sampled based on the adjusted priority: Where N is the total number of experiences in the experience pool, is the total priority of all experiences in the experience pool, P(τ i ) is the empirical τ i Probability of being adopted.

6. The microgrid dual-layer coordinated scheduling method according to claim 5, characterized in that: Solving the upper model includes the following steps: Step 1: Based on the upper model, define the Markov decision process MDP; Define an MDP as a five-tuple<S,A,P,r,γ> , where S is the set of state spaces, A is the set of action spaces, P is the state transition probability, r is the reward function, and γ is the discount factor; the state transition probability is learned by the agent through its interaction with the environment, and the discount factor weighs the short-term and long-term effects of the current action; In the state space, the state variables of the upper model are defined as the amount of electricity purchased, the power generation of each sub-microgrid, the energy storage status of each sub-microgrid, the load demand of each sub-microgrid, and the balance of power supply and demand of each sub-microgrid. The state space vector is expressed as: Where, P buy The amount of electricity purchased indicates the amount of electricity purchased by the microgrid group; is the power generation, which represents the amount of electricity generated by all power generation equipment in sub-microgrid i; is the energy storage state, which indicates the amount of electric energy stored in the energy storage system of sub-microgrid i; is the load demand, which represents the power demand of sub-microgrid i; is the power supply and demand balance, which represents the power supply and demand status of sub-microgrid i; Action space, the action variables of the upper model are defined as the adjustment of each sub-microgrid's power purchase, power generation, energy storage, and power exchange between sub-microgrids. The action space vector is expressed as: Where, is the electricity purchase adjustment of microgrid i; is the power generation adjustment of microgrid i; is the energy storage adjustment of sub-microgrid i; is the power exchange amount adjustment of microgrid i; The designed reward function is: ξ(k)=ξ max (1-e -μk ) Where c buy 、c gen and c storage are the cost coefficients of unit electricity purchase cost, unit power generation and unit energy storage respectively; n is the number of sub-microgrids included in the microgrid group; is the amount of electricity purchased by the microgrid group at the current moment t; is the power generation power of microgrid i; is the energy storage capacity of microgrid i; is the power supply and demand balance of microgrid i; ξ(k) is the penalty coefficient, where ξ max is the maximum penalty coefficient, μ is the adjustment speed parameter, which controls the growth rate of the penalty coefficient, k is the number of iterations, and the initial value is 0; Step 1: Define the Markov decision process; Step 2: Initialize the parameters of the policy network, value network, and target value network, and set the experience replay pool for the agent; Step 3: According to the current state s t and the strategy selects action a t And execute, get the new state s t+1 and reward r t , transfer samples (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool; Step 4: Randomly extract historical running data collected from the experience replay pool as sample data, calculate the target Q value and policy loss, and update the value network parameters and policy network parameters by minimizing the loss function; Step 5: Repeat the process of data collection and policy optimization until the policy network converges and outputs the optimized scheduling results of the upper-level model.

7. The microgrid dual-layer coordinated scheduling method according to claim 6, characterized in that: The improved ADMM algorithm comprises the following steps: By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted; The calculation formulas for the relative original residual and the relative dual residual include: r (k) =Ax (k) +Bz (k) -c Where x∈R n and z∈R m are all variables in the optimization problem; c∈R p is a constant vector of linear constraints, representing the offset of the constraints; A∈R p×n and B∈R p×m are coefficient matrices of linear constraints; x and y are original and dual variables respectively; r (k) is the original residual, is the relative original residual; s (k) is the dual residual, is the relative dual residual; is the two-norm relative to the original residual, is the binorm of the relative dual residual; τ and γ are constants, is the penalty parameter; Adaptively adjust the penalty parameter by calculating the relative original residual and the relative dual residual 8. The microgrid dual-layer coordinated scheduling method according to claim 7, characterized in that: The dual-layer optimization scheduling strategy for the microgrid group with the best output includes the following steps: Step 1: Collect historical operating data of the microgrid group, including power demand and energy storage status, and clean and standardize the collected historical operating data of the microgrid group; Step 2: Initialize the environment, establish a two-layer collaborative scheduling model for the microgrid cluster, and define the Markov decision process; Step 3: Initialize the parameters of the improved SAC algorithm and the improved ADMM algorithm, and set the experience replay pool to store the experience samples collected from the environment; Step 4: Based on the current state, the agent selects an action based on the output of the policy network and executes it, and stores the current state, action, reward, and new state into the experience replay pool; Step 5: Randomly select a batch of collected historical operation data of the microgrid group from the experience replay pool as sample data, input it into the value network and policy network, calculate the target Q value and loss function, and update the network parameters; Step 6: Determine whether the loss function converges. If so, output the obtained power exchange amount. Otherwise, return to step 4 and execute the strategy optimization process. Step 7: Iteratively update the primal residual and the dual residual according to the calculation formula of the relative primal residual and the relative dual residual; when the convergence condition is met, the optimal power generation and energy storage state of each sub-microgrid are obtained; Step 8: The power exchange amount solved by the upper model is passed to the lower model, and the power supply and demand balance of each sub-microgrid is calculated and fed back to the upper model; Step 9: If the power supply and demand balance satisfies the constraints, the optimal scheduling result of the lower model is output; if the constraints are not met, the number of iterations k=k+1 is adjusted, the state space is updated, the reward function is adjusted, and the process returns to step 4 to re-execute the policy optimization process until the power supply and demand balance satisfies the constraints. The optimized scheduling result after the adjustment of the upper model is obtained; Step 10: Combine the optimized scheduling results after adjustment of the upper-layer model and the optimal scheduling results of the lower-layer model to output the optimal microgrid group two-layer optimization scheduling strategy.

9. A microgrid dual-layer collaborative dispatching system, characterized in that: include: Data collection module, used to collect historical operating data of the microgrid group's power generation demand and energy storage status; A microgrid cluster two-layer optimization scheduling model construction module is used to construct a microgrid cluster two-layer optimization scheduling model; the microgrid cluster two-layer optimization scheduling model includes an upper-layer model and a lower-layer model; the upper-layer model is established with the goal of minimizing operating costs; the lower-layer model is established with the goal of minimizing operating costs of sub-microgrids, and the balance of power supply and demand is used as an indicator to measure overall scheduling decisions and fed back to the upper-layer model; The solution algorithm improvement module is used to use the low-dimensional latent variables of the variational autoencoder to replace the original state variables, improve the policy network and value network of the SAC algorithm, and introduce a time decay factor in the priority experience playback mechanism to dynamically adjust the priority of experience and balance the sampling frequency of experience entering the experience pool at different times. This results in an improved SAC algorithm and solves the upper-level model. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain the improved ADMM algorithm and solve the underlying model. The data output module is used to input the collected historical operating data of the microgrid group's power generation demand and energy storage status into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after adjustment of the upper model and the optimized scheduling results of the lower model; The optimized scheduling results after adjustment of the upper-layer model and the optimized scheduling results of the lower-layer model are integrated to output the optimal two-layer optimized scheduling strategy for the microgrid group.

Citation Information

Patent Citations

  • Microgrid group optimization scheduling strategy based on niche chaos particle swarm algorithm

    CN112821470A

  • Multi-microgrid collaborative optimization scheduling strategy considering energy complementation and power transmission loss

    CN119029842A