Micro-grid group double-layer cooperative scheduling method and system
By constructing a two-layer optimization scheduling model for microgrid groups and improving related algorithms, the problem of low efficiency of collaborative scheduling in the existing technology is solved, and more efficient energy scheduling and resource allocation are achieved.
Patent Information
- Application Number
- CN202510093308.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The scope of research on deep reinforcement learning algorithms in the prior art is limited to a single microgrid, and the problem of coordinated scheduling after multiple microgrids are interconnected to form microgrid groups is not fully considered, resulting in slow convergence speed and inaccurate results in optimization.
A two-layer collaborative scheduling method for microgrid group is proposed. By constructing a two-layer optimization scheduling model for microgrid group, the SAC algorithm is improved using a variational autoencoder, the time attenuation factor is introduced, the priority of experience is dynamically adjusted, and the ADMM algorithm is improved based on the residual balance method, and the penalty parameters are adaptively adjusted.
It significantly improves the sample utilization efficiency and training convergence speed, improves the learning efficiency and adaptability of the algorithm, and can handle the optimization operation and resource allocation problems within the sub-micronet more quickly and stably.
Smart Images

Figure CN119944649A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of microgrid group coordinated scheduling, and more specifically to a microgrid group double-layer coordinated scheduling method and system. Background Art
[0002] With the large-scale popularization of distributed clean energy, microgrids have been widely used as an effective way to absorb renewable energy. However, the scheduling of a single microgrid is difficult to cope with the increasingly complex energy demand. By interconnecting geographically adjacent microgrids to form a microgrid cluster, a variety of distributed energy sources can be more effectively integrated and utilized. The formation of a microgrid cluster can achieve energy sharing and coordinated scheduling among microgrids, and balance the power supply and demand in the region. Therefore, studying the coordinated scheduling method of microgrid clusters can not only improve energy utilization efficiency and reduce operating costs, but also optimize resource allocation and enhance the robustness of the system.
[0003] At present, the microgrid group optimization scheduling methods studied in the prior art mainly include heuristic algorithms and reinforced deep learning network algorithms. Among them, traditional heuristic algorithms are difficult to effectively handle the uncertainty of distributed power generation, resulting in insufficient robustness of the scheduling scheme. The microgrid coordinated optimization scheduling model based on the reinforced deep learning network algorithm adopts the experience replay mechanism and the frozen network parameter mechanism to improve the performance of the deep reinforcement learning algorithm, and realizes the energy management and optimization of the micro energy network with the goal of economy. The simulation results verify that compared with the heuristic algorithm, the deep reinforcement learning algorithm can find the optimal solution faster and can continuously optimize the scheduling scheme.
[0004] However, the research scope of the deep reinforcement learning algorithm adopted by the above-mentioned existing technologies is limited to a single microgrid, and fails to fully consider the coordinated scheduling problem after multiple microgrids are interconnected to form a microgrid group, resulting in slow convergence speed and inaccurate optimization results. Summary of the invention
[0005] In response to the problems existing in the above-mentioned fields, the present invention proposes a two-layer collaborative scheduling method and system for a microgrid group, which can solve the technical problems that the research scope of the deep reinforcement learning algorithm adopted in the prior art is limited to a single microgrid, and fails to fully consider the collaborative scheduling problem after multiple microgrids are interconnected to form a microgrid group, resulting in slow convergence speed and inaccurate optimization results.
[0006] In order to solve the above technical problems, the present invention discloses a microgrid group double-layer coordinated scheduling method, comprising the following steps:
[0007] Collect historical operating data of the microgrid’s power generation demand and energy storage status;
[0008] Constructing a two-layer optimization dispatching model for a microgrid group; the two-layer optimization dispatching model for a microgrid group includes an upper model and a lower model; the upper model is established with the goal of minimizing the operating cost; the lower model is established with the goal of minimizing the operating cost of the sub-microgrid, and the balance of power supply and demand is used as an indicator to measure the overall dispatching decision, which is fed back to the upper model;
[0009] The low-dimensional latent variables of the variational autoencoder are used to replace the original state variables, and the policy network and value network of the SAC algorithm are improved. The time decay factor is introduced into the priority experience playback mechanism to dynamically adjust the priority of the experience and balance the sampling frequency of the experience entering the experience pool at different times. The improved SAC algorithm is obtained to solve the upper model. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain the improved ADMM algorithm to solve the lower model.
[0010] The collected historical operating data of the power generation demand and energy storage status of the microgrid group are input into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after the adjustment of the upper model and the optimized scheduling results of the lower model;
[0011] The optimized dispatching results after adjustment of the upper-layer model are integrated with the optimized dispatching results of the lower-layer model to output the optimal two-layer optimized dispatching strategy for the microgrid group.
[0012] Preferably, the upper model is established with the goal of minimizing operating costs, comprising the following steps:
[0013] The objective function of the upper model is:
[0014]
[0015]
[0016] In the formula, Z MMG is the total operating cost of the microgrid; C g is the operating cost of all power generation units in the microgrid; C ess is the operating cost of all energy storage systems in the microgrid; C dn The transaction cost between the microgrid and the upper grid; For microgrid group and sub-microgrid MG n Transaction costs; N T 、N g 、N ess and N MG are the time period, the number of power generation units, the number of energy storage devices and the number of sub-microgrids respectively; ρ fuel and ρ maintare respectively the fuel cost and maintenance cost per unit of electricity generated; and Respectively, unit g i startup and shutdown costs; For the unit g i Output during period t; and They are start and stop states respectively. Indicates startup; Indicates shutdown; and They are charging cost, discharging cost and maintenance cost per unit of electricity; and They are energy storage systems j The charging and discharging power; and are the purchase price and sales price of electricity when the microgrid group trades with the upper power grid during period t; and They are respectively the purchased power and sold power when the microgrid group trades with the upper power grid during period t; and They are microgrid group and sub-microgrid MG respectively. n The purchase and sale prices of electricity during the trading period t; and They are microgrid group and sub-microgrid MG respectively. n The purchase and sale of electricity during the t period of trading;
[0017] The constraints of the objective function of the upper model are as follows:
[0018] Microgrid power balance constraints:
[0019]
[0020] Energy storage constraints:
[0021]
[0022] Interactive power constraints:
[0023]
[0024] Power supply and demand balance constraints:
[0025]
[0026] In the formula, is the load demand of the microgrid group in period t; and are the maximum charging and discharging powers of energy storage system j, respectively; is the energy state of energy storage system j in period t; η essj,c and η essj,d are the charging and discharging efficiencies of energy storage system j, respectively; is the maximum energy state of energy storage system j; and are the minimum and maximum values of the state of charge, respectively; and are the minimum and maximum values of the interaction power between the microgrid and the upper grid, respectively; and They are microgrid group and sub-microgrid MG respectively. n The minimum and maximum values of the interaction power; N represents the number of microgrids, and are the actual power generation and storage capacity of microgrid i respectively; P exchange,i is the total interactive power of microgrid i; L i is the load demand of microgrid i; σ is the power balance error.
[0027] Preferably, the lower layer model is established with the goal of minimizing the operating cost of the sub-microgrid, and includes the following steps:
[0028] The objective function of the lower model is:
[0029]
[0030] In the formula, MG n The operating cost of ld is the operating cost of all loads; C MMG is the transaction cost between the sub-microgrid and the microgrid group; is the demand cost coefficient of load k; is the required power of load k; and The distribution is the purchase price and sale price of electricity when the sub-microgrid and the microgrid group trade in the t period; and They are the purchased power and sold power of the sub-microgrid and microgrid group during the transaction in period t, respectively;
[0031] The constraints of the objective function of the lower model include:
[0032] Sub-microgrid power balance constraints:
[0033]
[0034] Sub-microgrid switching power constraints:
[0035]
[0036] Unit output climbing constraints:
[0037]
[0038] In the formula, is the interaction power between the sub-microgrid and the microgrid group; is the total load demand in period t; is the interaction power between sub-microgrids; and are the minimum and maximum values of the interaction power between the sub-microgrid and the microgrid group, respectively; and are the minimum and maximum values of the interaction power between sub-microgrids, respectively; and are the minimum and maximum output power of the fuel cell, respectively; and are the maximum values of the climbing output of the fuel cell and the diesel generator respectively; and are the minimum and maximum output power of the diesel generator respectively.
[0039] Preferably, the improvement of the strategy network and value network of the SAC algorithm specifically includes:
[0040] The original state variable s obtained from the environment is input into the trained VAE model to obtain a low-dimensional latent variable z. The low-dimensional latent variable z is used in the SAC algorithm to replace the original state variable s for strategy optimization and value evaluation.
[0041] The expression of the policy network function of the improved SAC algorithm is:
[0042]
[0043] In the formula, q(z t |s t ) is a given state s t The latent variable z under t The probability distribution of π θ (a t |z t ) is the policy network in a given state z t Output action a t The probability of; α is the temperature parameter, which controls the weight of the entropy term in the objective function; Q φ (z t ,a t ) is the output value of the value network; is the latent variable z generated by the conditional probability distribution t The mathematical expectation of To select action a according to the current strategy tThe mathematical expectation of
[0044] The expression of the value network function of the improved SAC algorithm is:
[0045]
[0046] Where D is the experience replay pool; a t+1 and z t+1 are the state and action of the agent at time t+1; r(z t ,a t ) is the agent in a given state z t Output action a t The reward after Q φ (z t ,a t ) is the output value of the target value function.
[0047] Preferably, balancing the sampling frequency of experiences entering the experience pool at different times comprises the following steps:
[0048] The empirical priority formula after introducing the time decay factor is:
[0049] δ i = r + γQ (z t+1 ,a t+1 )-Q(z t ,a t )
[0050]
[0051] Where β is the time attenuation factor, and its value range is (0,1); δ i is the TD error value of state transition i; ε is a constant to prevent the priority from being zero; t i is the experience τ i The number of time steps into the experience pool; r is the action a taken by the agent at the current time step t The reward obtained later; γ is the discount factor, which measures the importance of future rewards and current rewards, and its value range is [0,1]; Q(z t+1 ,a t+1 ) is the next state z t+1 and action a t+1 Q value; Q(z t ,a t ) is the current state z t and action a t Q value; p(τ i ) is the empirical τ i The adjusted priority of
[0052] According to the adjusted priority, calculate the probability of each experience being sampled:
[0053]
[0054] Where N is the total amount of experience in the experience pool, is the total priority of all experiences in the experience pool, P(τ i ) is the empirical τ i Probability of being adopted.
[0055] Preferably, solving the upper model comprises the following steps:
[0056] Step 1: Define the Markov decision process MDP according to the upper model;
[0057] Define the MDP as a five-tuple<S,A,P,r,γ> , where S is the set of state spaces, A is the set of action spaces, P is the state transition probability, r is the reward function, and γ is the discount factor; the state transition probability is learned by the agent itself in the interaction with the environment, and the discount factor weighs the short-term and long-term effects of the current action;
[0058] In the state space, the state variables of the upper model are defined as the amount of electricity purchased, the power generation of each sub-microgrid, the energy storage status of each sub-microgrid, the load demand of each sub-microgrid, and the balance of power supply and demand of each sub-microgrid. The state space vector is expressed as:
[0059]
[0060] Where P buy The amount of electricity purchased indicates the amount of electricity purchased by the microgrid group; is the power generation, which represents the amount of electricity generated by all power generation equipment in the microgrid i; is the energy storage state, indicating the amount of electric energy stored in the energy storage system of sub-microgrid i; is the load demand, which represents the power demand of sub-microgrid i; is the balance degree of power supply and demand, which indicates the power supply and demand status of sub-microgrid i;
[0061] Action space, the action variables of the upper model are defined as the adjustment of power purchase, power generation, energy storage and power exchange between sub-microgrids. The action space vector is expressed as:
[0062]
[0063] In the formula, The power purchase adjustment of microgrid i; is the power generation adjustment of microgrid i; is the energy storage adjustment of sub-microgrid i; is the power exchange amount adjustment of sub-microgrid i;
[0064] The designed reward function is:
[0065]
[0066] ξ(k)=ξ max (1-e -μk )
[0067] In the formula, c buy 、c gen and c storage are the cost coefficients of unit electricity purchase cost, unit power generation and unit energy storage respectively; n is the number of sub-microgrids included in the microgrid group; is the amount of electricity purchased by the microgrid group at the current time t; is the power generation of microgrid i; is the energy storage capacity of microgrid i; is the power supply and demand balance of microgrid i; ξ(k) is the penalty coefficient, where ξ max is the maximum penalty coefficient, μ is the adjustment speed parameter, which controls the speed of the penalty coefficient growth, k is the number of iterations, and the initial value is 0;
[0068] Step 1: Define the Markov decision process;
[0069] Step 2: Initialize the parameters of the policy network, value network, and target value network, and set the experience replay pool for the agent;
[0070] Step 3: According to the current state s t and the strategy selects action a t And execute, get the new state s t+1 and reward r t , transfer samples (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool;
[0071] Step 4: Randomly extract the collected historical running data from the experience replay pool as sample data, calculate the target Q value and strategy loss, and update the value network parameters and strategy network parameters by minimizing the loss function;
[0072] Step 5: Repeat the process of data collection and strategy optimization until the strategy network converges and outputs the optimized scheduling results of the upper-level model.
[0073] Preferably, the improved ADMM algorithm comprises the following steps:
[0074] By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted;
[0075] Among them, the calculation formulas of relative primal residual and relative dual residual include:
[0076] r (k) =Ax (k) +Bz (k) -c
[0077]
[0078] Where x∈R n and z∈R m are all variables in the optimization problem; c∈R p is a constant vector of linear constraints, representing the offset of the constraints; A∈R p×n and B∈R p×m are coefficient matrices of linear constraints; x and y are the original and dual variables respectively; r (k) is the original residual, is the relative original residual; s (k) is the dual residual, is the relative dual residual; is the second norm relative to the original residual, is the bi-norm of the relative dual residual; τ and γ are constants, is the penalty parameter;
[0079] The penalty parameter is adaptively adjusted by the calculation formula of relative original residual and relative dual residual
[0080] Preferably, the output optimal microgrid group dual-layer optimization scheduling strategy includes the following steps:
[0081] Step 1: Collect historical operation data of the microgrid group, including power demand and energy storage status, and clean and standardize the collected historical operation data of the microgrid group;
[0082] Step 2: Initialize the environment, establish a two-layer collaborative scheduling model for microgrid clusters, and define the Markov decision process;
[0083] Step: 3: Initialize the parameters of the improved SAC algorithm and the improved ADMM algorithm, and set the experience replay pool to store the experience samples collected from the environment;
[0084] Step 4: Based on the current state, the agent selects and executes an action based on the output of the policy network, and stores the current state, action, reward, and new state in the experience replay pool;
[0085] Step 5: Randomly select a batch of historical operation data of the microgrid group collected from the experience replay pool as sample data, input the value network and the policy network, calculate the target Q value and loss function, and update the network parameters;
[0086] Step 6: Determine whether the loss function converges. If so, output the obtained power exchange amount; otherwise, return to step 4 to execute the strategy optimization process;
[0087] Step 7: According to the calculation formula of relative original residual and relative dual residual, iteratively update the original residual and dual residual; when the convergence condition is met, the optimal power generation power and energy storage state of each sub-microgrid are obtained;
[0088] Step 8: The power exchange amount solved by the upper model is passed to the lower model, and the power supply and demand balance of each sub-microgrid is calculated and fed back to the upper model;
[0089] Step 9: If the balance between power supply and demand meets the constraints, the optimal scheduling result of the lower model is output; if the constraints are not met, the number of iterations k=k+1 is adjusted, the state space is updated, the reward function is adjusted, and the process of policy optimization is re-executed until the balance between power supply and demand meets the constraints, and the optimal scheduling result after the adjustment of the upper model is obtained;
[0090] Step 10: Combine the optimized dispatching results after adjustment of the upper-layer model and the optimal dispatching results of the lower-layer model to output the optimal microgrid group two-layer optimized dispatching strategy.
[0091] Preferably, a microgrid group dual-layer collaborative dispatching system is also included, including:
[0092] A data collection module is used to collect historical operating data of the power generation demand and energy storage status of the microgrid group;
[0093] A microgrid group double-layer optimization dispatching model construction module is used to construct a microgrid group double-layer optimization dispatching model; the microgrid group double-layer optimization dispatching model includes an upper model and a lower model; the upper model is established with the goal of minimizing the operating cost; the lower model is established with the goal of minimizing the operating cost of the sub-microgrid, and the balance of power supply and demand is used as an indicator to measure the overall dispatching decision, which is fed back to the upper model;
[0094] The solution algorithm improvement module is used to use the low-dimensional latent variables of the variational autoencoder to replace the original state variables, improve the policy network and value network of the SAC algorithm, and introduce the time decay factor in the priority experience playback mechanism to dynamically adjust the priority of the experience and balance the sampling frequency of the experience entering the experience pool at different times to obtain the improved SAC algorithm and solve the upper model; the ADMM algorithm is improved based on the residual balance method, and the penalty parameters are adaptively adjusted by comparing the relative original residual and the relative dual residual to obtain the improved ADMM algorithm and solve the lower model;
[0095] The data output module is used to input the collected historical operating data of the power generation demand and energy storage status of the microgrid group into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after adjustment of the upper model and the optimized scheduling results of the lower model; the optimized scheduling results after adjustment of the upper model and the optimized scheduling results of the lower model are integrated to output the optimal two-layer optimized scheduling strategy for the microgrid group.
[0096] Compared with the prior art, the present invention has the following beneficial effects:
[0097] The present invention proposes a two-layer collaborative scheduling method for a microgrid group, which constructs a two-layer optimization scheduling model for a microgrid group. The SAC algorithm is improved by using VAE, and low-dimensional potential variables are used to replace the original state variables. The policy network and value network of the SAC algorithm are improved, which can significantly improve the utilization efficiency of samples and the convergence speed of training, thereby showing better results in complex tasks. In addition, a time decay factor is introduced into the priority experience playback of the SAC algorithm to reduce the priority of experience entering the experience pool early and increase the chance of experience entering the experience pool later. The priority of experience can be dynamically adjusted, and the sampling frequency of experience entering the experience pool at different times is effectively balanced, thereby improving the sample adoption efficiency and the learning efficiency of the algorithm, and solving the upper model. The ADMM algorithm is improved based on the residual balance method, and the penalty parameter is adaptively adjusted. By balancing the original residual and the dual residual, and by adaptively adjusting the penalty parameter, the improved ADMM algorithm converges faster and more stably, and can effectively handle the optimization operation and resource allocation problems within the sub-microgrid. At the same time, no matter how the scale and numerical range of the problem change, this strategy can maintain the effectiveness and convergence of the algorithm, simplify the process of parameter adjustment, and improve the adaptability and robustness of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 It is a flow chart of the double-layer coordinated scheduling method of microgrid group proposed by the present invention;
[0099] Figure 2 It is the optimized dispatching result of a microgrid group using the microgrid group double-layer collaborative dispatching method of the present invention. DETAILED DESCRIPTION
[0100] The following will be combined with the attached embodiment of the present invention Figure 1-Figure 2 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms described in the present invention are only used to describe specific implementation methods and are not used to limit the present invention.
[0101] like Figure 1 As shown, a microgrid dual-layer collaborative scheduling method proposed in the present invention comprises the following steps:
[0102] The present invention collects historical operating data of the power generation demand and energy storage status of the microgrid group;
[0103] Taking into account the operation characteristics and control requirements of the microgrid group, a two-layer optimization scheduling model of the microgrid group is constructed, including: an upper model and a lower model.
[0104] The upper model builds a model with the goal of minimizing the operating cost, determines the optimal energy scheduling solution and passes it to the lower model. The lower model builds a model with the goal of minimizing the operating cost of the sub-microgrid, determines the specific operating strategy of each sub-microgrid, and at the same time, executes the current strategy to calculate the power supply and demand balance of each sub-microgrid, and feeds it back to the upper model to help the upper model optimize the scheduling strategy.
[0105] The low-dimensional latent variables of the variational autoencoder (VAE) are used to replace the original state variables, the policy network and value network of the SAC algorithm are improved, and the time decay factor is introduced into the priority experience playback mechanism of the SAC algorithm to dynamically adjust the priority of the experience and balance the sampling frequency of the experience entering the experience pool at different times. The improved SAC algorithm is used to solve the upper-level model to realize the overall energy scheduling decision of the microgrid group and improve the system efficiency and generalization ability.
[0106] The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to solve the lower model. The improved ADMM algorithm can effectively handle the optimization operation and resource allocation problems within the sub-microgrid.
[0107] According to the collected historical operation data of the power generation demand and energy storage status of the microgrid group, a batch of sample data is selected from the historical operation data, and the improved SAC algorithm and the improved ADMM algorithm are input to obtain the optimized scheduling results of the upper model and the lower model.
[0108] The optimized scheduling results of the upper-level model are input into the lower-level model to obtain the optimized scheduling results after adjustment of the upper-level model; the optimized scheduling results after adjustment of the upper-level model and the optimized scheduling results of the lower-level model are optimized and integrated to obtain the optimal collaborative scheduling optimization strategy.
[0109] Specifically, the objective function of the upper model is:
[0110]
[0111] In the formula, Z MMG is the total operating cost of the microgrid; C g is the operating cost of all power generation units in the microgrid; C ess is the operating cost of all energy storage systems in the microgrid; C dn The transaction cost between the microgrid and the upper grid; For microgrid group and sub-microgrid MG n Transaction costs; N T 、N g 、N ess and N MG are the time period, the number of power generation units, the number of energy storage devices and the number of sub-microgrids respectively; ρ fuel and ρ maint are respectively the fuel cost and maintenance cost per unit of electricity generated; and Respectively, unit g i startup and shutdown costs; For the unit g i Output during period t; and They are start and stop states respectively. Indicates startup; Indicates shutdown; and They are charging cost, discharging cost and maintenance cost per unit of electricity; and They are energy storage systems j The charging and discharging power; and are the purchase price and sales price of electricity when the microgrid group trades with the upper power grid during period t; and They are respectively the purchased power and sold power when the microgrid group trades with the upper power grid during period t; and They are microgrid group and sub-microgrid MG respectively. n The purchase and sale prices of electricity during the trading period t; and They are microgrid group and sub-microgrid MG respectively. nThe purchase and sale of electricity during the t period of trading.
[0112] The constraints of the objective function of the upper model are as follows:
[0113] Microgrid power balance constraints:
[0114]
[0115] Energy storage constraints:
[0116]
[0117] Interactive power constraints:
[0118]
[0119] Power supply and demand balance constraints:
[0120]
[0121] In the formula, is the load demand of the microgrid group in period t; and are the maximum charging and discharging powers of energy storage system j, respectively; is the energy state of energy storage system j in period t; η essj,c and η essj,d are the charging and discharging efficiencies of energy storage system j, respectively; is the maximum energy state of energy storage system j; and are the minimum and maximum values of the state of charge, respectively; and are the minimum and maximum values of the interaction power between the microgrid and the upper grid, respectively; and They are microgrid group and sub-microgrid MG respectively. n The minimum and maximum values of the interaction power; N represents the number of microgrids, and are the actual power generation and storage capacity of microgrid i respectively; P exchange,i is the total interactive power of microgrid i; L i is the load demand of microgrid i; σ is the power balance error.
[0122] The objective function of the lower model is:
[0123]
[0124] In the formula, MG n The operating cost of ldis the operating cost of all loads; C MMG is the transaction cost between the sub-microgrid and the microgrid group; is the demand cost coefficient of load k; is the required power of load k; and The distribution is the purchase price and sale price of electricity when the sub-microgrid and the microgrid group trade in the t period; and They are respectively the purchased power and sold power of the sub-microgrid and the microgrid group during transactions in period t.
[0125] The constraints of the objective function of the lower model include:
[0126] Sub-microgrid power balance constraints:
[0127]
[0128] Sub-microgrid switching power constraints:
[0129]
[0130] Unit output climbing constraints:
[0131]
[0132] In the formula, is the interaction power between the sub-microgrid and the microgrid group; is the total load demand in period t; is the interaction power between sub-microgrids; and are the minimum and maximum values of the interaction power between the sub-microgrid and the microgrid group, respectively; and are the minimum and maximum values of the interaction power between sub-microgrids, respectively; and are the minimum and maximum output power of the fuel cell, respectively; and are the maximum values of the climbing output of the fuel cell and the diesel generator respectively; and are the minimum and maximum output power of the diesel generator respectively.
[0133] The SAC algorithm is improved based on variational autoencoder VAE. At the same time, the time decay factor is introduced to improve the priority experience replay mechanism of the SAC algorithm to solve the upper-level model.
[0134] The SAC algorithm is an advanced reinforcement learning algorithm that aims to improve the exploratory nature of the policy by maximizing the entropy of the policy. However, the SAC algorithm still has shortcomings such as low sample utilization when dealing with high-dimensional states.
[0135] Therefore, the present invention improves the SAC algorithm based on the variational autoencoder VAE, improves the utilization efficiency of samples, and further improves the convergence speed of the algorithm.
[0136] (1) VAE
[0137] VAE is a generative model that captures the potential structure of data by learning the potential representation of the state. Collect variables from the environment, including the current state, action, reward, and next state; design the network structure of VAE, including encoder and decoder; the encoder compresses high-dimensional data into a low-dimensional latent space, and the decoder reconstructs the data in the low-dimensional latent space into high-dimensional data; define its loss function, which optimizes the parameters of VAE by minimizing the reconstruction error and KL divergence, so that the compressed low-dimensional latent space can retain as much information as possible from the original state. The specific expression of VAE is:
[0138]
[0139] In the formula, s is the original data; is the reconstructed data obtained by the decoder; β is the parameter that controls the trade-off between reconstruction error and KL divergence; D KL is the KL divergence between the approximate posterior distribution q(z|s) and the prior distribution p(z) of the latent variable z.
[0140] The VAE is trained using the collected historical running data, and the optimal encoder and decoder parameters are obtained by minimizing the loss function. The trained VAE can retain the main features in the high-dimensional data and provide effective low-dimensional representation for subsequent tasks.
[0141] The original state variable s obtained from the environment is input into the trained VAE model to obtain a low-dimensional latent variable z. The low-dimensional latent variable z is used in the SAC algorithm to replace the original state variable s for strategy optimization and value evaluation.
[0142] The expression of the policy network function of the improved SAC algorithm is:
[0143]
[0144] In the formula, q(z t |s t ) is a given state s t The latent variable z under t The probability distribution of π θ (a t |z t ) is the policy network in a given state z t Output action a tThe probability of; α is the temperature parameter, which controls the weight of the entropy term in the objective function; Q φ (z t ,a t ) is the output value of the value network; is the latent variable z generated by the conditional probability distribution t The mathematical expectation of To select action a according to the current strategy t The mathematical expectation of .
[0145] The expression of the value network function of the improved SAC algorithm is:
[0146]
[0147]
[0148] Where D is the experience replay pool; a t+1 and z t+1 are the state and action of the agent at time t+1; r(z t ,a t ) is the agent in a given state z t Output action a t The reward after Q φ (z t ,a t ) is the output value of the target value function.
[0149] By combining VAE and SAC algorithms and using low-dimensional latent variables to improve the policy network and value network of the SAC algorithm, the utilization efficiency of samples and the convergence speed of training can be significantly improved, thereby achieving better results in complex tasks.
[0150] (2) Improve the priority experience replay mechanism of the SAC algorithm
[0151] The current preferential experience playback measures mainly focus on obtaining high-value experience in the experience pool, but do not consider the time order in which the experience enters the experience pool, which may lead to an imbalance in the frequency of adoption. Specifically, the experience that enters the experience pool earlier is more likely to be sampled frequently, while the experience that enters later may be ignored.
[0152] However, the experience that enters later may be more valuable than the experience that enters the experience pool earlier and is frequently sampled. Therefore, the present invention introduces a time decay factor based on the priority experience playback, which reduces the priority of the experience that enters the experience pool earlier and increases the chance of the experience that enters the experience pool later.
[0153] Set the time decay factor to β, the value range is (0,1), and the empirical priority formula after introducing the time decay factor is:
[0154] δ i = r + γQ (z t+1 ,a t+1 )-Q(z t ,a t )
[0155]
[0156] In the formula, δ i is the TD error value of state transition i; ε is a constant to prevent the priority from being zero; t i is the experience τ i The number of time steps into the experience pool; r is the action a taken by the agent at the current time step t The reward obtained later; γ is the discount factor, which measures the importance of future rewards and current rewards, and its value range is [0,1]; Q(z t+1 ,a t+1 ) is the next state z t+1 and action a t+1 Q value; Q(z t ,a t ) is the current state z t and action a t Q value; p(τ i ) is the empirical τ i The adjusted priority.
[0157] According to the adjusted priority, calculate the probability of each experience being sampled:
[0158]
[0159] Where N is the total amount of experience in the experience pool, is the total priority of all experiences in the experience pool, P(τ i ) is the empirical τ i Probability of being adopted.
[0160] By introducing the time decay factor, the priority of experience can be adjusted dynamically, effectively balancing the sampling frequency of experiences entering the experience pool at different times, thereby improving the sample adoption efficiency and the learning efficiency of the algorithm.
[0161] The SAC algorithm is improved based on VAE, and the time decay factor is introduced to improve the priority experience playback mechanism to achieve the optimization solution of the upper model. The specific steps include:
[0162] Step 1: According to the upper model, define the Markov decision process MDP. Generally, the MDP is defined as a five-tuple<S,A,P,r,γ> , where S is the set of state spaces, A is the set of action spaces, P is the state transition probability, r is the reward function, and γ is the discount factor, where:
[0163] The state transition probabilities are usually learned by the agent itself in its interactions with the environment, while the discount factor weighs the short-term and long-term effects of the current action.
[0164] State space, state space is usually a variable that affects the operating state of the microgrid group. The present invention defines the state variables of the upper model as the amount of electricity purchased, the power generation power of each sub-microgrid, the energy storage state of each sub-microgrid, the load demand of each sub-microgrid, and the balance of power supply and demand of each sub-microgrid. The state space vector is expressed as:
[0165]
[0166] Where P buy The amount of electricity purchased indicates the amount of electricity purchased by the microgrid group; is the power generation, which represents the amount of electricity generated by all power generation equipment in the microgrid i; is the energy storage state, indicating the amount of electric energy stored in the energy storage system of sub-microgrid i; is the load demand, which represents the power demand of sub-microgrid i; is the balance degree of power supply and demand, which indicates the power supply and demand status of sub-microgrid i.
[0167] Action space, the present invention defines the action variables of the upper model as the adjustment of the power purchase of each sub-microgrid, the adjustment of the power generation, the adjustment of the energy storage and the adjustment of the power exchange between sub-microgrids. The action space vector is expressed as:
[0168]
[0169] In the formula, The power purchase adjustment of microgrid i; is the power generation adjustment of microgrid i; is the energy storage adjustment of sub-microgrid i; is the power exchange amount adjustment of microgrid i.
[0170] The optimization goal of the upper model established in the present invention is to minimize the operating cost. Based on this optimization goal, the designed reward function is:
[0171]
[0172] ξ(k)=ξ max (1-e -μk )
[0173] In the formula, c buy 、c gen and c storage are the cost coefficients of unit electricity purchase cost, unit power generation and unit energy storage respectively; n is the number of sub-microgrids included in the microgrid group; is the amount of electricity purchased by the microgrid group at the current time t; is the power generation of microgrid i; is the energy storage capacity of microgrid i; is the power supply and demand balance of microgrid i; ξ(k) is the penalty coefficient, where ξ max is the maximum penalty coefficient, μ is the adjustment speed parameter, which controls the speed of the penalty coefficient growth, k is the number of iterations, and the initial value is 0.
[0174] Using the improved SAC algorithm, the basic process of solving the upper model is as follows:
[0175] Step 1: Define the Markov decision process;
[0176] Step 2: Initialize the parameters of the policy network, value network, and target value network, and set the experience replay pool for the agent;
[0177] Step 3: According to the current state s t and the strategy selects action a t And execute, get the new state s t+1 and reward r t , transfer samples (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool;
[0178] Step 4: Randomly extract the collected historical running data from the experience replay pool as sample data, calculate the target Q value and strategy loss, and update the value network parameters and strategy network parameters by minimizing the loss function;
[0179] Step 5: Repeat the process of data collection and strategy optimization until the strategy network converges and outputs the optimal solution of the upper model.
[0180] The present invention also improves the ADMM algorithm based on the residual balance method to obtain an improved ADMM algorithm for solving the lower layer model.
[0181] The ADMM algorithm is a commonly used optimization algorithm, but its convergence speed is highly dependent on the choice of penalty parameters. Inappropriate penalty parameters may lead to slow convergence or even no convergence.
[0182] Therefore, the present invention proposes an adaptive penalty parameter strategy based on residual balance, which adaptively adjusts the penalty parameter by balancing the original residual and the dual residual, thereby improving the convergence speed and performance of the ADMM algorithm.
[0183] Specifically, the present invention realizes adaptive adjustment of the penalty parameter by comparing the relative original residual and the relative dual residual. The use of relative residual can improve the robustness of the algorithm and automatically adjust the penalty parameter when the formula of the optimization problem changes. The calculation formula of the relative original residual and the relative dual residual includes:
[0184] r (k) =Ax (k) +Bz (k) -c
[0185]
[0186] Where x∈R n and z∈R m are all variables in the optimization problem; c∈R p is a constant vector of linear constraints, representing the offset of the constraints; A∈R p×n and B∈R p×m are coefficient matrices of linear constraints; x and y are the original and dual variables respectively; r (k) is the original residual, is the relative original residual; s (k) is the dual residual, is the relative dual residual; is the second norm relative to the original residual, is the bi-norm of the relative dual residual; τ and γ are constants, is the penalty parameter.
[0187] The penalty parameter is adaptively adjusted by the calculation formula of relative original residual and relative dual residual This makes the improved ADMM algorithm converge faster and more stably. At the same time, no matter how the scale and numerical range of the problem change, this strategy can maintain the effectiveness and convergence of the algorithm, simplify the process of parameter adjustment, and improve the adaptability and robustness of the algorithm.
[0188] By optimizing and integrating the optimized dispatching results adjusted by the upper model and the optimized dispatching results of the lower model, the final optimal microgrid group coordinated dispatching strategy is obtained, which includes the following steps:
[0189] Step 1: Collect historical operation data of the microgrid group, including power demand and energy storage status, and clean and standardize the collected historical operation data of the microgrid group;
[0190] Step 2: Initialize the environment, establish a two-layer collaborative scheduling model for microgrid clusters, and define the Markov decision process;
[0191] Step: 3: Initialize the parameters of the improved SAC algorithm and ADMM algorithm, and set the experience replay pool to store the experience samples collected from the environment;
[0192] Step 4: Based on the current state, the agent selects and executes an action based on the output of the policy network, and stores the current state, action, reward, and new state in the experience replay pool;
[0193] Step 5: Randomly select a batch of historical operation data of the microgrid group collected from the experience replay pool as sample data, input the value network and the policy network, calculate the target Q value and loss function, and update the network parameters;
[0194] Step 6: Determine whether the loss function converges. If so, output the obtained power exchange amount; otherwise, return to step 4 to execute the strategy optimization process;
[0195] Step 7: According to the calculation formula of relative original residual and relative dual residual, iteratively update the original residual and dual residual; when the convergence condition is met, the optimal power generation power and energy storage state of each sub-microgrid are obtained;
[0196] Step 8: The power exchange amount solved by the upper model is passed to the lower model, and the power supply and demand balance of each sub-microgrid is calculated and fed back to the upper model;
[0197] Step 9: If the balance between power supply and demand meets the constraints, the optimal scheduling result of the lower model is output; if the constraints are not met, the number of iterations k=k+1 is adjusted, the state space is updated, the reward function is adjusted, and the process of policy optimization is re-executed until the balance between power supply and demand meets the constraints, and the optimal scheduling result after the adjustment of the upper model is obtained;
[0198] Step 10: Combine the optimized dispatching results after adjustment of the upper-layer model and the optimal dispatching results of the lower-layer model to output the optimal microgrid group two-layer optimized dispatching strategy.
[0199] The present invention proposes a microgrid group two-layer collaborative dispatching system, comprising:
[0200] A data collection module is used to collect historical operating data of the power generation demand and energy storage status of the microgrid group;
[0201] A microgrid group double-layer optimization dispatching model construction module is used to construct a microgrid group double-layer optimization dispatching model; the microgrid group double-layer optimization dispatching model includes an upper model and a lower model; the upper model is established with the goal of minimizing the operating cost; the lower model is established with the goal of minimizing the operating cost of the sub-microgrid, and the balance of power supply and demand is used as an indicator to measure the overall dispatching decision, which is fed back to the upper model;
[0202] The solution algorithm improvement module is used to use the low-dimensional latent variables of the variational autoencoder to replace the original state variables, improve the policy network and value network of the SAC algorithm, and introduce the time decay factor in the priority experience playback mechanism to dynamically adjust the priority of the experience and balance the sampling frequency of the experience entering the experience pool at different times to obtain the improved SAC algorithm and solve the upper model; the ADMM algorithm is improved based on the residual balance method, and the penalty parameters are adaptively adjusted by comparing the relative original residual and the relative dual residual to obtain the improved ADMM algorithm and solve the lower model;
[0203] The data output module is used to input the collected historical operating data of the power generation demand and energy storage status of the microgrid group into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after adjustment of the upper model and the optimized scheduling results of the lower model; the optimized scheduling results after adjustment of the upper model and the optimized scheduling results of the lower model are integrated to output the optimal two-layer optimized scheduling strategy for the microgrid group.
[0204] The microgrid group double-layer collaborative scheduling method proposed in the present invention constructs a microgrid group double-layer optimization scheduling model according to the operation characteristics and control requirements of the microgrid group. The upper model is established with the goal of minimizing the operation cost, and the lower model is established with the goal of minimizing the operation cost of the sub-microgrid. The balance of power supply and demand is fed back to the upper model as an indicator to measure the overall scheduling decision, and the decision of the upper model is reasonably adjusted. The SAC algorithm is improved based on VAE, and the time decay factor is introduced on the basis of the priority experience playback mechanism to improve the experience playback mechanism of the SAC algorithm. The upper model is solved by using the improved VAE-SAC algorithm to realize the overall energy scheduling decision of the microgrid group and improve the system efficiency and generalization ability. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to solve the lower model. The improved ADMM algorithm can effectively handle the optimization operation and resource allocation problems within the sub-microgrid. By inputting the optimization results of the upper-level model into the lower-level model, the upper- and lower-level scheduling results are optimized and integrated to obtain the final microgrid group collaborative scheduling plan.
[0205] Example
[0206] like Figure 1As shown, it is a flow chart of the two-layer collaborative scheduling method of the microgrid group proposed in the present invention. The method includes: constructing a two-layer optimization scheduling model of the microgrid group according to the operating characteristics and control requirements of the microgrid group, the upper model is established with the goal of minimizing the operating cost, and the lower model is established with the goal of minimizing the operating cost of the sub-microgrid, and the balance of power supply and demand is fed back to the upper model as an indicator to measure the overall scheduling decision, and the upper decision is reasonably adjusted. The SAC algorithm is improved based on VAE, and the time decay factor is introduced on the basis of the priority experience playback mechanism, and the experience playback mechanism of the SAC algorithm is improved. The upper model is solved by using the improved SAC algorithm to realize the overall energy scheduling decision of the microgrid group and improve the system efficiency and generalization ability. The ADMM algorithm is improved based on the residual balance method for solving the lower model. The improved ADMM algorithm can effectively handle the optimization operation and resource allocation problems within the sub-microgrid. The optimization results of the upper model are input into the lower model, and the upper and lower scheduling results are optimized and integrated to obtain the final collaborative scheduling solution.
[0207] like Figure 2 As shown in the figure, it is the optimization scheduling result of each sub-microgrid in the microgrid group of the method proposed in the present invention. By analyzing the data in the figure, it can be seen that Figure (a) shows that sub-microgrid 1 mainly behaves as the power demand side, while Figure (b) and Figure (c) respectively show that sub-microgrid 2 and sub-microgrid 3 have the dual roles of power supply and demand. During the period of low photovoltaic power generation and high load demand, each sub-microgrid can release electricity through the energy storage system; during the peak period of photovoltaic power generation, the energy storage system can be charged, thereby improving the utilization rate of energy. In addition, during the period of large photovoltaic power generation, sub-microgrid 2 and sub-microgrid 3 can sell excess electricity to sub-microgrid 1, and the reasonable distribution of electric energy between different sub-microgrids is achieved through power sales and power purchases, further improving the stability and reliability of the overall power grid. It can be obtained that the method proposed in the present invention can reasonably optimize the operation of the microgrid group and improve the operation efficiency and economic benefits of the microgrid group.
[0208] As shown in Table 1, for the comparison of operating indicators of different microgrid group collaborative scheduling methods, the present invention sets up 4 different schemes to verify the efficiency of the microgrid group double-layer collaborative scheduling method proposed in the present invention.
[0209] Scheme 1 is a microgrid group optimization scheduling method based on the PSO algorithm, Scheme 2 is a microgrid group optimization scheduling method based on the ADMM algorithm, Scheme 3 is a microgrid group optimization scheduling method based on the DQN algorithm, and Scheme 4 is a microgrid group optimization scheduling method of the method proposed in the present invention.
[0210] Table 1 Comparison of operating indicators of different microgrid coordinated dispatching methods
[0211]
[0212]
[0213] By analyzing the data in the table, it can be seen that the operating cost of Scheme 4, that is, the microgrid group optimization scheduling method of the method proposed in the present invention, is 28,016.36 yuan, which is reduced by 21.94%, 18.38% and 14.45% respectively compared with Schemes 1, 2 and 3; the energy utilization rate of Scheme 4 is 97.2%, which is increased by 13.85%, 9.72% and 5.43% respectively compared with Schemes 1, 2 and 3; the load balancing rate of Scheme 4 is 98.9%, which is increased by 4.55%, 2.70% and 1.44% respectively compared with Schemes 1, 2 and 3; the scheduling time of Scheme 4 is 13.7 seconds, which is reduced by 32.51%, 26.34% and 11.04% respectively compared with Schemes 1, 2 and 3.
[0214] It can be concluded that the method proposed in the present invention shows significant advantages in terms of operating cost, energy utilization, load balancing rate and scheduling time.
[0215] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
[0216] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.
Claims
1. A two-layer coordinated scheduling method for a microgrid group, characterized in that: The following steps are involved: Collect historical operating data of the microgrid’s power generation demand and energy storage status; Constructing a two-layer optimization dispatching model for a microgrid group; the two-layer optimization dispatching model for a microgrid group includes an upper model and a lower model; the upper model is established with the goal of minimizing the operating cost; the lower model is established with the goal of minimizing the operating cost of the sub-microgrid, and the balance of power supply and demand is used as an indicator to measure the overall dispatching decision, which is fed back to the upper model; The low-dimensional latent variables of the variational autoencoder are used to replace the original state variables, and the policy network and value network of the SAC algorithm are improved. The time decay factor is introduced into the priority experience playback mechanism to dynamically adjust the priority of the experience and balance the sampling frequency of the experience entering the experience pool at different times. The improved SAC algorithm is obtained to solve the upper model. The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain the improved ADMM algorithm to solve the underlying model. The collected historical operating data of the power generation demand and energy storage status of the microgrid group are input into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after the adjustment of the upper model and the optimized scheduling results of the lower model; The optimized dispatching results after adjustment of the upper-layer model are integrated with the optimized dispatching results of the lower-layer model to output the optimal two-layer optimized dispatching strategy for the microgrid group.
2. The microgrid dual-layer collaborative scheduling method according to claim 1 is characterized in that: The upper model is established with the goal of minimizing the operating cost, and includes the following steps: The objective function of the upper model is: In the formula, Z MMG is the total operating cost of the microgrid; C g is the operating cost of all power generation units in the microgrid; C ess is the operating cost of all energy storage systems in the microgrid; C dn The transaction cost between the microgrid and the upper grid; For microgrid group and sub-microgrid MG n Transaction costs; N T 、N g 、N ess and N MG are the time period, the number of power generation units, the number of energy storage devices and the number of sub-microgrids respectively; ρ fuel and ρ maint are respectively the fuel cost and maintenance cost per unit of electricity generated; and Respectively, unit g i startup and shutdown costs; For the unit g i Output during period t; and They are start and stop states respectively. Indicates startup; Indicates shutdown; and They are charging cost, discharging cost and maintenance cost per unit of electricity; and They are energy storage systems j The charging and discharging power; and are the purchase price and sales price of electricity when the microgrid group trades with the upper power grid during period t; and They are respectively the purchased power and sold power when the microgrid group trades with the upper power grid during period t; and They are microgrid group and sub-microgrid MG respectively. n The purchase and sale prices of electricity during the trading period t; and They are microgrid group and sub-microgrid MG respectively. n The purchase and sale of electricity during the t period of trading; The constraints of the objective function of the upper model are as follows: Microgrid power balance constraints: Energy storage constraints: Interactive power constraints: Power supply and demand balance constraints: In the formula, is the load demand of the microgrid group in period t; and are the maximum charging and discharging powers of energy storage system j, respectively; is the energy state of energy storage system j in period t; η essj,c and η essj,d are the charging and discharging efficiencies of energy storage system j, respectively; is the maximum energy state of energy storage system j; and are the minimum and maximum values of the state of charge, respectively; and are the minimum and maximum values of the interaction power between the microgrid and the upper grid, respectively; and They are microgrid group and sub-microgrid MG respectively. n The minimum and maximum values of the interaction power; N represents the number of microgrids, and are the actual power generation and storage capacity of microgrid i respectively; P exchange,i is the total interactive power of microgrid i; L i is the load demand of microgrid i; σ is the power balance error.
3. The microgrid group dual-layer collaborative scheduling method according to claim 2 is characterized in that: The lower layer model is established with the goal of minimizing the operating cost of the sub-microgrid, and includes the following steps: The objective function of the lower model is: In the formula, MG n The operating cost of ld is the operating cost of all loads; C MMG is the transaction cost between the sub-microgrid and the microgrid group; is the demand cost coefficient of load k; is the required power of load k; and The distribution is the purchase price and sale price of electricity when the sub-microgrid and the microgrid group trade in the t period; and They are the purchased power and sold power of the sub-microgrid and microgrid group during the transaction in period t, respectively; The constraints of the objective function of the lower model include: Sub-microgrid power balance constraints: Sub-microgrid switching power constraints: Unit output climbing constraints: In the formula, is the interaction power between the sub-microgrid and the microgrid group; is the total load demand in period t; is the interaction power between sub-microgrids; and are the minimum and maximum values of the interaction power between the sub-microgrid and the microgrid group, respectively; and are the minimum and maximum values of the interaction power between sub-microgrids, respectively; and are the minimum and maximum output power of the fuel cell, respectively; and are the maximum values of the climbing output of the fuel cell and the diesel generator respectively; and are the minimum and maximum output power of the diesel generator respectively.
4. The microgrid group dual-layer collaborative scheduling method according to claim 3 is characterized in that: The improvement of the strategy network and value network of the SAC algorithm specifically includes: The original state variable s obtained from the environment is input into the trained VAE model to obtain a low-dimensional latent variable z. The low-dimensional latent variable z is used in the SAC algorithm to replace the original state variable s for strategy optimization and value evaluation. The expression of the policy network function of the improved SAC algorithm is: In the formula, q(z t |s t ) is a given state s t The latent variable z under t The probability distribution of π θ (a t |z t ) is the policy network in a given state z t Output action a t The probability of; α is the temperature parameter, which controls the weight of the entropy term in the objective function; Q φ (z t ,a t ) is the output value of the value network; is the latent variable z generated by the conditional probability distribution t The mathematical expectation of To select action a according to the current strategy t The mathematical expectation of The expression of the value network function of the improved SAC algorithm is: Where D is the experience replay pool; a t+1 and z t+1 are the state and action of the agent at time t+1; r(z t ,a t ) is the agent in a given state z t Output action a t The reward after Q φ (z t ,a t ) is the output value of the target value function.
5. The microgrid dual-layer coordinated scheduling method according to claim 4 is characterized in that: The method of balancing the sampling frequency of experiences entering the experience pool at different times comprises the following steps: The empirical priority formula after introducing the time decay factor is: δ i =r+γQ(z t+1 ,a t+1 )-Q(z t ,a t ) Where β is the time attenuation factor, and its value range is (0,1); δ i is the TD error value of state transition i; ε is a constant to prevent the priority from being zero; t i is the experience τ i The number of time steps into the experience pool; r is the action a taken by the agent at the current time step t The reward obtained later; γ is the discount factor, which measures the importance of future rewards and current rewards, and its value range is [0,1]; Q(z t+1 ,a t+1 ) is the next state z t+1 and action a t+1 Q value; Q(z t ,a t ) is the current state z t and action a t Q value; p(τ i ) is the empirical τ i The adjusted priority of According to the adjusted priority, calculate the probability of each experience being sampled: Where N is the total amount of experience in the experience pool, is the total priority of all experiences in the experience pool, P(τ i ) is the empirical τ i Probability of being adopted.
6. The microgrid dual-layer coordinated scheduling method according to claim 5 is characterized in that: Solving the upper model includes the following steps: Step 1: Define the Markov decision process MDP according to the upper model; Define the MDP as a five-tuple<S,A,P,r,γ> , where S is the set of state spaces, A is the set of action spaces, P is the state transition probability, r is the reward function, and γ is the discount factor; the state transition probability is learned by the agent itself in the interaction with the environment, and the discount factor weighs the short-term and long-term effects of the current action; In the state space, the state variables of the upper model are defined as the amount of electricity purchased, the power generation of each sub-microgrid, the energy storage status of each sub-microgrid, the load demand of each sub-microgrid, and the balance of power supply and demand of each sub-microgrid. The state space vector is expressed as: Where P buy The amount of electricity purchased indicates the amount of electricity purchased by the microgrid group; is the power generation, which represents the amount of electricity generated by all power generation equipment in the microgrid i; is the energy storage state, indicating the amount of electric energy stored in the energy storage system of sub-microgrid i; is the load demand, which represents the power demand of sub-microgrid i; is the balance degree of power supply and demand, which indicates the power supply and demand status of sub-microgrid i; Action space, the action variables of the upper model are defined as the adjustment of power purchase, power generation, energy storage and power exchange between sub-microgrids. The action space vector is expressed as: In the formula, The power purchase adjustment of microgrid i; is the power generation adjustment of microgrid i; is the energy storage adjustment of sub-microgrid i; is the power exchange amount adjustment of sub-microgrid i; The designed reward function is: ξ(k)=ξ max (1-e -μk ) In the formula, c buy 、c gen and c storage are the cost coefficients of unit electricity purchase cost, unit power generation and unit energy storage respectively; n is the number of sub-microgrids included in the microgrid group; is the amount of electricity purchased by the microgrid group at the current time t; is the power generation power of microgrid i; is the energy storage capacity of microgrid i; is the power supply and demand balance of microgrid i; ξ(k) is the penalty coefficient, where ξ max is the maximum penalty coefficient, μ is the adjustment speed parameter, which controls the speed of the penalty coefficient growth, k is the number of iterations, and the initial value is 0; Step 1: Define the Markov decision process; Step 2: Initialize the parameters of the policy network, value network, and target value network, and set the experience replay pool for the agent; Step 3: According to the current state s t and the strategy selects action a t And execute, get the new state s t+1 and reward r t , transfer samples (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool; Step 4: Randomly extract the collected historical running data from the experience replay pool as sample data, calculate the target Q value and strategy loss, and update the value network parameters and strategy network parameters by minimizing the loss function; Step 5: Repeat the process of data collection and strategy optimization until the strategy network converges and outputs the optimized scheduling results of the upper-level model.
7. The microgrid group dual-layer coordinated scheduling method according to claim 6 is characterized in that: The improved ADMM algorithm comprises the following steps: By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted; Among them, the calculation formulas of relative primal residual and relative dual residual include: r (k) =Ax (k) +Bz (k) -c Where x∈R n and z∈R m are all variables in the optimization problem; c∈R p is a constant vector of linear constraints, representing the offset of the constraints; A∈R p×n and B∈R p×m are coefficient matrices of linear constraints; x and y are the original and dual variables respectively; r (k) is the original residual, is the relative original residual; s (k) is the dual residual, is the relative dual residual; is the second norm relative to the original residual, is the bi-norm of the relative dual residual; τ and γ are constants, is the penalty parameter; The penalty parameter is adaptively adjusted by the calculation formula of relative original residual and relative dual residual 8. The microgrid group dual-layer collaborative scheduling method according to claim 7 is characterized in that: The dual-layer optimization scheduling strategy for the microgrid group with the best output comprises the following steps: Step 1: Collect historical operation data of the microgrid group, including power demand and energy storage status, and clean and standardize the collected historical operation data of the microgrid group; Step 2: Initialize the environment, establish a two-layer collaborative scheduling model for microgrid clusters, and define the Markov decision process; Step: 3: Initialize the parameters of the improved SAC algorithm and the improved ADMM algorithm, and set the experience replay pool to store the experience samples collected from the environment; Step 4: Based on the current state, the agent selects and executes an action based on the output of the policy network, and stores the current state, action, reward, and new state in the experience replay pool; Step 5: Randomly select a batch of historical operation data of the microgrid group collected from the experience replay pool as sample data, input the value network and the policy network, calculate the target Q value and loss function, and update the network parameters; Step 6: Determine whether the loss function converges. If so, output the obtained power exchange amount; otherwise, return to step 4 to execute the strategy optimization process; Step 7: According to the calculation formula of relative original residual and relative dual residual, iteratively update the original residual and dual residual; when the convergence condition is met, the optimal power generation power and energy storage state of each sub-microgrid are obtained; Step 8: The power exchange amount solved by the upper model is passed to the lower model, and the power supply and demand balance of each sub-microgrid is calculated and fed back to the upper model; Step 9: If the balance between power supply and demand meets the constraints, the optimal scheduling result of the lower model is output; if the constraints are not met, the number of iterations k=k+1 is adjusted, the state space is updated, the reward function is adjusted, and the process of policy optimization is re-executed until the balance between power supply and demand meets the constraints, and the optimal scheduling result after the adjustment of the upper model is obtained; Step 10: Combine the optimized dispatching results after adjustment of the upper-layer model and the optimal dispatching results of the lower-layer model to output the optimal microgrid group two-layer optimized dispatching strategy.
9. A microgrid group dual-layer collaborative dispatching system, characterized in that: include: A data collection module is used to collect historical operating data of the power generation demand and energy storage status of the microgrid group; A microgrid group double-layer optimization dispatching model construction module is used to construct a microgrid group double-layer optimization dispatching model; the microgrid group double-layer optimization dispatching model includes an upper model and a lower model; the upper model is established with the goal of minimizing the operating cost; the lower model is established with the goal of minimizing the operating cost of the sub-microgrid, and the balance of power supply and demand is used as an indicator to measure the overall dispatching decision, which is fed back to the upper model; The solution algorithm improvement module is used to use the low-dimensional latent variables of the variational autoencoder to replace the original state variables, improve the policy network and value network of the SAC algorithm, introduce the time decay factor in the priority experience playback mechanism, dynamically adjust the priority of the experience, balance the sampling frequency of the experience entering the experience pool at different times, obtain the improved SAC algorithm, and solve the upper model; The ADMM algorithm is improved based on the residual balance method. By comparing the relative original residual and the relative dual residual, the penalty parameter is adaptively adjusted to obtain the improved ADMM algorithm to solve the underlying model. The data output module is used to input the collected historical operation data of the power generation demand and energy storage status of the microgrid group into the improved SAC algorithm and the improved ADMM algorithm to obtain the optimized scheduling results after the upper model is adjusted and the optimized scheduling results of the lower model; The optimized dispatching results after adjustment of the upper-layer model are integrated with the optimized dispatching results of the lower-layer model to output the optimal two-layer optimized dispatching strategy for the microgrid group.
Citation Information
Patent Citations
Microgrid group optimization scheduling strategy based on niche chaos particle swarm algorithm
CN112821470A
Power distribution network multi-target distributed optimization scheduling method based on ADMM algorithm
CN117458620A
Multi-microgrid collaborative optimization scheduling strategy considering energy complementation and power transmission loss
CN119029842A
Methods and systems for computation of bilevel mixed integer programming problems
US20160335223A1
Cited By
Micro-grid group distributed energy management method, system and equipment based on event driving and medium
CN120281020A
Micro-grid energy management method and system based on reinforcement learning
CN121689158A