Microgrid group low-carbon optimization operation method based on improved MAPPO algorithm

By improving the MAPPO algorithm, combining carbon governance costs and collaborative rewards, the negative sampling experience sharing framework and LSTM network are introduced, and the problem of collaborative operation of multiple micronets in the micronet group system is solved, and the low-carbon optimization operation of micronet groups is achieved, reducing operating costs and carbon emissions are reduced.

CN120073853AActive Publication Date: 2025-05-30CHINA THREE GORGES UNIV

Patent Information

Application Number
CN202510004129.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-30
Estimated Expiration
2045-01-02

AI Technical Summary

Technical Problem

In the prior art, it is difficult to achieve the coordinated operation of multiple micronets in the micronet group system, resulting in difficulty in effectively reducing system operation costs and carbon emissions.

Method used

The low-carbon optimization operation method of micronet groups based on the improved MAPPO algorithm is adopted. By building a micronet group optimization operation model, nonlinear constraints on carbon governance costs and collaborative rewards between micronets are introduced, and combined with the negative sampling experience sharing framework and LSTM network, the decision-making ability and training efficiency of the agent are improved.

Benefits of technology

Effectively reduce the operating costs of each micronet, while reducing system carbon emissions, improving energy utilization efficiency, and enhancing the collaborative scheduling capabilities of micronet groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120073853A_ABST
    Figure CN120073853A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-grid group low-carbon optimization operation method based on an improved MAPPO algorithm, and the method comprises the steps: building a micro-grid group optimization operation model with the minimum total cost in a single micro-grid operation period as an objective function; constructing a partially observable Markov decision process of micro-grid optimization operation, and establishing an observation space, an action space and a reward function corresponding to each micro-grid agent; non-linear constraints of carbon treatment cost and collaborative rewards among micro-grids are introduced into a reward function, and a negative sampling experience sharing framework is introduced into an MAPPO algorithm; the LSTM network is embedded into the MAPPO algorithm, and a differential learning rate attenuation strategy is designed to further improve the training speed of the LSTM-MAPPO model; and training the intelligent agent based on an improved MAPPO algorithm to obtain an optimal microgrid group low-carbon optimization operation scheme. The optimization method provided by the invention can effectively reduce the operation cost of each micro-grid and reduce the carbon emission of the system at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of optimal operation of microgrid clusters, and particularly relates to a low-carbon optimal operation method for microgrid clusters based on an improved MAPPO algorithm. Background Art

[0002] With the rapid growth of renewable energy and distributed energy, the power system is facing increasingly complex management challenges. As an intelligent system integrating multiple energy resources, microgrid clusters can achieve efficient operation of the power system through optimal scheduling. However, due to the complex structure of microgrid clusters, how to achieve the coordinated operation of multiple microgrids in the microgrid cluster system and reduce the system operation cost and carbon emissions has become a current hot issue.

[0003] In the prior art: Document [1]: "Energy Optimization Method for Microgrid Clusters Based on Improved Bat Algorithm" (Zeng Zhihui, Li Xueqiang, Yin Lulu, etc. Energy Optimization Method for Microgrid Clusters Based on Improved Bat Algorithm [J]. Electronic Measurement Technology, 2023, 46(10): 53-60.) studied the energy optimization of microgrid clusters based on an improved bat algorithm for issues such as how to stably operate microgrid clusters and minimize the operation cost of the grid cluster to the greatest extent, improving the economic benefits of the system. Document [2]: "Multi-energy Microgrid Multi-objective Optimal Scheduling Based on Improved Particle Swarm Algorithm" (Wang Yu, Hao Yi, Wang Lei, etc. Multi-energy Microgrid Multi-objective Optimal Scheduling Based on Improved Particle Swarm Algorithm [J]. Electrical Measurement & Instrumentation, 2023, 60(11): 29-36+59.) proposed a multi-energy microgrid multi-objective optimal scheduling method based on an improved particle swarm method to achieve flexible utilization of internal energy in the multi-energy microgrid while reducing the carbon emission pressure on the system.

[0004] However, the optimization methods adopted in the above documents belong to heuristic algorithms, which are prone to falling into local optimal solutions and are difficult to adapt to the complex dynamic environment changes of multi-microgrids.

[0005] Document [3]: "Microgrid Energy Trading Based on Multi-agent Reinforcement Learning" (Wei Guixi, Liu Xianggang, Chi Ming, etc. Microgrid Energy Trading Based on Multi-agent Reinforcement Learning [J]. Control Engineering, 2023, 30(12): 2274-2279+2296.) proposed an energy trading method based on multi-agent reinforcement learning to promote energy trading among microgrid users, reducing the peak load of the microgrid and the energy consumption cost of users. However, this document uses the traditional MAPPO algorithm to optimize the energy trading between microgrids, resulting in difficulty in fully capturing time-dependent relationships when dealing with time series tasks, and the data utilization efficiency needs to be improved. Summary of the Invention

[0006] To address the above deficiencies, the present invention discloses a low-carbon optimal operation method for a microgrid group based on an improved MAPPO algorithm. This optimization method can effectively reduce the operation cost of each microgrid while reducing the system carbon emissions.

[0007] The technical solution adopted by the present invention is as follows:

[0008] A low-carbon optimal operation method for a microgrid group based on an improved MAPPO algorithm, comprising the following steps:

[0009] Step 1: Taking the minimum total cost within the operation cycle of a single microgrid as the objective function, construct an optimal operation model for the microgrid group;

[0010] Step 2: Construct a partially observable Markov decision process for the optimal operation of the microgrid, and establish the observation space, action space, and reward function corresponding to each microgrid agent;

[0011] Step 3: Introduce the non-linear constraint of carbon governance cost and the collaborative reward among microgrids into the reward function to reduce the carbon emissions of the microgrid group and promote the electricity trading among microgrids;

[0012] Step 4: To solve the problems of mutual dependence and privacy protection among agents, and at the same time improve the training efficiency and sample utilization rate, introduce a negative sampling experience sharing framework into the MAPPO algorithm;

[0013] Step 5: Embed the LSTM network into the MAPPO algorithm to help the model better capture the temporal dependence to enhance the decision-making ability of the agent, and design a differential learning rate decay strategy to further improve the training speed of the LSTM-MAPPO model;

[0014] Step 6: Train the agent based on the improved MAPPO algorithm to obtain the optimal low-carbon optimal operation plan for the microgrid group. In the above Step 1, constructing the optimal operation model for the microgrid group includes:

[0015] (1) Objective function:

[0016] In addition to the power production costs of various distributed energy sources, the operation costs of energy storage systems, and the power interaction costs of electric energy, the present invention also introduces the carbon governance cost into the objective function of each microgrid to measure the carbon emission impact in the microgrid operation, and its objective function is:

[0017] C i,MG =minC i,DG +C i,ESS +C i,cbn +C i,grid +C i,j ;

[0018] In the formula, C i,MG is the total cost of the microgrid; C i,DGis the power generation cost of each distributed power source in microgrid i; C i,ESS is the operating cost of the energy storage system; C i,cbn is the carbon governance cost of microgrid i; C i,grid is the cost of the interactive power between microgrid i and the superior power grid; C i,j is the cost of the interactive power between microgrid i and microgrid j, and i ≠ j.

[0019] 1) Distributed power generation cost:

[0020] Each distributed power source in the microgrid includes wind power, photovoltaic and gas turbines, and its cost is shown in the following formula:

[0021] C i,DG = C i,WT + C i,PV + C i,MT ;

[0022] C i,WT = K WT P i,WT ;

[0023] C i,PV = K PV P i,PV ;

[0024] C i,MT = a i P i,MT ;

[0025] In the formula, C i,WT , C i,PV and C i,MT are the operating costs of the wind turbine, photovoltaic unit and gas turbine in microgrid i respectively; K WT , K PV are the use and maintenance cost coefficients of the fan and the photovoltaic system respectively; a i is the gas turbine cost coefficient of microgrid i; P i,WT , P i,PV and P i,MT are the gas turbine output, fan output and photovoltaic system output of microgrid i respectively.

[0026] 2) Energy storage device operating cost:

[0027] The energy storage device can store energy when the power supply is excessive and release energy during peak demand, helping the microgrid achieve supply-demand balance. Its operating cost is shown in the following formula:

[0028] C i,ESS = K i,ESS (η i,on P i,on + η i,off Pi,off );

[0029] Wherein, K i,ESS is the cost coefficient of the energy storage device of microgrid i; η i,on , η i,off are the charge and discharge efficiencies of the energy storage device respectively; P n,on , P n,off are the charge and discharge powers of the energy storage device respectively.

[0030] 3) Carbon governance cost:

[0031] By introducing the carbon governance cost, it is possible to encourage the microgrid cluster to optimize the energy allocation, improve the energy utilization efficiency, and promote the consumption of new energy. The carbon governance cost directly reflects the carbon emission situation of the system, and its calculation formula is as follows:

[0032] C i,cbn = Q i,D η t T - Q i,MG η T P;

[0033] Q i,D = Q i,WT + Q i,PV + Q i,MT + Q i,ESS ;

[0034] Q i,MG = Q i,D + Q i,B ;

[0035]

[0036] P i,DSG = P i,WT + P i,PV + P i,MT ;

[0037] Wherein, Q i,D is the carbon dioxide emission amount generated by microgrid i; η t is the carbon tax levy rate; T is the unit carbon emission tax; Q i,MG is the sum of the carbon dioxide emission amount generated by microgrid i and the carbon dioxide emission amount caused by purchasing electricity from the superior power grid; η T is the carbon dioxide emission reduction ratio; P is the trading price; Q i,WT , Q i,PV , Q i,MT and Q i,ESS are the carbon dioxide emission amounts caused by wind power generation, photovoltaic power generation, gas turbine power generation and energy storage device power generation respectively; Q i,B is the carbon dioxide emission amount caused by microgrid i purchasing electricity from the superior power grid; L i,Bis the total load demand within microgrid i; P i,ESS is the power generation output of the energy storage device; sgn is the sign function; P i,DSG is the cumulative power generation of all distributed generation devices without energy storage; D i,B is the total power purchase from the superior power grid; P i,WT 、P i,PV and P i,MT are the outputs of the wind turbine, photovoltaic unit, and micro gas turbine, respectively.

[0038] 4) Transaction costs between the microgrid and the superior power grid and between microgrids:

[0039] There may be power transactions between the microgrid and the superior power grid and between microgrids during the dispatching process. The transaction costs are as follows:

[0040] C i,grid = q gridbuy P i,gridbuy - q gridsell P i,gridsell ;

[0041]

[0042] In the formula, q gridbuy 、q gridsell are the power purchase and power sale prices of the microgrid from / to the superior power grid, respectively; P i,gridbuy 、P i,gridsell are the power purchase and power sale quantities of microgrid i from / to the superior power grid, respectively; q MG is the transaction electricity price between microgrids; P i,mbuy 、P i,msell are the power purchase and sale quantities of microgrid i from / to microgrid j, respectively.

[0043] (2) Constraint conditions:

[0044] To ensure the effectiveness and practicality of the microgrid group optimal dispatching model, each microgrid must comply with the following power constraint conditions to ensure that the microgrid system is technically feasible and economically efficient;

[0045] 1) Microgrid power balance constraint:

[0046]

[0047] 2) Upper and lower limits of power interaction between microgrids:

[0048]

[0049] In the formula, P i,j 、 are the transaction power and the maximum transaction power between microgrids, respectively.

[0050] 3) Output upper and lower limit constraints of the micro gas turbine:

[0051]

[0052] In the formula, and are the upper and lower limits of the output of the gas turbine respectively.

[0053] 4) Ramp rate constraints of the micro gas turbine:

[0054]

[0055] In the formula, are the output power of the gas turbine at time t and time t - 1 respectively; are the upper and lower limits of the change in the output of the unit per unit time.

[0056] 5) Energy storage device constraints:

[0057]

[0058] In the formula, are the state of charge of the energy storage at time t and time t - 1 respectively; u is the charge and discharge coefficient; are the upper and lower limits of the state of charge of the energy storage at time t; are the charge and discharge power of the energy storage at time t respectively; are the upper limits of the charge and discharge power of the energy storage at time t respectively.

[0059] 6) Interactive power constraints between the microgrid and the upstream power grid:

[0060]

[0061] In the formula, is the trading power and the trading power limit between the microgrid i and the upstream power grid.

[0062] 7) Trading electricity price constraints:

[0063] To promote each sub - microgrid in the microgrid cluster system to give priority to internal trading in the system, it should be ensured that the electricity price trading between microgrids is between the selling electricity price and the repurchase electricity price of the upstream network. Thus, the trading electricity price constraint can be obtained as:

[0064] q gridsell ≤q MG ≤q gridbuy ;

[0065] 8) Carbon emission constraints:

[0066]

[0067] In the formula, is the maximum allowable carbon emission of device m in microgrid i during period t; is the carbon emission of device m in microgrid i during period t; is the total carbon emission of device m within a scheduling period; is the maximum allowable total carbon emission of device m within a scheduling period; is the total carbon emission of microgrid i within a scheduling period; is the maximum allowable total carbon emission of microgrid i during period t; is the maximum allowable total carbon emission of microgrid i within a scheduling period. In step 2, the partially observable Markov decision process for the optimal operation of the microgrid is as follows:

[0068] Constructing the partially observable Markov decision process for the optimal operation of the power grid is actually a process of clarifying what each microgrid agent can observe (observation space), the actions it can take (action space), and the rewards obtained according to the behavior (reward function) for each microgrid agent. That is, a process of establishing the observation space, action space, and reward function corresponding to each microgrid agent. Specifically, it includes the following:

[0069] (1) Selection of the observation space O:

[0070] The observations of microgrid i include the load data at time t photovoltaic power generation wind power generation the purchase and sale electricity price q of the superior power grid grid and the electricity trading price q between microgrids MG as well as the state of charge of the energy storage For the optimal dispatching of the microgrid group, its observations can be expressed as:

[0071]

[0072] In the formula, represents the observation of microgrid agent i at time t;

[0073] (2) Selection of the action space A:

[0074] The goal of the economic and low-carbon optimal dispatching of the microgrid group is to determine the optimal output of each unit and the electricity trading situation. Therefore, the output power of the micro gas turbine at time t the charge and discharge power of the energy storage the trading power between the microgrid and the superior network and the trading power between microgrids are used as the action values:

[0075]

[0076] In the formula, a it Denote the action of microgrid agent \(i\) at time \(t\);

[0077] (3) Reward function:

[0078] For the optimal scheduling problem of the microgrid group, the designed reward function should reflect the interests of each agent while considering the constraints of system operation. Therefore, in addition to the cost function considered, the operation constraint formula of the energy storage device, the interactive power constraint formula with the superior power grid, and the carbon emission constraint formula need to be added to the reward function as penalty functions, and the penalty terms are respectively expressed as Then the reward function of each microgrid can be expressed as:

[0079]

[0080] In the formula, \(r\) i MG represents the reward function of each microgrid agent; \(\eta\) i,ESS , \(\eta\) i,ex , \(\eta\) i,co2 are all penalty coefficients.

[0081] In step 3, the non - linear penalty can ensure that when the carbon emission is too high, the penalty increases rapidly, thus promoting the agent to make low - carbon decisions. Set the non - linear function of carbon emission penalty as follows:

[0082]

[0083] In the formula, \(\alpha\) and \(\beta\) are coefficients to adjust the penalty intensity.

[0084] Designing a collaborative reward mechanism can encourage power trading between microgrids to achieve the optimization of collaborative scheduling of the microgrid group. When two microgrids conduct power trading, the reward can be calculated according to the trading power and the economy of the trading. Set the collaborative reward \(R\) co between microgrids as follows:

[0085]

[0086] In the formula, \(\gamma\) co is the coefficient of collaborative reward, controlling the intensity of the reward; \(P\) ij is the power trading volume between microgrid \(i\) and microgrid \(j\); \(\Delta P\) ij is the economy of power trading between microgrid \(i\) and microgrid \(j\), that is, the price difference of power trading.

[0087] Based on the original reward function, the complete reward function formula after combining the non - linear penalty of carbon governance cost and the collaborative reward between microgrids is:

[0088]

[0089] In the formula, represents the reward function of each microgrid agent after introducing the non - linear penalty of carbon governance cost and the collaborative reward among microgrids.

[0090] In step 4, to solve the problem of mutual dependence among agents and protect privacy, a parameter - sharing framework is introduced. Under this framework, all agents contribute their experience samples to a central shared experience pool for other agents to use jointly. When an agent extracts samples from the experience pool for learning, it can benefit from the experiences of other agents, thereby improving the training efficiency and sample utilization rate. In addition, by introducing negative sampling in the shared experience pool, overfitting and local optimal solutions can be effectively avoided. It includes the following steps:

[0091] (1) Construct a shared experience pool:

[0092] In a multi - agent environment, store the experience sequence generated by each agent i into a shared experience pool D:

[0093]

[0094] In the formula, is the observation of agent i at time t; a i t represents the action of microgrid agent i at time t; is the reward obtained by agent i at time t; is the observation of agent i at time t + 1.

[0095] (2) Negative sample judgment based on the advantage function:

[0096] 1) Advantage function The calculation formula of:

[0097]

[0098] In the formula, is the observed action value function of agent i, representing the expected total return that agent i can obtain by taking action under the current observation ; is the observed value function of agent i, representing the expected return of agent i under the current observation .

[0099] 2) Negative sample judgment:

[0100]

[0101] In the formula, B neg is the negative sample batch; B pos is the positive sample batch; θ(t) is the dynamically adjusted threshold; θmin is the initial threshold, usually set to a negative value; θ max is the maximum threshold, usually set to 0 or a positive value; t is the current training step; T is the total number of training steps.

[0102] θ(t) sets a relatively low threshold initially to sample more negative samples; while in the later stage of training, as the policy is gradually optimized, the threshold is gradually increased to reduce the selection of negative samples and maintain the stability of training.

[0103] (3) Negative sample sampling:

[0104] After introducing negative samples, it is necessary to set the negative sample sampling ratio p neg to control the proportion of negative samples in training to ensure that it will not cause too much interference to training. Then the final training sample batch B final is expressed as:

[0105] B final = (1 - p neg )·B pos + p neg ·B neg .

[0106] In step 5, it includes:

[0107] 5.1: LSTM - MAPPO model:

[0108] In the low - carbon optimal scheduling problem of the microgrid cluster, the system has temporal characteristics, that is, the decision of a certain microgrid agent not only depends on the current observation, but is also closely related to historical observations, past decisions and long - term behaviors. For example, the charge - discharge strategy of the energy storage system and the start - stop decision of the gas turbine are often affected by previous operations. As a recurrent neural network for processing temporal data, LSTM (Long Short - Term Memory network) can effectively capture this temporal dependence. Therefore, introducing the LSTM network into the MAPPO algorithm can help the model better identify temporal relationships and improve the learning ability of the multi - agent system.

[0109] The specific steps are as follows:

[0110] 1) LSTM network hidden state input:

[0111] In the traditional MAPPO algorithm, the observation of each agent is input into the policy network and the value network. After introducing the LSTM network, the output of the LSTM network will be passed as new observation information to the policy network for decision - making, so that each agent can consider the influence of historical observations while considering the current - moment observation. This will affect the observation representation used in the advantage function and the target function. Define the output of the LSTM network as

[0112]

[0113] wherein, to are the historical observation information of agent i at the past k time steps; represents the historical observation information of agent i at time t - k + 1;

[0114] 2) Update of the advantage function and the objective function:

[0115] In the traditional MAPPO algorithm, the advantage function is calculated based on the current observation and the estimation of the value network. After introducing the LSTM network, the calculation of the advantage function will not only depend on the current observation but also depend on the hidden state output by the LSTM network

[0116]

[0117] wherein, is the advantage function considering h t ; and are the action-value function and the observation-value function considering the hidden state respectively; N represents the total number of samples; L clip (θ, h t ) is the objective function embedded in the LSTM network; p t (θ) is the policy ratio considering the hidden state . ε represents the hyperparameter of the clipping amplitude, usually taking 0.1 or 0.2; clip(p t (θ), 1 - ε, 1 + ε) represents the clipping operation, which is used to prevent the policy update amplitude from being too large and causing training instability.

[0118] 5.2: Differential learning rate decay strategy:

[0119] To improve the training efficiency, stability, and generalization ability of LSTM in complex multi-agent systems, a differential learning rate decay strategy is introduced in the LSTM-MAPPO model. There are differences in the tasks and learning requirements of the LSTM network and the MAPPO algorithm in the joint training. In view of this characteristic, adopting different learning rate strategies can better meet their respective needs: the LSTM network part requires a stable change in the learning rate to accurately extract temporal features; while the MAPPO algorithm part requires a faster policy update to better adapt to the exploration and optimization process of the environment.

[0120] 1) Cosine Restart Decay:

[0121] Introduce a cosine restart learning rate decay strategy in the LSTM network to help the LSTM network periodically adjust the learning rate during training, thus avoiding premature convergence and promoting exploration. Each period can be gradually increased as needed, so that the LSTM network can repeatedly "explore" and "converge" during training, avoiding getting stuck in local optimal solutions prematurely. The specific formula is as follows:

[0122]

[0123] In the formula, η LSTM (t) is the learning rate of the LSTM network at the t-th step; η max is the maximum learning rate for each period; η min is the minimum learning rate; T max is the maximum number of steps in a period; t is the number of steps of the current training; T restart is the starting time of each period.

[0124] 2) Linear Decay:

[0125] The linear decay learning rate strategy is suitable for optimizing the learning rate of the MAPPO algorithm, helping to conduct larger exploration in the initial stage and gradually converging in the later stage to improve stability. The specific formula is as follows:

[0126]

[0127] In the formula, η MAPPO (t) is the learning rate of the MAPPO algorithm at the t-th step; η initial is the initial learning rate; T total is the total number of training steps; t is the number of steps of the current training.

[0128] The step 6 includes the following steps:

[0129] S6.1: Initialize the low-carbon optimal operation model of the microgrid group, and define the observation space, action space, reward function and corresponding constraint conditions for each microgrid agent.

[0130] S6.2: Each agent selects an action according to the current policy, records the current observation, action and reward, stores these experiences in the shared experience pool, and balances exploration and exploitation at the same time.

[0131] S6.3: When the agent samples experience samples, a negative sampling method is adopted to optimize the diversity and representativeness of training samples. S6.4: Use the LSTM-MAPPO model with a differential learning rate strategy. Calculate the advantage function and objective function of each microgrid agent. Calculate the policy gradient based on the objective function and update the policy of the microgrid agent to ensure the collaborative optimization of the policy. S6.5: Repeat S6.2 to S6.4 until the performance of the LSTM-MAPPO model is stable or the maximum number of training rounds is reached.

[0132] S6.6: After the training is completed, export the optimal policy, conduct offline verification and deploy it to the low-carbon optimal scheduling system of the microgrid cluster to ensure its effectiveness in the actual environment.

[0133] A low-carbon optimal operation method for a microgrid cluster based on an improved MAPPO algorithm of the present invention has the following technical effects:

[0134] 1) Step 1 of the present invention constructs an optimal operation model for the microgrid cluster with the minimum total cost within a single microgrid operation cycle as the objective function. This model innovatively combines economic optimization with the goal of green development, not only clarifying the optimization direction but also providing a key theoretical support and technical basis for the subsequent steps.

[0135] 2) Step 2 of the present invention is to construct a partially observable Markov decision process for the optimal operation of the microgrid, and establish the observation space, action space and reward function corresponding to each microgrid agent. This step decomposes the complex microgrid operation problem into a learnable multi-agent decision process by clarifying the state-behavior mapping relationship of the agent in a partially observable environment. At the same time, this step lays a theoretical foundation for the construction of the MAPPO model, ensuring that the agent can achieve efficient learning and optimal decision-making in a dynamic and incomplete information environment, which is an important prerequisite for the entire optimization algorithm.

[0136] 3) Step 3 of the present invention is to introduce the non-linear constraint of carbon governance cost and the collaborative reward among microgrids into the reward function to reduce the carbon emissions of the microgrid cluster and promote the electricity trading among microgrids. The non-linear penalty mechanism ensures that when the carbon emissions are too high, the penalty intensity increases rapidly by designing the characteristic that the carbon emission cost increases rapidly with the increase of emissions, thus effectively promoting the agent to give priority to low-carbon operation strategies in the decision-making process and actively reducing carbon emissions.

[0137] Meanwhile, a design collaboration reward mechanism is developed to encourage power trading among microgrids, thereby optimizing the collaborative scheduling of the microgrid cluster. This mechanism calculates appropriate reward values by analyzing the traded electricity volume, trading economy, and the impact of trading behavior on the overall system benefits between two microgrids. It not only motivates the collaborative behavior of microgrids but also improves energy utilization efficiency and reduces energy waste. Through the introduction of this dual mechanism, the balance ability of the model between low-carbon goals and economic optimization is further enhanced, providing innovative support for the overall operation optimization of the microgrid cluster.

[0138] 4) Step 4 of the present invention is to solve the problems of mutual dependence and privacy protection among agents, while improving the training efficiency and learning effect. A negative sampling experience sharing framework is introduced into the MAPPO algorithm. Specifically, through the mechanism of parameter sharing, the experience samples of all agents are pooled into a central shared experience pool for other agents to use jointly. This design not only enhances the collaborative ability among agents but also fully exploits the learning potential of multi-agents while ensuring privacy.

[0139] When an agent extracts samples from the shared experience pool for training, it can utilize the experiences of other agents to make up for the deficiency of its own observation information, thus significantly improving the sample utilization rate and training efficiency. At the same time, by introducing a negative sampling mechanism into the shared experience pool, the overfitting of the model to sub-optimal strategies is effectively avoided, reducing the risk of falling into local optimal solutions. This framework design fully balances the collaboration and independence among agents, providing key support for the stability and convergence of the optimization algorithm.

[0140] 5) Step 5 of the present invention embeds the LSTM network into the MAPPO algorithm to help the model better capture temporal dependencies to enhance the decision-making ability of agents, and designs a differential learning rate decay strategy to further improve the training speed of LSTM-MAPPO.

[0141] In the low-carbon optimal scheduling problem of the microgrid cluster, the system has obvious temporal characteristics, that is, the decision-making of a certain microgrid agent not only depends on the current observation but is also closely related to historical observations, past decisions, and long-term behaviors. For example, the charge and discharge strategies of the energy storage system and the start-stop decisions of gas turbines are often strongly influenced by previous operations. As a recurrent neural network specialized in processing temporal data, LSTM (Long Short-Term Memory network) can effectively capture such temporal dependencies, thereby helping the model identify complex time-related relationships and enhancing the learning ability of agents in dynamic environments. Therefore, embedding the LSTM network into the MAPPO algorithm can significantly improve the decision-making effect of the multi-agent system in complex temporal scenarios.

[0142] In addition, to further improve the training efficiency, stability, and generalization ability of the LSTM-MAPPO model in complex multi-agent systems, a differential learning rate decay strategy is proposed. Since there are differences in the task requirements of LSTM and MAPPO during joint training, it is particularly important to design independent learning rate decay strategies for different parts. Specifically: For the LSTM part: A more stable learning rate change is required to ensure accurate extraction of temporal features and avoid feature extraction errors caused by excessive learning rate fluctuations.

[0143] For the MAPPO part: Faster policy updates are needed to accelerate the exploration and optimization process of the environment, enabling the agent to quickly adapt to the dynamic environment.

[0144] The design of this differential learning rate strategy fully meets the different requirements of LSTM and MAPPO during joint training, thereby improving the overall training efficiency and adaptability of the model and providing more efficient technical support for solving complex multi-agent optimization problems.

[0145] 6) Step 6 of the present invention trains the agent based on the improved MAPPO to obtain the optimal microgrid cluster optimized operation plan. As the final key step, after optimizing and improving the MAPPO algorithm, it needs to be comprehensively trained. By learning the complex dynamic behavior and multi-agent cooperation mechanism of the microgrid cluster through the model, the optimal operation strategy that meets the low-carbon goal and economic requirements is finally output. This step ensures that the improved algorithm can effectively play its role in practical applications and provides a reliable solution for the efficient and green operation of the microgrid cluster. BRIEF DESCRIPTION OF THE DRAWINGS

[0146] The present invention will be further described below in conjunction with the drawings and examples;

[0147] Figure 1 is the flow chart of the optimization method proposed by the present invention.

[0148] Figure 2 is the schematic diagram of the negative sampling shared experience pool structure.

[0149] Figure 3 is the schematic diagram of the LSTM-MAPPO framework structure.

[0150] Figure 4 is the reward convergence result of three reinforcement learning algorithms.

[0151] Figure 5 is the scheduling result of Microgrid 1. DETAILED DESCRIPTION OF THE INVENTION

[0152] A low-carbon optimized operation method for a microgrid cluster based on an improved MAPPO algorithm Figure 1This is the flow chart of the optimization method proposed in the present invention. First, the present invention takes the minimum total cost within a single microgrid operation cycle as the objective function to establish a microgrid optimal operation model; secondly, constructs a partially observable Markov decision process for microgrid optimal operation, and establishes the observation space, action space and reward function corresponding to each microgrid agent; then, introduces the non-linear constraint of carbon governance cost and the collaborative reward between microgrids into the reward function to reduce the carbon emissions of the microgrid group and promote the electricity trading between microgrids; to solve the problems of mutual dependence and privacy protection between agents, and at the same time improve the training efficiency and sample utilization rate, a negative sampling experience sharing framework is introduced into the MAPPO algorithm; embeds the LSTM network into the MAPPO algorithm to help the model better capture the temporal dependence to enhance the decision-making ability of the agent, and designs a differential learning rate decay strategy to further improve the training speed of the LSTM-MAPPO model; finally, trains the agent based on the improved MAPPO algorithm to obtain the optimal low-carbon operation plan for the microgrid group. The optimization method proposed in the present invention can effectively reduce the operation cost of each microgrid and at the same time reduce the carbon emissions of the system. The improved MAPPO algorithm is not a simple LSTM-MAPPO model. The present invention makes three improvements to the MAPPO algorithm, namely steps 3, 4, and 5, that is:

[0153] 1) Introduce the non-linear constraint of carbon governance cost and the collaborative reward between microgrids into the reward function to reduce the carbon emissions of the microgrid group and promote the electricity trading between microgrids.

[0154] 2) To solve the problems of mutual dependence and privacy protection between agents, and at the same time improve the training efficiency and sample utilization rate, a negative sampling experience sharing framework is introduced into the MAPPO algorithm.

[0155] 3) Embed the LSTM network into the MAPPO algorithm to help the model better capture the temporal dependence to enhance the decision-making ability of the agent, and design a differential learning rate decay strategy to further improve the LSTM-MAPPO training speed. Therefore, the improved MAPPO algorithm combines the innovations in the above three aspects.

[0156] Figure 2 is the negative sampling shared experience pool structure. As can be seen from Figure 2 All agents share their respective experience samples into a central experience pool for all agents to access and use. When the agent extracts samples from the experience pool for learning, by introducing a negative sample sampling mechanism, those negative samples that are unhelpful or inefficient for policy optimization are selectively screened out, thus effectively reducing unnecessary noise interference, improving the efficiency of the learning process and the quality of the training results. Through this mechanism, the agent can focus more on valuable experiences, accelerate convergence, and optimize the final policy performance.

[0157] Figure 3 It is the structural diagram of the LSTM-MAPPO framework. Figure 3 In it, Critic and Actor represent the policy network and action network of MAPPO respectively. After introducing the LSTM network, the LSTM network will input the current observation and past observations into the MAPPO algorithm network, helping the model better capture temporal dependencies, thereby enhancing the agent's temporal perception ability in the decision-making process and improving the model's performance and decision optimization in dynamic environments.

[0158] Figure 4 It is the reward convergence results of three reinforcement learning algorithms. It can be seen from Figure 4 that the improved MAPPO algorithm has a rapid increase in cumulative rewards in the initial stage of training and tends to be stable after about 700 training rounds, and finally stabilizes at about 0.95, indicating that the algorithm has achieved good learning effects and policy optimization in training. Although compared with the PPO algorithm, the MAPPO algorithm has a faster reward growth rate and a higher final cumulative reward value. However, both the MAPPO algorithm and the PPO algorithm have problems of insufficient stability during the training process, and the final rewards are much lower than those of the improved MAPPO algorithm. In summary, the improved MAPPO algorithm has the characteristics of fast convergence speed, high reward value, and strong stability.

[0159] Figure 5 It is the scheduling result of Microgrid 1. During the period from 00:00 to 04:00, due to the low electricity price of the power grid and large load demand of Microgrid 1, at this time, the new energy generation is insufficient to meet its demand, so Microgrid 1 tends to purchase electricity from other microgrids or the superior distribution network rather than dispatching gas turbines to increase power generation. During the period from 05:00 to 07:00, Microgrid 1 purchases low-price electricity and stores it for subsequent use. During the period from 08:00 to 14:00, the photovoltaic power generation increases and the load demand gradually increases. Microgrid 1 increases its output by dispatching gas turbines and purchases electricity from other microgrids to meet the demand. During the period from 18:00 to 20:00, when the electricity price is high, Microgrid 1 chooses not to trade with the distribution network, but to discharge through the electric energy storage device, increase power generation and conduct electric energy trading with other microgrids to maintain balance. During the scheduling process of the microgrid, it only purchases electricity from the superior distribution network when other microgrids cannot meet its load. In addition, the three microgrids do not sell electricity to the distribution network during the entire scheduling process, but are more inclined to conduct electric energy transactions among microgrids to improve economic benefits.

[0160] Table 1 Comparison of Optimization Schemes

[0161]

[0162] Table 1 shows the comparison of optimization schemes. Three schemes are set up to verify the effectiveness and superiority of the proposed microgrid group optimization scheduling method based on improved MAPPO in the present invention.

[0163] Scheme 1: Microgrid group optimal operation method based on PPO;

[0164] Scheme 2: Microgrid group optimal operation method based on MAPPO;

[0165] Scheme 3: Microgrid group optimal operation method based on the method proposed in the present invention.

[0166] The results show that the operation cost of each microgrid in Scheme 3 is the lowest. At the same time, the average carbon governance cost of the microgrid group is reduced by 13.4% and 25.2% compared with Scheme 1 and Scheme 2 respectively. This result proves that the proposed method can not only effectively reduce the operation cost of the microgrid, but also significantly reduce the system carbon emissions.

Claims

1. A low-carbon optimization operation method for microgrids based on an improved MAPPO algorithm, characterized in that The following steps are involved: Step 1: Taking the minimum total cost within the operation cycle of a single microgrid as the objective function, a microgrid group optimization operation model is constructed; Step 2: Construct a partially observable Markov decision process for the optimal operation of the microgrid, and establish the observation space, action space, and reward function corresponding to each microgrid agent; Step 3: Introduce nonlinear constraints on carbon governance costs and collaborative rewards between microgrids into the reward function; Step 4: Introduce the negative sampling experience sharing framework into the MAPPO algorithm; Step 5: Embed the LSTM network into the MAPPO algorithm and design a differentiated learning rate decay strategy to further improve the training speed of the LSTM-MAPPO model; Step 6: Train the intelligent agent based on the improved MAPPO algorithm to obtain the optimal low-carbon optimization operation plan for the microgrid group.

2. The low-carbon optimization operation method of a microgrid group based on the improved MAPPO algorithm according to claim 1 is characterized in that: In step 1, the constructed microgrid optimization operation model includes the objective function: In addition to the electricity production costs of each distributed energy source, the operating costs of the energy storage system, and the mutual cost of electric power, the carbon governance cost is introduced into the objective function of each microgrid to measure the carbon emission impact of the microgrid operation. The objective function is: C i,MG =minC i,DG +C i,ESS +C i,cbn +C i,grid +C i,j ; In the formula, C i,MG is the total cost of the microgrid; C i,DG is the power generation cost of each distributed generation in microgrid i; C i,ESS is the operating cost of the energy storage system; C i,cbn Carbon governance cost of microgrid i; C i,grid is the power cost of interaction between microgrid i and the upper grid; C i,j is the interaction power cost between microgrid i and microgrid j, and i≠j; 1) Distributed power generation cost: The distributed power sources in the microgrid include wind power, photovoltaic power and gas turbines, and their costs are shown in the following formula: C i,DG =C i,WT +C i,PV +C i,MT ; C i,WT =K WT P i,WT ; C i,PV =K PV P i,PV ; C i,MT =a i P i,MT ; In the formula, C i,WT , C i,PV and C i,MT are the wind turbine operating cost, photovoltaic unit operating cost and gas turbine operating cost of microgrid i respectively; K WT , K PV are the wind turbine use and maintenance cost coefficient and the photovoltaic system use and maintenance cost coefficient respectively; a i is the gas turbine cost coefficient of microgrid i; P i,WT , P i,PV and P i,MT They are the gas turbine output, wind turbine output and photovoltaic system output of microgrid i respectively; 2) Operating cost of energy storage device: The operating cost of the energy storage device is shown as follows: C i,ESS HK i,ESS (η i,on P.S i,on +η i,off P.S i,off )4 In the formula, K i,ESS is the cost coefficient of energy storage device in microgrid i; η i,on , η i,off are the charging and discharging efficiency of the energy storage device; P n,on , P n,off are the charging and discharging power of the energy storage device respectively; 3) Carbon governance costs: The carbon governance cost directly reflects the carbon emissions of the system, and its calculation formula is as follows: C i,cbn =Q i,D the t TQ i,MG the T P; Q i,D =Q i,WT +Q i,PV +Q i,MT +Q i,ESS ; Q i,MG =Q i,D +Q i,B ; P i,DSG =P i,WT +P i,PV +P i,MT ; In the formula, Q i,D The carbon dioxide emissions generated by microgrid i; η t is the carbon tax rate; T is the unit carbon emission tax; Q i,MG It is the sum of the carbon dioxide emissions generated by microgrid i and the carbon dioxide emissions caused by purchasing electricity from the upper power grid; η T is the percentage of carbon dioxide emission reduction; P is the transaction price; Q i,WT , Q i,PV , Q i,MT and Q i,ESS are the carbon dioxide emissions caused by wind power generation, photovoltaic power generation, gas turbine power generation and energy storage device power generation; Q i,B L is the carbon dioxide emissions caused by microgrid i purchasing electricity from the upper power grid; i,B is the total load demand in microgrid i; P i,ESS is the power generation output of the energy storage device; sgn is the sign function; P i,DSG D is the cumulative power generation of all distributed generation equipment excluding energy storage; i,B P is the total amount of electricity purchased from the upper power grid; i,WT , P i,PV and P i,MT They are the outputs of wind turbines, photovoltaic units and micro gas turbines respectively; 4) Transaction costs between microgrids and upper-level power grids and between microgrids: There may be power transactions between the microgrid and the upper grid and between microgrids during the dispatching process. The transaction costs are as follows: C i,grid =q gridbuy P i,gridbuy -q gridsell P i,gridsell ; In the formula, q gridbuy ,q gridsell are the electricity purchase and sales prices of the microgrid to the upper grid respectively; P i,gridbuy , P i,gridsell are the power purchased and sold by microgrid i to the upper grid respectively; q MG is the transaction electricity price between microgrids; P i,mbuy , P i,msell They are the electricity purchased and sold from microgrid i to microgrid j respectively.

3. The low-carbon optimization operation method of a microgrid group based on the improved MAPPO algorithm according to claim 2 is characterized by: The constraints of the objective function include: 1) Microgrid power balance constraints: 2) Upper and lower limits of power interaction between microgrids: Where P i,j , They are the transaction power and the maximum transaction power between microgrids respectively; 3) Upper and lower limits of micro gas turbine output: In the formula, and are the upper and lower limits of the gas turbine output respectively; 4) Micro gas turbine climbing constraints: In the formula, are the gas turbine output power at time t and time t-1 respectively; They are the upper and lower limits of the unit output change per unit time respectively; 5) Energy storage device constraints: In the formula, are the charge states of the energy storage at time t and time t-1 respectively; u is the charge and discharge coefficient; are the upper and lower limits of the energy storage charge state at time t; are the charging and discharging power of energy storage at time t respectively; They are the upper limits of the charging and discharging power of the energy storage at time t; 6) Interaction power constraints between microgrid and upper grid: Where P i,grid , is the transaction power and upper limit of transaction power between microgrid i and the upper grid; 7) Transaction price constraints: In order to promote each microgrid in the microgrid group system to give priority to the transactions within the system, the electricity price transaction between microgrids should be guaranteed to be between the selling price and the repurchase price of the upper network, and the transaction price constraint is obtained as follows: q gridsell ≤q MG ≤q gridbuy ; 8) Carbon emission constraints: In the formula, is the maximum allowable carbon emission of device m in microgrid i during period t; is the carbon emission of device m in microgrid i during period t; is the total carbon emission of equipment m in a scheduling cycle; is the maximum allowable total carbon emission of equipment m within a scheduling cycle; is the total carbon emissions of microgrid i in a scheduling cycle; is the maximum allowable total carbon emission of microgrid i during period t; is the maximum allowable total carbon emissions of microgrid i in a scheduling cycle.

4. The low-carbon optimization operation method of a microgrid group based on the improved MAPPO algorithm according to claim 2 is characterized by: In step 2, the partially observable Markov decision process for the optimal operation of the microgrid includes the process of establishing the observation space, action space and reward function corresponding to each microgrid agent; specifically, it includes: (1) Selection of observation space O: The observation of microgrid i includes the load data at time t Photovoltaic power generation Wind power generation capacity The electricity purchase and sale price of the upper grid grid , microgrid transaction electricity price q MG and energy storage state of charge For the optimal dispatch of microgrid groups, the observation can be expressed as: In the formula, represents the observation of microgrid agent i at time t; (2) Selection of action space A: The goal of economic and low-carbon optimization dispatch of microgrid groups is to determine the optimal output of each unit and the electricity trading situation. Therefore, the output power of the micro gas turbine at time t is Energy storage charging and discharging power Microgrid and upper network trading power And microgrid transaction power As action value: In the formula, represents the action of microgrid agent i at time t; (3) Reward function: For the optimization scheduling problem of microgrid groups, the designed reward function not only reflects the interests of each intelligent agent but also takes into account the constraints of system operation; therefore, in addition to the cost function, the energy storage device operation constraint, the power constraint for interaction with the upper grid, and the carbon emission constraint are added to the reward function as penalty functions. The penalty terms are expressed as Then the reward function of each microgrid is expressed as: In the formula, r i MG represents the reward function of each microgrid agent; η i,ESS , η i,ex , Both are penalty coefficients.

5. The low-carbon optimization operation method of a microgrid group based on the improved MAPPO algorithm according to claim 1 is characterized in that: In step 3, a nonlinear function of carbon emission penalty is set as follows: In the formula, α and β are coefficients for adjusting the penalty intensity; When two microgrids trade electricity, they can calculate rewards based on the amount of electricity traded and the economic efficiency of the transaction. Set the inter-microgrid coordination reward R co as follows: In the formula, γ co is the coefficient of the synergy reward, controlling the intensity of the reward; P ij is the amount of electricity traded between microgrid i and microgrid j; △P ij is the economic efficiency of the electricity traded between microgrid i and microgrid j, i.e., the price difference of the transaction; On the basis of the original reward function, the complete reward function formula after combining the nonlinear penalty of carbon governance cost and the collaborative reward between microgrids is: In the formula, It represents the reward function of each microgrid agent after introducing nonlinear penalties for carbon governance costs and collaborative rewards between microgrids.

6. The low-carbon optimization operation method of a microgrid group based on the improved MAPPO algorithm according to claim 1 is characterized by: The step 4 comprises the following steps: (1) Building a shared experience pool: In a multi-agent environment, the experience sequence generated by each agent i is stored in a shared experience pool D: In the formula, is the observation of agent i at time t; represents the action of microgrid agent i at time t; is the reward obtained by agent i at time t; is the observation of agent i at time t+1; (2) Negative sample judgment based on advantage function: 1) Advantage function The calculation formula is: In the formula, is the observed action value function of agent i, which indicates that agent i is currently observing Take action The expected total return that can be obtained; is the observation value function of agent i, which indicates that agent i is currently observing Expected return under 2) Negative sample judgment: In the formula, B neg is a negative sample batch; B pos is the positive sample batch; θ(t) is the dynamically adjusted threshold; θ min is the initial threshold, usually set to a negative value; θ max is the maximum threshold, set to 0 or a positive value; t is the current training step number; T is the total training step number; θ(t) sets a lower threshold in the early stage so that more negative samples are sampled; in the later stage of training, as the strategy is gradually optimized, the threshold is gradually increased to reduce the selection of negative samples and maintain the stability of training; (3) Negative sample sampling: After introducing negative samples, set the negative sample sampling ratio p neg Control the proportion of negative samples in training to ensure that it does not cause too much interference to the training; then the final training sample batch B final It is expressed as: B final =(1-p neg )·B pos +p neg ·B neg 。 7. The low-carbon optimization operation method of a microgrid group based on the improved MAPPO algorithm according to claim 1 is characterized in that: The step 5 includes: 5.1: LSTM-MAPPO model: Introduce the LSTM network into the MAPPO algorithm. The specific steps are as follows: 1) LSTM network hidden state input: In the MAPPO algorithm, each agent’s observation As input to the policy network and value network; after the introduction of the LSTM network, the output of the LSTM network The new observation information is passed to the policy network for decision making, so that each agent can refer to the influence of historical observations when considering the current observation; this will affect the observation representation used in the advantage function and the objective function; the output of the LSTM network is defined as In the formula, arrive is the historical observation information of agent i at the past k moments; Represents the historical observation information of agent i at time t-k+1; 2) Update of advantage function and objective function: In the MAPPO algorithm, the advantage function is obtained by and the value network to calculate; after the introduction of the LSTM network, the calculation of the advantage function will not only depend on the current observation It also needs to rely on the hidden state output of the LSTM network To consider the timing dependency; the updated advantage function and objective function are: In the formula, To consider h t The advantage function after and Consider the hidden state The action value function and observation value function after that; N represents the total number of samples; L clip (θ,h t ) is the objective function embedded in the LSTM network; p t (θ) is the hidden state The post-strategy ratio; ε represents the hyperparameter of the clipping range; clip(p t (θ), 1-ε, 1+ε) represents the clipping operation, which is used to prevent the update amplitude of the strategy from being too large, resulting in unstable training; 5.2: A differentiated learning rate decay strategy is introduced in the LSTM-MAPPO model: 1) Cosine restart attenuation: The cosine restart learning rate decay strategy is introduced into the LSTM network to help the LSTM network periodically adjust the learning rate during training, thereby avoiding premature convergence and promoting exploration; the specific formula is as follows: Where η LSTM (t) is the learning rate of the LSTM network at the tth step; η max is the maximum learning rate for each cycle; η min is the minimum learning rate; T max is the maximum number of steps in a cycle; t is the number of steps in the current training; T restart is the start time of each cycle; 2) Linear attenuation: The linear decay learning rate strategy is suitable for optimizing the learning rate of the MAPPO algorithm, helping to conduct greater exploration in the early stage and gradually converge in the later stage to improve stability; the specific formula is as follows: Where η MAPPO (t) is the learning rate of the MAPPO algorithm at step t; η initial is the initial learning rate; T total is the total number of training steps; t is the number of current training steps.

8. The low-carbon optimization operation method of a microgrid group based on the improved MAPPO algorithm according to claim 7 is characterized in that: The step 6 comprises the following steps: S6.1: Initialize the low-carbon optimization operation model of the microgrid group, and define the observation space, action space, reward function and corresponding constraints for each microgrid agent; S6.2: Each agent selects actions based on the current strategy, records current observations, actions, and rewards, and stores these experiences in a shared experience pool, while balancing exploration and exploitation; S6.3: When the agent samples experience, a negative sampling method is used to optimize the diversity and representativeness of training samples; S6.4: Use the LSTM-MAPPO model with a differentiated learning rate strategy; Calculate the advantage function and objective function of each microgrid agent; Calculate the policy gradient based on the objective function and update the policy of the microgrid agent to ensure the coordinated optimization of the policy; S6.5: Repeat S6.2 to S6.4 until the performance of the LSTM-MAPPO model is stable or the maximum number of training rounds is reached; S6.6: After training is completed, the optimal strategy is exported, offline verified and deployed to the low-carbon optimization scheduling system of the microgrid cluster to ensure its effectiveness in the actual environment.

Citation Information

Patent Citations

  • Economic dispatching method for multi-park integrated energy system

    CN114417695A

  • Wireless routing optimization method and network system based on deep contrast reinforcement learning

    CN117749692A

  • Microgrid space-time perception energy management method based on secure deep reinforcement learning

    WO2024108817A1

Cited By

  • Multi-train dynamic scheduling method and system based on MAPPO algorithm

    CN120716796A

  • A method and system for dynamic scheduling of multiple trains based on the MAPPO algorithm

    CN120716796B

  • Microgrid demand side management method and system based on group enhancement strategy optimization

    CN121192715A