Micro-grid collaborative optimization scheduling method and related device
By employing a hybrid multi-agent soft-actor critic algorithm and generative adversarial learning in distribution microgrids, a multi-agent reinforcement learning model was designed. This model addresses the challenges of dense communication resources and difficult model convergence in the collaborative optimization of distribution microgrids, achieving efficient and stable control strategy solutions and ensuring system safety and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2026-03-27
AI Technical Summary
Existing collaborative optimization methods for distribution microgrids suffer from problems such as dense communication resources, large data acquisition volume, inability to guarantee privacy, difficulty in convergence of model-driven methods for non-convex optimization problems with discrete variables, and low computational efficiency. The complexity increases, especially when dealing with uncertainties.
A hybrid multi-agent soft actor critic algorithm is adopted. Based on a pre-trained policy network, the cooperative optimization model of the distribution network and microgrid is designed as a reinforcement learning model in a multi-agent environment. The policy network of the reinforcement learning model is pre-trained through generative adversarial learning. Combined with the state observation data of the distribution network and microgrid, the optimization model is solved to obtain the control strategy.
It effectively solves the problem of collaborative optimization and scheduling of high-dimensional, non-convex and nonlinear distribution microgrids, with fast convergence speed and training results. It is suitable for collaborative optimization of discrete equipment in complex non-convex models and under multiple uncertainties, meets the needs of real-time decision-making, and improves the safety, stability and collaborative optimization efficiency of the system.
Smart Images

Figure CN119787370B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of power distribution networks, and relates to a power distribution micro-grid collaborative optimization scheduling method and related device. BACKGROUND
[0002] As an important public infrastructure, the power distribution network plays an important role in ensuring power supply, supporting economic and social development, and serving to improve people's livelihood. With the wide access of distributed resources and energy storage devices, the coupling between the power distribution micro-grid is increasingly close. However, due to the inherent strong uncertainty of space-time coupling of the power distribution micro-grid system, the diversity of controllable resources, heavy communication burden and concurrent competitive local behavior, the regulation and coordination face serious challenges. In order to ensure the safe and stable operation of the power distribution micro-grid system, improve the convergence efficiency of large-scale grid model training, reduce the system operation cost, and protect privacy, it is of great significance to fully utilize the various flexible resources in the power distribution micro-grid and the information interaction and collaborative optimization capabilities between regions.
[0003] At present, the collaborative optimization problem of the power distribution micro-grid mainly adopts the model-driven centralized optimization method. The traditional centralized optimization requires intensive communication resources, has a large amount of data collection, and cannot guarantee the privacy of multiple subjects. Therefore, distributed or decentralized optimization is proposed, which uses limited information exchange to reach consensus and global optimization under the coordination of a weak central entity, and has better flexibility and privacy protection. However, the power distribution micro-grid side usually contains discrete devices such as on-load voltage regulating transformers and switched capacitor banks to reduce voltage over-limit and reduce economic losses. However, the model-driven method has convergence difficulties and low computational efficiency for non-convex optimization problems containing discrete variables. Moreover, the model-driven method usually uses stochastic optimization or robust optimization methods to handle uncertain factors, further increasing the complexity of the construction and solution of the power distribution micro-collaborative optimization model. SUMMARY
[0004] The purpose of the present application is to overcome the above-mentioned shortcomings of the prior art, and to provide a power distribution micro-grid collaborative optimization scheduling method and related device.
[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0006] In a first aspect, the application provides a micro-grid coordinated optimization scheduling method, comprising: obtaining state observation data of a power grid agent and each micro-grid agent; using a hybrid multi-agent soft actor critic algorithm, based on a pre-trained policy network, solving a preset micro-grid coordinated optimization model according to the state observation data of the power grid agent and each micro-grid agent to obtain a control strategy of the power grid and each micro-grid; wherein the pre-trained policy network is obtained by: designing the micro-grid coordinated optimization model as a reinforcement learning model in a multi-agent environment, and pre-training a policy network of the reinforcement learning model based on a generative adversarial imitation learning method to obtain the pre-trained policy network; and the micro-grid coordinated optimization model comprises: a micro-grid optimization scheduling model constructed with a minimum total production cost as an optimization objective, and a power grid optimization scheduling model constructed with a minimum total operation cost as an optimization objective.
[0007] Optionally, the state observation data of the power grid agent comprises: observation data of a static var compensator in the power grid, observation data of a renewable energy unit, observation data of a micro gas turbine, observation data of an energy storage system, observation data of an on-load tap-changing transformer, observation data of a switched capacitor bank, a number of times that the on-load tap-changing transformer has been operated and a number of times that the switched capacitor bank has been operated, an active load of the power grid superimposed with a normal distribution error, a set of power purchase and sale prices of the power grid, and boundary information from each micro-grid agent; and the state observation data of each micro-grid agent comprises: observation data of a renewable energy unit in the micro-grid, observation data of a micro gas turbine, and observation data of an energy storage system, an active load of the micro-grid superimposed with a normal distribution error, and boundary information from the power grid agent.
[0008] Optionally, the power grid agent is set by: virtually dividing the power grid into a plurality of autonomous regions, setting a continuous agent for each autonomous region, setting a separate agent for discrete devices in the power grid that are independent of each other and have multiple positions, and combining the continuous agents and the separate agents to obtain the power grid agent; and each micro-grid agent is set by: setting an autonomous continuous agent for each micro-grid to obtain each micro-grid agent.
[0009] Optionally, the objective function of the micro-grid optimization scheduling model is:
[0010]
[0011] wherein, F MG,i is the total production cost of the micro-grid i; T is the scheduling duration; is the number of micro gas turbines in the micro-grid i; is the power generation cost of the jth micro gas turbine at the tth time period; is the number of energy storage units in the micro-grid i; is the generation cost of the jth energy storage unit in the tth time period; is the purchase and sale cost of the ith microgrid to the distribution network in the tth time period.
[0012] The constraint conditions of the microgrid optimization scheduling model include renewable energy unit output constraints, micro gas turbine output constraints, energy storage system constraints, microgrid internal power flow balance constraints, and microgrid transmission power constraints.
[0013] The objective function of the distribution network optimization scheduling model is:
[0014]
[0015] wherein F ADN is the total operation cost of the distribution network; and N MG are the number of micro gas turbines, energy storage systems, switched capacitor banks, and microgrids in the distribution network, respectively; and are the generation cost of the jth micro gas turbine and energy storage unit in the distribution network in the tth time period, respectively; t OLTC and are the action cost of the on-load voltage regulating transformer and the jth switched capacitor bank in the distribution network in the tth time period, respectively; is the purchase and sale cost of the ith microgrid to the distribution network in the tth time period, and F t B′-S′ is the purchase and sale cost of the distribution network to the main network in the tth time period.
[0016] The constraint conditions of the distribution network optimization scheduling model include on-load voltage regulating transformer action frequency constraints, switched capacitor bank action frequency constraints, static var compensator output constraints, distribution network power flow balance constraints, node voltage constraints, and load flow constraints.
[0017] Optionally, the microgrid-distribution network collaborative optimization model is designed as a reinforcement learning model in a multi-agent environment, which includes designing the microgrid optimization scheduling model and the distribution network optimization scheduling model as reinforcement learning models according to the observation space, the action space, and the reward function.
[0018] The observation space of the reinforcement learning model of the microgrid optimization scheduling model is:
[0019]
[0020] wherein, and are the observation data of the energy storage system and the observation data of the micro gas turbine of the ith microgrid in the tth time period, respectively, and The observation data of the renewable energy units of the microgrid i and the active load superimposed with the normally distributed error of the time period t, respectively. The boundary information from the power distribution network agent for the time period t-1.
[0021] The action space of the reinforcement learning model of the microgrid optimal scheduling model For:
[0022]
[0023] Wherein, The reactive power prediction value of the renewable energy unit in the microgrid i for the time period t, And The active power and the reactive power of the micro gas turbine in the microgrid i for the time period t, respectively, The active power of the energy storage system charging and discharging in the microgrid i for the time period t.
[0024] The reward function of the reinforcement learning model of the microgrid optimal scheduling model For:
[0025]
[0026] Wherein, And The generation cost of the jth micro gas turbine and the generation cost of the jth energy storage unit for the time period t, respectively; The cost of buying and selling electricity from the power distribution network by the microgrid i for the time period t; F t TL The transmission power constraint over-limit penalty function between microgrids; ζ 1,2 The over-limit penalty weight; δ i The willingness coefficient of the microgrid i to support the safe operation of the power distribution network; F t penalty The over-limit penalty function of the power distribution network safe operation constraint.
[0027] The observation space of the reinforcement learning model of the power distribution network optimal scheduling model For:
[0028]
[0029] Wherein, And The observation data of the on-load voltage regulating transformer and the observation data of the switched capacitor bank, respectively, And The observation data of the static reactive power compensator of the microgrid i, the observation data of the energy storage system and the observation data of the micro gas turbine, And The observed data of the renewable energy units of the microgrid i and the active load superimposed with the normally distributed error for the t period; The set of power purchase and sale prices of the distribution network for the t period; The boundary information from the microgrid i microgrid agent.
[0030] The action space of the reinforcement learning model of the distribution network optimal scheduling model is:
[0031]
[0032] Wherein, And are the discrete action space and the continuous action space in the region i of the distribution network, And are the on-load voltage regulating transformer tap and the switched capacitor bank tap of the region i of the distribution network for the t period; is the reactive power prediction value of the renewable energy unit of the microgrid i for the t period, And are the active power and the reactive power of the micro gas turbine of the microgrid i for the t period, is the charging and discharging active power of the energy storage system of the microgrid i for the t period, is the reactive power output by the static var compensator of the microgrid i for the t period.
[0033] The reward function of the reinforcement learning model of the distribution network optimal scheduling model is:
[0034]
[0035] Wherein, F t OLTC And are the action cost of the on-load voltage regulating transformer and the action cost of the jth switched capacitor bank for the t period, And N MG are the number of micro gas turbines, energy storage systems, switched capacitor banks and microgrids in the distribution network; And are the power generation cost of the jth micro gas turbine and the power generation cost of the jth energy storage unit for the t period, F t B′-S′ is the power purchase and sale cost of the distribution network to the main network, is the power purchase and sale cost of the microgrid i to the distribution network for the t period, is the out-of-limit penalty function of the voltage and branch flow constraints of the distribution network.
[0036] Optionally, in the solving of the preset micro-grid collaborative optimization model based on the pre-trained policy network and according to the state observation data of the power grid agent and each micro-grid agent by using the hybrid multi-agent soft actor-critic algorithm, the micro-grid collaborative optimization model adopts a double-layer architecture of a power grid layer and a micro-grid layer; wherein the power grid optimization scheduling model of the power grid layer is set as a discrete-continuous agent containing a Critic network, a discrete Actor network and a continuous Actor network, and the Critic network is used to guide the update of the discrete Actor network and the continuous Actor network; each micro-grid optimization scheduling model of the micro-grid layer is set as a continuous agent containing a double Critic network and a continuous Actor network, and the double Critic network is used to guide the update of the continuous Actor network; the Critic network of the micro-grid optimization scheduling model updates parameters by minimizing Bellman residual error, and the Actor network and temperature parameters are updated by using a policy gradient with maximum entropy; the Critic network of the power grid optimization scheduling model updates parameters by using a centralized joint regression loss function, introduces OU noise in the continuous Actor network of the power grid optimization scheduling model, and introduces a multi-head attention mechanism in the Critic network of the power grid optimization scheduling model for selectively paying attention to information from each micro-grid optimization scheduling model.
[0037] In the second aspect of the present application, a micro-grid collaborative optimization scheduling system is provided, comprising: a data acquisition module for acquiring state observation data of a power grid agent and each micro-grid agent; a strategy generation module for solving a preset micro-grid collaborative optimization model based on a pre-trained policy network and according to the state observation data of the power grid agent and each micro-grid agent by using a hybrid multi-agent soft actor-critic algorithm, to obtain control strategies of the power grid and each micro-grid; wherein the pre-trained policy network is obtained by: designing the micro-grid collaborative optimization model as a reinforcement learning model in a multi-agent environment, and pre-training the policy network of the reinforcement learning model based on a generative adversarial imitation learning method to obtain the pre-trained policy network; and the micro-grid collaborative optimization model comprises: each micro-grid optimization scheduling model constructed with the minimum total production cost as the optimization target, and a power grid optimization scheduling model constructed with the minimum total operation cost as the optimization target.
[0038] Optionally, the state observation data of the power distribution network agent comprises observation data of static var compensators, observation data of renewable energy units, observation data of micro gas turbines, observation data of energy storage systems, observation data of on-load tap-changing transformers, observation data of switched capacitor banks, the number of times that on-load tap-changing transformers have been operated and the number of times that switched capacitor banks have been operated, active loads of the power distribution network superimposed with normal distribution errors, a set of power purchase and sale prices of the power distribution network, and boundary information from each micro-grid agent; the state observation data of the micro-grid agent comprises observation data of renewable energy units, observation data of micro gas turbines, and observation data of energy storage systems in the micro-grid, active loads of the micro-grid superimposed with normal distribution errors, and boundary information from the power distribution network agent.
[0039] Optionally, the power distribution network agent is set by: virtually dividing the power distribution network into a plurality of autonomous regions, setting a continuous agent for each autonomous region, and setting a separate agent for discrete devices in the power distribution network that are independent of each other and have multiple gears, and combining the agents to obtain the power distribution network agent; and the micro-grid agent is set by: setting an autonomous continuous agent for each micro-grid to obtain the micro-grid agent.
[0040] Optionally, the objective function of the micro-grid optimal dispatching model is:
[0041]
[0042] wherein, F MG,i is the total production cost of the micro-grid i; T is the dispatching duration; is the number of micro gas turbines in the micro-grid i; is the power generation cost of the jth micro gas turbine in the tth period; is the number of energy storage units in the micro-grid i; is the power generation cost of the jth energy storage unit in the tth period; is the power purchase and sale cost of the micro-grid i to the power distribution network in the tth period.
[0043] The constraint conditions of the micro-grid optimal dispatching model comprise renewable energy unit output constraints, micro gas turbine output constraints, energy storage system constraints, micro-grid internal power flow balance constraints, and micro-grid transmission power constraints.
[0044] The objective function of the power distribution network optimal dispatching model is:
[0045]
[0046] wherein, F ADN is the total operation cost of the power distribution network; and N MGrespectively represent the number of micro gas turbines, energy storage systems, switched capacitor banks and microgrids in the distribution network; and respectively represent the generation cost of the jth micro gas turbine and energy storage unit in the distribution network at time t; t OLTC and respectively represent the action cost of the on-load voltage regulating transformer and the jth switched capacitor bank in the distribution network at time t; represents the purchase and sale cost of electricity of the microgrid i to the distribution network at time t, F t B′-S′ represents the purchase and sale cost of electricity of the distribution network to the main network at time t.
[0047] The constraint conditions of the distribution network optimization scheduling model include: on-load voltage regulating transformer action frequency constraint, switched capacitor bank action frequency constraint, static var compensator output constraint, distribution network power flow balance constraint, node voltage constraint and load flow constraint.
[0048] Optionally, the microgrid coordination optimization model is designed as a reinforcement learning model in a multi-agent environment, including: the microgrid optimization scheduling model and the distribution network optimization scheduling model are both designed as reinforcement learning models according to the observation space, the action space and the reward function.
[0049] The observation space of the reinforcement learning model of the microgrid optimization scheduling model is :
[0050]
[0051] wherein, and respectively represent the observation data of the energy storage system and the observation data of the micro gas turbine of the microgrid i at time t, and respectively represent the observation data of the renewable energy unit and the active load of the microgrid i at time t superimposed with a normally distributed error; represents the boundary information from the distribution network agent at time t-1.
[0052] The action space of the reinforcement learning model of the microgrid optimization scheduling model is :
[0053]
[0054] wherein, represents the reactive power prediction value of the renewable energy unit in the microgrid i at time t, and respectively represent the active power and the reactive power of the micro gas turbine in the microgrid i at time t, represents the active power of the energy storage system in the microgrid i at time t.
[0055] Reward function of reinforcement learning model of microgrid optimal scheduling model For:
[0056]
[0057] Wherein, And Respectively, the jth micro gas turbine power generation cost and the jth energy storage unit power generation cost at t period; The purchase and sale of electricity cost of microgrid i to the power grid at t period; F t TL The transmission power constraint out-of-limit penalty function between microgrids; ζ 1,2 The out-of-limit penalty weight; δ i The willingness coefficient of microgrid i to support the safe operation of the power grid; Ft penalty The out-of-limit penalty function of the power grid safe operation constraint.
[0058] Observation space of reinforcement learning model of power grid optimal scheduling model For:
[0059]
[0060] Wherein, And The observation data of on-load voltage regulating transformer and the observation data of switched capacitor group, And The observation data of static var compensator, the observation data of energy storage system and the observation data of micro gas turbine of microgrid i, And The observation data of renewable energy unit and active load of microgrid i superimposed with normally distributed error at t period; The purchase and sale price set of the power grid at t period; The boundary information from the microgrid intelligent agent of microgrid i.
[0061] The action space of the reinforcement learning model of the power grid optimal scheduling model is:
[0062]
[0063] Wherein, And The discrete action space and continuous action space in the region i of the power grid, And The on-load voltage regulating transformer position and switched capacitor group position of the region i of the power grid at t period; The reactive power prediction value of renewable energy unit of microgrid i at t period, and are the active power and reactive power of the micro gas turbine of the microgrid i at time period t, respectively, is the charge-discharge active power of the energy storage system of the microgrid i at time period t, is the reactive power output of the static var compensator of the microgrid i at time period t.
[0064] reward function of the reinforcement learning model of the distribution network optimal dispatching model is:
[0065]
[0066] wherein F t OLTC and are the action cost of the on-load tap-changing transformer and the action cost of the jth switched capacitor bank at time period t, respectively, and N MG are the number of micro gas turbines, energy storage systems, switched capacitor banks and microgrids in the distribution network, respectively; and are the generation cost of the jth micro gas turbine and the jth energy storage unit at time period t, F t B′-S′ is the purchase and sale electricity cost of the distribution network to the main network, is the purchase and sale electricity cost of the microgrid i to the distribution network at time period t, is the out-of-limit penalty function of the voltage and branch power flow constraints of the distribution network.
[0067] Optionally, in the solving of the preset micro-grid collaborative optimization model based on the state observation data of the micro-grid agent and the micro-grid agent, the micro-grid collaborative optimization model adopts a double-layer structure of a power grid layer and a micro-grid layer; the power grid optimization scheduling model of the power grid layer is set as a discrete-continuous agent containing a Critic network, a discrete Actor network and a continuous Actor network, and the Critic network is used to guide the update of the discrete Actor network and the continuous Actor network; each micro-grid optimization scheduling model of the micro-grid layer is set as a continuous agent containing a double Critic network and a continuous Actor network, and the double Critic network is used to guide the update of the continuous Actor network; the Critic network of the micro-grid optimization scheduling model updates parameters by minimizing Bellman residual error, and the Actor network and temperature parameters are updated by introducing a maximum entropy policy gradient; the Critic network of the power grid optimization scheduling model updates parameters by a centralized joint regression loss function, introduces an OU noise in the continuous Actor network of the power grid optimization scheduling model, and introduces a multi-head attention mechanism in the Critic network of the power grid optimization scheduling model for selectively paying attention to information from each micro-grid optimization scheduling model.
[0068] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the micro-grid collaborative optimization scheduling method when executing the computer program.
[0069] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program implements the steps of the micro-grid collaborative optimization scheduling method when executed by a processor.
[0070] Compared with the prior art, the present application has the following beneficial effects:
[0071] The present invention provides a distribution microgrid collaborative optimization scheduling method. By designing the distribution microgrid collaborative optimization model as a reinforcement learning model in a multi-agent environment, and obtaining a pre-trained policy network based on the policy network of the reinforcement learning model pre-trained by generative adversarial learning, the method can fully consider the complex state space and action space of the multi-microgrid and distribution network collaborative model, thereby alleviating the problem of low training efficiency in the early stage of reinforcement learning. Next, a hybrid multi-agent soft-actor critic algorithm is adopted. Based on a pre-trained policy network, and using the state observation data of the distribution network agents and microgrid agents, a pre-defined distribution-microgrid collaborative optimization model is solved to obtain the control strategies of the distribution network and each microgrid. By using the hybrid multi-agent soft-actor critic algorithm, the optimization problem is designed as a distributed optimization problem in a multi-agent reinforcement learning environment. This method can effectively solve high-dimensional, non-convex, and nonlinear distribution-microgrid collaborative optimization scheduling problems. It is suitable for complex non-convex models and distribution-microgrid collaborative optimization with discrete equipment under multiple uncertainties. It has a fast convergence speed and good training results, and has a good ability to cope with environmental non-stationarity. It can also meet the real-time decision-making requirements of the scheduling system. It has obvious advantages in dealing with complex environmental problems and effectively maintains the safety and stability of microgrids and distribution networks. Attached Figure Description
[0072] Figure 1 This is a flowchart of the microgrid collaborative optimization scheduling method according to an embodiment of the present invention.
[0073] Figure 2 This is a schematic diagram of the intelligent agent region division in the IEEE 33-node standard system according to an embodiment of the present invention.
[0074] Figure 3 This is a block diagram of the microgrid collaborative optimization scheduling system according to an embodiment of the present invention. Detailed Implementation
[0075] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0076] It is to be understood that the terms "first", "second", and the like used in the description and the claims of the present application as well as the above-described drawings do not necessarily have to connote any ordinal, sequential or chronological order, but are merely used to distinguish a different set of objects. It is to be understood that the terms so used in the description and the claims are interchangeable under appropriate circumstances and embodiments of the application described herein are capable of operating in other sequences than the one explicitly given in the description and claims. Furthermore, the terms "comprise", "comprising", "include", "including", and the like used in the description and the claims of the present application are used in the sense of "including but not limited to", such that the processes, methods, articles, or apparatuses described herein should not be limited to the steps or elements of the processes, methods, articles, or apparatuses explicitly described in the description and claims.
[0077] The application will be further described in detail below with reference to the accompanying drawings:
[0078] Referring to Figure 1 In an embodiment of the present application, a micro-grid collaborative optimization scheduling method is provided, specifically a micro-grid collaborative optimization scheduling method based on efficient reinforcement learning training technology, which can effectively solve the problem that existing micro-grid collaborative optimization scheduling methods for high-proportion distributed resource access are not economical, efficient and reliable, and improve the collaborative optimization efficiency and effect of micro-grid.
[0079] Specifically, the micro-grid collaborative optimization scheduling method comprises the following steps:
[0080] S1: Obtain state observation data of the power distribution network agent and each micro-grid agent.
[0081] S2: Use a hybrid multi-agent soft actor critic algorithm to solve a preset micro-grid collaborative optimization model based on the state observation data of the power distribution network agent and each micro-grid agent and a pre-trained policy network, to obtain control strategies of the power distribution network and each micro-grid.
[0082] The pre-trained policy network is obtained by the following method: the micro-grid collaborative optimization model is designed as a reinforcement learning model in a multi-agent environment, and a policy network of the reinforcement learning model is pre-trained based on a generative adversarial imitation learning method to obtain the pre-trained policy network; the micro-grid collaborative optimization model comprises a micro-grid optimization scheduling model constructed with the minimum total production cost as the optimization target and a power distribution network optimization scheduling model constructed with the minimum total operation cost as the optimization target.
[0083] The micro-grid coordination optimization scheduling method can fully consider the complex state space and action space of the micro-grid and distribution network coordination model, and further relieve the problem of low training efficiency in the initial stage of reinforcement learning. Then, the hybrid multi-agent soft actor critic algorithm is adopted, and based on the pre-trained policy network, the state observation data of the distribution network agent and each micro-grid agent are used to solve the preset micro-grid coordination optimization model to obtain the control strategy of the distribution network and each micro-grid. The optimization problem is designed as a distributed optimization problem in the multi-agent reinforcement learning environment by using the hybrid multi-agent soft actor critic algorithm, which can effectively solve the high-dimensional, non-convex and nonlinear micro-grid coordination optimization scheduling problem, and is suitable for the micro-grid coordination optimization of the distribution network containing discrete devices under the condition of complex non-convex model and considering multiple uncertainties. The method has fast convergence speed and good training result, has good ability to cope with non-stationary environment, can meet the demand of real-time decision of the scheduling system, has obvious advantages in dealing with complex environmental problems, and effectively maintains the safety and stability of the micro-grid and the distribution network.
[0084] In a possible implementation, the state observation data of the distribution network agent includes observation data of a static var compensator, observation data of a renewable energy unit, observation data of a micro gas turbine, observation data of an energy storage system, observation data of an on-load tap-changing transformer, observation data of a switched capacitor bank, the number of times that the on-load tap-changing transformer has been operated and the number of times that the switched capacitor bank has been operated, an active load of the distribution network superimposed with a normal distribution error, a set of power purchase and sale prices of the distribution network, and boundary information from each micro-grid agent. The state observation data of the micro-grid agent includes observation data of a renewable energy unit, observation data of a micro gas turbine, and observation data of an energy storage system in the micro-grid, an active load of the micro-grid superimposed with a normal distribution error, and boundary information from the distribution network agent.
[0085] Explanatorily, the observation data of the renewable energy unit and the observation data of the micro gas turbine are generally output data, the observation data of the static var compensator is reactive power, the observation data of the energy storage system is charging and discharging power, and the observation data of the on-load tap-changing transformer and the observation data of the switched capacitor bank are generally the current gear. In addition, the normal distribution error is superimposed in the active load to fit the uncertainty, so as to ensure the adaptability of the final result.
[0086] In a possible implementation, the power distribution network agent is set by: virtually dividing the power distribution network into a plurality of autonomous regions, setting a continuous agent for each autonomous region, and setting a separate agent for discrete devices in the power distribution network that are independent of each other and have multiple gears, and combining the agents to obtain the power distribution network agent; and the micro-grid agent is set by: setting an autonomous continuous agent for each micro-grid to obtain the micro-grid agent.
[0087] Explanatorily, the agents are divided for the power grid: the setting of the multi-agent model in the micro-grid system mainly depends on two points: 1, the type of adjustable resources (discrete / continuous adjustment devices) in the system; 2, the decision dimension of each agent should not be too much. Based on this, an autonomous continuous agent is set for the micro-grid side, and considering that there are many dispersed controllable resources on the power distribution network side, in order to ensure that the decision dimension is appropriate, the power distribution network can be virtually divided into a plurality of autonomous regions, and a continuous agent is set for each region, and a separate agent is set for discrete devices that are independent of each other and have multiple gears.
[0088] Exemplarily, referring to Figure 2 Taking the IEEE33 node standard system as an example, the IEEE33 node standard system is divided into three autonomous regions as shown in the figure, three micro-grids are connected to nodes 17, 21 and 32, and a decentralized training-decentralized execution framework is adopted to realize the collaborative optimization operation of the multi-micro-grid-power distribution network system.
[0089] In a possible implementation, the objective function of the micro-grid optimization scheduling model is:
[0090]
[0091] Wherein, F MG,i is the total production cost of the micro-grid i; T is the scheduling time length; is the number of micro gas turbines in the micro-grid i; is the power generation cost of the jth micro gas turbine in the t period; is the number of energy storage units in the micro-grid i; is the power generation cost of the jth energy storage unit in the t period; is the power purchase and sale cost of the micro-grid i to the power distribution network in the t period.
[0092] The objective function of the power distribution network optimization scheduling model is:
[0093]
[0094] Wherein, F ADN is the total operation cost of the power distribution network; and N MG are the number of micro gas turbines, energy storage systems, switched capacitor banks and micro-grids in the power distribution network, respectively; and respectively are the generation cost of the jth micro gas turbine and energy storage unit in the distribution network at time period t; F t OLTC and respectively are the action cost of the on-load tap-changing transformer and the jth switched capacitor bank in the distribution network at time period t; is the buying and selling cost of the microgrid i to the distribution network at time period t, F t B′-S′ is the buying and selling cost of the distribution network to the main grid at time period t.
[0095] Explanatorily, since the distribution network and each microgrid run in a decentralized manner in a sequential decision-making process, only partial state information of the environment can be observed, and the hybrid multi-agent on the distribution network side adopts zoning regulation and control of discrete-continuous equipment, each agent only observes local information of the corresponding area, therefore the studied microgrid coordination optimization problem is modeled as a partially observable Markov decision process.
[0096] In constructing the microgrid coordination optimization model containing a high proportion of distributed resources, the microgrid i takes the minimum total production cost as the optimization objective, and the objective function of the microgrid optimization scheduling model is as follows:
[0097]
[0098] The optimization objective of the distribution network is to minimize the total operation cost, that is, to minimize the sum of the generation cost of the units and the interaction cost between the main grid and the multiple microgrids, and the objective function of the distribution network optimization scheduling model is as follows:
[0099]
[0100] According to the constructed microgrid coordination optimization model containing a high proportion of distributed resources, the constraint conditions of the microgrid coordination optimization model are determined, including the constraint conditions of the microgrid optimization scheduling model and the constraint conditions of the distribution network optimization scheduling model.
[0101] Optionally, the constraint conditions of the microgrid optimization scheduling model include renewable energy unit output constraints, micro gas turbine output constraints, energy storage system constraints, microgrid internal power flow balance constraints, and microgrid transmission power constraints.
[0102] Exemplarily, for the renewable energy unit (RES) constraint, in order to improve the economy, the RES is regulated to generate power in the maximum power point tracking mode, and only the reactive power capacity is used for regulation and control, which is specifically:
[0103]
[0104] wherein, are the active and reactive power prediction values of the jth RES in the ith microgrid at time t, respectively; are the active power of the maximum tracking point operation mode of the jth RES in the ith microgrid at time t, is the number of RES units in the ith microgrid; is the maximum apparent power of the jth RES.
[0105] At the same time, in order to simulate the randomness of the prediction process, the prediction data of the RES and the load are superimposed with prediction errors subject to normal distribution:
[0106]
[0107] wherein, are the active and reactive power prediction values of the jth RES in the ith microgrid at time t, respectively; is the RES prediction error at time t; μ RES , σ RES are the expectation and standard deviation of the RES prediction error.
[0108] For the output constraint of the micro gas turbine (MT), specifically:
[0109]
[0110] wherein, is the active power of the MT, are the minimum and maximum values of the active power of the MT, respectively; are the maximum downward and upward ramping power of the MT, respectively; are the reactive power and the maximum apparent power of the MT, respectively.
[0111] For the constraint of the energy storage system (ESS), specifically:
[0112]
[0113] wherein, are the minimum and maximum values of the active power of the ESS charge and discharge, respectively; soc j,t is the state of charge at time t; soc j,min , soc j,max are the minimum and maximum values of the state of charge, respectively; are the charging power and the discharging power, respectively; is the maximum storage capacity; is the charging and discharging efficiency.
[0114] For the internal power flow balance constraint of the microgrid, specifically:
[0115]
[0116] wherein, respectively are the active and reactive power of the microgrid i at time t; respectively are the active and reactive power of the jth RES in the microgrid i at time t; respectively are the active and reactive power of the MT; are the active and reactive power between the microgrid and the distribution network,
[0117] For the transmission power constraint of the microgrid, specifically:
[0118]
[0119] wherein, are the active and reactive power between the microgrid and the distribution network; respectively are the minimum and maximum values of the transmission active power between the microgrids; is the maximum apparent power.
[0120] Optionally, the constraint conditions of the distribution network optimization scheduling model include: action frequency constraint of on-load tap-changing transformer, action frequency constraint of switched capacitor bank, output constraint of static var compensator, power flow balance constraint of distribution network, node voltage constraint and load flow constraint. The action frequency constraint of on-load tap-changing transformer (OLTC) and the action frequency constraint of switched capacitor bank (CB) are specifically:
[0121]
[0122]
[0123] wherein, respectively record whether the OLTC and the jth CB are in action, the action frequency, the upper limit of the action frequency, V0, V 1,t respectively are the upper and lower end voltages of the OLTC feeder, is the unit step voltage adjustment step of the OLTC; is the middle step of the OLTC; is the unit adjustable reactive power of the jth CB; is the reactive power output by the jth CB at time t.
[0124] The output constraint of the static var compensator (SVC) is specifically:
[0125]
[0126] wherein, is the reactive power output by the jth SVC at time t; respectively are the minimum and maximum values of the reactive power output by the jth SVC; is the number of SVCs in the distribution network.
[0127] The power distribution network power flow balance constraint is specifically:
[0128]
[0129] wherein, P k,t and Q k,t are the active power and the reactive power injected into the power distribution network at node k at time t; V k,t and V l,t are the voltage amplitudes at nodes k and l at time t; G kl , B kl , and θ kl,t are the conductance, the susceptance, and the phase angle difference between nodes k and l.
[0130] The node voltage constraint is specifically:
[0131] V min ≤ V k,t ≤ V max
[0132] wherein, V k,t is the voltage amplitude at node k at time t, and V min and V max are the upper and lower limits of the node voltage constraint.
[0133] The current-carrying capacity constraint is specifically:
[0134] |I kl,t |≤ I kl,max
[0135] wherein, I kl,t is the current flowing through the branch between nodes k and l at time t; and I kl,max is the upper limit of the branch current.
[0136] In a possible implementation, the microgrid coordination optimization model is designed as a reinforcement learning model in a multi-agent environment, which includes: the microgrid optimization scheduling model and the power distribution network optimization scheduling model are both designed as reinforcement learning models according to the observation space, the action space, and the reward function.
[0137] Specifically, the observation space of the reinforcement learning model of the microgrid optimization scheduling model is:
[0138]
[0139] wherein, and are the observation data of the energy storage system and the micro gas turbine of the microgrid i at time t, and The observation data of the renewable energy unit of the microgrid i and the active load superimposed with the error of the normal distribution for the t period of time, respectively; The boundary information from the power distribution network agent for the t-1 period of time.
[0140] The action space of the reinforcement learning model of the microgrid optimization scheduling model Composed of continuous device actions, specifically:
[0141]
[0142] Among them, The reactive power prediction value of the renewable energy unit in the microgrid i for the t period of time, And The active power and the reactive power of the micro gas turbine in the microgrid i for the t period of time, respectively, The active power of the energy storage system charging and discharging in the microgrid i for the t period of time.
[0143] The reward function of the reinforcement learning model of the microgrid optimization scheduling model Is:
[0144]
[0145] Among them, And The generation cost of the jth micro gas turbine and the generation cost of the jth energy storage unit for the t period of time, respectively; The purchase and sale cost of the microgrid i to the power distribution network for the t period of time; F t TL The transmission power constraint out-of-limit penalty function between microgrids; ζ 1,2 The out-of-limit penalty weight; δ i The willingness coefficient of the microgrid i to support the safe operation of the power distribution network; F t penalty The out-of-limit penalty function of the power distribution network safe operation constraint.
[0146] Explanatory, when the above reward function is applied, the microgrid-based agent can support the upper-layer power distribution network, which may affect the safe operation of the power distribution network.
[0147] The observation space of the reinforcement learning model of the power distribution network optimization scheduling model Is:
[0148]
[0149] Among them, And The observation data of the on-load voltage regulating transformer and the observation data of the switched capacitor bank, respectively, And are the observed data of the static var compensator, the observed data of the energy storage system and the observed data of the micro gas turbine of microgrid i respectively, and are the observed data of the renewable energy units of microgrid i and the active load superimposed with normally distributed errors at time period t to fit the uncertainty state and enhance the robustness of the algorithm, is the set of power purchase and sale prices of the distribution network at time period t, is the boundary information from the microgrid i microgrid agent.
[0150] The action space of the reinforcement learning model of the distribution network optimal scheduling model is composed of discrete-continuous device groups, specifically:
[0151]
[0152] wherein, and are the discrete action space and the continuous action space in region i of the distribution network, and are the tap position of the on-load voltage regulating transformer and the switching capacitor bank position of region i of the distribution network at time period t, is the reactive power prediction value of the renewable energy units of microgrid i at time period t, and are the active power and the reactive power of the micro gas turbine of microgrid i at time period t, is the charging and discharging active power of the energy storage system of microgrid i at time period t, is the reactive power output by the static var compensator of microgrid i at time period t.
[0153] The reward function of the reinforcement learning model of the distribution network optimal scheduling model is specifically:
[0154]
[0155] wherein, F t OLTC and are the action cost of the on-load voltage regulating transformer and the action cost of the jth switching capacitor bank at time period t, and N MG are the number of micro gas turbines, energy storage systems, switching capacitor banks and microgrids in the distribution network; and are the power generation cost of the jth micro gas turbine and the power generation cost of the jth energy storage unit at time period t, F t B′-S′ is the power purchase and sale cost of the distribution network to the main network, is the power purchase and sale cost of microgrid i to the distribution network at time period t, A penalty function for the voltage and branch power flow constraints of the distribution network.
[0156] where the reward value of the discrete agent in region i is composed of the negative global discrete device action cost and the global continuous device action cost .
[0157] Explanatorily, the reinforcement learning model in the multi-agent environment constructed above is pre-trained by using the generative adversarial imitation learning (GAIL) to the policy network of the agent. The GAIL adopts the generative adversarial method in the generative adversarial networks (GANs). The algorithm introduces a state-action occupancy measure of the imitator, which is similar to the relevant characteristics of the demonstrator. It uses a discriminator in the GAN to give an action-value function estimate based on the demonstration data. For the reinforcement learning process, the action-value can be obtained from the demonstration by a generative method:
[0158]
[0159] where, are the state spaces of the microgrid and the distribution network, respectively, is the action space of the microgrid i, is the discrete-continuous action space in the distribution network region i; Γ i represents the sample set explored at the i-th iteration, and D is the output value from the discriminator, the parameter of the discriminator is ω, and ω represents that the Q value is estimated after updating the parameter of the discriminator for one step, so the iteration number is i+1.
[0160] The loss function of the discriminator is defined as the general form:
[0161]
[0162] where, i , Γ E are the sample sets from exploration and expert demonstration, respectively.
[0163] In the evaluator network, the following objective function is used to make the evaluator accurately judge the state value function of the current generator policy:
[0164] minE τ [(r t (s t )+V φ (s t+1 )-Vφ (s t )) 2 ]
[0165] where τ represents the sampled policy trajectory from the expert library, r represents the reward function corresponding to the current state, V φ represents the evaluator model output, and φ represents the evaluator parameters.
[0166] In a possible implementation, when the hybrid multi-agent soft actor critic algorithm is used to solve the preset micro-grid collaborative optimization model based on the pre-trained policy network according to the state observation data of the power grid agent and each micro-grid agent, the micro-grid collaborative optimization model adopts a double-layer architecture of a power grid layer and a micro-grid layer; the power grid optimization scheduling model of the power grid layer is set as a discrete-continuous agent containing a Critic network, a discrete Actor network and a continuous Actor network, and the Critic network is used to guide the update of the discrete Actor network and the continuous Actor network; each micro-grid optimization scheduling model of the micro-grid layer is set as a continuous agent containing a double Critic network and a continuous Actor network, and the double Critic network is used to guide the update of the continuous Actor network; the Critic network of the micro-grid optimization scheduling model updates parameters by minimizing the Bellman residual error, and the Actor network and the temperature parameter are updated by using the policy gradient with the maximum entropy; the Critic network of the power grid optimization scheduling model updates parameters by using a centralized joint regression loss function, and the continuous Actor network of the power grid optimization scheduling model introduces an OU noise, and the Critic network of the power grid optimization scheduling model introduces a multi-head attention mechanism for selectively focusing on information from each micro-grid optimization scheduling model.
[0167] Explanatorily, a hybrid multi-agent soft actor critic (HMASAC) algorithm is used to solve the above-mentioned micro-grid collaborative optimization model based on the above-mentioned pre-trained policy network. The HMASAC follows an AC architecture and is composed of a power grid layer and a micro-grid layer. The power grid layer has a single MHAM-Critic network guiding the update of the discrete-continuous Actor network, and the micro-grid layer has a double Critic network guiding the update of the continuous Actor network. The MHAM-Critic network is a Critic network introducing a multi-head attention mechanism.
[0168] The Critic network of each micro-grid agent updates parameters by minimizing the Bellman residual error, and the loss value is as follows:
[0169]
[0170] where, is the double Q network established to avoid overestimation of Q value; τ i is the sampling segment; D is the experience pool; φ 1,2 , are the estimated value and target Critic network parameters, respectively; is the microgrid local observation space, represents the boundary information from the power grid agent; y i,t is the cumulative reward value, and γ is the discount factor; a i,t is the action of the microgrid agent, and α i is the temperature parameter; is the entropy value.
[0171] The Actor network and temperature parameter of each microgrid agent are updated using the policy gradient with maximum entropy, as shown in the following formula:
[0172]
[0173] where τ is the sampling segment, and D is the experience pool; is the microgrid local observation space, represents the boundary information from the power grid agent; α i is the temperature parameter; is the entropy value; θ i is the policy network parameter; is the target entropy value.
[0174] The Critic network of the ADN discrete-continuous agent is updated by a centralized joint regression loss function, as shown in the following formula:
[0175]
[0176] where N ADN is the total number of discrete-continuous agents in the ADN; τ is the sampling segment; and D is the experience pool; is the power grid local observation space, represents the boundary information from the microgrid agent; a t is the action of the agent, and y i,t is the cumulative reward value, γ is the discount factor, α is the temperature parameter, and the entropy function is the target policy network parameter; and H is the number of multi-head attention.
[0177] The loss function for updating the Actor network of the power grid discrete-continuous agent is as shown in the following formula:
[0178]
[0179] where A i (·) is the advantage function, which is used to judge the pros and cons of the current action; α is the temperature parameter, τ is the sampling segment; D is the experience pool; is the microgrid local observation space, represents the boundary information from the distribution network agent, a t is the action of the agent, is the distribution network local observation space, represents the boundary information from the microgrid agent; \i is all agents except agent i; b(o t ,a \i,t ) is the state baseline, which aims to highlight the agents that are more helpful to the target task, so as to obtain a better learning strategy.
[0180] OU noise is introduced into the Actor network of continuous agents to enhance the exploration ability, and the decision of the strategy can be rewritten as:
[0181]
[0182] where, is the observation space of the microgrid, represents the boundary information from the distribution network agent; β is the clipped Gaussian noise added by the Actor network, and OU noise, i.e. Ornstein-Uhlenbeck noise, is a random process with regression characteristics. In reinforcement learning, by introducing time-correlated OU noise, the agent can better explore the environment and find a better strategy.
[0183] According to the solution principle of the HMASAC algorithm, the overall algorithm flow of the microgrid-distribution grid collaborative optimization model based on the HMASAC algorithm is as follows: first, initialize the parameters and the experience pool. In a parallel environment, after each agent makes a decision based on the observation, the action is applied to the power grid environment, and after the power flow calculation, each agent will obtain a reward value. The stored in time t is connected to the new observation in the next time t+1. All agents are sampled from the experience pool D to update the neural network parameters, and the microgrid agent soft updates the double Critic network parameters
[0184]
[0185] The parameters of the Critic network with multi-head attention mechanism are soft-updated by the discrete-continuous agents of the distribution network and the parameters of the Actor network as shown in the following formula:
[0186]
[0187] wherein υ is a soft update factor, υ∈(0,1).
[0188] For example, when solving the preset micro-grid collaborative optimization model, the model can be programmed and solved based on the PyTorch framework. Compared with the un-pretrained reinforcement learning optimization solution, the pre-trained model has faster training convergence speed and better training results. Through the comparison between the multi-agent deep reinforcement learning method with and without the multi-head attention mechanism, not only does it have faster convergence speed and better training results in the training process, but it also has better ability to cope with environmental non-stationarity and meets the demand of real-time decision-making of the scheduling system, and has obvious advantages in dealing with complex environmental problems. The HMASAC algorithm selectively focuses on information from other agents, reducing the dimensionality explosion problem, and has obvious advantages in scalable learning of flexible access of distributed resources on the distribution network side.
[0189] The micro-grid collaborative optimization scheduling method based on the efficient training technology of reinforcement learning provided by the application mainly aims at the current micro-grid collaborative operation which needs to consider multi-agent privacy protection and non-convex model calculation, resulting in great challenges for model-based optimization operation methods, and the problem of uncertain factors such as renewable energy is also paid more and more attention, therefore, the application constructs a micro-grid collaborative optimization model containing a high proportion of distributed resources, adopts a hybrid multi-agent soft actor critic algorithm, and designs the optimization problem as a distributed optimization problem in a multi-agent reinforcement learning environment, so that the high-dimensional, non-convex and nonlinear optimization problem can be effectively solved.
[0190] The micro-grid collaborative optimization scheduling method has the following advantages: (1) the pre-training technology based on the generative adversarial imitation learning method can fully consider the complex state space and action space of the micro-grid collaborative optimization model, and can alleviate the problem of low training efficiency in the early stage of reinforcement learning; (2) the hybrid multi-agent soft actor critic algorithm is used to set the distribution network discrete-continuous agent and the micro-grid continuous agent for the micro-grid collaborative optimization model, and a hybrid multi-agent model for high-proportion distributed resources is constructed; (3) the multi-head attention mechanism is used to realize the coordination between the distribution network discrete-continuous agent and the micro-grid continuous agent, and selectively focuses on information from other agents, effectively alleviating the problem of slow convergence speed caused by dimension explosion of the model. In addition, the method also supports flexible access of distributed resources on the distribution network side, and realizes scalable learning.
[0191] The following is an apparatus embodiment of the application, which can be used to execute the method embodiment of the application. For details not disclosed in the apparatus embodiment, please refer to the method embodiment of the application.
[0192] See Figure 3 In another embodiment of the present invention, a distribution microgrid collaborative optimization scheduling system is provided, which can be used to implement the above-mentioned distribution microgrid collaborative optimization scheduling method. Specifically, the distribution microgrid collaborative optimization scheduling system includes a data acquisition module and a strategy generation module.
[0193] The data acquisition module is used to acquire state observation data of the distribution network agent and each microgrid agent; the strategy generation module is used to use a hybrid multi-agent soft actor critic algorithm, based on a pre-trained policy network, to solve a preset distribution-microgrid collaborative optimization model based on the state observation data of the distribution network agent and each microgrid agent, and obtain the control strategy of the distribution network and each microgrid.
[0194] The pre-trained policy network is obtained as follows: the distribution microgrid collaborative optimization model is designed as a reinforcement learning model in a multi-agent environment, and the policy network of the reinforcement learning model is pre-trained based on the generative adversarial imitation learning method to obtain the pre-trained policy network; the distribution microgrid collaborative optimization model includes: each microgrid optimization scheduling model constructed with the goal of minimizing total production cost and the distribution network optimization scheduling model constructed with the goal of minimizing total operating cost.
[0195] In one possible implementation, the state observation data of the distribution network intelligent agent includes: observation data of static var compensators in the distribution network, observation data of renewable energy units, observation data of micro gas turbines, observation data of energy storage systems, observation data of on-load tap-changing transformers, observation data of switching capacitor banks, the number of times on-load tap-changing transformers have been activated and the number of times switching capacitor banks have been activated, as well as the active load of the distribution network superimposed with normal distribution error, the set of electricity purchase and sale prices of the distribution network, and boundary information from each microgrid intelligent agent; the state observation data of the microgrid intelligent agent includes: observation data of renewable energy units in the microgrid, observation data of micro gas turbines and observation data of energy storage systems, as well as the active load of the microgrid superimposed with normal distribution error and boundary information from the distribution network intelligent agent.
[0196] In one possible implementation, the distribution network intelligent agent is configured as follows: the distribution network is virtually divided into several autonomous regions, a continuous intelligent agent is configured for each autonomous region, and individual intelligent agents are configured for discrete devices in the distribution network that are independent of each other and have multiple levels, and these are combined to obtain the distribution network intelligent agent; the microgrid intelligent agents are configured as follows: an autonomous continuous intelligent agent is configured for each microgrid to obtain each microgrid intelligent agent.
[0197] In one possible implementation, the objective function of the microgrid optimal scheduling model is:
[0198]
[0199] wherein F MG,i is the total production cost of the microgrid i; T is the scheduling duration; is the number of micro gas turbines in the microgrid i; is the generation cost of the jth micro gas turbine in the tth period; is the number of energy storage units in the microgrid i; is the generation cost of the jth energy storage unit in the tth period; is the purchase and sale cost of the microgrid i to the distribution network in the tth period.
[0200] The constraint conditions of the microgrid optimization scheduling model include renewable energy unit output constraints, micro gas turbine output constraints, energy storage system constraints, microgrid internal power flow balance constraints, and microgrid transmission power constraints.
[0201] The objective function of the distribution network optimization scheduling model is:
[0202]
[0203] wherein F ADN is the total operation cost of the distribution network; and N MG are the number of micro gas turbines, energy storage systems, switched capacitor banks, and microgrids in the distribution network, respectively; and are the generation cost of the jth micro gas turbine and energy storage unit in the tth period in the distribution network, respectively; F t OLTC and are the action cost of the on-load voltage regulating transformer and the jth switched capacitor bank in the tth period in the distribution network, respectively; is the purchase and sale cost of the microgrid i to the distribution network in the tth period, F t B′-S′ is the purchase and sale cost of the distribution network to the main network in the tth period.
[0204] The constraint conditions of the distribution network optimization scheduling model include on-load voltage regulating transformer action frequency constraints, switched capacitor bank action frequency constraints, static var compensator output constraints, distribution network power flow balance constraints, node voltage constraints, and load flow constraints.
[0205] In a possible implementation, the microgrid-distribution network collaborative optimization model is designed as a reinforcement learning model in a multi-agent environment, which includes: designing the microgrid optimization scheduling model and the distribution network optimization scheduling model as reinforcement learning models according to the observation space, the action space, and the reward function.
[0206] The observation space of the reinforcement learning model of the microgrid optimization scheduling model is:
[0207]
[0208] wherein, and are the observation data of the energy storage system and the micro gas turbine of the microgrid i at time t, respectively, and are the observation data of the renewable energy unit and the active load of the microgrid i at time t with normally distributed errors superimposed; is the boundary information from the power grid agent at time t-1.
[0209] Action space of the reinforcement learning model of the microgrid optimal scheduling model is:
[0210]
[0211] wherein, is the reactive power prediction value of the renewable energy unit in the microgrid i at time t, and are the active power and the reactive power of the micro gas turbine in the microgrid i at time t, t is the active power of the energy storage system charging and discharging in the microgrid i at time t.
[0212] Reward function of the reinforcement learning model of the microgrid optimal scheduling model is:
[0213]
[0214] wherein, and are the power generation cost of the jth micro gas turbine and the power generation cost of the jth energy storage unit at time t, respectively; is the cost of buying and selling electricity from the power grid by the microgrid i at time t; F t TL is the transmission power constraint over-limit penalty function between microgrids; ζ 1,2 is the over-limit penalty weight; δ i is the willingness coefficient of the microgrid i to support the safe operation of the power grid; F t penalty is the over-limit penalty function of the power grid safe operation constraint.
[0215] Observation space of the reinforcement learning model of the power grid optimal scheduling model is:
[0216]
[0217] wherein, and are the observation data of on-load tap-changing transformers and the observation data of switched capacitor banks, respectively, and are the observation data of static var compensators, the observation data of energy storage systems and the observation data of micro gas turbines of microgrid i, respectively, and are the observation data of renewable energy units of microgrid i and active load superimposed with errors of normal distribution at time period t; is the set of electricity purchase and sale prices of distribution network at time period t; is the boundary information from microgrid i microgrid agent.
[0218] The action space of the reinforcement learning model of the distribution network optimal scheduling model is:
[0219]
[0220] wherein, and are the discrete action space and the continuous action space in region i of the distribution network, and are the on-load tap-changing transformer tap and the switched capacitor bank tap of region i of the distribution network at time period t; is the reactive power prediction value of renewable energy units of microgrid i at time period t, and are the active power and the reactive power of micro gas turbines of microgrid i at time period t, is the charge and discharge active power of energy storage systems of microgrid i at time period t, is the reactive power output by static var compensators of microgrid i at time period t.
[0221] The reward function of the reinforcement learning model of the distribution network optimal scheduling model is :
[0222]
[0223] wherein, F t OLTC and are the action cost of on-load tap-changing transformers and the action cost of the jth switched capacitor bank at time period t, and N MG are the number of micro gas turbines, energy storage systems, switched capacitor banks and microgrids in the distribution network; and are the power generation cost of the jth micro gas turbine and the power generation cost of the jth energy storage unit at time period t, F t B′-S′The cost of purchasing and selling electricity from the main grid to the distribution network. Let t be the cost of purchasing and selling electricity from the distribution network to microgrid i during time period t. This is the penalty function for exceeding the limits of voltage and branch power flow constraints in the distribution network.
[0224] In one possible implementation, when employing the hybrid multi-agent soft actor critic algorithm, based on a pre-trained policy network, and solving the preset distribution-microgrid collaborative optimization model according to the state observation data of the distribution network agents and each microgrid agent, the distribution-microgrid collaborative optimization model adopts a two-layer architecture of the distribution network layer and the microgrid layer. Specifically, the distribution network optimization scheduling model of the distribution network layer is configured as a discrete-continuous agent containing a Critic network, a discrete Actor network, and a continuous Actor network, and the Critic network guides the updates of the discrete and continuous Actor networks. The microgrid optimization scheduling model of each microgrid in the microgrid layer is configured to contain dual Critic networks. The model employs continuous agents in both ic networks and continuous Actor networks, with dual Critic networks guiding the updates of the continuous Actor network. The Critic network of the microgrid optimal scheduling model updates parameters by minimizing Bellman residuals, while both the Actor network and temperature parameters use policy gradients that introduce maximum entropy to update parameters. The Critic network of the distribution network optimal scheduling model updates parameters through a centralized joint regression loss function. Furthermore, it introduces OU noise into the continuous Actor network of the distribution network optimal scheduling model and a multi-head attention mechanism in the Critic network of the distribution network optimal scheduling model to selectively focus on information from each microgrid optimal scheduling model.
[0225] All relevant content of each step involved in the aforementioned embodiments of the distribution microgrid collaborative optimization scheduling method can be referenced from the functional description of the corresponding functional module of the distribution microgrid collaborative optimization scheduling system in the embodiments of the present invention, and will not be repeated here.
[0226] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional modules in the various embodiments of the invention can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0227] In still another embodiment of the present application, a computer device is provided, which comprises a processor and a memory, the memory being configured to store a computer program, the computer program comprising program instructions, and the processor being configured to execute the program instructions stored in the computer storage medium. The processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc., which are the computing core and control core of the terminal, and are suitable for implementing one or more instructions, and are particularly suitable for loading and executing one or more instructions in the computer storage medium to implement a corresponding method flow or a corresponding function; the processor in the embodiments of the present application can be used for the operation of the micro-grid collaborative optimization scheduling method.
[0228] In still another embodiment of the present application, the present application further provides a storage medium, specifically a computer readable storage medium (Memory), which is a memory device in the computer device, and is configured to store programs and data. It can be understood that the computer readable storage medium herein can include an internal storage medium in the computer device, and of course can also include an expansion storage medium supported by the computer device. The computer readable storage medium provides a storage space, and the storage space stores an operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor are also stored in the storage space, and the instructions can be one or more computer programs (including program codes). It should be noted that the computer readable storage medium herein can be a high-speed RAM memory, or a non-volatile memory such as at least one disk memory. One or more instructions stored in the computer readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the micro-grid collaborative optimization scheduling method in the above embodiments.
[0229] Those skilled in the art will appreciate that embodiments of the application can be devised for a method, a system, or a computer program product. Accordingly, the present application can be embodied in the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.
[0230] The present application is described in reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.
[0231] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.
[0232] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in the flowchart illustrations and / or block diagrams.
[0233] Finally, it should be noted that the above-mentioned embodiments are merely intended for describing and illustrating, not limiting the technical solutions of the present application. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the present application, and any modifications or equivalent replacements without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.
Claims
1. A method for coordinated optimization scheduling of distribution microgrids, characterized in that, include: Acquire state observation data of distribution network intelligent agents and microgrid intelligent agents; A hybrid multi-agent soft actor critic algorithm is adopted. Based on a pre-trained policy network, the algorithm solves the preset distribution-microgrid collaborative optimization model according to the state observation data of the distribution network agent and each microgrid agent, and obtains the control strategy of the distribution network and each microgrid. The pre-trained policy network is obtained in the following way: The collaborative optimization model of the distribution microgrid is designed as a reinforcement learning model in a multi-agent environment, and the policy network of the reinforcement learning model is pre-trained based on the generative adversarial imitation learning method to obtain the pre-trained policy network. The microgrid collaborative optimization model includes: The optimal scheduling models for each microgrid are constructed with the goal of minimizing total production cost, and the optimal scheduling model for the distribution network is constructed with the goal of minimizing total operating cost. The collaborative optimization model for distribution microgrids adopts a two-layer architecture of distribution network layer and microgrid layer; Among them, the distribution network optimization scheduling model of the distribution network layer is set as a discrete-continuous agent containing Critic network, discrete Actor network and continuous Actor network, and the Critic network is used to guide the update of discrete Actor network and continuous Actor network; the microgrid optimization scheduling model of each microgrid layer is set as a continuous agent containing dual Critic network and continuous Actor network, and the dual Critic network is used to guide the update of continuous Actor network. The Critic network of the microgrid optimal scheduling model updates parameters by minimizing the Bellman residual, while the Actor network and temperature parameters are updated using a policy gradient that introduces maximum entropy. The Critic network of the distribution network optimal scheduling model updates parameters through a centralized joint regression loss function, and introduces OU noise into the continuous Actor network of the distribution network optimal scheduling model, as well as a multi-head attention mechanism in the Critic network of the distribution network optimal scheduling model for selectively focusing on information from each microgrid optimal scheduling model.
2. The microgrid collaborative optimization scheduling method according to claim 1, characterized in that, The state observation data of the power distribution network intelligent agent includes: The data includes observation data of static var compensators, renewable energy units, micro gas turbines, energy storage systems, on-load tap changers, capacitor banks, the number of times on-load tap changers and capacitor banks have been activated, the active load of the distribution network with superimposed normal distribution error, the set of electricity purchase and sale prices of the distribution network, and boundary information from each microgrid agent. The state observation data of the microgrid intelligent agent includes: observation data of renewable energy units in the microgrid, observation data of micro gas turbines and energy storage systems, as well as active load of the microgrid superimposed with normal distribution error and boundary information from the distribution network intelligent agent.
3. The microgrid collaborative optimization scheduling method according to claim 1, characterized in that, The distribution network intelligent agent is set up in the following way: the distribution network is virtually divided into several autonomous regions, a continuous intelligent agent is set up for each autonomous region, and a separate intelligent agent is set up for the discrete devices with multiple levels that are independent of each other in the distribution network, and they are combined to obtain the distribution network intelligent agent; the microgrid intelligent agent is set up in the following way: an autonomous continuous intelligent agent is set up for each microgrid to obtain each microgrid intelligent agent.
4. The microgrid collaborative optimization scheduling method according to claim 1, characterized in that, The objective function of the microgrid optimal scheduling model is: in, For microgrids Total production cost; For scheduling duration; For microgrids The number of micro and small gas turbines; for Time period Cost of generating electricity using a micro gas turbine; For microgrids The number of medium-sized energy storage units; for Time period The cost of electricity generation from Taiwan's energy storage units; for Time-of-use microgrids Cost of buying and selling electricity from the distribution network; The constraints of the microgrid optimal scheduling model include: Constraints on the output of renewable energy units, the output of micro gas turbines, the constraints on energy storage systems, the power flow balance within microgrids, and the transmission power constraints of distribution microgrids; The objective function of the power distribution network optimization scheduling model is: in, The total operating cost of the distribution network; , , and These refer to the number of micro gas turbines, energy storage systems, switched capacitor banks, and microgrids in the power distribution network; and They are respectively In the time-sharing distribution network, the first The cost of generating electricity using a miniature gas turbine and energy storage unit; and They are respectively The distribution network has on-load tap-changing transformers and the first Operating costs of switching capacitor banks from Taiwan; for Time-of-use microgrids The cost of buying and selling electricity from the distribution network for The cost of purchasing and selling electricity from the distribution network to the main grid during a given time period; The constraints of the power distribution network optimization scheduling model include: Constraints include the number of on-load tap-changing transformer operations, the number of capacitor bank switching operations, the output constraint of static var compensators, the power flow balance constraint of the distribution network, the node voltage constraint, and the current carrying capacity constraint.
5. The microgrid collaborative optimization scheduling method according to claim 1, characterized in that, The design of the microgrid collaborative optimization model as a reinforcement learning model in a multi-agent environment includes: Both the microgrid optimal scheduling model and the distribution network optimal scheduling model are designed as reinforcement learning models according to the observation space, action space, and reward function. Observation space of reinforcement learning model for microgrid optimal scheduling for: in, and They are respectively Time-of-use microgrids Observational data of energy storage systems and micro gas turbines, and They are respectively Microgrids with time-varying errors superimposed on normal distribution Observational data and active load of renewable energy units; for -1 time period comes from the boundary information of the distribution network intelligent agent; Action space of reinforcement learning model for microgrid optimal scheduling for: in, for Time-of-use microgrids Reactive power prediction values of renewable energy units in China and They are respectively Time-of-use microgrids The active and reactive power of micro and medium-sized gas turbines for Time-of-use microgrids Active power of the medium-sized energy storage system during charging and discharging; Reward function of reinforcement learning model for microgrid optimal scheduling for: in, and They are respectively Time period The cost of generating electricity from a micro gas turbine and the first The cost of electricity generation from Taiwan's energy storage units; for Time-of-use microgrids Cost of buying and selling electricity from the distribution network; To provide a penalty function for exceeding the power transmission constraint between microgrids; For microgrids Willingness coefficient to support the safe operation of the power distribution network; The function for constraining and penalizing limits for safe operation of the distribution network; The observation space of the reinforcement learning model of the distribution network optimization scheduling model for: in, and These are observation data for on-load tap-changing transformers and observation data for switching capacitor banks, respectively. , and microgrids Observational data of static var compensators, energy storage systems, and micro gas turbines. and They are respectively Microgrids with time-varying errors superimposed on normal distribution Observational data and active load of renewable energy units; for Time-based distribution network electricity purchase and sale prices aggregated; For microgrids Boundary information of microgrid smart agents; The action space of the reinforcement learning model for the distribution network optimization scheduling model is: in, and For distribution network area Discrete action space and continuous action space in and They are respectively Time-of-use distribution network area On-load tap-changing transformer tap positions and capacitor bank switching tap positions; for Time-of-use microgrids Reactive power prediction values of renewable energy units and They are respectively Time-of-use microgrids The active and reactive power of the micro gas turbine. for Time-of-use microgrids The active power of the energy storage system during charging and discharging. for Time-of-use microgrids The reactive power output of the static var compensator; Reward function of reinforcement learning model for distribution network optimization scheduling for: in, and They are respectively Operating costs of on-load tap-changing transformers and the first The operating cost of switching capacitor banks in Taiwan. , , and These refer to the number of micro gas turbines, energy storage systems, switched capacitor banks, and microgrids in the power distribution network; and for Time period The cost of generating electricity from a micro gas turbine and the first The cost of electricity generation from Taiwan's energy storage units The cost of purchasing and selling electricity from the main grid to the distribution network. for Time-of-use microgrids The cost of buying and selling electricity from the distribution network This is the penalty function for exceeding the limits of distribution network voltage and branch power flow constraints; Cost of continuous global device actions.
6. A microgrid collaborative optimization scheduling system, characterized in that, include: The data acquisition module is used to acquire state observation data of distribution network intelligent agents and microgrid intelligent agents; The strategy generation module is used to solve the preset distribution-microgrid collaborative optimization model based on the state observation data of the distribution network agent and each microgrid agent by employing the hybrid multi-agent soft actor critic algorithm and a pre-trained policy network, thereby obtaining the control strategy of the distribution network and each microgrid. The pre-trained policy network is obtained in the following way: The collaborative optimization model of the distribution microgrid is designed as a reinforcement learning model in a multi-agent environment, and the policy network of the reinforcement learning model is pre-trained based on the generative adversarial imitation learning method to obtain the pre-trained policy network. The microgrid collaborative optimization model includes: The optimal scheduling models for each microgrid are constructed with the goal of minimizing total production cost, and the optimal scheduling model for the distribution network is constructed with the goal of minimizing total operating cost. The collaborative optimization model for distribution microgrids adopts a two-layer architecture of distribution network layer and microgrid layer; Among them, the distribution network optimization scheduling model of the distribution network layer is set as a discrete-continuous agent containing Critic network, discrete Actor network and continuous Actor network, and the Critic network is used to guide the update of discrete Actor network and continuous Actor network; the microgrid optimization scheduling model of each microgrid layer is set as a continuous agent containing dual Critic network and continuous Actor network, and the dual Critic network is used to guide the update of continuous Actor network. The Critic network of the microgrid optimal scheduling model updates parameters by minimizing the Bellman residual, while the Actor network and temperature parameters are updated using a policy gradient that introduces maximum entropy. The Critic network of the distribution network optimal scheduling model updates parameters through a centralized joint regression loss function, and introduces OU noise into the continuous Actor network of the distribution network optimal scheduling model, as well as a multi-head attention mechanism in the Critic network of the distribution network optimal scheduling model for selectively focusing on information from each microgrid optimal scheduling model.
7. The microgrid collaborative optimization scheduling system according to claim 6, characterized in that, The state observation data of the power distribution network intelligent agent includes: The data includes observation data of static var compensators, renewable energy units, micro gas turbines, energy storage systems, on-load tap changers, capacitor banks, the number of times on-load tap changers and capacitor banks have been activated, the active load of the distribution network with superimposed normal distribution error, the set of electricity purchase and sale prices of the distribution network, and boundary information from each microgrid agent. The state observation data of the microgrid intelligent agent includes: observation data of renewable energy units in the microgrid, observation data of micro gas turbines and energy storage systems, as well as active load of the microgrid superimposed with normal distribution error and boundary information from the distribution network intelligent agent.
8. The distribution microgrid collaborative optimization scheduling system according to claim 6, characterized in that the distribution network intelligent agent is set up in the following manner: the distribution network is virtually divided into several autonomous regions, a continuous intelligent agent is set up for each autonomous region, and a separate intelligent agent is set up for the discrete devices in the distribution network that are independent of each other and have multiple levels, and the intelligent agents are combined to obtain the distribution network intelligent agent; the microgrid intelligent agent is set up in the following manner: an autonomous continuous intelligent agent is set up for each microgrid to obtain each microgrid intelligent agent.
9. The microgrid collaborative optimization scheduling system according to claim 6, characterized in that, The objective function of the microgrid optimal scheduling model is: in, For microgrids Total production cost; For scheduling duration; For microgrids The number of micro and small gas turbines; for Time period Cost of generating electricity using a micro gas turbine; For microgrids The number of medium-sized energy storage units; for Time period The cost of electricity generation from Taiwan's energy storage units; for Time-of-use microgrids Cost of buying and selling electricity from the distribution network; The constraints of the microgrid optimal scheduling model include: Constraints on the output of renewable energy units, the output of micro gas turbines, the constraints on energy storage systems, the power flow balance within microgrids, and the transmission power constraints of distribution microgrids; The objective function of the power distribution network optimization scheduling model is: in, The total operating cost of the distribution network; , , and These refer to the number of micro gas turbines, energy storage systems, switched capacitor banks, and microgrids in the power distribution network; and They are respectively In the time-sharing distribution network, the first The cost of generating electricity using a miniature gas turbine and energy storage unit; and They are respectively The distribution network has on-load tap-changing transformers and the first Operating costs of switching capacitor banks from Taiwan; for Time-of-use microgrids The cost of buying and selling electricity from the distribution network for The cost of purchasing and selling electricity from the distribution network to the main grid during a given time period; The constraints of the power distribution network optimization scheduling model include: Constraints include the number of on-load tap-changing transformer operations, the number of capacitor bank switching operations, the output constraint of static var compensators, the power flow balance constraint of the distribution network, the node voltage constraint, and the current carrying capacity constraint.
10. The microgrid collaborative optimization scheduling system according to claim 6, characterized in that, The design of the microgrid collaborative optimization model as a reinforcement learning model in a multi-agent environment includes: Both the microgrid optimal scheduling model and the distribution network optimal scheduling model are designed as reinforcement learning models according to the observation space, action space, and reward function. Observation space of reinforcement learning model for microgrid optimal scheduling for: in, and They are respectively Time-of-use microgrids Observational data of energy storage systems and micro gas turbines, and They are respectively Microgrids with time-varying errors superimposed on normal distribution Observational data and active load of renewable energy units; for -1 time period comes from the boundary information of the distribution network intelligent agent; Action space of reinforcement learning model for microgrid optimal scheduling for: in, for Time-of-use microgrids Reactive power prediction values of renewable energy units in China and They are respectively Time-of-use microgrids The active and reactive power of micro and medium-sized gas turbines for Time-of-use microgrids Active power of the medium-sized energy storage system during charging and discharging; Reward function of reinforcement learning model for microgrid optimal scheduling for: in, and They are respectively Time period The cost of generating electricity from a micro gas turbine and the first The cost of electricity generation from Taiwan's energy storage units; for Time-of-use microgrids Cost of buying and selling electricity from the distribution network; To provide a penalty function for exceeding the power transmission constraint between microgrids; For microgrids Willingness coefficient to support the safe operation of the power distribution network; The function for constraining and penalizing limits for safe operation of the distribution network; The observation space of the reinforcement learning model of the distribution network optimization scheduling model for: in, and These are observation data for on-load tap-changing transformers and observation data for switching capacitor banks, respectively. , and microgrids Observational data of static var compensators, energy storage systems, and micro gas turbines. and They are respectively Microgrids with time-varying errors superimposed on normal distribution Observational data and active load of renewable energy units; for Time-based distribution network electricity purchase and sale prices aggregated; For microgrids Boundary information of microgrid smart agents; The action space of the reinforcement learning model for the distribution network optimization scheduling model is: in, and For distribution network area Discrete action space and continuous action space in and They are respectively Time-of-use distribution network area On-load tap-changing transformer tap positions and capacitor bank switching tap positions; for Time-of-use microgrids Reactive power prediction values of renewable energy units and They are respectively Time-of-use microgrids The active and reactive power of the micro gas turbine. for Time-of-use microgrids The active power of the energy storage system during charging and discharging. for Time-of-use microgrids The reactive power output of the static var compensator; Reward function of reinforcement learning model for distribution network optimization scheduling for: in, and They are respectively Operating costs of on-load tap-changing transformers and the first The operating cost of switching capacitor banks in Taiwan. , , and These refer to the number of micro gas turbines, energy storage systems, switched capacitor banks, and microgrids in the power distribution network; and for Time period The cost of generating electricity from a micro gas turbine and the first The cost of electricity generation from Taiwan's energy storage units The cost of purchasing and selling electricity from the main grid to the distribution network. for Time-of-use microgrids The cost of buying and selling electricity from the distribution network This is the penalty function for exceeding the limits of distribution network voltage and branch power flow constraints; Cost of continuous global device actions.
11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the microgrid collaborative optimization scheduling method as described in any one of claims 1 to 5.
12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the microgrid collaborative optimization scheduling method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Power distribution network-microgrid group master-slave game optimization scheduling method based on multi-agent reinforcement learning algorithm
CN118611067A