Micro-grid group optimal operation strategy generation method, system, device and storage medium
By employing a deep deterministic policy gradient algorithm with multiple edge agents, the problems of excessive computational and communication burdens and information security in the optimized operation of microgrids are solved. This enables distributed energy management and user privacy protection, thereby improving the operational efficiency and economic benefits of microgrids.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2026-03-27
AI Technical Summary
Existing research on microgrid cluster optimization and operation suffers from problems such as excessive computational and communication burdens in centralized optimization methods, high risk of single point of failure, and difficulty in ensuring information security. In particular, it is difficult to achieve effective energy management and user privacy protection under diversified microgrid entities.
Employing a multi-edge agent deep deterministic policy gradient algorithm, this algorithm generates microgrid cluster optimization operation strategies through energy auction mutual assistance and cooperative game among edge agents, thereby achieving distributed energy management and collaborative optimization while protecting user privacy.
It enables optimized energy management under incomplete information, protects user privacy and security, and improves the overall operating efficiency and economic benefits of microgrids.
Smart Images

Figure CN114977160B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of optimal operation of microgrid groups, and in particular to a microgrid group optimal operation strategy generation method, system, device and storage medium. BACKGROUND
[0002] Microgrids solve a series of problems such as distributed new energy consumption and grid connection, load optimization in a reliable, efficient and economical way, but a single microgrid has limited capacity for new energy consumption and load integration optimization due to its small size, so research on microgrid groups has become a hot topic. Compared with a single microgrid, a microgrid group is more flexible and reliable, and has more advantages. First, since the microgrid group is composed of microgrids in close geographical proximity, energy can be shared between microgrids to improve new energy utilization and reduce the cost of traditional generator power generation and carbon emissions. Second, when a fault occurs in a certain area of the distribution network, the microgrid at that location can be isolated from the fault by splitting and providing auxiliary services through other microgrids. Finally, energy transfer and coordination can be more effective, and there is great potential to achieve load optimization and peak shaving, which helps to improve the stability and reliability of the power grid. Although the microgrid group system has great technical advantages, it is complex in structure, so research on the optimal operation of microgrid groups has become the focus of scholars.
[0003] Currently, the research on the optimal operation of microgrid groups mainly adopts a centralized optimization approach considering a single subject. The single centralized approach has problems such as excessive burden on the convergence point, difficulty in bearing high computational and communication burden for the optimization scheme, easy occurrence of single point failure, difficulty in protecting the information security of each subject, etc. Moreover, existing optimization methods are modeled and solved on the premise of complete mastery of internal parameter information of each microgrid. With the diversification of microgrid subjects, the emergence of a large number of information islands and the increasing demand for user privacy, internal parameter information of each microgrid is often difficult to obtain and share. SUMMARY
[0004] In order to solve the problem of household data privacy in the prior art, the purpose of the present application is to provide a microgrid group optimal operation strategy generation method, system, device and storage medium, which is based on a multi-edge intelligent agent deep deterministic policy gradient algorithm and can achieve distributed energy management and collaborative optimal operation of microgrid groups under incomplete information on the basis of protecting user privacy.
[0005] To achieve the above purpose, the technical scheme is adopted as follows:
[0006] A microgrid group optimal operation strategy generation method, comprising:
[0007] Based on the single microgrid optimization scheduling model, the unbalanced power status of each edge agent is obtained, and the unbalanced power status of adjacent edge agents in the microgrid group and the market clearing information are used as internal unbalanced power information.
[0008] The internal unbalanced power information is input into the pre-trained microgrid group optimization operation strategy generation model to generate the first microgrid group optimization operation strategy.
[0009] Based on the first microgrid group optimization operation strategy, a second microgrid group optimization operation strategy is generated through a cooperative game of energy auction and mutual assistance among edge agents in the microgrid group.
[0010] As a further improvement of the present invention, the edge agent is divided into micro-networks of the micro-network group, and each edge agent interacts with its neighboring edge agents.
[0011] As a further improvement of the present invention, the step of obtaining the unbalanced power state of each edge agent based on the single microgrid optimized scheduling model includes:
[0012] The unbalanced power status of adjacent edge agents in the microgrid group is obtained. The unbalanced power status is combined with the known distributed renewable energy output, total user load and flexible load in the microgrid to construct a single microgrid optimization scheduling model.
[0013] Solving the single microgrid optimization scheduling model yields the optimization strategy for user demand side of a single microgrid and information on excess or missing power in the microgrid.
[0014] As a further improvement of the present invention, after obtaining the optimization strategy for user demand side and the information on excess or missing power of microgrids, the model gradient information is obtained by encapsulating each microgrid separately.
[0015] As a further improvement of the present invention, when there are microgrids in the microgrid group with a new energy ratio higher than a preset threshold and / or microgrids with large load demand, the power transmitted to the outside by the output microgrid is obtained by using a single microgrid optimization scheduling model, and the output microgrids are used to price the power transmitted to the outside.
[0016] As a further improvement of the present invention, the microgrid group optimization operation strategy generation model is pre-trained, and the specific method is as follows:
[0017] The internal unbalanced power information of each edge agent is obtained, and the model gradient information encapsulated in each microgrid is exchanged.
[0018] Based on the optimization strategy and model gradient information, a microgrid cluster optimization operation strategy generation model is obtained by training the model using a deep reinforcement learning algorithm.
[0019] The theoretical global optimal value is obtained based on the strategy generation model and used as the optimization strategy for the first microgrid group; then the theoretical global optimal value is sent back to the edge agents of each microgrid.
[0020] As a further improvement of the present invention, the microgrid group optimization operation strategy generation model includes:
[0021] State Space: A 4-dimensional state space is established for each edge agent, including the current time period t, the electricity price p sold by the power grid company, and the current time period t. grid,s The electricity purchase price p of the power grid company grid,b Electricity p used for trading trans , represented as s i,t ={t、p grid,s p grid,b p trans};
[0022] Action Space: A 2D action space is established for each edge agent, including the microgrid operator setting the electricity price η and trading electricity p with the grid company. ex , represented as a i,t ={η、p ex};
[0023] Reward function: It consists of two parts, namely the profit function. Punishment for abandoning wind and light The reward function is:
[0024]
[0025]
[0026]
[0027] In the formula: p w To actually absorb wind power, p w,f To predict wind power output, p PV To actually absorb photovoltaic power, p PV,f To predict photovoltaic power.
[0028] As a further improvement of the present invention, the training method for the microgrid cluster optimization operation strategy generation model includes:
[0029] Obtain the state information of adjacent edge agents in the microgrid group, input the state information into the single microgrid optimization scheduling model, and output the decision of each microgrid edge agent at the current time;
[0030] The decision of each microgrid edge agent at the current moment is sent to each microgrid edge agent to execute action instructions;
[0031] After each micro-grid edge agent performs an action, based on market clearing requirements, the reward return and next period state information of each micro-grid edge agent are obtained;
[0032] The state information of each micro-grid is stored in an experience replay pool, and it is judged whether the training round number reaches the maximum training threshold, if it reaches, the training is exited, if it does not reach, the next round of training is continued;
[0033] A batch of samples are taken out from the experience pool, and the model is trained by using the stochastic gradient descent method, and the policy network and value network of each micro-grid edge agent are updated;
[0034] It is judged whether the maximum training round number M set is reached, if it is met, the training is ended, if it is not met, the next round of network parameter update is returned.
[0035] As a further improvement of the application, the first micro-grid group optimization operation strategy is generated by a cooperative game mode of energy auction mutual aid interaction between edge agents in the micro-grid group, specifically including:
[0036] Task allocation and communication cooperation are performed for all edge agents in the micro-grid group, and each micro-grid edge agent is divided into two sub-task edge agents, one type of output micro-grid edge agent is responsible for pricing the excess energy and completing complete consumption of new energy, and the other type of injection micro-grid edge agent purchases the required energy according to the energy transmission network loss to complete the micro-grid source load matching;
[0037] Each micro-grid edge agent exchanges and cooperates through an information sharing mechanism, each edge agent obtains the excess energy or missing energy and pricing strategy of adjacent edge agents, and a second micro-grid group optimization operation strategy is generated by a cooperative game mode of energy auction mutual aid interaction between edge agents in the micro-grid group.
[0038] A micro-grid group optimization operation strategy generation system, comprising:
[0039] An optimization module is used for obtaining the unbalanced power state of each edge agent based on a single micro-grid optimization scheduling model, and taking the unbalanced power state of adjacent edge agents in the micro-grid group and market clearing information as internal imbalance power information;
[0040] A generation module is used for inputting the internal imbalance power information into a pre-trained micro-grid group optimization operation strategy generation model to generate a first micro-grid group optimization operation strategy;
[0041] The control module is used for generating a second micro-grid group optimal operation strategy based on the micro-grid group optimal operation strategy through the cooperative game mode of energy auction interaction between edge intelligent agents in the micro-grid group, and realizing overall consumption of distributed new energy and maximization of overall economic benefits of the micro-grid group and each subject.
[0042] A distributed energy management system comprises:
[0043] The micro-grid group end comprises the micro-grid group optimal operation strategy generation system.
[0044] The cloud end interacts with the micro-grid group end, and is used for training the micro-grid group optimal operation strategy generation model based on internal unbalanced power information obtained by the single micro-grid optimal scheduling model, and delivering the micro-grid group optimal operation strategy generation model to the micro-grid group end.
[0045] An electronic device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the micro-grid group optimal operation strategy generation method when executing the computer program.
[0046] A computer readable storage medium stores a computer program, and the computer program implements the steps of the micro-grid group optimal operation strategy generation method when executed by a processor.
[0047] Compared with the prior art, the present application has the following beneficial effects:
[0048] The method generates a micro-grid group optimal operation strategy based on a multi-edge intelligent agent deep deterministic policy gradient method, and compared with the existing micro-grid group optimal operation method, the trained optimal model can realize energy optimization management under incomplete information and protect the privacy and security of users. The method realizes a decentralized distributed energy management architecture, and simultaneously considers incomplete information and the data privacy and security of users. BRIEF DESCRIPTION OF DRAWINGS
[0049] Figure 1 A micro-grid group optimal operation strategy generation method flowchart of the present application;
[0050] Figure 2 A distributed energy management system schematic diagram of the present application;
[0051] Figure 3 A micro-grid group collaborative optimization flowchart framework schematic diagram of the present application;
[0052] Figure 4 A micro-grid group optimal operation strategy model training schematic diagram of the present application;
[0053] Figure 5A multi-edge-end intelligent agent deep reinforcement learning model training flowchart of the application;
[0054] Figure 6 A microgrid group optimal operation strategy generation system block diagram of the application;
[0055] Figure 7 An electronic device schematic diagram of the application. DETAILED DESCRIPTION
[0056] In order for those skilled in the art to better understand the application scheme, the technical solutions in the embodiments of the application will be described clearly and completely below in conjunction with the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the application.
[0057] It should be noted that the terms "first", "second", etc. in the specification and claims of the application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0058] At present, most researches on energy management and distribution are based on a centralized management unit at the upper layer, and the electricity price is formulated through the way of game between the upper layer distribution network or operator and the lower layer microgrid, such as the double-layer Stackelberg game model of power supply company, microgrid operator and user in the paper "Micro Energy Network Energy Management Optimization Based on Double-layer Stackelberg Game", which transmits all microgrid and user information to the leader power supply company to conduct game under complete information (including the number of photovoltaic wind machines and output power in each microgrid, the number of users and user daily electricity data, the number of nodes and node power, voltage, etc.). The patent "Optimization method of multi-layer subject transaction strategy of microgrid group" needs to upload the energy purchasing price and load demand response information in each microgrid to the microgrid group operator completely. In the research of such problems, centralized, complete information and user data privacy are not considered in the collaborative optimization of microgrid group.
[0059] To further adapt to the scene of micro-grid group collaborative operation with a large number of subjects and the need to consider user privacy, a multi-edge-end intelligent agent distributed algorithm is used to carry out related research on the optimal operation and energy management of the micro-grid group, and the construction of a distributed control mode of edge-cloud collaboration and the communication mechanism between multi-edge-end intelligent agents directly affect the effect of the final optimization strategy.
[0060] Embodiment 1
[0061] As shown in Figure 1 The present application provides a micro-grid group optimal operation strategy generation method, comprising:
[0062] S100, based on the single micro-grid optimal scheduling model, obtaining the unbalanced power state of each edge-end intelligent agent, and taking the unbalanced power state of adjacent edge-end intelligent agents in the micro-grid group and the market clearing information as internal unbalanced power information;
[0063] S200, inputting the internal unbalanced power information into the pre-trained micro-grid group optimal operation strategy generation model to generate a first micro-grid group optimal operation strategy;
[0064] S300, based on the micro-grid group optimal operation strategy, generating a second micro-grid group optimal operation strategy through the cooperative game mode of energy auction mutual aid interaction between edge-end intelligent agents in the micro-grid group.
[0065] The present application provides a micro-grid group collaborative optimization operation strategy based on multi-edge-end intelligent agent deep deterministic policy gradient, realizes a decentralized distributed energy management architecture, and simultaneously considers the non-complete information condition and the data privacy and security of users.
[0066] Embodiment 2
[0067] The present application provides a micro-grid group optimal operation strategy generation method, applied to the micro-grid group end, in detail comprising:
[0068] S100, based on the single micro-grid optimal scheduling model, obtaining the unbalanced power state of each edge-end intelligent agent, and taking the unbalanced power state of adjacent edge-end intelligent agents in the micro-grid group and the market clearing information as internal unbalanced power information;
[0069] S200, the edge-end intelligent agent inputs the optimal strategy into the pre-trained micro-grid group optimal operation strategy generation model, and generates a first micro-grid group optimal operation strategy in combination with the theoretical global optimal value control strategy issued by the cloud;
[0070] S300, the edge end agent generates a second micro-grid group optimal operation strategy based on the first micro-grid group optimal operation strategy through a cooperative game mode of energy auction interaction between edge end agents in the micro-grid group, and realizes overall consumption of distributed new energy and maximization of overall economic benefits of the micro-grid group and each subject.
[0071] In an optional embodiment of the present application, in S100, the edge end agent is divided from each micro-grid of the micro-grid group, and each edge end agent interacts information with adjacent edge end agents.
[0072] In an optional embodiment of the present application, the unbalanced power state of each edge end agent based on the single micro-grid optimal scheduling model comprises:
[0073] The state and action of adjacent edge end agents in the micro-grid group are obtained, and a single micro-grid optimal scheduling model is constructed according to the distributed new energy output condition in the known micro-grid, the total user load and the flexible load amount.
[0074] The single micro-grid optimal scheduling model is solved by using a CPLEX solver, and the optimal strategy of the user demand side of a single micro-grid and the information of excess or missing power of the micro-grid are obtained, the optimal strategy including the lowest micro-grid operation cost, the new energy consumption rate and the user satisfaction.
[0075] In an optional embodiment of the present application, in S200, the micro-grid group optimal operation strategy generation model is obtained by pre-training in the cloud, and the cloud training process comprises the following steps:
[0076] The internal unbalanced power information of each edge end agent is obtained, and the model gradient information of each micro-grid encapsulation is interacted.
[0077] According to the optimal strategy and the model gradient information, a micro-grid group optimal operation strategy generation model is trained by using a deep reinforcement learning algorithm model.
[0078] The cloud obtains a theoretical global optimal value according to the strategy generation model, as the first micro-grid group optimal operation strategy, and the theoretical global optimal value is returned to each micro-grid edge end agent.
[0079] As an optional scheme of the present application, the micro-grid group optimal operation strategy generation model can be trained in other subjects, and the present application gives cloud training, which is not limited to the cloud.
[0080] Embodiment 3
[0081] The present application will be further described in detail below with reference to the accompanying drawings and in combination with examples.
[0082] Figure 2As shown in the schematic diagram of the distributed energy management system, each microgrid edge-end agent can transmit power flow and information flow between adjacent agents, and the edge-end agent can transmit information of the agents in the region to the cloud. The application provides a distributed energy management system for optimizing operation of a microgrid group, as shown in Figure 2 The application mainly comprises:
[0083] 1. Edge-cloud collaborative control mode:
[0084] Each microgrid in the microgrid group is divided into an edge-end agent, each edge-end agent can interact with adjacent edge-end agents, and the state and action of the adjacent edge-end agents are used to optimize the strategy of the edge-end agent.
[0085] The edge-end agents are formed between the microgrids, and the edge-end agents can transmit power flow and information flow between each other. The cloud and each edge-end agent can also transmit information flow.
[0086] The optimized strategy is uploaded to the cloud, the model gradient information of each microgrid is exchanged, the cloud group intelligence system calculates the global optimal value control strategy, and the random gradient descent method is used to return the updated gradient to each microgrid edge-end agent, thereby supporting the edge-cloud collaborative microgrid group optimization operation strategy calculation.
[0087] 2. The collaborative optimization framework of the microgrid group is shown in Figure 3 The optimization process of the optimization strategy is as follows:
[0088] (1) First, the user demand side of the single microgrid edge-end agent is optimized and managed, the output of the known microgrid distributed new energy, the total user load and the flexible load are used to construct a single microgrid optimization scheduling model, including the lowest microgrid operation cost, new energy consumption rate and user satisfaction, the decision variables include transferable load, shiftable load and interruptible load, and the CPLEX solver is used for model solving.
[0089] The single microgrid optimization scheduling model can predict the new energy output and analyze the user flexible load to obtain the new energy data and load data of each microgrid.
[0090] (2) When there are microgrids with high new energy proportion and / or microgrids with high load demand in the microgrid group, the output microgrid can export power to the outside and the injection microgrid needs to purchase power from the outside by using the above model, and the output microgrid can price the power that can be exported to the outside, so that the maximum benefit of each subject and the consumption of new energy in the whole microgrid group can be realized through the game under the condition of incomplete information.
[0091] The single micro-grid is optimized to obtain net energy Xnet, Xnet>0, the injection type micro-grid, Xnet<0, and the output type micro-grid.
[0092] Among them, the high proportion of new energy refers to that the new energy generation capacity can reach more than a preset threshold (for example, twenty percent), such as Xinjiang, Gansu and other provinces.
[0093] 3. Micro-grid group optimization strategy generation based on multi-edge intelligent agent deep deterministic policy gradient, Figure 4 The micro-grid group optimization operation strategy model training schematic diagram, the environment in the model training is the market environment of the micro-grid group transaction, the environment state S k (t) includes electric energy a m (t) that can be used for transaction, grid company price, time period, action space includes pricing strategy of each micro-grid intelligent agent and electric energy quantity traded with the grid, data generated by the model in the training process (s(t), a(t), r(t), s(t+1)) is stored in the experience pool, which is used for updating network parameters. The specific strategy is as follows:
[0094] (1) The micro-grid group optimization operation strategy is generated through the cooperative game mode of energy auction mutual aid interaction between micro-grid edge intelligent agents, realizing the overall consumption of distributed new energy and the maximization of the overall economic benefits of the micro-grid group and each subject.
[0095] For all edge intelligent agents in the micro-grid group, task allocation and exchange cooperation are carried out, and the edge intelligent agents of each micro-grid are divided into two sub-task edge intelligent agents, one type of output type micro-grid edge intelligent agent is responsible for pricing the excess electric energy, and completes the complete consumption of new energy; another type of injection type micro-grid edge intelligent agent purchases the needed electric energy according to the electric energy transmission network loss, and completes the micro-grid source load matching. The edge intelligent agents of each micro-grid exchange and cooperate through the information sharing mechanism, and each edge intelligent agent can know the excess electric energy or missing electric energy and pricing strategy of the adjacent edge intelligent agent, and the second micro-grid group optimization operation strategy is generated through the cooperative game mode of energy auction mutual aid interaction between the edge intelligent agents in the micro-grid group, so as to continuously optimize the own strategy.
[0096] (2) In order to protect the micro-grid user data privacy in the micro-grid group optimization operation, by constructing a single micro-grid optimization scheduling model, the outputtable electric energy and the needed injection electric energy can be directly optimized according to the internal state information of the micro-grid, and are encapsulated separately, without the need to upload all the information of the micro-grid internal nodes, voltage, user load, generator output and new energy output, so as to protect the micro-grid internal information and user data privacy.
[0097] Embodiment 4
[0098] The scheme gives a specific scheme for training of micro-grid group optimal operation strategy generation model based on multi-edge end intelligent agent deep deterministic policy gradient, including but not limited to multi-edge end intelligent agent deep reinforcement learning algorithm. For easy understanding, a training example based on multi-edge end intelligent agent deep reinforcement learning method is given.
[0099] The training of micro-grid group optimal operation strategy generation model based on multi-edge end intelligent agent deep deterministic policy gradient adopts a deep reinforcement learning algorithm model, which is described as follows:
[0100] (1) State space: a 4-dimensional state space is established for each edge end intelligent agent, including the current time period t, the power grid company electricity selling price p grid,s , the power grid company electricity purchasing price p grid,b , and the electricity available for transaction p trans , which can be represented as s i,t ={t, p grid,s , p grid,b , p trans}
[0101] (2) Action space: a 2-dimensional action space is established for each edge end intelligent agent, including the micro-grid operator's electricity price η and the electricity traded with the power grid company p ex , which can be represented as a i,t ={η, p ex}
[0102] (3) Reward function: the reward function is divided into two parts, namely the revenue function and the wind and light punishment The reward function is:
[0103]
[0104]
[0105]
[0106] In the formula: p w is the actual wind power consumption, p w,f is the predicted wind power, p PV is the actual photovoltaic power consumption, and p PV,f is the predicted photovoltaic power.
[0107] Figure 5 is the multi-edge end intelligent agent deep reinforcement learning model training flowchart, first determine the dispatching period T of the micro-grid, set the training round number, initialize the network parameters, etc. Basic parameter settings, assign network parameters. The deep reinforcement learning algorithm model training method is as follows in combination with the drawings:
[0108] Step 1: Input the status information of each microgrid into the single microgrid optimization scheduling model, output the excess or missing power information of the microgrid, and assign the power s1, s2, ... s n Input into the neural network;
[0109] Step 2: Using ReLU as the activation function, the neural network outputs the decisions a1, a2, ... a1 of each micro-network edge agent at the current time step. n ;
[0110] Step 3: Based on the generated decisions, each microgrid edge agent executes action instructions;
[0111] Step 4: After each microgrid edge agent executes its action, determine the reward r1, r2, ..., r for each microgrid edge agent based on market clearing requirements. n and the status information for the next time period s1, s2, ... s n ;
[0112] Step 5: Set {s t ,a t ,r t ,s t+1 Store the episode in the experience replay pool and determine whether the episode has ended. Specifically, check if t ≥ T. If not, the episode has not ended. If the episode has not ended, set t = t + 1 and continue to interact with the environment to complete the next step.
[0113] Step 6: If yes, then end. After a certain amount of experience storage is completed, a batch of samples are taken from the experience pool and the model is trained using stochastic gradient descent to update the policy network and value network of each micro-network edge agent.
[0114] Step 7: Determine whether the set maximum number of training rounds M has been reached, i.e., m≥M. If yes, end the training; if not, set m=m+1. If not, return to the next round of network parameter updates.
[0115] Example 5
[0116] like Figure 6 As shown, the present invention also provides a microgrid group optimization operation strategy generation system, comprising:
[0117] The optimization module is used to obtain the unbalanced power status of each edge agent based on the single microgrid optimization scheduling model, and to use the unbalanced power status of adjacent edge agents in the microgrid group and market clearing information as internal unbalanced power information.
[0118] The generation module is used to input the internal unbalanced power information into the pre-trained microgrid cluster optimization operation strategy generation model to generate the first microgrid cluster optimization operation strategy.
[0119] The control module is used to generate a second microgrid group optimization operation strategy based on the microgrid group optimization operation strategy and through a cooperative game of energy auction and mutual assistance among edge agents in the microgrid group, so as to realize the overall consumption of distributed new energy and maximize the overall economic benefits of the microgrid group and the interests of each entity.
[0120] Example 6
[0121] like Figure 2 As shown, the present invention also provides a distributed energy management system, comprising:
[0122] The microgrid cluster terminal includes the aforementioned microgrid cluster optimized operation strategy generation system;
[0123] The cloud interacts with the microgrid cluster terminal to train the microgrid cluster optimization operation strategy generation model based on the unbalanced power state of each edge agent obtained from the single microgrid optimization scheduling model; and then distributes it to the microgrid cluster terminal.
[0124] Example 7
[0125] like Figure 7 As shown, a third objective of this invention is to provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the microgrid cluster optimization operation strategy generation method.
[0126] The method for generating microgrid group optimization operation strategies includes the following steps:
[0127] Based on the single microgrid optimization scheduling model, the unbalanced power status of each edge agent is obtained, and the unbalanced power status of adjacent edge agents in the microgrid group and the market clearing information are used as internal unbalanced power information.
[0128] The internal unbalanced power information is input into the pre-trained microgrid group optimization operation strategy generation model to generate the first microgrid group optimization operation strategy.
[0129] Based on the aforementioned microgrid cluster optimization operation strategy, a second microgrid cluster optimization operation strategy is generated through a cooperative game of energy auction and mutual assistance among edge agents in the microgrid cluster. This achieves the overall consumption of distributed new energy and maximizes the overall economic benefits of the microgrid cluster and the interests of each entity.
[0130] Example 8
[0131] A fourth objective of this invention is to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the microgrid cluster optimization operation strategy generation method.
[0132] The micro-grid group optimal operation strategy generation method comprises the following steps:
[0133] Based on the single micro-grid optimal scheduling model, the unbalanced power state of each edge agent is obtained, and the unbalanced power state of adjacent edge agents in the micro-grid group and market clearing information are used as internal imbalance power information;
[0134] The internal imbalance power information is input into the pre-trained micro-grid group optimal operation strategy generation model to generate a first micro-grid group optimal operation strategy;
[0135] Based on the micro-grid group optimal operation strategy, a second micro-grid group optimal operation strategy is generated through the cooperative game mode of energy auction mutual aid interaction between edge agents in the micro-grid group, and the overall consumption of distributed new energy and the maximization of the overall economic benefits of the micro-grid group and each subject are realized.
[0136] Those skilled in the art will understand that embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.
[0137] The present application is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device that implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks
[0138] These computer program instructions can also be stored in a computer-readable memory that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer-readable memory produce a product including instruction means, which implements the functions specified in the flowcharts and / or block diagrams. Figure 1 The functions specified in one or more flows and / or blocks Figure 1 The functions specified in one or more flows and / or blocks
[0139] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flowcharts and / or blocks Figure 1 one or more flowcharts and / or blocks
[0140] Finally, it should be noted that the above-mentioned embodiments are merely used to illustrate the technical solutions of the present application, rather than limit the technical solutions of the present application. Although the present application has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that the specific embodiments of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and any modification or equivalent replacement without departing from the spirit and scope of the present application should be covered in the protection scope of the claims of the present application.
Claims
1. A method for generating a microgrid cluster optimization operation strategy, characterized in that, include: The unbalanced power state of each edge agent is obtained based on the single-microgrid optimized scheduling model, and the unbalanced power state of adjacent edge agents in the microgrid group and market clearing information are used as internal unbalanced power information; wherein, the unbalanced power state of each edge agent obtained based on the single-microgrid optimized scheduling model includes: The state and actions of adjacent edge agents in the microgrid cluster are obtained. The state and actions are combined with the known distributed renewable energy output, total user load and flexible load in the microgrid to construct a single microgrid optimization scheduling model. Solve the single microgrid optimization scheduling model to obtain the optimization strategy for user demand side of a single microgrid and the information on excess or missing power of the microgrid, which serves as the unbalanced power state of each edge agent; The internal unbalanced power information is input into the pre-trained microgrid group optimization operation strategy generation model to generate the first microgrid group optimization operation strategy. Based on the first microgrid group optimization operation strategy, a second microgrid group optimization operation strategy is generated through a cooperative game of energy auction and mutual assistance among edge agents in the microgrid group. The microgrid group optimization operation strategy generation model is pre-trained, and the specific method is as follows: The internal unbalanced power information of each edge agent is obtained, and the model gradient information encapsulated in each microgrid is exchanged. Based on the optimization strategy and model gradient information, a microgrid cluster optimization operation strategy generation model is obtained by training the model using a deep reinforcement learning algorithm. The theoretical global optimal value is obtained based on the strategy generation model and used as the optimization operation strategy for the first microgrid group; the theoretical global optimal value is then fed back to the edge agents of each microgrid. The training method for the microgrid group optimization operation strategy generation model includes: Obtain the state information of adjacent edge agents in the microgrid group, input the state information into the single microgrid optimization scheduling model, and output the decision of each microgrid edge agent at the current time; The decision of each microgrid edge agent at the current moment is sent to each microgrid edge agent to execute action instructions; After each microgrid edge agent performs an action, based on market clearing requirements, the reward and status information for the next time period are obtained for each microgrid edge agent. The status information of each micronet is stored in the experience replay pool, and it is determined whether the number of training rounds has reached the maximum training threshold. If it has, training is terminated; otherwise, training continues to the next round. A batch of samples is taken from the experience pool and the model is trained using stochastic gradient descent to update the policy network and value network of each micro-network edge agent; Determine whether the set maximum number of training rounds M has been reached. If so, end the training; otherwise, return to perform the next round of network parameter updates.
2. The method for generating a microgrid group optimization operation strategy according to claim 1, characterized in that, The edge agent is composed of micro-networks in the micro-network group, and each edge agent interacts with its neighboring edge agents.
3. The method for generating a microgrid group optimization operation strategy according to claim 1, characterized in that, After obtaining the optimization strategy for user demand side and the information on excess or missing power of a single microgrid, the model gradient information is obtained by encapsulating each microgrid separately.
4. The method for generating a microgrid group optimization operation strategy according to claim 1, characterized in that, When there are microgrids in the microgrid group with a higher proportion of new energy sources than a preset threshold and / or microgrids with high load demand, the single microgrid optimization scheduling model is used to obtain the power transmitted to the outside by the output microgrid, and the output microgrid is used to price the power transmitted to the outside by itself.
5. The method for generating a microgrid group optimization operation strategy according to claim 1, characterized in that, The microgrid cluster optimization operation strategy generation model includes: State space: A 4-dimensional state space is established for each edge agent, including the current time period. t Electricity sales price of power grid company p grid,s Electricity purchase price by power grid company p grid,b Electricity used for trading p trans , represented as s i , t ={ t、 p grid,s , p grid,b , p trans }; Action Space: A 2D action space was established for each edge agent, including the microgrid operator's electricity pricing mechanism. η Trading electricity with the power grid company p ex , represented as a i , t ={ η, p ex }; Reward function: It consists of two parts, namely the profit function. Punishment for abandoning wind and light The reward function is: (1) η*(p trans - p ex ) + p grid,b* p ex (2) (3) In the formula: To actually absorb wind power, To predict wind power output, To actually absorb photovoltaic power, To predict photovoltaic power.
6. The method for generating a microgrid group optimization operation strategy according to claim 1, characterized in that, The optimized operation strategy based on the first microgrid group utilizes a cooperative game theory approach involving energy auctions and mutual support among edge agents within the microgrid group. Specifically, it includes: Task allocation and communication collaboration are carried out for all edge agents in the microgrid cluster. Each microgrid edge agent is divided into two sub-task edge agents: one type of output-type microgrid edge agent is responsible for pricing excess electricity and completing the full consumption of new energy; the other type of injection-type microgrid edge agent purchases the required electricity according to the power transmission network loss and completes the source-load matching within the microgrid. Each microgrid edge agent communicates and collaborates through an information sharing mechanism. Each edge agent obtains the surplus or shortage of power and pricing strategies from neighboring edge agents. Through a cooperative game of energy auctions and mutual assistance among edge agents in the microgrid cluster, a second microgrid cluster optimization operation strategy is generated.
7. A microgrid cluster optimization operation strategy generation system, implementing the microgrid cluster optimization operation strategy generation method according to any one of claims 1-6, characterized in that, include: The optimization module is used to obtain the unbalanced power status of each edge agent based on the single microgrid optimization scheduling model, and to use the unbalanced power status of adjacent edge agents in the microgrid group and market clearing information as internal unbalanced power information. The generation module is used to input the internal unbalanced power information into the pre-trained microgrid group optimization operation strategy generation model to generate the first microgrid group optimization operation strategy. The control module is used to generate a second microgrid optimization operation strategy based on the microgrid optimization operation strategy and through a cooperative game of energy auction and mutual assistance among edge agents in the microgrid.
8. A distributed energy management system, characterized in that, include: The microgrid cluster terminal includes the microgrid cluster optimization operation strategy generation system as described in claim 7; The cloud interacts with the microgrid cluster terminal to train the microgrid cluster optimization operation strategy generation model based on the internal unbalanced power information obtained from the single microgrid optimization scheduling model. And then distribute it to the microgrid group terminal.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the microgrid cluster optimization operation strategy generation method according to any one of claims 1-6.
10. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the microgrid cluster optimization operation strategy generation method according to any one of claims 1-6.
Citation Information
Patent Citations
Game theory-based multi-micro-grid interconnection running optimization method
CN107545325A
Cloud collaboration-edge collaboration optimization scheduling method for comprehensive energy service provider
CN109993419A