Multi-microgrid distributed energy scheduling method and system based on energy flow information entropy

By introducing a multi-micronet distributed energy scheduling method with energy flow information entropy, the reinforcement learning algorithm is used to minimize the energy flow information entropy, and the communication overhead and network blocking problems of centralized scheduling are solved, and the orderly and efficient energy scheduling of the multi-micronet system is realized.

CN120185109BActive Publication Date: 2025-08-12HOHAI UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510654427.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-12
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing centralized microgrid energy scheduling methods have problems such as large communication overhead, difficulty in achieving near-level power balance, frequent equipment operation, large computing volume and network blockage. The existing distributed control technology has failed to effectively guide the multi-microgrid system to evolve into an orderly state.

Method used

A multi-micronet distributed energy scheduling method based on energy flow information entropy is adopted. By collecting multi-micronet grid structure and operation information, a reinforcement learning network is built to achieve a three-level balance within the micronet, between adjacent micronets and the global micronet groups. The reinforcement learning algorithm is used to minimize the energy flow information entropy, and combined with the energy storage system capacity reward items to reduce network blockage.

Benefits of technology

The orderly improvement of the micronet system is achieved, the frequency of equipment operation is reduced, the local consumption efficiency of distributed energy is improved, and the probability of network blocking is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185109B_ABST
    Figure CN120185109B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-microgrid distributed energy scheduling method and system based on energy flow information entropy. The method comprises: collecting multi-microgrid grid structure information; collecting multi-microgrid operation information; introducing energy flow information entropy to construct a multi-microgrid control target model; constructing a reinforcement learning network that outputs a scheduling strategy based on the multi-microgrid control target model; training the reinforcement learning network; inputting the real-time collected multi-microgrid operation information into the trained reinforcement learning network to output a scheduling strategy; and executing the multi-microgrid system scheduling strategy. The present invention can achieve three-level balance within a microgrid, between adjacent microgrids, and globally across a microgrid group. By minimizing energy flow information entropy, the system's orderliness is improved, the frequency of equipment operation in the multi-microgrid system is reduced, scheduling requirements are reduced, and the local consumption efficiency of distributed energy is improved. The capacity bonus of the energy storage system takes into account the capacity limitations of the transmission channel, reducing the probability of network congestion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power systems and relates to microgrid dispatching and control technology, and specifically to a multi-microgrid distributed energy dispatching method and system based on energy flow information entropy. Background Art

[0002] Currently, the energy dispatch and control model is gradually evolving from traditional centralized control to decentralized control. The centralized dispatch and control architecture relies on the centralized controller to collect and process information from the entire network, including data such as distributed power output and load demand. After global optimization calculations are performed through the centralized energy management system, the dispatch instructions are then issued to each local execution unit. The decentralized dispatch and control model, on the other hand, places greater emphasis on the autonomy and flexibility of the system, giving each control unit greater decision-making autonomy while ensuring the economic operation of the microgrid. In this model, real-time information exchange is maintained between the central controller and the local controllers. Each local controller not only provides feedback on the current operating status and future forecast data to the central controller, but also proactively requests power generation or load adjustment based on local needs.

[0003] Existing centralized microgrid energy scheduling methods rely on high-frequency global information exchange. Traditional algorithms, however, are limited by the high-dimensionality of control variables and non-convex optimization objectives when solving load scheduling problems, making them ineffective. These algorithms face the following challenges: 1. High communication overhead, requiring frequent global data exchange; 2. The addition of new distributed generation sources requires changes to the overall control design; 3. Low energy scheduling efficiency, making it difficult to achieve local power balancing, leading to line losses and network congestion caused by long-distance transmission; 4. Centralized control, coupled with the massive number of devices in multiple microgrids, requires processing a large amount of information, placing a significant computational burden on the central controller.

[0004] While existing technologies have proposed a framework for applying reinforcement learning to distributed microgrid control, each microgrid is considered an intelligent agent, restricted to information exchange only with agents in neighboring microgrids. After training, each agent operates in a distributed manner, enabling the local microgrid to be self-sufficient from distributed generation devices and to trade with grids with power needs, thus achieving local energy autonomy and privacy protection. However, existing microgrid distributed control technologies focus solely on the operating status of each microgrid while ignoring the overall operating status of the multi-microgrid system. This fails to effectively guide the system towards an orderly state, potentially reducing the lifespan of power electronic equipment in the microgrid due to frequent positive and negative power interactions. Furthermore, ignoring the overall needs of the multi-microgrid system can make it difficult to coordinate the output plans of distributed units during the pre-dispatch phase. Summary of the Invention

[0005] Purpose of the invention: In order to overcome the deficiencies in the prior art, a multi-microgrid distributed energy scheduling method and system based on energy flow information entropy is provided, which can achieve three-level balance within the microgrid, between adjacent microgrids, and globally within the microgrid group; by minimizing energy flow information entropy, the orderliness of the system is improved, the operation frequency of equipment in the multi-microgrid system is reduced, the scheduling demand is reduced, and the local consumption efficiency of distributed energy is improved; the transmission channel capacity limitation is taken into account in the energy storage system capacity reward item, thereby reducing the probability of network congestion.

[0006] Technical solution: To achieve the above objectives, the present invention provides a multi-microgrid distributed energy scheduling method based on energy flow information entropy, comprising the following steps:

[0007] S1: Collecting multi-microgrid structure information;

[0008] S2: Based on the multi-microgrid grid structure information, collect multi-microgrid operation information;

[0009] S3: Based on the multi-microgrid operation information, the energy flow information entropy is introduced to construct the multi-microgrid control target model;

[0010] S4: Based on the multi-microgrid control target model, a reinforcement learning network is constructed to output the scheduling strategy;

[0011] S5: Train the reinforcement learning network;

[0012] S6: Input the real-time collected multi-microgrid operation information into the trained reinforcement learning network and output the scheduling strategy;

[0013] S7: Execute the multi-microgrid system scheduling strategy.

[0014] Furthermore, the multi-microgrid grid structure information in step S1 includes: the number n of microgrid nodes in the target multi-microgrid area, the system node impedance matrix Z of the target multi-microgrid area, the electrical distance matrix D between each microgrid node in the target multi-microgrid area, the maximum capacity matrix CAP of the transmission channel between each microgrid node in the target multi-microgrid area, and the maximum energy storage capacity matrix Esc of each microgrid node in the target multi-microgrid area.

[0015] Furthermore, the expression of the system node impedance matrix Z of the target multi-microgrid area in step S1 is:

[0016]

[0017] The element Z in row i and column j in the matrix Z ij is the system equivalent impedance viewed from ports i and j;

[0018] According to formula (1), the electrical distance matrix D between each microgrid node in the target multi-microgrid area is obtained, which is expressed as:

[0019]

[0020] The element D in row i and column j in matrix D ij The meaning of is the electrical distance from microgrid node i to microgrid node j, which is used to characterize the closeness of the electrical connection between the two microgrids. ij =Z ii +Z jj -2Z ij , when i=j, the electrical distance D ii =0;

[0021] The expression of the maximum capacity matrix CAP of the transmission channel between microgrid nodes in the target multi-microgrid area is:

[0022]

[0023] The element C in row i and column j in the matrix CAP ij is the maximum transmission capacity of the transmission channel from microgrid node i to microgrid node j. CAP is a symmetric matrix, namely C ij =C ji ;

[0024] The expression of the maximum energy storage capacity matrix Esc of each microgrid node in the target multi-microgrid area is:

[0025] Esc=[Esc1 … Esc i … Esc n ] (4)

[0026] Among them, Esc i Represents the maximum energy storage capacity of microgrid node i.

[0027] Furthermore, the multi-microgrid operation information in step S2 includes the energy storage capacity state SOC of each microgrid. m , Inter-microgrid interactive electricity Eex m , Multi-microgrid user load demand Pc m , Multi-microgrid distributed power output Pg m , Operational status of transmission channels between multiple microgrids TL m .

[0028] Furthermore, the energy storage state SOC of each microgrid in step S2 is m The expression is as follows:

[0029]

[0030] in, is the percentage of available state of the remaining capacity of energy storage in microgrid i at the time of the mth acquisition;

[0031] Interactive electricity between microgrids Eex m The expression is as follows:

[0032]

[0033] in, is the amount of electricity sent from microgrid i to microgrid j during the mth collection;

[0034] Multi-microgrid user load demand Pc m The expression is as follows:

[0035]

[0036] in, is the user load demand in microgrid i at the time of the mth collection;

[0037] Multi-microgrid distributed power output Pg m The expression is as follows:

[0038]

[0039] in, is the total output of distributed generation in microgrid i at the time of the mth acquisition;

[0040] Operation status of transmission channel between multiple microgrids TL m The expression is as follows:

[0041]

[0042] in, Is the transmission line from microgrid i to microgrid j in the mth data collection direction in operation? If it is in operation, If not in running state

[0043] Furthermore, the construction of the multi-microgrid control target model in step S3 includes:

[0044] The energy interaction target is set to minimize the energy flow information entropy of each microgrid node, that is:

[0045] Min H i (10)

[0046] The expression of energy flow information entropy for microgrid i at the mth acquisition moment is:

[0047]

[0048] Among them, abs(·) is the absolute value function, p ijis the energy flow probability, which indicates that the energy interaction between microgrid i and microgrid j accounts for the total interaction power Ee of microgrid i i The proportion of ε is the logarithmic control parameter.

[0049] Furthermore, in the multi-microgrid control target model of step S3, the power balance constraint is used to ensure the stable operation of the entire system and the matching of source and load. The expression of the power balance constraint is:

[0050]

[0051] Furthermore, the construction of the reinforcement learning network in step S4 includes:

[0052] A1: Constructing the multi-microgrid system state space S k , the mathematical expression is as follows:

[0053] S k ={SOC k ,Eex k ,Pc k ,Pg k ,TL k} (15)

[0054] Among them, k represents the iteration round, SOC k Represents the microgrid battery state set, SOC k ={soc 1,k ,…,soc i,k ,…,soc n,k}, soc i,k is the percentage of available remaining capacity of energy storage in microgrid i at the kth iteration; Eex k Represents the interactive electricity collection between microgrids, Eex k ={Eex 11,k ,…,Eex ij,k ,…,Eex nn,k}, Eex ij,k is the amount of electricity sent from microgrid i to microgrid j at the kth iteration; Pc k represents the user load set in the microgrid, Pc k ={Pc 1,k ,…,Pc i,k ,…,Pc n,k}, Pc i,k is the user load demand in microgrid i at the kth iteration; Pg k Represents the output of distributed power sources in the microgrid, Pg k ={Pg 1,k ,…,Pg i,k ,…,Pg n,k}, Pg i,kis the output of distributed generation in microgrid i at the kth iteration, including fossil energy and new energy units; TL k represents the set of operating states of the transmission channel between microgrids, TL k ={TL 11,k ,…,TL ij,k ,…,TL nn,k}, TL ij,k Is the transmission line from microgrid i to microgrid j in the kth iteration whether it is running. If it is running, then TL ij,k =1, if not in running state, TL ij,k =0;

[0055] A2: Construct action space A k , the mathematical expression is as follows:

[0056] A k ={Eex' k+1 ,Pg' i,k+1 ,Pbes' i,k+1} (16)

[0057] Among them, Eex' k+1 Represents the interactive power adjustment plan set of the microgrid in the next iteration, Eex' k+1 ={Eex' ii,k+1 ,…,Eex' ij,k+1 ,…,Eex' nn,k+1}, Eex' ij,k+1 is the power transmission plan of microgrid i to microgrid j at the k+1th iteration; Pg' i,k+1 represents the output plan set of each microgrid distributed generation device in the next iteration, Pg' i,k+1 ={Pg' 1,k+1 ,…,Pg' i,k+1 ,…,Pg' n,k+1}, Pg' i,k+1 is the output plan of distributed generation in microgrid i at the k+1th iteration; Pbes' i,k+1 Indicates the charging and discharging power of each microgrid energy storage at the next iteration, Pbes' i,k+1 ={Pbes' 1,k+1 ,…,Pbes' i,k+1 ,…,Pbes' n,k+1}, Pbes' i,k+1 is the charging and discharging power of the energy storage in microgrid i at the k+1th iteration. The power is positive during charging and negative during discharging.

[0058] A3: Constructing the reward function R k , mapping the multi-microgrid control constraint model into a reward function.

[0059] Furthermore, the reward function R in step A3 k The build includes:

[0060] A3-1: Constructing the energy flow information entropy reward term H i,k :

[0061]

[0062] Among them, p ij is the energy flow probability, which indicates that the energy interaction between microgrid i and microgrid j accounts for the total interaction power Ee of microgrid i i The proportion of , ε is the logarithmic control parameter;

[0063] A3-2: Constructing Power Balance Penalty Item D i,k :

[0064]

[0065] Among them, EB i,k For the kth iteration, microgrid i is The total unbalanced electricity caused by the fluctuation of distributed energy and load output within a certain period of time; Pc i,k User load demand;

[0066] A3-3: Building Energy Storage System Capacity Incentives

[0067]

[0068] This reward is to prevent the charging and discharging plan from exceeding the capacity limit of the energy storage, where the upper bound of the energy storage elasticity range of microgrid i at the kth iteration is the maximum charging capacity The lower limit of the energy storage elastic range is the maximum discharge capacity They are all determined by the energy storage charge state and energy storage capacity;

[0069] A3-4: Construct the reward function belonging to microgrid i:

[0070]

[0071] Among them, r1, r2, and r3 are weight coefficients.

[0072] The present invention also provides a multi-microgrid distributed energy scheduling system based on energy flow information entropy, comprising:

[0073] A multi-microgrid grid structure information acquisition module is used to obtain grid structure information of the target multi-microgrid area, including the number of microgrid nodes in the target multi-microgrid area, the system node impedance of the target multi-microgrid area, the maximum capacity of the transmission channel between each microgrid node in the target multi-microgrid area, and the maximum energy storage capacity of each microgrid node in the target multi-microgrid area;

[0074] The multi-microgrid operation information collection module is used to collect the operation information of multiple microgrids in real time, including the energy storage power status of each microgrid, the interactive power between microgrids, the load demand of multiple microgrid users, the output of distributed power sources in multiple microgrids, and the operation status of transmission channels between multiple microgrids;

[0075] Multi-microgrid control target model construction module, used to build scheduling control targets and related constraints;

[0076] The reinforcement learning agent module that outputs the scheduling strategy is used to establish a reinforcement learning framework, train the reinforcement learning agent based on historical data, and output the scheduling strategy based on real-time collected data;

[0077] The multi-microgrid system scheduling strategy execution module is used to execute the multi-microgrid system energy scheduling strategy.

[0078] Beneficial effects: Compared with the existing technology, the present invention introduces the concept of information entropy into the dispatching and control of microgrids in power systems, which can achieve three-level balance within the microgrid, between adjacent microgrids, and globally in the microgrid group; by minimizing the information entropy of energy flow, the system orderliness is improved, the operation frequency of equipment in the multi-microgrid system is reduced, the dispatching demand is reduced, and the local consumption efficiency of distributed energy is improved; the transmission channel capacity limitation is taken into account in the capacity reward item of the energy storage system, reducing the probability of network congestion. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 Schematic diagram of the process of the present invention;

[0080] Figure 2 Schematic diagram of the structure of the system of the present invention. DETAILED DESCRIPTION

[0081] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0082] Example 1:

[0083] like Figure 1 As shown, this embodiment provides a multi-microgrid distributed energy scheduling method based on energy flow information entropy, including the following steps:

[0084] S1: Collect and obtain multi-microgrid grid structure information based on design drawing information;

[0085] The multi-microgrid grid structure information includes: the number of microgrid nodes n in the target multi-microgrid area, the system node impedance matrix Z in the target multi-microgrid area, the electrical distance matrix D between each microgrid node in the target multi-microgrid area, the maximum capacity matrix CAP of the transmission channel between each microgrid node in the target multi-microgrid area, and the maximum energy storage capacity matrix Esc of each microgrid node in the target multi-microgrid area.

[0086] The expression of the system node impedance matrix Z in the target multi-microgrid area is:

[0087]

[0088] The element Z in row i and column j in the matrix Z ij is the system equivalent impedance viewed from ports i and j;

[0089] According to formula (1), the electrical distance matrix D between each microgrid node in the target multi-microgrid area is obtained, which is expressed as:

[0090]

[0091] The element D in row i and column j in matrix D ij The meaning of is the electrical distance from microgrid node i to microgrid node j, which is used to characterize the closeness of the electrical connection between the two microgrids. ij =Z ii +Z jj -2Z ij , when i=j, the electrical distance D ii =0;

[0092] The expression of the maximum capacity matrix CAP of the transmission channel between microgrid nodes in the target multi-microgrid area is:

[0093]

[0094] The element D in row i and column j in the matrix CAP ij is the maximum transmission capacity of the transmission channel from microgrid node i to microgrid node j. CAP is a symmetric matrix, namely C ij =C ji ;

[0095] The expression of the maximum energy storage capacity matrix Esc of each microgrid node in the target multi-microgrid area is:

[0096] Esc=[Esc1 … Esc i … Esc n ] (4)

[0097] Among them, Esc i Represents the maximum energy storage capacity of microgrid node i.

[0098] S2: Data is collected in real time through sensors installed at the energy input and output ports, energy storage, users, and distributed power sources of each microgrid node. The collection frequency is f acq Should be equal to the control frequency f con , let m be the number of acquisitions;

[0099] Multi-microgrid operation information includes the energy storage capacity state SOC of each microgrid m , Inter-microgrid interactive electricity Eex m , Multi-microgrid user load demand Pc m , Multi-microgrid distributed power output Pg m , Operational status of transmission channels between multiple microgrids TL m .

[0100] Energy storage capacity state SOC of each microgrid m The expression is as follows:

[0101]

[0102] in, is the percentage of available state of the remaining capacity of energy storage in microgrid i at the time of the mth acquisition;

[0103] Interactive electricity between microgrids Eex m The expression is as follows:

[0104]

[0105] in, is the amount of electricity sent from microgrid i to microgrid j during the mth collection;

[0106] Multi-microgrid user load demand Pc m The expression is as follows:

[0107]

[0108] in, is the user load demand in microgrid i at the time of the mth collection;

[0109] Multi-microgrid distributed power output Pg m The expression is as follows:

[0110]

[0111] in, is the total output of distributed generation in microgrid i at the time of the mth acquisition;

[0112] Operation status of transmission channel between multiple microgrids TL m The expression is as follows:

[0113]

[0114] in, Is the transmission line from microgrid i to microgrid j in the mth data collection direction in operation? If it is in operation, If not in running state

[0115] S3: Based on the multi-microgrid operation information, energy flow information entropy is introduced to construct the multi-microgrid control target model;

[0116] The control goal is to ensure the stable operation of the entire system and achieve source-load matching while minimizing the energy interaction between multiple microgrids, and ultimately achieve the effect of "local autonomy" of multi-microgrid energy.

[0117] The construction of the multi-microgrid control target model includes:

[0118] The energy interaction target is set to minimize the energy flow information entropy of each microgrid node, that is:

[0119] Min H i (10)

[0120] The expression of energy flow information entropy for microgrid i at the mth acquisition moment is:

[0121]

[0122] Where abs(·) is the absolute value function; p ij is the energy flow probability, which indicates that the energy interaction between microgrid i and microgrid j accounts for the total interaction power Ee of microgrid i i ; ε is a logarithmic control parameter, whose value affects the training step size and convergence speed of the reinforcement learning agent. In this embodiment, a fixed value of 10 is taken.

[0123] For a pure output microgrid with only one transmission line, the probability of energy flow is zero, that is, the energy flow information entropy is zero. The purpose of defining the energy flow information entropy is not to prevent energy interaction between microgrids, but to carry out reasonable energy interaction and ensure the local balance of electricity.

[0124] In the multi-microgrid control target model, the power balance constraint is used to ensure the stability of the entire system operation and the matching of source and load. The expression of the power balance constraint is:

[0125]

[0126] S4: Based on the multi-microgrid control target model, a reinforcement learning network is constructed to output the scheduling strategy;

[0127] This embodiment uses a multi-agent deep Q-network, with a DNN (Deep Neural Network) as the neural network core. To achieve distributed energy scheduling, a reinforcement learning agent is deployed at each microgrid node. This agent perceives information from the state space, calculates a reward function, and adjusts each agent's control operations in the action space by minimizing the reward function.

[0128] The construction of the reinforcement learning network includes:

[0129] A1: Constructing the multi-microgrid system state space S k , the mathematical expression is as follows:

[0130] S k ={SOC k ,Eex k ,Pc k ,Pg k ,TL k} (15)

[0131] Among them, k represents the iteration round, SOC k Represents the microgrid battery state set, SOC k ={soc 1,k ,…,soc i,k ,…,soc i,k}, soc i,k is the percentage of available remaining capacity of energy storage in microgrid i at the kth iteration; Eex k Represents the interactive electricity collection between microgrids, Eex k ={Eex 11,k ,…,Eex ij,k ,…,Eex nn,k}, Eex ij,k is the amount of electricity sent from microgrid i to microgrid j at the kth iteration; Pc k represents the user load set in the microgrid, Pc k ={Pc 1,k ,…,Pc i,k ,…,Pc n,k}, Pc i,k is the user load demand in microgrid i at the kth iteration; Pg k Represents the output of distributed power sources in the microgrid, Pg k ={Pg 1,k ,…,Pg i,k ,…,Pg n,k}, Pg i,kis the output of distributed generation in microgrid i at the kth iteration, including fossil energy and new energy units; TL k represents the set of operating states of the transmission channel between microgrids, TL k ={TL 11,k ,…,TL ij,k ,…,TL nn,k}, TL ij,k Is the transmission line from microgrid i to microgrid j in the kth iteration whether it is running. If it is running, then TL ij,k =1, if not in running state, TL ij,k =0; It is worth noting that when the reinforcement learning agent is trained, the state space uses the historical operating data of the target multi-microgrid system, not the real-time collected data;

[0132] A2: Construct action space A k , the mathematical expression is as follows:

[0133] A k ={Eex' k+1 ,Pg' i,k+1 ,Pbes' i,k+1} (16)

[0134] Among them, Eex' k+1 Represents the interactive power adjustment plan set of the microgrid in the next iteration, Eex' k+1 ={Eex' ii,k+1 ,…,Eex' ij,k+1 ,…,Eex' m,k+1}, Eex' ij,k+1 is the power transmission plan of microgrid i to microgrid j at the k+1th iteration; Pg' i,k+1 represents the output plan set of each microgrid distributed generation device in the next iteration, Pg' i,k+1 ={Pg' 1,k+1 ,…,Pg' i,k+1 ,…,Pg' n,k+1}, Pg' i,k+1 is the output plan of distributed generation in microgrid i at the k+1th iteration; Pbes' i,k+1 Indicates the charging and discharging power of each microgrid energy storage at the next iteration, Pbes' i,k+1 ={Pbes' 1,k+1 ,…,Pbes' i,k+1 ,…,Pbes' n,k+1}, Pbes' i,k+1 is the charging and discharging power of the energy storage in microgrid i at the k+1th iteration. The power is positive during charging and negative during discharging.

[0135] A3: Constructing the reward function R k , mapping the multi-microgrid control constraint model into a reward function;

[0136] The core role of the reward function is to guide the agent towards its pre-set goal. The agent executes actions in a specific state according to its current strategy. The environment generates state transitions based on the feedback from these actions and simultaneously provides an immediate reward signal. This reward signal essentially constructs an incentive mechanism, the goal of which is to maximize the cumulative sum of these reward signals. In this invention, each agent has its own reward function.

[0137] Reward function R k The build includes:

[0138] A3-1: Constructing the energy flow information entropy reward term H i,k :

[0139]

[0140] Among them, p ij is the energy flow probability, which indicates that the energy interaction between microgrid i and microgrid j accounts for the total interaction power Ee of microgrid i i The proportion of; ε is a logarithmic control parameter, which takes a fixed value of 10 in this embodiment;

[0141] A3-2: Constructing Power Balance Penalty Item D i,k :

[0142]

[0143] Among them, EB i,k For the kth iteration, microgrid i is The unbalanced total power caused by the fluctuation of distributed energy and load output within a certain period of time is calculated using the user load demand PC. i,k The penalty item has been revised;

[0144] A3-3: Building Energy Storage System Capacity Incentives

[0145]

[0146] This reward is to prevent the charging and discharging plan from exceeding the capacity limit of the energy storage, where the upper bound of the energy storage elasticity range of microgrid i at the kth iteration is the maximum charging capacity The lower limit of the energy storage elastic range is the maximum discharge capacity They are all determined by the energy storage charge state and energy storage capacity;

[0147] A3-4: Construct the reward function belonging to microgrid i:

[0148]

[0149] Among them, r1, r2, and r3 are weight coefficients, which are set to r1 = 0.3, r2 = 0.5, and r3 = 0.2 in this embodiment.

[0150] S5: Train the reinforcement learning network;

[0151] S6: Input the real-time collected multi-microgrid operation information into the trained reinforcement learning network and output the scheduling strategy;

[0152] S7: Execute the multi-microgrid system scheduling strategy.

[0153] Each microgrid node agent performs actions according to the scheduling strategy generated in step S6, including actuating the microgrid node input and output port controller to adjust the disconnection state of the transmission channel between microgrids at the next moment; actuating the distributed power generation device controller within the microgrid to adjust the output plan of each microgrid distributed power generation device at the next moment; and actuating the energy storage system controller to set the energy storage charging and discharging power of each microgrid at the next moment.

[0154] A multi-microgrid system is composed of multiple microgrids, which are connected by transmission lines and are at the same voltage level. In actual operation, due to the volatility of the output of distributed power generation devices within the multi-microgrid, the magnitude and even direction of the line flow frequently change, which poses a challenge to the stable and orderly operation of the multi-microgrid system. In order to guide the multi-microgrid system to evolve towards an ordered state, the method provided in this embodiment introduces the concept of information entropy into the microgrid dispatching control of the power system. Information entropy is the expected value of the amount of information about all possible events and measures the uncertainty of the system. The greater the information entropy, the more uncertain and complex the system. For a multi-microgrid system, the greater the information entropy of the interactive energy flow, the more difficult it is for the system to achieve local energy autonomy. The more ordered a system is, the lower the information entropy is. The method of the present invention uses a reinforcement learning algorithm to achieve the control target of minimizing the information entropy of energy flow. The ultimate goal is to make the multi-microgrid system an ordered system.

[0155] Based on the collected multi-microgrid grid structure information, the method of the present invention uses historical multi-microgrid operation information to perform reinforcement learning training in the target multi-microgrid scheduling and control area. Under the guidance of a set reward function, the optimal strategy is interactively learned with the environmental state to train the intelligent agent network, where the scheduling control objectives and constraints: energy flow information entropy, power balance constraints, and energy storage system capacity constraints are reflected by the reward function. After training is completed, the real-time collected multi-microgrid operation information is input to the reinforcement learning intelligent agent, and the action information of the multi-microgrid system at the next moment is output, including the multi-microgrid system adjusting the disconnection state of the transmission channel between microgrids at the next moment; each sub-microgrid node sets the output of its distributed generation device and the energy storage charge and discharge power at the next moment.

[0156] Example 2:

[0157] Based on the method of Example 1, Figure 2 As shown, this embodiment provides a multi-microgrid distributed energy scheduling system based on energy flow information entropy, including:

[0158] A multi-microgrid grid structure information acquisition module is used to obtain grid structure information of the target multi-microgrid area, including the number of microgrid nodes in the target multi-microgrid area, the system node impedance of the target multi-microgrid area, the maximum capacity of the transmission channel between each microgrid node in the target multi-microgrid area, and the maximum energy storage capacity of each microgrid node in the target multi-microgrid area;

[0159] The multi-microgrid operation information collection module is used to collect the operation information of multiple microgrids in real time, including the energy storage power status of each microgrid, the interactive power between microgrids, the load demand of multiple microgrid users, the output of distributed power sources in multiple microgrids, and the operation status of transmission channels between multiple microgrids;

[0160] Multi-microgrid control target model construction module, used to build scheduling control targets and related constraints;

[0161] The reinforcement learning agent module that outputs the scheduling strategy is used to establish a reinforcement learning framework, train the reinforcement learning agent based on historical data, and output the scheduling strategy based on real-time collected data;

[0162] The multi-microgrid system scheduling strategy execution module is used to execute the multi-microgrid system energy scheduling strategy.

Claims

1. A multi-microgrid distributed energy scheduling method based on energy flow information entropy, characterized in that: The steps include: S1: Collecting multi-microgrid structure information; S2: Based on the multi-microgrid grid structure information, collect multi-microgrid operation information; S3: Based on the multi-microgrid operation information, the energy flow information entropy is introduced to construct the multi-microgrid control target model; S4: Based on the multi-microgrid control target model, a reinforcement learning network is constructed to output the scheduling strategy; S5: Train the reinforcement learning network; S6: Input the real-time collected multi-microgrid operation information into the trained reinforcement learning network and output the scheduling strategy; S7: Execute multi-microgrid system scheduling strategy; The construction of the multi-microgrid control target model in step S3 includes: The energy interaction target is set to minimize the energy flow information entropy of each microgrid node, that is: My H i (10) The expression of energy flow information entropy for microgrid i at the mth acquisition moment is: Among them, abs(·) is the absolute value function, p ij is the energy flow probability, which indicates that the energy interaction between microgrid i and microgrid j accounts for the total interaction power Ee of microgrid i i The proportion of , ε is the logarithmic control parameter; The construction of the reinforcement learning network in step S4 includes: A1: Constructing the multi-microgrid system state space S k , the mathematical expression is as follows: S k ={SOC k ,Eex k ,PC k ,Pg k ,TL k } (15) Among them, k represents the iteration round, SOC k Represents the microgrid battery state set, SOC k ={soc 1,k ,…,soc i,k ,…,soc n,k }, soc i,k is the percentage of available remaining capacity of energy storage in microgrid i at the kth iteration; Eex k Represents the interactive electricity collection between microgrids, Eex k ={Eex 11,k ,…,Eex ij,k ,…,Eex nn,k }, Eex ij,k is the amount of electricity sent from microgrid i to microgrid j at the kth iteration; Pc k represents the user load set in the microgrid, Pc k ={Pc 1,k ,…,Pc i,k ,…,Pc n,k }, Pc i,k is the user load demand in microgrid i at the kth iteration; Pg k Represents the output of distributed power sources in the microgrid, Pg k ={Pg 1,k ,…,Pg i,k ,…,Pg n,k }, Pg i,k is the output of distributed generation in microgrid i at the kth iteration, including fossil energy and new energy units; TL k represents the set of operating states of the transmission channel between microgrids, TL k ={TL 11,k ,…,TL ij,k ,…,TL nn,k }, TL ij,k Is the transmission line from microgrid i to microgrid j in the kth iteration whether it is running. If it is running, then TL ij,k =1, if not in running state, TL ij,k =0; A2: Construct action space A k , the mathematical expression is as follows: A k ={Eex' k+1 ,Pg' i,k+1 ,Pbes' i,k+1 } (16) Among them, Eex' k+1 Represents the interactive power adjustment plan set of the microgrid in the next iteration, Eex' k+1 ={Eex' ii,k+1 ,…,Eex' ij,k+1 ,…,Eex' nn,k+1 }, Eex' ij,k+1 is the power transmission plan of microgrid i to microgrid j at the k+1th iteration; Pg' i,k+1 represents the output plan set of each microgrid distributed generation device in the next iteration, Pg' i,k+1 ={Pg' 1,k+1 ,…,Pg' i,k+1 ,…,Pg' n,k+1 }, Pg' i,k+1 is the output plan of distributed generation in microgrid i at the k+1th iteration; Pbes' i,k+1 Indicates the charging and discharging power of each microgrid energy storage at the next iteration, Pbes' i,k+1 ={Pbes' 1,k+1 ,…,Pbes' i,k+1 ,…,Pbes' n,k+1 }, Pbes' i,k+1 is the charging and discharging power of the energy storage in microgrid i at the k+1th iteration; A3: Constructing the reward function R k , mapping the multi-microgrid control constraint model into a reward function; The reward function R in step A3 k The build includes: A3-1: Constructing the energy flow information entropy reward term H i,k : Among them, p ij is the energy flow probability, which indicates that the energy interaction between microgrid i and microgrid j accounts for the total interaction power Ee of microgrid i i The proportion of , ε is the logarithmic control parameter; A3-2: Constructing Power Balance Penalty Item D i,k : Among them, EB i,k For the kth iteration, microgrid i is The unbalanced total power caused by the fluctuation of distributed energy output and user load demand within a certain period of time; Pc i,k User load demand; A3-3: Building Energy Storage System Capacity Incentives Among them, the upper limit of the energy storage elasticity range of microgrid i at the kth iteration is the maximum charging capacity The lower limit of the energy storage elastic range is the maximum discharge capacity f acq is the acquisition frequency; A3-4: Construct the reward function belonging to microgrid i: Among them, r1, r2, and r3 are weight coefficients.

2. A multi-microgrid distributed energy scheduling method based on energy flow information entropy according to claim 1, characterized in that: The multi-microgrid grid structure information in step S1 includes: the number n of microgrid nodes in the target multi-microgrid area, the system node impedance matrix Z of the target multi-microgrid area, the electrical distance matrix D between each microgrid node in the target multi-microgrid area, the maximum capacity matrix CAP of the transmission channel between each microgrid node in the target multi-microgrid area, and the maximum energy storage capacity matrix Esc of each microgrid node in the target multi-microgrid area.

3. The multi-microgrid distributed energy scheduling method based on energy flow information entropy according to claim 2 is characterized in that: The expression of the system node impedance matrix Z of the target multi-microgrid area in step S1 is: The element Z in row i and column j in the matrix Z ij is the system equivalent impedance viewed from ports i and j; According to formula (1), the electrical distance matrix D between each microgrid node in the target multi-microgrid area is obtained, which is expressed as: The element D in row i and column j in matrix D ij The meaning of is the electrical distance from microgrid node i to microgrid node j, which is used to characterize the closeness of the electrical connection between the two microgrids. ij =Z ii +Z jj -2Z ij ; The expression of the maximum capacity matrix CAP of the transmission channel between microgrid nodes in the target multi-microgrid area is: The element C in row i and column j in the matrix CAP ij is the maximum transmission capacity of the transmission channel from microgrid node i to microgrid node j. CAP is a symmetric matrix, namely C ij =C ji ; The expression of the maximum energy storage capacity matrix Esc of each microgrid node in the target multi-microgrid area is: Esc=[Esc1 … Esc i … Esc n ] (4) Among them, Esc i Represents the maximum energy storage capacity of microgrid node i.

4. The multi-microgrid distributed energy scheduling method based on energy flow information entropy according to claim 3 is characterized in that: The multi-microgrid operation information in step S2 includes the energy storage capacity state SOC of each microgrid. m , Inter-microgrid interactive electricity Eex m , Multi-microgrid user load demand Pc m , Multi-microgrid distributed power output Pg m , Operational status of transmission channels between multiple microgrids TL m .

5. The multi-microgrid distributed energy scheduling method based on energy flow information entropy according to claim 4 is characterized in that: The energy storage state SOC of each microgrid in step S2 m The expression is as follows: in, is the percentage of available state of the remaining capacity of energy storage in microgrid i at the time of the mth acquisition; Interactive electricity between microgrids Eex m The expression is as follows: in, is the amount of electricity sent from microgrid i to microgrid j during the mth collection; Multi-microgrid user load demand Pc m The expression is as follows: in, is the user load demand in microgrid i at the time of the mth collection; Multi-microgrid distributed power output Pg m The expression is as follows: in, is the total output of distributed generation in microgrid i at the time of the mth acquisition; Operation status of transmission channel between multiple microgrids TL m The expression is as follows: in, Is the transmission line from microgrid i to microgrid j in the mth data collection direction in operation? If it is in operation, If not in running state 6. The multi-microgrid distributed energy scheduling method based on energy flow information entropy according to claim 1 is characterized in that: In the multi-microgrid control target model of step S3, the power balance constraint is used to ensure the stable operation of the entire system and the matching of source and load. The expression of the power balance constraint is:

Citation Information

Patent Citations

  • Multi-microgrid system layered reinforcement learning optimization method and system, and storage medium

    CN115115211A