A multi-microgrid system coordination control method fusing q-learning and potential game

By integrating Q-learning and potential game theory, a distributed game architecture is constructed and Nash equilibrium is directly calculated using parameter passing. This solves the problem of economic relations and interest balance in multi-microgrid systems, realizes the maximization of microgrid revenue and output balance, and improves the system's economy and computational efficiency.

CN115411728BActive Publication Date: 2026-05-05NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF SCI & TECH
Filing Date
2022-09-26
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Traditional methods are difficult to effectively coordinate economic relationships and balance interests within multi-microgrid systems. Traditional centralized control methods are difficult to meet the control requirements of multi-microgrid systems, and existing game optimization methods have high computational complexity, making it difficult to maximize microgrid revenue and balance output.

Method used

A coordinated control method for multi-microgrid systems integrating Q-learning and potential game theory is proposed. By constructing a distributed game architecture, designing local payoff functions and global situation functions, and utilizing parameter transfer to integrate potential game theory and Q-learning algorithms, the Nash equilibrium is directly calculated, reducing computational complexity and maximizing microgrid revenue and achieving power output balance.

Benefits of technology

It improves the economic efficiency and balance of interests of multi-microgrid systems, reduces computational complexity, achieves a balance between the overall system and individual interests, and optimizes the output coordination of microgrids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115411728B_ABST
    Figure CN115411728B_ABST
Patent Text Reader

Abstract

This paper presents a coordinated control method for multi-microgrid systems that integrates Q-learning and potential game theory, belonging to the field of microgrid coordinated control technology. It addresses the problem of achieving coordinated control of multi-microgrid systems with the goals of maximizing microgrid revenue and balancing output among microgrids. Based on a distributed coordination architecture and potential game optimization strategy, a coordinated control method integrating reinforcement learning and potential game theory is constructed. Leveraging the distributed nature of potential game theory, each microgrid is treated as an intelligent agent. A distributed coordination control structure is adopted, and a potential game model is established to maximize and balance the economic efficiency of individual microgrids and the overall multi-microgrid system. Then, using the Q-learning algorithm of reinforcement learning as a carrier, the potential game theory and reinforcement learning algorithms are integrated through parameter transfer to obtain the optimal Nash equilibrium solution. This improves optimization performance, enhances the economic efficiency of the multi-microgrid system, and achieves a balance between the interests of the overall system and its individual components.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of microgrid coordinated control technology, and relates to a coordinated control method for multi-microgrid systems that integrates Q-learning and potential game theory. Background Technology

[0002] With the rapid development of renewable energy technologies and the widespread penetration of distributed energy in power distribution networks, single microgrid systems are gradually transforming into multi-microgrid systems. Multi-microgrids not only have higher reliability but also effectively improve the local consumption capacity of renewable energy. However, due to their large scale, high complexity, and diversified investment entities, traditional centralized control methods are difficult to meet their control needs, and it is difficult to balance the overall interests of the system with the interests of individuals within the system (see the literature "A multiagent-based hierarchical energy management strategy for multi-microgrids considering adjustable power and demand response" (VHBui, etc., IEEE Transactions on Smart Grid 9.2(2018):1323-1333). Therefore, it is urgent to study a distributed coordinated control method for multi-microgrids that effectively coordinates the economic relationship between the whole and the individuals and improves the economic efficiency of the system.

[0003] Reinforcement learning primarily involves agents interacting with their environment to continuously improve their behavior. The agent selects actions to act on the environment, receives rewards or penalties, and chooses the next action based on the feedback and changes in the environment. Actions beneficial to the goal are retained, while actions detrimental to the goal are eliminated. Q-learning is an offline control algorithm in reinforcement learning based on value function iteration. Its principle is to use a Q-value table containing prior experience as the initial value for subsequent iterations, thereby shortening the algorithm's convergence time. Potential games (PGs) are a subclass of non-cooperative games, first proposed by Modener and Shapely in 1996. They map changes in individual payoffs to a potential function. When an individual increases their payoff by adjusting their strategy, the value of the potential function also increases. By solving for the maximum or maxima of the potential function, the Nash equilibrium can be indirectly obtained. Potential games have distributed characteristics, making them suitable for solving distributed optimization problems. They also possess finite improvement properties (FIP), as every finite potential game has a pure policy Nash equilibrium. Therefore, potential games have significant advantages in terms of algorithm complexity and computational cost.

[0004] In existing technologies, coordination game optimization of multi-microgrid systems often employs traditional master-slave game theory and Cuno oligopoly game theory. For example, the paper "Economic optimization method of multi-stakeholder in a multi-microgrid system based on Stackelberg game theory" (Q.Wu, etc., Energy Reports 8(2022):345-351) proposes a microgrid system energy management optimization method based on Stackelberg game theory; and the paper "Cournot oligopoly game-based local energy trading considering renewable energy uncertainty costs" (YJZhang, etc., Renewable Energy 159.3(2020):1117-1127) applies Cuno oligopoly game theory to the electricity market to improve transactions between power generation companies and customs or balance the profits among multiple suppliers. However, these methods all suffer from problems such as difficulty in fitting distributed optimization control methods or the complexity of Nash equilibrium solution processes. The paper "A Potential Game Approach to Distributed Operational Optimization for Microgrid Energy Management with Renewable Energy and Demand Response" (J. Zeng, etc., IEEE Transactions on Industrial Electronics 66.6(2019):4479-4489) applies potential game theory to the fully distributed operation optimization of microgrid energy management systems. However, when there are many game participants and a large set of strategies, the computational load of this method is still very large, and the algorithm's solution performance still needs to be improved.The paper "Coordinated Scheduling of Grid-Connected Integrated Energy Microgrids Based on Multi-Agent Game Theory and Reinforcement Learning" (Liu Hong et al., Key Laboratory of Smart Grid, Ministry of Education, Tianjin University, January 2019) addresses the challenges of traditional centralized optimization scheduling methods failing to fully reflect the interests of different agents within an integrated energy microgrid, and the need for further exploration of the application of artificial intelligence in integrated energy scheduling. It proposes a coordinated scheduling model and method for grid-connected integrated energy microgrids based on multi-agent game theory and reinforcement learning. However, the technical problem addressed in this paper is achieving coordinated scheduling of the microgrid with the goal of balancing the interests of multiple agents. The technical solution adopted in this paper involves establishing a multi-agent game-based coordinated scheduling model based on joint game theory, first selecting state-action values ​​that satisfy Nash equilibrium, and then using the Nash Q-learning algorithm for iterative calculation to solve for the optimal Nash equilibrium. The process of selecting Nash equilibrium values ​​is relatively complex and computationally intensive. Summary of the Invention

[0005] The technical problem to be solved by this invention is how to achieve coordinated control of multiple microgrids with the goal of maximizing microgrid revenue and balancing the output between microgrids.

[0006] The present invention solves the above-mentioned technical problems through the following technical solutions:

[0007] A coordinated control method for multi-microgrid systems integrating Q-learning and potential game theory includes the following steps:

[0008] S1. Construct a target optimization decision model for maximizing microgrid output revenue and achieving output balance under a distributed game architecture of multiple microgrids, and set power balance constraints and microgrid output constraints.

[0009] S2. The local payoff function is obtained by linearly weighting the target optimization decision. Then, the global situation function and local utility function that satisfy the potential equation are designed to establish the potential game strategy set and construct a potential game model with distributed characteristics.

[0010] S3. The potential game control and Q-learning algorithm are integrated by parameter passing to solve the potential game model, obtain the game optimization results, and analyze them.

[0011] The technical solution of this invention is based on a distributed coordination architecture and potential game optimization for multi-microgrid systems. It constructs a coordinated control method for multi-microgrid systems that integrates reinforcement learning and potential game theory. By fully utilizing the distributed characteristics of potential game theory, each microgrid is treated as an intelligent agent. A distributed coordination control structure is adopted to establish a potential game model with the aim of maximizing and balancing the economic efficiency of individual microgrids and the overall multi-microgrid system. Then, using the Q-learning algorithm as a carrier, the potential game theory and reinforcement learning algorithms are integrated through parameter transfer to obtain the optimal Nash equilibrium solution, thereby improving optimization performance. This achieves the dual objectives of maximizing microgrid revenue and balancing output among microgrids, improving the economic efficiency of the multi-microgrid system and achieving a balance of interests between the overall system and individual microgrids. Furthermore, it eliminates the need for filtering state action values; the game utility function value is passed to the reward value and directly substituted into the Q-learning iterative formula to calculate the Nash equilibrium and determine whether it is the optimal Nash equilibrium, further reducing computational complexity.

[0012] Furthermore, the method for constructing the optimization decision model described in step S1 is as follows:

[0013] 1) The net benefit of maximizing the power output of the microgrid is:

[0014] maxF 1,i =(ρ-m i )P i (1)

[0015] Among them, F 1,i The net revenue from power output of the microgrid, P i Let ρ be the output of microgrid i in a multi-microgrid system, ρ be the unit electricity price, and m be the output of microgrid i. i The output cost coefficient of microgrid i;

[0016] 2) Minimize the power difference between each microgrid and its neighboring microgrids in a multi-microgrid system to balance the output of each microgrid. The objective function is:

[0017]

[0018] Among them, F 2,i I represents the power difference between microgrid i and its neighboring microgrid j. i Let P be the neighbor set of microgrid i. j The output of microgrid i is for its neighboring microgrid j.

[0019] Furthermore, the power balance constraint and microgrid output constraint conditions described in step S1 are as follows:

[0020]

[0021] Among them, P loadLet P be the total load of the multi-microgrid system, N be the set of potential game participants, and P be the total load of the multi-microgrid system. i,max n is the rated capacity of microgrid i; MG This refers to the number of microgrids in a multi-microgrid system.

[0022] Furthermore, the linear weighting process described in step S2 is as follows:

[0023]

[0024] Among them, F i (P i ,P -i Let P be the local payment function of microgrid i. -i For the output of other microgrids besides microgrid i in a multi-microgrid system, λ1 and λ2 are the weighting coefficients for different objective functions.

[0025] Furthermore, the formula for the overall situation function φ mentioned in step S2 is as follows:

[0026]

[0027] The formula for the local utility function is as follows:

[0028]

[0029] Among them, U i (P i ,P -i F is a local utility function. j (P i ,P -i Let be the local payment function of microgrid j, a neighboring microgrid of microgrid i.

[0030] Furthermore, the design method for the potential game strategy set described in step S2 is as follows:

[0031] (1) Design a potential game strategy set based on the microgrid output constraints, the potential game strategy set Y i Specifically:

[0032] Y i ={P i :0≤P i ≤P i,max} (7)

[0033] (2) The potential game strategy obtained by solving must be within the microgrid capacity limit and also satisfy the power balance constraint of the multi-microgrid system.

[0034] Furthermore, the method described in step S3 for fusing potential game control with the Q-learning algorithm via parameter passing, solving the potential game model, obtaining the game optimization results, and analyzing them is as follows:

[0035] (a) First, initialize the game parameters and Q-value, discretize the potential game strategy set, and pass it to the state set learned by Q.

[0036] (b) Considering the rated capacity of the microgrid and to avoid excessive power fluctuations that could cause system instability, design a Q-learning action set consisting of the power change value ΔP;

[0037] (c) Collect information about neighboring microgrids, calculate the utility function of each microgrid, pass the utility function value to the immediate reward in the Q-learning algorithm, and update the Q value in the Q-learning algorithm;

[0038] (d) Use a greedy strategy to select the optimal action, update the state value according to the selected action, and pass the state value to the game optimization strategy.

[0039] (e) Determine whether Nash equilibrium has been reached. If it has, proceed to the next step; otherwise, return to step (c).

[0040] (f) Determine whether the convergence condition is met. If it is met, obtain the final microgrid output plan; otherwise, return to step (c).

[0041] Furthermore, the discrete interval length ΔP of the potential game strategy set described in step (a) s for:

[0042]

[0043] Where M is the number of intervals divided; P max and P min Determined by the upper and lower bounds of the strategy set in a potential game;

[0044] Furthermore, the formula for updating the Q-value in the Q-learning algorithm described in step (c) is as follows:

[0045]

[0046] Among them, P i ∈A represents the action value at each step in Q-learning, α∈[0,1] is the learning rate of the Q-learning algorithm, and γ∈[0,1] is the discount parameter. This is the Q-iteration value at the (k+1)th iteration. Let ΔP be the Q-iteration value at the k-th iteration. i Let be the output change value of the i-th microgrid. Let ΔP be the utility function value of the i-th microgrid at the k-th time. i'P' represents the output change value corresponding to the maximum Q value in the k-th iteration of the i-th microgrid. i 'For the i-th microgrid, after passing through ΔP i 'The changed output value.'

[0047] Furthermore, the formula for selecting the optimal action using a greedy strategy described in step (d) is as follows:

[0048]

[0049] in, This refers to the optimal action chosen using a greedy strategy.

[0050] The advantages of this invention are:

[0051] The technical solution of this invention is based on a distributed coordination architecture and potential game optimization for multi-microgrid systems. It constructs a coordinated control method for multi-microgrid systems that integrates reinforcement learning and potential game theory. By fully utilizing the distributed characteristics of potential game theory, each microgrid is treated as an intelligent agent. A distributed coordination control structure is adopted to establish a potential game model with the aim of maximizing and balancing the economic efficiency of individual microgrids and the overall multi-microgrid system. Then, using the Q-learning algorithm as a carrier, the potential game theory and reinforcement learning algorithms are integrated through parameter transfer to obtain the optimal Nash equilibrium solution, thereby improving optimization performance. This achieves the dual objectives of maximizing microgrid revenue and balancing output among microgrids, improving the economic efficiency of the multi-microgrid system and achieving a balance of interests between the overall system and individual microgrids. Furthermore, it eliminates the need for filtering state action values; the game utility function value is passed to the reward value and directly substituted into the Q-learning iterative formula to calculate the Nash equilibrium and determine whether it is the optimal Nash equilibrium, further reducing computational complexity. Attached Figure Description

[0052] Figure 1 This is the distributed game coordination architecture for multi-microgrids according to Embodiment 1 of the present invention;

[0053] Figure 2 This is a flowchart of the coordinated control method for multi-microgrid systems that integrates Q-learning and potential game theory according to Embodiment 1 of the present invention;

[0054] Figure 3 This is a flowchart of the fusion of Q-learning algorithm and potential game theory in Embodiment 1 of the present invention;

[0055] Figure 4 This is a simulation model structure diagram of the coordinated control of a multi-microgrid system integrating Q-learning and potential game theory according to Embodiment 1 of the present invention.

[0056] Figure 5 This is a comparison diagram of the system output before and after the coordinated control of a multi-microgrid system integrating Q-learning and potential game theory according to Embodiment 1 of the present invention.

[0057] Figure 6 This is a comparison diagram of system benefits before and after the coordinated control of a multi-microgrid system integrating Q-learning and potential game theory according to Embodiment 1 of the present invention;

[0058] Figure 7 This is a verification diagram of the results of coordinated control of a multi-microgrid system integrating Q-learning and potential game theory according to Embodiment 1 of the present invention.

[0059] Figure 8 This is a comparison diagram of the system output of the multi-microgrid system coordinated control integrating Q-learning and potential game theory in Embodiment 1 of the present invention and the traditional potential game control.

[0060] Figure 9 This is a comparison diagram of the system benefits of coordinated control of a multi-microgrid system integrating Q-learning and potential game theory, and traditional potential game control, according to Embodiment 1 of the present invention.

[0061] Figure 10 This is a comparison chart of the convergence of the algorithms for coordinated control of multi-microgrid systems integrating Q-learning and potential game theory in Embodiment 1 of the present invention and traditional potential game theory control. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments:

[0064] Example 1

[0065] like Figure 1 As shown, the distributed game-theoretic coordination architecture of the multi-microgrid in this embodiment treats each microgrid as an intelligent agent. The agents interact with each other through virtual communication lines to exchange information among neighbors. After collecting neighbor information, each agent takes into account its own utility and the utility of its neighbors to carry out game-theoretic coordination and control, thereby obtaining the optimized output of each microgrid, thus maximizing the benefits of each microgrid and the overall benefits of the multi-microgrid system.

[0066] like Figure 2 As shown in this embodiment, the coordinated control method for multi-microgrid systems that integrates Q-learning and potential game theory includes the following steps:

[0067] Step 1: With the aim of maximizing the total system revenue and balancing the economic relationships within the system, construct an optimization decision model under the distributed game architecture of multiple microgrids, which takes into account the objectives of maximizing microgrid output revenue and balancing output. Set power balance and microgrid output constraints to coordinate and optimize the economy and fairness of multiple microgrids.

[0068] Step 1.1: Construct an optimization decision model under a distributed game architecture involving multiple microgrids, considering both the goal of maximizing microgrid output revenue and the goal of achieving output balance. First, consider maximizing the net revenue obtained from microgrid output, which can be written as:

[0069] maxF 1,i =(ρ-m i )P i (1)

[0070] In formula (1), P i Let m be the output of the microgrid i in the system; ρ be the unit electricity price; where m i The output cost coefficient of microgrid i.

[0071] Step 1.2: Next, consider minimizing the power difference between each microgrid and its neighboring microgrids to balance the output of each microgrid, increase fairness, and avoid resource waste. The objective function can be written as:

[0072]

[0073] In formula (2), I i Let P be the neighbor set of microgrid i; j Let J be the output of microgrid j, which is a neighbor of microgrid i in the system.

[0074] Step 1.3: To ensure the safe and stable operation of the power system, all variables must be within the specified range. The power balance constraints and microgrid output constraints are as follows:

[0075]

[0076] In formula (3), P load The total load of the multi-microgrid system is N; the set of game participants is P. i,max The rated capacity of the microgrid; n MG This refers to the number of microgrids in a multi-microgrid system.

[0077] Step 2: Based on the optimization decision objective constructed in Step 1, linear weighting is performed to obtain the local payoff function of each microgrid. Then, the local utility function and global situation function that satisfy the potential equation are designed, the potential game strategy set is established, and a game coordination model with distributed characteristics is constructed to realize the distributed coordination control function among multiple microgrids.

[0078] Step 2.1: The linear weighted method is used to process the optimization decision objective in Step 1. The process is as follows:

[0079]

[0080] In formula (4), P -i For the system to provide power output to other microgrids besides microgrid i; and

[0081] These are the weighting coefficients for different objective functions.

[0082] Step 2.2: Design the global situation function and local utility function that satisfy the potential equation. Based on the principle of maximizing the overall benefit of the system, the global situation function is established as follows:

[0083]

[0084] The design of the local utility function considers not only the payoff obtained by the player's own strategy, but also the impact of the neighbor's strategy on the player's own payoff. Its formula is as follows:

[0085]

[0086] In formula (6), F i (P i ,P -i F is the local payment function of microgrid i; j (P i ,P -i Let be the local payoff function of the neighbors of microgrid i. This formula fully reflects the distributed characteristics of potential game theory, which is consistent with the distributed optimization concept of multiple microgrids and improves optimization performance.

[0087] Step 2.3: The game strategy set can be designed based on the microgrid output constraints, and can be written as:

[0088] Y i ={P i :0≤P i ≤P i,max} (7)

[0089] The final game strategy obtained must be within the microgrid capacity limit and also satisfy the system power balance constraint.

[0090] Step 3: Combining potential game theory and reinforcement learning principles, a fusion of potential game control and Q-learning algorithm is proposed using a parameter transfer method. This leads to a coordinated control algorithm for multi-microgrid systems that integrates reinforcement learning and potential game theory. The distributed potential game model obtained in Step 2 is then solved, and its optimization performance and convergence are effectively improved. The fusion algorithm flow is as follows: Figure 3As shown, the final game optimization results are obtained and analyzed.

[0091] Step 3.1: First, initialize the game parameters and Q-value, discretize the game strategy set, and pass it to the state set for Q-learning. The game strategy set is discretized into interval form to correspond to the discrete state set for Q-learning. The interval length can be written as:

[0092]

[0093] In formula (9), M is the number of intervals divided; P max and P min It is determined by the upper and lower limits of the game strategy set.

[0094] Step 3.2: Considering the rated capacity of the microgrid and to avoid excessive power fluctuations that could cause system instability, design a Q-learning action set consisting of power change values ​​P.

[0095] Step 3.3: Collect information about neighboring microgrids, calculate the utility function of each microgrid, and pass the utility function value to the immediate reward in the Q-learning algorithm. Update the Q value according to the following formula:

[0096]

[0097] In formula (9), P i ∈A represents the action value at each step in Q-learning; α∈[0,1] is the learning rate of Q-learning; γ∈[0,1] is the discount parameter.

[0098] Step 3.4: Use a greedy strategy to select the optimal action as shown in formula (10), update the state value according to the selected action, and pass the state value to the game optimization strategy.

[0099]

[0100] Step 3.5: Determine if Nash equilibrium has been reached. If it has, proceed to the next step; otherwise, return to step 3.3.

[0101] Step 3.6: Determine whether the convergence condition is met. If it is met, obtain the final microgrid output plan; otherwise, return to step 3.3.

[0102] like Figure 4As shown, a simulation model of a multi-microgrid system is established, containing three microgrids. Each microgrid contains several distributed power sources and loads. The rated active power of loads 1 and 3 is 0.5MW, and the rated active power of loads 2 and 4 is 0.6MW and 0.4MW, respectively. The rated capacity of each microgrid is 1MW. At t=0s, switch S1 is open, and the multi-microgrid system operates in islanded mode. At t=1s, a load surge occurs in the multi-microgrid system, at which point reinforced game theory control is implemented. The unit electricity price is set at 1.2 yuan / kW, and the output cost coefficients m1, m2, and m3 of each microgrid are 0.7, 0.6, and 0.8 yuan / kW, respectively.

[0103] like Figure 5 and Figure 6 As shown, the output of each microgrid is significantly closer after game-theoretic coordination than during autonomous operation, indicating a significant improvement in the output balance of the microgrids. Under the reinforced game-theoretic control mode, the overall revenue of the multi-microgrid system increases by 16.84 yuan compared to the autonomous operation mode, indicating that the overall system benefit is effectively optimized. Secondly, when using the reinforced game-theoretic method, the stable output and revenue of microgrid 1 and microgrid 2 are both higher than during autonomous operation, while the output and revenue of microgrid 3 are lower. This is because microgrid 3 has the highest output cost coefficient, and increasing the output of microgrid 3 does not easily yield higher returns. To better balance the individual interests of microgrids with the overall system benefit, the individual interests of microgrid 3 are appropriately sacrificed.

[0104] like Figure 7 As shown, the game strategies of each agent obtained by the solution are verified by introducing parameters c1, c2, and c3 ∈ [0.2 3] to control each agent to change its strategy individually based on the final solution. When c1 = 1, c2 = 1, and c3 = 1, it represents the output result obtained by the reinforcement game method. The utility function of each microgrid after changing its strategy individually is as follows: Figure 6 As shown in the figure, by observing the changing trends of their respective utility functions, it is clear from the figure that the utility function values ​​of each agent are the largest when c1=1, c2=1 and c3=1. Therefore, the correctness of the obtained Nash equilibrium results can be proven.

[0105] like Figure 8 and Figure 9 As shown, compared with traditional potential game control, the reinforced game control method results in a higher degree of similarity in the output of each microgrid and better stability. Compared with the autonomous operation mode, when using traditional potential game control, the revenue of microgrid 1 remains basically unchanged, the revenue of microgrid 2 increases, and the revenue of microgrid 3 decreases, meaning that only the individual benefit of one microgrid is improved. Furthermore, the overall system revenue of traditional potential game control is 1.96 yuan lower than that of reinforced game control. Therefore, the optimization effect of the reinforced game control method is clearly superior to that of traditional potential game control.

[0106] like Figure 10 As shown, after the system adopts two different control methods at t=1s, the reinforcement game control algorithm enters the convergence range after t=1.3s, while the traditional potential game control algorithm only enters the convergence range after t=1.5s. This indicates that the reinforcement game control algorithm has better convergence and higher algorithm efficiency.

[0107] The technical solution of this invention is based on a distributed coordination architecture for multi-microgrid systems and a potential game optimization strategy, constructing a coordinated control method for multi-microgrid systems that integrates reinforcement learning and potential game theory. By fully utilizing the distributed characteristics of potential game theory, each microgrid is treated as an intelligent agent. A distributed coordination control structure is adopted to establish a potential game model with the aim of maximizing and balancing the economic efficiency of individual microgrids and the overall multi-microgrid system. Then, using the Q-learning algorithm of reinforcement learning as a carrier, the potential game theory and reinforcement learning algorithms are integrated through parameter transfer to obtain the optimal Nash equilibrium solution, thereby improving optimization performance, enhancing the economic efficiency of the multi-microgrid system, and achieving a balance between the interests of the overall system and the individuals within the system.

[0108] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A coordinated control method for multi-microgrid systems integrating Q-learning and potential game theory, characterized in that, Includes the following steps: S1. Construct a target optimization decision model for maximizing microgrid output revenue and achieving output balance under a distributed game architecture of multiple microgrids, and set power balance constraints and microgrid output constraints. S2. The local payoff function is obtained by linearly weighting the target optimization decision. Then, the global situation function and local utility function that satisfy the potential equation are designed to establish the potential game strategy set and construct a potential game model with distributed characteristics. S3. The potential game control and Q-learning algorithm are integrated through parameter passing to solve the potential game model, obtain the game optimization results, and analyze them. The specific method is as follows: (a) First, initialize the game parameters and Q-value, discretize the potential game strategy set, and pass it to the state set learned by Q. (b) Considering the rated capacity of the microgrid and to avoid excessive power fluctuations that could cause system instability, the design should be based on power variation values. The Q learning action set composed of P; (c) Collect information about neighboring microgrids, calculate the utility function of each microgrid, pass the utility function value to the immediate reward in the Q-learning algorithm, and update the Q value in the Q-learning algorithm; (d) Use a greedy strategy to select the optimal action, update the state value according to the selected action, and pass the state value to the game optimization strategy; (e) Determine if Nash equilibrium has been reached. If it has, proceed to the next step; otherwise, return to step (c). (f) Determine whether the convergence condition is met. If it is met, obtain the final microgrid output plan; otherwise, return to step (c).

2. The coordinated control method for a multi-microgrid system integrating Q-learning and potential game theory as described in claim 1, characterized in that, The method for constructing the optimization decision model described in step S1 is as follows: 1) The net benefit of maximizing the power output of the microgrid is: (1) Among them, F 1,i The net revenue from power output of the microgrid, P i For the output of microgrid i in a multi-microgrid system, For unit electricity price, m i The output cost coefficient of microgrid i; 2) Minimize the power difference between each microgrid and its neighboring microgrids in a multi-microgrid system to balance the output of each microgrid. The objective function is: (2) Among them, F 2,i I represents the power difference between microgrid i and its neighboring microgrid j. i Let P be the neighbor set of microgrid i. j The output of microgrid i is for its neighboring microgrid j.

3. The method for coordinated control of a multi-microgrid system integrating Q-learning and potential game theory as described in claim 2, characterized in that, The power balance constraints and microgrid output constraints mentioned in step S1 are as follows: (3) Among them, P load Let P be the total load of the multi-microgrid system, N be the set of potential game participants, and P be the total load of the multi-microgrid system. i,max n is the rated capacity of microgrid i; MG This refers to the number of microgrids in a multi-microgrid system.

4. The coordinated control method for a multi-microgrid system integrating Q-learning and potential game theory as described in claim 3, characterized in that, The linear weighting process described in step S2 is as follows: (4) Among them, F i (P i ,P -i Let P be the local payment function of microgrid i. -i For power output to other microgrids besides microgrid i in a multi-microgrid system, and These are the weighting coefficients for different objective functions.

5. The coordinated control method for a multi-microgrid system integrating Q-learning and potential game theory as described in claim 4, characterized in that, The global situation function described in step S2 The formula is as follows: (5) The formula for the local utility function is as follows: (6) in, F is a local utility function. j (P i ,P -i Let be the local payment function of microgrid j, a neighboring microgrid of microgrid i.

6. The coordinated control method for a multi-microgrid system integrating Q-learning and potential game theory as described in claim 5, characterized in that, The design method for the potential game strategy set mentioned in step S2 is as follows: (1) Design a potential game strategy set based on the output constraints of the microgrid. Specifically: (7) (2) The potential game strategy obtained by solving must be within the microgrid capacity limit and also satisfy the power balance constraint of the multi-microgrid system.

7. The coordinated control method for a multi-microgrid system integrating Q-learning and potential game theory as described in claim 6, characterized in that, The discrete interval length of the potential game strategy set described in step (a) for: (8) Where M is the number of intervals divided; P max and P min The upper and lower bounds of the strategy set in the power game are determined.

8. The method for coordinated control of a multi-microgrid system integrating Q-learning and potential game theory as described in claim 7, characterized in that, The formula for updating the Q-value in the Q-learning algorithm described in step (c) is as follows: (9) Among them, P i ∈A represents the action value at each step in Q-learning, α∈[0,1] is the learning rate of the Q-learning algorithm, and γ∈[0,1] is the discount parameter; This is the Q-iteration value at the (k+1)th iteration. The Q-iteration value is the value of the k-th iteration. Let be the output change value of the i-th microgrid. Let be the utility function value of the i-th microgrid at the k-th time. This represents the output change value corresponding to the maximum Q value during the k-th iteration of the i-th microgrid. For the i-th microgrid The changed output value.

9. The coordinated control method for a multi-microgrid system integrating Q-learning and potential game theory as described in claim 8, characterized in that, The formula for selecting the optimal action using a greedy strategy in step (d) is as follows: (10) in, This refers to the optimal action chosen using a greedy strategy.

Citation Information

Patent Citations

  • Potential game-based multi-unmanned aerial vehicle cooperative search method

    CN105700555A

  • Micro-grid energy management system distributed optimization algorithm based on potential game

    CN106451552A