Distributed Energy System Game Optimization Scheduling Method, System, Device and Medium

By using WoLF-PHC algorithm to construct a multi-subject game model and Q-value table in a distributed energy system, the problems of strong initial value dependence and insufficient privacy protection in traditional methods are solved, and the Nash equilibrium solution in an incomplete information environment is achieved, scheduling accuracy and privacy protection are improved, and the interests of each entity are coordinated.

CN115313520BActive Publication Date: 2025-08-01CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211128856.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-16
Publication Date
2025-08-01
Estimated Expiration
2042-09-16

AI Technical Summary

Technical Problem

In the distributed energy system, the traditional game optimization scheduling method has strong dependence on the initial value, which is easy to fall into the local optimal solution, and it is difficult to ensure the consistency of the Nash equilibrium solution in an incomplete information environment, and insufficient privacy protection.

Method used

The WoLF-PHC algorithm is used for reinforcement learning, a multi-subject game model and Q-value table is constructed, and the Nash equilibrium solution of their respective game optimization scheduling is realized through the design of the state parameters and action space of each agent, and the privacy of each subject's strategies and benefit functions is protected.

Benefits of technology

Without complete information, the coordination of the interests of various entities in the distributed energy system is achieved, the accuracy of solving scheduling problems and privacy protection are improved, load fluctuations can be suppressed, and new energy consumption can be promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115313520B_ABST
    Figure CN115313520B_ABST
Patent Text Reader

Abstract

The present invention discloses a game optimization scheduling method, system, device and medium for a distributed energy system, including: obtaining the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent and a load aggregator agent; based on the state parameters, reinforcement learning is performed to construct a multi-agent game model and a Q-value table; the WoLF-PHC algorithm is used for agent training and updating the Q-value table of each agent, and each agent obtains the Nash equilibrium solution of its respective game optimization scheduling based on the Q-value table; the Nash equilibrium solutions of their respective game optimization scheduling are output for the day-ahead optimization scheduling of each agent. The present invention can effectively improve the solution accuracy of the game optimization scheduling problem of the distributed energy system, promote the implementation of relevant artificial intelligence technologies, and drive the intelligence of power optimization scheduling decisions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power grid dispatching, and particularly relates to a game optimization dispatching method, system, device and medium for a distributed energy system. Background Art

[0002] A large number of distributed power sources and energy storage devices in the distributed energy system are invested and constructed by social capital. As independent interest entities, distributed power source operators enable various devices to participate in the system operation in an integrated form. At the same time, a large number of demand response users are integrated by load aggregators and participate in the system optimization dispatching to achieve the optimal allocation of power resources. Under the market mechanism, each entity has its own power generation and consumption demands, and there are relatively independent or even conflicting optimization goals among the entities. Therefore, it is necessary to coordinate the interests of each entity on the premise of ensuring the safe and efficient operation of the overall system.

[0003] With the gradual opening of the power grid to market competition, the entities participating in the operation of the distributed energy system are becoming increasingly diverse. Under the market mechanism, each entity has its own power generation and consumption demands, and there are relatively independent or even conflicting optimization goals among the entities in the distributed energy system. Therefore, it is necessary to coordinate the interests of each entity on the premise of ensuring the safe and efficient operation of the overall system. Game theory provides a solution to the game dispatching problem of multiple interest entities, but the solution of the game model generally adopts the mathematical derivation method and the heuristic algorithm. The mathematical derivation method has a strong dependence on the initial value and may not converge in practical applications; the heuristic algorithm is prone to falling into a local optimal solution. The multi-agent reinforcement learning algorithm organically combines the reinforcement learning method and game theory, which makes up for the limitations of the traditional method to a certain extent. Therefore, the existing technologies have the following problems:

[0004] (1) The traditional game optimization dispatching solution method has a strong dependence on the initial value, and may not converge in practical applications, or is prone to falling into a local optimum and cannot guarantee the consistency with the Nash equilibrium solution.

[0005] (2) The traditional game optimization dispatching method takes the complete information environment as a premise assumption, which is not conducive to protecting the privacy of each entity's strategy and benefit function, etc. Summary of the Invention

[0006] In order to solve the problem of multi-entity interest coordination in the distributed energy system, the present invention provides a game optimization dispatching method, system, device and medium for a distributed energy system. In the field of distributed energy system optimization dispatching, the present invention can effectively improve the solution accuracy of the game optimization dispatching problem of the distributed energy system, promote the implementation of related artificial intelligence technologies, and drive the intelligence of power optimization dispatching decisions.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A game optimization scheduling method for a distributed energy system, comprising:

[0009] Obtaining the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent;

[0010] Based on the state parameters, reinforcement learning is performed to construct a multi-agent game model and a Q-value table;

[0011] The WoLF-PHC algorithm is used to train the agents and update the Q-value tables of each agent, and each agent obtains the Nash equilibrium solution of its respective game optimization scheduling based on the Q-value table;

[0012] Output the Nash equilibrium solutions of their respective game optimization scheduling for the day-ahead optimization scheduling of each agent.

[0013] As a further improvement of the present invention, the performing reinforcement learning to construct a multi-agent game model includes: construction of a state space, an action space, and a reward function;

[0014] The joint state space at time t is expressed as:

[0015]

[0016] In the formula, P t pv , P t load and are respectively the photovoltaic power generation, the load power, and the electricity storage capacity in the system at time t; is the power of the micro gas turbine at time t-1;

[0017] The action space of the system operator agent is:

[0018]

[0019] In the formula, is the electricity selling price of the system operator to users at time t; is the electricity purchasing price of the system operator from the distributed power operator at time t;

[0020] The constraint conditions of the action space of the system operator agent are:

[0021]

[0022]

[0023] In the formula, are respectively the upper and lower limits of the electricity purchasing price at time t; They are the upper and lower limits of the electricity selling price in period t, respectively.

[0024] The action space of the distributed power generation operator agent is:

[0025]

[0026] In the formula, R t is the ramp power of the micro gas turbine in period t; represents the reactive power output of the micro gas turbine; respectively represent the active and reactive power outputs of the electrical energy storage;

[0027] The action space of the load aggregator agent only includes its load shedding power P t il , and the formula is

[0028]

[0029] The reward function of the system operator is:

[0030] r t SO = C sell (t) - C buy (t) - C grid (t) (7)

[0031] In the formula, C sell (t), C buy (t), C grid (t) are the electricity selling revenue to users, the electricity purchase cost from the distributed power generation operator, and the interaction cost with the superior power grid of the system operator, respectively;

[0032] The decision variables of the distributed power generation operator are the active and reactive power outputs of the micro gas turbine and the active and reactive power outputs of the electrical energy storage. Its optimization goal is to maximize the electricity selling revenue, and the reward function is:

[0033]

[0034] P t d = P t pv + P t mt + P t es (12)

[0035] In the formula, P t pv , P t mt , P t es are the photovoltaic power generation, the micro gas turbine power, and the electrical energy storage discharge power, respectively; Cmt (t) and C b (t) are the operating costs of the micro gas turbine and the electric energy storage respectively;

[0036] The benefit function of the load aggregator is:

[0037]

[0038] In the formula, is the user's electricity consumption utility function, representing the user's satisfaction with electricity purchase, and is simulated by the quadratic function shown in Equation (14):

[0039]

[0040] In the formula, both d and e are coefficients;

[0041] The actual load demand P t load satisfies:

[0042] P t load = P t l0 - P t il (15)

[0043] In the formula, P t l0 is the fixed load; P t il is the curtailed load, with an upper limit constraint:

[0044]

[0045] In the formula, is the maximum curtailable load.

[0046] As a further improvement of the present invention, the specific calculation methods of the said C sell (t), C buy (t), C grid (t) are:

[0047]

[0048] In the formula, P t load is the actual electricity consumption power of the user at time t;

[0049]

[0050] In the formula, P t d is the power sold by the distributed power generation operator at time t.

[0051]

[0052] In the formula, and are respectively the selling electricity price and the grid-connected electricity price of the superior power grid.

[0053] As a further improvement of the present invention, the Q-value table Q(s p , a k ) is:

[0054]

[0055] The Q-value table is a function table formed by states and actions, expressed as:

[0056] Q(s p , a k )

[0057] where the subscripts p and k respectively represent the number of states and the number of actions of the agent.

[0058] As a further improvement of the present invention, the WoLF-PHC algorithm is used to train the agents and update the Q-value table of each agent, including:

[0059] Initialize the Q-value table Q n (s, a n );

[0060] Initialize the joint state space to obtain the joint state space s0;

[0061] The system operator agent, the distributed power generation operator agent, and the load aggregator agent respectively determine their respective action spaces according to the ε-greedy strategy;

[0062] According to the decisions of each agent, the corresponding rewards are obtained from their respective reward functions, and the joint operation state s of the system at the next time period t+1 , and update the Q-value table of each agent; the maximum Q-value obtained by traversing the action space.

[0063] As a further improvement of the present invention, the following method is used to update the Q-value table of each agent:

[0064]

[0065]

[0066] In the formula, π n (s, a n ) represents the strategy of agent n, |A n | represents the number of actions of agent n, and δ represents the variable learning rate. The variable learning rate is obtained by the following method:

[0067]

[0068]

[0069] In the formula, δ w is the learning rate when the agent performs well, and δ l is the learning rate when the agent performs poorly, and δ l > δ w ; is the average policy of agent n, and C(s) represents the number of times the state s appears.

[0070] As a further improvement of the present invention, the maximum Q value obtained by traversing the action space includes:

[0071] Judge whether the current update step reaches T. If it reaches T, proceed to the next step; otherwise, return to the step of initializing the joint state space to obtain the joint state space s0.

[0072] Judge whether the current learning round reaches the maximum learning round M. If it reaches M, end the training; otherwise, return to the step of initializing the Q-value table.

[0073] Update the obtained Q-value table according to the action space and state space that reach the maximum learning round M.

[0074] As a further improvement of the present invention, each agent obtains its own Nash equilibrium solution for game optimization scheduling based on the Q-value table, including:

[0075] Each agent outputs its own Nash equilibrium strategy

[0076] As a further improvement of the present invention, the state parameters include:

[0077] The operating parameters of photovoltaic, micro gas turbine, and electric energy storage in the distributed energy system, and the usage parameters of the load.

[0078] A game optimization scheduling system for a distributed energy system, including:

[0079] An acquisition module for acquiring the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent;

[0080] A construction module for constructing a multi-agent game model and a Q-value table based on the state parameters through reinforcement learning;

[0081] An update module, which is used to train agents by using the WoLF-PHC algorithm and update the Q-value tables of each agent, and each agent obtains the Nash equilibrium solution of its game optimization scheduling based on the Q-value table;

[0082] An output module, which is used to output the Nash equilibrium solution of each game optimization scheduling for the day-ahead optimization scheduling of each agent.

[0083] As a further improvement of the present invention, in the construction module, the construction of the multi-agent game model by reinforcement learning includes: the construction of the state space, the action space, and the reward function;

[0084] The joint state space at time t is expressed as:

[0085]

[0086] In the formula, P t pv , P t load and are respectively the photovoltaic power generation power, the load power, and the electricity storage capacity of the electricity storage in the system at time t; is the power of the micro gas turbine at time t-1;

[0087] The action space of the system operator agent is:

[0088]

[0089] In the formula, is the electricity selling price of the system operator to users at time t; is the electricity purchasing price of the system operator from the distributed power generation operator at time t;

[0090] The constraint conditions of the action space of the system operator agent are:

[0091]

[0092]

[0093] In the formula, are respectively the upper and lower limits of the electricity purchasing price at time t; are respectively the upper and lower limits of the electricity selling price at time t;

[0094] The action space of the distributed power generation operator agent is:

[0095]

[0096] In the formula, R t is the ramp power of the micro gas turbine at time t; represents the reactive power output of the micro gas turbine; respectively represent the active and reactive power outputs of the electrical energy storage;

[0097] The action space of the load aggregator agent only contains its load shedding power P t il , and the method is as follows:

[0098]

[0099] The reward function of the system operator is:

[0100] r t SO = C sell (t) - C buy (t) - C grid (t) (7)

[0101] In the formula, C sell (t), C buy (t), C grid (t) are respectively the electricity selling revenue of the system operator to users, the electricity purchase cost from the distributed power generation operator, and the interaction cost with the superior power grid;

[0102] The decision variables of the distributed power generation operator are the active and reactive power outputs of the micro gas turbine and the active and reactive power outputs of the electrical energy storage. The optimization goal is to maximize the electricity selling revenue, and the reward function is:

[0103]

[0104] P t d = P t pv + P t mt + P t es (12)

[0105] In the formula, P t pv , P t mt , P t es are respectively the photovoltaic power generation power, the micro gas turbine power, and the electrical energy storage discharge power; C mt (t) and C b (t) are respectively the operating costs of the micro gas turbine and the electrical energy storage;

[0106] The benefit function of the load aggregator is:

[0107]

[0108] In the formula, It is the user's electricity consumption utility function, representing the user's satisfaction with electricity purchase, and is simulated by a quadratic function as shown in Equation (14):

[0109]

[0110] In the formula, both d and e are coefficients;

[0111] Actual load demand Satisfy:

[0112] P t load = P t l0 - P t il (15)

[0113] In the formula, P t l0 is the fixed load; P t il is the curtailed load, with an upper limit constraint:

[0114]

[0115] In the formula, is the maximum curtailable load.

[0116] As a further improvement of the present invention, in the update module, the WoLF-PHC algorithm is used to train the agents and update the Q-value tables of each agent, including:

[0117] Initialize the Q-value table Q n (s,a n );

[0118] Initialize the joint state space to obtain the joint state space s0;

[0119] The system operator agent, the distributed power generation operator agent, and the load aggregator agent respectively determine their respective action spaces according to the ε-greedy strategy;

[0120] According to the decisions of each agent, obtain the corresponding rewards from their respective reward functions, and the joint operation state s of the system at the next time period t+1 , and update the Q-value tables of each agent; the maximum Q-value obtained by traversing the action space.

[0121] As a further improvement of the present invention, in the update module, each agent obtains the Nash equilibrium solution of its respective game optimization scheduling based on the Q-value table, including:

[0122] Each agent outputs its respective Nash equilibrium strategy

[0123] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the game optimization scheduling method for the distributed energy system are implemented.

[0124] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the game optimization scheduling method for the distributed energy system are implemented.

[0125] Compared with the prior art, the present invention has the following beneficial effects:

[0126] The game optimization scheduling method for the distributed energy system based on WoLF-PHC of the present invention solves the problem of multi-agent interest coordination in the distributed energy system. Each agent constructed based on the WoLF-PHC method can achieve the solution of the Nash equilibrium in a non-complete information game environment without obtaining the strategy space and benefit function of other agents, by continuously exploring the operation state of the distributed energy system. Therefore, this method can effectively protect the privacy of each agent's strategy and benefit function. Moreover, this method has high application value in terms of solution accuracy. The present invention introduces reinforcement learning technology and game theory into the distributed energy system, and this optimization scheduling method can coordinate the interests of each participating agent in the system.

[0127] Furthermore, the multi-agent training method based on WoLF-PHC enables each agent to solve the optimization scheduling problem of the distributed energy system through repeated exploration and trial-and-error in an incomplete information environment.

[0128] Furthermore, the constructed multi-agent game model can guide the output of distributed power sources and adjust the user's energy consumption plan through price signals, which is beneficial to suppressing load fluctuations and promoting the consumption of new energy. BRIEF DESCRIPTION OF THE DRAWINGS

[0129] Figure 1 is a flowchart of a game optimization scheduling method for a distributed energy system of the present invention;

[0130] Figure 2 is a framework diagram of the game optimization scheduling based on WoLF-PHC constructed by the present invention;

[0131] Figure 3 is the game optimization scheduling algorithm flow based on WoLF-PHC;

[0132] Figure 4 is a game optimization scheduling system for a distributed energy system provided by the present invention;

[0133] Figure 5 is a schematic diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0134] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0135] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0136] In the variable learning rate and policy hill climbing (WoLF-PHC) algorithm, each agent can learn and converge to an optimal policy relative to the policies of other agents by updating its own Q function, and this policy is the Nash equilibrium solution. This method has achieved good convergence results in practical applications.

[0137] To solve the problem of multi-agent interest coordination in a distributed energy system, the present invention provides a game optimization scheduling method for a distributed energy system based on WoLF-PHC. This method realizes the solution of the game equilibrium strategy in a non-complete information game environment where each agent does not need to obtain the strategies of other agents.

[0138] As Figure 1 shown, a game optimization scheduling method for a distributed energy system proposed by the present invention includes:

[0139] Obtain the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent;

[0140] Based on the state parameters, perform reinforcement learning to construct a multi-agent game model and a Q-value table;

[0141] The WoLF-PHC algorithm is used to train the agents and update the Q-value tables of each agent. Each agent obtains the Nash equilibrium solution of its game optimization scheduling based on the Q-value table;

[0142] Output the Nash equilibrium solution of each agent's game optimization scheduling for the day-ahead optimal scheduling of each agent.

[0143] This method first models each game participant as an agent and constructs a multi-agent game model including a system operator agent, a distributed power generation operator agent, and a load aggregator agent. Then, it designs an agent training process based on the WoLF-PHC method. Finally, each agent can perform day-ahead optimal scheduling according to the trained Q-value table to obtain the Nash equilibrium solution.

[0144] A game optimization scheduling method for a distributed energy system based on WoLF-PHC of the present invention particularly relates to the field of distributed energy system optimization scheduling. Each stakeholder can achieve the solution of the Nash equilibrium through continuous exploration of the operation state of the distributed power system by itself in a non-complete information game environment without obtaining the strategy space and benefit function of other agents, and has high application value in terms of solution accuracy.

[0145] The present invention realizes the above object of the technical solution through steps Step 0 to Step 9:

[0146] Step 0: Obtain the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power generation operator agent, and a load aggregator agent;

[0147] The state parameters include: the operation parameters of photovoltaic, micro gas turbine, and electrical energy storage in the distributed energy system, and the usage parameters of the load.

[0148] Step 1: First, construct a reinforcement learning model, mainly including the construction of the state space, action space, and the design of the reward function.

[0149] 1) State space

[0150] In the optimization scheduling method based on WoLF-PHC, each agent makes decisions by observing the joint state space. This joint state space includes the operation states of various devices in the system, so the joint state space at time t is expressed as:

[0151]

[0152] In the formula, P t n,pv , P t n,load And They are the photovoltaic power generation, load power, and electricity storage power in the system during period t, respectively. is the power of the micro gas turbine during period t-1.

[0153] 2) Action space

[0154] The action space of each agent is the relevant decision variable. The action space of the system operator agent is set as:

[0155]

[0156] In the formula, is the electricity selling price from the system operator to users during period t; is the electricity purchasing price from the distributed power generation operator by the system operator during period t.

[0157] In addition, constraints as shown in formulas (3) and (4) need to be set for the electricity purchasing and selling prices to avoid the distribution network maliciously reducing the electricity purchasing price or increasing the electricity selling price to improve its own benefits.

[0158]

[0159]

[0160] In the formula, are the upper and lower limits of the electricity purchasing price during period t, respectively; are the upper and lower limits of the electricity selling price during period t, respectively.

[0161] The action space of the distributed power generation operator agent is set as:

[0162]

[0163] In the formula, R t is the ramp power of the micro gas turbine during period t; represents the reactive power output of the micro gas turbine; represent the active and reactive power outputs of the electricity storage, respectively.

[0164] ]>The action space of the load aggregator agent only includes its load shedding power P t il .

[0165]

[0166] 3) Reward function

[0167] The reward function of the system operator is:

[0168] r t SO = C sell (t)-C buy (t)-C grid(t) (7)

[0169] Wherein, C sell (t), C buy (t), C grid (t) are respectively the electricity selling revenue from the system operator to the user, the electricity purchasing cost from the distributed power generation operator, and the cost of interacting with the superior power grid. The specific expressions are shown in formulas (8) to (10):

[0170]

[0171] Wherein, P t load is the actual electricity consumption power of the user at time t.

[0172]

[0173] Wherein, P t d is the sold power of the distributed power generation operator at time t.

[0174]

[0175] Wherein, and are respectively the electricity selling price of the superior power grid and the grid connection price.

[0176] The decision variables of the distributed power generation operator are the active and reactive power outputs of the micro gas turbine and the active and reactive power outputs of the electrical energy storage. Its optimization goal is to maximize the electricity selling revenue, and the reward function is:[[]]

[0177]

[0178]

[0179] Wherein, P t pv , P t n,mt , P t n,es are respectively the photovoltaic power generation, the micro gas turbine power, and the electrical energy storage discharge power; C mt (t) and C b (t) are respectively the operating costs of the micro gas turbine and the electrical energy storage.

[0180] The users participating in the demand response maximize the consumer surplus by adjusting the curtailable load power. The consumer surplus is expressed as the difference between the user's electricity consumption utility and the electricity purchasing cost. The benefit function of the load aggregator is:[[]]

[0181]

[0182] Wherein, It is the electricity consumption utility function of users, representing the satisfaction of users with electricity purchase, and is simulated by the quadratic function shown in Equation (14):

[0183]

[0184] In the formula, both d and e are coefficients.

[0185] The actual load demand P t load Satisfies:

[0186] P t load =P t l0 -P t il (15)

[0187] In the formula, P t l0 Is the fixed load; P t il Is the curtailed load, with an upper limit constraint:

[0188]

[0189] In the formula, Is the maximum curtailable load.

[0190] Step 2: Construct a game optimization scheduling framework based on the WoLF-PHC algorithm, as Figure 1 Shown. Model each stakeholder as an agent. The system operator, distributed generation operator, and load aggregator correspond to the SO agent, DGO agent, and LA agent respectively. Based on Step 1, design the joint state space, action space, and reward function for each agent, and update the Q-value table of each agent with the help of the WoLF-PHC algorithm. Each stakeholder obtains the Nash equilibrium solution of the game optimization scheduling based on this table.

[0191] The Q-value table is shown in Table 1 below.

[0192] Table 1 Q-value table

[0193]

[0194]

[0195] In the table, the subscripts p and k represent the number of states and the number of optional actions of the agent respectively.

[0196] Step 3: Initialize the Q-value table, set all elements in the Q-value table of each agent to 0; initialize the strategy π of each agent n (s,a n ) and the average strategy Let Let C(s) be 0;

[0197] Step 4: Initialize the combined state space s0 shown in Equation (1).

[0198] Step 5: The SO agent, DGO agent, and LA agent respectively determine the actions shown in Equations (2), (5), and (6) according to the ε-greedy policy, that is, the agent randomly selects an action from the set of optional actions with a probability of ε, and selects the action that can maximize the Q value with a probability of 1-ε.

[0199] Step 6: Determine the rewards shown in Equations (11) to (13) and the combined operating state s of the system in the next time period according to the decisions of each agent t+1 , and update the Q value tables of each agent according to Equations (17) to (20):

[0200]

[0201]

[0202]

[0203]

[0204] where π n (s,a n ) represents the policy of agent n, |A n | represents the number of actions of agent n, δ represents the variable learning rate, δ w is the learning rate when the agent performs well, δ l is the learning rate when the agent performs poorly, and δ l >δ w , is the average policy of agent n, and C(s) represents the number of times the state s appears.

[0205] Step 7: Determine whether the number of update steps has reached T. If it has reached T, go to Step 8; otherwise, return to Step 4.

[0206] Step 8: Determine whether the maximum number of learning rounds M has been reached. If it has reached M, end the training and go to Step 9; otherwise, return to Step 3.

[0207] Step 9: Update the obtained Q value tables according to Steps 3 to 8, and each agent outputs its own Nash equilibrium policy

[0208] As Figure 4 shown, the present invention also provides a game optimization scheduling system for a distributed energy system, including:

[0209] An acquisition module for acquiring the state parameters of each agent in a distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent;

[0210] A construction module for constructing a multi-agent game model and a Q-value table through reinforcement learning based on the state parameters;

[0211] An update module for training the agents using the WoLF-PHC algorithm and updating the Q-value table of each agent, and each agent obtains the Nash equilibrium solution of its respective game-optimized scheduling based on the Q-value table;

[0212] An output module for outputting the Nash equilibrium solutions of their respective game-optimized scheduling for the day-ahead optimal scheduling of each agent.

[0213] Among them, in the construction module, the construction of the multi-agent game model through reinforcement learning includes: the construction of the state space, the action space, and the reward function;

[0214] 1) State space

[0215] The combined state space at time t is expressed as:

[0216]

[0217] Where P t n,pv , P t n,load and are respectively the photovoltaic power generation, load power, and electrical energy storage capacity in the system at time t; is the power of the micro gas turbine at time t-1;

[0218] 2) Action space

[0219] The action space of the system operator agent is:

[0220]

[0221] Where is the electricity selling price of the system operator to users at time t; is the electricity purchasing price of the system operator from the distributed power operator at time t;

[0222] The constraint conditions of the action space of the system operator agent are:

[0223]

[0224]

[0225] Where are the upper and lower limits of the electricity purchase price during period t respectively; are the upper and lower limits of electricity sales price in period t respectively;

[0226] The action space of the distributed power operator agent is:

[0227]

[0228] Where R t is the ramp power of the micro gas turbine during period t; Indicates the reactive output of the micro gas turbine; Respectively represent the active and reactive output of electric energy storage;

[0229] The action space of the load aggregator agent only contains its load reduction power The formula is

[0230]

[0231] 3) Reward Function

[0232] The system operator reward function is:

[0233] r t SO =C sell (t)-C buy (t)-C grid (t) (7)

[0234] Where C sell (t), C buy (t), C grid (t) are the revenue from electricity sales to users, the cost of electricity purchase from distributed generation operators, and the cost of interaction with the upper-level power grid;

[0235] The decision variables of the distributed power generation operator are the active and reactive output of the micro gas turbine and the active and reactive output of the electric energy storage. The optimization goal is to maximize the revenue from electricity sales. The reward function is:

[0236]

[0237] P t d =P t pv +P t mt +P t es (12)

[0238] Where, P t pv 、P t n,mt 、Pt n,es are the photovoltaic power generation, the micro gas turbine power, and the electric energy storage discharge power respectively; C mt (t) and C b (t) are the operating costs of the micro gas turbine and the electric energy storage respectively;

[0239] The benefit function of the load aggregator is:

[0240]

[0241] In the formula, f u t is the user's electricity consumption utility function, representing the user's satisfaction with electricity purchase, and is simulated by the quadratic function shown in Equation (14):

[0242]

[0243] In the formula, d and e are both coefficients;

[0244] The actual load demand P t load satisfies:

[0245] P t load = P t l0 - P t il (15)

[0246] In the formula, P t l0 is the fixed load; P t il is the curtailed load, with an upper limit constraint:

[0247]

[0248] In the formula, is the maximum curtailable load.

[0249] In the update module, the WoLF-PHC algorithm is used to train the agents and update the Q-value tables of each agent, including:

[0250] Initialize the Q-value table Q n (s, a n ), and set all elements in the Q-value table of each agent to 0; Initialize each agent's policy π n (s, a n ) and the average policy Let Let C(s) be 0;

[0251] Initialize the joint state space to obtain the joint state space s0;

[0252] The system operator agent, the distributed power operator agent, and the load aggregator agent respectively determine their respective action spaces according to the ε-greedy strategy;

[0253] According to the decisions of each agent, the corresponding rewards are obtained by their respective reward functions, and the joint operation state s of the system in the next time period t+1 , and the Q-value tables of each agent are updated according to the formula; the maximum Q-value obtained by traversing the action space.

[0254] Each of the agents obtains the Nash equilibrium solution of its respective game optimization scheduling based on the Q-value table, including:

[0255] Each agent outputs its own Nash equilibrium strategy

[0256] As Figure 5 shown, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the game optimization scheduling method for the distributed energy system are implemented.

[0257] The game optimization scheduling method for the distributed energy system includes the following steps:

[0258] Obtain the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent;

[0259] Based on the state parameters, reinforcement learning is performed to construct a multi-agent game model and a Q-value table;

[0260] The WoLF-PHC algorithm is used for agent training and the Q-value tables of each agent are updated. Each agent obtains the Nash equilibrium solution of its respective game optimization scheduling based on the Q-value table;

[0261] Output the Nash equilibrium solutions of their respective game optimization scheduling for the day-ahead optimization scheduling of each agent.

[0262] The present invention also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the game optimization scheduling method for the distributed energy system are implemented.

[0263] The game optimization scheduling method for the distributed energy system includes the following steps:

[0264] Obtain the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent;

[0265] Based on the state parameters, reinforcement learning is performed to construct a multi-agent game model and a Q-value table;

[0266] The WoLF-PHC algorithm is used to train the agents and update the Q-value tables of the agents. Each agent obtains the Nash equilibrium solution of its game optimization scheduling based on the Q-value table;

[0267] Output the Nash equilibrium solutions of their respective game optimization scheduling for the day-ahead optimization scheduling of each agent.

[0268] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0269] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in one Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0270] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in one Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.

[0271] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for realizing the functions specified in one Figure 1 one or more flows and / or blocks Figure 1Steps of the functions specified in one or more boxes.

[0272] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific implementation manners of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the claims of the present invention.

Claims

1. A game-optimized scheduling method for a distributed energy system, characterized in that Including: Obtain the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent; Based on the state parameters, perform reinforcement learning to construct a multi-agent game model and a Q-value table; Use the WoLF-PHC algorithm to train the agents and update the Q-value tables of each agent. Each agent obtains the Nash equilibrium solution of its game-optimized scheduling based on the Q-value table; Output the Nash equilibrium solutions of their respective game-optimized scheduling for the day-ahead optimal scheduling of each agent; The performing reinforcement learning to construct a multi-agent game model includes: the construction of the state space, the action space, and the reward function; t The time period combined state space is expressed as: (1) In the formula, , and are respectively t the photovoltaic power generation power, the load power and the electricity storage capacity of the electricity storage in the system during the is t the micro gas turbine power at the -1 time period. The action space of the system operator agent is: (2) In the formula, is t the electricity selling price from the time-period system operator to the user; is t the electricity purchasing price from the distributed power operator by the time-period system operator; The constraint conditions of the action space of the system operator agent are: (3) (4) Wherein, and are respectively t the upper and lower limits of the electricity purchase price for the and are respectively t the upper and lower limits of the electricity selling price for the The action space of the distributed power operator agent is: (5) In the formula, is t the ramp power of the micro gas turbine during a period; represents the reactive power output of the micro gas turbine; , respectively represent the active and reactive power outputs of the electrical energy storage; The action space of the load aggregator agent only contains its load shedding power , and the method is as follows: (6) The reward function of the system operator is: (7) In the formula, , , are respectively the electricity sales revenue from users by the system operator, the electricity purchase cost from the distributed power generation operator, and the cost of interacting with the superior power grid; The decision variables of the distributed power operator are the active and reactive power outputs of the micro gas turbine and the active and reactive power outputs of the electrical energy storage. The optimization goal is to maximize the electricity sales revenue, and the reward function is: (11) (12) In the formula, , , are the photovoltaic power generation power, the micro gas turbine power, and the electric energy storage discharge power respectively; and are the operating costs of the micro gas turbine and the electric energy storage respectively; The benefit function of the load aggregator is: (13) In the formula, is the user's electricity consumption utility function, representing the user's satisfaction with electricity purchase, and is simulated by the quadratic function shown in Equation (14): (14) wherein, and are both coefficients; Actual load demand Meet: (15) In the formula, is the fixed load; is the curtailed load, with an upper limit constraint: (16) In the formula, is the maximum load that can be reduced.

2. The game optimization scheduling method for a distributed energy system according to claim 1, characterized in that The said , , The specific calculation method is as follows: (8) Wherein, is t the actual power consumption of the user during the period; (9) Wherein, is t the power sold by the distributed power operator in a time period. (10) In the formula, and are the selling electricity price and the on-grid electricity price of the superior power grid, respectively.

3. The game-optimized scheduling method for a distributed energy system according to claim 1, wherein The Q-value table is a function table formed by states and actions, expressed as: Among them, p and k represent the number of states and actions of the agent respectively.

4. The game optimization scheduling method for a distributed energy system according to claim 1, wherein The using the WoLF-PHC algorithm to train the agents and update the Q-value tables of each agent includes: Initialize the Q-value table ; Initialize the joint state space to obtain the joint state space ; The system operator agent, the distributed power operator agent, and the load aggregator agent respectively determine their respective action spaces according to the greedy strategy; According to the decisions of each agent, the corresponding rewards are obtained by their respective reward functions, as well as the combined operation state of the system in the next time period , and update the Q-value tables of each agent; the maximum Q-value obtained by traversing the action space.

5. The game optimization scheduling method for a distributed energy system according to claim 4, wherein The updating the Q-value tables of each agent adopts the following method: (17) (18) In the formula, represents the agent n policy, represents the number of actions of the agent n , represents a variable learning rate, and the variable learning rate is obtained by the following method: (19) (20) Wherein, is the learning rate when the agent performs well, is the learning rate when the agent performs poorly, and ; is the average policy of the agent n , and C( s ) represents the number of times the state s appears.

6. The game-optimized scheduling method for a distributed energy system according to claim 4, wherein The maximum Q-value obtained by traversing the action space includes: Determine whether the current update step reaches T. If it reaches T, proceed to the next step; otherwise, return the initialized joint state space to obtain the joint state space Step; Judge whether the current learning round reaches the maximum learning round M; if it reaches M, end the training, otherwise return to the step of initializing the Q-value table; Update the obtained Q-value table according to the action space and state space that reach the maximum learning round M.

7. The game optimization scheduling method for a distributed energy system according to claim 1, wherein The each agent obtaining the Nash equilibrium solution of its game-optimized scheduling based on the Q-value table includes: Each agent outputs its own Nash equilibrium strategy 。 8. The game optimization scheduling method for a distributed energy system according to claim 1, wherein The state parameters include: The operating parameters of the photovoltaic, micro gas turbine, and electrical energy storage in the distributed energy system, and the usage parameters of the load.

9. A game optimization scheduling system for a distributed energy system, characterized in that Including: An acquisition module, used to obtain the state parameters of each agent in the distributed energy system; each agent includes a system operator agent, a distributed power operator agent, and a load aggregator agent; A construction module, used to perform reinforcement learning to construct a multi-agent game model and a Q-value table based on the state parameters; An update module, used to use the WoLF-PHC algorithm to train the agents and update the Q-value tables of each agent. Each agent obtains the Nash equilibrium solution of its game-optimized scheduling based on the Q-value table; An output module, used to output the Nash equilibrium solutions of their respective game-optimized scheduling for the day-ahead optimal scheduling of each agent; In the construction module, the performing reinforcement learning to construct a multi-agent game model includes: the construction of the state space, the action space, and the reward function; t The time period combined state space is expressed as: (1) In the formula, , and are respectively t the photovoltaic power generation power, the load power, and the electricity storage capacity of the electricity storage in the system during the period; is t the power of the micro gas turbine at the -1 period; The action space of the system operator agent is: (2) Wherein, is t the electricity selling price of the time period system operator to the user; is t the electricity purchasing price of the time period system operator from the distributed power operator; The constraint conditions of the action space of the system operator agent are: (3) (4) Wherein, and are respectively t the upper and lower limits of the electricity purchase price for the and are respectively t the upper and lower limits of the electricity selling price for the The action space of the distributed power operator agent is: (5) In the formula, is t the ramp power of the micro gas turbine during a time period; represents the reactive power output of the micro gas turbine; , respectively represent the active and reactive power outputs of the electrical energy storage; The action space of the load aggregator agent only contains its load shedding power , and the method is as follows: (6) The reward function of the system operator is: (7) Wherein, , , are respectively the electricity sales revenue from users by the system operator, the electricity purchase cost from the distributed power generation operator, and the cost of interacting with the superior power grid; The decision variables of the distributed power operator are the active and reactive power outputs of the micro gas turbine and the active and reactive power outputs of the electrical energy storage. The optimization goal is to maximize the electricity sales revenue, and the reward function is: (11) (12) Wherein, , , are the photovoltaic power generation power, the micro gas turbine power and the electric energy storage discharge power respectively; and are the operating costs of the micro gas turbine and the electric energy storage respectively; The benefit function of the load aggregator is as follows: (13) In the formula, is the user's electricity consumption utility function, representing the user's satisfaction with electricity purchase, and is simulated by a quadratic function as shown in Equation (14): (14) wherein, , are all coefficients; Actual load demand Satisfy: (15) In the formula, is the fixed load; is the curtailed load, with an upper limit constraint: (16) In the formula, is the maximum load that can be reduced.

10. The game optimization scheduling system of the distributed energy system according to claim 9, characterized in that, In the updating module, the WoLF-PHC algorithm is used to train the agents and update the Q-value tables of the agents, including: Initialize the Q-value table ; Initialize the joint state space to obtain the joint state space ; The system operator agent, the distributed power operator agent, and the load aggregator agent respectively determine their respective action spaces according to the greedy strategy; According to the decisions of each agent, the corresponding rewards are obtained by their respective reward functions, as well as the combined operating state of the system in the next time period , and update the Q-value tables of each agent; the maximum Q-value obtained by traversing the action space.

11. The game-optimized scheduling system for a distributed energy system according to claim 9, characterized in that In the updating module, each agent obtains its Nash equilibrium solution for game-optimized scheduling based on the Q-value table, including: Each agent outputs its own Nash equilibrium strategy .

12. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the distributed energy system game-optimized scheduling method according to any one of claims 1-8.

13. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the distributed energy system game-optimized scheduling method according to any one of claims 1-8.

Citation Information

Patent Citations

  • An intelligent generation control method based on deep reinforcement learning with self-optimizing ability

    CN109217306A

  • Multi-agent power generation optimal scheduling method based on reinforcement learning

    CN110728406A