Power grid dynamic management and control method, device and equipment based on multi-market subject game
By constructing a multi-market player hybrid game model and the MADDPG algorithm, dynamic collaborative game is carried out between power generation companies and independent system operators, which solves the problem of inaccurate decision-making in the power market and achieves stable market operation and efficient allocation of resources.
Patent Information
- Application Number
- CN202510624614.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies make it difficult to effectively consider the impact of the carbon trading market in the electricity market, resulting in insufficient decision-making accuracy for power generation companies. Traditional methods suffer from instability and incomplete information in multi-agent interaction scenarios.
A multi-market player hybrid game model is constructed, and the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is used in combination with the Stackelberg game relationship to conduct dynamic collaborative games between power generation companies and independent system operators to formulate power supply strategies and market clearing prices, and optimize profit maximization and electricity purchase costs.
It has improved the decision-making accuracy and adaptability of power generation companies in the power market, optimized resource allocation, and achieved an effective balance between stable market operation and profits.
Smart Images

Figure CN120672365A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of power system management, and in particular to a method, device and equipment for dynamic power grid control based on multi-market player game. Background Art
[0002] Against the backdrop of global efforts to address climate change and pursue the "dual carbon" goals, the power industry is undergoing profound transformation. Building a new power system dominated by clean energy and promoting the large-scale development of renewable energy have become key pathways to achieving a sustainable energy transition. However, current research on the participation of renewable energy in power market transactions remains deficient. Some studies have solely used the forecasted output of renewable energy as a constraint on market clearing, ignoring the crucial role of carbon trading markets (CET markets). Furthermore, research examining the impact of CET markets on power markets often fails to fully consider the participation of renewable energy.
[0003] Traditional methods for studying bidding strategies for power suppliers each have their limitations. Cost analysis fails to fully consider market supply and demand and the decisions of other suppliers, making it difficult to maximize their own interests. Clearing price forecasting and competitor bid analysis rely on large amounts of historical data. In the early stages of power market reform, data scarcity and constantly changing market structure rules made accurate price forecasts difficult. Game theory analysis is ineffective when dealing with multi-person and incomplete information game problems. Traditional intelligent optimization algorithms, due to the uncertainty of game outcomes, face problems such as environmental instability and increased variance in multi-agent interaction scenarios, which affect the accuracy of power generation companies' decisions. Therefore, there is an urgent need to provide a solution to improve the accuracy of power generation companies' decisions. Summary of the Invention
[0004] The embodiments of the present application provide a method, device and equipment for dynamic power grid control based on multi-market player game to solve the problem of how to improve the decision-making accuracy of power generation enterprises.
[0005] In a first aspect, an embodiment of the present application provides a method for dynamic power grid management and control based on multi-market player game, including:
[0006] Obtaining market demand data and price fluctuation data, and obtaining a pre-built multi-agent hybrid game model, wherein the multi-agent hybrid game model includes a bidding decision layer and a market clearing layer; the multi-agent hybrid game model has power generation enterprises as leaders and independent system operators as followers;
[0007] At the bidding decision-making level, each power generation company selects a corresponding game model based on market demand data and price fluctuation data, and formulates a power supply strategy with the goal of maximizing profits. The game models include cooperative and non-cooperative game models. Each thermal power company formulates a power supply strategy based on the power market supply-demand ratio and the thermal power company's market share based on the Multi-agent Deep Deterministic Policy Gradient (MADDPG) algorithm. The power supply strategy includes: declared power consumption and declared electricity price.
[0008] In the market clearing layer, the independent system operator aims to minimize the cost of purchasing electricity and determines the planned power generation and market clearing price of each power generation enterprise based on load constraints and unit constraints; among which, the unit constraints are determined according to the unit type corresponding to the power generation enterprise.
[0009] In a possible implementation, each power generation enterprise selects a corresponding game mode according to market demand data and price fluctuation data, including:
[0010] When market demand is higher than the set demand range and price fluctuations exceed the set fluctuation range, the cooperative game mode is selected;
[0011] When market demand is lower than the set demand range and price fluctuations are within the set fluctuation range, a non-cooperative game mode is selected.
[0012] In one possible implementation, each power generation enterprise selects a corresponding game model based on market demand data and price fluctuation data, and formulates a power supply strategy with the goal of maximizing profits, including:
[0013] In the cooperative game model, each power generation company formulates its power supply strategy with the goal of maximizing its own profits;
[0014] Under the non-cooperative game model, each power generation company formulates a power supply strategy with the goal of maximizing alliance profits.
[0015] In a possible implementation, in the non-cooperative game mode, the Shapley value method is used to distribute alliance profits according to the contribution of each enterprise to the alliance.
[0016] In one possible implementation, the MADDPG algorithm includes an Actor strategy network model and a Critic value network model;
[0017] Each thermal power company formulates a power supply strategy based on the MADDPG algorithm, according to the power market supply and demand ratio and the thermal power company's market share, including:
[0018] The power market supply-demand ratio and the market share of thermal power companies are input into the Actor strategy network model, and the expected power supply strategy of the thermal power companies is output; the expected power supply strategy includes: expected declared electricity volume and expected declared electricity price; wherein, it is expressed as a declared electricity volume decision coefficient and a declared electricity price decision coefficient, and optionally, the declared electricity volume decision coefficient ranges from [0,1], and the declared electricity price decision coefficient ranges from [1,1.2].
[0019] The power market supply-demand ratio and the market share of thermal power companies are input into the Critic value network model to output a state value function that evaluates the quality of the power supply expectation strategy;
[0020] A power supply strategy for the thermal power company is formulated according to the state value function.
[0021] In a possible implementation, the cooperative game model corresponding to the cooperative game mode is:
[0022]
[0023] Among them, Π C is the total profit of the cooperative alliance C, γ C and δ C is the coefficient, ΔΠ C (t) is the change in alliance profit;
[0024] The non-cooperative game model corresponding to the non-cooperative game mode is:
[0025]
[0026] Among them, Π j is the profit of power generation enterprise j, P jt is the power generation at time t, B jt is the bid price at time t, C j (P jt ) is the power generation cost function, α j and β j is the coefficient, I jk (t) is the information interaction impact value between enterprises j and k at time t, ΔP jt is the change in power generation.
[0027] In a possible implementation, the unit constraints include: output limits and bidding quantity constraints of thermal power units and wind and solar power stations, thermal power unit ramp rate constraints, and energy storage system operation constraints;
[0028] The load constraint conditions include: day-ahead power market supply and demand balance constraints and bidding price constraints.
[0029] In one possible implementation, the day-ahead power market supply and demand balance constraint is:
[0030]
[0031] in, are the output of thermal power, wind power and photovoltaic power at time t, D t is the load demand at time t;
[0032] The output limit and bidding quantity constraints of thermal power units are as follows:
[0033]
[0034] in, are the minimum and maximum output of thermal power units, are the minimum and maximum bidding prices of thermal power units respectively;
[0035] The output limits and bidding constraints for wind power and photovoltaic power stations are as follows:
[0036]
[0037] in, are the predicted outputs of wind power and photovoltaic power respectively, The output after adjustment of the energy storage systems equipped for wind power and photovoltaic power, are the minimum and maximum bidding prices for wind power and PV, respectively;
[0038] The thermal power unit ramp rate constraint is:
[0039]
[0040] Among them, V t,domn is the downward climbing rate, V t,up Δt is the upward climbing rate, Δt is the time interval;
[0041] The energy storage system operation constraints are:
[0042]
[0043] SOC min E max ≤E r ≤SOC max E max
[0044]
[0045] in, are the charging and discharging power of the energy storage system, is the maximum charge and discharge power of the energy storage system, SOC min , SOC maxare the minimum and maximum charge states of the energy storage system, E r The amount of electricity stored in the energy storage system, E t is the amount of energy storage system at the moment, θ c ,θ d are the charging and discharging efficiency of the energy storage system respectively.
[0046] In a second aspect, an embodiment of the present application provides a power grid dynamic management and control device based on multi-market player game, including:
[0047] An acquisition module is used to acquire market demand data and price fluctuation data, and acquire a pre-built multi-agent hybrid game model, wherein the multi-agent hybrid game model includes a bidding decision layer and a market clearing layer;
[0048] A game optimization module is used at the bidding decision-making level for each power generation enterprise to select a corresponding game mode based on market demand data and price fluctuation data, and to formulate a power supply strategy with the goal of maximizing profits. The game modes include cooperative game mode and non-cooperative game mode. Each thermal power enterprise formulates a power supply strategy based on the power market supply and demand ratio and the thermal power enterprise's market share based on the MADDPG algorithm. The power supply strategy includes: declared power consumption and declared electricity price.
[0049] In the market clearing layer, the independent system operator aims to minimize the cost of purchasing electricity and determines the planned power generation and market clearing price of each power generation enterprise based on load constraints and unit constraints; among which, the unit constraints are determined according to the unit type corresponding to the power generation enterprise.
[0050] In a third aspect, an embodiment of the present application provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method in the first aspect or any possible implementation of the first aspect is implemented.
[0051] In the embodiment of the present application, a multi-agent hybrid game model including a bidding decision layer and a market clearing layer is constructed to dynamically optimize the operation of the electricity market. At the bidding decision layer, power generation companies choose cooperative or non-cooperative game modes based on market demand and price fluctuations, and use the MADDPG algorithm to accurately formulate strategies for declared electricity volume and electricity prices, thereby improving the decision-making efficiency and adaptability of thermal power companies. At the market clearing layer, independent system operators determine power generation plans and clearing prices with the goal of minimizing electricity purchase costs, taking into account load and unit constraints. The embodiment of the present application coordinates the interests of market entities through a game mechanism. At the same time, based on the intelligent optimization of electricity prices and benefits by the MADDPG algorithm, it effectively balances supply and demand, optimizes resource allocation, and improves the accuracy of power generation company decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1This is a flowchart of an implementation method of a power grid dynamic control method based on multi-market subject game provided by an embodiment of the present application;
[0053] Figure 2 A market mechanism diagram of a power grid dynamic control method based on multi-market player game provided in an embodiment of the present application;
[0054] Figure 3 This is a schematic diagram of the interaction process between the agent and the environment in the MADDPG algorithm reinforcement learning provided in an embodiment of the present application;
[0055] Figure 4 The implementation framework of the MADDPG algorithm provided in the embodiments of this application is provided;
[0056] Figure 5 The embodiment of the present application provides a solution framework for the MADDPG algorithm;
[0057] Figure 6 This is a structural diagram of a power grid dynamic control device based on multi-market subject game provided by an embodiment of the present application;
[0058] Figure 7 Schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] This application constructs a comprehensive multi-agent hybrid game model with reference to relevant research. This model covers wind, photovoltaic, and thermal power generation companies, and comprehensively considers the complex competitive relationships and bidding behaviors among various entities in the electricity market and CET market. Specifically, the Stackelberg game model is used to characterize the leader-follower relationship between renewable energy power generation companies and the day-ahead electricity market; the bidding behaviors of different power generation companies are described using cooperative and non-cooperative game models, respectively, thereby carefully depicting the dynamic evolution of renewable energy power generation companies participating in market transactions. At the same time, a dynamic CET and China Certified Emission Reduction (CCER) price system is established to clearly reflect the impact of market supply and demand changes on the expected profits of power generation companies. By setting a variety of initial CET prices, CCER prices, and free carbon quotas, the mechanism of the impact of different development stages of the CET market on renewable energy trading is deeply explored. This model provides an effective tool for in-depth understanding of the interactive logic of various entities in the electricity market, laying a solid foundation for the subsequent dynamic adjustment of electricity price mechanisms and innovation in revenue distribution methods.
[0060] As the penetration rate of new energy in the power system continues to increase, the price of electricity spot market fluctuates frequently and significantly, and the profits of thermal power companies are damaged. In the process of continuous advancement of power market reform, medium- and long-term power transactions have undergone major changes, and thermal power companies are in urgent need of exploring quotation strategies that adapt to market changes. Traditional reinforcement learning methods are difficult to solve multi-agent incomplete information game models, while the MADDPG algorithm can effectively meet this challenge. By updating the parameters of the neural network to simulate the bounded rational process of the game, it ensures that the game process is close to reality, and provides a new solution for the optimization of bidding strategies of thermal power companies in the medium- and long-term power market. This application aims to provide a method for dynamic control of electricity prices and profits based on collaborative games and intelligent optimization of multiple market players to achieve efficient and stable operation of the power market.
[0061] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0062] See also Figure 1 , which shows a flowchart for implementing a method for dynamic power grid control based on multi-market subject game provided by an embodiment of the present application, including the following steps:
[0063] S101: Obtain market demand data and price fluctuation data, and obtain a pre-built multi-agent hybrid game model. The multi-agent hybrid game model includes a bidding decision layer and a market clearing layer.
[0064] The execution entities of each embodiment of the present application can be servers, processors, microprocessors and other devices with data processing functions. In the actual implementation process, the specific implementation method of the execution entity can be selected according to actual needs. This embodiment does not impose any special restrictions on this, as long as it is a device with data processing functions.
[0065] The multi-agent hybrid game model uses power generation companies as leaders and independent system operators as followers. A Stackelberg game relationship is established between power generation companies and independent system operators, and the basic hierarchical structure of the market is constructed based on this Stackelberg game relationship.
[0066] The multiple entities include power generation companies and independent system operators. Power generation companies can include different types of power generation companies based on actual application scenarios. Optionally, power generation companies include thermal power, wind power, and photovoltaic power. There can be one or more different types of power generation companies.
[0067] This application aims to leverage a multi-player collaborative game mechanism to accurately characterize the behavior of market players and optimize bidding decisions. By leveraging the complex game relationships between the various players in this model, the application accurately presents the competitive landscape and bidding behavior characteristics of different types of power generation companies in the electricity market. By deeply analyzing the strategic interactions between these players, the application provides a scientific basis for power generation companies to develop precise bidding strategies, enabling them to make optimal decisions in a complex and volatile market environment and enhance their market competitiveness.
[0068] S102, at the bidding decision-making level, each power generation enterprise selects the corresponding game model based on market demand data and price fluctuation data, and formulates a power supply strategy with the goal of maximizing profits; at the market clearing level, the independent system operator aims to minimize the cost of purchasing electricity, and determines the planned power generation and market clearing price of each power generation enterprise based on load constraints and unit constraints.
[0069] Among them, the game modes include: cooperative game mode and non-cooperative game mode; each thermal power company formulates a power supply strategy based on the MADDPG algorithm, according to the power market supply and demand ratio and the market share of thermal power companies; the power supply strategy includes: declared power consumption and declared electricity price; the unit constraint conditions are determined according to the unit type corresponding to the power generation company.
[0070] At the bidding decision-making level, power generation companies choose between cooperative and non-cooperative game models based on their own power generation costs, resource reserves, market demand forecasts, and other factors. Different game models correspond to different strategy-making logics, which determine the decision variables and feasible domain for subsequent optimization methods. For example, in a non-cooperative game model, power generation companies aim to maximize their own profits. Their profit function includes factors such as power generation, bid price, power generation costs, and information interaction with other companies. These factors constitute key elements that the optimization algorithm needs to consider when searching for the optimal decision. Based on this, the optimization method adjusts decision variables (such as bid price and power generation) to achieve the goal of profit maximization.
[0071] In the computational process of a hybrid game model involving power generation companies, traditional methods may struggle to find a solution that satisfies all constraints and achieves optimality, given the numerous constraints. To address the unique circumstances of thermal power companies in the medium- and long-term electricity market, the MADDPG algorithm is employed within the model framework for optimization, fully accounting for multiple factors such as market supply and demand, competitor behavior, and the company's own costs. Leveraging the MADDPG algorithm's powerful search and optimization capabilities, it efficiently finds optimal solutions within complex constraint spaces, optimizes thermal power companies' bidding strategies, improves the scientific nature and accuracy of their decision-making, and ensures reasonable returns in market competition. For example, when determining a power generation company's planned power generation and market-clearing price, the market-clearing layer aims to minimize the cost of electricity purchases. The optimization method comprehensively considers the game results and various constraints at the bidding decision-making layer to achieve efficient resource allocation, ensuring the feasibility and optimality of the game model's results in actual market operations.
[0072] In addition, the introduction of the MADDPG algorithm enables dynamic adjustments to electricity pricing mechanisms based on intelligent optimization, promoting a clean market transition. The MADDPG algorithm model closely tracks dynamic changes in market supply and demand, reflecting their impact on power generation companies' expected profits in real time. It conducts in-depth analysis of the competitive advantages and disadvantages of thermal power companies under different market conditions, develops targeted bidding strategies for them, and helps them maintain stable profits amidst market fluctuations while avoiding the negative impact of irrational market behavior on market efficiency, thereby achieving the coordinated development of thermal power companies and other power generation companies. It also improves the response speed and decision-making accuracy of thermal power companies to market changes, enhances their market adaptability, optimizes the allocation of power resources, and comprehensively improves the overall operational efficiency of the power market.
[0073] In this embodiment, a multi-agent hybrid game model including a bidding decision layer and a market clearing layer is constructed to dynamically optimize the operation of the electricity market. At the bidding decision layer, power generation companies choose cooperative or non-cooperative game modes based on market demand and price fluctuations, and use the MADDPG algorithm to accurately formulate strategies for declared electricity volume and electricity prices, thereby improving the decision-making efficiency and adaptability of thermal power companies. At the market clearing layer, independent system operators determine power generation plans and clearing prices based on load and unit constraints with the goal of minimizing electricity purchase costs. The embodiment of the present application coordinates the interests of market entities through a game mechanism. At the same time, based on the intelligent optimization of electricity prices and benefits by the MADDPG algorithm, it effectively balances supply and demand, optimizes resource allocation, and improves the accuracy of decision-making by power generation companies.
[0074] In one possible implementation, each power generation company selects a corresponding game model based on market demand data and price fluctuation data, including:
[0075] When market demand is higher than the set demand range and price fluctuations exceed the set fluctuation range, the cooperative game mode is selected;
[0076] When market demand is lower than the set demand range and price fluctuations are within the set fluctuation range, a non-cooperative game mode is selected.
[0077] At the bidding decision-making level, power generation companies will choose cooperative or non-cooperative game strategies based on factors such as their own power generation costs, resource reserves, and market demand forecasts. Faced with changes in market demand and price fluctuations, power generation companies will also adjust their strategy formulation accordingly. The cooperative and non-cooperative game models used for the bidding behavior of different power generation companies detailed the dynamic evolution of renewable energy power generation companies participating in market transactions, providing a rich set of market behavior scenarios for optimization methods. When the market is in a state of high demand and volatile prices, power generation companies are more inclined to adopt cooperative game strategies. In this case, the optimization method needs to coordinate the decisions of each company based on the rules and constraints of the cooperative game to maximize the alliance's profits. For example, in the cooperative game model, the optimization strategy is formulated by considering factors such as the change in alliance profits and the power generation costs of each company.
[0078] When market demand is low or price volatility is minimal, power generation companies may opt for a non-cooperative game strategy. In this scenario, each company prioritizes maximizing its own interests, and optimization methods must consider the competitive dynamics inherent in non-cooperative games. Companies strive to gain market share by predicting competitor behavior and developing differentiated bidding strategies. In a non-cooperative game model, companies independently make decisions based on their own generation costs, market demand forecasts, and potential competitor reactions to maximize profits.
[0079] In this embodiment, by setting demand ranges and price fluctuation ranges as switching conditions for game modes, dynamic adaptation to market changes is achieved. The cooperative game mode enhances collaborative efficiency among power generation companies and stabilizes market supply; the non-cooperative game mode stimulates the competitive potential of individual companies and enhances market vitality. This rule-based selection mechanism simplifies decision-making logic in complex market environments and reduces the complexity of strategy formulation.
[0080] In one possible implementation, the cooperative game model corresponding to the cooperative game mode is:
[0081]
[0082] Among them, Π C is the total profit of the cooperative alliance C, γ C and δ C is the coefficient, ΔΠ C (t) is the change in alliance profit;
[0083] The non-cooperative game model corresponding to the non-cooperative game mode is:
[0084]
[0085] Among them, Π j is the profit of power generation enterprise j, P jt is the power generation at time t, B jt is the bid price at time t, C j (P jt ) is the power generation cost function, α j and β j is the coefficient, I jk (t) is the information interaction impact value between enterprises j and k at time t, ΔP jt is the change in power generation.
[0086] In this embodiment, key influencing factors such as power generation cost, bidding price, and information interaction are quantified through a profit model of cooperative and non-cooperative games.
[0087] In one possible implementation, each power generation company selects a corresponding game model based on market demand data and price fluctuation data, and formulates a power supply strategy with the goal of maximizing profits, including:
[0088] In the cooperative game model, each power generation company formulates its power supply strategy with the goal of maximizing its own profits;
[0089] Under the non-cooperative game model, each power generation company formulates a power supply strategy with the goal of maximizing alliance profits.
[0090] When power generation companies follow the principle of individual rationality and aim to maximize their own profits, a non-cooperative game model is adopted. Each participant formulates a strategy simultaneously without information exchange. The model is expressed as:
[0091]
[0092] Among them, k1, k2, and k3 are enterprise type identifiers, TH, W, and PV represent the enterprise technology types of traditional thermal power, wind power generation, and photovoltaic power generation, respectively, and t is the time period, indicating the time point or period of strategy formulation. For enterprise type k i output or generating capacity, For enterprise type k i The electricity price, is the profit function of the enterprise, representing the enterprise type k i profit.
[0093] Among them, the profit functions of different types of power generation enterprises are as follows:
[0094] (1) The profit function of thermal power companies needs to take into account peak-valley regulation compensation and cost structure, further subdivide fuel costs, set different cost coefficients according to different fuel sources and qualities, and increase consideration of environmental costs, such as pollution emission control costs and fines for exceeding carbon emission standards. The piecewise function is:
[0095]
[0096] The power generation costs of power companies include fuel costs, carbon emission costs and the levelized cost of carbon capture and storage (CCS).
[0097] Among them, Π TH is the profit of thermal power enterprises, P THt is thermal power generation, B THt is the bid price, C fuel (P THt ) is the fuel cost function, C CCS (P THt ) is the cost of carbon capture and storage, C env (P THt ) is the environmental cost, ∈ TH is the coefficient.
[0098] (2) The profit function of wind power enterprises is: combining the treatment of uncertainty factors, adding quantitative indicators of resource uncertainty such as wind speed, introducing probability distribution function to describe resource uncertainty, and incorporating it into cost and profit calculations. The piecewise function is:
[0099]
[0100] The costs of wind power companies include the levelized cost of wind power energy storage systems (ESS) and CCER costs.
[0101] Among them, Π W is the profit of wind power enterprises, P Wt is the wind power generation, B Wt is the bid price, C ESS (P Wt ) is the energy storage system cost, C CCER (P Wt ) is China’s certified emission reductions, C uncertainty (P Wt ) is the resource uncertainty cost, η W is the coefficient.
[0102] (3) The profit function of photovoltaic enterprises is:
[0103]
[0104] Among them, W tis the actual power generation of the photovoltaic enterprise in time period t, is the expected income or subsidy of the photovoltaic enterprise in period t.
[0105] The cost of photovoltaic enterprises includes the levelized cost of photovoltaic ESS and CCER cost.
[0106] During implementation, information exchange mechanisms between agents will be added to cooperative and non-cooperative game models. In addition to existing information exchange for alliances, non-cooperative games will also allow for a certain degree of market signal sharing, such as price trend forecasts and information on other companies' capacity changes, to restructure the game model at the bidding decision-making level.
[0107] When power generation companies follow the principle of group rationality and aim to maximize alliance profits, a cooperative game model is adopted. Participants cooperate to formulate strategies and there is clear information exchange. The model is expressed as:
[0108]
[0109] Among them, R i represents the payoff function of the i-th participant, c i,t represents the power generation capacity of the i-th participant in the time period, p i,t To represent the electricity price of the i-th participant in the time period, It is to introduce theoretical benefits under Pareto optimality or Nash equilibrium to quantify the loss of strategy efficiency.
[0110] This example clearly defines the goal-oriented differences between cooperative and non-cooperative games, precisely matching the behavioral logic of enterprises in different market scenarios. By setting hierarchical goals, we avoid strategic conflicts and ensure that the model can both meet individual interests in a dynamic market and enhance overall market efficiency through alliance cooperation.
[0111] In one possible implementation, under a non-cooperative game model, the Shapley value method is used to distribute alliance profits according to the contribution of each enterprise to the alliance.
[0112] In this embodiment, under a non-cooperative game model, the Shapley value method is used to distribute alliance profits. This method quantifies benefits based on each enterprise's contribution to the alliance, ensuring fair and reasonable profit distribution. This mechanism reduces market friction caused by uneven profit distribution, enhances alliance stability, incentivizes enterprises to actively participate in cooperative games, and promotes the stable operation of the power grid system.
[0113] Figure 2 It is a market mechanism diagram of the power grid dynamic control method based on multi-market player game provided in the embodiment of the present application. Figure 2The figure shows the architecture of the hybrid game model between power generation enterprises and independent system operators under the electricity-carbon integrated market mechanism in the dynamic control scheme of the power grid based on multi-agent hybrid game. At the bidding decision-making level, power generation enterprises choose cooperative or non-cooperative game mode according to their own situation, and formulate strategies with the goal of maximizing profits. Their strategies are affected by factors such as power generation costs and unit characteristics. At the market clearing level, independent system operators determine the planned power generation and market clearing prices of power generation enterprises based on constraints such as load demand and with the goal of minimizing the cost of purchasing electricity. The figure clearly presents the hierarchical relationship and decision-making logic between the various subjects in the model, which helps to understand how this application can achieve effective analysis and optimization of power market trading behavior through a dynamic control scheme of the power grid based on multi-agent hybrid game.
[0114] In one possible implementation, the MADDPG algorithm includes an Actor policy network model and a Critic value network model;
[0115] Based on the MADDPG algorithm, each thermal power company formulates a power supply strategy according to the power market supply and demand ratio and the thermal power company's market share, including:
[0116] The power market supply-demand ratio and the market share of thermal power companies are input into the actor strategy network model, which then outputs the power supply forecast strategy for the thermal power companies. The forecast strategy includes the expected declared power volume and the expected declared electricity price, expressed as a declared power volume decision coefficient and a declared price decision coefficient. The declared power volume decision coefficient can optionally be in the range [0, 1], and the declared price decision coefficient can be in the range [1, 1.2].
[0117] The power market supply-demand ratio and thermal power company market share are input into the Critic value network model, and the state value function that evaluates the quality of power supply forecast strategy is output.
[0118] Formulate power supply strategies for thermal power companies based on the state value function.
[0119] This application adopts a tiered pricing mechanism to reflect the impact of market supply and demand on carbon trading prices. As the demand for carbon trading quotas increases, the carbon trading price rises; as the supply of CCER quotas increases, the CCER price falls. t CET ) and CCER price (T t CCER ) is calculated as follows:
[0120]
[0121] in, To price the carbon trading benchmark, μ CET is the carbon trading price adjustment coefficient, l CET is the benchmark threshold for carbon trading quota demand, is the CCER benchmark price, μ CCER is the CCER price adjustment coefficient, carbon trading quota demand and CCER quota supply Respectively expressed as:
[0122]
[0123] in, is the carbon emission coefficient of thermal power, which indicates the carbon emission per unit of power generation, β k,Free is the proportion of free carbon quota, represents the power generation of the thermal power company at time t, and ε represents other adjustment factors.
[0124] By setting CET and CCER prices, we can closely track market supply and demand dynamics and reflect their impact on power generation companies' expected profits in real time. By establishing a flexible electricity price adjustment mechanism, we can guide power generation companies to actively respond to market signals and optimize their energy structure.
[0125] In a multi-agent hybrid game model, a MADDPG algorithm-based model is constructed to address the bidding strategies of thermal power companies in the medium- and long-term electricity market. In this model, thermal power companies formulate trading strategies based on their own operational data and incomplete market information, aiming to maximize profits. Based on the cost differences among thermal power companies, two scenarios, non-cooperative and cooperative, are set up, and corresponding game models are constructed for each.
[0126] In a non-cooperative scenario, the profit of thermal power companies in the medium- and long-term market (taking the monthly centralized bidding market as an example) is the profit from electricity sales minus the cost of power generation. The cost of power generation is calculated based on the marginal cost of power generation in the bidding decision of thermal power suppliers. In the non-cooperative game scenario of the dynamic power grid control scheme based on multi-agent hybrid game, the non-cooperative profit model is expressed as:
[0127]
[0128] in, is the monthly centralized bidding transaction profit of thermal power companies, p mc is the market clearing price, a G 、b G 、c G is the production cost coefficient of thermal power enterprises, For trading electricity, is the monthly decomposition of annual contracted electricity, T Mon The number of hours per month.
[0129] Thermal power companies use their market power and adopt certain bidding strategies (including physical and economic holdings) to influence the market clearing price in order to increase their own profits. In the monthly centralized bidding transactions, the decision variables for the power volume and price declared by thermal power companies are as follows:
[0130]
[0131] Among them, α G is the decision coefficient for declared electricity quantity, β G is the decision coefficient for the declared electricity price, The maximum monthly surplus electricity of thermal power companies, is the monthly maximum power generation capacity, is the monthly average marginal power generation cost. At the same time, the decision model must also meet the following conditions: thermal power output constraint, ramp constraint, minimum continuous start-stop time constraint, maximum start-stop number constraint, power limit and quotation limit.
[0132] In the cooperative scenario, the thermal power companies cooperate to form an alliance, with the maximization of the overall profit of the alliance as the decision-making goal. The profit function is:
[0133]
[0134] The Shapley value method is used to distribute alliance profits, and the benefits are distributed according to the contribution of each enterprise to the alliance.
[0135] The decision-making model for cooperation is subject to the same constraints as the non-cooperative game. The goal is to maximize the total profit of the alliance. The decision-making model can be described as follows:
[0136]
[0137] The MADDPG algorithm is an Actor-Critic network architecture. The input of the Actor strategy network model is the power market status and the state characteristics of the thermal power companies themselves, including the power market supply and demand ratio and the thermal power companies' market share. The output is the thermal power companies' declared power and declared electricity price decision-making behavior, with the declared power decision coefficient α G and the declared electricity price decision coefficient β G In which, α G The range is [0,1], β G The range is [1,1.2].
[0138] The input of the Critic value network model is the same state characteristics as the Actor policy network model, and the output is a state value function, which is used to evaluate the quality of the strategy.
[0139] The MADDPG algorithm uses the MLMF model for calculation. During the training process, each thermal power company (Agent) uses a joint strategy to interact with the environment. It can act as an actor (strategy) network and output continuous actions based on the current local observation and strategy. When the Critic network performs strategy evaluation, it improves the strategy by generating an estimated q-value function. In the MADDPG algorithm, the observation values of each agent are completely different, and the global equilibrium electricity price λ MG,t To achieve full observations, the improved centralized critic network takes all agent behaviors and a single agent's local observations as input and outputs a q-value function. The critic network is a neural network consisting of an input layer, a hidden layer, and an output layer. Relu and Linear are used as activation functions for the neural network. The MADDPG algorithm leverages some of the techniques of the DDPG algorithm. Specifically, its implementation is as follows:
[0140] (1) Initialization: Create the following network for each agent (MG):
[0141] Actor network (strategy network μ i ): Input local observation o i (5-dimensional), output continuous action a i ∈[-1,1](charge and discharge instructions).
[0142] Critic network (Q value network Q i ): Input local observation o i and all agents’ actions a1,...,a n , output Q value.
[0143] Target Actor Network and Target Critic Network: Parameters are the same as the online network to stabilize training. Initialize the experience pool D to store the interaction experience of all agents.
[0144] The experience tuple (o i ,a,r i ,o′ i ), the environment returns a reward r i and the next moment observation o′ i Deposit into experience pool D.
[0145] (2) Each agent i observes o according to the current i , generate actions through the Actor network:
[0146]
[0147] Among them, when the agent chooses an action, a random Gaussian distribution is used to achieve an appropriate balance between exploration and exploitation. Used to encourage the exploration of different strategies, where Gaussian noise can be expressed as:
[0148]
[0149] Wherein, χ is the initial value of the standard deviation of the Gaussian distribution, and the decrease rate of the standard deviation of the Gaussian distribution is a decrease rate less than 1.
[0150] Convert standardized actions into actual control signals:
[0151] P ES,i,t =a i ·P ES,max
[0152] (3) Based on the supply and demand relationship, the equilibrium electricity price is calculated through a distributed algorithm:
[0153] λ MG,t =f(∑P ES,i,t ,∑P AL,i,t )
[0154] Among them, ∑P AL,i,t is the demand response amount. At the same time, the equilibrium electricity price λ MG,t Feedback is given to each agent as the global observation input for the next period.
[0155] (3) Experience storage:
[0156] Reward function design:
[0157]
[0158] Among them, δ1 is the weight coefficient of the market coordination goal. The larger the coefficient, the more inclined the intelligent agent is to prioritize reducing the supply-demand deviation; δ2 is the weight coefficient of the user satisfaction goal. The larger the coefficient, the more inclined the intelligent agent is to lower the electricity price or adjust the electricity consumption plan to maintain user satisfaction.
[0159] The current observation, action, reward, and next observation are stored in the experience pool for subsequent updates.
[0160] (4) Network Update:
[0161] ①Critic Network Update:
[0162] Calculate the target Q value:
[0163] y i =r i +γQ′ i (o′ i ,μ′1(o′1),...,μ′ n (o′ n ))
[0164] Minimize the Critic's mean squared error loss:
[0165]
[0166] Actor Network Updates:
[0167] Maximize the Q value by gradient ascent and update the strategy:
[0168]
[0169] Target network soft update:
[0170]
[0171] ②Distributed execution:
[0172] After training is completed, each agent only relies on local observations o i Generate action a i =μ i (o i ), no other agent information is required.
[0173] The model construction process comprehensively considers the competitive and cooperative relationships among different types of power generation companies in the power market, as well as the impact of the carbon market on the decision-making of power generation companies. At the same time, targeting the special circumstances of thermal power companies in the medium and long-term market, the MADDPG algorithm is used to optimize their bidding strategies, aiming to achieve stable operation of the power market and efficient allocation of resources.
[0174] In this example, the MADDPG algorithm's actor-critic network architecture enables intelligent optimization of thermal power companies' bidding strategies. Based on the power market's supply-demand ratio and market share, the actor network outputs continuous decision parameters for bid power and price, improving strategy generation efficiency. The critic network evaluates strategies and provides feedback, guiding the algorithm to quickly converge to the optimal solution. This algorithm transcends traditional methods' reliance on historical data, enabling high-precision, low-latency, real-time decision-making in dynamic markets.
[0175] Figure 3 The figure shows the interaction process between the agent and the environment in the MADDPG algorithm reinforcement learning. Figure 3 As shown in Figure 2, the agent is the executor of the decision in this interaction process. It receives the current state (state, s) from the environment. t ), this environment is the market environment set by the power grid dynamic control scheme based on multi-agent hybrid game, based on which a decision is made and an action is output (action, a t ) is fed back to the environment. When the environment receives the action of the agent (a t ) will update the state according to its own rules and mechanisms, from the current state (st ) changes to the next state (s t+1 At the same time, the environment will give corresponding rewards (reward, r) according to the agent's actions. t and r t+1 The reward mechanism is a key factor in guiding agents to learn optimal strategies in reinforcement learning. The agent's goal is to maximize the cumulative reward by continuously interacting with the environment. This process is achieved within the overall framework of a dynamic power grid control solution based on multi-agent hybrid games.
[0176] Figure 4 The implementation framework of the MADDPG algorithm is shown in Figure 4 As shown, in a multi-agent system, multiple thermal power companies (Agents) interact with the power market environment set by the dynamic power grid control scheme based on multi-agent hybrid game through joint strategies. The strategy network of each Agent outputs decision-making behavior based on local observation information, and at the same time evaluates the joint behavior value function and updates the strategy parameters based on its gradient. The value network evaluates the pros and cons of the strategy based on the same state characteristics, providing a basis for the adjustment of the strategy network. The environmental model is a medium- and long-term power market clearing model, which outputs rewards and next states based on the behavior of the Agent. This figure clearly shows how the MADDPG algorithm optimizes the bidding strategy of thermal power companies in this application, reflects the advantages of the algorithm in dealing with multi-agent decision-making problems, and helps to understand how this application uses the algorithm to improve the decision-making ability and market adaptability of thermal power companies in the power market.
[0177] Figure 5 The solution framework of the MADDPG algorithm is shown in Figure 2. Figure 5 As shown in Figure 1, there are n thermal power companies in the multi-agent system, and each agent has a policy network. The MADDPG algorithm uses centralized training and distributed execution.
[0178] First, during the training process, n agents Use joint strategies to interact with the environment set by the power grid dynamic control scheme based on multi-agent hybrid game. At the same time, the joint behavior value function Q of each agent i is i (o1,a1,o2,a2,…,o n ,a n ) is evaluated and the strategy of each agent is updated based on the gradient of the joint behavior value function relative to the strategy parameters. The strategy input of each agent i is the local observation value o i , the output is the action a of agent i i Secondly, in the execution phase, the input of agent i is the local observation o i , the output is the action a of agent i iThe entire process is completed within the framework of the power grid dynamic control scheme based on multi-agent hybrid game.
[0179] To address the unique challenges faced by thermal power companies in medium- and long-term electricity markets, the MADDPG algorithm was used within a multi-agent hybrid game-based dynamic grid management and control framework to optimize the bidding strategies of thermal power companies. This algorithm fully considers the influence of multiple factors, including market supply and demand, competitor behavior, and the company's own costs, to optimize the bidding strategies of these companies. Experimental results show that using this algorithm to optimize bidding strategies significantly improves the profits of thermal power companies. For example, in a simulated monthly centralized bidding market, one thermal power company achieved a 5.2% profit increase compared to traditional methods through optimized strategies. This effectively improves the scientific nature and accuracy of thermal power companies' decision-making, enabling them to obtain more reasonable returns in market competition.
[0180] As renewable energy penetration increases in the power market, the dynamic electricity pricing mechanism is encouraging power generation companies to prioritize clean energy utilization. Experimental data shows that after implementing the dynamic electricity pricing mechanism, the proportion of renewable energy generation in the power market has increased, effectively promoting the widespread adoption of renewable energy, facilitating the sustainable development of the power market towards clean, low-carbon development, and providing strong support for achieving the "dual carbon" goals. Furthermore, reasonable carbon trading prices incentivize thermal power companies to actively participate in carbon reduction, resulting in a reduction in their carbon emissions, further promoting the green transformation of the power market.
[0181] In one possible implementation, unit constraints include: output limits and bidding quantity constraints for thermal power units and wind and solar power stations, thermal power unit ramp rate constraints, and energy storage system operation constraints;
[0182] Load constraints include: day-ahead power market supply and demand balance constraints and bidding price constraints.
[0183] In this example, unit constraints prevent grid instability caused by sudden changes in generator power. Load constraints improve energy storage resource utilization and mitigate fluctuations in renewable energy. The systematic integration of constraints ensures the practical feasibility of the model and the engineering feasibility of the market-clearing results.
[0184] In one possible implementation, the day-ahead power market supply and demand balance constraint is:
[0185]
[0186] in, are the output of thermal power, wind power and photovoltaic power at time t, D t is the load demand at time t;
[0187] Thermal power unit output limits and bidding quantity constraints:
[0188]
[0189] in, are the minimum and maximum output of thermal power units, are the minimum and maximum bidding prices of thermal power units respectively;
[0190] The output limits and bidding constraints for wind power and photovoltaic power stations are as follows:
[0191]
[0192] in, are the predicted outputs of wind power and photovoltaic power respectively, The output after adjustment of the energy storage systems equipped for wind power and photovoltaic power, are the minimum and maximum bidding prices for wind power and PV, respectively;
[0193] The ramp rate constraint of thermal power units is:
[0194]
[0195] Among them, V t,domn is the downward climbing rate, V t,up Δt is the upward climbing rate, Δt is the time interval;
[0196] The operating constraints of the energy storage system are:
[0197]
[0198] SOC min E max ≤E r ≤SOC max E max
[0199]
[0200] in, are the charging and discharging power of the energy storage system, is the maximum charge and discharge power of the energy storage system, SOC min , SOC max are the minimum and maximum charge states of the energy storage system, E r The amount of electricity stored in the energy storage system, E t is the amount of energy storage system at the moment, θ c ,θ d are the charging and discharging efficiency of the energy storage system respectively.
[0201] In this embodiment, the day-ahead power market supply and demand balance formula ensures real-time matching of total power generation with load demand. Thermal power unit output limits and bidding constraints, as well as wind and photovoltaic power plant output limits and bidding constraints, prevent malicious bidding from disrupting market order. Thermal power unit ramp rate constraints prevent grid instability caused by sudden power changes. Energy storage system operation constraints enhance power supply stability. Formulated constraints enhance model transparency and reduce the risk of biased implementation of market rules.
[0202] Based on the above embodiments, this application overcomes the limitations of traditional research methods for power supplier bidding strategies. Compared with the cost analysis method, this application's dynamic power grid control scheme based on multi-agent hybrid game fully considers market supply and demand and other supplier decisions, and can maximize its own interests; compared with the clearing price prediction method and competitor quotation analysis method, this application does not rely on a large amount of historical data and can still make accurate decisions in the early stages of power market reform when data is scarce and market structure rules are constantly changing; compared with game theory analysis methods, this application is more effective in dealing with multi-person and incomplete information game problems; compared with traditional intelligent optimization algorithms, this application has a more stable environment and smaller variance in multi-agent interaction scenarios, effectively improving the accuracy of power generation enterprise decision-making.
[0203] With the advancement of power market reform, the penetration rate of renewable energy continues to increase, electricity spot market prices fluctuate frequently, thermal power companies' profits are damaged, and medium- and long-term power transactions are undergoing significant changes. This application provides thermal power companies with an effective solution to adapt to market changes by constructing a dynamic power grid control scheme based on a multi-agent hybrid game and applying the MADDPG algorithm. This helps thermal power companies maintain their competitiveness in a complex and changing market environment, while promoting the stable and efficient operation of the entire power market.
[0204] In actual applications, with the continuous practice and in-depth research of the relevant technologies of this application, more accurate and detailed data will be obtained to further accurately quantify the various effects brought about by the invention, thereby more comprehensively demonstrating its value in the field of power market management.
[0205] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0206] The following are device embodiments of the present application. For details not fully described therein, please refer to the corresponding method embodiments described above.
[0207] Figure 6 A schematic diagram of the structure of a power grid dynamic control device based on multi-market subject game provided by an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown, which are detailed as follows:
[0208] like Figure 6 As shown in FIG, the dynamic power grid control device based on the game of multiple market players includes:
[0209] An acquisition module 601 is used to acquire market demand data and price fluctuation data, and acquire a pre-built multi-agent hybrid game model, wherein the multi-agent hybrid game model includes a bidding decision layer and a market clearing layer;
[0210] The game optimization module acquisition module 62 is used to select a corresponding game mode at the bidding decision-making level based on market demand data and price fluctuation data, and formulate a power supply strategy with the goal of maximizing profits. The game modes include: cooperative game mode and non-cooperative game mode. Each thermal power company formulates a power supply strategy based on the power market supply and demand ratio and the thermal power company's market share based on the MADDPG algorithm. The power supply strategy includes: declared power consumption and declared electricity price.
[0211] At the market clearing level, the independent system operator aims to minimize the cost of electricity purchases and determines the planned power generation and market clearing price of each power generation enterprise based on load constraints and unit constraints; among them, the unit constraints are determined according to the unit type corresponding to the power generation enterprise.
[0212] In the embodiment of the present application, a multi-agent hybrid game model including a bidding decision layer and a market clearing layer is constructed to dynamically optimize the operation of the electricity market. At the bidding decision layer, power generation companies choose cooperative or non-cooperative game modes based on market demand and price fluctuations, and use the MADDPG algorithm to accurately formulate strategies for declared electricity volume and electricity prices, thereby improving the decision-making efficiency and adaptability of thermal power companies. At the market clearing layer, independent system operators determine power generation plans and clearing prices with the goal of minimizing electricity purchase costs, taking into account load and unit constraints. The embodiment of the present application coordinates the interests of market entities through a game mechanism. At the same time, based on the intelligent optimization of electricity prices and benefits by the MADDPG algorithm, it effectively balances supply and demand, optimizes resource allocation, and improves the accuracy of power generation company decisions.
[0213] Figure 7 Schematic diagram of an electronic device provided in an embodiment of the present application. Figure 7 As shown, the electronic device 7 of this embodiment includes a processor 70 and a memory 71. The memory 71 stores a computer program 72. When the processor 70 executes the computer program 72, the steps of the above-described method embodiments are implemented. Alternatively, when the processor 70 executes the computer program 72, the functions of the modules / units in the above-described device embodiments are implemented.
[0214] Exemplarily, the computer program 72 may be divided into one or more modules / units, which are stored in the memory 71 and executed by the processor 70 to implement the present application. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program 72 in the electronic device 7.
[0215] The electronic device 7 may include, but is not limited to, a processor 70 and a memory 71. Those skilled in the art will appreciate that Figure 7 It is only an example of the electronic device 7 and does not constitute a limitation of the electronic device 7. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the electronic device 7 may also include input and output devices, network access devices, buses, etc.
[0216] For the sake of convenience and brevity, the division of the above functional modules / units is only used as an example. In actual applications, the above functions can be assigned to different functional modules / units as needed. The above modules / units can be implemented in the form of hardware, software, or a combination of hardware and software.
[0217] In the above embodiments, the descriptions of each embodiment have their own focus. For parts not described or recorded in detail in one embodiment, please refer to the relevant descriptions of other embodiments. Unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features of different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0218] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for dynamic power grid control based on multi-market subject game, characterized in that: include: Obtaining market demand data and price fluctuation data, and obtaining a pre-built multi-agent hybrid game model, wherein the multi-agent hybrid game model includes a bidding decision layer and a market clearing layer; At the bidding decision-making level, each power generation enterprise selects a corresponding game mode based on market demand data and price fluctuation data, and formulates a power supply strategy with the goal of maximizing profits. The game modes include cooperative game mode and non-cooperative game mode. Each thermal power enterprise formulates a power supply strategy based on the power market supply and demand ratio and the thermal power enterprise's market share based on the multi-agent deep deterministic policy gradient MADDPG algorithm. The power supply strategy includes: declared power consumption and declared electricity price. In the market clearing layer, the independent system operator aims to minimize the cost of electricity purchase and determines the planned power generation and market clearing price of each power generation enterprise based on load constraints and unit constraints.
2. The method for dynamic power grid control based on multi-market subject game according to claim 1, characterized in that: Each power generation enterprise selects a corresponding game model based on market demand data and price fluctuation data, including: When market demand is higher than the set demand range and price fluctuations exceed the set fluctuation range, the cooperative game mode is selected; When market demand is lower than the set demand range and price fluctuations are within the set fluctuation range, a non-cooperative game mode is selected.
3. The method for dynamic power grid control based on multi-market subject game according to claim 1, characterized in that: Each power generation enterprise selects a corresponding game model based on market demand data and price fluctuation data, and formulates a power supply strategy with the goal of maximizing profits, including: In the cooperative game model, each power generation company formulates its power supply strategy with the goal of maximizing its own profits; Under the non-cooperative game model, each power generation company formulates a power supply strategy with the goal of maximizing alliance profits.
4. The method for dynamic power grid control based on multi-market subject game according to claim 3 is characterized in that: In the non-cooperative game model, the Shapley value method is used to distribute alliance profits according to the contribution of each enterprise to the alliance.
5. The method for dynamic power grid control based on multi-market subject game according to claim 1, characterized in that: The MADDPG algorithm includes an Actor strategy network model and a Critic value network model; Each thermal power company formulates a power supply strategy based on the MADDPG algorithm, according to the power market supply and demand ratio and the thermal power company's market share, including: Input the power market supply-demand ratio and the market share of thermal power companies into the Actor strategy network model, and output the power supply expectation strategy of the thermal power companies; the power supply expectation strategy includes: expected declared power volume and expected declared power price; The power market supply-demand ratio and the market share of thermal power companies are input into the Critic value network model to output a state value function that evaluates the quality of the power supply expectation strategy; A power supply strategy for the thermal power company is formulated according to the state value function.
6. The method for dynamic power grid control based on multi-market subject game according to claim 1, characterized in that: The cooperative game model corresponding to the cooperative game mode is: Among them, Π C is the total profit of the cooperative alliance C, γ C and δ C is the coefficient, ΔΠ C (t) is the change in alliance profit; The non-cooperative game model corresponding to the non-cooperative game mode is: Among them, Π j is the profit of power generation enterprise j, P jt is the power generation at time t, B jt is the bid price at time t, C j (P jt ) is the power generation cost function, α j and β j is the coefficient, I jk (t) is the information interaction impact value between enterprises j and k at time t, ΔP jt is the change in power generation.
7. The method for dynamic power grid control based on multi-market subject game according to claim 1, characterized in that: The unit constraints include: output limits and bidding quantity constraints of thermal power units and wind and solar power stations, thermal power unit ramp rate constraints, and energy storage system operation constraints; The load constraint conditions include: day-ahead power market supply and demand balance constraints and bidding price constraints.
8. The method for dynamic power grid control based on multi-market subject game according to claim 7, characterized in that: The day-ahead power market supply and demand balance constraint is: in, are the output of thermal power, wind power and photovoltaic power at time t, D t is the load demand at time t; The output limit and bidding quantity constraints of thermal power units are as follows: in, are the minimum and maximum output of thermal power units, are the minimum and maximum bidding prices of thermal power units respectively; The output limits and bidding constraints for wind power and photovoltaic power stations are as follows: in, are the predicted outputs of wind power and photovoltaic power respectively, The output after adjustment of the energy storage systems equipped for wind power and photovoltaic power, are the minimum and maximum bid prices for wind power and PV, respectively; The thermal power unit ramp rate constraint is: Among them, V t,domn is the downward climbing rate, V t,up Δt is the upward climbing rate, Δt is the time interval; The energy storage system operation constraints are: in, are the charging and discharging power of the energy storage system, is the maximum charge and discharge power of the energy storage system, SOC min , SOC max are the minimum and maximum charge states of the energy storage system, E r The amount of electricity stored in the energy storage system, E t is the amount of energy storage system at the moment, θ c ,θ d are the charging and discharging efficiency of the energy storage system respectively.
9. A power grid dynamic control device based on multi-market subject game, characterized in that: include: An acquisition module is used to acquire market demand data and price fluctuation data, and acquire a pre-built multi-agent hybrid game model, wherein the multi-agent hybrid game model includes a bidding decision layer and a market clearing layer; A game optimization module is used at the bidding decision-making level for each power generation enterprise to select a corresponding game mode based on market demand data and price fluctuation data, and to formulate a power supply strategy with the goal of maximizing profits. The game modes include cooperative game mode and non-cooperative game mode. Each thermal power enterprise formulates a power supply strategy based on the power market supply and demand ratio and the thermal power enterprise's market share based on the MADDPG algorithm. The power supply strategy includes: declared power consumption and declared electricity price. In the market clearing layer, the independent system operator aims to minimize the cost of purchasing electricity and determines the planned power generation and market clearing price of each power generation enterprise based on load constraints and unit constraints; among which, the unit constraints are determined according to the unit type corresponding to the power generation enterprise.
10. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Cited By
A peak shaving auxiliary service cooperation game cost allocation method and system considering environmental externality
CN122509641A