Evolutionary game method of power-carbon-green certificate coupled market considering revenue function
Patent Information
- Application Number
- CN202610922222.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]本发明的目的在于提供考虑收益函数的电力-碳-绿证耦合市场进化博弈方法,以解决上述背景技术中提出的现有方法难以在同一框架下兼顾多主体异构收益建模、多市场约束处理、策略动态演化、均衡策略求解和策略反馈修正的问题
[0051](1)本发明通过将电力-碳-绿证耦合市场参与主体划分为发电企业智能体、用电用户智能体、可再生能源企业智能体和监管机构智能体,并分别构建与各类智能体市场角色和决策目标对应的异构收益函数,使收益模型能够同时表征跨市场收益耦合关系和主体策略互动关系;相较于传统单一市场收益叠加模型,本发明能够更准确地反映三市场联动场景下不同主体的真实收益变化,为后续策略演化和策略优化提供更可靠的收益计算基础;
Smart Images

Figure CN122840987A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of coordinated operation technology of electricity market, carbon market and green certificate market, specifically to an evolutionary game method for coupled electricity-carbon-green certificate market considering the payoff function. Background Technology
[0002] Against the backdrop of the "dual carbon" goals and the construction of a new power system, the linkage between the electricity market, carbon market, and green certificate market is gradually strengthening. The electricity market reflects the supply and demand relationship of electricity through electricity prices, the carbon market constrains high-carbon emission behavior through carbon quotas and carbon prices, and the green certificate market guides the production and consumption of green electricity through the issuance, trading, and holding mechanisms of green certificates. As power generation companies, electricity users, renewable energy companies, and regulatory agencies are simultaneously affected by multiple market rules, the strategy optimization methods under a single market are no longer sufficient to meet the operational needs of the coupled electricity-carbon-green certificate market.
[0003] Existing optimization methods for the electricity, carbon, and green certificate markets often focus on single-market modeling or strategy solving under static constraints. They are difficult to simultaneously characterize the heterogeneous returns of different stakeholders, cross-market return coupling, market constraints, and dynamic evolution of strategies within the same framework. In coupled markets with multiple stakeholders, the lack of unified modeling of return functions, strategy spaces, feasible regions of constraints, and equilibrium solution processes can easily lead to a disconnect between strategy optimization results and actual market operation constraints, affecting the executability of strategies and the effectiveness of coordinated regulation.
[0004] Therefore, it is essential to design an evolutionary game theory approach that considers the payoff function in the coupled electricity-carbon-green certificate market. Summary of the Invention
[0005] The purpose of this invention is to provide an evolutionary game theory method for the coupled market of electricity, carbon, and green certificates that considers the payoff function, in order to solve the problem that existing methods proposed in the background art are difficult to simultaneously consider multi-agent heterogeneous payoff modeling, multi-market constraint handling, dynamic strategy evolution, equilibrium strategy solution, and strategy feedback correction within the same framework.
[0006] To achieve the above objectives, this invention provides the following technical solution: a market evolution game method considering the payoff function of electricity-carbon-green certificates, comprising the following steps:
[0007] S1. Construct a multi-agent heterogeneous revenue model: Divide the market participants in the electricity-carbon-green certificate coupling into power generation enterprise agents, electricity user agents, renewable energy enterprise agents, and regulatory agency agents; construct heterogeneous revenue functions according to the market roles and decision-making objectives of each type of agent;
[0008] S2, Define the agent policy space and construct the pre-feasibility domain constraints: Based on the heterogeneous benefit functions and decision objectives of various agents in S1, determine the policy variables of the four types of agents and form corresponding policy spaces; establish a multi-constraint coupled equation system around the policy space, which includes at least hard constraints on power supply and demand, total carbon emission constraints, renewable energy consumption constraints, green certificate supply and demand balance constraints, and price difference anti-collusion constraints; set constraint relaxation and adopt a dynamic contraction mechanism to adjust the feasible domain range so that the candidate strategy satisfies the pre-feasibility domain constraints before entering the policy evolution stage;
[0009] S3: Construct a multi-agent evolutionary game dynamic equation with Lagrange constraint correction: Calculate the policy payoffs of various agents based on the heterogeneous payoff function in S1, and form a policy distribution based on the policy space in S2; introduce a heterogeneous payoff weight matrix, and introduce the multi-constraint coupling equation set in S2 into the Lagrange function; add a Lagrange dual constraint correction term to the replicated dynamic equation to form a reconstructed dynamic equation, so that the policy distribution evolves in the prior feasible region according to payoff-driven and constraint correction, and adaptively updates the Lagrange dual variables;
[0010] S4, Solving the equilibrium policy distribution by integrating ESS and MAPPO algorithms: Based on the policy evolution relationship represented by the reconstructed dynamic equation in S3, an ESS prior model is constructed and a stable equilibrium prior policy distribution is output. The stable equilibrium prior policy distribution is then weighted and fused with the policy distribution output by the MAPPO policy network. A multi-market coupled attention mechanism is introduced into the MAPPO value network, and an equilibrium stability penalty term is added to the policy optimization objective. After iterative training, the equilibrium policy distribution of four types of agents is output.
[0011] S5, Output the optimal strategy and establish a feedback loop: Based on the equilibrium strategy distribution output in S4, determine the strategy optimization search range and initialization conditions, and according to the strategy variable types and optimization objectives of the four types of intelligent agents, respectively execute the bidding and trading strategy optimization of the power generation enterprise intelligent agent, the energy consumption and green electricity consumption strategy optimization of the electricity user intelligent agent, the time-of-use power output strategy optimization of the renewable energy enterprise intelligent agent, and the multi-objective regulation strategy optimization of the regulatory agency intelligent agent, to obtain the executable optimal strategy for the four types of intelligent agents; based on the executable optimal strategy, construct a multi-market collaborative operation strategy for electricity-carbon-green certificates, and reuse the multi-constraint coupled equation system in S2 for full-constraint feasibility verification; establish a feedback loop, where short-term feedback updates the strategy based on real-time operating status, and long-term feedback updates the revenue model parameters and constraint parameters based on cumulative operating data.
[0012] As a further technical solution of the present invention, the step S1 of constructing a multi-agent heterogeneous benefit model includes the following steps:
[0013] S1.1, Classification and Decision-Making Objectives of Coupled Market Intelligent Agents: The participants in the coupled electricity-carbon-green certificate market are classified into four heterogeneous intelligent agents: power generation enterprise intelligent agents, electricity user intelligent agents, renewable energy enterprise intelligent agents, and regulatory agency intelligent agents. Among them: power generation enterprise intelligent agents participate in electricity market, carbon market, and green certificate market transactions simultaneously, and their decision-making objective is to maximize the comprehensive benefits of the three markets; electricity user intelligent agents, as the demand-side entities of electricity and green certificates, aim to maximize the comprehensive benefits of electricity consumption utility, carbon emission reduction benefits, and green electricity consumption rewards; renewable energy enterprise intelligent agents, as the entities responsible for green electricity supply and green certificate issuance, aim to maximize the comprehensive benefits of green electricity generation and green certificate revenue; and regulatory agency intelligent agents, as the entities responsible for market rule-making and regulation, aim to collaboratively achieve multiple public objectives, including stable electricity prices, carbon emission reduction, green electricity consumption, and maximizing social welfare.
[0014] S1.2, Construction of Heterogeneous Revenue Functions for Four Types of Intelligent Agents: For the decision-making objectives and trading scope of the four types of intelligent agents, heterogeneous revenue functions with differentiated structures are constructed respectively. Specifically: the revenue function for power generation enterprises comprehensively covers electricity sales revenue, power generation costs, carbon quota trading profits and losses, and green certificate holding revenue, and embeds cross-market coupled revenue terms to characterize the impact of the interaction between electricity prices, carbon prices, and green certificate prices on the enterprise's overall revenue; the revenue function for electricity users comprehensively covers electricity utility, electricity expenditure, and green certificate holding revenue, and introduces two types of strategic interaction revenue terms: implicit carbon emission reduction revenue and excess green electricity consumption rewards; the revenue function for renewable energy enterprises comprehensively covers green electricity sales revenue and green certificate issuance revenue, and introduces a green electricity scarcity premium strategy interaction revenue term; the revenue function for regulatory agencies adopts a multi-objective reward and punishment mechanism, setting penalty terms for electricity price fluctuations, excessive system carbon emissions, and deviations in green electricity consumption, and setting reward terms for the social welfare corresponding to the total market revenue.
[0015] As a further technical solution of the present invention, the cross-market coupling revenue term of the revenue function of the power generation enterprise intelligent agent in S1.2 includes a carbon-green certificate price linkage sensitivity coefficient, which is calibrated through the following steps:
[0016] S1.2.1, Data Acquisition and Preprocessing: Collect daily price data for the electricity market, carbon market, and green certificate market covering three complete calendar years; perform stationarity tests on the price series; and perform first-order differencing on non-stationary series to obtain stationary series.
[0017] S1.2.2, Model Construction: Construct a third-order lagged vector autoregressive model (VAR(3) model) containing three variables: electricity price, carbon price, and green certificate price, to capture the linkage transmission law of price lag over multiple periods;
[0018] S1.2.3, Parameter Calculation: Model parameters are obtained through regression estimation. The marginal impact of carbon price and green certificate price on electricity price is extracted separately. The carbon-green certificate price linkage sensitivity coefficient is obtained by combining the coupling effect of the two, which is used to quantify the linkage effect of carbon price and green certificate price on electricity price under the joint action of carbon price and green certificate price.
[0019] As a further technical solution of the present invention, the S2 definition of the agent policy space includes the following steps:
[0020] S2.1 Determination of the strategy space of the power generation enterprise's intelligent agent: The strategy space of the power generation enterprise's intelligent agent includes three types of decision variables: bidding coefficient, carbon quota purchase volume, and green certificate holding volume, and correspondingly sets a reasonable bidding range, carbon quota purchase upper limit, and green certificate holding upper limit;
[0021] S2.2 Determination of the strategy space for the intelligent agent of electricity users: The strategy space for the intelligent agent of electricity users includes three types of decision variables: 24-hour time-of-use load curve, green certificate holding amount, and green electricity consumption ratio, and corresponding daily total electricity consumption constraints and single-period load capacity upper limits are set.
[0022] S2.3 Determination of the strategy space of the intelligent agent of renewable energy enterprises: The strategy space of the intelligent agent of renewable energy enterprises includes two types of decision variables: time-of-use power output plan and green certificate issuance volume, and corresponding upper limit of installed capacity and tolerance range of power output prediction error are set.
[0023] S2.4 Determination of the regulatory agency's intelligent agent strategy space: The regulatory agency's intelligent agent strategy space includes three types of decision variables: annual carbon quota allocation, green electricity consumption target, and bid collusion prevention threshold, and corresponding adjustment ranges allowed by policy are set; among them, the bid collusion prevention threshold is dynamically adjusted with market concentration. The higher the market concentration, the stricter the bid collusion prevention threshold and the stronger the constraint on bid coordination.
[0024] As a further technical solution of the present invention, the step of S2 constructing the pre-feasible domain constraint includes the following steps:
[0025] S2.5, Establishment of Multi-Constraint Coupled Equations: A multi-constraint coupled equation set is established around the strategy space of four types of intelligent agents, covering hard constraints on power supply and demand, total carbon emission constraints, renewable energy consumption constraints, green certificate supply and demand balance constraints, and anti-collusion constraints on price differences. Specifically: hard constraints on power supply and demand ensure the matching relationship between total power generation output and total load on the power consumption side, as well as grid transmission losses; total carbon emission constraints ensure that the actual carbon emissions of power generation enterprises do not exceed the combined scale of their own carbon quotas and purchased carbon quotas; renewable energy consumption constraints ensure that the proportion of green electricity consumption in the entire market is not lower than the minimum consumption target requirement; green certificate supply and demand balance constraints ensure the matching relationship of the quantity of green certificates issued, traded, and held throughout the entire market chain; and anti-collusion constraints on price differences determine the risk of collusion by combining market concentration and price differences of power generation enterprises, suppressing market manipulation behavior such as abnormally consistent or abnormally divergent pricing.
[0026] S2.6, Implementation of the Dynamic Contraction Mechanism for the Feasible Region: A constraint relaxation level is set for each type of core constraint. A higher relaxation level is set at the initial iteration stage, corresponding to a wider feasible region. During the iteration process, a dynamic decay rule is used to adjust the relaxation level, adaptively narrowing the feasible region based on the deviation between the actual constraint value and the threshold. The higher the deviation, the more significant the relaxation narrowing. The feasible region gradually shrinks during a single iteration. In the long-term parameter iteration stage, the feasible region is reset or adjusted based on the updated constraint threshold and market boundary. All candidate strategies must fall within the feasible region before proceeding to the subsequent evolutionary game calculation stage.
[0027] As a further technical solution of the present invention, the S3 method for constructing the dynamic equations of a multi-agent evolutionary game with Lagrange constraint correction includes the following steps:
[0028] S3.1 Generation of policy payoff vector and policy distribution vector: Based on heterogeneous payoff functions, calculate the payoff levels corresponding to different policies for various agents to form policy payoff vectors; based on policy space, represent the proportion of different agents to choose different policies in the form of probability distribution to form policy distribution vectors.
[0029] S3.2, Heterogeneous Revenue Weight Matrix Calculation: A heterogeneous revenue weight matrix is introduced to quantify the differences in market influence among different intelligent agents. The heterogeneous revenue weight matrix is a diagonal structure, with off-diagonal elements set to zero. The weights of various intelligent agents are calculated by coupling three indicators: market share, information transparency, and strategy adjustment frequency. This comprehensively reflects the market power, information acquisition ability, and strategy response speed of the main body, and all weights meet the normalization requirements.
[0030] S3.3, Reconstructing the Dynamic Equations of Evolutionary Game Theory: The multi-constraint coupled equation set is transformed into a Lagrange constraint penalty mechanism, and the traditional replication dynamic equation is reconstructed to form a reconstructed dynamic equation that integrates payoff-driven and constraint-guided approaches, allowing the strategy distribution to evolve within the pre-feasible domain; wherein: the payoff-driven mechanism gradually increases the proportion of strategies with payoff levels higher than the group average level in the group, and gradually decreases the proportion of strategies with payoff levels lower than the group average level; the Lagrange constraint correction mechanism transforms the constraints into directional guiding forces for strategy evolution, and applies a corrective effect to strategies that are close to the constraint boundary or have a tendency to violate the constraint.
[0031] S3.4, Lagrange Dual Variable Update: The dual variable is updated with an adaptive step size to quantify the penalty of the constraint. When a constraint is violated or close to the violation boundary, the corresponding dual variable is increased synchronously to strengthen the penalty effect of the constraint on the policy evolution. When the constraint is stably satisfied, the dual variable is kept stable or moderately reduced to avoid excessive constraints compressing the policy exploration space.
[0032] As a further technical solution of the present invention, the S4 fusion of ESS and MAPPO algorithms to solve the equilibrium strategy distribution includes the following steps:
[0033] S4.1, ESS Prior Model Construction and Initialization: Based on the strategy evolution law represented by the reconstructed dynamic equation, an evolutionarily stable strategy prior model is constructed. The evolutionarily stable strategy prior model adopts a long short-term memory network structure. The input is a sequence of historical multi-round strategy distributions, and the output is an evolutionarily stable equilibrium prior strategy distribution. When running for the first time and without historical convergence results, cold start initialization is completed using historical market states and strategy samples, uniform distributions, or industry experience rules. In subsequent runs, the prior model is updated on a rolling basis using the converged strategy distributions generated by historical iterations.
[0034] S4.2, Policy Distribution Weighted Fusion Mechanism: The final policy distribution is formed by fusing the original output of the MAPPO policy network with the stable equilibrium prior distribution of the ESS. In the early stage of training, the guidance weight of the ESS prior is relatively high, which prioritizes guiding the policy to explore the historical stable equilibrium region. During the training process, the weight ratio is gradually adjusted to strengthen the leading role of MAPPO autonomous learning and achieve a balance between stable prior guidance and reinforcement learning autonomous optimization. The ESS prior network is trained by minimizing the distribution difference to keep the prior distribution consistent with the historical stable equilibrium policy distribution.
[0035] S4.3, Multi-market Coupled Attention Value Network: Introduce a multi-market coupled attention mechanism into the MAPPO value network, and construct value assessment networks corresponding to the three sub-markets of electricity market, carbon market and green certificate market respectively. Quantify the linkage effect between different sub-markets through attention weights, so that the global value estimate reflects the comprehensive benefit level under the coupling effect of the three markets.
[0036] S4.4, Optimization objective with equilibrium stability penalty term: The strategy optimization objective introduces an equilibrium stability penalty mechanism on the basis of the clipping optimization framework to impose constraints on the strategy iteration that deviates from the stable equilibrium state and suppress the oscillation phenomenon in the strategy iteration process; The advantage function is calculated using the generalized advantage estimation method to balance the evaluation weights of short-term and long-term returns.
[0037] As a further technical solution of the present invention, the S4 fusion of ESS and MAPPO algorithms for solving the equilibrium strategy distribution further includes the following steps:
[0038] S4.5 Iterative Training Initialization: Set the network parameters of the policy network, value network, and ESS prior model; initialize the dual variables and heterogeneous reward weight matrix; and set the maximum number of iterations and the number of iterations per round.
[0039] S4.6, State Acquisition: Each agent makes decisions based on the current policy distribution and market state, collects state, policy, profit, and next state samples, and stores them in the trajectory sampling cache pool used to store the current round's sampling trajectory.
[0040] S4.7, Network Update: Extract a small batch of samples from the trajectory sampling buffer pool and update the parameters of the ESS prior model, policy network and value network synchronously.
[0041] S4.8, Convergence Determination: The degree of change in policy distribution between adjacent iterations is measured by multi-dimensional comprehensive distance. When the change in policy distribution in multiple consecutive iterations is not greater than the preset convergence threshold, the algorithm is determined to have reached evolutionary stable equilibrium and outputs the balanced policy distribution of the four types of agents.
[0042] As a further technical solution of the present invention, the S5 output optimal strategy includes the following steps:
[0043] S5.1, Search Prior Setting Guided by Balanced Distribution: Using the balanced policy distribution as the search prior, the search range and initialization conditions for policy optimization are determined; the high-probability policy interval corresponding to the balanced policy distribution is used as the priority search neighborhood of the optimization algorithm, and the central tendency value of the distribution is used as the initial solution or initial population center of the optimization algorithm, so that the final optimal policy falls within the evolutionary stable region.
[0044] S5.2, Differentiated Optimization Solution for Different Agents: For the decision-making scenarios of four types of intelligent agents, appropriate optimization algorithms are matched to solve the executable optimal strategy within their own strategy space; Specifically: For the power generation enterprise intelligent agent, for the optimization objective of pricing and trading strategies, particle swarm optimization algorithm is used to solve the executable optimal strategy that maximizes comprehensive benefits; For the electricity user intelligent agent, for the optimization objective of energy consumption and green electricity consumption strategies, genetic algorithm is used to solve the executable optimal strategy that maximizes utility and additional benefits; For the renewable energy enterprise intelligent agent, for the optimization objective of time-of-use power output strategies, dynamic programming algorithm is used to solve the executable optimal strategy that maximizes the benefits of green electricity and green certificates; For the regulatory agency intelligent agent, for the optimization objective of multi-objective regulation strategies, third-generation non-dominated sorting genetic algorithm is used to solve the executable regulation strategy that achieves synergistic optimization of multiple common objectives.
[0045] As a further technical solution of the present invention, the multi-market collaborative operation strategy, full-constraint feasibility verification, and feedback closed loop in S5 include the following steps:
[0046] S5.3, Construction of a Multi-Market Collaborative Operation Strategy: Based on the executable optimal strategies of four types of intelligent agents, a multi-market collaborative operation strategy for electricity, carbon, and green certificates is constructed. The carbon price linkage equilibrium level is determined according to the carbon-green certificate price linkage sensitivity coefficient and the current average transaction price in the green certificate market. Positive and negative deviations are distinguished by the degree of deviation between the actual carbon price and the equilibrium level. Based on a benchmark arbitrage threshold, the trigger threshold is dynamically adjusted in conjunction with the electricity price volatility level. The higher the electricity price volatility, the higher the trigger threshold; the lower the electricity price volatility, the lower the trigger threshold. When the actual carbon price is higher than the equilibrium level and the deviation exceeds the threshold, the regulatory agency increases the supply of carbon allowances and guides renewable energy companies to increase the scale of green certificate issuance, thus promoting the return of carbon prices to equilibrium. When the actual carbon price is lower than the equilibrium level and the deviation exceeds the threshold, the regulatory agency reduces the supply of carbon allowances and guides power generation companies to increase the scale of carbon allowance procurement and electricity users to increase the scale of green certificate holdings, thus promoting the return of carbon prices to equilibrium.
[0047] S5.4, Full-constraint feasibility verification: Reuse the multi-constraint coupled equation set corresponding to the previous feasible domain to perform full-constraint feasibility verification on all executable optimal strategies to ensure that the final output strategy simultaneously meets all constraints of power supply and demand, total carbon emissions, renewable energy consumption, green certificate supply and demand balance and anti-collusion of price differences;
[0048] S5.5, Short-term deviation correction mechanism: Real-time market operation data is collected on a 24-hour cycle. After normalizing the multi-dimensional state indicators, the comprehensive deviation between the actual operation state and the equilibrium state is calculated. The deviation judgment threshold is dynamically set according to the intensity of market fluctuations. The stronger the market fluctuations, the higher the sensitivity of feedback recognition. When the deviation does not exceed the threshold, the current strategy is maintained and monitoring continues. When the deviation exceeds the threshold, the current actual operation state is used as the new initial value, and the game dynamic equation is returned to the iteration stage to update the equilibrium strategy and the optimal execution strategy.
[0049] S5.6, Long-term parameter iteration mechanism: Statistical analysis is performed on the accumulated running data on a weekly basis to update the profit model parameters, price linkage sensitivity coefficient, constraint threshold and feasible domain range; After the parameter update is completed, the game dynamic equation stage of S3 is returned to start the full-link re-solution to ensure that the model and strategy continuously adapt to long-term changes in market structure and policy rules.
[0050] Compared with existing technologies, the beneficial effects of this electricity-carbon-green certificate coupled market evolution game method that considers the payoff function are:
[0051] (1) This invention divides the participants in the electricity-carbon-green certificate coupling market into intelligent agents of power generation enterprises, intelligent agents of electricity users, intelligent agents of renewable energy enterprises, and intelligent agents of regulatory agencies, and constructs heterogeneous benefit functions corresponding to the market roles and decision-making objectives of each type of intelligent agent. This enables the benefit model to simultaneously represent the cross-market benefit coupling relationship and the interaction relationship of the agent strategies. Compared with the traditional single-market benefit superposition model, this invention can more accurately reflect the real benefit changes of different agents in the three-market linkage scenario, and provide a more reliable benefit calculation basis for subsequent strategy evolution and strategy optimization.
[0052] (2) In this invention, a cross-market coupling benefit term is set in the benefit function of the power generation enterprise's intelligent agent, which is between the electricity market, carbon market and green certificate market. The linkage sensitivity coefficient of carbon-green certificate price is used to characterize the linkage effect between electricity price, carbon price and green certificate price, so that the comprehensive benefit of the power generation enterprise when participating in electricity trading, carbon quota trading and green certificate trading can reflect the multi-market price transmission relationship, thereby improving the matching degree between the power generation enterprise's strategy optimization results and the actual market operation status.
[0053] (3) This invention establishes a set of multi-constraint coupled equations around the strategy space of four types of intelligent agents, and dynamically shrinks the feasible region through constraint relaxation, so that the candidate strategies are subject to the constraints of the pre-feasible region such as power supply and demand, carbon emissions, renewable energy consumption, green certificate supply and demand balance and bid collusion prevention before entering the game evolution. This method can reduce the entry of invalid candidate strategies into the subsequent evolution process, and gradually improve the feasibility of later strategies while ensuring the early strategy exploration space, thereby improving the efficiency of strategy optimization and the feasibility of the results.
[0054] (4) Based on the replication dynamic equation, this invention introduces a heterogeneous payoff weight matrix and a Lagrange dual constraint correction term. The heterogeneous payoff weight matrix represents the difference in influence caused by different agents’ market share, information transparency and policy adjustment frequency. The Lagrange dual constraint correction term applies correction to policies that are close to the constraint boundary or have a tendency to violate the constraint. This makes the policy evolution driven by both payoff improvement and constraint satisfaction, thereby improving the adaptability of the game dynamic equation to the complex constraint environment of the coupled market.
[0055] (5) This invention weightedly fuses the stable equilibrium prior policy distribution output by the ESS prior model with the policy distribution output by the MAPPO policy network, and introduces a multi-market coupling attention mechanism in the MAPPO value network, adding an equilibrium stability penalty term to the policy optimization objective; thereby, it can use stable priors to guide policy search in the early stage of training, reducing invalid exploration; during the training process, the multi-market coupling attention mechanism improves the ability of value estimation to express the linkage relationship of the three markets of electricity, carbon, and green certificates; and the equilibrium stability penalty term suppresses policy oscillations, improving the convergence speed and stability of the equilibrium policy distribution solution.
[0056] (6) This invention uses the equilibrium strategy distribution as a search prior to determine the search range and initialization conditions for strategy optimization. Then, it combines the strategy variable types and optimization objectives of the four types of agents to solve the executable optimal strategy. This method avoids the problem that the output of only probabilistic equilibrium results is difficult to implement directly. It can transform the evolutionary stable equilibrium results into the bidding and trading strategies of power generation enterprises, the energy consumption and green electricity consumption strategies of electricity users, the time-sharing power output strategies of renewable energy enterprises, and the multi-objective regulation strategies of regulatory agencies, thereby enhancing the executability of the strategy output.
[0057] (7) This invention constructs a multi-market collaborative operation strategy for electricity, carbon and green certificates based on an executable optimal strategy, and reuses the multi-constraint coupled equation set corresponding to the previous feasible domain for full constraint feasibility verification. This ensures that the final output strategy simultaneously meets the constraints of electricity supply and demand, total carbon emissions, renewable energy consumption, green certificate supply and demand balance and bid collusion prevention, thereby improving the compliance and implementation reliability of the multi-market collaborative regulation strategy.
[0058] (8) This invention establishes a feedback loop that combines short-term deviation correction with long-term parameter iteration; short-term feedback calculates the deviation between the actual operating state and the equilibrium state based on the real-time operating state, and returns to the game dynamic equation stage for re-iteration when the deviation exceeds the threshold, updating the equilibrium strategy and the optimal execution strategy; long-term feedback updates the profit model parameters, price linkage sensitivity coefficient, constraint threshold and feasible domain range based on the cumulative operating data; through the above feedback mechanism, this invention can continuously adapt to changes in market structure, policy rules and operating state, and improve the long-term adaptability and stability of the strategy. Attached Figure Description
[0059] Figure 1 This is a schematic diagram of the method flow of the present invention. Detailed Implementation
[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0061] Please see Figure 1 The present invention provides an embodiment of an evolutionary game theory method for a coupled market of electricity, carbon, and green certificates, considering the payoff function, comprising the following steps:
[0062] S1. Construct a multi-agent heterogeneous revenue model: Divide the market participants in the electricity-carbon-green certificate coupling into power generation enterprise agents, electricity user agents, renewable energy enterprise agents, and regulatory agency agents; construct heterogeneous revenue functions according to the market roles and decision-making objectives of each type of agent;
[0063] S2, Define the agent policy space and construct the pre-feasibility domain constraints: Based on the heterogeneous benefit functions and decision objectives of various agents in S1, determine the policy variables of the four types of agents and form corresponding policy spaces; establish a multi-constraint coupled equation system around the policy space, which includes at least hard constraints on power supply and demand, total carbon emission constraints, renewable energy consumption constraints, green certificate supply and demand balance constraints, and price difference anti-collusion constraints; set constraint relaxation and adopt a dynamic contraction mechanism to adjust the feasible domain range so that the candidate strategy satisfies the pre-feasibility domain constraints before entering the policy evolution stage;
[0064] S3: Construct a multi-agent evolutionary game dynamic equation with Lagrange constraint correction: Calculate the policy payoffs of various agents based on the heterogeneous payoff function in S1, and form a policy distribution based on the policy space in S2; introduce a heterogeneous payoff weight matrix, and introduce the multi-constraint coupling equation set in S2 into the Lagrange function; add a Lagrange dual constraint correction term to the replicated dynamic equation to form a reconstructed dynamic equation, so that the policy distribution evolves in the prior feasible region according to payoff-driven and constraint correction, and adaptively updates the Lagrange dual variables;
[0065] S4, Solving the equilibrium policy distribution by integrating ESS and MAPPO algorithms: Based on the policy evolution relationship represented by the reconstructed dynamic equation in S3, an ESS prior model is constructed and a stable equilibrium prior policy distribution is output. The stable equilibrium prior policy distribution is then weighted and fused with the policy distribution output by the MAPPO policy network. A multi-market coupled attention mechanism is introduced into the MAPPO value network, and an equilibrium stability penalty term is added to the policy optimization objective. After iterative training, the equilibrium policy distribution of four types of agents is output.
[0066] S5, Output the optimal strategy and establish a feedback loop: Based on the equilibrium strategy distribution output in S4, determine the strategy optimization search range and initialization conditions, and according to the strategy variable types and optimization objectives of the four types of intelligent agents, respectively execute the bidding and trading strategy optimization of the power generation enterprise intelligent agent, the energy consumption and green electricity consumption strategy optimization of the electricity user intelligent agent, the time-of-use power output strategy optimization of the renewable energy enterprise intelligent agent, and the multi-objective regulation strategy optimization of the regulatory agency intelligent agent, to obtain the executable optimal strategy for the four types of intelligent agents; based on the executable optimal strategy, construct a multi-market collaborative operation strategy for electricity-carbon-green certificates, and reuse the multi-constraint coupled equation system in S2 for full-constraint feasibility verification; establish a feedback loop, where short-term feedback updates the strategy based on real-time operating status, and long-term feedback updates the revenue model parameters and constraint parameters based on cumulative operating data;
[0067] Furthermore, S1 constructs a multi-agent heterogeneous benefit model including the following steps:
[0068] S1.1, Classification and Decision-Making Objectives of Coupled Market Intelligent Agents: The participants in the coupled electricity-carbon-green certificate market are classified into four heterogeneous intelligent agents: power generation enterprise intelligent agents, electricity user intelligent agents, renewable energy enterprise intelligent agents, and regulatory agency intelligent agents. Among them: power generation enterprise intelligent agents participate in electricity market, carbon market, and green certificate market transactions simultaneously, and their decision-making objective is to maximize the comprehensive benefits of the three markets; electricity user intelligent agents, as the demand-side entities of electricity and green certificates, aim to maximize the comprehensive benefits of electricity consumption utility, carbon emission reduction benefits, and green electricity consumption rewards; renewable energy enterprise intelligent agents, as the entities responsible for green electricity supply and green certificate issuance, aim to maximize the comprehensive benefits of green electricity generation and green certificate revenue; and regulatory agency intelligent agents, as the entities responsible for market rule-making and regulation, aim to collaboratively achieve multiple public objectives, including stable electricity prices, carbon emission reduction, green electricity consumption, and maximizing social welfare.
[0069] S1.2, Construction of Heterogeneous Revenue Functions for Four Types of Intelligent Agents: For the decision-making objectives and trading scope of the four types of intelligent agents, heterogeneous revenue functions with differentiated structures are constructed respectively. Specifically: the revenue function for power generation enterprises comprehensively covers electricity sales revenue, power generation costs, carbon quota trading profits and losses, and green certificate holding revenue, and embeds cross-market coupled revenue terms to characterize the impact of the interaction between electricity prices, carbon prices, and green certificate prices on the enterprise's overall revenue; the revenue function for electricity users comprehensively covers electricity utility, electricity expenditure, and green certificate holding revenue, and introduces two types of strategic interaction revenue terms: implicit carbon emission reduction revenue and excess green electricity consumption rewards; the revenue function for renewable energy enterprises comprehensively covers green electricity sales revenue and green certificate issuance revenue, and introduces a green electricity scarcity premium strategy interaction revenue term; the revenue function for regulatory agencies adopts a multi-objective reward and punishment mechanism, setting penalty terms for electricity price fluctuations, excessive system carbon emissions, and deviations in green electricity consumption, and setting reward terms for the social welfare corresponding to the total market revenue;
[0070] Furthermore, in S1.2, the cross-market coupling revenue term of the power generation enterprise's intelligent agent revenue function includes a carbon-green certificate price linkage sensitivity coefficient, which is calibrated through the following steps:
[0071] S1.2.1, Data Acquisition and Preprocessing: Collect daily price data for the electricity market, carbon market, and green certificate market covering three complete calendar years; perform stationarity tests on the price series; and perform first-order differencing on non-stationary series to obtain stationary series.
[0072] S1.2.2, Model Construction: Construct a third-order lagged vector autoregressive model (VAR(3) model) containing three variables: electricity price, carbon price, and green certificate price, to capture the linkage transmission law of price lag over multiple periods;
[0073] S1.2.3, Parameter Calculation: Model parameters are obtained through regression estimation. The marginal impact of carbon price and green certificate price on electricity price is extracted separately. The carbon-green certificate price linkage sensitivity coefficient is obtained by combining the coupling effect of the two, which is used to quantify the linkage effect of carbon price and green certificate price on electricity price under the joint effect of carbon price and green certificate price.
[0074] Furthermore, S2 defines the agent's policy space by the following steps:
[0075] S2.1 Determination of the strategy space of the power generation enterprise's intelligent agent: The strategy space of the power generation enterprise's intelligent agent includes three types of decision variables: bidding coefficient, carbon quota purchase volume, and green certificate holding volume, and correspondingly sets a reasonable bidding range, carbon quota purchase upper limit, and green certificate holding upper limit;
[0076] S2.2 Determination of the strategy space for the intelligent agent of electricity users: The strategy space for the intelligent agent of electricity users includes three types of decision variables: 24-hour time-of-use load curve, green certificate holding amount, and green electricity consumption ratio, and corresponding daily total electricity consumption constraints and single-period load capacity upper limits are set.
[0077] S2.3 Determination of the strategy space of the intelligent agent of renewable energy enterprises: The strategy space of the intelligent agent of renewable energy enterprises includes two types of decision variables: time-of-use power output plan and green certificate issuance volume, and corresponding upper limit of installed capacity and tolerance range of power output prediction error are set.
[0078] S2.4 Determination of the regulatory agency's intelligent agent strategy space: The regulatory agency's intelligent agent strategy space includes three types of decision variables: annual carbon quota allocation, green electricity consumption target, and bid collusion prevention threshold, and corresponding adjustment ranges allowed by policy are set; among them, the bid collusion prevention threshold is dynamically adjusted with market concentration. The higher the market concentration, the stricter the bid collusion prevention threshold and the stronger the constraint on bid coordination.
[0079] Furthermore, the S2 construction of the preceding feasible region constraints includes the following steps:
[0080] S2.5, Establishment of Multi-Constraint Coupled Equations: A multi-constraint coupled equation set is established around the strategy space of four types of intelligent agents, covering hard constraints on power supply and demand, total carbon emission constraints, renewable energy consumption constraints, green certificate supply and demand balance constraints, and anti-collusion constraints on price differences. Specifically: hard constraints on power supply and demand ensure the matching relationship between total power generation output and total load on the power consumption side, as well as grid transmission losses; total carbon emission constraints ensure that the actual carbon emissions of power generation enterprises do not exceed the combined scale of their own carbon quotas and purchased carbon quotas; renewable energy consumption constraints ensure that the proportion of green electricity consumption in the entire market is not lower than the minimum consumption target requirement; green certificate supply and demand balance constraints ensure the matching relationship of the quantity of green certificates issued, traded, and held throughout the entire market chain; and anti-collusion constraints on price differences determine the risk of collusion by combining market concentration and price differences of power generation enterprises, suppressing market manipulation behavior such as abnormally consistent or abnormally divergent pricing.
[0081] S2.6, Implementation of the Dynamic Contraction Mechanism for the Feasible Region: A constraint relaxation level is set for each type of core constraint. A higher relaxation level is set at the initial iteration stage, corresponding to a wider feasible region. During the iteration process, a dynamic decay rule is used to adjust the relaxation level, adaptively narrowing the feasible region based on the deviation between the actual constraint value and the threshold. The higher the deviation, the more significant the relaxation narrowing. The feasible region gradually shrinks during a single iteration. In the long-term parameter iteration stage, the feasible region is reset or adjusted based on the updated constraint threshold and market boundary. All candidate strategies must fall within the feasible region before proceeding to the subsequent evolutionary game calculation stage.
[0082] Furthermore, S3 constructs the dynamic equations for a multi-agent evolutionary game with Lagrange constraint corrections, including the following steps:
[0083] S3.1 Generation of policy payoff vector and policy distribution vector: Based on heterogeneous payoff functions, calculate the payoff levels corresponding to different policies for various agents to form policy payoff vectors; based on policy space, represent the proportion of different agents to choose different policies in the form of probability distribution to form policy distribution vectors.
[0084] S3.2, Heterogeneous Revenue Weight Matrix Calculation: A heterogeneous revenue weight matrix is introduced to quantify the differences in market influence among different intelligent agents. The heterogeneous revenue weight matrix is a diagonal structure, with off-diagonal elements set to zero. The weights of various intelligent agents are calculated by coupling three indicators: market share, information transparency, and strategy adjustment frequency. This comprehensively reflects the market power, information acquisition ability, and strategy response speed of the main body, and all weights meet the normalization requirements.
[0085] S3.3, Reconstructing the Dynamic Equations of Evolutionary Game Theory: The multi-constraint coupled equation set is transformed into a Lagrange constraint penalty mechanism, and the traditional replication dynamic equation is reconstructed to form a reconstructed dynamic equation that integrates payoff-driven and constraint-guided approaches, allowing the strategy distribution to evolve within the pre-feasible domain; wherein: the payoff-driven mechanism gradually increases the proportion of strategies with payoff levels higher than the group average level in the group, and gradually decreases the proportion of strategies with payoff levels lower than the group average level; the Lagrange constraint correction mechanism transforms the constraints into directional guiding forces for strategy evolution, and applies a corrective effect to strategies that are close to the constraint boundary or have a tendency to violate the constraint.
[0086] S3.4, Lagrange Dual Variable Update: The dual variable is used to quantify the penalty of the constraint, and the dual variable is updated with an adaptive step size; when a constraint is violated or close to the violation boundary, the corresponding dual variable is increased synchronously to strengthen the penalty effect of the constraint on the policy evolution; when the constraint is stably satisfied, the dual variable is kept stable or moderately reduced to avoid excessive constraints compressing the policy exploration space.
[0087] Furthermore, the S4 algorithm, which integrates ESS and MAPPO algorithms, to solve for the equilibrium policy distribution includes the following steps:
[0088] S4.1, ESS Prior Model Construction and Initialization: Based on the strategy evolution law represented by the reconstructed dynamic equation, an evolutionarily stable strategy prior model is constructed. The evolutionarily stable strategy prior model adopts a long short-term memory network structure. The input is a sequence of historical multi-round strategy distributions, and the output is an evolutionarily stable equilibrium prior strategy distribution. When running for the first time and without historical convergence results, cold start initialization is completed using historical market states and strategy samples, uniform distributions, or industry experience rules. In subsequent runs, the prior model is updated on a rolling basis using the converged strategy distributions generated by historical iterations.
[0089] S4.2, Policy Distribution Weighted Fusion Mechanism: The final policy distribution is formed by fusing the original output of the MAPPO policy network with the stable equilibrium prior distribution of the ESS. In the early stage of training, the guidance weight of the ESS prior is relatively high, which prioritizes guiding the policy to explore the historical stable equilibrium region. During the training process, the weight ratio is gradually adjusted to strengthen the leading role of MAPPO autonomous learning and achieve a balance between stable prior guidance and reinforcement learning autonomous optimization. The ESS prior network is trained by minimizing the distribution difference to keep the prior distribution consistent with the historical stable equilibrium policy distribution.
[0090] S4.3, Multi-market Coupled Attention Value Network: Introduce a multi-market coupled attention mechanism into the MAPPO value network, and construct value assessment networks corresponding to the three sub-markets of electricity market, carbon market and green certificate market respectively. Quantify the linkage effect between different sub-markets through attention weights, so that the global value estimate reflects the comprehensive benefit level under the coupling effect of the three markets.
[0091] S4.4, Optimization objective with equilibrium stability penalty term: The strategy optimization objective introduces an equilibrium stability penalty mechanism on the basis of the clipping optimization framework to impose constraints on the strategy iteration that deviates from the stable equilibrium state and suppress the oscillation phenomenon in the strategy iteration process; The advantage function is calculated using the generalized advantage estimation method to balance the evaluation weights of short-term and long-term returns.
[0092] Furthermore, the S4 fusion of ESS and MAPPO algorithms for solving the equilibrium policy distribution also includes the following steps:
[0093] S4.5 Iterative Training Initialization: Set the network parameters of the policy network, value network, and ESS prior model; initialize the dual variables and heterogeneous reward weight matrix; and set the maximum number of iterations and the number of iterations per round.
[0094] S4.6, State Acquisition: Each agent makes decisions based on the current policy distribution and market state, collects state, policy, profit, and next state samples, and stores them in the trajectory sampling cache pool used to store the current round's sampling trajectory.
[0095] S4.7, Network Update: Extract a small batch of samples from the trajectory sampling buffer pool and update the parameters of the ESS prior model, policy network and value network synchronously.
[0096] S4.8, Convergence Determination: The degree of change in policy distribution between adjacent iterations is measured by multi-dimensional comprehensive distance. When the change in policy distribution in multiple consecutive iterations is not greater than the preset convergence threshold, the algorithm is determined to have reached evolutionary stable equilibrium and outputs the balanced policy distribution of the four types of agents.
[0097] Furthermore, the optimal strategy output by S5 includes the following steps:
[0098] S5.1, Search Prior Setting Guided by Balanced Distribution: Using the balanced policy distribution as the search prior, the search range and initialization conditions for policy optimization are determined; the high-probability policy interval corresponding to the balanced policy distribution is used as the priority search neighborhood of the optimization algorithm, and the central tendency value of the distribution is used as the initial solution or initial population center of the optimization algorithm, so that the final optimal policy falls within the evolutionary stable region.
[0099] S5.2, Differentiated Optimization Solution for Different Agents: For the decision-making scenarios of four types of intelligent agents, appropriate optimization algorithms are matched to solve the executable optimal strategy within their own strategy space; Specifically: For the power generation enterprise intelligent agent, for the optimization objective of pricing and trading strategies, particle swarm optimization algorithm is used to solve the executable optimal strategy that maximizes comprehensive benefits; For the electricity user intelligent agent, for the optimization objective of energy consumption and green electricity consumption strategies, genetic algorithm is used to solve the executable optimal strategy that maximizes utility and additional benefits; For the renewable energy enterprise intelligent agent, for the optimization objective of time-of-use power output strategies, dynamic programming algorithm is used to solve the executable optimal strategy that maximizes the benefits of green electricity and green certificates; For the regulatory agency intelligent agent, for the optimization objective of multi-objective regulation strategies, third-generation non-dominated sorting genetic algorithm is used to solve the executable regulation strategy that achieves synergistic optimization of multiple common objectives.
[0100] Furthermore, the multi-market collaborative operation strategy, full-constraint feasibility verification, and feedback loop in S5 include the following steps:
[0101] S5.3, Construction of a Multi-Market Collaborative Operation Strategy: Based on the executable optimal strategies of four types of intelligent agents, a multi-market collaborative operation strategy for electricity, carbon, and green certificates is constructed. The carbon price linkage equilibrium level is determined according to the carbon-green certificate price linkage sensitivity coefficient and the current average transaction price in the green certificate market. Positive and negative deviations are distinguished by the degree of deviation between the actual carbon price and the equilibrium level. Based on a benchmark arbitrage threshold, the trigger threshold is dynamically adjusted in conjunction with the electricity price volatility level. The higher the electricity price volatility, the higher the trigger threshold; the lower the electricity price volatility, the lower the trigger threshold. When the actual carbon price is higher than the equilibrium level and the deviation exceeds the threshold, the regulatory agency increases the supply of carbon allowances and guides renewable energy companies to increase the scale of green certificate issuance, thus promoting the return of carbon prices to equilibrium. When the actual carbon price is lower than the equilibrium level and the deviation exceeds the threshold, the regulatory agency reduces the supply of carbon allowances and guides power generation companies to increase the scale of carbon allowance procurement and electricity users to increase the scale of green certificate holdings, thus promoting the return of carbon prices to equilibrium.
[0102] S5.4, Full-constraint feasibility verification: Reuse the multi-constraint coupled equation set corresponding to the previous feasible domain to perform full-constraint feasibility verification on all executable optimal strategies to ensure that the final output strategy simultaneously meets all constraints of power supply and demand, total carbon emissions, renewable energy consumption, green certificate supply and demand balance and anti-collusion of price differences;
[0103] S5.5, Short-term deviation correction mechanism: Real-time market operation data is collected on a 24-hour cycle. After normalizing the multi-dimensional state indicators, the comprehensive deviation between the actual operation state and the equilibrium state is calculated. The deviation judgment threshold is dynamically set according to the intensity of market fluctuations. The stronger the market fluctuations, the higher the sensitivity of feedback recognition. When the deviation does not exceed the threshold, the current strategy is maintained and monitoring continues. When the deviation exceeds the threshold, the current actual operation state is used as the new initial value, and the game dynamic equation is returned to the iteration stage to update the equilibrium strategy and the optimal execution strategy.
[0104] S5.6, Long-term parameter iteration mechanism: Statistical analysis is performed on the accumulated running data on a weekly basis to update the profit model parameters, price linkage sensitivity coefficient, constraint threshold and feasible domain range; After the parameter update is completed, the game dynamic equation stage of S3 is returned to start the full-link re-solution to ensure that the model and strategy continuously adapt to long-term changes in market structure and policy rules.
[0105] An application example provided: S1 Constructing a multi-agent heterogeneous benefit model: This step is used to clarify the market roles, decision-making objectives, and benefit composition of each participant in the electricity-carbon-green certificate coupled market, providing a basis for benefit calculation for subsequent strategy evolution and strategy optimization;
[0106] S1.1 Classification and Decision-Making Objectives of Coupled Market Agents:
[0107] The participants in the electricity-carbon-green certificate coupling market are divided into four heterogeneous intelligent entities: power generation enterprise intelligent entities, electricity user intelligent entities, renewable energy enterprise intelligent entities, and regulatory agency intelligent entities.
[0108] Among them, the power generation enterprise intelligent agent participates in electricity market, carbon market and green certificate market transactions at the same time. Its decision-making objective is to maximize the comprehensive benefits of the three markets while taking into account power generation costs, carbon emission constraints and green certificate holding requirements.
[0109] As the demand-side entity for electricity and green certificates, the intelligent agent of electricity users aims to maximize the comprehensive benefits of electricity utility, carbon emission reduction benefits and green electricity consumption rewards by adjusting time-of-use electricity behavior, green certificate holdings and green electricity consumption ratio.
[0110] As the main body for green electricity supply and green certificate issuance, the decision-making goal of renewable energy enterprise intelligent agents is to maximize the combined revenue from green electricity generation and green certificate issuance by optimizing time-of-use power output plans and green certificate issuance.
[0111] As the main body for market rule-making and regulation, the decision-making objective of the regulatory agency is to collaboratively achieve multiple public goals, including stable electricity prices, carbon emission reduction, green electricity consumption, and maximization of social welfare.
[0112] S1.2 Construction of Heterogeneous Reward Functions for Four Types of Agents:
[0113] For the decision-making objectives and trading scope of the four types of intelligent agents, heterogeneous revenue functions with different structures are constructed respectively. The revenue function of the power generation enterprise intelligent agent is set with a cross-market coupling revenue term to characterize the impact of the interaction between electricity price, carbon price and green certificate price on the comprehensive revenue of the power generation enterprise. The revenue functions of the electricity user intelligent agent, renewable energy enterprise intelligent agent and regulatory agency intelligent agent are respectively set with strategy interaction revenue terms corresponding to their respective strategy behavior or regulatory objectives.
[0114] The revenue function of a power generation company's intelligent agent includes electricity sales revenue, power generation costs, carbon quota trading profits and losses, green certificate holding revenue, and cross-market coupling revenue items, specifically:
[0115]
[0116] In the formula, For the first The revenue of an intelligent agent of a power generation enterprise; To clear electricity prices in the electricity market; For the first Each power generation company contributed its efforts; For power generation companies, the power generation cost function is used. For carbon trading prices; Carbon emission quotas for power generation companies; Carbon emission factor per unit of electricity generation; The price of a green certificate; Green certificate holdings; The carbon-green certificate price linkage sensitivity coefficient; , These represent the degree of correlation between carbon prices and green certificate prices and electricity prices, respectively.
[0117] The power generation cost function uses a quadratic cost function:
[0118]
[0119] In the formula, , , The cost coefficient for power generation companies can be obtained by fitting historical power generation data.
[0120] The revenue function for electricity users' intelligent agents includes electricity utility, electricity cost expenditure, green certificate holding revenue, implicit carbon emission reduction revenue, and excess green electricity consumption reward, specifically:
[0121]
[0122] In the formula, For the first The benefits of an individual electricity user's intelligent agent; This is a function for the efficiency of electricity use; For the first Time-of-use load of individual electricity users; For users' green certificate holdings; The implicit value of carbon emission reduction per unit; Carbon emission reductions for users; This refers to the excess consumption reward coefficient. The actual green electricity consumption ratio for users; The minimum green energy consumption ratio;
[0123] The electricity utility function is in logarithmic form:
[0124]
[0125] In the formula, This is the electricity efficiency coefficient;
[0126] User carbon emission reductions can be expressed as:
[0127]
[0128] In the formula, For users' green electricity consumption ratio, The average carbon emission factor for power generation across the entire grid;
[0129] The revenue function of a renewable energy enterprise includes revenue from green electricity sales, revenue from green certificate issuance, and revenue from the scarcity premium of green electricity, specifically:
[0130]
[0131] In the formula, For the first The revenue of a renewable energy enterprise intelligent agent; Provide time-sharing services for renewable energy companies; This refers to the number of green certificates issued. The premium coefficient for the scarcity of green electricity; For indicator functions; This represents the actual proportion of green electricity consumed. The target is the proportion of green electricity consumption;
[0132] The number of green certificates issued corresponds to the amount of green electricity output, expressed as follows:
[0133]
[0134] In the formula, This refers to the green certificate issuance coefficient.
[0135] The green electricity scarcity premium coefficient can be dynamically adjusted according to the green electricity consumption gap:
[0136]
[0137] In the formula, Based on the base premium coefficient, For adjustment coefficients; when hour, This triggers a green electricity scarcity premium; when hour, ;
[0138] The regulatory agency's intelligent agent revenue function adopts a multi-objective reward and punishment mechanism, including penalties for electricity price fluctuations, system carbon emissions, green electricity consumption deviations, and social welfare rewards, specifically:
[0139]
[0140] In the formula, For the benefit of the regulatory agency's intelligent agent; This represents the variance of electricity prices; Total carbon emissions of the system; , , , For multi-objective preference weights, satisfying ; It is the sum of the benefits of all types of intelligent agents, used to characterize the level of social welfare;
[0141] The variance of electricity prices can be expressed as:
[0142]
[0143] The total carbon emissions of the system can be expressed as:
[0144]
[0145] S1.2.1 Data Acquisition and Preprocessing:
[0146] Carbon-Green Certificate Price Linkage Sensitivity Coefficient As the core calculation parameter of the cross-market coupling benefit term of the power generation enterprise intelligent agent, it is calibrated by the third-order vector autoregression VAR(3) model;
[0147] First, daily price data for the electricity market, carbon market, and green certificate market were collected covering three complete calendar years, including electricity market clearing prices. Carbon trading prices And green certificate price Perform ADF stationarity test on the price series, and perform first-order differencing on the non-stationary price series to obtain the stationary price series. , and ;
[0148] S1.2.2 Model Construction:
[0149] A third-order lag vector autoregressive model (VAR(3) model) was constructed, which includes three variables: electricity price, carbon price, and green certificate price.
[0150]
[0151] In the formula, A vector of constant terms. , , The coefficient matrix, The vector is the perturbation term.
[0152] S1.2.3 Parameter Calculation:
[0153] By estimating the model parameters using OLS regression, the degree of linkage between carbon price and green certificate price and electricity price is extracted, and the carbon-green certificate price linkage sensitivity coefficient is calculated according to the following formula:
[0154]
[0155] From this, we obtain It is used as a cross-market coupling revenue term in the revenue function of the power generation enterprise's intelligent agent, and can be used in the calculation of carbon price linkage equilibrium value in subsequent multi-market collaborative operation strategies;
[0156] S2 defines the agent's policy space and constructs pre-feasible domain constraints: This step follows the heterogeneous payoff model in S1, which is used to clarify the range of decision-making variables for the four types of agents and establish pre-hard constraint boundaries so that candidate strategies meet the requirements of power system physical rules, carbon emission rules, green certificate trading rules and market competition order before entering the game evolution.
[0157] S2.1 Determining the strategy space of the power generation enterprise's intelligent agent:
[0158] The strategy space of the power generation enterprise's intelligent agent is:
[0159]
[0160] In the formula, For the first The quotation coefficient for each power generation company; For carbon allowance purchases; Green certificate holdings; Set within a preset reasonable price range. Purchases shall not exceed the preset carbon quota limit. Not exceeding the preset green certificate holding limit;
[0161] S2.2 Determination of the policy space for electricity user intelligent agents:
[0162] The policy space for the electricity user's intelligent agent is:
[0163]
[0164] In the formula, For the first Individual electricity users 24-hour time-of-use load curve for a given period; For users' green certificate holdings; The proportion of green electricity consumption; To meet the daily total electricity consumption constraints and the upper limit of single-period load capacity. satisfy ;
[0165] S2.3 Determination of the strategy space for renewable energy enterprise intelligent agents:
[0166] The strategy space for renewable energy enterprise intelligent agents is:
[0167]
[0168] In the formula, For the first A renewable energy company's time-sharing power output plan; This refers to the number of green certificates issued. It shall not exceed the upper limit of the installed capacity and shall meet the tolerance range of the power output prediction error;
[0169] S2.4 Determining the policy space of the regulatory agency's intelligent agent:
[0170] The policy space of the regulatory agency's intelligent agent is:
[0171]
[0172] In the formula, Carbon quota allocation; To achieve the goal of green energy consumption; A threshold for preventing collusion in pricing; The market concentration is dynamically adjusted. The higher the market concentration, the stricter the threshold for preventing collusion in bidding, and the stronger the constraint on bid coordination.
[0173] Market concentration is represented by the HHI index:
[0174]
[0175] In the formula, For the first The market share of each power generation company; the bid anti-collusion threshold is expressed as:
[0176]
[0177] In the formula, Based on the threshold, For adjustment coefficients;
[0178] S2.5 Establishment of the multi-constraint coupled equation system:
[0179] A set of multi-constraint coupled equations is established around the policy space of four types of intelligent agents. The set of multi-constraint coupled equations covers hard constraints on power supply and demand, total carbon emission constraints, renewable energy consumption constraints, green certificate supply and demand balance constraints, and price difference anti-collusion constraints.
[0180] The hard constraints of electricity supply and demand are:
[0181]
[0182] In the formula, For power grid transmission losses;
[0183] The total carbon emission limit is:
[0184]
[0185] The constraints for renewable energy consumption are:
[0186]
[0187] The supply and demand balance constraint for green certificates is:
[0188]
[0189] In the formula, The sales volume of green certificates by power generation companies or renewable energy companies. For electricity users' green certificate purchase volume, This represents the total number of green certificates issued.
[0190] The anti-collusion constraint for price discrepancies is:
[0191]
[0192] In the formula, To exclude the first The average bid coefficient of power generation companies other than individual power generation companies;
[0193] S2.6 Implementation of the feasible domain dynamic shrinkage mechanism:
[0194] Set constraint slack for each type of core constraint. ,in For constraint sequence number, For the iteration time; a high relaxation level is set in the initial stage, for example... This provides a wider space for strategy exploration;
[0195] During the iterative process, a dynamic decay rule is used to adjust the relaxation:
[0196]
[0197] In the formula, The shrinkage attenuation coefficient; For the first A constraint in The actual value at any given time is determined by the agent's own policy. Other intelligent agent strategies Calculated; For the first The threshold of each constraint;
[0198] During a single iteration, the feasible region gradually shrinks according to the degree of constraint deviation. The higher the degree of constraint deviation, the more obvious the relaxation narrows. In the long-term parameter iteration stage, the feasible region is reset or adjusted according to the updated constraint threshold and market boundary. All candidate strategies must fall within the previous feasible region before they can enter the subsequent game evolution calculation stage.
[0199] S3 Constructs the dynamic equations of multi-agent evolutionary game with Lagrange constraint correction: This step follows the heterogeneous payoff function of S1 and the prior feasible region constraint of S2, and is used to characterize the evolution of the four types of agent strategies under the combined effect of payoff drive and constraint correction.
[0200] S3.1 Generation of strategy payoff vector and strategy distribution vector:
[0201] Based on the heterogeneous payoff functions of the four types of agents in S1, the payoff levels corresponding to different strategies chosen by each type of agent are calculated, forming a strategy payoff vector:
[0202]
[0203] Based on the policy space of the four types of agents in S2, the proportion of each type of agent choosing different policies is represented by a probability distribution, forming a policy distribution vector:
[0204]
[0205] In the formula, , , , Let represent the policy distributions of the power generation company agent, the electricity user agent, the renewable energy company agent, and the regulatory agency agent, respectively, and satisfy the following:
[0206]
[0207] S3.2 Calculation of Heterogeneous Revenue Weight Matrix:
[0208] Introducing a heterogeneous revenue weight matrix This is used to quantify the differences in market influence among different agents in a coupled market; the heterogeneous revenue weight matrix is a diagonal matrix, with off-diagonal elements set to zero. diagonal elements Calculate using the following formula:
[0209]
[0210] In the formula, For the first Market share of intelligent agents; For the first Information transparency of intelligent agents; For the first The policy adjustment frequency of the agent-like system; all weights satisfy the normalization requirement:
[0211]
[0212] S3.3 Reconstructing the dynamic equations of the evolutionary game:
[0213] By introducing the multi-constraint coupled equations in S2 into the Lagrangian function to form a constraint penalty mechanism, and adding a Lagrangian dual constraint correction term to the traditional replicated dynamic equations, the reconstructed dynamic equations are obtained:
[0214]
[0215] In the formula, This is a heterogeneous revenue weight matrix; For A diagonal matrix with diagonal elements; This represents the strategy payoff vector; For the average benefit of the group; for A vector of order all 1s; This is a correction factor; For the Lagrange function with respect to the policy distribution The gradient; Let Lagrange be the dual variable vector;
[0216] Among them, revenue drivers The Lagrange dual constraint correction term is used to gradually increase the proportion of strategies with returns higher than the group average and gradually decrease the proportion of strategies with returns lower than the group average; the Lagrange dual constraint correction term is used to apply a correction effect to strategies that are close to the constraint boundary or have a violation trend, so that the strategy distribution evolves within the previous feasible region.
[0217] The Lagrange function can be expressed as:
[0218]
[0219] In the formula, The objective function is the profit function determined by the profit functions of the four types of agents; This represents the constraint function vector corresponding to the multi-constraint coupled equation system; For the constraint threshold vector;
[0220] S3.4 Lagrange dual variable update:
[0221] The Lagrange dual variable is updated using an adaptive step size; let Indicates the first The degree of violation of a constraint is then determined as follows:
[0222]
[0223] In the formula, , They are respectively , Time of the first The dual variables of the constraints; For adaptive step size, ; This is the initial step size;
[0224] When a constraint is violated or approaches the boundary of violation, the corresponding dual variable increases, strengthening the penalty effect of the constraint on policy evolution; when the constraint is stably satisfied, the dual variable remains stable or decreases moderately, avoiding excessive constraints that compress the policy exploration space.
[0225] S4: Solving the equilibrium policy distribution by fusing ESS and MAPPO algorithms: This step follows the policy evolution relationship in S3 and solves the equilibrium policy distribution of four types of agents by fusing evolutionary stable policy priors with multi-agent reinforcement learning.
[0226] S4.1 ESS Prior Model Construction and Initialization:
[0227] Based on the policy evolution law represented by the reconstructed dynamic equation of S3, an ESS prior model is constructed. The ESS prior model adopts a long short-term memory network structure, with the input being a sequence of historical multi-round policy distributions and the output being an evolutionarily stable equilibrium prior policy distribution.
[0228]
[0229] In the formula, For the first The market state corresponding to intelligent agents. These are the parameters for the ESS prior model;
[0230] When running for the first time and without historical convergence results, cold start initialization is completed using historical market state and strategy samples, uniform distribution or industry experience rules; in subsequent runs, the convergence strategy distribution generated by historical iterations is used to continuously update the ESS prior model.
[0231] The ESS prior model is trained by minimizing the distribution difference:
[0232]
[0233] In the formula, For historically stable equilibrium strategy distribution, Let KL divergence be a metric.
[0234] S4.2 Policy Distribution Weighted Fusion Mechanism:
[0235] The MAPPO policy network outputs the original policy distribution based on the current state, while the ESS prior model outputs a stable equilibrium prior policy distribution. The two are then weighted and fused to obtain the final policy distribution.
[0236] The original policy network output of MAPPO is:
[0237]
[0238] The policy distribution after the ESS-MAPPO weighted fusion is as follows:
[0239]
[0240] In the formula, For the first intelligent agents in state Select action The strategy distribution; The policy network weight matrix; For bias terms; For ESS prior fusion weights; in the early stages of training, Set to 0.3 so that the ESS prior distribution accounts for 30% and the MAPPO raw output accounts for 70%; adjust gradually during training. To strengthen the leading role of MAPPO in autonomous learning;
[0241] S4.3 Multi-Market Coupled Attention Value Network:
[0242] In the MAPPO value network, a multi-market coupling attention mechanism is introduced to construct value assessment networks for the three sub-markets of electricity market, carbon market and green certificate market, respectively, and the linkage effect between different sub-markets is quantified by attention weight;
[0243] The global value function is:
[0244]
[0245] In the formula, For global value function; , , These are the sub-market value functions corresponding to the electricity market, carbon market, and green certificate market, respectively. , For the first , Individual market value function; Attention weights;
[0246] Attention weights are represented as follows:
[0247]
[0248]
[0249] In the formula, This is the attention weight matrix. , These represent the characteristics of different sub-markets;
[0250] S4.4 Optimization objective including equilibrium stability penalty term:
[0251] A stability penalty term is added to the MAPPO policy optimization objective to suppress oscillations during policy iteration; The objective function for optimizing an agent-like system is:
[0252]
[0253] In the formula, For the first The objective function for policy optimization in an agent-like system; For expectation operators; For strategy ratio; The dominant function; The threshold for strategy clipping; To achieve a balanced and stable penalty coefficient; Let the square norm of the Lagrange function gradient be denoted as .
[0254] The strategy ratio is:
[0255]
[0256] The advantage function is calculated using the generalized advantage estimation method:
[0257]
[0258]
[0259] In the formula, This refers to timing difference error; Discount factor; The coefficients are estimates of the generalized advantage.
[0260] S4.5 Iterative Training Initialization:
[0261] Set the network parameters of the policy network, value network, and ESS prior model; initialize the Lagrange dual variables and heterogeneous payoff weight matrix; and set the maximum number of iterations and the number of iterations per round.
[0262] S4.6 Status Acquisition:
[0263] Each agent makes decisions based on the current policy distribution and market state, collects state, policy, profit and next state samples, and stores them in a trajectory sampling cache pool used to store the current round of sampling trajectories;
[0264] S4.7 Network Update:
[0265] Mini-batch samples are extracted from the trajectory sampling cache pool, and the parameters of the ESS prior model, policy network, and value network are updated synchronously.
[0266] S4.8 Convergence Criterion:
[0267] Calculate the change in policy distribution between adjacent iterations:
[0268]
[0269] when Furthermore, if this condition is met for 10 consecutive iterations, the algorithm is considered to have reached an evolutionarily stable equilibrium, and the equilibrium policy distribution of the four types of agents is output:
[0270]
[0271] S5 outputs the optimal strategy and establishes a feedback loop: This step follows the equilibrium strategy distribution output by S4, transforms the probabilistic equilibrium result into an executable optimal strategy, and establishes a feedback loop for short-term deviation correction and long-term parameter iteration.
[0272] S5.1 Equivalent Distribution-Guided Search Prior Settings:
[0273] Equalization strategy distribution based on S4 output To establish search priors, the search range and initialization conditions for policy optimization are determined. The high-probability policy interval corresponding to the equilibrium policy distribution is used as the priority search neighborhood of the optimization algorithm, and the central tendency value of the distribution is used as the initial solution or initial population center of the optimization algorithm, so that the optimal policy output is located in the evolutionary stable region and avoids deviating from the game equilibrium.
[0274] S5.2 Sub-subject Differentiation Optimization Solution:
[0275] For the decision-making scenarios of four types of intelligent agents, appropriate optimization algorithms are matched and the optimal executable policy is solved in its own policy space.
[0276] The optimal strategy for the power generation company's intelligent agent is:
[0277]
[0278] Specifically, the distribution of equilibrium strategies Substituting the revenue function of the power generation company's intelligent agent into the strategy space... The algorithm uses particle swarm optimization to solve for the pricing coefficient, carbon quota purchase volume, and green certificate holding volume, so as to maximize the overall benefits of power generation companies.
[0279] The optimal strategy for the electricity user's intelligent agent is:
[0280]
[0281] Specifically, the distribution of equilibrium strategies Substituting the revenue function of the electricity user agent into the policy space... The system uses a genetic algorithm to solve for the time-of-use load curve, green certificate holdings, and green electricity consumption ratio, thereby maximizing user utility and additional benefits.
[0282] The optimal strategy for a renewable energy enterprise intelligent agent is:
[0283]
[0284] Specifically, the distribution of equilibrium strategies Substituting the revenue function of the renewable energy enterprise intelligent agent into the policy space... The system employs dynamic programming to solve the time-of-use power output plan and the amount of green certificate issuance, thereby maximizing the revenue from green electricity and the revenue from green certificates.
[0285] The optimal strategy for the regulatory agency's intelligent agent is:
[0286]
[0287] Specifically, the distribution of equilibrium strategies Substituting the revenue function of the regulatory agency's intelligent agent into the policy space... The system employs a third-generation non-dominated sorting genetic algorithm to solve for carbon quota allocation, green electricity consumption targets, and bid anti-collusion thresholds, thereby achieving coordinated optimization of multiple objectives, including electricity price stability, carbon emission reduction, green electricity consumption, and social welfare.
[0288] S5.3 Multi-Market Collaborative Operation Strategy Construction:
[0289] Based on the executable optimal strategies of four types of intelligent agents, a multi-market collaborative operation strategy for electricity, carbon, and green certificates is constructed.
[0290] First, calculate the carbon price equilibrium value based on the carbon-green certificate price linkage sensitivity coefficient:
[0291]
[0292] In the formula, This is the equilibrium value linked to carbon prices; This represents the current average transaction price in the green certificate market. The carbon-green certificate price linkage sensitivity coefficient obtained from S1;
[0293] Then, calculate the signed deviation between the actual carbon price and the carbon price equilibrium value:
[0294]
[0295] The trigger threshold is dynamically adjusted based on the benchmark arbitrage threshold and the level of electricity price fluctuations.
[0296]
[0297] In the formula, Use the benchmark arbitrage threshold; This represents the variance of electricity prices; the greater the fluctuation in electricity prices, the higher the trigger threshold, and the smaller the fluctuation in electricity prices, the lower the trigger threshold, in order to reduce frequent adjustments during periods of high market volatility.
[0298] when When the actual carbon price is higher than the linked equilibrium level and the deviation exceeds the threshold, the regulatory agency increases the supply of carbon quotas and guides renewable energy companies to increase the scale of green certificate issuance, so as to promote the return of carbon prices to the equilibrium level.
[0299] when When the actual carbon price is lower than the equilibrium level and the deviation exceeds the threshold, the regulatory agency will reduce the supply of carbon allowances, guide power generation companies to increase the scale of carbon allowance procurement, and power users to increase the scale of green certificate holdings, so as to promote the return of carbon prices to the equilibrium level.
[0300] S5.4 Fully Constrained Feasibility Verification:
[0301] By reusing the multi-constraint coupled equation set corresponding to the previous feasible region in S2, a full-constraint feasibility check is performed on all executable optimal strategies to ensure that the final output strategy simultaneously meets all constraints of electricity supply and demand, total carbon emissions, renewable energy consumption, green certificate supply and demand balance, and anti-collusion of price differences.
[0302] In a multi-market collaborative operation scenario, the following set of collaborative constraint equations is used for verification:
[0303]
[0304]
[0305]
[0306]
[0307] In the formula, the one with " The variables in "" are all executable optimal policy parameters obtained from S5.2;
[0308] S5.5 Short-term deviation correction mechanism:
[0309] Real-time market operation data is collected on a 24-hour cycle, including real-time electricity price, real-time carbon price, real-time green certificate price, actual strategy execution data of each intelligent agent, total carbon emissions of the system and green electricity consumption ratio; after normalizing the multi-dimensional state indicators, the comprehensive deviation between the actual operating state and the equilibrium state is calculated.
[0310] Assume the actual operating state is The equilibrium state is The deviation is:
[0311]
[0312] Specifically, it can be expressed as:
[0313]
[0314] In the formula, , , These are the actual electricity price, the actual carbon price, and the actual green certificate price, respectively. , , These are the electricity price, carbon price, and green certificate price under equilibrium conditions, respectively.
[0315] The deviation judgment threshold is dynamically set based on the intensity of market volatility.
[0316]
[0317]
[0318] In the formula, The baseline deviation threshold; This is the market volatility coefficient; , , These are the normalized volatility of electricity price, carbon price, and green certificate price, respectively; the stronger the market volatility, the higher the sensitivity of the feedback identification.
[0319] when When, maintain the current optimal executable strategy and continue monitoring; when When the current actual operating state is used as the new initial state, the game dynamic equation of S3 is returned to iterate again, and the equilibrium strategy distribution, executable optimal strategy and multi-market collaborative operation strategy are updated in turn.
[0320] S5.6 Long-Term Parameter Iteration Mechanism:
[0321] On a weekly basis, statistical analysis is performed on the cumulative operating data to update the parameters of the revenue model, the sensitivity coefficient of carbon-green certificate price linkage, the constraint threshold, and the feasible domain. The revenue model parameters include the generation cost coefficient, the electricity consumption efficiency coefficient, the green electricity scarcity premium coefficient, and the multi-objective preference weight, etc. The constraint parameters include the total carbon emission threshold, the green electricity consumption target, the green certificate supply and demand balance threshold, and the bidding anti-collusion threshold, etc.
[0322] After the parameters are updated, the entire chain is re-solved starting from the game dynamic equation stage of S3, so that the model and strategy can continuously adapt to long-term changes in market structure, policy rules and operating status.
[0323] Through the above steps, this embodiment unifies the heterogeneous payoff functions of four types of intelligent agents, the dynamic shrinking mechanism of the pre-feasible region, the dynamic equation of the evolutionary game with Lagrange constraint correction, the ESS-MAPPO equilibrium solution mechanism, the optimal strategy output mechanism of the sub-agents, and the dual-cycle feedback correction mechanism into the same technical process, so as to realize the multi-agent strategy collaborative optimization of the electricity market, carbon market, and green certificate market.
[0324] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. A market evolutionary game method considering the payoff function of electricity-carbon-green certificates, characterized in that, Includes the following steps: S1. Construct a multi-agent heterogeneous revenue model: Divide the market participants in the electricity-carbon-green certificate coupling into power generation enterprise agents, electricity user agents, renewable energy enterprise agents, and regulatory agency agents; construct heterogeneous revenue functions according to the market roles and decision-making objectives of each type of agent; S2, Define the agent policy space and construct the pre-feasible domain constraints: Based on the heterogeneous benefit functions and decision objectives of various agents in S1, determine the policy variables of the four types of agents and form the corresponding policy spaces; A set of multi-constraint coupled equations is established around the strategy space, including at least hard constraints on power supply and demand, total carbon emission constraints, renewable energy consumption constraints, green certificate supply and demand balance constraints, and price difference anti-collusion constraints. Constraint relaxation is set and a dynamic contraction mechanism is used to adjust the feasible region range so that the candidate strategy satisfies the pre-feasible region constraints before entering the strategy evolution. S3: Construct a multi-agent evolutionary game dynamic equation with Lagrange constraint correction: Calculate the policy payoffs of various agents based on the heterogeneous payoff function in S1, and form a policy distribution based on the policy space in S2; introduce a heterogeneous payoff weight matrix, and introduce the multi-constraint coupling equation set in S2 into the Lagrange function; add a Lagrange dual constraint correction term to the replicated dynamic equation to form a reconstructed dynamic equation, so that the policy distribution evolves in the prior feasible region according to payoff-driven and constraint correction, and adaptively updates the Lagrange dual variables; S4, Solving the equilibrium policy distribution by integrating ESS and MAPPO algorithms: Based on the policy evolution relationship represented by the reconstructed dynamic equation in S3, an ESS prior model is constructed and a stable equilibrium prior policy distribution is output. The stable equilibrium prior policy distribution is then weighted and fused with the policy distribution output by the MAPPO policy network. A multi-market coupled attention mechanism is introduced into the MAPPO value network, and an equilibrium stability penalty term is added to the policy optimization objective. After iterative training, the equilibrium policy distribution of four types of agents is output. S5, Output the optimal strategy and establish a feedback loop: Based on the equilibrium strategy distribution output in S4, determine the strategy optimization search range and initialization conditions, and according to the strategy variable types and optimization objectives of the four types of intelligent agents, respectively execute the bidding and trading strategy optimization of the power generation enterprise intelligent agent, the energy consumption and green electricity consumption strategy optimization of the electricity user intelligent agent, the time-of-use power output strategy optimization of the renewable energy enterprise intelligent agent, and the multi-objective regulation strategy optimization of the regulatory agency intelligent agent, to obtain the executable optimal strategy for the four types of intelligent agents; based on the executable optimal strategy, construct a multi-market collaborative operation strategy for electricity-carbon-green certificates, and reuse the multi-constraint coupled equation system in S2 for full-constraint feasibility verification; establish a feedback loop, where short-term feedback updates the strategy based on real-time operating status, and long-term feedback updates the revenue model parameters and constraint parameters based on cumulative operating data.
2. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function as described in claim 1, characterized in that, The S1 step of constructing a multi-agent heterogeneous benefit model includes the following steps: S1.1, Classification of Intelligent Agents in the Coupled Market and Definition of Decision-Making Objectives: The participants in the electricity-carbon-green certificate coupled market are classified into four heterogeneous intelligent agents: power generation enterprise intelligent agents, electricity user intelligent agents, renewable energy enterprise intelligent agents, and regulatory agency intelligent agents; among which: The intelligent agent of the power generation enterprise participates in the electricity market, carbon market and green certificate market trading at the same time, and the decision-making objective is to maximize the comprehensive benefits of the three markets. As the demand-side entity for electricity and green certificates, the intelligent agent of electricity users aims to maximize the comprehensive benefits of electricity consumption efficiency, carbon emission reduction benefits and green electricity consumption rewards. As the main body for green electricity supply and green certificate issuance, the decision-making objective of renewable energy enterprise intelligent agents is to maximize the combined revenue from green electricity generation and green certificate issuance. As the main body for market rule-making and regulation, the regulatory agency's intelligent agent has the decision-making goal of collaboratively achieving multiple public objectives such as stable electricity prices, carbon emission reduction, green electricity consumption, and maximizing social welfare. S1.2, Construction of Heterogeneous Revenue Functions for Four Types of Intelligent Agents: For the decision-making objectives and transaction scope of the four types of intelligent agents, heterogeneous revenue functions with differentiated structures are constructed respectively; where: The revenue function of the power generation enterprise's intelligent agent comprehensively covers electricity sales revenue, power generation costs, carbon quota trading profits and losses, and green certificate holding revenue, and embeds cross-market coupled revenue terms to characterize the impact of the interaction between electricity price, carbon price, and green certificate price on the enterprise's overall revenue. The revenue function of the electricity user's intelligent agent comprehensively covers electricity utility, electricity expenditure, and green certificate holding revenue, and introduces two types of strategic interactive revenue items: implicit carbon emission reduction revenue and excess green electricity consumption reward. The revenue function of the intelligent agent of renewable energy enterprises comprehensively covers the revenue from green electricity sales and the revenue from green certificate issuance, and introduces the interactive revenue item of green electricity scarcity premium strategy; The regulatory agency's intelligent agent revenue function adopts a multi-objective reward and punishment mechanism, setting penalty items for electricity price fluctuations, excessive system carbon emissions, and deviations in green electricity consumption, and setting reward items for social welfare corresponding to the total market revenue.
3. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function as described in claim 2, characterized in that, The cross-market coupling revenue term of the power generation enterprise intelligent agent revenue function in S1.2 includes the carbon-green certificate price linkage sensitivity coefficient, which is calibrated through the following steps: S1.2.1, Data Acquisition and Preprocessing: Collect daily price data for the electricity market, carbon market, and green certificate market covering three complete calendar years; perform stationarity tests on the price series; and perform first-order differencing on non-stationary series to obtain stationary series. S1.2.2, Model Construction: Construct a third-order lagged vector autoregressive model containing three variables: electricity price, carbon price, and green certificate price, to capture the linkage transmission pattern of prices with multiple lags; S1.2.3, Parameter Calculation: Model parameters are obtained through regression estimation. The marginal impact of carbon price and green certificate price on electricity price is extracted separately. The carbon-green certificate price linkage sensitivity coefficient is obtained by combining the coupling effect of the two, which is used to quantify the linkage effect of carbon price and green certificate price on electricity price under the joint action of carbon price and green certificate price.
4. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function as described in claim 1, characterized in that, The S2 definition of the agent policy space includes the following steps: S2.1 Determination of the strategy space of the power generation enterprise's intelligent agent: The strategy space of the power generation enterprise's intelligent agent includes three types of decision variables: bidding coefficient, carbon quota purchase volume, and green certificate holding volume, and correspondingly sets a reasonable bidding range, carbon quota purchase upper limit, and green certificate holding upper limit; S2.2 Determination of the strategy space for the intelligent agent of electricity users: The strategy space for the intelligent agent of electricity users includes three types of decision variables: 24-hour time-of-use load curve, green certificate holding amount, and green electricity consumption ratio, and corresponding daily total electricity consumption constraints and single-period load capacity upper limits are set. S2.3 Determination of the strategy space of the intelligent agent of renewable energy enterprises: The strategy space of the intelligent agent of renewable energy enterprises includes two types of decision variables: time-of-use power output plan and green certificate issuance volume, and corresponding upper limit of installed capacity and tolerance range of power output prediction error are set. S2.4 Determination of the regulatory agency's intelligent agent strategy space: The regulatory agency's intelligent agent strategy space includes three types of decision variables: annual carbon quota allocation, green electricity consumption target, and bid anti-collusion threshold, and corresponding adjustment ranges allowed by policy are set. Among them, the bid collusion prevention threshold is dynamically adjusted with the market concentration. The higher the market concentration, the stricter the bid collusion prevention threshold and the stronger the constraint on bid coordination.
5. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function as described in claim 1, characterized in that, The S2 step of constructing the pre-feasible region constraint includes the following steps: S2.5, Establishment of Multi-Constraint Coupled Equations: A multi-constraint coupled equation system is established around the policy space of four types of intelligent agents, covering hard constraints on power supply and demand, total carbon emission constraints, renewable energy consumption constraints, green certificate supply and demand balance constraints, and anti-collusion constraints on price differences; among which: Hard constraints on power supply and demand are used to ensure the matching relationship between total power output on the generation side and total load on the consumption side, as well as power grid transmission losses; The total carbon emission limit is used to ensure that the actual carbon emissions of power generation companies do not exceed the total amount of carbon allowances they hold and purchase. Renewable energy consumption constraints are used to ensure that the proportion of green electricity consumption in the entire market is not lower than the minimum consumption target requirement; The supply and demand balance constraint for green certificates is used to ensure the quantity matching relationship of green certificates throughout the entire chain of issuance, trading, and holding in the market; The pricing difference anti-collusion constraint is used to determine the risk of collusion by combining market concentration and the pricing differences of power generation companies, and to suppress market manipulation behavior with abnormally consistent or abnormally divergent pricing. S2.6, Implementation of the dynamic shrinkage mechanism of feasible region: Set the constraint relaxation for each type of core constraint, and set a higher relaxation in the initial stage of iteration, corresponding to a wider feasible region range; During the iterative process, a dynamic decay rule is used to adjust the relaxation degree. The feasible region is adaptively narrowed according to the deviation between the actual constraint value and the threshold. The higher the deviation, the more obvious the relaxation degree narrowing. During a single iteration, the feasible region gradually shrinks. In the long-term parameter iteration phase, the range of the feasible region is reset or adjusted based on the updated constraint thresholds and market boundaries. All candidate strategies must fall within the feasible region before they can proceed to the subsequent evolutionary game calculation stage.
6. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function as described in claim 1, characterized in that, The S3 method for constructing the dynamic equations of a multi-agent evolutionary game with Lagrange constraint correction includes the following steps: S3.1, Generation of policy payoff vector and policy distribution vector: Based on heterogeneous payoff functions, calculate the payoff levels corresponding to different policies chosen by various agents to form policy payoff vectors; Based on the policy space, the proportion of each type of agent choosing different policies is represented in the form of a probability distribution, forming a policy distribution vector; S3.2, Heterogeneous Revenue Weight Matrix Calculation: A heterogeneous revenue weight matrix is introduced to quantify the differences in market influence among different intelligent agents. The heterogeneous revenue weight matrix is a diagonal structure, and the off-diagonal elements are zero. The weights of various intelligent agents are calculated by coupling three indicators: market share, information transparency, and strategy adjustment frequency. This comprehensively reflects the agent's market power, information acquisition ability, and strategy response speed, and all weights meet the normalization requirements. S3.3, Reconstructing the Evolutionary Game Dynamic Equation: The multi-constraint coupled equation set is transformed into a Lagrange constraint penalty mechanism, reconstructing the traditional replicative dynamic equation to form a reconstructed dynamic equation that integrates payoff-driven and constraint-guided approaches, allowing the strategy to evolve within the pre-existing feasible domain; where: The payoff-driven mechanism gradually increases the proportion of strategies with payoffs higher than the group average and gradually decreases the proportion of strategies with payoffs lower than the group average. The Lagrange constraint correction mechanism transforms constraints into directional guiding forces for policy evolution, applying corrective effects to policies that are close to the constraint boundary or exhibit a tendency to violate the constraint. S3.4, Lagrange dual variable update: The constraint penalty strength is quantified by the dual variable, and the dual variable is updated with an adaptive step size; When a constraint is violated or approaches the boundary of violation, the corresponding dual variable is simultaneously increased to strengthen the penal effect of the constraint on policy evolution. When the constraints are satisfied stably, the dual variable remains stable or is moderately reduced to avoid excessive constraints that compress the strategy exploration space.
7. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function as described in claim 1, characterized in that, The S4 fusion of ESS and MAPPO algorithms to solve for the equilibrium strategy distribution includes the following steps: S4.1, ESS Prior Model Construction and Initialization: Based on the policy evolution law represented by the reconstructed dynamic equation, an evolutionarily stable policy prior model is constructed. The evolutionarily stable policy prior model adopts a long short-term memory network structure. The input is a historical multi-round policy distribution sequence, and the output is an evolutionarily stable equilibrium prior policy distribution. When running for the first time and without historical convergence results, cold start initialization is completed using historical market conditions and strategy samples, uniform distribution, or industry experience rules. In subsequent operation, the prior model is updated on a rolling basis using the convergence strategy distribution generated from historical iterations; S4.2, Policy Distribution Weighted Fusion Mechanism: The final policy distribution is formed by fusing the original output of the MAPPO policy network with the stable equilibrium prior distribution of ESS; In the early stages of training, the prior guidance weight of ESS is relatively high, prioritizing the strategy to explore the historical stable equilibrium region. During the training process, the weight ratio is gradually adjusted to strengthen the dominant role of MAPPO autonomous learning and achieve a balance between stable prior guidance and reinforcement learning autonomous optimization. The ESS prior network is trained by minimizing the distribution difference, so that the prior distribution is consistent with the historical stable equilibrium policy distribution. S4.3, Multi-market Coupled Attention Value Network: Introduce a multi-market coupled attention mechanism into the MAPPO value network, and construct value assessment networks corresponding to the three sub-markets of electricity market, carbon market and green certificate market respectively. Quantify the linkage effect between different sub-markets through attention weights, so that the global value estimate reflects the comprehensive benefit level under the coupling effect of the three markets. S4.4, Optimization objective with equilibrium stability penalty term: The policy optimization objective introduces an equilibrium stability penalty mechanism on the basis of the clipping optimization framework, which imposes constraints on policy iteration that deviates from the stable equilibrium state and suppresses the oscillation phenomenon in the policy iteration process; The advantage function is calculated using the generalized advantage estimation method to balance the evaluation weights of short-term and long-term returns.
8. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function according to claim 1, characterized in that, The S4 fusion of ESS and MAPPO algorithms for solving the equilibrium strategy distribution also includes the following steps: S4.5 Iterative Training Initialization: Set the network parameters of the policy network, value network, and ESS prior model; initialize the dual variables and heterogeneous reward weight matrix; and set the maximum number of iterations and the number of iterations per round. S4.6, State Acquisition: Each agent makes decisions based on the current policy distribution and market state, collects state, policy, profit, and next state samples, and stores them in the trajectory sampling cache pool used to store the current round's sampling trajectory. S4.7, Network Update: Extract a small batch of samples from the trajectory sampling buffer pool and update the parameters of the ESS prior model, policy network and value network synchronously. S4.8, Convergence Determination: The degree of change in policy distribution between adjacent iterations is measured by multi-dimensional comprehensive distance. When the change in policy distribution in multiple consecutive iterations is not greater than the preset convergence threshold, the algorithm is determined to have reached evolutionary stable equilibrium and outputs the balanced policy distribution of the four types of agents.
9. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function as described in claim 1, characterized in that, The optimal strategy for S5 output includes the following steps: S5.1, Search Prior Setting Guided by Balanced Distribution: Using balanced strategy distribution as the search prior, determine the search range and initialization conditions for strategy optimization; The high-probability policy interval corresponding to the equilibrium policy distribution is used as the priority search neighborhood of the optimization algorithm, and the central tendency value of the distribution is used as the initial solution or initial population center of the optimization algorithm, so that the final optimal policy falls within the evolutionary stable region. S5.2, Differentiated Optimization Solution for Different Agents: For the decision-making scenarios of the four types of agents, appropriate optimization algorithms are matched to solve for the executable optimal policy within their own policy space; where: For the purpose of optimizing pricing and trading strategies, the intelligent agent of the power generation company uses the particle swarm optimization algorithm to solve the executable optimal strategy that maximizes overall benefits. For the optimization objectives of energy consumption and green electricity consumption strategies, the intelligent agent of electricity users uses a genetic algorithm to solve the executable optimal strategy that maximizes utility and additional benefits. For the optimization objective of time-of-use power output strategy, the intelligent agent of renewable energy enterprises uses dynamic programming algorithm to solve the executable optimal strategy for maximizing the benefits of green electricity and green certificates; The regulatory agency's intelligent agent optimizes the objectives of multi-objective regulation strategies by employing a third-generation non-dominated sorting genetic algorithm to solve for the executable regulation strategy that achieves the best synergy among multiple common objectives.
10. The electricity-carbon-green certificate coupled market evolution game method considering the payoff function according to claim 1, characterized in that, The multi-market collaborative operation strategy, full-constraint feasibility verification, and feedback closed loop in S5 include the following steps: S5.3, Construction of Multi-Market Collaborative Operation Strategy: Based on the executable optimal strategy of four types of intelligent agents, a multi-market collaborative operation strategy of electricity-carbon-green certificates is constructed; The carbon price linkage equilibrium level is determined based on the carbon-green certificate price linkage sensitivity coefficient and the current average transaction price in the green certificate market. Positive and negative deviations are distinguished by the degree of deviation between the actual carbon price and the equilibrium level. Based on the benchmark arbitrage threshold, the trigger threshold is dynamically adjusted in combination with the level of electricity price fluctuation. The greater the electricity price fluctuation, the higher the trigger threshold, and the smaller the electricity price fluctuation, the lower the trigger threshold. When the actual carbon price is higher than the equilibrium level and the deviation exceeds the threshold, the regulatory authorities will increase the supply of carbon quotas and guide renewable energy companies to increase the scale of green certificate issuance, so as to promote the return of carbon prices to equilibrium. When the actual carbon price is lower than the equilibrium level and the deviation exceeds the threshold, the regulatory authorities will reduce the supply of carbon quotas, guide power generation companies to increase the scale of carbon quota procurement, and power users to increase the scale of green certificate holdings, so as to promote the return of carbon prices to equilibrium. S5.4, Full-constraint feasibility verification: Reuse the multi-constraint coupled equation set corresponding to the previous feasible domain to perform full-constraint feasibility verification on all executable optimal strategies to ensure that the final output strategy simultaneously meets all constraints of power supply and demand, total carbon emissions, renewable energy consumption, green certificate supply and demand balance and anti-collusion of price differences; S5.5, Short-term deviation correction mechanism: Collect real-time market operation data on a 24-hour cycle, normalize multi-dimensional state indicators, and calculate the comprehensive deviation between the actual operation state and the equilibrium state. The deviation judgment threshold is dynamically set according to the intensity of market fluctuations; the stronger the market fluctuations, the higher the sensitivity of the feedback recognition. If the deviation does not exceed the threshold, the current strategy is maintained and monitoring continues. If the deviation exceeds the threshold, the current actual running state is used as the new initial value, and the game dynamic equation is returned to iterate again to update the equilibrium strategy and the optimal execution strategy. S5.6, Long-term parameter iteration mechanism: Statistical analysis is performed on the cumulative running data on a weekly basis to update the profit model parameters, price linkage sensitivity coefficient, constraint threshold and feasible domain range; After the parameters are updated, the game dynamic equations in S3 are returned to begin full-link re-solution, ensuring that the model and strategy continue to adapt to long-term changes in market structure and policy rules.