A virtual power plant distributed bidding method based on MADDPG algorithm

By introducing a tiered carbon trading mechanism and a time-of-use pricing mechanism into virtual power plants, and combining them with the MADDPG algorithm, a carbon-electricity coupling price is constructed and P2P trading is conducted. This solves the problem that traditional bidding strategies are difficult to adapt to the carbon-electricity interaction mechanism, and achieves synergistic optimization of the economic benefits of virtual power plants and carbon emission reduction targets.

CN122335343APending Publication Date: 2026-07-03STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
Filing Date
2026-04-09
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Traditional virtual power plant bidding strategies driven by a single electricity price are ill-suited to the complexities of the electricity-carbon interaction mechanism. Especially with the increasing proportion of distributed energy, they fail to effectively consider changes in renewable energy output and carbon trading costs, resulting in insufficient economic benefits and low-carbon operation levels.

Method used

A distributed bidding method for virtual power plants based on the MADDPG algorithm is adopted, which combines a tiered carbon trading mechanism and a time-of-use pricing mechanism to construct an electricity-carbon coupled price. Resource allocation is optimized through P2P trading, and autonomous collaborative decision-making of multiple virtual power plant alliances is realized using the MADDPG algorithm.

Benefits of technology

It enhances the synergistic optimization of the economic benefits of virtual power plants in the electricity-carbon coupling market with carbon emission reduction targets, realizes the precise quantification of carbon emission costs and the flexibility of resource allocation, and strengthens the bidding adaptability and comprehensive benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335343A_ABST
    Figure CN122335343A_ABST
Patent Text Reader

Abstract

This invention discloses a distributed bidding method for virtual power plants based on the MADDPG algorithm, belonging to the field of electricity market trading and intelligent optimization technology. The method includes the following steps: Step 1) Virtual power plant system modeling and environment construction: constructing a virtual power plant resource characteristic model; Step 2) Distributed bidding model construction: considering P2P transactions between multiple virtual power plants, constructing a distributed bidding model for the electricity-carbon coupling market; Step 3) Reinforcement learning-based model solving: addressing the multi-agent and strongly coupled characteristics of the proposed distributed bidding model, a solution method based on the MADDPG algorithm is used to solve the model. This invention can achieve efficient collaborative decision-making in bidding problems with strong coupling among multiple agents, high dimensionality, and obvious continuous decision-making characteristics, improving the economic benefits and strategy adaptability of virtual power plants, and has good scalability and engineering application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual power plants, specifically a distributed bidding method for virtual power plants based on the MADDPG algorithm. Background Technology

[0002] The market mechanism is gradually evolving towards a dual-driven model of "electricity price and carbon price." Carbon emission costs are no longer a post-hoc accounting factor, but rather a crucial component directly affecting marginal resource costs, power dispatch decisions, and bidding strategies. Especially against the backdrop of the continuous increase in the proportion of distributed energy, virtual power plants need to simultaneously consider factors such as renewable energy output and changes in carbon trading costs, making traditional single-price-driven bidding strategies difficult to adapt to the complexity of the electricity-carbon interaction mechanism. Summary of the Invention

[0003] This invention addresses the shortcomings of existing technologies by proposing a distributed bidding method for virtual power plants based on the MADDPG algorithm, aiming to improve the economic benefits and low-carbon operation of virtual power plants. This invention addresses the trend of deep coupling between the electricity market and the carbon market, researching bidding strategies for virtual power plants in an electricity-carbon coupling market. First, a tiered carbon trading mechanism is introduced, combined with carbon emission flow theory and time-of-use pricing mechanisms, to obtain an electricity-carbon coupling price reflecting carbon costs, time-of-use characteristics, and the impact of node carbon potential. Then, based on this, a multi-virtual power plant alliance bidding model for the electricity-carbon coupling market is constructed. Finally, addressing the problems of the bidding model's complexity, high dimensionality, and obvious continuous decision-making characteristics, a solution method based on the MADDPG algorithm is adopted to achieve autonomous and collaborative decision-making of resources within the virtual power plant.

[0004] This invention is achieved through the following technical solution: This invention discloses a distributed bidding method for virtual power plants based on the MADDPG algorithm, comprising the following steps: Step 1) Virtual power plant system modeling and environment construction: Construct a virtual power plant resource characteristic model; Step 2) Construction of Distributed Bidding Model: Considering P2P transactions between multiple virtual power plants, construct a distributed bidding model for virtual power plants oriented towards the electricity-carbon coupling market; Step 3) Solving the model based on reinforcement learning: In view of the multi-agent and strongly coupled characteristics of the proposed distributed bidding model, the model is solved by using the MADDPG algorithm.

[0005] As a further aspect of the present invention, step 1) (1) Distributed photovoltaic: power characteristics: In the formula, for The output power of distributed photovoltaic units at all times; for The actual solar irradiance in the working environment of the distributed photovoltaic unit at any given time; This refers to the rated power of the distributed photovoltaic (PV) unit under standard operating environmental conditions. The temperature coefficient of a photovoltaic panel is typically taken as 1. ; for The operating ambient temperature of the distributed photovoltaic unit at all times.

[0006] Economic characteristics: In the formula, The deviation penalty cost for distributed photovoltaic units; This refers to the deviation penalty coefficient for distributed photovoltaic (PV) units. for The predicted power output of distributed photovoltaic (PV) units recently declared.

[0007] Carbon market trading characteristics In the formula, The number of CCER certifications that can be obtained for distributed photovoltaic power generation; CCER certification coefficient for renewable energy generator sets: (2) Distributed wind power: power characteristics: In the formula, for The output power of distributed wind turbine units at all times; This refers to the rated power of the distributed wind turbine generator set. for The wind speed in the working environment of the distributed wind turbine at all times; For distributed wind turbine units; The cut-out wind speed for distributed wind turbines; The wind speed in the working environment corresponding to the output rated power of a distributed wind turbine.

[0008] Cost characteristics: In the formula, The deviation penalty cost for distributed wind turbine units; This refers to the deviation penalty coefficient for distributed wind turbine units. for The predicted power output of distributed wind turbines was recently declared.

[0009] Carbon market trading characteristics: In the formula, This refers to the number of CCER certifications that can be obtained for distributed wind power.

[0010] (3) Energy storage: Power characteristics: In the formula, for State of charge (SOC) that stores energy at all times; The energy loss coefficient for energy storage; , They are respectively The charging and discharging power of the stored energy at any given time; The charging and discharging efficiency of energy storage; This refers to the capacity of the energy storage device.

[0011] Cost characteristics: In the formula, The operating cost of energy storage; This is the energy storage operating cost coefficient; This is the energy storage charge loss cost coefficient.

[0012] (4) Gas turbine: Power characteristics: In the formula, for The output power of the gas turbine at any given time; and These are the coefficients of the linear term and the constant term of the piecewise linear function, respectively. for The amount of gas entering the gas turbine at any given time; The segment number is the segment number of the piecewise linear function; Let be the set of pieces of a piecewise linear function.

[0013] Cost characteristics: In the formula, The operating cost of the gas turbine; , and These are the quadratic coefficient, linear coefficient, and constant term of the gas turbine cost characteristic function, respectively.

[0014] Carbon market trading characteristics: In the formula, The amount of free carbon allowance that can be allocated to a gas turbine unit; The allocation factor for allocating free carbon allowances using the baseline method.

[0015] (5) Building air conditioning: Power characteristics: In the formula, for The indoor temperature of the building's air conditioning system at all times; for The outdoor temperature of the working environment of the building's air conditioning unit at all times; Building air conditioning first stage Equivalent heat capacity in the model; Building air conditioning level 1 Equivalent thermal resistance in the model; for The cooling or heating capacity of the building's air conditioning at all times; for The cooling or heating capacity of the building's air conditioning system at all times; and These are the coefficients of the first-order term and the constant term, respectively, representing the linear relationship between the cooling or heating capacity of a building air conditioner and its cooling or heating power.

[0016] Cost characteristics: In the formula, Economic compensation for the reduction of building air conditioning load; Compensation for reduced load on building air conditioning units; for Reduce the load on building air conditioning at all times.

[0017] As a further aspect of the present invention, step 2) (1) Tiered carbon trading The core idea of ​​the tiered carbon trading mechanism is to divide the trading volume of virtual power plants in the carbon market into multiple tiers based on carbon emission allowances, with different tiers corresponding to different carbon price levels. When the actual carbon trading volume of a virtual power plant falls into a lower tier, the enterprise only needs to bear a lower carbon cost; while when the carbon trading volume exceeds a certain threshold, it automatically enters a higher tier, and the corresponding carbon price increases accordingly. Through this mechanism, the system can guide resources to prioritize the consumption of new energy sources, reduce the operating time of high-emission resources, and curb excessive emission behavior while maintaining total carbon emission constraints. The specific expression is shown in the following formula.

[0018] In the formula, For tiered carbon trading costs; The benchmark carbon trading price; This refers to the carbon trading volume of virtual power plants in the carbon market. The growth rate of carbon trading prices; The range length for carbon trading volume.

[0019] (2) Price of electrocarbon coupling Based on nodal carbon potential and traditional carbon trading prices, this section proposes a dynamic carbon price coefficient, which can reflect the corresponding marginal carbon emission cost according to the size of the nodal carbon potential, thereby reflecting the actual carbon emission pressure corresponding to the use of electricity under different carbon emission levels. The calculation method of the dynamic carbon price coefficient is shown in the equation.

[0020] In the formula, for The dynamic carbon price coefficient at any given time; for The carbon potential at a given moment; This is the carbon potential threshold, used to classify the level of carbon emissions. The carbon price increase coefficient under high carbon emission levels; This represents the carbon price reduction coefficient under low carbon emission levels.

[0021] Secondly, based on the tiered carbon trading mechanism, the average carbon price of virtual power plants can be calculated after considering carbon quotas, free allocation, excess emissions, and CCER offsetting. Because the tiered carbon trading mechanism has a distinct segmented nature, the transaction costs exhibit different marginal curves when emissions fall within different segments. By weighting the carbon costs of virtual power plants at each stage, the average carbon price at the virtual power plant level can be obtained, thus reflecting its overall economic pressure from carbon emissions. The calculation method for the average carbon price is shown in the following formula.

[0022] In the formula, The average carbon price of a virtual power plant.

[0023] Then, by combining the aforementioned dynamic carbon price coefficient and average carbon price, a dynamic carbon price for virtual power plants can be constructed. This price reflects both the marginal incremental cost of carbon emissions from virtual power plants and retains the local carbon emission responsibility allocation characteristics. The specific calculation method for the dynamic carbon price is shown in the following formula.

[0024] In the formula, The dynamic carbon price for virtual power plants.

[0025] Finally, the dynamic carbon price is coupled with the time-of-use electricity price to construct an electricity-carbon coupled price, as shown in the following equation.

[0026] In the formula, The price of carbon coupling for a virtual power plant.

[0027] (3) Virtual power plant distributed bidding model considering peer-to-peer transactions Under the P2P mechanism, virtual power plants are not only market participants but also potential bilateral trading entities. Each virtual power plant can autonomously decide to sell or buy electricity from other virtual power plants based on its own external characteristics, marginal costs, carbon emission levels, and electricity-carbon coupling prices, in order to maximize economic benefits and minimize carbon costs. When a virtual power plant has surplus low-carbon electricity, it can sell electricity to neighboring virtual power plants; conversely, when a virtual power plant faces rising marginal costs due to high load, high carbon prices, or emission restrictions at a certain time, it can purchase external low-carbon electricity through the P2P market, thereby improving its competitive position. In this process, the electricity-carbon coupling price provides a unified value measure for P2P transactions, making the trading process not only economically driven but also involving the flow and redistribution of carbon emission responsibilities. Each virtual power plant makes independent decisions in a distributed market environment, forming trading prices and trading volumes through mutual game theory, ultimately achieving synergistic optimization of the overall economic benefits and carbon emission reduction targets of the system.

[0028] Based on the above ideas, a distributed bidding model with multiple virtual power plants is constructed, incorporating both electricity-carbon coupling prices and P2P bilateral transactions into the bidding model. The specific bidding model is shown in the following equation.

[0029] In the formula, For virtual power plants Consider the net revenue from P2P transactions in the electric carbon coupling market; for Virtual Power Plant The price of electro-carbon coupling; for Virtual Power Plant The winning bid volume; For virtual power plants The total revenue from P2P transactions between the company and other virtual power plants; For virtual power plants Deviation penalty cost of distributed photovoltaic units; For virtual power plants Deviation penalty cost of distributed wind turbine units; For virtual power plants The operating cost of the gas turbine; For virtual power plants The economic compensation cost for the reduction of building air conditioning load; For virtual power plants The operating cost of energy storage.

[0030] When virtual power plants bid in the electricity-carbon coupling market, they also need to meet the following constraints: 1) Power constraints for virtual power plants participating in the electricity-carbon coupling market To ensure that the virtual power plant can safely and stably execute its output according to market instructions after winning the bid, the internal operation of the virtual power plant needs to meet the power constraints of the bid, as shown in the following formula.

[0031] In the formula, for Virtual Power Plant The winning bid volume for distributed photovoltaic (PV) units in China; for Virtual Power Plant The winning bid volume of distributed wind turbine units in China; for Virtual Power Plant The winning bid volume of China Gas Turbine; for Virtual Power Plant China Energy Storage's winning bid volume; for Virtual Power Plant The amount of electricity traded through P2P platforms.

[0032] 2) Constraints on the amount of electricity used by virtual power plants in P2P transactions To ensure the safe and stable operation of the virtual power plant itself, the virtual power plant must meet the P2P transaction electricity constraints when participating in P2P transactions, as shown in the following formula.

[0033] In the formula, for The maximum amount of electricity traded between virtual power plants in a P2P transaction at any given time.

[0034] 3) Balance constraints of electricity volume for virtual power plants participating in P2P transactions In order to maintain the conservation of energy flow within the regional market and the feasibility and fairness of P2P transactions, virtual power plants participating in P2P transactions need to meet the power balance constraint, as shown in the following formula.

[0035] In the formula, This is a collection of all virtual power plants participating in P2P transactions.

[0036] After each virtual power plant submits its bidding strategy, the electricity-carbon coupling market needs to determine the final winning power of each virtual power plant through a unified clearing mechanism. The electricity-carbon coupling market clearing model aims to maximize social welfare, i.e., minimize the overall market energy purchase cost, and its specific expression is shown in the following formula.

[0037] In the formula, This represents the overall market energy purchase cost.

[0038] To ensure the feasibility of the clearing results, the electric carbon coupling market must meet the output constraints and ramp-up constraints of each virtual power plant during the clearing process, as shown in equations 1 and 2.

[0039] In the formula, for Virtual Power Plant Maximum output; for Virtual Power Plant Minimum output; for Virtual Power Plant Maximum downhill climbing rate; for Virtual Power Plant The maximum uphill climbing rate.

[0040] As a further aspect of the present invention, step 3) (1) MADDPG algorithm principle The MADDPG algorithm is an extension of the Deep Deterministic Policy Gradient (DDPG) algorithm, meaning it involves multiple agents in the environment, each implementing the DDPG algorithm. Similar to the DDPG framework, MADDPG also employs an actor-critic network architecture. For each agent, its configured actor-critic network architecture is independent and trained separately. Each actor-critic network architecture includes an online actor network. A single Actor target network A Critic online network A Critic target network That is, for those containing The MADDPG algorithm has a total of [number] agents. There are two networks. The Actor network is responsible for generating continuous actions from local states; that is, the goal of the Actor network is to maximize the state-action value function (Q-value) output by the Critic network, and iteratively updates it using deterministic policy gradients. The Critic network, on the other hand, estimates the Q-value of the joint actions based on complete information during training. That is, the Critic network receives the actions of all agents and the global state as input, thereby calculating a more accurate gradient signal, which in turn guides the Actor network's updates.

[0041] The core idea of ​​the MADDPG algorithm is centralized training and distributed execution. Specifically, during the training phase, i.e., the forward propagation phase of computing the Critic network, each agent's Critic network can concatenate the global state information observations and the actions of all agents into an observation vector. and action vectors and will The Q-value is then used as input to the Critic online network and output as output. This allows the agent to complete the training of its own Critic network; however, during the execution phase, i.e., the forward propagation phase of computing the Actor network, the agent relies solely on its own local observation vectors. As input to the Actor online network, the output action This enables distributed execution.

[0042] The specific training process of the MADDPG algorithm is as follows: 1) Environment and agent initialization Initialize the Actor online network for each agent. and target network and Critic online network and target network Initialize the experience replay pool Initialize the environment .

[0043] 2) The agent interacts with the environment and collects experience. Each agent Obtain local observations of itself Then, noise is added based on the current Actor's output action. As shown in the following formula.

[0044] Then the combined action of all agents The joint observation is applied to the environment, and then the environment returns to the next moment. and the rewards for each agent .

[0045] 3) Storing multi-agent joint experience Add the current interaction sample to the experience replay pool. In the formula shown below.

[0046] 4) Sample batch data from the playback pool From the experience replay pool Randomly sample a batch of empirical samples .

[0047] 5) Centralized training of the Critic network For intelligent agents Joint observation and joint action of all intelligent agents Input the Critic network and output the agent's Q-value under the joint state-action condition.

[0048] The update method for the Critic network is shown in the following equation.

[0049] 6) Update the Actor network policy for each agent. The updated Critic network guides the update of the Actor network policy; that is, each agent's Actor network policy is updated by maximizing the expected value of its Critic network function. The specific update method is shown in the following equation.

[0050] 7) Soft update target network To enhance training stability, a soft update method is used to update the target network parameters. The specific update method is shown in the following equation.

[0051] 8) Iterate repeatedly until convergence or the maximum number of training iterations is reached.

[0052] (2) Model solving method based on MADDPG algorithm In the distributed bidding model for virtual power plants oriented towards the electricity-carbon coupling market, each virtual power plant represents an agent, the electricity-carbon coupling market represents the algorithm environment, the bidding strategy of the virtual power plant represents the agent's action, the market clearing result represents the state of the environment, and the revenue of the virtual power plant represents the agent's reward. The correspondence between the MADDPG algorithm and the distributed bidding model for virtual power plants proposed in this invention is shown in the table below.

[0053] The specific process of solving the model in this chapter based on the MADDPG algorithm, i.e., the pseudocode for solving the model using the MADDPG algorithm, is shown in the table below.

[0054] Compared to existing technologies, the beneficial effects of this invention are as follows: This invention achieves precise quantification of carbon emission costs through an electricity-carbon coupling pricing mechanism, and enhances resource allocation flexibility by combining P2P trading; the application of the MADDPG algorithm effectively solves the complex decision-making problem of multi-entity bidding, achieving synergistic optimization of economic benefits and carbon emission reduction targets. This method possesses good scalability and engineering application value, and can significantly improve the bidding adaptability and overall benefits of virtual power plants in the electricity-carbon coupling market.

[0055] This invention addresses the trend of deep coupling between the electricity market and the carbon market, researching bidding strategies for virtual power plants in an electricity-carbon coupling market. First, a tiered carbon trading mechanism is introduced, combined with carbon emission flow theory and time-of-use pricing mechanisms, to obtain an electricity-carbon coupling price that reflects carbon costs, time-of-use characteristics, and the impact of nodal carbon potential. Then, based on this, a multi-virtual power plant alliance bidding model is constructed for the electricity-carbon coupling market. Finally, addressing the problems of the bidding model's complexity, high dimensionality, and significant continuous decision-making characteristics, a solution method based on the MADDPG algorithm is adopted to achieve autonomous and collaborative decision-making of resources within the virtual power plant. Attached Figure Description

[0056] Figure 1 This is a flowchart of an implementation example; Figure 2 A framework diagram for distributed bidding of virtual power plants targeting the electricity-carbon coupling market; Figure 3 This is a schematic diagram of tiered carbon trading. Figure 4 A price chart for electro-carbon coupling; Figure 5 This is a framework diagram of the MADDPG algorithm; Figure 6 For the total reward training curve; Figure 7 Training curves for rewards for each virtual power plant; Figure 8 The winning bid volume for each virtual power plant; Figure 9 The bidding results for the internal resources of Virtual Power Plant 1; Figure 10 The bidding results for internal resources of Virtual Power Plant 2; Figure 11 The bidding results for internal resources of Virtual Power Plant 3; Detailed Implementation This embodiment provides a distributed bidding method for virtual power plants based on the MADDPG algorithm, including the following steps: Step 1): Virtual power plant system modeling and environment construction.

[0057] (1) Distributed photovoltaic: power characteristics: In the formula, for The output power of distributed photovoltaic units at all times; for The actual solar irradiance in the working environment of the distributed photovoltaic unit at any given time; This refers to the rated power of the distributed photovoltaic (PV) unit under standard operating environmental conditions. The temperature coefficient of a photovoltaic panel is typically taken as 1. ; for The operating ambient temperature of the distributed photovoltaic unit at all times.

[0058] Economic characteristics: In the formula, The deviation penalty cost for distributed photovoltaic units; This refers to the deviation penalty coefficient for distributed photovoltaic (PV) units. for The predicted power output of distributed photovoltaic (PV) units recently declared.

[0059] Carbon market trading characteristics In the formula, The number of CCER certifications that can be obtained for distributed photovoltaic power generation; CCER certification coefficient for renewable energy generator sets: (2) Distributed wind power: power characteristics: In the formula, for The output power of distributed wind turbine units at all times; This refers to the rated power of the distributed wind turbine generator set. for The wind speed in the working environment of the distributed wind turbine at all times; For distributed wind turbine units; The cut-out wind speed for distributed wind turbines; The wind speed in the working environment corresponding to the output rated power of a distributed wind turbine.

[0060] Cost characteristics: In the formula, The deviation penalty cost for distributed wind turbine units; This refers to the deviation penalty coefficient for distributed wind turbine units. for The predicted power output of distributed wind turbines was recently declared.

[0061] Carbon market trading characteristics: In the formula, This refers to the number of CCER certifications that can be obtained for distributed wind power.

[0062] (3) Energy storage: Power characteristics: In the formula, for State of charge (SOC) that stores energy at all times; The energy loss coefficient for energy storage; , They are respectively The charging and discharging power of the stored energy at any given time; The charging and discharging efficiency of energy storage; This refers to the capacity of the energy storage device.

[0063] Cost characteristics: In the formula, The operating cost of energy storage; This is the energy storage operating cost coefficient; This is the energy storage charge loss cost coefficient.

[0064] (4) Gas turbine: Power characteristics: In the formula, for The output power of the gas turbine at any given time; and These are the coefficients of the linear term and the constant term of the piecewise linear function, respectively. for The amount of gas entering the gas turbine at any given time; The segment number is the segment number of the piecewise linear function; Let be the set of pieces of a piecewise linear function.

[0065] Cost characteristics: In the formula, The operating cost of the gas turbine; , and These are the quadratic coefficient, linear coefficient, and constant term of the gas turbine cost characteristic function, respectively.

[0066] Carbon market trading characteristics: In the formula, The amount of free carbon allowance that can be allocated to a gas turbine unit; The allocation factor for allocating free carbon allowances using the baseline method.

[0067] (5) Building air conditioning: Power characteristics: In the formula, for The indoor temperature of the building's air conditioning system at all times; for The outdoor temperature of the working environment of the building's air conditioning unit at all times; Building air conditioning first stage Equivalent heat capacity in the model; Building air conditioning level 1 Equivalent thermal resistance in the model; for The cooling or heating capacity of the building's air conditioning at all times; for The cooling or heating capacity of the building's air conditioning system at all times; and These are the coefficients of the first-order term and the constant term, respectively, representing the linear relationship between the cooling or heating capacity of a building air conditioner and its cooling or heating power.

[0068] Cost characteristics: In the formula, Economic compensation for the reduction of building air conditioning load; Compensation for reduced load on building air conditioning units; for Reduce the load on building air conditioning at all times.

[0069] Step 2): Construction of the distributed bidding model.

[0070] (1) Tiered carbon trading The core idea of ​​the tiered carbon trading mechanism is to divide the trading volume of virtual power plants in the carbon market into multiple tiers based on carbon emission allowances, with different tiers corresponding to different carbon price levels. When the actual carbon trading volume of a virtual power plant falls into a lower tier, the enterprise only needs to bear a lower carbon cost; while when the carbon trading volume exceeds a certain threshold, it automatically enters a higher tier, and the corresponding carbon price increases accordingly. Through this mechanism, the system can guide resources to prioritize the consumption of new energy sources, reduce the operating time of high-emission resources, and curb excessive emission behavior while maintaining total carbon emission constraints. The specific expression is shown in the following formula.

[0071] In the formula, For tiered carbon trading costs; The benchmark carbon trading price; This refers to the carbon trading volume of virtual power plants in the carbon market. The growth rate of carbon trading prices; The range length for carbon trading volume.

[0072] (2) Price of electrocarbon coupling Based on nodal carbon potential and traditional carbon trading prices, this section proposes a dynamic carbon price coefficient, which can reflect the corresponding marginal carbon emission cost according to the size of the nodal carbon potential, thereby reflecting the actual carbon emission pressure corresponding to the use of electricity under different carbon emission levels. The calculation method of the dynamic carbon price coefficient is shown in the equation.

[0073] In the formula, for The dynamic carbon price coefficient at any given time; for The carbon potential at a given moment; This is the carbon potential threshold, used to classify the level of carbon emissions. The carbon price increase coefficient under high carbon emission levels; This represents the carbon price reduction coefficient under low carbon emission levels.

[0074] Secondly, based on the tiered carbon trading mechanism, the average carbon price of virtual power plants can be calculated after considering carbon quotas, free allocation, excess emissions, and CCER offsetting. Because the tiered carbon trading mechanism has a distinct segmented nature, the transaction costs exhibit different marginal curves when emissions fall within different segments. By weighting the carbon costs of virtual power plants at each stage, the average carbon price at the virtual power plant level can be obtained, thus reflecting its overall economic pressure from carbon emissions. The calculation method for the average carbon price is shown in the following formula.

[0075] In the formula, The average carbon price of a virtual power plant.

[0076] Then, by combining the aforementioned dynamic carbon price coefficient and average carbon price, a dynamic carbon price for virtual power plants can be constructed. This price reflects both the marginal incremental cost of carbon emissions from virtual power plants and retains the local carbon emission responsibility allocation characteristics. The specific calculation method for the dynamic carbon price is shown in the following formula.

[0077] In the formula, The dynamic carbon price for virtual power plants.

[0078] Finally, the dynamic carbon price is coupled with the time-of-use electricity price to construct an electricity-carbon coupled price, as shown in the following equation.

[0079] In the formula, The price of carbon coupling for a virtual power plant.

[0080] (3) Virtual power plant distributed bidding model considering peer-to-peer transactions Under the P2P mechanism, virtual power plants are not only market participants but also potential bilateral trading entities. Each virtual power plant can autonomously decide to sell or buy electricity from other virtual power plants based on its own external characteristics, marginal costs, carbon emission levels, and electricity-carbon coupling prices, in order to maximize economic benefits and minimize carbon costs. When a virtual power plant has surplus low-carbon electricity, it can sell electricity to neighboring virtual power plants; conversely, when a virtual power plant faces rising marginal costs due to high load, high carbon prices, or emission restrictions at a certain time, it can purchase external low-carbon electricity through the P2P market, thereby improving its competitive position. In this process, the electricity-carbon coupling price provides a unified value measure for P2P transactions, making the trading process not only economically driven but also involving the flow and redistribution of carbon emission responsibilities. Each virtual power plant makes independent decisions in a distributed market environment, forming trading prices and trading volumes through mutual game theory, ultimately achieving synergistic optimization of the overall economic benefits and carbon emission reduction targets of the system.

[0081] Based on the above ideas, a distributed bidding model with multiple virtual power plants is constructed, incorporating both electricity-carbon coupling prices and P2P bilateral transactions into the bidding model. The specific bidding model is shown in the following equation.

[0082] In the formula, For virtual power plants Consider the net revenue from P2P transactions in the electric carbon coupling market; for Virtual Power Plant The price of electro-carbon coupling; for Virtual Power Plant The winning bid volume; For virtual power plants The total revenue from P2P transactions between the company and other virtual power plants; For virtual power plants Deviation penalty cost of distributed photovoltaic units; For virtual power plants Deviation penalty cost of distributed wind turbine units; For virtual power plants The operating cost of the gas turbine; For virtual power plants The economic compensation cost for the reduction of building air conditioning load; For virtual power plants The operating cost of energy storage.

[0083] When virtual power plants bid in the electricity-carbon coupling market, they also need to meet the following constraints: 1) Power constraints for virtual power plants participating in the electricity-carbon coupling market To ensure that the virtual power plant can safely and stably execute its output according to market instructions after winning the bid, the internal operation of the virtual power plant needs to meet the power constraints of the bid, as shown in the following formula.

[0084] In the formula, for Virtual Power Plant The winning bid volume for distributed photovoltaic (PV) units in China; for Virtual Power Plant The winning bid volume of distributed wind turbine units in China; for Virtual Power Plant The winning bid volume of China Gas Turbine; for Virtual Power Plant China Energy Storage's winning bid volume; for Virtual Power Plant The amount of electricity traded through P2P platforms.

[0085] 2) Constraints on the amount of electricity used by virtual power plants in P2P transactions To ensure the safe and stable operation of the virtual power plant itself, the virtual power plant must meet the P2P transaction electricity constraints when participating in P2P transactions, as shown in the following formula.

[0086] In the formula, for The maximum amount of electricity traded between virtual power plants in a P2P transaction at any given time.

[0087] 3) Balance constraints of electricity volume for virtual power plants participating in P2P transactions In order to maintain the conservation of energy flow within the regional market and the feasibility and fairness of P2P transactions, virtual power plants participating in P2P transactions need to meet the power balance constraint, as shown in the following formula.

[0088] In the formula, This is a collection of all virtual power plants participating in P2P transactions.

[0089] After each virtual power plant submits its bidding strategy, the electricity-carbon coupling market needs to determine the final winning power of each virtual power plant through a unified clearing mechanism. The electricity-carbon coupling market clearing model aims to maximize social welfare, i.e., minimize the overall market energy purchase cost, and its specific expression is shown in the following formula.

[0090] In the formula, This represents the overall market energy purchase cost.

[0091] To ensure the feasibility of the clearing results, the electric carbon coupling market must meet the output constraints and ramp-up constraints of each virtual power plant during the clearing process, as shown in equations 1 and 2.

[0092] In the formula, for Virtual Power Plant Maximum output; for Virtual Power Plant Minimum output; for Virtual Power Plant Maximum downhill climbing rate; for Virtual Power Plant The maximum uphill climbing rate.

[0093] Step 3): Solve the model based on reinforcement learning.

[0094] (1) MADDPG algorithm principle The MADDPG algorithm is an extension of the Deep Deterministic Policy Gradient (DDPG) algorithm, meaning it involves multiple agents in the environment, each implementing the DDPG algorithm. Similar to the DDPG framework, MADDPG also employs an actor-critic network architecture. For each agent, its configured actor-critic network architecture is independent and trained separately. Each actor-critic network architecture includes an online actor network. A single Actor target network A Critic online network A Critic target network That is, for those containing The MADDPG algorithm has a total of [number] agents. There are two networks. The Actor network is responsible for generating continuous actions from local states; that is, the goal of the Actor network is to maximize the state-action value function (Q-value) output by the Critic network, and iteratively updates it using deterministic policy gradients. The Critic network, on the other hand, estimates the Q-value of the joint actions based on complete information during training. That is, the Critic network receives the actions of all agents and the global state as input, thereby calculating a more accurate gradient signal, which in turn guides the Actor network's updates.

[0095] The core idea of ​​the MADDPG algorithm is centralized training and distributed execution. Specifically, during the training phase, i.e., the forward propagation phase of computing the Critic network, each agent's Critic network can concatenate the global state information observations and the actions of all agents into an observation vector. and action vectors and will The Q-value is then used as input to the Critic online network and output as output. This allows the agent to complete the training of its own Critic network; however, during the execution phase, i.e., the forward propagation phase of computing the Actor network, the agent relies solely on its own local observation vectors. As input to the Actor online network, the output action This enables distributed execution.

[0096] The specific training process of the MADDPG algorithm is as follows: 1) Environment and agent initialization Initialize the Actor online network for each agent. and target network and Critic online network and target network Initialize the experience replay pool Initialize the environment .

[0097] 2) The agent interacts with the environment and collects experience. Each agent Obtain local observations of itself Then, noise is added based on the current Actor's output action. As shown in the following formula.

[0098] Then the combined action of all agents The joint observation is applied to the environment, and then the environment returns to the next moment. and the rewards for each agent .

[0099] 3) Storing multi-agent joint experience Add the current interaction sample to the experience replay pool. In the formula shown below.

[0100] 4) Sample batch data from the playback pool From the experience replay pool Randomly sample a batch of empirical samples .

[0101] 5) Centralized training of the Critic network For intelligent agents Joint observation and joint action of all intelligent agents Input the Critic network and output the agent's Q-value under the joint state-action condition.

[0102] The update method for the Critic network is shown in the following equation.

[0103] 6) Update the Actor network policy for each agent. The updated Critic network guides the update of the Actor network policy; that is, each agent's Actor network policy is updated by maximizing the expected value of its Critic network function. The specific update method is shown in the following equation.

[0104] 7) Soft update target network To enhance training stability, a soft update method is used to update the target network parameters. The specific update method is shown in the following equation.

[0105] 8) Iterate repeatedly until convergence or the maximum number of training iterations is reached.

[0106] (2) Model solving method based on MADDPG algorithm In the distributed bidding model for virtual power plants oriented towards the electricity-carbon coupling market, each virtual power plant represents an agent, the electricity-carbon coupling market represents the algorithm environment, the bidding strategy of the virtual power plant represents the agent's action, the market clearing result represents the state of the environment, and the revenue of the virtual power plant represents the agent's reward. The correspondence between the MADDPG algorithm and the distributed bidding model for virtual power plants proposed in this invention is shown in the table below.

[0107] The specific process of solving the model in this chapter based on the MADDPG algorithm, i.e., the pseudocode for solving the model using the MADDPG algorithm, is shown in the table below.

[0108] To further investigate the impact of electricity-carbon coupling prices and P2P transactions on the bidding of virtual power plants in the electricity-carbon coupling market, the following four scenarios are set up for the example: Scenario 1: Ignoring the price of electricity-carbon coupling and P2P transactions; Scenario 2: Considering the price of electricity-carbon coupling, but not P2P transactions; Scenario 3: Ignoring the price of electricity-carbon coupling, but considering P2P transactions; Scenario 4: Considering the price of electricity-carbon coupling and P2P transactions.

[0109] The model solving method based on the MADDPG algorithm described above was used to solve the above four scenarios respectively, and the virtual power plant revenue results under each scenario are shown in the table below.

[0110] The above-described specific implementations can be partially adjusted by those skilled in the art in different ways without departing from the principles and purpose of the present invention. The scope of protection of the present invention is defined by the claims and is not limited to the above-described specific implementations. All implementation schemes within the scope of the claims are bound by the present invention.

Claims

1. A distributed bidding method for virtual power plants based on the MADDPG algorithm, characterized in that, Includes the following steps: Step 1) Virtual power plant system modeling and environment construction: Construct a virtual power plant resource characteristic model that includes distributed photovoltaic, wind power, energy storage, gas turbine and building air conditioning; Step 2) Construction of distributed bidding model: Based on the tiered carbon trading mechanism and the generation of electricity-carbon coupling price by node carbon potential, combined with multi-virtual power plant P2P trading, a distributed bidding model for the electricity-carbon coupling market is constructed, and the winning bid power, P2P trading power and power balance constraints are set. Step 3) Solving the model based on reinforcement learning: The MADDPG algorithm is used to solve the distributed bidding model.

2. The distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 1, characterized in that, The virtual power plant resource characteristic model in step 1) includes: (1) Distributed photovoltaic: Power characteristics: ; In the formula, for The output power of distributed photovoltaic units at all times; for The actual solar irradiance in the working environment of the distributed photovoltaic unit at any given time; This refers to the rated power of the distributed photovoltaic (PV) unit under standard operating environmental conditions. Let be the temperature coefficient of the photovoltaic panel, and take it as . ; for The operating ambient temperature of the distributed photovoltaic (PV) unit at all times; Economic characteristics: ; In the formula, The deviation penalty cost for distributed photovoltaic units; This refers to the deviation penalty coefficient for distributed photovoltaic (PV) units. for The predicted power output of distributed photovoltaic (PV) units recently declared; Carbon market trading characteristics ; In the formula, The number of CCER certifications that can be obtained for distributed photovoltaic power generation; CCER certification coefficient for renewable energy generator sets: (2) Distributed wind power: Power characteristics: In the formula, for The output power of distributed wind turbine units at all times; This refers to the rated power of the distributed wind turbine generator set. for The wind speed in the working environment of the distributed wind turbine at all times; For distributed wind turbine units; The cut-out wind speed for distributed wind turbines; The wind speed in the working environment corresponding to the output rated power of the distributed wind turbine; Cost characteristics: In the formula, The deviation penalty cost for distributed wind turbine units; This refers to the deviation penalty coefficient for distributed wind turbine units. for The predicted power output of distributed wind turbine units as recently declared; Carbon market trading characteristics: ; In the formula, The number of CCER certifications that can be obtained for distributed wind power; (3) Energy storage: Power characteristics: In the formula, for State of charge (SOC) that stores energy at all times; The energy loss coefficient for energy storage; , They are respectively The charging and discharging power of the stored energy at any given time; The charging and discharging efficiency of energy storage; For energy storage device capacity; Cost characteristics: ; In the formula, The operating cost of energy storage; This is the energy storage operating cost coefficient; This is the energy storage charge loss cost coefficient; (4) Gas turbine: Power characteristics: In the formula, for The output power of the gas turbine at any given time; and These are the coefficients of the linear term and the constant term of the piecewise linear function, respectively. for The amount of gas entering the gas turbine at any given time; The segment number is the segment number of the piecewise linear function; It is the set of pieces of a piecewise linear function; Cost characteristics: In the formula, The operating cost of the gas turbine; , and These are the coefficients of the quadratic term, the coefficients of the linear term, and the constant term of the gas turbine cost characteristic function, respectively. Carbon market trading characteristics: In the formula, The amount of free carbon allowance that can be allocated to a gas turbine unit; Allocation coefficients for allocating free carbon allowances using the baseline method; (5) Building air conditioning: Power characteristics: In the formula, for The indoor temperature of the building's air conditioning system at all times; for The outdoor temperature of the working environment of the building's air conditioning unit at all times; Building air conditioning first stage Equivalent heat capacity in the model; Building air conditioning level 1 Equivalent thermal resistance in the model; for The cooling or heating capacity of the building's air conditioning at all times; for The cooling or heating capacity of the building's air conditioning system at all times; and These are the coefficients of the first term and the constant term of the linear relationship between the cooling or heating capacity of a building air conditioner and its cooling or heating power, respectively. Cost characteristics: In the formula, Economic compensation for the reduction of building air conditioning load; Compensation for reduced load on building air conditioning units; for Reduce the load on building air conditioning at all times.

3. The distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 1, characterized in that, Step 2) includes: (1) Tiered carbon trading Based on carbon emission allowances, the trading volume of virtual power plants in the carbon market is divided into multiple intervals, with different intervals corresponding to different carbon price levels. The specific expression is shown in the following formula: In the formula, For tiered carbon trading costs; The benchmark carbon trading price; This refers to the carbon trading volume of virtual power plants in the carbon market. The growth rate of carbon trading prices; The range length for carbon trading volume; (2) Price of electrocarbon coupling Based on nodal carbon potential and traditional carbon trading prices, a dynamic carbon price coefficient is proposed. This coefficient reflects the marginal carbon emission cost corresponding to the size of the nodal carbon potential, and embodies the actual carbon emission pressure corresponding to the use of electricity at different carbon emission levels. The calculation method of the dynamic carbon price coefficient is shown in the following formula: In the formula, for The dynamic carbon price coefficient at any given time; for The carbon potential at a given moment; This is the carbon potential threshold, used to classify the level of carbon emissions. The carbon price increase coefficient under high carbon emission levels; This represents the carbon price reduction coefficient under low carbon emission levels. Secondly, based on the tiered carbon trading mechanism, the average carbon price of virtual power plants is calculated after considering carbon quotas, free allocation, excess emissions, and CCER offsetting. The calculation method for the average carbon price is shown in the following formula: In the formula, The average carbon price of a virtual power plant; Then, by combining the dynamic carbon price coefficient and the average carbon price, the dynamic carbon price of the virtual power plant is constructed. The specific calculation method of the dynamic carbon price is shown in the following formula. In the formula, The dynamic carbon price for virtual power plants; Finally, the dynamic carbon price is coupled with the time-of-use electricity price to construct the electricity-carbon coupled price, as shown in the following formula: In the formula, The price of electricity-carbon coupling for a virtual power plant; (3) Virtual power plant distributed bidding model considering peer-to-peer transactions A distributed bidding model with multiple virtual power plants is constructed, incorporating both electricity-carbon coupling prices and P2P bilateral transactions into the bidding model. The specific bidding model is shown in the following equation: In the formula, For virtual power plants Consider the net revenue from P2P transactions in the electric carbon coupling market; for Virtual Power Plant The price of electro-carbon coupling; for Virtual Power Plant The winning bid volume; For virtual power plants The total revenue from P2P transactions between the company and other virtual power plants; For virtual power plants Deviation penalty cost of distributed photovoltaic units; For virtual power plants Deviation penalty cost of distributed wind turbine units; For virtual power plants The operating cost of the gas turbine; For virtual power plants The economic compensation cost for the reduction of building air conditioning load; For virtual power plants The operating cost of energy storage.

4. The distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 3, characterized in that, When virtual power plants bid in the electricity-carbon coupling market, they must meet the following constraints: 1) Power constraints for virtual power plants participating in the electricity-carbon coupling market The virtual power plant operates within a system that meets the power constraints stipulated in the bid, as shown in the following formula: In the formula, for Virtual Power Plant The winning bid volume for distributed photovoltaic (PV) units in China; for Virtual Power Plant The winning bid volume of distributed wind turbine units in China; for Virtual Power Plant The winning bid volume of China Gas Turbine Co., Ltd.; for Virtual Power Plant China Energy Storage's winning bid volume; for Virtual Power Plant P2P transaction volume; 2) Constraints on the amount of electricity used by virtual power plants in P2P transactions When a virtual power plant participates in P2P transactions, it must meet the P2P transaction electricity constraints, as shown in the following formula: In the formula, for The maximum transaction volume of P2P transactions between virtual power plants at any given time; 3) Balance constraints of electricity volume for virtual power plants participating in P2P transactions Virtual power plants participating in P2P transactions must meet power balance constraints, as shown in the following formula: In the formula, This is a collection of all virtual power plants participating in P2P transactions.

5. The distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 4, characterized in that, After each virtual power plant submits its bidding strategy, the electricity-carbon coupling market determines the final winning power of each virtual power plant through a unified clearing mechanism. The electricity-carbon coupling market clearing model aims to maximize social welfare, i.e., minimize the overall market energy purchase cost, and the specific expression is shown in the following formula: In the formula, This represents the overall market energy purchase cost.

6. The distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 5, characterized in that, During the clearing process, the electric-carbon coupling market must satisfy the output constraints and ramp-up constraints of each virtual power plant, as shown in the following formula: In the formula, for Virtual Power Plant Maximum output; for Virtual Power Plant Minimum output; for Virtual Power Plant Maximum downhill climbing rate; for Virtual Power Plant The maximum uphill climbing rate.

7. The distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 6, characterized in that, Step 3) involves using the MADDPG algorithm to solve the distributed bidding model, which includes the following steps: The MADDPG algorithm is based on the Deep Deterministic Policy Gradient (DDPG) algorithm, meaning there are multiple agents in the environment, and each agent implements the DDPG algorithm. The MADDPG algorithm employs an Actor-Critic network architecture. For each agent, its configured Actor-Critic network architecture is independent and trained separately. Each Actor-Critic network architecture includes an online Actor network. A single Actor target network A Critic online network A Critic target network That is, for those containing The MADDPG algorithm has a total of [number] agents. The Actor network is responsible for generating continuous actions from local states. That is, the goal of the Actor network is to maximize the state-action value function Q value output by the Critic network. It is iteratively updated through deterministic policy gradients. The function of the Critic network is to estimate the Q value of joint actions based on complete information during training. That is, the Critic network receives the actions of all agents and the global state as input, thereby calculating a more accurate gradient signal and back-guiding the Actor network to update. In the training phase of the MADDPG algorithm, which is the forward propagation phase of computing the Critic network, each agent's Critic network can concatenate the global state information observations and the actions of all agents into an observation vector. and action vectors and will The Q-value is then used as input to the Critic online network and output as output. This allows the agent to complete the training of its own Critic network; and during the execution phase, i.e., the forward propagation phase of computing the Actor network, the agent bases its training on its own local observation vectors. As input to the Actor online network, the output action This enables distributed execution.

8. The distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 7, characterized in that, The specific training process of the MADDPG algorithm is as follows: 1) Environment and agent initialization Initialize the Actor online network for each agent. and target network and Critic online network and target network Initialize the experience replay pool Initialize the environment ; 2) The agent interacts with the environment and collects experience. Each agent Obtain local observations of itself Then, noise is added based on the current Actor's output action. As shown in the following formula: Then the combined action of all agents The joint observation is applied to the environment, and then the environment returns to the next moment. and the rewards for each agent ; 3) Storing multi-agent joint experience Add the current interaction sample to the experience replay pool. In the following formula: 4) Sample batch data from the playback pool From the experience replay pool Randomly sample a batch of empirical samples ; 5) Centralized training of the Critic network For intelligent agents Joint observation and joint action of all intelligent agents Input the Critic network and output the agent's Q-value under the joint state-action condition; The update method for the Critic network is shown in the following equation: 6) Update the Actor network policy for each agent. The updated Critic network guides the update of the Actor network policy. Specifically, each agent's Actor network policy is updated by maximizing the expected value of its Critic network function, as shown in the following equation: 7) Soft update target network The target network parameters are updated using a soft update method, as shown in the following formula: 8) Iterate repeatedly until convergence or the maximum number of training iterations is reached.

9. A distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 8, characterized in that, In the MADDPG algorithm, each component corresponds one-to-one with an element in the virtual power plant distributed bidding model, and the specific correspondence is as follows: Environment: Corresponding to the electricity-carbon coupling market, it represents the external system of virtual power plant interaction, and market rules and clearing mechanisms constitute the environmental dynamics; Intelligent agent: Corresponding to each virtual power plant, as an independent decision-making entity, responsible for formulating bidding strategies to maximize its own profits; Actions: The bidding strategy corresponding to the virtual power plant, including specific decision-making behaviors such as electricity pricing and carbon quota bidding; Action space: The set of bidding strategies, which defines the range of all possible actions of the agent, including the upper and lower limits of the bid. Observation: The bidding results corresponding to the virtual power plant, i.e., the local feedback information obtained by the agent from the environment, including the winning bid amount and price; Observation space: The set of winning bids, representing the range of all possible observations; Status: Corresponds to the market clearing results, including overall market information such as clearing price and supply and demand balance, used to describe the overall environmental situation; State space: The set of market clearing outcomes, encompassing all possible market states; Reward: The revenue corresponding to the virtual power plant is a quantitative representation of the agent's optimization objective, calculated based on the winning bid and cost.

10. A distributed bidding method for virtual power plants based on the MADDPG algorithm according to claim 9, characterized in that, Model solving methods based on the MADDPG algorithm include: (1) Initialization configuration: For each virtual power plant participating in the bidding, initialize the parameters of the Actor online network, Actor target network, Critic online network, and Critic target network respectively, and initialize the experience replay pool to store the experience samples generated by multi-agent interaction. (2) Iterative training start: Set the total number of episodes for iteration, and repeat the following process until the preset number of episodes is reached: (2.1) Market state initialization: At the start of each episode, the initial clearing results of the electric carbon coupling market are initialized to provide the initial environment state for subsequent bidding interactions; (2.2) Time step interaction loop: Set the total time steps for each episode, and repeat the following process until the preset number of time steps is reached: (2.2.1) Bidding Strategy Formulation: Each virtual power plant observes the market clearing result at the current moment and formulates a corresponding bidding strategy based on the decision output of its own Actor online network; (2.2.2) Market Clearing and Revenue Calculation: After each virtual power plant submits its bidding strategy, the electricity-carbon coupling market completes one clearing cycle according to the clearing rules, generates the market clearing result for the next time step, and calculates the actual revenue of each virtual power plant; (2.2.3) Experience Sample Storage: The state transition sample consisting of the current state, action, reward, and next state is stored in the experience replay pool; (2.2.4) Network parameter update: Randomly sample a batch of experience samples from the experience replay pool. For each virtual power plant, update the parameters of its Actor online network and Critic online network based on the sampled samples. Use the Q-value estimation of the Critic network to guide the optimization of the Actor network strategy. (2.2.5) Target Network Soft Update: The parameters of the Actor and Critic target networks of each virtual power plant are updated using a soft update method to ensure the stability of the training process; (2.3) Iteration Termination: When the preset maximum number of episodes is reached or the network training converges, the iteration is terminated and the final virtual power plant bidding strategy model is output.