A virtual power plant multi-type electric power commodity joint transaction game method and system

By constructing a multi-level market price difference arbitrage framework and the DDPG_AE algorithm to optimize the trading strategy of virtual power plants, the complexity problem of virtual power plants in multi-type electricity commodity trading is solved, the global optimal electricity trading is achieved, and the market trading efficiency and resource allocation benefits are improved.

CN119919175BActive Publication Date: 2025-10-17STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411983454.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-10-17
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Virtual power plants face challenges such as uncertainty in electricity supply and demand, complexity of market rules, and balancing economic and environmental benefits when participating in various types of electricity commodity transactions. An effective game theory approach is needed to optimize the decision-making process to ensure their efficient operation and maximum benefit in the electricity market.

Method used

A multi-layer interaction framework between VPPOs and VPPs for various types of electricity commodities is constructed based on the principle of multi-level market price difference arbitrage. Combined with the data-driven method of the DDPG_AE algorithm, a two-layer model is designed to optimize trading strategies. By minimizing cost and utility and maximizing market share, the globally optimal electricity trading strategy is achieved.

Benefits of technology

It effectively solved the complexity problem of virtual power plants in the joint trading of multiple types of electricity commodities, realized the global optimal electricity trading strategy, and improved market transaction efficiency and resource allocation benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919175B_ABST
    Figure CN119919175B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of power transaction, in particular to a virtual power plant multi-type power commodity joint transaction game method and system, the method comprising: constructing a multi-layer interaction framework between VPPO and VPP of multi-type power commodities; constructing a single VPP decision model with the minimum cost and utility and the maximum market share as the target; constructing a multi-virtual power plant operation optimization model based on different targets, and obtaining local market transaction results by using distributed solution; constructing a VPPO decision model with the minimum cost as the target based on the multi-VPP operation optimization results and the local market transaction results; designing the transaction process of VPPO and VPP as a double-layer model, solving the model by using the data-driven method of DDPG_AE algorithm, and obtaining the globally optimal power transaction strategy. Through the present application, the complexity problem of virtual power plant in multi-type power commodity joint transaction is effectively solved, the globally optimal power transaction strategy is realized, and the market transaction efficiency and resource allocation benefit are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power transaction, in particular to a virtual power plant multi-type power commodity joint transaction game method and system. BACKGROUND

[0002] With the rapid development of renewable energy and the widespread application of distributed energy resources (DERs), virtual power plant (VPP) as a new energy management technology plays an increasingly important role in the electricity market. Virtual power plant aggregates various DERs such as photovoltaic, wind power, energy storage system and controllable load, forming an energy supply system similar to traditional power plant, and can participate in various power market transactions such as electricity transaction, auxiliary service transaction and demand response transaction. This aggregation not only improves the economy and stability of DERs, but also enhances the flexibility and reliability of the power system.

[0003] However, virtual power plant faces many challenges when participating in multi-type power commodity transactions, such as uncertainty of power supply and demand, complexity of market rules, balance of economic benefits and environmental benefits, etc. In order to cope with these challenges, an effective game method is needed to optimize the decision-making process of virtual power plant, ensuring its efficient operation and maximum benefit in the electricity market. SUMMARY

[0004] The present application provides a virtual power plant multi-type power commodity joint transaction game method and system, thereby effectively solving the problems pointed out in the background art.

[0005] In order to achieve the above purpose, the technical solution adopted by the present application is:

[0006] A virtual power plant multi-type power commodity joint transaction game method, comprising:

[0007] Based on the multi-level market price difference arbitrage principle, a multi-layer interaction framework between VPPO and VPP of multi-type power commodity is constructed;

[0008] According to the resource configuration and market transaction optimization of VPP and based on the multi-layer interaction framework, a single VPP decision model is constructed with the goal of minimizing cost and utility and maximizing market share;

[0009] Based on the single VPP decision model, a multi-virtual power plant operation optimization model based on different targets is constructed, and a distributed solution to the multi-VPP multi-type power commodity sharing problem is adopted to obtain the local market transaction result;

[0010] Based on the multi-VPP operation optimization results and local market transaction results, a VPPO decision model is constructed with the objective of cost minimization;

[0011] Based on the peer-to-peer non-cooperative game theory, the transaction process of VPPO and VPP is designed as a double-layer model, and the data-driven method of DDPG_AE algorithm is adopted to solve the model based on the way of one party pricing and the other party bidding, so as to obtain the globally optimal power transaction strategy.

[0012] Further, based on the multi-level market price difference arbitrage principle, a multi-level interaction framework between VPPO and VPP of multiple types of electric power commodities is constructed, including:

[0013] According to the hierarchical structure of the power market, the types of electric power commodities in different market levels are determined;

[0014] Based on the price fluctuation of the multi-level market, the potential arbitrage opportunity is identified by calculating the price difference between each market level;

[0015] VPPO allocates different types of electric power commodities between each market level by formulating market participation strategies and pricing mechanisms;

[0016] An interaction framework between VPPO and VPP is constructed, and a bidirectional information transmission mechanism is established, VPPO transmits the signal of electric power commodity transaction to VPP through price signal, and adjusts the market strategy according to the feedback of VPP.

[0017] Further, the single VPP decision model is:

[0018]

[0019] In the formula, C i is the total cost of VPP, is the MT generation cost, is the purchase power cost of VPP and VPPO transaction, is the purchase power cost of P2P transaction between VPPs, is the purchase reserve cost of VPP i, is the purchase reserve cost of P2P transaction of VPP i, is the purchase frequency modulation cost of VPP i, is the purchase frequency modulation cost of P2P transaction of VPP i, is the energy efficiency of VPP i, is the reserved reserve cost of VPP i.

[0020] Further, in the initial stage of market development, VPP expands market share, and the single VPP model for occupying local market is:

[0021]

[0022] wherein U i is the objective function of market share acquisition, is the price traded by VPP i with VPPO in time period t; is the energy traded by VPP i with VPPO in time period t; is the up reserve, down reserve price traded by VPP i with VPPO in time period t; is the up reserve, down reserve price traded by VPP i with VPPO in time period t; is the up reserve, down reserve amount traded by VPP i with VPPO in time period t; is the up reserve, down reserve amount traded by VPP i with VPPO in time period t; is the frequency regulation price traded by VPP i with VPPO in time period t; is the frequency regulation amount traded by VPP i with VPPO in time period t.

[0023] Further, the augmented Lagrangian function of the optimization problem is:

[0024]

[0025] wherein, wherein C i (x i,t ) is the total cost of the ith VPP under decision x i,t ; x i,t is the variable vector of the VPP decision model; y i,t = Y ij,t is the auxiliary variable, Y ij,t is the energy, up reserve amount, down reserve amount and frequency regulation amount traded by VPP i with VPP j; is the Lagrangian multiplier corresponding to the P2P transaction constraint; ρ t is the iteration step size; ||·||2 is the 2-norm.

[0026] Further, the VPPO decision model is:

[0027]

[0028] wherein, is the ESS wear cost of VPPO, is the reserve cost reserved by VPPO, is the wholesale market purchase and sale electricity cost of VPPO, is the wholesale market purchase reserve cost of VPPO, Cost of frequency regulation for wholesale market, Revenue of selling electricity for VPPO local market, Revenue of selling reserve for VPPO local market, Revenue of selling frequency regulation for local market.

[0029] Further, VPPO adjusts pricing strategy in the interaction process using reinforcement learning, which can be formalized as an MDP, the MDP includes:

[0030] State space, for describing the current situation of the environment, the state space is represented as a vector, reflecting the current state of the system;

[0031] Action space, representing all possible actions or strategies that the agent can choose, the action space is a vector, combining scheduling strategy and pricing strategy, for guiding the specific actions taken by the agent in the decision-making process;

[0032] Reward function, for reflecting the optimization goal, calculating the immediate revenue or cost of the agent after performing an action in a certain state;

[0033] State transition probability, describing the probability of the environment being affected and transitioning to the next state after performing an action in a given state.

[0034] Further, the data-driven method of DDPG_AE algorithm is used to solve the model, and the globally optimal power trading strategy is obtained, including:

[0035] Initialize the network structure and parameters of DDPG_AE algorithm, including the parameters of actor network and critic network, experience replay pool and adaptive exploration coefficient;

[0036] The agent selects appropriate trading actions based on the current policy and Gaussian noise by obtaining the current system state, and executes the actions according to market feedback, transitions to the next state and obtains the corresponding reward, and stores the four-tuple in the experience replay pool;

[0037] When the number of samples in the experience replay pool reaches the preset threshold, the agent randomly samples batch samples from the replay pool, calculates the target value of the critic network and the timing difference error, and updates the parameters of the actor network and the critic network, and updates the target network parameters using the soft update mechanism;

[0038] The agent repeatedly samples and trains, gradually adjusts the exploration coefficient to optimize the strategy, and outputs the final actor network and critic network parameters when the stop condition is met, to obtain the globally optimal power trading strategy.

[0039] Further, the DDPG_AE has a two-layer loop structure, including:

[0040] The outer loop represents one iteration of the algorithm;

[0041] The inner loop controls the parameter training in each timestamp in each iteration.

[0042] A virtual power plant multi-type power commodity joint transaction game system, the system comprises:

[0043] A multi-layer interactive framework construction module constructs a multi-layer interactive framework between VPPO and VPP of multi-type power commodities based on the multi-layer market price difference arbitrage principle;

[0044] A single decision model construction module constructs a single VPP decision model according to the resource configuration and market transaction optimization of the VPP and based on the multi-layer interactive framework, with the minimum cost and utility and the maximum market share as the target;

[0045] A market transaction result acquisition module constructs a multi-virtual power plant operation optimization model based on different targets based on the single VPP decision model, and solves the multi-VPP multi-type power commodity sharing problem in a distributed manner to obtain the local market transaction result;

[0046] A VPPO decision model construction module constructs a VPPO decision model based on the multi-VPP operation optimization result and the local market transaction result, with the minimum cost as the target;

[0047] An optimal transaction strategy acquisition module designs the transaction process of VPPO and VPP as a double-layer model based on the equal non-cooperative game theory, and solves the model by using the data-driven method of the DDPG_AE algorithm in the way of one party pricing and the other party bidding to obtain the globally optimal power transaction strategy.

[0048] Through the technical scheme of the present application, the following technical effects can be achieved:

[0049] The complexity problem of virtual power plants in multi-type power commodity joint transaction is effectively solved, the globally optimal power transaction strategy is realized, and the market transaction efficiency and resource allocation benefit are improved. DETAILED DESCRIPTION

[0050] In order to more clearly illustrate the technical scheme in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments described in the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0051] Figure 1 It is a flowchart of a virtual power plant multi-type power commodity joint transaction game method.

[0052] Figure 2 It is a multi-layer interaction framework diagram of a wholesale market, a virtual power plant aggregator and multiple virtual power plants.

[0053] Figure 3 It is a solution framework diagram of the DDPG_AE algorithm. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the description of the application herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.

[0056] Embodiment one

[0057] As shown in the drawings, the present application provides a virtual power plant multi-type power commodity joint transaction game method, the method comprising: Figure 1 S1: based on the multi-level market price difference arbitrage principle, a multi-layer interaction framework between VPPO and VPP of multiple types of power commodities is constructed;

[0058] Specifically, by analyzing the price difference between different power markets, the price difference opportunities in the hierarchical markets such as the spot market, the power futures market, and the ancillary service market are identified. VPPO can arbitrage in these price differences to obtain economic benefits. The construction of the multi-layer interaction framework can ensure the decision-making coordination between VPPO and VPP, so that they can effectively transmit information and adjust strategies between different market levels, and realize the optimal scheduling and transaction of power resources. This framework provides a new way for VPPO to jointly trade different types of power commodities according to market conditions, thereby maximizing economic benefits.

[0059] S2: according to the resource configuration and market transaction optimization of VPP and based on the multi-layer interaction framework, a single VPP decision-making model is constructed with the minimum cost and utility and the maximum market share as the target;

[0060]

[0061] ​Specifically, based on the optimization of resource allocation and market transactions, a decision model is built for each individual VPP under the support of a multi-layer interaction framework, aiming to minimize cost and utility and maximize market share. In this process, VPP needs to reasonably allocate and schedule resources according to market demand and available resources, optimize power trading strategies to reduce operating costs and improve resource use efficiency. At the same time, the decision model aims to help VPPs improve their competitiveness in the electricity market and strive for greater market share through precise resource scheduling and optimized market pricing strategies.

[0062] S3: Based on the single VPP decision model, a multi-VPP operation optimization model based on different objectives is constructed, and a distributed solution is used to solve the multi-VPP multi-type power commodity sharing problem to obtain local market transaction results.

[0063] Specifically, by considering the resource allocation and market demand between different VPPs, a comprehensive optimization framework is constructed to achieve efficient operation of the entire virtual power plant system. The optimization model aims to minimize cost, maximize utility, and maximize market share, and aims to maximize the overall efficiency of the virtual power plant by reasonably adjusting the operation strategy of each VPP. In addition, due to the complexity of interaction and resource sharing between multiple VPPs, a distributed solution method is used to efficiently handle the sharing problem of multi-VPP multi-type power commodities. Each VPP performs local optimization calculation based on its own resources and market conditions, and solves the final global optimization problem through coordination mechanism to obtain the transaction results of the local market.

[0064] S4: Based on the multi-VPP operation optimization results and local market transaction results, a VPPO decision model is constructed with the goal of minimizing cost.

[0065] Specifically, VPPO integrates the operation optimization results of multiple VPPs and combines the transaction conditions of each VPP in the local market to comprehensively evaluate the resource allocation and transaction strategy of the entire virtual power plant. By analyzing these results, VPPO can optimize its resource scheduling and market strategy to ensure the best cost-effectiveness in the overall electricity market. The decision model of VPPO will adjust the operation strategy based on this information to minimize resource scheduling costs, market transaction costs, and other operating expenses, thereby improving the economic efficiency of the virtual power plant.

[0066] S5: Based on the equal non-cooperative game theory, the transaction process between VPPO and VPP is designed as a double-layer model, and the DDPG_AE algorithm is used to solve the model based on the way of one party pricing and the other party bidding to obtain the globally optimal power trading strategy.

[0067] Specifically, based on the peer-to-peer non-cooperative game theory, the transaction process between the virtual power plant operator (VPPO) and the virtual power plant unit (VPP) is designed as a double-layer game model, and the interaction and decision-making of both parties are described through this model. This step adopts the game strategy of one party pricing and the other party bidding, and combines the data-driven method of DDPG_AE algorithm (deep deterministic policy gradient algorithm combined with autoencoder) to solve the problem, so as to realize the globally optimal power transaction strategy. In this process, VPPO and VPP are regarded as the two parties of the game, VPPO is responsible for pricing, and VPP reports the transaction volume according to the market price. Through the peer-to-peer non-cooperative game theory, the model can simulate the optimal decision-making behavior of both parties without cooperation constraints, and ensure the maximization of the interests of both parties. The DDPG_AE algorithm is used to process the high-dimensional continuous space problem in this game model, and the reinforcement learning is used to continuously adjust the decision-making strategy, and the data-driven method is used to find the optimal power transaction decision; finally, based on the double-layer game model and the data-driven method, a globally optimal power transaction strategy can be obtained, so that VPPO and VPP can realize the optimal economic benefits in the game process, thereby improving the operation efficiency and market competitiveness of the entire virtual power plant.

[0068] Through the present application, the complexity problem of virtual power plant in multi-type power commodity joint transaction is effectively solved, and the globally optimal power transaction strategy is realized, and the market transaction efficiency and resource allocation benefit are improved.

[0069] As a preferred embodiment of the above embodiment, based on the multi-level market price difference arbitrage principle, a multi-layer interaction framework between VPPO and VPP of multi-type power commodities is constructed, including:

[0070] S11: according to the hierarchical structure of the power market, the type of power commodity in different market levels is determined;

[0071] S12: based on the price fluctuation of multi-level market, the potential arbitrage opportunity is identified by calculating the price difference between each market level;

[0072] S13: VPPO distributes different types of power commodities among various market levels by formulating market participation strategies and pricing mechanisms;

[0073] S14: an interaction framework between VPPO and VPP is constructed, and a bidirectional information transmission mechanism is established, VPPO transmits the signal of power commodity transaction to VPP through price signal, and adjusts the market strategy according to the feedback of VPP.

[0074] Specifically, the virtual power plant (VPP) is taken as the research object. These power plants are equipped with distributed renewable energy sources such as photovoltaic (PV) and wind turbine (WT), as well as adjustable generators such as micro turbine (MT), and have certain flexible load (FL) adjustment potential. Due to capacity and policy restrictions, these VPPs cannot directly participate in wholesale market transactions. Therefore, the virtual power plant aggregator (VPPO) acts as an upper agent to represent these VPPs to participate in wholesale market transactions. To cope with the energy fluctuation risk brought by market price fluctuations, the VPPO manages a certain capacity of energy storage system (ESS).

[0075] In terms of market architecture, the VPPO does not directly control the resources of the VPP, but raises demand-side resources through local markets, participates in the transaction of the upper wholesale market, and makes arbitrage using the price difference of the multi-level market. Since the VPPO provides a way for the VPP to participate in the wholesale market, it has the pricing power of the VPP's electricity commodities such as electricity purchase and sale, frequency regulation and reserve. The interaction between VPPO and VPP can be summarized as a bi-level optimization problem, as shown in Figure 2 .

[0076] In Figure 2 , the upper layer is VPPO and the lower layer is VPP. VPPO finds the optimal purchase and sale prices under the existing VPP response potential based on the prediction results of the wholesale market day-ahead price. These prices are transmitted to the VPP at the bottom as signals reflecting the system supply and demand balance state, guiding them to optimize their own power generation and consumption behavior, and reasonably allocate the proportion of local market and VPPO transaction to minimize the management cost of energy, frequency regulation and reserve.

[0077] As a preferred embodiment of the above embodiment, the single VPP decision model is:

[0078]

[0079] In the formula, C i is the total cost of VPP, is the MT power generation cost, is the purchase electricity cost of VPP and VPPO transaction, is the purchase electricity cost of P2P transaction between VPPs, is the purchase reserve cost of VPP i, is the purchase reserve cost of P2P transaction of VPP i, is the purchase frequency regulation cost of VPP i, is the purchase frequency regulation cost of P2P transaction of VPP i, is the energy utility of VPP i, is the reserved reserve cost of VPP i.

[0080] As a preferred embodiment of the above, in the initial stage of market development, the VPP expands market share, and the single VPP captures the local market model is:

[0081]

[0082] In the formula, U i is the target function of VPP I for capturing market share, is the price traded by VPP i with VPPO in the t period; is the electricity energy traded by VPP i with VPPO in the t period; is the upper reserve, lower reserve price traded by VPP i with VPPO in the t period; is the upper reserve, lower reserve price traded by VPP i with VPPO in the t period; is the upper reserve, lower reserve amount traded by VPP i with VPPO in the t period; is the upper reserve, lower reserve amount traded by VPP i with VPPO in the t period; is the purchase and sale frequency modulation price traded by VPP i with VPPO in the t period; is the purchase and sale frequency modulation amount traded by VPP i with VPPO in the t period.

[0083] Considering the adjustable resources of equipment, operating characteristics, purchase and sale of electricity commodity cost and type, etc., a single virtual power plant decision model is constructed with the minimum cost and utility and the maximum market share as the target.

[0084] 1) Single VPP participates in local market transaction cost and utility model

[0085] ① MT generation cost

[0086]

[0087] In the formula, i∈I,I is the VPP set; t∈T,T is the time period set; is the active power output of MT of VPP i in the t period; a i , b i , c i is the MT cost coefficient.

[0088] ② VPP i purchase and sale electricity cost

[0089] To meet the joint sharing demand of energy, frequency regulation and reserve of VPP with the minimum cost, two kinds of transaction channels are provided for VPP: first, VPP can trade with other VPPs in point-to-point (P2P) transaction, and the transaction price is determined by the clearing of distribution market; second, VPP can trade with VPPO, and the transaction price is determined by VPPO. The revenue obtained by VPP from selling energy, frequency regulation or reserve to VPPO or other VPPs is counted as cost in reverse, and then

[0090] The purchase cost of VPP trading with VPPO is:

[0091]

[0092] wherein, is the purchase cost of VPP trading with VPPO; is the price of VPP i trading with VPPO in period t; is the amount of energy purchased by VPP i from VPPO in period t.

[0093] The purchase cost of VPPs trading with each other in P2P transaction is:

[0094]

[0095] wherein, is the purchase cost of VPPs trading with each other in P2P transaction; is the price of VPP i trading with VPP j in P2P energy transaction in period t; is the amount of VPP i trading with VPP j in P2P energy transaction in period t.

[0096] ③VPP i purchase reserve cost

[0097] VPP i can also trade with upper VPPO or lower VPP in P2P reserve transaction, and the purchase reserve cost of VPP i trading with VPPO is:

[0098]

[0099] wherein, is the purchase reserve cost of VPP i trading with VPPO; is the price of upper reserve purchased by VPP i from VPPO in period t; is the price of lower reserve purchased by VPP i from VPPO in period t; is the amount of upper reserve purchased by VPP i from VPPO in period t; is the amount of lower reserve purchased by VPP i from VPPO in period t.

[0100] The purchase reserve cost of VPP i for P2P transaction is:

[0101]

[0102] In the formula, The purchase reserve cost of VPP i for P2P transaction; The upper reserve price and lower reserve price of VPP i for transaction with VPP j in time period t; The upper reserve amount and lower reserve amount of VPP i for transaction with VPP j in time period t.

[0103] ④Purchase frequency adjustment cost of VPP i

[0104] VPP i can also conduct P2P reserve transaction with upper VPPO or lower VPP, and the purchase frequency adjustment cost of VPP i for transaction with VPPO is:

[0105]

[0106] In the formula, The purchase frequency adjustment cost of VPP i for transaction with VPPO; The purchase and sale frequency adjustment price of VPP i for transaction with VPPO in time period t; The purchase and sale frequency adjustment amount of VPP i for transaction with VPPO in time period t.

[0107] The purchase frequency adjustment cost of VPP i for P2P transaction is:

[0108]

[0109] In the formula, The purchase frequency adjustment cost of VPP i for P2P transaction; The frequency adjustment price of VPP i for transaction with VPP j in time period t; The frequency adjustment amount of VPP i for transaction with VPP j in time period t.

[0110] ⑤Utility of VPP i

[0111] Considering that FL can provide flexibility by adjusting load, this paper adopts a quadratic function with increasing marginal utility to represent the utility corresponding to the dissatisfaction degree of VPP energy experience. It is expressed as follows:

[0112]

[0113] In the formula, The utility corresponding to the dissatisfaction degree of VPP energy experience; The utility coefficient; The running power of FL of VPP i in time period t; The upper and lower limits of FL running power of VPP i, respectively, and the mean value thereof represents the most comfortable running power set by VPP i.

[0114] ⑥The reserve cost reserved by VPP i

[0115] In order to solve the problem of real-time power deviation caused by the uncertainty of distributed photovoltaic, wind turbine and load, each VPP should reserve or purchase reserve capacity to avoid potential economic risks.

[0116] The reserve cost reserved by VPP i is:

[0117]

[0118] In the formula, The reserve cost reserved by VPP i; The unit cost of upper and lower reserve provided by MT and FL of VPP i, respectively; The upper and lower reserve capacity that can be provided by MT and FL of VPP i in time period t.

[0119] In summary, the decision model of each VPP is as follows:

[0120]

[0121] In the formula, C i The total cost of VPP i.

[0122] Each VPP needs to meet the following device running constraint conditions:

[0123]

[0124]

[0125] In the formula, The active power output of PV and WT of VPP i in time period t, respectively; The active power output limit of MT of VPP i and the maximum predicted power output of PV and WT, respectively; The Boolean variable representing the transaction identity of VPP i in time period t and VPPO; constraint formula (11) to formula (12) indicates that the reserve provided by MT and FL should not exceed the surplus power generation capacity of the corresponding device. Constraint formula (13) to formula (18) indicates the DERs running characteristic constraint. Constraint formula (19) to formula (26) stipulates that VPP i can only participate in transactions in one identity in the same market in the same time period.

[0126] In addition, the VPP decision model should also meet the power balance constraint condition and the P2P transaction constraint:

[0127]

[0128]

[0129] where, is the rigid load forecast of VPP i at time period t.

[0130] Therefore, the variable vector x of VPP decision model is i which can be expressed as (32).

[0131]

[0132] 2) Single VPP pre-empting local market model

[0133] Some VPPs wish to expand market share, especially in the early stages of market development.

[0134] The decision model of each VPP is expressed as follows:

[0135]

[0136] where, U i is the objective function of VPP i to pre-empt market share.

[0137] The technical constraints are the same as above, as shown in (11)-(31).

[0138] Therefore, the variable vector x of VPP decision model is i which can be expressed as (34).

[0139]

[0140] As a preferred embodiment of the above, the augmented Lagrangian function of the optimization problem is:

[0141]

[0142] where, where, C i (x i,t ) is the total cost of the ith VPP under decision x i,t ; x i,t is the variable vector of VPP decision model; y i,t = Y ij,t is the auxiliary variable, Y ij,t is the electricity energy, upward reserve, downward reserve and frequency modulation traded between VPP i and VPP j; is the Lagrange multiplier corresponding to the P2P transaction constraint; ρ tis the iteration step size; ||·||2 is the 2-norm.

[0143] Specifically, a multi-virtual power plant (VPP) operation optimization model based on different targets is constructed, and a distributed method is used to solve the multi-VPP multi-type power commodity sharing problem. Multi-center calculation is used to reduce the computational burden, and local market transaction results are obtained under the premise of protecting the privacy of VPPs.

[0144] 1) An optimization model is established to minimize the total cost and utility of multi-VPP

[0145]

[0146] where C I is the total cost of multi-VPP.

[0147] In addition, the optimization model should satisfy the following power balance constraint condition:

[0148]

[0149] 2) An optimization model is established to quickly seize market share of multi-VPP

[0150]

[0151] where U I is the total purchase cost of multi-VPP.

[0152] In the initial stage, VPPs choose the optimization model shown in equation (39) to quickly seize market share, and choose the optimization model shown in equation (35) after stabilization.

[0153] A distributed method is used to solve the multi-objective multi-VPP multi-type power commodity sharing problem. Multi-center calculation is used to reduce the computational burden, and local market transaction results are obtained under the premise of protecting the privacy of VPPs. It is known that among all the constraint conditions, except for the P2P transaction constraint which belongs to a global coupling constraint, the remaining constraint conditions are only related to the decision variables of each VPP itself. Therefore, auxiliary variables are introduced to rewrite the coupling constraint, set X ij,t ∈x i,t , y i,t = Y ij,t , and

[0154]

[0155] where Y ij,t is the energy, upper reserve, lower reserve, and frequency modulation traded between VPP i and VPP j.

[0156] From the perspective of economics, the Lagrange multiplier corresponding to the P2P transaction constraint is The augmented Lagrangian function of the optimization problem is:

[0157]

[0158] where,

[0159]

[0160] where, p t is the iteration step; ||·||2 is the 2-norm.

[0161] Let k be the iteration number, then the solution process is as follows.

[0162] ① Update the local decision variable

[0163] VPP i obtains Solve the local optimization decision problem (44) and pass the updated to VPP j.

[0164]

[0165] ② Update the auxiliary variable

[0166] Solve the following global optimization problem to update the auxiliary variable:

[0167]

[0168] In order to realize the decentralized transaction between VPPs, while considering the solving speed and privacy protection, the coupling part in the above optimization problem (43) is separated to obtain the analytical solution of Y ij,t

[0169]

[0170] VPP i can independently update its auxiliary variable without relying on the central platform for global optimization, thus realizing completely distributed solution.

[0171] ③ Update the Lagrange multiplier

[0172] The original residual r t (k) of P2P transaction volume and the dual residual The calculation formula is as follows. In order to solve the problem of uneven convergence speed of residual, the adaptive parameter δ and τ strategy is adopted, and the step size is adjusted flexibly according to the residual state. When the original residual is large, the step size is increased to speed up the convergence speed and promote the consistency of the original variable and the auxiliary variable; while when the original residual is small, the step size is reduced to reduce the oscillation of the objective function. ​

[0173]

[0174] Finally, VPP i updates the Lagrange multiplier independently and delivers the price signal to the corresponding VPP j, i.e.

[0175]

[0176] The iteration convergence criterion is as follows:

[0177]

[0178] wherein ε1 is the set error precision.

[0179] As a preferred embodiment of the above embodiment, the VPPO decision model is:

[0180]

[0181] wherein, is the ESS loss cost of VPPO, is the VPPO reserved standby cost, is the VPPO wholesale market purchase and sale electricity cost, is the VPPO wholesale market purchase standby cost, is the wholesale market purchase frequency modulation cost, is the VPPO local market electricity sale revenue, is the VPPO local market standby sale revenue, is the local market frequency modulation sale revenue.

[0182] Specifically, the VPPO decision model is constructed with the objective of cost minimization, the transaction process between VPPO and VPP is designed as a double-layer model based on the equal non-cooperative game theory, and the data-driven method of DDPG_AE algorithm is adopted to solve the model based on the way of one party pricing and the other party bidding.

[0183] In order to ensure the maximization of the interests of both parties, the transaction process between VPPO and VPP is designed as a double-layer model. VPPO has the right to price the energy, frequency modulation and standby of VPP, but in order to avoid unreasonable pricing leading to reduced transaction volume and affecting total revenue, the transaction process adopts the DDPG_AE algorithm to solve the model.

[0184] VPPO decision model:

[0185] In a transaction cycle, the cost and revenue of VPPO participating in wholesale market and local market transactions can be expressed as follows.

[0186] ①ESS loss cost of VPPO ​

[0187]

[0188] where, and are the charging and discharging power of the ESS of VPPO at time period t; η ESS is the loss coefficient of the ESS.

[0189] ②VPPO reserved backup cost

[0190] Since the reserved part of the ESS charging and discharging capacity cannot be traded, the relevant cost is:

[0191]

[0192] where, are the reserved charging and discharging capacity of the ESS of VPPO at time period t; π ESS is the unit cost of the upper and lower backup capacity provided by the ESS.

[0193] ③VPPO wholesale market electricity purchase and sale cost

[0194] VPPO can participate in the wholesale market transaction as an independent market participant. Assuming that VPPO is the price acceptor of the wholesale market, its electricity purchase cost in the wholesale market can be expressed as:

[0195]

[0196] where, is the predicted value of the wholesale market electricity price of VPPO at time period t; P t VPPO-DSO is the demand power of VPPO at time period t.

[0197] ④VPPO wholesale market backup purchase cost

[0198]

[0199] where, are the predicted values of the upper and lower backup prices of VPPO in the wholesale market at time period t; are the upper and lower backup demand of VPPO at time period t.

[0200] ⑤VPPO wholesale market frequency modulation cost

[0201]

[0202] ⑥VPPO local market electricity sale revenue

[0203] As an intermediary between the wholesale market and the local market, the VPPO can determine the transaction price in the local market and resell the energy purchased from the wholesale market. The VPPO's revenue from selling electricity in the local market can be expressed as:

[0204]

[0205] ⑦VPPO local market reserve sale revenue

[0206] Similarly, the VPPO's reserve sale revenue in the local market is:

[0207]

[0208] ⑧VPPO local market frequency modulation sale revenue

[0209]

[0210] Based on the above analysis, the VPPO's decision model can be expressed as:

[0211]

[0212] To encourage VPPs to trade with the VPPO and balance the impact of trading electricity and trading benefits on total revenue, the VPP should satisfy the following constraints when formulating the purchase and sale electricity price:

[0213]

[0214]

[0215] where w buy and w sell are the purchase and sale pricing adjustment parameters, respectively.

[0216] The variable vector U corresponding to the VPPO decision model is shown in equation (81):

[0217]

[0218] where SOC t is the state of charge of the VPPO's ESS at time period t.

[0219] The VPPO decision model should also satisfy the following constraints:

[0220]

[0221] SOC min ≤ SOC t ≤ SOC max (91)

[0222]

[0223] B ESS and respectively, are the ESS capacity of VPPO and the limit of charge-discharge power; SOC t , SOC min , SOC max respectively, represent the state of charge of VPPO and its limit at time period t; u ESS is the Boolean variable of the state of charge-discharge of ESS at time period t. Constraints (82) to (84) represent the power and reserve balance of VPPO respectively. Constraints (85) to (88) represent that the reserve provided by ESS should not exceed its surplus generation capacity. Constraints (89) to (92) represent the operating characteristics constraints of ESS.

[0224] As a preferred embodiment of the above-mentioned embodiment, the VPPO continuously adjusts the pricing strategy in the interaction process using reinforcement learning, which can be formalized as a MDP, including:

[0225] a state space, used to describe the current situation of the environment, the state space being represented as a vector reflecting the current state of the system;

[0226] an action space, representing all possible actions or strategies that the agent can choose, the action space being a vector combining the scheduling strategy and the pricing strategy, used to guide the specific actions taken by the agent in the decision-making process;

[0227] a reward function, used to reflect the optimization objective, calculating the immediate income or cost of the agent after performing an action in a certain state;

[0228] a state transition probability, describing the probability of the environment being affected and transitioning to the next state after performing an action in a given state.

[0229] Specifically, the agent is defined as the VPPO, and the environment is defined as the world observed by the VPPO, including the upper grid, various generating units and VPPs. It is assumed that the environment can interact with the VPPO multiple times, and the VPPO can continuously adjust the pricing strategy in these interactions using reinforcement learning to obtain the optimal pricing scheme. This interaction process can be formalized as a Markov decision process (MDP), where the MDP is represented by a four-tuple (S, A, R, P), corresponding to the state space, the action space, the reward function and the state transition probability respectively.

[0230] 1) State Space: The state is used to describe the current state of the environment in the MDP. The state space defined in this paper includes the current time period, the current storage capacity of the energy storage device, the dispatch transaction volume of each VPP, and the transaction price between the VPPO and the upper-level power grid. This state space can be represented as an 18-dimensional vector, as shown in Equation (93).

[0231]

[0232] 2) Action space: The action space combines the scheduling strategy and pricing strategy of VPPO.

[0233] It can be expressed as a seventeen-dimensional vector, as shown in Equation (94).

[0234]

[0235] 3) Reward function: The reward function should directly reflect the goal of the optimization problem. This paper defines the instantaneous reward function at time t as the negative of the cost of VPPO, as shown in Equation (95):

[0236]

[0237] 4) State transition probability: The state transition probability describes the state s t Next, perform action a t After that, the environment is affected and VPPO transitions to the next state s t+1 =s' probability, as shown in formula (96):

[0238] P a (s,s')=P(s t+1 =s'|s t =s,a t =a) (96)

[0239] As a preferred embodiment of the above embodiment, the data-driven method of the DDPG_AE algorithm is used to solve the model to obtain the globally optimal power trading strategy, including:

[0240] A10: Initialize the network structure and parameters of the DDPG_AE algorithm, including the parameters of the actor network and critic network, the experience replay pool, and the adaptive exploration coefficient;

[0241] A20: The agent obtains the current system state, selects appropriate trading actions based on the current strategy and Gaussian noise, executes the actions based on market feedback, transitions to the next state, obtains the corresponding rewards, and stores the quadruple in the experience replay pool.

[0242] A30: When the number of samples in the experience replay pool reaches a preset threshold, the agent randomly samples a batch of samples from the replay pool, calculates the target value and temporal difference error of the critic network, and updates the parameters of the actor network and the critic network. The soft update mechanism is used to update the target network parameters.

[0243] A40: The agent repeatedly samples and trains, gradually adjusting the exploration coefficient to optimize the strategy, and outputs the final actor network and critic network parameters when the stopping condition is met to obtain the globally optimal electricity trading strategy.

[0244] Specifically, the DDPG_AE algorithm solution framework is as follows Figure 3 As shown in the figure, the framework first constructs four neural networks, namely the actor network μ(s|θ), the critic network Actor target network μ(s|θ') and critic target network In the sampling phase, the state s is randomly selected. t Input to the actor network, the actor network according to the current state s t Adopting adaptive action exploration mechanism to select action a t ; Then, the agent interacts with the environment according to the selected action and outputs the corresponding reward r t and the next state s t+1 , forming a four-tuple sample (s t ,a t ,r t ,s t+1 ) and stored in the experience replay pool D. This process is repeated until the number of samples in the experience replay pool reaches the predetermined size L, after which the algorithm enters the training phase. At the beginning of the training phase, the state s is first taken. t A four-tuple sample is formed and stored in the experience replay pool D. If D has reached its maximum capacity, the first four-tuple sample in D is replaced and added to the end; then, M four-tuple samples (s k ,a k ,r k ,s k+1 ), which is input into the critic network to output the Q value function of the current state-action The Q-value function represents the expected discounted reward value that can be obtained by performing an action in a given state, as shown in Equation (97); Next, the actor target network provides the next state s k+1 The estimated optimal action μ(s k+1 |θ'), critic target network to μ(s k+1 |θ') estimates the Q-value function of the state-action at time k+1 Finally, the target value y of the state-action at the current time k is calculated according to the Bellman equation k as shown in equation (98).

[0245]

[0246] wherein γ represents the discount factor of the Bellman equation; r represents the reward value corresponding to the terminal state. T

[0247] When the target value y is obtained k Then, the four network parameters can be updated. The parameters in the critic network are updated by the gradient descent method as shown in equation (99).

[0248]

[0249] wherein: represents the time difference error of the sample; represents the learning rate of the critic network.

[0250] The actor network generates the optimal action μ(s k according to the current state s k |θ), and inputs these actions into the critic network to generate the Q value function and feedback to the actor network. The actor network updates the parameters by the gradient ascent method, as shown in equation (100).

[0251]

[0252] wherein ω θ represents the learning rate of the actor network.

[0253] In the parameter updating module of the critic target network and the actor target network, a soft updating mechanism is adopted to perform weighted average on the current network parameters and the target network parameters, so as to ensure slow updating of the parameters and improve the stability of the training. The specific updating formula is shown in equations (101) and (102).

[0254]

[0255] θ'←τθ+(1-τ)θ'(102)

[0256] wherein τ∈(0, 1) represents a soft updating factor.

[0257] ​In DDPG or other reinforcement learning algorithms, the action selection strategy is very important in the optimization problem with continuous action space, which should ensure that the agent can explore the action space widely. The action selection strategy in the original DDPG algorithm adopts a normal distribution sampling method. Specifically, the agent uses the current optimal action a' = μ(s|θ) output by the actor network as the mean of the normal distribution, and introduces a standard deviation parameter δ to construct the normal distribution. Subsequently, the agent randomly samples a new action a from the normal distribution, and replaces the original action a' with the new action a and outputs from the actor network. The implementation is shown in equation (103).

[0258] a = μ ε (s) = μ(s|θ) + ε, ε ~ N(0, σ) (103)

[0259] In the equation, ε represents exploration noise, and N(0, σ) is a normal distribution with mean 0 and variance σ.

[0260] In the system optimization process, the problem model is a continuous multi-step MDP. Even if the variance σ of the exploration noise ε is set to be small, due to the effect of multi-step concatenation, the final action (s, a) distribution variance may still be too large, resulting in a lack of development in the action selection process and affecting the convergence stability and overall effect. Therefore, the embodiment proposes an adaptive exploration action selection strategy, which introduces an adaptive exploration coefficient A before the noise in equation (103), as shown in equation (104).

[0261] a = μ Aε (s) = μ(s|θ) + A·ε, ε ~ N(0, σ) (104)

[0262] In the equation, the coefficient A combines the average loss value of the critic network in the current training episode and the variance of the cumulative reward in the previous n training episodes, as shown in equations (105) and (106).

[0263]

[0264] In the equation, λ is the noise coefficient; var(·) represents the variance function; R episode-i represents the cumulative reward value generated after the agent interacts with the environment in the episode-i training episode, i = 1, 2,..., n; r t episode-i represents the immediate reward of the timestamp t in the episode-i training episode. critic network loss value calculated by the last sampled sample in the current training episode, t = 1, 2,..., T. A is a decreasing function of the sum of the reward variance and the inverse of the average loss value of the critic network, and the larger the value of A is when the sum of the two is smaller.

[0265] To solve the dimension problem of the two terms in equation (105), it is necessary to first normalize the cumulative reward value of the previous n training episodes:

[0266]

[0267] In the formula: R min and R max represent the minimum and maximum values in the cumulative reward of the previous n training episodes.

[0268] Through mathematical proof, the variance range of the reward normalization is between [0, 0.25]. In order to make the loss value of the critic network also fall within this interval, we normalize the loss value of the critic network to [4, 100], as shown in equation (108).

[0269]

[0270] In the formula: and represent the minimum and maximum values in the cumulative reward of the previous n training episodes.

[0271] According to equations (104) and (105), when the agent is in the early stage of training, the average loss value of the critic network is larger. At this time, if the change range of the reward is small (i.e. the variance of the reward is small), the adaptive exploration coefficient A is increased to enhance the exploration ability of the action space. In the later stage of training, when the average loss value of the critic network is small, if the change range of the reward is large (i.e. the variance of the reward is large), A is reduced to improve the use of the current optimal action.

[0272] As a preferred embodiment of the above embodiment, DDPG_AE has a two-layer loop structure, including:

[0273] The outer loop represents one iteration of the algorithm;

[0274] The inner loop controls the parameter training in each timestamp in each iteration.

[0275] Specifically, first, the network parameters, the experience replay pool and the adaptive exploration coefficient are initialized. Then, the sampling stage is entered, the agent obtains the initial observation state, and selects a suitable action according to the Gaussian noise. After the action is executed, the agent will be transferred to the next state according to the state transition probability and obtain the reward, and these four-tuple samples are stored at the end of the experience replay pool D. If the number of samples in the experience replay pool D reaches the preset value, the sampling stage is switched to the training stage. The agent repeats the operation of the sampling stage, and if the number of samples in D reaches the maximum capacity, the first sample in D is replaced and a new four-tuple sample is added. Next, a batch of samples are randomly sampled from D, the target value of the critic network, the time difference (TD) error and the respective gradient are calculated. The parameters of the critic network and the actor network are updated according to formula (99) and (100), and the critic target network and the actor target network are regularly updated using the soft update mechanism (such as formula (101) and (102)). Finally, the adaptive exploration coefficient is calculated according to formula (105), and a new round of iteration is started. The iteration process continues until the stop condition is met, and the final actor network parameters and critic network parameters are output.

[0276] Embodiment two

[0277] Based on the same inventive concept as the virtual power plant multi-type power commodity joint transaction game method in the foregoing embodiments, the application further provides a virtual power plant multi-type power commodity joint transaction game system, which comprises:

[0278] A multi-layer interaction framework construction module constructs a multi-layer interaction framework between VPPO and VPP of multi-type power commodities based on the multi-layer market price difference arbitrage principle;

[0279] A single decision model construction module constructs a single VPP decision model based on the resource configuration and market transaction optimization of the VPP and the multi-layer interaction framework, with the cost and utility minimization and market share maximization as the target;

[0280] A market transaction result acquisition module constructs a multi-virtual power plant operation optimization model based on different targets based on the single VPP decision model, and solves the multi-VPP multi-type power commodity sharing problem in a distributed manner to obtain the local market transaction result;

[0281] A VPPO decision model construction module constructs a VPPO decision model based on the multi-VPP operation optimization result and the local market transaction result, with the cost minimization as the target;

[0282] The optimal transaction strategy acquisition module designs the transaction process of VPPO and VPP as a double-layer model based on peer-to-peer non-cooperative game theory, and solves the model by using a data-driven method of DDPG_AE algorithm based on the mode of one party pricing and the other party bidding, so as to obtain a globally optimal power transaction strategy.

[0283] The game system in the application can effectively realize the virtual power plant multi-type power commodity joint transaction game method, and can achieve the technical effects as described in the above embodiments, which will not be repeated here.

[0284] Although the present application has been described in connection with specific features and embodiments thereof, it is obvious that it can be varied in a variety of ways without departing from the spirit and scope of the application. Accordingly, the description and drawings are to be regarded simply as illustrative in nature and are to be construed only as limiting the scope of the application insofar as it is defined in the appended claims. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalent technology, the present application is intended to include these modifications and variations.

Claims

1. A virtual power plant multi-type electricity commodity joint trading game method, characterized by: include: Based on the principle of multi-level market price arbitrage, a multi-level interactive framework between VPPOs and VPPs for various types of electricity commodities is constructed; Based on the VPP's resource allocation and market transaction optimization and based on the multi-layer interactive framework, a single VPP decision model is constructed with the goal of minimizing cost and utility and maximizing market share; Based on the single VPP decision model, a multi-virtual power plant operation optimization model with different objectives is constructed, and a distributed solution is used to solve the multi-VPP multi-type electricity commodity sharing problem to obtain local market transaction results; Based on the optimization results of multiple VPP operations and local market transaction results, a VPPO decision model is constructed with the goal of minimizing costs. Based on the theory of peer-to-peer non-cooperative game, the transaction process of VPPO and VPP is designed as a two-layer model. The data-driven method of DDPG_AE algorithm is used to solve the model based on the method of one party setting prices and the other party reporting quantities, and the globally optimal electricity trading strategy is obtained.

2. The virtual power plant multi-type electricity commodity joint trading game method according to claim 1 is characterized in that: Based on the principle of multi-level market price arbitrage, a multi-level interactive framework between VPPOs and VPPs for various types of electricity commodities is constructed, including: According to the hierarchical structure of the electricity market, determine the types of electricity commodities in different levels of the market; Based on the price fluctuations of multi-tier markets, potential arbitrage opportunities are identified by calculating the price differences between each market tier; VPPO allocates different types of electricity commodities among various market tiers by formulating market participation strategies and pricing mechanisms; Build an interactive framework between VPPO and VPP, and establish a two-way information transmission mechanism. VPPO transmits electricity commodity trading signals to VPP through price signals, and adjusts market strategies based on VPP feedback.

3. The virtual power plant multi-type electricity commodity joint trading game method according to claim 1 is characterized in that: The single VPP decision model is: Where C i is the total cost of VPP, is the MT power generation cost, is the electricity purchase cost of VPP and VPPO transactions, The cost of electricity purchased through P2P transactions between VPPs. The cost of purchasing spare parts for VPP is The purchase and standby cost of VPP i for P2P transactions, The cost of purchasing frequency modulation for VPP i, The purchase and modulation cost of VPP i for P2P transactions, Energy efficiency for VPP i, Reserve backup costs for VPP i.

4. The virtual power plant multi-type electricity commodity joint trading game method according to claim 3 is characterized in that: In the early stages of market development, VPPs expand their market share, and the model for a single VPP to seize the local market is: Where U i is VPP I is the objective function for seizing market share, is the price at which VPP i trades with VPPO during period t; The power purchase and sale between VPP i and VPPO in period t; are the upper reserve and lower reserve prices purchased by VPP i in transactions with VPPO during period t; are the upper reserve and lower reserve prices sold by VPP i in transactions with VPPO during period t; The upper and lower reserves purchased by VPP i from VPPO during period t; The upper and lower reserve quantities sold by VPP i in transactions with VPPO during period t; is the purchase and sale FM price of VPP i in transactions with VPPO during period t; is the purchase and sale frequency modulation volume of VPP i and VPPO during period t.

5. The virtual power plant multi-type electricity commodity joint trading game method according to claim 1 is characterized in that: The augmented Lagrangian function of the optimization problem is: in, Where C i (x i,t ) is the i-th VPP in decision x i,t Total cost under x i,t is the variable vector of the VPP decision model; y i,t =Y ij,t is an auxiliary variable, Y ij,t is the electric energy, upper reserve, lower reserve and frequency regulation amount traded between VPPi and VPP j; is the Lagrange multiplier corresponding to the P2P transaction constraints; ρ t is the iteration step size; ||·||2 is the 2-norm.

6. The virtual power plant multi-type electricity commodity joint trading game method according to claim 1 is characterized in that: The VPPO decision model is: Where, is the ESS loss cost of VPPO, Reserve spare costs for VPPO, The cost of electricity purchase and sale in the VPPO wholesale market, Purchase spare costs for the VPPO wholesale market, The cost of purchasing frequency modulation for the wholesale market, The income from electricity sales in the VPPO local market, To provide spare income for VPPO local market sales, Sell ​​FM revenue for the local market.

7. The virtual power plant multi-type electricity commodity joint trading game method according to claim 1 is characterized in that: VPPO uses reinforcement learning to continuously adjust pricing strategies during the interaction process, which can be formalized as an MDP. The MDP includes: The state space is used to describe the current situation of the environment. The state space is represented as a vector that reflects the current state of the system. The action space represents all possible actions or strategies that the agent can choose. The action space is a vector that combines the scheduling strategy and the pricing strategy to guide the specific actions that the agent takes during the decision-making process. The reward function is used to reflect the optimization goal and calculate the immediate benefit or cost of the agent after performing a certain action in a certain state; The state transition probability describes the probability that in a given state, after performing a certain action, the environment is affected and transitions to the next state.

8. The virtual power plant multi-type electricity commodity joint trading game method according to claim 1 is characterized in that: The data-driven approach of the DDPG_AE algorithm is used to solve the model and obtain the globally optimal power trading strategy, including: Initialize the network structure and parameters of the DDPG_AE algorithm, including the parameters of the actor network and critic network, the experience replay pool, and the adaptive exploration coefficient; The agent obtains the current system state, selects appropriate trading actions based on the current strategy and Gaussian noise, executes the actions based on market feedback, moves to the next state, obtains corresponding rewards, and stores the quadruple in the experience replay pool; When the number of samples in the experience replay pool reaches a preset threshold, the agent randomly samples batches of samples from the replay pool, calculates the target value and temporal difference error of the critic network, and updates the parameters of the actor network and the critic network, using a soft update mechanism to update the target network parameters; The intelligent agent repeatedly performs sampling and training, gradually adjusts the exploration coefficient to optimize the strategy, and outputs the final actor network and critic network parameters when the stopping condition is met to obtain the globally optimal power trading strategy.

9. The virtual power plant multi-type electricity commodity joint trading game method according to claim 1 is characterized in that: The DDPG_AE has a two-layer loop structure, including: The outer loop represents one iteration of the algorithm; The inner loop controls the parameter training at each timestamp in each iteration.

10. A virtual power plant multi-type electricity commodity joint trading game system, characterized by: The system comprises: A multi-layer interactive framework construction module, based on the principle of multi-level market price difference arbitrage, builds a multi-layer interactive framework between VPPOs and VPPs for various types of electricity commodities; A single decision model building module, which builds a single VPP decision model based on the VPP's resource allocation and market transaction optimization and the multi-layer interactive framework, with the goal of minimizing cost and utility and maximizing market share; The market transaction result acquisition module constructs a multi-virtual power plant operation optimization model based on a single VPP decision model and different objectives, and adopts a distributed solution to the multi-VPP multi-type power commodity sharing problem to obtain local market transaction results. The VPPO decision model construction module builds a VPPO decision model based on the optimization results of multiple VPP operations and local market transaction results, with the goal of minimizing costs. The optimal trading strategy acquisition module designs the trading process of VPPO and VPP into a two-layer model based on the theory of peer-to-peer non-cooperative game, and solves the model using the data-driven method of the DDPG_AE algorithm based on the method of one party setting prices and the other party reporting quantities, thereby obtaining the globally optimal electricity trading strategy.

Citation Information

Patent Citations

  • Multi-virtual power plant dynamic game transaction behavior analysis method based on finite rationality

    CN112001752A

  • Virtual power plant optimization operation and P2P transaction method based on combined game

    CN119129994A