Bidding optimization method and system for virtual power plant

By combining reinforcement learning algorithms with deep deterministic policy gradients and simulated annealing-genetic algorithms, the bidding strategies for virtual power plants and distribution networks are optimized, solving the problem of insufficient robustness in existing technologies, achieving a win-win situation for both virtual power plants and distribution networks, and improving the global optimality and dynamic adaptability of the bidding strategy.

CN120996240APending Publication Date: 2025-11-21XJ ELECTRIC CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510903952.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing virtual power plant bidding strategies employ heuristic search algorithms, resulting in low robustness and a tendency to get trapped in local optima, making it difficult to achieve a win-win situation for both virtual power plants and distribution networks.

Method used

A reinforcement learning algorithm based on deep deterministic policy gradients combined with simulated annealing-genetic algorithm is adopted to optimize the bidding strategy of virtual power plants and distribution networks through a two-layer bidding optimization model. The two-stage solution is performed using reinforcement learning algorithm, and a reward function is designed with economic benefits, risks and grid stability as objectives.

Benefits of technology

The robustness of the virtual power plant bidding strategy has been improved, enabling it to shift from local optima to global optima and achieve a common optimal value point for both the virtual power plant and the distribution network. This enhances the strategy's dynamic adaptability and convergence efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996240A_ABST
    Figure CN120996240A_ABST
Patent Text Reader

Abstract

The invention relates to a bidding optimization method and system for a virtual power plant, and belongs to the technical field of virtual power plant bidding. The method comprises the following steps: obtaining known parameters related to a double-layer bidding optimization model for resolving a virtual power plant and a power distribution network, inputting the known parameters into the model, and carrying out iterative calculation by using the model to obtain an optimal bidding strategy of the virtual power plant; a virtual power plant optimization layer in the double-layer bidding optimization model obtains a bidding strategy of a virtual power plant according to an objective function of the virtual power plant optimization layer and a clearing result fed back by a power distribution network optimization layer, and the solving process comprises the following steps: carrying out first-stage solving by using a heuristic search algorithm; and taking a first-stage solution result as an initial strategy of a reinforcement learning algorithm, and carrying out second-stage solution by utilizing the reinforcement learning algorithm. On the basis of adopting a heuristic search algorithm in the prior art, a reinforcement learning algorithm is adopted for further solving, so that the obtained optimal bidding strategy is adjusted from local optimum to global optimum, and the robustness is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a bidding optimization method and system for virtual power plants, belonging to the field of virtual power plant bidding technology. Background Technology

[0002] The distribution-side retail electricity market products include energy market products and ancillary service market products. Among them, energy market products refer to active power services used to ensure system power balance, while ancillary service market products include reactive power services used to ensure node voltage range and reserve services used to maintain system reserve balance.

[0003] The daily operation and clearing scheme of the distribution-side retail electricity market are the responsibility of the distribution system operator (DSO). The main process of market operation can be divided into the following three stages:

[0004] (1) The virtual power plant agent reports the prices and output range of various market products to the DSO;

[0005] (2) The DSO collects bidding information of virtual power plants, takes into account the output range and cost characteristics of various units, and takes into account power flow safety and the prices of various market products in the wholesale market, performs market clearing optimization, and determines the winning bid price and winning bid quantity of virtual power plant agents;

[0006] (3) Each virtual power plant agent conducts internal optimization to keep up with the number of winning bids for various products.

[0007] Virtual power plant agents determine the aggregated adjustment range of active power, reactive power, and reserve capacity based on the power load and renewable energy forecasts within their plants, taking into account the cost characteristics and output limitations of their internal generator units. This determines the type of market service they participate in and the pricing decision.

[0008] Bidding for a Virtual Power Plant (VPP) includes: active power output pricing. Range of active power reactive power output quotation reactive power output range Alternative quotes and maximum spare capacity

[0009] The pricing of virtual power plants should take into account the optimization and clearing process of the distribution network, so that both the virtual power plant and the distribution network can reach the optimal value point in the end, thus achieving a win-win situation for both.

[0010] Although virtual power plant agents and distribution network operators have different sources of revenue, their actions can influence each other. Therefore, the pricing of virtual power plants should take into account the optimized operation of the distribution network and the market clearing process.

[0011] Chinese invention patent application CN111222917A, published on June 2, 2020, discloses a virtual power plant bidding strategy interacting with a diversified retail market on the distribution side. This bidding strategy uses a hybrid simulated annealing (SA)-genetic algorithm (GA) to solve a two-layer bidding decision model for virtual power plants interacting with a diversified distribution market. The solution output includes the type of retail market the virtual power plant participates in, the output range of active power, reactive power, and spinning reserve at different time-of-use nodes, and the bidding price. The hybrid simulated annealing-genetic algorithm used in this scheme is an improvement on the traditional genetic algorithm. Essentially, it is a heuristic search algorithm. Using only a heuristic search algorithm can easily lead to getting trapped in local optima, resulting in low robustness. Summary of the Invention

[0012] The purpose of this invention is to provide a bidding optimization method and system for virtual power plants, in order to solve the problem of low robustness caused by the use of heuristic search algorithms to obtain bidding strategies for existing virtual power plants.

[0013] To achieve the above objectives, the present invention includes:

[0014] The present invention provides a bidding optimization method for a virtual power plant, comprising the following steps:

[0015] The known parameters related to the two-level bidding optimization model of virtual power plants and distribution networks are obtained and input into the model. The optimal bidding strategy of virtual power plants is obtained by iterative calculation using the model.

[0016] The two-layer bidding optimization model includes a virtual power plant optimization layer and a distribution network optimization layer;

[0017] The virtual power plant optimization layer is used to solve for the bidding strategy of the virtual power plant based on the objective function of the virtual power plant optimization layer and the clearing results fed back by the distribution network optimization layer, and then feeds it back to the distribution network optimization layer. The process of solving for the bidding strategy of the virtual power plant includes: firstly, using a heuristic search algorithm to perform a first-stage solution, and then using the first-stage solution result as the initial strategy of the reinforcement learning algorithm to perform a second-stage solution.

[0018] The distribution network optimization layer is used to solve for the clearing result of the distribution network based on the objective function of the distribution network optimization layer and the bidding strategy fed back by the virtual power plant optimization layer, and then feed it back to the virtual power plant optimization layer.

[0019] Furthermore, the reinforcement learning algorithm employs a deep deterministic policy gradient.

[0020] Furthermore, the reward function for the deep deterministic strategy gradient is set to balance economic benefits, risks, and grid stability.

[0021] Economic benefits refer to the economic benefits of a virtual power plant, while risks include the risk of loss due to extreme market volatility or insufficient renewable energy output.

[0022] Furthermore, the reward function is expressed as:

[0023] R t =α·Profit t -β·Risk t +γ·Stability t

[0024] α+β+γ=1

[0025] In the formula, R t Let α be the reward function value, β be the weight of economic returns, γ be the weight of risk, and γ be the weight of grid stability. t The economic benefit is the difference between the product of the winning bid and the bid amount, and the generation cost. t For risk, Stability t For grid stability, grid stability refers to the power imbalance in the grid caused by bidding strategies.

[0026] Furthermore, α>β>γ.

[0027] Furthermore, the heuristic search algorithm is a simulated annealing-genetic algorithm, which is obtained by improving the genetic algorithm. The improvement is that after performing the genetic operation, a simulated annealing operation is added, and then the determination of whether the iteration termination condition of the genetic algorithm is met is performed.

[0028] Furthermore, the clearing results include the clearing price and the winning bid volume of the virtual power plant at each node.

[0029] Furthermore, the bidding strategy for virtual power plants includes the active, reactive, and reserve output prices and output range of the virtual power plant at each node.

[0030] The present invention provides a bidding optimization system for a virtual power plant, comprising a processor for executing a computer program to implement the steps of the bidding optimization method for a virtual power plant as described above.

[0031] The beneficial effects of this invention are:

[0032] This invention is an improved invention, providing a bidding optimization method for virtual power plants. Based on the existing method of obtaining bidding strategies for virtual power plants using heuristic search algorithms, this method further employs reinforcement learning algorithms for solving the problem, which can adjust the obtained optimal bidding strategy from local optimum to global optimum, thereby improving robustness. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the existing two-layer optimized structure of virtual power plants and distribution networks;

[0034] Figure 2 This is a flowchart of the model optimization solution based on SA-GA-RL. Detailed Implementation

[0035] To address the problems in the background technology, this invention utilizes the SA-GA-RL algorithm to optimize and solve the two-layer bidding optimization model. While retaining the global jump advantage of SA and the parallel search advantage of GA, it employs reinforcement learning algorithms to improve the dynamic adaptability and convergence efficiency of the policy, thereby enhancing robustness.

[0036] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0037] Implementation of a bidding optimization method for virtual power plants:

[0038] A bidding optimization method for virtual power plants is disclosed, relating to the field of power system technology, and particularly to a bidding optimization method for virtual power plants based on the SA-GA-RL hybrid algorithm. The specific content of this method is as follows: it can take into account the respective interests of virtual power plants and distribution networks and achieve a win-win situation, so that both virtual power plants and distribution networks can reach the optimal value point while balancing risks.

[0039] I. First layer, distribution network operator optimization layer, refer to Figure 1 The distribution network optimization layer in the system.

[0040] The optimization objective of a distribution network operator is to minimize the total operating cost of the distribution network, which includes the operating costs of traditional generator units and the costs of purchasing electricity from the wholesale market and virtual power plants. The specific mathematical model for the distribution network operator's optimization objective can be described as follows:

[0041]

[0042] Where: Φ DN ΔT represents the total operating cost of the distribution network; T represents the total number of operating segments (h); ΔT represents the duration of the optimized time interval (h). N represents the number of traditional generator sets. VPPThis represents the number of virtual power plant grid-connected nodes. and These represent the prices of active power, reactive power, and reserve power purchased from the main grid by the distribution side, in $ / MWh. and These represent the active power, reactive power, and reserve capacity purchased by the distribution network from the main grid, in MW (MVar). and Let $ be the bid price for active power, reactive power, and reserve power of the virtual power plant at node i at time t, respectively, in MWh. and Let be the active power, reactive power, and reserve capacity purchased by the distribution network from the main virtual power plant at node i at time t, respectively, in MW (MVar).

[0043] The cost characteristics of traditional generator sets can be expressed as follows:

[0044]

[0045] In the formula: Let the operating cost function of the i-th distributed generator set be denoted as . and Let be the active power and reserve capacity of generator unit i at time t, respectively, in MW; The probability of being called at time t, %; a i b i and c i All of these are operating cost parameters for traditional generator set i.

[0046] The decision variables for this optimization problem are: Constraints for distribution network operation optimization include output constraints of various generating units and virtual power plants (VPPs), node voltage range constraints, branch power flow capacity limitations, and system reserve balance, etc., as detailed below:

[0047] 1) A linearized distribution-side power flow equation is used to describe the distribution network topology and power flow model as follows:

[0048]

[0049] In the formula: N B P represents the total number of nodes in the distribution network. i Q i θ i and V i The injected active power, injected reactive power, phase angle, and voltage magnitude at node i are pu; S N Reference power, MW; P ij Q ij r ij and xij Let pu represent the active power flow, reactive power flow, resistance, and reactance of the line between node i and node j, respectively. and Let θ represent the active power output (MW) of the wind farm and the photovoltaic power station at node i, respectively; j Let V be the voltage phase angle at node j. j P represents the voltage amplitude at node j. L,i,t Q L,i,t These represent the active and reactive loads of node i during time period t, respectively.

[0050] In addition, during system operation, the voltage at each node and the power flow in each branch should be limited to a certain range:

[0051]

[0052] In the formula: pu; V represents the upper limit of the apparent power of the line between node i and node j. i and Let pu be the lower and upper voltage limits at node i, respectively.

[0053] 2) The output range of each generator set and renewable energy unit in the system should be limited to a certain range:

[0054]

[0055] In the formula: P i DG , and Let be the upper and lower limits of active power output and reactive power output of generator i at time t, respectively, in MW (MVar); Let i be the rated capacity of generator i, in MW; λ represents the upper limit of generator i's ramp output, in MW / h; DG This is the power factor limit value for the generator set; and Let be the predicted average power output of wind farm i and photovoltaic power station i at time t, respectively, in MW.

[0056] To describe the uncertainty of renewable energy in the system, it is assumed that the prediction errors of wind and solar power output have a mean of 0 (μ = 0) and a standard deviation of 5% of the prediction mean (σ = 5%·P). j,t,mean Gaussian probability density distribution function ε j,t ~N(μ,σ 2 ), where P j,t,mean This is represented as the baseline value of the prediction error. The fluctuation range of the error (σ) is proportional to the prediction mean (e.g., the larger the prediction mean, the larger the possible range of absolute error).

[0057] This invention uses confidence levels to describe the risk attitudes of distribution network operators and virtual power plant agents towards their internal renewable energy output. Different confidence levels lead to different bidding decision results for virtual power plants. Confidence level C level The relationship between the expected output of renewable energy can be described as follows:

[0058]

[0059] In the formula: P i,t,fore Let P be the predicted output of renewable energy i at time t, in MW; i,t,mean Let ε be the expected mean output of renewable energy i at time t, in MW; i,t Let be the power output prediction error of renewable energy i at time t, in MW; P r (P i,t,fore >P i,t,mean +ε i,t ) represents the calculation of the actual predicted output P i,t,fore Exceeding threshold P i,t,mean +ε i,t The probability of f ε (ε i,t ) represents the probability density function of the error term.

[0060] 3) The total reserve in the system should be higher than the reserve requirement:

[0061]

[0062] In the formula: and Total reserve requirements for load, wind power, and solar power, respectively, in MW.

[0063] 4) Virtual Power Plant Output Limitations. During the distribution network clearing phase, the service types and output ranges provided by virtual power plants are known, and their active, reactive, and reserve outputs should be limited to the bidding scope:

[0064]

[0065] In the formula: and Let be the upper and lower limits of the active power output of the virtual power plant at node i at time t, in MW. and Let MVar represent the upper and lower limits of the reactive power output of the virtual power plant at node i at time t. Let λ be the upper limit of the reserve capacity bid for the virtual power plant at node i at time t, in MW; VPP,i This represents the power factor limit value for the virtual power plant at node i.

[0066] The above-mentioned power distribution network operation optimization model is a typical convex optimization problem, and its compact form can be expressed as follows:

[0067] min Φ(x)

[0068]

[0069] f n (x)≤0,n=1,2,...N

[0070] In the formula: M is the total number of equality constraints; N is the total number of inequality constraints; Φ(x) is the objective function, which is the minimization objective function; x is the decision variable vector, which contains all the variables that need to be optimized; For equality constraints, where a m It is the coefficient vector of the constraint, b m It is a constant term, m is the index of the equality constraint, ranging from 1 to M; f n (x)≤0 is an inequality constraint, where f n (x) is a function of the inequality constraints, and n is the index of the inequality constraints, ranging from 1 to N.

[0071] By introducing Lagrange multipliers, we can obtain the Lagrange function of the original problem as follows:

[0072]

[0073] In the formula: λ m μ is the Lagrange multiplier corresponding to the equality constraint m; n Let n be the Lagrange multiplier corresponding to the inequality constraint n.

[0074] The problem can be solved using the Karush-Kuhn-Tucker (KKT) conditions, yielding an economical operating plan for the distribution network station (DSO). Furthermore, according to the envelope theorem, the Lagrange multipliers of the equality constraints correspond to the cleared active power price, reactive power price, and reserve price at node i of the distribution network.

[0075] II. Second layer, virtual power plant agent optimization layer, see reference. Figure 1 The virtual power plant optimization layer in the system.

[0076] Virtual power plants exist in multiple forms, including renewable energy sources, traditional generators, and power loads. Furthermore, virtual power plants can connect to the distribution network from multiple nodes and simultaneously participate in the distribution network's active, reactive, and reserve markets. The decision variables for the virtual power plant agency bidding optimization problem are the virtual power plant's bid prices and output range for active, reactive, and reserve power at each node: The optimization objective of virtual power plant agency is to maximize the revenue of the virtual power plant. The optimal bidding strategy obtained by the virtual power plant optimization layer includes the market types in which the virtual power plant participates at each node, the output range of various products, and the bid price. The optimization objective of virtual power plant agency can be expressed as:

[0077]

[0078] In the formula: Φ VPP The revenue of the virtual power plant; T is the total number of running segments, h; ΔT is the optimization time interval, h; N VPP This represents the total number of nodes connecting the virtual power plant to the distribution network. The total number of generators owned by the virtual power plant; and These represent the active power price, reactive power price, and reserve power price obtained at node j at time t through distribution network clearing, respectively, $ / MWh ($ / MVarh). and Let the winning active power, winning reactive power, and winning reserve capacity of the virtual power plant at node j at time t be MW (MVar); and Let be the active power output and reserve power provided by generator unit j within the virtual power plant at time t, respectively, in MW; Let be the operating cost function of the j-th distributed generator unit.

[0079] Among them, the electricity price of each node of the virtual power plant and the winning bid volume of various products can be obtained by the distribution network optimization and clearing. The operating constraints of the virtual power plant include: internal power and reserve balance constraints of the virtual power plant, output limits of various units and quotation range constraints, etc.

[0080] 1) The active power, reactive power, and reserve balance at each node of the virtual power plant should be satisfied as follows:

[0081]

[0082] In the formula: and Let be the wind power and solar power output of the virtual power plant at node j at time t, respectively, in MW; and Let be the active and reactive loads (in MW) of the virtual power plant at node j at time t; and Let be the load at node j at time t, and the total reserve demand for wind power and photovoltaic power, respectively, in MW.

[0083] 2) The bidding price for virtual power plants should meet the government's restrictions, and the bidding output range should meet the actual aggregated available range of the virtual power plant.

[0084]

[0085] In the formula: These represent the lower and upper limits of the active power output of virtual power plant j in time period t, respectively. The lower limit of active power output is the minimum active power that must be provided (such as the contractual commitment value), and the upper limit of active power output is the maximum active power that can be dispatched (limited by the capacity of distributed resources). These represent the lower and upper limits of reactive power output of virtual power plant j in time period t, respectively, which can provide the range of reactive power adjustment to support grid voltage stability. This represents the maximum reserve capacity of virtual power plant j during time period t, which is the upper limit of reserve power that can respond quickly (such as the reserve capacity of energy storage or interruptible loads). Let the feasible region of aggregated power output of the virtual power plant at node j at time t be defined. and These are the upper and lower limits for active power pricing, reactive power pricing, and standby pricing, respectively, $ / MWh ($ / MVar).

[0086] 3) The traditional generator sets and renewable energy units within the virtual power plant should meet the following output range constraints. Furthermore, the method for describing the uncertainty of renewable energy output in the virtual power plant optimization layer is exactly the same as that in the distribution network optimization layer, and will not be repeated here.

[0087]

[0088] In the formula, This means that the actual active power output of the traditional generator sets inside the virtual power plant j must be within its minimum and maximum allowable ranges. This means that the actual reactive power output of the traditional generator sets inside the virtual power plant j must be within its minimum and maximum allowable ranges. This means that the total active and reactive power (including reserve capacity) of the traditional generator sets inside the virtual power plant j does not exceed its apparent power capacity. This means that the reserve capacity of virtual power plant j during time period t must be within its allowable range; This means that the wind power output of the virtual power plant j in time period t must be within its predicted output range. This means that the photovoltaic and wind power output of the virtual power plant j during time period t must be within its predicted output range.

[0089] III. Solving the model based on SA-GA-RL

[0090] In the above optimization model, the distribution network optimization layer is a typical convex optimization problem, which is solved using the commercial software CPLEX. However, for the virtual power plant optimization layer, since the revenue of the virtual power plant is affected by the distribution network clearing result, there is no explicit analytical expression between the decision variables and the objective function value. Therefore, a hybrid algorithm, SA-GA-RL, is used to solve it. This method integrates the advantages of simulated annealing, genetic algorithms, and reinforcement learning algorithms. In the early stage, the global search capability of the SA-GA algorithm is used to quickly find feasible solutions, and in the later stage, the local fine-tuning capability of the RL algorithm is used to optimize the strategy details. After the genetic algorithm converges to a near-optimal solution, the deep deterministic policy gradient (DDPG) algorithm in reinforcement learning is used to dynamically adjust the bidding strategy. Its reward function is designed as follows:

[0091] R t =0.6·Profit t -0.3·CV a R 0.95 +0.1·σ min (ΔP loss )

[0092] In the formula, R t The reward function value; Profit t Economic benefit (profit) is the difference between the product of the awarded electricity volume and the bid price, and the generation cost; CV a R 0.95 For risk; σ min (ΔP loss ) represents grid stability, which is the power imbalance in the grid caused by bidding strategies.

[0093] In the formula, 0.6 represents the profit weight. In the reward function, the profit at the current moment is given a relatively large weight (60%). This means that the algorithm will prioritize maximizing profit during the optimization process.

[0094] In the formula, 0.3 represents the risk control weight, and in the reward function, the conditional value at risk (CVaR) is given a negative weight (30%). This means that the algorithm will try to reduce potential high-risk losses during the optimization process, that is, while pursuing high returns, it should also pay attention to risk management.

[0095] In the formula, 0.1 represents the stability weight. In the reward function, the fluctuation of power loss is given a small positive weight (10%). This means that the algorithm also takes into account the stability of the system during the optimization process, and tries to minimize the fluctuation of power loss.

[0096] This reward function focuses on the operational stability of the VPP itself, improving market credibility by reducing output fluctuations (e.g., avoiding penalties for deviations). If the VPP's wind power forecasting error is large, this reward function will preferentially reduce σ. min (ΔP loss This reduces output jumps.

[0097] The model optimization solution process based on SA-GA-RL, such as Figure 2 As shown, see the steps below for details.

[0098] Step 1: Input the distribution network, virtual power plant internal unit and topology parameters, and renewable energy forecast information.

[0099] Step 2: Generate the initial population and initialize the virtual power plant bidding information, including the types of market products provided by the virtual power plant from each node, the bid price, and the output range.

[0100] Step 3: The distribution network optimizes its operation based on the bidding information of the virtual power plant to obtain the clearing price and winning bid volume of the virtual power plant at each node.

[0101] Step 4: The virtual power plant performs internal optimization based on the clearing information released by the distribution network to obtain its own optimal benefit and calculate the population fitness.

[0102] Step 5: Update the optimization variables for each individual in the population to form a new virtual power plant bidding plan, and then proceed to Step 3 for iteration until the following iteration termination condition is met:

[0103]

[0104] Where: Φ VPP (l) represents the optimal fitness in the l-th generation population, $; $ represents the iteration termination error; K is the pre-set convergence assessment step size, where K represents the range of iterations considered when calculating convergence, and convergence is determined by the change in profit of the most recent K iterations; k represents the index of the current iteration number, marking that the algorithm has run to the kth step.

[0105] Step 6: Load the optimal solution of SA-GA into the RL algorithm as the initial policy.

[0106] Step 7: Initialize DDPG network parameters:

[0107] 1) State Space: Define the state space of the optimization problem, including the bid prices and output ranges of the virtual power plant at each node for active, reactive and reserve power, as well as the market environment and competitors' pricing strategies.

[0108] The state space design ensures that the RL algorithm can perceive market dynamics, VPP resource status, and grid constraints, providing comprehensive input for strategy adjustment.

[0109] For example, when the energy storage SOC is low, RL will reduce electricity sales to preserve backup capacity. The state space is the foundation for RL algorithms to perceive the environment and must comprehensively reflect the operating status of the VPP and the market environment.

[0110] In VPP bidding, the state space includes the following key elements:

[0111] ① Market environment:

[0112] Real-time market electricity price p t market This directly affects the bidding revenue of VPPs and requires real-time monitoring of market price fluctuations.

[0113] Load demand Dt: determines the output demand of VPP and affects the range of electricity to be bid.

[0114] Renewable energy output forecast Uncertainty in wind and solar power output is a source of risk and needs to be incorporated into the status quo to optimize robustness.

[0115] ②VPP internal state:

[0116] Energy storage SOCEt: The charge and discharge status of energy storage affects the dispatchable capacity and determines the flexibility of bidding.

[0117] Traditional unit output The operating costs and output limitations of traditional generators need to be considered in the bidding process.

[0118] Reserve capacity Rt: Participation in the reserve market requires consideration of reserve capacity and call probability.

[0119] ③ Distribution network feedback:

[0120] Node clearing price This reflects the market's acceptance of VPP bids.

[0121] Winning bid electricity This directly impacts the actual returns of the VPP.

[0122] Line safety margin σt: measures grid stability and avoids line overload caused by bidding strategies.

[0123] 2) Action Space: Define the action space of the optimization problem, including adjusting the bid price and output range of the active, reactive and reserve power of the virtual power plant at each node.

[0124] The action space is directly mapped to the bidding behavior of VPPs. For example, during peak electricity price periods (real-time market electricity price p...t market (Higher), RL may increase p t bid To increase revenue; when renewable energy output is sufficient, lower bids to improve the success rate. The action space corresponds to the VPP's bidding strategy and needs to cover all market participation dimensions. Bidding strategy vector a t Represented as:

[0125]

[0126] Where, p t bid The active power bid: the price declared per unit of electricity (yuan / MWh), directly affects the probability of winning the bid.

[0127] q t bid Reactive power pricing: Participation strategy in the ancillary services market (RMB / MVarh).

[0128] r t bid For standby pricing: Standby capacity price (RMB / MW), which needs to be optimized in conjunction with the probability of standby call-up.

[0129] 3) Reward function R t A reward function is designed based on the optimization objective to evaluate the effectiveness of the strategy. The reward function can include the total revenue of the virtual power plant, market share, and competitors' pricing strategies. The reward function guides the RL algorithm to pursue high returns while avoiding high-risk and unstable strategies.

[0130] For example, if a bidding process leads to line overload (ΔP) loss If the reward function value is greater than 0, the reward function will decrease significantly, prompting subsequent policy adjustments. The reward function is the driving force of RL learning and needs to be considered in conjunction with economic benefits, risk control, and grid stability.

[0131] R t =α·Profit t -β·Risk t +γ·Stability t

[0132] Here, α (reward weight) represents the importance of the VPP's economic returns in the total reward. For example, if α is large, the RL algorithm will be more inclined to increase the bid to pursue higher profits; if α is small, it may sacrifice some returns to reduce risk or improve stability.

[0133] Wherein, β (risk weight): represents the weight of risk penalty terms (such as Conditional Value at Risk, CVaR). The larger the β, the more strictly the algorithm avoids market volatility or uncertainty in renewable energy output.

[0134] Wherein, γ (stability weight): represents the weight of grid stability (such as line power balance and voltage security). When γ is high, RL will prioritize avoiding line overload or voltage exceeding limits caused by bidding strategies.

[0135] α+β+γ is usually normalized to 1, reflecting the trade-offs among multiple objectives. As a hyperparameter, it needs to be manually set or optimized through parameter tuning, and directly affects the strategy preference (economic benefits vs. risk vs. stability). The weight of economic benefits, risk control and grid stability is: α>β>γ. Typical values ​​are: α=0.6 (profit priority), β=0.3 (risk control), γ=0.1 (stability guarantee).

[0136] Benefits:

[0137] Winning bid volume (q) t clear ) and quote (p t bid The product of ( ) and the power generation cost (Cgen) is deducted.

[0138] Risk item: t =CVaR α (Loss)

[0139] Conditional Value at Risk (CVaR) quantifies the risk of loss from extreme market volatility or insufficient renewable energy output.

[0140] Stability term: Stability t =-log(max(ΔP) loss ,0))

[0141] Penalty for grid power imbalance (ΔP) caused by bidding strategy loss ), such as line overload.

[0142] This reward function prioritizes the physical security of the power grid, ensuring that the bidding strategy does not trigger grid failures (such as congestion management). If VPP bidding causes a power flow exceeding limits on a certain line, this reward function will directly penalize ΔP. loss Forced adjustment of bid volume.

[0143] The difference between the two reward functions mentioned above is essentially due to the different optimization levels:

[0144] VPP-level stability: Optimize internal resource scheduling and reduce interaction deviations with the market; Grid-level stability: Ensure that market clearing results do not conflict with grid physical limitations.

[0145] This layered design reflects the need for synergistic optimization and overall security in the electricity market.

[0146] 4) RL Algorithm Selection: Deep Deterministic Policy Gradient (DDPG)

[0147] DDPG can handle continuous variables (such as electricity price and volume) in VPP bidding and learn patterns from historical market data through offline experience replay (ReplayBuffer). For example, the Actor network is responsible for learning deterministic policies, directly outputting action values ​​given a state, and can learn to increase bids during peak load periods, while the Critic network is responsible for learning state-value functions, evaluating the value of the current state, and assessing the long-term benefits of this strategy.

[0148] Actor Network: Outputs a deterministic action a t =μ(s) t |θ μ For example, a bid vector can be used to directly generate a bidding strategy. t For state s t The output is a deterministic action, such as a bid vector. μ is the policy function of the Actor network, used to output the action. t This represents the current state, which could include information such as peak load. θ μ These are the weight parameters of the Actor network.

[0149] Critic Network: Evaluating the Q-value Q(s) t a t |θ Q The Critic network evaluates the long-term value (Q-value) of actions to guide the Actor's optimization strategy. Q is the evaluation function of the Critic network, used to assess the long-term value (Q-value) of actions. t This represents the current state. t For action. θ Q These are the weight parameters of the Critic network.

[0150] Target network update: Soft update parameters θ′←τθ+(1-τ)θ′, stabilizing the training process through soft updates. θ′ represents the updated weight parameters of the target network. τ is the soft update coefficient, used to smoothly update the target network parameters. θ represents the weight parameters of the target network to be updated.

[0151] ·θ μ (Actor parameter): Used to generate deterministic bidding strategies (such as active bid p). t bid reactive power quotation q t bid ).

[0152] Input status (market electricity price, energy storage SOC, etc.), output action (bidding strategy).

[0153] ·θ Q(Critic parameter): Used to evaluate the long-term value (Q value) of an action and guide the actor's optimization strategy.

[0154] Input state and action, output expected cumulative reward.

[0155] 5) Key formulas and strategies updated:

[0156] Actor updates via policy gradient:

[0157]

[0158] Parameter meaning:

[0159] Regarding the Actor network parameter θ μ The gradient of the performance index J. The expectation operator represents the average over all samples. The Q-value function Q(s,a|θ) for action a Q The gradient of ). a=μ(s) : This indicates the case where action a takes the value of policy μ(s). Regarding the Actor network parameter θ μ strategy μ(s|θ) μ The gradient of ).

[0160] This formula is used to update the parameters θ of the Actor network. μ It calculates the performance index J with respect to θ. μ This is achieved by calculating the gradient of the Q-value function. Specifically, it first calculates the gradient of the Q-value function with respect to action a, then calculates the gradient when action a takes the value of policy μ(s), and finally multiplies it by the policy μ(s|θ). μ Regarding θ μ The gradient is calculated. The purpose of this is to enable the Actor network to learn action policies that maximize the Q-value.

[0161] Critic updates by minimizing the time-series difference error:

[0162] y t =r t +γQ′(s t+1 ,μ′(s t+1 |θ μ′ )|θ Q′ )

[0163]

[0164] Parameter meaning:

[0165] y t : The target Q value at step t. t: The immediate reward obtained at step t. γ: A discount factor used to balance immediate and future rewards. Q′(s) t+1 ,μ′(s t+1 |θ μ′ θ|θ Q′ ): In the next state s t+1 Below, the Q-value is calculated using the target Actor network policy μ′. L: Loss function, used to measure the difference between the predicted Q-value and the true Q-value. N: Number of samples. Q(s) t ,a t |θ Q ): In state s at step t t and action a t Below, using the Critic network parameter θ Q The calculated Q value.

[0166] This formula is used to update the parameter θ of the Critic network. Q It achieves this by minimizing the timing difference error. Specifically, it first calculates the target Q-value y at step t. t Then, the loss function L is calculated, which is the average of the sum of squares of the differences between the predicted Q-value and the true Q-value. Finally, the parameters θ of the Critic network are updated using the backpropagation algorithm. Q This is done to reduce the value of the loss function. The goal is to enable the Critic network to accurately estimate the value of actions, thereby guiding the Actor network to learn better policies.

[0167] 6) Training process and parameter settings:

[0168] By using OU processes (Ornstein-Uhlenbeck noise) to add random perturbations to the action space, a balance is struck between exploring new strategies and exploiting known high-yield strategies.

[0169] The key parameters are as follows:

[0170] Actor network learning rate: 10 -4 (To avoid policy mutation); Critic network learning rate: 10 -3 (Fast convergence Q-value); Discount factor: 0.99 (emphasizing long-term returns); Replay buffer size: 10 6 (Covering diverse market scenarios.)

[0171] Parameter settings directly affect learning efficiency and stability. For example, a larger playback buffer (10...) 6 It can store diverse market scenarios (such as a sharp drop in electricity prices or a sudden decrease in wind and solar power output), helping RL algorithms generalize to unknown scenarios.

[0172] Step 8: In the optimized environment, interact with the current policy and observe state transitions and reward signals.

[0173] Step 9: Store each experience in the replay buffer.

[0174] Step 10: Update the policy parameters and optimize the policy details based on the reward signal and value function. Update the Actor / Critic network.

[0175] Step 11: Repeat the interactive learning and policy update until the termination conditions are met, such as reaching the maximum number of iterations or policy convergence.

[0176] Step 12: Output the optimal result.

[0177] Building upon the near-optimal solution output by SA-GA, RL further optimizes the strategy through interactive learning. For example, SA-GA might present a high-return but high-risk policy, which RL can mitigate by dynamically adjusting the risk exposure. The final RL policy not only pursues theoretical optimality but also needs to be validated through practical constraints. For instance, a policy might have high returns during training but lead to insufficient backup capacity; the backup bid (rtbid) needs to be adjusted before deployment to meet system requirements.

[0178] This invention balances the interests of both virtual power plant and distribution network owners, considering the various flexible resource characteristics and system network topology within the virtual power plant. It proposes a two-layer optimization model for the interaction between the virtual power plant and the distribution network in a distribution-side retail market environment, and utilizes the SA-GA-RL algorithm to optimize and solve this two-layer model. SA-GA-RL, by introducing RL (a policy learning method that optimizes decision-making through trial and error), retains the global leap advantage of SA and the parallel search advantage of GA, while enhancing dynamic decision-making, noise resistance, and complex problem handling capabilities, thus improving the dynamic adaptability and convergence efficiency of the policy.

[0179] The core value of RL in VPP bidding lies in:

[0180] 1) Dynamically adapt to market changes: Adjust bidding strategies by sensing electricity prices, load and renewable energy output in real time;

[0181] 2) Multi-objective trade-offs: Achieving a balance between returns, risks, and stability to avoid the drawbacks of optimizing a single objective;

[0182] 3) Learning from data: Utilize historical market data and simulated environments to extract complex market patterns, overcoming the limitations of traditional optimization models.

[0183] RL upgrades VPP bidding from static optimization to a dynamic learning process, significantly improving the strategy's economy and robustness. The proposed method can assist virtual power plants in participating in the active, reactive, and reserve markets when acting as agents. The algorithm used in this scheme can effectively solve the model, providing guidance for the formulation of virtual power plant bidding strategies.

[0184] An implementation method for a bidding optimization system for virtual power plants:

[0185] A bidding optimization system for a virtual power plant includes a processor, which executes a computer program to implement the steps of a bidding optimization method for a virtual power plant. The specific steps of the bidding optimization method for a virtual power plant have been described in detail in an implementation of the method and will not be repeated here.

Claims

1. A bidding optimization method for virtual power plants, characterized in that, Includes the following steps: The known parameters related to the two-level bidding optimization model of virtual power plants and distribution networks are obtained and input into the model. The optimal bidding strategy of virtual power plants is obtained by iterative calculation using the model. The two-layer bidding optimization model includes a virtual power plant optimization layer and a distribution network optimization layer; The virtual power plant optimization layer is used to solve the bidding strategy of the virtual power plant based on the objective function of the virtual power plant optimization layer and the clearing results fed back by the distribution network optimization layer, and then feed it back to the distribution network optimization layer. The process of obtaining the bidding strategy for the virtual power plant includes: firstly, using a heuristic search algorithm to perform a first-stage solution, and then using the first-stage solution result as the initial strategy for a reinforcement learning algorithm to perform a second-stage solution. The distribution network optimization layer is used to solve for the clearing result of the distribution network based on the objective function of the distribution network optimization layer and the bidding strategy fed back by the virtual power plant optimization layer, and then feed it back to the virtual power plant optimization layer.

2. The bidding optimization method for virtual power plants according to claim 1, characterized in that, The reinforcement learning algorithm employs a deep deterministic policy gradient.

3. The bidding optimization method for virtual power plants according to claim 2, characterized in that, The reward function of the deep deterministic strategy gradient is set to balance economic benefits, risks, and grid stability. The economic benefits refer to the economic benefits of the virtual power plant, and the risks include the risk of loss due to extreme market fluctuations or insufficient output of renewable energy.

4. The bidding optimization method for virtual power plants according to claim 3, characterized in that, The reward function is expressed as follows: R t =a·Profit t -b·Risk t +γ·Stability t α+β+γ=1 In the formula, R t Let α be the reward function value, β be the weight of economic returns, γ be the weight of risk, and γ be the weight of grid stability. t The economic benefit is the difference between the product of the winning bid volume and the bid price, and the generation cost. t For the aforementioned risk, Stability t The grid stability refers to the grid power imbalance caused by bidding strategies.

5. The bidding optimization method for virtual power plants according to claim 4, characterized in that, α>β>γ.

6. The bidding optimization method for virtual power plants according to claim 1, characterized in that, The heuristic search algorithm is a simulated annealing-genetic algorithm, which is obtained by improving the genetic algorithm. The improvement is that after performing the genetic operation, a simulated annealing operation is added, and then the determination of whether the iteration termination condition of the genetic algorithm is met is performed.

7. The bidding optimization method for virtual power plants according to claim 1, characterized in that, The clearing results include the clearing price and winning bid volume of the virtual power plant at each node.

8. The bidding optimization method for virtual power plants according to claim 1, characterized in that, The bidding strategy for virtual power plants includes the active, reactive, and reserve output prices and output range of the virtual power plant at each node.

9. A bidding optimization system for a virtual power plant, comprising a processor, characterized in that, The processor is used to execute a computer program to implement the steps of the bidding optimization method for a virtual power plant as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • A virtual power plant bidding strategy interacting with power distribution side multivariate retail market

    CN111222917A