Virtual power plant resource allocation method, device and equipment

By constructing an initial reinforcement model and optimizing network parameters and value functions using historical state observations and action quantities, the problem of inaccurate power allocation caused by the heterogeneity of power plant resources in virtual power plants is solved, realizing dynamic and accurate allocation of power plant resources and improving the accuracy and stability of regulation results.

CN120728729BActive Publication Date: 2026-01-27EAST CHINA BRANCH OF STATE GRID CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510628645.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2026-01-27
Estimated Expiration
2045-05-15

AI Technical Summary

Technical Problem

Existing virtual power plant resource allocation models cannot effectively cope with the heterogeneity of frequency regulation rates, accuracy, and response delays among different units, resulting in inaccurate power allocation regulation results.

Method used

By constructing an initial reinforcement model, utilizing historical state observations and action quantities of resources from various power plants, optimizing network parameters and value functions, predicting current action quantities, and achieving dynamic allocation of power plant resources, the allocation strategy is optimized by combining comprehensive frequency regulation performance evaluation indicators and net frequency regulation revenue as reward functions.

Benefits of technology

It improves the accuracy of power allocation and adjustment results in virtual power plants, achieves good results in complex environments, reduces instability during training, and ensures the rationality and stability of allocation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120728729B_ABST
    Figure CN120728729B_ABST
Patent Text Reader

Abstract

The application discloses a virtual power plant resource allocation method, device and equipment, and relates to the technical field of electric power resources, aiming at the problem of low accuracy of power allocation adjustment result caused by internal frequency modulation resource heterogeneity of a VPP. The method comprises the following steps: determining historical state observation and historical action of each power plant resource in a virtual power plant; the virtual power plant comprises multiple power plant resources, i.e. a gas turbine, a wind turbine, a photovoltaic turbine, an energy storage battery and an electric vehicle. According to the historical state observation and the historical action, the last actual observation, the last actual action and the current action under the current state observation are predicted; the historical action represents the resource scheduling allocation result of the each power plant resource under the historical state observation. According to the current action, the resource scheduling allocation result of the each power plant resource is represented, and the operation power of the each power plant resource is controlled and adjusted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power resource technology, and in particular to a method, apparatus and equipment for allocating virtual power plant resources. Background Technology

[0002] Driven by the "dual-carbon" strategic goals, the power system is accelerating its transformation towards high-proportion renewable energy integration and high intelligence. Virtual power plants (VPPs) participating in AGC (automatic generation control) frequency regulation ancillary services can fully leverage their resource aggregation and collaborative optimization advantages, improving frequency regulation resource utilization efficiency and reducing system frequency regulation costs. However, due to the heterogeneity in the adjustment rate, accuracy, and response delay among various units within a VPP, existing fixed-proportion allocation models cannot effectively address the impact of these differences, resulting in inaccurate VPP power allocation regulation results. Summary of the Invention

[0003] This invention provides a virtual power plant resource allocation method, resource allocation device, and equipment to at least solve the problem of low accuracy in power allocation adjustment results caused by the heterogeneity of frequency regulation resources within a virtual power plant (VPP). The technical solution of this invention is as follows:

[0004] According to a first aspect of the present invention, a virtual power plant resource allocation method is provided. The method includes: determining historical state observations and historical action quantities for each power plant resource in the virtual power plant. The virtual power plant includes multiple power plant resources such as gas turbines, wind turbines, photovoltaic units, energy storage batteries, and electric vehicles. Based on the historical state observations and historical action quantities, the method predicts the current action quantity under the previous actual observation, the previous actual action, and the current state observation. The historical action quantity represents the resource scheduling and allocation result for each power plant resource under the historical state observation. Based on the current action quantity representing the resource scheduling and allocation result for each power plant resource, the method controls and adjusts the operating power of each power plant resource.

[0005] In this implementation, the previous actual observation refers to the power plant resource status observed before the current time of the current state observation and the closest acquisition time to the current time; the previous actual action refers to the frequency modulation action acquired before the current time of the current state observation and the closest acquisition time to the current time.

[0006] In one implementation, based on historical state observations and historical action quantities, the current action quantity under the previous actual observation, the previous actual action, and the current state observation is predicted. This includes: constructing an initial reinforcement model that characterizes the mapping relationship between state observations and action quantities, using the weighted sum of the comprehensive frequency regulation performance evaluation index of each power plant resource and the net frequency regulation revenue as the reward function of the initial reinforcement model. The network parameters and value function of the initial reinforcement model are optimized based on the historical state observations and historical action quantities to obtain the target reinforcement model. The previous actual observation, the previous actual action, and the current state observation are input into the target reinforcement model to obtain the current action. The historical action quantities and current action quantities include the same number of actions as the number of power plant resources included in the virtual power plant.

[0007] In this implementation, the resource allocation method divides the reinforcement model update into two stages, which allows for more efficient use of the collected empirical data. Utilizing updates to network parameters and the value function helps reduce instability during training, enabling the allocation results to achieve better performance even in complex environments.

[0008] In another implementation, before constructing an initial reinforcement model that characterizes the mapping relationship between state observations and action quantities, using the comprehensive frequency regulation performance evaluation index of each power plant resource and the net frequency regulation revenue as the reward function of the initial reinforcement model, the method further includes: determining the net frequency regulation revenue of the virtual power plant based on the total frequency regulation revenue of the virtual power plant and the frequency regulation cost of each power plant resource; and determining the comprehensive frequency regulation performance evaluation index for the regulation of each power plant resource according to the regulation rate, regulation accuracy, and response time of the scheduling and regulation of each power plant resource.

[0009] In this implementation, the frequency regulation performance evaluation indicators and net frequency regulation revenue of each virtual power plant resource are integrated as the reward function, ensuring the rationality of the allocation strategy training behavior.

[0010] In another implementation, the virtual power plant resource allocation method further includes: collecting historical state observations and historical action quantities according to a preset period; determining the historical state observations for the next preset period corresponding to each historical state observation and each historical action quantity; solving for the comprehensive frequency regulation performance evaluation index and net frequency regulation revenue generated by each historical state observation and each historical action quantity up to the historical state observations for the next preset period according to the reward function, and obtaining the corresponding historical reward; and using each historical state observation, the corresponding historical action quantity, the corresponding historical state observations for the next preset period, and the corresponding historical reward as a set of training data in the experience pool.

[0011] In this implementation, the parameters of the allocation strategy are continuously adjusted and the structure of the algorithm is optimized through iterative updates, so that the allocation results are more reasonable.

[0012] In another implementation, the virtual power plant resource allocation method also includes: determining the probability of occurrence of each historical action quantity under each historical state observation based on the training data of each group in the experience pool.

[0013] In another implementation, the virtual power plant resource allocation method also includes: according to the convergence of the historical rewards corresponding to different historical action quantities of two historical state observations corresponding to adjacent preset periods, the training data of each group in the experience pool is screened to obtain the screened training data.

[0014] In this implementation, historical rewards are ensured to converge during training, guaranteeing that the model adopts effective strategies to maximize long-term rewards in a stable environment, thus achieving a relatively stable training result.

[0015] In another implementation, the function variables of the value function include network parameters, used to evaluate the network parameters of the initial reinforcement model. The network parameters and value function of the initial reinforcement model are optimized according to historical state observations and historical action quantities to obtain the target reinforcement model. This includes: inputting each group of training data from the experience pool into the initial reinforcement model to obtain the output action quantities for each group; the number of output action quantities includes actions equal to the number of power plant resources included in the virtual power plant. For any group of output action quantities, the model adjustment reward for the historical state observations in the next preset period is adjusted according to the reward function based on the input historical action quantities. Also, the difference between the first reward of the output action quantity and the second reward of the corresponding historical action quantity is determined. Based on the preset network objective, the network parameters and value function are optimized according to the model adjustment reward, the difference in rewards, and the probability of occurrence of each group of output action quantities.

[0016] In this implementation, the network parameters and value function are optimized by taking into account the difference in revenue from power plant resource actions and the probability of occurrence, which will further improve the performance and stability of the allocation results.

[0017] In another implementation, historical and current state observations include: gas turbine regulation capacity, wind turbine regulation capacity, photovoltaic unit regulation capacity, energy storage battery capacity, electric vehicle state-of-charge capacity, energy storage battery charge / discharge rate, electric vehicle charge / discharge rate, and the winning bid price corresponding to the frequency regulation capacity for each time period.

[0018] According to a second aspect of the present invention, a virtual power plant resource allocation apparatus includes:

[0019] The determination unit is configured to determine the historical state observations and historical action quantities of each power plant resource in the virtual power plant; the virtual power plant includes multiple of the following power plant resources: gas turbines, wind turbines, photovoltaic units, energy storage batteries, and electric vehicles.

[0020] The prediction unit is configured to predict the current action quantity under the previous actual observation, the previous actual action, and the current state observation based on historical state observations and historical action quantities; the historical action quantities represent the resource scheduling and allocation results of each power plant resource under the historical state observations.

[0021] The regulating unit is configured to control and regulate the operating power of each power plant resource according to the resource scheduling and allocation results of each power plant resource represented by the current action quantity.

[0022] According to a third aspect of the present invention, a power resource allocation device is configured to perform a virtual power plant resource allocation method as described in the first aspect and any possible implementation thereof.

[0023] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which instructions are stored, such that when the instructions in the computer-readable storage medium are executed by a processor of a computer device, the computer device is able to perform a virtual power plant resource allocation method as described in the first aspect and any possible implementation thereof.

[0024] According to a fifth aspect of the embodiments of this application, a computer program product is provided, the computer program product including computer instructions, which, when executed on a computer device, cause the computer device to perform the virtual power plant resource allocation method of the first aspect and any possible implementation thereof.

[0025] The technical solutions provided by the embodiments of the present invention bring at least the following beneficial effects:

[0026] By utilizing historical state observations (i.e., different power plant resource states) and historical action quantities (i.e., different frequency regulation actions) of various power plant resources, the interaction between changes in different allocation actions and different power plant resource states can be fully reflected. Based on this, by using the relationship between power plant resource states and resource allocation actions represented by the aforementioned historical state observations and historical action quantities, accurate predictions can be made of the transition from the previous actual action under the previous actual observation to the current resource allocation action under the current state observation, thus achieving dynamic allocation of power plant resources. Furthermore, based on the interaction between the action changes and power plant resource state changes represented by historical state observations and historical action quantities, and taking into account the impact of the heterogeneity of power plant resources on the allocation results during the resource allocation process, dynamic and accurate allocation of power plant resources is achieved, improving the accuracy of VPP power allocation regulation results.

[0027] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0028] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application, and do not constitute an undue limitation of this application.

[0029] Figure 1 This is a flowchart illustrating a virtual power plant resource allocation method according to an exemplary embodiment. Figure 1 ;

[0030] Figure 2 This is a flowchart illustrating a virtual power plant resource allocation method according to an exemplary embodiment. Figure 2 ;

[0031] Figure 3 This is a block diagram illustrating a virtual power plant resource allocation device according to an exemplary embodiment;

[0032] Figure 4 This is a schematic diagram of a power resource distribution device according to an exemplary embodiment. Detailed Implementation

[0033] To enable those skilled in the art to better understand the technical solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0034] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0035] Before providing a detailed description of the virtual power plant resource allocation method provided in the embodiments of this application, let's briefly introduce the application scenarios involved in the embodiments of this application.

[0036] Currently, the main AGC power command allocation strategy adopts the PRPO (proportional, fixed ratio) allocation strategy, where the control entity allocates AGC power commands to each frequency regulation unit based on the adjustable reserve capacity ratio. However, this strategy has many limitations. On the one hand, the PRPO strategy cannot respond to real-time changes in system demand in a timely manner, easily leading to problems such as insufficient regulation accuracy, resource waste, and equipment wear and tear. On the other hand, due to the significant heterogeneity of frequency regulation resources within the VPP in terms of regulation rate, accuracy, and response delay, it is difficult to effectively coordinate these differences, posing a significant challenge to AGC power command allocation.

[0037] Research has found that optimizing AGC frequency regulation and maximizing the expected net benefit of the VPP are the optimal goals. An improved genetic algorithm is used to configure the capacity and power of the hybrid energy storage system and optimize the scheduling of the output of each part of the VPP. Another approach involves an improved quantum genetic algorithm-based strategy where the VPP participates in AGC optimal scheduling across multiple time scales to maximize the overall benefit of the VPP in AGC frequency regulation and achieve the best comprehensive frequency regulation effect. Yet another approach aims to maximize VPP benefits by using a branch-and-bound algorithm to solve the optimal strategy in a two-stage optimal scheduling model of virtual power plants participating in frequency regulation ancillary services.

[0038] However, these methods still have shortcomings in handling market uncertainty, practical operational constraints, and user response behavior. Furthermore, the solution efficiency and robustness of the algorithms used need further improvement. When facing complex operating conditions, current optimization allocation methods are prone to getting trapped in local optima, making it difficult to achieve global optimization. The randomness of system load disturbances further exacerbates the complexity of frequency regulation, while the power grid has extremely high real-time requirements for AGC (Automatic Generation Control). Traditional power command allocation strategies are significantly inadequate in terms of rapid response and adaptation to dynamic changes.

[0039] The above findings reveal the following issues: While using VPP to participate in AGC frequency regulation auxiliary services allows for flexible control of the power generation of each unit and reasonable allocation of the total power command issued by the superior to each power output unit, the traditional fixed-ratio allocation cannot effectively address the impact of these differences due to the heterogeneity in the adjustment rate, accuracy, and response delay among the various units in the VPP. For example, if a power output unit in the virtual power plant experiences a problem and cannot output power in a timely manner, the traditional fixed-ratio allocation strategy cannot respond quickly to this change, resulting in insufficient power supply and other problems.

[0040] To address the aforementioned issues, this application provides a virtual power plant resource allocation method. It uses net frequency regulation revenue and comprehensive frequency regulation performance evaluation indicators as the reward function of the initial reinforcement model, optimizes the network parameters and value function of the initial reinforcement model to obtain the target reinforcement model, determines historical state observations and historical actions, and inputs the previous actual observation, the previous actual action, and the current state observation into the target reinforcement model to predict the current action. Based on the current action, the resource scheduling allocation results for each power plant are determined, and the operating power of each power plant is controlled and regulated.

[0041] The relevant terms used in this application are explained as follows.

[0042] Heterogeneity refers to the differences that exist between individuals, groups, things, etc.

[0043] Virtual Power Plant (VPP) participation in AGC (Automatic Generation Control) frequency regulation ancillary services refers to the involvement of virtual power plants in the power system's automatic generation control frequency regulation ancillary services. Leveraging their flexible control over distributed resources, virtual power plants can rapidly adjust the generation capacity or load size of their aggregated resources based on changes in grid frequency. When the grid frequency decreases, virtual power plants can increase generation capacity or decrease controllable load; conversely, when the grid frequency increases, virtual power plants can decrease generation capacity or increase controllable load. This helps the grid maintain frequency stability, improves grid reliability and stability, and simultaneously generates economic benefits from providing this ancillary service.

[0044] The power command allocation strategy determines how to rationally allocate these power adjustment commands to the various power generation units participating in AGC regulation.

[0045] For ease of understanding, the virtual power plant resource allocation method provided in this application will be described in detail below with reference to the accompanying drawings.

[0046] Figure 1 This is a flowchart illustrating a virtual power plant resource allocation method according to an exemplary embodiment, such as... Figure 1 As shown, the virtual power plant resource allocation method includes the following steps.

[0047] S11, determine the historical state observations and historical action quantities of each power plant resource in the virtual power plant.

[0048] Historical state observations refer to the different resource states of each power plant during historical monitoring periods.

[0049] Historical action data refers to the frequency regulation actions of various power plant resources during historical monitoring periods.

[0050] The virtual power plant includes multiple gas turbines, wind turbines, photovoltaic units, energy storage batteries, and electric vehicles from the following power plant resources.

[0051] Historical state observations include AGC regulation capacity upper and lower limit datasets for gas turbines, wind turbines, and photovoltaic units, as well as SOC state of charge and capacity upper and lower limit datasets for energy storage batteries and electric vehicles.

[0052] Historical action quantities represent the resource scheduling and allocation results of each power plant under the historical state observation.

[0053] Historical state observations and historical action quantities of resources at each power plant can fully reflect the interaction between different resource allocation actions and the states of resources at different power plants. Furthermore, the power plant resource characteristics characterized by the aforementioned historical state observations can be used to reflect the differences in changes in power plant resource characteristics based on changes in state observations, i.e., the heterogeneous impact of resources.

[0054] In one embodiment, the process of determining historical state observations involves considering the VPP real-time AGC power command signal. The bid price for the corresponding frequency modulation capacity, the upper and lower limits of the state adjustment capacity of each frequency modulation unit, and the design of the intelligent agent state observation space are detailed in the following formula (1).

[0055] (1).

[0056] For state observation space; This is the VPP real-time AGC power command signal; For VPP number The winning bid price for the next requested frequency modulation capacity; , , These are the minimum AGC regulation capacities for gas turbines, wind turbines, and photovoltaic units, respectively. , , These are the maximum AGC regulation capacity values ​​for gas turbines, wind turbines, and photovoltaic units, respectively. The state of charge of the energy storage battery at time t; Maximum regulating capacity of electric vehicles Minimum regulating capacity for electric vehicles.

[0057] VPP participated in frequency modulation and was called upon to generate a multi-day AGC power command signal dataset. Dataset of winning bid prices for frequency modulation capacity at corresponding time points The corresponding time-based unit dataset obtained from the acquired VPP-initiated AGC power command signal dataset is used as historical state observation.

[0058] The process of determining the historical action quantity is as follows: the power regulation commands allocated to the gas turbine, wind turbine, photovoltaic unit, energy storage battery and electric vehicle are used as the continuous action space, as detailed in the following formula (2).

[0059] (2).

[0060] in, For action space; , , , and VPP number The power regulation command signals allocated to the gas turbine, wind turbine, photovoltaic power station, energy storage power station and electric vehicle are invoked.

[0061] S12, based on historical state observations and historical action quantities, predict the current action quantity under the previous actual observation, the previous actual action, and the current state observation.

[0062] The previous actual observation refers to the power plant resource status observed before the current moment of the current state observation and at the closest acquisition moment to the current moment.

[0063] The previous actual action refers to the frequency modulation action acquired at the time closest to the current acquisition time of the current state observation.

[0064] In step S12 above, the process of predicting the current action quantity under the previous actual observation, the previous actual action, and the current state observation is to train the phasic policy gradient (PPG) algorithm based on historical state observations and historical action quantities, and alternately update the network parameters and value function until the reward value converges to obtain the current action quantity.

[0065] S13, based on the current action quantity representing the resource scheduling and allocation results of each power plant resource, controls and adjusts the operating power of each power plant resource.

[0066] The current action quantity is obtained by accurately predicting the relationship between the power plant resource status and resource allocation action represented by the historical state observations and historical action quantities in steps S11 and S12, and transitioning from the previous actual action under the previous actual observation to the current resource allocation action under the current state observation. The current frequency regulation status of each power plant resource is determined by the resource scheduling and allocation results of each power plant resource represented by the current action quantity, thereby adjusting the operating power of each power plant resource.

[0067] Through the above implementation methods, by utilizing historical state observations (i.e., different power plant resource states) and historical action quantities (i.e., different frequency regulation actions) of each power plant resource, the interaction between the changes in different allocation actions of each power plant resource and the different power plant resource states can be fully reflected. Based on this, by using the relationship between power plant resource states and resource allocation actions represented by the above-mentioned historical state observations and historical action quantities, accurate predictions can be made of the transition from the previous actual action under the previous actual observation to the current resource allocation action under the current state observation, thus realizing the dynamic allocation of each power plant resource. Furthermore, based on the interaction between the changes in action and changes in power plant resource states represented by historical state observations and historical action quantities, and taking into account the impact of the heterogeneity of each power plant resource on the allocation results during the resource allocation process, dynamic and accurate allocation of each power plant resource is achieved, improving the accuracy of VPP power allocation regulation results.

[0068] In one implementation, such as Figure 2 As shown, step S12 is specifically implemented through the following steps S121 to S123.

[0069] S121 uses the weighted sum of the comprehensive frequency regulation performance evaluation index of each power plant resource and the net frequency regulation revenue as the reward function of the initial reinforcement model, and constructs an initial reinforcement model that represents the mapping relationship between state observations and action quantities.

[0070] In step S121 above, the specific process for determining the reward function is as follows: Based on the total frequency regulation revenue of the virtual power plant and the frequency regulation cost of each power plant resource, determine the net frequency regulation revenue of the virtual power plant. Based on the regulation rate, regulation accuracy, and response time of the resource scheduling and regulation of each power plant, determine the comprehensive frequency regulation performance evaluation index for the resource scheduling of each power plant.

[0071] Furthermore, a multi-objective optimization model incorporating VPP frequency regulation revenue and performance, including gas turbines, renewable energy, and energy storage systems, is constructed as the reward function for the initial reinforcement model. The specific process for establishing the multi-objective optimization model is as follows.

[0072] First, determine the net frequency regulation revenue of the virtual power plant and the comprehensive frequency regulation performance evaluation index of each power plant's resource regulation.

[0073] Specifically, the net frequency regulation revenue of the virtual power plant is determined by formulas (3) to (12).

[0074] First, the total frequency regulation revenue of the virtual power plant is determined by formulas (3) to (6).

[0075] According to the virtual power plant in the The frequency modulation compensation income is determined by the bid price of the frequency modulation mileage and frequency modulation capacity provided by the next call.

[0076] (3).

[0077] (4).

[0078] in, Frequency modulation compensation income; For virtual power plants in the first The frequency modulation mileage provided by the next call is the absolute value of the difference between the actual total output value of each frequency modulation unit at the end of each frequency modulation command and the output value at the time of response to the command. For the virtual power plant The winning bid price for the next requested frequency modulation capacity; , , , and The virtual power plant is the first The actual active power output changes of the gas turbine, wind turbine, photovoltaic power station, energy storage power station and electric vehicle are called upon next; , , , and The virtual power plant is the first The power regulation command signals allocated to the gas turbine, wind turbine, photovoltaic power station, energy storage power station and electric vehicle are invoked.

[0079] Consider the losses incurred by virtual power plants in the spot market due to their participation in frequency regulation ancillary services that increase or decrease electricity volume, and establish a spot adjustment compensation fee.

[0080] (5).

[0081] in, Spot price adjustment compensation fee; For calling the time constant; To automatically adjust the compensation standard for spot market access in order to control power generation; For virtual power plants in the first The frequency modulation mileage provided by the next call.

[0082] The total frequency regulation revenue of the virtual power plant is determined based on frequency regulation compensation revenue and spot market adjustment compensation fees. .

[0083] (6).

[0084] in, Total revenue from frequency regulation of virtual power plants; Frequency modulation compensation income; Spot price adjustment compensation fee.

[0085] Secondly, the frequency regulation cost of each power plant resource is determined by formulas (7) to (11).

[0086] Considering the wear and tear on the gas turbine caused by frequent ramp-ups due to participation in AGC (Automatic Generation Control) services, determine the gas turbine frequency regulation cost.

[0087] (7).

[0088] in, For the cost of frequency regulation of gas turbines; This serves as the baseline value for the wear coefficient; The impact factor coefficient; The sampling time period; For the first The initial power value of the gas turbine is called upon next; For VPP number The change in the actual active power output of the gas turbine that was called upon.

[0089] The frequency regulation cost of energy storage batteries is determined by considering the operation and maintenance costs of energy storage batteries participating in frequency regulation and the aging costs of battery charging and discharging.

[0090] (8).

[0091] (9).

[0092] in, For the frequency regulation cost of energy storage batteries; The unit capacity operation and maintenance cost of energy storage batteries; The sampling time period; This represents the actual change in active power output of the energy storage power station. The aging coefficient of the energy storage battery; The state of charge of the energy storage battery at time t; This serves as the reference state of charge for the energy storage battery. The charging / discharging power of the energy storage battery; This refers to the rated capacity of the energy storage battery. This refers to the charge / discharge efficiency of the energy storage battery.

[0093] The cost of electric vehicle frequency regulation is determined by considering the depreciation cost of electric vehicles participating in frequency regulation.

[0094] (10).

[0095] in, For the cost of frequency modulation for electric vehicles; This is the depreciation factor; The actual change in the active power output of electric vehicles.

[0096] Determine the frequency regulation costs of each power plant's resources.

[0097] (11).

[0098] in, Frequency regulation costs for resources at various power plants; For the cost of frequency regulation of gas turbines; For the frequency regulation cost of energy storage batteries; The cost of frequency modulation for electric vehicles.

[0099] Finally, the net frequency regulation revenue of the virtual power plant is determined by formula (12).

[0100] (12).

[0101] in, Net revenue from frequency regulation of virtual power plants; Total revenue from frequency regulation of virtual power plants; Frequency regulation costs for resources at each power plant.

[0102] The comprehensive frequency regulation performance evaluation index for resource regulation of each power plant is determined by formulas (13) to (16).

[0103] The comprehensive frequency regulation performance evaluation index for each power plant resource is determined by comprehensively considering three aspects: regulation rate, regulation accuracy, and response time.

[0104] (13).

[0105] (14).

[0106] (15).

[0107] (16).

[0108] Where K is the comprehensive frequency regulation performance evaluation index of each power plant resource; , and Weighting coefficients;

[0109] , and These are the adjustment rate coefficient, adjustment accuracy coefficient, and response time coefficient, respectively. The virtual power plant's response frequency regulation control command rate; Standard adjustment rate; The frequency modulation unit adjustment error refers to the deviation between the actual output value and the control command value after the frequency modulation unit responds to the frequency modulation control command. The total rated power for the virtual power plant; The frequency modulation unit response delay time refers to the delay between the frequency modulation unit's action and receiving the frequency modulation command. This represents the maximum allowable response delay time.

[0110] Secondly, an optimization objective function model is established based on the net frequency regulation revenue of the virtual power plant and the comprehensive frequency regulation performance evaluation index of resource regulation of each power plant.

[0111] The objective function model is determined by formula (17).

[0112] (17).

[0113] in, The total number of FM tunings within the period of the official daily clearing results for the FM market; This is the net revenue model for VPP frequency modulation; Ki is the comprehensive frequency modulation performance evaluation index for VPP.

[0114] Third, based on the objective function model, considering the upper and lower limits of AGC regulation capacity and the unit ramp-up rate limit, constraints are established for gas turbines, wind turbines, and photovoltaic units to construct a multi-objective optimization model.

[0115] The unit constraints are determined by formulas (18) to (22).

[0116] (18).

[0117] (19).

[0118] (20).

[0119] (twenty one).

[0120] (twenty two).

[0121] in, , , These are the minimum AGC regulation capacities for gas turbines, wind turbines, and photovoltaic units, respectively. , , These are the maximum AGC regulation capacity values ​​for gas turbines, wind turbines, and photovoltaic units, respectively. , , These are the ramp rates for gas turbines, wind turbines, and photovoltaic units, respectively. , This represents the minimum SOC capacity for energy storage batteries and electric vehicles. , This represents the maximum SOC capacity for energy storage batteries and electric vehicles. , This represents the optimal charging / discharging rate for energy storage batteries and electric vehicles.

[0122] Fourth, the reward function of the initial reinforcement model is determined based on the multi-objective optimization model.

[0123] The reward function is determined by formulas (23) and (24).

[0124] (twenty three).

[0125] (twenty four).

[0126] in, For the reward function; This is the net frequency regulation benefit model for VPP; Ki is the comprehensive frequency regulation performance evaluation index for virtual power plants. For AGC power commands; , , , and The virtual power plant is the first The power regulation command signals allocated to the gas turbine, wind turbine, photovoltaic power station, energy storage power station and electric vehicle are called upon next; This is the VPP real-time AGC power command signal. This is the proportional amplification factor for the comprehensive frequency modulation performance evaluation index; This is the amplification factor for the penalty ratio between the action command and the AGC power command.

[0127] In this implementation, the rationality of the allocation strategy training behavior is ensured by integrating the frequency regulation performance evaluation indicators and net frequency regulation revenue of each virtual power plant resource as the reward function.

[0128] S122, based on historical state observations and historical action quantities, optimize the network parameters and value function of the initial reinforcement model to obtain the target reinforcement model.

[0129] In one embodiment, the specific process of optimizing the network parameters and value function of the initial reinforcement model to obtain the target reinforcement model is as follows.

[0130] First, the data required for the PPG algorithm iterative update is determined based on historical state observations and historical action quantities.

[0131] Optionally, historical state observations and historical action quantities are collected according to a preset period. The historical state observations for the next preset period corresponding to each historical state observation and each historical action quantity are determined. Based on the reward function, the comprehensive frequency modulation performance evaluation index and net frequency modulation revenue generated for each historical state observation and each historical action quantity up to the next preset period are solved to obtain the corresponding historical reward. Each historical state observation, its corresponding historical action quantity, the corresponding historical state observation for the next preset period, and the corresponding historical reward are used as a set of training data in the experience pool.

[0132] Specifically, one day's data is randomly selected from the historical dataset. , at t=1 As the initial state input. As input to the actor network of the PPG algorithm, the actor network perceives the state input data and outputs action values. Based on the generated action values And the reward function, calculate the reward value for this step. Execute t=t+1 to read the status data for the next AGC cycle from the historical dataset. .Will Store the data in memory pool D as experience needed for iterative training of the algorithm. Repeat the above steps until t=M.

[0133] In this implementation, the parameters of the allocation strategy are continuously adjusted and the structure of the algorithm is optimized through iterative updates, so that the allocation results are more reasonable.

[0134] Secondly, the policy network and value function are optimized during the policy update phase.

[0135] Optionally, the training data from each group in the experience pool are input into the initial reinforcement model to obtain the output action quantities for each group. The number of output action quantities is the same as the number of power plant resources included in the virtual power plant. For any group of output action quantities, the model adjustment reward for the historical state observations in the next preset period is adjusted according to the reward function based on the input historical action quantities. Furthermore, the difference between the first reward of the output action quantity and the second reward of the corresponding historical action quantity is determined. Based on the preset network objective, the network parameters and value function are optimized according to the model adjustment reward, the difference in rewards, and the occurrence probability of each group of output action quantities.

[0136] In this implementation, the network parameters and value function are optimized by taking into account the difference in revenue from power plant resource actions and the probability of occurrence, which will further improve the performance and stability of the allocation results.

[0137] Specifically, formulas (25) to (27) represent the optimization process of network parameters and value functions.

[0138] Based on any set of output action quantities in step S122 above, the reward function determines the model adjustment reward and the difference in returns for action quantities.

[0139] (25).

[0140] in, For state The following is a value estimate; Reward value.

[0141] Optionally, based on the training data from each group in the experience pool, determine the probability of occurrence of each historical action quantity under each historical state observation. See formula (26) below for details.

[0142] (26).

[0143] in, For hyperparameters; Adjust the reward for the model.

[0144] Determine the mean squared error loss of the optimized value function network.

[0145] (27).

[0146] in, The objective is the value function calculated using generalized advantage estimation.

[0147] State and corresponding Store in the buffer for use in the auxiliary phase.

[0148] Finally, an auxiliary phase update is performed, in which the shared features of the policy network are extracted through auxiliary tasks to further optimize the value function network. See formula (28) below for details.

[0149] (28).

[0150] Using the data in the buffer, Equation (27) is executed to further optimize the loss of the independent value function network.

[0151] In this implementation, based on historical state observations and historical action quantities, the interaction between the represented action changes and power plant resource state changes is input into the initial reinforcement model, which optimizes the network parameters and value function in a complex environment, enabling the model to more accurately identify power plant resource state changes and realize the dynamic allocation of resources for each power plant.

[0152] S123, input the previous actual observation, the previous actual action, and the current state observation into the target reinforcement model to obtain the current action.

[0153] In one implementation, the optimization process of step S122 is repeated to iteratively train the PPG algorithm until the reward value converges, thereby obtaining the resource scheduling and allocation results of each power plant resource.

[0154] Optionally, the historical action volume and current action volume include the same number of actions as the number of power plant resources included in the virtual power plant.

[0155] It is understandable that the historical action volume represents the scheduling and allocation results of power plant resources included in the virtual power plant, and the appropriate number of actions corresponds one-to-one with the power plant resources.

[0156] Optionally, the training data in the experience pool is filtered based on the convergence of historical rewards corresponding to different historical action quantities for two historical state observations corresponding to adjacent preset periods, so as to obtain the filtered training data.

[0157] The convergence of historical rewards during training ensures that the model adopts effective strategies to maximize long-term rewards in a stable environment, resulting in a relatively stable training outcome.

[0158] To achieve the above functions, the virtual power plant resource allocation device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art will readily recognize that, based on the algorithmic steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0159] This application embodiment also provides a method such as Figure 3 The virtual power plant resource allocation device shown includes: a determination unit 301, a prediction unit 302, and an adjustment unit 303.

[0160] Unit 301 is configured to determine the historical state observations and historical action quantities of each power plant resource in the virtual power plant. The virtual power plant includes multiple of the following power plant resources: gas turbines, wind turbines, photovoltaic units, energy storage batteries, and electric vehicles.

[0161] The prediction unit 302 is configured to predict the current action quantity under the previous actual observation, the previous actual action, and the current state observation based on historical state observations and historical action quantities. The historical action quantities represent the resource scheduling and allocation results of the various power plant resources under the historical state observations.

[0162] The adjustment unit 303 is configured to control and adjust the operating power of each power plant resource according to the resource scheduling and allocation results of each power plant resource represented by the current action quantity.

[0163] As one implementation method, the prediction unit 302 is specifically configured to predict the current action quantity under the previous actual observation, the previous actual action, and the current state observation, based on historical state observations and historical action quantities. This includes: constructing an initial reinforcement model that characterizes the mapping relationship between state observations and action quantities by using the weighted sum of the comprehensive frequency regulation performance evaluation index of each power plant resource and the net frequency regulation revenue as the reward function of the initial reinforcement model; optimizing the network parameters and value function of the initial reinforcement model based on historical state observations and historical action quantities to obtain the target reinforcement model; and inputting the previous actual observation, the previous actual action, and the current state observation into the target reinforcement model to obtain the current action. The historical action quantities and current action quantities include the same number of actions as the number of power plant resources included in the virtual power plant.

[0164] As one implementation method, the prediction unit 302 is specifically configured such that, before constructing an initial reinforcement model representing the mapping relationship between state observations and action quantities, using the comprehensive frequency regulation performance evaluation index of each power plant resource and the net frequency regulation revenue as the reward function of the initial reinforcement model, the method further includes: determining the net frequency regulation revenue of the virtual power plant based on the total frequency regulation revenue of the virtual power plant and the frequency regulation cost of each power plant resource; and determining the comprehensive frequency regulation performance evaluation index for the regulation of each power plant resource according to the regulation rate, regulation accuracy, and response time of the scheduling and regulation of each power plant resource.

[0165] As one implementation method, the prediction unit 302 is specifically configured such that the virtual power plant resource allocation method further includes: collecting historical state observations and historical action quantities according to a preset period; determining the historical state observations for the next preset period corresponding to each historical state observation and each historical action quantity; solving for the comprehensive frequency regulation performance evaluation index and net frequency regulation revenue generated from each historical state observation and each historical action quantity to the historical state observations for the next preset period according to the reward function, and obtaining the corresponding historical reward; and using each historical state observation, the corresponding historical action quantity, the corresponding historical state observations for the next preset period, and the corresponding historical reward as a set of training data in the experience pool.

[0166] As one implementation method, the prediction unit 302 is specifically configured such that the virtual power plant resource allocation method further includes: determining the probability of occurrence of each historical action quantity under each historical state observation according to the training data of each group in the experience pool.

[0167] As one implementation method, the prediction unit 302 is specifically configured such that the virtual power plant resource allocation method further includes: according to the convergence of the historical rewards corresponding to different historical action quantities of two historical state observations corresponding to adjacent preset periods, the training data of each group in the experience pool is screened to obtain the screened training data.

[0168] As one implementation method, the prediction unit 302 is specifically configured such that the function variables of the value function include network parameters, used to evaluate the network parameters of the initial reinforcement model; the network parameters and value function of the initial reinforcement model are optimized according to historical state observations and historical action quantities to obtain the target reinforcement model; this includes: inputting each group of training data from the experience pool into the initial reinforcement model in groups to obtain each group of output action quantities; the output action quantities include the same number of actions as the number of power plant resources included in the virtual power plant. For any group of output action quantities, the model adjustment reward for the historical state observations of the next preset period is adjusted according to the reward function based on the input historical action quantities. Also, the difference between the first reward of the output action quantity and the second reward of the corresponding historical action quantity is determined. Based on the preset network objective, the network parameters and value function are optimized according to the model adjustment reward, the difference in rewards, and the probability of occurrence of each group of output action quantities.

[0169] As one implementation method, the prediction unit 302 is specifically configured such that the historical state observations and current state observations include: gas turbine regulation capacity, wind turbine regulation capacity, photovoltaic unit regulation capacity, energy storage battery capacity, electric vehicle state of charge capacity, energy storage battery charging / discharging rate, electric vehicle charging / discharging rate, and the winning bid price corresponding to the frequency regulation capacity for each time period.

[0170] Figure 4 This is a schematic diagram of a power resource distribution device provided in this application. Figure 4 The allocation device 40 may include at least one processor 401 and a memory 403 for storing processor-executable instructions. The processor 401 is configured to execute the instructions in the memory 403 to implement the virtual power plant resource allocation method in the following embodiments.

[0171] In addition, the prediction device 40 may also include a communication bus 402, at least one communication interface 404, an input device 406, and an output device 405.

[0172] Processor 401 may be a processor (central processing unit, CPU), microprocessor unit, ASIC, or one or more integrated circuits for controlling the execution of programs according to the present application.

[0173] The communication bus 402 may include a path for transmitting information between the aforementioned components.

[0174] Communication interface 404 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0175] Input device 406 is used to receive input signals and output device 605 is used to output signals.

[0176] Memory 403 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. Memory may exist independently and be connected to the processing unit via a bus. Memory may also be integrated with the processing unit.

[0177] The memory 403 stores instructions for executing the scheme of this application, and the processor 401 controls the execution. The processor 401 executes the instructions stored in the memory 403 to realize the functions of the method of this application.

[0178] In a specific implementation, as one example, processor 401 may include one or more CPUs, for example... Figure 4 CPU0 and CPU1 in the CPU.

[0179] In a specific implementation, as one example, the prediction device 40 may include multiple processors, such as... Figure 4 Processors 401 and 407 are described in the text. Each of these processors may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. A processor here may refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0180] The predictive device, such as Figure 4 The diagram includes a processor 401 and a memory 403 for storing executable instructions of the processor 401; wherein the processor 401 is configured to execute executable instructions to implement the virtual power plant resource allocation method as described in any of the possible embodiments above. Furthermore, it achieves the same technical effect, and to avoid repetition, will not be described further here.

[0181] This application also provides a computer-readable storage medium, wherein when the instructions in the computer-readable storage medium are executed by a processor of a control device or control apparatus, the control device or control apparatus is able to perform a virtual power plant resource allocation method as described in any of the possible implementations above. And it can achieve the same technical effect; to avoid repetition, it will not be described again here.

[0182] This application also provides a computer program product, including a computer program or instructions, which are executed by a processor as a virtual power plant resource allocation method according to any of the possible implementations described above. Furthermore, it achieves the same technical effects, and to avoid repetition, it will not be described again here.

[0183] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0184] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A method for allocating virtual power plant resources, characterized in that, The method includes: Determine the historical state observations and historical action quantities of each power plant resource in the virtual power plant; the virtual power plant includes multiple of the following power plant resources: gas turbine, wind turbine, photovoltaic unit, energy storage battery, and electric vehicle; Based on the historical state observations and the historical action quantities, the prediction of the previous actual observations, the previous actual actions, and the current action quantities under the current state observations is carried out, including: using the weighted sum of the comprehensive frequency regulation performance evaluation index of each power plant resource and the net frequency regulation revenue as the reward function of the initial reinforcement model, and constructing the initial reinforcement model that characterizes the mapping relationship between state observations and action quantities. Based on the historical state observations and the historical action quantities, the network parameters and value function of the initial reinforcement model are optimized to obtain the target reinforcement model; The previous actual observation, the previous actual action, and the current state observation are input into the target enhancement model to obtain the current action; The historical action quantity and the current action quantity include the same number of actions as the number of power plant resources included in the virtual power plant; the historical action quantity represents the resource scheduling and allocation results of each power plant resource under the historical state observation. Based on the current action quantity representing the resource scheduling and allocation results of each power plant resource, the operating power of each power plant resource is controlled and adjusted.

2. The virtual power plant resource allocation method according to claim 1, characterized in that, Before constructing the initial reinforcement model that characterizes the mapping relationship between state observations and action quantities by using the comprehensive frequency regulation performance evaluation index and net frequency regulation revenue of each power plant resource as the reward function of the initial reinforcement model, the method further includes: The net frequency regulation revenue of the virtual power plant is determined based on the total frequency regulation revenue of the virtual power plant and the frequency regulation cost of each power plant resource. Based on the regulation rate, regulation accuracy, and response time of the resource scheduling and regulation of each power plant, a comprehensive frequency regulation performance evaluation index for the resource scheduling of each power plant is determined.

3. The virtual power plant resource allocation method according to claim 1, characterized in that, The method further includes: Collect the historical state observations and the historical action quantities according to a preset cycle; Determine the historical state observation for each historical state observation and the historical action quantity for the next preset period; According to the reward function, the comprehensive frequency modulation performance evaluation index and frequency modulation net revenue generated for each of the historical state observations and each of the historical action quantities to the next preset period are solved to obtain the corresponding historical reward. Each historical state observation, the corresponding historical action quantity, the corresponding historical state observation for the next preset period, and the corresponding historical reward are used as a set of training data in the experience pool.

4. The virtual power plant resource allocation method according to claim 3, characterized in that, The method further includes: Based on the training data of each group in the experience pool, determine the probability of occurrence of each historical action quantity under each historical state observation.

5. The virtual power plant resource allocation method according to claim 3, characterized in that, The method further includes: Based on the convergence of historical rewards corresponding to different historical action quantities for two historical state observations corresponding to adjacent preset periods, the training data of each group in the experience pool is filtered to obtain the filtered training data.

6. The virtual power plant resource allocation method according to claim 4, characterized in that, The value function includes network parameters as its function variables, used to evaluate the network parameters of the initial reinforcement model; the optimization of the network parameters and value function of the initial reinforcement model according to the historical state observations and the historical action quantities to obtain the target reinforcement model includes: The training data from each group in the experience pool are input into the initial reinforcement model to obtain the output action quantity for each group; the output action quantity includes the same number of actions as the number of power plant resources included in the virtual power plant; For any set of output action quantities, according to the reward function, determine the model adjustment reward for adjusting the historical state observation quantity in the next preset period based on the input historical action quantity; and determine the difference between the first reward of the output action quantity and the second reward of the corresponding historical action quantity. Based on the preset network objective, the network parameters and the value function are optimized according to the model adjustment reward, the profit difference, and the occurrence probability of the output action quantity in each group.

7. The virtual power plant resource allocation method according to any one of claims 1 to 6, characterized in that, The historical and current state observations include: gas turbine regulation capacity, wind turbine regulation capacity, photovoltaic unit regulation capacity, energy storage battery capacity, electric vehicle state-of-charge capacity, energy storage battery charge / discharge rate, electric vehicle charge / discharge rate, and the winning bid price corresponding to the frequency regulation capacity for each time period.

8. A virtual power plant resource allocation device, characterized in that, The device includes: The determining unit is configured to determine the historical state observations and historical action quantities of each power plant resource in the virtual power plant; the virtual power plant includes multiple of the following power plant resources: gas turbines, wind turbines, photovoltaic units, energy storage batteries, and electric vehicles; The prediction unit is configured to predict the current action quantity under the previous actual observation, the previous actual action, and the current state observation based on the historical state observations and the historical action quantities. This includes: constructing an initial reinforcement model representing the mapping relationship between state observations and action quantities, using the weighted sum of the comprehensive frequency regulation performance evaluation index and net frequency regulation revenue of each power plant resource as the reward function of the initial reinforcement model; optimizing the network parameters and value function of the initial reinforcement model according to the historical state observations and the historical action quantities to obtain a target reinforcement model; and inputting the previous actual observation, the previous actual action, and the current state observation into the target reinforcement model to obtain the current action. The historical action quantities and the current action quantities include the same number of actions as the number of power plant resources included in the virtual power plant. The historical action quantities represent the resource scheduling and allocation results of each power plant resource under the historical state observations. The regulating unit is configured to control and regulate the operating power of each power plant resource according to the resource scheduling and allocation result of each power plant resource represented by the current action quantity.

9. A power resource distribution device, characterized in that, It is configured to perform the virtual power plant resource allocation method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual power plant online optimization scheduling method based on deep reinforcement learning algorithm

    CN119204546A