Virtual power plant cooperative control method and system and electronic equipment
By employing a hierarchical multi-agent deep reinforcement learning-based collaborative control strategy for virtual power plants, combined with an intelligent fusion terminal, the problems of distributed resource heterogeneity and system complexity in virtual power plants are solved. This achieves multi-objective collaborative optimization of macro-energy balance and micro-control, thereby improving the system's adaptability and efficiency.
Patent Information
- Application Number
- CN202511440759.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-10
AI Technical Summary
The heterogeneity and dynamism of distributed resources in virtual power plants, the complexity and uncertainty of system operation, and the challenges of modeling and control make traditional control methods computationally complex, burdensome in communication, and difficult to meet real-time regulation requirements.
A hierarchical multi-agent deep reinforcement learning (HMADRL) collaborative control strategy is adopted, combined with the application of intelligent fusion terminals at the edge of the virtual power plant. Through hierarchical learning and collaborative mechanisms, the control strategy is pushed down to the edge and intelligence is achieved. Bayesian deep reinforcement learning models and robust optimization models are used for decision-making.
It achieves multi-objective collaborative optimization of virtual power plants among macro-energy balance, market interaction strategies and micro-distributed resource refined operation control, improves the scalability and robustness of the system, and adapts to the ever-changing power market environment and grid operating conditions.
Smart Images

Figure CN120914892A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power distribution Internet of Things, in particular to a virtual power plant cooperative control method, a virtual power plant cooperative control system, an electronic device, a storage medium and a computer program product. BACKGROUND
[0002] Virtual power plant (VPP) as an advanced energy management system, by aggregating and optimizing control of geographically dispersed distributed energy resources (DERs) such as distributed photovoltaic, wind power, energy storage system, controllable load and electric vehicle, etc., in the form of a whole to participate in the electricity market and power grid operation, plays a crucial role in improving the flexibility of the power grid, accommodating large-scale renewable energy, and ensuring the safe and stable operation of the power system.
[0003] However, VPP still faces many challenges in actual promotion and efficient operation. Specifically, (1) the heterogeneity and dynamics of distributed resources (DERs): VPP integrates a variety of DERs with different characteristics. Renewable energy (such as photovoltaic and wind power) has significant volatility and intermittency; energy storage systems have diverse operating modes, and need to consider charging and discharging efficiency, life management, etc.; the response ability and willingness of controllable load also differ. These factors together constitute a highly complex and uncertain environment for VPP internal operation. (2) Complexity and uncertainty of system operation: VPP optimization scheduling involves multiple time scales such as day-ahead, intra-day, real-time, etc. How to effectively coordinate between different scales is a difficult problem. At the same time, the fluctuation of electricity market price, the deviation of load demand prediction, the uncertainty of new energy output prediction, and potential equipment failures all bring great challenges to the stable and economic operation of VPP. (3) Modeling and control problems: Traditional VPP control relies on accurate mathematical modeling of internal massive DERs. However, DERs are numerous, diverse in type, and time-varying in parameters, making it extremely difficult to establish an accurate and unified aggregation model. In addition, traditional centralized optimization control methods have high computational complexity and heavy communication burden when facing large-scale VPP, and there is a risk of single point failure, which is difficult to meet the demand of real-time regulation.
[0004] Although the traditional optimization methods based on mathematical programming (such as mixed integer programming, robust optimization, stochastic programming, etc.) are theoretically mature, they are highly dependent on the accuracy of the model and are difficult to effectively handle the inherent strong uncertainty and high nonlinearity of VPP. Although some existing hierarchical control or distributed control architectures can alleviate the drawbacks of centralized control to some extent, there is still much room for improvement and innovation in the coordination efficiency between intelligent agents, the accuracy of upper-level instruction decomposition, the consistency of local autonomy of lower-level intelligent agents and global target guarantee, etc. SUMMARY
[0005] The purpose of the embodiments of the present application is to provide a virtual power plant cooperative control method and system and electronic equipment, a cooperative control strategy of a virtual power plant based on HMADRL (Hierarchical Multi-Agent Deep Reinforcement Learning), combined with the application of intelligent fusion terminals on the edge side of the virtual power plant, realizing the sinking of the control strategy and edge intelligence, overcoming the limitations of traditional control methods in modeling accuracy and computational efficiency through hierarchical learning and cooperative mechanism, to at least solve some of the problems in the background art.
[0006] In order to achieve the above-mentioned purpose, a virtual power plant cooperative control method is provided in the present application, the method comprising: a virtual power plant cooperative control method, the method comprising: obtaining a control target, and selecting a global state parameter affecting the control target according to the control target; calculating a fluctuation coefficient according to historical data of the global state parameter, determining an uncertainty set, and taking it as a hard constraint for decision-making of a high-level intelligent agent based on the control target; establishing a robust optimization model based on the uncertainty set, and solving the robust optimization model to obtain an optimal solution of a robust objective function; training a Bayesian deep reinforcement learning model with the historical data, optimizing variational parameters, and the training target is to maximize the lower bound of the evidence; inputting current state data of the global state parameter into the trained Bayesian deep reinforcement learning model, generating an optimal action, and calculating a corresponding confidence; according to the relationship between the confidence and a preset confidence threshold, selecting one of the optimal action and the optimal solution of the robust objective function to generate a first control instruction and issuing it to a low-level intelligent agent.
[0007] Preferably, the method further comprises: converting, by the low-level intelligent agent, the first control instruction into a second control instruction executable by a distributed resource according to local state parameters, and issuing the second control instruction to the distributed resource; and acquiring, by the low-level intelligent agent, an execution effect of the second control instruction on the distributed resource, and feeding back to the high-level intelligent agent.
[0008] Preferably, the high-level intelligent agent is deployed in a cloud platform or a regional coordination control center of the virtual power plant, and the low-level intelligent agent is deployed at an edge side close to the distributed resource.
[0009] Preferably, the feedback to the high-level intelligent agent comprises updating historical data according to the feedback data and periodically adjusting the variation parameter of the Bayesian deep reinforcement learning model and the fluctuation coefficient of the robust optimization.
[0010] Preferably, a single high-level intelligent agent is arranged in each virtual power plant; the number of low-level intelligent agents in each virtual power plant is determined according to the number of jurisdictional ranges divided by the user and is constrained by the number of intelligent fusion terminals in the edge side.
[0011] Preferably, if the low-level intelligent agent belongs to a low-level intelligent agent cluster, the low-level intelligent agent cluster to which the low-level intelligent agent belongs receives and jointly responds to the first control instruction whose target is the low-level intelligent agent.
[0012] Preferably, the global state parameter comprises a market price parameter, a power grid demand parameter, and a virtual power plant aggregated resource parameter; the market price parameter comprises a market quotation, a historical quotation, a market rule, and an expected quotation; the power grid demand parameter comprises an overall power generation plan, an overall power consumption plan, an operation cost, an operation benefit, a consumption strategy, a charging and discharging scheduling strategy, and a controllable load coordination strategy; and the virtual power plant aggregated resource parameter comprises a resource distribution, a resource state, and a resource division of the virtual power plant.
[0013] Preferably, the first control instruction comprises an action instruction and a target object; the action instruction comprises one of an overall power curve, a cluster frequency modulation capacity, a market participation parameter, a regional coordination target, and an economic dispatching signal; and the target object comprises a low-level intelligent agent or a low-level intelligent agent cluster.
[0014] Preferably, the second control instruction comprises an action instruction, and the action instruction has a bottom-layer communication protocol and a data format that can be recognized and executed by a corresponding distributed resource.
[0015] Preferably, the local state parameter comprises a real-time running state of a local distributed resource, a local environment and prediction information, a deviation between a predicted value and an actual value, a constraint condition of a local power grid, and a device self-running constraint.
[0016] Preferably, the high-level agent adopts a deep reinforcement learning model to generate a first control instruction according to global state parameters; when managing a single distributed resource or multiple weakly coupled distributed resources, the low-level agent adopts an independent deep reinforcement learning model to convert the first control instruction into a second control instruction executable by the distributed resource according to local state parameters; when managing multiple strongly coupled distributed resources, the low-level agent adopts a multi-agent reinforcement learning model to convert the first control instruction into a second control instruction executable by the distributed resource according to local state parameters.
[0017] Preferably, the high-level agent adopts a deep reinforcement learning model to generate a first control instruction according to global state parameters, including: the high-level agent adopts a deep reinforcement learning model to generate corresponding first control instructions based on different global state parameters in different decision stages; the decision stages include real-time control stage, short-term prediction stage and long-term prediction stage.
[0018] Preferably, after the low-level agent obtains the execution effect of the second control instruction on the distributed resource and feeds back to the high-level agent, the method further includes: the high-level agent evaluates the effectiveness of the strategy based on the feedback execution effect and the key running state of the low-level agent and calculates the reward signal, and / or the feedback execution effect and the key running state of the low-level agent are included in the state space of the next decision cycle of the high-level agent.
[0019] Preferably, the low-level agent is further configured to generate a third control instruction targeting other low-level agents or low-level agent clusters, the third control instruction being used for adjacent area coordination, resource mutual aid and local balance optimization.
[0020] In the present application, a virtual power plant collaborative control system is also provided, which includes a high-level agent and a low-level agent; the high-level agent and the low-level agent are separately arranged and connected through communication; the high-level agent is configured to generate a first control instruction according to global state parameters and issue it to the low-level agent; the low-level agent is configured to convert the first control instruction into a second control instruction executable by the distributed resource according to local state parameters and issue it to the distributed resource; and obtain the execution effect of the second control instruction on the distributed resource and feed it back to the high-level agent.
[0021] In the present application, an electronic device is also provided, which includes: at least one processor; a memory connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the virtual power plant collaborative control method described above by executing the instructions stored in the memory.
[0022] A machine readable storage medium is also provided in the present application, and the machine readable storage medium stores instructions which, when executed by a processor, cause the processor to be configured to implement the virtual power plant collaborative control method.
[0023] A computer program product is also provided in the present application, and the computer program product comprises a computer program which, when executed by a processor, implements the virtual power plant collaborative control method.
[0024] The above technical solution has the following beneficial effects: The present application provides a virtual power plant collaborative control strategy based on hierarchical multi-agent deep reinforcement learning, which combines the application of intelligent fusion terminals on the edge side of the virtual power plant, realizes the sinking of the control strategy and edge intelligence, overcomes the limitations of traditional control methods in modeling accuracy and computational efficiency through hierarchical learning and collaborative mechanism, realizes the multi-objective collaborative optimization of the virtual power plant between macro energy balance, market interaction strategy and micro-distributed resource (DER) fine operation control, and responds to the variable power market environment and power grid operation conditions.
[0025] Other features and advantages of the present application will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0026] The accompanying drawings are included to provide a further understanding of the present application, and constitute a part of the specification, and are used together with the following specific embodiments to explain the present application, but do not constitute a limitation on the present application. In the drawings: Figure 1 The steps of the virtual power plant collaborative control method according to the present application are schematically shown; Figure 2 The implementation framework of the virtual power plant collaborative control method according to the present application is schematically shown; Figure 3 The scheduling decision flowchart of the virtual power plant collaborative control method according to the present application is schematically shown; Figure 4 The structure of the virtual power plant collaborative control system according to the present application is schematically shown; Figure 5 The internal structure of the electronic device according to the present application is schematically shown. DETAILED DESCRIPTION
[0027] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application.
[0028] Figure 1 The steps of the virtual power plant collaborative control method according to the embodiments of the present application are schematically shown. As shown in the figure, a virtual power plant collaborative control method comprises the following steps: Figure 1 acquiring a control target and selecting a global state parameter affecting the control target according to the control target; the control target here refers to a target expected to be achieved by the virtual power plant collaborative control.
[0029] calculating a fluctuation coefficient according to historical data of the global state parameter, determining an uncertainty set, and taking the uncertainty set as a hard constraint for decision-making of a high-level agent based on the control target; establishing a robust optimization model based on the uncertainty set, and solving the robust optimization model to obtain an optimal solution of a robust objective function; training a Bayesian deep reinforcement learning model with the historical data, optimizing variational parameters, and taking the maximization of the lower bound of the evidence as the training target; inputting current state data of the global state parameter into the trained Bayesian deep reinforcement learning model, generating an optimal action, and calculating a corresponding confidence; according to the relationship between the confidence and a preset confidence threshold, selecting one of the optimal action and the optimal solution of the robust objective function to generate a first control instruction and issuing the first control instruction to a low-level agent.
[0030] Considering that market price fluctuations and photovoltaic / wind power output prediction deviations will cause the global strategy of the HLA to fail, robust optimization is introduced and a Bayesian DRL agent is used instead of a traditional DRL, the fluctuation range of the uncertainty parameter is statistically calculated using historical data, an uncertainty set is constructed, and the uncertainty set is taken as a hard constraint for decision-making of the HLA to ensure that the global strategy of the HLA is still feasible in the worst scenario. An embodiment is as follows: 1. Uncertainty boundary modeling (robust optimization) 1) Uncertainty parameter definition Let the historical data sample set be , (n is the historical data volume, N is the time step) and the fluctuation boundary is determined based on the 3σ principle (99.7% confidence interval): t Market electricity price fluctuation is: (1) wherein, is the electricity price prediction value, and the electricity price fluctuation coefficient .
[0031] Photovoltaic power output fluctuation is: (2) wherein, the photovoltaic power output fluctuation coefficient
[0032] Wind power output fluctuation is: (3) Where, wind power output fluctuation coefficient .
[0033] 2) Uncertainty set and robust objective function Define the uncertainty set of HLA decision Contains all possible parameter combinations: .
[0034] The core decision of HLA is "bid power Q", the goal is to maximize the revenue in the worst case scenario within the uncertainty set Robust optimization model: .
[0035] Where: is the operating cost of HLA, the constraint condition is: .
[0036] is the maximum aggregated output, , is the energy storage discharge power, is the energy storage discharge cost, is the operating cost, is the load power.
[0037] 2. Bayesian DRL (Deep Reinforcement Learning) 1) Bayesian DRL network training Policy network parameters Subject to probability distribution (traditional DRL is deterministic parameter), use variational Bayesian inference to approximate parameter posterior , use variational distribution to fit the posterior, the training goal is to maximize the lower bound of evidence: (4) Where, is the likelihood term (the fitting degree of policy on historical data D, s is the global state, a is the first control instruction.
[0038] is the KL divergence, used to constrain the difference between the variational distribution and the prior , to avoid overfitting.
[0039] Use variational distribution Gaussian distribution , For variational parameters, equation (4) is modified as: (5) where, is the historical state-action pair, is the global state, is the historical first control instruction, M is the batch size, is the parameter prior.
[0040] 2) Strategy confidence calculation For the current state , by sampling the strategy parameters multiple times, ~ , generate k candidate actions … The confidence is: , where, is the variance, and the value range .
[0041] 3) Strategy triggering logic Set the confidence threshold to , if , execute the Bayesian DRL optimal action, ; if , execute the optimal solution of the robust objective function: .
[0042] The following embodiments take the control target of intraday market bidding as an example, and the overall steps are as follows: Step 1: According to the historical data of global state parameters (market price, photovoltaic / wind power output), calculate the fluctuation coefficient to determine the uncertainty set .
[0043] Step 2: Train the Bayesian DRL model with historical data D to optimize the variational parameters , so that converge.
[0044] Step 3: HLA obtains the current global state , Bayesian DRL samples K sets of parameters , generates candidate bidding power , and calculates the confidence .
[0045] Step 4: If , output the preferred decision , if , execute the optimal solution of the robust objective function, and output the suboptimal decision , where, The desired function is.
[0046] Step 5: The bidding power is issued as a "first control instruction" to the LLA.
[0047] Step 6: LLA feedback execution effect.
[0048] Step 7: Update the historical data set D according to the feedback data, and periodically adjust the DRL variation parameter and the volatility coefficient a of robust optimization.
[0049] Through the above implementation, the intelligence level of HLA decision and the accuracy of decision are realized.
[0050] In some embodiments of the present application, in view of the challenges of internal resource diversity of virtual power plant (VPP), operation uncertainty, and difficulty of traditional control method to balance global optimization and local fine control, the present application proposes a collaborative control framework of hierarchical multi-agent deep reinforcement learning (HMADRL), which decouples the complex VPP control problem into two interrelated but well-defined levels: the high-level is responsible for macro-strategy planning, and the low-level is responsible for micro-strategy execution. Specifically, the method includes: The first control instruction is issued to the low-level agent by the high-level agent according to the global state parameter, and the specific implementation can use the steps described above.
[0051] The first control instruction is converted into a second control instruction executable by the distributed resource by the low-level agent according to the local state parameter, and is issued to the distributed resource; The execution effect of the second control instruction on the distributed resource is obtained by the low-level agent and fed back to the high-level agent.
[0052] In the above implementation, first, a double-layer intelligent control architecture composed of a single high-level agent (High-Level Agent, hereinafter also referred to as HLA) and multiple low-level agents (Low-Level Agents, hereinafter also referred to as LLA) is constructed. This structured design aims to effectively decompose and manage the complex decision-making process of virtual power plant (VPP), improving the scalability and robustness of the system. The above three steps roughly correspond to the three steps of target issuance, instruction execution and local optimization, and feedback and learning.
[0053] In some optional embodiments, the feedback to the high-level agent includes updating the historical data according to the feedback data, and periodically adjusting the variation parameter of the Bayesian deep reinforcement learning model and periodically adjusting the volatility coefficient of robust optimization. That is, the method also includes the aforementioned Step 7.
[0054] Figure 2 An implementation framework of the virtual power plant collaborative control method is shown schematically. As shown in the figure, in the present embodiment, the high-level agent is deployed in the cloud platform or regional coordination control center of the virtual power plant, and the low-level agent is deployed near the edge side of the distributed resource. The HLA is usually deployed in the cloud platform or regional coordination control center of the VPP, and is responsible for the global and long-period (such as 24 to 48 hours) macroscopic strategy making task. It focuses on the energy balance of the VPP as a whole, the market interaction benefit, and the response to the superior grid instruction. The LLA is deployed near the edge side of the distributed resource (DER), and is specifically dependent on the intelligent fusion terminal. Each LLA is responsible for the local and short-period fine control and optimization of one or more DERs in its jurisdiction. Figure 2
[0055] In some embodiments of the present application, a single high-level agent is provided in each virtual power plant; the number of low-level agents in each virtual power plant is determined according to the number of jurisdictional ranges divided by the user, and is constrained by the number of intelligent fusion terminals in the edge side. In order to unify the control in the virtual power plant, a single high-level agent is provided in each virtual power plant. The deployment position of the LLA includes that the core logic of the LLA runs on the intelligent fusion terminal installed at each DER access point or small DER aggregation point. One intelligent fusion terminal can carry one LLA, so the upper limit of the number of LLAs can be determined according to the number of intelligent fusion terminals. The LLA manages one or more DERs directly connected thereto.
[0056] In some embodiments, a low-level agent cluster composed of low-level agents is also provided. If multiple fusion terminals and the DERs managed thereby are geographically adjacent or functionally tightly coupled, they can also constitute an LLA cluster, which collectively responds to the instruction of the HLA. If the low-level agent belongs to a low-level agent cluster, the low-level agent cluster to which the low-level agent belongs receives and collectively responds to the first control instruction whose target is the low-level agent.
[0057] The embodiment describes the aforementioned global state parameters. The global state parameters include: market price parameters, grid demand parameters, and virtual power plant aggregated resource parameters. The core responsibilities of HLA include: market interaction strategy formulation, macro energy balance and economic dispatch, and regional / cluster target setting and instruction issuance. Market interaction strategy formulation includes: HLA is responsible for participating in various power markets on behalf of the entire VPP, such as energy markets and ancillary service markets (frequency modulation, backup, etc.). It needs to intelligently formulate bidding strategies, offer curves, and declared capacities according to market price parameters and the adjustable resource conditions within the VPP. The market price parameters include: market offers, historical offers, market rules, and expected offers. Macro energy balance and economic dispatch include: one of the core tasks of HLA is to formulate the overall power generation and power consumption plan of VPP to ensure internal energy supply and demand balance and strive for the lowest operating cost or maximum revenue. This includes optimizing renewable energy consumption strategies (such as coordinating energy storage charging to absorb excess photovoltaic / wind power), managing energy storage system charging and discharging dispatch at the global level (such as participating in peak-valley arbitrage), and coordinating controllable load response, so the grid demand parameters include: overall power generation plan, overall power consumption plan, operating cost, operating revenue, consumption strategy, charging and discharging dispatch strategy, and controllable load coordination strategy. The virtual power plant aggregated resource parameters include: resource distribution, resource state, and resource division of the virtual power plant.
[0058] In some embodiments of the present application, the first control instruction of HLA, although issued to LLA or LLAs, is actually directed to the DER or DER aggregation managed by one or more LLAs. The DER aggregation can be divided by geographical location, such as all DERs in an industrial park, or by resource type, such as all controllable air conditioning clusters. The first control instruction includes action instructions and target objects; the action instructions include: overall power curve: the total active power output / consumption curve and reactive power curve that the VPP should reach in one or more future dispatch periods (such as the next 24 hours, every 15 minutes a point). Cluster frequency modulation capacity: requires the kth DER cluster, such as the energy storage cluster, to provide upward / downward frequency modulation capacity in a certain period. Market participation parameters: such as in the day-ahead energy market, the bidding capacity and corresponding offer for a specific period. Regional coordination target: for example, to solve the voltage overrun problem of the local power grid, LLA requires the LLA to adjust the reactive power output of the DER under its jurisdiction to achieve a regional reactive power compensation target or voltage control target. Economic dispatch signal: issue internal "shadow price" or incentive signal to LLAs to guide them to optimize in the direction of reducing the total cost of VPP while meeting power balance.
[0059] In some embodiments of the present application, the second control instruction includes action instructions, which have underlying communication protocols and data formats that can be recognized and executed by the corresponding distributed resources. The action instructions herein vary according to different distributed resources. The actions of the action instructions output by the LLA are as follows. For a photovoltaic inverter j: active power given value, reactive power given value, or power factor setting, or reference voltage value in voltage control mode. For a storage system k: charge / discharge power (positive for discharging, negative for charging), or charge / discharge current, and operating mode (such as peak shaving mode, output smoothing mode, instruction tracking mode). For controllable load m (such as air conditioner, water heater): start / stop instruction, power adjustment percentage, temperature set point adjustment. For electric vehicle charging and discharging pile n: maximum allowed charging power, discharge power instruction in V2G (technology for bidirectional energy flow between electric vehicles and power grid) mode. These action instructions will be converted by the fusion terminal into underlying communication protocols and data formats that can be recognized and executed by the corresponding DER.
[0060] In some embodiments of the present application, the local state parameters include: real-time operating state of the local distributed resource, such as: instantaneous power generation of photovoltaic array, inverter efficiency, SOC of energy storage unit, temperature, state of health (SOH), current energy consumption level and adjustable range of controllable load, etc. Local environment and prediction information, such as: local light intensity, environmental temperature collected by the fusion terminal, and super short-term (such as future 5-60 minutes) DER output prediction and load prediction generated based thereon. Deviation between predicted value and actual value, such as prediction deviation obtained by real-time tracking. Constraints of local power grid, such as: under the premise that the intelligent fusion terminal can be perceived, whether the voltage of the local grid-connected point is within the allowed range, whether the feeder power flow is close to the limit value, etc. Constraints of device itself, such as: such as upper and lower limits of SOC safety of storage, maximum charge / discharge rate, daily cycle limit, power factor adjustment range of inverter, climbing rate limit, etc.
[0061] In some embodiments of the present application, in the HLA, the HLA obtains global state parameters and deep reinforcement learning (DRL) model outputs based on its perception of the global state, formulates macro operation targets or guiding strategies. In each LLA, or in each intelligent fusion terminal, an intelligent agent running therein generally adopts a deep reinforcement learning algorithm to learn its control strategy, and the specific algorithm selection depends on the complexity and coupling degree of the DERs it manages. For example: if the fusion terminal manages one DER or the coupling between multiple DERs it manages is weak, i.e. controlling one DER does not significantly affect the optimal strategy of other DERs, a single-agent DRL algorithm can be adopted. If the fusion terminal manages multiple DERs that are coupled to each other, i.e. like a microgrid of generation and storage units, or an EV cluster that needs to coordinate charging and discharging, a multi-agent reinforcement learning algorithm is adopted. In this case, the fusion terminal can contain multiple sub-agents, each corresponding to a DER, which cooperatively learn and act under the framework of LLA. The input information of the single-agent DRL algorithm or the multi-agent reinforcement learning algorithm in this case is shown in the following examples, including: instructions / targets from the HLA: for example, the target power allocated to the i-th LLA, i.e. LLA i , or the capacity participating in frequency modulation. Real-time state of the local j-th DER, i.e. DER j : actual output of photovoltaic, actual output of wind turbine, SOC, voltage, current of energy storage k, current power and adjustable state of controllable load m. Local comprehensive prediction information: local aggregated photovoltaic output prediction, local aggregated load prediction, and deviation of these predictions in the past period of time. Local grid information: including grid point voltage, frequency, etc. if available, and if the fusion terminal has corresponding measurement or estimation capability, also including local line flow.
[0062] In some embodiments of the present application, after the execution effect of the second control instruction on the distributed resources is obtained by the low-level agent and fed back to the high-level agent, the method further comprises: the execution effect of the LLAs, such as actual aggregated output and target tracking error, and local key operating states such as resource adjustable potential and abnormal events, are processed and uploaded to the HLA as feedback information. These feedback information are used for the HLA to evaluate the effectiveness of its strategy and calculate the reward signal on the one hand, and become part of the state space of the next decision cycle of the HLA on the other hand. Through this continuous interaction and feedback, both the HLA and the LLAs can continuously optimize their respective strategy models, thereby improving the collaborative performance and adaptive ability of the entire VPP system.
[0063] In some embodiments of the present application, the low-level agent is further configured to generate a third control instruction targeting other low-level agents or a cluster of low-level agents, the third control instruction being used for adjacent area coordination, resource mutual aid, and local balance optimization. In the aforementioned system comprising a high-level agent and a low-level agent, in addition to the interaction between the high-level agent and the low-level agent, an interaction mechanism between the low-level agents is designed, i.e., the low-level agent interacts with other low-level agents or a cluster of low-level agents through the third control instruction to achieve horizontal collaboration. The third control instruction realizes the aforementioned horizontal collaboration through adjacent area coordination, resource mutual aid, and local balance optimization.
[0064] In some embodiments of the present application, the high-level agent adopts a deep reinforcement learning model to generate the first control instruction according to the global state parameter, comprising: the high-level agent adopts a deep reinforcement learning model to generate a corresponding first control instruction based on different global state parameters in different decision stages; the decision stages include a real-time control stage, a short-term prediction stage, and a long-term prediction stage. Due to the change of the first control instruction, the second control instruction on the LLA side also changes accordingly. For example, with "day" as the granularity, the short-term prediction stage is also called the intra-day stage, the long-term prediction stage is selected as the previous day, also called the day-ahead stage, and the real-time control stage is also called the real-time stage. Figure 3 A scheduling decision flowchart of the virtual power plant collaborative control method according to an embodiment of the present application is schematically shown. As shown in Figure 3 the day-ahead stage, the system mainly performs long-term prediction and planning to formulate market participation strategies; in the intra-day stage, rolling optimization and plan adjustment are performed; in the real-time stage, precise control and rapid response are executed, and the system performance is continuously optimized through continuous execution monitoring, feedback, and learning.
[0065] The detailed flow is described as follows: 1. Day-ahead stage. HLA data collection and prediction: the high-level agent collects historical data, meteorological data, market price data, and power grid dispatching information within the VPP network, and uses a deep learning model to perform day-ahead load prediction, renewable energy output prediction, and electricity price prediction. HLA day-ahead strategy formulation: based on the prediction results, the HLA formulates energy balance strategies and market participation strategies, considering multiple dimensions such as economy, reliability, and environmental protection. HLA target issuance: the HLA decomposes the preliminary control target into specific target values for each LLA cluster, and issues them to each intelligent fusion terminal through a secure communication channel. LLA resource pre-allocation: after receiving the target, each LLA combines local DER characteristics and constraint conditions to perform preliminary resource allocation planning and formulate a day-ahead execution plan.
[0066] 2. Intraday stage. HLA rolling optimization: based on updated short-term prediction, market information changes and system actual running state, HLA performs rolling optimization once an hour to dynamically adjust VPP operation strategy. HLA target update: HLA calculates and updates control targets for each LLA cluster, responding to external environment and internal state changes in a timely manner. LLA plan adjustment: each LLA re-optimizes local DER output plan according to updated target values and the latest local short-term prediction, ensuring efficient execution of HLA-issued targets.
[0067] 3. Real-time stage. HLA global monitoring: HLA monitors VPP overall running state in real time, including key parameters such as grid frequency, voltage, power balance, etc., for minute-level short-term scheduling. LLA fine-grained control: LLA performs second-level fine-grained control on distributed energy sources according to real-time state monitoring and HLA-issued instructions, combined with local ultra-short-term prediction. LLA autonomous response: when local random events occur, LLA can make response decisions autonomously without waiting for HLA instructions, improving the system's ability to respond to random events. DER execution and feedback: distributed energy devices execute control instructions, and real-time monitoring of execution effectiveness is performed through intelligent fusion terminals and feedback to LLA and HLA.
[0068] Model online update: HLA and LLA periodically update their respective reinforcement learning models based on execution effectiveness evaluation and historical data accumulation, forming a closed-loop optimization.
[0069] Key data interaction explanation. Bottom-up state reporting: each DER device → LLA: real-time running state, adjustable capacity, constraint conditions. LLA → HLA: regional aggregated state, execution capability evaluation, abnormal situation report. Top-down control: HLA → LLA: overall control target, adjustment strategy, priority setting; LLA → DER: precise control instruction, parameter setting, running mode switching. Horizontal coordination: LLA LLA: coordination with adjacent regions, resource mutual aid, local balance optimization. Closed-loop feedback: execution results → LLA → HLA: instruction execution effectiveness, deviation analysis, model optimization signal.
[0070] Considering that market price fluctuations and photovoltaic / wind power output prediction deviations can cause the global strategy of HLA to fail, robust optimization is introduced and a Bayesian DRL agent is used instead of traditional DRL. The fluctuation range of uncertainty parameters is statistically calculated using historical data, and an uncertainty set is constructed as a hard constraint for HLA decision-making, ensuring that the global strategy of HLA is still feasible in the worst-case scenario.
[0071] As can be seen from the above implementation methods, the implementation methods of this application provide a VPP collaborative control strategy based on hierarchical multi-agent deep reinforcement learning. Combined with the application of intelligent fusion terminals at the edge of the VPP, the control strategy is pushed down and edge intelligence is realized. Through hierarchical learning and collaborative mechanism, the limitations of traditional control methods in terms of modeling accuracy and computational efficiency are overcome. The VPP achieves multi-objective collaborative optimization between macro-energy balance, market interaction strategy and micro-level DER refined operation control, so as to cope with the ever-changing power market environment and grid operation conditions.
[0072] Based on the same inventive concept, this application also provides a virtual power plant collaborative control system. Figure 4 A schematic diagram illustrating the structure of a virtual power plant collaborative control system according to an embodiment of this application is shown. Figure 4 As shown, the system includes a high-level intelligent agent and a low-level intelligent agent; the high-level intelligent agent and the low-level intelligent agent are deployed separately but connected by communication; the high-level intelligent agent is configured to: generate a first control command based on global state parameters and send it to the low-level intelligent agent, preferably by performing the aforementioned steps: obtaining a control objective, and selecting global state parameters that affect the control objective based on the control objective; calculating the fluctuation coefficient based on historical data of the global state parameters, determining the uncertainty set, and using it as a hard constraint for the high-level intelligent agent's decision based on the control objective; establishing a robust optimization model based on the uncertainty set, solving the robust optimization model to obtain the optimal solution of the robust objective function; and training a Bayesian algorithm using the historical data. A deep reinforcement learning model is used to optimize variational parameters, with the training objective being to maximize the lower bound of evidence. The current state data of the global state parameters are input into the trained Bayesian deep reinforcement learning model to generate an optimal action and calculate the corresponding confidence level. Based on the relationship between the confidence level and a preset confidence threshold, one of the optimal action and the optimal solution of the robust objective function is selected to generate a first control command, which is then sent to a low-level agent. The low-level agent is configured to: convert the first control command into a second control command executable by distributed resources based on local state parameters and send it to the distributed resources; and obtain the execution effect of the second control command on the distributed resources and feed it back to the high-level agent.
[0073] In some alternative implementations, the high-level intelligent agents are deployed on the cloud platform of the virtual power plant or in a regional coordination and control center, while the low-level intelligent agents are deployed at the edge of the distributed resources.
[0074] In some alternative implementations, each virtual power plant is equipped with a single high-level intelligent agent; the number of low-level intelligent agents in each virtual power plant is determined according to the number of jurisdictions divided by the user and is constrained by the number of intelligent fusion terminals on the edge side.
[0075] In some optional embodiments, if the low-level agent belongs to a low-level agent cluster, the low-level agent cluster to which the low-level agent belongs receives and jointly responds to the first control instruction for the low-level agent.
[0076] In some optional embodiments, the global state parameter includes a market price parameter, a power grid demand parameter, and a virtual power plant aggregated resource parameter; the market price parameter includes a market quotation, a historical quotation, a market rule, and an expected quotation; the power grid demand parameter includes an overall power generation plan, an overall power consumption plan, an operation cost, an operation benefit, a consumption strategy, a charging and discharging scheduling strategy, and a controllable load coordination strategy; and the virtual power plant aggregated resource parameter includes a resource distribution, a resource state, and a resource division of the virtual power plant.
[0077] In some optional embodiments, the first control instruction includes an action instruction and a target object; the action instruction includes one of an overall power curve, a cluster frequency modulation capacity, a market participation parameter, a regional coordination target, and an economic scheduling signal; and the target object includes a low-level agent or a low-level agent cluster.
[0078] In some optional embodiments, the second control instruction includes an action instruction, and the action instruction has a bottom-layer communication protocol and a data format that can be recognized and executed by a corresponding distributed resource.
[0079] In some optional embodiments, the local state parameter includes a real-time operation state of a local distributed resource, a local environment and prediction information, a deviation between a predicted value and an actual value, a constraint condition of a local power grid, and a device self-operation constraint.
[0080] In some optional embodiments, the high-level agent adopts a deep reinforcement learning model to generate a first control instruction according to a global state parameter; when managing a single distributed resource or multiple weakly coupled distributed resources, the low-level agent adopts an independent deep reinforcement learning model to convert the first control instruction into a second control instruction that can be executed by a distributed resource according to a local state parameter; and when managing multiple strongly coupled distributed resources, the low-level agent adopts a multi-agent reinforcement learning model to convert the first control instruction into a second control instruction that can be executed by a distributed resource according to a local state parameter.
[0081] In some optional embodiments, the high-level agent adopts a deep reinforcement learning model to generate a first control instruction according to a global state parameter, including that the high-level agent adopts a deep reinforcement learning model to generate a corresponding first control instruction based on different global state parameters in different decision stages; and the decision stages include a real-time control stage, a short-term prediction stage, and a long-term prediction stage.
[0082] In some optional embodiments, after the low-level agent acquires the execution effect of the second control instruction on the distributed resources and feeds back to the high-level agent, the system further comprises: the high-level agent evaluates the effectiveness of the strategy and calculates the reward signal based on the feedback execution effect and the key running state of the low-level agent, and / or the feedback execution effect and the key running state of the low-level agent are included in the state space of the next decision cycle of the high-level agent.
[0083] In some optional embodiments, the low-level agent is further configured to generate a third control instruction targeting other low-level agents or a cluster of low-level agents, the third control instruction being used for adjacent area coordination, resource mutual aid and local balance optimization.
[0084] The specific definitions of the various functional modules in the virtual power plant collaborative control system described above can be referred to the definitions of the virtual power plant collaborative control method in the foregoing, which will not be repeated here. Each module in the system can be realized by software, hardware and combinations thereof, in whole or in part. Each module described above can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module. It is also based on the hierarchical multi-agent deep reinforcement learning technology, which realizes the sinking of the control strategy and the edge intelligence, overcomes the limitations of the traditional control method in modeling accuracy and computational efficiency through the hierarchical learning and collaborative mechanism, and achieves the beneficial effects.
[0085] In some embodiments of the present application, an electronic device is also provided, comprising: at least one processor; a memory connected with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor executes the foregoing virtual power plant collaborative control method. Its internal structure diagram can be as shown in Figure 5 Figure 5 The internal structure diagram of the electronic device according to the embodiments of the present application is schematically shown. The electronic device comprises a processor A01, a network interface A02, a memory (not shown in the figure) and a database (not shown in the figure) connected by a system bus. Among them, the processor A01 of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes an internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02 and a database (not shown in the figure). The internal memory A03 provides an environment for the running of the operating system B01 and the computer program B02 in the non-volatile storage medium A04. The network interface A02 of the electronic device is used to communicate with the external terminal through the network connection. The computer program B02 is executed by the processor A01 to implement a virtual power plant collaborative control method.
[0086] Those skilled in the art can understand that Figure 5 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0087] In an embodiment provided by the present application, a machine readable storage medium is provided, and the machine readable storage medium has instructions stored thereon. The instructions, when executed by a processor, cause the processor to be configured to perform the virtual power plant collaborative control method.
[0088] In an embodiment provided by the present application, a computer program product is provided, and the computer program product includes a computer program. The computer program, when executed by a processor, implements the virtual power plant collaborative control method.
[0089] Those skilled in the art can understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) containing computer-usable program code.
[0090] The present application is described with reference to flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing apparatus to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing apparatus produce a device implemented in accordance with the flowcharts and / or block diagrams. Figure 1 The function specified in one or more flows and / or blocks. Figure 1 The device that implements the function specified in one or more flows and / or blocks.
[0091] These computer program instructions can also be stored in a computer readable storage medium that can direct the computer or other programmable data processing apparatus to work in a specific manner, so that the instructions stored in the computer readable storage medium produce a manufactured product including instruction devices that implement the flowcharts and / or block diagrams. Figure 1 The function specified in one or more flows and / or blocks. Figure 1 The device that implements the function specified in one or more flows and / or blocks.
[0092] These computer program instructions can also be loaded into a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 Figure 1
[0093] In one typical configuration, the computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0094] The memory can include non-persistent memory and / or volatile memory, such as random access memory (RAM) about which the processor can execute instructions. The memory can also include non-volatile memory, such as read only memory (ROM), electrically programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), flash memory, or other memory technologies, about which the processor can execute instructions. The memory is an example of computer readable media.
[0095] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically programmable read only memory (EEPROM), flash memory or other memory technologies, compact disc read only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.
[0096] It should also be noted that the terms "comprising", "comprises", "including", "includes" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or apparatus that comprises a list of elements does not include only those elements recited, but can also include other elements not expressly listed or inherent to such process, method, article or apparatus. Without limitation, an element preceded by "comprises a" does not, without more constraints, foreclose the existence of additional identical elements in the process, method, article or apparatus that includes the element.
[0097] The above merely provides an example of the present application, and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall fall into the scope of claims of the present application.
Claims
1. A virtual power plant cooperative control method, characterized by, The method comprises: acquiring a control target and selecting a global state parameter affecting the control target according to the control target; calculating a fluctuation coefficient according to historical data of the global state parameter, determining an uncertainty set, and taking the uncertainty set as a hard constraint for decision-making of a high-level agent based on the control target; establishing a robust optimization model based on the uncertainty set, and solving the robust optimization model to obtain an optimal solution of a robust objective function; training a Bayesian deep reinforcement learning model with the historical data, optimizing variational parameters, and taking maximizing the lower bound of evidence as a training target; inputting current state data of the global state parameter into the trained Bayesian deep reinforcement learning model, generating an optimal action, and calculating a corresponding confidence; selecting one of the optimal action and the optimal solution of the robust objective function according to a relationship between the confidence and a preset confidence threshold, and generating a first control instruction to be sent to a low-level agent.
2. The method of claim 1, wherein, The method further comprises: converting, by the low-level agent, the first control instruction into a second control instruction executable by a distributed resource according to local state parameters, and sending the second control instruction to the distributed resource; acquiring, by the low-level agent, an execution effect of the second control instruction on the distributed resource, and feeding back to the high-level agent.
3. The method of claim 2, wherein, The high-level agent is deployed on a cloud platform of a virtual power plant or a regional coordination control center, and the low-level agent is deployed on an edge side close to the distributed resource.
4. The method of claim 3, wherein, The feedback to the high-level agent comprises: updating the historical data according to the feedback data, and periodically adjusting the variational parameters of the Bayesian deep reinforcement learning model and the fluctuation coefficient of the robust optimization.
5. The method of claim 2, wherein, A single high-level agent is arranged in each virtual power plant; the number of low-level agents in each virtual power plant is determined according to the number of jurisdictional ranges divided by a user, and is constrained by the number of intelligent fusion terminals on the edge side.
6. The method of claim 2, wherein, If a low-level agent belongs to a low-level agent cluster, the low-level agent cluster to which the low-level agent belongs receives and jointly responds to the first control instruction for the low-level agent.
7. The method of claim 2, wherein, The global state parameters comprise: market price parameters, power grid demand parameters, and virtual power plant aggregated resource parameters; The market price parameters comprise: market quotes, historical quotes, market rules, and expected quotes; The power grid demand parameters comprise: overall power generation plans, overall power consumption plans, operation costs, operation benefits, consumption strategies, charging and discharging scheduling strategies, and controllable load coordination strategies; The virtual power plant aggregated resource parameters comprise: resource distribution, resource state, and resource division of the virtual power plant.
8. The method of claim 6, wherein, The first control instruction comprises an action instruction and a target object; The action instruction comprises one of an overall power curve, a cluster frequency modulation capacity, a market participation parameter, a regional coordination target, and an economic dispatching signal; The target object comprises a low-level agent or a low-level agent cluster.
9. The method of claim 2, wherein, The second control instruction comprises an action instruction, and the action instruction has a corresponding bottom-layer communication protocol and data format that can be recognized and executed by a distributed resource.
10. The method of claim 2, wherein, The local state parameters comprise: real-time running state of a local distributed resource, local environment and prediction information, deviation between predicted values and actual values, constraint conditions of a local power grid, and self-running constraints of equipment.
11. The method of claim 2, wherein, The high-level agent adopts a deep reinforcement learning model to generate a first control instruction according to global state parameters; When managing a single distributed resource or multiple weakly coupled distributed resources, the low-level agent adopts an independent deep reinforcement learning model to convert the first control instruction into a second control instruction executable by the distributed resource according to local state parameters; When managing multiple strongly coupled distributed resources, the low-level agent adopts a multi-agent reinforcement learning model to convert the first control instruction into a second control instruction executable by the distributed resource according to local state parameters.
12. The method of claim 11, wherein, The high-level agent adopts a deep reinforcement learning model to generate a first control instruction according to global state parameters, including: The high-level agent adopts a deep reinforcement learning model to generate corresponding first control instructions based on different global state parameters at different decision stages; The decision stages include real-time control stage, short-term prediction stage and long-term prediction stage.
13. The method of claim 2, wherein, After the low-level agent obtains the execution effect of the second control instruction on the distributed resource and feeds back to the high-level agent, the method further includes: The high-level agent assesses the effectiveness of the strategy and calculates the reward signal based on the feedback execution effect and the key running state of the low-level agent, And / or the feedback execution effect and the key running state of the low-level agent are included in the state space of the next decision cycle of the high-level agent.
14. The method of claim 8, wherein, The low-level agent is also configured to generate a third control instruction targeting other low-level agents or clusters of low-level agents, which is used for adjacent area coordination, resource mutual aid and local balance optimization.
15. A virtual power plant coordinated control system, characterized by, The system includes a high-level agent and a low-level agent; the high-level agent and the low-level agent are separately arranged and connected through communication; The high-level agent is configured to perform the virtual power plant collaborative control method of claim 1; The low-level agent is configured to convert the first control instruction into a second control instruction executable by the distributed resource according to local state parameters, and issue the second control instruction to the distributed resource; and obtain the execution effect of the second control instruction on the distributed resource and feed back to the high-level agent.
16. An electronic device, comprising: Including: At least one processor; Memory connected with the at least one processor; Wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the virtual power plant collaborative control method of any one of claims 1 to 14 by executing the instructions stored in the memory.
17. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to implement the virtual power plant collaborative control method of any one of claims 1 to 14.
18. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to implement the virtual power plant collaborative control method of any one of claims 1 to 14.
Citation Information
Patent Citations
Collaborative optimization scheduling method, device and equipment for multiple virtual power plants, and storage medium
CN114036825A
Robust optimization virtual power plant optimization control system based on stochastic programming
CN115422728A
Virtual power plant online optimization scheduling method based on deep reinforcement learning algorithm
CN119204546A
Virtual power plant multi-stage coordinated regulation and control method and system
CN119275940A
Virtual power plant intelligent control method and system based on multiple agents
CN120474103A
Cited By
Power grid intelligent decision-making method, system and equipment oriented to source network load storage low-carbon operation, medium and computer program product
CN122178455A