Virtual power plant collaborative control method and system, and electronic device

By employing a collaborative control strategy based on hierarchical multi-agent deep reinforcement learning, combined with an intelligent fusion terminal, the heterogeneity of distributed resources and system complexity in virtual power plants are addressed. This enables collaborative optimization of virtual power plants at both macro and micro levels, enhancing the system's adaptability and computational efficiency.

CN120914892BActive Publication Date: 2026-02-10BEIJING SMARTCHIP MICROELECTRONICS TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511440759.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2026-02-10
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

The heterogeneity and dynamism of distributed resources in virtual power plants, the complexity and uncertainty of system operation, and the challenges of modeling and control are all factors that traditional control methods cannot effectively handle in terms of modeling accuracy and computational efficiency, and are difficult to effectively handle strong uncertainty and highly nonlinear characteristics.

Method used

A hierarchical multi-agent deep reinforcement learning (HMADRL) collaborative control strategy is adopted, combined with the application of intelligent fusion terminals at the edge of the virtual power plant. Through hierarchical learning and collaborative mechanisms, the control strategy is pushed down to the edge and intelligence is achieved. Bayesian deep reinforcement learning models and robust optimization models are used for decision support.

Benefits of technology

It achieves multi-objective collaborative optimization of virtual power plants among macro-energy balance, market interaction strategies and micro-distributed resource refined operation control, improves the system's adaptability and computing efficiency, and copes with the ever-changing power market environment and grid operation conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120914892B_ABST
    Figure CN120914892B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a virtual power plant cooperative control method and system and an electronic device, and relate to the technical field of power distribution Internet of Things. The method comprises: obtaining a control target; calculating a fluctuation coefficient according to historical data of a global state parameter, and determining an uncertainty set; establishing a robust optimization model based on the uncertainty set, solving the robust optimization model to obtain an optimal solution of a robust objective function; training a Bayesian deep reinforcement learning model using the historical data; inputting current state data of the global state parameter into the trained Bayesian deep reinforcement learning model, generating an optimal action, and calculating a corresponding confidence; according to the relationship between the confidence and a preset confidence threshold, selecting one of the optimal action and the optimal solution of the robust objective function to generate a first control instruction and issuing the first control instruction to a low-level intelligent agent. The embodiments provided in the present application can promote the development of virtual power plants in the direction of finer, more intelligent and more autonomous.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power distribution Internet of Things (IoT) technology, specifically to a virtual power plant collaborative control method, a virtual power plant collaborative control system, an electronic device, a storage medium, and a computer program product. Background Technology

[0002] Virtual power plants (VPPs), as an advanced energy management system, aggregate and optimize the control of geographically dispersed distributed energy resources (DERs), such as distributed photovoltaics, wind power, energy storage systems, controllable loads, and electric vehicles, to participate in the electricity market and grid operation as a whole. They play a crucial role in improving grid flexibility, accommodating large-scale renewable energy, and ensuring the safe and stable operation of the power system.

[0003] However, VPP still faces many challenges in its actual promotion and efficient operation. Specifically, these include: (1) Heterogeneity and dynamism of distributed resources (DERs): VPP integrates a wide variety of DERs with different characteristics. The output of renewable energy (such as photovoltaic and wind power) has significant volatility and intermittency; the operation mode of energy storage systems is diverse, and factors such as charging and discharging efficiency and lifespan management need to be considered; the responsiveness and willingness of controllable loads also vary. These factors together constitute a highly complex and uncertain environment for the operation of VPP. (2) Complexity and uncertainty of system operation: The optimal scheduling of VPP involves multiple time scales such as day-ahead, intraday, and real-time. How to achieve effective coordination between different scales is a difficult problem. At the same time, fluctuations in electricity market prices, deviations in load demand forecasting, uncertainty in the forecasting of renewable energy output, and potential equipment failures all pose huge challenges to the stable and economical operation of VPP. (3) Modeling and control challenges: Traditional VPP control relies on accurate mathematical modeling of the massive number of internal DERs. However, the number of DERs is huge, the types are diverse, and the parameters are time-varying, making it extremely difficult to establish an accurate and unified aggregation model. Furthermore, traditional centralized optimization control methods suffer from high computational complexity, heavy communication burden, and single-point failure risk when dealing with large-scale VPPs, making it difficult to meet the needs of real-time regulation.

[0004] Traditional mathematical programming-based optimization methods (such as mixed-integer programming, robust optimization, and stochastic programming) are theoretically mature, but they are highly dependent on the accuracy of the model and struggle to effectively handle the inherent strong uncertainty and high nonlinearity of VPP. While some existing hierarchical or distributed control architectures have alleviated the drawbacks of centralized control to some extent, there is still significant room for improvement and innovation in areas such as the collaborative efficiency among agents, the accuracy of upper-level instruction decomposition, and the consistency between the local autonomy of lower-level agents and global objectives. Summary of the Invention

[0005] The purpose of this application is to provide a collaborative control method, system, and electronic device for virtual power plants. Based on the collaborative control strategy of virtual power plants using HMADRL (Hierarchical Multi-Agent Deep Reinforcement Learning), and combined with the application of intelligent fusion terminals at the edge of the virtual power plant, the method achieves the sinking of control strategies and edge intelligence. Through hierarchical learning and collaborative mechanisms, it overcomes the limitations of traditional control methods in terms of modeling accuracy and computational efficiency, thereby at least solving some of the problems in the background art.

[0006] To achieve the above objectives, this application provides a virtual power plant collaborative control method, comprising: acquiring a control objective and selecting global state parameters affecting the control objective based on the control objective; calculating fluctuation coefficients based on historical data of the global state parameters, determining an uncertainty set, and using it as a hard constraint for decision-making by a higher-level agent based on the control objective; establishing a robust optimization model based on the uncertainty set, and solving the robust optimization model to obtain the optimal solution of the robust objective function; training a Bayesian deep reinforcement learning model using the historical data, optimizing variational parameters, with the training objective being to maximize the lower bound of evidence; inputting the current state data of the global state parameters into the trained Bayesian deep reinforcement learning model, generating an optimal action, and calculating the corresponding confidence level; and, based on the relationship between the confidence level and a preset confidence threshold, selecting one of the optimal action and the optimal solution of the robust objective function to generate a first control command and issuing it to a lower-level agent.

[0007] Preferably, the method further includes: a low-level agent converting the first control instruction into a second control instruction that can be executed by the distributed resource based on local state parameters, and sending it to the distributed resource; the low-level agent obtaining the execution effect of the second control instruction on the distributed resource and feeding it back to the high-level agent.

[0008] Preferably, the high-level intelligent agent is deployed on the cloud platform or regional coordination and control center of the virtual power plant, and the low-level intelligent agent is deployed at the edge of the distributed resource.

[0009] Preferably, the feedback to the high-level intelligent agent includes: updating historical data based on the feedback data, periodically adjusting the variational parameters of the Bayesian deep reinforcement learning model, and periodically adjusting the fluctuation coefficient of the robust optimization.

[0010] Preferably, each virtual power plant is equipped with a single high-level intelligent agent; the number of low-level intelligent agents in each virtual power plant is determined according to the number of jurisdictions divided by the user, and is constrained by the number of intelligent fusion terminals on the edge side.

[0011] Preferably, if a low-level agent belongs to a low-level agent cluster, the low-level agent cluster to which the low-level agent belongs receives and jointly responds to the first control command whose target is the low-level agent.

[0012] Preferably, the global state parameters include: market price parameters, grid demand parameters, and virtual power plant aggregated resource parameters; the market price parameters include: market quotations, historical quotations, market rules, and expected quotations; the grid demand parameters include: overall power generation plan, overall power consumption plan, operating costs, operating revenue, absorption strategy, charging and discharging dispatching strategy, and controllable load coordination strategy; the virtual power plant aggregated resource parameters include: the resource distribution, resource status, and resource allocation of the virtual power plant.

[0013] Preferably, the first control command includes an action command and a target object; the action command includes one of the following: overall power curve, cluster frequency modulation capacity, market participation parameters, regional coordination target, and economic dispatch signal; the target object includes: a low-level agent or a cluster of low-level agents.

[0014] Preferably, the second control command includes an action command, which has an underlying communication protocol and data format that can be recognized and executed by the corresponding distributed resources.

[0015] Preferably, the local status parameters include: the real-time operating status of the local distributed resources, local environment and prediction information, the deviation between the predicted value and the actual value, the constraints of the local power grid, and the operating constraints of the equipment itself.

[0016] Preferably, the high-level agent uses a deep reinforcement learning model to generate a first control instruction based on global state parameters; when managing a single distributed resource or multiple loosely coupled distributed resources, the low-level agent uses its own independent deep reinforcement learning model to convert the first control instruction into a second control instruction that the distributed resource can execute based on local state parameters; when managing multiple tightly coupled distributed resources, the low-level agent uses a multi-agent reinforcement learning model to convert the first control instruction into a second control instruction that the distributed resource can execute based on local state parameters.

[0017] Preferably, the high-level intelligent agent employs a deep reinforcement learning model to generate a first control command based on global state parameters, including: the high-level intelligent agent employs a deep reinforcement learning model to generate corresponding first control commands based on different global state parameters according to different decision stages; the decision stages include a real-time control stage, a short-term prediction stage, and a long-term prediction stage.

[0018] Preferably, after the lower-level agent obtains the execution effect of the second control instruction on the distributed resources and feeds it back to the higher-level agent, the method further includes: the higher-level agent evaluating the effectiveness of the strategy based on the feedback execution effect and the key operating state of the lower-level agent and calculating a reward signal, and / or incorporating the feedback execution effect and the key operating state of the lower-level agent into the state space of the next decision cycle of the higher-level agent.

[0019] Preferably, the low-level agent is further configured to generate a third control instruction targeting other low-level agents or a cluster of low-level agents, the third control instruction being used for neighboring region coordination, resource sharing, and local balance optimization.

[0020] This application also provides a virtual power plant collaborative control system, the system comprising a high-level intelligent agent and a low-level intelligent agent; the high-level intelligent agent and the low-level intelligent agent are deployed separately but connected by communication; the high-level intelligent agent is configured to: generate a first control command based on global state parameters and send it to the low-level intelligent agent; the low-level intelligent agent is configured to: convert the first control command into a second control command that can be executed by distributed resources based on local state parameters and send it to the distributed resources; and obtain the execution effect of the second control command on the distributed resources and feed it back to the high-level intelligent agent.

[0021] This application also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the at least one processor implements the aforementioned virtual power plant collaborative control method by executing the instructions stored in the memory.

[0022] This application also provides a machine-readable storage medium storing instructions that, when executed by a processor, configure the processor to perform the aforementioned virtual power plant collaborative control method.

[0023] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned virtual power plant collaborative control method.

[0024] The above technical solution has the following beneficial effects:

[0025] This application provides a collaborative control strategy for virtual power plants based on hierarchical multi-agent deep reinforcement learning. By combining the application of intelligent fusion terminals at the edge of the virtual power plant, the control strategy is brought down to the edge and intelligent at the edge. Through hierarchical learning and collaborative mechanisms, the limitations of traditional control methods in terms of modeling accuracy and computational efficiency are overcome. This enables multi-objective collaborative optimization of virtual power plants in terms of macro-energy balance, market interaction strategies and micro-distributed resource (DER) fine-grained operation control, so as to cope with the ever-changing power market environment and grid operation conditions.

[0026] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description

[0027] The accompanying drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the following detailed description to explain the embodiments of this application, but do not constitute a limitation on the embodiments of this application. In the drawings:

[0028] Figure 1 The schematic diagram illustrates the steps of the virtual power plant collaborative control method according to the embodiments of this application;

[0029] Figure 2 The illustration schematically shows a framework diagram of the implementation of the virtual power plant collaborative control method according to the embodiments of this application;

[0030] Figure 3 A schematic diagram illustrates the scheduling decision-making flowchart of the virtual power plant collaborative control method according to an embodiment of this application;

[0031] Figure 4 A schematic diagram of the structure of the virtual power plant collaborative control system according to an embodiment of this application is shown.

[0032] Figure 5 The diagram schematically illustrates the internal structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0033] The specific embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the embodiments of this application.

[0034] Figure 1 The diagram illustrates the steps of a virtual power plant collaborative control method according to an embodiment of this application. For example... Figure 1 As shown, a virtual power plant collaborative control method includes:

[0035] Obtain the control objective and select global state parameters that affect the control objective based on the control objective; here, the control objective refers to the objective that the virtual power plant collaborative control expects to achieve.

[0036] The fluctuation coefficient is calculated based on the historical data of the global state parameters, and the uncertainty set is determined, which is used as the hard constraint for the high-level intelligent agent to make decisions based on the control objective.

[0037] A robust optimization model is established based on the aforementioned uncertain set, and the optimal solution of the robust objective function is obtained by solving the robust optimization model.

[0038] The historical data was used to train a Bayesian deep reinforcement learning model, and the variational parameters were optimized. The training objective was to maximize the lower bound of evidence.

[0039] The current state data of the global state parameters are input into the trained Bayesian deep reinforcement learning model to generate the optimal action and calculate the corresponding confidence.

[0040] Based on the relationship between the confidence level and the preset confidence threshold, one of the optimal action and the optimal solution of the robust objective function is selected to generate a first control command and send it to the lower-level agent.

[0041] Considering that market price fluctuations and forecasting errors in photovoltaic / wind power output can cause the global strategy of HLA to fail, robust optimization is introduced and Bayesian DRL is used to surrogate traditional DRL. Historical data is used to statistically analyze the fluctuation range of uncertainty parameters, constructing an uncertainty set, which serves as a hard constraint for HLA decision-making, ensuring that the global strategy of HLA remains feasible even in the worst-case scenario. An example of its implementation is as follows:

[0042] 1. Uncertainty Boundary Modeling (Robust Optimization)

[0043] 1) Definition of uncertainty parameters

[0044] Let the historical data sample set be , ( N For historical data volume, t The fluctuation boundary is determined based on the 3σ principle (99.7% confidence interval) for the time step:

[0045] Market electricity price fluctuations are as follows: (1)

[0046] in, Electricity price forecast, electricity price fluctuation coefficient .

[0047] The fluctuation in photovoltaic power output is as follows: (2)

[0048] Among them, the photovoltaic power output fluctuation coefficient

[0049] Wind power output fluctuation is: (3)

[0050] Among them, wind power output fluctuation coefficient .

[0051] 2) Uncertainty set and robust objective function

[0052] Define the set of uncertainties in HLA decision-making. Includes all possible parameter combinations:

[0053] .

[0054] HLA's core decision-making is based on the objective of "bidding quantity Q" within a set of uncertainties. The goal is to maximize the return in the worst-case scenario. The robust optimization model is as follows:

[0055] .

[0056] in: The operating cost of HLA is subject to the following constraints: .

[0057] To maximize the aggregation output, , For energy storage discharge power, For the cost of energy storage and discharge, For operating costs, This represents the load power.

[0058] 2. Bayesian DRL (Deep Reinforcement Learning)

[0059] 1) Bayesian DRL Network Training

[0060] Policy network parameters Following a probability distribution (traditional DRL uses deterministic parameters), variational Bayesian inference is used to approximate the posterior parameters. Using variational distribution The posterior fit is fitted, and the training objective is to maximize the lower bound of evidence.

[0061] (4)

[0062] in, For the likelihood term (strategy) The fit on historical data D, where s is the global state and a is the first control command.

[0063] KL divergence is used to constrain the variational distribution and priors. Differences should be considered to avoid overfitting.

[0064] Using variational distribution Gaussian distribution , If the parameter is a variational parameter, then equation (4) is modified as follows:

[0065] (5)

[0066] in, For historical state-action pairs, This is the global state. This is the first control instruction in history, where M is the batch size. The parameters are prior.

[0067] 2) Calculation of strategy confidence

[0068] Regarding the current state By sampling strategy parameters multiple times, ~ Generate k candidate actions …,

[0069] The confidence level is: ,

[0070] in, The variance is given, and its range is given. .

[0071] 3) Strategy Triggering Logic

[0072] Set the reliability threshold to ,like Execute the optimal Bayesian DRL action. ;like The optimal solution for executing the robust objective function:

[0073] .

[0074] The following implementation method takes intraday market bidding as an example, and its overall steps are illustrated below:

[0075] Step 1: Calculate the volatility coefficient and determine the uncertainty set based on historical data of global state parameters (market price, photovoltaic / wind power output). .

[0076] Step 2: Train a Bayesian DRL model using historical data D and optimize the variational parameters. ,make convergence.

[0077] Step 3: HLA obtains the current global state Bayesian DRL sampling of K sets of parameters Generate candidate bidding electricity Calculate confidence level .

[0078] Step 4: If Output optimal decision ,like Execute the optimal solution of the robust objective function and output the suboptimal decision. ,in, Let be the expected function.

[0079] Step 5: Send the bid electricity amount as the "first control command" to the LLA.

[0080] Step 6: LLA provides feedback on the execution results.

[0081] Step 7: Update the historical dataset D based on the feedback data and periodically adjust the DRL variational parameters. And the robust optimization of the volatility coefficient α.

[0082] Through the above implementation methods, the intelligence level and accuracy of HLA decision-making have been improved.

[0083] In some embodiments of this application, addressing the challenges of resource diversity, operational uncertainty, and the difficulty of traditional control methods in simultaneously achieving global optimization and fine-grained local control within virtual power plants (VPPs), this application proposes a hierarchical multi-agent deep reinforcement learning (HMADRL) collaborative control framework. This framework decouples the complex VPP control problem into two interconnected but clearly defined layers: a higher layer responsible for macro-level policy planning, and a lower layer responsible for micro-level policy execution. Specifically, the method includes:

[0084] The higher-level intelligent agent generates a first control command based on global state parameters and sends it to the lower-level intelligent agent. The specific implementation method can adopt the steps described above.

[0085] The lower-level intelligent agent converts the first control instruction into a second control instruction that can be executed by the distributed resources based on the local state parameters, and then sends it to the distributed resources.

[0086] The lower-level intelligent agent obtains the execution effect of the second control command on the distributed resources and feeds it back to the higher-level intelligent agent.

[0087] In the above implementation, the first step is to construct a two-tier intelligent control architecture consisting of a single high-level agent (HLA) and multiple low-level agents (LLAs). This structured design aims to effectively decompose and manage the complex decision-making process of the virtual power plant (VPP), improving the system's scalability and robustness. These three steps roughly correspond to the steps of goal assignment, instruction execution and local optimization, and feedback and learning.

[0088] In some optional implementations, the feedback to the higher-level agent includes: updating historical data based on the feedback data, and periodically adjusting the variational parameters of the Bayesian deep reinforcement learning model and the fluctuation coefficient of the robust optimization. That is, the method also includes the aforementioned Step 7.

[0089] Figure 2 The illustration schematically shows a framework diagram of the implementation of the virtual power plant collaborative control method according to an embodiment of this application. For example... Figure 2 As shown, in this embodiment, the high-level intelligent agents are deployed on the cloud platform or regional coordination and control center of the virtual power plant, while the low-level intelligent agents are deployed at the edge near the distributed resources. HLAs are typically deployed on the cloud platform or regional coordination and control center of the VPP, undertaking global, long-term (e.g., 24 to 48 hours) macro-level strategy formulation tasks. They focus on the overall energy balance of the VPP, market interaction benefits, and response to commands from the upper-level grid. LLAs are deployed at the edge near the distributed resources (DERs), specifically relying on intelligent converged terminals. Each LLA is responsible for the local, short-term, fine-grained control and optimization of one or more DERs within its jurisdiction.

[0090] In some embodiments of this application, each virtual power plant is equipped with a single high-level intelligent agent; the number of low-level intelligent agents in each virtual power plant is determined according to the number of jurisdictions divided by the user, and is constrained by the number of intelligent converged terminals on the edge side. For the sake of control uniformity within the virtual power plant, each virtual power plant is equipped with a single high-level intelligent agent. LLA deployment locations include: the core logic of the LLA runs on intelligent converged terminals installed at various DER access points or small DER aggregation points. One intelligent converged terminal can carry one LLA, therefore the upper limit of the number of LLAs can be determined based on the number of intelligent converged terminals. The LLA manages one or more DERs directly connected to it.

[0091] In some implementations, a low-level agent cluster composed of low-level agents is also provided. If multiple converged terminals and their managed DERs are geographically adjacent or functionally closely coupled, they can also form an LLA cluster to jointly respond to HLA commands. If a low-level agent belongs to a certain low-level agent cluster, the low-level agent cluster to which the low-level agent belongs receives and jointly responds to the first control command whose target is the low-level agent.

[0092] This implementation describes the aforementioned global state parameters. These global state parameters include: market price parameters, grid demand parameters, and virtual power plant aggregated resource parameters. The core responsibilities of the HLA include: market interaction strategy formulation, macro-level energy balance and economic dispatch, and regional / cluster target setting and instruction issuance. Market interaction strategy formulation includes: the HLA is responsible for representing the entire VPP in various electricity markets, such as the energy market and ancillary service markets (frequency regulation, reserve, etc.). It needs to intelligently formulate bidding strategies, bid curves, and declared capacity based on market price parameters and the VPP's internal adjustable resources. The market price parameters include: market bids, historical bids, market rules, and expected bids. Macro-level energy balance and economic dispatch include: one of the core tasks of the HLA is to formulate the VPP's overall power generation and consumption plans to ensure internal energy supply and demand balance and strive for the lowest operating costs or maximum revenue. This includes optimizing renewable energy consumption strategies (such as coordinating energy storage charging to absorb excess photovoltaic / wind power), managing the charging and discharging scheduling of energy storage systems at the global level (such as participating in peak-valley arbitrage), and coordinating the response of controllable loads. Therefore, grid demand parameters include: overall generation plan, overall electricity consumption plan, operating costs, operating revenue, consumption strategy, charging and discharging scheduling strategy, and controllable load coordination strategy. Virtual power plant aggregated resource parameters include: virtual power plant resource distribution, resource status, and resource allocation.

[0093] In some embodiments of this application, although the first control command of the HLA is issued to the LLA or LLAs, it is actually aimed at the DER or DER aggregate managed by one or more LLAs. DER aggregates can be divided by geographical location, such as all DERs within an industrial park, or by resource type, such as all adjustable air conditioning clusters. The first control command includes action instructions and target objects; the action instructions include: Overall power curve: the total active power output / consumption curve and reactive power curve that the VPP should achieve in one or more scheduling cycles (e.g., every 15 minutes in the next 24 hours). Cluster frequency regulation capacity: the required up / down frequency regulation capacity provided by the k-th DER cluster, such as an energy storage cluster, during a specific time period. Market participation parameters: for example, the bidding volume and corresponding price for a specific time period in the day-ahead energy market. Regional coordination target: for example, to solve the voltage over-limit problem of a local power grid, requiring LLAs in a certain region to coordinate the reactive power output of the DERs under their jurisdiction to achieve a regional reactive power compensation target or voltage control target. Economic dispatch signals: Send internal "shadow prices" or incentive signals to LLAs to guide them to spontaneously optimize in the direction of reducing the total cost of VPP while meeting power balance requirements.

[0094] In some embodiments of this application, the second control command includes action commands, which have underlying communication protocols and data formats that can be recognized and executed by the corresponding distributed resources. These action commands vary depending on the distributed resources. Examples of actions output by the LLA are as follows: For photovoltaic inverter j: active power setpoint, reactive power setpoint, power factor setting, or reference voltage value in voltage control mode. For energy storage system k: charging / discharging power (positive for discharging, negative for charging), or charging / discharging current, and operating mode (e.g., peak shaving and valley filling mode, smoothing output mode, command tracking mode). For controllable load m (e.g., air conditioner, water heater): start / stop command, power adjustment percentage, temperature setpoint adjustment. For electric vehicle charging / discharging pile n: maximum allowed charging power, discharge power command in V2G (vehicle-to-grid) mode. These action commands are converted by the fusion terminal into underlying communication protocols and data formats that can be recognized and executed by the corresponding DER.

[0095] In some embodiments of this application, the local state parameters include: real-time operating status of local distributed resources, such as: instantaneous power generation of photovoltaic arrays, inverter efficiency, SOC, temperature, state of health (SOH) of energy storage units, current energy consumption level and adjustable range of controllable loads, etc.; local environmental and forecast information, such as: local irradiance and ambient temperature collected by the fusion terminal, and ultra-short-term (e.g., 5-60 minutes) DER output forecasts and load forecasts generated based on these; deviation between predicted and actual values, such as obtaining prediction deviations through real-time tracking; constraints of the local power grid, such as: whether the voltage at the local grid connection point is within the allowable range, and whether the feeder power flow is close to the limit, provided that the smart fusion terminal can perceive it; and equipment operating constraints, such as: upper and lower limits of SOC safety for energy storage, maximum charge and discharge rate, daily cycle limit, inverter power factor adjustment range, ramp rate limit, etc.

[0096] In some embodiments of this application, within the HLA, the HLA acquires global state parameters and Deep Reinforcement Learning (DRL) model outputs based on its perception of the global state, and formulates macro-level operational goals or guiding strategies. The agents operating in each LLA, or rather each intelligent converged terminal, typically employ deep reinforcement learning algorithms to learn their control strategies. The specific algorithm selection depends on the complexity and coupling degree of the DERs it manages. For example, if the converged terminal manages one DER or the coupling between the managed multiple DERs is weak—that is, the optimal strategy for controlling one DER without significantly affecting other DERs—a single-agent DRL algorithm can be used. If the converged terminal manages multiple coupled DERs, such as a power generation, storage, and load unit within a microgrid, or an EV cluster requiring coordinated charging and discharging, a multi-agent reinforcement learning algorithm is used. In this case, the converged terminal may contain multiple sub-agents, each corresponding to one DER, which learn and act collaboratively within the framework of the LLA. Examples of input information for the single-agent DRL algorithm or multi-agent reinforcement learning algorithm include: instructions / goals from the HLA: for example, instructions assigned to the i-th LLA, i.e., the LLA... i The target power, or the capacity participating in frequency modulation. The local j-th DER, i.e., DER j Real-time status: Actual output of photovoltaic power, actual output of wind turbines, SOC, voltage, and current of energy storage k, and current power and adjustable status of controllable load m. Local comprehensive forecast information: Local aggregated photovoltaic output forecast, local aggregated load forecast, and deviations of these forecasts over a past period. Local grid information: Where available, this includes grid connection point voltage, frequency, etc. If the fusion terminal has corresponding measurement or estimation capabilities, it also includes local line power flow.

[0097] In some embodiments of this application, after the lower-level agent obtains the execution effect of the second control command on the distributed resources and feeds it back to the higher-level agent, the method further includes: the execution effects of LLAs, such as actual aggregated output and target tracking errors, as well as local key operating states such as resource adjustability potential and abnormal events, are processed and then uploaded to HLA as feedback information. This feedback information is used by HLA to evaluate the effectiveness of its strategy and calculate reward signals, and also becomes part of the state space of HLA's next decision cycle. Through this continuous interaction and feedback, both HLA and LLAs can continuously optimize their respective policy models, thereby improving the collaborative performance and adaptive capability of the entire VPP system.

[0098] In some embodiments of this application, the low-level agent is further configured to generate a third control command targeting other low-level agents or clusters of low-level agents. This third control command is used for neighboring region coordination, resource sharing, and local balance optimization. In the aforementioned system including high-level and low-level agents, in addition to the interaction between high-level and low-level agents, an interaction mechanism between low-level agents is also designed. That is, low-level agents interact with other low-level agents or clusters of low-level agents through the third control command to achieve horizontal collaboration. This third control command achieves the aforementioned horizontal collaboration through neighboring region coordination, resource sharing, and local balance optimization.

[0099] In some embodiments of this application, the high-level agent employs a deep reinforcement learning model to generate a first control command based on global state parameters. This includes: the high-level agent using a deep reinforcement learning model to generate corresponding first control commands based on different global state parameters for different decision-making stages; the decision-making stages include a real-time control stage, a short-term prediction stage, and a long-term prediction stage. Due to the changes in the first control command, the second control command on the LLA side also changes accordingly. For example, using "day" as the granularity, the short-term prediction stage is also called the intraday stage, the long-term prediction stage is selected as the previous day and is also called the pre-day stage, and the real-time control stage is also called the real-time stage. Figure 3 A schematic diagram illustrating the scheduling decision-making flowchart of the virtual power plant collaborative control method according to an embodiment of this application is shown. Figure 3 As shown, in the day-ahead phase, the system mainly conducts long-term forecasting and planning, and formulates market participation strategies; in the intraday phase, it performs rolling optimization and plan adjustments; in the real-time phase, it executes precise control and rapid response, and continuously optimizes system performance through continuous execution monitoring, feedback and learning.

[0100] The detailed process is described below:

[0101] 1. Day-ahead Phase. HLA Data Acquisition and Forecasting: The high-level intelligent agent collects historical data, meteorological data, market price data, and grid dispatch information across the entire VPP network, and uses deep learning models to forecast day-ahead load, renewable energy output, and electricity prices. HLA Day-ahead Strategy Formulation: Based on the forecast results, HLA formulates energy balance strategies and market participation strategies, comprehensively considering multiple dimensions such as economic efficiency, reliability, and environmental protection. HLA Target Distribution: HLA decomposes the initial control target into specific target values ​​for each LLA cluster and distributes them to each intelligent converged terminal through a secure communication channel. LLA Resource Pre-allocation: After receiving the target, each LLA, combined with its local DER characteristics and constraints, performs preliminary resource allocation planning and formulates a day-ahead execution plan.

[0102] 2. Intraday Phase. HLA Rolling Optimization: Based on updated short-term forecasts, changes in market information, and the actual operating status of the system, HLA performs rolling optimization every hour, dynamically adjusting the VPP operating strategy. HLA Target Update: HLA calculates and updates the control targets for each LLA cluster, responding promptly to changes in the external environment and internal status. LLA Plan Adjustment: Each LLA re-optimizes its local DER output plan based on updated target values ​​and the latest local short-term forecasts, ensuring efficient execution of the targets issued by HLA.

[0103] 3. Real-time Stage. HLA Global Monitoring: The HLA monitors the overall operating status of the VPP in real time, including key parameters such as grid frequency, voltage, and power balance, and performs short-term scheduling at the minute level. LLA Fine-grained Control: Based on real-time status monitoring and commands issued by the HLA, combined with local ultra-short-term forecasts, the LLA performs fine-grained control of distributed energy resources at the second level. LLA Autonomous Response: When local random events occur, the LLA can make autonomous response decisions without waiting for HLA commands, improving the system's ability to cope with random events. DER Execution and Feedback: Distributed energy devices execute control commands, and the execution effect is monitored in real time through intelligent converged terminals and fed back to the LLA and HLA.

[0104] Online model updates: HLA and LLA periodically update their respective reinforcement learning models based on performance evaluation and historical data accumulation, forming a closed-loop optimization.

[0105] Key data interaction descriptions. Bottom-up status reporting: Each DER device → LLA: Real-time operating status, adjustable capacity, constraints. LLA → HLA: Regional aggregate status, execution capability assessment, anomaly reports. Top-down control: HLA → LLA: Overall control objectives, adjustment strategies, priority settings; LLA → DER: Precise control commands, parameter settings, operating mode switching. Horizontal collaboration: LLA LLA: Neighboring area coordination, resource sharing, and local balance optimization. Closed-loop feedback: Execution result → LLA → HLA: Instruction execution effect, deviation analysis, and model optimization signals.

[0106] Considering that market price fluctuations and photovoltaic / wind power output forecasting errors can cause HLA's global strategy to fail, robust optimization is introduced and Bayesian DRL is used to replace traditional DRL. By using historical data to statistically analyze the fluctuation range of uncertainty parameters, an uncertainty set is constructed and used as a hard constraint for HLA decision-making, ensuring that HLA's global strategy remains feasible in the worst-case scenario.

[0107] As can be seen from the above implementation methods, the implementation methods of this application provide a VPP collaborative control strategy based on hierarchical multi-agent deep reinforcement learning. Combined with the application of intelligent fusion terminals at the edge of the VPP, the control strategy is pushed down and edge intelligence is realized. Through hierarchical learning and collaborative mechanism, the limitations of traditional control methods in terms of modeling accuracy and computational efficiency are overcome. The VPP achieves multi-objective collaborative optimization between macro-energy balance, market interaction strategy and micro-level DER refined operation control, so as to cope with the ever-changing power market environment and grid operation conditions.

[0108] Based on the same inventive concept, this application also provides a virtual power plant collaborative control system. Figure 4 A schematic diagram illustrating the structure of a virtual power plant collaborative control system according to an embodiment of this application is shown. Figure 4As shown, the system includes a high-level intelligent agent and a low-level intelligent agent; the high-level intelligent agent and the low-level intelligent agent are deployed separately but connected by communication; the high-level intelligent agent is configured to: generate a first control command based on global state parameters and send it to the low-level intelligent agent, preferably by performing the aforementioned steps: obtaining a control objective, and selecting global state parameters that affect the control objective based on the control objective; calculating the fluctuation coefficient based on historical data of the global state parameters, determining the uncertainty set, and using it as a hard constraint for the high-level intelligent agent's decision based on the control objective; establishing a robust optimization model based on the uncertainty set, solving the robust optimization model to obtain the optimal solution of the robust objective function; and training a Bayesian algorithm using the historical data. A deep reinforcement learning model is used to optimize variational parameters, with the training objective being to maximize the lower bound of evidence. The current state data of the global state parameters are input into the trained Bayesian deep reinforcement learning model to generate an optimal action and calculate the corresponding confidence level. Based on the relationship between the confidence level and a preset confidence threshold, one of the optimal action and the optimal solution of the robust objective function is selected to generate a first control command, which is then sent to a low-level agent. The low-level agent is configured to: convert the first control command into a second control command executable by distributed resources based on local state parameters and send it to the distributed resources; and obtain the execution effect of the second control command on the distributed resources and feed it back to the high-level agent.

[0109] In some alternative implementations, the high-level intelligent agents are deployed on the cloud platform of the virtual power plant or in a regional coordination and control center, while the low-level intelligent agents are deployed at the edge of the distributed resources.

[0110] In some alternative implementations, each virtual power plant is equipped with a single high-level intelligent agent; the number of low-level intelligent agents in each virtual power plant is determined according to the number of jurisdictions divided by the user and is constrained by the number of intelligent fusion terminals on the edge side.

[0111] In some alternative implementations, if a low-level agent belongs to a cluster of low-level agents, the cluster of low-level agents to which the low-level agent belongs receives and jointly responds to the first control command whose target is the low-level agent.

[0112] In some optional implementations, the global state parameters include: market price parameters, grid demand parameters, and virtual power plant aggregated resource parameters; the market price parameters include: market quotations, historical quotations, market rules, and expected quotations; the grid demand parameters include: overall power generation plan, overall power consumption plan, operating costs, operating revenue, absorption strategy, charge and discharge dispatch strategy, and controllable load coordination strategy; the virtual power plant aggregated resource parameters include: the resource distribution, resource status, and resource allocation of the virtual power plant.

[0113] In some alternative implementations, the first control command includes an action command and a target object; the action command includes one of the following: overall power curve, cluster frequency modulation capacity, market participation parameters, regional coordination target, and economic dispatch signal; the target object includes: a low-level agent or a cluster of low-level agents.

[0114] In some alternative implementations, the second control command includes an action command having an underlying communication protocol and data format that can be recognized and executed by the corresponding distributed resources.

[0115] In some optional implementations, the local status parameters include: the real-time operating status of the local distributed resources, local environment and prediction information, the deviation between the predicted value and the actual value, the constraints of the local power grid, and the operating constraints of the equipment itself.

[0116] In some alternative implementations, the high-level agent employs a deep reinforcement learning model to generate a first control instruction based on global state parameters; when managing a single distributed resource or multiple loosely coupled distributed resources, the low-level agent employs its own independent deep reinforcement learning model to convert the first control instruction into a second control instruction that the distributed resource can execute based on local state parameters; when managing multiple tightly coupled distributed resources, the low-level agent employs a multi-agent reinforcement learning model to convert the first control instruction into a second control instruction that the distributed resource can execute based on local state parameters.

[0117] In some optional implementations, the high-level intelligent agent employs a deep reinforcement learning model to generate a first control command based on global state parameters, including: the high-level intelligent agent employs a deep reinforcement learning model to generate corresponding first control commands based on different global state parameters according to different decision stages; the decision stages include a real-time control stage, a short-term prediction stage, and a long-term prediction stage.

[0118] In some optional implementations, after the lower-level agent obtains the execution effect of the second control instruction on the distributed resources and feeds it back to the higher-level agent, the system further includes: the higher-level agent evaluating the effectiveness of the strategy based on the feedback execution effect and the key operating state of the lower-level agent and calculating a reward signal, and / or incorporating the feedback execution effect and the key operating state of the lower-level agent into the state space of the next decision cycle of the higher-level agent.

[0119] In some alternative implementations, the low-level agent is further configured to generate third control instructions targeting other low-level agents or clusters of low-level agents, the third control instructions being used for neighboring region coordination, resource sharing, and local balance optimization.

[0120] The specific limitations of each functional module in the aforementioned virtual power plant collaborative control system can be found in the limitations of the virtual power plant collaborative control method described above, and will not be repeated here. Each module in the system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independently of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module. It is also based on hierarchical multi-agent deep reinforcement learning technology, realizing the sinking of control strategies and edge intelligence. Through hierarchical learning and collaborative mechanisms, it overcomes the limitations of traditional control methods in terms of modeling accuracy and computational efficiency.

[0121] In some embodiments of this application, an electronic device is also provided, comprising: at least one processor; and a memory connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which executes the aforementioned virtual power plant collaborative control method. Its internal structure diagram can be shown as follows. Figure 5 As shown. Figure 5 The diagram schematically illustrates the internal structure of an electronic device according to an embodiment of this application. The electronic device includes a processor A01, a network interface A02, a memory (not shown), and a database (not shown) connected via a system bus. The processor A01 provides computing and control capabilities. The memory includes internal memory A03 and a non-volatile storage medium A04. The non-volatile storage medium A04 stores an operating system B01, a computer program B02, and a database (not shown). The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 stored in the non-volatile storage medium A04. The network interface A02 is used for communication with external terminals via a network connection. When the computer program B02 is executed by the processor A01, it implements a virtual power plant collaborative control method.

[0122] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0123] In one embodiment provided in this application, a machine-readable storage medium is provided, on which instructions are stored, which, when executed by a processor, cause the processor to be configured to perform the aforementioned virtual power plant collaborative control method.

[0124] In one embodiment provided in this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the aforementioned virtual power plant collaborative control method.

[0125] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0126] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0127] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0128] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0129] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0130] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0131] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0132] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0133] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for collaborative control of a virtual power plant, characterized in that, The method includes: Obtain the control target, and select global state parameters that affect the control target based on the control target; The fluctuation coefficient is calculated based on the historical data of the global state parameters, and the uncertainty set is determined, which is used as the hard constraint for the high-level intelligent agent to make decisions based on the control objective. A robust optimization model is established based on the aforementioned uncertain set, and the optimal solution of the robust objective function is obtained by solving the robust optimization model. The historical data was used to train a Bayesian deep reinforcement learning model, and the variational parameters were optimized. The training objective was to maximize the lower bound of evidence. The current state data of the global state parameters are input into the trained Bayesian deep reinforcement learning model to generate the optimal action and calculate the corresponding confidence. Based on the relationship between the confidence level and a preset confidence threshold, one of the optimal solutions to the optimal action and the robust objective function is selected to generate a first control command and issue it to the lower-level agent; wherein, selecting one of the optimal solutions to the optimal action and the robust objective function includes: setting a confidence threshold of... If confidence level Execute the optimal Bayesian DRL action; if Execute the optimal solution of the robust objective function; The high-level intelligent agents are deployed on the cloud platform of the virtual power plant or in the regional coordination and control center, while the low-level intelligent agents are deployed at the edge of the distributed resources.

2. The method according to claim 1, characterized in that, The method further includes: a low-level agent converting the first control instruction into a second control instruction that can be executed by the distributed resource based on local state parameters, and then sending it to the distributed resource; The lower-level intelligent agent obtains the execution effect of the second control command on the distributed resources and feeds it back to the higher-level intelligent agent.

3. The method according to claim 1, characterized in that, Feedback is sent to the higher-level intelligent agent, including: updating historical data based on the feedback data, periodically adjusting the variational parameters of the Bayesian deep reinforcement learning model, and periodically adjusting the fluctuation coefficient of the robust optimization.

4. The method according to claim 2, characterized in that, Each virtual power plant has a single high-level intelligent agent; the number of low-level intelligent agents in each virtual power plant is determined according to the number of jurisdictions divided by the user, and is constrained by the number of intelligent fusion terminals on the edge side.

5. The method according to claim 2, characterized in that, If a low-level agent belongs to a certain low-level agent cluster, then the low-level agent cluster to which the low-level agent belongs receives and jointly responds to the first control command whose target is the low-level agent.

6. The method according to claim 2, characterized in that, The global state parameters include: market price parameters, power grid demand parameters, and virtual power plant aggregated resource parameters; The market price parameters include: market quotes, historical quotes, market rules, and expected quotes; Grid demand parameters include: overall power generation plan, overall power consumption plan, operating costs, operating revenue, absorption strategy, charging and discharging dispatch strategy, and controllable load coordination strategy; The aggregated resource parameters of a virtual power plant include: the resource distribution, resource status, and resource allocation of the virtual power plant.

7. The method according to claim 5, characterized in that, The first control command includes an action command and a target object; The action instructions include one of the following: overall power curve, cluster frequency modulation capacity, market participation parameters, regional coordination objectives, and economic dispatch signals. The target objects include: low-level intelligent agents or clusters of low-level intelligent agents.

8. The method according to claim 2, characterized in that, The second control command includes action commands, which have underlying communication protocols and data formats that can be recognized and executed by the corresponding distributed resources.

9. The method according to claim 2, characterized in that, The local status parameters include: the real-time operating status of local distributed resources, local environment and prediction information, deviation between prediction and actual values, constraints of the local power grid, and operating constraints of the equipment itself.

10. The method according to claim 2, characterized in that, The high-level intelligent agent uses a deep reinforcement learning model to generate the first control command based on global state parameters; When a low-level intelligent agent manages a single distributed resource or multiple loosely coupled distributed resources, it uses its own independent deep reinforcement learning model to convert the first control instruction into a second control instruction that can be executed by the distributed resource based on the local state parameters. When managing multiple tightly coupled distributed resources, the low-level agent uses a multi-agent reinforcement learning model to convert the first control instruction into a second control instruction that can be executed by the distributed resources based on the local state parameters.

11. The method according to claim 10, characterized in that, The high-level intelligent agent employs a deep reinforcement learning model to generate the first control command based on global state parameters, including: The high-level intelligent agent uses a deep reinforcement learning model to generate corresponding first control commands based on different global state parameters at different decision stages. The decision-making stage includes a real-time control stage, a short-term forecasting stage, and a long-term forecasting stage.

12. The method according to claim 2, characterized in that, After the lower-level agent obtains the execution effect of the second control command on the distributed resources and feeds it back to the higher-level agent, the method further includes: The high-level agent evaluates the effectiveness of the strategy and calculates reward signals based on the feedback of the execution results and the key operational states of the low-level agent. And / or incorporate the feedback execution results and key operational states of lower-level agents into the state space of the next decision cycle of higher-level agents.

13. The method according to claim 7, characterized in that, The low-level agent is also configured to generate a third control instruction targeting other low-level agents or clusters of low-level agents, the third control instruction being used for neighboring region coordination, resource sharing, and local balance optimization.

14. A virtual power plant collaborative control system, characterized in that, The system includes high-level intelligent agents and low-level intelligent agents; the high-level intelligent agents and low-level intelligent agents are deployed separately but connected by communication. The high-level intelligent agent is configured to execute the virtual power plant collaborative control method as described in claim 1; The lower-level agent is configured to: convert the first control instruction into a second control instruction that can be executed by the distributed resource based on local state parameters, and send it to the distributed resource; and obtain the execution effect of the second control instruction on the distributed resource and feed it back to the higher-level agent.

15. An electronic device, characterized in that, include: At least one processor; A memory connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the at least one processor implements the virtual power plant collaborative control method according to any one of claims 1 to 13 by executing the instructions stored in the memory.

16. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the virtual power plant collaborative control method as described in any one of claims 1 to 13.

17. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the virtual power plant collaborative control method as described in any one of claims 1 to 13.

Citation Information

Patent Citations

  • Virtual power plant online optimization scheduling method based on deep reinforcement learning algorithm

    CN119204546A

  • Method and apparatus for constructing adjustable capacity of virtual power plant, electronic device, storage medium, program, and program product

    WO2024060413A1