Virtual power plant scheduling method and related device
By constructing an undirected graph and a deep reinforcement learning model, and combining market game theory and electrical connection constraints, the virtual power plant scheduling strategy is optimized, solving the reliability and accuracy problems of virtual power plant scheduling and achieving efficient scheduling in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CENT SOUTH UNIV
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-29
AI Technical Summary
Existing virtual power plant scheduling methods suffer from poor reliability and accuracy in complex dynamic environments. In particular, they struggle to meet real-time decision-making requirements when faced with massive heterogeneous resources and high uncertainty. Furthermore, existing methods are difficult to balance physical feasibility with strategy optimization.
By constructing an undirected graph and extracting topological relationships based on the graph feature matrix, a multi-objective reward function and an optimization objective function are designed. Combined with a deep reinforcement learning model, considering market game relations and electrical connection constraints, the scheduling strategy is optimized.
It improves the reliability and accuracy of virtual power plant dispatching, enabling millisecond-level response in complex environments and meeting the reliability requirements of market competition strategies and the feasibility of physical constraints.
Smart Images

Figure CN122114480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power plant control technology, and in particular to a virtual power plant dispatching method and related equipment. Background Technology
[0002] The penetration rate of distributed energy resources (DERs) in distribution networks has increased significantly. However, the inherent randomness and volatility of DERs make them difficult to directly adapt to the dispatching needs of traditional electricity markets, which has become a major bottleneck restricting their large-scale grid connection. Virtual power plants (VPPs), as an efficient energy aggregation and management platform, play a crucial role in building new power systems. Through advanced communication and control technologies, VPPs can not only effectively smooth out the fluctuations in renewable energy and improve system flexibility, but also fully tap the economic value of demand-side resources by participating in spot trading and ancillary service markets.
[0003] Despite the promising prospects of VPP (Virtual Power Utilization), its optimal scheduling in complex dynamic environments still faces numerous challenges. Traditional economic dispatch methods are mostly based on deterministic models or scenario analysis, such as Mixed Integer Linear Programming (MILP) and Stochastic Programming (SP). However, these model-driven methods have significant limitations in practical applications: First, facing massive heterogeneous resources and high uncertainty, scenario-based methods often lead to an exponential increase in computational complexity due to the need to characterize random variables, making it difficult to meet the timeliness requirements of real-time decision-making in the electricity spot market; second, robust optimization and other methods tend to adopt overly conservative strategies to avoid extreme risks, thus sacrificing potential market returns; furthermore, the non-convexity of the physical power flow constraints of the distribution network makes it difficult for traditional convex optimization algorithms to solve directly, while linear approximation may introduce physical biases, affecting the safe operation of the system.
[0004] Due to the shortcomings of traditional methods in terms of real-time performance and model dependence, Deep Reinforcement Learning (DRL) has been widely used in fields such as frequency regulation, economic dispatch, energy management and trading, thanks to its powerful high-dimensional perception capability and "model-free" adaptive characteristics.
[0005] While DRL provides a new paradigm for intelligent scheduling of VPPs, existing research largely focuses on solving single technical challenges, lacking a holistic approach that considers game-theoretic interaction, large-scale collaboration, and system security. Specifically, existing methods suffer from three core limitations in practical deployment: 1) Existing studies often assume that the VPP is a price taker or treat market prices as a static stochastic process. Although existing studies have solved the scheduling problem of continuous action space using the DDPG algorithm, they have neglected the dynamic game relationship between the VPP as a "price setter" and other market participants. This results in the strategy lacking robustness when facing unknown competitors and is prone to failure.
[0006] 2) Regarding resource coordination, existing multi-agent reinforcement learning (MARL) methods often neglect the physical topology of the power grid. Existing research in peer-to-peer transactions assumes a fully connected communication topology, ignoring the actual electrical connection constraints of the distribution network, making it difficult to implement coordination strategies in engineering.
[0007] 3) Existing methods struggle to strike a balance between physical feasibility and strategy optimization. Current research introduces a projection-based safety shield to enforce constraints on agent actions. While this ensures hard constraints, it cuts off gradient backpropagation, limiting the agent's exploration efficiency in boundary regions and making it prone to getting trapped in local optima. On the other hand, methods based on soft penalties struggle to strictly guarantee hard constraints and are sensitive to penalty coefficients, easily leading to training oscillations.
[0008] This shows that the current virtual power plant dispatching has problems with poor reliability and accuracy. Summary of the Invention
[0009] This application provides a virtual power plant scheduling method and related equipment, which can solve the problems of poor reliability and accuracy in virtual power plant scheduling.
[0010] In a first aspect, embodiments of this application provide a virtual power plant scheduling method, which includes: Acquire the power data and market data of the target virtual power plant, as well as the market data of multiple competing power plants, and construct a virtual power plant decision model; The virtual power plant decision-making model was initially trained using all market data to obtain the trained virtual power plant decision-making model. Construct an undirected graph of the target virtual power plant and extract the graph feature matrix based on the undirected graph; multiple nodes in the undirected graph correspond one-to-one with multiple power nodes in the target virtual power plant, and the edges between nodes represent the topological relationship between the corresponding two power nodes; Design an optimization objective function, and input the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model to solve the optimization objective function and obtain the scheduling strategy of the target virtual power plant; the optimization objective function is used to describe the scheduling constraints of the target virtual power plant. The target virtual power plant is scheduled and controlled based on the scheduling strategy.
[0011] Optionally, the virtual power plant decision-making model includes a state space, an action space, and a multi-objective reward function; The state space includes internal state, market state, and physical state; The action space includes the adjustment ratio of each aggregator in the virtual power plant; The multi-objective reward function is:
[0012] in, This represents the value of the multi-objective reward function. , , , , These are the weighting coefficients. As a conditional risk value penalty, Penalty for operating costs, For instruction tracking rewards, For physical security penalties, Rewards for market gains:
[0013]
[0014]
[0015]
[0016]
[0017] in, For energy market revenue, To assist market returns, For incremental operating costs, To adjust mileage costs, Indicates the number of aggregators. Indicates the tracking penalty coefficient. This indicates a higher-level adjustment instruction. This represents the actual power output of the aggregator. For critical path power flow. This is the upper limit of the line's transmission capacity. For the voltage amplitude at critical nodes, To allow for voltage deviation, For risk aversion coefficient, To be at confidence level Conditional Value at Risk (VaR) estimation This refers to the deviation between actual and expected returns.
[0018] Optionally, the virtual power plant decision-making model can be initially trained using all market data to obtain a trained virtual power plant decision-making model, including: Multiple market task environments are constructed based on market data of the target virtual power plant and market data of all competing power plants, and a strategy loss function is constructed accordingly. Using a virtual power plant decision model, decision inferences are performed and the value of the strategy loss function is calculated for each market task environment; For each market task environment, the temporary model parameters for that market task environment are calculated based on the value of the strategy loss function corresponding to that market task environment. The virtual power plant decision model is updated with parameters based on all temporary model parameters to obtain the trained virtual power plant decision model.
[0019] Optionally, based on the value of the strategy loss function corresponding to the market task environment, temporary model parameters under the market task environment are calculated, including: Through the formula:
[0020] Calculate the first Temporary model parameters for each market task environment ; in, The parameters represent the virtual power plant decision-making model. Indicates the inner learning rate. Indicates the first The value of the strategy loss function corresponding to each market task environment; The virtual power plant decision model is updated based on all temporary model parameters, including: Through the formula:
[0021]
[0022] Update the parameters of the virtual power plant decision model; in, This represents the model parameters of the trained virtual power plant decision-making model. This represents gradient operation. This represents the learning rate of the outermost element.
[0023] Optionally, the graph feature matrix can be extracted based on the undirected graph, including: Based on the undirected graph, extract the high-level feature representation of each node in the undirected graph; The high-level feature representations of all nodes are integrated into a matrix to obtain the graph feature matrix.
[0024] Optionally, based on the undirected graph, extract high-level feature representations for each node in the undirected graph, including: Through the formula:
[0025] Calculate the first The first layer update Feature representation of each node ; Among them, when hour, For the first High-level feature representation of each node, For the updated layer number, In an undirected graph, the first... A set of vector nodes of nodes. For activation function, Indicates the first The node and the first Normalized attention weights between nodes, Indicates the first Layer The updated node representation.
[0026] Optionally, the objective function to be optimized is:
[0027] in, This indicates the value of the objective function to be optimized. Represents the Lagrange multipliers. Represents risk variables, Indicates the first A constrained Lagrange multiplier, Indicates the first The expected violation of a constraint. This represents the risk aversion coefficient. This indicates that CVaR replaces the objective function.
[0028] Secondly, embodiments of this application provide a virtual power plant dispatching device, comprising: The acquisition module is used to acquire the power data and market data of the target virtual power plant, as well as the market data of multiple competing power plants, and to build a virtual power plant decision model. The training module is used to perform preliminary training on the virtual power plant decision-making model using all market data, resulting in the trained virtual power plant decision-making model. The construction module is used to construct an undirected graph of the target virtual power plant and extract the graph feature matrix based on the undirected graph; multiple nodes in the undirected graph correspond one-to-one with multiple power nodes in the target virtual power plant, and the edges between nodes represent the topological relationship between the corresponding two power nodes. The solution module is used to design the optimization objective function and input the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model to solve the optimization objective function and obtain the scheduling strategy of the target virtual power plant; the optimization objective function is used to describe the scheduling constraints of the target virtual power plant. The control module is used to perform scheduling control on the target virtual power plant based on the scheduling strategy.
[0029] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the virtual power plant scheduling method described above.
[0030] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned virtual power plant scheduling method.
[0031] The above-mentioned solution in this application has the following beneficial effects: In the embodiments of this application, by acquiring the power data and market data of the target virtual power plant, as well as the market data of multiple competing power plants, a virtual power plant decision model is constructed. Then, the virtual power plant decision model is initially trained using all the market data to obtain the trained virtual power plant decision model. Next, an undirected graph of the target virtual power plant is constructed, and a graph feature matrix is extracted based on the undirected graph. Then, an optimization objective function is designed, and the graph feature matrix and the power data of the target virtual power plant are input into the trained virtual power plant decision model to solve the optimization objective function, thereby obtaining the scheduling strategy of the target virtual power plant. Finally, the target virtual power plant is scheduled and controlled based on the scheduling strategy. Among these methods, market data from competing power plants is used to initially train the virtual power plant decision-making model, taking into account the game relationship with other market players to improve the performance and reliability of the virtual power plant decision-making model in terms of market competition strategies. An undirected graph is constructed and the scheduling strategy is solved based on the features of the undirected graph, taking into account the actual electrical connection constraints of the virtual power plant, which increases the feasibility and accuracy of the scheduling strategy. An optimized objective function is designed to impose certain constraints on the scheduling control of the virtual power plant, further improving the reliability and accuracy of virtual power plant scheduling.
[0032] Other beneficial effects of this application will be described in detail in the following detailed description section. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1 A flowchart illustrating a virtual power plant scheduling method provided in an embodiment of this application; Figure 2 A schematic diagram illustrating the framework of a virtual power plant scheduling method provided in an embodiment of this application; Figure 3 This is a schematic diagram of a hierarchical control relationship provided in an embodiment of this application; Figure 4 This is a schematic diagram of a meta-game process provided in an embodiment of this application; Figure 5 This is a schematic diagram of an undirected graph provided in an embodiment of this application; Figure 6 A constraint diagram provided for one embodiment of this application; Figure 7 This is a schematic diagram of performance curves provided for one embodiment of this application; Figure 8 This is a schematic diagram of the structure of a virtual power plant dispatching device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0035] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0036] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0037] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0038] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0039] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0040] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0041] To address the issues of poor reliability and accuracy in existing virtual power plant scheduling, this application provides a virtual power plant scheduling method. This method utilizes market data from competing power plants to initially train a virtual power plant decision-making model, considers game relationships with other market players to improve the performance and reliability of the virtual power plant decision-making model in terms of market competition strategies, constructs an undirected graph and solves scheduling strategies based on the features of the undirected graph, considers the actual electrical connection constraints of the virtual power plant, thereby increasing the feasibility and accuracy of the scheduling strategy, and designs and optimizes the objective function to impose certain constraints on the scheduling control of the virtual power plant, further improving the reliability and accuracy of virtual power plant scheduling.
[0042] To facilitate understanding of the method described in this application, some terms used in this application are explained below: Aggregators are intermediary organizations that aggregate resources in virtual power plants to participate in electricity market transactions and provide grid auxiliary services. Their core function is to integrate, coordinate, and optimize the scheduling of a large number of dispersed small-scale, distributed energy resources (such as distributed photovoltaics, energy storage, electric vehicles, adjustable loads, etc.) or load resources through advanced information and communication technologies and intelligent control platforms, forming a unified and controllable "virtual power source" or "virtual load".
[0043] A distributed energy resource cluster refers to a group of distributed energy resources (such as residential / commercial distributed photovoltaic, small wind power, micro gas turbines, energy storage systems, etc.) that are geographically concentrated or have the same management attributes in a virtual power plant.
[0044] The marginal electricity price at a node is the minimum marginal cost that the power grid pays to supply electricity to meet the needs of a specific node in the power grid for each additional unit of electricity load (usually 1 megawatt-hour).
[0045] The market-regulated mileage price is the core billing unit price for frequency regulation services. It refers to the cumulative change in output of frequency regulation resources (i.e., the increase or decrease in output of the unit / energy storage and other frequency regulation entities from the benchmark value) for each unit of regulation mileage completed, usually measured in megawatts. The service fee paid by the power grid or power trading center (for every minute) is based on the actual workload of the frequency regulation service, rather than just the frequency regulation capacity.
[0046] Renewable energy penetration rate refers to the ratio of renewable energy installed capacity (or power generation) to the maximum load (or total power generation) of the power system, reflecting the system's level of acceptance of clean energy.
[0047] The droop factor refers to the ratio of the percentage change in unit frequency to the percentage change in output. The smaller the droop factor, the more sensitive the unit is to frequency changes, the larger the output adjustment range, and the more effectively it can suppress frequency fluctuations.
[0048] The virtual power plant scheduling method provided in this application will be illustrated below.
[0049] like Figure 1 As shown, the virtual power plant scheduling method provided in this application includes the following steps: Step 11: Obtain the power data and market data of the target virtual power plant, as well as the market data of multiple competing power plants, and construct a virtual power plant decision model.
[0050] The aforementioned target virtual power plant refers to the virtual power plant that requires dispatch and control. Competing power plants refer to other virtual power plants that compete with the target virtual power plant in the electricity trading market. The aforementioned electricity data refers to the electricity-related data of the virtual power plant, such as node voltage amplitude, surplus energy, and system net load. The aforementioned market data refers to the electricity market-related data of the virtual power plant, such as node marginal price, market-regulated mileage price, historical quotations, user load patterns, renewable energy penetration rate, and droop coefficient.
[0051] For example, one can obtain power data and market data by accessing the management system of the target virtual power plant, and obtain market data of competing power plants by accessing their management systems.
[0052] The aforementioned virtual power plant decision-making model can be a Markov decision process model, including state space, action space, multi-objective reward function, etc.
[0053] The state space includes internal state, market state, and physical state.
[0054] The action space includes the adjustment ratio of each aggregator in the virtual power plant.
[0055] The multi-objective reward function is:
[0056] in, This represents the value of the multi-objective reward function. , , , , These are the weighting coefficients. As a conditional risk value penalty, Penalty for operating costs, For instruction tracking rewards, For physical security penalties, Rewards for market gains:
[0057]
[0058]
[0059]
[0060]
[0061] in, For energy market revenue, To assist market returns, For incremental operating costs, To adjust mileage costs, Indicates the number of aggregators. Indicates the tracking penalty coefficient. This indicates a higher-level adjustment instruction. This represents the actual power output of the aggregator. For critical path power flow. This is the upper limit of the line's transmission capacity. For the voltage amplitude at critical nodes, To allow for voltage deviation, For risk aversion coefficient, To be at confidence level Conditional Value at Risk (VaR) estimation This refers to the deviation between actual and expected returns.
[0062] For example, according to the Markov decision process model, the quintuple of this model ,in, For state space, For the action space, Let be the state transition probability. For the reward function, Discount factor: To address some observability issues, the state vector... It needs to comprehensively reflect the internal operating status of the VPP, the external market environment, and the physical status of the power grid, as defined below:
[0063] Among them, internal state Including the power output of distributed energy resource (DER) clusters ( (For aggregator index), Group Remaining Energy The plan is to exert effort within the day. and the current adjustment range Market Status Including nodal marginal electricity price superior's adjustment order Market-regulated mileage pricing Historical quotes and competitor quotes; physical condition Including network topology connection matrix and voltage amplitude of key nodes. Critical path power flow and system net load .
[0064] To facilitate agent exploration and ensure the physical feasibility of actions, a standardized proportional adjustment method is used to define the action space. The action vectors output by the agent... Composed of the adjustment ratios of each aggregator:
[0065] For each aggregator Its actual adjustment instructions Obtained through linear interpolation:
[0066] The final reference power command issued to aggregator K is This design standardizes the continuous motion space and naturally satisfies... Physical regulation limits.
[0067] Based on the above definition, the scheduling objective of a virtual power plant is to find an optimal strategy while satisfying a set of physical constraints (such as power flow constraints and voltage constraints). This maximizes the expected cumulative discount return:
[0068]
[0069] in, For a moment No. The amount of a constraint violation, For allowed constraints to violate the upper limit.
[0070] The constrained optimization problem is transformed into an unconstrained problem using the Lagrange relaxation method. The Lagrange function is constructed as follows:
[0071] in These are Lagrange multipliers. The original problem is equivalent to solving for saddle points. :
[0072] Within the RL framework, strategy From neural network parameters Parameterization, multipliers These parameters also become learnable. By alternately optimizing and updating the policy parameters and multipliers, the optimal policy learning under the constraints is ultimately achieved.
[0073] Step 12: Use all market data to perform preliminary training on the virtual power plant decision model to obtain the trained virtual power plant decision model.
[0074] In some embodiments of this application, the steps described above for initially training the virtual power plant decision-making model using all market data to obtain the trained virtual power plant decision-making model include: The first step is to construct multiple market task environments based on the market data of the target virtual power plant and the market data of all competing power plants, and to construct a strategy loss function.
[0075] Specifically, multiple market task environments are generated by randomly combining user load patterns, renewable energy penetration rates, and droop coefficients from the market data of the target virtual power plant and the market data of all competing power plants.
[0076] Defined by the following parameter tuple:
[0077] in, For the market task environment, Indicates the user load pattern. Indicates the penetration rate of renewable energy. The droop coefficient represents the combination of different user load patterns, renewable energy penetration rates, and the droop coefficient to form a specific market task environment.
[0078] For example, the above policy loss function for:
[0079] It should be noted that the market bidding process between the target virtual power plant and competing power plants can be modeled as a multi-agent partially observable Markov game (POSG). This simulator includes the protagonist intelligence to be trained; one or more adversary agents controlled by neural networks, whose objective is to minimize the protagonist's cumulative reward, thus providing the protagonist with the most challenging training environment; and a market clearing model employing a simplified unified clearing mechanism to calculate the nodal marginal price (LMP) and the winning bid amount for each VPP based on the bid curves of all VPPs and system demand. Target Virtual Power Plant Strategy Competitive power plant strategies The following training was conducted using minimax games:
[0080] in, For state vectors, Let the motion vector be the target virtual power plant. For the action vector of the competing power plant.
[0081] The second step involves using a virtual power plant decision model to make decision inferences and calculate the value of the strategy loss function in each market task environment.
[0082] For example, based on each market task environment The parameters of the power market simulator are configured (i.e., specific load curves, photovoltaic penetration rates, and competitor strategies are set), enabling the virtual power plant decision model to perform interactive sampling in the simulation environment, obtain state-action trajectories, and use these trajectories to calculate the value of the strategy loss function.
[0083] The third step is to calculate the temporary model parameters for each market task environment based on the value of the strategy loss function corresponding to that market task environment.
[0084] Specifically, through the formula:
[0085] Calculate the first Temporary model parameters for each market task environment .
[0086] in, The parameters represent the virtual power plant decision-making model. Indicates the inner learning rate. Indicates the first The value of the strategy loss function corresponding to each market task environment.
[0087] The fourth step is to update the parameters of the virtual power plant decision model based on all the temporary model parameters to obtain the trained virtual power plant decision model.
[0088] Specifically, through the formula:
[0089]
[0090] Update the parameters of the virtual power plant decision model.
[0091] in, This represents the model parameters of the trained virtual power plant decision-making model. This represents gradient operation. This represents the learning rate of the outermost element.
[0092] For example, the minimax game problem described above can be solved using a meta-learning algorithm. This algorithm consists of two stages: inner adaptation and outer update. First, the temporary model parameters for each task are calculated using the formula described above. This step simulates the rapid learning process of an agent facing a new market environment. After obtaining all the temporary model parameters, the performance of these adapted strategies is evaluated using new query data. The goal of outer optimization is to find a globally optimal initial parameter that minimizes the expected cumulative loss across all tasks after inner adaptation. Finally, the model parameters are updated using stochastic gradient descent.
[0093] Step 13: Construct an undirected graph of the target virtual power plant and extract the graph feature matrix based on the undirected graph.
[0094] In the undirected graph described above, multiple nodes in the target virtual power plant correspond one-to-one. The edges between nodes represent the topological relationships between corresponding power nodes. Circuit nodes include nodes in the target virtual power plant such as aggregators and load devices. The graph feature matrix described above includes high-level feature representations of all power nodes in the undirected graph.
[0095] In some embodiments of this application, the steps of constructing an undirected graph of the target virtual power plant and extracting a graph feature matrix based on the undirected graph include: The first step is to construct an undirected graph of the target virtual power plant.
[0096] For example, a corresponding node is constructed for each power node in the target virtual power plant. If there is a power connection line between two power nodes, they are considered to have a topological relationship, and an edge is generated between the corresponding two nodes. Otherwise, they do not have a topological relationship, and no edge is generated. This judgment process is performed for every two nodes to obtain an undirected graph.
[0097] The second step is to extract the high-level feature representation of each node in the undirected graph.
[0098] Specifically, through the formula:
[0099] Calculate the first The first layer update Feature representation of each node .
[0100] Among them, when hour, For the first High-level feature representation of each node, For the updated layer number, In an undirected graph, the first... A set of vector nodes of nodes. For activation function, Indicates the first The node and the first Normalized attention weights between nodes, Indicates the first Layer The updated node representation.
[0101] For example, the expressions for the above node representation, normalized attention weights, etc., are as follows:
[0102]
[0103]
[0104] in, Indicates the first Layer The updated node representation. Represents the shared weight matrix. Indicates the first The first layer update The feature representation of each node. For the first The node and the first Attention weights between nodes For attention vectors, Indicates the first The node and the first Attention weights between nodes.
[0105] It should be noted that the weighted aggregation process in the above formula has a profound physical isomorphism with the nodal power flow equations in a power system. In the AC power flow physical model, the injected power of a node depends on the voltage magnitude and phase angle of its neighboring nodes, and its coupling strength is determined by the admittance matrix. Decide.
[0106] The third step is to integrate the high-level feature representations of all nodes into a matrix to obtain the graph feature matrix.
[0107] Step 14: Design the optimization objective function, and input the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model to solve the optimization objective function and obtain the scheduling strategy of the target virtual power plant.
[0108] The aforementioned objective function aims to maximize the overall scheduling benefits while ensuring physical safety constraints. The output of the aforementioned scheduling strategy is the power regulation coefficient of each aggregator in the target virtual power plant, specifically defined as the interpolation ratio of the aggregator within the adjustable power range (i.e., between minimum and maximum output) at the current moment.
[0109] In some embodiments of this application, the steps of designing and optimizing the objective function, inputting the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model, and solving the objective function to obtain the scheduling strategy of the target virtual power plant include: The first step is to design and optimize the objective function.
[0110] Specifically, the objective function to be optimized is:
[0111] in, This indicates the value of the objective function to be optimized. Represents the Lagrange multipliers. Represents risk variables, Indicates the first A constrained Lagrange multiplier, Indicates the first The expected violation of a constraint. This represents the risk aversion coefficient. This represents the CVaR alternative objective function. Policy parameters are updated alternately using stochastic gradient ascent or descent. : Lagrange multipliers : and risk variables : .in This indicates a projection operation that ensures the multipliers are non-negative and bounded.
[0112] For example, to address the strict physical constraints that must be met, such as power limitations and energy storage capacity boundaries, a differentiable projection layer is designed and embedded at the output of the neural network:
[0113] in, Let j be the j-th physical hard constraint function. The original action commands output by the neural network. Let n be the n-dimensional real action space.
[0114] To maintain the smooth flow of gradients while performing hard truncation, a differentiable computation graph is constructed using duality theory and the Karush-Kuhn-Tucker conditions. Based on the implicit function theorem, the Jacobian matrix of the projected output with respect to the input action can be analytically computed, thereby backpropagating the safety correction signal to the preceding policy network and guiding the agent to actively learn and efficiently explore the feasible region boundary.
[0115] Conditional Value at Risk (CVaR) is introduced to address the uncertainty of returns caused by market price fluctuations. It serves as a risk metric in finance to quantify and control the extreme risk at the tails of the return distribution. CVaR is discussed at confidence levels. Defined as the conditional expectation of the tail of the loss distribution, it is more effective than variance in capturing extreme market risk. In strategy... Below, random returns of The definition is as follows:
[0116] in Value at Risk (CVaR). To optimize CVaR in RL, we introduce an auxiliary variable. Let VaR be represented, and optimize the following alternative objectives:
[0117] The Lagrange dual update integrates the aforementioned hard constraint handling with soft risk control, constructing a unified Lagrange optimization objective (i.e., the optimization objective function mentioned above). The system no longer simply pursues the maximization of expected return, but automatically searches for the optimal balance between risk and return while satisfying physical constraints.
[0118] The second step involves inputting the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model, solving the objective function, and obtaining the scheduling strategy of the target virtual power plant.
[0119] For example, the graph feature matrix is embedded into the critic network in the trained virtual power plant decision model, and then the power data of the target virtual power plant is substituted into the state space. The trained virtual power plant decision model is used to make decision inferences, with the goal of minimizing the optimization objective function, and the scheduling strategy of the target virtual power plant is output.
[0120] For example: Commentator Network after embedding graph feature matrix for:
[0121]
[0122] in, For state space, The adjustment ratio for the first aggregator. For the first The adjustment ratio of each aggregator For the first The adjustment ratio of each aggregator The feature matrix of the graph. For policy networks, For local observation, Indicates the first High-level feature representation of the nodes corresponding to each aggregator Indicates the first High-level feature representation of the nodes corresponding to each aggregator For the first The set of neighboring nodes of each aggregator node.
[0123] No. The aggregator's executor maximizes its critics' estimates. It's worth updating:
[0124] The commentator updates the time series difference (TD) error by minimizing it:
[0125]
[0126] in and The target network parameters are kept stable through soft updates.
[0127] It should be noted that before performing this step or connecting the virtual power plant decision model to the real power grid, multiple virtual power plants as samples need to be selected to train the model. First, a high-fidelity digital twin simulation environment needs to be constructed. Based on the historical load data, meteorological data, and power grid physical parameters of the virtual power plants as samples, a Markov Decision Process (MDP) simulation environment for the virtual power plants is built. Using the virtual power plant decision model, large-scale iterative training is performed in the simulation environment through training algorithms such as gradient descent, assigning higher sampling weights to rare samples with large prediction biases. :
[0128] in For sampling probability, For sample priority, To normalize the denominator, For buffer size, This is the bias correction index. Through this mechanism, the agent can quickly grasp the basic game strategy and cooperative logic under complex boundary conditions, and obtain initial network parameters with a certain generalization ability.
[0129] Step 15: Perform scheduling control on the target virtual power plant based on the scheduling strategy.
[0130] Specifically, for each aggregator in the target virtual power plant, a reference power command is generated for that aggregator based on the adjustment ratio corresponding to that aggregator in the scheduling strategy. The reference power command is then executed in the aggregator to achieve scheduling control.
[0131] For example, the expression for generating the reference power command is: ,in, For reference power command, For aggregators k exist t The adjustable power limit at any given time. For aggregators k exist t The adjustable power lower limit at any given time. For scheduling strategy, This represents the adjustment ratio corresponding to the aggregator in the scheduling strategy.
[0132] It is worth mentioning that the virtual power plant decision-making model was initially trained using market data from competing power plants, taking into account the game relationship with other market players, thereby improving the performance and reliability of the virtual power plant decision-making model in terms of market competition strategies. An undirected graph was constructed and the scheduling strategy was solved based on the features of the undirected graph, taking into account the actual electrical connection constraints of the virtual power plant, which improved the feasibility and accuracy of the scheduling strategy. An optimized objective function was designed to impose certain constraints on the scheduling control of the virtual power plant, further improving the reliability and accuracy of virtual power plant scheduling.
[0133] The method of this application will be illustrated below with a specific example.
[0134] The specific framework of the method in this application is as follows: Figure 2 As shown, the vertical layering includes the following: The first layer: Market Game Theory Layer (Meta-learning and Opponent Modeling), which specifically includes: meta-reinforcement learning agents, opponent modeling (i.e., market simulators), strategic bidding, and expected return signals. This layer corresponds to step 12. The second layer: Topological Collaboration Layer (Graph Neural Networks and Multi-Agents), which specifically includes: constructing an undirected graph and acquiring graph features through graph attention mechanisms, multi-agent collaboration (Grah-MADDPG), and aggregating raw actions. This layer corresponds to step 13. The third layer: Security Constraints and Risk Management Layer, which specifically includes: physical constraints (Lagrange relaxation), financial constraints (CVaR index), action projection / correction, and outputting the final safe and executable action. This layer corresponds to step 14. After obtaining the final safe and executable action, it interacts with the external environment (power grid and market), outputting returns and violation signals, and using the next state as input for real-time states (including market, physical, and internal states).
[0135] The three-tiered hierarchical control relationship between virtual power plants, system operators, and distributed energy resources is as follows: Figure 3 As shown, the system operator is responsible for network-wide scheduling and market organization. It initiates scheduling requests to the Virtual Power Plant (VPP) agent. The VPP processes the data through communication, data acquisition, forecasting center, and optimized scheduling to obtain decomposition instructions, which are then transmitted to distributed energy resources. Distributed energy resources include distributed photovoltaics, energy storage systems, controllable loads, and other equipment. After decomposing the instructions, the distributed energy resources provide status feedback to the VPP. Based on the status feedback, the VPP transmits the tracking error to the system operator.
[0136] exist Figure 2 Within the framework shown, the specific process of the first-level game is as follows: Figure 4 As shown, multiple market states are taken as input, enabling the protagonist agent (i.e., the virtual power plant VPP agent) to engage in adversarial interactions with the opposing agents (i.e., market participants: power generators and aggregators), and outputting the protagonist agent's strategy. Strategies with competitors Through strategy Perform inner adaptation (inner loop) through strategy and strategy Meta-game training and meta-learning are performed, and the game expression is: The parameters are updated as follows: This implements outer loop updates and ultimately outputs the initial policy parameters. .
[0137] Undirected graphs Figure 5 As shown, Figure 5 It includes four nodes, namely node i (battery, feature). ), Node j photovoltaic (characteristics) ), node k load, node l electric vehicle (characteristics) The features of each node include corresponding active power P, reactive power Q, and voltage V data. Then, the attention between nodes is calculated, and a weighted summation and nonlinear activation are applied to output the aggregated high-level features, expressed as: .
[0138] Soft and hard constraints, such as Figure 6 As shown, input the original action. Differentiable projections are performed on it, including hard-constrained projections: The soft constraint assessment CVaR modifies the action based on these constraints and outputs a safe action. And perform gradient backpropagation .
[0139] The above framework was simulated, with the simulation scenario set as a virtual power plant (VPP) aggregating various heterogeneous resources. Its physical topology was modified from the IEEE 33-node radial distribution system. The hyperparameter settings in the algorithm are shown in Table 1.
[0140] Table 1
[0141] The distributed power sources within the VPP include 5 photovoltaic (PV) units and 3 wind turbine (WT) units. Their output curves are based on sampling data from a company in Hunan Province, with 5% Gaussian white noise added to simulate prediction errors. The energy storage system is equipped with 4 sets of lithium-ion battery energy storage systems, each with a rated capacity of 1 MWh, a maximum charge / discharge power of 250 kW, and a charge / discharge efficiency of 95%. Historical annual load data from a certain electricity market in 2024 was selected as a benchmark and distributed proportionally to each load node.
[0142] To rigorously evaluate the generalization ability of the proposed evaluation model, the annual data was divided into training and test sets. Data from January 1st to September 30th was selected for model training and parameter updates, covering the mild spring season and the peak summer electricity consumption period. Data from October 1st to December 31st was selected as the test set. This period includes load fluctuations caused by winter heating, and this data is completely invisible to the model during training, used to verify the model's performance in unknown market environments. To comprehensively evaluate the generalization ability and robustness of the proposed framework under complex operating conditions, four heterogeneous test scenarios with different statistical characteristics were constructed in the implementation case, aiming to simulate various extreme environments that VPPs may face in actual operation. The characteristics of different datasets are shown in Table 2.
[0143] Table 2
[0144] Different settings were made for the average load demand, peak-valley difference, renewable energy penetration rate, and environmental random fluctuation rate for each scenario.
[0145] Simulation experiments were conducted on the method of this application using the above settings, and the overall performance is as follows: Figure 7 As shown, Figure 7 The horizontal axis represents the number of rounds, and the different curves correspond to the baseline scenario, high photovoltaic penetration, high load, and high volatility scenarios, respectively. Figure 7 'a' represents the curve showing the change in reward value during the iteration period, with the vertical axis representing the reward value. Figure 7 b represents the curve showing the change in power imbalance rate during the iteration, with the vertical axis representing the power imbalance rate. Figure 7 c represents the curve showing the change in operating costs during the iteration period, with the vertical axis representing operating costs. Figure 7 d represents the curve showing the change in constraint violations during iterations, with the vertical axis representing the constraint violation values. This figure reflects the dynamic convergence characteristics of the agent during training in four typical operating scenarios (baseline scenario, high photovoltaic scenario, high load scenario, and high volatility scenario), including reward gains, power imbalance, system operating costs, and constraint violations. The solid lines in the figure represent the average values of all time steps within each training round, while the shaded areas represent the fluctuation range formed by the minimum and maximum values of the indicators at each time step within the same round. This demonstrates the convergence stability and robustness of the algorithm under different operating environments.
[0146] In summary, the proposed method can effectively address the economic dispatch problem of VPPs in various scenarios. It not only demonstrates stable convergence across various performance indicators but also exhibits good generalization ability and adaptability under different operating conditions, providing a reliable technical approach for virtual power plants to participate in intelligent collaborative dispatch within the electricity market environment.
[0147] The virtual power plant dispatching device provided in this application is described below as an example.
[0148] like Figure 8 As shown, this application embodiment provides a virtual power plant dispatching device, the virtual power plant dispatching device 800 including: The acquisition module 801 is used to acquire the power data and market data of the target virtual power plant, as well as the market data of multiple competing power plants, and to build a virtual power plant decision model. Training module 802 is used to perform preliminary training on the virtual power plant decision model using all market data to obtain the trained virtual power plant decision model. Module 803 is used to construct an undirected graph of the target virtual power plant and extract the graph feature matrix based on the undirected graph; multiple nodes in the undirected graph correspond one-to-one with multiple power nodes in the target virtual power plant, and the edges between nodes represent the topological relationship between the corresponding two power nodes. The solver module 804 is used to design the optimization objective function and input the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model to solve the optimization objective function and obtain the scheduling strategy of the target virtual power plant; the optimization objective function is used to describe the scheduling constraints of the target virtual power plant. The control module 805 is used to perform scheduling control on the target virtual power plant based on the scheduling strategy.
[0149] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.
[0150] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0151] like Figure 9As shown, an embodiment of this application provides a terminal device, wherein the terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 9 The diagram shows only one processor, a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 executes the computer program D102 to implement the steps in any of the above method embodiments.
[0152] Specifically, when the processor D100 executes the computer program D102, it performs preliminary training on the virtual power plant decision model using market data from competing power plants, considers the game relationship with other market players, improves the performance and reliability of the virtual power plant decision model in terms of market competition strategies, constructs an undirected graph and solves the scheduling strategy based on the features of the undirected graph, considers the actual electrical connection constraints of the virtual power plant, thereby increasing the feasibility and accuracy of the scheduling strategy, and designs and optimizes the objective function to impose certain constraints on the scheduling control of the virtual power plant, further improving the reliability and accuracy of the virtual power plant scheduling.
[0153] The processor D100 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0154] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may be an external storage device of the terminal device D10, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the terminal device D10. Furthermore, the memory D101 may include both internal and external storage units of the terminal device D10. The memory D101 is used to store the operating system, applications, bootloader, data, and other programs, such as the program code of the computer program. The memory D101 can also be used to temporarily store data that has been output or will be output.
[0155] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.
[0156] This application provides a computer program product that, when run on a terminal device, enables the terminal device to implement the steps described in the various method embodiments above.
[0157] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to the virtual power plant dispatching method apparatus / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, such as a USB flash drive, a portable hard drive, a magnetic disk, or an optical disk.
[0158] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0159] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0160] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention.
Claims
1. A virtual power plant dispatching method, characterized in that, include: Acquire the power data and market data of the target virtual power plant, as well as the market data of multiple competing power plants, and construct a virtual power plant decision model; The virtual power plant decision-making model was initially trained using all market data to obtain the trained virtual power plant decision-making model. Construct an undirected graph of the target virtual power plant and extract a graph feature matrix based on the undirected graph; multiple nodes in the undirected graph correspond one-to-one with multiple power nodes in the target virtual power plant, and the edges between nodes represent the topological relationship between two corresponding power nodes; Design an optimization objective function, and input the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model to solve the optimization objective function to obtain the scheduling strategy of the target virtual power plant; the optimization objective function is used to describe the scheduling constraints of the target virtual power plant. The target virtual power plant is scheduled and controlled based on the aforementioned scheduling strategy.
2. The virtual power plant scheduling method according to claim 1, characterized in that, The virtual power plant decision-making model includes a state space, an action space, and a multi-objective reward function; The state space includes internal state, market state, and physical state; The action space includes the adjustment ratio of each aggregator in the virtual power plant; The multi-objective reward function is: in, This represents the value of the multi-objective reward function. , , , , These are the weighting coefficients. As a conditional risk value penalty, Penalty for operating costs, For instruction tracking rewards, For physical security penalties, Rewards for market gains: in, For energy market revenue, To assist market returns, For incremental operating costs, To adjust mileage costs, Indicates the number of aggregators. Indicates the tracking penalty coefficient. This indicates a higher-level adjustment instruction. This represents the actual power output of the aggregator. For critical path power flow. This is the upper limit of the line's transmission capacity. For the voltage amplitude at critical nodes, To allow for voltage deviation, For risk aversion coefficient, To be at confidence level Conditional Value at Risk (VaR) estimation This refers to the deviation between actual and expected returns.
3. The virtual power plant scheduling method according to claim 2, characterized in that, The preliminary training of the virtual power plant decision-making model using all market data to obtain the trained virtual power plant decision-making model includes: Multiple market task environments are constructed based on market data of the target virtual power plant and market data of all competing power plants, and a strategy loss function is constructed accordingly. Using a virtual power plant decision model, decision inference is performed and the value of the strategy loss function is calculated for each market task environment. For each market task environment, based on the value of the strategy loss function corresponding to the market task environment, calculate the temporary model parameters for that market task environment; The virtual power plant decision model is updated with parameters based on all temporary model parameters to obtain the trained virtual power plant decision model.
4. The virtual power plant scheduling method according to claim 3, characterized in that, The calculation of temporary model parameters under the market task environment based on the value of the strategy loss function corresponding to the market task environment includes: Through the formula: Calculate the first Temporary model parameters for each market task environment ; in, The parameters represent the virtual power plant decision-making model. Indicates the inner learning rate. Indicates the first The value of the strategy loss function corresponding to each market task environment; The step of updating the parameters of the virtual power plant decision model based on all temporary model parameters includes: Through the formula: Update the parameters of the virtual power plant decision model; in, This represents the model parameters of the trained virtual power plant decision-making model. This represents gradient operation. This represents the learning rate of the outermost element.
5. The virtual power plant dispatching method according to claim 4, characterized in that, The extraction of the graph feature matrix based on the undirected graph includes: Based on the undirected graph, extract the high-level feature representation of each node in the undirected graph; The high-level feature representations of all nodes are integrated into a matrix to obtain the graph feature matrix.
6. The virtual power plant scheduling method according to claim 5, characterized in that, The step of extracting high-level feature representations of each node in the undirected graph includes: Through the formula: Calculate the first The first layer update Feature representation of each node ; Among them, when hour, For the first High-level feature representation of each node, For the updated layer number, In an undirected graph, the first... A set of vector nodes of nodes. For activation function, Indicates the first The node and the first Normalized attention weights between nodes, Indicates the first Layer The updated node representation.
7. The virtual power plant scheduling method according to claim 6, characterized in that, The optimization objective function is: in, This indicates the value of the objective function to be optimized. Represents the Lagrange multipliers. Represents risk variables, Indicates the first A constrained Lagrange multiplier, Indicates the first The expected violation of a constraint. This represents the risk aversion coefficient. This indicates that CVaR replaces the objective function.
8. A virtual power plant dispatching device, characterized in that, include: The acquisition module is used to acquire the power data and market data of the target virtual power plant, as well as the market data of multiple competing power plants, and to build a virtual power plant decision model. The training module is used to perform preliminary training on the virtual power plant decision model using all market data, so as to obtain the trained virtual power plant decision model. The construction module is used to construct an undirected graph of the target virtual power plant and extract a graph feature matrix based on the undirected graph; multiple nodes in the undirected graph correspond one-to-one with multiple power nodes in the target virtual power plant, and the edges between nodes represent the topological relationship between two corresponding power nodes. The solution module is used to design an optimization objective function and input the graph feature matrix and the power data of the target virtual power plant into the trained virtual power plant decision model to solve the optimization objective function and obtain the scheduling strategy of the target virtual power plant; the optimization objective function is used to describe the scheduling constraints of the target virtual power plant. The control module is used to perform scheduling control on the target virtual power plant based on the scheduling strategy.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the virtual power plant scheduling method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the virtual power plant scheduling method as described in any one of claims 1 to 7.