Multi-modal resource collaborative optimization scheduling method, system and equipment of virtual power plant
By unifying the modeling of multimodal resources within a virtual power plant and applying deep learning agents, collaborative scheduling instructions and joint bidding strategies are generated, solving the problem of insufficient multimodal resource coordination and improving the operational economy and market competitiveness of the virtual power plant.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-10
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, virtual power plants suffer from insufficient multimodal resource coordination, disconnect between response assessment and scheduling, and lack of multi-market coordination mechanisms. This makes it difficult to fully tap the potential for synergistic complementarity of distributed resources, and results in insufficient operational returns when facing complex power market environments.
By uniformly modeling and dynamically aggregating multimodal distributed resources within a virtual power plant, a standardized resource pool is generated. Deep reinforcement learning agents are used to generate collaborative scheduling instructions and multi-market joint bidding strategies. Entropy methods and graph neural networks are combined for resource evaluation and feature extraction. A multi-objective optimization reward function is constructed, and the strategy is iteratively optimized.
It improves the operational economy, reliability, and market competitiveness of virtual power plants, and achieves efficient collaborative scheduling of multimodal resources and maximizes market benefits.
Smart Images

Figure CN122026503A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of virtual power plant optimization scheduling, specifically to a multimodal resource collaborative optimization scheduling method, system, and equipment for virtual power plants. Background Technology
[0002] Distributed power generation is characterized by intermittency, volatility, and dispersion, posing significant challenges to the real-time balance, safety, stability, and economic operation of power systems. Virtual power plants (VPS), as a crucial means of aggregating distributed energy resources, flexible loads, and energy storage systems, possess significant potential in enhancing grid flexibility, promoting renewable energy consumption, and participating in the electricity market. Currently, VPS dispatching and market participation primarily focus on single-type resources or markets, such as optimizing the charging and discharging of energy storage systems or employing bidding strategies for the electricity market. However, with the increasing diversification of distributed resources, encompassing multiple modes including photovoltaics, wind power, energy storage, electric vehicles, and flexible loads, traditional dispatching methods struggle to fully leverage their synergistic and complementary potential. Furthermore, the electricity market mechanism is becoming increasingly complex, with multiple market types coexisting, including energy markets and ancillary service markets, making it difficult for market bidding strategies to maximize the operational benefits of VPS. In addition, the external environment and market conditions exhibit significant uncertainties, such as renewable energy output, electricity price fluctuations, and load changes, requiring VPS to possess the ability to efficiently predict and dynamically respond to multi-source information.
[0003] Therefore, current technologies suffer from technical problems such as insufficient multimodal resource coordination, disconnect between response assessment and scheduling, and lack of multi-market coordination mechanisms. Summary of the Invention
[0004] This application provides a multimodal resource collaborative optimization scheduling method, system, and equipment for virtual power plants, which solves the technical problems of insufficient multimodal resource collaboration, disconnect between response assessment and scheduling, and lack of multi-market collaboration mechanisms in the existing technology. It achieves the technical effect of improving the economy, reliability, and market competitiveness of virtual power plant operation by generating collaborative scheduling instructions and multi-market joint bidding strategies through deep learning agents.
[0005] This application provides a multimodal resource collaborative optimization scheduling method for virtual power plants. The method includes: performing unified modeling and dynamic aggregation of multimodal distributed resources within the virtual power plant to generate a standardized resource pool representing the aggregation response capability; generating prediction sequences of multiple future scenarios based on external market and environmental information; using the state of the standardized resource pool and the prediction sequences as the state observation space, making decisions through a deep reinforcement learning agent, and outputting scheduling instructions and joint bidding strategies for multiple electricity markets; issuing the scheduling instructions to the underlying physical devices for execution, and simultaneously submitting the joint bidding strategies to multiple electricity markets, and continuously iteratively optimizing the strategies based on the responses.
[0006] In a possible implementation, the multimodal resource collaborative optimization scheduling method for the virtual power plant further performs the following processing: modeling the independent response capabilities of multi-source heterogeneous resources within the virtual power plant; constructing a unified feature extraction model based on a graph neural network to extract the spatiotemporal dynamic features of the multi-source heterogeneous resources; establishing a unified evaluation index system; assigning weights to each index in the evaluation index system based on the spatiotemporal dynamic features using the entropy method; and mapping and aggregating the multi-source heterogeneous resources into three standard types of resource pools—interruptible, movable, and adjustable—according to different weight combinations, thereby generating the standardized resource pool.
[0007] In a possible implementation, the multimodal resource collaborative optimization scheduling method for the virtual power plant also performs the following processing: the evaluation index system includes response time, duration, power upper and lower limits, ramp rate, and average response time.
[0008] In a possible implementation, the multimodal resource collaborative optimization scheduling method for the virtual power plant further performs the following processing: the deep reinforcement learning agent adopts a deep Q-network algorithm, and the action space includes power setting and adjustment instructions for multimodal distributed resources and bidding strategies for multiple electricity markets; the reward function of the deep reinforcement learning agent is a multi-objective optimization form, and the multi-objectives include total market revenue, penalty for default, operating costs, and conditional risk value, wherein the total market revenue is the sum of electricity market revenue, ancillary service market revenue, and carbon trading revenue, and the conditional risk value is used to quantify the tail risk under multi-market coupling.
[0009] In a possible implementation, the multimodal resource collaborative optimization scheduling method for the virtual power plant also performs the following processing: the conditional value of risk is used to measure multi-market coupling risk and includes an expected return term and a risk adjustment term.
[0010] In a possible implementation, the multimodal resource collaborative optimization scheduling method for the virtual power plant further performs the following processing: each market clears based on a master-slave game, where the virtual power plant, as the leader, submits its application first, and the electricity market, ancillary services market, and carbon trading market, as followers, clear their markets and return clearing prices; the virtual power plant, based on the returned clearing price information and the collected actual resource output and user response behavior information, along with the state and actions in the deep reinforcement learning agent, constitutes an experience tuple, which is stored in an experience replay buffer; the network parameters of the deep reinforcement learning agent are updated based on the experience tuple in the experience replay buffer, and the strategy is continuously iteratively optimized.
[0011] In possible implementations, the multimodal resource collaborative optimization scheduling method for the virtual power plant also performs the following processing: the master-slave game also includes physical constraints, market constraints, and time coupling constraints; wherein, physical constraints include power balance constraints, reserve capacity constraints, and carbon emission constraints; market constraints include power-ancillary service capacity coupling constraints and cross-market arbitrage constraints.
[0012] This application also provides a multimodal resource collaborative optimization scheduling system for a virtual power plant. The system includes: a power plant resource aggregation module, used to uniformly model and dynamically aggregate multimodal distributed resources within the virtual power plant to generate a standardized resource pool representing the aggregation response capability; a prediction sequence generation module, used to generate prediction sequences for various future scenarios based on external market and environmental information; a bidding strategy output module, used to make decisions through a deep reinforcement learning agent using the state of the standardized resource pool and the prediction sequence as the state observation space, and output scheduling instructions and joint bidding strategies for multiple electricity markets; and a strategy optimization module, used to issue the scheduling instructions to the underlying physical devices for execution, and simultaneously submit the joint bidding strategy to multiple electricity markets, continuously iteratively optimizing the strategy based on the response.
[0013] This application also provides an electronic device, including: a memory for storing executable instructions; and a processor for implementing a multimodal resource collaborative optimization scheduling method for a virtual power plant when executing the executable instructions stored in the memory.
[0014] This application proposes a multimodal resource collaborative optimization scheduling method, system, and equipment for virtual power plants. This method unifies and dynamically aggregates diverse distributed resources within the virtual power plant, constructing a standardized resource pool. Based on external market and environmental information, it generates predictive sequences for various future scenarios, forming a state observation space. A deep reinforcement learning agent is used for decision-making, simultaneously generating real-time scheduling instructions and joint bidding strategies. The scheduling instructions are executed, and market bids are submitted, with continuous iteration and optimization based on actual market responses and environmental feedback. This addresses the technical problems of insufficient multimodal resource collaboration, disconnect between response assessment and scheduling, and the lack of multi-market collaboration mechanisms in existing technologies. It achieves the technical effect of improving the economy, reliability, and market competitiveness of virtual power plant operation through the generation of collaborative scheduling instructions and multi-market joint bidding strategies using deep learning agents. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments of this disclosure will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0016] Figure 1 A schematic diagram of the multimodal resource collaborative optimization scheduling method for a virtual power plant provided in this application embodiment.
[0017] Figure 2 A schematic diagram of the structure of a multimodal resource collaborative optimization scheduling system for a virtual power plant provided in this application embodiment.
[0018] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0019] Explanation of reference numerals in the attached drawings: Electric field resource aggregation module 10, prediction sequence generation module 20, bidding strategy output module 30, strategy optimization module 40, input device 401, processor 402, memory 403, output device 404. Detailed Implementation
[0020] To further illustrate the technical means and effects adopted by the present invention in order to achieve the intended purpose, the following detailed description is provided in conjunction with the accompanying drawings and preferred embodiments, based on the specific implementation methods, structures, features and effects of the present invention.
[0021] This application provides a multimodal resource collaborative optimization scheduling method for virtual power plants, such as... Figure 1 As shown, the method includes: Step S100: Perform unified modeling and dynamic aggregation of multimodal distributed resources within the virtual power plant to generate a standardized resource pool that characterizes the aggregation response capability.
[0022] Step S100 further includes step S110, which involves modeling the independent response capabilities of the multi-source heterogeneous resources within the virtual power plant, constructing a unified feature extraction model based on a graph neural network, and extracting the spatiotemporal dynamic features of the multi-source heterogeneous resources; step S120, which involves establishing a unified evaluation index system, assigning weights to each index in the evaluation index system based on the spatiotemporal dynamic features using the entropy method, and mapping and aggregating the multi-source heterogeneous resources into three standard types of resource pools: interruptible, movable, and adjustable, according to different weight combinations, thereby generating the standardized resource pool.
[0023] Preferably, the virtual power plant contains various types of resources, including photovoltaic power generation units, wind power generation units, battery energy storage systems, electric vehicle charging piles, industrial adjustable loads, and commercial temperature-controlled loads. Independent response capability modeling is performed on these multi-source heterogeneous resources to quantify their response capabilities. Response capability refers to the resource's ability to change its power output after receiving a dispatch command. Key parameters typically include response delay (the time from receiving the command to initiating action); adjustment rate (the magnitude of power change per unit time, i.e., ramp rate); duration (the duration of continuous operation or response at a specific power level); and adjustment capacity (the maximum magnitude by which power can be increased or decreased). A unified feature extraction model is constructed based on graph neural networks. All resources within the virtual power plant are considered as nodes in a graph structure. The edges connecting nodes represent their geographical proximity, electrical connection relationships, or commercial affiliation. Graph neural networks are used to process the graph structure data. The initial features of each node resource are its independently modeled response capability parameters. By aggregating information from adjacent nodes, spatiotemporal dynamic features are extracted, including spatial and temporal features. Spatial features characterize the local mutual influence of resources in the power grid, while temporal features characterize the pattern of resource response capability changes over time.
[0024] Furthermore, step S100 also includes the evaluation index system including response time, duration, power upper and lower limits, ramp rate, and average response time.
[0025] Preferably, quantitative indicators are defined for the final measurement and comparison of standardized resource types, establishing a unified evaluation indicator system. These include: response time, which measures the time from receiving an instruction to initiating a response, reflecting the resource's rapid response capability; duration, which assesses the resource's continuous availability during the response process, reflecting its stability; upper and lower power limits, which determine the maximum and minimum adjustment capabilities of resources after aggregation, providing a basis for grid dispatch; ramp rate, which represents the rate of power change of a resource per unit time, reflecting its dynamic adjustment capability; and average response time, which comprehensively considers the response times of all resources, providing an average value for the overall response capability. Based on spatiotemporal dynamic characteristics, the entropy method is used to assign weights to each indicator in the evaluation indicator system. That is, the weight is determined according to the dispersion of each indicator's data. The greater the data difference and the smaller the entropy value of an indicator among all resources, the greater the information provided by that indicator in distinguishing resource types, and thus its higher weight. The feature vectors extracted from all resources through a graph neural network are substituted into the evaluation indicator system for calculation, obtaining a numerical matrix for each resource on each indicator. The entropy method is then used on this matrix to automatically calculate the weight of each indicator, thereby avoiding bias caused by subjective weight setting.
[0026] Preferably, based on the weights determined by the entropy method, the scores of each resource on each indicator are weighted and calculated to obtain the tendency score or membership degree of each resource corresponding to the three standard types of "interruptible", "transferable", and "adjustable". Then, the multi-source heterogeneous resources are mapped and aggregated into resource pools of the three standard types: interruptible, transferable, and adjustable. The interruptible resource pool mainly includes resources that can be completely stopped or started in a short period of time, but have strict limitations on the duration or are sensitive to the number of interruptions, such as certain industrial process loads. The transferable resource pool mainly includes resources with fixed total power consumption or power generation, but whose operating period can be flexibly adjusted within a certain time window, such as charging piles and certain manufacturing loads. The adjustable resource pool mainly includes resources that can continuously and smoothly adjust the output or power consumption level within a certain range and maintain it for a period of time, such as energy storage systems, gas turbines, and air conditioning loads. Finally, all heterogeneous resources in the virtual power plant are classified and aggregated to generate standardized resource pools.
[0027] Step S200: Based on external market and environmental information, generate prediction sequences for various future scenarios.
[0028] Preferably, forecast information from external markets and the environment is integrated, including short-term wind and solar power output forecasts, ancillary service price trends, and market electricity price fluctuations. Through historical scenario clustering and multi-source data fusion, forecast sequences for various future scenarios are generated. Each sequence represents a market and environmental state at 96 future time points. For example, under the market sequence, there are electricity price sequences, ancillary service price sequences, and carbon price sequences; under the environmental sequence, there are wind and solar power output sequences and temperature sequences; and it may also include net load sequences or price difference sequences under the coupling sequence.
[0029] Step S300: Using the state of the standardized resource pool and the predicted sequence as the state observation space, a deep reinforcement learning agent makes decisions and outputs scheduling instructions and joint bidding strategies for multiple electricity markets.
[0030] Preferably, the state of the standardized resource pools is the internal state, representing the real-time status of the three standardized resource pools that are controllable by the virtual power plant itself. This includes the total available capacity, average response time, sustainability, ramp-up capability, and state of charge of the energy storage system for each resource pool at the current moment. The prediction sequence is a description of the uncertainty of the external environment, representing predicted data such as market prices and wind and solar power output over a future period. Using the state of the standardized resource pools and the prediction sequence as the state observation space, a deep reinforcement learning agent makes decisions. The deep reinforcement learning agent is a trained deep neural network that can select the optimal action that maximizes long-term cumulative rewards based on the currently observed state. Specifically, the deep reinforcement learning agent attempts actions, such as commanding the adjustable pool to increase output and bid in the day-ahead market. Executing this action leads to the next moment, where the market... The agent receives feedback from resources, such as actual economic benefits, operating costs, penalties for breach of constraints, and conditional risk value. It records the experience of "taking action A in state S, receiving reward R, and transitioning to a new state S'". Through massive trial and error experience, the agent continuously adjusts its neural network parameters to determine the reward function, which is designed in the form of multi-objective optimization: R = coefficient 1 × market return + coefficient 2 × penalty for breach of contract - coefficient 3 × operating cost - coefficient 4 × risk metric. It can output the optimal action that brings the highest long-term return in the face of any complex state. By adjusting the weights of each item in the reward function, multi-objective optimization can be achieved, such as maximizing returns, minimizing risks, and minimizing deviations. Finally, it outputs dispatch instructions and joint bidding strategies for multiple electricity markets, that is, specific control commands for the internal standardized resource pool and bid quantity and price combinations for multiple electricity markets.
[0031] Furthermore, step S300 also includes the following: the deep reinforcement learning agent adopts a deep Q-network algorithm, and the action space includes power setting and adjustment instructions for multimodal distributed resources and bidding strategies for multiple electricity markets; the reward function of the deep reinforcement learning agent is a multi-objective optimization form, and the multi-objectives include total market revenue, penalty for default, operating costs, and conditional risk value, wherein the total market revenue is the sum of electricity market revenue, ancillary service market revenue, and carbon trading revenue, and the conditional risk value is used to quantify the tail risk under multi-market coupling.
[0032] Furthermore, step S300 also includes the conditional value at risk being used to measure multi-market coupling risk, comprising an expected return term and a risk adjustment term.
[0033] Preferably, the Deep Q-Network algorithm is a powerful algorithm in the field of deep reinforcement learning. Its core is a deep neural network, known as a "Q-Network." It calculates the Q-value for each possible action based on a given description of a standardized resource pool state and a sequence of predicted future scenarios. The Q-value represents the estimated long-term expected cumulative reward for performing that action in the current state and subsequently following the optimal policy. The deep reinforcement learning agent simply needs to choose the action with the highest Q-value in the current state. The action space includes power setting and adjustment instructions for multimodal distributed resources, directly corresponding to the control of the standardized resource pool, for example... Actions could include setting the total output of the adjustable resource pool to 50MW within the next 15 minutes and ultimately allocating it to specific wind turbines, energy storage, loads, and other equipment; and bidding strategies targeting multiple electricity markets to define the market application status of virtual power plants. For example, actions could include applying to sell 30MWh of electricity between 10:00 and 11:00 the following day in the day-ahead electricity market at a price of 500 yuan / MWh; applying to provide 10MW of up-regulation capacity the following day in the frequency regulation ancillary services market at a price of 80 yuan / MW; and applying to sell 100 tons of carbon emission allowances in the carbon market.
[0034] Preferably, the reward function of the deep reinforcement learning agent is a multi-objective optimization form. These multi-objectives include total market revenue, penalty for default, operating costs, and conditional risk value. Total market revenue is the sum of electricity market revenue, ancillary service market revenue, and carbon trading revenue. Specifically, electricity market revenue is the difference between electricity sales revenue and electricity purchase cost; ancillary service market revenue is the reward obtained by providing services such as frequency regulation and backup power; and carbon trading revenue is the revenue obtained from selling surplus carbon allowances or the cost of purchasing allowances. Penalty for default is a constraint guarantee; if the actual output of the virtual power plant does not reach the amount it promised in the market bid or fails to fulfill the instructions issued by the dispatching agency, it will be penalized. Operating costs include fuel costs for distributed generation, equipment start-up and shutdown losses, energy storage cycle degradation, and response compensation fees paid to flexible load users. The total revenue Rtotal = Re + Ra + Rc includes: electricity market revenue. Ancillary service revenue Carbon trading revenue: Specifically, Re represents the electricity market revenue, λtePe,tb represents the electricity sales revenue, λte represents the electricity market price at time t, Pe,tb refers to the electricity bidding volume of the virtual power plant at time t, and Ce(Pe,tb) represents the electricity supply cost, including DER generation variable costs, demand response compensation costs, energy storage cycle losses, etc., typically in the form Ce(P) = aP² + bP + c (quadratic cost function). For example, if a virtual power plant bids for 50MW during peak hours at an electricity price of 650 yuan / MWh and a supply cost of 18,000 yuan, then the net electricity revenue for that period is 650 × 50 - 18,000 = 14,500 yuan, where Rc represents carbon trading revenue, λc represents the carbon quota market price, and Eb... ase represents the baseline carbon emission allowance, which is usually determined by the regulatory authorities based on historical emission intensity. ∑tetPe,tb represents the actual carbon emissions. et represents the marginal emission factor for time period t, with et=0 for clean energy such as wind and solar. Ec,tb represents the net trading volume in the carbon market, with positive values representing the sale of allowances and negative values representing the purchase of allowances. Ra represents the ancillary service market revenue. Ca(Pa,tb) represents the ancillary service opportunity cost, which is the reduced electricity market opportunity revenue due to reserved standby capacity, including the preparation cost of standby capacity, such as energy storage standby losses. λta+ represents the increase in standby service price for time period t, λta- represents the decrease in standby service price for time period t, and rt+ and rt- represent the increase / decrease in ancillary service capacity declared and won in time period t.
[0035] Preferably, conditional value at risk is used to quantify tail risk in a multi-market coupling environment. This refers to the risk that occurs under extremely unfavorable market scenarios, such as a sharp drop in electricity prices leading to both booming business and a subsequent sharp decline in profits or even losses. The objective function incorporates the idea of profit-risk equilibrium optimization and includes the following two key parts: the expected profit term... E[Rtotal] represents the expected total revenue of a virtual power plant participating in multi-market operations, and ρ∈[0,1] is the risk preference coefficient, reflecting the decision-maker's degree of risk aversion. Weighting reflects the degree of emphasis on expected returns, while the risk adjustment item... CVaRα represents the conditional value of risk at confidence level α. This means that the gains are treated as negative losses for applying CVaR calculations, which quantifies the average loss level under extremely adverse scenarios. When ρ=0, it is a pure gain maximization strategy, and when ρ=1, it is a completely risk-averse strategy. By introducing CVaR as a negative reward term, the agent, when learning to maximize long-term rewards, will actively avoid strategies that may have high expected returns but lead to catastrophic losses in extreme situations.
[0036] In step S400, the scheduling instruction is sent to the underlying physical device for execution, and the joint bidding strategy is submitted to multiple electricity markets. The strategy is continuously iterated and optimized based on the responses.
[0037] Preferably, a master-follower game framework is constructed, placing the virtual power plant in the position of a core decision-maker, interacting with multiple electricity markets. Here, the master-follower game is an economic model used to describe a specific type of leader-follower game, where one participant makes a decision first, and then another participant makes its own decision knowing the leader's decision. The virtual power plant, as the leader, is used to formulate and declare its resource aggregation's energy-ancier service joint output curve to the electricity energy market and the frequency regulation ancillary service market, that is, how much capacity is used for electricity sales and how much capacity is reserved for providing frequency regulation services at different times. At the same time, it calculates its carbon footprint and participates in carbon market trading. Followers include the electricity energy market, the frequency regulation ancillary service market, and the carbon trading market. The electricity energy market includes the medium- and long-term trading market and the spot trading market. Based on the declaration information of the virtual power plant and other market participants, with the goal of maximizing social welfare or minimizing system operating costs, market clearing is carried out to form the nodal marginal electricity price, the frequency regulation ancillary service clearing price, and the carbon trading price, and the price signals are fed back to the virtual power plant. This game is a dynamic iterative process. The virtual power plant formulates its bidding strategy based on price forecasts and market rules. Each market clears and returns a price. The virtual power plant readjusts its bidding strategy based on the clearing price. This continues until a Nash equilibrium is reached. At the equilibrium point, the virtual power plant's profit is maximized, and no single entity can benefit by changing its strategy alone.
[0038] Step S400 further includes the following: each market clears itself based on a master-slave game, with the virtual power plant acting as the leader and submitting its declaration first, while the electricity market, ancillary services market, and carbon trading market act as followers and clear their markets respectively, returning clearing prices; the virtual power plant, based on the returned clearing price information and the collected actual resource output and user response behavior information, along with the state and actions in the deep reinforcement learning agent, forms an experience tuple, which is stored in an experience replay buffer; the network parameters of the deep reinforcement learning agent are updated based on the experience tuple in the experience replay buffer, and the strategy is continuously iteratively optimized.
[0039] Preferably, dispatch instructions are issued to the underlying physical equipment to execute dynamic simulations. Each market clears itself based on a master-slave game, with the virtual power plant acting as the leader and submitting its bid first. The electricity market, ancillary services market, and carbon trading market act as followers and clear their respective markets. Specifically, the virtual power plant, based on the decisions of a deep reinforcement learning agent, submits its bidding strategies to multiple electricity trading markets, such as electricity quantity and price, ancillary service capacity, and carbon quota trading volume, predicting the market's reaction to its bidding strategies. The electricity market, ancillary services market, and carbon trading market, upon receiving the bids from the virtual power plant and other participating markets, then proceed with the market clearing process. After market participants submit their bids, clearing calculations are performed according to their respective independent market rules. For example, the electricity market typically uses safe-constrained economic dispatch with the goal of maximizing social welfare, calculating the marginal electricity price and trading volume at each node; the ancillary services market clears based on bid ranking; and the carbon market determines the trading price based on supply and demand matching. The clearing price and trading volume are then returned as a direct quantitative feedback to the environment on the bidding actions taken by virtual power plants. For example, if a virtual power plant's bid price is too high, it may fail to win the bid for its electricity volume, resulting in zero revenue; if the bid price is too low, it may win the bid but with meager profits.
[0040] Preferably, the virtual power plant uses the returned clearing price information to verify the effectiveness of the bidding strategy, while simultaneously collecting actual resource output, such as the actual power generation / charging power of wind turbines, photovoltaics, and energy storage; user response behavior information, such as the actual reduction or shifting of flexible loads, to verify the execution of dispatch instructions and reflect the deviation between the actual response and the expected model; and then acquiring the state and actions of the deep reinforcement learning agent, which are the standardized resource pool state observed by the agent when making decisions and the future scenario prediction, as well as the specific dispatch instructions and bidding strategies output by the agent, to jointly constitute an experience tuple, which is stored in the experience replay buffer. Its function includes randomly sampling samples from the buffer for learning, breaking the strong temporal correlation of the original data, making the neural network training more stable, and allowing valuable interaction data to be used for learning multiple times, thereby improving data efficiency.
[0041] Preferably, the network parameters of the deep reinforcement learning agent are updated based on the experience tuples in the experience replay buffer. Specifically, a small batch of experience tuples is randomly selected from the experience replay buffer. For each sample, the target Q-network is used to calculate the estimated value of the best long-term reward that can be obtained in the next new state after performing an action. The predicted Q value of the current main Q-network for the current state is compared with the calculated target value, and the mean squared error is calculated as the loss function. The error is backpropagated according to the loss function through the gradient descent algorithm to update the weight parameters of the main Q-network so that its predicted Q value is closer to the real long-term reward. Then, the policy is continuously iterated and optimized. That is, the new policy generates new actions and interacts with the environment and market in a new way, which generates new experience tuples and stores them in the experience replay buffer. The new round of learning draws samples from the buffer to continue updating the network. This cycle repeats, and the agent's policy evolves, adjusts and optimizes continuously in the continuous interaction with the complex and dynamic environment, gradually approaching the optimum.
[0042] Furthermore, step S400 also includes the master-slave game including physical constraints, market constraints, and time coupling constraints; wherein, physical constraints include power balance constraints, reserve capacity constraints, and carbon emission constraints; market constraints include power-ancillary service capacity coupling constraints and cross-market arbitrage constraints.
[0043] Preferably, the master-slave game also includes physical constraints, market constraints, and time coupling constraints. Physical constraints include power balance constraints. It must meet the 15-minute granularity balance requirement, allowing for 5% short-term imbalance (<5 minutes) through real-time market adjustments, and has reserve capacity constraints. Carbon emission constraints Where, Pdis / ch,t represents the discharge / charge power of energy storage, Dt represents the net load demand, including uncontrollable and adjustable loads, Pimax represents the maximum power generation of the i-th DER, Pimin represents the minimum power generation of the i-th DER, Pi,t represents the power generation of the i-th DER at time t, λc represents the carbon allowance market price, Ebase represents the benchmark carbon emission allowance, usually determined by the regulatory authorities based on historical emission intensity, ∑tetPe,tb represents the actual carbon emissions, et represents the marginal emission factor for time period t, and et=0 for clean energy such as wind and solar. Ec,tb represents the net trading volume in the carbon market; market constraints include the electricity-ancillary service capacity coupling constraint Pe,tb+Pa,tb≤Paggmax, and the cross-market arbitrage constraint λta+≥γλte, where Pe,tb represents the electricity bid volume at time t, Pa,tb represents the ancillary service bid volume at time t, and Paggmax represents the maximum total aggregated electricity of the virtual power plant; time coupling constraints Where Rup / down represents the ramp rate limit, Δt represents the market time interval, and Pa,tb represents the ancillary service bid volume at time t.
[0044] In the above text, refer to Figure 1 A multimodal resource collaborative optimization scheduling method for a virtual power plant according to embodiments of the present invention is described in detail. Next, reference will be made to... Figure 2 A multimodal resource collaborative optimization scheduling system for a virtual power plant according to an embodiment of the present invention is described.
[0045] The multimodal resource collaborative optimization scheduling system for virtual power plants according to embodiments of the present invention addresses the technical problems of insufficient multimodal resource collaboration, disconnect between response assessment and scheduling, and lack of multi-market collaborative mechanisms in existing technologies. It achieves the technical effect of improving the economy, reliability, and market competitiveness of virtual power plant operation by generating collaborative scheduling instructions and multi-market joint bidding strategies through deep learning agents. Figure 2 As shown, the multimodal resource collaborative optimization scheduling system of the virtual power plant includes: a power field resource aggregation module 10, a prediction sequence generation module 20, a bidding strategy output module 30, and a strategy optimization module 40.
[0046] The electric field resource aggregation module 10 is used to uniformly model and dynamically aggregate multimodal distributed resources within the virtual power plant, generating a standardized resource pool that characterizes the aggregation response capability; the prediction sequence generation module 20 is used to generate prediction sequences for various future scenarios based on external market and environmental information; the bidding strategy output module 30 is used to make decisions through a deep reinforcement learning agent using the state of the standardized resource pool and the prediction sequence as the state observation space, and output scheduling instructions and joint bidding strategies for multiple electricity markets; the strategy optimization module 40 is used to send the scheduling instructions to the underlying physical devices for execution, and simultaneously submit the joint bidding strategy to multiple electricity markets, continuously iterating and optimizing the strategy based on the response.
[0047] The specific configuration of the electric field resource aggregation module 10 will be described in detail below. The electric field resource aggregation module 10 further includes: modeling the independent response capabilities of multi-source heterogeneous resources within the virtual power plant; constructing a unified feature extraction model based on a graph neural network to extract the spatiotemporal dynamic features of the multi-source heterogeneous resources; establishing a unified evaluation index system; assigning weights to each index in the evaluation index system based on the spatiotemporal dynamic features using the entropy method; and mapping and aggregating the multi-source heterogeneous resources into three standard types of resource pools—interruptible, movable, and adjustable—according to different weight combinations, generating the standardized resource pool.
[0048] The specific configuration of the electric field resource aggregation module 10 will be described in detail below. The electric field resource aggregation module 10 further includes: the evaluation index system includes response time, duration, power upper and lower limits, ramp rate, and average response time.
[0049] The specific configuration of the bidding strategy output module 30 will be described in detail below. The bidding strategy output module 30 further includes: the deep reinforcement learning agent employs a deep Q-network algorithm; its action space includes power setting and adjustment instructions for multimodal distributed resources and bidding strategies for multiple electricity markets; the reward function of the deep reinforcement learning agent is a multi-objective optimization form, with multiple objectives including total market revenue, penalty for default, operating costs, and conditional risk value. The total market revenue is the sum of electricity market revenue, ancillary service market revenue, and carbon trading revenue; the conditional risk value is used to quantify tail risk under multi-market coupling.
[0050] The specific configuration of the bidding strategy output module 30 will be described in detail below. The bidding strategy output module 30 further includes: the conditional value at risk (VaR) used to measure multi-market coupling risk, comprising an expected return term and a risk adjustment term.
[0051] The specific configuration of the strategy optimization module 40 will be described in detail below. The strategy optimization module 40 further includes: each market clearing based on a master-slave game, where the virtual power plant, as the leader, submits its declaration first, and the electricity market, ancillary services market, and carbon trading market, as followers, clear their markets and return clearing prices; the virtual power plant, based on the returned clearing price information and the collected actual resource output and user response behavior information, along with the state and actions in the deep reinforcement learning agent, constitutes an experience tuple, which is stored in an experience replay buffer; the network parameters of the deep reinforcement learning agent are updated based on the experience tuple in the experience replay buffer, and the strategy is continuously iteratively optimized.
[0052] The specific configuration of the strategy optimization module 40 will be described in detail below. The strategy optimization module 40 further includes: the master-slave game also includes physical constraints, market constraints and time coupling constraints; among which, physical constraints include power balance constraints, reserve capacity constraints and carbon emission constraints; market constraints include power-ancillary service capacity coupling constraints and cross-market arbitrage constraints.
[0053] The multimodal resource collaborative optimization scheduling system for virtual power plants provided in this embodiment of the invention can execute the multimodal resource collaborative optimization scheduling method for virtual power plants provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0054] Figure 3This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention, showing a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present invention. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention. This electronic device is in the form of a general-purpose computing device, and its components may include, but are not limited to, an input device 401, a processor 402, a memory 403, and an output device 404. The processor 402 may be one or more; the memory 403 may include a computer-readable medium and at least one program product having a set (at least one) of program modules configured to perform the functions of the embodiments of this application.
[0055] The memory 403 shown in this embodiment of the invention can be any combination of one or more computer-readable media. The computer-readable storage media can be, but is not limited to, infrared, semiconductor systems, devices or components, or any combination thereof, used to store software programs, computer-executable programs and modules, such as the program instructions / modules corresponding to the multimodal resource collaborative optimization scheduling method of the virtual power plant in this embodiment of the invention. The processor 402 executes various functional applications and data processing of the computer device by running the software programs, instructions and modules stored in the memory 403, thereby realizing the multimodal resource collaborative optimization scheduling method of the virtual power plant described above.
[0056] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0057] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multimodal resource collaborative optimization scheduling method for virtual power plants, characterized in that, include: The multimodal distributed resources within the virtual power plant are modeled and dynamically aggregated in a unified manner to generate a standardized resource pool that represents the aggregation response capability. Based on external market and environmental information, a series of predictions for various future scenarios are generated. Using the state of the standardized resource pool and the predicted sequence as the state observation space, a deep reinforcement learning agent makes decisions and outputs scheduling instructions and joint bidding strategies for multiple electricity markets. The scheduling instructions are sent to the underlying physical devices for execution, and the joint bidding strategy is submitted to multiple electricity markets. The strategy is continuously iterated and optimized based on the responses.
2. The multimodal resource collaborative optimization scheduling method for virtual power plants as described in claim 1, characterized in that, The system performs unified modeling and dynamic aggregation of multimodal distributed resources within the virtual power plant, generating a standardized resource pool that characterizes the aggregation response capability, including: Independent response capability modeling is performed on multi-source heterogeneous resources within a virtual power plant. A unified feature extraction model is constructed based on graph neural networks to extract the spatiotemporal dynamic features of the multi-source heterogeneous resources. A unified evaluation index system is established. Based on the spatiotemporal dynamic characteristics, the entropy method is used to assign weights to each index in the evaluation index system. According to different weight combinations, the multi-source heterogeneous resources are mapped and aggregated into three standard types of resource pools: interruptible, transferable, and adjustable, thus generating the standardized resource pool.
3. The multimodal resource collaborative optimization scheduling method for virtual power plants as described in claim 2, characterized in that, The evaluation index system includes response time, duration, power upper and lower limits, ramp rate, and average response time.
4. The multimodal resource collaborative optimization scheduling method for virtual power plants as described in claim 1, characterized in that, The deep reinforcement learning agent employs a deep Q-network algorithm, and its action space includes power setting and adjustment commands for multimodal distributed resources as well as bidding strategies for multiple electricity markets. The reward function of the deep reinforcement learning agent is a multi-objective optimization form, which includes total market revenue, penalty for default, operating costs, and conditional risk value. The total market revenue is the sum of electricity market revenue, ancillary service market revenue, and carbon trading revenue. The conditional risk value is used to quantify the tail risk under multi-market coupling.
5. The multimodal resource collaborative optimization scheduling method for virtual power plants as described in claim 4, characterized in that, The conditional value at risk is used to measure multi-market coupling risk and includes an expected return term and a risk adjustment term.
6. The multimodal resource collaborative optimization scheduling method for virtual power plants as described in claim 1, characterized in that, The scheduling instructions are issued to the underlying physical devices for execution, and the joint bidding strategy is submitted to multiple electricity markets. The strategy is continuously iterated and optimized based on the responses, including: Each market clears out based on a leader-follower game, with virtual power plants acting as leaders and submitting declarations first, while the electricity market, ancillary services market, and carbon trading market act as followers and clear out their respective markets and return clearing prices. The virtual power plant, based on the returned clearing price information and the collected actual resource output and user response behavior information, together with the state and actions in the deep reinforcement learning agent, constitutes an experience tuple, which is stored in the experience replay buffer. The network parameters of the deep reinforcement learning agent are updated based on the experience tuples in the experience replay buffer to continuously iterate and optimize the policy.
7. The multimodal resource collaborative optimization scheduling method for virtual power plants as described in claim 6, characterized in that, Master-slave game also includes physical constraints, market constraints, and time coupling constraints; Physical constraints include power balance constraints, reserve capacity constraints, and carbon emission constraints; market constraints include power-ancillary service capacity coupling constraints and cross-market arbitrage constraints.
8. A multimodal resource collaborative optimization scheduling system for virtual power plants, characterized in that, The system is used to implement the multimodal resource collaborative optimization scheduling method for virtual power plants according to any one of claims 1 to 7, and the system includes: The electric field resource aggregation module is used to perform unified modeling and dynamic aggregation of multimodal distributed resources within the virtual power plant, generating a standardized resource pool that characterizes the aggregation response capability. The prediction sequence generation module is used to generate prediction sequences for various future scenarios based on external market and environmental information. The bidding strategy output module is used to make decisions through a deep reinforcement learning agent using the state of the standardized resource pool and the predicted sequence as the state observation space, and output scheduling instructions and joint bidding strategies for multiple electricity markets. The strategy optimization module is used to send the scheduling instructions to the underlying physical devices for execution, and submit the joint bidding strategy to multiple electricity markets, and continuously iterate and optimize the strategy based on the responses.
9. An electronic device, characterized in that, The electronic device includes: Memory, used to store executable instructions; The processor, when executing executable instructions stored in the memory, implements the multimodal resource collaborative optimization scheduling method for the virtual power plant as described in any one of claims 1-7.