Dynamic regulation and control system, method and equipment for virtual power plant based on digital twinning
The virtual power plant dynamic control system, which utilizes digital twin technology and deep reinforcement learning, solves the problems of insufficient prediction accuracy and static scheduling in virtual power plants, and achieves dynamic collaborative optimization of multiple resources, thereby improving the stability and economy of the power grid.
Patent Information
- Application Number
- CN202512051877.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-07
AI Technical Summary
Existing virtual power plant systems suffer from insufficient forecasting accuracy, static scheduling strategies, and difficulties in collaborative optimization, leading to resource waste and market revenue losses, and impacting grid stability and economy.
A virtual power plant dynamic control system based on digital twins is adopted. Real-time data is acquired through the data acquisition and sensing module, the digital twin modeling module constructs and updates the digital twin model, the intelligent decision optimization module generates joint action instructions using a deep reinforcement learning strategy network, and closed-loop control is achieved through the instruction execution and feedback module.
It improves prediction accuracy, enables dynamic adaptive scheduling, optimizes multi-resource collaboration, and enhances the economy and reliability of virtual power plants.
Smart Images

Figure CN121813558A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of power system dispatching, in particular to a virtual power plant dynamic regulation system, method and equipment based on digital twinning. BACKGROUND
[0002] As an advanced energy management system, the virtual power plant integrates geographically dispersed distributed energy resources, including solar photovoltaic, wind power, energy storage systems and controllable loads, through information communication technology and intelligent software platform, to form a unified controllable whole participating in power grid operation and power market transaction.
[0003] However, the prior art has significant defects in actual application. The problem of insufficient prediction accuracy is particularly prominent. The traditional virtual power plant system relies on statistical or shallow machine learning methods for new energy output and load prediction, but renewable energy such as photovoltaic and wind power is highly intermittent and volatile due to weather influence, and load demand also has random variation characteristics, resulting in a large deviation between the prediction result and the actual operation state, and further causing the dispatching plan to deviate from the actual demand, causing resource waste and market revenue loss. The problem of static dispatching strategy also restricts system performance. The current dispatching model mostly uses deterministic optimization or rolling optimization method, the core of which is based on a simplified linear model, which cannot adapt to power grid frequency fluctuations, load mutations and price signal instantaneous changes in real time, and cannot accurately depict the nonlinear operation characteristics and aging degradation mechanism of distributed energy, making it difficult to achieve global optimization of the generated dispatching instructions, affecting the stability and economy of the power grid. In addition, the difficulty of collaborative optimization has become a technical bottleneck. The virtual power plant contains heterogeneous resources with different response speeds, adjustment costs and operating lives, such as energy storage devices with fast response but limited life, and photovoltaic systems with large output fluctuations. How to balance the economic interests of each resource party while meeting the power grid dispatching instructions and considering the device health status to achieve multi-objective dynamic collaborative optimization is still a difficult problem to be solved. These defects seriously hinder the efficient operation of virtual power plants in complex power grid environments and the improvement of market competitiveness. SUMMARY
[0004] The purpose of the embodiments of the present application is to provide a virtual power plant dynamic regulation system, method and equipment based on digital twinning, which has the advantages of improving prediction accuracy, realizing dynamic adaptive dispatching and optimizing multi-resource collaboration.
[0005] To solve the above technical problems, the embodiments of the present application provide a virtual power plant dynamic regulation system based on digital twinning, comprising: a data acquisition and perception module, configured to acquire the operating state data, environment data and external power grid and market data of each physical distributed energy unit in the virtual power plant in real time; A digital twin modeling module is configured to construct and update in real time a digital twin model corresponding to each physical distributed energy unit in the virtual power plant. An intelligent decision optimization module is configured to take the environment formed by all the digital twin models in the digital twin modeling layer as an interactive object, generate a joint action instruction for regulating each physical distributed energy unit based on the system state output by the digital twin environment module, and through a pre-trained deep reinforcement learning policy network. An instruction execution and feedback module is configured to issue the joint action instruction to the corresponding physical distributed energy unit for execution, and collect actual execution results and feed back to the data acquisition and perception module.
[0006] To solve the above technical problems, the embodiment of the present application provides a virtual power plant dynamic regulation method based on digital twinning, comprising: Real-time acquisition of operation state data, environment data, and external power grid and market data of each physical distributed energy unit in the virtual power plant; Based on the operation state data, the environment data, and the external power grid and market data, a digital twin model corresponding to each physical distributed energy unit in the virtual power plant is constructed and updated in real time; Based on the current state of all the digital twin models, the environment data, and the external power grid and market data, a joint action instruction for regulating each physical distributed energy unit is generated through a pre-trained deep reinforcement learning policy network; The joint action instruction is issued to the corresponding physical distributed energy unit for execution, and the actual execution results are collected to form a closed-loop feedback.
[0007] To solve the above technical problems, the present application adopts one technical solution: to provide a computer device, comprising one or more processors; a memory for storing one or more programs, so that one or more processors implement the virtual power plant dynamic regulation method based on digital twinning described in any one of the above.
[0008] The embodiment of the present application provides a virtual power plant dynamic regulation system, method and equipment based on digital twinning. The system comprises: a data acquisition and perception module, which is used for acquiring the running state data, environment data and external power grid and market data of each physical distributed energy unit in the virtual power plant in real time; a digital twinning modeling module, which is used for constructing and updating the digital twinning model corresponding to each physical distributed energy unit in the virtual power plant in real time; an intelligent decision optimization module, which is used for taking the environment formed by all the digital twinning models in the digital twinning modeling layer as an interactive object, generating a joint action instruction for regulating each physical distributed energy unit based on the system state output by the digital twinning environment module through a pre-trained deep reinforcement learning strategy network; and an instruction execution and feedback module, which is used for issuing the joint action instruction to the corresponding physical distributed energy unit for execution and collecting actual execution results and feeding back to the data acquisition and perception module. The embodiment of the present application realizes a dynamic regulation closed loop through real-time data acquisition, digital twinning modeling and intelligent decision optimization, has the ability to improve the prediction accuracy, realize dynamic adaptive scheduling and optimize multi-resource cooperation, thereby effectively solving the problems of prediction deviation, static scheduling and cooperation difficulty, and improving the economy and reliability of the virtual power plant. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the scheme in the present application, the drawings needed in the description of the embodiments of the present application will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0010] Figure 1 It is a digital twinning-based virtual power plant dynamic regulation system provided by the embodiment of the present application; Figure 2 It is a structure schematic diagram of a digital twinning environment module provided by the embodiment of the present application; Figure 3 It is an implementation flowchart of a digital twinning-based virtual power plant dynamic regulation method flow provided by the embodiment of the present application; Figure 4 It is an implementation flowchart of a first sub-flow in the digital twinning-based virtual power plant dynamic regulation method provided by the embodiment of the present application; Figure 5 It is an implementation flowchart of a second sub-flow in the digital twinning-based virtual power plant dynamic regulation method provided by the embodiment of the present application; Figure 6 It is an implementation flowchart of a third sub-flow in the digital twinning-based virtual power plant dynamic regulation method provided by the embodiment of the present application; Figure 7is a fourth sub-flow implementation flowchart of a virtual power plant dynamic regulation method based on digital twinning provided by the embodiment of the application; Figure 8 is a schematic diagram of a computer device provided by the embodiment of the application. DETAILED DESCRIPTION
[0011] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the specification herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the description and claims of the application as well as the above abstract are intended to cover all alternatives, modifications, and equivalents of the application falling within the scope of the application; the terms "comprising", "having", "including", and "containing" used in the specification herein are intended to be open-ended and do not exclude the presence of other elements or steps; the terms "first", "second" and the like used in the specification herein do not necessarily denote any ordinal, sequential or chronological significance, but are used to distinguish different features.
[0012] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase that the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of other embodiments. It is explicitly contemplated that embodiments described herein can be combined with other embodiments.
[0013] In order to enable persons skilled in the art to better understand the scheme of the application, the technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings.
[0014] Reference should be made to Figure 1 The application provides an embodiment of a virtual power plant dynamic regulation system based on digital twinning, which can be applied to various electronic devices.
[0015] As shown in Figure 1 The virtual power plant dynamic regulation system based on digital twinning of the embodiment includes: The data acquisition and perception module 10 is used to acquire the operation state data, environmental data, and external grid and market data of each physical distributed energy unit in the virtual power plant in real time; the digital twin modeling module 20 is used to build and update the digital twin model corresponding to each physical distributed energy unit in the virtual power plant in real time; the intelligent decision optimization module 30 is used to take the environment formed by all the digital twin models in the digital twin modeling layer as the interactive object, generate joint action instructions for regulating each physical distributed energy unit based on the system state output by the digital twin environment module through the pre-trained deep reinforcement learning strategy network; and the instruction execution and feedback module 40 is used to issue the joint action instructions to the corresponding physical distributed energy unit for execution, and collect the actual execution results and feed them back to the data acquisition and perception module 10.
[0016] The traditional existing virtual power plant system has made progress in aggregation and preliminary scheduling, but still faces challenges in actual application. Specifically, due to the intermittency, volatility of new energy and randomness of load, the prediction accuracy is insufficient; the existing scheduling strategy is mostly static optimization, which is difficult to respond to grid and market changes in real time, and cannot accurately describe the complex operation characteristics of distributed energy; in addition, it is difficult to optimize the internal heterogeneous resources of the virtual power plant, and it is difficult to balance the economic benefits and equipment health status of each resource party.
[0017] To this end, the application proposes a virtual power plant dynamic regulation system based on digital twinning, which acquires multi-source data in real time through the data acquisition and perception module 10, builds and updates the digital twin model of each physical distributed energy unit in real time through the digital twin modeling module 20, generates joint action instructions based on the digital twin environment through the deep reinforcement learning strategy network of the intelligent decision optimization module 30, and realizes closed-loop regulation through the instruction execution and feedback module 40, thereby effectively improving the prediction accuracy, realizing dynamic scheduling and optimizing collaborative management.
[0018] For ease of understanding, some key terms in the present embodiment are explained as follows: Digital twinning refers to the digital mapping of a physical entity in a virtual space, which realizes the bidirectional interaction and synchronization between the physical entity and the virtual model through real-time data connection.
[0019] Virtual power plant refers to the aggregation of geographically dispersed distributed energy, energy storage systems, and controllable loads, which are coordinated and controlled through information communication technology and intelligent management systems, and participate in the energy management system as a whole in the power market and grid operation.
[0020] Physical distributed energy unit refers to the actual physical equipment that constitutes the virtual power plant, such as solar photovoltaic, wind turbine, energy storage battery, electric vehicle charging pile, and controllable load, etc.
[0021] Deep reinforcement learning policy network refers to an artificial intelligence model combining deep learning and reinforcement learning, which learns the optimal decision-making strategy through interaction with the environment to maximize long-term rewards.
[0022] Joint action instruction refers to a set of control instructions generated by the intelligent decision optimization module 30 for simultaneously regulating multiple physically distributed energy units in the virtual power plant, such as the power output, charging and discharging state, or load start and stop of each unit.
[0023] The data acquisition and perception module 10 is configured to acquire real-time operation state data, environmental data, and external grid and market data of each physically distributed energy unit in the virtual power plant. As an implementation, this module can use independent sensors and communication devices, directly connected to each physically distributed energy unit through wired or wireless means, and upload data to the central server regularly. For example, each photovoltaic panel can be equipped with a power sensor, uploading data every minute. As another implementation, this module can batch acquire aggregated data from existing energy management systems (EMS) or SCADA systems through their data interfaces. For example, acquire the state of charge data of all energy storage units from the existing EMS of the virtual power plant operator.
[0024] The digital twin modeling module 20 is configured to build and update real-time digital twin models corresponding to each physically distributed energy unit in the virtual power plant. Specifically, a mathematical model based on physical mechanisms can be established for each physically distributed energy unit, such as for energy storage batteries, establishing physical equation models of charging and discharging efficiency, capacity attenuation, etc., which are initialized through preset parameters. Alternatively, a purely data-driven approach can be used, training machine learning models with historical operation data, such as using neural network models to learn the relationship between photovoltaic output and light intensity, temperature. Further, a hybrid model can also be constructed, combining mechanism models and data-driven models, using mechanism models as the basic framework, and using data-driven models to calibrate uncertain parameters in mechanism models or model residuals.
[0025] The intelligent decision optimization module 30 is configured to interact with a digital twin environment constituted by all the digital twin models built by the digital twin modeling module 20, and generate joint action instructions for regulating each physical distributed energy unit based on the system state reflected by the digital twin environment through a pre-trained deep reinforcement learning policy network. As an implementation manner, the deep reinforcement learning policy network can be trained to maximize a simple reward function, for example, only considering the total power generation revenue of the virtual power plant. During the training process, the policy network learns to output the optimal joint action instructions under different system states through a large number of interactions with the digital twin environment. As another implementation manner, the policy network can be trained to minimize a simple cost function, for example, only considering the grid imbalance penalty. The policy network gradually adjusts its parameters through iterative optimization to find a sequence of actions that can effectively reduce the cost in the digital twin environment.
[0026] The instruction execution and feedback module 40 is configured to issue the joint action instructions to the corresponding physical distributed energy units for execution and collect the actual execution results to the data collection and perception module 10. Specifically, the joint action instructions can be sent to the local controllers of each physical distributed energy unit through a standard communication protocol (for example, Modbus TCP / IP, IEC 61850), and the local controllers are responsible for the parsing and execution of the instructions. The execution results are returned through the same communication link. As another implementation manner, a cloud control platform can be used, and the joint action instructions are first sent to the cloud platform, and then issued to each distributed energy unit by the cloud platform through a wide area network. The actual execution data is uploaded to the cloud platform by each unit, and then forwarded to the data collection and perception module 10 by the cloud platform.
[0027] The virtual power plant dynamic regulation system based on digital twin proposed in the present application can accurately reflect the running state of the physical distributed energy unit and the change of the external environment through real-time and multi-source data collection combined with high-fidelity digital twin models, thereby significantly improving the prediction accuracy of new energy output and load. The intelligent decision optimization module 30 uses deep reinforcement learning to make dynamic decisions in the digital twin environment, realizes real-time and adaptive regulation of the virtual power plant, and effectively solves the limitations of traditional static scheduling strategies. In addition, through the closed-loop instruction execution and feedback mechanism, the system can continuously optimize, meet the grid demand, and take into account the economic benefits and equipment health of each distributed energy unit, thereby overcoming the problem of collaborative optimization of heterogeneous resources.
[0028] In some embodiments of the present application, a digital twin environment module is proposed to output system states to support the intelligent decision optimization module 30 to generate joint action instructions. However, in the implementation process, due to the lack of specific state updating mechanism, performance prediction function and health assessment means, the digital twin model cannot accurately reflect the real-time state of the physical distributed energy unit, it is difficult to predict future performance changes and it is impossible to quantify the impact of control actions on equipment life, resulting in insufficient prediction accuracy, scheduling strategies that are difficult to dynamically respond to changes in the power grid and market, and the inability to effectively optimize economic benefits and equipment health.
[0029] Please refer to Figure 2 , Figure 2 The structure of the digital twin environment module provided by the embodiments of the present application is shown. The digital twin environment module includes: a state mapping unit 201 for receiving operating state data and external environment data from a physical distributed energy unit, and updating the real-time state of a digital twin model according to the operating state data and the external environment data; a performance simulation unit 202 for predicting changes in performance parameters of the physical distributed energy unit within a future set period based on the current state of the digital twin model and received control instructions, and obtaining a prediction result; and a health assessment unit 203 for quantifying the health state degradation of the physical distributed energy unit and the life loss cost caused by control actions according to historical operating data of the physical distributed energy unit and the prediction result.
[0030] The state mapping unit 201 is responsible for synchronizing the real-time operation data and external environment information of the distributed energy unit in the physical world to its corresponding digital twin model accurately and timely, ensuring that the state of the digital twin model is consistent with the physical entity. This ensures that the digital twin model can accurately reflect the current state of the physical entity in real time, providing a reliable real-time data foundation for subsequent performance simulation, health assessment and intelligent decision-making. Specifically, the state mapping unit 201 can connect with the controller or sensor network of the physical distributed energy unit through data interfaces such as Modbus TCP, OPC UA, etc., and collect real-time operation state data such as power, voltage, current, battery state of charge and health status, etc. At the same time, external environment data such as light and wind speed, ambient temperature, power grid frequency, etc. are obtained through weather stations, power grid data interfaces, etc. The collected data is preprocessed (such as data cleaning, format conversion) and directly updates the corresponding state variables in the digital twin model. Alternatively, the state mapping unit 201 can use an event-driven mechanism, which triggers data collection and updating when the operation state of the physical distributed energy unit or the external environment data changes significantly. For example, data thresholds or change rates can be set, and once the preset value is exceeded, state synchronization is performed immediately. In addition, a time stamp synchronization mechanism can be combined to ensure the time consistency of the physical world and the digital world.
[0031] The performance simulation unit 202 is responsible for predicting the performance parameter change trend of the physical distributed energy unit in a future specific time period based on the current state of the digital twin model and the received control instructions, using the built-in simulation model or algorithm. This can predict the impact of control instructions on the performance of the physical distributed energy unit, provide forward-looking performance prediction information for the intelligent decision optimization module 30, and support the system to make better scheduling decisions and avoid blind operation. Specifically, the performance simulation unit 202 can be built-in with a physical mechanism-based model, for example, for a photovoltaic unit, a physical model between light intensity, temperature and output power can be established; for the energy storage unit, an electrochemical model between charging and discharging current, temperature, state of charge and voltage can be established. When receiving control instructions (such as setting power output, charging and discharging rate), combined with the current state of the digital twin model, iterative calculation is performed through these mechanism models to predict the performance parameters such as power output, state of charge and voltage in the future period of time. Alternatively, the performance simulation unit 202 can use a data-driven simulation model, for example, a deep learning model (such as long short-term memory network LSTM, gated recurrent unit GRU) or a regression model trained based on historical operation data. These models can learn the performance response law of the physical distributed energy unit under different operating conditions and control instructions. When receiving control instructions and the current state, the model can predict the change of future performance parameters. For example, predicting the power curve, battery state of charge change, etc. in the next 15 minutes or 1 hour.
[0032] The health assessment unit 203 is responsible for quantitatively evaluating the degree of health state degradation of the equipment and the life consumption cost caused by the regulation action according to the historical operation data of the physically distributed energy unit and the prediction results provided by the performance simulation unit 202. This can take the health status and life consumption of the equipment into account in decision-making, avoid over-pursuing short-term economic benefits at the expense of the long-term operation life of the equipment, and achieve the balance optimization between economic benefits and equipment health. Specifically, the health assessment unit 203 can establish a health assessment algorithm based on experience or physical degradation model. For example, for a battery energy storage system, the health state degradation of the battery can be evaluated according to the number of charge and discharge cycles, depth, temperature and other parameters, combined with the SOH (State of Health) model or life prediction model. The life consumption cost caused by the regulation action can be quantified as the economic value corresponding to the unit health loss, for example, how much maintenance or replacement cost corresponds to a decrease of 1% SOH. Alternatively, the health assessment unit 203 can use a data-driven health assessment method, for example, using machine learning algorithms (such as support vector machine SVM, random forest RF) to learn the failure mode and performance degradation trend in the historical operation data, thereby predicting the health state of the equipment. Combined with the prediction results of the performance simulation unit 202, the impact of a specific regulation action on the life of the equipment can be further evaluated. For example, by analyzing the degree of life acceleration loss of the equipment caused by high-power charge and discharge, frequent start-stop and other operations, the life consumption cost is converted into specific life consumption cost, for example, X hours of equipment life are reduced per hour of high-load operation, and the economic cost is converted.
[0033] By introducing the state mapping unit 201, the performance simulation unit 202 and the health assessment unit 203, the digital twin environment module of the present application can more comprehensively and accurately reflect and predict the state and behavior of the physically distributed energy unit. Specifically, the state mapping unit 201 ensures real-time synchronization between the digital twin model and the physical entity, solves the problem of insufficient prediction accuracy of traditional systems, and provides a high- confidence real-time data basis for subsequent decision-making. The performance simulation unit 202 then predicts the performance response of the physically distributed energy unit based on these real-time states and control instructions, enabling the system to dynamically respond to changes in the power grid and the market, overcoming the limitations of static scheduling strategies. Further, the health assessment unit 203 takes into account the health state degradation and life consumption cost of the equipment, so that the intelligent decision optimization module 30 considers not only economic benefits but also the long-term operation health of the equipment when generating joint action instructions, thereby achieving the coordinated optimization of economic benefits and equipment life, effectively solving the challenge of heterogeneous resource collaborative optimization. This detailed and comprehensive digital twin environment construction significantly improves the intelligent level and operation efficiency of the virtual power plant dynamic regulation system.
[0034] In some embodiments of the present application, a deep reinforcement learning policy network is proposed to generate the regulation instructions, however, in the implementation process, the training target may only focus on a single factor such as market revenue, while ignoring the influence of device aging loss and grid scheduling deviation, resulting in that the scheduling strategy cannot achieve global optimization, the device life is easily damaged, and the grid demand is difficult to fully meet.
[0035] To this end, the present application further proposes that the training target of the deep reinforcement learning policy network in the intelligent decision optimization module 30 is to maximize a composite reward value, and the calculation of the composite reward value at least couples the following factors: market revenue calculated according to power market information and the joint action instruction; a penalty term calculated according to the deviation between the joint action instruction and the grid scheduling demand; and the device aging loss cost related to the joint action instruction output by the health assessment unit 203.
[0036] Specifically, the training target of the deep reinforcement learning policy network is set to maximize a composite reward value, which means that the policy network will learn how to make decisions to achieve the best balance between multiple interrelated goals. The composite reward value is a comprehensive indicator to measure the goodness of an action, which integrates multiple dimensions of performance (such as economy, device health, and grid compliance) into a single numerical value, thereby guiding the policy network to optimize in a complex multi-objective scenario. For example, a reward function R(s, a) = w1*R_market + w2*R_penalty + w3*R_health can be defined, where w1, w2, w3 are weight coefficients, and R_market, R_penalty, R_health represent market revenue, grid deviation penalty, and device aging loss cost, respectively. The goal of the policy network is to learn a policy π(a|s) that maximizes the expected cumulative composite reward after taking action a in a given state s. In addition, the method of multi-objective reinforcement learning (MORL) can also be used, which regards each target (market revenue, penalty term, and device aging loss) as an independent reward dimension, and then integrates it into a single composite reward through Pareto optimization or weighted sum, or directly optimizes the Pareto frontier in the training process. The grid deviation penalty is calculated according to the difference between the actual response power of the VPP and the instruction power; the device aging loss cost is quantified by the health assessment module of the digital twin modeling module 20 according to the loss of device life caused by the current action.
[0037] The calculation of the composite reward value at least couples the above factors. "Coupling" here refers to combining multiple independent factors through some mathematical relationship (such as weighted sum, product, nonlinear function, etc.) to form a comprehensive evaluation index. This coupling ensures that the impact of a decision (joint action instruction) in different dimensions is fully considered when evaluating its pros and cons, avoiding the negative effects that may be caused by single target optimization. The most common coupling method is weighted sum, i.e. composite reward value = Σ (weight i * factor i). By adjusting the weights, the relative importance of different factors in the overall optimization goal can be reflected. Hierarchical or sequential coupling methods can also be used, for example, first ensure that the grid demand is met (penalty term), then maximize market revenue under the premise of meeting this constraint, and minimize equipment aging loss.
[0038] Among them, the market revenue calculated according to the power market information and the joint action instruction is the economic benefit obtained by the virtual power plant participating in the power market transaction (such as the spot market, auxiliary service market). It directly reflects the economic value brought by the virtual power plant to the power grid by aggregating and optimizing the output of distributed energy units to provide power or auxiliary services to the power grid. Including it in the reward function can encourage the policy network to learn how to maximize the profitability of the virtual power plant. Market revenue can be calculated according to the amount of electricity sold by the virtual power plant to the grid (or the amount of auxiliary services provided) in a certain time period multiplied by the corresponding real-time electricity price (or auxiliary service price). For example, revenue = Σ (P_output * Price_market), where P_output is the net output of the virtual power plant, and Price_market is the real-time electricity price. More complex market mechanisms such as capacity market, day-ahead market, and intraday market can also be considered, and the transaction revenues in different markets can be weighted or accumulated.
[0039] The penalty term calculated according to the deviation between the joint action instruction and the grid dispatching demand is used to quantify the deviation between the actual execution of the joint action instruction by the virtual power plant and the load demand, frequency regulation demand, and other dispatching instructions proposed by the external grid (or dispatching center). The purpose is to force the policy network to learn how to strictly follow the dispatching requirements of the grid, ensuring that the virtual power plant as a reliable component of the grid avoids negative impacts on grid stability or penalties due to not meeting dispatching requirements. The penalty term can be designed as a function of the square or absolute value of the deviation. The larger the deviation, the greater the penalty, thereby reducing the composite reward value. Different penalty mechanisms can also be set according to the nature of the deviation, for example, asymmetric penalties can be applied for insufficient output and excessive output, or independent penalty terms can be set for different types of auxiliary service requirements such as frequency regulation and voltage support, and they are accumulated.
[0040] The equipment aging loss cost output by the health assessment unit 203, related to the joint action command, measures the economic loss caused by frequent or drastic control actions to shorten the lifespan or degrade the performance of each physical distributed energy unit (such as energy storage batteries, inverters, etc.) within the virtual power plant. Incorporating this cost into the reward function encourages the strategy network to consider the long-term healthy operation of equipment while pursuing economic benefits and meeting grid demands, avoiding short-term optimization that is "killing the goose that lays the golden eggs," thereby extending equipment lifespan and reducing long-term operation and maintenance costs. The health assessment unit 203 can predict the impact of each control action on the remaining lifespan of the equipment based on the operating data of the physical distributed energy units (such as charge / discharge depth, cycle count, temperature, current surge, etc.) and a preset lifespan model, and convert this into economic costs. For example, for batteries, the equivalent cycle count can be calculated based on their charge / discharge cycle count and depth, and then combined with the battery's lifespan cost to estimate the lifespan loss cost caused by a single control action. More sophisticated health assessment models, such as data-driven machine learning models, can be used to predict the rate of health degradation of equipment under specific control actions by analyzing historical operating data and fault data, and then link it to the replacement or maintenance costs of the equipment to quantify the cost of aging and wear.
[0041] By setting the training objective of the deep reinforcement learning policy network in the intelligent decision optimization module 30 to maximize a composite reward value, and ensuring that the calculation of this composite reward value is coupled with at least the market revenue, the penalty term for grid dispatch demand deviation, and the equipment aging and loss cost, this application effectively solves the problems in existing technologies where dispatch strategies cannot achieve global optimization, equipment lifespan is easily compromised, and grid demand is difficult to fully meet. Specifically, the introduction of market revenue incentivizes the policy network to obtain maximum economic benefits in the electricity market; the penalty term for grid dispatch demand deviation forces the policy network to learn how to accurately respond to grid commands, ensuring the stability and reliability of the virtual power plant; and the coupling of the equipment aging and loss cost output by the health assessment unit 203 enables the policy network to fully consider the impact of control actions on the long-term health of physical distributed energy units when making decisions, avoiding equipment lifespan shortening caused by overuse or improper operation. This multi-objective collaborative optimization mechanism enables the deep reinforcement learning policy network to generate joint action commands that take into account economy, grid compliance, and equipment health, thereby achieving globally optimal dynamic control of the virtual power plant and significantly improving the long-term operating efficiency and sustainability of the system.
[0042] In some of the embodiments described above in this application, a data acquisition and sensing module 10 is proposed to acquire operational status data, environmental data, and external power grid and market data. However, in this process, since specific data items are not clearly defined, data acquisition may be incomplete or inaccurate, thereby affecting the accuracy of digital twin modeling and the reliability of intelligent decision optimization, and further exacerbating prediction errors and scheduling coordination difficulties.
[0043] In this regard, this application further clarifies the data content acquired by the data acquisition and sensing module 10. Specifically, the operating status data includes power, voltage, current, battery state of charge, and health status; the environmental data includes illumination and wind speed; and the external power grid and market data includes power grid frequency, load demand information, and real-time electricity price information.
[0044] The operational status data is key information reflecting the real-time operating status of each physical distributed energy unit within the virtual power plant. Power refers to the rate at which a distributed energy unit outputs or consumes electrical energy at a given moment, which can be directly measured by a power sensor installed at the output end of the distributed energy unit, or calculated by measuring voltage and current. Voltage refers to the potential difference at the distributed energy unit or its connection point, which can be monitored in real-time by a voltage sensor or voltage transformer. Current refers to the amount of charge flowing through the distributed energy unit, which can be measured in real-time by a current sensor or current transformer. The State of Charge (SOC) refers to the percentage of remaining charge in the energy storage battery, which can be estimated using the Coulomb counting method, open-circuit voltage method, Kalman filtering algorithm, or machine learning-based models. The State of Health (SOH) refers to the degree of performance degradation of a distributed energy unit, especially an energy storage device, relative to its initial state, which can be assessed by analyzing historical operating data, performance degradation models, impedance spectral analysis, or artificial intelligence-based predictive models.
[0045] The environmental data refers to the external natural conditions that affect the output characteristics of distributed energy units. The sunlight refers to the intensity of solar radiation received per unit area, which can be measured in real time using a sunlight sensor (such as a solar intensity meter) installed near the photovoltaic power station, or obtained from a meteorological station. The wind speed refers to the speed of air movement, which can be measured in real time using an anemometer (such as a cup anemometer or ultrasonic anemometer), or obtained from a meteorological service agency.
[0046] The external power grid and market data are crucial for virtual power plants to participate in grid operation and electricity market transactions. The grid frequency refers to the frequency of alternating current in the power system, which can be directly obtained through grid frequency monitoring equipment or provided by the grid dispatch center. The load demand information refers to the total electricity or power required by the grid within a specific time period, which can be obtained from the grid dispatching agency or estimated through historical data analysis and predictive models. The real-time electricity price information refers to the immediate trading price of electricity in the electricity market, which can be obtained from electricity trading platforms or market operators.
[0047] Through the above technical solutions, this application enhances the targeting and completeness of data collection, ensuring the accuracy of subsequent digital twin modeling and intelligent decision-making from the source. Specifically, the basic electrical parameters such as power, voltage, and current in the operating status data provide an accurate physical state foundation for the digital twin model, avoiding model deviations caused by missing key parameters, thereby improving the accuracy of the digital twin model's mapping of the operating status of physical distributed energy units. Battery state of charge and health status, for the energy storage system, provide battery charging and discharging status and aging information, enabling the intelligent decision optimization module 30 to fully consider equipment lifespan degradation and operating efficiency when generating joint action commands, avoiding collaborative failures caused by neglecting equipment characteristics, thereby extending equipment lifespan and reducing maintenance costs. Light and wind speed in the environmental data are directly related to the output fluctuations of the new energy unit, providing key environmental factors as input for the prediction model, significantly improving the prediction accuracy of intermittent new energy output, and effectively reducing prediction errors. The grid frequency, load demand information, and real-time electricity price information from external grid and market data enable the intelligent decision optimization module 30 to dynamically respond to grid dispatching needs and market price changes, ensuring that the generated control commands maximize the economic benefits of the virtual power plant while meeting grid demands. Therefore, by clarifying and refining these data items, this application effectively solves the problems caused by incomplete or inaccurate data collection, providing a solid data foundation for the accuracy of digital twin modeling, the reliability of intelligent decision optimization, and the prediction accuracy and dispatch coordination capabilities of the entire system.
[0048] Please see Figure 3 , Figure 3 A specific implementation of a dynamic control method for a virtual power plant based on digital twins is illustrated. This method is applied to the aforementioned dynamic control system for a virtual power plant based on digital twins.
[0049] It should be noted that if substantially the same result is obtained, the method of this invention is not based on... Figure 3 Limited to the order of the processes shown, this method includes the following steps: S1: Real-time acquisition of operational status data, environmental data, and external power grid and market data for each physical distributed energy unit within the virtual power plant. S2: Construction and real-time updating of digital twin models corresponding to each physical distributed energy unit within the virtual power plant based on the operational status data, environmental data, and external power grid and market data. S3: Generation of joint action commands for regulating each physical distributed energy unit using a pre-trained deep reinforcement learning policy network, based on the current state of all digital twin models, the environmental data, and the external power grid and market data. S4: Issuance of the joint action commands to the corresponding physical distributed energy units for execution, and collection of actual execution results to form a closed-loop feedback loop.
[0050] Specifically, firstly, real-time acquisition of operational status data, environmental data, and external power grid and market data from each physical distributed energy unit within the virtual power plant is achieved. Specifically, operational status data can include power, voltage, current, battery state of charge, and health status. This data can be collected in real-time by sensors, smart meters, or SCADA systems deployed on each physical distributed energy unit and transmitted to the data processing center via a communication network. Environmental data can include light intensity, wind speed, and temperature, which can be obtained through weather stations, environmental monitoring equipment, or third-party data interfaces. External power grid and market data can include grid frequency, load demand information, and real-time electricity price information, which can be obtained through data interfaces with the power grid dispatch center and the power trading platform. By acquiring this multi-source heterogeneous data in real-time and comprehensively, an accurate and up-to-date information foundation is provided for subsequent digital twin modeling and intelligent decision-making.
[0051] Secondly, based on the operational status data, environmental data, and external power grid and market data, digital twin models corresponding to each physical distributed energy unit within the virtual power plant are constructed and updated in real time. Specifically, a digital twin model integrating mechanism and data-driven approaches is constructed for each physical distributed energy unit. For example, for energy storage systems, their charge-discharge characteristics and lifetime degradation can be predicted by combining their electrochemical mechanism model and a machine learning model trained based on historical operational data; for photovoltaic units, their output can be estimated by combining their photoelectric conversion mechanism model and a prediction model trained based on environmental data. Based on this, a digital twin simulation environment is constructed and maintained, which includes multiple simulable models corresponding to the physical distributed energy units. The state of the digital twin model is synchronized with the physical distributed energy unit in real time to ensure that the digital twin model accurately reflects the real-time state of the physical entity. Based on the digital twin model and input control commands, the performance response information and lifetime loss cost of the physical distributed energy unit are simulated and predicted, providing a comprehensive evaluation basis for subsequent decision-making.
[0052] Furthermore, based on the current states of all the digital twin models, the environmental data, and the external power grid and market data, a pre-trained deep reinforcement learning policy network generates joint action commands for regulating each of the physical distributed energy units. Specifically, the deep reinforcement learning policy network can be implemented using algorithms such as Deep Q-Network (DQN), Proximal Policy Optimization (PPO), or Asynchronous Advantageous Actor-Critic (A3C). This policy network uses the digital twin simulation environment as its interaction object, receiving the current state (e.g., predicted output, health status), environmental data (e.g., illumination, wind speed), and market data (e.g., real-time electricity price) from the digital twin models as input. Through its internally learned policies, it outputs joint action commands for all physical distributed energy units within the virtual power plant, such as power setpoints, charging / discharging commands, or start / stop states for each unit. By engaging in extensive interaction and learning within the digital twin simulation environment, the policy network aims to maximize a composite reward value as its training objective, thereby gaining the ability to make optimal decisions in complex dynamic environments.
[0053] Finally, the joint action command is sent to the corresponding physical distributed energy unit for execution, and the actual execution results are collected to form a closed-loop feedback. Specifically, the joint action command is sent to the local controller or energy management system of each physical distributed energy unit through a communication interface and protocol (such as Modbus, IEC 61850, or MQTT). Each physical distributed energy unit performs actual operations according to the received command, such as adjusting power output or charging and discharging. During the execution of the command, the system continuously collects the actual operating status data of each physical distributed energy unit, such as actual output, voltage, and current, and feeds these actual execution results back to the data acquisition and sensing module, thereby forming a complete closed-loop control loop.
[0054] Through the aforementioned technical solution, this application provides a solid data foundation for the dynamic control of virtual power plants by acquiring real-time and comprehensive data, effectively solving the prediction error problem caused by data lag in traditional methods. Based on this data, a digital twin model is constructed and updated in real time, enabling dynamic simulation of the complex behavior of each physical distributed energy unit. This significantly improves the prediction accuracy of the physical entity's state and performance, providing a reliable simulation environment for intelligent decision-making. Furthermore, utilizing a pre-trained deep reinforcement learning policy network, the system can adaptively generate optimal joint action commands based on the real-time state and external information of the digital twin environment. This solves the problem of traditional scheduling strategies being static and unable to respond to grid and market changes in real time, achieving better economic benefits and operational stability. Finally, by issuing and executing commands and collecting actual results to form a closed-loop feedback, the system can continuously learn and optimize. This not only verifies the control effect but also provides a basis for the calibration of the digital twin model and the online fine-tuning of the reinforcement learning strategy. This achieves refined collaborative optimization of heterogeneous distributed energy resources, taking into account the economic interests and equipment health of each unit, effectively solving the problem of difficult multi-objective collaborative optimization.
[0055] Please see Figure 4 , Figure 4 The specific implementation method of step S2 is shown, including: S21: Based on the operating status data, the environmental data, and the external power grid and market data, construct a fusion mechanism and data-driven digital twin model for each physical distributed energy unit; S22: Build and maintain a digital twin simulation environment based on the digital twin model, wherein the digital twin simulation environment includes multiple simulable models corresponding to the physical distributed energy units; S23: Synchronize the state of the digital twin model with that of the physical distributed energy unit in real time; S24: Based on the digital twin model and the input control commands, simulate and predict the performance response information and lifetime loss cost of the physical distributed energy unit.
[0056] Specifically, a digital twin model integrating mechanistic and data-driven approaches is constructed for each physical distributed energy unit. This aims to combine the structured knowledge and generalization capabilities provided by mechanistic models such as physical laws and chemical reactions with the ability of data-driven models, such as machine learning and deep learning, to learn complex nonlinear relationships and adapt to data. This ensures that the model accurately reflects the inherent laws of the physical system while adapting to complex and variable phenomena in actual operation that are difficult to fully describe by mechanistic methods, thereby improving prediction accuracy and robustness. For example, a hybrid modeling approach can be used, where physical mechanism equations serve as constraints or regularization terms for neural networks, or the output of the mechanism model serves as the input features of the data-driven model. For battery energy storage systems, electrochemical mechanism models can be used to describe their charging and discharging processes, while neural networks trained with cycle life data can be used to predict capacity decay. Another approach is a hierarchical modeling method, employing different types of models at different levels of abstraction. For example, data-driven models can be used at the macroscopic level to capture overall behavior, while mechanistic models can be used at the microscopic level to describe the physical processes of key components.
[0057] Based on this, a digital twin simulation environment is constructed and maintained using the aforementioned digital twin model. This simulation environment includes multiple simulable models corresponding to physical distributed energy units. This simulation environment is a virtual, high-fidelity simulation platform that integrates the digital twin models of all physical distributed energy units and can simulate their interactions and interactions with the external environment. This provides a safe and controllable sandbox environment for testing and verifying control strategies, allowing for the evaluation of system performance under different scenarios without affecting the actual physical system. For example, professional simulation software such as MATLAB / Simulink and OpenDSS can be used to encapsulate each digital twin model as an independent simulation module, and data exchange and collaborative simulation can be performed through a unified interface. Alternatively, a distributed microservice architecture can be adopted, deploying each digital twin model as an independent microservice, communicating through message queues or API gateways to build a scalable, high-concurrency simulation environment suitable for large-scale, heterogeneous virtual power plants.
[0058] Simultaneously, the state of the digital twin model is synchronized with that of the physical distributed energy unit in real time, ensuring that the state of the digital twin model remains highly consistent with the actual operating state of the corresponding physical distributed energy unit at any given moment. This eliminates the "digital divide" between the model and the physical entity, guaranteeing that the digital twin model accurately reflects the real-time status of the physical system and providing a reliable foundation for subsequent prediction, evaluation, and decision-making. For example, high-speed data acquisition systems (such as SCADA and IoT sensors) can transmit the operating status data (such as power, voltage, current, and temperature) of the physical distributed energy unit to the digital twin platform in real time, triggering the state update logic of the digital twin model. Alternatively, an event-driven synchronization mechanism can be used, immediately triggering the corresponding state update of the digital twin model when a critical state change occurs in the physical distributed energy unit or when new control commands are received, ensuring synchronous response to critical events.
[0059] Furthermore, based on the digital twin model and the input control commands, the performance response information and lifetime attrition cost of the physical distributed energy unit are simulated and predicted. This utilizes the synchronized digital twin model to predict the performance (e.g., output, energy consumption, efficiency) and lifetime attrition of the physical distributed energy unit over a future period after receiving control commands. This provides forward-looking information to the intelligent decision-making optimization module, enabling it to comprehensively consider short-term economic benefits and long-term equipment health when generating control commands, achieving multi-objective optimization. For example, the digital twin model can run its internal prediction algorithm to simulate the dynamic response of the physical unit under the action of commands over a future period and output performance parameter curves. Simultaneously, a dedicated lifetime prediction sub-model (e.g., based on fatigue accumulation, capacity decay curves, etc.) is integrated into the digital twin model to quantitatively assess the impact of control commands on equipment lifetime and output predictions of lifetime attrition rate or lifetime attrition cost.
[0060] Through the above technical solutions, this application addresses the accuracy and real-time issues in model construction and updating by refining the specific steps of building and updating the digital twin model. The model construction approach, which integrates mechanism and data-driven methods, combines the deterministic advantages of physical mechanisms with the adaptability of data-driven approaches, ensuring that the model accurately captures the complex dynamic characteristics of distributed energy units and avoids prediction biases caused by traditional simplified models, thereby improving prediction accuracy. Building and maintaining a digital twin simulation environment provides a unified and scalable interactive platform for the integrated simulation of the entire virtual power plant, effectively solving the infrastructure problem of heterogeneous resource collaborative optimization. Real-time synchronization of the digital twin model's state with the physical distributed energy units ensures consistency between the model's state and the actual units through continuous data feedback, eliminating accumulated errors caused by model lag and maintaining real-time response capability of regulation. Finally, based on the digital twin model's simulation prediction of performance response information and lifetime loss costs, key performance changes and lifetime loss costs are provided for optimization decisions, enabling regulation commands to simultaneously consider short-term economics and long-term equipment health, solving the challenge of multi-objective collaborative optimization. These steps are interconnected and together enhance the practicality of the digital twin model and the overall efficiency of dynamic control. They provide a high-fidelity, real-time synchronous interactive environment for subsequent deep reinforcement learning policy networks, enabling the generated joint action instructions to be more accurate and efficient, and truly achieving dynamic optimization and control of the virtual power plant.
[0061] Please see Figure 5 , Figure 5 One specific implementation method prior to step S3 is shown, including: S3A: Use the digital twin simulation environment as the interaction environment for the deep reinforcement learning policy network. S3B: Output action instructions based on the state of the deep reinforcement learning policy network. S3C: Simulate and execute the action instructions based on the digital twin simulation environment, generating the next state and a composite reward value. S3D: Update the parameters of the deep reinforcement learning policy network using the composite reward value to obtain the pre-trained deep reinforcement learning policy network.
[0062] Specifically, the digital twin simulation environment is used as the interaction environment for the deep reinforcement learning policy network, and then action instructions are output based on the state of the deep reinforcement learning policy network. The action instructions are executed in the digital twin simulation environment to generate the next state and a composite reward value. The parameters of the deep reinforcement learning policy network are updated using the composite reward value to obtain the pre-trained deep reinforcement learning policy network.
[0063] The training objective of the deep reinforcement learning policy network is to maximize a composite reward value, the calculation of which is coupled with at least the following factors: market revenue calculated based on electricity market information and the joint action command; penalty term calculated based on the deviation between the joint action command and grid dispatch demand; and equipment aging and depreciation costs related to the joint action command output by the health assessment unit.
[0064] Please see Figure 6 , Figure 6 One specific implementation method following step S4 is shown, including: S41: Compare the actual execution results with the prediction results of the digital twin simulation environment, and calibrate the parameters of the corresponding model in the digital twin simulation environment. S42: Incorporate the interaction experience information containing the actual execution results into the training dataset of the deep reinforcement learning policy network for online fine-tuning or periodic retraining of the deep reinforcement learning policy network.
[0065] Specifically, the comparison between the actual execution results and the prediction results of the digital twin simulation environment aims to identify the deviation between the predictions of the digital twin simulation environment and the actual operation of the physical entity, providing a basis for subsequent model calibration. This comparison can be achieved by calculating error indicators between the actual execution results and the prediction results, such as mean squared error (MSE), mean absolute error (MAE), or percentage error, to quantify the comparison results. Alternatively, statistical methods, such as hypothesis testing or confidence interval analysis, can be used to determine whether the actual results significantly deviate from the prediction results, thereby triggering the calibration mechanism.
[0066] Based on this, the parameters of the corresponding model in the digital twin simulation environment are calibrated to correct the parameters within the digital twin simulation model, making it more accurately reflect the real behavior and characteristics of the physical distributed energy unit, thereby improving the prediction accuracy of the simulation environment. Parameter calibration can employ optimization algorithm-based methods, such as least squares, gradient descent, or Kalman filtering, to adjust model parameters based on comparison results to minimize the error between reality and prediction. Alternatively, machine learning methods, such as adaptive learning algorithms or online learning models, can be used to dynamically update model parameters based on new real-world data to adapt to changes in the physical system over time.
[0067] Simultaneously, interactive experience information containing actual execution results is incorporated into the training dataset of the deep reinforcement learning policy network. This aims to enrich the training samples of the deep reinforcement learning policy network using real-world physical system operation data, enabling it to learn from real-world experience and improve its generalization ability and adaptability to real-world environments. This can be achieved by adding actual execution results (e.g., the reward obtained after performing an action in a specific state and the next state) as new experience tuples (state, action, reward, next_state) to the policy network's experience replay buffer. Alternatively, this real-world interaction data can be mixed with simulation data to form a more comprehensive training dataset for the policy network's learning.
[0068] Furthermore, online fine-tuning or periodic retraining of the deep reinforcement learning policy network is possible to ensure its continuous adaptation to changes in the virtual power plant operating environment and to continuously optimize its decision-making capabilities based on actual operating experience, thereby generating more effective and accurate control commands. Online fine-tuning can be achieved through incremental learning, where, after the policy network is deployed, newly collected actual interaction experience data is used to continuously make small adjustments to the network parameters with a small learning rate. Periodic retraining involves comprehensively retraining the policy network at preset time intervals (e.g., daily, weekly, or monthly) or when performance indicators drop to a certain threshold, using accumulated actual interaction experience data to ensure its performance remains at an optimal level.
[0069] Through the above technical solution, this application introduces a closed-loop feedback mechanism, solving the problem of inaccurate control commands caused by prediction deviations in the digital twin simulation environment, thereby improving the adaptability and economy of dynamic control of the virtual power plant. Specifically, by comparing the actual execution results with the prediction results of the digital twin simulation environment, prediction errors can be detected in real time, providing a direct basis for parameter calibration; calibrating the parameters of the corresponding model in the digital twin simulation environment corrects model deviations, ensuring that the simulation environment is closer to physical reality and reducing subsequent prediction errors; incorporating interactive experience information containing actual execution results into the training dataset of the deep reinforcement learning policy network enriches the learning samples with real data, avoiding the policy network's over-reliance on simulation data and insufficient generalization; and performing online fine-tuning or periodic retraining of the deep reinforcement learning policy network dynamically optimizes the policy network parameters, enabling it to continuously adapt to external changes and generate more accurate control commands, thereby effectively improving the economic dispatch and equipment health management level of the virtual power plant.
[0070] Please see Figure 7 , Figure 7 Another specific implementation method following step S4 is shown, including: S4A: When a new type of physical distributed energy unit is added to the virtual power plant, a corresponding digital twin model is constructed for the new physical distributed energy unit type and connected to the digital twin simulation environment to obtain an updated digital twin simulation environment. S4B: The deep reinforcement learning policy network is incrementally learned within the updated digital twin simulation environment.
[0071] Specifically, when a new type of physical distributed energy unit is connected to the virtual power plant, this technical feature describes a dynamic change scenario faced by the system, namely, the physical composition of the virtual power plant changes, and a new, previously unincluded type of distributed energy unit is introduced. This "new type" means that its operating characteristics, response modes, aging and degradation patterns may differ significantly from existing units, requiring the system to have the ability to identify and adapt to these new characteristics. The system can identify the connection of a new type of physical distributed energy unit by monitoring changes in the topology of the virtual power plant or by receiving configuration update notifications from the management platform. For example, when a new energy storage system, photovoltaic array, or electric vehicle charging station is physically connected and registered to the virtual power plant management platform, the system will trigger the corresponding processing flow. Alternatively, the system can analyze the metadata of the connected devices, such as device model, manufacturer, and rated parameters, and compare it with the existing known device type library. If a mismatch or unregistered type is found, it is determined to be a new type of physical distributed energy unit.
[0072] Subsequently, a corresponding digital twin model is constructed for the new physical distributed energy unit type and integrated into the digital twin simulation environment to obtain an updated digital twin simulation environment. This step aims to create an accurate mapping in digital space for the newly integrated physical distributed energy unit type with unique operating characteristics and integrate it into the existing digital twin simulation environment. Constructing a digital twin model is fundamental to understanding and predicting the behavior of the new unit, while integrating it into the simulation environment ensures that the digital representation of the entire virtual power plant fully reflects the latest state of the physical entity. A hybrid modeling approach can be used to construct the digital twin model, combining mechanistic models (based on physical laws and engineering principles, such as the electrochemical model of a battery and the PV curve model of a photovoltaic system) and data-driven models (trained using historical operating data through machine learning algorithms, such as neural networks and support vector machines, to capture complex nonlinear relationships). For new unit types, a general mechanistic model can be pre-established, and parameter identification and data-driven model training can be performed using a small amount of initial operating data. Integrating into the digital twin simulation environment typically involves registering the newly constructed digital twin model into the simulation platform's model library and updating the simulation environment's topology and data interface configuration. This may include defining the inputs (such as control commands and environmental data) and outputs (such as performance parameters and health status) of the new model, as well as the interaction logic between it and other existing digital twin models. The updated digital twin simulation environment will be able to simulate the operation of the entire virtual power plant, including the new units.
[0073] Building upon this, incremental learning is performed on the deep reinforcement learning policy network within the updated digital twin simulation environment. This step aims to efficiently update and expand the existing deep reinforcement learning policy network's capabilities using the newly constructed digital twin simulation environment that integrates the new units, enabling it to adapt to and optimize the overall operation of the virtual power plant, including the new units. Incremental learning avoids training from scratch, significantly improving learning efficiency and system response speed. Incremental learning can be achieved through an "experience replay" mechanism, where new experience data containing interactions with the new units is generated in the updated digital twin simulation environment and mixed with old experience data for training the policy network. Simultaneously, a "knowledge distillation" technique can be employed, allowing the newly trained policy network to learn from the old policy network, maintaining its optimization capabilities for the original units. Another incremental learning method employs "transfer learning" or "meta-learning" techniques. For example, the low-level parameters related to general feature extraction in the policy network can be frozen, with only the upper-level parameters related to decision output or specific behaviors of the new units being fine-tuned. Alternatively, a meta-learning algorithm can be used to enable the policy network to quickly adapt to the characteristics of the new units, achieving effective policy updates with only a small amount of interaction data.
[0074] Through the aforementioned technical solution, when a new type of physical distributed energy unit is introduced into the virtual power plant, the system can quickly identify and construct an accurate digital twin model for that new type of unit, seamlessly integrating it into the existing digital twin simulation environment. This results in an updated digital twin simulation environment that comprehensively reflects the physical structure and operational characteristics of the current virtual power plant. Based on this, by incrementally learning the deep reinforcement learning policy network, the decision network can efficiently absorb the operational characteristics and optimization objectives of the new unit without requiring time-consuming and resource-intensive re-training. This enables the system to quickly adapt to the dynamic changes of the virtual power plant, ensuring that optimal joint action commands are continuously generated even after the introduction of new types of distributed energy units. This maintains or improves the overall control performance, economic efficiency, and equipment health of the virtual power plant, significantly enhancing the system's flexibility, scalability, and real-time response capabilities.
[0075] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 8 , Figure 8 This is a basic structural block diagram of the computer device in this embodiment.
[0076] Computer device 5 includes a memory 51, a processor 52, and a network interface 53 that are interconnected via a system bus. It should be noted that... Figure 8 Only a computer device 5 with three components—memory 51, processor 52, and network interface 53—is shown. It should be understood that implementing all shown components is not required; more or fewer components may be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0077] Computer devices can include desktop computers, laptops, handheld computers, and cloud servers. These devices allow for human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0078] The memory 51 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 51 may be an internal storage unit of the computer device 5, such as the hard disk or memory of the computer device 5. In other embodiments, the memory 51 may also be an external storage device of the computer device 5, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 5. Of course, the memory 51 may include both the internal storage unit and the external storage device of the computer device 5. In this embodiment, the memory 51 is typically used to store the operating system and various application software installed on the computer device 5, such as the program code of a virtual power plant dynamic control method based on digital twins. In addition, the memory 51 can also be used to temporarily store various types of data that have been output or will be output.
[0079] In some embodiments, processor 52 may be a central processing unit (CPU), controller, microcontroller, microprocessor, or other data processing chip. This processor 52 is typically used to control the overall operation of computer device 5. In this embodiment, processor 52 is used to run program code stored in memory 51 or process data, for example, to run the program code of the aforementioned dynamic control method for a virtual power plant based on digital twins, to implement various embodiments of the dynamic control method for a virtual power plant based on digital twins.
[0080] The network interface 53 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 5 and other electronic devices.
[0081] This application also provides another embodiment, namely, a computer-readable storage medium storing a computer program that can be executed by at least one processor to cause the at least one processor to perform the steps of the above-described method for dynamic control of a virtual power plant based on a digital twin.
[0082] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of this application.
[0083] Obviously, the embodiments described above are merely some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the scope of this application. This application can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of protection of this application.
Claims
1. A dynamic control system for a virtual power plant based on digital twins, characterized in that, include: The data acquisition and sensing module is used to acquire in real time the operating status data, environmental data, and external power grid and market data of each physical distributed energy unit in the virtual power plant; The digital twin modeling module is used to construct and update in real time the digital twin models corresponding to each physical distributed energy unit in the virtual power plant; The intelligent decision optimization module is used to take the environment composed of all the digital twin models in the digital twin modeling layer as the interaction object, and generate joint action instructions for regulating each of the physical distributed energy units through a pre-trained deep reinforcement learning policy network based on the system state output by the digital twin environment module. The instruction execution and feedback module is used to issue the joint action instruction to the corresponding physical distributed energy unit for execution, and collect the actual execution results and feed them back to the data acquisition and sensing module.
2. The virtual power plant dynamic control system based on digital twins according to claim 1, characterized in that, The digital twin environment module includes: A state mapping unit is used to receive operating status data and external environment data from the physical distributed energy unit, and update the real-time state of the digital twin model according to the operating status data and the external environment data. The performance simulation unit is used to predict the performance parameter changes of the physical distributed energy unit within a future set time period based on the current state of the digital twin model and the received control commands, and to obtain the prediction results. A health assessment unit is used to quantify the health degradation of the physical distributed energy unit and the life loss cost caused by the control actions, based on the historical operating data of the physical distributed energy unit and the prediction results.
3. The virtual power plant dynamic control system based on digital twins according to claim 2, characterized in that, The training objective of the deep reinforcement learning policy network in the intelligent decision optimization module is to maximize a composite reward value, the calculation of which is coupled with at least the following factors: Market revenue calculated based on electricity market information and the aforementioned joint action instructions; Penalty calculated based on the deviation between the joint action command and the power grid dispatching requirements; And the equipment aging and depreciation costs associated with the joint action commands, output by the health assessment unit.
4. The virtual power plant dynamic control system based on digital twins according to claim 1, characterized in that, The operational status data includes power, voltage, current, battery state of charge, and health status; the environmental data includes illumination and wind speed; and the external power grid and market data includes power grid frequency, load demand information, and real-time electricity price information.
5. A method for dynamic control of a virtual power plant based on digital twins, characterized in that, The method, applied to the virtual power plant dynamic control system based on digital twins as described in any one of claims 1 to 4, comprises: Real-time acquisition of operational status data, environmental data, and external power grid and market data for each physical distributed energy unit within the virtual power plant; Based on the operational status data, the environmental data, and the external power grid and market data, a digital twin model corresponding to each physical distributed energy unit in the virtual power plant is constructed and updated in real time. Based on the current state of all the digital twin models, the environmental data, and the external power grid and market data, a pre-trained deep reinforcement learning policy network is used to generate joint action commands for regulating each of the physical distributed energy units. The joint action command is sent to the corresponding physical distributed energy unit for execution, and the actual execution results are collected to form a closed-loop feedback.
6. The method for dynamic control of a virtual power plant based on digital twins according to claim 5, characterized in that, The process of constructing and updating in real-time digital twin models corresponding to each physical distributed energy unit within the virtual power plant based on the operational status data, environmental data, and external power grid and market data includes: Based on the operational status data, the environmental data, and the external power grid and market data, a digital twin model integrating the mechanism and data-driven approach is constructed for each physical distributed energy unit. A digital twin simulation environment is constructed and maintained based on the aforementioned digital twin model, wherein the digital twin simulation environment includes multiple simulable models corresponding to physical distributed energy units; The state of the digital twin model is synchronized with that of the physical distributed energy unit in real time. Based on the digital twin model and the input control commands, the performance response information and lifetime loss cost of the physical distributed energy unit are simulated and predicted.
7. The method for dynamic control of a virtual power plant based on digital twins according to claim 6, characterized in that, Before generating joint action commands for regulating each of the physical distributed energy units using a pre-trained deep reinforcement learning policy network based on the current state of all the digital twin models, the environmental data, and the external power grid and market data, the method further includes: The digital twin simulation environment is used as the interaction environment for the deep reinforcement learning policy network; Based on the state output action instructions of the deep reinforcement learning policy network; The action instructions are executed based on a digital twin simulation environment, generating the next state and a composite reward value. The parameters of the deep reinforcement learning policy network are updated using the composite reward value to obtain the pre-trained deep reinforcement learning policy network.
8. The method for dynamic control of a virtual power plant based on digital twins according to claim 7, characterized in that, After issuing the joint action command to the corresponding physical distributed energy unit for execution and collecting the actual execution results to form a closed-loop feedback, the method further includes: The actual execution results are compared with the prediction results of the digital twin simulation environment, and the parameters of the corresponding model in the digital twin simulation environment are calibrated. Interaction experience information containing the actual execution results is incorporated into the training dataset of the deep reinforcement learning policy network to perform online fine-tuning or periodic retraining of the deep reinforcement learning policy network.
9. The method for dynamic control of a virtual power plant based on digital twins according to claim 6, characterized in that, After issuing the joint action command to the corresponding physical distributed energy unit for execution and collecting the actual execution results to form a closed-loop feedback, the method further includes: When a new type of physical distributed energy unit is connected to the virtual power plant, a corresponding digital twin model is constructed for the new type of physical distributed energy unit and connected to the digital twin simulation environment to obtain an updated digital twin simulation environment. The deep reinforcement learning policy network is incrementally learned in the updated digital twin simulation environment.
10. A computer device, characterized in that, It includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method for dynamic control of a virtual power plant based on digital twins as described in any one of claims 5 to 9.