Virtual power plant intelligent scheduling method and system based on digital twinning
Patent Information
- Application Number
- CN202511122053.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
Smart Images

Figure CN120638336A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital power grid technology, and specifically to a virtual power plant intelligent scheduling method based on digital twins, a virtual power plant intelligent scheduling system based on digital twins, a computer-readable storage medium, an electronic device and a computer program product. Background Art
[0002] With the rapid development of renewable energy consumption, the electricity spot market, and the coordinated development of power generation, grid-load, and storage, virtual power plants (VPPs), as an important vehicle for aggregating distributed power resources, are becoming a key component in supporting flexible grid dispatch and stable operation. In this context, achieving precise dispatch of diverse distributed power equipment within VPPs has become a crucial technical foundation for building a new power system.
[0003] Existing virtual power plant dispatch methods primarily rely on policy models or heuristic optimization algorithms trained on historical operating data. These methods simulate objectives such as power balance, cost optimization, or supply-demand matching to output corresponding dispatch control actions. However, these methods often overlook the fundamental characteristics of power systems as physically coupled systems and lack dynamic feedback on grid topology, electrical distribution, and inter-node coupling. As a result, dispatch actions can cause local voltage fluctuations, power flow instability, or inter-node interference amplification at the electrical level, making it difficult to effectively meet the physical safety requirements of grid operation.
[0004] In addition, although some scheduling methods attempt to introduce safety constraint rules, such as node voltage limits and line capacity limits, they are often attached to the scheduling model in the form of static boundary conditions, and cannot achieve dynamic perception and linkage adjustment during the scheduling strategy generation process. As a result, the generated strategy responds slowly to changes in constraint boundaries, making it difficult to maintain the physical feasibility of the action and system stability during high-frequency scheduling.
[0005] In summary, the current virtual power plant dispatching scheme still has obvious deficiencies in its ability to couple dispatching strategies with grid operation constraints. It is urgent to establish a dispatching modeling path that can integrate physical response characteristics and has real-time constraint perception capabilities to improve the adaptability and robustness of dispatching strategies to the actual operating status of the grid. Summary of the Invention
[0006] The purpose of the embodiments of the present invention is to provide a method and system for intelligent scheduling of virtual power plants based on digital twins, so as to at least solve the problems in existing solutions that the scheduling strategies lack the ability to dynamically perceive the operating constraints of the power grid, and the scheduling process fails to fully integrate the physical coupling response laws.
[0007] In order to achieve the above-mentioned objectives, the first aspect of the present invention provides a virtual power plant intelligent scheduling method based on digital twins, the method comprising: based on a digital twin agent matrix constructed by the grid nodes corresponding to each distributed power equipment participating in the scheduling, determining the electrical response change of each grid node caused by each historical control action to form a corresponding action response training data set; constructing a grid operation constraint response model based on the action response training data set, and generating constraint feedback information reflecting the feasibility of the current target control action based on the grid operation constraint response model; using the constraint feedback information as coupling information of a pre-trained scheduling strategy model, performing scheduling strategy training based on the current target control action, and obtaining a scheduling strategy model with constraint perception capability; generating corresponding scheduling control actions based on the scheduling strategy model.
[0008] Optionally, the construction rules of the digital twin agent matrix are as follows: for each distributed power equipment participating in the scheduling, a digital twin agent including a physical modeling layer, a data-driven layer and a state deduction layer is established respectively; the digital twin agents corresponding to the distributed power equipment are indexed and organized according to the grid nodes to which the corresponding distributed power equipment is connected to form a digital twin agent matrix; wherein, the physical modeling layer is used to represent the operating mechanism of the corresponding distributed power equipment; the data-driven layer is used to learn the behavioral characteristics of the distributed power equipment based on historical power and environmental data; the state deduction layer is used to output the change in node electrical response under the target control action. Optionally, based on the digital twin agent matrix constructed by the grid nodes corresponding to the distributed power equipment participating in the scheduling, the electrical response change of each grid node caused by each historical control action is determined, including: each digital twin agent in the digital twin agent matrix locally executes each electrical response prediction task; wherein the electrical response prediction task is a prediction task that includes historical control actions; each digital twin agent outputs the electrical response change of its corresponding grid node, and obtains the electrical response change of each grid node caused by the target control action based on the prediction results of multiple digital twin agents. Optionally, each digital twin agent in the digital twin agent matrix performs an electrical response prediction task locally, including: based on the electrical response prediction task, outputting device mechanism parameters through the physical modeling layer of the digital twin agent, and outputting historical behavior characteristics through the data-driven layer; inputting the device mechanism parameters and the historical behavior characteristics into the state deduction layer to perform the electrical response prediction task, and outputting the electrical response change of the corresponding power grid node under the corresponding historical control action; uploading the electrical response change to the digital twin agent matrix to perform electrical response change aggregation of each power grid node through the digital twin agent matrix. Optionally, a power grid operation constraint response model is constructed based on the node electrical response change, including: traversing all power grid nodes that have a linkage relationship with historical control actions, and respectively executing the binding of the electrical response change of each power grid node with the target control action to obtain multiple data pairs for characterizing the electrical disturbance response relationship; based on the multiple data pairs, combined with the preset power grid physical topology structure and the control coupling relationship between nodes, the actual electrical response change of each linked power grid node is used as the supervision target to construct a constraint response model that reflects the response law of the operating status of each power grid node to the historical control action. Optionally, the physical topology structure of the power grid is constructed based on the grid access relationship of each distributed power equipment participating in the scheduling, and is used to characterize the physical connection path and control influence path between each grid node when constructing the constraint response model; the control coupling relationship between nodes is constructed based on the state coordinated change law of each grid node under the influence of historical control actions, and is used to characterize the response propagation path and interference superposition characteristics of the target control action between nodes when constructing the constraint response model. Optionally, the current target control action is a target control action input by the user, or a target control action generated based on the current cycle / state; the rule for generating the target control action based on the current cycle is: obtain the typical load curve of the power grid corresponding to the current cycle, and match the preset cycle and control action mapping rules, and extract the reference control action matching the current cycle as the target control action; the rule for generating the target control action based on the current operating state is: collect the current power grid state parameters, match them in combination with the preset state and action mapping rules, and extract the initial target control action that meets the current operating state requirements as the target control action. Optionally, constraint feedback information reflecting the feasibility of the current target control action is generated based on the grid operation constraint response model, including: inputting the current target control action into the grid operation constraint response model, simulating the voltage change, feeder power fluctuation rate and / or inter-node power flow change rate of each grid node under the action of the control action as a node-level electrical disturbance index; based on the electrical disturbance index, comparing with the pre-constructed grid operation constraint judgment standard, performing feasibility level classification, and outputting constraint feedback information for identifying the feasibility of the target control action; wherein, the grid operation constraint judgment standard is constructed based on the rated voltage, safety margin threshold, topology stability boundary and load fluctuation tolerance of each grid node. Optionally, the method also includes executing scheduling strategy model training, and the training rules are: constructing a basic scheduling strategy training data set based on each historical control action and its corresponding target response state; using the basic scheduling strategy training data set as a training sample, and using the scheduling effect as a supervision target to execute strategy behavior learning to obtain a scheduling strategy model.
[0009] Optionally, the constraint feedback information is used as coupling information of a pre-trained scheduling strategy model, and scheduling strategy training is performed based on the current target control action to obtain a scheduling strategy with constraint perception capability, including: taking the current target control action and the corresponding constraint feedback information as input, embedding them into the reward function structure of the pre-trained scheduling strategy model, performing scheduling strategy model training, and outputting a scheduling strategy with constraint perception capability; wherein, during the training process, a negative reward is set for the target control action that violates the constraint feedback information, and a positive reward is set for the target control action that satisfies the constraint feedback information, and the range of the optional action space is compressed based on the accumulated reward and punishment results. Optionally, corresponding scheduling control actions are generated based on the scheduling strategy model, including: in the current cycle / state, calling the obtained scheduling strategy model with constraint perception capability; inputting the target control action corresponding to the current cycle / state, performing reasoning through the scheduling strategy model to generate the optimal scheduling control action in the current cycle, and distributing the scheduling control action to each corresponding distributed power equipment to execute control of each distributed power equipment.
[0010] The second aspect of the present invention provides a virtual power plant intelligent dispatching system based on digital twins, and the system includes: an initial unit, which is used to determine the electrical response change of each grid node caused by each historical control action based on the digital twin agent matrix constructed by the grid nodes corresponding to each distributed power equipment participating in the dispatch, so as to form a corresponding action response training data set; a processing unit, which is used to construct a grid operation constraint response model based on the action response training data set, and generate constraint feedback information reflecting the feasibility of the current target control action based on the grid operation constraint response model; a model training unit, which is used to use the constraint feedback information as coupling information of the pre-trained dispatching strategy model, perform dispatching strategy training based on the current target control action, and obtain a dispatching strategy model with constraint perception capability; a dispatching unit, which is used to generate corresponding dispatching control actions based on the dispatching strategy model. A third aspect of the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enables the computer to execute the above-mentioned intelligent scheduling method for a virtual power plant based on digital twins.
[0011] The fourth aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the above-mentioned digital twin-based virtual power plant intelligent scheduling method is implemented.
[0012] A fifth aspect of the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the above-mentioned digital twin-based virtual power plant intelligent scheduling method.
[0013] Through the above-mentioned technical solution, the present invention introduces a digital twin agent matrix to accurately capture the changes in the electrical response of grid nodes under each historical control action. This allows the construction of a grid operation constraint response model that reflects the node's operating patterns. Based on this model, constraint feedback information is generated to describe the feasibility of the current target control action. Furthermore, by embedding constraint feedback information into the scheduling strategy training process, the scheduling strategy dynamically perceives and responds to grid operation constraints, significantly improving the feasibility and safety of scheduling control actions and resolving the issue of insufficient response of scheduling strategies to grid physical constraints in existing technologies.
[0014] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present invention, but do not constitute a limitation of the embodiments of the present invention. In the accompanying drawings: Figure 1 This is a flowchart of the steps of a virtual power plant intelligent scheduling method based on digital twins provided by one embodiment of the present invention; Figure 2 This is a schematic diagram of digital twin agent matrix construction and linkage mapping provided by one embodiment of the present invention; Figure 3 This is a system structure diagram of a virtual power plant intelligent dispatching system based on digital twins provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following describes the specific embodiments of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention.
[0017] Figure 1 This is a flowchart of the steps of the intelligent scheduling method of virtual power plant based on digital twins provided by one embodiment of the present invention. Figure 1 As shown, an embodiment of the present invention provides a virtual power plant intelligent scheduling method based on digital twins, the method comprising: Step S10: Based on the digital twin agent matrix constructed by the grid nodes corresponding to the distributed power equipment participating in the scheduling, the electrical response change of each grid node caused by each historical control action is determined to form a corresponding action response training data set.
[0018] Specifically, the construction rules of the digital twin agent matrix are as follows: for each distributed power equipment participating in the scheduling, a digital twin agent including a physical modeling layer, a data-driven layer and a state deduction layer is established respectively; the digital twin agents corresponding to the distributed power equipment are indexed and organized according to the grid nodes to which the corresponding distributed power equipment is connected to form a digital twin agent matrix; wherein, the physical modeling layer is used to represent the operating mechanism of the corresponding distributed power equipment; the data-driven layer is used to learn the behavioral characteristics of the distributed power equipment based on historical power and environmental data; the state deduction layer is used to output the change in the node electrical response under the target control action.
[0019] Furthermore, based on the digital twin agent matrix constructed by the grid nodes corresponding to the distributed power equipment participating in the scheduling, the electrical response change of each grid node caused by each historical control action is determined, including: each digital twin agent in the digital twin agent matrix locally executes each electrical response prediction task; wherein the electrical response prediction task is a prediction task that includes historical control actions; each digital twin agent outputs the electrical response change of its corresponding grid node, and obtains the electrical response change of each grid node caused by the target control action based on the prediction results of multiple digital twin agents.
[0020] Furthermore, each digital twin agent in the digital twin agent matrix performs an electrical response prediction task locally, including: based on the electrical response prediction task, outputting device mechanism parameters through the physical modeling layer of the digital twin agent, and outputting historical behavior characteristics through the data-driven layer; inputting the device mechanism parameters and the historical behavior characteristics into the state deduction layer to perform the electrical response prediction task, and outputting the electrical response change of the corresponding power grid node under the corresponding historical control action; uploading the electrical response change to the digital twin agent matrix to perform electrical response change aggregation of each power grid node through the digital twin agent matrix.
[0021] In an embodiment of the present invention, in an intelligent dispatching scenario for a distributed power system, traditional dispatching models often rely on simplified physical models or empirical rules, which makes it difficult to accurately characterize the multi-scale dynamic response characteristics of different types of power equipment under the intervention of control actions. In particular, when faced with a complex power grid environment consisting of heterogeneous equipment such as photovoltaics, batteries, heat pumps, and wind turbines, the dispatching strategy's predictive ability and constraint adaptability for electrical responses are significantly insufficient. To this end, a training data construction mechanism based on a digital twin agent matrix is proposed. The core idea is to establish a digital twin agent for each distributed power device, and on this basis, extract the causal relationship between control actions and electrical responses, thereby realizing the controllable generation of action response training data sets, and providing a high-quality data foundation for the subsequent dispatching strategy model construction and constraint response modeling from the source.
[0022] Specifically, digital twin agents are constructed for each distributed power device involved in scheduling. Each digital twin agent contains three functional layers: a physical modeling layer, a data-driven layer, and a state deduction layer. The physical modeling layer is used to express the physical mechanism characteristics of the device and construct a basic physical model of the device's operation, including electrical topology relationships, component structure, energy conversion path, and parameter settings. For example, for a battery energy storage device, the physical modeling layer must clearly define its charge and discharge efficiency, capacity, SOC curve, maximum power limit, etc.; for a heat pump device, it must include parameters such as its thermodynamic cycle model, operating condition curve, and load response curve.
[0023] Furthermore, the data-driven layer, based on historical operating data, learns and supplements the behavioral characteristics exhibited by the device in actual operation, compensating for the physical modeling's weak adaptability to external disturbances. This layer primarily extracts features, performs time series modeling, and performs association learning on historical power output, operating status, voltage and current records, and environmental data (such as temperature, light intensity, and wind speed). This layer generates a set of feature vectors that reflect the dynamic behavior of the device, assisting in subsequent state calculations. These feature vectors can be generated by fitting lightweight neural network models such as RNN and GRU, or by constructing a high-dimensional behavioral representation space based on statistical indicators.
[0024] The state deduction layer, as the final output module, is responsible for performing electrical response prediction tasks based on given target control actions, combined with the mechanism parameters and historical behavior characteristics generated by the upper layer. The deduction task uses the grid node to which the device is connected as the output index, predicting the electrical response changes that the device may cause to its connected node under the action of a specific control action, such as voltage offset, current fluctuation, power change, phase adjustment, etc., and represents the output results in a vectorized structure. It should be emphasized that the deduction process should take into account the nonlinear characteristics and time delay effects of the device response. Time series prediction structures such as LSTM and variational autoencoders can be used to construct the state calculation network.
[0025] After constructing digital twin agents for all distributed power devices, they must be further organized into a unified data structure, the "digital twin agent matrix," indexed by the grid nodes to which they are connected. Each row of this matrix represents a grid node, and each column represents the digital twin agent unit associated with that node. This matrix structure supports response analysis requirements at both the node and device levels, enabling unified modeling of distributed power grid responses.
[0026] To acquire electrical response data, each digital twin agent must locally execute the electrical response prediction task. This task explicitly specifies a historical control action as an input variable. This means that the prediction does not use real-time control signals. Instead, it uses a known scheduling action at a certain point in the past as a prerequisite, and predicts the device's response changes under the influence of this historical action. "Local execution" here means that each twin agent runs independently in a self-consistent manner, without relying on the status of other nodes or devices, thus supporting parallel computing and edge deployment.
[0027] When each digital twin agent performs response prediction, the physical modeling layer first outputs the device mechanism parameters, which are static inputs. The data-driven layer then outputs the dynamic behavior characteristics within the time window. These two types of information are input into the state deduction layer, triggering the deduction calculation process, ultimately determining the change in electrical response at the device node under the specified historical control action. This response should include a complete set of indicators, such as voltage change value, current peak value, frequency disturbance rate, and power flow direction switching, to facilitate subsequent constraint verification and model training.
[0028] After calculating all device responses, the prediction results are uploaded to the digital twin agent matrix in a standard format for unified archiving and aggregation of electrical response data. During the aggregation process, the control action label, grid node ID, device ID, response time, and prediction indicator set for each set of response data are retained to construct a multidimensional action response training dataset. This dataset serves as crucial data support for subsequent training of grid operation constraint response models and dispatch strategy models, ensuring high controllability, traceability, and scenario adaptability.
[0029] Based on the solution of the present invention, the process can accurately establish a causal mapping relationship between control actions and electrical responses, truly reflect the feedback characteristics of the equipment on the operating status of the power grid under various control decisions, and significantly improve the authenticity and interpretability of model training; at the same time, through the matrix organizational structure, it effectively solves the problems of difficult management of device heterogeneity and weak response coupling in traditional distributed modeling, and realizes batch prediction and data integration of multi-device and multi-node response behaviors before scheduling, laying a solid foundation for building an intelligent scheduling system with constraint perception capabilities.
[0030] In one possible implementation, Figure 2 For a virtual power plant that includes heterogeneous power equipment such as photovoltaic panels, energy storage batteries and charging piles, a corresponding digital twin system is built, and the scheduling control strategy training and iteration of each distributed power equipment are realized through joint parameter learning, strategy optimization and knowledge transfer modules.
[0031] At the physical layer, photovoltaic panels, energy storage batteries, charging stations, and other devices are deployed across multiple sites. Sensors collect real-time operational data on each device's operation, which is then synchronously transmitted to its corresponding digital twin generation agent. At the digital twin agent layer, each agent represents a specific device and is responsible for modeling device mechanisms, extracting data-driven behaviors, and performing state deduction. The resulting group of agents forms the digital twin agent matrix for the entire virtual power plant.
[0032] During parameter learning, existing model parameters are calibrated using real-world device operating data from the physical layer. For example, the photovoltaic device agent adjusts the photovoltaic characteristic model based on the power output curve under actual sunlight and temperature, while the energy storage battery agent updates its charge and discharge efficiency function based on state-of-charge (SOC) changes and load response. After all agents are updated, model state feedback is generated and passed to the upper-layer "model update" module to complete parameter synchronization.
[0033] During the policy learning phase, the digital twin agent uses the target control action generated during the current operating cycle as input. The agent predicts the electrical response value of each node under this action and uploads the response data to the policy optimization module. Using a reinforcement learning algorithm, the policy optimization module compares the digital twin response results with system objectives (such as peak shaving, valley filling, optimal electricity prices, and maximized equipment life) using a reward function to generate a policy update signal to guide agent behavior. This reward signal is transmitted back to each agent within the policy layer for iteration of the current scheduling strategy. Furthermore, to improve the efficiency of policy training, knowledge transfer can be implemented between different cycles or sites. The knowledge base sharing module supports the interoperability of policy structure parameters, response behavior patterns, and other content between digital twin agents. If two agents are found to be operating in highly similar scenarios, the established policy can be migrated to the new agent, shortening training time.
[0034] The entire process supports a decentralized trading platform interface at the virtual power plant level, enabling each agent to bid for energy transactions based on predicted electrical behavior and external electricity price curves, achieving adaptive economic optimal control. This implementation demonstrates that the intelligent scheduling system built based on the schematic structure integrates cross-cycle adaptability, local device intelligence, and global optimization capabilities, improving the operational stability and economic efficiency of distributed energy systems.
[0035] Step S20: constructing a power grid operation constraint response model based on the action response training data set, and generating constraint feedback information reflecting the feasibility of the current target control action based on the power grid operation constraint response model.
[0036] Specifically, a power grid operation constraint response model is constructed based on the node electrical response change, including: traversing all power grid nodes that have a linkage relationship with historical control actions, and respectively executing the binding of the electrical response change of each power grid node with the target control action to obtain multiple data pairs for characterizing the electrical disturbance response relationship; based on the multiple data pairs, combined with the preset power grid physical topology structure and the control coupling relationship between nodes, the actual electrical response change of each linked power grid node is used as the supervision target to construct a constraint response model that reflects the response law of the operating status of each power grid node to the historical control action.
[0037] Furthermore, the physical topology structure of the power grid is constructed based on the grid access relationship of each distributed power equipment participating in the scheduling, and is used to characterize the physical connection path and control influence path between each power grid node when constructing the constraint response model; the control coupling relationship between nodes is constructed based on the state coordinated change law of each power grid node under the influence of historical control actions, and is used to characterize the response propagation path and interference superposition characteristics of the target control action between nodes when constructing the constraint response model.
[0038] Specifically, the current target control action is a target control action input by the user, or a target control action generated based on the current cycle / state; the rule for generating the target control action based on the current cycle is: obtain the typical load curve of the power grid corresponding to the current cycle, and match the preset cycle and control action mapping rules, and extract the reference control action matching the current cycle as the target control action; the rule for generating the target control action based on the current operating state is: collect the current power grid state parameters, match them in combination with the preset state and action mapping rules, and extract the initial target control action that meets the current operating state requirements as the target control action.
[0039] In an embodiment of the present invention, in a virtual power plant scheduling scenario, due to the highly complex grid operating environment and the tight coupling of equipment states, if the control strategy does not fully consider the physical constraints and grid response laws, it is very easy to cause scheduling failure, equipment overload or system instability. Therefore, in order to ensure the executability and electrical safety of the scheduling strategy, it is necessary to introduce a grid operation constraint response model before scheduling, simulate the response of the current target control action, and generate constraint feedback information reflecting its feasibility. The following describes in detail the construction method of the model and the feedback generation logic in combination with the relevant processing flow in the embodiment of the present invention.
[0040] Specifically, the grid operation constraint response model is constructed based on the acquired action response training dataset. This dataset consists of the changes in node electrical responses generated by each grid node corresponding to the distributed power equipment participating in the dispatch, under the conditions of executing historical control actions. In practice, it is necessary to traverse all grid nodes that have a linkage relationship with historical control actions. This means that it is not limited to the nodes directly affected by the target control action, but also includes nodes that may be indirectly affected by it, such as upstream and downstream nodes and power flow intersection nodes.
[0041] For each grid node, the changes in its electrical response under multiple historical control actions (such as node voltage fluctuations, feeder power change rates, and inter-node power flow transfers) must be bound to the corresponding control actions to generate data pairs that reflect the node's response patterns. Each data pair consists of the characteristic parameters of the control action (such as start / stop commands, power regulation amplitude, and execution time window) and the corresponding node's electrical response parameters. The latter are numerical vectors covering physical variables in multiple dimensions.
[0042] Furthermore, building on the vast amount of data pairs already generated, the physical topology of the power grid and the control coupling relationships between nodes are introduced as important structural prior information for model construction. The physical topology of the power grid is constructed based on the access node information of each distributed power device participating in the dispatch. It is usually expressed in the form of an adjacency matrix or connection graph, recording the physical connectivity between nodes and reflecting the power flow transmission paths and action boundaries.
[0043] The modeling of inter-node control coupling relationships is based on the coordinated response patterns of grid nodes under the influence of historical control actions. For example, if a node performs a load switch and its upstream and downstream nodes experience a simultaneous voltage rebound or frequency disturbance, this indicates a strong coupling effect on the surrounding nodes. By counting a large number of such coordinated change events, a response propagation path matrix between nodes is formed, which is used to constrain the propagation strength and directionality of each response channel during model training.
[0044] When constructing a grid operation constraint response model, it is recommended to use a neural network model with graph structure processing capabilities (such as GCN and GAT). Its input is the combined characteristics of the aforementioned control actions and node states, and the edge weights are jointly regulated by the topological structure and the coupling relationship matrix. Using the actual electrical response changes of each interconnected grid node as the supervision target, supervised learning is performed. This enables the constructed constraint response model to accurately predict the electrical response trends of each node under any target control action, especially in complex scenarios such as nonlinear perturbations and large-scale equipment coordinated scheduling.
[0045] The constructed grid operation constraint response model is used to determine the feasibility of the current target control action. This target control action can be directly input by the user or dynamically generated by the system based on the current cycle or current operating status. For cycle-driven scenarios, the typical load curve of the grid can be matched according to the current cycle (such as peak, trough, and spike), and the corresponding action mapping rule library can be searched to extract reference actions that meet the operating characteristics of the cycle. For state-driven scenarios, the current operating status parameters of each node (such as load rate, voltage level, frequency deviation, etc.) need to be collected in real time and matched with the state-action mapping rule library to determine the most appropriate initial target control action.
[0046] This target control action is input into the grid operation constraint response model, and a complete response prediction process is executed to simulate the state changes of each grid node driven by this action, with particular attention paid to three types of electrical disturbance indicators: node voltage change, feeder power fluctuation rate, and inter-node power flow change rate. It should be noted that each electrical disturbance indicator is not observed in isolation but needs to be compared with equipment rated parameters, safe operating boundaries, physical topology boundaries, and load fluctuation tolerance to form a system-level judgment.
[0047] To this end, a grid operation constraint judgment standard is pre-established. This standard consists of multiple rules, including voltage tolerance thresholds, power flow variation limits, equipment load safety margins, and topology stability judgment logic. These rules can be implemented using interval judgments, logical combinations, and fuzzy intervals. The electrical disturbance indicators predicted by simulation are compared item by item with these judgment standards to determine whether the target control action exceeds any safety boundaries.
[0048] In one possible implementation, taking the distributed scheduling of a virtual power plant during the summer peak period as an example, the plant's historical control actions and the resulting node voltage, current, and power fluctuations are first collected to construct an action response training dataset. Subsequently, based on this dataset, combined with the physical grid topology formed by device access and the control coupling relationships between nodes, a constrained response model is constructed. During this period, target control actions, such as "activate two energy storage systems and reduce PV feed-in by 20%," are derived by matching the actual grid load curve and input into the constructed model. The model then determines the response of each node based on historical disturbance patterns. If a boundary feeder power exceeds the limit, it generates constraint feedback at the "boundary compensation required" level, prompting adjustments to the strategy. This feedback is used to optimize the scheduling strategy, preventing scheduling actions from violating operational safety boundaries and ensuring the enforceability of the scheduling strategy and grid operational stability.
[0049] Furthermore, constraint feedback information reflecting the feasibility of the current target control action is generated based on the grid operation constraint response model, including: inputting the current target control action into the grid operation constraint response model, simulating the voltage change, feeder power fluctuation rate and / or inter-node power flow change rate of each grid node under the action of the control action as a node-level electrical disturbance index; based on the electrical disturbance index, compared with the pre-constructed grid operation constraint judgment standard, performing feasibility level classification, and outputting constraint feedback information for identifying the feasibility of the target control action; wherein, the grid operation constraint judgment standard is constructed based on the rated voltage, safety margin threshold, topology stability boundary and load fluctuation tolerance of each grid node.
[0050] In this embodiment of the present invention, when generating constraint feedback information reflecting the feasibility of the current target control action based on the grid operation constraint response model, the current target control action is input into the trained grid operation constraint response model. This model has learned the patterns of disturbances caused by various actions on each node of the grid based on a large amount of historical action response data, thus providing predictive capabilities. In this step, the target control action can be user-specified or automatically generated by matching the load characteristics of the current operating cycle or grid state. Once input, the model performs response prediction.
[0051] Furthermore, the grid operation constraint response model simulates the electrical state changes at each grid node based on the input target control action. This includes predicting indicators such as the voltage change at each node, the power fluctuation rate of the connected feeder, and the rate of change of power flow with surrounding nodes. These electrical disturbance indicators constitute a set of action-induced disturbance characteristics, which are used to characterize the scope and intensity of the electrical impact of the target control action under the current structure.
[0052] These disturbance indicators are compared and analyzed against pre-set grid operation constraint judgment criteria. These criteria, pre-established based on grid design specifications, grid topology, and equipment operating data, primarily include the following: the permissible rated voltage fluctuation range for each node (e.g., ±5%), voltage safety margin thresholds (e.g., nodes marked as risky if below 90% of rated voltage), topology stability boundaries (e.g., no circuit breaker or circulating current can be triggered), and load fluctuation tolerances determined based on historical data (e.g., feeder power fluctuations exceeding 10% are marked as exceeding the limit). By comparing the disturbance indicators against these boundary conditions, it is possible to determine whether the current target control action violates any operational boundary.
[0053] Based on the comparison results, the target control action is classified into different feasibility levels, such as "completely feasible," "near the boundary," or "unfeasible," and corresponding constraint feedback information is output. This information serves as a key input for subsequent dispatch strategy training or control strategy revision, effectively avoiding the risk of ignoring operational boundaries during strategy generation. It achieves precise coupling between strategy generation and physical constraints, ensuring the feasibility of the dispatch plan and the safety and stability of the overall power grid operation.
[0054] Step S30: using the constraint feedback information as coupling information of the pre-trained scheduling strategy model, performing scheduling strategy training based on the current target control action, and obtaining a scheduling strategy model with constraint perception capability.
[0055] Specifically, the method also includes executing scheduling strategy model training, and the training rules are: constructing a basic scheduling strategy training data set based on each historical control action and its corresponding target response state; using the basic scheduling strategy training data set as a training sample, and using the scheduling effect as a supervision target to execute strategy behavior learning to obtain a scheduling strategy model.
[0056] Furthermore, the constraint feedback information is used as coupling information of the pre-trained scheduling strategy model, and scheduling strategy training is performed based on the current target control action to obtain a scheduling strategy with constraint perception capability, including: taking the current target control action and the corresponding constraint feedback information as input, embedding them into the reward function structure of the pre-trained scheduling strategy model, performing scheduling strategy model training, and outputting a scheduling strategy with constraint perception capability; wherein, during the training process, a negative reward is set for the target control action that violates the constraint feedback information, and a positive reward is set for the target control action that satisfies the constraint feedback information, and the range of the optional action space is compressed based on the accumulated reward and punishment results.
[0057] In the implementation of the present invention, the constraint feedback information is used as coupling information for a pre-trained dispatching strategy model. Dispatching strategy training is performed based on the current target control action to obtain a dispatching strategy model with constraint-aware capabilities. Specifically, to ensure that the dispatching strategy is able to perceive the grid's operating constraints, constraint feedback information is introduced during dispatching strategy model training as a reward function adjustment factor in the reinforcement learning process, thereby enabling quantitative evaluation of the feasibility of control actions and feedback on the reward and penalty mechanism. This training process consists of two phases: basic model training and constraint feedback coupled training, which respectively undertake the functions of dispatching behavior learning and constraint-aware capability development.
[0058] During the basic dispatch policy model training phase, a dispatch policy training dataset must be constructed. This dataset should focus on actual control actions taken during the historical operation cycle and their corresponding grid operational status responses. Each training sample includes the control action number, action execution time, control objective (such as node power injection, reactive power regulation, voltage regulation, etc.), action duration, and node operational status parameters after the action, such as voltage offset, power flow distribution, frequency variation, and line load. These status parameters are obtained by combining SCADA or PMU sampling data from the dispatch cycle with digital twin node simulation results. Next, the training objective should be to maximize the dispatch effectiveness evaluation function. Typical evaluation metrics include minimizing grid-wide active power loss, minimizing voltage offset, and maximizing load balance. Based on this training objective, policy models such as policy gradient, deep Q-network (DQN), or proximal policy optimization (PPO) from reinforcement learning methods are used to train the initial dispatch policy, resulting in a basic policy model.
[0059] After the basic policy model is constructed, the constraint feedback information coupling phase begins. The core idea of this phase is to integrate constraint feedback information into the reward and penalty mechanism of the reinforcement learning model, so that the dispatch strategy not only optimizes economic efficiency or operational efficiency but also meets the physical feasibility constraints of the power grid. Specifically, the current target control action is bound to the constraint feedback information output by the power grid operation constraint response model and used as the input variable of the reward function in the reinforcement learning process. The following structure should be considered in the reward function design: when a target control action causes a voltage offset at a grid node to exceed a safety threshold, feeder power fluctuations to exceed a set tolerance, frequent power flow changes, or power reversals, a significant negative reward should be set. Conversely, when a control action maintains voltage stability, balanced power flow distribution, and avoids overload, a positive reward should be given. Based on this, the control policy is updated based on the accumulated expected reward. Negative rewards are used to compress the action space containing default actions, gradually eliminating control behavior paths that are not feasible to execute.
[0060] During training, the target control action and constraint feedback information are input into the policy model via feature encoding. For example, a two-branch neural network structure is used: one branch processes the control action encoding, and the other processes the constraint feedback encoding. The two branches are then fused and fed into the main layer of the policy network. Alternatively, the two can be concatenated into a joint feature vector, which is then mapped to the policy space via a feedforward network, outputting an optimal action probability distribution or deterministic action outcome. Furthermore, to avoid learning failures caused by frequent constraint violations during training, a delayed penalty strategy is employed. Penalties are only applied after the cumulative number of violations exceeds a certain threshold, enhancing the robustness of policy learning.
[0061] To enhance the dispatching strategy's generalizability to diverse grid operating conditions, a multi-state sampling mechanism can be introduced. Training samples can be constructed to encompass a variety of typical load curves, distributed generation output scenarios, and operating conditions influenced by weather changes. Furthermore, a target control action perturbation mechanism can be introduced. This involves introducing fine-tuning perturbations based on the initial target action to guide the strategy model in learning the optimal behavior under boundary conditions. These perturbations can include adjustments to the control variable percentage, delay simulation, or target node switching.
[0062] Training convergence criteria should be based on a combination of the action space compression ratio, the scheduling feasibility achievement rate, and the reward function convergence curve. Scheduling strategy training is considered converged when the compressed action space is stable, the achievement rate is above the set threshold, and the reward function fluctuation range is within a tolerable range. The final output scheduling strategy model is a constraint-aware model. When the target control action is input, it can output the optimal scheduling control action that satisfies both the scheduling objective and the operational constraints.
[0063] Traditional reinforcement learning scheduling schemes are prone to theoretically optimal but physically ineffective performance. However, by using constraint feedback as a reward modifier for reinforcement training, the policy model internalizes its understanding of operational boundaries during training, proactively avoiding high-risk operations during execution. Furthermore, the action space compression mechanism not only improves the efficiency of policy generation, but also enhances the model's generalization capabilities for multi-state adaptability, laying the foundation for robust control in variable load scenarios and emergencies.
[0064] Step S40: Generate corresponding scheduling control actions based on the control strategy model.
[0065] Specifically, in the current cycle / state, the obtained scheduling strategy model with constraint perception capability is called; the target control action corresponding to the current cycle / state is input, and the scheduling strategy model is used to perform reasoning to generate the optimal scheduling control action in the current cycle, and the scheduling control action is distributed to each corresponding distributed power device to execute the control of each distributed power device.
[0066] In an embodiment of the present invention, the specific process of generating corresponding dispatch control actions based on the control strategy model is based on the premise that the dispatch strategy model has adaptability to the current cycle or current operating state. That is, the dispatch strategy model is a strategy model that is retrained or adapted based on the current cycle / state characteristics and grid constraint feedback information. The key to this step is that instead of using a universal, long-term unchanging strategy model, a corresponding strategy model is dynamically generated for each operating cycle or real-time monitored operating state characteristics, so that the control actions issued have stronger environmental compatibility and physical constraint consistency.
[0067] Furthermore, upon entering a specific operating cycle or detecting that the current operating state meets the policy reconstruction conditions, the target control action corresponding to the current cycle / state is obtained and combined with historical data to form a policy training sample. Using the obtained grid operation constraint feedback information, a constraint-aware dispatch policy model is generated through the policy training mechanism. This model is able to identify the characteristics of the current cycle / state and understand which actions are feasible, optimal, or highly profitable under the current physical constraints.
[0068] During the real-time scheduling phase, the constraint-aware scheduling policy model obtained above is used as the basis for generating effective scheduling behaviors for the current cycle / state. The target control action corresponding to the current cycle / state is input. This action serves as the starting point for the policy model's behavior, guiding the model to output the optimal scheduling control action for the current cycle or state based on its internal policy mapping rules, the learned reward structure, and the compressed action space. This process is completed quickly within the scheduling timeframe through forward reasoning, ensuring high real-time consistency between model output and system response.
[0069] Furthermore, to achieve lower-level control granularity, the optimal dispatch control action includes specific instructions for each participating distributed power device, such as power adjustment amplitude, start / stop signals, or reaction delay intervals. This control action is broken down by device type and node index based on the model output, and ultimately distributed to each corresponding device to complete control execution. Because the construction of the dispatch strategy model fully considers the coupling relationship and physical topology between grid nodes, it can avoid chain disturbances or feedback instability caused by single-point control in actual application, thereby improving the stability of the overall grid operation and the physical feasibility of the dispatch action.
[0070] In summary, by using the corresponding dispatching strategy model to generate and distribute control actions in the current cycle or current operating state, not only the cycle matching and state adaptation of the dispatching strategy are ensured, but also a dynamic response to constraint feedback information is achieved in the actual regulation process, effectively enhancing the adaptability of the virtual power plant to the actual power grid constraint environment, and significantly improving the scientificity, accuracy and feasibility of the dispatching behavior.
[0071] In one possible implementation, a regional virtual power plant, for example, includes multiple distributed power devices involved in scheduling, including photovoltaic power generation devices, energy storage units, and distributed load control devices. When performing intelligent scheduling tasks for this virtual power plant, the scheduling strategy training method and strategy generation mechanism described in this embodiment are employed.
[0072] A digital twin agent matrix is constructed based on the grid nodes to which each distributed power device is connected. The construction rules are as follows: a digital twin agent consisting of a physical modeling layer, a data-driven layer, and a state-deduction layer is established for each device, and then organized into a matrix structure based on node mapping relationships. The physical modeling layer is used to express the device's operating mechanism, the data-driven layer learns device behavioral characteristics from historical operating data, and the state-deduction layer outputs node-level electrical response changes after inputting target control actions.
[0073] During the initial modeling phase, the system collects data on control actions, equipment status, grid loads, and environmental parameters from historical operating cycles to form an action response training dataset. This dataset is used to construct a grid operation constraint response model. By modeling the mapping between electrical response changes and target control actions, regularity information reflecting the physical constraint boundaries is extracted.
[0074] The system, in conjunction with the grid operation constraint response model, performs constraint feedback information deduction for the target control action generated for the current cycle or current operating state. Constraint feedback information includes, but is not limited to, indicators such as voltage variation, feeder power fluctuation rate, and power flow disturbance. This information is then compared with grid constraint judgment criteria to generate a rating evaluation of the feasibility of the current target control action.
[0075] To improve the constraint-awareness of the scheduling policy model, the current target control action and the corresponding constraint feedback information are fed into the pre-trained policy model structure during the policy model training phase. By embedding a reward function mechanism, the action space is converged and optimized within the reinforcement learning framework: if an action violates a constraint, a negative reward is applied; if an action meets the constraint boundary, a positive incentive is given, and the non-preferred action space is continuously compressed accordingly. As a result, the trained policy model possesses the ability to recognize and respond to physical boundaries and constraints in real time.
[0076] During actual operation, the scheduling cycle can be dynamically divided according to the operating status, such as executing policy updates in units of hours, or triggering a model reconstruction mechanism when the operating status changes significantly. In this embodiment, a dynamically evolving machine learning mechanism is introduced to periodically retrain and evaluate the performance of the original policy model. The scheduling system archives and organizes the scheduling actions executed daily and their response effects, and automatically evaluates the deviation between the output of the policy model and the actual operating results. If the deviation continues to accumulate or increases significantly, the training process is automatically triggered, and the policy model is incrementally updated using the latest operating data, and representative historical samples for a period of time are retained to maintain the model's generalization ability.
[0077] The system also incorporates a rule-based learning mechanism. This means that if certain typical scheduling scenarios occur during operation (such as sudden equipment unavailability or an abnormal load surge), the system records the corresponding scheduling actions and their corresponding feedback, extracts them as potential scheduling rules, and submits them to experts for review. Once approved, these rules are incorporated into the rule base, providing empirical guidance for subsequent strategy training.
[0078] Figure 3 This is a system structure diagram of a virtual power plant intelligent dispatching system based on digital twins provided by an embodiment of the present invention. Figure 3 As shown, an embodiment of the present invention provides a virtual power plant intelligent dispatching system based on digital twins, and the system includes: an initial unit, which is used to determine the electrical response change of each grid node caused by each historical control action based on the digital twin agent matrix constructed by the grid nodes corresponding to each distributed power equipment participating in the dispatch, so as to form a corresponding action response training data set; a processing unit, which is used to construct a grid operation constraint response model based on the action response training data set, and generate constraint feedback information reflecting the feasibility of the current target control action based on the grid operation constraint response model; a model training unit, which is used to use the constraint feedback information as coupling information of a pre-trained dispatching strategy model, perform dispatching strategy training based on the current target control action, and obtain a dispatching strategy model with constraint perception capability; a dispatching unit, which is used to generate corresponding dispatching control actions based on the dispatching strategy model. A third aspect of the present invention provides a computer-readable storage medium having instructions stored thereon, which, when executed on a computer, enables the computer to execute the above-mentioned intelligent scheduling method for a virtual power plant based on digital twins.
[0079] The fourth aspect of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, the above-mentioned digital twin-based virtual power plant intelligent scheduling method is implemented.
[0080] A fifth aspect of the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the above-mentioned digital twin-based virtual power plant intelligent scheduling method.
[0081] Those skilled in the art will appreciate that all or part of the steps in the methods described in the aforementioned embodiments can be performed by instructing the relevant hardware through a program. The program, stored in a storage medium, includes instructions for causing a microcontroller, chip, or processor to execute all or part of the steps in the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0082] The above describes in detail the optional embodiments of the present invention in conjunction with the accompanying drawings. However, the embodiments of the present invention are not limited to the specific details in the above embodiments. Within the technical concept of the embodiments of the present invention, a variety of simple modifications can be made to the technical solutions of the embodiments of the present invention, and these simple modifications all fall within the scope of protection of the embodiments of the present invention. It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner unless there is any contradiction. In order to avoid unnecessary repetition, the embodiments of the present invention will no longer describe the various possible combinations separately.
[0083] In addition, the various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the embodiments of the present invention, they should also be regarded as the contents disclosed in the embodiments of the present invention.
Claims
1. A virtual power plant intelligent scheduling method based on digital twins, characterized in that: The method comprises: Based on the digital twin agent matrix constructed by the grid nodes corresponding to each distributed power device participating in the dispatch, the change in the electrical response of each grid node caused by each historical control action is determined to form a corresponding action response training dataset; Building a power grid operation constraint response model based on the action response training data set, and generating constraint feedback information reflecting the feasibility of the current target control action based on the power grid operation constraint response model; The constraint feedback information is used as coupling information of a pre-trained scheduling strategy model, and scheduling strategy training is performed based on the current target control action to obtain a scheduling strategy model with constraint perception capability; Generate corresponding scheduling control actions based on the scheduling strategy model.
2. The method according to claim 1, characterized in that The construction rules of the digital twin agent matrix are: For each distributed power device involved in scheduling, a digital twin agent is established, which includes a physical modeling layer, a data-driven layer, and a state deduction layer. The digital twin agents corresponding to the distributed power equipment are indexed and organized according to the grid nodes to which the corresponding distributed power equipment is connected to form a digital twin agent matrix; The physical modeling layer is used to represent the operating mechanism of the corresponding distributed power equipment; The data-driven layer is used to learn the behavioral characteristics of the distributed power equipment based on historical power and environmental data; The state deduction layer is used to output the change in node electrical response under the target control action.
3. The method according to claim 2, characterized in that Based on the digital twin agent matrix constructed by the grid nodes corresponding to each distributed power device participating in the dispatch, the electrical response changes of each grid node caused by each historical control action are determined, including: Each digital twin agent in the digital twin agent matrix performs each electrical response prediction task locally; wherein, The electrical response prediction task is a prediction task including historical control actions; Each digital twin agent outputs the electrical response change of the corresponding power grid node respectively, and based on the prediction results of multiple digital twin agents, the electrical response change of each power grid node caused by the target control action is obtained.
4. The method according to claim 3, characterized in that Each digital twin agent in the digital twin agent matrix locally performs an electrical response prediction task, including: Based on the electrical response prediction task, the physical modeling layer of the digital twin agent outputs the device mechanism parameters, and the data-driven layer outputs the historical behavior characteristics; Input the device mechanism parameters and the historical behavior characteristics into the state deduction layer to perform the electrical response prediction task, and output the electrical response change of the corresponding power grid node under the corresponding historical control action; The electrical response variation is uploaded to the digital twin agent matrix to perform electrical response variation aggregation of each power grid node through the digital twin agent matrix.
5. The method according to claim 1, wherein Building a power grid operation constraint response model based on the node electrical response variation includes: Traverse all grid nodes that have linkage relationships with historical control actions, and perform binding of electrical response changes of each grid node with target control actions to obtain multiple data pairs for characterizing electrical disturbance response relationships; Based on the multiple data pairs, combined with the preset physical topology of the power grid and the control coupling relationship between nodes, a constraint response model is constructed to reflect the response law of the operating status of each power grid node to the historical control action, taking the actual electrical response change of each linked power grid node as the supervision target.
6. The method according to claim 5, characterized in that The physical topology of the power grid is constructed based on the grid access relationship of each distributed power device participating in the scheduling, and is used to characterize the physical connection path and control influence path between each grid node when constructing the constraint response model; The inter-node control coupling relationship is constructed based on the state cooperative change law of each grid node under the influence of historical control actions, and is used to characterize the response propagation path and interference superposition characteristics of the target control action between nodes when constructing the constraint response model.
7. The method according to claim 1, characterized in that The current target control action is a target control action input by a user, or a target control action generated based on a current cycle / state; The rules for generating target control actions based on the current cycle are: Obtain the typical load curve of the power grid corresponding to the current cycle, match it with the preset mapping rules between the cycle and the control action, and extract the reference control action that matches the current cycle as the target control action; The rules for generating target control actions based on the current operating status are: The current grid state parameters are collected, matched with the preset state and action mapping rules, and the initial target control action that meets the current operating state requirements is extracted as the target control action.
8. The method according to claim 1, characterized in that Generating constraint feedback information reflecting the feasibility of the current target control action based on the power grid operation constraint response model includes: Inputting the current target control action into the grid operation constraint response model, simulating the voltage change, feeder power fluctuation rate, and / or inter-node power flow change rate of each grid node under the control action as a node-level electrical disturbance indicator; Based on the electrical disturbance index, the feasibility level classification is performed in comparison with the pre-built grid operation constraint judgment standard, and constraint feedback information for identifying the feasibility of the target control action is output; wherein, The grid operation constraint judgment standard is constructed based on the rated voltage, safety margin threshold, topology stability boundary and load fluctuation tolerance of each grid node.
9. The method according to claim 1, characterized in that The method further includes executing scheduling strategy model training, wherein the training rules are: Based on each historical control action and its corresponding target response state, a basic scheduling strategy training dataset is constructed; The basic scheduling strategy training data set is used as a training sample, and the scheduling effect is used as a supervision target to perform strategy behavior learning to obtain a scheduling strategy model.
10. The method according to claim 1, characterized in that The constraint feedback information is used as coupling information of the pre-trained scheduling strategy model, and scheduling strategy training is performed based on the current target control action to obtain a scheduling strategy with constraint perception capability, including: The current target control action and the corresponding constraint feedback information are used as input and embedded into the reward function structure of the pre-trained scheduling strategy model. The scheduling strategy model training is performed and the scheduling strategy with constraint perception capability is output. During the training process, negative rewards are set for target control actions that violate the constraint feedback information, positive rewards are set for target control actions that satisfy the constraint feedback information, and the range of the optional action space is compressed based on the accumulated reward and punishment results.
11. The method according to claim 1, wherein Generating corresponding scheduling control actions based on the scheduling strategy model includes: In the current cycle / state, call the obtained scheduling strategy model with constraint awareness capability; Input the target control action corresponding to the current cycle / state, perform reasoning through the scheduling strategy model to generate the optimal scheduling control action in the current cycle, and distribute the scheduling control action to each corresponding distributed power device to execute the control of each distributed power device.
12. A virtual power plant intelligent dispatching system based on digital twins, characterized by: The system comprises: An initialization unit is used to determine the electrical response changes of each grid node caused by each historical control action based on the digital twin agent matrix constructed by the grid nodes corresponding to each distributed power device participating in the dispatch, so as to form a corresponding action response training data set; a processing unit, configured to construct a power grid operation constraint response model based on the action response training data set, and generate constraint feedback information reflecting the feasibility of a current target control action based on the power grid operation constraint response model; A model training unit, configured to use the constraint feedback information as coupling information of a pre-trained scheduling strategy model, perform scheduling strategy training based on a current target control action, and obtain a scheduling strategy model with constraint perception capability; The scheduling unit is used to generate corresponding scheduling control actions based on the scheduling strategy model.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the intelligent scheduling method for a virtual power plant based on digital twins as described in any one of claims 1 to 11.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements the intelligent scheduling method of virtual power plant based on digital twins as described in any one of claims 1 to 11.
15. A computer program product comprising a computer program, characterized in that When executed by a processor, the computer program implements the intelligent scheduling method for virtual power plants based on digital twins according to any one of claims 1 to 11.
Citation Information
Patent Citations
Optimized regulation and control method for participation of virtual power plant in multi-market transaction based on digital twinning
CN114219186A
Scheduling method for virtual power plant to participate in power market transaction based on digital twinning
CN116957294A
Virtual power plant load prediction and dynamic adjustment optimization system and method
CN120338563A
Virtual power plant green energy consumption cooperation method and system fusing digital twinning and reinforcement learning
CN120409220A
Method and apparatus for constructing digital twin hybrid model of main device of power system
WO2024103934A1
Cited By
New energy consumption-oriented power grid digital twinborn optimization scheduling method and system
CN121529850A
Optimal dispatching method and system for power grid digital twin oriented to new energy consumption
CN121529850B
Efficient power-saving optimizer control method and device suitable for power grid and medium
CN121688963A
Power grid dispatching method and device of artificial intelligence Transform model based on physical constraint
CN121886337A