Photovoltaic grid-connected flexible direct current power transmission and energy storage collaborative control system and method
By using digital twins, a forward-looking risk assessment module, and a deep reinforcement learning control strategy, the problems of transient stability control lag and economic dispatch disconnect in photovoltaic grid-connected flexible DC transmission and energy storage systems were solved, achieving proactive defense and optimized collaborative control of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG KECHUANG POWER TECH CO LTD
- Filing Date
- 2025-11-17
- Publication Date
- 2026-05-05
AI Technical Summary
In existing technologies, the transient stability control of photovoltaic grid-connected flexible DC transmission and energy storage systems lacks foresight, resulting in delayed response. Furthermore, the economic dispatch and transient security requirements are disconnected, making it impossible to achieve unified and coordinated optimization.
By establishing digital twins and forward-looking risk quantification indicators, a deep reinforcement learning control strategy is constructed, and a parameter correction closed loop is established to achieve intelligent collaborative optimization across multiple time scales.
It has enabled the system to have proactive defense capabilities, improved its ability to resist unknown disturbances and its stability, optimized the accuracy and rationality of economic scheduling decisions, and ensured the safe and economical operation of the system at multiple time scales.
Smart Images

Figure CN121150334B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system control technology, and in particular to a photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system and method. Background Technology
[0002] With the transformation of the global energy structure, the penetration rate of renewable energy, represented by photovoltaics, in the power system is growing at an unprecedented rate. However, the inherent intermittency and randomness of photovoltaic power generation, as well as the reduction in system rotational inertia caused by its grid connection through power electronic devices, pose significant challenges to the frequency and voltage stability of the power grid. To promote the large-scale consumption of photovoltaic energy and maintain the safe and stable operation of the power grid, flexible direct current transmission technology (VSC-HVDC) is typically used for long-distance, large-capacity power transmission, and energy storage systems (ESS) are configured to smooth power fluctuations and provide ancillary services. Therefore, how to effectively coordinate and control photovoltaic units, flexible direct current transmission units, and energy storage units has become a key technical issue for ensuring the safe and economical operation of the new power system.
[0003] In existing technologies, the control of such systems typically employs a layered, decoupled architecture. At the transient stability control level, it mainly relies on the local controllers of each unit, using strategies such as droop control and virtual synchronous generator (VSG) control. These control strategies simulate the external characteristics of synchronous generators to provide virtual inertia and damping support to the grid, thus coping with sudden frequency or voltage disturbances. At the economic operation level, typically on a minute- or hourly timescale, based on forecasts of electricity prices, load, and photovoltaic output, model predictive control (MPC) or other optimal scheduling algorithms are used to determine the power allocation plan for each unit, aiming to minimize operating costs or maximize economic benefits.
[0004] While existing technologies have achieved the basic control functions of the system to a certain extent, some shortcomings still exist:
[0005] Existing transient stability control methods are essentially passive response methods. Whether it's droop control or virtual synchronous generator control, their control logic is based on real-time measurement of the system's frequency or voltage deviation at the current moment. This means that the control system only begins to act after the system stability has been disturbed and a deviation has occurred. This "post-event response" model lacks foresight. For rapidly developing or complex faults caused by multiple concurrent factors, it may miss the optimal control opportunity due to response delays or insufficient coordination, failing to fundamentally prevent instability events. The root cause is that these control methods lack a mechanism that can anticipate future system dynamics and predict instability risks in real time.
[0006] Furthermore, existing technologies lack an objective and dynamic decision-making basis when addressing the conflict between system transient security and economic operation. Typically, to ensure transient security, flexible resources such as energy storage need to reserve substantial standby capacity, but this directly sacrifices their opportunity to participate in the electricity market and gain economic benefits. This reservation strategy is often based on offline, static worst-case analysis, setting fixed, conservative standby thresholds. This approach cannot dynamically quantify the true value of "reserving flexibility for unforeseen needs." Conversely, optimal scheduling with economic efficiency as the sole objective may over-utilize flexible resources, resulting in insufficient standby to provide effective transient support when the system encounters sudden disturbances. The root of the problem lies in the fact that transient security requirements and economic dispatch objectives are separated into two independent models, lacking a unified framework that dynamically and endogenously integrates the "risk cost" of transient events into the economic dispatch decision-making process. Summary of the Invention
[0007] The purpose of this invention is to provide a photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system and method, which solves the problems in the prior art where transient stability control is slow to respond due to lack of foresight, and economic dispatch and transient security requirements cannot be unified and coordinated for optimization due to the separation of decision-making frameworks.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] The first aspect of this invention provides a photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system:
[0010] This system achieves intelligent collaborative optimization across multiple time scales by establishing a real-time mapping between digital twins and physical systems, introducing forward-looking risk quantification indicators, constructing a deep reinforcement learning control strategy to enhance risk, and establishing a parameter correction closed loop that feeds back from underlying control experience to the upper-level economic model.
[0011] To achieve the above objectives, the photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system provided by the present invention may specifically include:
[0012] The system status awareness module is configured as follows:
[0013] The system collects and integrates operational data from photovoltaic units, flexible DC transmission units, energy storage units, and grid connection points via a communication network, synchronizing and generating a multi-dimensional original state vector using a unified timestamp. .
[0014] The original state vector At any moment Its composition can be:
[0015] ;
[0016] In the formula, This refers to the voltage at the grid connection point. For grid connection frequency; The rate of change of frequency; , The active and reactive power output of the photovoltaic unit; , The active and reactive power exchanged by the flexible DC transmission converter station; , The active and reactive power exchanged for the energy storage unit; This refers to the state of charge of the energy storage unit; This refers to the health status of the energy storage unit.
[0017] A digital twin and forward-looking risk assessment module, which is connected to the output of the system state awareness module, is configured as follows:
[0018] Receive the original state vector Furthermore, it utilizes a built-in, high-fidelity digital twin model synchronized with the physical system to perform parallel simulations under ultra-real-time conditions, thereby calculating and outputting a forward-looking risk gradient vector characterizing the future stability risk of the system. The specific calculation process of this module is as follows:
[0019] First, define a comprehensive stability index. This indicator is a function for evaluating the system's security level; secondly, it is used to control the current system's control vector within the digital twin model. Apply a preset virtual micro-perturbation Calculate the comprehensive stability index The change in the value of the risk is then approximated by the numerical difference method to obtain the forward-looking risk gradient vector:
[0020] ;
[0021] In the formula, For the current moment The aforementioned forward-looking risk gradient vector; This is a pre-defined comprehensive stability index function used to evaluate the system's security level. For the current moment The control vector of the system; This is a preset virtual micro-perturbation applied to the control vector; Calculate the norm of a vector.
[0022] A transient collaborative control module, the inputs of which are connected to the outputs of the system state perception module and the digital twin and forward-looking risk assessment module, respectively. This module is configured as follows:
[0023] First, the received original state vector With the aforementioned forward-looking risk gradient vector To perform fusion to form a state vector with increased risk. :
[0024] ;
[0025] In the formula, For the current moment The risk-enhanced state vector; For the current moment The original state vector; For the current moment The aforementioned forward-looking risk gradient vector; An operator that combines two vectors;
[0026] Secondly, the state vector with the risk augmentation Input into a pre-trained deep reinforcement learning policy network To output a differentiated transient cooperative action vector in real time. The action vector may include control commands for the virtual moment of inertia, damping coefficient, active power reference value, and reactive power reference value of the photovoltaic unit, flexible DC transmission unit, and energy storage unit.
[0027] Finally, the experience data generated from each interaction—a sequence containing information such as state, action, and reward—is stored in an experience replay pool. .
[0028] The energy optimization management module, whose input is connected to the experience playback pool of the transient collaborative control module. Connection. This module is configured as follows:
[0029] First, the system's flexibility resources, such as the reserve capacity of energy storage, are treated as a European put option, and its option value is calculated using the Black-Scholes option pricing model. This is used to assess the economic value of retaining this flexibility resource. The calculation formula is as follows:
[0030] ;
[0031] In the formula, For the current moment The option value of the aforementioned flexible resources; For the current moment The strike price of the option; For the current moment The immediate economic benefits of the aforementioned flexible resources are used as the present value of the underlying asset of the option; is the base of the natural logarithm; The risk-free rate; This represents the option's expiration time, corresponding to the scheduling time window; The cumulative distribution function of the standard normal distribution; and It is an intermediate variable.
[0032] Secondly, from the aforementioned experience replay pool Extract historical interaction data and analyze the key parameter of the option pricing model, namely the strike price. and volatility Dynamic adjustments are made. Among these, the execution price... The correction formula is:
[0033] ;
[0034] In the formula, For the current moment The strike price of the option pricing model described above; An operator for calculating mathematical expectation; This is the experience playback pool; The state is the empirical data extracted from the empirical replay pool; Actions in the experience data extracted from the experience replay pool; The reward is the experience data extracted from the experience replay pool; A set of pre-defined actions used for transient support.
[0035] volatility It can be obtained by calculating the standard deviation of the historical reward signals in the experience replay pool.
[0036] Finally, by comparing the instantaneous benefits of immediately using the aforementioned flexibility resources. With the value of the option Select the strategy corresponding to the one with the highest value to generate the economic operating power baseline for that scheduling period. .
[0037] A hierarchical collaborative instruction execution module is included, with its inputs connected to the outputs of both the transient collaborative control module and the energy optimization management module. This module is configured to integrate the transient collaborative action vectors. and the economic operating power baseline And generate the final control command in accordance with the principle of safety first.
[0038] When the system is running stably, the final control command is mainly based on tracking the economic operating power baseline; when the system is disturbed or a risk is foreseen, the module will assign the transient cooperative action vector the highest execution priority.
[0039] A second aspect of this invention provides a method for coordinated control of photovoltaic grid-connected flexible DC transmission and energy storage:
[0040] This method corresponds to the system described in the first aspect. By executing a series of steps, it achieves a deep synergy between transient stability and economic operation. Specifically, the method includes:
[0041] S1. State Awareness Step: Continuously collect operational data from photovoltaic units, flexible DC transmission units, energy storage units, and grid connection points, and synchronously integrate them to generate a multi-dimensional original state vector. .
[0042] S2. Risk assessment step: The original state vector... The input is fed into a high-fidelity digital twin model synchronized with the physical system; this model is used for ultra-real-time parallel simulation, and the sensitivity of a preset comprehensive stability index to changes in the current control vector is calculated, thereby obtaining a forward-looking risk gradient vector characterizing the future stability risk of the system. .
[0043] S3. Transient control step: Convert the original state vector... With the aforementioned forward-looking risk gradient vector Fusion to form a state vector with increased risk. The risk-enhanced state vector is input into a pre-trained deep reinforcement learning policy network to generate differentiated transient cooperative action vectors in real time. Simultaneously, the interaction data, including status, actions, and rewards, is stored in an experience replay pool. .
[0044] S4. Energy Management Steps: Within a minute-level scheduling cycle, firstly, historical interaction data is extracted from the experience replay pool, and the key parameters (including strike price and volatility) of the preset real option pricing model are dynamically corrected accordingly; then, the option value of the system's flexibility resources is calculated using the corrected model; finally, by comparing the option value with the instantaneous returns of immediately usable resources, an optimal decision is made to generate the economic operating power baseline for that scheduling cycle. .
[0045] S5. Cooperative Execution Step: Integrate the transient cooperative action vectors and the economic operating power baseline When the system is running stably, control commands are issued primarily to track the economic operating power baseline. When the system experiences disturbances or risks are anticipated, the transient cooperative action vector is issued as the highest priority command to ensure system safety. Once the disturbances are eliminated, the control commands smoothly return to tracking the economic operating power baseline.
[0046] In summary, the present invention has at least one of the following beneficial technical effects:
[0047] 1. This invention, by setting up a digital twin and forward-looking risk assessment module, can calculate a forward-looking risk gradient vector representing future risks based on high-fidelity real-time extrapolation. This vector expands the decision-making basis of the transient collaborative control module from merely responding to the current system state to quantitatively predicting future stability risks. This mechanism transforms the system's control mode from traditional passive response to active defense, enabling more targeted control measures to be taken at the initial stage or even before a disturbance occurs, thereby enhancing the system's resilience and stability in the face of unknown disturbances.
[0048] 2. This invention solves the problem of key parameters relying on static estimation or empirical tuning in traditional economic scheduling by setting up an energy optimization management module and introducing a parameter correction mechanism based on historical data from an experience replay pool. The exercise price of the real options model is dynamically calibrated using a formula. This allows the opportunity cost upon which scheduling decisions are based to truly and dynamically reflect the actual cost of underlying transient support control. This makes economic scheduling decisions more accurate and rational, avoiding economic losses caused by excessive conservatism or excessive aggressiveness, and improving the overall operational efficiency of the system.
[0049] 3. This invention establishes an information feedback channel from the experience playback pool of the transient collaborative control module to the parameter correction unit of the energy optimization management module, and combines this with the dynamic integration of the priorities of the two layers of instructions by the hierarchical collaborative instruction execution module. This breaks down the separation between transient safety control and economic optimization scheduling in the system architecture. This design forms an organic whole where transient physical constraints can be dynamically transmitted and corrected to the upper-level economic model, and economic objectives can guide the steady-state operation of the system. Thus, while ensuring absolute system safety, it maximizes the exploitation of the system's economic potential and achieves multi-timescale, multi-objective collaborative control. Attached Figure Description
[0050] Figure 1 This is a schematic diagram of the system structure of the present invention;
[0051] Figure 2 This is a schematic diagram of the method flow of the present invention;
[0052] Figure 3This is a schematic diagram of the internal functions of the system state perception module of the present invention;
[0053] Figure 4 This is a schematic diagram illustrating the internal functions of the digital twin and forward-looking risk assessment module of the present invention;
[0054] Figure 5 This is a schematic diagram of the internal functions of the energy optimization management module of the present invention;
[0055] Figure 6 This is a schematic diagram of the internal functions of the transient collaborative control module of the present invention. Detailed Implementation
[0056] The following is in conjunction with the appendix Figure 1 - Appendix Figure 6 The present invention will be further described in detail below.
[0057] like Figure 1 As shown, Figure 1 This is a schematic diagram of the system structure of a photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to an embodiment of the present invention. The system may include:
[0058] The system includes a system status awareness module 10, a digital twin and forward-looking risk assessment module 20, a transient collaborative control module 30, an energy optimization management module 40, and a hierarchical collaborative instruction execution module 50.
[0059] The system state awareness module 10 is configured to continuously collect various electrical and non-electrical operating data from the photovoltaic unit, flexible DC transmission unit, energy storage unit, and public grid connection point via data acquisition equipment and communication network. This module performs timestamp alignment, cleaning, and synchronization processing on the collected multi-source heterogeneous data to generate a unified, multi-dimensional original state vector. The original state vector The specific composition can be as follows:
[0060] ;
[0061] In the formula, This refers to the voltage at the grid connection point. For grid connection frequency; The rate of change of frequency; , The active and reactive power output of the photovoltaic unit; , The active and reactive power exchanged by the flexible DC transmission converter station; , The active and reactive power exchanged for the energy storage unit; This refers to the state of charge of the energy storage unit; This refers to the health status of the energy storage unit.
[0062] The digital twin and forward-looking risk assessment module 20 has its input connected to the output of the system state awareness module 10. This module is configured to receive the raw state vector. Based on this, ultra-real-time parallel simulations are performed in the built-in high-fidelity digital twin model to calculate and output a forward-looking risk gradient vector characterizing the future stability risk of the system. The vector is approximated by the following formula:
[0063] ;
[0064] In the formula, For the current moment Forward-looking risk gradient vector; This is a pre-defined comprehensive stability index function used to evaluate the system's security level. For the current moment The system's control vector; This refers to a preset virtual small perturbation applied to the control vector; Calculate the norm of a vector.
[0065] The transient collaborative control module 30 has its inputs connected to the outputs of the system state perception module 10 and the digital twin and forward-looking risk assessment module 20, respectively. This module is configured to process the original state vector... With forward-looking risk gradient vector By merging, a state vector with increased risk is formed:
[0066] ;
[0067] In the formula, For the current moment The risk-enhanced state vector; For the current moment The original state vector; For the current moment Forward-looking risk gradient vector; This is an operator that combines two vectors.
[0068] Subsequently, the module augments the state vector with risk. The input is fed into a pre-trained deep reinforcement learning policy network to output a differentiated transient cooperative action vector in real time. Simultaneously, the interaction data (status, actions, rewards, etc.) is stored in an experience replay pool. .
[0069] The energy optimization management module 40 has its input connected to the experience playback pool in the transient collaborative control module 30. Data interaction connections exist. This module is configured to establish an economic operating power baseline based on real options theory within a minute-level scheduling cycle. .
[0070] Specifically, this module treats the system's flexibility resources as put options and uses the Black-Scholes model to calculate their option value. The calculation formula is as follows:
[0071] ;
[0072] In the formula, For the current moment The option value of flexible resources; For the current moment The strike price of the option; For the current moment The immediate economic benefits of using flexible resources are used as the present value of the underlying asset of the option. is the base of the natural logarithm; The risk-free rate; This represents the option's expiration time, corresponding to the scheduling time window; The cumulative distribution function of the standard normal distribution; and It is an intermediate variable.
[0073] Meanwhile, this module draws from the experience replay pool. Extract the data and use the following formula to calculate the strike price of the option pricing model. Perform dynamic correction:
[0074] ;
[0075] In the formula, For the current moment The strike price of the option pricing model; An operator for calculating mathematical expectation; For experience replay pool; The state in the experience data extracted from the experience replay pool; Actions in the experience data extracted from the experience replay pool; Rewards are derived from experience data extracted from the experience replay pool. A set of pre-defined actions used for transient support.
[0076] The hierarchical collaborative instruction execution module 50 has its inputs connected to the outputs of the transient collaborative control module 30 and the energy optimization management module 40, respectively. This module is configured to integrate transient collaborative action vectors. and economic operating power baseline It generates and issues final control instructions to each execution unit of the system based on preset priority logic.
[0077] See attached document Figure 2 , Figure 2 This is a schematic flowchart of a method for coordinated control of photovoltaic grid-connected flexible DC transmission and energy storage according to an embodiment of the present invention. The present invention also provides a coordinated control method corresponding to the above system, which may include the following steps:
[0078] Step S100, Perform state perception: Continuously collect and integrate the operating data of each unit in the photovoltaic grid-connected flexible DC transmission and energy storage collaborative control system to generate the original state vector. .
[0079] Step S200, Perform risk assessment: Convert the original state vector Input a digital twin model for ultra-real-time simulation to calculate and output a forward-looking risk gradient vector characterizing the future stability risk of the system. .
[0080] Step S300, Perform transient control: fuse the original state vector With forward-looking risk gradient vector To form a state vector with increased risk It generates differentiated transient cooperative action vectors based on deep reinforcement learning. At the same time, the interaction data is stored in the experience replay pool. .
[0081] Step S400, Perform energy management: Establish an economic operating power baseline based on real options theory. and utilize the experience replay pool The data in the model is used to dynamically correct the key parameters of the theoretical model.
[0082] Step S500, Perform cooperative execution: Integrate and execute transient cooperative action vectors and economic operating power baseline This process generates and issues final control commands. When the system is running stably, the commands primarily track the economic operating power baseline; when the system experiences disturbances or risks are foreseen, the transient cooperative action vectors are assigned the highest execution priority.
[0083] The method of the present invention is implemented by the system of the present invention through the sequential execution of the above steps. System state perception module 10 executes state perception step S100. Digital twin and forward-looking risk assessment module 20 executes risk assessment step S200. Transient collaborative control module 30 executes transient control step S300. Energy optimization management module 40 executes energy management step S400. Hierarchical collaborative instruction execution module 50 executes collaborative execution step S500. Through the collaborative work of each module and step, unified control of the transient safety and economic operation of the photovoltaic grid-connected flexible DC transmission and energy storage system is achieved.
[0084] See attached document Figure 3 , Figure 3 This is a schematic diagram illustrating the internal functions of a system status awareness module 10 according to an embodiment of the present invention. The purpose of the system status awareness module 10 is to provide real-time, accurate, and synchronous system status information for subsequent risk assessment and control decisions.
[0085] In one specific implementation, the system status sensing module 10 includes multiple data acquisition interfaces, which are respectively connected to measuring devices located at the grid common point of connection (PCC), photovoltaic power generation unit, flexible DC transmission converter station (VSC), and energy storage unit (ESS). These measuring devices may include, but are not limited to: current transformers for measuring voltage and current, energy meters for measuring active and reactive power, and battery management systems (BMS) for monitoring the state of charge (SOC) and state of health (SOH) of the energy storage unit. Data is transmitted to the central processing unit of the system status sensing module 10 via reliable communication media such as fiber optics or industrial Ethernet, using standard industrial communication protocols such as IEC61850 and Modbus TCP / IP.
[0086] To ensure absolute temporal consistency of data collected from different physical locations, the system status awareness module 10 is equipped with a high-precision clock synchronization unit. This unit can synchronize with a standard time server by receiving pulse-of-seconds (PPS) signals from the Global Positioning System (GPS) or via the Network Time Protocol (NTP), assigning a uniform, high-precision timestamp to each frame of collected data. For the received raw measurement data, this module also performs necessary preprocessing operations.
[0087] Preprocessing may include: using data verification algorithms (such as CRC check) to identify and discard damaged data packets during transmission; using digital filters (such as Kalman filters or low-pass filters) to smooth noisy signals (such as grid connection point frequencies); and applying methods such as data interpolation or previous value hold to handle temporary communication interruptions or data loss in order to ensure the continuity of the output data stream.
[0088] After completing data acquisition and preprocessing, the system state awareness module 10 combines all synchronized and cleaned data items according to a predefined structure to generate the data for that moment. The original state vector The vector is formatted as a numerical array or a structured data object, consistent with the formulas shown in the preceding general description.
[0089] Ultimately, this module will generate the original state vector. The data is simultaneously sent to the input terminals of the digital twin and forward-looking risk assessment module 20 and the transient collaborative control module 30 via the internal data bus or network interface, serving as the real-time basis for subsequent calculations and decisions by these two modules.
[0090] See attached document Figure 4 , Figure 4 This is a schematic diagram illustrating the internal functions of a digital twin and forward-looking risk assessment module 20 according to an embodiment of the present invention. The digital twin and forward-looking risk assessment module 20 receives an original state vector from the system state awareness module 10 at its input terminal. Its core function is to quantify the stability risk of the system within a very short time window in the future based on this real-time state, and output a forward-looking risk gradient vector. To the transient collaborative control module 30.
[0091] In one specific implementation, the core of this module is a high-fidelity digital twin model. This model consists of a set of differential algebraic equations (DAEs) that accurately describe the dynamic characteristics of a photovoltaic grid-connected flexible DC transmission and energy storage system.
[0092] Specifically, the model includes: an inverter model for the photovoltaic unit, which accurately characterizes the dynamic process of maximum power point tracking (MPPT) control and voltage-current dual closed-loop control; a detailed model of the flexible DC-DC converter station (VSC), which includes mathematical expressions for constant active / reactive power outer loop control and dq-axis current inner loop decoupling control; and a bidirectional DC-DC converter and battery equivalent circuit model for the energy storage unit.
[0093] The parameters of the digital twin model (such as line impedance, controller parameters, etc.) are configured during system initialization and can be periodically identified and corrected offline based on the historical operating data of the physical system to ensure a high degree of consistency between the model and the physical entity.
[0094] This module is further configured with an uncertainty perturbation simulation unit. This unit generates a set of typical future perturbation scenarios based on statistical analysis of historical data or short-term perturbation predictions (such as photovoltaic power fluctuation predictions). Upon receiving each new original state vector... Then, the module uses the vector as an initial condition and loads the above set of disturbance scenarios to perform parallel, ultra-real-time simulation calculations in a high-fidelity digital twin model, thereby obtaining multiple future state trajectories of the system under different disturbances.
[0095] Based on the parallel simulation results described above, a risk gradient calculation unit is responsible for quantitatively calculating forward-looking risks. First, a comprehensive stability index function is defined. This metric is used to evaluate the security level of a system. It can be constructed using a weighted function, for example:
[0096] ;
[0097] In the formula, This represents the absolute value of the maximum deviation of the power grid frequency within the simulation time window; This represents the absolute value of the maximum deviation of the grid connection point voltage. The maximum absolute value of the rate of change of frequency; , , These are preset weight coefficients used to characterize the importance of different stability dimensions.
[0098] Subsequently, the risk gradient calculation unit employs a numerical difference approximation method to calculate the forward-looking risk gradient vector by performing the following steps. :
[0099] Step 1: Using the current system's control vector and state vector Using these as initial conditions, a simulation is run in the digital twin model to obtain a baseline comprehensive stability index value. .
[0100] Step 2: Control Vector Apply a preset, extremely small virtual perturbation to one or more components. To form a new control vector .
[0101] Step 3: Using the new control vector and the same state vector Using these as initial conditions, the simulation is run again to obtain a new comprehensive stability index value. .
[0102] Step 4: Using the aforementioned formula:
[0103] ;
[0104] In the formula, For the current moment Forward-looking risk gradient vector; This is a pre-defined comprehensive stability index function used to evaluate the system's security level. For the current moment The system's control vector; This refers to a preset virtual small perturbation applied to the control vector; Calculate the norm of a vector.
[0105] The sensitivity of the stability index to current control changes, i.e., the forward-looking risk gradient, is calculated. The magnitude of each component of this vector directly reflects the degree of impact of adjusting the corresponding control variable on the future stability of the system. After calculation, this module outputs this vector to the transient cooperative control module 30.
[0106] See attached document Figure 5 , Figure 5 This is a schematic diagram illustrating the internal functions of a transient collaborative control module 30 according to an embodiment of the present invention. The purpose of the transient collaborative control module 30 is to generate and output real-time control commands that can ensure the dynamic safety of the system based on the current system state and future risk predictions.
[0107] In one specific implementation, the input of module 30 receives the raw state vector from system state sensing module 10. and the forward-looking risk gradient vector from the digital twin and forward-looking risk assessment module 20 The state fusion unit within the module first concatenates or stacks these two vectors to form a higher-dimensional, more information-rich risk augmentation state vector. The risk-enhanced state vector It will serve as the sole state input for the deep reinforcement learning agent within the module.
[0108] Deep reinforcement learning agents employ an actor-critic architecture, such as Deep Deterministic Policy Gradient (DDPG) or Soft Actor-Critic (SAC) algorithms. This architecture consists of a policy network (ActorNetwork) and a value network (CriticNetwork). Both networks can be constructed from multi-layer fully connected neural networks, i.e., multilayer perceptrons (MLPs).
[0109] The function of the policy network is to receive risk-enhanced state vectors. As input, it directly outputs a deterministic, multi-dimensional transient cooperative action vector. The function of a value network is to receive state vectors. Action vectors output by the policy network As input, it outputs a scalar Q value, which is used to evaluate the long-term value of performing the action in this state.
[0110] Transient cooperative action vector The vector is multi-dimensional, directly corresponding to multiple controllable degrees of freedom in the system. Specifically, this vector can be defined as:
[0111] ;
[0112] In the formula, and These represent the adjustment amounts of the virtual moment of inertia and damping coefficient in the control parameters of photovoltaic or energy storage inverters, respectively. For the current moment The transient cooperative action vector; The adjustment amount set for the virtual moment of inertia in the control parameters of the photovoltaic unit inverter; The adjustment amount set for the damping coefficient in the control parameters of the photovoltaic unit inverter; This is the adjustment amount for the active power reference value of the photovoltaic unit; This is the adjustment amount for the reactive power reference value of the photovoltaic unit; The adjustment amount set for the virtual moment of inertia in the control parameters of the energy storage unit converter; The adjustment amount set for the damping coefficient in the control parameters of the energy storage unit converter; This is the adjustment amount for the reference value of the active power of the energy storage unit; This is the adjustment amount for the reactive power reference value of the energy storage unit.
[0113] The components of this action vector cover the control commands of one or more of the photovoltaic unit, flexible DC transmission unit, and energy storage unit.
[0114] To guide the deep reinforcement learning agent to learn effective control strategies, a dedicated reward function was designed within the module. At each control moment An immediate reward value is calculated based on the system's performance. This reward function is constructed in the following form:
[0115] ;
[0116] In the formula: and These represent the grid connection point frequency and voltage at the current moment, respectively. and These are the rated frequency and rated voltage of the power grid, respectively. The control action vector output at the previous moment; , , These are the preset weighting coefficients for positive constants.
[0117] The design goal of this reward function is to encourage the agent to maintain the stability of the system voltage and frequency by applying negative weights to the quadratic terms of the frequency and voltage deviations; and to constrain the control cost by applying negative weights to the norm of the control action vector, thereby avoiding unnecessary or overly drastic control actions.
[0118] The module's operation consists of two phases: offline training and online execution. During offline training, the agent explores numerous interactions within a high-fidelity digital twin model. In each interaction, the agent outputs an action based on its current state, and the digital twin model returns the next state and an immediate reward. The experience samples generated from these interactions are known as quadruplets. Large quantities are stored in an experience replay pool. middle.
[0119] During training, a small batch of experience samples is randomly drawn from this experience replay pool to simultaneously update the network parameters of the policy network and the value network. Specifically, the value network is updated by minimizing the temporal difference error (TD-error), while the policy network is updated according to the gradient information provided by the value network in the direction that can obtain a higher Q value.
[0120] After sufficient offline training, the trained policy network is deployed in the module for online execution. During online execution, the module uses only the policy network, augmenting the state vector based on the real-time input risk. Generate transient cooperative action vectors with extremely low computational latency. The results are then output to the hierarchical collaborative instruction execution module 50. Simultaneously, new experience samples generated during online operation are continuously stored in the experience replay pool. It can be used for subsequent incremental training or fine-tuning of the model.
[0121] See attached document Figure 6 , Figure 6 This is a schematic diagram illustrating the internal functions of an energy optimization management module 40 according to an embodiment of the present invention. The energy optimization management module 40 operates on a minute- or hour-scale timescale, and its purpose is to establish an economically optimal operating power baseline that guides the system's energy exchange in steady state. It also aims to achieve a balance between transient security requirements and long-term economic benefits by quantifying the value of system flexibility resources.
[0122] In one specific implementation, this module equates the system's dispatchable flexibility resources, such as the reserve capacity in energy storage units designated for emergency power support, to a European put option in finance. In this equivalent model, immediately using the flexibility resource to obtain instantaneous economic benefits (such as participating in energy market arbitrage) is considered selling the underlying asset corresponding to the option; while retaining the flexibility resource for future unforeseen needs is equivalent to continuing to hold the put option. The value of this option represents the economic value of retaining the flexibility.
[0123] A real options pricing unit is configured to use the Black-Scholes option pricing formula to calculate the option value of a flexible resource. The calculation process first requires determining several key input parameters: the present value of the underlying asset. This refers to the instantaneous economic benefit that could be obtained by immediately using the flexible resource at the current moment; this value can be calculated based on the real-time market electricity price and dispatchable capacity; option expiration time. It is set to the current energy dispatch cycle, such as 15 minutes or 1 hour; and the risk-free interest rate. Its value can be either the government bond interest rate or the interbank lending rate for the same period.
[0124] A core aspect of this invention is that module 40 also includes a DRL experience-driven parameter correction unit. The input of this unit is connected to the experience playback pool of the transient collaborative control module 30. Establish data connections to dynamically and adaptively determine the two most critical parameters in the Black-Scholes model: execution price. and the volatility of the underlying asset .
[0125] Execution Price In the technical solution of this invention, this is defined as the opportunity cost or price incurred when the system is forced to utilize its flexibility resources for emergency transient support. The parameter correction unit calculates this value by performing the following steps:
[0126] First, from the experience replay pool Select all empirical samples that meet specific conditions. The condition is the action in the sample. Belongs to the preset set of transient support actions (For example, the action of energy storage urgently outputting active power to the grid). Then, the unit calculates the reward value for these selected samples. The expected value (or arithmetic mean) of the negative number is used as the execution price for the current scheduling cycle. This calculation is performed using the aforementioned formula. This is how it is accomplished. In this way, the execution price is no longer a static, subjectively set parameter, but dynamically reflects the actual "average cost" of the underlying physical control system in response to transient events.
[0127] underlying asset volatility In this scheme, the uncertainty or risk of the system's operating state is quantified. The parameter correction unit calculates the empirical playback pool. All recent reward signals The standard deviation is used to estimate this value. Drastic fluctuations in the reward signal indicate that the system frequently deviates from its steady state, i.e., high volatility. In this case, the value of options that retain flexibility resources also increases accordingly.
[0128] At the beginning of each scheduling cycle, an optimal scheduling decision unit will use dynamic parameters provided by the parameter correction unit. and and real-time acquisition Calculate the current option value. Subsequently, the unit... and Compare. If Greater than This indicates that the potential value of retaining this flexibility resource outweighs the benefits of using it immediately, and the decision-making unit will generate a more conservative economic operating power baseline. For example, the energy storage unit is instructed to maintain a high state of charge.
[0129] Conversely, if Greater than or equal to This indicates that using resources immediately is more economical, and the decision-making unit will generate a more aggressive power baseline. For example, it can instruct energy storage units to discharge in order to participate in market arbitrage. Ultimately, the module will generate an economic operating power baseline. Output to the hierarchical collaborative instruction execution module 50.
[0130] The hierarchical collaborative instruction execution module 50 is the final instruction output unit of this control system. It is responsible for integrating control instructions from different time scales and generating and issuing a unified and conflict-free final control instruction set to each physical execution unit of the system based on the current operating status of the system.
[0131] In one specific implementation, the input terminals of module 50 are connected to the output terminals of transient cooperative control module 30 and energy optimization management module 40, respectively. It receives two types of input signals: one is the economic operating power baseline generated by energy optimization management module 40 and updated on a minute-by-minute basis. Second, the transient cooperative action vector generated by the transient cooperative control module 30 and updated in milliseconds. .
[0132] This module 50 internally includes an instruction decomposition unit. For the input economic operating power baseline... The unit decomposes the power into steady-state active power reference values for each of the photovoltaic unit, flexible DC transmission unit, and energy storage unit according to a preset allocation strategy. For the input transient cooperative action vector This unit takes its components, such as and This is directly mapped to the dynamic adjustment of the active power reference value and virtual moment of inertia parameter of the energy storage unit.
[0133] After instruction decomposition, an instruction fusion and arbitration unit is responsible for merging the two types of instructions. The core of this unit is a dynamic priority switching mechanism based on the system's operating state. Specifically, this mechanism monitors the raw state vector from the system state awareness module 10. To determine the mode the system is in.
[0134] In stable system operation mode, where key indicators such as grid connection point voltage and frequency fluctuate within preset normal ranges, this unit will assign the highest execution priority to the economic operating power baseline. At this time, the final control commands issued will track the changes in power consumption. The power reference values of each unit obtained from the decomposition are the primary objective, while the transient cooperative action vector... Each component is set to zero or given a very small weight to ensure the economic efficiency of system operation.
[0135] When the system enters transient or early warning mode, that is, when it detects that the grid connection point voltage or frequency exceeds the safety threshold, or receives a high-risk early warning signal from the digital twin and forward-looking risk assessment module 20, the unit immediately switches priorities.
[0136] At this point, the transient cooperative action vector It is assigned the highest execution priority. The final control command is generated by superimposing the reference value obtained from the economic operating power baseline decomposition with the adjustment amount of the transient cooperative action vector. For example, the final active power command of the energy storage unit. It can be calculated using the following formula:
[0137] ;
[0138] In the formula, This is the final active power command issued to the energy storage unit; Based on the economic operating power baseline The reference value of the active power of the energy storage unit obtained by decomposition; For transient cooperative action vectors The provided adjustment amount for the active power reference value of the energy storage unit;
[0139] Meanwhile, other transient control parameters, such as virtual moment of inertia Also based on the transient cooperative action vector The corresponding components are updated in real time.
[0140] To ensure smooth switching between different operating modes, module 50 is also equipped with an instruction smoothing transition unit. When the system recovers from transient mode to stable mode, this unit does not immediately cancel the transient cooperative action vector. Instead of acting as a function, it uses a preset ramp function or time constant to adjust the transient regulation (such as...) over a certain period of time (e.g., several seconds). It smoothly decays to zero, while simultaneously allowing the control command to be smoothly handed over from transient control to the economic operating power baseline.
[0141] This mechanism avoids secondary system oscillations caused by sudden changes in control mode. Finally, the module sends the complete instruction set, after fusion, arbitration, and smoothing, to the underlying controllers of each unit (such as the drive boards of inverters or converters) for execution via digital or analog signal interfaces.
[0142] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system, characterized in that, include: The system state perception module is used to collect and integrate the operating data of each unit in the photovoltaic grid-connected flexible DC transmission and energy storage collaborative control system to generate the original state vector. The digital twin and forward-looking risk assessment module is connected to the system state perception module. It is used to receive the original state vector and perform ultra-real-time extrapolation in the digital twin model to calculate and output a forward-looking risk gradient vector that characterizes the future stability risk of the system. The transient collaborative control module is connected to the system state perception module and the digital twin and forward-looking risk assessment module. It is used to fuse the original state vector and the forward-looking risk gradient vector to form a risk-enhanced state vector, and generate differentiated transient collaborative action vectors based on deep reinforcement learning, while storing the interaction data in the experience playback pool. The energy optimization management module, connected to the transient collaborative control module, is used to formulate an economic operating power baseline based on the real options theory model and to dynamically correct the key parameters of the theoretical model using data from the experience replay pool. A hierarchical collaborative instruction execution module, connected to the transient collaborative control module and the energy optimization management module, is used to integrate and execute the transient collaborative action vector and the economic operating power baseline to generate and issue the final control instruction; The digital twin and forward-looking risk assessment module includes: The digital twin parallel extrapolation unit is used to extrapolate the future trajectories of multiple systems in parallel, after receiving the original state vector, in a built-in dynamic model composed of differential algebraic equations, combined with uncertainty disturbance prediction. The risk gradient calculation unit is used to calculate the sensitivity of the comprehensive stability index to current control changes based on the inference results of the digital twin parallel inference unit, so as to quantify the forward-looking risk gradient vector.
2. The photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to claim 1, characterized in that, The risk gradient calculation unit employs a risk gradient calculation method based on numerical difference approximation, and calculates the forward-looking risk gradient vector using the following formula. : ; In the formula, For the current moment The aforementioned forward-looking risk gradient vector; This is a pre-defined comprehensive stability index function used to evaluate the system's security level. For the current moment The control vector of the system; This is a preset virtual micro-perturbation applied to the control vector; Calculate the norm of a vector.
3. The photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to claim 1, characterized in that, The transient collaborative control module is specifically used for: The original state vector is combined with the forward-looking risk gradient vector to obtain the risk-enhanced state vector; The risk-enhanced state vector is input into a pre-trained deep reinforcement learning policy network to output the differentiated transient cooperative action vector in real time.
4. The photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to claim 3, characterized in that, The transient collaborative control module is specifically used for: The original state vector is combined with the forward-looking risk gradient vector to obtain the risk-enhanced state vector; The risk-enhanced state vector is input into a pre-trained deep reinforcement learning policy network to output the differentiated transient cooperative action vector in real time.
5. The photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to claim 1, characterized in that, The energy optimization management module adopts an optimization scheduling architecture with adaptive parameter correction, and the module specifically includes: The real options pricing unit is used to treat the system's flexibility resources as put options and calculate their option value using an option pricing model to assess the economic value of retaining the flexibility resources. The DRL experience-driven parameter correction unit is used to extract data from the experience replay pool and dynamically correct the key parameters of the option pricing model accordingly. The optimal scheduling decision unit is used to compare the instantaneous benefit of immediately using the flexibility resources with the option value, and select the strategy corresponding to the one with the highest value to generate the economic operating power baseline.
6. The photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to claim 5, characterized in that, The DRL experience-driven parameter correction unit employs a dynamic strike price calibration method, dynamically correcting the strike price of the option pricing model using the following formula. : ; In the formula, For the current moment The strike price of the option pricing model described above; An operator for calculating mathematical expectation; This is the experience playback pool; The state is the empirical data extracted from the empirical replay pool; Actions in the experience data extracted from the experience replay pool; The reward is the experience data extracted from the experience replay pool; A set of pre-defined actions used for transient support.
7. The photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to claim 1, characterized in that, The hierarchical collaborative instruction execution module employs a priority dynamic switching logic based on runtime status. Specifically, this module is used for: When the system is running stably, the final control commands should primarily track the economic operating power baseline. When a disturbance occurs in the system or a risk is anticipated, the transient cooperative action vector is given the highest execution priority, and its instructions are superimposed or overridden on the economic operating power baseline to ensure system safety.
8. The photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to claim 6, characterized in that, The input of the DRL experience-driven parameter correction unit establishes a data connection with the experience playback pool of the transient collaborative control module.
9. A method for coordinated control of photovoltaic grid-connected flexible DC transmission and energy storage, characterized in that, Using the photovoltaic grid-connected flexible DC transmission and energy storage coordinated control system according to any one of claims 1-8 includes the following steps: S1. State perception step: Collect and integrate the operating data of each unit in the photovoltaic grid-connected flexible DC transmission and energy storage collaborative control system to generate the original state vector; S2. Risk assessment step: Input the original state vector into the digital twin model for real-time simulation to calculate and output a forward-looking risk gradient vector that characterizes the future stability risk of the system. S3. Transient control step: The original state vector and the prospective risk gradient vector are fused to form a risk-enhanced state vector, and a differentiated transient cooperative action vector is generated based on deep reinforcement learning. At the same time, the interaction data is stored in the experience replay pool. S4. Energy Management Steps: Based on real options theory, establish an economic operating power baseline and use the data in the experience replay pool to dynamically correct the key parameters of the theoretical model; S5. Cooperative execution steps: Integrate and execute the transient cooperative action vector and the economic operating power baseline to generate and issue the final control command.
Citation Information
Patent Citations
Regional resource collaborative optimization configuration method, device, equipment and medium
CN119005400A
Data processing method and device and electronic equipment
CN119126961A