An online optimization control strategy for a flexible DC system with multiple landing points in the receiving end area
Patent Information
- Application Number
- CN202311678707.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-08
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2043-12-08
AI Technical Summary
[0009]本发明提供了一种面向受端分区多落点柔性直流系统在线优化控制策略,通过受端分区内多落点柔性直流的功率协调控制技术,实现分区内功率支援与紧急支撑,有效解决了分区内部功率不平衡的问题,有效提高分区安全稳定性,增加资源利用效率
Smart Images

Figure CN117674190B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power technology, and in particular relates to an online optimization control strategy for a flexible DC system with multiple landing points in the receiving end area. Background Technology
[0002] For active power coordination optimization from day-ahead planning to intraday rolling timescales, domestic and international research mainly focuses on economic dispatch and unit combination models that consider power transmission constraints. Research has been conducted from aspects such as decentralized coordination between multiple regions, modeling of uncertainties in renewable energy output, and analytical expression of safety and stability constraints. The aim of multi-regional active power coordination optimization is to adapt to decentralized control modes and reduce the scale of the problem solution.
[0003] 1) The first type of method focuses on reducing the number of iterations between two levels of scheduling. Based on multi-parameter linear programming or simplified models within partitions, it obtains feasible tie-line power optimization schemes with a small number of iterations or non-iterations.
[0004] 2) The second type of method uses methods such as mixed integer programming value function to explore active co-optimization algorithms that are adapted to non-convex discrete models, but such algorithms are usually more complex.
[0005] 3) The third type of method focuses on modeling the power characteristics of DC tie lines, which improves the economic efficiency of interconnected grid operation while ensuring the feasibility of DC power operation. However, the model and method are only applicable to two regional grids.
[0006] Besides classification and time-scale refinement of reserves, existing research mainly explores uncertainty optimization theory and models, including stochastic optimization, robust optimization, and partial robust optimization. In addition to traditional unit combination and economic dispatch, uncertainty optimization is also widely used in multi-regional coordinated dispatch models and rolling dispatch models. While uncertainty optimization models can guarantee reasonable reserve retention and real-time availability, the computational cost of solving such models is usually high. The analytical expression of safety and stability constraints aims to establish the algebraic relationship between the state variables and operational decision variables of the system transient process, realizing the embedded expression of system stability criteria in the optimization model. Taking frequency stability constraints as an example, piecewise linearization and average frequency change rate estimation methods can be used to establish an approximate linear relationship between the frequency extremum point and the unit output level before the disturbance in a multi-machine system under large disturbances. There are also related studies on the analytical correlation expression between frequency stability within a two-region power system and the inter-regional power exchange, but current methods are not applicable to multi-regional systems.
[0007] In the primary and secondary frequency regulation timescales, utilizing DC to provide auxiliary services such as frequency security and stability for the power grid is a key means of operation and control for asynchronous interconnected power grids, and has also become a technological development trend. Theoretical research on DC participation in power grid frequency control mainly focuses on the emergency frequency control and primary frequency regulation stages. While there is existing research on the characteristic analysis of different types of DC frequency converters, coordination with unit primary frequency regulation, and coordination of DC frequency regulation strategies for asynchronous power grids, a systematic theoretical foundation and guiding methods are still lacking for frequency coordination and optimization control of multi-regional asynchronous power grids.
[0008] In summary, current research on active power optimization in multi-asynchronous zoned power grids still lacks efficient modeling theories and algorithms for multi-zone coordination; similarly, research on frequency coordination control in multi-asynchronous zoned power grids lacks methods for coordinating multiple DC frequency regulation control strategies and coordinating DC and zone frequency control strategies. Therefore, research on active power coordination control methods for high-proportion renewable energy interconnected asynchronous zoned power grids is urgently needed to address the challenges of safe and economical operation of interconnected asynchronous zoned power systems. Summary of the Invention
[0009] This invention provides an online optimization control strategy for multi-point flexible DC systems in receiving-end zones. By using power coordination control technology for multi-point flexible DC systems in receiving-end zones, it achieves power support and emergency support within the zones, effectively solving the problem of power imbalance within the zones, effectively improving the safety and stability of the zones, and increasing resource utilization efficiency.
[0010] An online optimization control strategy for a flexible DC system with multiple landing points at the receiving end includes:
[0011] (1) Determine the form of power supply for the receiving end system, including centralized wind power generation at the sending end, centralized photovoltaic power generation at the sending end, conventional generator power generation at the sending end, and local conventional generator power generation. All energy at the sending end is transmitted to the receiving end through flexible DC digital power technology.
[0012] (2) Determine the optimization objective and constraints of the system control; the optimization objective is as follows:
[0013]
[0014] Among them, w n This represents the weight corresponding to the nth sub-objective function. This represents the load rate of line j at time t. This represents the reference value for the load rate of line j. This indicates the voltage compliance of bus i at time t. N represents the voltage compliance reference value for bus i; line N bus Let T represent the set of lines and the set of buses, respectively, and T represent the total number of moments in each optimization cycle. This represents the short-circuit current assessment value of bus i at time t. This represents the maximum short-circuit current limit of bus i;
[0015] The constraints include weight allocation constraints, real-time active power balance constraints, bus voltage over-limit constraints, line load rate constraints, spinning reserve capacity constraints, and bus short-circuit current constraints.
[0016] (3) Construct a Safe-SAC reinforcement learning agent to solve the above optimization objectives; wherein, the Safe-SAC reinforcement learning agent includes an Actor network, a Critic network, and a Safe network.
[0017] The Actor network is based on the current state s t Perform the action 'a' at the current moment. t ; The Critic network adjusts state-action pairs (s) t ,a t Output the corresponding Q value Q. θ ;Safe network acts on state (s) t ,a t Output the corresponding Q value. Furthermore, Safe networks minimize To update its neural network to the target, the Actor maximizes... Critic updates its neural network to minimize Q. θ With Q target The difference is used as the target to update its neural network, Q. target This represents the target Q-value of the neural network.
[0018] (4) Set the reward function for the Safe-SAC reinforcement learning agent during training. During training, the Actor network, Critic network and Safe network are all updated with gradients according to the loss function.
[0019] (5) Use the trained Safe-SAC reinforcement learning agent for the control of flexible DC systems with multiple landing points in the receiving end partition.
[0020] In step (2), the weight allocation constraint is:
[0021]
[0022] In the formula, N w This indicates the number of sub-objective functions.
[0023] The real-time active power balance constraint is:
[0024]
[0025] In the formula, This represents the active power generated by the local conventional generating units in the receiving-end partition at time t. This represents the active power input from the sending-end section at bus i to the receiving-end section via VSC at time t. This represents the active power input to the receiving-end zone by the centralized photovoltaic power station at bus i via VSC at time t. This represents the active power fed into the receiving-end section by the centralized wind power station at bus i via VSC at time t. and Let represent the discharge and charging power emitted by the receiving-end zoned energy storage system at bus i at time t, respectively. This represents the active power required by the load at section bus i at time t. N represents the network loss of the receiving end system at time t; load N represents the set of busbars where the equivalent load is located. LG N represents the set of feed buses for conventional units at the receiving end. ES N represents the bus cluster where the receiving-end energy storage station is located. VG N represents the set of feed buses for the conventional sending-end partition. VP N represents the collection of feed buses for a centralized photovoltaic power station. VW This refers to the collection of feed busbars for centralized wind power generation stations.
[0026] Bus voltage over-limit constraint is:
[0027]
[0028]
[0029] In the formula, V represents the effective voltage value of bus i at time t. i bus,ref This represents the reference voltage value for bus i. Indicates the compliance of bus voltage. This indicates the degree of voltage compliance corresponding to the voltage reference value of bus i. This indicates the maximum deviation in bus voltage compliance.
[0030] The line load factor constraint is:
[0031]
[0032] In the formula, This represents the load rate of line j at time t. This represents the reference value for the load rate of line j. This indicates the maximum deviation of the load rate of line j.
[0033] The spinning reserve capacity constraint is:
[0034]
[0035] In the formula, λ r This is the rotating reserve capacity factor for conventional units. The short-circuit current constraint for the bus is:
[0036]
[0037] In the formula, This represents the short-circuit current assessment value of bus i at time t. This represents the maximum short-circuit current limit of bus i.
[0038] In step (4), the Actor network, Critic network, and Safe network are all updated with gradients based on the loss function, as follows:
[0039] J θ,t =Q θ (s t ,a t )-[r(s t ,a t )+V r (s t+1 ,a t )]
[0040]
[0041] J π,t =Q π (s t ,a t )+λQ d (s t ,a t )-αlogπ(a t )
[0042] In the formula, π, θ, d(s) represents the parameters of the Actor, Critic, and Safety neural networks in the Safe-SAC algorithm, respectively; J represents the loss function of the neural network; and Q represents the Q-value of the neural network output. Compared to the traditional SAC algorithm, Safe-SAC requires additional gradient updates for the Safe network, d(s) t ,a t ) represents a state-action pair (s) t ,a t The immediate security reward obtained under ) is related to r(s) t ,a t Similar to; V d (s t+1 ,a t) indicates taking action a t Then, the expected value of the security rewards that can be obtained in the future, and its relationship with V. r (s t+1 ,a t Similar to r(s); t ,a t ) indicates that the agent is taking action a. t Arrival in state s t The instant reward obtained afterward, V r (s t+1 ,a t ) indicates that the agent is taking action a. t Arrival in state s t+1 The expected value of the reward received, Q d (s t ,a t ) indicates the action pair (s) t ,a t The safe Q value under () is λ, where λ represents the safe Q value coefficient and α represents the reward decay coefficient.
[0043] In step (5), during the control process, control commands corresponding to the control actions are output according to the power grid status;
[0044] The power grid status is as follows: in, This represents the active power required by the load at bus i in the receiving end section at time t. This represents the effective voltage value of bus i at time t. This represents the load rate of line j at time t;
[0045] The control commands are: in, This represents the active power generated by the local conventional generating units in the receiving-end partition at time t. This represents the active power input from the sending-end section at bus i to the receiving-end section via VSC at time t. This represents the active power input to the receiving-end zone by the centralized photovoltaic power station at bus i via VSC at time t. This represents the active power fed into the receiving-end section by the centralized wind power station at bus i via VSC at time t. and These represent the discharge and charging power emitted by the receiving-end zone energy storage system at bus i at time t, respectively.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] 1. This invention considers the power supply brought by flexible DC transmission to the receiving-end regional power grid, realizes power balance within the receiving-end regional grid, balances the line load rate within the receiving-end regional grid, optimizes the bus voltage level, and improves the spinning reserve capacity of the boosting system, thereby enhancing the emergency power support capability within the regional grid.
[0048] 2. This invention proposes a reinforcement learning control algorithm based on Safe-SAC to solve this control problem. This algorithm is an improvement on the Soft Actor-Critic (SAC) algorithm and is used to solve the safety problem in reinforcement learning. Its characteristic is that during task execution, it can ensure that the agent's policy will not lead to unsafe behavior or unstable state, improving the safety and stability of the receiving end partition while ensuring solution efficiency. Attached Figure Description
[0049] Figure 1 This is a schematic diagram of the power supply of the receiving-end system in an embodiment of the present invention;
[0050] Figure 2 This is a schematic diagram of the training architecture of the Safe-SAC reinforcement learning agent in an embodiment of the present invention. Detailed Implementation
[0051] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that the embodiments described below are intended to facilitate the understanding of the present invention and do not constitute any limitation thereof.
[0052] A schematic diagram of a receiving-end system with multiple flexible direct feeds is shown below. Figure 1 As shown, the power supply of the receiving-end system mainly takes four forms: a) centralized wind power generation at the sending end, b) centralized photovoltaic power generation at the sending end, c) conventional power generation at the sending end, and d) local conventional power generation. All energy from the sending end is transmitted to the receiving end via flexible DC digital power technology, and all receiving-end converters are voltage source converters (VSC).
[0053] The optimization objective of the control problem is as follows:
[0054]
[0055] Among them, w n This represents the weight corresponding to the nth sub-objective function. This represents the load rate of line j at time t. This represents the reference value for the load rate of line j. This indicates the voltage compliance of bus i at time t. N represents the voltage compliance reference value for bus i; line N bus Let T represent the set of lines and the set of buses, respectively, and T represent the total number of moments in each optimization cycle. This represents the short-circuit current assessment value of bus i at time t. This represents the maximum short-circuit current limit of bus i.
[0056] The constraints involved in this invention mainly include: a) weight allocation constraints, b) real-time active power balance constraints, c) bus voltage over-limit constraints, d) line load rate constraints, e) spinning reserve capacity constraints, and f) short-circuit current constraints.
[0057] a) Weight allocation constraints
[0058]
[0059] In the formula, N w This indicates the number of sub-objective functions.
[0060] b) Real-time active power balance constraint
[0061]
[0062] In the formula, This represents the active power generated by the local conventional generating units in the receiving-end partition at time t. This represents the active power input from the sending-end section at bus i to the receiving-end section via VSC at time t. This represents the active power input to the receiving-end zone by the centralized photovoltaic power station at bus i via VSC at time t. This represents the active power fed into the receiving-end section by the centralized wind power station at bus i via VSC at time t. and Let represent the discharge and charging power emitted by the receiving-end zoned energy storage system at bus i at time t, respectively. P represents the active power required by the load at the receiving-end bus i at time t. t loss N represents the network loss of the receiving end system at time t; load N represents the set of busbars where the equivalent load is located. LG N represents the set of feed buses for conventional units at the receiving end. ES N represents the bus cluster where the receiving-end energy storage station is located. VG N represents the set of feed buses for the conventional sending-end partition. VP N represents the collection of feed buses for a centralized photovoltaic power station. VW This refers to the collection of feed busbars for centralized wind power generation stations.
[0063] For conventional generator sets at the receiving end, their power constraints and ramping constraints can be expressed as follows:
[0064]
[0065]
[0066] For non-new energy power plants that centrally send out flexible direct transmission, the default maximum transmission capacity of the flexible direct transmission does not change over time.
[0067]
[0068] Considering that centralized renewable energy power plants may benefit multiple receiving-end systems, and the real-time allocation limit of their transmitted power is determined by local dispatch decisions, there is a renewable energy output limit constraint for a single receiving-end system. In this case, the active power output reference value of the receiving-end VSC cannot exceed this limit, i.e.
[0069]
[0070]
[0071] The energy storage constraints of the receiving-end system are as follows:
[0072]
[0073]
[0074]
[0075]
[0076]
[0077] In the formula, equations (9) to (11) represent the active power constraints for charging and discharging of the energy storage system at bus i, E i,t P represents the capacity of the energy storage system at bus i at time t. i ES,max This represents the maximum power of the energy storage system at bus i. The SOC represents the maximum capacity of the energy storage system at bus i. i,t This represents the state of charge of the energy storage system at bus i at time t. These represent the minimum and maximum values of the state of charge of the energy storage system at bus i, respectively.
[0078] c) Bus voltage over-limit constraint
[0079]
[0080]
[0081] In the formula, This represents the effective voltage value of bus i at time t. This represents the reference voltage value for bus i. Indicates the compliance of bus voltage. This indicates the degree of voltage compliance corresponding to the voltage reference value of bus i. This indicates the maximum deviation in bus voltage compliance.
[0082] d) Line load factor constraints
[0083]
[0084] In the formula, This represents the load rate of line j at time t. This represents the reference value for the load rate of line j. This indicates the maximum deviation of the load rate of line j.
[0085] e) Spinning Reserve Capacity Constraints
[0086] To provide emergency power support within the receiving-end zone, the spinning reserve capacity of local conventional units must be considered. The receiving-end system must possess a certain emergency support capability, i.e.
[0087]
[0088] In the formula, λ r This is the rotating reserve capacity factor for conventional units.
[0089] f) Short-circuit current constraint
[0090] Factors such as the power flow distribution of the receiving-end system, the load conditions of the lines, and the output distribution of the generating units all affect the short-circuit current level of the bus. The short-circuit current constraint of the bus is set as follows:
[0091]
[0092] This invention proposes to solve the aforementioned optimization problem using the Safe-SAC algorithm. Safe-SAC (Safe Soft Actor-Critic) is an enhanced Soft Actor-Critic (SAC) reinforcement learning algorithm used to address the safety problem in reinforcement learning. The goal of Safe-SAC is to ensure that the agent's policy does not lead to unsafe behavior or unstable states while performing tasks. It introduces a safety constraint to limit the agent's behavior, preventing it from performing dangerous actions or falling into unstable states.
[0093] The traditional SAC algorithm consists of two core parts: the Actor and the Critic. The Actor determines the current state s based on the Critic. t Perform the action 'a' at the current moment. t ;Critic acts on (s) based on state. t ,a tOutput the corresponding Q value Q. θ Subsequently, the Actor maximizes Q. θ Critic updates its neural network to minimize Q. θ With Q target The difference is used as the target to update its own neural network.
[0094] In addition to the Actor and Critic components, the Safe-SAC algorithm adds a safety-oriented Safe network. The Safe network is based on the state-action pair (s... t ,a t Output the corresponding Q value. And minimize The Actor updates its neural network to target the desired outcome. The Actor's update method then becomes maximization. The Critic part is consistent with the SAC algorithm. It is precisely because the Actor update considers the safety characteristics of Safe networks that Safe-SAC makes safer and more reliable decisions compared to traditional SAC. The training architecture of the Safe-SAC reinforcement learning agent proposed in this invention is as follows: Figure 2 As shown.
[0095] This invention selects line load rate and bus voltage level as the reward function for the Critic module, both of which are generally considered routine indicators of steady-state operation of the power system. For the Safety module, the magnitude of the short-circuit current of the power system bus is considered a safety indicator, so the difference between the short-circuit current and its allowable value is selected as the reward function. When a short circuit occurs, it is generally desirable for the short-circuit current to be as small as possible below its maximum value. The reward functions for both the Critic and Safe modules can be expressed in the following form:
[0096]
[0097] In the formula, M i The weight w in the optimization objective represents the weight w. i The calculated value of the sub-objective function, Describing the sub-objective function M i The maximum allowed value, A i,1 A i,2 A i,3 These represent the reward values under the three scenarios.
[0098] All three types of neural networks mentioned above require gradient updates based on the loss function, as detailed below:
[0099] J θ,t =Q θ (s t ,a t)-[r(s t ,a t )+V r (s t+1 ,a t (20)
[0100]
[0101] J π,t =Q π (s t ,a t )+λQ d (s t ,a t )-αlogπ(a t ) (twenty two)
[0102] In the formula, π, θ, d(s) represents the parameters of the Actor, Critic, and Safety neural networks in the Safe-SAC algorithm, respectively; J represents the loss function of the neural network; and Q represents the Q-value of the neural network output. Compared to the traditional SAC algorithm, Safe-SAC requires additional gradient updates for the Safe network, d(s) t ,a t ) represents a state-action pair (s) t ,a t The "security reward" immediately obtained under ) is related to r(s) t ,a t Similar to; V d (s t+1 ,a t ) indicates taking action a t Then, the expected value of the security rewards that can be obtained in the future, and its relationship with V. r (s t+1 ,a t )similar.
[0103] The embodiments described above provide a detailed explanation of the technical solutions and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An online optimization control strategy for a flexible DC system with multiple landing points in a receiving-end partition, characterized in that, include: (1) Determine the form of power supply for the receiving end system, including centralized wind power generation at the sending end, centralized photovoltaic power generation at the sending end, conventional unit power generation at the sending end, and local conventional unit power generation. All energy at the sending end is transmitted to the receiving end through flexible DC transmission technology. (2) Determine the optimization objective and constraints of the system control; the optimization objective is as follows: Among them, w n This represents the weight corresponding to the nth sub-objective function. This represents the load rate of line j at time t. This represents the reference value for the load rate of line j. This indicates the voltage compliance of bus i at time t. N represents the voltage compliance reference value for bus i; line N bus Let T represent the set of lines and the set of buses, respectively, and let T represent the total number of moments in each optimization cycle. This represents the short-circuit current assessment value of bus i at time t. This represents the maximum short-circuit current limit of bus i; The constraints include weight allocation constraints, real-time active power balance constraints, bus voltage over-limit constraints, line load rate constraints, spinning reserve capacity constraints, and bus short-circuit current constraints. (3) Construct a Safe-SAC reinforcement learning agent to solve the above optimization objectives; wherein, the Safe-SAC reinforcement learning agent includes an Actor network, a Critic network and a Safe network; The Actor network is based on the current state s t Perform the action 'a' at the current moment. t ; The Critic network adjusts state-action pairs (s) t ,a t Output the corresponding Q value Q. θ ;Safe network acts on state (s) t ,a t Output the corresponding Q value. Furthermore, Safe networks minimize To update its neural network to the target, the Actor maximizes... Critic updates its neural network to minimize Q. θ With Q target The difference is used as the target to update its neural network, Q. target This represents the target Q-value of the neural network; (4) Set the reward function for the Safe-SAC reinforcement learning agent during training. During training, the Actor network, Critic network and Safe network are all updated with gradients according to the loss function. (5) Use the trained Safe-SAC reinforcement learning agent for the control of flexible DC systems with multiple landing points in the receiving end partition.
2. The online optimization control strategy for a multi-point flexible DC system with regionalized receiving ends as described in claim 1, characterized in that, In step (2), the weight allocation constraint is: In the formula, N w This indicates the number of sub-objective functions.
3. The online optimization control strategy for a multi-point flexible DC system with partitioned receiving end as described in claim 2, characterized in that, In step (2), the real-time active power balance constraint is: In the formula, This represents the active power generated by the local conventional generating units in the receiving-end partition at time t. This represents the active power input from the sending-end section at bus i to the receiving-end section via VSC at time t. This represents the active power input to the receiving-end zone by the centralized photovoltaic power station at bus i via VSC at time t. This represents the active power fed into the receiving-end section by the centralized wind power station at bus i via VSC at time t. and Let represent the discharge and charging power emitted by the receiving-end zoned energy storage system at bus i at time t, respectively. P represents the active power required by the load at the receiving-end bus i at time t. t loss N represents the network loss of the receiving end system at time t; load N represents the set of busbars where the equivalent load is located. LG N represents the set of feed buses for conventional units at the receiving end. ES N represents the bus cluster where the receiving-end energy storage station is located. VG N represents the set of feed buses for the conventional sending-end partitions. VP N represents the collection of feed buses for a centralized photovoltaic power station. VW This refers to the collection of feed busbars for centralized wind power generation stations.
4. The online optimization control strategy for a multi-point flexible DC system with receiving-end partitioning according to claim 3, characterized in that, In step (2), the bus voltage over-limit constraint is: In the formula, V represents the effective voltage value of bus i at time t. i bus,ref This represents the reference voltage value for bus i. Indicates the compliance of the bus voltage. This indicates the degree of voltage compliance corresponding to the voltage reference value of bus i. This indicates the maximum deviation in bus voltage compliance.
5. The online optimization control strategy for a multi-point flexible DC system with partitioned receiving end as described in claim 4, characterized in that, In step (2), the line load rate constraint is: In the formula, This represents the load rate of line j at time t. This represents the reference value for the load rate of line j. This indicates the maximum deviation of the load rate of line j.
6. The online optimization control strategy for a multi-point flexible DC system with partitioned receiving end as described in claim 5, characterized in that, In step (2), the rotating reserve capacity constraint is: In the formula, λ r This is the rotating reserve capacity factor for conventional units.
7. The online optimization control strategy for a multi-point flexible DC system with partitioned receiving end as described in claim 6, characterized in that, In step (2), the short-circuit current constraint of the busbar is: In the formula, This represents the short-circuit current assessment value of bus i at time t. This represents the maximum short-circuit current limit of bus i.
8. The online optimization control strategy for a multi-point flexible DC system with partitioned receiving end as described in claim 1, characterized in that, In step (4), the Actor network, Critic network, and Safe network are all updated with gradients based on the loss function, as follows: J θ,t =Q θ (s t ,a t )-[r(s t ,a t )+V r (s t+1 ,a t )] J π,t =Q π (s t ,a t )+λQ d (s t ,a t )-αlogπ(a t ) In the formula, π, θ, d(s) represents the parameters of the Actor, Critic, and Safety neural networks in the Safe-SAC algorithm, respectively; J represents the loss function of the neural network; Q represents the Q-value of the neural network output; compared to the traditional SAC algorithm, Safe-SAC requires additional gradient updates for the Safe network, d(s) t ,a t ) represents a state-action pair (s) t ,a t The immediate security reward obtained under ) is related to r(s) t ,a t Similar to; V d (s t+1 ,a t ) indicates taking action a t Then, the expected value of the security rewards that can be obtained in the future, and its relationship with V. r (s t+1 ,a t Similar to r(s); t ,a t ) indicates that the agent is taking action a. t Arrival state s t The instant reward obtained afterward, V r (s t+1 ,a t ) indicates that the agent is taking action a. t Arrival in state s t+1 The expected value of the reward received, Q d (s t ,a t ) represents a state-action pair (s) t ,a t The safe Q value under () is λ, where λ represents the safe Q value coefficient and α represents the reward decay coefficient.
9. The online optimization control strategy for a multi-point flexible DC system with partitioned receiving end as described in claim 1, characterized in that, In step (5), during the control process, control commands corresponding to the control actions are output according to the power grid status; The power grid status is as follows: in, This represents the active power required by the load at bus i in the receiving end section at time t. This represents the effective voltage value of bus i at time t. This represents the load rate of line j at time t; The control commands are: in, This represents the active power generated by the local conventional generating units in the receiving-end partition at time t. This represents the active power input from the sending-end section at bus i to the receiving-end section via VSC at time t. This represents the active power input to the receiving-end zone by the centralized photovoltaic power station at bus i via VSC at time t. This represents the active power fed into the receiving-end section by the centralized wind power station at bus i via VSC at time t. and These represent the discharge and charging power emitted by the receiving-end zone energy storage system at bus i at time t, respectively.
Citation Information
Patent Citations
Multi-device cooperative broadband oscillation suppression method for offshore wind power flexible direct current grid-connected system
CN114512995A
Power system operation mode intelligent generation method based on deep reinforcement learning
CN115912367A