Response-driven load shedding control method for receiving-end power grid considering transient voltage recovery requirement

By employing the Double DQN adaptive load shedding control method and utilizing a Markov sequential model and agent training, the problem of transient voltage recovery requirements of the receiving-end power grid was solved, achieving the adaptability and voltage recovery effect of distributed load shedding control.

CN119726757BActive Publication Date: 2026-02-10NORTHEAST DIANLI UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411767844.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2026-02-10
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

Existing load shedding control schemes are difficult to effectively respond to transient voltage recovery needs in receiving-end power grids, leading to voltage instability. Furthermore, distributed control strategies lack adaptability and cannot meet the TVRC requirements of IEEE 1547-2020.

Method used

An adaptive load shedding control method based on Double DQN is adopted. Transient voltage control is modeled using a Markov sequential model. Load shedding agents and time-delay agents are designed and trained offline using a reward function. The methods are then deployed at various load sites to achieve distributed load shedding control and adaptively determine the location, amount, and timing of load shedding.

Benefits of technology

It enables real-time adaptive decision-making based on the severity of transient voltage in the receiving-end power grid, meets the load shedding control requirements of TVRC, reduces the cost of load shedding, and improves the voltage recovery capability and stability of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119726757B_ABST
    Figure CN119726757B_ABST
Patent Text Reader

Abstract

The application provides a kind of load shedding control method of receiving end power grid considering transient voltage recovery demand, belongs to response driven type load shedding control, solve the technical problem of insufficient system voltage regulation capability caused by large-scale access of new energy to receiving end power grid and large number of replacement of traditional fossil energy synchronous unit.Its technical scheme is as follows: containing the following steps: (1) establish the Markov sequential decision process model of transient voltage control;(2) design the state, action and reward in Markov process;(3) establish the reinforcement learning model based on Double DQN;(4) carry out off-line centralized learning and training of corresponding agent;(5) form decentralized adaptive load shedding control method.The beneficial results of the application are as follows: the load shedding controller can be deployed at each load site, and the closed-loop decentralized load shedding control strategy of "monitoring-determination-decision-control" can be realized to promote the overall voltage recovery of power grid.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of response-driven load shedding control technology, and in particular to a receiving-end grid response-driven load shedding control method that takes into account transient voltage recovery requirements. Background Technology

[0002] Transient voltage stability is one of the key issues affecting the safe operation of the receiving-end power grid in large load centers. National economic development has led to a continuous increase in electricity load, especially in induction motor loads such as commercial and residential air conditioning and industrial loads, which account for over 60% of the total load. During voltage dips, induction motor loads decelerate or even stall, resulting in a significant increase in dynamic reactive power consumption. This easily leads to positive feedback voltage deterioration, causing transient voltage instability and even triggering major power outages. To prevent induction motor loads from stalling and disconnecting from the grid after a fault, and to avoid transient voltage safety breaches, the North American Electric Reliability Council (NEC) has specified a segmented transient voltage recovery criterion (TVRC) based on the "voltage threshold-recovery time" in its IEEE 1547-2020 reliability guidelines.

[0003] Currently, the receiving-end power grid has phased out a large number of local fossil fuel synchronous generator units due to the acceptance of clean energy feed-in from distant sources, resulting in severe hollowing out of load centers and a significant decrease in the voltage support capacity of load center power sources. Given the insufficient voltage support capacity and the scarcity of dynamic reactive power compensation equipment, load shedding control has become a necessary measure to prevent significant voltage drops and avoid transient voltage instability. Existing load shedding control schemes are broadly divided into event-driven and response-driven modes. Response-driven load shedding schemes do not require specific matching of operating scenarios and faults; instead, they provide load shedding decisions based on the dynamic voltage change process to promote voltage recovery. Response-driven load shedding schemes can be further divided into centralized and decentralized types. Decentralized load shedding control schemes only utilize locally monitored voltage information to make on-site load shedding decisions, making them easier to implement under current engineering conditions compared to centralized schemes. Summary of the Invention

[0004] The purpose of this invention is to provide a receiving-end grid response-driven load shedding control method that considers transient voltage recovery requirements, proposing an adaptive load shedding transient voltage control strategy for the receiving-end grid based on Double DQN response-driven approach. First, a Markov sequential model for transient voltage load shedding control is established. Transient voltage evolution trajectory information extracted from the grid using TVRC is used as the state, and the load shedding amount and time delay between rounds are selected as actions. A reward function is designed based on the control effect. Second, load shedding agents and time delay agents based on DQN representation are designed according to the actions, forming a Double DQN training structure. The agents are trained offline using a state-action-reward time-series dataset generated by the Markov model during interaction with the environment, learning the optimal load shedding control strategy. Then, the trained agents are deployed at each bus in the system, forming a distributed load shedding control strategy. This strategy can adaptively decide the load shedding location, amount, time, and rounds in real time based on the severity of the current transient voltage, promoting voltage recovery to meet TVRC requirements with a more efficient load shedding cost.

[0005] The inventive concept of this invention is as follows: Transient voltage control is modeled as a Markov sequential process. The severity of the transient voltage is extracted as the environmental state variable, and the load reduction amount and the time delay between cycles are selected as actions. A reward function is designed based on the control effect. Secondly, load reduction agents and delay agents based on DQN representation are designed according to the actions, forming a Double DQN training structure. The agents are trained offline using a state-action-reward time-series dataset generated based on the Markov model during interaction with the environment, learning the optimal load reduction control strategy. Then, the trained agents are deployed at each bus in the system, forming a distributed load reduction control strategy. This strategy can adaptively decide the load reduction location, amount, time, and cycle in real time based on the current transient voltage severity, realizing a closed-loop distributed load reduction control strategy of "monitoring-judgment-decision-control" to promote the overall voltage recovery of the power grid.

[0006] To achieve the aforementioned objectives, the present invention employs the following technical solution: a receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements, comprising the following steps:

[0007] (1) The transient voltage control is modeled as a Markov sequential decision process model;

[0008] (2) Based on the transient voltage recovery requirement (TVRC) of the power grid and the real-time voltage evolution trajectory of the load node, the environmental state variables in the Markov sequential decision-making process are extracted;

[0009] (3) Based on the factors affecting voltage recovery, design the load reduction amount and the time delay between rounds as actions in the Markov sequential decision-making process;

[0010] (4) Based on the action design, load reduction agents and time delay agents were designed; and a differentiated reward function was designed;

[0011] (5) The agent uses the state-action-reward dataset obtained from the Markov model during the interaction with the environment for centralized offline training.

[0012] (6) The trained load shedding agents and time delay agents are deployed at load sites to form a distributed load shedding control structure. They can adaptively decide the load shedding location, load shedding amount and load shedding time, so as to promote voltage recovery with a smaller load shedding amount and meet the transient voltage recovery requirements.

[0013] The specific implementation method of step (1) consists of (2), (3) and (4).

[0014] As a further optimization scheme of the receiving-end grid response-driven load reduction control method considering transient voltage recovery requirements provided by the present invention, in step (2), environmental state quantities are extracted;

[0015] Based on the TVRC requirements of the power grid and the real-time voltage evolution trajectory of the load nodes, the negative voltage amplitude deviation ΔV is extracted. m and recovery time deviation ΔT r Voltage deviation ΔV m Represented as:

[0016] ΔV m (t k ) = V th -V m (t k (1)

[0017] Where: ΔV m (t k ) indicates the current load node at time t k Voltage deviation at time, V th Indicates that at the current t k The voltage threshold that needs to be restored to at different times according to TVRC requirements, V m (t k ) represents the current load node at time t k The voltage value at that moment;

[0018] Voltage recovery time deviation ΔT r Calculated by the following formula:

[0019]

[0020] ΔT r (t k )=min{t k +T es (tk )-t c -T req T r,max} (3)

[0021] In equation (2): T es (t k ) for in t k The equation represents the estimated time required for the voltage to recover to the corresponding safe voltage threshold at time t, where Δt is the time step. This equation indicates that at time t... k At any given moment, if the voltage shows an upward trend, T es (t k If T is a finite positive value, and the voltage shows a decreasing trend, then T es (t k The value of ΔT is positive infinity, indicating that the estimated voltage trajectory is difficult to recover. In equation (3): ΔT r (t k ) for t k Time recovery time deviation, t k Indicates the current time, t c T represents the time when the short-circuit fault is cleared. req This indicates that the voltage has recovered to V after the fault. th Maximum allowable recovery time, T r,max Indicates ΔT r The maximum value of can be taken as the time scale of the transient voltage recovery problem of interest. It should be noted that TVRC has piecewise characteristics, and ΔV varies at different times during the voltage recovery process. m and ΔT r It should be based on V at the corresponding time. th and T req To calculate.

[0022] ΔV extracted based on TVRC requirements m (t k ) and ΔT r (t k This can be used to characterize the severity of the current voltage evolution state. The necessary condition for triggering load shedding control is: ΔV m (t k )>0 and ΔT r (t k )>0. When a load node bus needs to perform load reduction control, ΔV is extracted based on the voltage trajectory evolution characteristics of each load bus. m and ΔT r Two-dimensional state as environmental state quantity s t :

[0023] s t ={ΔV m (t),ΔT r (t)} (4)

[0024] As a further optimization scheme of the receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements provided by the present invention, in step (3), based on the factors affecting voltage recovery, the load shedding amount and the time delay between rounds are designed as actions in the Markov sequential decision-making process, including:

[0025] The magnitude of the load shedding in each round and the time delay between rounds (affecting the timing of the load shedding action) both influence the transient voltage load shedding control effect. Furthermore, considering the discrete nature of the substation load outgoing lines, the load shedding amount is set as a discrete quantity. The magnitude of the load shedding amount and the time delay between rounds are selected as decision actions a, respectively. Ct and a Tt Its corresponding discrete action space A Cd and A Td They are represented as follows:

[0026] A Cd ={ΔP d1 ,ΔP d2 ,…,Δp dm} (5)

[0027] A Td ={ΔT d1 ,ΔT d2 ,…,ΔT dn} (6)

[0028] Where: ΔP d1 ,ΔP d2 ,…,ΔP dm ΔT represents the action sequence consisting of the percentage of load reduction at load stations. d1 ,ΔT d2 ,…,ΔT dn This represents the sequence of actions consisting of the time intervals between load shedding rounds, where m and n represent the action space A, respectively. Cd and A Td Dimensions.

[0029] Two agents, a load-reduction agent and a time-delay agent, were designed to select load-reduction actions. These agents can respond differently to the environment and achieve optimal performance through training. The load-reduction agent determines the amount of load reduction based on environmental changes, while the time-delay agent determines the time interval between each load-reduction action. The agents select actions and respond using a value estimation network. The value estimation network Q for both the load-reduction agent and the time-delay agent is described. C and Q T This is used to establish a state-action value mapping relationship. The value estimation network is a multi-layer neural network, and the input is the transient voltage evolution state ΔV of the load bus extracted based on TVRC. m and ΔT rThe two-dimensional state information is fed forward by multiple hidden layer neurons to obtain the output value. The output value is the value of each discrete action in that state, and its dimension is consistent with the defined action space dimension. The neural network can be used to evaluate the merits of different actions under the current transient voltage severity.

[0030] As a further optimization scheme of the receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements provided by the present invention, in step (4), a load shedding agent and a time-delay agent are designed according to the action, and a differentiated reward function is designed:

[0031] The purpose of implementing load reduction control in this invention is to minimize the load reduction while ensuring that the voltage recovery of the load bus after a disturbance meets the TVRC requirements. Therefore, rewards are set based on the voltage recovery effect and the magnitude of the load reduction loss.

[0032] For the load shedding agent, a reward function r is designed based on the load bus voltage recovery effect. Ct The expression is as follows:

[0033]

[0034] In the formula: α1 and α2 are positive constants, U t Let U be the load bus voltage at time t, and U0 be the initial value of the bus voltage before the disturbance.

[0035] For a time-delay agent, a reward function r is designed based on the magnitude of the load shedding loss. Tt The expression is as follows:

[0036]

[0037] In the formula: β1 and β2 are positive constants, ΔP Ct This represents the current load reduction amount at the load station.

[0038] The reward function r of the above design Ct and r Tt By introducing conditional statements, it is possible to clearly define whether the load bus status meets the TVRC requirements after the agent's action is implemented, and to provide targeted rewards. In the dynamic interaction between the load reduction agent and the delay agent and the power grid environment, the value of the action implementation can be effectively evaluated, which can be used to guide the agent to perform reinforcement learning to optimize the action strategy.

[0039] As a further optimization scheme of the receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements provided by the present invention, in step (5), the gradient descent algorithm is used to train the agent offline in a centralized manner to further improve the control performance:

[0040] Offline intensive training is performed based on the "state-action-reward-next state" experience quadruples generated from the interaction between the load reduction agent and the delay agent and the power grid environment.

[0041] (5-1) Construction of the dataset

[0042] The agent selects load shedding control actions based on the transient voltage evolution state perceived at its location to respond appropriately to the power grid environment, and receives corresponding rewards based on the control effects of the actions. This dynamic process is described using Markov modeling. t A t ,R t ,S t+1 A sequential dataset composed of four tuples. For the transient voltage evolution trajectory at load sites, combined with TVRC, the corresponding transient voltage state—voltage amplitude deviation ΔV—is extracted. m and recovery time deviation ΔT r As state S t Two agents make action decisions A based on the voltage state, namely, the amount of load reduction in each round and the time delay between rounds. t And obtain the corresponding reward R from the power grid environment. t and the next state S to which it transitions. t+1 These constitute the load-reducing intelligent agent and the time-delay intelligent agent, respectively. t A t ,R t ,S t+1 The quadruple dataset is stored in the load reduction experience pool and the latency experience pool as training samples.

[0043] When an agent interacts with the power grid to collect an experience quadruple dataset, action a Ct and a Tt The ε-greedy algorithm is used for selection, and the expression is:

[0044]

[0045] Where: ε C and ε T To reduce the exploration rate of both load-bearing and time-delayed agents, ω C and ω T Q C and Q T The weight parameter, a C and a T The actions that can be selected in the corresponding action space for both load-reducing and time-delayed agents.

[0046] (5-2) Training process

[0047] ​​A loss function is established, and batch data is selected from the experience pool for agent training to obtain load reduction agents and delay agents that maximize rewards. Based on the above process and the four-tuple dataset of each load site obtained interactively, a unified offline centralized reinforcement learning training is performed on the load reduction agents and delay agents. A temporal difference gradient descent algorithm is used to continuously increase the rewards obtained by the agent's actions, enabling the agent to select the optimal value action strategy in response to the environmental state.

[0048] The agent learns and trains to optimize the state-action value of the Q-network, enabling it to make the best action in response to a state. This is based on Markov processes and uses the optimal Bellman equation to characterize the Q-network. C and Q T The maximum expected value, combined with the estimation of the target network, the state-action value function formed by the optimal Bellman equation based on Double DQN can be expressed as:

[0049]

[0050] In the formula: R represents expectation; Ct and R Tt For the agent's reward that has not yet been observed at the current moment, it can be calculated from the reward function in the next state; γ C and γ T These are the discount factors for the corresponding agents; maxQ C (S t+1 ,a C ;ω C ) is the next state S t+1 The following corresponds to Q C The corresponding maximum action value, maxQ T (S t+1 ,a T ;ω T ) is the next state S t+1 Q T The corresponding maximum action value.

[0051] Monte Carlo methods can be used to approximate the expectation of random variables; therefore, quadruples obtained using time series processes can be utilized. t ,a t ,r t ,s t+1 To express the expectation, the above formulas (11) and (12) are approximately:

[0052] Q C (s t ,a Ct ;ω C )≈r Ct +γ C ​maxQ C (s t+1 ,a C ;ω C (13)

[0053] Q T (s t ,a Tt ;ω T )≈r Tt +γ T maxQ T (s t+1 ,a T ;ω T (14)

[0054] Where: maxQ C (s t+1 ,a C ;ω C ) is the next state s t+1 The following corresponds to Q C The corresponding maximum action value, maxQ T (s t+1 ,a T ;ω T ) is the next state s t+1 Q T The corresponding maximum action value.

[0055] Combining equations (13) and (14), the right side of the equation represents the learning objective, i.e., the label quantity used in the reinforcement learning process, for Q. C and Q T Establish loss functions respectively and The expression is as follows:

[0056]

[0057] In the formula: y Ct and y Tt Represents Q C and Q T The target value of the network at time t is composed of the reward from the target network and the actual observation, where T is the end time of the round.

[0058] Estimate network weight parameters ω using gradient descent algorithm C and ω T The update formula can be expressed as:

[0059]

[0060] In the formula: Indicates gradient calculation; α C and α TThese are the learning rates for the unloaded agent and the delayed agent, respectively.

[0061] The agent explores the action space through equations (9) to (10), accumulates the data of the four-tuples in the Markov process to build an experience pool. When the data samples stored in the experience pool reach the predetermined storage capacity, the agent randomly extracts a small batch of sample data from the experience pool for offline centralized training. The agent updates the weight parameters of the estimation network and the target network through equations (13) to (22) to minimize the loss function and make the state-action value estimation network approach the optimum.

[0062] As a further optimization of the receiving-end grid response-driven load shedding control method provided by this invention, which considers the transient voltage recovery requirements, a distributed adaptive load shedding control strategy is formed by comprehensively deploying each load site of the trained agent. This strategy can respond to local voltage information to perform closed-loop load shedding actions and make local adaptive decisions on the load shedding location, timing, and amount, forming a "monitoring-judgment-decision-control" closed-loop transient voltage load shedding control strategy, thereby promoting the overall voltage recovery of the power grid.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0064] (1) The present invention specifically designs a receiving-end grid response-driven load shedding control method that considers the transient voltage recovery requirements, so that the voltage recovery meets TVRC and can minimize the load shedding cost. The controller at each load site can respond to local voltage information to perform adaptive load shedding actions, so as to realize the overall voltage recovery of the grid after the disturbance.

[0065] (2) This paper models the transient voltage control process as a Markov decision process and solves it using the DDQN algorithm. It designs a load shedding agent and a time delay agent to learn the optimal control method in the process of interacting with the environment, and obtains an adaptive load shedding transient voltage control method based on D2DQN. This strategy can adaptively decide the load shedding location, load shedding amount, load shedding rounds and load shedding time according to the severity of the transient voltage evolution trajectory in the power grid, and respond to TVRC with a smaller load shedding cost.

[0066] (3) The transient voltage load shedding control structure based on D2DQN designed in this paper adopts the execution mode of offline centralized training-online distributed control, which enables the agent to effectively take into account the global control performance of the system while performing real-time autonomous control based on local information.

[0067] (4) The distributed load shedding control method proposed in this invention consists of load shedding agents and time delay agents installed and deployed at each load site. It can make real-time adaptive decisions on load shedding elements based on the severity of the transient voltage evolution trajectory of the load node, so as to make the voltage recovery meet the TVRC requirements with a smaller load shedding requirement, thereby promoting the overall voltage recovery of the power grid. Attached Figure Description

[0068] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof.

[0069] Figure 1 This is a flowchart illustrating the overall implementation of a receiving-end grid response-driven load shedding control method that considers transient voltage recovery requirements, as proposed in this invention.

[0070] Figure 2 This is a schematic diagram illustrating the feedback signal extraction based on TVRC requirements and real-time voltage change status in this invention.

[0071] Figure 3 This is a schematic diagram of the offline centralized training part of the interaction between the intelligent agent and the environment designed in this invention.

[0072] Figure 4 This is a diagram of the distributed real-time adaptive load shedding control structure of the power grid in this invention.

[0073] Figure 5 This is a structural diagram of the Nordic system used in this invention to verify the proposed load reduction control strategy.

[0074] Figure 6 For the present invention Figure 5 The diagram shows the bus voltage cluster curves under four basic fault scenarios of the power grid setup when no load shedding action is implemented.

[0075] Among them, (a) is the bus voltage cluster curve of fault scenario 1 without load shedding control; (b) is the bus voltage cluster curve of fault scenario 2 without load shedding control; (c) is the bus voltage cluster curve of fault scenario 3 without load shedding control; and (d) is the bus voltage cluster curve of fault scenario 4 without load shedding control.

[0076] Figure 7 This is a bus voltage cluster curve under the proposed load reduction control method applied in this invention;

[0077] Among them, (a) is the bus voltage cluster curve after applying load reduction control in fault scenario 1; (b) is the bus voltage cluster curve after applying load reduction control in fault scenario 2; (c) is the bus voltage cluster curve after applying load reduction control in fault scenario 3; and (d) is the bus voltage cluster curve after applying load reduction control in fault scenario 4. Detailed Implementation

[0078] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. Of course, the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0079] Example 1

[0080] To implement a response-driven distributed load shedding control strategy for the receiving-end power grid and achieve voltage recovery control, this embodiment provides a technical solution: a response-driven load shedding control method for the receiving-end power grid that considers transient voltage recovery requirements, comprising the following steps:

[0081] (1) The transient voltage control is modeled as a Markov sequential decision process model;

[0082] (2) Based on the transient voltage recovery requirement (TVRC) of the power grid and the real-time voltage evolution trajectory of the load node, the environmental state variables in the Markov sequential decision-making process are extracted;

[0083] (3) Based on the factors affecting voltage recovery, design the load reduction amount and the time delay between rounds as actions in the Markov sequential decision-making process;

[0084] (4) Based on the action design, load reduction agents and time delay agents were designed, and a differentiated reward function was designed;

[0085] (5) The agent uses the state-action-reward dataset generated during the interaction with the environment based on the Markov model for centralized offline training.

[0086] (5) The trained load shedding agents and time delay agents are deployed at load sites to form a distributed load shedding control structure. They can adaptively decide the load shedding location, load shedding amount and load shedding time, so as to promote voltage recovery with a smaller load shedding amount and meet the transient voltage recovery requirements.

[0087] Specifically, the implementation method of step (1) consists of (2), (3), and (4). In step (2), based on the grid transient voltage recovery requirements and the real-time voltage evolution trajectory of the load nodes, the environmental state variables in the Markov sequential decision-making process are extracted:

[0088] With attachment Figure 2In the example of voltage trajectory 2, where the voltage recovers to the safe range with a delay, the negative voltage amplitude deviation ΔV is extracted based on the power grid TVRC requirements and the real-time voltage evolution trajectory of the load node. m and recovery time deviation ΔT r Voltage deviation ΔV m Represented as:

[0089] ΔV m (t k ) = V th1 -V m (t k (1)

[0090] Where: ΔV m (t k ) indicates the current load node at time t k Voltage deviation at time, V th1 Indicates that at the current t k The voltage threshold that needs to be restored to at different times according to TVRC requirements, V m (t k ) represents the current load node at time t k The voltage value at that moment;

[0091] Voltage recovery time deviation ΔT r Calculated by the following formula:

[0092]

[0093] ΔT r (t k )=min{t k +T es (t k )-t c -T req1 T r,max} (3)

[0094] In equation (2): T es (t k ) for in t k The formula represents the time required for the estimated voltage to recover to the corresponding safe voltage threshold at time t, where Δt is the time step of 0.02 s. This formula indicates that at time t k At any given moment, if the voltage shows an upward trend, T es (t k If T is a finite positive value, and the voltage shows a decreasing trend, then T es (t k The value of ΔT is positive infinity, indicating that the estimated voltage trajectory is difficult to recover. In equation (3): ΔT r (t k ) for t k Time recovery time deviation, tc t represents the time when the short-circuit fault is cleared. k Indicates the current time, T req This indicates that the voltage has recovered to V after the fault. th Maximum allowable recovery time, T r,max Indicates ΔT r The maximum value of can be taken as the time scale of the transient voltage recovery problem of interest. It should be noted that TVRC has piecewise characteristics, and ΔV varies at different times during the voltage recovery process. m and ΔT r It should be based on V at the corresponding time. th and T req To calculate.

[0095] ΔV extracted based on TVRC requirements m (t k ) and ΔT r (t k This can be used to characterize the severity of the current voltage evolution state. The necessary condition for triggering load shedding control is: ΔV m (t k )>0 and ΔT r (t k )>0. When a load node bus needs to perform load reduction control, ΔV is extracted based on the voltage trajectory evolution characteristics of each load bus. m and ΔT r Two-dimensional state as environmental state quantity s t :

[0096] s t ={ΔV m (t),ΔT r (t)} (4)

[0097] Specifically, in step (3), based on the factors affecting voltage recovery, the magnitude of the load reduction and the time delay between rounds are designed as actions in the Markov sequential decision-making process, including:

[0098] The magnitude of the load shedding in each round and the time delay between rounds (affecting the timing of the load shedding action) both influence the transient voltage load shedding control effect. Furthermore, considering the discrete nature of the substation load outgoing lines, the load shedding amount is set as a discrete quantity. The magnitude of the load shedding amount and the time delay between rounds are selected as decision actions a, respectively. Ct and a Tt Its corresponding discrete action space A Cd and A Td They are represented as follows:

[0099] A Cd ={ΔP d1 ,ΔP d2 ,…,Δp dm} (5)

[0100] A Td ={ΔT d1 ,ΔT d2 ,…,ΔT dn} (6)

[0101] Where: ΔP d1 ,ΔP d2 ,…,ΔP dm ΔT represents the action sequence consisting of the percentage of load reduction at load stations. d1 ,ΔT d2 ,…,ΔT dn This represents the sequence of actions consisting of the time intervals between load shedding rounds, where m and n represent the action space A, respectively. Cd and A Td Dimensions.

[0102] Two agents, a load-reduction agent and a time-delay agent, were designed to select load-reduction actions. These agents can respond differently to the environment and achieve optimal performance through training. The load-reduction agent determines the amount of load reduction based on environmental changes, while the time-delay agent determines the time interval between each load-reduction action. The agents select actions and respond using a value estimation network. The value estimation network Q for both the load-reduction agent and the time-delay agent is described. C and Q T This is used to establish a state-action value mapping relationship. The value estimation network is a multi-layer neural network, and the input is the transient voltage evolution state ΔV of the load bus extracted based on TVRC. m and ΔT r The two-dimensional state information is fed forward by multiple hidden layer neurons to obtain the output value. The output value is the value of each discrete action in that state, and its dimension is consistent with the defined action space dimension. The neural network can be used to evaluate the merits of different actions under the current transient voltage severity.

[0103] Specifically, in step (4), a load-reducing agent and a delay agent are designed based on the action design, and a differentiated reward function is designed:

[0104] The purpose of implementing load shedding control in this paper is to minimize the load shedding amount while ensuring that the voltage recovery of the load bus after the disturbance meets the TVRC requirements. Therefore, rewards are set based on the voltage recovery effect and the magnitude of the load shedding loss.

[0105] For the load shedding agent, a reward function r is designed based on the load bus voltage recovery effect. Ct The expression is as follows:

[0106]

[0107] In the formula: α1 and α2 are both set to 1, Ut Let U be the load bus voltage at time t, and U0 be the initial value of the bus voltage before the disturbance.

[0108] For a time-delay agent, a reward function r is designed based on the magnitude of the load shedding loss. Tt The expression is as follows:

[0109]

[0110] In the formula: β1=1, β2=10, ΔP Ct This represents the current load reduction amount at the load station.

[0111] The reward function r of the above design Ct and r Tt By introducing conditional statements, it is possible to clearly define whether the load bus status meets the TVRC requirements after the agent's action is implemented, and to provide targeted rewards. In the dynamic interaction between the load reduction agent and the delay agent and the power grid environment, the value of the action implementation can be effectively evaluated, which can be used to guide the agent to perform reinforcement learning to optimize the action strategy.

[0112] Specifically, in step (5), the gradient descent algorithm is used to train the agent offline in a centralized manner to further improve the control performance:

[0113] Offline intensive training is performed based on the "state-action-reward-next state" experience quadruples generated from the interaction between the load reduction agent and the delay agent and the power grid environment.

[0114] (5-1) Construction of the dataset

[0115] The agent selects load shedding control actions based on the transient voltage evolution state perceived at its location to respond appropriately to the power grid environment, and receives corresponding rewards based on the control effects of the actions. This dynamic process is described using Markov modeling. t A t ,R t ,S t+1 A sequential dataset composed of four tuples. For the transient voltage evolution trajectory at load sites, combined with TVRC, the corresponding transient voltage state—voltage amplitude deviation ΔV—is extracted. m and recovery time deviation ΔT r As state S t Two agents make action decisions A based on the voltage state, namely, the amount of load reduction in each round and the time delay between rounds. t And obtain the corresponding reward R from the power grid environment. t and the next state S to which it transitions. t+1 These constitute the load-reducing intelligent agent and the time-delay intelligent agent, respectively. t A t ,R​​t ,S t+1 The quadruple dataset is stored in the load reduction experience pool and the latency experience pool as training samples.

[0116] When an agent interacts with the power grid to collect an experience quadruple dataset, action a Ct and a Tt The ε-greedy algorithm is used for selection, and the expression is:

[0117]

[0118] Where: ε C and ε T To reduce the exploration rate ε of the load agent and the time-delay agent C =0.3, ε T =0.3; ω C and ω T Q C and Q T The weighting parameter; a C and a T The actions that can be selected in the corresponding action space for both load-reducing and time-delayed agents.

[0119] (5-2) Training process

[0120] A loss function is established, and batch data is selected from the experience pool for agent training to obtain load reduction agents and delay agents that maximize rewards. Based on the above process and the four-tuple dataset of each load site obtained interactively, a unified offline centralized reinforcement learning training is performed on the load reduction agents and delay agents. A temporal difference gradient descent algorithm is used to continuously increase the rewards obtained by the agent's actions, enabling the agent to select the optimal value action strategy in response to the environmental state.

[0121] The agent learns and trains to optimize the state-action value of the Q-network, enabling it to make the best action in response to a state. This is based on Markov processes and uses the optimal Bellman equation to characterize the Q-network. C and Q T The maximum expected value, combined with the estimation of the target network, the state-action value function formed by the optimal Bellman equation based on Double DQN can be expressed as:

[0122]

[0123] In the formula: R represents expectation; Ct and R Tt For the agent's reward that has not yet been observed at the current moment, it can be calculated from the reward function in the next state; γ C and γ TThese are the discount factors for the corresponding agents; maxQ C (S t+1 ,a C ;ω C ) is the next state S t+1 The following corresponds to Q C The corresponding maximum action value, maxQ T (S t+1 ,a T ;ω T ) is the next state S t+1 Q T The corresponding maximum action value.

[0124] Monte Carlo methods can be used to approximate the expectation of random variables; therefore, quadruples obtained using time series processes can be utilized. t ,a t ,r t ,s t+1 To express the expectation, the above formulas (11) and (12) are approximately:

[0125] Q C (s t ,a Ct ;ω C )≈r Ct +γ C maxQ C (s t+1 ,a C ;ω C (13)

[0126] Q T (s t ,a Tt ;ω T )≈r Tt +γ T maxQ T (s t+1 ,a T ;ω T (14)

[0127] Where: maxQ C (s t+1 ,a C ;ω C ) is the next state s t+1 The following corresponds to Q C The corresponding maximum action value, maxQ T (s t+1 ,a T ;ω T ) is the next state s t+1 Q T The corresponding maximum action value. ​

[0128] Combining equations (13) and (14), the right side of the equation represents the learning objective, i.e., the label quantity used in the reinforcement learning process, for Q. C and Q T Establish loss functions respectively and The expression is as follows:

[0129]

[0130] In the formula: y Ct and y Tt Represents Q C and Q T The target value of the network at time t is composed of the reward from the target network and the actual observation, where T is the end time of the round.

[0131] Estimate network weight parameters ω using gradient descent algorithm C and ω T The update formula can be expressed as:

[0132]

[0133] In the formula: Indicates gradient calculation; α C and α T The learning rates α and α for the unloaded agent and the delayed agent, respectively. C =0.001, α T =0.001.

[0134] The agent explores the action space through equations (9)-(10), accumulates the data of the four-tuples in the Markov process to build an experience pool. When the data samples stored in the experience pool reach the predetermined storage capacity, the agent randomly extracts a small batch of sample data from the experience pool for offline centralized training. The agent updates the weight parameters of the estimation network and the target network through equations (13)-(22) to minimize the loss function and make the state-action value estimation network approach the optimum.

[0135] Specifically, in step (6), a distributed adaptive load shedding control strategy is formed by comprehensively deploying the trained intelligent agents at various load sites. This strategy uses local voltage information to respond to closed-loop load shedding actions, making local adaptive decisions on the load shedding location, timing, and amount. This forms a "monitoring-judgment-decision-control" closed-loop transient voltage load shedding control strategy, thereby promoting the overall voltage recovery of the power grid. The distributed load shedding control structure is shown in the attached figure. Figure 4 As shown.

[0136] To fully verify the effectiveness of the load reduction control strategy proposed in this embodiment, an additional method is adopted. Figure 5The simulation analysis is conducted using the Nordic power grid. A CLM load model was established for 11 load sites in the heavily loaded Central region for simulation. Specifically, three-phase short-circuit faults at sites a, b, c, and d were simulated as basic fault scenarios 1, 2, 3, and 4. The voltage curves before and after applying transient control are shown in the attached figures. Figure 6 and attached Figure 7 As shown.

[0137] It can be seen that the bus voltage trajectory without load reduction control in all four scenarios exceeded the safe voltage recovery range specified by TVRC. However, the voltage trajectory after the deployed and trained agent applied load reduction control could effectively meet the TVRC requirements. The total load reduction in the four scenarios was 805MW.

[0138] In summary, the simulation results further verify the sufficiency and effectiveness of the load reduction control method proposed in this embodiment, which is of great significance for improving the voltage safety level of the receiving-end power grid and ensuring the safe and stable operation of the system.

[0139] Example 2

[0140] To further verify the superiority of the method in this embodiment, the method proposed in this embodiment is compared with the traditional distributed load shedding control method. The specific load shedding control method is as follows:

[0141] (1) Method 1: The adaptive load shedding transient voltage control method based on Double DQN response driving proposed in this embodiment;

[0142] (2) Method 2: Traditional distributed load reduction control strategy, with load reduction control start thresholds of 0.9pu, 0.8pu and 0.7pu respectively, load reduction amount of 10% in each round, and time delay between rounds of 0.2s.

[0143] Table 1 shows the control effects and load reduction of the two control methods under fault scenarios 1-4. The results show that after control by method 1, all scenarios meet the TVRC requirements, while after control by method 4, the voltage over-limit situation is particularly severe, and the TVRC requirements cannot be met in any of the four fault scenarios. Furthermore, in all scenarios, method 1 can promote voltage recovery with less load reduction, and the overall load reduction is 69.7% smaller than that of method 2. In summary... Figure 6 , Figure 7 The simulation results in Table 1 show that the proposed load shedding control strategy can respond to TVRC at a lower load shedding cost when adaptively performing transient voltage load shedding control, thus promoting the recovery of voltage levels. This verifies the superiority of the proposed adaptive load shedding control strategy.

[0144] Table 1. Control effect and load reduction under different control methods

[0145] Tab.1Control performance and load shedding amount with different control methods in fault scenario 1-4

[0146]

[0147]

[0148] Note: ΔP All This represents the load reduction (MW) for each fault scenario; ΣΔP All Represents all fault scenarios ΔP All The sum; P erf_ind This represents the transient voltage control performance index, where 1 indicates that TVRC is met and 0 indicates that TVRC is not met.

[0149] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements, characterized in that, Includes the following steps: (1) The transient voltage control is modeled as a Markov sequential decision process model; (2) Based on the transient voltage recovery requirements (TVRC) of the power grid and the real-time voltage evolution trajectory of the load node, the environmental state variables in the Markov sequential decision-making process are extracted. (3) Based on the factors affecting voltage recovery, design the load reduction amount and the time delay between rounds as actions in the Markov sequential decision-making process; (4) Based on the action design, load reduction agents and time delay agents were designed, and a differentiated reward function was designed; (5) The agent uses the state-action-reward dataset generated during the interaction with the environment based on the Markov model for centralized offline training. In step (5), offline centralized training is performed based on the "state-action-reward-next state" experience quadruples generated from the interaction between the load reduction agent and the delay agent with the power grid environment, including: The first step is to construct the dataset. Based on the transient voltage evolution perceived at its location, the agent selects load shedding control actions to respond to the power grid environment and observes the control effects under these actions, receiving corresponding rewards. This dynamic process is described using Markov modeling. t A t ,R t ,S t+1 The sequential dataset composed of four tuples is used to extract the corresponding transient voltage state-voltage amplitude deviation ΔV based on the transient voltage evolution trajectory at load sites, combined with TVRC. m and recovery time deviation ΔT r As state S t Two agents make action decisions A based on the voltage state, namely, the amount of load reduction in each round and the time delay between rounds. t And obtain the corresponding reward R from the power grid environment. t and the next state S to which it transitions. t+1 These constitute the load-reducing intelligent agent and the time-delay intelligent agent, respectively. t A t ,R t ,S t+1 The quadruple dataset is stored in the load reduction experience pool and the latency experience pool as training samples.​​ When an agent interacts with the power grid to collect an experience quadruple dataset, action a Ct and a Tt Using an ε-greedy algorithm for selection, the expression is: Where: ε C and ε T To reduce the exploration rate of both load-bearing and time-delayed agents, ω C and ω T Q C and Q T The weight parameter, a C and a T The actions that can be selected by the load-reducing agent and the time-delay agent in their respective action spaces; Secondly, a loss function is established and a batch of data is selected from the experience pool for agent training to obtain load reduction agents and delay agents with optimal action values. Based on the four-tuple dataset of each load site obtained through interaction, the above process is used to conduct unified offline centralized reinforcement learning training for load reduction agents and delay agents. The temporal difference gradient descent algorithm is used to continuously increase the reward obtained by the agent's actions, so that the agent can choose the optimal value action strategy in response to the environmental state. The agent learns and is trained to make optimal decisions in response to states. This is based on Markov processes, and the optimal Bellman equation characterizes Q. C and Q T The maximum expected value, combined with the estimation network and the target network, is expressed by the state-action value function based on the optimal Bellman equation of DoubleDQN as follows: In the formula: R represents expectation; Ct and R Tt The reward for the agent has not yet been observed at the current moment; it is calculated by the reward function in the next state. C and γ T These are the discount factors for the corresponding agents; maxQ C (S t+1 ,a C ;ω C ) is the next state S t+1 The following corresponds to Q C The corresponding maximum action value, maxQ T (S t+1 ,a T ;ω T ) is the next state S t+1 Q T The corresponding maximum action value; Monte Carlo methods are used to approximate the expectation of random variables, utilizing quadruples obtained through time series processes. t ,a t ,r t ,s t+1 > To express expectations, the above formulas (11) and (12) are expressed as:​ Q C (s t ,a Ct Oh C )≈r Ct +g C maxQ C (s t+1 ,a C Oh C (13) Q T (s t ,a Tt Oh T )≈r Tt +g T maxQ T (s t+1 ,a T Oh T (14) Where: maxQ C (s t+1 ,a C ;ω C ) is the next state s t+1 The following corresponds to Q C The corresponding maximum action value, maxQ T (s t+1 ,a T ;ω T ) is the next state s t+1 Q T The corresponding maximum action value; Combining equations (13) and (14), the right side of the equation represents the learning objective, i.e., the label quantity used in the reinforcement learning process, for Q. C and Q T Establish loss functions separately and The expression is as follows: In the formula: y Ct and y Tt Represents Q C and Q T The target value of the network at time t is composed of the reward of the target network and the actual observation, where T is the end time of the round; Estimate network weight parameters ω using gradient descent algorithm C and ω T The update formula is expressed as: In the formula: Indicates gradient calculation; α C and α T These are the learning rates for the unloaded agent and the delayed agent, respectively. The agent is trained to achieve the optimal state-action value function based on the Double DQN algorithm. The action space is explored through equations (9) to (10), and the experience pool is built by accumulating four-tuple data in the Markov process. When the data samples stored in the experience pool reach the predetermined storage capacity, the agent randomly selects a small batch of sample data from the experience pool for offline centralized training. The estimation network and target network weight parameters of the agent are updated through equations (13) to (22) to minimize the loss function and make the state-action value estimation network approach the optimal value. (6) The trained load shedding agents and time delay agents are deployed at the load sites to form a distributed load shedding control structure. They adaptively decide the load shedding location, load shedding amount and load shedding time to minimize the load shedding amount and promote voltage recovery to meet the transient voltage recovery requirements.

2. The receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements according to claim 1, characterized in that, In step (2), based on the power grid transient voltage recovery requirements and the real-time voltage evolution trajectory of the load nodes, the environmental state variables in the Markov sequential decision-making process are extracted, including: Based on the TVRC requirements of the power grid and the real-time voltage evolution trajectory of the load nodes, the voltage amplitude deviation ΔV is extracted. m and recovery time deviation ΔT r : Voltage deviation ΔV m Represented as: ΔV m (t k )=V th -V m (t k ) (1) Where: ΔV m (t k ) indicates the current load node at time t k Voltage deviation at time, V th V represents the voltage threshold to which the voltage recovers at different time periods according to TVRC requirements. m (t k ) represents the current load node at time t k The voltage value at that moment; Voltage recovery time deviation ΔT r Calculated by the following formula: ΔT r (t k )=min{t k +T es (t k )-t c -T req ,T r,max } (3) In equation (2): T es (t k ) for in t k The formula shows that at time t, the estimated time required for the voltage to recover to the corresponding safe voltage threshold is given. k At any given moment, if the voltage shows an upward trend, T es (t k If T is a finite positive value, and the voltage shows a decreasing trend, then T es (t k The value is positive infinity, used to indicate that the predicted voltage trajectory is difficult to recover, and Δt is the time step. In equation (3): ΔT r (t k ) for t k Time recovery time deviation, t k Indicates the current time, t c T represents the time when the short-circuit fault is cleared. req This indicates that the voltage recovered to V after the fault. th Maximum allowable recovery time, T r,max Indicates ΔT r The maximum value is taken as the time scale of the transient voltage recovery problem of interest. TVRC has piecewise characteristics, and ΔV varies at different times during the voltage recovery process. m and ΔT r It should be based on V at the corresponding time. th and T req To calculate; ΔV extracted based on TVRC requirements m (t k ) and ΔT r (t k This is used to characterize the severity of the current voltage evolution state. The necessary condition for triggering load shedding control is: ΔV m (t k )>0 and ΔT r (t k When the load node busbar performs load shedding control, ΔV is extracted based on the voltage trajectory evolution characteristics of each load busbar, where ΔV > 0. m and ΔT r Two-dimensional state as environmental state quantity s t : s t ={ΔV m (t),ΔT r (t)} (4)。 3. The receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements according to claim 1, characterized in that, In step (3), based on the factors affecting voltage recovery, the magnitude of the load reduction and the time delay between rounds are designed as actions in the Markov sequential decision-making process: The magnitude of the load shedding in each round and the time delay between rounds both affect the transient voltage load shedding control effect. Furthermore, considering the discrete nature of the substation load outgoing lines, the load shedding amount is set as a discrete quantity, and the magnitude of the load shedding amount and the time delay between rounds are selected as decision actions. Ct and a Tt Its corresponding discrete action space A Cd and A Td They are represented as follows: A Cd ={ΔP d1 ,ΔP d2 , …,Δp dm } (5) A Td ={ΔT d1 ,ΔT d2 ,…,ΔT dn } (6) Where: ΔP d1 ,ΔP d2 ,…,ΔP dm ΔT represents the action sequence consisting of the percentage of load reduction at load stations. d1 ,ΔT d2 ,…,ΔT dn This represents the sequence of actions consisting of the time intervals between load shedding rounds, where m and n represent the action space A, respectively. Cd and A Td Dimensions.

4. The receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements according to claim 1, characterized in that, In step (5), two agents, a load reduction agent and a delay agent, are designed based on the selection of the load reduction action: The agent responds to different actions based on the environment through a value estimation network. After learning and training, the agent reaches its optimal state. The load reduction agent decides the amount of load reduction based on environmental changes, while the delay agent decides the time interval between each load reduction action. The value estimation network Q for both the load reduction agent and the delay agent... C and Q T To establish the state-action value mapping relationship, the value estimation network is a multi-layer neural network, and the input is the load bus transient voltage evolution state ΔV extracted based on TVRC. m and ΔT r The two-dimensional state information is fed forward by multiple hidden layer neurons to obtain the output value. The output value is the value of each discrete action in the state. Its dimension is consistent with the defined action space dimension. The neural network evaluates the merits of different actions under the current transient voltage severity. Based on the control objectives and performance sought by transient voltage load shedding control, a reward function is set for the agent. Load shedding control refers to minimizing the load shedding amount while ensuring that the load bus voltage recovers to meet the TVRC requirements after the disturbance. Rewards are set based on the voltage recovery effect and the magnitude of the load shedding loss, respectively. For the load shedding agent, a reward function r is designed based on the load bus voltage recovery effect. Ct The expression is as follows: In the formula: α1 and α2 are positive constants, U t Let U be the load bus voltage at time t, and U0 be the initial value of the bus voltage before the disturbance. For a time-delay agent, a reward function r is designed based on the magnitude of the load shedding loss. Tt The expression is as follows: In the formula: β1 and β2 are positive constants, ΔP Ct The amount of load reduction for the current wheel at the load station; The reward function r of the above design Ct and r Tt All of them explicitly define whether the load bus status meets the TVRC requirements after the agent's action is implemented by introducing conditional statements, and give targeted rewards. The value of the action implementation is evaluated in the dynamic interaction between the load reduction agent and the delay agent and the power grid environment, which is used to guide the agent to optimize the action strategy through reinforcement learning.

5. The receiving-end grid response-driven load shedding control method considering transient voltage recovery requirements according to claim 1, characterized in that, In step (5), the load reduction agent and the time delay agent, which have been reinforced and trained, are deployed at each load reduction site. Based on the transient voltage trajectory evolution state monitored by the load site, the agent makes local adaptive decisions on the load reduction location, load reduction time and load reduction amount, forming a transient voltage load reduction control strategy in a closed loop of "monitoring-judgment-decision-control". The agent updates the parameters in a centralized manner based on the global experience dataset, so that the distributed control takes into account the overall load reduction optimization control performance on the basis of local efficient autonomous load reduction control.

Citation Information

Patent Citations

  • Power distribution network dynamic reconstruction method and system based on dynamic space screening

    CN115051362A