Emergency load shedding control method, system and medium for transient voltage stability of power system
By adopting deep reinforcement learning agents and linear decision space in power systems, the problems of low training efficiency and decision quality of agents in emergency load control are solved, and efficient emergency load control is achieved in the case of transient voltage instability.
Patent Information
- Application Number
- CN202211325051.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-10-27
AI Technical Summary
In the emergency load control, the existing technology has problems such as difficulty in designing Markov decision-making process, excessive decision space, and insufficient knowledge integration, resulting in low training efficiency and decision-making quality of agents.
Deep reinforcement learning (DRL) agents are used to design linear decision spaces and integrate domain knowledge, establish knowledge-enhanced MDPs, use fault condition samples for training, select effective load cutting control measures, and perform emergency load cutting control in the power system.
It improves the training efficiency and decision-making quality of the agent, and can provide effective emergency load control measures in the case of transient voltage instability, reduces expert workload and improves the accuracy and efficiency of decision-making.
Smart Images

Figure CN115566690B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power system stability control, and more specifically, relates to a transient voltage stability emergency load shedding control method, system and medium for a power system. Background Art
[0002] After a large disturbance in the power system, the transient voltage instability is a common form of instability. For transient voltage instability, the most commonly used control measure is emergency load shedding. The emergency load shedding control measure is generally put into operation quickly after a fault occurs to restore the safe and stable operation of the power system. The emergency load shedding control measure in the off-line pre-decision - on-line matching mode is widely used in the actual power grid. The traditional off-line pre-decision is generally completed by power grid experts, which is time-consuming and laborious. At present, the off-line pre-decision emergency load shedding control measure is usually formulated based on power flow equations, heuristic algorithms, sensitivity analysis of transient process control variables, etc. The decision measures obtained are often relatively rough and also time-consuming.
[0003] With the development of artificial intelligence, some studies have applied deep reinforcement learning to power system power flow adjustment, generator control, energy management, etc., and achieved good results. Applying deep reinforcement learning to the field of emergency load shedding control is expected to improve the efficiency of formulating off-line pre-decision load shedding measures. However, when the current deep reinforcement learning is applied to emergency load shedding control, there are problems such as difficult design of the Markov Decision Process (MDP), too large decision space, and insufficient knowledge integration. The intelligent agent is difficult to train and the decision-making efficiency is relatively low. Therefore, how to design an effective MDP, design a reasonable decision space and integrate domain knowledge to improve the training efficiency and decision-making quality of the intelligent agent is a technical problem to be solved urgently at present. Summary of the Invention
[0004] Aiming at the defects and improvement requirements of the prior art, the present invention provides a transient voltage stability emergency load shedding control method, system and medium for a power system, aiming to improve the training efficiency and decision-making quality of the intelligent agent.
[0005] To achieve the above object, according to one aspect of the present invention, a transient voltage stability emergency load shedding control method for a power system is provided, including: S1, selecting a fault condition sample from a sample set, where the sample set contains multiple groups of fault conditions under transient voltage instability modes; S2, simulating to obtain a first state, a first completion flag, and a first reward according to the fault condition sample and the previous load shedding control measure; S3, the DRL agent selects the current load shedding control measure in the linear decision space according to the previous load shedding control measure, and simulates to obtain a second state according to the fault condition sample and the current load shedding control measure; S4, using the first state, the first completion flag, the first reward, the current load shedding control measure, and the second state, updating the policy network parameters of the DRL agent, and taking the current load shedding control measure as the new previous load shedding control measure; S5, repeating S2 - S4 until an effective load shedding control measure is obtained or a preset number of iterations is reached; S6, repeating S1 - S5 until a preset number of training rounds is reached, and using the final DRL agent to determine the corresponding relationship between the fault conditions and the load shedding control measures under the transient voltage instability mode.
[0006] Furthermore, the selection of the current load shedding control measure in S3 is restricted by given constraint conditions, and the given constraint conditions include: the selected current load shedding control measure is different from all previously selected load shedding control measures; and when the previous load shedding control measure makes the transient voltage instability more serious, the selected current load shedding control measure does not include any actions in the previous load shedding control measure.
[0007] Furthermore, the decision space of the DRL agent is a linear decision space, and the number of neurons in the output layer of the neural network is:
[0008] N l =q×m
[0009] where, N l is the number of neurons in the output layer of the DRL agent neural network, m is the number of controllable loads, and q is the total number of actions of each controllable load.
[0010] Furthermore, the selection of the current load shedding control measure in S3 includes: in the linear decision space, the DRL agent adopts an exploration - greedy strategy to obtain the value of the output neuron with the largest value in the policy network, and converts the value of the output neuron with the largest value into the sequence number of the load to be shed and the load shedding action value; calculating the load shedding amount of the load corresponding to the sequence number of the load to be shed according to the load shedding action value, and determining the load shedding amounts of the loads corresponding to the remaining sequence numbers according to the previous load shedding control measure to form the current load shedding control measure.
[0011] Furthermore, the load shedding amount in the current load shedding control measure is:
[0012]
[0013]
[0014]
[0015]
[0016] Among them, are the load shedding amounts of the j-th load in the current load shedding control measure and the previous load shedding control measure respectively, is n k The load shedding amount of the corresponding load, n k is the sequence number of the load to be shed, l k is the action value of the load to be shed, l max is the maximum load shedding amount, l min is the minimum load shedding amount, q is the total number of actions of each controllable load, N i is the value of the output neuron with the largest value.
[0017] Furthermore, the S2 includes: simulating the bus voltage amplitude and the relative rotor angle value of the generator under the current decision according to the fault condition sample and the previous load shedding control measure, where the first state includes the bus voltage amplitude, the relative rotor angle value of the generator, and the previous load shedding control measure; judging whether the current decision is effective according to each bus voltage amplitude, and generating the first completion flag and the first reward according to the judgment result.
[0018] Furthermore, the update method of the policy network parameters of the DRL agent in the S4 is:
[0019]
[0020]
[0021] Among them, e TD () is the temporal difference error between the policy network and the target network, θ w 、θ′ w are the policy network parameters and the target network parameters obtained by the w-th update respectively, Ψ is the number of samples selected from the prioritized experience replay pool, γ is the decay factor, Q′() is the target network value, Q() is the policy network value, s k 、s k+1 are the environmental state values of the current training and the previous training respectively, a is the action corresponding to the maximum Q value, r k is the first reward, is the first completion flag, a k is the current load shedding control measure, α is the learning rate, Derive with respect to the parameter θ.
[0022] Furthermore, the method further includes: during the operation of the power grid, when a fault condition is detected, the emergency load shedding control is performed by using the load shedding control measure corresponding to the fault condition in the corresponding relationship.
[0023] According to another aspect of the present invention, there is provided a power system transient voltage stability emergency load shedding control system, including: a sample selection module for selecting a fault condition sample from a sample set, the sample set including multiple groups of fault conditions under transient voltage instability modes; a simulation module for simulating a first state, a first completion flag, and a first reward according to the fault condition sample and the previous load shedding control measure; a measure selection and simulation module for enabling the DRL agent to select the current load shedding control measure in the linear decision space according to the previous load shedding control measure, and simulating a second state according to the fault condition sample and the current load shedding control measure; an update module for updating the policy network parameters of the DRL agent by using the first state, the first completion flag, the first reward, the current load shedding control measure, and the second state, and taking the current load shedding control measure as the new previous load shedding control measure; a repeated execution module for repeatedly executing the simulation module, the measure selection and simulation module, and the update module until an effective load shedding control measure is obtained or a preset number of iterations is reached; a decision module for repeatedly executing the sample selection module, the simulation module, the measure selection and simulation module, the update module, and the repeated execution module until a preset number of training rounds is reached, and using the final DRL agent to determine the corresponding relationship between the fault conditions and the load shedding control measures under the transient voltage instability mode.
[0024] According to another aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the power system transient voltage stability emergency load shedding control method as described above is implemented.
[0025] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0026] (1) Provide a power system transient voltage stability emergency load shedding control method, establish an MDP of a knowledge-enhanced deep reinforcement learning (DRL) agent. Different from the traditional MDP based on response expansion, this MDP is based on events. The DRL agent directly selects the emergency load shedding control measure according to the environmental state, which can better guide the training of the DRL agent. The DRL agent can give an effective emergency load shedding control measure under a new transient voltage instability event;
[0027] (2) Design a linear decision space, which effectively reduces the size of the decision space while ensuring the decision accuracy, thereby reducing the training difficulty of the DRL agent and improving the training efficiency of the DRL agent;
[0028] (3) During the training process of the DRL agent, incorporate constraints to remove duplicate actions and negative-effect actions to constrain the agent's decisions. Compared with pure data-driven DRL, it has higher training efficiency and decision-making quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of the power system transient voltage stability emergency load shedding control method provided by an embodiment of the present invention;
[0030] Figure 2 is Figure 1 an MDP schematic diagram of the method shown;
[0031] Figure 3 is a comparison diagram of the linear decision space and the exponential decision space provided by an embodiment of the present invention;
[0032] Figure 4 is a block diagram of the knowledge-enhanced DRL provided by an embodiment of the present invention;
[0033] Figure 5 is a single-line diagram of the China Electric Power Research Institute 8-machine 36-node system provided by an embodiment of the present invention;
[0034] Figure 6A and Figure 6B are respectively schematic diagrams before and after decision-making of the knowledge-enhanced DRL agent provided by an embodiment of the present invention under a certain transient voltage instability condition;
[0035] Figure 7A and Figure 7B and Figure 7C are respectively schematic diagrams of the total number of iterations, total rewards, and total number of successes of all samples on the test set each time the knowledge-enhanced or not DRL agent is tested provided by an embodiment of the present invention;
[0036] Figure 8 is a block diagram of the power system transient voltage stability emergency load shedding control system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0038] In the present invention, terms such as "first" and "second" in the present invention and the accompanying drawings (if any) are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.
[0039] Figure 1 It is a flowchart of the emergency load shedding control method for transient voltage stability of the power system provided by the embodiment of the present invention. Refer to Figure 1 , and in combination with Figures 2 - 7C , the emergency load shedding control method for transient voltage stability of the power system in this embodiment will be described in detail. The method includes operations S1 - S6, and its overall implementation process is as Figure 2 shown.
[0040] Operation S1: Select a fault condition sample from the sample set, where the sample set contains multiple groups of fault conditions under transient voltage instability modes.
[0041] Before performing operation S1, it is also necessary to design the sample set. Specifically, different fault conditions can be set, including different power flow operation conditions, the proportion of load motor components, fault lines, fault locations, and fault durations. Based on different fault conditions, transient stability simulations are carried out to identify the dominant instability mode, and the fault conditions corresponding to the transient voltage instability mode are selected to form the sample set.
[0042] For different fault conditions, for example, the load level can be set to fluctuate between 90% and 110% of the standard condition; the proportion of motor components in the load is 60%, 70%, 80%; the fault lines traverse all non-transformer lines in the system; the fault locations are 2%, 50%, 98% of the full length of each line; the fault durations are set to 0.05 s and 0.15 s. For the identification of the dominant instability mode, for example, in terms of the time scale, the power angle running away first is usually power angle instability, the voltage collapsing first is usually voltage instability, and generator terminal faults are often power angle instability, etc.
[0043] Furthermore, the sample set is divided into a training set and a test set. In operation S1, a fault condition sample is selected from the training set as a training example.
[0044] Operation S2: According to the fault condition sample and the previous load shedding control measure, simulate to obtain the first state, the first completion flag, and the first reward.
[0045] According to the embodiment of the present invention, operation S2 includes sub-operations S21 - S22.
[0046] In sub-operation S21, based on the fault condition samples and all previous load control measures, the bus voltage amplitude and the relative generator power angle value under the current decision are obtained through simulation. The first state includes the bus voltage amplitude, the relative generator power angle value, and all previous load control measures.
[0047] Specifically, according to the selected fault condition samples and all previous load control measures a k-1 (initial no control measure), the Power System Analysis Software Package (PSASP) simulation tool is called to simulate and obtain the generator state s f from t0 to t k , that is, the first state:
[0048] s k = [V k , θ k , a k-1
[0049] where V k is the bus voltage amplitude at the current (i.e., the k-th step) decision, θ k is the relative generator power angle value at the current decision, and a k-1 is the previous (i.e., the (k - 1)-th step) load shedding control measure.
[0050]
[0051]
[0052]
[0053] where represents the amplitude of the B-th bus voltage at the t s -th second of the k-th step decision, represents the relative power angle value of the G-th generator at the t s -th second of the k-th step decision. represents the control measure (i.e., the load shedding amount) of the m-th load at the (k - 1)-th step decision.
[0054] V k and a k-1 do not need preprocessing, while θ k needs to be mapped to the range [-1, 1]:
[0055]
[0056] In sub-operation S22, based on the bus voltage amplitudes of each bus, it is judged whether the current decision is valid, and a first completion flag and a first reward are generated according to the judgment result.
[0057] Judge the effectiveness of the control measures for the current decision according to the following criteria:
[0058]
[0059] Among them, is the first completion flag, indicating the flag of stability after the current decision. 1 means the current control measure is effective and the transient voltage recovers to stability, and 0 means the current control measure is ineffective. N max represents the maximum allowed number of iterations (i.e., the preset number of iterations).
[0060] TVSI is defined as:
[0061]
[0062]
[0063] Among them, V i,t represents the voltage amplitude of the i-th bus at the t-th moment, t2 is the fault clearing time, and β is the transient voltage stability threshold, generally taking 0.1.
[0064] The environment gives the DRL agent the first reward r according to the effectiveness of the current control measure k :
[0065]
[0066] Among them, r T is the reward for the effective load shedding control measure; is the load shedding amount of load i at the (k - 1)-th step; r Δ is the load shedding penalty, used to limit over-load shedding.
[0067] For operation S3, the DRL agent selects the current load shedding control measure in the linear decision space according to the previous load shedding control measure, and simulates to obtain the second state according to the fault condition sample and the current load shedding control measure.
[0068] According to the embodiments of the present invention, operation S3 includes sub-operations S31 - S32.
[0069] In sub-operation S31, the DRL agent selects the current load shedding control measure in the linear decision space according to the previous load shedding control measure. Specifically as follows:
[0070] (1) In the linear decision space, the DRL agent adopts an exploration-greedy strategy to obtain the value of the output neuron with the largest value in the policy network:
[0071]
[0072] where ε0 is the initial exploration probability, and ε e is the minimum exploration probability, is the set for shielding repeated actions, is the set for shielding negative-effect actions, n k represents the k-th action, and n decay is the attenuation coefficient.
[0073] The selection of the current load shedding control measure is restricted by given constraint conditions, which include: the current load shedding control measure selected is different from all previously selected load shedding control measures; and when the previous load shedding control measure makes the transient voltage instability more serious, the current load shedding control measure selected does not include any action in the previous load shedding control measure.
[0074] Referring to Figure 4 , it can be seen that the output of the linear decision space is restricted. Specifically, the given constraint conditions can be expressed as:
[0075]
[0076]
[0077] where a τ represents the historical load shedding control measures; t(a k ) represents the moment when the lowest bus voltage is lower than 0.5 p.u. after implementing the action a k . If this time is advanced, it indicates that the action a k deteriorates the transient voltage.
[0078] In the embodiment of the present invention, the decision space of the DRL agent is a linear decision space, and the number of neurons in the output layer of the neural network is:
[0079] N l = q × m
[0080] where N l is the number of neurons in the output layer of the DRL agent neural network, m is the number of controllable loads, and q is the total number of actions for each controllable load.
[0081] The number of neurons in the output layer is equal to the number of control measures. The selection of a neuron represents the selection of a control measure. For example, if the number of loads is 9 and each load has 11 action values (cut off 0%, 10%, 20%,..., 100%). For the exponential decision space, any combination of action values of all loads can be selected at one time, and the total number of possible control measures is 11 9。It can be seen that with such a design, the number of neurons increases exponentially as the number of loads increases, and it is almost untrainable with an excessive number of neurons. For a linear decision space, only one action value of one load is selected each time, so there are only 9 * 11 = 99 possible control measures in total. It can be seen that the number of neurons increases linearly as the number of loads increases. When reaching the same load shedding control measures, the number of iterations in the linear decision space may be more, but its training difficulty is significantly reduced, as Figure 3 shown.
[0082] (2) Convert the value of the output neuron with the largest value into the load shedding sequence number and the load shedding action value to be shed:
[0083]
[0084]
[0085] where n k is the load shedding sequence number to be shed, l k is the load shedding action value to be shed, q is the total number of actions of each controllable load, and N i is the value of the output neuron with the largest value.
[0086] (3) Calculate the load shedding amount of the load corresponding to the load shedding sequence number according to the load shedding action value to be shed, and determine the load shedding amounts of the loads corresponding to the remaining sequence numbers according to the previous load shedding control measure to form the current load shedding control measure. The load shedding amounts in the current load shedding control measure are:
[0087]
[0088]
[0089] where are the load shedding amounts of the j-th load in the current load shedding control measure and the previous load shedding control measure respectively, is the load shedding amount of the load corresponding to n k , l max is the maximum load shedding amount, and l min is the minimum load shedding amount.
[0090] In sub-operation S32, according to the fault condition sample and the current load shedding control measure, the second state is simulated. This process is the same as the method for calculating the first state in sub-operation S21 and will not be elaborated here.
[0091] Operation S4: Use the first state, the first completion flag, the first reward, the current load shedding control measure, and the second state to update the policy network parameters of the DRL agent, and use the current load shedding control measure as the new previous load shedding control measure.
[0092] The first state s k , the first completion flag the first reward r k , the current load shedding control measure and the second state a k and the second state s k+1 are combined to form an interaction sample and put into the prioritized experience replay pool. Each interaction sample M has the following form:
[0093]
[0094] Preferably, the knowledge-enhanced DRL agent periodically extracts interaction samples from the prioritized experience replay pool to update its policy network parameters, and its target network parameters are synchronized from the policy network periodically. The update method of the policy network parameters of the DRL agent is as follows:
[0095]
[0096]
[0097] where e TD () is the temporal difference error between the policy network and the target network, θ w , θ′ w are the policy network parameters and target network parameters obtained by the w-th update respectively, Ψ is the number of samples selected from the prioritized experience replay pool, γ is the decay factor, Q′() is the target network value, Q() is the policy network value, s k , s k+1 are the environmental state values of the current training and the previous training respectively, a is the action corresponding to the maximum Q value, r k is the first reward, is the first completion flag, a k is the current load shedding control measure, α is the learning rate, take the derivative of the parameter θ.
[0098] Operation S5, repeatedly execute Operation S2 - Operation S4 until an effective load shedding control measure is obtained or a preset number of iterations is reached.
[0099] Operation S6, repeatedly execute Operation S1 - Operation S5 until a preset number of training rounds is reached, and use the final DRL agent to determine the corresponding relationship between the fault conditions and the load shedding control measures in the transient voltage instability mode.
[0100] According to an embodiment of the present invention, the method further includes: during the operation of the power grid, when a fault condition is detected, use the load shedding control measure corresponding to the fault condition in the corresponding relationship for emergency load shedding control.
[0101] Taking the 8-machine 36-bus test system of China Electric Power Research Institute as an example, the above method will be specifically described below. The single-line diagram of the system is as shown in Figure 5 . In the sample set generation stage, different fault conditions are first set. The pre-fault operating conditions include three power flow levels: 90%, 100%, and 110%. Three-phase metallic short-circuit faults are set on all 26 non-transformer lines. The short-circuit positions are 2%, 50%, and 98% of the line positions respectively, and the faults last for 0.05 s and 0.15 s respectively. The simulation duration is set to 10 s. Then, the relative rotor angle values of the generators and the bus voltage amplitudes within the simulation duration are obtained through PSASP simulation. The dominant instability mode is discriminated, and finally 212 transient voltage instability samples are obtained. The transient voltage instability sample set is randomly divided. 166 samples are used for training, and 46 samples are used for testing.
[0102] Figure 6A and Figure 6B show a typical transient voltage instability condition. The power flow level is 110%, the proportion of load motors is 80%. A three-phase short-circuit fault occurs at 2% of the line position between BUS16 and BUS20 at 0.1 s, and the faulty line is removed at 0.15 s. The number of controllable loads is 9, and the action value of each load ranges from 0 to 1, with an interval of 0.1. Among them, the number of output neurons in the linear decision space is 99, which is easy to train. For the exponential decision space, the number of output neurons is 11 9 , and it is almost impossible to train.
[0103] Referring to Figure 6A , it can be seen that after the fault occurs, the voltages of some buses continue to drop, and the system experiences transient voltage instability, and the bus voltage at BUS16 node is the lowest. The agent is a DRL agent that removes duplicate actions and negative-effect actions. The output of the linear decision space is 21, and the converted actual load-shedding action is: 50% of the load carried by BUS16 node is removed at 0.45 s, with a total of 275 MW removed. Subsequently, the transient voltage recovers to stability, as shown in Figure 6B . According to the experience of power grid experts, load is also shed at BUS16 where the bus voltage is the lowest. Therefore, the above event-based MDP can effectively guide the training of the DRL agent and learn the mapping relationship between the power grid transient voltage instability event and the emergency load-shedding control. The knowledge-enhanced DRL agent can give effective emergency load-shedding measures and reduce the workload of experts.
[0104] Figure 7A 、 Figure 7B 、 Figure 7CThe decision-making performance of DRL agents with and without knowledge enhancement was compared. First, three types of agents were defined. A1 represents a knowledge-enhanced DRL agent that removes duplicate actions and negative-effect actions. A2 represents a knowledge-enhanced DRL agent that removes duplicate actions. A3 represents a DRL agent without knowledge enhancement. A total of 30,000 training rounds were conducted. The decision-making performance of DRL was tested on the test set every 200 training rounds, and the upper limit of each iteration was 50.
[0105] The total reward represents the sum of the rewards obtained by the DRL agent in all test cases for each test. The larger the reward value, the higher the decision-making quality of the agent. The total number of iterations represents the sum of the number of iterations of the DRL agent in all test cases for each test. The smaller the number of iterations, the higher the decision-making quality of the agent. The total number of successes represents the number of successes of the DRL agent in all test cases for each test. The more the number of successes, the higher the decision-making quality of the agent. The faster the agent reaches the convergence state during the training process, the higher the training efficiency of the agent.
[0106] Comparing A2 and A3, the knowledge-enhanced DRL agent that removes duplicate actions performs better in all three metrics and converges faster iteratively, indicating that the knowledge of removing duplicate actions can effectively improve the training efficiency and decision-making quality of the agent. Comparing A1 and A2, the knowledge-enhanced DRL agent that further incorporates the knowledge of removing negative-effect actions performs better in all three metrics and converges faster iteratively, indicating that the knowledge of removing duplicate actions can further improve the training efficiency and decision-making quality of the agent.
[0107] Figure 8 This is the block diagram of the emergency load shedding control system for power system transient voltage stability provided by the embodiment of the present invention. Refer to Figure 8 As shown in, the emergency load shedding control system 800 for power system transient voltage stability includes a sample selection module 810, a simulation module 820, a measure selection and simulation module 830, an update module 840, a repeated execution module 850, and a decision module 860.
[0108] The sample selection module 810, for example, executes operation S1 to select a fault condition sample from the sample set, and the sample set contains multiple groups of fault conditions under transient voltage instability modes.
[0109] The simulation module 820, for example, executes operation S2 to simulate and obtain a first state, a first completion flag, and a first reward according to the fault condition sample and the previous load shedding control measure.
[0110] The measure selection and simulation module 830, for example, executes operation S3 to enable the DRL agent to select the current load shedding control measure in the linear decision space according to the previous load shedding control measure, and simulate to obtain the second state based on the fault condition sample and the current load shedding control measure.
[0111] The update module 840, for example, executes operation S4 to update the policy network parameters of the DRL agent by using the first state, the first completion flag, the first reward, the current load shedding control measure, and the second state, and use the current load shedding control measure as the new previous load shedding control measure.
[0112] The repeated execution module 850, for example, executes operation S5 to repeatedly execute the simulation module 820, the measure selection and simulation module 830, and the update module 840 until an effective load shedding control measure is obtained or a preset number of iterations is reached.
[0113] The decision-making module 860, for example, executes operation S6 to repeatedly execute the sample selection module 810, the simulation module 820, the measure selection and simulation module 830, the update module 840, and the repeated execution module 850 until a preset number of training rounds is reached, and use the final DRL agent to determine the corresponding relationship between the fault condition and the load shedding control measure in the transient voltage instability mode.
[0114] The power system transient voltage stability emergency load shedding control system 800 is used to execute the above Figures 1 - 7C shown power system transient voltage stability emergency load shedding control method.
[0115] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the power system transient voltage stability emergency load shedding control method as shown above Figures 1 - 7C shown.
[0116] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. An emergency load shedding control method for transient voltage stability of a power system, characterized in that, Including: S1. Select a fault condition sample from the sample set, where the sample set contains multiple groups of fault conditions under transient voltage instability modes; S2. According to the fault condition sample and the previous load shedding control measure, simulate to obtain a first state, a first completion flag, and a first reward; S3. The DRL agent selects a current load shedding control measure in the linear decision space according to the previous load shedding control measure, and simulates to obtain a second state according to the fault condition sample and the current load shedding control measure; The selection of the current load shedding control measure in S3 includes: in the linear decision space, the DRL agent adopts an exploration-greedy strategy to obtain the value of the output neuron with the largest value in the policy network, and converts the value of the output neuron with the largest value into the sequence number of the load to be shed and the action value of the load to be shed; calculate the load shedding amount of the load corresponding to the sequence number of the load to be shed according to the action value of the load to be shed, and determine the load shedding amounts of the loads corresponding to the remaining sequence numbers according to the previous load shedding control measure to form the current load shedding control measure; the load shedding amount in the current load shedding control measure is: Among them, are the load shedding amounts of the j-th load in the current load shedding control measure and the previous load shedding control measure respectively, is n k is the load shedding amount of the corresponding load, n k is the serial number of the load to be shed, l k is the action value of the load to be shed, l max is the maximum load shedding amount, l min is the minimum load shedding amount, q is the total action number of each controllable load, N i is the value of the output neuron with the largest value; S4. Use the first state, the first completion flag, the first reward, the current load shedding control measure, and the second state to update the policy network parameters of the DRL agent, and use the current load shedding control measure as the new previous load shedding control measure; S5. Repeat S2-S4 until an effective load shedding control measure is obtained or a preset number of iterations is reached; S6. Repeat S1-S5 until a preset number of training rounds is reached, and use the final DRL agent to determine the corresponding relationship between the fault conditions and the load shedding control measures under the transient voltage instability mode.
2. The method according to claim 1, characterized in that The selection of the current load shedding control measure in S3 is restricted by given constraint conditions, and the given constraint conditions include: The currently selected load shedding control measure is different from all previously selected load shedding control measures; and when the previous load shedding control measure makes the transient voltage instability more serious, the currently selected load shedding control measure does not include any action in the previous load shedding control measure.
3. The method according to claim 1, characterized in that, The decision space of the DRL agent is a linear decision space, and the number of neurons in the output layer of the neural network is: N l = q × m Among them, N l is the number of neurons in the output layer of the DRL agent neural network, m is the number of controllable loads, and q is the total number of actions of each controllable load.
4. The method according to claim 1, wherein S2 includes: According to the fault condition sample and the previous load shedding control measure, simulate to obtain the bus voltage amplitude and the relative power angle value of the generator under the current decision. The first state includes the bus voltage amplitude, the relative power angle value of the generator, and the previous load shedding control measure; Judge whether the current decision is effective according to the bus voltage amplitudes of each bus, and generate the first completion flag and the first reward according to the judgment result.
5. The method according to claim 1, characterized in that, The update method of the policy network parameters of the DRL agent in S4 is: where, e TD ( ) is the temporal difference error between the policy network and the target network, θ w , θ′ w are the policy network parameters and target network parameters obtained by the w-th update respectively, Ψ is the number of samples selected from the prioritized experience replay pool, γ is the decay factor, Q′( ) is the target network value, Q() is the policy network value, s k , s k+1 are the environmental state values of the current training and the previous training respectively, a is the action corresponding to the maximum Q value, r k is the first reward, is the first completion flag, a k is the current load shedding control measure, α is the learning rate, derivative of the parameter θ.
6. The method according to claim 1, wherein The method further includes: during the operation of the power grid, when a fault condition is detected, use the load shedding control measure corresponding to the fault condition in the corresponding relationship for emergency load shedding control.
7. An emergency load shedding control system for transient voltage stability of a power system, characterized in that, For implementing the power system transient voltage stability emergency load shedding control method according to any one of claims 1-6, including: A sample selection module for selecting a fault condition sample from the sample set, where the sample set contains multiple groups of fault conditions under transient voltage instability modes; A simulation module, configured to simulate a first state, a first completion flag, and a first reward according to the fault condition sample and the previous all-load control measures; A measure selection and simulation module, configured to enable the DRL agent to select the current load shedding control measure in the linear decision space according to the previous all-load control measures, and simulate a second state according to the fault condition sample and the current load shedding control measure; An update module, configured to update the policy network parameters of the DRL agent by using the first state, the first completion flag, the first reward, the current load shedding control measure, and the second state, and use the current load shedding control measure as the new previous all-load control measure; A repeated execution module, configured to repeatedly execute the simulation module, the measure selection and simulation module, and the update module until an effective load shedding control measure is obtained or a preset number of iterations is reached; A decision-making module, configured to repeatedly execute the sample selection module, the simulation module, the measure selection and simulation module, the update module, and the repeated execution module until a preset number of training rounds is reached, and use the final DRL agent to decide the corresponding relationship between the fault condition and the load shedding control measure in the transient voltage instability mode.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the emergency load shedding control method for power system transient voltage stability according to any one of claims 1-6.
Citation Information
Patent Citations
Interruptible load optimization method based on deep reinforcement learning
CN111428903A
Construction method of dominant instability mode discrimination model and dominant instability mode discrimination method
CN112215722A