Online updating method for power system safety and stability emergency control strategy and related device

By combining deep reinforcement learning with a wide-area measurement system, an intelligent agent is constructed to achieve online generation and updating of emergency control strategies for power systems, solving the problems of slow computing speed and mismatch in existing technologies and improving the stability and adaptability of power systems.

CN119674997BActive Publication Date: 2025-10-10ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411790368.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-10-10
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

The existing power system emergency control decision-making method cannot meet the increasingly complex and changeable operation modes of modern power systems. It requires high computing speed and is prone to mismatch. There is a need for an emergency control decision-making method that is suitable for online application and is not affected by the complexity of the model.

Method used

Deep reinforcement learning methods are used to train deep neural networks. Combined with wide-area measurement systems and transient stability simulation, an intelligent agent is constructed. By online generating and updating emergency control strategies for power systems, reinforcement learning algorithms are used to optimize emergency control measures.

Benefits of technology

It improves the adaptability and computational efficiency of emergency control strategies, reduces conservatism, reduces labor costs, speeds up stabilization and control decision-making, and promotes the development of power systems towards intelligence and automation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119674997B_ABST
    Figure CN119674997B_ABST
Patent Text Reader

Abstract

The application discloses a power system safety and stability emergency control strategy online updating method and related device, comprising: constructing an emergency control strategy corresponding to an instability scene of a stability control system, and saving the emergency control strategy to an offline strategy table of the stability control system; using a reinforcement learning method, constructing an agent for online generation of stability control measures of the power system, when the agent is applied online, obtaining an online operation mode of the power system by a wide-area measurement system according to a preset period, and comparing the online operation mode with an original typical operation mode, if the difference exceeds a certain threshold, checking whether the control strategy in the offline strategy table is effective; if the strategy is invalid, using the agent to iteratively update the strategy for the mismatched scene; further, after checking the updated strategy, the strategy is issued to each stability control site. The application trains the agent through the reinforcement learning method, so that the agent can quickly obtain effective emergency control measures, and the strategy table is updated online.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power systems, and in particular to an online updating method for emergency control strategies of power system security and stability and a related device. BACKGROUND

[0002] In recent years, new energy power generation such as wind power and photovoltaic power has developed rapidly. Under these new trends, the scale of power systems is becoming increasingly large, the operating conditions are becoming more variable, and the dynamic behavior of power systems is becoming more and more complex, making it increasingly difficult to maintain the safe and stable operation of power systems.

[0003] When a fault occurs or a disturbance is received in a power grid, the system operating state will change. In most cases, the power system after fault removal will gradually recover to a new balance point for stable operation. However, when the disturbance is too large, the system state can change dramatically, and if effective security and stability control measures are not taken in time, the stability of the power grid can be destroyed or even the system can collapse. To ensure the safe and stable operation of the power system under various accident disturbances, the security and stability standards of the power system under large disturbances are divided into three levels, and three defense lines are set accordingly. Emergency control belongs to the second defense line of the power system, and is the key to reducing the risk of system instability and ensuring the stable operation of the system as a whole.

[0004] Currently, there are three application modes for emergency control of power systems: "offline pre-decision, real-time matching", "online pre-decision, real-time matching", and "real-time decision, real-time control". The "offline pre-decision, real-time matching" mode is widely used. It refers to obtaining possible combinations of system operating modes and anticipated faults in an offline state, calculating emergency control measures that can restore safe operation for each combination, and then forming an emergency control strategy table by combining the combinations of system operating modes and anticipated faults and their corresponding emergency control measures. This method has no requirements for decision-making time, but the strategies developed are generally conservative and prone to mismatching; the strategy table of the "online pre-decision, real-time matching" mode is not fixed, and the system obtains power grid information every certain period of time, automatically performs an emergency control pre-decision, and forms a new strategy table. This implementation method tracks the system operating conditions in real time and constantly refreshes the strategy table, greatly improving the robustness to system operating condition changes and significantly reducing the mismatch probability; compared with the offline pre-decision and online pre-decision methods, the measures obtained by real-time decision are closer to the actual situation of the system and the control is more effective, but both methods have very high requirements for calculation speed.

[0005] In summary, currently used emergency control decision-making methods for power systems are unable to meet the increasingly complex and volatile operating conditions of modern power systems. There is a need for emergency control decision-making methods that are suitable for online applications and are not affected by model complexity. Therefore, we consider applying deep reinforcement learning methods, combined with traditional pre-decision models, to develop an online update method for power system emergency control strategies that is computationally efficient and can accurately determine control measures. Summary of the Invention

[0006] The present application provides an online updating method and related device for the emergency control strategy of power system safety and stability, which is used to train a deep neural network through reinforcement learning method, so that it can quickly obtain effective emergency control measures according to the current state of the power system and preset faults, and fully avoid the conservatism of the obtained strategy, thereby realizing the online generation of the emergency control strategy of the power system and the online updating of the strategy table.

[0007] In view of this, the first aspect of the present application provides a method for online updating of an emergency control strategy for safety and stability of a power system, the method comprising:

[0008] Based on transient stability simulation, construct an emergency control strategy corresponding to the stability control system in an instability scenario, and save the emergency control strategy to the offline strategy table of the stability control system;

[0009] Constructing an intelligent agent for online generation of stabilization measures for the power system. When the intelligent agent is applied online, the online operation mode of the power system is obtained according to a preset period through a wide-area measurement system.

[0010] Comparing the online operation mode with a preset typical operation mode and obtaining a difference; when the difference is not less than a preset threshold, performing a transient stability simulation on the instability scenario and analyzing the control status of the emergency control strategy of the offline strategy table;

[0011] When the control situation is mismatch, the intelligent agent generates an emergency control action corresponding to the instability scenario when the mismatch occurs and saves it to the offline strategy table, thereby updating the emergency control strategy of the offline strategy table;

[0012] When the control situation is that there is no mismatch, the intelligent agent generates an emergency control action corresponding to the instability scenario when there is no mismatch, and compares the control effect of the emergency control action with the control effect of the emergency control strategy, and determines whether to update the emergency control strategy of the offline strategy table through the emergency control action based on the control effect.

[0013] Optionally, the method further includes: sending the updated offline strategy table and the online operation mode to each control site in the stabilization control system.

[0014] Optionally, constructing an emergency control strategy corresponding to an instability scenario of a stabilization control system based on transient stability simulation, and saving the emergency control strategy to an offline strategy table of the stabilization control system includes:

[0015] Obtain the topological relationship and grid parameters between the nodes of the power system where the stability control system is located to establish a power system simulation model;

[0016] generating a set of anticipated faults including several fault scenarios based on the power system simulation model;

[0017] Under several working conditions, transient stability simulation of the power system is performed according to each of the fault scenarios, and the fault scenario corresponding to the instability of the power system is obtained to obtain instability scenario samples;

[0018] Based on operating experience and in combination with simulation tools, an emergency control strategy corresponding to the instability scenario is generated and verified, and the verified emergency control strategy is saved in the offline strategy table of the stability control system.

[0019] Optionally, the constructing of an intelligent agent for online generation of stabilization measures for a power system includes:

[0020] A deep neural network is written as an intelligent agent, and a reinforcement learning method is used to interact with the power system simulation model to train the intelligent agent to obtain a model of the intelligent agent; the intelligent agent is tested, and the intelligent agent that passes the test is put into online application of the power system.

[0021] Optionally, the step of programming a deep neural network as an intelligent agent, using a reinforcement learning method to interact with the power system simulation model, training the intelligent agent, and obtaining an intelligent agent model; testing the intelligent agent, and putting the tested intelligent agent into online application in the power system includes:

[0022] Initializing the intelligent agent and the power system simulation model, and setting training parameters;

[0023] When making a decision at each step in each round of training, steps S211 to S214 are executed:

[0024] S211. Performing transient stability simulation based on the power system simulation model to obtain a state observation of the power system;

[0025] S212, inputting the state observation into the agent, so that the agent returns an emergency control action;

[0026] S213: Input the emergency control action into the power system simulation model to obtain a reward function and an observation state at the next moment, and store them in an experience pool;

[0027] S214: updating the policy network and the value network, and updating the state observation value, through a reinforcement learning algorithm according to the state observation value, the emergency control action, and the reward function;

[0028] When the training is completed, the agent model is saved and the stabilization measures are output;

[0029] The stabilization measures are tested by using the fault scenario as a test scenario, and when the test passes, the intelligent agent is put into online application of the power system.

[0030] Optionally, it is characterized in that the reinforcement learning method is: Actor-Critic algorithm.

[0031] Optionally, comparing the online operation mode with a preset typical operation mode and obtaining a difference, and when the difference is not less than a preset threshold, performing a transient stability simulation on the instability scenario and analyzing the control status of the emergency control strategy of the offline strategy table, includes:

[0032] Comparing the generator output and load of each node of the power system in the online operation mode with the generator output and load of each node of the power system in the preset typical operation mode, and calculating the Euclidean distance value;

[0033] When the Euclidean distance value is not less than a preset threshold, a transient stability simulation is performed on the instability scenario to determine whether a mismatch occurs in the emergency control strategy of the offline strategy table during operation of the power system.

[0034] Optionally, comparing the emergency control action with the control effect of the emergency control strategy includes:

[0035] The total amount of generator and load shedding in the power system under the emergency control action is compared with the total amount of generator and load shedding in the power system under the emergency control strategy.

[0036] Optionally, the determining, according to the control effect, whether to update the emergency control strategy in the offline strategy table through the emergency control action includes:

[0037] If both the emergency control action and the emergency control strategy can restore the power system to stability after a fault, but the control cost of the emergency control action is a preset percentage of the emergency control strategy, then the emergency control action is determined to be superior to the emergency control strategy, and the emergency control strategy of the offline strategy table is updated through the emergency control action.

[0038] A second aspect of the present application provides an online updating system for emergency control strategies for power system security and stability, the system comprising:

[0039] A first construction unit is configured to construct an emergency control strategy corresponding to an instability scenario of a stabilization control system based on a transient stability simulation, and save the emergency control strategy to an offline strategy table of the stabilization control system;

[0040] A second construction unit is configured to construct an intelligent agent for online generation of stabilization measures for the power system. When the intelligent agent is applied online, the intelligent agent obtains the online operation mode of the power system according to a preset period through the wide-area measurement system.

[0041] an analyzing unit, configured to compare the online operation mode with a preset typical operation mode and obtain a difference; when the difference is not less than a preset threshold, perform a transient stability simulation on the instability scenario and analyze the control status of the emergency control strategy of the offline strategy table;

[0042] a first updating unit configured to, when a mismatch occurs in the control situation, generate, by the agent, an emergency control action for the instability scenario corresponding to the mismatch and save the action to the offline strategy table, thereby updating the emergency control strategy in the offline strategy table;

[0043] The second updating unit is used to generate an emergency control action corresponding to the instability scenario when there is no mismatch through the intelligent agent when the control situation is that there is no mismatch, and compare the control effect of the emergency control action with the control effect of the emergency control strategy, and determine whether to update the emergency control strategy of the offline strategy table through the emergency control action according to the control effect.

[0044] Optionally, the method further includes: a sending unit configured to:

[0045] A third aspect of the present application provides an online update device for an emergency control strategy for safety and stability of a power system, the device comprising a processor and a memory:

[0046] The memory is used to store program code and transmit the program code to the processor;

[0047] The processor is used to execute the steps of the online updating method of the power system security and stability emergency control strategy as described in the first aspect according to the instructions in the program code.

[0048] In a fourth aspect, the present application provides a computer-readable storage medium for storing program code, wherein the program code is used to execute the online update method for the emergency control strategy for safety and stability of the power system described in the first aspect.

[0049] It can be seen from the above technical solutions that this application has the following advantages:

[0050] The present application provides an online updating method for emergency control strategies for power system safety and stability, which can make full use of the real-time operating mode of the power system obtained by the wide-area measurement system, compare its difference with the original typical operating mode, and update the strategy table if it exceeds a certain threshold. This helps to improve the adaptability of the emergency control strategy table to changes in the operating mode of the power system, so that the stability control strategy can better help the power system to restore stability after a fault; and through the deep reinforcement learning method, an intelligent agent is trained to make stability control decisions, which can reduce the dependence on manually adjusted strategy tables, reduce labor costs, and greatly improve the speed of stability control decisions, and can fully avoid the conservatism of the obtained strategy; finally, the proposal and application of this method promotes the development of power systems towards intelligence and automation, and provides strong technical support for the construction of new power systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of an online updating method for an emergency control strategy for power system safety and stability provided in an embodiment of the present application;

[0052] Figure 2 This is a topology diagram of the selected IEEE39 node test system provided in the embodiment of the present application, in which the generator and the removable load are specially marked;

[0053] Figure 3 A comparison curve of the maximum power angle difference of the generator after taking different control measures provided in the embodiment of the present application;

[0054] Figure 4 This is a structural diagram of an online update system for emergency control strategies for power system safety and stability provided in an embodiment of the present application. DETAILED DESCRIPTION

[0055] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0056] See also Figure 1 , an embodiment of the present application provides an online updating method for a power system safety and stability emergency control strategy, comprising:

[0057] Step 101: Based on transient stability simulation, construct an emergency control strategy corresponding to the stability control system in an instability scenario, and save the emergency control strategy to an offline strategy table of the stability control system.

[0058] In one embodiment, step 101 includes the following steps:

[0059] Obtain the topological relationship and grid parameters between the nodes of the power system where the stability control system is located to establish a power system simulation model;

[0060] Generate a set of expected faults containing several fault scenarios based on the power system simulation model;

[0061] Under several working conditions, the transient stability of the power system is simulated according to various fault scenarios to obtain the fault scenarios corresponding to the instability of the power system and obtain instability scenario samples;

[0062] Based on operating experience and combined with simulation tools, emergency control strategies corresponding to instability scenarios are generated and verified, and the verified emergency control strategies are saved to the offline strategy table of the stability control system.

[0063] It should be noted that for step 101, the fault scenarios include AC line faults such as three-phase short circuit to ground, two-phase short circuit to ground, two-phase short circuit, single-phase short circuit to ground, and DC line faults such as single-pole lockout, double-pole lockout, and commutation failure; the lines where faults may occur include AC lines with higher voltage levels in the power system grid and all DC lines; the possible fault locations on each line include 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, and 90% from the near end of the line. In this step, an expected fault set is generated according to the type of the above fault scenario, for example, Figure 2 As shown in the IEEE 39-node system, there are 100 generators, 39 nodes, and 46 AC lines. The final generated expected fault set includes A fault.

[0064] Transient stability simulation refers to the use of simulation technology to evaluate the stability of the power system after a sudden fault or disturbance. In this application, transient stability simulation is used to evaluate the stability of the power system after a fault scenario in the expected fault set. During the simulation, the power system is set to fail at 0.1s, and 5 different fault removal times are set for each fault scenario, namely, removal at 0.1s, 0.14s, 0.20s, 0.24s or 0.30s respectively. Combined with the obtained fault set, under each working condition, there are a total of A simulation scene.

[0065] The generation of the offline strategy table first needs to determine various emergency situations (such as line faults, generator unit failures, system frequency or voltage abnormalities, etc.) and establish a system model based on the topology of the power grid. Through simulation tools, different emergency situations are analyzed to identify weak links in the system and design corresponding emergency control strategies. These strategies include fast load shedding, generator dispatch adjustment, power flow redistribution, and frequency and voltage regulation. Subsequently, based on the simulation results, a strategy table is generated to clearly define the triggering conditions, specific control measures, and priorities for each emergency situation, ensuring that the strategies can be effectively implemented in actual operations. Finally, the reliability of the strategies is ensured through verification and optimization, and operators are trained to ensure that they can respond quickly in emergency situations. In the embodiments of the present application, the strategy table for some fault scenarios is as shown in Table 1:

[0066] Table 1 Offline strategy table

[0067]

[0068] Step 102, construct an agent for online generation of stability control measures of a power system. When the agent is applied online, the online operation mode of the power system is obtained through the wide-area measurement system at a preset period.

[0069] In one embodiment, step 102 includes the following steps:

[0070] A deep neural network is written as an agent, which is trained using a reinforcement learning method and a power system simulation model to obtain a model of the agent; the agent is tested, and the agent that passes the test is put into online application of the power system;

[0071] Specifically:

[0072] Initialize the agent and the power system simulation model, and set the training parameters;

[0073] When making a decision at each step in each round of training, steps S211 to S214 are performed:

[0074] S211, perform transient stability simulation based on the power system simulation model to obtain state observations of the power system;

[0075] S212, input the state observations to the agent, so that the agent returns an emergency control action;

[0076] S213, input the emergency control action to the power system simulation model to obtain a reward function and an observation state at the next time, and store them in an experience pool;

[0077] S214, based on the state observation, emergency control action and reward function, update the policy network and value network, and update the state observation through the reinforcement learning algorithm;

[0078] When the training is completed, the agent model is saved and the stabilization measures are output;

[0079] The stabilization measures are tested using fault scenarios as test scenarios. When the test passes, the intelligent agent is put into online application in the power system.

[0080] When the intelligent agent is applied online, the online operation mode of the power system is obtained according to the preset cycle through the wide-area measurement system.

[0081] It's important to note that the Wide Area Measurement System (WAMS) refers to a new generation of dynamic power grid monitoring and control systems based on synchronized phasor technology. Featuring remote, high-precision synchronized phasor measurement, high-speed communication, and rapid response, the WAMS is ideally suited for real-time monitoring of dynamic processes in large-span power grids.

[0082] Furthermore, it should be noted that the state observations, emergency control actions, and reward functions of steps S211-S214 are as follows:

[0083] The state observation of the power system refers to:

[0084] The voltage amplitude and phase of each node in the power system. In transient stability analysis, the voltage characteristics of each node in the system can reflect the transient stability of the system. The voltage amplitude and phase of each node are a time series, which includes observation data before, during, and before control. If the system has N nodes, the state observation S can be expressed as:

[0085] ;

[0086] Where, and is the voltage amplitude and phase angle before the fault of the i-th node, ; and is the voltage amplitude and phase angle during the fault process of the i-th node, each of which is a time series with a length of 10, , ; and is the voltage amplitude and phase angle of the i-th node after the fault is cleared and before this control, each of which is a time series with a length of 10. , In this embodiment, the test system contains 39 nodes in total, so the state observation is a 39-dimensional vector. .

[0087] The emergency control action refers to:

[0088] The position and percentage of generator or load to be cut off in the control process given by the neural network, if the system has n power stations and m nodes of cuttable load, the emergency control action A can be expressed as:

[0089] ;

[0090] In the formula, is the cut-off percentage of the i-th generator, , is the cut-off percentage of the i-th cuttable load, , the control variables are discrete values, which can be selected from 0, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1. In this embodiment, the test system contains 10 generators in total, and nodes 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 are selected as cuttable loads, so the control action is a 28-dimensional vector.

[0091] The reward function refers to:

[0092] The reward function is calculated according to the state observation of the power system and the control variables, and there are different calculation methods according to the transient stability of the power system after control, which are as follows:

[0093] If the power system recovers stability after control, the reward function is positive, and the smaller the control variable, the larger the reward function, if the system has n power stations and m nodes of cuttable load, the reward function at this time is:

[0094] ;

[0095] In the formula, is the cut-off percentage of the i-th generator, , is the cut-off percentage of the i-th cuttable load, .

[0096] If the power system is still unstable after control, the reward function is negative, since the degree of instability of the power system can be measured by the speed of change of the power angle of the accelerating group, so the reward function is taken as the reciprocal of the absolute value of the maximum value of the power angle acceleration of each generator, i.e. the reciprocal of the absolute value of the maximum value of the slope of the power angle curve of each generator, the smaller the reward function, the greater the degree of instability of the power system, as follows:

[0097] ;

[0098] Where, is the power angle of the i-th generator, , n is the total number of generators.

[0099] Among them, the reinforcement learning method used in this application is: Actor-Critic algorithm.

[0100] It should be noted that the reinforcement learning algorithm of the present application adopts the Actor-Critic algorithm, which can be divided into two parts: Actor (policy network) and Critic (value network). What the Actor has to do is to interact with the environment and learn a better strategy using policy gradient under the guidance of the Critic value function. What the Critic has to do is to learn a value function through the data collected by the interaction between the Actor and the environment. This value function will be used to judge what actions are good and what actions are not good in the current state, and then help the Actor to update the strategy. Through this algorithm, it is possible to converge to the optimal strategy more quickly and obtain a neural network that can accurately generate emergency control actions. Preferably, the Actor-Critic algorithm used in the reinforcement learning can also be replaced by the TRPO algorithm, the PPO algorithm, the SAC algorithm, etc.

[0101] Step 103: Compare the online operation mode with the preset typical operation mode and obtain the difference. When the difference is not less than the preset threshold, perform transient stability simulation on the instability scenario and analyze the control status of the emergency control strategy in the offline strategy table.

[0102] In one embodiment, step 103 includes the following steps:

[0103] Compare the generator output and load of each node in the power system in the online operation mode with the generator output and load of each node in the power system in the preset typical operation mode, and calculate the Euclidean distance value;

[0104] When the Euclidean distance value is not less than the preset threshold, a transient stability simulation is performed on the instability scenario to determine whether a mismatch occurs in the emergency control strategy of the offline strategy table during the operation of the power system.

[0105] Regarding step 103, it should be noted that in the online application of this embodiment of the present application, the online power system operating mode is obtained through a wide-area measurement system at a cycle of 5-10 minutes. This is then compared with the original typical operating mode. If the difference is not less than a preset threshold, a transient simulation is performed for the instability scenario (fault scenarios within the expected fault set) to check whether the emergency control strategy in the offline strategy table is effective. Furthermore, if a mismatch is found, an alarm is issued and relevant information about the mismatch scenario is stored in a historical database.

[0106] Among them, the operation mode of the power system includes key parameters such as node voltage amplitude and phase angle, current of transmission lines and buses, active and reactive power, system frequency, etc. This information reflects the operation status of the power grid, such as power distribution, load changes, frequency fluctuations, etc. In addition, it also includes system topology, such as line switch status, generator and load operation, as well as equipment information such as power flow distribution, transformer load and tap position. When comparing with the original typical operation mode, the main comparison is made on the generator output and load in the power system. If the Euclidean distance between the current system operation mode and the original typical operation mode is not less than the preset threshold, it is necessary to check whether the emergency control strategy in the offline strategy table is valid:

[0107] ;

[0108] Where, It is the generator output and load of each node in the system under the current system operation mode, is the generator output and load of each node in the system under the original typical operation mode, both of which are a length of Array of The threshold value can be set to 0.1.

[0109] In this embodiment, the selected IEEE-39 node system includes 18 loads and 10 generators, so the array length is , it can be calculated that under a certain operation mode, the Euclidean distance between the current operation mode of the power system and the original typical operation mode reaches 0.627, which is greater than , so the policy table needs to be updated.

[0110] Mismatch in power system emergency control refers to the discrepancy between the actual operating state and the intended control strategy during power system operation, due to various reasons (such as equipment failures and load changes). This mismatch can lead to a decrease in system stability or even cause a power system failure or collapse. In these cases, emergency control strategies must be re-established based on the system state to ensure power system stability. In this embodiment, no mismatch was observed.

[0111] Step 104: When a control mismatch occurs, the intelligent agent generates an emergency control action for the instability scenario corresponding to the mismatch and saves it to the offline strategy table, thereby updating the emergency control strategy of the offline strategy table.

[0112] It should be noted that when the control situation is a mismatch, based on the offline strategy table, the state observation quantity corresponding to the instability scenario simulation process when the mismatch occurs is extracted, and the emergency control action is obtained after input into the intelligent agent. After verification, the power system operation mode and control strategy are written into the strategy table, thereby updating the emergency control strategy of the offline strategy table.

[0113] Step 105: When the control situation shows that there is no mismatch, the intelligent agent generates an emergency control action corresponding to the instability scenario when there is no mismatch, and compares the control effect of the emergency control action with the emergency control strategy. According to the control effect, it is determined whether to update the emergency control strategy of the offline strategy table through the emergency control action.

[0114] In one embodiment, step 105 compares the control effects of the emergency control action and the emergency control strategy, and determines whether to update the emergency control strategy in the offline strategy table through the emergency control action according to the control effect, including the following steps:

[0115] Compare the total amount of power system generator and load shedding under emergency control action with the total amount of power system generator and load shedding under emergency control strategy.

[0116] If both the emergency control action and the emergency control strategy can restore the power system to stability after a fault, but the control cost of the emergency control action is a preset percentage of the emergency control strategy, then the emergency control action is determined to be superior to the emergency control strategy, and the emergency control strategy in the offline strategy table is updated through the emergency control action.

[0117] It can be understood that for the fault scenario without mismatch, the agent generates an emergency control action, and compares its control cost and control effect with the offline strategy. If the online strategy is much better than the offline strategy, it is written into the strategy table, otherwise the strategy table remains unchanged.

[0118] The online generation strategy is far superior to the offline strategy if both the online and offline strategies can restore the power system to stability after a fault, but the control cost of the online strategy is less than 75% of that of the offline strategy. The control cost is the total amount of generator and load shedding, that is:

[0119]

[0120] Where, is the removal amount of the i-th generator, , is the shedding amount of the i-th shearable load, .

[0121] In this embodiment, there is no mismatch in the offline strategy table, but some of the control measures can be further modified to achieve the optimal value and reduce the conservatism of the control measures:

[0122] Table 2

[0123]

[0124] As shown in Table 2, we can see that after the system state changes, the amount of machine switching required by the strategy generated by the agent is smaller than that of the original offline strategy table. Figure 3 This is a curve comparison of the maximum power angle difference of the generator after the original offline strategy and the intelligent agent generation strategy are adopted for stabilization after a three-phase short circuit fault occurs on the node 28-node 29 line. It can be seen that both the original offline strategy and the intelligent agent generation strategy can restore the stability of the power system. The intelligent agent generation strategy achieves a smaller maximum power angle difference of the generator with a smaller amount of machine cutting.

[0125] In one embodiment, the online updating method of the power system safety and stability emergency control strategy of the present application further includes:

[0126] The updated offline strategy table and online operation mode are distributed to each control site in the stabilization control system.

[0127] It should be noted that in this embodiment, after further verification of the updated offline strategy table, it is sent to each control site. When the offline strategy table is sent, it needs to include both the operating mode and the control strategy for each fault scenario. The control sites mainly include each switchable generator site and switchable load site.

[0128] In summary, the embodiment of the present application provides an online update method for emergency control strategies for power system safety and stability, including the following steps: obtaining the power system topology where the stability control system is located, generating an expected fault set, performing transient stability simulation on the fault scenarios in the expected fault set under multiple working conditions, and screening out transient instability samples; generating an offline strategy table corresponding to the instability scenario based on historical experience; using a reinforcement learning method to interact with the power system simulation model and train a stability control intelligent agent; obtaining power system status information through a wide-area measurement system, and comparing it with the original typical operating mode. If the difference exceeds a certain threshold, performing transient simulation on the fault scenarios in the expected fault set to check whether the control strategy in the offline strategy table is effective; extracting state observations in the mismatch fault scenario, inputting them into the intelligent agent to obtain emergency control actions, and then updating the strategy table after verification through simulation. The updated strategy table is verified and sent to each stability control site. The technical solution provided in the embodiment of the present application trains an intelligent agent through a reinforcement learning method, so that it can quickly obtain effective emergency control measures based on the current state of the system, and can fully avoid the conservatism of the obtained strategy, update the strategy table online, and improve the adaptability of the strategy table to the current operation mode of the system.

[0129] The above is an online update method for the emergency control strategy of safety and stability of a power system provided in an embodiment of the present application. The following is an online update system for the emergency control strategy of safety and stability of a power system provided in an embodiment of the present application.

[0130] See also Figure 4 , an embodiment of the present application provides an online update system for an emergency control strategy for a power system safety and stability, comprising:

[0131] The first constructing unit 201 is configured to construct an emergency control strategy corresponding to an instability scenario of the stabilization control system based on transient stability simulation, and save the emergency control strategy to an offline strategy table of the stabilization control system.

[0132] The second construction unit 202 is used to construct an intelligent agent for online generation of stabilization measures for the power system. When the intelligent agent is applied online, the online operation mode of the power system is obtained according to a preset period through the wide area measurement system.

[0133] The analysis unit 203 is used to compare the online operation mode with the preset typical operation mode and obtain the difference. When the difference is not less than the preset threshold, a transient stability simulation is performed on the instability scenario, and the control status of the emergency control strategy of the offline strategy table is analyzed.

[0134] The first updating unit 204 is configured to, when the control condition is mismatch, generate an emergency control action corresponding to the unstable scene when mismatch occurs by the agent and save the emergency control action to the offline policy table, so as to update the emergency control strategy of the offline policy table.

[0135] The second updating unit 205 is configured to, when the control condition is no mismatch, generate an emergency control action corresponding to the unstable scene when no mismatch occurs by the agent, compare the control effects of the emergency control action and the emergency control strategy, and determine whether to update the emergency control strategy of the offline policy table by the emergency control action according to the control effects.

[0136] In an embodiment, the power system safety and stability emergency control strategy online updating system also comprises a delivery unit configured to:

[0137] deliver the updated offline policy table and the online operation mode to each control site in the stability control system.

[0138] Further, an embodiment of the present application further provides a power system safety and stability emergency control strategy online updating device, which comprises a processor and a memory:

[0139] The memory is configured to store program code and transmit the program code to the processor.

[0140] The processor is configured to execute the steps of the power system safety and stability emergency control strategy online updating method according to the instructions in the program code.

[0141] Further, an embodiment of the present application further provides a computer readable storage medium configured to store program code, and the program code is configured to execute the power system safety and stability emergency control strategy online updating method.

[0142] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and the unit described above can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0143] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or apparatus.

[0144] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.

[0145] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0146] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0147] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0148] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name: Read-Only Memory, English abbreviation: ROM), random access memory (full name: Random Access Memory, English abbreviation: RAM), disk or optical disk, and other media that can store program code.

[0149] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for online updating of emergency control strategy for power system security and stability, characterized in that: include: Based on transient stability simulation, construct an emergency control strategy corresponding to the stability control system in an instability scenario, and save the emergency control strategy to the offline strategy table of the stability control system; Constructing an intelligent agent for online generation of stabilization measures for the power system. When the intelligent agent is applied online, the online operation mode of the power system is obtained according to a preset period through a wide-area measurement system. Comparing the online operation mode with a preset typical operation mode and obtaining a difference; when the difference is not less than a preset threshold, performing a transient stability simulation on the instability scenario and analyzing the control status of the emergency control strategy of the offline strategy table; When the control situation is mismatch, the intelligent agent generates an emergency control action corresponding to the instability scenario when the mismatch occurs and saves it to the offline strategy table, thereby updating the emergency control strategy of the offline strategy table; When the control situation is that there is no mismatch, the intelligent agent generates an emergency control action corresponding to the instability scenario when there is no mismatch, and compares the control effect of the emergency control action with the control effect of the emergency control strategy, and determines whether to update the emergency control strategy of the offline strategy table through the emergency control action based on the control effect.

2. The online updating method for emergency control strategy of power system security and stability according to claim 1 is characterized in that: Also includes: The updated offline strategy table and the online operation mode are sent to each control site in the stabilization control system.

3. The online updating method for the power system safety and stability emergency control strategy according to claim 1 is characterized in that: The method of constructing an emergency control strategy corresponding to an instability scenario of a stabilization control system based on transient stability simulation and saving the emergency control strategy to an offline strategy table of the stabilization control system includes: Obtain the topological relationship and grid parameters between the nodes of the power system where the stability control system is located to establish a power system simulation model; generating a set of anticipated faults including several fault scenarios based on the power system simulation model; Under several working conditions, transient stability simulation of the power system is performed according to each of the fault scenarios, and the fault scenario corresponding to the instability of the power system is obtained to obtain instability scenario samples; Based on operating experience and in combination with simulation tools, an emergency control strategy corresponding to the instability scenario is generated and verified, and the verified emergency control strategy is saved in the offline strategy table of the stability control system.

4. The online updating method for the power system safety and stability emergency control strategy according to claim 3 is characterized in that: The intelligent agent for online generation of stabilization measures for a power system comprises: A deep neural network is written as an intelligent agent, and a reinforcement learning method is used to interact with the power system simulation model to train the intelligent agent to obtain a model of the intelligent agent; the intelligent agent is tested, and the intelligent agent that passes the test is put into online application of the power system.

5. The online updating method for emergency control strategy of power system security and stability according to claim 4 is characterized in that: The deep neural network is written as an intelligent agent, and a reinforcement learning method is used to interact with the power system simulation model to train the intelligent agent to obtain a model of the intelligent agent; Testing the intelligent agent and putting the intelligent agent that has passed the test into online application in the power system, including: Initializing the intelligent agent and the power system simulation model, and setting training parameters; When making a decision at each step in each round of training, steps S211 to S214 are executed: S211. Performing transient stability simulation based on the power system simulation model to obtain a state observation of the power system; S212, inputting the state observation into the agent, so that the agent returns an emergency control action; S213: Input the emergency control action into the power system simulation model to obtain a reward function and an observation state at the next moment, and store them in an experience pool; S214: updating the policy network and the value network, and updating the state observation value, through a reinforcement learning algorithm according to the state observation value, the emergency control action, and the reward function; When the training is completed, the agent model is saved and the stabilization measures are output; The stabilization measures are tested by using the fault scenario as a test scenario, and when the test passes, the intelligent agent is put into online application of the power system.

6. The online updating method for the power system security and stability emergency control strategy according to claim 4 or 5, characterized in that: The reinforcement learning method is: Actor-Critic algorithm.

7. The method for online updating of emergency control strategy for power system security and stability according to claim 1, characterized in that: The online operation mode is compared with a preset typical operation mode to obtain a difference. When the difference is not less than a preset threshold, a transient stability simulation is performed on the instability scenario, and a control condition of the emergency control strategy of the offline strategy table is analyzed, including: Comparing the generator output and load of each node of the power system in the online operation mode with the generator output and load of each node of the power system in the preset typical operation mode, and calculating the Euclidean distance value; When the Euclidean distance value is not less than a preset threshold, a transient stability simulation is performed on the instability scenario to determine whether a mismatch occurs in the emergency control strategy of the offline strategy table during operation of the power system.

8. The method for online updating of emergency control strategy for power system security and stability according to claim 1, characterized in that: The comparing the emergency control action with the control effect of the emergency control strategy includes: The total amount of generator and load shedding in the power system under the emergency control action is compared with the total amount of generator and load shedding in the power system under the emergency control strategy.

9. The method for online updating of the power system security and stability emergency control strategy according to claim 8, characterized in that: The determining, according to the control effect, whether to update the emergency control strategy of the offline strategy table through the emergency control action includes: If both the emergency control action and the emergency control strategy can restore the power system to stability after a fault, but the control cost of the emergency control action is a preset percentage of the emergency control strategy, then the emergency control action is determined to be superior to the emergency control strategy, and the emergency control strategy of the offline strategy table is updated through the emergency control action.

10. An online update system for emergency control strategy of power system security and stability, characterized by: include: A first construction unit is configured to construct an emergency control strategy corresponding to an instability scenario of a stabilization control system based on a transient stability simulation, and save the emergency control strategy to an offline strategy table of the stabilization control system; A second construction unit is configured to construct an intelligent agent for online generation of stabilization measures for the power system. When the intelligent agent is applied online, the intelligent agent obtains the online operation mode of the power system according to a preset period through the wide-area measurement system. an analyzing unit, configured to compare the online operation mode with a preset typical operation mode and obtain a difference; when the difference is not less than a preset threshold, perform a transient stability simulation on the instability scenario and analyze the control status of the emergency control strategy of the offline strategy table; a first updating unit configured to, when a mismatch occurs in the control situation, generate, by the agent, an emergency control action for the instability scenario corresponding to the mismatch and save the action to the offline strategy table, thereby updating the emergency control strategy in the offline strategy table; The second updating unit is used to generate an emergency control action corresponding to the instability scenario when there is no mismatch through the intelligent agent when the control situation is that there is no mismatch, and compare the control effect of the emergency control action with the control effect of the emergency control strategy, and determine whether to update the emergency control strategy of the offline strategy table through the emergency control action according to the control effect.

11. The online update system for emergency control strategy of power system security and stability according to claim 10 is characterized in that: Also includes: The issuing unit is used to: The updated offline strategy table and the online operation mode are sent to each control site in the stabilization control system.

12. An online update device for emergency control strategy of power system safety and stability, characterized by: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the online updating method of the power system security and stability emergency control strategy according to any one of claims 1 to 9 according to the instructions in the program code.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the online updating method of the power system security and stability emergency control strategy according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Online setting method of safety and stability emergency control fixed value of electric power system

    CN102638040A

  • Intelligent agent training method and pre-setting method for setting emergency generator tripping control measures

    CN114967649A