Transient stability emergency control method and device based on Rainbow reinforcement learning
By applying Rainbow reinforcement learning method in power systems, the agent uses deep reinforcement learning Rainbow algorithm decision-making and outputs emergency control strategies to solve the problem of lack of flexibility and adaptability of the existing emergency control methods in power systems, and achieves rapid response and efficient power supply.
Patent Information
- Application Number
- CN202510282798.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing emergency control methods of power systems lack flexibility and adaptability, and cannot respond quickly and effectively to sudden failures or load changes, resulting in long power outages and low power supply reliability.
The transient stable emergency control method based on Rainbow reinforcement learning is adopted to make decisions through the deep reinforcement learning Rainbow algorithm of the agent, output emergency control strategies, optimize resource configuration, and reduce power outage time.
It has achieved rapid adaptation to emergencies, optimized resource allocation, reduced power outage time, improved power supply reliability, and improved the efficiency and accuracy of emergency control strategy output.
Smart Images

Figure CN120222334A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of emergency control of power grid systems, and particularly to a transient stability emergency control method and device based on Rainbow reinforcement learning. Background Art
[0002] When facing disturbances such as sudden faults or sudden changes in load, ensuring the transient stability of the system is crucial for the reliable operation of the power system. Transient instability can lead to the loss of synchronization between generators, causing large-scale power outages and even serious economic and social consequences.
[0003] In recent decades, various methods have been used to solve the emergency control problem in power systems. Early methods include heuristic algorithms for controlling generator tripping strategies, potential energy boundary surface (PEBS) methods, boundary control of stability regions (BCU) methods, and equal area criterion (EEAC), etc. With the progress of computing power, it has been possible to perform more accurate time-domain simulations of detailed mathematical models of power systems. Therefore, the emergency control problem can be solved as an optimal control problem (OCP), and the focus of research has become how to solve the OCP problem. This method can minimize the control cost while maintaining system stability. In addition, the optimal power flow (OPF) problem with transient stability constraints can be transformed into a finite-dimensional optimization problem to make it feasible in large-scale systems. The application of sensitivity analysis methods further improves the computational efficiency of problems such as optimal load shedding. However, the method modeled as OCP faces challenges in solving the differential-algebraic equations (DAEs) that describe system dynamics. The numerical methods of OCP are generally divided into indirect methods and direct methods. Although indirect methods were widely used in the early stage, they have convergence problems in path constraint problems. Direct methods discretize the state variables and transform the problem into a non-linear programming problem (NLP), showing greater promise. Although they are effective, full discretization may lead to too high dimensions, especially in large power systems.
[0004] Currently, traditional emergency control methods rely on experience and rules, lack flexibility and adaptability, and cannot respond to emergencies in a timely and effective manner; methods of simplifying the simulation process are adopted to speed up the calculation, but it is difficult to take into account the economy of distribution network operation; most machine learning-based methods have poor convergence during model training and long training time.
[0005] Therefore, how to invent a transient stability emergency control method based on reinforcement learning, which can quickly adapt to emergencies, optimize resource allocation, reduce power outage time, and improve power supply reliability, has become an urgent problem to be solved. Summary of the Invention
[0006] To this end, the present invention provides a transient stability emergency control method and device based on Rainbow reinforcement learning, which can quickly adapt to emergencies, optimize resource allocation, reduce power outage time, improve power supply reliability, and enhance the efficiency and accuracy of the output of the emergency control strategy for the distribution network. At the same time, using deep learning technology, the agent can process high-dimensional complex data and provide more scientific decision-making support.
[0007] To achieve the above object, the present invention provides the following technical solutions: A transient stability emergency control method based on Rainbow reinforcement learning, comprising:
[0008] Training the agent by setting a training strategy to obtain a trained agent;
[0009] Monitoring the power system by setting monitoring devices to obtain monitoring data of the power system;
[0010] Obtaining a fault disturbance through the monitoring data;
[0011] Judging the transient stability of the fault disturbance through a transient stability index. If the transient stability is achieved, the power system resumes stability; if the transient stability is not achieved, the state data of the power system is obtained and emergency control processing is performed;
[0012] Inputting the state data of the power system into the trained agent; the agent makes a decision through the deep reinforcement learning Rainbow algorithm and outputs an emergency control strategy;
[0013] The power system performs emergency control processing according to the emergency control strategy and resumes stability.
[0014] As a preferred solution of a transient stability emergency control method based on Rainbow reinforcement learning, during the process of training the agent by setting a training strategy, the training steps of the agent are:
[0015] Initializing the environmental parameters according to the training objective;
[0016] The agent obtains state observations from the power system environment;
[0017] Inputting the state observations into the agent's estimation network model, and estimating and processing through the agent's estimation network model to output an emergency control action;
[0018] Interacting the emergency control action with the power system environment to obtain a feedback reward and an observation state at the next moment, and storing them in an experience library;
[0019] Updating the policy network according to the observation state, the emergency control action and the feedback reward, and updating the state observations;
[0020] The power system performs emergency control processing according to the emergency control action, and determines whether the power system resumes stability; if the power system resumes stability, subsequent processing is performed; if the power system does not resume stability, the state observation quantity is retrieved from the power system environment again, and the next iterative emergency processing is performed until the number of steps reaches the set maximum number of steps, and the iteration stops;
[0021] When the power system resumes stability, the parameters of the agent estimation network model are updated;
[0022] The number of agent training rounds is judged. If the set number of training rounds is not reached, the state observation quantity is retrieved from the power system environment again for iterative training; if the set number of training rounds is reached, the agent evaluation network model is saved and the final control strategy is output;
[0023] The final control strategy is tested through a test scenario, and the test result is output.
[0024] As a preferred solution of a transient stability emergency control method based on Rainbow reinforcement learning, in the process of initializing the environment parameters according to the training target, the environment parameters include: the number of training rounds, the maximum number of steps, the learning rate, the discount factor, and the greedy ratio.
[0025] As a preferred solution of a transient stability emergency control method based on Rainbow reinforcement learning, in the process of judging the transient stability of the fault disturbance through the transient stability index, the expression of the transient stability index is:
[0026]
[0027] In the formula, Δδ max is the maximum angle of attack difference between any two generators within a period of time.
[0028] As a preferred solution of a transient stability emergency control method based on Rainbow reinforcement learning, in the process of interacting the emergency control action with the power system environment to obtain feedback rewards, the short-term reward function is defined as: energy constraint reward function, control cost constraint reward function, control step constraint reward function, active power constraint reward function, and voltage constraint reward function; the expression of the short-term reward function is:
[0029] R F = r e + r c + r g + r v + r n
[0030] where r e is the energy constraint reward function; r c is the control cost constraint reward function; r g is the control step constraint reward function; r v is the active power constraint reward function; r n is the voltage constraint reward function.
[0031] The present invention also provides a transient stability emergency control device based on Rainbow reinforcement learning. Based on the above transient stability emergency control method based on Rainbow reinforcement learning, it includes:
[0032] An agent training module, configured to train the agent by setting a training strategy to obtain a trained agent;
[0033] A monitoring data acquisition module, configured to monitor the power system by setting monitoring devices to obtain the monitoring data of the power system;
[0034] A fault disturbance acquisition module, configured to obtain a fault disturbance through the monitoring data;
[0035] A transient stability judgment module, configured to judge the transient stability of the fault disturbance through a transient stability index. If the transient stability is satisfied, the power system returns to stability; if the transient stability is not satisfied, the state data of the power system is obtained and emergency control processing is performed;
[0036] An agent emergency control strategy output module, configured to input the state data of the power system into the trained agent; the agent makes a decision through the deep reinforcement learning Rainbow algorithm and outputs an emergency control strategy;
[0037] An emergency control processing module, configured to perform emergency control processing on the power system according to the emergency control strategy to restore stability.
[0038] As a preferred solution of a transient stability emergency control device based on Rainbow reinforcement learning, in the agent training module, the training sub-module includes:
[0039] An environment parameter initialization sub-module, configured to initialize the environment parameters according to the training objective;
[0040] A state observation quantity acquisition sub-module, configured to enable the agent to obtain state observation quantities from the power system environment;
[0041] An estimation network model processing sub-module, configured to input the state observation quantities into the agent estimation network model, and output an emergency control action through the estimation and processing of the agent estimation network model;
[0042] An interaction feedback sub-module, configured to interact the emergency control action with the power system environment, obtain a feedback reward and an observation state at the next moment, and store them in an experience library;
[0043] A policy network update sub-module, configured to update the policy network according to the observation state, the emergency control action and the feedback reward, and update the state observation quantity;
[0044] An iterative training sub-module, configured to perform emergency control processing on the power system according to the emergency control action, and determine whether the power system returns to stability; if the power system returns to stability, perform subsequent processing; if the power system does not return to stability, re-obtain the state observation quantity from the power system environment, and perform the next iterative emergency processing until the number of steps reaches the set maximum number of steps, and stop the iteration;
[0045] A model parameter update sub-module, configured to update the parameters of the intelligent agent estimation network model when the power system returns to stability;
[0046] An intelligent agent evaluation network model storage sub-module, configured to judge the number of training rounds of the intelligent agent. If the set number of training rounds is not reached, re-obtain the state observation quantity from the power system environment and perform iterative training; if the set number of training rounds is reached, save the intelligent agent evaluation network model and output the final control strategy;
[0047] A test verification sub-module, configured to test the final control strategy through a test scenario and output a test result.
[0048] As a preferred solution of a transient stability emergency control device based on Rainbow reinforcement learning, in the environment parameter initialization sub-module of the intelligent agent training module, in the process of initializing the environment parameters according to the training target, the environment parameters include: the number of training rounds, the maximum number of steps, the learning rate, the discount factor and the greedy ratio.
[0049] As a preferred solution of a transient stability emergency control device based on Rainbow reinforcement learning, in the transient stability judgment module, in the process of judging the transient stability of the fault disturbance through the transient stability index, the expression of the transient stability index is:
[0050]
[0051] In the formula, Δδ max is the maximum angle of attack difference between any two generators within a period of time.
[0052] As an optimal solution of a transient stability emergency control device based on Rainbow reinforcement learning, in the interactive feedback sub-module of the agent training module, during the process of interacting the emergency control actions with the power system environment to obtain feedback rewards, the short-term reward function is defined as: energy constraint reward function, control cost constraint reward function, control step constraint reward function, active power constraint reward function, and voltage constraint reward function; the expression of the short-term reward function is:
[0053] R F = r e + r c + r g + r v + r n
[0054] In the formula, r e is the energy constraint reward function; r c is the control cost constraint reward function; r g is the control step constraint reward function; r v is the active power constraint reward function; r n is the voltage constraint reward function.
[0055] The present invention has the following advantages: By setting a training strategy, an agent is trained to obtain a trained agent; By setting monitoring devices to monitor the power system, monitoring data of the power system is obtained; Through the monitoring data, a fault disturbance is obtained; Through a transient stability index, a transient stability judgment is made on the fault disturbance. If the transient stability is achieved, the power system returns to stability; If the transient stability is not achieved, the power system state data is obtained and emergency control processing is performed; The power system state data is input into the trained agent; The agent makes a decision through the deep reinforcement learning Rainbow algorithm and outputs an emergency control strategy; The power system performs emergency control processing according to the emergency control strategy and returns to stability. The present invention introduces a reinforcement learning method to model and output a strategy for the emergency control of the power system. The decision-making time is short and real-time control can be achieved; At the same time, as an advanced algorithm of deep reinforcement learning, the Rainbow algorithm combines multiple advanced technologies (such as prioritized experience replay, double Q-learning, distributed learning, etc.), making the learning process more efficient and stable. With the help of the Rainbow algorithm, more intelligent and faster emergency control decisions can be made. Combining advanced machine learning technologies, the deficiencies in traditional load restoration methods are solved. This method has good adaptability and intelligence, provides new ideas and a practical basis for the development of smart grids, and has broad application prospects. The present invention improves the efficiency and accuracy of the output of the emergency control strategy for the distribution network, can adapt to emergencies more quickly, optimize resource allocation, reduce power outage time, and improve power supply reliability. At the same time, using deep learning technology, the agent can process high-dimensional complex data and provide more scientific decision-making support. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, other implementation drawings can be obtained according to the provided drawings without creative efforts.
[0057] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change in the ratio relationship, or adjustment of the size should still fall within the scope covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.
[0058] Figure 1 It is a schematic flow chart of a transient stability emergency control method based on Rainbow reinforcement learning provided in Embodiment 1 of the present invention;
[0059] Figure 2 Schematic diagram of the training and specific implementation framework of a transient stability emergency control method based on Rainbow reinforcement learning provided in Embodiment 1 of the present invention;
[0060] Figure 3 Schematic diagram of the training process of the intelligent agent Rainbow reinforcement learning emergency control in a transient stability emergency control method based on Rainbow reinforcement learning provided in Embodiment 1 of the present invention;
[0061] Figure 4 Schematic diagram of the interaction relationship between the power grid environment and the intelligent agent in a transient stability emergency control method based on Rainbow reinforcement learning provided in Embodiment 1 of the present invention;
[0062] Figure 5 Schematic diagram of the update process of the estimation network in a transient stability emergency control method based on Rainbow reinforcement learning provided in Embodiment 1 of the present invention;
[0063] Figure 6 Schematic diagram of the power system network topology structure in a possible embodiment provided in Embodiment 1 of the present invention;
[0064] Figure 7 Schematic diagrams of the training results of various algorithms in a possible embodiment provided in Embodiment 1 of the present invention; among them, (a) is the training result of the DQN algorithm; (b) is the training result of the Double DQN algorithm; (c) is the training result of the Dueling DQN algorithm; (d) is the training result of the Rainbow algorithm;
[0065] Figure 8 Schematic diagram of the architecture of a transient stability emergency control device based on Rainbow reinforcement learning provided in Embodiment 2 of the present invention. Detailed implementation manners
[0066] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.
[0067] Embodiment 1
[0068] See Figure 1 and Figure 2, Embodiment 1 of the present invention provides a transient stability emergency control method based on Rainbow reinforcement learning, including the following steps:
[0069] S1. Train the agent by setting a training strategy to obtain a trained agent;
[0070] S2. Monitor the power system by setting monitoring devices to obtain monitoring data of the power system;
[0071] S3. Obtain the fault disturbance from the monitoring data;
[0072] S4. Judge the transient stability of the fault disturbance through the transient stability index. If the transient stability is satisfied, the power system returns to stability; if the transient stability is not satisfied, obtain the power system state data and perform emergency control processing;
[0073] S5. Input the power system state data into the trained agent; the agent makes a decision through the deep reinforcement learning Rainbow algorithm and outputs an emergency control strategy;
[0074] S6. The power system performs emergency control processing according to the emergency control strategy and returns to stability.
[0075] In this embodiment, in step S1, the agent is trained by setting a training strategy to obtain a trained agent;
[0076] Specifically, as Figure 3 shown, in the process of training the agent by setting a training strategy, the training steps of the agent are:
[0077] S11. Initialize the environmental parameters according to the training objective;
[0078] Specifically, the environmental parameters include: the number of training episodes, the maximum number of steps, the learning rate, the discount factor, and the greedy ratio.
[0079] S12. The agent obtains state observations from the power system environment;
[0080] Specifically, the agent obtains state observations from the power system environment, including the angle of attack, the rotor speed, the voltage amplitude, and the phase angle.
[0081] S13. Input the state observations into the agent's estimation network model, and through the estimation processing of the agent's estimation network model, output an emergency control action;
[0082] Specifically, in the agent during training, in order to enable the agent to jump out of the local optimal solution and have the ability to conduct global exploration, the ε-greedy random greedy strategy is adopted. That is, a random number x is taken. If x < ε, the action with the highest evaluation value Q is selected as the current action. If x > ε, a random action is selected from all actions. And ε increases continuously with the number of training rounds. When the number of training times is sufficient, the parameters in the deep neural network hardly change anymore. At this time, ε is 1, and the best action is selected every time.
[0083] S14. Interact the emergency control action with the power system environment to obtain the feedback reward and the observation state at the next moment, and store them in the experience library.
[0084] Specifically, interact this action with the power system environment, feedback the reward and the observation state at the next moment, and store them in the experience library.
[0085] Among them, the short-term reward function is defined as the energy constraint reward function r e , the control cost constraint reward function r c , the control step number constraint reward function r n , the active power constraint reward function r g , the voltage constraint reward function r v These five parts.
[0086] R F = r e + r c + r g + r v + r n
[0087] The present invention defines the potential energy index F of the energy function Vp , and uses a single-valued function variable value to describe the energy change before and after the control of the initial working point. According to the definition of the transient energy function of the multi-machine system, the system rotor position potential energy and the system magnetic field potential energy are selected to form the potential energy index, as shown below:
[0088] F Vp = λ1V p1 + λ2V p2
[0089]
[0090] In the formula: V p1 and V p2 respectively represent the rotor position potential energy index and the magnetic field potential energy index; λ1 and λ2 represent the weights of the two, and in the present invention, 0.1 and 1 are respectively taken; P mi is the mechanical power input by the i-th generator; E i is the electromotive force of the i-th generator; G iiand B ij represent the real and imaginary parts of the shunt admittance matrix of the contraction node; δ ij is the relative power angle between the i-th and j-th generators; and represent the power angle of the i-th generator and the relative power angle between the i-th and j-th generators at the initial operating point, respectively.
[0091] In addition, the energy function reward will have a large penalty value after the system becomes unstable. If this value remains too high for a long time, the agent may directly choose an action that makes the system unstable to end the operation prematurely, making it difficult to continue training. Therefore, the present invention imposes a maximum and minimum constraint on the energy reward function, and the actual reward function used is:
[0092]
[0093] In the formula: r e0 is a positive number representing the limit value of the forced constraint.
[0094] The control cost reward function reflects the costs of generator tripping and load shedding, and imposes a penalty based on the weighted sum of the tripping amounts. The expression is as follows:
[0095]
[0096] In the formula: c G represents the generator tripping penalty coefficient, P g,i represents the tripping amount of the i-th generator, c L represents the load shedding penalty coefficient, P l,j represents the load shedding amount of the j-th load.
[0097] The active power constraint of the generator is to limit the output of each generator after control so that it is within the upper and lower limit constraints. The present invention imposes a penalty based on the magnitude of the excess limit. The expression is as follows:
[0098]
[0099] In the formula: c g is the active power overlimit penalty coefficient of the generator; r pg,i represents the power overlimit value of the i-th generator; and represent the upper and lower limits of the power of the i-th generator, respectively.
[0100] The voltage constraint reward function is to limit the voltage of each node after control and impose a penalty based on the magnitude of the excess of the upper and lower limits. The expression is as follows:
[0101]
[0102] In the formula, c vis the voltage violation penalty coefficient, r nv,j represents the voltage violation value of the j-th node, respectively represent the upper and lower limits of the voltage amplitude of the j-th node.
[0103] The control step constraint reward function is to limit the total number of control actions each time, guide the agent to complete the control goal with the least number of actions, and give penalties according to the number of control times. The expression is as follows:
[0104] r n = c n N step , N step = 1, 2,..., N m
[0105] In the formula, c n is the step penalty coefficient, N step is the number of control times, N m is the maximum number of control times.
[0106] The short-term reward function is updated and accumulated after each action, and the long-term reward function is calculated only at the end of each episode. The positive reward value of the reward function should be greater than the negative reward value, that is, the long-term stable positive reward is greater than the negative penalty during control, so as to ensure that each successful control action can be learned by the agent. In order to intuitively display the change of the reward function during training, the sliding average is used to calculate the cumulative training reward value after each episode.
[0107] S15. Update the policy network according to the observed state, the emergency control action and the feedback reward, and update the state observation value;
[0108] Specifically, update the policy network using the state, action and reward, update the state observation value, step+1.
[0109] S16. The power system performs emergency control processing according to the emergency control action, and judges whether the power system returns to stability; if the power system returns to stability, perform subsequent processing; if the power system does not return to stability, re-obtain the state observation value from the power system environment and perform the next iterative emergency processing until the number of steps reaches the set maximum number of steps, and stop the iteration;
[0110] S17. When the power system returns to stability, update the parameters of the agent estimation network model;
[0111] Specifically, update the evaluation network parameters on time, episode+1.
[0112] S18. Determine the number of training rounds of the agent. If the set number of training rounds is not reached, obtain the state observation from the power system environment again and perform iterative training. If the set number of training rounds is reached, save the agent evaluation network model and output the final control strategy.
[0113] S19. Test the final control strategy through a test scenario and output the test result.
[0114] In this embodiment, in step S2, a monitoring device is set to monitor the power system and obtain the monitoring data of the power system.
[0115] Specifically, define the parameter H of the power system:
[0116]
[0117] In the formula, S B3 represents the system rated three-phase power (MVA);
[0118] Define the general form of the rotor motion equation:
[0119]
[0120] In the formula, δ is the angle of attack, in rad, relative to the synchronous rotating reference frame; ω is the angular velocity, in rad / s; T a is the accelerating torque, in N·m.
[0121] Define:
[0122]
[0123] Among them, there are diagonal elements and non-diagonal elements According to the definition:
[0124]
[0125] Therefore, in a multi-machine system, the rotor motion equation set is:
[0126]
[0127] In this embodiment, the state space of the power system is used to describe the environmental information perceived by the agent. In transient stability analysis, the angle of attack of the generator and the voltage of the system network nodes can reflect the transient stability of the system. Therefore, the present invention selects the generator angle of attack, the generator rotor angular velocity, the network node voltage amplitude and phase as the state space of the agent. If the system has n nodes, m generators, and l loads, the state space S of the agent can be expressed as:
[0128] S = [δ Gi , ω Gi , V Nj , θ Nj
[0129] Where δ Gi is the angle of attack of generator i, i = 1, ..., m; ω Gi represents the rotor angular velocity of generator i, i = 1, ..., m; V Nj represents the voltage amplitude of node j, j = 1, ..., n; θ Nj represents the voltage phase of node j, j = 1, ..., n.
[0130] In this embodiment, the action space of the power system is used to describe all the control strategies of the agent.
[0131] In transient stability emergency control, disconnecting generators can reduce the mechanical power input to the system. When the system load is too heavy, loads need to be disconnected. The control of the agent includes generator tripping and load shedding actions. Therefore, the action space A is as follows:
[0132] A = [P G1 ,...P Gm , P L1 ,...P Ll
[0133] Where, P Gi is the generator tripping amount; m represents the number of generators; P Lk represents the load shedding amount; l represents the load number.
[0134] In reinforcement learning, the DRL agent can determine the optimal control action strategy combination through repeated trial and error. For the convenience of reinforcement learning training, the present invention discretizes the control action space of generator tripping and load shedding, and uniformly numbers the discretizable actions. Generator tripping control generally trips out the whole generator, so each generator is set with an action serial number; load shedding control usually sheds load in a certain proportion, so each load is set with an equal-interval number of actions; at the same time, the serial number of no action is set to 0. Then the serial number of the agent after discretization is expressed as:
[0135]
[0136] Where, c is the action serial number represented by natural numbers; P G represents the control of the tripped generator; represents the control of the shed load; u represents the load cutting ratio, u = 1, 2, ... h.
[0137] In this embodiment, in step S3, the fault disturbance is obtained through the monitoring data.
[0138] In this embodiment, in step S4, the transient stability of the fault disturbance is judged by the transient stability index. If the transient state is stable, the power system resumes stability; if the transient state is unstable, the power system state data is acquired and emergency control processing is performed.
[0139] Among them, the expression of the transient stability index is:
[0140]
[0141] In the formula, Δδ max is the maximum angle of attack difference between any two generators within a period of time.
[0142] In this embodiment, in step S5, the power system state data is input into the trained agent; the agent makes decisions through the deep reinforcement learning Rainbow algorithm and outputs an emergency control strategy.
[0143] Specifically, the interaction relationship between the power grid environment and the agent is as Figure 4 shown.
[0144] Among them, Rainbow is an improved algorithm that combines several different techniques based on the traditional DQN (Deep Q-Network) algorithm.
[0145] In the traditional DQN algorithm, the update method of the estimation network is as Figure 5 shown. First, the agent model obtains the power system environment state during the interaction. According to this state s t and action a t , calculate the estimated network action value function Q(s t , a t ), and select the action that maximizes Q(s t+1 , a) according to the next moment state s t+1 . Obtain the action value function under this action and form the learning target together with the reward value. Then, according to the estimated network action value function and the Q learning target, calculate the TD error, and use the square of the TD error as the loss function for updating the parameters of the neural network, and update the internal parameters by the stochastic gradient descent method to complete an iterative update of the action value function. As the number of interactions increases, the action value function of the estimation network will approach the true value, so that the optimal strategy can be obtained. The setting of the target network is to increase the stability of the DQN algorithm and remain unchanged within a certain training time interval. In this way, the learning target during this period will also be relatively stable and conducive to convergence. The implementation method is to copy the estimation network parameters to the target network to replace the original parameters at a certain training step interval to complete the temporary "freezing" of the estimation network parameters.
[0146] In this embodiment, Double Q-learning reduces the overestimation error by separating action selection from action evaluation.
[0147] In the traditional DQN algorithm, the action value function uses the maximum value in the evaluation network to update the network parameters. In the maximization calculation, the values of some actions will be overestimated, which is not conducive to the selection of the optimal action. The Double Q-learning method decouples action selection from value estimation, thus reducing the impact of overestimation. When updating the network, Double DQN first finds the action with the largest Q value from the estimation network, and then finds the output value of the target network corresponding to this action. In other words, the agent does not select the action with the largest value through the target network, but selects the action with the largest next state value through the estimation network. Then, the training target is generated according to the value of the action in the target network for evaluating the selected action.
[0148] Since Double Q-learning needs to construct two action value functions, one for estimating actions and the other for estimating the values of actions. Considering that the DQN algorithm already has an evaluation network and a target network, Double DQN will determine actions through the evaluation network and determine action values using the target network, without the need to reconstruct the network.
[0149] y t ← r t+1 + γ max a q(s t+1 , a; w t )
[0150] y t = r t+1 + q(s t+1 , argmax a q(s t+1 , a; w e ) ; w t )
[0151] In the formula, y t is the target Q value; γ ∈ [0, 1] is a discount factor; q(s t+1 , a; w t ) is the Q value of the state-action pair; where, w t represents the parameters of the target network; w e represents the parameters of the estimation network.
[0152] Prioritized experience replay improves the sample training efficiency by replaying more information more frequently.
[0153] TD (Time Error) error (time difference) measures the difference between the current state and the desired state in reinforcement learning. It is typically used to update the parameters of the policy network so that the policy can learn faster in the direction of the optimal policy. Specifically, the TD error calculates the difference between the reward in the current state and the expected reward in the next state. This difference reflects the difference between the current state and the expected state, and thus is often used to update the parameters of the policy network. Simply put, the TD error is used to help the reinforcement learning model quantify the gap between the generated optimal policy and the actual optimal policy, thereby helping the model learn faster.
[0154] Traditional greedy algorithms would select data with larger TD errors for training. Although this would theoretically accelerate convergence, it would also lead to some data being difficult to sample, or problems such as local optimality and overestimation. Therefore, the PER mechanism uses TD-error as a metric to evaluate the priority of sampled data, and determines the sampling probability p based on |TD-error| t , and improves data efficiency by replaying more frequently, so that there are more samples to learn from.
[0155]
[0156] where ω is a hyperparameter that determines the distribution; p t is the sampling probability determined according to the TD error.
[0157] In this embodiment, the Dueling network separates the estimation of the state value and the optimal action, and has better learning effects in multiple scenarios with similar value actions.
[0158] Dueling DQN is an improvement to the DQN neural network. In some states, the influence of actions on state transitions is small, while the state transitions generated by some actions have a greater impact on the state itself. Therefore, it is necessary to consider the value of the state itself, and thus a dual network is introduced. The Dueling DQN method introduces a state target value, and changes the output layer of the estimation network into two branches: the state target value and the advantage action target value, in order to improve the convergence effect of the algorithm.
[0159] A * (s,a) = Q * (s,a) - V * (s)
[0160]
[0161] where A * (s,a) is the optimal advantage function, Q * (s,a) is the optimal action value function; V *(s) is the optimal state value function; w represents the parameters of the neural network.
[0162] In this embodiment, multi-step learning improves the learning effect by calculating the rewards at multiple time steps.
[0163] The multi-step learning algorithm, namely multi-step DQN, its core idea is to define the return value with a fixed number of steps n. Different from the traditional DQN algorithm that uses one step, that is, only considering the current reward and the Q value of the next state at each time step, the multi-step DQN algorithm calculates the return value by considering the rewards of multiple future time steps. That is, at each time step, n actions are executed starting from the current state. The corresponding rewards are accumulated, and the Q value of the state after n steps is used as the target value for training, thus making full use of the reward signal of environmental delay, accelerating learning, and improving sample efficiency.
[0164]
[0165] In this embodiment, distributed reinforcement learning models the distribution of the return value to obtain more structural information about the rewards.
[0166] Distributed Reinforcement Learning (Distributional RL) is a class of value-based reinforcement learning algorithms. Classical value-based reinforcement learning methods use the expected value to establish a cumulative return model, and the expected value is represented as a value function or a state-action value function. In this modeling process, a large amount of complete distribution information is lost. Distributional reinforcement learning aims to solve this problem by modeling the distribution Z(x,a) of the cumulative return random variable, rather than just modeling its expectation.
[0167] d t =(z,p θ (S t ,A t ))
[0168]
[0169] D KL (Φ z d’ t ||d t )
[0170] In the formula, z is a vector defined on ; d t is the approximate distribution at time t; p θ is the probability; Φ z is the L2 norm projection of the target distribution on z, is the greedy action in state S t+1The average action value below.
[0171] In this embodiment, the noise network introduces randomness into the network weights to better explore.
[0172] Exploration like the ε-greedy algorithm adds noise to the action space, but there is a better method called the noisy network that adds noise to the parameter space. For example, given the same state, if the agent uses the ε-greedy algorithm, it may perform different actions, but the actual policy is not like this. A real-world policy should have the same response given the same state. If, in the same state, the agent sometimes selects a specific action based on the maximum value of the Q function and sometimes randomly selects an action, this will contradict reality. However, if noise is added to the parameters of the Q-function network, this will not happen. Because if noise is added to the network parameters of the Q function, during the entire interaction, in the same round, its network parameters are always fixed, so when in the same or similar states, the agent will take the same actions, which is more normal.
[0173] y = (b + Wx) + (b noisy ⊙ ∈ b + (W noisy ⊙ ∈ w )x)
[0174] In the formula, ∈ b and ∈ w are random variables, and ⊙ represents the Element-wise operation.
[0175] In this embodiment, the Rainbow agent replaces the one-step loss with a multi-step loss value variable. According to the cumulative discount factor, the target value is calculated through the n-step reward value, thereby constructing the target distribution.
[0176]
[0177] In the formula, Φ z is the projection to z.
[0178] Combining the multi-step distribution loss with DDQN, the greedy algorithm is used to select the action as t+n according to the online network state value S and evaluate this action using the target network.
[0179] All distributed Rainbow agents prioritize samples through the KL loss:
[0180]
[0181] The KL loss, as a priority, has better robustness against a noisy random environment because the loss can continue to decrease even when the reward is uncertain.
[0182] The network structure adopts the Dueling network structure, which is suitable for use in conjunction with the reward distribution.
[0183] For each z i , the action function and the advantage function are added together, just as in Dueling DQN, and then the normalized parameter distribution of the reward distribution estimate is obtained through the softmax layer:
[0184]
[0185] where φ = f ξ (s), and
[0186] In the noisy linear layer, Gaussian noise is used to reduce the number of independent noise variables.
[0187] In this embodiment, in step S6, the power system performs emergency control processing according to the emergency control strategy and resumes stability.
[0188] In a possible embodiment, an example of the emergency control of the power grid system is provided as follows:
[0189] The example power grid system is the New England 10-machine 39-bus system, and its network topology is as Figure 6 shown. The system includes 10 generators, 39 buses, 19 loads, and 46 branches. The generators all adopt the classical subtransient model, considering the action of the excitation system, and the loads adopt the constant impedance model. The environment of the power system in this embodiment is built using the PSS / E simulation software. The offline training integration platform is programmed in the Python language, and the interaction learning process between the DRL agent and the environment is realized through the interaction interface between the PSS / E simulation software and the Python program. The fault scenarios for emergency control are written based on the PSS / E module package, and the measured data of WAMS and PMU are simulated with the dynamic response data of the environment in the dynamic simulation. The influence of the uncertainty brought by the time delay is considered during the interaction process.
[0190] According to the deep reinforcement learning model, the emergency control problem is transformed into a Markov decision process. In this example, the state space of the agent includes the power angles and rotor angular velocities of 10 generators, and the bus voltages and phase angles of 39 nodes. The actions of the agent include generator tripping and load shedding actions. The proportion of staged load shedding is set to 10%, that is, each generator contains one control action, and each load shedding contains 10 actions. The reward function includes long-term rewards and short-term rewards. The long-term rewards are calculated based on whether the system remains stable after the control ends, and the short-term rewards are calculated after each step of control. The hyperparameter settings are shown in Table 1.
[0191] Parameter <![CDATA[R c > <![CDATA[R F > <![CDATA[C e > <![CDATA[C G > <![CDATA[C L > <![CDATA[C g > <![CDATA[C v > <![CDATA[C n > Value 100 -10 0.1 -0.003 -0.02 -0.003 -0.20 -1.0
[0192] Table 1 Hyperparameter settings
[0193] The training results of the DQN algorithm, Double DQN algorithm, Dueling DQN algorithm, and Rainbow algorithm are compared respectively, as Figure 7 shown. During the training process, the training results of the DQN algorithm are as Figure 7 (a) shows that its convergence is poor, and there are still large reward fluctuations around the 8000th episode. The overestimated value in training leads to the learned stable strategy not being optimal. Therefore, the DQN algorithm is not suitable for problems with high-dimensional state and action spaces. As an improved algorithm of DQN, the training results of Double DQN are as Figure 7 (b) shows that due to avoiding the problem of overestimation to a certain extent, its convergence is better than that of the DQN algorithm, and it converges at about the 6000th episode. The training results of the Dueling DQN algorithm are as Figure 7 (c) shows that it also converges at about the 6000th episode and considers the value of the state itself, so the training process is more stable. Finally, the Rainbow algorithm adopted in the present invention is as Figure 7 (d) shows that it converges at about the 4000th episode. It can be seen that the Rainbow algorithm has better convergence and smaller fluctuations in rewards during the training process.
[0194] In summary, the present invention trains an agent through a set training strategy to obtain a trained agent; monitors the power system through a set monitoring device to obtain monitoring data of the power system; obtains a fault disturbance through the monitoring data; judges the transient stability of the fault disturbance through a transient stability index. If the transient stability is achieved, the power system returns to stability. If the transient stability is not achieved, the power system state data is obtained and emergency control processing is performed; the power system state data is input into the trained agent; the agent makes a decision through the deep reinforcement learning Rainbow algorithm and outputs an emergency control strategy; the power system performs emergency control processing according to the emergency control strategy and returns to stability. The present invention introduces a reinforcement learning method to model and output a strategy for the emergency control of the power system, with a short decision-making time and the ability to achieve real-time control. At the same time, as an advanced algorithm of deep reinforcement learning, the Rainbow algorithm combines multiple advanced technologies (such as prioritized experience replay, double Q-learning, distributed learning, etc.), making the learning process more efficient and stable. With the help of the Rainbow algorithm, more intelligent and faster emergency control decisions can be achieved. Combining advanced machine learning technologies, it solves the deficiencies in traditional load restoration methods. This method has good adaptability and intelligence, provides new ideas and practical bases for the development of smart grids, and has broad application prospects. The present invention improves the efficiency and accuracy of the output of the emergency control strategy for the distribution network, can adapt to emergencies more quickly, optimize resource allocation, reduce power outage time, and improve power supply reliability. At the same time, using deep learning technology, the agent can process high-dimensional complex data and provide more scientific decision-making support.
[0195] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by the cooperation of multiple devices. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0196] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0197] Embodiment 2
[0198] See Figure 8, Embodiment 2 of the present invention further provides a transient stability emergency control device based on Rainbow reinforcement learning, including:
[0199] An agent training module 001, configured to train an agent by setting a training strategy to obtain a trained agent;
[0200] A monitoring data acquisition module 002, configured to monitor a power system by setting monitoring devices to obtain monitoring data of the power system;
[0201] A fault disturbance acquisition module 003, configured to obtain a fault disturbance through the monitoring data;
[0202] A transient stability judgment module 004, configured to judge the transient stability of the fault disturbance through a transient stability index. If the transient state is stable, the power system resumes stability; if the transient state is unstable, the state data of the power system is obtained and emergency control processing is performed;
[0203] An agent emergency control strategy output module 005, configured to input the state data of the power system into the trained agent; the agent makes a decision through the deep reinforcement learning Rainbow algorithm and outputs an emergency control strategy;
[0204] An emergency control processing module 006, configured to perform emergency control processing on the power system according to the emergency control strategy to resume stability.
[0205] In this embodiment, in the agent training module 001, the training sub-module includes:
[0206] An environment parameter initialization sub-module 011, configured to initialize the environment parameters according to the training objective;
[0207] A state observation quantity acquisition sub-module 012, configured for the agent to obtain state observation quantities from the power system environment;
[0208] An estimation network model processing sub-module 013, configured to input the state observation quantities into the agent estimation network model, perform estimation processing through the agent estimation network model, and output an emergency control action;
[0209] An interaction feedback sub-module 014, configured to interact the emergency control action with the power system environment, obtain a feedback reward and an observation state at the next moment, and store them in an experience library;
[0210] A policy network update sub-module 015, configured to update the policy network according to the observation state, the emergency control action, and the feedback reward, and update the state observation quantities;
[0211] The iterative training sub-module 016 is used for the power system to perform emergency control processing according to the emergency control action, and determine whether the power system resumes stability; if the power system resumes stability, subsequent processing is performed; if the power system does not resume stability, the state observation quantity is retrieved from the power system environment again, and the next iterative emergency processing is performed until the number of steps reaches the set maximum number of steps, and the iteration is stopped;
[0212] The model parameter update sub-module 017 is used to update the parameters of the agent estimation network model when the power system resumes stability;
[0213] The agent evaluation network model storage sub-module 018 is used to judge the number of agent training rounds. If the set number of training rounds is not reached, the state observation quantity is retrieved from the power system environment again for iterative training; if the set number of training rounds is reached, the agent evaluation network model is saved and the final control strategy is output;
[0214] The test and verification sub-module 019 is used to test the final control strategy through a test scenario and output the test result.
[0215] In this embodiment, in the environment parameter initialization sub-module 011 of the agent training module 001, in the process of initializing the environment parameters according to the training target, the environment parameters include: the number of training rounds, the maximum number of steps, the learning rate, the discount factor, and the greedy ratio.
[0216] In this embodiment, in the transient stability judgment module 004, in the process of judging the transient stability of the fault disturbance through the transient stability index, the expression of the transient stability index is:
[0217]
[0218] In the formula, Δδ max is the maximum angle of attack difference between any two generators within a period of time.
[0219] In this embodiment, in the interaction feedback sub-module 014 of the agent training module 001, in the process of interacting the emergency control action with the power system environment to obtain the feedback reward, the short-term reward function is defined as: the energy constraint reward function, the control cost constraint reward function, the control step constraint reward function, the active power constraint reward function, and the voltage constraint reward function; the expression of the short-term reward function is:
[0220] R F =r e +r c +r g +r v +rn
[0221] Wherein, r e is the energy constraint reward function; r c is the control cost constraint reward function; r g is the control step number constraint reward function; r v is the active power constraint reward function; r n is the voltage constraint reward function.
[0222] It should be noted that for the information interaction, execution process, etc. between the above-mentioned system modules, since they are based on the same concept as the method embodiment in Embodiment 1 of the present application, the technical effects brought by them are the same as those of the method embodiment of the present application. For specific content, reference can be made to the description in the method embodiment shown above in the present application, and details will not be repeated here.
[0223] Embodiment 3
[0224] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code of a transient stability emergency control method based on Rainbow reinforcement learning is stored, and the program code includes instructions for executing a transient stability emergency control method based on Rainbow reinforcement learning according to Embodiment 1 or any possible implementation manner thereof.
[0225] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0226] Embodiment 4
[0227] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0228] The processor and the memory communicate with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute a transient stability emergency control method based on Rainbow reinforcement learning according to Embodiment 1 or any possible implementation manner thereof by calling the program instructions.
[0229] Specifically, the processor can be implemented by hardware or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor that realizes its functions by reading software code stored in a memory. The memory can be integrated in the processor or exist independently outside the processor.
[0230] In the above embodiments, they can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means.
[0231] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general computing system. They can be concentrated on a single computing system or distributed on a network composed of multiple computing systems. Optionally, they can be implemented by program code executable by the computing system. Thus, they can be stored in a storage system and executed by the computing system. And in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules respectively, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0232] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A transient stability emergency control method based on Rainbow reinforcement learning, characterized in that: include: By setting the training strategy, the intelligent agent is trained to obtain a trained intelligent agent; Monitor the power system by setting monitoring equipment and obtain monitoring data of the power system; Obtaining fault disturbance through the monitoring data; The transient stability of the fault disturbance is judged by the transient stability index. If the transient state is stable, the power system is restored to stability; if the transient state is unstable, the power system state data is obtained and emergency control processing is performed; Inputting the power system state data into a trained agent; The agent makes decisions through deep reinforcement learning Rainbow algorithm and outputs emergency control strategies; The power system performs emergency control processing according to the emergency control strategy and restores stability.
2. The transient stability emergency control method based on Rainbow reinforcement learning according to claim 1 is characterized in that: In the process of training the agent by setting the training strategy, the training steps for the agent are: Initialize the environment parameters according to the training objectives; The agent obtains state observations from the power system environment; Inputting the state observation into an intelligent agent estimation network model, performing estimation processing by the intelligent agent estimation network model, and outputting an emergency control action; Interacting the emergency control action with the power system environment, obtaining feedback rewards and the observation state at the next moment, and storing them in an experience database; According to the observed state, the emergency control action and the feedback reward, the policy network is updated, and the state observation is updated; The power system performs emergency control processing according to the emergency control action to determine whether the power system has recovered stability; If the power system returns to stability, follow-up processing will be carried out; If the power system does not recover stability, the state observation is obtained from the power system environment again, and the next iterative emergency processing is performed until the number of steps reaches the set maximum number of steps, and the iteration is stopped; When the power system recovers stability, updating the parameters of the agent estimation network model; The number of training rounds of the intelligent agent is judged. If the set number of training rounds is not reached, the state observation quantity is re-acquired from the power system environment to perform iterative training; if the set number of training rounds is reached, the intelligent agent evaluation network model is saved and the final control strategy is output; The final control strategy is tested through a test scenario, and the test results are output.
3. The transient stability emergency control method based on Rainbow reinforcement learning according to claim 2 is characterized in that: In the process of initializing the environmental parameters according to the training objectives, the environmental parameters include: the number of training rounds, the maximum number of steps, the learning rate, the discount factor and the greedy ratio.
4. The transient stability emergency control method based on Rainbow reinforcement learning according to claim 3 is characterized in that: In the process of judging the transient stability of the fault disturbance by using the transient stability index, the expression of the transient stability index is: In the formula, Δδ max is the maximum angle of attack difference between any two generators over a period of time.
5. The transient stability emergency control method based on Rainbow reinforcement learning according to claim 4 is characterized in that: In the process of interacting the emergency control action with the power system environment to obtain feedback rewards, the short-term reward function is defined as: energy constraint reward function, control cost constraint reward function, control step constraint reward function, active power constraint reward function and voltage constraint reward function; the expression of the short-term reward function is: R F =r e +r c +r g +r v +r n In the formula, r e is the energy constraint reward function; r c To control the cost constraint reward function; r g To control the number of steps constrained reward function; r v is the active power constraint reward function; r n is the voltage constraint reward function.
6. A transient stability emergency control device based on Rainbow reinforcement learning, adopting a transient stability emergency control method based on Rainbow reinforcement learning according to any one of claims 1 to 5, characterized in that: include: The agent training module is used to train the agent by setting the training strategy to obtain a trained agent; A monitoring data acquisition module is used to monitor the power system by setting monitoring equipment and acquire monitoring data of the power system; A fault disturbance acquisition module, used to acquire the fault disturbance through the monitoring data; A transient stability judgment module is used to judge the transient stability of the fault disturbance through a transient stability index. If the transient state is stable, the power system is restored to stability; if the transient state is unstable, the power system state data is obtained and emergency control processing is performed; An intelligent agent emergency control strategy output module, used to input the power system state data into a trained intelligent agent; The agent makes decisions through deep reinforcement learning Rainbow algorithm and outputs emergency control strategies; The emergency control processing module is used for the power system to perform emergency control processing according to the emergency control strategy to restore stability.
7. The transient stability emergency control device based on Rainbow reinforcement learning according to claim 6 is characterized in that: In the agent training module, the training submodule includes: The environment parameter initialization submodule is used to initialize the environment parameters according to the training objectives; The state observation acquisition submodule is used by the intelligent agent to obtain state observations from the power system environment; An estimation network model processing submodule, used for inputting the state observation into the intelligent agent estimation network model, and outputting an emergency control action through estimation processing by the intelligent agent estimation network model; An interactive feedback submodule, used to interact the emergency control action with the power system environment, obtain feedback rewards and the observation state at the next moment, and store them in an experience database; A policy network update submodule, used to update the policy network and the state observation according to the observed state, the emergency control action and the feedback reward; An iterative training submodule is used for the power system to perform emergency control processing according to the emergency control action to determine whether the power system has recovered stability; if the power system has recovered stability, then subsequent processing is performed; if the power system has not recovered stability, then the state observation quantity is re-acquired from the power system environment to perform the next iterative emergency processing until the number of steps reaches the set maximum number of steps, and then the iteration is stopped; A model parameter updating submodule, used for updating the parameters of the agent estimation network model when the power system recovers stability; The intelligent agent evaluation network model storage submodule is used to judge the number of intelligent agent training rounds. If the set number of training rounds is not reached, the state observation quantity is re-acquired from the power system environment for iterative training; if the set number of training rounds is reached, the intelligent agent evaluation network model is saved and the final control strategy is output; The test verification submodule is used to test the final control strategy through a test scenario and output the test results.
8. The transient stability emergency control device based on Rainbow reinforcement learning according to claim 7 is characterized in that: In the environment parameter initialization submodule of the agent training module, in the process of initializing the environment parameters according to the training objectives, the environment parameters include: number of training rounds, maximum number of steps, learning rate, discount factor and greedy ratio.
9. The transient stability emergency control device based on Rainbow reinforcement learning according to claim 8, characterized in that: In the transient stability judgment module, in the process of judging the transient stability of the fault disturbance by using the transient stability index, the expression of the transient stability index is: In the formula, Δδ max is the maximum angle of attack difference between any two generators over a period of time.
10. The transient stability emergency control device based on Rainbow reinforcement learning according to claim 9, characterized in that: In the interactive feedback submodule of the agent training module, in the process of interacting the emergency control action with the power system environment to obtain feedback rewards, the short-term reward function is defined as: energy constraint reward function, control cost constraint reward function, control step constraint reward function, active power constraint reward function and voltage constraint reward function; the expression of the short-term reward function is: R F =r e +r c +r g +r v +r n In the formula, r e is the energy constraint reward function; r c To control the cost constraint reward function; r g To control the number of steps constrained reward function; r v is the active power constraint reward function; r n is the voltage constraint reward function.