Power grid risk disposal plan generation method and system based on reinforcement learning
Through the method based on reinforcement learning, the operating status and sensitivity of the power grid are analyzed and the grid accident handling plan is generated, which solves the problem of insufficient real-time and flexibility of the power grid accident handling plan in traditional methods, and achieves efficient and accurate accident response and plan optimization.
Patent Information
- Application Number
- CN202411829574.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-13
AI Technical Summary
Traditional grid accident handling plans are difficult to adapt to the complex and changeable grid operation state, and their real-time and flexibility are limited, resulting in insufficient efficiency and accuracy of accident handling.
Using reinforcement learning-based methods, the priority of key variables is determined through the analysis of the operating status of the power grid and the sensitivity analysis, and a grid accident handling plan is generated through the policy network. The method includes building an agent, dynamically adjusting the action selection range, and designing a reward function to balance the current constraints and generator adjustment costs.
It has achieved accurate and rapid response to power grid accidents, has the ability to optimize itself, and continuously improves the quality of accident handling plans through continuous learning, can adapt to changes in power grid data in real time, and improves the ability to deal with power grid accidents and system operation stability.
Smart Images

Figure CN119990793A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system operation control, and in particular to a method and system for generating a power grid risk disposal plan based on reinforcement learning. Background Art
[0002] The power grid system faces a variety of complex emergencies, such as transmission line tripping, equipment failure, etc., for which effective accident handling plans need to be formulated. Traditional plans rely on expert experience and fixed rules, which are difficult to adapt to the complex and changeable power grid operation status, and have limited real-time and flexibility. Therefore, an intelligent and dynamic response method is needed to improve the efficiency and accuracy of power grid accident handling. Summary of the invention
[0003] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method for generating a power grid risk disposal plan based on reinforcement learning, which can solve the problems mentioned in the background technology.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for generating a power grid risk disposal plan based on reinforcement learning, comprising analyzing the power grid operation status and determining the priority of key variables;
[0008] The intelligent agent is constructed through the preset algorithm, and the power grid accident handling plan is generated through the strategy network.
[0009] As a preferred solution of the method for generating a power grid risk disposal plan based on reinforcement learning of the present invention, it also includes verifying the generated power grid accident disposal plan.
[0010] As a preferred solution of the method for generating a power grid risk disposal plan based on reinforcement learning of the present invention, the power grid operation status is analyzed and the priority of key variables is determined, including:
[0011] Analyze the real-time operating status of the power grid to identify key variables and equipment;
[0012] Calculate the impact of each generator in the power grid on the active power flow of a specific line to determine the priority.
[0013] As a preferred solution of the method for generating a power grid risk disposal plan based on reinforcement learning of the present invention, wherein: an intelligent agent is constructed by a preset algorithm, including:
[0014] Define the state space and action space of the agent through a preset algorithm;
[0015] Dynamically adjust the action selection range of the agent, and adjust the action selection range of the agent according to the results of sensitivity analysis;
[0016] The reward function is designed to balance the power flow constraints of the grid and the adjustment cost of the generators.
[0017] As a preferred solution of the method for generating a power grid risk disposal plan based on reinforcement learning of the present invention, wherein: generating a power grid accident disposal plan through a strategy network includes:
[0018] The policy network takes the current state of the power grid as input, generates a Gaussian distribution and samples it to obtain an action;
[0019] This action is performed to adjust the power grid status and generate and execute the optimal emergency plan.
[0020] As a preferred solution of the method for generating a power grid risk disposal plan based on reinforcement learning of the present invention, the generated power grid accident disposal plan is verified, including:
[0021] Use the real-time operation data of the power grid to compare the status before and after the accident and verify the effectiveness of the plan.
[0022] As a preferred solution of the method for generating a power grid risk disposal plan based on reinforcement learning of the present invention, the preset algorithm is a Soft Actor-Critic (SAC) deep reinforcement learning algorithm.
[0023] In a second aspect, the present invention provides a system for generating a power grid risk disposal plan based on reinforcement learning, comprising: an analysis and determination module for analyzing the power grid operation status and determining the priority of key variables;
[0024] A generation module is constructed to build an intelligent agent through a preset algorithm and generate a power grid accident disposal plan through a policy network.
[0025] In a third aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0026] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the above method when the computer program is executed by a processor.
[0027] Compared with the prior art, the present invention has the following beneficial effects: by implementing sensitivity analysis to determine the priority of key variables in the power grid, and integrating reinforcement learning models to dynamically generate efficient accident response strategies. This method can not only respond to power grid accidents accurately and quickly, but also has the ability to self-optimize, and continuously improve the quality of accident handling plans through continuous learning. Its dynamic adjustment characteristics enable this method to adapt to changes in power grid data in real time, effectively serving the rapid response in emergency situations and the optimization of daily power grid operations, thereby comprehensively improving the power grid accident handling capabilities and the stability of system operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0029] Figure 1 This is a schematic diagram of the overall process of the power grid risk disposal plan generation method based on reinforcement learning.
[0030] Figure 2 This is a schematic diagram showing the interface of the power grid risk disposal plan generation system based on reinforcement learning.
[0031] Figure 3 This is a graph showing reward changes.
[0032] Figure 4 A schematic diagram of the internal structure of a computer device. DETAILED DESCRIPTION
[0033] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0034] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0035] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0036] Example 1
[0037] Reference Figure 1-2 , which is the first embodiment of the present invention, and provides a method for generating a power grid risk disposal plan based on reinforcement learning, which includes:
[0038] S1. Analyze the grid operation status and prioritize key variables.
[0039] Furthermore, the grid operation status is analyzed and key variables are prioritized, including:
[0040] Analyze the real-time operating status of the power grid to identify key variables and equipment;
[0041] Calculate the impact of each generator in the power grid on the active power flow of a specific line to determine the priority.
[0042] It should be noted that for each transmission line in the real power grid, its power flow sensitivity is calculated:
[0043]
[0044] Among them, S ij,Pg The active power of generator g on the active power flow P on line ij is ij (Line ij represents the line between bus i and bus j connecting two nodes), P ij is the power of line ij, P g is the active power of generator g.
[0045] Preferably, the level of power flow sensitivity can reflect the degree of influence of the adjustment of a generator on a specific line. Through this sensitivity analysis, the generators with greater influence on the key lines can be adjusted first to improve the efficiency of accident handling.
[0046] S2. Build an intelligent agent through a preset algorithm and generate a power grid accident handling plan through a strategy network.
[0047] Furthermore, the preset algorithm is the Soft Actor-Critic (SAC) deep reinforcement learning algorithm.
[0048] Furthermore, the intelligent agent is constructed through the preset algorithm, including,
[0049] The state space and action space of the intelligent agent are defined by a preset algorithm, wherein the state space includes the grid parameters such as the active power, reactive power, bus voltage and active output of the unit;
[0050] It should be noted that the SAC deep reinforcement learning agent training state S includes:
[0051] S={P line , Q line , V bus , P gen}
[0052] In the formula, the initial state includes the active power of the line in the power grid P line 、Line reactive power Q line , bus voltage V bus and the active power P of the unit gen .
[0053] Initialize the experience replay buffer It is expressed as:
[0054]
[0055] In the formula, S represents the current state of the environment, A represents the action selected according to the strategy in the current environment, that is, the adjustment amount of each unit in the section. R represents the reward given after the action is selected, S' represents the next power grid environment after the action is executed, and done represents the flag of whether this round is over.
[0056] For the cross-section state S after adjustment t times t , select action a t :
[0057] a t =π θ (S t )
[0058] In the formula, action a t Represented as a vector containing the recommended active value adjustment for all units in the selected section sample, π θ (S t ) represents the strategy network, according to the current power grid state S t Get the corresponding action.
[0059] Dynamically adjust the action selection range of the agent, and adjust the action selection range of the agent according to the results of sensitivity analysis;
[0060] It should be noted that for generators with high sensitivity, their action range is adjusted so that they are more likely to choose large-scale actions when the strategy network outputs. θ (S t ) before outputting, use sensitivity priority to dynamically adjust the range of certain actions. For example:
[0061] a′ g =a g ×(1-λS g )
[0062] In the formula, a g is the action chosen by the agent for the generator g, S g is the sensitivity of the generator, λ is the adjustment coefficient, which is used to control the influence of sensitivity on the action amplitude, a′ g It is the final action after sensitivity priority adjustment.
[0063] Preferably, by directly scaling the policy network output, units with high sensitivity are more likely to be adjusted more.
[0064] It is further explained that action offset is added to the policy network. Specifically, the output of the policy network is guided by sensitivity analysis, so that the actions corresponding to the units with high sensitivity are more likely to be biased in a specific direction. Specifically, the mean of the action distribution is guided:
[0065] μ′ g =μ g +KS g
[0066] In the formula, μ g is the original action mean output by the policy network, k is the offset coefficient used to control the impact of sensitivity on the action mean, S g is the generator sensitivity, μ' g is the mean of the actions after adding the offset.
[0067] The reward function is designed to balance the power flow constraints of the grid and the adjustment cost of the generators.
[0068] It should be noted that the reward function gives additional rewards or penalties to units with high sensitivity, so that the intelligent agent will be more inclined to choose units with greater impact on the power grid for adjustment. Specifically,
[0069]
[0070] In the formula, R is the reward value, i∈CriticaILines represents the set of critical lines in the system, and we need to pay attention to whether these critical lines are out of limit, α i is the weight of line i, indicating the importance of the line, P i actual and P i Iimit They represent the actual power flow and power flow limit of line i respectively, β is the weight coefficient of the adjustment cost of the generator, which controls the degree of emphasis on the adjustment cost of the generator in the reward function, j∈Generators is the set of all generators in the system, ω j =1-λSj, where S jRepresents the sensitivity of unit j, λ is a non-negative coefficient used to control the influence of sensitivity on weight. ΔP j is the adjustment amount of the active output of generator j, which is the action selected by the SAC agent and represents the increase or decrease in the output of the generator at the current time step.
[0071] Preferably, the design of the reward function can balance the power flow constraints of the power grid and the cost of generator adjustment. By adding the weight of the combined sensitivity, the agent can give priority to adjusting the generator that has a greater impact on the coefficient, thereby improving the efficiency of the entire strategy of the SAC agent during training. i and β are the trade-offs between line safety and generator regulation. i When β is larger, the agent pays more attention to keeping the line power flow within a safe range, and a larger β means reducing the frequency of generator adjustment, which can improve the economy of the strategy.
[0072] It should be further explained that the technical solution also includes the generation of reinforcement learning strategies, specifically, initializing the environment to obtain the initial state S of the system. Strategy network π θ (S t ) receives the current state S as input, outputs a probability distribution, and samples an action a from the distribution t SAC uses Gaussian distribution to describe the continuous action space. The policy network outputs a mean and standard deviation, and then samples based on this to get the action of the current time step.
[0073] Execute action a t , the environment changes accordingly according to this action and returns to the new state S′ t , immediate reward R′ t , and a flag done that indicates whether the environment is finished.
[0074] S′ t , R′ t ,done=Environment(S t , a t )
[0075] The current state, action, reward, and new state are stored in the experience replay pool as an experience sample
[0076] Update the policy network (Actor) and the value network (Critic)
[0077] Storing experience, updating model weights including replaying the experience buffer A small batch of samples (S t ,a t ,r t,S′ t , done), calculate the target value y:
[0078]
[0079] Among them, y represents the target Q value, which refers to the expected cumulative reward after the model selects action a in state s, and is used to update the Q function. r represents the immediate reward, which means the reward obtained after taking action a in the current state s. γ represents the discount factor, which is used to weigh the relative importance of future rewards and current rewards. a’~π represents the expected value of action a' generated according to strategy π, Qφ' (s',a') represents the target Q function, which represents the Q value obtained by taking action a' in the new state s'. φ' represents the parameter of the Q function, α represents the temperature parameter, which controls the randomness of the strategy, and logπ(a'|s') is the logarithm of the probability that strategy π takes action a' in state s'.
[0080] The loss function of minimizing the mean square error is expressed as:
[0081]
[0082] Among them, Qφ(s,a) represents the expected cumulative reward after taking action a in state s under the parameter φ of the Q function.
[0083] Update strategy network π θ To maximize the objective function:
[0084]
[0085] Among them, π θ represents a parameterized strategy, represents the sum of the cumulative rewards, γ t Represents the discount factor, which is used to weigh the relative importance of future rewards and current rewards. t represents the time step, which means that the reward after each step is decayed by γ times. t ,a t ) represents the immediate reward function, which is in the state s at time step t t Take action a t Instant rewards received.
[0086] Update the value function:
[0087]
[0088] Where Qφ(s,a) represents the Q function, which represents the expected cumulative reward for taking action a in state s. Qφ' (s',a')represents the expected cumulative reward under the next state s' and action a', logπ(a'|s') represents the logarithm of the probability of selecting action a' in state s' under strategy π, and represents the entropy of the strategy.
[0089] Preferably, entropy regularization is used to maintain the exploratory nature of the strategy and improve the stability and convergence speed of the model.
[0090] Furthermore, the strategy network generates a power grid accident handling plan, including:
[0091] The policy network takes the current state of the power grid as input, generates a Gaussian distribution and samples it to obtain an action;
[0092] This action is performed to adjust the power grid status and generate and execute the optimal emergency plan.
[0093] It should be noted that after the model training is completed, the policy network is used to generate a plan for handling power grid accidents. Specifically, the current real-time operation data is obtained from the power grid monitoring system, including parameters such as the apparent power of the line, bus voltage, and active output of the generator. These data constitute the current power grid state, and at the same time, potential risk accidents during the operation of the power grid, such as line tripping, transformer failure, and generator shutdown, are loaded as inputs to the SAC model. The policy network receives this state and outputs an action, which represents the adjustment of the active power output of the generator. This action is then sent to the power grid system to adjust the corresponding equipment.
[0094] Furthermore, it also includes verification of the generated power grid accident handling plan.
[0095] Furthermore, the generated power grid accident disposal plan is verified, including:
[0096] Use the real-time operation data of the power grid to compare the status before and after the accident and verify the effectiveness of the plan.
[0097] It should be noted that the effectiveness of the verification plan includes indicators such as accident handling efficiency, system recovery time, and safety of power equipment. If the power grid can be effectively restored, the dispatching and control process will be recorded to form a standard plan for accident handling, providing a reference for subsequent responses to similar accidents.
[0098] In summary, the beneficial effect of a method for generating power grid risk disposal plans based on reinforcement learning is to implement sensitivity analysis to determine the priority of key variables in the power grid, and integrate reinforcement learning models to dynamically generate efficient accident response strategies. This method can not only respond to power grid accidents accurately and quickly, but also has the ability to self-optimize, and continuously improve the quality of accident handling plans through continuous learning. Its dynamic adjustment characteristics enable this method to adapt to changes in power grid data in real time, effectively serving the rapid response in emergency situations and the optimization of daily power grid operations, thereby comprehensively improving the power grid accident handling capabilities and the stability of system operations.
[0099] Example 2
[0100] This embodiment provides a system for generating a power grid risk disposal plan based on reinforcement learning, which includes an analysis and determination module for analyzing the power grid operation status and determining the priority of key variables;
[0101] A generation module is constructed to build an intelligent agent through a preset algorithm and generate a power grid accident disposal plan through a policy network.
[0102] The above-mentioned unit modules may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to the above-mentioned modules.
[0103] Example 3
[0104] This embodiment provides a computer device, which may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. When the computer program is executed by the processor, a method for generating a power grid risk disposal plan based on reinforcement learning is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0105] This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements: analyzing the operating status of the power grid and determining the priority of key variables; constructing an intelligent body through a preset algorithm, and generating a power grid accident handling plan through a strategy network.
[0106] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for generating a power grid risk disposal plan based on reinforcement learning, characterized in that: include, Analyze the grid operation status and prioritize key variables; The intelligent agent is constructed through the preset algorithm, and the power grid accident handling plan is generated through the strategy network.
2. The method for generating a power grid risk disposal plan based on reinforcement learning according to claim 1, characterized in that: It also includes verifying the generated power grid accident handling plan.
3. The method for generating a power grid risk disposal plan based on reinforcement learning according to claim 2, characterized in that: The analysis of the grid operation status and determination of priorities of key variables include: Analyze the real-time operating status of the power grid to identify key variables and equipment; Calculate the impact of each generator in the power grid on the active power flow of a specific line to determine the priority.
4. The method for generating a power grid risk disposal plan based on reinforcement learning according to claim 3, characterized in that: The intelligent agent is constructed by a preset algorithm, including: Define the state space and action space of the agent through a preset algorithm; Dynamically adjust the action selection range of the agent, and adjust the action selection range of the agent according to the results of sensitivity analysis; The reward function is designed to balance the power flow constraints of the grid and the adjustment cost of the generators.
5. The method for generating a power grid risk disposal plan based on reinforcement learning according to claim 4, characterized in that: The generation of a power grid accident handling plan through a strategy network includes: The policy network takes the current state of the power grid as input, generates a Gaussian distribution and samples it to obtain an action; This action is performed to adjust the power grid status and generate and execute the optimal emergency plan.
6. The method for generating a power grid risk disposal plan based on reinforcement learning according to claim 5, characterized in that: The verification of the generated power grid accident handling plan includes: Use the real-time operation data of the power grid to compare the status before and after the accident and verify the effectiveness of the plan.
7. The method for generating a power grid risk disposal plan based on reinforcement learning according to any one of claims 1 to 6, characterized in that: The preset algorithm is the Soft Actor-Critic (SAC) deep reinforcement learning algorithm.
8. A power grid risk disposal plan generation system based on reinforcement learning, characterized in that: include: An analysis and determination module is used to analyze the operation status of the power grid and determine the priority of key variables; A generation module is constructed to build an intelligent agent through a preset algorithm and generate a power grid accident disposal plan through a policy network.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Overload line emergency control strategy continuous learning method suitable for changing scene
CN120545987A