Power grid static security risk prevention and control method and system based on reinforcement learning

By superimposing security constraints on reinforcement learning objectives, a secure embedded intelligent agent is constructed. A security reward function is set, comprising risk elimination, security incentive, and stability margin terms. Through projection verification and effect evaluation, the security of power grid decision-making is ensured, solving the security risk problem of online decision-making in existing technologies and realizing the secure and efficient embedding of reinforcement learning in the power grid.

CN121097643APending Publication Date: 2025-12-09CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511191759.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing reinforcement learning methods lack a security constraint embedding mechanism in power grid control, which leads to security risks in online decision-making and makes it difficult to effectively address static security risks of the power grid.

Method used

By overlaying security constraints on reinforcement learning objectives, a secure embedded intelligent agent is constructed. A security reward function is set with risk elimination, security incentive, and stability margin terms. Through projection verification and effect pre-evaluation, the original actions are corrected by projection methods. A security projection verification and effect evaluation process is added to ensure the security of the decision.

Benefits of technology

It enables the safe and efficient embedding of reinforcement learning agents in actual power grids, and can respond to static safety risks in real time, avoiding safety accidents caused by improper decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121097643A_ABST
    Figure CN121097643A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid static security risk prevention and control method and system based on reinforcement learning, and the method comprises the steps: obtaining a power grid parameter, superposing a security constraint on a reinforcement learning target, and constructing a security embedded reinforcement learning agent; setting a security reward function including a risk elimination item, a security incentive item, a stability margin item and a control cost item, and constructing a reinforcement learning strategy network; safety projection check is added after the reinforcement learning strategy network is output, the safety projection check forcibly corrects original actions through projection constraint optimization, and corrected feasible actions are obtained; and setting a legality verification and effect pre-evaluation link for the corrected feasible action, and adjusting action execution according to a legality verification and effect pre-evaluation result to realize static security risk prevention and control of the power grid. According to the method, the reinforcement learning agent can be really embedded into a real-time online closed-loop control process which is strictly required by an actual power grid, and the static safety risk of the power grid is safely and efficiently dealt with.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of power grid safety prevention and control, and particularly relates to a power grid static safety risk prevention and control method and system based on reinforcement learning. BACKGROUND

[0002] With the gradual expansion of the power system, the proportion of new energy and power electronic equipment accessing the power system is rising, and the pre-plan relying on offline calculation is difficult to adapt to the real-time fluctuation of the power grid working condition, for example, the rapid change of new energy output and load. The calculation time of the online optimal power flow optimization method under large-scale complex constraints is long, which may not meet the real-time requirements of online decision-making, and may fall into local optimum. In order to improve the safety of power grid regulation, many researchers have carried out research on the application of reinforcement learning method in the regulation field, however, at present, it mainly stays in the simulation stage, and the key problem of decision safety difficulty guarantee in online application has not been solved.

[0003] The prior art, such as the patent application with publication number CN118886341A, proposes an "auxiliary decision-making method for power grid static security risk and related device", which includes: obtaining power grid model parameters and operation data, and using a trained agent to perform simulation operation to obtain an auxiliary decision-making result; the agent training process includes: initializing the environment state of the power grid power flow simulation environment; the agent generates an action strategy according to the environment state and a reward function; the environment executes the action strategy and updates the state according to the action strategy, calculates the reward score based on the reward function, and feeds back the state and the reward score to the agent; the action strategy is verified by using the power grid regulation simulation function, the adjusted device load rate is calculated, and the training is ended when all the overruns are eliminated. This technical solution solves the problems of large optimization problem solving difficulty, insufficient consideration of actual regulation experience, and decision-making results exceeding the stable boundary in the prior art. For example, the patent application with publication number CN119921301A proposes an "emergency control method, device and equipment for power grid transient frequency based on reinforcement learning", which relates to the fields of transient frequency and emergency control. The method includes: if the power grid is in a fault state, obtaining the current operation data of the power grid; inputting the current operation data into a target transient frequency emergency control model to match an actual emergency control strategy from a plurality of emergency control strategy tables in the target transient frequency emergency control model, wherein the target transient frequency emergency control model is trained based on power grid fault data samples using a preset improved deep deterministic policy gradient algorithm; and performing emergency stability control on the transient frequency of the power grid according to the actual emergency control strategy. Thus, the problem of frequency fluctuation caused by large-scale power grid disturbance is solved, the reliability and economy of power grid stable operation are improved, online real-time emergency control strategies for generator tripping and load shedding can be quickly and accurately generated, and the rapid recovery of transient frequency stability after a fault is achieved. The above existing methods mainly remain in the offline simulation verification stage, focusing on the training effect of the algorithm itself, and lack of key technology research for online deployment, such as safety constraint embedding mechanism. There is no current research on how to prevent operation accidents caused by inappropriate decision-making of the agent in the production environment, and there are deficiencies. SUMMARY

[0004] The purpose of the present application is to solve the problems in the above-mentioned prior art, and to provide a power grid static security risk prevention control method and system based on reinforcement learning, which embeds safety constraints into reinforcement learning agents and sets up a dual safety check link of legality verification and effect pre-evaluation, so that the reinforcement learning agent can be truly embedded into the real-time online closed-loop control process strictly required by the actual power grid, thereby safely and efficiently dealing with the static security risk of the power grid.

[0005] In order to achieve the above-mentioned purpose, the present application has the following technical solutions:

[0006] Firstly, a reinforcement learning-based method for preventing and controlling static security risks in power grids is provided, including:

[0007] Obtain power grid parameters, superimpose security constraints on reinforcement learning objectives, and construct a security embedded reinforcement learning agent;

[0008] For a secure embedded reinforcement learning agent, a security reward function is set up, which includes risk elimination, security incentive, stability margin and control cost, and a reinforcement learning policy network is constructed.

[0009] After the output of the reinforcement learning policy network, a safe projection check is added. The safe projection check is used to force the correction of the original action through projection constraint optimization to obtain the corrected action.

[0010] For the revised feasible actions, a legality verification and effect pre-evaluation process is set up. The execution of the actions is adjusted according to the results of the legality verification and effect pre-evaluation to achieve prevention and control of static safety risks in the power grid.

[0011] As a preferred embodiment, the process of acquiring power grid parameters, superimposing security constraints on the reinforcement learning objective, and constructing a security-embedded reinforcement learning agent includes:

[0012] Set the following nonlinear inequality constraint as a safety boundary:

[0013]

[0014] In the formula, V i V represents the real-time voltage amplitude of the i-th bus. i min This is the minimum allowable voltage limit for the i-th busbar; F is the voltage safety margin function; j For the real-time active power flow of the j-th transmission line / transformer; This represents the maximum allowed transmission capacity of the j-th line; For power flow safety margin function; For the safety margin of the k-th type of stable problem;

[0015] Apply the following formula to the reinforcement learning objective:

[0016]

[0017] In the formula, θ represents the policy network parameters; τ represents finding the maximum value of the policy network parameter θ; τ represents the operating history of the power grid under the control of the agent, which is a state-action-reward sequence. For strategy π θThe expected value of the generated trajectory τ is used to average the effect of all possible power grid operation paths; γ is a discount factor used to quantify the importance of future rewards; r(s) t ,a t ) is the immediate reward function, representing the reward in state s. t Next, execute action a t The benefits; Indicates cumulative discount rewards; The margin function representing the safety constraints. Indicates the expected safety margin; a t For the control action at time t, A feas (s t ) represents the set of actions allowed in the current state, a t ∈A feas (s t () indicates that the action is within the feasible region.

[0018] As a preferred embodiment, in the step of constructing a reinforcement learning policy network by setting a safety reward function that includes a risk elimination term, a safety incentive term, a stability margin term, and a control cost term for the safety embedded reinforcement learning agent, the expression of the safety reward function is as follows:

[0019] r(s t ,a t )=ω1r risk (s t )+ω2r cost (a t )+ω3r safe (s t )+ω4r stab (s t )

[0020] Where ω1, ω2, ω3, and ω4 represent the weight coefficients of the corresponding items;

[0021] The calculation expression for the risk elimination item is:

[0022]

[0023] In the formula, The voltage safety margin of node i is expressed as follows: Among them, V i V is the actual voltage at node i. i min This is the lower limit of the voltage at node i; The line load rate is represented by the following expression: Among them, F j For the actual power flow of line j, This represents the thermal stability limit of line j.

[0024] The calculation expression for the control cost item is as follows:

[0025] r cost (a t )=-(∑||ΔP g,k ||+λ||u discrete ||0)

[0026] In the formula, ΔP g,k Let u be the change in output of the k-th generator. discrete For the number of discrete control operations;

[0027] The calculation expression for the security incentive term is:

[0028]

[0029] In the formula, The stability margin function is expressed as follows: Where, η k Let k be the current value of the stable index. δ is the safety threshold for the kth stable index, and δ is the margin protection threshold.

[0030] The calculation expression for the stability margin term is as follows:

[0031]

[0032] In the formula, ζm represents the damping ratio of the m-th oscillation mode; ζ0 represents the reference damping ratio; and β is the penalty intensity.

[0033] As a preferred embodiment, the modified feasible action is calculated using the following expression:

[0034]

[0035] In the formula, a' t For the revised action, A feas (s t ) represents the feasible set after projection, and π θ (s t ) represents the uncorrected control commands output by the policy network;

[0036] The expression for the security projection verification is as follows:

[0037]

[0038] In the formula, a represents the possible action to be solved, A is the linear constraint matrix, and b is the linear constraint boundary; g k (a t ,s t() represents a nonlinear constraint function determined according to topological and logically related security rules;

[0039] When projection fails, the feasible set A after projection is... feas (s t If ) is empty, the action is corrected using the following formula:

[0040]

[0041] In the formula, κ is a coefficient that balances safety risk and economy.

[0042] As a preferred approach, the legality verification step checks whether the action complies with the hard operational constraints of all devices by using a preset rule base after the action is generated; the effect pre-evaluation step uses a real-time sensitivity matrix to simulate the short-term impact of the corresponding action on key state quantities within milliseconds, and decides whether to accept or reject the corresponding action based on the effect pre-evaluation results.

[0043] Secondly, a power grid static security risk prevention and control system based on reinforcement learning is provided, including:

[0044] The reinforcement learning agent construction module is used to acquire power grid parameters, superimpose security constraints on the reinforcement learning objective, and construct a secure embedded reinforcement learning agent.

[0045] The safety reward function setting module is used to set a safety reward function for a safety embedded reinforcement learning agent, which includes a risk elimination term, a safety incentive term, a stability margin term, and a control cost term, and to construct a reinforcement learning policy network.

[0046] The secure projection verification module is used to add secure projection verification after the output of the reinforcement learning policy network. The secure projection verification forcefully corrects the original action through projection constraint optimization to obtain the corrected action.

[0047] The dual safety verification module is used to set up legality verification and effect pre-evaluation for the modified feasible actions. Based on the results of legality verification and effect pre-evaluation, the execution of the actions is adjusted to achieve prevention and control of static safety risks in the power grid.

[0048] As a preferred embodiment, the reinforcement learning agent construction module sets the following nonlinear inequality constraint as a safety boundary:

[0049]

[0050] In the formula, V i V represents the real-time voltage amplitude of the i-th bus. i min This is the minimum allowable voltage limit for the i-th busbar; F is the voltage safety margin function; j For the real-time active power flow of the j-th transmission line / transformer; This represents the maximum allowed transmission capacity of the j-th line; For power flow safety margin function; For the safety margin of the k-th type of stable problem;

[0051] Apply the following formula to the reinforcement learning objective:

[0052]

[0053] In the formula, θ represents the policy network parameters; τ represents finding the maximum value of the policy network parameter θ; τ represents the operating history of the power grid under the control of the agent, which is a state-action-reward sequence. For strategy π θ The expected value of the generated trajectory τ is used to average the effect of all possible power grid operation paths; γ is a discount factor used to quantify the importance of future rewards; r(s) t ,a t ) is the immediate reward function, representing the reward in state s. t Next, execute action a t The benefits; Indicates cumulative discount rewards; The margin function representing the safety constraints. Indicates the expected safety margin; a t For the control action at time t, A feas (s t ) represents the set of actions allowed in the current state, a t ∈A feas (s t () indicates that the action is within the feasible region.

[0054] As a preferred embodiment, the security reward function expression set by the security reward function setting module is as follows:

[0055] r(s t ,a t )=ω1r risk (s t )+ω2r cost (a t )+ω3r safe (s t )+ω4r stab (s t )

[0056] Where ω1, ω2, ω3, and ω4 represent the weight coefficients of the corresponding items;

[0057] The calculation expression for the risk elimination item is:

[0058]

[0059] In the formula, The voltage safety margin of node i is expressed as follows: Among them, V i V is the actual voltage at node i. i min This is the lower limit of the voltage at node i; The line load rate is represented by the following expression: Among them, F j For the actual power flow of line j, This represents the thermal stability limit of line j.

[0060] The calculation expression for the control cost item is as follows:

[0061] r cost (a t )=-(∑||ΔP g,k ||+λ||u discrete ||0)

[0062] In the formula, ΔP g,k Let u be the change in output of the k-th generator. discrete For the number of discrete control operations;

[0063] The calculation expression for the security incentive term is:

[0064]

[0065] In the formula, The stability margin function is expressed as follows: Where, η k Let k be the current value of the stable index. δ is the safety threshold for the kth stable index, and δ is the margin protection threshold.

[0066] The calculation expression for the stability margin term is as follows:

[0067]

[0068] In the formula, ζ m ζm represents the damping ratio of the m-th oscillation mode; ζ0 represents the reference damping ratio; and β is the penalty intensity.

[0069] As a preferred embodiment, the corrected feasible action of the safety projection verification module is calculated using the following expression:

[0070]

[0071] In the formula, a't For the revised action, A feas (s t ) represents the feasible set after projection, and π θ (s t ) represents the uncorrected control commands output by the policy network;

[0072] The expression for the security projection verification is as follows:

[0073]

[0074] In the formula, a represents the possible action to be solved, A is the linear constraint matrix, and b is the linear constraint boundary; g k (a t ,s t () represents a nonlinear constraint function determined according to topological and logically related security rules;

[0075] When projection fails, the feasible set A after projection is... feas (s t If ) is empty, the action is corrected using the following formula:

[0076]

[0077] In the formula, κ is a coefficient that balances safety risk and economy.

[0078] As a preferred embodiment, the legality verification step of the dual security verification module checks whether the action complies with the hard operation constraints of all devices through a preset rule base after the action is generated; the effect pre-evaluation step of the dual security verification module uses a real-time sensitivity matrix to simulate the short-term impact of the corresponding action on key state quantities within milliseconds, and decides whether to accept or reject the corresponding action based on the effect pre-evaluation results.

[0079] Thirdly, an electronic device is provided, including a processor and a memory, the processor being used to execute a computer program stored in the memory to implement the power grid static security risk prevention and control method based on reinforcement learning as described in the first aspect.

[0080] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing at least one instruction, which, when executed by a processor, implements the power grid static security risk prevention and control method based on reinforcement learning as described in the first aspect.

[0081] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects:

[0082] Static security risks refer to potential safety hazards caused by factors such as equipment performance, grid structure, and operating methods under normal power system operation, which may lead to power outages or system stability problems. This invention superimposes security constraints onto standard reinforcement learning objectives, incorporating risk elimination terms, security incentive terms, and stability margin terms into the reward function. This guides the agent to generate action decisions that better improve grid security. A security projection check is added after the reinforcement learning policy network output to forcibly correct the original actions, ensuring that the actions meet physical feasibility and basic security constraints, thus blocking physically illegal actions at the source. Finally, a dual security verification process of legality checking and effect pre-evaluation is set up to further verify the agent's actions, ensuring that the power system is under control, thereby avoiding safety accidents caused by improper decision-making when the reinforcement learning agent is applied online. This invention enables reinforcement learning agents to be truly embedded into the real-time online closed-loop control process strictly required by actual power grids, safely and efficiently addressing static security risks of the power grid.

[0083] It is understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0084] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0085] Figure 1 Flowchart of the power grid static security risk prevention and control method based on reinforcement learning according to an embodiment of the present invention;

[0086] Figure 2 A schematic diagram of the power grid static safety risk prevention and control system based on reinforcement learning in an embodiment of the present invention. Detailed Implementation

[0087] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0088] Please see Figure 1This invention presents a power grid static safety risk prevention and control method based on reinforcement learning, integrating a high-performance reinforcement learning decision engine with multiple safety assurance mechanisms. By constructing a secure embedded reinforcement learning agent, the reward function imposes severe penalties on high-risk states, provides incentives for good recovery margins, and appropriately penalizes control costs, ensuring higher security for online decision-making. Simultaneously, constraints are added to the output layer of the agent's policy network to ensure that the initially generated actions are within the basic operational range allowed by the physical equipment. Furthermore, a dual safety verification process is implemented, using action legality verification and effect pre-evaluation to further ensure the security of the agent's online application.

[0089] Specifically, the power grid static security risk prevention and control method based on reinforcement learning in this embodiment of the invention includes:

[0090] S1. Obtain power grid parameters, superimpose security constraints on the reinforcement learning objective, and construct a security embedded reinforcement learning agent;

[0091] S2. For a secure embedded reinforcement learning agent, a security reward function is set up, which includes risk elimination, security incentive, stability margin and control cost, and a reinforcement learning policy network is constructed.

[0092] S3. After the output of the reinforcement learning policy network, a safe projection check is added. The safe projection check is used to force the correction of the original action through projection constraint optimization to obtain the corrected action.

[0093] S4. Set up a legality verification and effect pre-evaluation step for the revised feasible actions, and adjust the execution of actions according to the results of legality verification and effect pre-evaluation to achieve prevention and control of static safety risks of the power grid.

[0094] In one possible implementation, step S1 sets a nonlinear inequality constraint as a safety boundary, expressed as follows:

[0095]

[0096] In the formula, V i V represents the real-time voltage amplitude of the i-th bus. i min This is the minimum allowable voltage limit for the i-th busbar; F is the voltage safety margin function; j For the real-time active power flow of the j-th transmission line / transformer; This represents the maximum allowed transmission capacity of the j-th line; For power flow safety margin function; The safety margin for the k-th type of stability problem depends on the weak links of the system, such as the static voltage stability margin and the power angle stability margin.

[0097] Apply the following formula to the reinforcement learning objective:

[0098]

[0099] In the formula, θ represents the policy network parameters, i.e., how the reinforcement learning agent generates control actions based on the power grid state. τ represents finding the maximum value of the policy network parameter θ; τ represents the operating history of the power grid under the control of the agent, which is a state-action-reward sequence. For strategy π θ The mathematical expectation of the generated trajectory τ averages the effect of all possible power grid operation paths (including uncertainties such as load fluctuations and faults); γ is a discount factor, γ∈[0,1), used to quantify the importance of future rewards. When γ≈0, only immediate rewards are considered, such as immediately eliminating over-limits; when γ≈1, long-term safety is also taken into account, such as preventing voltage collapse; r(s t ,a t ) is the immediate reward function, representing the reward in state s. t Next, execute action a t The benefits are determined by both the degree of safety risk elimination and the cost of control. This represents the cumulative discount reward, which is the agent's overall performance over T steps; The margin function representing the safety constraints. This represents the expected safety margin, which needs to be kept non-negative; a t For the control action at time t, A feas (s t (a) represents the set of actions allowed under the current state, defined by the physical laws of the power grid and the limits of the equipment, such as the generator's ramp rate not exceeding the upper limit and capacitors not being continuously switched on and off; t ∈A feas (s t This indicates that the action is within the feasible domain, thus avoiding the generation of physically unexecutable instructions.

[0100] In one possible implementation, step S2 encourages the agent to move toward safer actions by setting a safety reward function that includes risk elimination, safety incentive, stability margin, and control cost.

[0101] The expression for the security reward function is as follows:

[0102] r(s t ,a t )=ω1r risk (s t )+ω2r cost (a t )+ω3r safe (st )+ω4r stab (s t )

[0103] Where ω1, ω2, ω3, and ω4 represent the weight coefficients of the corresponding items;

[0104] The calculation expression for the risk elimination item is:

[0105]

[0106] In the formula, The voltage safety margin of node i is expressed as follows: Among them, V i V is the actual voltage at node i. i min This is the lower limit of the voltage at node i; The line load rate is represented by the following expression: Among them, F j For the actual power flow of line j, This represents the thermal stability limit of line j.

[0107] The calculation expression for the control cost item is as follows:

[0108] r cost (a t )=-(∑||ΔP g,k ||+λ||u discrete ||0)

[0109] In the formula, ΔP g,k Let u be the change in output of the k-th generator. discrete For the number of discrete control operations;

[0110] The calculation expression for the security incentive term is:

[0111]

[0112] In the formula, The stability margin function is expressed as follows: Where, η k Let k be the current value of the stable index. δ is the safety threshold for the kth stable index, and δ is the margin protection threshold.

[0113] The calculation expression for the stability margin term is as follows:

[0114]

[0115] In the formula, ζ mζm represents the damping ratio of the m-th oscillation mode; ζ0 represents the reference damping ratio, which is a constant value; and β is the penalty intensity, which is also a constant.

[0116] In one possible implementation, step S3 adds a security projection check after the output of the reinforcement learning policy network to block physical violations at the source.

[0117] The revised possible actions are calculated using the following expression:

[0118]

[0119] In the formula, a' t For the revised action, A feas (s t ) represents the feasible set after projection, and π θ (s t ) represents the uncorrected control commands output by the policy network;

[0120] The expression for the security projection verification is as follows:

[0121]

[0122] In the formula, a represents the possible action to be solved, A is the linear constraint matrix, and b is the linear constraint boundary; g k (a t ,s t () represents a nonlinear constraint function determined according to topological and logically related security rules;

[0123] When projection fails, the feasible set A after projection is... feas (s t If ) is empty, the action is corrected using the following formula:

[0124]

[0125] In the formula, κ is a coefficient that balances safety risk and economy.

[0126] In one possible implementation, step S4 enhances the security of online applications of reinforcement learning agents by setting up a legality verification and effect pre-evaluation process.

[0127] In the legality verification stage, after the action is generated, the system checks whether the action complies with the hard operating constraints of all equipment, such as on / off status, maximum / minimum output, and ramp rate, using a preset rule base. In the effect pre-evaluation stage, a fast power flow model based on a real-time sensitivity matrix is ​​used to simulate the short-term impact of the action on key state variables within milliseconds. If the pre-evaluation results indicate that the system state will deteriorate after the action is executed, such as a further drop in voltage, further exceedance of power flow limits, or failure to effectively eliminate static safety risks, the action is rejected, and the system switches to a preset traditional reliable backup control strategy, such as a manually preset strategy or a sensitivity-based generator / load shedding strategy, to ensure that the power system is always under control.

[0128] This invention proposes a systematic solution that integrates a high-performance reinforcement learning decision engine with multiple security mechanisms. By embedding security constraints into the reinforcement learning agent and implementing dual security checks, the reinforcement learning agent can be truly embedded into the real-time online closed-loop control process that is strictly required by the actual power grid, thus safely and efficiently addressing the static security risks of the power grid.

[0129] Please see Figure 2 Another embodiment of the present invention also proposes a power grid static security risk prevention and control system based on reinforcement learning, comprising:

[0130] The reinforcement learning agent construction module 201 is used to acquire power grid parameters, superimpose security constraints on the reinforcement learning objective, and construct a secure embedded reinforcement learning agent.

[0131] The safety reward function setting module 202 is used to set a safety reward function for a safety embedded reinforcement learning agent, which includes a risk elimination term, a safety incentive term, a stability margin term, and a control cost term, and to construct a reinforcement learning policy network.

[0132] The secure projection verification module 203 is used to add secure projection verification after the output of the reinforcement learning policy network. The secure projection verification forcefully corrects the original action through projection constraint optimization to obtain the corrected action.

[0133] The dual safety verification module 204 is used to set up a legality verification and effect pre-evaluation stage for the modified action. The execution of the action is adjusted according to the results of the legality verification and effect pre-evaluation to realize the prevention and control of static safety risks of the power grid.

[0134] In one possible implementation, the reinforcement learning agent building module 201 sets a nonlinear inequality constraint as a safety boundary, expressed as follows:

[0135]

[0136] In the formula, Vi V represents the real-time voltage amplitude of the i-th bus. i min This is the minimum allowable voltage limit for the i-th busbar; F is the voltage safety margin function; j For the real-time active power flow of the j-th transmission line / transformer; This represents the maximum allowed transmission capacity of the j-th line; For power flow safety margin function; For the safety margin of the k-th type of stable problem;

[0137] Apply the following formula to the reinforcement learning objective:

[0138]

[0139] In the formula, θ represents the policy network parameters; τ represents finding the maximum value of the policy network parameter θ; τ represents the operating history of the power grid under the control of the agent, which is a state-action-reward sequence. For strategy π θ The expected value of the generated trajectory τ is used to average the effect of all possible power grid operation paths; γ is a discount factor used to quantify the importance of future rewards; r(s) t ,a t ) is the immediate reward function, representing the reward in state s. t Next, execute action a t The benefits; Indicates cumulative discount rewards; The margin function representing the safety constraints. Indicates the expected safety margin; a t For the control action at time t, A feas (s t ) represents the set of actions allowed in the current state, a t ∈A feas (s t () indicates that the action is within the feasible region.

[0140] In one possible implementation, the security reward function expression set by the security reward function setting module 202 is as follows:

[0141] r(s t ,a t )=ω1r risk (s t )+ω2r cost (a t )+ω3r safe (s t )+ω4r stab (s t )

[0142] Where ω1, ω2, ω3, and ω4 represent the weight coefficients of the corresponding items;

[0143] The calculation expression for the risk elimination item is:

[0144]

[0145] In the formula, The voltage safety margin of node i is expressed as follows: Among them, V i V is the actual voltage at node i. i min This is the lower limit of the voltage at node i; The line load rate is represented by the following expression: Among them, F j For the actual power flow of line j, This represents the thermal stability limit of line j.

[0146] The calculation expression for the control cost item is as follows:

[0147] r cost (a t )=-(∑||ΔP g,k ||+λ||u discrete ||0)

[0148] In the formula, ΔP g,k Let u be the change in output of the k-th generator. discrete For the number of discrete control operations;

[0149] The calculation expression for the security incentive term is:

[0150]

[0151] In the formula, The stability margin function is expressed as follows: Where, η k Let k be the current value of the stable index. δ is the safety threshold for the kth stable index, and δ is the margin protection threshold.

[0152] The calculation expression for the stability margin term is as follows:

[0153]

[0154] In the formula, ζ m ζm represents the damping ratio of the m-th oscillation mode; ζ0 represents the reference damping ratio; and β is the penalty intensity.

[0155] In one possible implementation, the modified feasible action of the safety projection verification module 203 is calculated using the following expression:

[0156]

[0157] In the formula, a' t For the revised action, A feas (s t ) represents the feasible set after projection, and π θ (s t ) represents the uncorrected control commands output by the policy network;

[0158] The expression for the security projection verification is as follows:

[0159]

[0160] In the formula, a represents the possible action to be solved, A is the linear constraint matrix, and b is the linear constraint boundary; g k (a t ,s t () represents a nonlinear constraint function determined according to topological and logically related security rules;

[0161] When projection fails, the feasible set A after projection is... feas (s t If ) is empty, the action is corrected using the following formula:

[0162]

[0163] In the formula, κ is a coefficient that balances safety risk and economy.

[0164] In one possible implementation, the dual security verification module sets up a legality verification step 204. After the action is generated, it checks whether the action complies with the hard operation constraints of all devices through a preset rule base. The effect pre-evaluation step set up by the dual security verification module uses a real-time sensitivity matrix to simulate the short-term impact of the corresponding action on key state quantities within milliseconds, and decides whether to accept or reject the corresponding action based on the effect pre-evaluation results.

[0165] Another embodiment of the present invention also proposes an electronic device, including a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the reinforcement learning-based power grid static security risk prevention and control method.

[0166] Another embodiment of the present invention also proposes a computer-readable storage medium storing at least one instruction, which, when executed by a processor, implements the reinforcement learning-based power grid static security risk prevention and control method.

[0167] The computer program includes computer program code, which can be in the form of source code, object code, executable file, or some intermediate form. The computer-readable storage medium can include any entity or device capable of carrying the computer program code, a medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals. For ease of explanation, the above content only shows the parts related to the embodiments of the present invention; for specific technical details not disclosed, please refer to the method section of the embodiments of the present invention. This computer-readable storage medium is non-transitory and can be stored in storage devices formed by various electronic devices, enabling the execution process described in the method of the embodiments of the present invention.

[0168] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0169] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0170] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0171] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for preventing and controlling static security risks in power grids based on reinforcement learning, characterized in that, include: Obtain power grid parameters, superimpose security constraints on reinforcement learning objectives, and construct a security embedded reinforcement learning agent; For a secure embedded reinforcement learning agent, a security reward function is set up, which includes risk elimination, security incentive, stability margin and control cost, and a reinforcement learning policy network is constructed. After the output of the reinforcement learning policy network, a safe projection check is added. The safe projection check is used to force the correction of the original action through projection constraint optimization to obtain the corrected action. For the revised feasible actions, a legality verification and effect pre-evaluation process is set up. The execution of the actions is adjusted according to the results of the legality verification and effect pre-evaluation to achieve prevention and control of static safety risks in the power grid.

2. The power grid static security risk prevention and control method based on reinforcement learning according to claim 1, characterized in that, The process of acquiring power grid parameters, superimposing security constraints on the reinforcement learning objective, and constructing a security-embedded reinforcement learning agent includes: Set the following nonlinear inequality constraint as a safety boundary: In the formula, V i V represents the real-time voltage amplitude of the i-th bus. i min This is the minimum allowable voltage limit for the i-th busbar; F is the voltage safety margin function; j For the real-time active power flow of the j-th transmission line / transformer; This represents the maximum allowed transmission capacity of the j-th line; For power flow safety margin function; For the safety margin of the k-th type of stable problem; Apply the following formula to the reinforcement learning objective: In the formula, θ represents the policy network parameters; τ represents finding the maximum value of the policy network parameter θ; τ represents the operating history of the power grid under the control of the agent, which is a state-action-reward sequence. For strategy π θ The expected value of the generated trajectory τ is used to average the effect of all possible power grid operation paths; γ is a discount factor used to quantify the importance of future rewards; r(s) t ,a t ) is the immediate reward function, representing the reward in state s. t Next, execute action a t The benefits; Indicates cumulative discount rewards; The margin function representing the safety constraints. Indicates the expected safety margin; a t For the control action at time t, A feas (s t ) represents the set of actions allowed in the current state, a t ∈A feas (s t () indicates that the action is within the feasible region.

3. The power grid static security risk prevention and control method based on reinforcement learning according to claim 1, characterized in that, In the step of constructing a reinforcement learning policy network by setting a safety reward function that includes a risk elimination term, a safety incentive term, a stability margin term, and a control cost term for a safety embedded reinforcement learning agent, the expression of the safety reward function is as follows: r(s t ,a t )=ω1r risk (s t )+ω2r cost (a t )+ω3r safe (s t )+ω4r stab (s t ) Where ω1, ω2, ω3, and ω4 represent the weight coefficients of the corresponding items; The calculation expression for the risk elimination item is: In the formula, The voltage safety margin of node i is expressed as follows: Among them, V i V is the actual voltage at node i. i min This is the lower limit of the voltage at node i; The line load factor is expressed as follows: Among them, F j For the actual power flow of line j, This represents the thermal stability limit of line j. The calculation expression for the control cost item is as follows: r cost (a t )=-(∑||ΔP g,k ||+λ||u discrete ||0) In the formula, ΔP g,k Let u be the change in output of the k-th generator. discrete For the number of discrete control operations; The calculation expression for the security incentive term is: In the formula, The stability margin function is expressed as follows: Where, η k Let k be the current value of the stable index. δ is the safety threshold for the kth stable index, and δ is the margin protection threshold. The calculation expression for the stability margin term is as follows: In the formula, ζ m ζm represents the damping ratio of the m-th oscillation mode; ζ0 represents the reference damping ratio; and β is the penalty intensity.

4. The power grid static security risk prevention and control method based on reinforcement learning according to claim 1, characterized in that, The revised feasible action is calculated using the following expression: In the formula, a' t For the revised action, A feas (s t ) represents the feasible set after projection, and π θ (s t ) represents the uncorrected control commands output by the policy network; The expression for the security projection verification is as follows: stAa≤b g k (a t ,s t )≥0 In the formula, a represents the possible action to be solved, A is the linear constraint matrix, and b is the linear constraint boundary; g k (a t ,s t () represents a nonlinear constraint function determined according to topological and logically related security rules; When projection fails, the feasible set A after projection is... feas (s t If ) is empty, the action is corrected using the following formula: In the formula, κ is a coefficient that balances safety risk and economy.

5. The power grid static security risk prevention and control method based on reinforcement learning according to claim 1, characterized in that, In the legality verification stage, after the action is generated, the action is checked against the hard operation constraints of all devices using a preset rule base. In the effect pre-evaluation stage, the short-term impact of the corresponding action on key state variables is simulated within milliseconds using a real-time sensitivity matrix, and the decision to accept or reject the corresponding action is made based on the effect pre-evaluation results.

6. A power grid static security risk prevention and control system based on reinforcement learning, characterized in that, include: The reinforcement learning agent construction module is used to acquire power grid parameters, superimpose security constraints on the reinforcement learning objective, and construct a secure embedded reinforcement learning agent. The safety reward function setting module is used to set a safety reward function for a safety embedded reinforcement learning agent, which includes a risk elimination term, a safety incentive term, a stability margin term, and a control cost term, and to construct a reinforcement learning policy network. The secure projection verification module is used to add secure projection verification after the output of the reinforcement learning policy network. The secure projection verification forcefully corrects the original action through projection constraint optimization to obtain the corrected action. The dual safety verification module is used to set up legality verification and effect pre-evaluation for the modified feasible actions. Based on the results of legality verification and effect pre-evaluation, the execution of the actions is adjusted to achieve prevention and control of static safety risks in the power grid.

7. The power grid static security risk prevention and control system based on reinforcement learning according to claim 6, characterized in that, The reinforcement learning agent construction module sets the following nonlinear inequality constraint as a safety boundary: In the formula, V i V represents the real-time voltage amplitude of the i-th bus. i min This is the minimum allowable voltage limit for the i-th busbar; F is the voltage safety margin function; j For the real-time active power flow of the j-th transmission line / transformer; This represents the maximum allowed transmission capacity of the j-th line; For power flow safety margin function; For the safety margin of the k-th type of stable problem; Apply the following formula to the reinforcement learning objective: In the formula, θ represents the policy network parameters; τ represents finding the maximum value of the policy network parameter θ; τ represents the operating history of the power grid under the control of the agent, which is a state-action-reward sequence. For strategy π θ The expected value of the generated trajectory τ is used to average the effect of all possible power grid operation paths; γ is a discount factor used to quantify the importance of future rewards; r(s) t ,a t ) is the immediate reward function, representing the reward in state s. t Next, execute action a t The benefits; Indicates cumulative discount rewards; The margin function representing the safety constraints. Indicates the expected safety margin; a t For the control action at time t, A feas (s t ) represents the set of actions allowed in the current state, a t ∈A feas (s t () indicates that the action is within the feasible region.

8. The power grid static security risk prevention and control system based on reinforcement learning according to claim 6, characterized in that, The security reward function expression set by the security reward function setting module is as follows: r(s t ,a t )=ω1r risk (s t )+ω2r cost (a t )+ω3r safe (s t )+ω4r stab (s t ) Where ω1, ω2, ω3, and ω4 represent the weight coefficients of the corresponding items; The calculation expression for the risk elimination item is: In the formula, The voltage safety margin of node i is expressed as follows: Among them, V i V is the actual voltage at node i. i min This is the lower limit of the voltage at node i; The line load factor is expressed as follows: Among them, F j For the actual power flow of line j, This represents the thermal stability limit of line j. The calculation expression for the control cost item is as follows: r cost (a t )=-(∑||ΔP g,k ||+λ||u discrete ||0) In the formula, ΔP g,k Let u be the change in output of the k-th generator. discrete For the number of discrete control operations; The calculation expression for the security incentive term is: In the formula, The stability margin function is expressed as follows: Where, η k Let k be the current value of the stable index. δ is the safety threshold for the kth stable index, and δ is the margin protection threshold. The calculation expression for the stability margin term is as follows: In the formula, ζ m ζm represents the damping ratio of the m-th oscillation mode; ζ0 represents the reference damping ratio; and β is the penalty intensity.

9. The power grid static security risk prevention and control system based on reinforcement learning according to claim 6, characterized in that, The corrected feasible actions of the safety projection verification module are calculated using the following expression: In the formula, a' t For the revised action, A feas (s t ) represents the feasible set after projection, and π θ (s t ) represents the uncorrected control commands output by the policy network; The expression for the security projection verification is as follows: stAa≤b g k (a t ,s t )≥0 In the formula, a represents the possible action to be solved, A is the linear constraint matrix, and b is the linear constraint boundary; g k (a t ,s t () represents a nonlinear constraint function determined according to topological and logically related security rules; When projection fails, the feasible set A after projection is... feas (s t If ) is empty, the action is corrected using the following formula: In the formula, κ is a coefficient that balances safety risk and economy.

10. The power grid static security risk prevention and control system based on reinforcement learning according to claim 6, characterized in that, The legality verification step set in the dual security verification module checks whether the action complies with the hard operation constraints of all devices after the action is generated, using a preset rule base. The dual security verification module is set up with an effect pre-evaluation step, which uses a real-time sensitivity matrix to simulate the short-term impact of the corresponding action on key state quantities within milliseconds, and decides whether to accept or reject the corresponding action based on the effect pre-evaluation results.

11. An electronic device, characterized in that, It includes a processor and a memory, the processor being used to execute a computer program stored in the memory to implement the power grid static security risk prevention and control method based on reinforcement learning as described in any one of claims 1 to 5.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, which, when executed by a processor, implements the power grid static security risk prevention and control method based on reinforcement learning as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Assistant decision-making method for static security risk of power grid and related device

    CN118886341A

  • Power grid transient frequency emergency control method, device and equipment based on reinforcement learning

    CN119921301A