A transient voltage stability multi-period prevention control method, system, medium, device and product

By adopting a multi-period transient voltage stability prevention and control method based on DDPG reinforcement learning network, the problem of insufficient identification of multi-period safety hazards in the power grid in the existing technology is solved, and transient voltage stability control with rapid response to power grid changes is realized, thereby improving the flexibility and reliability of power grid operation.

CN122136833APending Publication Date: 2026-06-02STATE GRID ELECTRIC POWER RES INST +2

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID ELECTRIC POWER RES INST
Filing Date
2025-12-24
Publication Date
2026-06-02

Smart Images

  • Figure CN122136833A_ABST
    Figure CN122136833A_ABST
Patent Text Reader

Abstract

This invention discloses a method, system, medium, equipment, and product for multi-period transient voltage stability prevention and control in the field of power system automation technology. The method includes: constructing a multi-period transient voltage stability prevention and control model based on the power grid topology; acquiring multi-period power grid operation mode data based on real-time power grid operating status information, dispatch plans, renewable energy forecast information, and load forecast information; and generating a multi-period prevention and control strategy using the trained multi-period transient voltage stability prevention and control model based on the multi-period power grid operation mode data, and adjusting the power grid operating status in real time. This invention can address transient voltage stability problems under highly uncertain power grid scenarios, quickly generate transient voltage stability prevention and control auxiliary decisions, and provide a reliable basis for reactive power optimization and stability control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system automation technology, and in particular to a method, system, medium, equipment and product for multi-period prevention and control of transient voltage stability. Background Technology

[0002] With the construction of new power systems, the grid-connected capacity of new energy units is constantly increasing, and a large number of conventional units are being replaced, resulting in a significant decrease in the reactive power and voltage support capacity of the power grid. Severe AC / DC faults may lead to transient voltage instability.

[0003] Currently, existing transient voltage stability prevention and control technologies only target a single time segment, failing to fully utilize predictive information and making it difficult to effectively identify and eliminate power grid safety hazards across multiple time periods. There is a contradiction between the control strategy "not keeping up" with system state changes and "not being able to cover" extreme random fluctuations in the system, which may lead to excessive control costs, complex strategy forms, or even the inability to find a feasible solution.

[0004] One type of technology significantly improves the accuracy and robustness of transient voltage stability assessment by modifying the loss function and sample weights to train deep neural networks, but has not yet formed a preventive control framework for multi-period complex scenarios that can proactively optimize control decisions. Another type of technology fits the implicit mapping of control variables in the transient voltage stability index into an explicit polynomial using the collocation method, embedding transient stability constraints into a nonlinear programming model. The preventive control strategy can be solved quickly using the interior-point method. Examples show that the transient voltage stability index of the system approaches zero after correction. However, the polynomial needs to be calibrated offline, and the selection of the order is sensitive to accuracy and computational cost. At present, there is no effective method to solve the problem of rapid changes in transient voltage state in the future scenario of the power grid. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method, system, medium, device and product for multi-period prevention and control of transient voltage stability, which can deal with the transient voltage stability problem under the strong uncertainty of the power grid, quickly generate auxiliary decision for transient voltage stability prevention and control, and provide a reliable basis for reactive power optimization and stability control.

[0006] To solve the above-mentioned technical problems, the present invention is implemented using the following technical solution:

[0007] In a first aspect, the present invention provides a method for preventing and controlling transient voltage stabilization over multiple time periods, comprising:

[0008] Based on the power grid topology, a multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network is constructed.

[0009] Based on real-time power grid operation status information, dispatching plans, new energy forecast information and load forecast information, power system simulation software is used to obtain real-time power grid operation mode data for multiple time periods.

[0010] Based on the real-time operation data of the power grid in multiple time periods, a multi-time period prevention and control strategy is generated using the trained multi-time period transient voltage stability prevention and control model.

[0011] The operating status of the power grid is adjusted in real time according to the multi-period prevention and control strategy.

[0012] Optionally, the multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network includes a state space, an action space, and a reward function;

[0013] The expression for the state space is as follows:

[0014] ,

[0015] in, Representing the state space, Indicates the first The voltage amplitude at each node, Indicates the first The reactive power output of the generator Indicates the first The reactive power output of the Taiwanese capacitor Indicates the first The reactive power output of the Taiwan power reactor Indicates the first Reactive power of each load node Indicates the first The reactive power output of the SVG in Taiwan Indicates the first The reactive power output of the Taiwan SVC;

[0016] The expression for the action space is as follows:

[0017] ,

[0018] in, Represents the action space. Indicates the first The reactive power adjustment of the generator. Indicates the first The reactive power adjustment of the capacitor bank. Indicates the first Reactive power adjustment of the Taiwan reactor Indicates the first Reactive power adjustment of each load node Indicates the first The reactive power adjustment of the SVG. Indicates the first The reactive power adjustment of the SVC;

[0019] The expression for the reward function is as follows:

[0020] ,

[0021] in, Represents the reward function, Indicates the decision-making window, This represents a multi-period transient voltage stability risk function. This represents the prevention and control constraint function. This signifies the punishment for an unchecked trend.

[0022] Optionally, the multi-time transient voltage stability risk function The expression is as follows:

[0023] ,

[0024] in, Represents the set of scheduling periods. Represents a set of scenes. Display decision window Inside Scheduling period in the scenario Cost indicators of transient voltage stabilization control measures; Decision-making window Inside Scheduling period in the scenario Power outage loss indicators.

[0025] Optionally, the prevention and control constraint function It includes at least one of the following: power flow constraints, voltage and phase angle magnitude constraints at each node, generator output constraints, reactive power adjustment constraints of capacitive reactors, active and reactive power constraints of load receiving, reactive power adjustment constraints of SVC and reactive power adjustment constraints of SVG.

[0026] Optionally, the expression for the power flow constraint is as follows:

[0027] ,

[0028] in, This represents the set of nodes excluding the balancing node. Represents a node The active power of power generation, Represents a node The merits, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents the cosine function. Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them Represents the sine function. Represents the set of PQ nodes. Represents a node reactive power, Represents a node reactive power demand, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them;

[0029] The expressions for the voltage and phase angle magnitude constraints at each node are as follows:

[0030] ,

[0031] in, Represents a node voltage, Represents a node Rated voltage, Represents a node phase angle, Represents a node The maximum phase angle, Represents a node The minimum phase angle;

[0032] The expression for the generator output constraint is as follows:

[0033] ,

[0034] in, Indicates the first The operating power of the generators, Indicates the first The generator operates at its minimum output power. Indicates the first The generator's maximum operating output power Indicates the first Reactive power of generators Indicates the first The generator operates at its minimum reactive power output. Indicates the first The generator's maximum reactive power output is [missing information]. Indicates the total number of generators;

[0035] The expression for the reactive power adjustment constraint of the capacitive reactor is as follows:

[0036] ,

[0037] in, This indicates the total number of capacitors. Indicates the total number of reactors. Indicates the first Taiwan capacitor outputs reactive power. Indicates the first The minimum output reactive power of the capacitor. Indicates the first The maximum reactive power output of the capacitor is... Indicates the first Taiwan reactor output reactive power Indicates the first Minimum output reactive power of the reactor Indicates the first The maximum reactive power output of the Taiwan reactor;

[0038] The expression for the active and reactive power constraints of the load is as follows:

[0039] ,

[0040] in, Indicates the first Active power received by each load node Indicates the first Minimum power received by each load node, Indicates the first The maximum operating power received by each load node, Indicates the first Reactive power of each load node Indicates the first The minimum reactive power required for the operation of each load node. Indicates the first The maximum reactive power of each load node during operation. Indicates the total number of load nodes;

[0041] The expression for the SVC reactive power adjustment constraint is as follows:

[0042] ,

[0043] in, Indicates the total number of SVCs. For the first The SVC outputs reactive power. For the first Minimum output reactive power of the SVC unit For the first Maximum reactive power output of the SVC unit;

[0044] The expression for the SVG reactive power adjustment constraint is as follows:

[0045] ,

[0046] in, Indicates the total number of SVGs. Indicates the first The SVG outputs reactive power. Indicates the first Minimum output reactive power of SVG Indicates the first The maximum reactive power output of the SVG.

[0047] Optionally, the multi-time transient voltage stability prevention and control model is trained offline based on the constructed reinforcement learning training sample set to obtain the trained multi-time transient voltage stability prevention and control model, including:

[0048] Initialize the parameters of the decision network and evaluation network in the multi-time transient voltage stability prevention and control model. Based on the reinforcement learning training sample set, repeat the following steps until the preset termination condition is met:

[0049] Randomly sampled from the reinforcement learning training sample set There are 10 samples, each including the current state. ,action Reward Value and the next state ;in, Indicates the sample sequence number;

[0050] Use a decision network to obtain the next state for each sample The following action ;in, Representation Decision Network Parameters;

[0051] Use the evaluation network to obtain the next state Next action Action value Q value ;in, Indicates evaluation network Parameters;

[0052] Using an evaluation network to obtain state Next action Action value Q value ;

[0053] According to the state Next action Action value Q value In state Next action Actual reward value and in the next state Next action Action value Q value Calculate the loss value of the evaluation network. ;

[0054] Based on the loss value of the evaluation network The parameters of the evaluation network are updated using the gradient descent method to obtain the updated evaluation network parameters. ;

[0055] Using decision networks to determine the state of each sample Generate new actions ;

[0056] Use the updated evaluation network to obtain the state Next action Action value Q value ;

[0057] According to the state Next action Action value Q value Calculate the loss of the decision network ;

[0058] According to the loss of the decision network The parameters of the decision network are updated using the gradient descent method to obtain the updated evaluation network parameters. .

[0059] Optionally, the reinforcement learning training sample set is constructed through the following steps:

[0060] Repeat the following steps until a predetermined number of reinforcement learning training samples are obtained, thus obtaining the reinforcement learning training sample set:

[0061] Active load at nodes The reactive load at the node is within the preset normal level range of [0, k1]. Active power of node-linked units within the preset normal level range of [0, k2] times. The grid operation status information is formed by randomly selecting values ​​under the constraints of the rated power range of the generator unit in the range of [0,1] times the rated power range and the reactive power regulation equipment CQ being within the upper and lower limits of the reactive power output range; where k1 and k2 are preset parameters.

[0062] Using power system simulation software, simulations are performed based on preset scheduling plans, new energy forecast information, load forecast information, and the power grid operation status information to obtain power grid operation mode data for multiple time periods.

[0063] Power flow calculation and simulation verification are performed on the power grid's multi-period operation mode data. If the power flow verification converges, the power grid's multi-period operation mode data is taken as the current state of the power grid's operation mode. ;in, Indicates the sample sequence number;

[0064] Based on a pre-set set of anticipated faults, a voltage safety and stability prevention and control strategy is implemented to determine the current state of the power grid operation mode. Select Action By executing actions in power system simulation software To obtain the next state and set the current state. Execute action Reward value =1;

[0065] Based on the current state ,action Reward Value and the next state Constructing training samples .

[0066] Secondly, the present invention provides a transient voltage stabilization multi-period prevention and control system, comprising:

[0067] The model building module is used to: construct a multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network according to the power grid topology;

[0068] The simulation module is used to: obtain real-time data on the power grid's operation mode in multiple time periods using power system simulation software, based on real-time power grid operation status information, scheduling plans, new energy forecast information, and load forecast information.

[0069] The strategy generation module is used to generate multi-period prevention and control strategies based on the real-time operation mode data of the power grid in multiple time periods and using the trained multi-period transient voltage stability prevention and control model.

[0070] The operation status adjustment module is used to adjust the operation status of the power grid in real time according to the multi-time period prevention and control strategy.

[0071] Optionally, the multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network includes a state space, an action space, and a reward function;

[0072] The expression for the state space is as follows:

[0073] ,

[0074] in, Representing the state space, Indicates the first The voltage amplitude at each node, Indicates the first The reactive power output of the generator Indicates the first The reactive power output of the Taiwanese capacitor Indicates the first The reactive power output of the Taiwan power reactor Indicates the first Reactive power of each load node Indicates the first The reactive power output of the SVG in Taiwan Indicates the first The reactive power output of the Taiwan SVC;

[0075] The expression for the action space is as follows:

[0076] ,

[0077] in, Represents the action space. Indicates the first The reactive power adjustment of the generator. Indicates the first The reactive power adjustment of the capacitor bank. Indicates the first Reactive power adjustment of the Taiwan reactor Indicates the first Reactive power adjustment of each load node Indicates the first The reactive power adjustment of the SVG. Indicates the first The reactive power adjustment of the SVC;

[0078] The expression for the reward function is as follows:

[0079] ,

[0080] in, Represents the reward function, Indicates the decision-making window, This represents a multi-period transient voltage stability risk function. This represents the prevention and control constraint function. This signifies the punishment for an unchecked trend.

[0081] Optionally, the multi-time transient voltage stability risk function The expression is as follows:

[0082] ,

[0083] in, Represents the set of scheduling periods. Represents a set of scenes. Display decision window Inside Scheduling period in the scenario Cost indicators of transient voltage stabilization control measures; Decision-making window Inside Scheduling period in the scenario Power outage loss indicators.

[0084] Optionally, the prevention and control constraint function It includes at least one of the following: power flow constraints, voltage and phase angle magnitude constraints at each node, generator output constraints, reactive power adjustment constraints of capacitive reactors, active and reactive power constraints of load receiving, reactive power adjustment constraints of SVC and reactive power adjustment constraints of SVG.

[0085] Optionally, the expression for the power flow constraint is as follows:

[0086] ,

[0087] in, This represents the set of nodes excluding the balancing node. Represents a node The active power of power generation, Represents a node The merits, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents the cosine function. Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them Represents the sine function. Represents the set of PQ nodes. Represents a node reactive power, Represents a node reactive power demand, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them;

[0088] The expressions for the voltage and phase angle magnitude constraints at each node are as follows:

[0089] ,

[0090] in, Represents a node voltage, Represents a node Rated voltage, Represents a node phase angle, Represents a node The maximum phase angle, Represents a node The minimum phase angle;

[0091] The expression for the generator output constraint is as follows:

[0092] ,

[0093] in, Indicates the first The operating power of the generators, Indicates the first The generator operates at its minimum output power. Indicates the first The generator's maximum operating output power Indicates the first Reactive power of generators Indicates the first The generator operates at its minimum reactive power output. Indicates the first The generator's maximum reactive power output is [missing information]. Indicates the total number of generators;

[0094] The expression for the reactive power adjustment constraint of the capacitive reactor is as follows:

[0095] ,

[0096] in, This indicates the total number of capacitors. Indicates the total number of reactors. Indicates the first Taiwan capacitor outputs reactive power. Indicates the first The minimum output reactive power of the capacitor. Indicates the first The maximum reactive power output of the capacitor is... Indicates the first Taiwan reactor output reactive power Indicates the first Minimum output reactive power of the reactor Indicates the first The maximum reactive power output of the Taiwan reactor;

[0097] The expression for the active and reactive power constraints of the load is as follows:

[0098] ,

[0099] in, Indicates the first Active power received by each load node Indicates the first Minimum power received by each load node, Indicates the first The maximum operating power received by each load node, Indicates the first Reactive power of each load node Indicates the first The minimum reactive power required for the operation of each load node. Indicates the first The maximum reactive power of each load node during operation. Indicates the total number of load nodes;

[0100] The expression for the SVC reactive power adjustment constraint is as follows:

[0101] ,

[0102] in, Indicates the total number of SVCs. For the first The SVC outputs reactive power. For the first Minimum output reactive power of the SVC unit For the first Maximum reactive power output of the SVC unit;

[0103] The expression for the SVG reactive power adjustment constraint is as follows:

[0104] ,

[0105] in, Indicates the total number of SVGs. Indicates the first The SVG outputs reactive power. Indicates the first Minimum output reactive power of SVG Indicates the first The maximum reactive power output of the SVG.

[0106] Optionally, the multi-time transient voltage stability prevention and control model is trained offline based on the constructed reinforcement learning training sample set to obtain the trained multi-time transient voltage stability prevention and control model, including:

[0107] Initialize the parameters of the decision network and evaluation network in the multi-time transient voltage stability prevention and control model. Based on the reinforcement learning training sample set, repeat the following steps until the preset termination condition is met:

[0108] Randomly sampled from the reinforcement learning training sample set There are 10 samples, each including the current state. ,action Reward Value and the next state ;in, Indicates the sample sequence number;

[0109] Use a decision network to obtain the next state for each sample The following action ;in, Representation Decision Network Parameters;

[0110] Use the evaluation network to obtain the next state Next action Action value Q value ;in, Indicates evaluation network Parameters;

[0111] Using an evaluation network to obtain state Next action Action value Q value ;

[0112] According to the state Next action Action value Q value In state Next action Actual reward value and in the next state Next action Action value Q value Calculate the loss value of the evaluation network. ;

[0113] Based on the loss value of the evaluation network The parameters of the evaluation network are updated using the gradient descent method to obtain the updated evaluation network parameters. ;

[0114] Using decision networks to determine the state of each sample Generate new actions ;

[0115] Use the updated evaluation network to obtain the state Next action Action value Q value ;

[0116] According to the state Next action Action value Q value Calculate the loss of the decision network ;

[0117] According to the loss of the decision network The parameters of the decision network are updated using the gradient descent method to obtain the updated evaluation network parameters. .

[0118] Optionally, the reinforcement learning training sample set is constructed through the following steps:

[0119] Repeat the following steps until a predetermined number of reinforcement learning training samples are obtained, thus obtaining the reinforcement learning training sample set:

[0120] Active load at nodes The reactive load at the node is within the preset normal level range of [0, k1]. Active power of node-linked units within the preset normal level range of [0, k2] times. The grid operation status information is formed by randomly selecting values ​​under the constraints of the rated power range of the generator unit in the range of [0,1] times the rated power range and the reactive power regulation equipment CQ being within the upper and lower limits of the reactive power output range; where k1 and k2 are preset parameters.

[0121] Using power system simulation software, simulations are performed based on preset scheduling plans, new energy forecast information, load forecast information, and the power grid operation status information to obtain power grid operation mode data for multiple time periods.

[0122] Power flow calculation and simulation verification are performed on the power grid's multi-period operation mode data. If the power flow verification converges, the power grid's multi-period operation mode data is taken as the current state of the power grid's operation mode. ;in, Indicates the sample sequence number;

[0123] Based on a pre-set set of anticipated faults, a voltage safety and stability prevention and control strategy is implemented to determine the current state of the power grid operation mode. Select Action By executing actions in power system simulation software To obtain the next state and set the current state. Execute action Reward value =1;

[0124] Based on the current state ,action Reward Value and the next state Constructing training samples .

[0125] Thirdly, the present invention provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the transient voltage stabilization multi-period prevention and control method described in any of the first aspects.

[0126] Fourthly, the present invention provides a computer device, comprising:

[0127] Memory, used to store computer instructions;

[0128] A processor for executing the computer instructions to implement the steps of the transient voltage stability multi-period prevention and control method according to any of the first aspects.

[0129] Fifthly, the present invention provides a computer program product, including computer instructions, characterized in that, when the computer instructions are executed by a processor, they implement the steps of the transient voltage stabilization multi-period prevention and control method described in any of the first aspects.

[0130] Compared with existing technologies, the beneficial effects achieved by this invention are as follows:

[0131] 1. The proposed method for multi-period transient voltage stability prevention and control comprehensively considers the cost of transient voltage stability prevention and control in the power system and the risk of power outage losses. It constructs a multi-period transient voltage stability prevention and control model based on reinforcement learning. Based on power grid simulation data and transient voltage safety and stability assessment results, it generates reinforcement learning model training samples to train the multi-period transient voltage stability prevention and control model. The trained reinforcement learning model is then used to generate online transient voltage stability prevention and control strategies for each period, allowing for real-time adjustments to the power grid's operating status. This method simultaneously considers the cost of prevention and control and the losses caused by power outages. By constructing a model through reinforcement learning, it can achieve more economical and effective voltage stability prevention and control in different time periods. It can address transient voltage stability problems under highly uncertain power grid scenarios and quickly generate auxiliary decisions for transient voltage stability prevention and control. This helps to quickly respond to changes in the power grid, improve the flexibility and reliability of power grid operation, and provide a reliable basis for reactive power optimization and stability control.

[0132] 2. The transient voltage stability multi-period prevention and control system proposed in this invention realizes transient voltage stability multi-period prevention and control by setting up a model building module, an offline training module, a simulation module, a strategy generation module, and an operating state adjustment module, and has good application prospects.

[0133] 3. The computer media, equipment and products provided by the present invention can execute the steps of the transient voltage stability multi-period prevention and control method provided by the present invention. Attached Figure Description

[0134] Figure 1 This is a flowchart of a multi-period prevention and control method for transient voltage stabilization provided in an embodiment of the present invention. Detailed Implementation

[0135] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments and specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations thereof. In the absence of conflict, the embodiments and technical features in the embodiments can be combined with each other.

[0136] It should be noted that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0137] Example 1

[0138] This invention discloses a multi-period prevention and control method for transient voltage stability, with reference to... Figure 1 As shown, the specific steps include the following:

[0139] S1. Based on the power grid topology, construct a multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network;

[0140] S2, based on the real-time operation status information of the power grid, the dispatch plan, the new energy forecast information and the load forecast information, the power system simulation software is used to obtain real-time operation mode data of the power grid in multiple time periods;

[0141] S3. Based on the real-time operation data of the power grid in multiple time periods, generate a multi-time period prevention and control strategy using the trained multi-time period transient voltage stability prevention and control model.

[0142] S4, adjust the power grid's operating status in real time according to the multi-period prevention and control strategy.

[0143] In step S1, DDPG is a deep reinforcement learning algorithm that combines a deep Q-network and a deterministic policy gradient, including an evaluation network and a policy network. By setting a reasonable state space, action space, and reward function, the reinforcement learning model learns in the direction of the target. In this embodiment, based on the power grid topology, the following state space, action space, and reward function are constructed: The expression of the state space is as follows:

[0144] ,

[0145] in, Representing the state space, Indicates the first The voltage amplitude at each node, Indicates the first The reactive power output of the generator Indicates the first The reactive power output of the Taiwanese capacitor Indicates the first The reactive power output of the Taiwan power reactor Indicates the first Reactive power of each load node Indicates the first The reactive power output of the SVG in Taiwan Indicates the first The reactive power output of the SVC; through the construction of the above state space, it is helpful to describe the voltage state of the power grid more comprehensively and improve the accuracy and effectiveness of multi-time transient voltage stability prevention and control strategies.

[0146] The expression for the action space is as follows:

[0147] ,

[0148] in, Represents the action space. Indicates the first The reactive power adjustment of the generator. Indicates the first The reactive power adjustment of the capacitor bank. Indicates the first Reactive power adjustment of the Taiwan reactor Indicates the first Reactive power adjustment of each load node Indicates the first The reactive power adjustment of the SVG. Indicates the first The reactive power adjustment of the SVC; by constructing the above action space, it is helpful to describe the transient voltage prevention and control strategy more specifically, and improve the accuracy and effectiveness of the multi-time transient voltage stability prevention and control strategy. The expression of the reward function is as follows:

[0149] ,

[0150] in, Represents the reward function, Indicates the decision-making window, This represents the prevention and control constraint function. This signifies punishment for an unchecked trend. The expression for the multi-period transient voltage stability risk function is as follows:

[0151] ,

[0152] in, Represents the set of scheduling periods. Represents a set of scenes. Display decision window Inside Scheduling period in the scenario Cost indicators of transient voltage stabilization control measures; Decision-making window Inside Scheduling period in the scenario The power outage loss index. The construction of the above reward function helps to more accurately describe the execution effect of the transient voltage prevention and control strategy, improving the accuracy and effectiveness of the multi-period transient voltage stability prevention and control strategy. In this embodiment, the prevention and control constraint function... This includes the following constraints: power flow constraints, voltage and phase angle magnitude constraints at each node, generator output constraints, reactive power adjustment constraints of capacitive reactors, active and reactive power constraints of load receiving, reactive power adjustment constraints of SVC and reactive power adjustment constraints of SVG:

[0153] The expression for power flow constraints is as follows:

[0154] ,

[0155] in, This represents the set of nodes excluding the balancing node. Represents a node The active power of power generation, Represents a node The merits, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents the cosine function. Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them Represents the sine function. Represents the set of PQ nodes. Represents a node reactive power, Represents a node reactive power demand, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them;

[0156] The expressions for the voltage and phase angle magnitude constraints at each node are as follows:

[0157] ,

[0158] in, Represents a node voltage, Represents a node Rated voltage, Represents a node phase angle, Represents a node The maximum phase angle, Represents a node The minimum phase angle;

[0159] The expression for the generator output constraint is as follows:

[0160] ,

[0161] in, Indicates the first The operating power of the generators, Indicates the first The generator operates at its minimum output power. Indicates the first The generator's maximum operating output power Indicates the first Reactive power of generators Indicates the first The generator operates at its minimum reactive power output. Indicates the first The generator's maximum reactive power output is [missing information]. Indicates the total number of generators;

[0162] The expression for the reactive power adjustment constraint of the capacitive reactor is as follows:

[0163] ,

[0164] in, This indicates the total number of capacitors. Indicates the total number of reactors. Indicates the first Taiwan capacitor outputs reactive power. Indicates the first The minimum output reactive power of the capacitor. Indicates the first The maximum reactive power output of the capacitor is... Indicates the first Taiwan reactor output reactive power Indicates the first Minimum output reactive power of the reactor Indicates the first The maximum reactive power output of the Taiwan reactor;

[0165] The expression for the active and reactive power constraints of the load is as follows:

[0166] ,

[0167] in, Indicates the first Active power received by each load node Indicates the first Minimum power received by each load node, Indicates the first The maximum operating power received by each load node, Indicates the first Reactive power of each load node Indicates the first The minimum reactive power required for the operation of each load node. Indicates the first The maximum reactive power of each load node during operation. Indicates the total number of load nodes;

[0168] The expression for the reactive power adjustment constraint of SVC is as follows:

[0169] ,

[0170] in, Indicates the total number of SVCs. For the first The SVC outputs reactive power. For the first Minimum output reactive power of the SVC unit For the first Maximum reactive power output of the SVC unit;

[0171] The expression for the reactive power adjustment constraint of SVG is as follows:

[0172] ,

[0173] in, Indicates the total number of SVGs. Indicates the first The SVG outputs reactive power. Indicates the first Minimum output reactive power of SVG Indicates the first The maximum reactive power output of the SVG.

[0174] The construction of the constraint function for multi-period transient voltage stability prevention and control helps guide the reinforcement learning model to learn more effective control strategies, reduces the computational resource consumption during the training process of the multi-period transient voltage stability prevention and control reinforcement learning model, and improves the accuracy and effectiveness of subsequent multi-period transient voltage stability prevention and control strategies.

[0175] In step S2, power system simulation software is used to perform simulations based on preset scheduling plans, new energy forecast information, load forecast information, and the real-time operating status information of the power grid, obtaining real-time operating mode data for multiple time periods of the power grid. This construction of real-time operating mode data for multiple time periods provides a data foundation for multi-time period transient voltage stability prevention and control strategies. In step S3, the real-time operating mode data for multiple time periods of the power grid is input into a trained multi-time period transient voltage stability prevention and control model. For each time period's real-time operating mode of the power grid, a transient voltage prevention and control sequence is sequentially generated. The new system state S' after the previous time period's prevention and control strategy is executed in the simulation environment, along with the remaining action space, serves as the input to the reinforcement learning model for the next time period, generating the transient voltage prevention and control strategy for the next time period. Finally, a multi-time period transient voltage stability prevention and control strategy sequence is generated. Through the generation process of the multi-time period transient voltage stability prevention and control strategy sequence, the sequential generation of the transient voltage prevention and control sequence ensures the adaptability of the transient voltage prevention and control strategy to the temporal changes in the power grid operating scenario.

[0176] The training objective of the evaluation network in the multi-time transient voltage stability prevention and control model is to ensure that the action value Q of any reactive power equipment control strategy satisfies the action value Bellman equation; the current reactive power equipment control strategy The action value Q-value equals the agent's reward value. Next state of the system Execute action The expected value of the sum of Q values; to enable the decision network to generate actions with the maximum Q value, the decision network guides its own network parameter training based on the Q value generated by the evaluation network; the offline training process of the multi-time transient voltage stability prevention and control model includes:

[0177] Initialize the parameters of the decision network and evaluation network in the multi-time transient voltage stability prevention and control model. Based on the reinforcement learning training sample set, repeat the following steps until the change in the reward function is within a preset range or the number of iterations reaches a preset update iteration number:

[0178] Randomly sampled from the reinforcement learning training sample set There are 10 samples, each including the current state. ,action Reward Value and the next state ;in, Indicates the sample sequence number;

[0179] Use a decision network to obtain the next state for each sample The following action ;in, Representation Decision Network Parameters;

[0180] Use the evaluation network to obtain the next state Next action Action value Q value ;in, Indicates evaluation network Parameters;

[0181] Using an evaluation network to obtain state Next action Action value Q value ;

[0182] According to the state Next action Action value Q value In state Next action Actual reward value and in the next state Next action Action value Q value Calculate the loss value of the evaluation network. ;

[0183] Based on the loss value of the evaluation network The parameters of the evaluation network are updated using the gradient descent method to obtain the updated evaluation network parameters. ;

[0184] Using decision networks to determine the state of each sample Generate new actions ;

[0185] Use the updated evaluation network to obtain the state Next action Action value Q value ;

[0186] According to the state Next action Action value Q value Calculate the loss of the decision network ;

[0187] According to the loss of the decision network The parameters of the decision network are updated using the gradient descent method to obtain the updated evaluation network parameters. .

[0188] Through the training process of the evaluation network and decision network described above, the reinforcement learning evaluation network and decision network can respond to changes in the power grid state in real time. The evaluation network can quickly evaluate the advantages and disadvantages of different decision schemes, helping the decision network to quickly select a better strategy, thereby improving decision efficiency and reducing computation time and resource consumption.

[0189] The reinforcement learning training sample set is constructed through the following steps:

[0190] Repeat the following steps until a predetermined number of reinforcement learning training samples are obtained, thus obtaining the reinforcement learning training sample set:

[0191] Active load at nodes The reactive load at the node is within the preset normal level range of [0, 1.2]. Active power of node-linked units within the preset normal level range of [0, 1.2] times. Random values ​​are taken within the range of [0,1] times the rated power of the generator unit and the reactive power regulation equipment CQ is within the upper and lower limits of the reactive power output range to form grid operation status information.

[0192] Using power system simulation software, simulations are performed based on preset scheduling plans, new energy forecast information, load forecast information, and the power grid operation status information to obtain power grid operation mode data for multiple time periods.

[0193] Power flow calculation and simulation verification are performed on the power grid's multi-period operation mode data. If the power flow verification converges, the power grid's multi-period operation mode data is taken as the current state of the power grid's operation mode. ;in, Indicates the sample sequence number;

[0194] Based on a pre-set set of anticipated faults, a voltage safety and stability prevention and control strategy is implemented to determine the current state of the power grid operation mode. Select Action By executing actions in power system simulation software To obtain the next state and set the current state. Execute action Reward value =1;

[0195] Based on the current state ,action Reward Value and the next state Constructing training samples .

[0196] The process of constructing the training sample set described above can effectively improve the training effect and decision quality of the power grid multi-time period prevention and control optimization decision model.

[0197] In summary, the multi-period transient voltage stability prevention and control method proposed in this embodiment comprehensively considers the costs of transient voltage stability prevention and control in the power system and the risk of power outage losses. It constructs a multi-period transient voltage stability prevention and control model based on reinforcement learning. Based on power grid simulation data and voltage security and stability assessment results, it generates training samples for the reinforcement learning model and then trains the multi-period transient voltage stability prevention and control model. The trained reinforcement learning model is used to generate periodic transient voltage stability prevention and control strategies online, allowing for real-time adjustments to the power grid's operating state. This method can address transient voltage stability problems under highly uncertain power grid scenarios, quickly generate auxiliary decisions for transient voltage stability prevention and control, and provide a reliable basis for reactive power optimization and stability control.

[0198] Example 2:

[0199] Based on the same inventive concept as Embodiment 1, this embodiment of the invention discloses a transient voltage stabilization multi-time period prevention and control system, comprising:

[0200] The model building module is used to: construct a multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network according to the power grid topology;

[0201] The simulation module is used to: obtain real-time data on the power grid's operation mode in multiple time periods using power system simulation software, based on real-time power grid operation status information, scheduling plans, new energy forecast information, and load forecast information.

[0202] The strategy generation module is used to generate multi-period prevention and control strategies based on the real-time operation mode data of the power grid in multiple time periods and using the trained multi-period transient voltage stability prevention and control model.

[0203] The operation status adjustment module is used to adjust the operation status of the power grid in real time according to the multi-time period prevention and control strategy.

[0204] DDPG is a deep reinforcement learning algorithm that combines deep Q-networks and deterministic policy gradients. It includes an evaluation network and a policy network, and by setting appropriate state space, action space, and reward function, the reinforcement learning model learns in the direction of the target. In this embodiment, based on the power grid topology, the following state space, action space, and reward function are constructed:

[0205] The state space expression is as follows:

[0206] ,

[0207] in, Representing the state space, Indicates the first The voltage amplitude at each node, Indicates the first The reactive power output of the generator Indicates the first The reactive power output of the Taiwanese capacitor Indicates the first The reactive power output of the Taiwan power reactor Indicates the first Reactive power of each load node Indicates the first The reactive power output of the SVG in Taiwan Indicates the first The reactive power output of the Taiwan SVC;

[0208] The expression for the action space is as follows:

[0209] ,

[0210] in, Represents the action space. Indicates the first The reactive power adjustment of the generator. Indicates the first The reactive power adjustment of the capacitor bank. Indicates the first Reactive power adjustment of the Taiwan reactor Indicates the first Reactive power adjustment of each load node Indicates the first The reactive power adjustment of the SVG. Indicates the first The reactive power adjustment of the SVC;

[0211] The expression for the reward function is as follows:

[0212] ,

[0213] in, Represents the reward function, Indicates the decision-making window, This represents the prevention and control constraint function. This signifies punishment for an unchecked trend. The expression for the multi-period transient voltage stability risk function is as follows:

[0214] ,

[0215] in, Represents the set of scheduling periods. Represents a set of scenes. Display decision window Inside Scheduling period in the scenario Cost indicators of transient voltage stabilization control measures; Decision-making window Inside Scheduling period in the scenario Power outage loss indicators.

[0216] In this embodiment, the prevention control constraint function This includes the following constraints: power flow constraints, voltage and phase angle magnitude constraints at each node, generator output constraints, reactive power adjustment constraints of capacitive reactors, active and reactive power constraints of load receiving, reactive power adjustment constraints of SVC and reactive power adjustment constraints of SVG:

[0217] The expression for power flow constraints is as follows:

[0218] ,

[0219] in, This represents the set of nodes excluding the balancing node. Represents a node The active power of power generation, Represents a node The merits, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents the cosine function. Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them Represents the sine function. Represents the set of PQ nodes. Represents a node reactive power, Represents a node reactive power demand, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them;

[0220] The expressions for the voltage and phase angle magnitude constraints at each node are as follows:

[0221] ,

[0222] in, Represents a node voltage, Represents a node Rated voltage, Represents a node phase angle, Represents a node The maximum phase angle, Represents a node The minimum phase angle;

[0223] The expression for the generator output constraint is as follows:

[0224] ,

[0225] in, Indicates the first The operating power of the generators, Indicates the first The generator operates at its minimum output power. Indicates the first The generator's maximum operating output power Indicates the first Reactive power of generators Indicates the first The generator operates at its minimum reactive power output. Indicates the first The generator's maximum reactive power output is [missing information]. Indicates the total number of generators;

[0226] The expression for the reactive power adjustment constraint of the capacitive reactor is as follows:

[0227] ,

[0228] in, This indicates the total number of capacitors. Indicates the total number of reactors. Indicates the first Taiwan capacitor outputs reactive power. Indicates the first The minimum output reactive power of the capacitor. Indicates the first The maximum reactive power output of the capacitor is... Indicates the first Taiwan reactor output reactive power Indicates the first Minimum output reactive power of the reactor Indicates the first The maximum reactive power output of the Taiwan reactor;

[0229] The expression for the active and reactive power constraints of the load is as follows:

[0230] ,

[0231] in, Indicates the first Active power received by each load node Indicates the first Minimum power received by each load node, Indicates the first The maximum operating power received by each load node, Indicates the first Reactive power of each load node Indicates the first The minimum reactive power required for the operation of each load node. Indicates the first The maximum reactive power of each load node during operation. Indicates the total number of load nodes;

[0232] The expression for the reactive power adjustment constraint of SVC is as follows:

[0233] ,

[0234] in, Indicates the total number of SVCs. For the first The SVC outputs reactive power. For the first Minimum output reactive power of the SVC unit For the first Maximum reactive power output of the SVC unit;

[0235] The expression for the reactive power adjustment constraint of SVG is as follows:

[0236] ,

[0237] in, Indicates the total number of SVGs. Indicates the first The SVG outputs reactive power. Indicates the first Minimum output reactive power of SVG Indicates the first The maximum reactive power output of the SVG.

[0238] The simulation module uses power system simulation software to perform simulations based on preset scheduling plans, new energy forecast information, load forecast information, and the real-time operation status information of the power grid, thereby obtaining real-time operation mode data of the power grid in multiple time periods.

[0239] The strategy generation module inputs the real-time operation data of the power grid in multiple time periods into the trained multi-time period transient voltage stability prevention and control model. For the real-time operation mode of the power grid in each time period, it sequentially generates a transient voltage prevention and control sequence. The new system state S' after the prevention and control strategy of the previous time period is executed in the simulation environment and the remaining action space are used as the input of the reinforcement learning model of the next time period to generate the transient voltage prevention and control strategy for the next time period. Finally, a multi-time period transient voltage stability prevention and control strategy sequence is generated.

[0240] The training objective of the evaluation network in the multi-time transient voltage stability prevention and control model is to ensure that the action value Q of any reactive power equipment control strategy satisfies the action value Bellman equation; the current reactive power equipment control strategy The action value Q-value equals the agent's reward value. Next state of the system Execute action The expected value of the sum of Q values; to enable the decision network to generate actions with the maximum Q value, the decision network guides its own network parameter training based on the Q value generated by the evaluation network; the offline training process of the multi-time transient voltage stability prevention and control model includes:

[0241] Initialize the parameters of the decision network and evaluation network in the multi-time transient voltage stability prevention and control model. Based on the reinforcement learning training sample set, repeat the following steps until the change in the reward function is within a preset range or the number of iterations reaches a preset update iteration number:

[0242] Randomly sampled from the reinforcement learning training sample set There are 10 samples, each including the current state. ,action Reward Value and the next state ;in, Indicates the sample sequence number;

[0243] Use a decision network to obtain the next state for each sample The following action ;in, Representation Decision Network Parameters;

[0244] Use the evaluation network to obtain the next state Next action Action value Q value ;in, Indicates evaluation network Parameters;

[0245] Using an evaluation network to obtain state Next action Action value Q value ;

[0246] According to the state Next action Action value Q value In state Next action Actual reward value and in the next state Next action Action value Q value Calculate the loss value of the evaluation network. ;

[0247] Based on the loss value of the evaluation network The parameters of the evaluation network are updated using the gradient descent method to obtain the updated evaluation network parameters. ;

[0248] Using decision networks to determine the state of each sample Generate new actions ;

[0249] Use the updated evaluation network to obtain the state Next action Action value Q value ;

[0250] According to the state Next action Action value Q value Calculate the loss of the decision network ;

[0251] According to the loss of the decision network The parameters of the decision network are updated using the gradient descent method to obtain the updated evaluation network parameters. .

[0252] The reinforcement learning training sample set is constructed through the following steps:

[0253] Repeat the following steps until a predetermined number of reinforcement learning training samples are obtained, thus obtaining the reinforcement learning training sample set:

[0254] Active load at nodes The reactive load at the node is within the preset normal level range of [0, 1.2]. Active power of node-linked units within the preset normal level range of [0, 1.2] times. Random values ​​are taken within the range of [0,1] times the rated power of the generator unit and the reactive power regulation equipment CQ is within the upper and lower limits of the reactive power output range to form grid operation status information.

[0255] Using power system simulation software, simulations are performed based on preset scheduling plans, new energy forecast information, load forecast information, and the power grid operation status information to obtain power grid operation mode data for multiple time periods.

[0256] Power flow calculation and simulation verification are performed on the power grid's multi-period operation mode data. If the power flow verification converges, the power grid's multi-period operation mode data is taken as the current state of the power grid's operation mode. ;in, Indicates the sample sequence number;

[0257] Based on a pre-set set of anticipated faults, a voltage safety and stability prevention and control strategy is implemented to determine the current state of the power grid operation mode. Select Action By executing actions in power system simulation software To obtain the next state and set the current state. Execute action Reward value =1;

[0258] Based on the current state ,action Reward Value and the next state Constructing training samples .

[0259] The specific functions of each module described above are explained in the relevant content of the method in Embodiment 1, and will not be repeated here.

[0260] Example 3:

[0261] This embodiment provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the transient voltage stability multi-period prevention and control method as described in any of the embodiments.

[0262] Example 4:

[0263] This embodiment provides a computer device, including:

[0264] Memory, used to store computer instructions;

[0265] A processor is configured to execute the computer instructions to implement the steps of the transient voltage stability multi-period prevention and control method as described in any one of Embodiment 1.

[0266] Example 5:

[0267] This embodiment provides a computer program product, including computer instructions, characterized in that, when the computer instructions are executed by a processor, they implement the steps of the transient voltage stability multi-period prevention and control method as described in any one of Embodiment 1.

[0268] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0269] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0270] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0271] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0272] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method for multi-period prevention and control of transient voltage stability, characterized in that, include: Based on the power grid topology, a multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network is constructed. Based on real-time power grid operation status information, dispatching plans, new energy forecast information and load forecast information, power system simulation software is used to obtain real-time power grid operation mode data for multiple time periods. Based on the real-time operation data of the power grid in multiple time periods, a multi-time period prevention and control strategy is generated using the trained multi-time period transient voltage stability prevention and control model. The operating status of the power grid is adjusted in real time according to the multi-period prevention and control strategy.

2. The transient voltage stability multi-period prevention and control method according to claim 1, characterized in that, The multi-period transient voltage stability prevention and control model based on DDPG reinforcement learning network includes a state space, an action space, and a reward function. The expression for the state space is as follows: , in, Representing the state space, Indicates the first The voltage amplitude at each node, Indicates the first The reactive power output of the generator Indicates the first The reactive power output of the Taiwanese capacitor Indicates the first The reactive power output of the Taiwan power reactor Indicates the first Reactive power of each load node Indicates the first The reactive power output of the SVG in Taiwan Indicates the first The reactive power output of the Taiwan SVC; The expression for the action space is as follows: , in, Represents the action space. Indicates the first The reactive power adjustment of the generator. Indicates the first The reactive power adjustment of the capacitor bank. Indicates the first Reactive power adjustment of the Taiwan reactor Indicates the first Reactive power adjustment of each load node Indicates the first The reactive power adjustment of the SVG. Indicates the first The reactive power adjustment of the SVC; The expression for the reward function is as follows: , in, Represents the reward function, Indicates the decision-making window, This represents a multi-period transient voltage stability risk function. This represents the prevention and control constraint function. This signifies the punishment for an unchecked trend.

3. The transient voltage stability multi-period prevention and control method according to claim 2, characterized in that, The multi-period transient voltage stability risk function The expression is as follows: , in, Represents the set of scheduling periods. Represents a set of scenes. Display decision window Inside Scheduling period in the scenario Cost indicators of transient voltage stabilization control measures; Decision-making window Inside Scheduling period in the scenario Power outage loss indicators.

4. The transient voltage stability multi-period prevention and control method according to claim 2, characterized in that, The prevention and control constraint function It includes at least one of the following: power flow constraints, voltage and phase angle magnitude constraints at each node, generator output constraints, reactive power adjustment constraints of capacitive reactors, active and reactive power constraints of load receiving, reactive power adjustment constraints of SVC and reactive power adjustment constraints of SVG.

5. The transient voltage stability multi-period prevention and control method according to claim 4, characterized in that, The expression for the power flow constraint is as follows: , in, This represents the set of nodes excluding the balancing node. Represents a node The active power of power generation, Represents a node The merits, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents the cosine function. Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them Represents the sine function. Represents the set of PQ nodes. Represents a node reactive power, Represents a node reactive power demand, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them; The expressions for the voltage and phase angle magnitude constraints at each node are as follows: , in, Represents a node voltage, Represents a node Rated voltage, Represents a node phase angle, Represents a node The maximum phase angle, Represents a node The minimum phase angle; The expression for the generator output constraint is as follows: , in, Indicates the first The operating power of the generators, Indicates the first The generator operates at its minimum output power. Indicates the first The generator's maximum operating output power Indicates the first Reactive power of generators Indicates the first The generator operates at its minimum reactive power output. Indicates the first The generator's maximum reactive power output is [missing information]. Indicates the total number of generators; The expression for the reactive power adjustment constraint of the capacitive reactor is as follows: , in, This indicates the total number of capacitors. Indicates the total number of reactors. Indicates the first Taiwan capacitor outputs reactive power. Indicates the first The minimum output reactive power of the capacitor. Indicates the first The maximum reactive power output of the capacitor is... Indicates the first Taiwan reactor output reactive power Indicates the first Minimum output reactive power of the reactor Indicates the first The maximum reactive power output of the Taiwan reactor; The expression for the active and reactive power constraints of the load is as follows: , in, Indicates the first Active power received by each load node Indicates the first Minimum power received by each load node, Indicates the first The maximum operating power received by each load node, Indicates the first Reactive power of each load node Indicates the first The minimum reactive power required for the operation of each load node. Indicates the first The maximum reactive power of each load node during operation. Indicates the total number of load nodes; The expression for the SVC reactive power adjustment constraint is as follows: , in, Indicates the total number of SVCs. For the first The SVC outputs reactive power. For the first Minimum output reactive power of the SVC unit For the first Maximum reactive power output of the SVC unit; The expression for the SVG reactive power adjustment constraint is as follows: , in, Indicates the total number of SVGs. Indicates the first The SVG outputs reactive power. Indicates the first Minimum output reactive power of SVG Indicates the first The maximum reactive power output of the SVG.

6. The transient voltage stability multi-period prevention and control method according to claim 1, characterized in that, Based on the constructed reinforcement learning training sample set, the multi-time transient voltage stability prevention and control model is trained offline to obtain the trained multi-time transient voltage stability prevention and control model, including: Initialize the parameters of the decision network and evaluation network in the multi-time transient voltage stability prevention and control model. Based on the reinforcement learning training sample set, repeat the following steps until the preset termination condition is met: Randomly sampled from the reinforcement learning training sample set There are 10 samples, each including the current state. ,action Reward Value and the next state ;in, Indicates the sample sequence number; Use a decision network to obtain the next state for each sample The following action ;in, Representation Decision Network Parameters; Use the evaluation network to obtain the next state Next action Action value Q value ;in, Indicates evaluation network Parameters; Using an evaluation network to obtain state Next action Action value Q value ; According to the state Next action Action value Q value In state Next action Actual reward value and in the next state Next action Action value Q value Calculate the loss value of the evaluation network. ; Based on the loss value of the evaluation network The parameters of the evaluation network are updated using the gradient descent method to obtain the updated evaluation network parameters. ; Using decision networks to determine the state of each sample Generate new actions ; Use the updated evaluation network to obtain the state Next action Action value Q value ; According to the state Next action Action value Q value Calculate the loss of the decision network ; According to the loss of the decision network The parameters of the decision network are updated using the gradient descent method to obtain the updated evaluation network parameters. .

7. The transient voltage stability multi-period prevention and control method according to claim 6, characterized in that, The reinforcement learning training sample set is constructed through the following steps: Repeat the following steps until a predetermined number of reinforcement learning training samples are obtained, thus obtaining the reinforcement learning training sample set: Active load at nodes The reactive load at the node is within the preset normal level range of [0, k1]. Active power of node-linked units within the preset normal level range of [0, k2] times. The grid operation status information is formed by randomly selecting values ​​under the constraints of the rated power range of the generator unit in the range of [0,1] times the rated power range and the reactive power regulation equipment CQ being within the upper and lower limits of the reactive power output range; where k1 and k2 are preset parameters. Using power system simulation software, simulations are performed based on preset scheduling plans, new energy forecast information, load forecast information, and the power grid operation status information to obtain power grid operation mode data for multiple time periods. Power flow calculation and simulation verification are performed on the power grid's multi-period operation mode data. If the power flow verification converges, the power grid's multi-period operation mode data is taken as the current state of the power grid's operation mode. ;in, Indicates the sample sequence number; Based on a pre-set set of anticipated faults, a voltage safety and stability prevention and control strategy is implemented to determine the current state of the power grid operation mode. Select Action By executing actions in power system simulation software To obtain the next state and set the current state. Execute action Reward value =1; Based on the current state ,action Reward Value and the next state Constructing training samples .

8. A transient voltage stabilization multi-time period prevention and control system, characterized in that, include: The model building module is used to: construct a multi-period transient voltage stability prevention and control model based on the DDPG reinforcement learning network according to the power grid topology; The simulation module is used to: obtain real-time data on the power grid's operation mode in multiple time periods using power system simulation software, based on real-time power grid operation status information, scheduling plans, new energy forecast information, and load forecast information. The strategy generation module is used to generate multi-period prevention and control strategies based on the real-time operation mode data of the power grid in multiple time periods and using the trained multi-period transient voltage stability prevention and control model. The operation status adjustment module is used to adjust the operation status of the power grid in real time according to the multi-time period prevention and control strategy.

9. The transient voltage stabilization multi-period prevention and control system according to claim 8, characterized in that, The multi-period transient voltage stability prevention and control model based on DDPG reinforcement learning network includes a state space, an action space, and a reward function. The expression for the state space is as follows: , in, Representing the state space, Indicates the first The voltage amplitude at each node, Indicates the first The reactive power output of the generator Indicates the first The reactive power output of the Taiwanese capacitor Indicates the first The reactive power output of the Taiwan power reactor Indicates the first Reactive power of each load node Indicates the first The reactive power output of the SVG in Taiwan Indicates the first The reactive power output of the Taiwan SVC; The expression for the action space is as follows: , in, Represents the action space. Indicates the first The reactive power adjustment of the generator. Indicates the first The reactive power adjustment of the capacitor bank. Indicates the first Reactive power adjustment of the Taiwan reactor Indicates the first Reactive power adjustment of each load node Indicates the first The reactive power adjustment of the SVG. Indicates the first The reactive power adjustment of the SVC; The expression for the reward function is as follows: , in, Represents the reward function, Indicates the decision-making window, This represents a multi-period transient voltage stability risk function. This represents the prevention and control constraint function. This signifies the punishment for an unchecked trend.

10. The transient voltage stabilization multi-period prevention and control system according to claim 9, characterized in that, The multi-period transient voltage stability risk function The expression is as follows: , in, Represents the set of scheduling periods. Represents a set of scenes. Display decision window Inside Scheduling period in the scenario Cost indicators of transient voltage stabilization control measures; Decision-making window Inside Scheduling period in the scenario Power outage loss indicators.

11. The transient voltage stabilization multi-period prevention and control system according to claim 9, characterized in that, The prevention and control constraint function It includes at least one of the following: power flow constraints, voltage and phase angle magnitude constraints at each node, generator output constraints, reactive power adjustment constraints of capacitive reactors, active and reactive power constraints of load receiving, reactive power adjustment constraints of SVC and reactive power adjustment constraints of SVG.

12. The transient voltage stabilization multi-period prevention and control system according to claim 11, characterized in that, The expression for the power flow constraint is as follows: , in, This represents the set of nodes excluding the balancing node. Represents a node The active power of power generation, Represents a node The merits, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents the cosine function. Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them Represents the sine function. Represents the set of PQ nodes. Represents a node reactive power, Represents a node reactive power demand, Represents a node voltage amplitude, express The total number of all nodes in the network. Represents a node voltage amplitude, Represents a node To the node The electrical conductance between them Represents a node To the node The voltage phase angle difference between them Represents a node To the node The susceptance between them; The expressions for the voltage and phase angle magnitude constraints at each node are as follows: , in, Represents a node voltage, Represents a node Rated voltage, Represents a node phase angle, Represents a node The maximum phase angle, Represents a node The minimum phase angle; The expression for the generator output constraint is as follows: , in, Indicates the first The operating power of the generators, Indicates the first The generator operates at its minimum output power. Indicates the first The generator's maximum operating output power Indicates the first Reactive power of generators Indicates the first The generator operates at its minimum reactive power output. Indicates the first The generator's maximum reactive power output is [missing information]. Indicates the total number of generators; The expression for the reactive power adjustment constraint of the capacitive reactor is as follows: , in, This indicates the total number of capacitors. Indicates the total number of reactors. Indicates the first Taiwan capacitor outputs reactive power. Indicates the first The minimum output reactive power of the capacitor. Indicates the first The maximum reactive power output of the capacitor is... Indicates the first Taiwan reactor output reactive power Indicates the first Minimum output reactive power of the reactor Indicates the first The maximum reactive power output of the Taiwan reactor; The expression for the active and reactive power constraints of the load is as follows: , in, Indicates the first Active power received by each load node Indicates the first Minimum power received by each load node, Indicates the first The maximum operating power received by each load node, Indicates the first Reactive power of each load node Indicates the first The minimum reactive power required for the operation of each load node. Indicates the first The maximum reactive power of each load node during operation. Indicates the total number of load nodes; The expression for the SVC reactive power adjustment constraint is as follows: , in, Indicates the total number of SVCs. For the first The SVC outputs reactive power. For the first Minimum output reactive power of the SVC unit For the first Maximum reactive power output of the SVC unit; The expression for the SVG reactive power adjustment constraint is as follows: , in, Indicates the total number of SVGs. Indicates the first The SVG outputs reactive power. Indicates the first Minimum output reactive power of SVG Indicates the first The maximum reactive power output of the SVG.

13. The transient voltage stabilization multi-period prevention and control system according to claim 8, characterized in that, Based on the constructed reinforcement learning training sample set, the multi-time transient voltage stability prevention and control model is trained offline to obtain the trained multi-time transient voltage stability prevention and control model, including: Initialize the parameters of the decision network and evaluation network in the multi-time transient voltage stability prevention and control model. Based on the reinforcement learning training sample set, repeat the following steps until the preset termination condition is met: Randomly sampled from the reinforcement learning training sample set There are 10 samples, each including the current state. ,action Reward Value and the next state ;in, Indicates the sample sequence number; Use a decision network to obtain the next state for each sample The following action ;in, Representation Decision Network Parameters; Use the evaluation network to obtain the next state Next action Action value Q value ;in, Indicates evaluation network Parameters; Using an evaluation network to obtain state Next action Action value Q value ; According to the state Next action Action value Q value In state Next action Actual reward value and in the next state Next action Action value Q value Calculate the loss value of the evaluation network. ; Based on the loss value of the evaluation network The parameters of the evaluation network are updated using the gradient descent method to obtain the updated evaluation network parameters. ; Using decision networks to determine the state of each sample Generate new actions ; Use the updated evaluation network to obtain the state Next action Action value Q value ; According to the state Next action Action value Q value Calculate the loss of the decision network ; According to the loss of the decision network The parameters of the decision network are updated using the gradient descent method to obtain the updated evaluation network parameters. .

14. The transient voltage stabilization multi-period prevention and control system according to claim 13, characterized in that, The reinforcement learning training sample set is constructed through the following steps: Repeat the following steps until a predetermined number of reinforcement learning training samples are obtained, thus obtaining the reinforcement learning training sample set: Active load at nodes The reactive load at the node is within the preset normal level range of [0, k1]. Active power of node-linked units within the preset normal level range of [0, k2] times. The grid operation status information is formed by randomly selecting values ​​under the constraints of the rated power range of the generator unit in the range of [0,1] times the rated power range and the reactive power regulation equipment CQ being within the upper and lower limits of the reactive power output range; where k1 and k2 are preset parameters. Using power system simulation software, simulations are performed based on preset scheduling plans, new energy forecast information, load forecast information, and the power grid operation status information to obtain power grid operation mode data for multiple time periods. Power flow calculation and simulation verification are performed on the power grid's multi-period operation mode data. If the power flow verification converges, the power grid's multi-period operation mode data is taken as the current state of the power grid's operation mode. ;in, Indicates the sample sequence number; Based on a pre-set set of anticipated faults, a voltage safety and stability prevention and control strategy is implemented to determine the current state of the power grid operation mode. Select Action By executing actions in power system simulation software To obtain the next state and set the current state. Execute action Reward value =1; Based on the current state ,action Reward Value and the next state Constructing training samples .

15. A computer-readable storage medium storing computer instructions thereon, characterized in that, When the computer instruction is executed by the processor, it implements the steps of the transient voltage stabilization multi-period prevention and control method according to any one of claims 1-7.

16. A computer device, characterized in that, Memory, used to store computer instructions; A processor for executing the computer instructions to implement the steps of the transient voltage stability multi-period prevention and control method according to any one of claims 1-7.

17. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the transient voltage stabilization multi-period prevention and control method according to any one of claims 1-7.