Power grid voltage multi-time scale stability control method considering network topology reconstruction
Patent Information
- Application Number
- CN202510734360.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-09
Smart Images

Figure CN120613741A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of grid voltage stability control, and in particular to a grid voltage multi-time scale stability control method taking into account network topology reconstruction. Background Art
[0002] The increasing integration of distributed photovoltaics (PV) has led to significant fluctuations in low-voltage grid voltages due to their inherent volatility and intermittency, significantly reducing the grid's static voltage stability. The continued increase in distributed PV penetration has led to frequent voltage overshoots in distribution networks, seriously impacting the grid's power quality and safety. Furthermore, the rapid reactive power regulation capabilities of current PV inverters give them the potential to contribute to grid reactive power and voltage control.
[0003] Currently, methods for reactive power and voltage control in low-voltage power grids fall into two main categories: model-based optimization methods and data-driven methods. Model-based methods treat reactive power and voltage control as a nonlinear optimization problem, explicitly describing the physical constraints and operating state of the power grid, and then using optimization algorithms to obtain the optimal operating solution.
[0004] However, these methods often rely on precise network models and global information. Due to limitations such as modeling difficulties and long solution times, they struggle to meet the demands of real-time optimization. Given the limited voltage and power data available to control centers, model-based reactive power and voltage control methods are difficult to apply in practice. Multi-agent deep reinforcement learning, with its advantages of model-free operation and rapid solution times, is more suitable for the distributed optimization of distribution networks. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects existing in the above-mentioned background technology. The present invention provides a multi-time-scale stable control method of grid voltage taking into account network topology reconstruction, which is used to solve the problem of difficult voltage control caused by traditional voltage control methods when dealing with distributed photovoltaic access to the grid.
[0006] The present invention adopts the following technical solutions to solve the above technical problems:
[0007] A multi-time-scale stabilization control method for power grid voltage taking into account network topology reconstruction includes the following steps:
[0008] Step S1: Build a graph neural network feature extraction model for the power grid, input the power-voltage electrical parameter matrix of the power grid nodes, and output a feature vector representing the topology and electrical characteristics;
[0009] Step S2: Based on the feature vectors in step S1, a K-means clustering algorithm is used to generate a typical topology set Ω = {G1, ..., G K}, where K is determined by the clustering convergence threshold, and Ω provides candidate topologies for network topology reconstruction;
[0010] Step S3: Establish a multi-agent control system and implement hierarchical training, including:
[0011] Step S3-1: Establish a network topology reconstruction agent A1, and select the optimal topology G based on the voltage deviation and over-limit status within the network topology reconstruction period τ1. i ∈Ω, and update the A1 parameters through reconstruction feedback;
[0012] Step S3-2: Establish a long-time-scale agent A2, calculate the optimal switching plan for the long-time-scale equipment based on the voltage deviation and over-limit status within the long-time-scale control period τ2, and update the A2 parameters through feedback;
[0013] Step S3-3: Establish a short-time-scale agent A3, adjust the short-time-scale device based on the voltage deviation and over-limit status within the short-time-scale control period τ3, and update the A3 parameters through feedback;
[0014] Step S4: Deploy the trained A1, A2, and A3 to the power grid control center to obtain the optimized network topology reconstruction decision, long-time scale device action decision, and short-time scale device action decision.
[0015] Furthermore, the specific steps of step S1 are as follows:
[0016] Step S1-1: Express the grid reactive power voltage sensitivity matrix in a graphical data format to form a grid node power-voltage sensitivity matrix, specifically:
[0017]
[0018] Where S is the reactive voltage sensitivity matrix of the power grid; Q is the reactive power matrix of the power grid; U is the node voltage matrix of the power grid; △ represents the partial derivative. A directed graph G = (X, A) is constructed to store the matrix S. The elements on the main diagonal of the sensitivity matrix are arranged in chronological order to construct the node information vector X, and the remaining elements are used to construct the edge information vector A.
[0019] Step S1-2: Use GIN to reduce the dimension of the above graph data. Input the graph data G into the GIN network for dimensionality reduction and output a feature vector H(G) representing the topological characteristics. The GIN is a graph isomorphic network.
[0020] Furthermore, the clustering convergence threshold in step S2 is the silhouette coefficient ε=0.46;
[0021] Each element G in the typical topology set Ω in step S2 i The power grid topology corresponding to the cluster center represents the network structure under a specific operation scenario.
[0022] Furthermore, in step S3, the network topology reconstruction period, the long time scale control period, and the short time scale control period are 24 hours, 6 hours, and 15 minutes, respectively.
[0023] Furthermore, the hierarchical collaborative training in step S3 is specifically that the agent receives the environment state, performs control actions, and iteratively updates its control parameters according to the reward function.
[0024] Furthermore, the agent receives the environment state as:
[0025]
[0026] Where s1 is the environment state vector received by the network topology reconstruction agent, s2 is the environment state vector received by the long-time scale device agent, and s3 is the environment state vector received by the short-time scale agent. Represents the gear position of the capacitor bank at time t, K t represents the grid topology at time t, is the cumulative value of voltage exceeding the limit of all nodes in the power grid at time t, is the cumulative value of the voltage offset of all nodes in the power grid at time t, max() means taking the larger value of the two, U i,t is the voltage amplitude of node i at time t, Respectively represent the active power output and reactive power output of photovoltaic at time t, They represent the active power and reactive power consumed by the load at time t respectively.
[0027] Furthermore, the control action performed by the agent is:
[0028]
[0029] Where a1 is the action space of the network topology reconstruction agent, a2 is the action space of the long-time scale agent, a3 is the action space of the short-time scale agent, and a K is the network topology reconstruction strategy, a CB is the gear control instruction of all capacitors in the power grid, a PV It is the reactive power instruction of all photovoltaics in the grid.
[0030] Furthermore, the agent reward function is:
[0031]
[0032] In the formula is the reward of the network topology reconstruction agent and the long-time scale agent at time t. When k = 95, it represents the reward of the network topology reconstruction agent. When k = 23, it represents the reward of the long-time scale agent. is the short-time-scale agent reward at time t, is the value of the power grid network loss at time t, I ij,t is the current amplitude of the line between nodes i and j at time t, R ij Represents the resistance of the line between nodes i and j; is the cumulative value of voltage exceeding the limit of all nodes in the power grid at time t, is the cumulative value of the voltage offset of all nodes in the power grid at time t, and α is the discount factor, which is taken as 0.5.
[0033] Furthermore, the network topology reconstruction agent and the long time scale agent in step S3 are each composed of a DDQN network, and the parameters are updated by performing gradient descent on the following loss function. The DDQN is a dual deep Q network:
[0034]
[0035] Where L represents the loss function of the DDQN network, θ t represents network parameters, E represents expected calculation, r DDQN represents the reward of the network topology reconstruction agent or the long-time scale agent, γ DDQN represents the conversion factor, Q DDQN represents the value evaluation value of the network topology reconstruction agent or the long-time scale agent, a represents the action decision of the network topology reconstruction agent or the long-time scale agent, s represents the environmental state received by the network topology reconstruction agent or the long-time scale agent, s' represents the state received by the agent after taking action decision a under the environmental state s, a' represents the agent action under the environmental state s', α represents the learning rate, Represents the gradient.
[0036] Furthermore, in step S3, the short-time-scale agent is composed of a value evaluation network and an action network. When the short-time-scale agent updates its parameters, it is necessary to update the parameters of the two networks simultaneously. The value evaluation network updates its parameters by performing gradient descent on the following loss function:
[0037]
[0038] The action network updates its parameters by gradient descent on the following loss function:
[0039]
[0040] Among them, L' represents the loss function of the value evaluation network, θ Q1 Represents the parameters of the value assessment network, r DDPG represents the reward of the short-time-scale device agent, γ DDPGrepresents the conversion coefficient of the short-time-scale device agent, Q Q1 represents the value evaluation value of the value evaluation network, s1 represents the environmental state received by the short-time scale agent, a1 represents the action of the short-time scale agent, s1' represents the environmental state received by the short-time scale agent after taking the action decision a1 under the environmental state s1, a1' represents the action decision under the environmental state s1', α Q1 represents the learning rate of the value evaluation network, J represents the loss function of the action network, θ π represents the parameters of the action network, α π represents the learning rate of the action network, represents the gradient, E represents the expected calculation, Indicates the gradient of action a1, represents the value of the action network parameter at time t, Represents the value of the value evaluation network parameters at time t.
[0041] Compared with the prior art, the present invention adopts the above technical solution and has the following beneficial effects:
[0042] (1) The multi-time-scale grid voltage stability control method proposed in this paper takes into account the network topology reconstruction. By combining graph neural network feature extraction with a multi-agent hierarchical reinforcement learning framework, it realizes multi-time-scale collaborative optimization control of the grid static voltage stability, providing a robust and economical voltage control solution for grids with high photovoltaic penetration.
[0043] (2) The multi-time-scale stable control method for grid voltage taking into account network topology reconstruction proposed in the present invention realizes multi-time-scale control of grid voltage by constructing a network topology reconstruction agent A1, a long-time-scale agent A2 and a short-time-scale agent A3, and processes network topology reconstruction, long-time-scale equipment switching and short-time-scale equipment adjustment in different control cycles respectively, which can handle different problems in grid operation more comprehensively and meticulously, and improve the stability and control accuracy of grid operation.
[0044] (3) The multi-time-scale stable control method of grid voltage taking into account network topology reconstruction proposed in the present invention uses a graph neural network to construct a feature extraction model, converts the grid reactive voltage sensitivity matrix into a graph data form, and obtains the characteristic vector representing the topological characteristics through GIN network dimensionality reduction. It can effectively extract relevant information of the grid topology and electrical characteristics, provide more accurate input for subsequent clustering analysis and topology reconstruction, and improve the rationality and effectiveness of grid topology reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of the structure of the power grid system improved based on the IEEE33 node calculation example in the embodiment;
[0046] Figure 2 The voltage distribution before optimization in the embodiment;
[0047] Figure 3 1 is the voltage distribution after optimization using the proposed method in the embodiment. DETAILED DESCRIPTION
[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0049] The multi-time-scale voltage stability control method for a power grid taking into account network topology reconstruction includes the following steps:
[0050] Step S1: Build a graph neural network feature extraction model for the power grid, input the power-voltage electrical parameter matrix of the power grid nodes, and output a feature vector representing the topology and electrical characteristics;
[0051] Step S1-1: Express the grid reactive power voltage sensitivity matrix in the form of graph data. The grid node power-voltage sensitivity matrix is specifically:
[0052]
[0053] Where S is the grid's reactive sensitivity matrix; Q is the grid's reactive power matrix; U is the grid's node voltage matrix, and △ represents the partial derivative. A directed graph G = (X, A) is constructed to store matrix S. The elements on the main diagonal of the sensitivity matrix are arranged in chronological order to construct the node information vector X. The remaining elements are used to construct the edge information vector A.
[0054] Step S1-2: Use GIN to reduce the dimensionality of the graph data. Input the graph data G into the GIN network for dimensionality reduction, and output a feature vector H(G) representing the topological characteristics. The GIN is a graph isomorphic network.
[0055] The power grid in this embodiment is an improved power grid system based on the IEEE 33 node calculation example. Figure 1 The system includes distributed photovoltaics connected at electrical nodes 13, 18, 19, 25, 28, and 33, each with a capacity of 0.6 MW. It also includes discrete switching capacitors connected at electrical nodes 5 and 32, with four adjustable levels and a capacity of 0.25 MW each. Four tie switches, TS1 through TS4, are also connected.
[0056] Step S2: Based on the feature vector, a K-means clustering algorithm is used to generate a typical topology set Ω = {G1, ..., G K}, where K is determined by the clustering convergence threshold, and Ω provides candidate topologies for network topology reconstruction;
[0057] The cluster convergence threshold in cluster analysis is the silhouette coefficient ε = 0.46. The candidate topologies selected according to the above method are shown in Table 1.
[0058] Table 1 Candidate topology set for network topology reconstruction
[0059]
[0060] Each element G in the typical topological set Ω i The power grid topology corresponding to the cluster center represents the network structure under a specific operation scenario.
[0061] Step S3: Establish a multi-agent control system and implement hierarchical training, including:
[0062] Step S3-1: Establish a network topology reconstruction agent A1, and select the optimal topology G based on the voltage deviation and over-limit status within the network topology reconstruction period τ1. i ∈Ω, and update the A1 parameters through reconstruction feedback;
[0063] Step S3-2: Establish a long-time-scale agent A2, calculate the optimal switching plan for the long-time-scale equipment based on the voltage deviation and over-limit status within the long-time-scale control period τ2, and update the A2 parameters through feedback;
[0064] Step S3-3: Establish a short-time-scale agent A3, adjust the short-time-scale device based on the voltage deviation and over-limit status within the short-time-scale control period τ3, and update the A3 parameters through feedback;
[0065] Hierarchical collaborative training involves the agent receiving the environment state, executing control actions, and iteratively updating its control parameters based on the reward function. The environment state received by the agent is:
[0066]
[0067] Where s1 is the environment state vector received by the network topology reconstruction agent. s2 is the environment state vector received by the long-time scale device agent. s3 is the environment state vector received by the short-time scale agent. Represents the gear position of the capacitor bank at time t, K t Represents the grid topology at time t. is the cumulative value of voltage exceeding the limit of all nodes in the power grid at time t, is the cumulative value of the voltage offset of all nodes in the power grid at time t, max() means taking the larger value of the two, U i,t is the voltage amplitude of node i at time t. Respectively represent the active power output and reactive power output of photovoltaic at time t, Represent the active power and reactive power consumed by the load at time t respectively. The control action performed by the agent is:
[0068]
[0069] Where a1 is the action space of the network topology reconstruction agent, a2 is the action space of the long-time scale agent, and a3 is the action space of the short-time scale agent. K is the network topology reconstruction strategy, a CB is the gear control instruction of all capacitors in the power grid, a PV is the reactive power instruction of all photovoltaics in the grid. The reward of the agent is:
[0070]
[0071] In the formula is the reward of the network topology reconstruction agent and the long-time scale agent at time t. When k = 95, it represents the reward of the network topology reconstruction agent, and when k = 23, it represents the reward of the long-time scale agent. is the short-time-scale agent reward at time t, is the value of the power grid network loss at time t, I ij,t is the current amplitude of the line between nodes i and j at time t, R ij Represents the resistance of the line between nodes i and j; is the cumulative value of voltage exceeding the limit of all nodes in the power grid at time t, is the cumulative value of the voltage offset of all nodes in the power grid at time t, and α is the discount factor, which is taken as 0.5.
[0072] The network topology reconstruction agent and the long-time scale agent are each composed of a DDQN network, and their parameters are updated by gradient descent of the following loss function:
[0073]
[0074] Among them, L represents the loss function of the DDQN network, θ represents the network parameters, E represents the expected calculation, r DDQN represents the reward of the network topology reconstruction agent or the long-time scale agent, γ DDQN represents the conversion factor, Q DDQNrepresents the value evaluation value of the network topology reconstruction agent or the long-time scale agent, a represents the action decision of the network topology reconstruction agent or the long-time scale agent, s represents the environmental state received by the network topology reconstruction agent or the long-time scale agent, s' represents the state received by the agent after taking action decision a under the environmental state s, a' represents the agent action under the environmental state s', α represents the learning rate, Represents the gradient.
[0075] The short-time-scale agent consists of a value evaluation network and an action network. When updating parameters, the short-time-scale agent needs to update the parameters of both networks simultaneously. The value evaluation network updates its parameters by performing gradient descent on the following loss function:
[0076]
[0077] The action network updates its parameters by gradient descent on the following loss function:
[0078]
[0079] Among them, L' represents the loss function of the value evaluation network, θ Q1 Represents the parameters of the value assessment network, r DDPG represents the reward of the short-time-scale device agent, γ DDPG represents the conversion coefficient of the short-time-scale device agent, Q Q1 represents the value evaluation value of the value evaluation network, s1 represents the environmental state received by the short-time scale agent, a1 represents the action of the short-time scale agent, s1' represents the environmental state received by the short-time scale agent after taking the action decision a1 under the environmental state s1, a1' represents the action decision under the environmental state s1', α Q1 represents the learning rate of the value evaluation network, J represents the loss function of the action network, θ π represents the parameters of the action network, α π represents the learning rate of the action network, Indicates the gradient of action a1, represents the value of the action network parameter at time t, Represents the value of the value evaluation network parameters at time t.
[0080] Step S4: Deploy the trained A1, A2, and A3 to the power grid control center to obtain the optimized network topology reconstruction decision, long-time scale device action decision, and short-time scale device action decision.
[0081] In this example, the long-timescale devices are discrete switched capacitors connected to electrical nodes 5 and 32, and the short-timescale devices are distributed photovoltaic devices connected to electrical nodes 13, 18, 19, 25, 28, and 33. The network topology reconstruction period, long-timescale control period, and short-timescale control period are 24 hours, 6 hours, and 15 minutes, respectively.
[0082] The grid voltage distribution before optimization is as follows: Figure 2 As shown in the figure, the voltage peak of the IEEE33 bus system before optimization is very high, close to 1.07pu. The voltage distribution of the grid after optimization by the proposed method is shown in the figure. Figure 3 As shown in the figure, after optimization using the proposed method, the voltage peak is significantly reduced, close to the per-unit value 1.0pu, and the voltage fluctuation is smaller, indicating that the proposed method has achieved good voltage control effect and can improve the power quality of the grid.
[0083] Table 2 summarizes the power quality indicators of the power grid before and after optimization. As shown in the table, the proposed method effectively eliminates voltage over-limit phenomena, reducing the voltage over-limit rate from 1.5% to 0% and the voltage deviation from 0.0102 pu to 0.0064 pu, significantly enhancing system reliability. Furthermore, the proposed method improves the economic efficiency of power grid operation and reduces average network losses by 26%.
[0084] Table 2 Power quality indicators of the power grid before and after optimization
[0085]
[0086] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
[0087] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.
Claims
1. A multi-time-scale stabilization control method for power grid voltage considering network topology reconstruction, characterized in that: The control method comprises the following steps: Step S1: Build a graph neural network feature extraction model for the power grid, input the power-voltage electrical parameter matrix of the power grid nodes, and output a feature vector representing the topology and electrical characteristics; Step S2: Based on the feature vectors in step S1, a K-means clustering algorithm is used to generate a typical topology set Ω = {G1, ..., G K }, where K is determined by the clustering convergence threshold, and Ω provides candidate topologies for network topology reconstruction; Step S3: Establish a multi-agent control system and implement hierarchical training, including: Step S3-1: Establish a network topology reconstruction agent A1, and select the optimal topology G based on the voltage deviation and over-limit status within the network topology reconstruction period τ1. i ∈Ω, and update the A1 parameters through reconstruction feedback; Step S3-2: Establish a long-time-scale agent A2, calculate the optimal switching plan for the long-time-scale equipment based on the voltage deviation and over-limit status within the long-time-scale control period τ2, and update the A2 parameters through feedback; Step S3-3: Establish a short-time-scale agent A3, adjust the short-time-scale device based on the voltage deviation and over-limit status within the short-time-scale control period τ3, and update the A3 parameters through feedback; Step S4: Deploy the trained A1, A2, and A3 to the power grid control center to obtain the optimized network topology reconstruction decision, long-time scale device action decision, and short-time scale device action decision.
2. The grid voltage multi-time scale stability control method considering network topology reconstruction according to claim 1 is characterized in that: The specific steps of step S1 are as follows: Step S1-1: Express the grid reactive power voltage sensitivity matrix in a graphical data format to form a grid node power-voltage sensitivity matrix, specifically: Where S is the reactive voltage sensitivity matrix of the power grid; Q is the reactive power matrix of the power grid; U is the node voltage matrix of the power grid; △ represents the partial derivative. A directed graph G = (X, A) is constructed to store the matrix S. The elements on the main diagonal of the sensitivity matrix are arranged in chronological order to construct the node information vector X, and the remaining elements are used to construct the edge information vector A. Step S1-2: Use GIN to reduce the dimension of the above graph data. Input the graph data G into the GIN network for dimensionality reduction and output a feature vector H(G) representing the topological characteristics. The GIN is a graph isomorphic network.
3. The grid voltage multi-time scale stability control method considering network topology reconstruction according to claim 1 is characterized in that: The cluster convergence threshold in step S2 is the silhouette coefficient ε=0.46; Each element G in the typical topology set Ω in step S2 i The power grid topology corresponding to the cluster center represents the network structure under a specific operation scenario.
4. The method for multi-time-scale stabilization control of power grid voltage taking into account network topology reconstruction according to claim 1, characterized in that: In step S3, the network topology reconstruction period, the long time scale control period, and the short time scale control period are 24 hours, 6 hours, and 15 minutes, respectively.
5. The grid voltage multi-time scale stability control method considering network topology reconstruction according to claim 1 is characterized in that: The hierarchical collaborative training in step S3 specifically includes the agent receiving the environment state, executing the control action, and iteratively updating its control parameters according to the reward function.
6. The method for multi-time-scale stabilization control of power grid voltage taking into account network topology reconstruction according to claim 5, characterized in that: The agent receives the environment state as: Where s1 is the environment state vector received by the network topology reconstruction agent, s2 is the environment state vector received by the long-time scale device agent, and s3 is the environment state vector received by the short-time scale agent. Represents the gear position of the capacitor bank at time t, K t represents the grid topology at time t, is the cumulative value of voltage exceeding the limit of all nodes in the power grid at time t, is the cumulative value of the voltage offset of all nodes in the power grid at time t, max() means taking the larger value of the two, U i,t is the voltage amplitude of node i at time t, Respectively represent the active power output and reactive power output of photovoltaic at time t, They represent the active power and reactive power consumed by the load at time t respectively.
7. The method for multi-time-scale stabilization control of power grid voltage taking into account network topology reconstruction according to claim 5, characterized in that: The control actions performed by the agent are: Where a1 is the action space of the network topology reconstruction agent, a2 is the action space of the long-time scale agent, a3 is the action space of the short-time scale agent, and a K is the network topology reconstruction strategy, a CB is the gear control instruction of all capacitors in the power grid, a PV It is the reactive power instruction of all photovoltaics in the grid.
8. The method for multi-time-scale stabilization control of power grid voltage taking into account network topology reconstruction according to claim 5, characterized in that: The agent reward function is: In the formula is the reward of the network topology reconstruction agent and the long-time scale agent at time t. When k = 95, it represents the reward of the network topology reconstruction agent. When k = 23, it represents the reward of the long-time scale agent. is the short-time-scale agent reward at time t, is the value of the power grid network loss at time t, I ij,t is the current amplitude of the line between nodes i and j at time t, R ij Represents the resistance of the line between nodes i and j; is the cumulative value of voltage exceeding the limit of all nodes in the power grid at time t, is the cumulative value of the voltage offset of all nodes in the power grid at time t, and α is the discount factor, which is taken as 0.
5.
9. The method for multi-time-scale stabilization control of power grid voltage taking into account network topology reconstruction according to claim 1, characterized in that: The network topology reconstruction agent and the long-time scale agent in step S3 are each composed of a DDQN network, and the parameters are updated by gradient descent of the following loss function. The DDQN is a dual deep Q network: Where L represents the loss function of the DDQN network, θ t represents network parameters, E represents expected calculation, r DDQN represents the reward of the network topology reconstruction agent or the long-time scale agent, γ DDQN represents the conversion factor, Q DDQN represents the value evaluation value of the network topology reconstruction agent or the long-time scale agent, a represents the action decision of the network topology reconstruction agent or the long-time scale agent, s represents the environmental state received by the network topology reconstruction agent or the long-time scale agent, s' represents the state received by the agent after taking action decision a under the environmental state s, a' represents the agent action under the environmental state s', α represents the learning rate, Represents the gradient.
10. The grid voltage multi-time scale stability control method considering network topology reconstruction according to claim 1, characterized in that: In step S3, the short-time-scale agent consists of a value evaluation network and an action network. When updating parameters, the short-time-scale agent needs to update the parameters of the two networks simultaneously. The value evaluation network updates the parameters by performing gradient descent on the following loss function: The action network updates its parameters by gradient descent on the following loss function: Among them, L' represents the loss function of the value evaluation network, θ Q1 Represents the parameters of the value assessment network, r DDPG represents the reward of the short-time-scale device agent, γ DDPG represents the conversion coefficient of the short-time-scale device agent, Q Q1 represents the value evaluation value of the value evaluation network, s1 represents the environmental state received by the short-time scale agent, a1 represents the action of the short-time scale agent, s1' represents the environmental state received by the short-time scale agent after taking the action decision a1 under the environmental state s1, a1' represents the action decision under the environmental state s1', α Q1 represents the learning rate of the value evaluation network, J represents the loss function of the action network, θ π represents the parameters of the action network, α π represents the learning rate of the action network, represents the gradient, E represents the expected calculation, Indicates the gradient of action a1, represents the value of the action network parameter at time t, Represents the value of the value evaluation network parameters at time t.