Reconfigurable soft switch-containing power distribution network voltage control method based on joint learning
By adopting a voltage control method with recombinant soft switches based on joint learning in the distribution network, coordinating traditional discrete equipment and recombinable SOP, the problem of voltage violations in the distribution network after photovoltaic energy integration is solved, and rapid voltage regulation and economical and flexibility of system operation are achieved.
Patent Information
- Application Number
- CN202510016971.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-23
AI Technical Summary
The distribution network faces voltage violations after photovoltaic energy integration. Traditional voltage regulation devices cannot respond quickly to real-time fluctuations in voltage, resulting in unstable system operation.
Using a voltage control method for distribution networks with recombinant soft switches (R-SOP) based on joint learning, a two-stage voltage control framework with joint learning of DDPG and DQN is constructed to coordinate the collaboration between traditional discrete devices and recombinant SOPs to achieve rapid voltage regulation.
This method can minimize total power loss, reduce rapid voltage deviations in the active distribution network, and improve the economic and flexibility of system operation.
Smart Images

Figure CN120033715A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power distribution system operation optimization, and in particular to a voltage control method for a power distribution network containing reconfigurable soft switches based on joint learning. Background Art
[0002] Highly penetrated renewable energy generation is gradually being integrated into distribution networks. Photovoltaic (PV) power generation, as a clean energy with great production potential, is gaining increasing attention worldwide to meet the growing demand for renewable energy. However, with the widespread application of photovoltaic energy in distribution networks, distribution networks face significant challenges, which have strong uncertainties and lead to serious voltage violations. Voltage regulation has become a key issue and has put forward higher requirements for the flexible operation of distribution networks. In terms of voltage control, various reactive devices are integrated into distribution networks. Traditional slow-action devices such as on-load tap changers (OLTCs) and capacitor banks (CBs) are widely used. However, these traditional voltage regulation devices can only be adjusted over a long time scale and are mostly limited by the number of adjustments, and cannot quickly respond to real-time voltage fluctuations. With the support of advanced power electronic devices such as smart soft switches (SOPs) and multi-terminal smart soft switches (MOPs), traditional distribution networks have gradually evolved into highly flexible flexible distribution networks (FDNs). Distribution networks with flexible interconnected devices can regulate voltage through flexible and accurate active power transmission and reactive power compensation, thereby alleviating voltage fluctuations caused by distributed power sources (DGs). Both discrete variables and continuous variables in the voltage control action space are independent. In the action space of R-SOP, there is a complex coupling relationship due to active power transmission constraints and capacity constraints. And the action space of R-SOP is also discrete and continuous, and there is also connectivity. In order to fill the research gap in the application of deep reinforcement learning DRL in R-SOP, the present invention proposes a voltage control method for distribution network with reconfigurable soft switch based on joint learning. It aims to minimize the total power loss and mitigate the rapid voltage deviation of active distribution network ADN. Summary of the invention
[0003] The purpose of the present invention is to provide a voltage control method for a distribution network containing reconfigurable soft switches based on joint learning, which can reasonably utilize the multi-controllable devices of the distribution system, minimize the total power loss, and alleviate the rapid voltage deviation in the ADN, ultimately enhancing the economy of the system operation.
[0004] In one aspect of the present invention, the present invention proposes a voltage control method for a distribution network with reconfigurable soft switches based on joint learning. According to an embodiment of the present invention, the method comprises the following steps:
[0005] Step 1: Construct a two-stage voltage control framework based on joint learning of DDPG and DQN;
[0006] Step 2: Establish a day-ahead optimization model. The optimization model-related constraints include system power flow constraints, voltage safety constraints, and discrete equipment OLTC and CBs operation constraints.
[0007] Step 3: The optimal OLTC and CBs control strategies for voltage regulation are obtained through training and learning;
[0008] Step 4: Establishing the R-SOP model from the real-time voltage control stage;
[0009] Step 4.1, establish the R-SOP model of the real-time voltage control problem under the joint agent of DDPG and DQN, and transform this optimization problem into a Markov decision process;
[0010] Step 4.2: Train the R-SOP model through the joint agent of DDPG and DQN.
[0011] In addition, the voltage control method for a distribution network with reconfigurable soft switches based on joint learning according to the above embodiment of the present invention may also have the following additional technical features:
[0012] In some embodiments of the present invention, step 1 specifically includes the following steps: setting two intelligent agents, respectively as a day-ahead agent and a real-time agent, for voltage control at different time scales; in the first stage, based on the photovoltaic PV power generation and load demand predicted on the day-ahead, conveying them as observations to the day-ahead DQN agent, and the day-ahead DQN agent is trained in the Markov decision process framework to learn the optimal OLTC and CBs control strategies for voltage regulation; in the second stage, constructing a real-time stage joint agent based on the DDPG and DQN algorithms, and taking the scheduling results of the OLTC and CBs in the first stage as the observation information of the real-time agent, training the R-SOP in the real-time agent MDP framework, adjusting the R-SOP port power and feeder selection to achieve voltage control on a fast time scale.
[0013] In some embodiments of the present invention, in step 2, the day-ahead optimization model is as follows:
[0014]
[0015] In formula (1), the objective function of system voltage security and economy is included, λ L and λ V Respectively represent the weight coefficients of economic cost and voltage safety risk; Ω B It is a branch collection; N T is a collection of time periods, N N is the collection of all nodes in the system; Δt is the duration of each period, r ij is the resistance value of branch ij, Iij,t is the current flowing through branch ij at time t, U i,t is the voltage amplitude of node i at time t.
[0016] In some embodiments of the present invention, in step 2:
[0017] The system power flow constraints and voltage security constraints are as follows:
[0018]
[0019] Formula (2) is the node voltage constraint, U i,t is the voltage amplitude of node i at time t, U max and U min are the upper and lower limits of the node voltage safe operating range respectively; Equations (3) to (9) are the power flow constraints based on second-order cone programming, which describe the node power balance constraints; r ij and x ij are the resistance and reactance of branch ij respectively; I ij,t is the current on branch ij at time t; P ij,t and Q ij,t is the active power and reactive power on the branch ij at time t; P jk,t and Q jk,t is the active power and reactive power on the branch jk at time t; is the active power of the photovoltaic connected to the node i at time t; and is the active power and reactive power at node i at time t; is the reactive power injected by the CBs connected to the node i at time t; S ij is the apparent power transmitted on branch ij, P i,t , Q i,t are the active and reactive powers injected into node i at time t respectively; are the squares of the voltages on nodes i and j at time t;
[0020] The operating constraints of discrete equipment OLTC and CBs are as follows:
[0021]
[0022] Formula (10) is the relationship between OLTC regulation voltage and gear position and operation constraints, U i,t is the voltage on node i at time t, k ij,t and K ij,t is the transformation ratio and gear position of OLTC at time t, k ij,0 and Δk ij are the initial ratio and gear increment of OLTC respectively; N T is the sum of the periods, NOLTC It is the upper limit of the number of switching times in one day. is the maximum value of the gear change; Equation (11) represents the relationship between the reactive power injected by CBs and the gear and the operation constraints, represents the unit reactive power capacity of CBs at node i, is the injected reactive power of CBs at node i at time t, is the number of CBs switched on at node i at time t, N CB It is the upper limit of the number of switching times in one day. is the maximum value of the switching quantity, is the square of the voltage on node j; K ij,t-1 is the gear position of OLTC at time t-1; is the number of CBs switched on at node i at time t-1.
[0023] In some embodiments of the present invention, step 3 is specifically as follows:
[0024] 3.1. Formulate the MDP of the day-ahead voltage control problem. The MDP of the day-ahead DQN agent includes the state space Action Space and the reward function r t d ; In formula (12), the state space is defined, including the distribution network node load data, photovoltaic power generation, and the number of times OLTC and CBs have been operated; formula (13) defines the action space, including the gear action values of OLTC and CBs; formula (14) shows the design reward value;
[0025]
[0026]
[0027] In formula (12) - formula (14) and is the active power and reactive power at node i at time t; is the photovoltaic active power output at node i at time t; N OLTC,t and N CB,t are the times that OLTC and CBs have been operated respectively; T OLTC,t and T CB,t is the gear action value of OLTC and CBs; Indicates the deviation of the voltage at node i from the per-unit value at time t; κ 1 and κ 2 are the power loss coefficient and voltage deviation penalty coefficient in the reward function respectively;
[0028] 3.2. After formulating the MDP, train the DQN agent.
[0029] In some embodiments of the present invention, in step 3.2, the current DQN agent training steps are as follows:
[0030] After initializing the experience replay pool D and the parameters θ of the DQN network, according to the current state Select Action Use the ε-greedy strategy to select actions, as shown in formula (15):
[0031]
[0032] In the formula, is the Q value network of DQN, θ = [W, b] is the weight and bias;
[0033] After executing the action, observe the reward r t d and the next state Experience Store them in the experience replay pool D; when the experience pool reaches a certain capacity, randomly sample a batch of experience of size N from it to update the network;
[0034] The update of the Q-value network is achieved by minimizing the mean square error loss function. The loss function L(θ) of DQN is defined based on the mean square error (MSE) between the current Q-value and the target Q-value, as shown in formula (16):
[0035]
[0036] In the formula, for each experience, y i is the target Q value, is the current Q value, N is the number of samples in a small batch;
[0037] The target Q value is estimated by the target network, and the calculation formula is shown in formula (17):
[0038]
[0039] In formula (17), Q′ represents the target Q value network of DQN; θ - is the target network parameter; r i d is the reward value of sampling experience; γ is the discount factor; is the maximum Q value of the next state, where is the state of the next moment of sampling experience, a d For sampling experience The action in the state, θ - is the target network parameter;
[0040] For the target Q network of DQN, soft update is used to update it, as shown in formula (18):
[0041] θ - =τθ+(1-τ)θ - (18)
[0042] Where τ is the soft update factor, θ - are the parameters of the target Q network; θ are the parameters of the Q network.
[0043] In some embodiments of the present invention, step 4 specifically includes the following steps:
[0044] Step 4.1, build R-SOP operation constraints;
[0045] Step 4.2: For voltage control with R-SOP, the state space S contains node loads and photovoltaic and day-ahead discrete device scheduling results; the action space A consists of active power transmission, reactive power support and feeder selection of R-SOP; the real-time distribution network operation optimization problem is transformed into a Markov decision process, represented by the tuple <S,A,P,R,γ>, where P represents the state transfer function, R is the reward obtained by the agent for executing the action, and γ is the discount factor.
[0046] In some embodiments of the present invention, in step 4.1, the R-SOP operation constraints are established as follows:
[0047]
[0048] In equations (19) to (25), equations (19) to (21) are power balance constraints, where: is the active power of the R-SOP DC side connected to node i at time t; is the active power actually transmitted by the R-SOP accessed by node i at time t; is the active power loss of R-SOP connected to node i at time t; is the loss factor of R-SOP; Ω R-SOP It is a set of R-SOP ports; represents the power capacity injected into the node i connected to the R-SOP; is the reactive power injected by the voltage source converter at node i at time t; B i,n represents the switch state of node i on the nth branch connected to the R-SOP; Equations (22) to (25) are the R-SOP capacity constraints, where: It represents the power transmission capacity of the nth branch connected to R-SOP; Table 1 shows the capacity of the R-SOP accessed by node i; B is the reactive power actually transmitted by the R-SOP connected to node i at time t; n Indicates the switch status on the nth branch connected to the R-SOP; For VSC i Reactive power output limit; N s is the total number of branches connected to R-SOP;
[0049] The environment of the DDPG and DQN joint agent is the distribution system power flow model in Equations (2) to (9), where Equations (5) and (6) are modified as follows to add the active transmission and reactive compensation of R-SOP, as shown in Equations (26) to (27):
[0050]
[0051] In formula (26)-formula (27), and is the active power and reactive power emitted by R-SOP on node i at time t.
[0052] In some embodiments of the present invention, step 4.2 specifically includes the following steps:
[0053] 1) Construct the state space of the DDPG and DQN joint agent using formula (28):
[0054] The state space consists of the injected active power and injected reactive power of each node in the distribution network, the active power injected by photovoltaics, and the discrete device actions that constitute the day before, which can be expressed as:
[0055]
[0056] In formula (28), and They represent the active load and reactive load of the node at time t respectively, Indicates the active power injected by the photovoltaic power station, tap t represents the tap position of the on-load tapchanger at time t, is the reactive power value of the capacitor bank switched on at time t;
[0057] 2) Construct the R-SOP action space of the DDPG and DQN joint agent from formula (29):
[0058] The action space is controlled by DDPG R-SOP (N S -1) Active power transmitted by each port, N S The reactive power provided by the ports and the N controlled by DQN S Feeder selection for each port:
[0059]
[0060] In formula (29), a t,DDPG and a t,DQN They represent the power output action and feeder selection action of the DDPG agent and the DQN agent at time t, respectively; and They represent the active transmission and reactive output of the R-SOP port at time t, represents the feeder selection of the R-SOP port at time t, and Ns represents the number of R-SOP ports in the distribution network;
[0061] 3) Construct the reward function of the DDPG and DQN joint agent from formula (30):
[0062] Because the deep reinforcement learning algorithm is a strategy designed to explore long-term reward maximization, when designing the reward function of the Markov decision process, the objective function is inverted to obtain:
[0063]
[0064] In formula (29), κ 1 and κ 2 Represent the penalty coefficients for power loss and voltage respectively.
[0065] In some embodiments of the present invention, step 5 specifically includes the following steps:
[0066] First, after the joint agent observes the input state, a deterministic strategy is obtained through the DDPG policy network; noise is added to construct the behavior strategy, and the DQN agent selects the action strategy through the ε-greedy strategy. The two actions together constitute the output strategy of R-SOP, as shown in Equation (31)-Equation (32):
[0067]
[0068] In formula (31)-formula (32) is the deterministic policy output by the DDPG policy network, where is the state at time t, θ u =[W,b] are the weights and biases in the DDPG policy network, N(0,σ t ) is a process of adding random quantities. The exploration of action values is completed by adding a noise sample to the network output value. The noise follows a normal distribution with a mean of zero and a standard deviation of σ t , parameter σ t The size of represents the degree of exploration and decreases with the decay rate during training; is the Q value network of DQN, θ = [W, b] is the weight and bias;
[0069] Combine the actions output by the DDPG and DQN agents and calculate the reward value r of the strategy by interacting with the environment t r and the state at the next moment Will Stored in the experience replay pool; when the samples in the experience replay pool reach the upper limit, a random small batch is sampled from the experience replay pool in each iteration, and the DDPG agent calculates the target Q value through the target value network, as shown in formula (33):
[0070]
[0071] In formula (33), y DDPG is the target Q value of the DDPG network; is the state at the next moment; θ μ′ is the target strategy network parameter; θ Q′ is the target value network parameter; Q′ DDPG is the target value network of DDPG; is the strategy of the target value network; r t r represents the reward; γ is the discount factor;
[0072] The update of the value network is achieved by minimizing the mean square error loss function: its loss function is expressed as minimizing the mean square error, as shown in formula (34):
[0073]
[0074] In formula (33), L DDPG (θ) is the loss function of the DDPG value network; Q DDPG is the value network about the state at time t and actions The output, θ Q is the value network parameter;
[0075] For the policy network, its parameter θ μ The update of is based on the deterministic policy gradient, as shown in formula (35):
[0076]
[0077] In the formula, is the gradient after the policy network parameters are updated, s i is the sampling experience state; is the gradient of the value network parameters; is the gradient of the policy network; s r and a r are states and actions respectively, θQ and θ μ are the parameters of the value network and the strategy network respectively;
[0078] The target value network and target policy network are updated using a soft update mechanism, as shown in Equation (36)-Equation (37):
[0079] θ Q′ =τθ Q +(1-τ)θ Q′ (36)
[0080] θ μ′ =τθ μ +(1-τ)θ μ′ (37)
[0081] In formula (36)-(37), τ is a constant, which indicates the update speed of the target network; θ Q′ and θ μ′ are the parameters of the target value network and the target strategy network respectively; θ Q and θ μ are the parameters of the value network and the strategy network;
[0082] For the update of the DQN network, the target Q value is first calculated, the gap between the Q network output and the target Q value is calculated using the mean square error loss function, and the network parameters are updated using gradient descent, as shown in equations (38)-(39):
[0083]
[0084] θ←θ-α▽ θ (Q(sr,ar;θ)-y) 2 (39)
[0085] In formula (38)-formula (39), L DQN (θ) is the loss function of the DQN network; r t r is the reward value of sampling experience; γ is the discount factor; is the maximum Q value of the next state, where Q′ DQN represents the target Q value network of DQN, To sample the state of the next moment, For sampling experience The action in the state, θ - is the target Q network parameter; is the current Q value, where is the sampled experience state, For sampling experience The action in the state, θ is the Q network parameter; α is the learning rate; ▽ θis the gradient of the Q network parameters;
[0086] The target Q network of DQN is also updated using a soft update mechanism, as shown in formula (40):
[0087] θ - =τθ+(1-τ)θ - (40)
[0088] In formula (40), θ is the Q network parameter; θ - is the target Q network parameter; τ represents the update speed of the target network.
[0089] Compared with the prior art, the present invention has the following beneficial effects:
[0090] 1. This paper proposes a two-stage voltage control strategy to coordinate the collaboration between traditional voltage regulation equipment and reconfigurable SOP, covering two time scales. In the day-ahead stage, a single agent based on the DQN algorithm is used to schedule OLTC and CB. In the real-time stage, a real-time joint agent combining DDPG and DQN algorithms is constructed to control R-SOP to cope with the regulation needs of rapid voltage violations. The goal of this coordinated optimization is to minimize the total energy consumption while balancing the power loss caused by rapid changes in photovoltaic power generation and voltage violations.
[0091] 2. Introducing asymmetric reconfigurable smart soft switch (R-SOP) in the real-time voltage control stage, this technology combines feeder selection switches and asymmetric AC / DC converters to achieve flexible power transmission between feeders. Compared with the traditional multi-terminal SOP, this design has significant advantages in the following aspects: first, it reduces the overall cost of the converter and improves the economy; second, by optimizing the feeder selection, it significantly improves the utilization of the system and ensures the efficient allocation of power resources; third, it reduces the power loss of the feeder during power transmission, thereby improving the overall energy efficiency; finally, this innovation enhances the flexibility of the distribution network in dealing with load changes and renewable energy access.
[0092] 3. A joint agent combining DDPG (Deep Deterministic Policy Gradient) and DQN (Deep Q Network) algorithms is used to simultaneously process continuous and discrete control variables in R-SOP, thereby achieving precise control of feeder active transmission and reactive compensation. This method optimizes the real-time voltage control process and significantly reduces network power loss and voltage deviation in online real-time control by dynamically adjusting feeder selection and port power output. The training mechanism of this joint algorithm not only improves the flexibility and efficiency of control, but also provides a new solution to multivariable control problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] Figure 1It is a two-stage voltage control framework diagram in Embodiment 1 of the present invention;
[0094] Figure 2 It is a topological structure diagram of R-SOP in Example 1 of the present invention;
[0095] Figure 3 It is a diagram of the IEEE33 node system in the application example of the present invention;
[0096] Figure 4 is a diagram of the training process of a real-time agent in an application example of the present invention;
[0097] Figure 5 It is a photovoltaic prediction and daily load curve diagram in the application example of the present invention;
[0098] Figure 6 It is a diagram of the gear position changes of OLTC and CB in the application example of the present invention;
[0099] Figure 7 It is a voltage distribution diagram within one hour (12:00-13:00) after intraday optimization in the application example of the present invention;
[0100] Figure 8 It is a heat value diagram of voltage distribution within one hour (12:00-13:00) after intraday optimization in the application example of the present invention;
[0101] Fig. 9 It is a diagram of active transmission of R-SOP in an application example of the present invention;
[0102] Fig.10 1 is a diagram of reactive power compensation of R-SOP in an application example of the present invention;
[0103] Fig.11 is a diagram of the voltage condition of the node 18 in different scenarios in the application example of the present invention;
[0104] Fig.12 It is a voltage distribution curve diagram under different scenarios at 4:00am in the application example of the present invention;
[0105] Fig.13 It is a flow chart of the voltage control method of the distribution network with reconfigurable soft switches based on joint learning in Example 1 of the present invention. DETAILED DESCRIPTION
[0106] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0107] Example 1
[0108] like Fig.13 As shown, in this embodiment, the voltage control method of the distribution network with reconfigurable soft switch based on joint learning has the following specific steps:
[0109] Step 1: Construct a two-stage voltage control framework based on joint learning of DDPG and DQN.
[0110] A two-stage voltage control framework based on the joint learning of deep deterministic policy gradient network DDPG and classic deep Q learning DQN is proposed. Figure 1 As shown in the figure, a novel intelligent control framework is proposed to mitigate voltage violations while minimizing the real-time power loss of the network. The framework coordinates the collaboration of traditional discrete devices (on-load tap changer OLTC and capacitor banks CBs) and voltage continuous control devices reconfigurable smart soft switches R-SOP at two time scales. To this end, two intelligent agents are designed, namely the day-ahead agent and the real-time agent, which are specialized for voltage control at different time scales.
[0111] In the first stage, the photovoltaic PV generation and load demand are predicted based on the day-ahead forecast and communicated to the day-ahead DQN agent as observations. The DQN agent is trained to learn the optimal OLTC and CBs control strategy for voltage regulation within the Markov decision process (MDP) framework. First, according to its initial state and the DQN network output action. Then, the next state is obtained by power flow calculation using the input actions and historical data of PV, load, OLTC and CB in the environment. After that, the input action, current state and next state are stored in the memory pool for centralized learning. When the memory pool reaches a certain capacity, samples are extracted from it to update the network parameters to obtain the trained day-ahead DQN agent.
[0112] In the second stage, a real-time joint agent is constructed based on the DDPG and DQN algorithms, and the scheduling results of the OLTC and CB in the first stage are used as the observation information of the real-time agent. First, the action of R-SOP is output according to its initial state and the joint network, where the DDPG and DQN joint agents process the R-SOP port power and feeder selection respectively. Then, the state of PV, load, OLTC and CB, the action of R-SOP and historical data are used in the environment to calculate the next state. After that, the input action, current state and next state are stored in the memory pool for centralized learning. When the memory pool reaches a certain capacity, samples are extracted from it to update the network parameters until the model converges. R-SOP is trained in the real-time agent MDP framework, and the R-SOP port power and feeder selection are adjusted to achieve voltage control on a fast time scale.
[0113] Step 2: Establish the day-ahead optimization model. The optimization model-related constraints include system power flow constraints, voltage safety constraints, and discrete equipment OLTC and CBs operation constraints. Specifically, the following steps are included:
[0114] Step 2.1: Construct the day-ahead optimization model based on formula (1):
[0115]
[0116] In formula (1), the objective function of system voltage security and economy is included, λ L and λ V Respectively represent the weight coefficients of economic cost and voltage safety risk; Ω B It is a collection of branches. T is a collection of time periods, N N is the collection of all nodes in the system; Δt is the duration of each period, r ij is the resistance value of branch ij, I ij,t is the current flowing through branch ij at time t, U i,t is the voltage amplitude of node i at time t.
[0117] Step 2.2: The system power flow constraint and voltage security constraint are constructed from equations (2) to (9) as follows:
[0118]
[0119]
[0120] Formula (2) is the node voltage constraint, U i,t is the voltage amplitude of node i at time t, U max and U min are the upper and lower limits of the node voltage safe operating range respectively; Equations (3) to (9) are the power flow constraints based on second-order cone programming, which describe the node power balance constraints; r ij and x ij are the resistance and reactance of branch ij respectively; I ij,t is the current on branch ij at time t; P ij,t and Q ij,t is the active power and reactive power on the branch ij at time t; P jk,t and Q jk,t is the active power and reactive power on the branch jk at time t; is the active power of the photovoltaic connected to the node i at time t; and is the active power and reactive power at node i at time t; is the reactive power injected by the CBs connected to the node i at time t; S ijis the apparent power transmitted on branch ij, P i,t , Q i,t are the active and reactive powers injected into node i at time t respectively; are the squares of the voltages on nodes i and j at time t;
[0121] Step 2.3: The relevant operation constraints of discrete equipment OLTC and CBs are constructed from (10) to (11) as follows:
[0122]
[0123] Formula (10) is the relationship between OLTC regulation voltage and gear position and operation constraints, U i,t is the voltage on node i at time t, k ij,t and K ij,t is the transformation ratio and gear position of OLTC at time t, k ij,0 and Δk ij are the initial ratio and gear increment of OLTC respectively; N T is the sum of the periods, N OLTC It is the upper limit of the number of switching times in one day. is the maximum value of the gear change; Equation (11) represents the relationship between the reactive power injected by CBs and the gear and the operation constraints, represents the unit reactive power capacity of CBs at node i, is the injected reactive power of CBs at node i at time t, is the number of CBs switched on at node i at time t, N CB It is the upper limit of the number of switching times in one day. is the maximum value of the switching quantity, is the square of the voltage on node j at time t; K ij,t-1 is the gear position of OLTC at time t-1; is the number of CBs switched on at node i at time t-1.
[0124] Step 3: After training, the optimal OLTC and CBs control strategies for voltage regulation are obtained.
[0125] Based on the Markov decision process for the day-ahead voltage control problem formulated in step 2, the day-ahead DQN agent is trained within the Markov decision process (MDP) framework to learn the optimal OLTC and CBs control strategies for voltage regulation, including the following steps:
[0126] To formulate the MDP for the day-ahead voltage control problem, the MDP of the day-ahead DQN agent consists of the state space Action Space and the reward function r t d; The DQN agent determines the gear values of OLTC and CBs according to the predicted load demand and PV output. At the same time, OLTC and CBs must also meet their own action times limit; therefore, the state space is defined in formula (12), including the distribution network node load data, PV power generation, and the number of times OLTC and CBs have been acted; formula (13) defines the action space, including the gear action values of OLTC and CBs; the reward value is designed as shown in formula (14), taking into account the power loss and voltage deviation, to ensure that the training process learns the optimal strategy to ensure the safety and economy of the distribution network;
[0127]
[0128] In formula (12) - formula (14) and is the active power and reactive power at node i at time t; is the photovoltaic active power output at node i at time t; N OLTC,t and N CB,t are the times that OLTC and CBs have been operated respectively; T OLTC,t and T CB,t is the gear action value of OLTC and CBs; Indicates the deviation of the voltage at node i from the per-unit value at time t; κ 1 and κ 2 are the power loss coefficient and voltage deviation penalty coefficient in the reward function respectively;
[0129] After formulating the MDP, the specific training process of the DQN agent today is as follows: After initializing the experience replay pool D and the parameters θ of the DQN network, according to the current state Select Action Specifically, the ε-greedy strategy is used to select actions to balance exploration and utilization, as shown in formula (15):
[0130]
[0131] In the formula, is the Q value network of DQN, θ = [W, b] is the weight and bias;
[0132] After executing the action, observe the reward r t d and the next state Experience Stored in the experience replay pool D. When the experience pool reaches a certain capacity, a batch of experience of size N is randomly sampled from it to update the network;
[0133] The update of the Q-value network is achieved by minimizing the mean square error loss function. The loss function L(θ) of DQN is defined based on the mean square error (MSE) between the current Q-value and the target Q-value, as shown in formula (16):
[0134]
[0135] In the formula, for each experience, y i is the target Q value, is the current Q value, N is the number of samples in a small batch;
[0136] The target Q value is estimated by the target network, and the calculation formula is shown in formula (17):
[0137]
[0138] In formula (17), y i is the target Q value; Q′ represents the target Q value network of DQN; r i d is the reward value of sampling experience; γ is the discount factor; is the maximum Q value of the next state, where is the state of the next moment of sampling experience, a d For sampling experience The action in the state, θ - is the target network parameter;
[0139] For the target Q network of DQN, soft update is used to update it, as shown in formula (18):
[0140] θ - =τθ+(1-τ)θ - (18)
[0141] Where τ is the soft update factor, θ - are the parameters of the target Q network; θ are the parameters of the Q network.
[0142] Step 4: Establish the R-SOP model from the real-time voltage control stage. Based on the day-ahead scheduling results of OLTC and CBs, a mathematical model of the real-time voltage control problem is established under the joint agent of DDPG and DQN, and this optimization problem is transformed into a Markov decision process.
[0143] In the real-time voltage control stage, the R-SOP topology diagram is established as follows Figure 2As shown in the figure, in the real-time stage, the voltage is controlled by controlling the active transmission and reactive compensation of R-SOP on the feeder, while reducing the network loss. Since the control variables of R-SOP not only include the power output of each port, but also involve the feeder selection of the port, that is, it contains both continuous and discrete control variables, DDPG and DQN joint agents are used to train R-SOP in the real-time stage. Based on the scheduling results of OLTC and CBs a day ago, a mathematical model of the real-time voltage control problem is established under the joint agent of DDPG and DQN. R-SOP is jointly trained by DDPG and DQN joint agents to minimize the network power loss and voltage deviation in online real-time voltage control; then the R-SOP operation model is constructed, and based on the scheduling results of OLTC and CBs a day ago, a mathematical model of the real-time voltage control problem is established under the joint agent of DDPG and DQN, and this optimization problem is converted into a Markov decision process.
[0144] A mathematical model of the real-time voltage control problem is established under the joint agent of DDPG and DQN, and this optimization problem is transformed into a Markov decision process.
[0145] Step 4.1: Construct R-SOP operation constraints from equations (19) to (25):
[0146]
[0147] In equations (19) to (25), equations (19) to (21) are power balance constraints, where: is the active power of the R-SOP DC side connected to node i at time t; is the active power actually transmitted by the R-SOP accessed by node i at time t; is the active power loss of R-SOP connected to node i at time t; is the loss factor of R-SOP; Ω R-SOP is a collection of R-SOP ports, represents the power capacity injected into the node i connected to the R-SOP; is the reactive power injected by the voltage source converter at node i at time t; B i,n represents the switch state of node i on the nth branch connected to the R-SOP; Equations (22) to (25) are the R-SOP capacity constraints, where: It represents the power transmission capacity of the nth branch connected to R-SOP; Table 1 shows the capacity of the R-SOP accessed by node i; B is the reactive power actually transmitted by the R-SOP connected to node i at time t; n Indicates the switch status on the nth branch connected to the R-SOP; For VSC i Reactive power output limit; N s is the total number of branches connected to R-SOP;
[0148] The environment of the DDPG and DQN joint agent is the distribution system power flow model in Equations (2) to (9), where Equations (5) and (6) are modified as follows to add the active transmission and reactive compensation of R-SOP, as shown in Equations (26) to (27):
[0149]
[0150] In formula (26)-formula (27), and is the active power and reactive power emitted by R-SOP on node i at time t.
[0151] Step 4.2: Based on the scheduling results of the day-ahead OLTC and CBs, a mathematical model of the real-time voltage control problem is established under the joint agent of DDPG and DQN, and this optimization problem is transformed into a Markov decision process, which is characterized by the following steps:
[0152] For voltage control with R-SOP, the state space S contains node loads and photovoltaic and day-ahead discrete device scheduling results; the action space A consists of active power transmission, reactive power support and feeder selection of R-SOP; thus, the real-time distribution network operation optimization problem is transformed into a Markov decision process, which can be represented by the tuple <S,A,P,R,γ>, where P represents the state transfer function, R is the reward obtained by the agent for executing the action, and γ is the discount factor;
[0153] 1) Construct the state space of the DDPG and DQN joint agent using formula (28):
[0154] The state space consists of the injected active power and injected reactive power of each node in the distribution network, the active power injected by photovoltaics, and the discrete device actions that constitute the day before, which can be expressed as
[0155]
[0156] In formula (28), and They represent the active load and reactive load of the node at time t respectively, Indicates the active power injected by the photovoltaic power station, tap t represents the tap position of the on-load tapchanger at time t, It is the reactive power value of the capacitor bank switched on at time t.
[0157] 2) Construct the R-SOP action space of the DDPG and DQN joint agent from formula (29):
[0158] The action space is controlled by DDPG R-SOP (N S -1) Active power transmitted by each port, N S The reactive power provided by the ports and the N controlled by DQN S Feeder selection for each port.
[0159]
[0160] In formula (29), a t,DDPG and a t,DQN They represent the power output action and feeder selection action of the DDPG agent and the DQN agent at time t, respectively; and They represent the active transmission and reactive output of the R-SOP port at time t, represents the feeder selection of the R-SOP port at time t, and Ns represents the number of R-SOP ports in the distribution network;
[0161] 3) Construct the reward function of the DDPG and DQN joint agent from formula (30):
[0162] Because the deep reinforcement learning algorithm is a strategy designed to explore long-term reward maximization, when designing the reward function of the Markov decision process, the objective function is inverted to obtain:
[0163]
[0164] In formula (29), κ 1 and κ 2 Represent the penalty coefficients of power loss and voltage respectively;
[0165] Step 5: R-SOP is trained jointly through DDPG and DQN agents to achieve voltage control at different time scales, effectively coordinating the collaboration between these traditional devices and the reconfigurable SOP to minimize network power loss and voltage deviation in online real-time voltage control.
[0166] Based on the Markov decision process of the formulated day-ahead voltage control problem, the R-SOP is jointly trained through the DDPG and DQN joint agents; thereby achieving voltage control at different time scales and effectively coordinating the collaboration between these traditional devices and the reconfigurable SOP to minimize network power loss and voltage deviation in online real-time voltage control.
[0167] The offline learning of the proposed DDPG and DQN joint agent training R-SOP method in the real-time voltage control stage can be regarded as a closed-loop trial-and-error process through interaction with the environment; in each iteration, the real-time DDPG agent takes action based on the observed state, reflecting the power output of the R-SOP, and the real-time DQN agent takes action based on the observation, which reflects the feeder selection of the R-SOP, combining the two actions to interact with the environment; the environment receives the action, performs power flow calculations, and returns the updated state information to the agent; in order to evaluate the effectiveness of the actions of this real-time agent, an appropriate reward function is set to minimize power loss and voltage deviation; the experience is stored in the experience pool, and when the experience pool capacity reaches the upper limit, the agent updates the DDPG and DQN neural network parameters;
[0168] Specifically, after the joint agent observes the input state, a deterministic strategy is obtained through the DDPG policy network. In order to increase the exploratory nature, noise is added to construct the behavior strategy. At the same time, the DQN agent selects the action strategy through the ε-greedy strategy. The two actions together constitute the output strategy of R-SOP, as shown in Equation (31)-Equation (32):
[0169]
[0170] In formula (31)-formula (32) is the deterministic policy output by the DDPG policy network, where is the state at time t, θ u =[W,b] are the weights and biases in the DDPG policy network, N(0,σ t ) is a process of adding random quantities. The exploration of action values is completed by adding a noise sample to the network output value. The noise follows a normal distribution with a mean of zero and a standard deviation of σ t , parameter σ t The size of represents the degree of exploration and decreases with the decay rate during training; is the Q value network of DQN, θ = [W, b] is the weight and bias;
[0171] Combine the actions output by the DDPG and DQN agents and calculate the reward value r of the strategy by interacting with the environment t r and the state at the next moment Will Stored in the experience replay pool. When the samples in the experience replay pool reach the upper limit, a random small batch is sampled from the experience replay pool in each iteration, and the DDPG agent calculates the target Q value through the target value network, as shown in formula (33):
[0172]
[0173] In formula (33), y DDPG is the target Q value of the DDPG network; is the state at the next moment; θ μ′ is the target strategy network parameter; θ Q′ is the target value network parameter; Q′ DDPG is the target value network of DDPG; is the strategy of the target value network; r t r represents the reward; γ is the discount factor;
[0174] The update of the value network is achieved by minimizing the mean square error loss function: its loss function can be expressed as minimizing the mean square error, as shown in formula (34):
[0175]
[0176] In formula (34), L DDPG (θ) is the loss function of the DDPG value network; Q DDPG is the value network about the state at time t and actions The output, θ Q is the value network parameter;
[0177] For the policy network, its parameter θ μ The update of is based on the deterministic policy gradient, as shown in formula (35):
[0178]
[0179] In the formula, is the gradient after the policy network parameters are updated, s i is the sampling experience state; is the gradient of the value network parameters; is the gradient of the policy network; s r and a r are states and actions respectively, θ Q and θ μ are the parameters of the value network and the strategy network respectively;
[0180] The target value network and target policy network are updated using a soft update mechanism, as shown in Equation (36)-Equation (37):
[0181] θ Q′ =τθ Q +(1-τ)θ Q′ (36)
[0182] θ μ′ =τθμ +(1-τ)θ μ′ (37)
[0183] In formula (36)-(37), τ is a constant, which indicates the update speed of the target network; θ Q′ and θ μ′ are the parameters of the target value network and the target strategy network respectively; θ Q and θ μ are the parameters of the value network and the strategy network;
[0184] For the update of the DQN network, the target Q value is first calculated, the gap between the Q network output and the target Q value is calculated using the mean square error loss function, and the network parameters are updated using gradient descent, as shown in equations (38)-(39):
[0185]
[0186] θ←θ-α▽ θ L DQN (θ)(39)
[0187] Formula (38)-In formula (38), L DQN (θ) is the loss function of the DQN network; r t r is the reward value of sampling experience; γ is the discount factor; is the maximum Q value of the next state, where Q′ DQN represents the target Q value network of DQN, To sample the state of the next moment, For sampling experience The action in the state, θ - is the target Q network parameter; is the current Q value, where is the sampled experience state, For sampling experience The action in the state, θ is the Q network parameter; α is the learning rate; ▽ θ is the gradient of the Q network parameters;
[0188] The parameters of the target Q network of DQN are θ - The soft update mechanism is also used for updating, as shown in formula (40):
[0189] θ - =τθ+(1-τ)θ - (40)
[0190] Where, θ is the Q network parameter; θ - is the target Q network parameter; τ represents the update speed of the target network.
[0191] Example 2
[0192] An electronic device includes a memory and a processor, wherein the memory is used to store a program that supports the processor to execute the method described in Example 1, and the processor is configured to execute the program stored in the memory.
[0193] Example 3
[0194] A computer-readable storage medium stores a computer program on the computer-readable storage medium, and the computer program executes the steps of the method described in Example 1 when executed by a processor.
[0195] Application Examples
[0196] Use Figure 3 The feasibility and effectiveness of the model and method of the embodiment 1 of the invention are verified by performing a simulation on the modified IEEE 33-node system. The simulation is performed on a 64-bit computer with a 3.00GHz CPU and 128GB RAM. The optimization program is written in the Python 3.8 environment, and PYPOWER 5.1.6 and Pytorch 1.8.1 are used to solve the model. The training process of the neural network is performed by Pytorch on an NVIDIA GeForce RTX 4090Ti GPU with 12GB RAM.
[0197] 1. Test system and parameter settings
[0198] The two-stage R-SOP voltage control strategy proposed in this invention is implemented on an improved IEEE33 node system with a rated voltage level of 12.66 kV. The network structure of the test system is shown in the figure. Figure 3 As shown. Five PV units are connected to the grid and installed at nodes 7, 10, 18, 24 and 27, with capacities of 300kWp, 700kWp, 300kWp, 500kWp and 700kWp respectively. The capacity, operating parameters and access location of the voltage control equipment are shown in Table 1. A set of 4-port R-SOPs with a capacity of 4kVA are installed between nodes 12, 22, 25 and 29. The loss coefficient of each converter is 0.02. There is an 11-speed OLTC between nodes 1 and 2, and each speed represents 1% regulation. There is a 6-speed CB on node 33, and each switching unit can compensate for 60kVar of reactive power. The number of actions is limited to N OLTC and N CB They are 4 and 6 times respectively. The safe range of node voltage is between 0.95pu and 1.05pu.
[0199] Table 1 Voltage control device parameters
[0200]
[0201] In the real-time voltage control stage, the dataset contains 115 days of measurement data, including 110 days for training and 5 days for testing. The scheduling period is 24 hours and the sampling interval is 5 minutes. To maximize the long-term cumulative benefits, the discount factor is set to 0.99. The number of input and output nodes of the neural network corresponds to the state dimension and action dimension of the agent, respectively. Table 2 lists the specific parameter settings of the DDPG and DQN joint training agents. The neural network parameters of the DQN agent are consistent with the real-time stage. The action network consists of one input layer, two hidden layers and one output layer. The number of input and output nodes corresponds to the state and action dimensions of the agent. The hidden layers are all fully connected structures, containing 128 and 56 neurons, and the activation function is ReLU. The critic network consists of one input layer, three hidden layers and one output layer. The hidden layers are also fully connected. The number of neurons is 256, 128 and 64, respectively, and the activation function is also ReLU. The maximum capacity of the priority experience replay pool is 10,000 samples. After reaching the upper limit, 32 samples are extracted from small and medium batches to update the parameters of the neural network. The learning rates of the action network and the critic network are 0.0001 and 0.001 respectively.
[0202] Table 2 Network learning hyperparameters and related parameter settings
[0203]
[0204] 2. Model training process
[0205] The DDPG and DQN joint agent used for training was built on the pytorch platform and trained using Python 3.8, pytorch1.8.1, and IDE Pycharm.
[0206] In the real-time training process, a total of 110 days of training data were generated based on the real PV and load outputs. The proposed real-time joint agent was trained 30,000 times to learn the optimal control strategy of R-SOP in the operation of the distribution system. The training process is as follows: Figure 4 As shown in the figure. The total time cost of offline training on IEEE33 nodes is about 12 minutes. The offline process only needs to be executed once to achieve long-term application, and no retraining is required when the distribution system grid structure does not change. After the number of samples in the experience replay pool reaches the upper limit, the multi-modal agent begins to learn and update the parameters of the neural network. As can be seen from the figure, the total reward remains fluctuating, but it shows an upward trend, and the corresponding neural network converges after about 12,000 iterations.
[0207] From the cumulative reward curve of training, we can observe that the reward value changes with the increase of training steps, and the curve gradually stabilizes and approaches the maximum value, which indicates that the performance of the model tends to stabilize after enough training. At first, the curve may show large fluctuations, which is because the model is exploring different strategies. As the training progresses, the fluctuation of the reward value gradually decreases and tends to converge, indicating that the model's strategy is gradually optimized and its decision-making ability has been improved.
[0208] 3. Current optimization results
[0209] Figure 5 It is the photovoltaic forecast data and daily load curve of one hour resolution in the test data set. The photovoltaic output is concentrated from 8:00 to 18:00. The photovoltaic active power is greater than the load active power from about 10:30 to 15:00, and the maximum output is reached around 12:00 noon. In the day-ahead dispatch stage, the dispatch results of OLTC and CB are as follows Figure 6 As shown. When the load level is high, both CB and OLTC are adjusted to a higher gear. Specifically, it can be seen that the OLTC tap is moved from gear '1' to gear '3' at 1:00, and is moved to gear '4' again at 13:00. It remains at gear '4' at other times, and operates 3 times in total, meeting the daily maximum adjustment limit. CB will gradually increase the reactive compensation power provided to the system to 240Kvar from 1:00 to 5:00, and then remain at 240Kvar unchanged, operating 3 times in total, also meeting the daily maximum adjustment limit.
[0210] 4. Real-time voltage control strategy analysis
[0211] In the real-time voltage control stage, Figure 7-8 The voltage distribution and variation of the IEEE33 nodes in the distribution network within one hour (12:00-13:00) after using the method proposed in this paper. As can be seen from the figure, all nodes are within the safety range of 0.95-1.05pu during this period. Fig. 9 and Fig.10 Show.
[0212] 5. Analysis of numerical results
[0213] A test day is used to evaluate the performance of the model trained by Integrated DDPG-DQNAgent (IDDA). In order to evaluate the effectiveness of the proposed method, this paper uses scenario 1 for comparison. Scenario 1: Without R-SOP control, the initial state of the distribution network is obtained as the basic scenario. Scenario 2: Use the method based on DDPG and DQN joint training proposed in this paper to optimize R-SOP operation.
[0214] Fig.11 The voltage curve of node 18 under different conditions is shown. In addition, the voltage optimization effect and power loss comparison under each scenario are shown in Table 3. The results show that in scenario I, without the support of R-SOP, the voltage distribution of some nodes is lower than the lower limit (0.95pu) due to the heavy load and small DG power generation. Compared with scenario I, the voltage curve in scenario II is significantly improved, the overall voltage is raised, and it is maintained within a safe range. Compared with scenario 1 without regulation, the average daily power loss using the method proposed in this paper is reduced by 16.71%.
[0215] Table 3 Comparison of voltage optimization effect and power loss in different scenarios
[0216]
[0217] As can be seen from the table, compared with scenario I, the maximum voltage deviation of scenario II is reduced by 51.43%. The minimum voltage value in scenario I is 0.9057, which is much lower than the voltage lower limit. In contrast, the voltage curves of scenario II are all kept within the safe voltage range of [0.95, 1.05], and the average voltage is 1.0010. The overall voltage situation is more concentrated and close to the per-unit value.
[0218] When t=4:00am, the voltage curves under different scenarios are as follows Fig.12 As shown. It can be observed that when no control strategy is used, that is, when R-SOP is not adopted for optimization control, nodes 7-12 exceed the lower limit. Here, it is observed that the proposed method can improve the voltage over-limit situation through the reactive support of R-SOP.
[0219] The above contents are merely examples and explanations of the present invention. Those skilled in the art may make various modifications or additions to the specific embodiments described or replace them in a similar manner. As long as they do not deviate from the structure of the present invention or exceed the scope defined by the claims, they shall all fall within the protection scope of the present invention.
Claims
1. A voltage control method for a distribution network with reconfigurable soft switches based on joint learning, characterized in that: The following steps are involved: Step 1: Construct a two-stage voltage control framework based on joint learning of DDPG and DQN; Step 2: Establish a day-ahead optimization model. The optimization model-related constraints include system power flow constraints, voltage safety constraints, and discrete equipment OLTC and CBs operation constraints. Step 3: The optimal OLTC and CBs control strategies for voltage regulation are obtained through training and learning; Step 4: Establish the R-SOP model in the real-time voltage control stage. Based on the day-ahead OLTC and CBs scheduling results, establish the R-SOP model of the real-time voltage control problem under the joint agent of DDPG and DQN, and transform this optimization problem into a Markov decision process. Step 5: Train the R-SOP model through the joint agent of DDPG and DQN.
2. The voltage control method of distribution network with reconfigurable soft switch based on joint learning according to claim 1 is characterized in that: The step 1 specifically includes the following steps: setting two intelligent agents, respectively as a day-ahead agent and a real-time agent, for voltage control at different time scales; in the first stage, based on the photovoltaic PV power generation and load demand predicted on the day-ahead, conveying them as observations to the day-ahead DQN agent, and the day-ahead DQN agent is trained in the Markov decision process framework to learn the optimal OLTC and CBs control strategies for voltage regulation; in the second stage, constructing a real-time stage joint agent based on the DDPG and DQN algorithms, and taking the scheduling results of the OLTC and CBs in the first stage as the observation information of the real-time agent, training the R-SOP in the real-time agent MDP framework, adjusting the R-SOP port power and feeder selection, so as to achieve voltage control on a fast time scale.
3. The voltage control method of distribution network with reconfigurable soft switch based on joint learning according to claim 1 is characterized in that: In step 2, the day-ahead optimization model is as follows: In formula (1), the objective function of system voltage security and economy is included, λ L and λ V Respectively represent the weight coefficients of economic cost and voltage safety risk; Ω B It is a branch collection; N T is a collection of time periods, N N It is the collection of all nodes in the system; Δt is the duration of each period, r ij is the resistance value of branch ij, I ij,t is the current flowing through branch ij at time t, U i,t is the voltage amplitude of node i at time t.
4. The voltage control method of distribution network with reconfigurable soft switch based on joint learning according to claim 1 is characterized in that: In step 2: The system power flow constraints and voltage safety constraints are as follows: Formula (2) is the node voltage constraint, U i,t is the voltage amplitude of node i at time t, U max and U min are the upper and lower limits of the node voltage safe operating range respectively; Equations (3) to (9) are the power flow constraints based on second-order cone programming, which describe the node power balance constraints; r ij and x ij are the resistance and reactance of branch ij respectively; I ij,t is the current on branch ij at time t; P ij,t and Q ij,t is the active power and reactive power on the branch ij at time t; P jk,t and Q jk,t is the active power and reactive power on the branch jk at time t; is the active power of the photovoltaic connected to the node i at time t; and is the active power and reactive power at node i at time t; is the reactive power injected by the CBs connected to the node i at time t; S ij is the apparent power transmitted on branch ij, P i,t , Q i,t are the active and reactive powers injected into node i at time t respectively; are the squares of the voltages on nodes i and j at time t respectively; The operating constraints of discrete equipment OLTC and CBs are as follows: Formula (10) is the relationship between OLTC regulation voltage and gear position and operation constraints, U i,t is the voltage on node i at time t, k ij,t and K ij,t is the transformation ratio and gear position of OLTC at time t, k ij,0 and Δk ij are the initial ratio and gear increment of OLTC respectively; N T is the sum of the periods, N OLTC It is the upper limit of the number of switching times in one day. is the maximum value of the gear change; Equation (11) represents the relationship between the reactive power injected by CBs and the gear and the operation constraints, represents the unit reactive power capacity of CBs at node i, is the injected reactive power of CBs at node i at time t, is the number of CBs switched on at node i at time t, N CB It is the upper limit of the number of switching times in one day. is the maximum value of the switching quantity, is the square of the voltage on node j; K ij,t-1 is the gear position of OLTC at time t-1; is the number of CBs switched on at node i at time t-1.
5. The voltage control method of distribution network with reconfigurable soft switch based on joint learning according to claim 1 is characterized in that: The step 3 is as follows: 3.
1. Formulate the MDP of the day-ahead voltage control problem. The MDP of the day-ahead DQN agent includes the state space Action Space And the reward function The state space is defined in formula (12), including the load data of the distribution network nodes, the PV power generation, and the number of times the OLTC and CBs have been operated; the action space is defined in formula (13), including the gear action values of the OLTC and CBs; the design reward value is shown in formula (14); In formula (12) - formula (14) and is the active power and reactive power at node i at time t; is the photovoltaic active power output at node i at time t; N OLTC,t and N CB,t are the times that OLTC and CBs have been operated respectively; T OLTC,t and T CB,t is the gear action value of OLTC and CBs; represents the deviation of the voltage of node i from the per-unit value at time t; κ1 and κ2 are the power loss coefficient and voltage deviation penalty coefficient in the reward function respectively; 3.
2. After formulating the MDP, train the DQN agent.
6. The voltage control method of distribution network with reconfigurable soft switches based on joint learning according to claim 5 is characterized in that: In step 3.2, the current DQN agent training steps are as follows: After initializing the experience replay pool D and the parameters θ of the DQN network, according to the current state Select Action Use the ε-greedy strategy to select actions, as shown in formula (15): In the formula, is the Q value network of DQN, θ = [W, b] is the parameter of the Q value network, including weights and biases; After executing the action, observe the reward r t d and the next state Experience Store them in the experience replay pool D; when the experience pool reaches a certain capacity, randomly sample a batch of experience of size N from it to update the network; The update of the Q-value network is achieved by minimizing the mean square error loss function. The loss function L(θ) of DQN is defined based on the mean square error (MSE) between the current Q-value and the target Q-value, as shown in formula (16): In the formula, for each experience, y i is the target Q value, is the current Q value, N is the number of samples in a small batch; The target Q value is estimated by the target Q network, and the calculation formula is shown in formula (17): In formula (17), y i is the target Q value; Q′ represents the target Q value network of DQN; r i d is the reward value of sampling experience; γ is the discount factor; is the maximum Q value of the next state, where is the state of the next moment of sampling experience, a d For sampling experience The action in the state, θ - is the target network parameter; For the target Q network of DQN, soft update is used to update it, as shown in formula (18): i - =τθ+(1-τ)θ - (18) Where τ is the soft update factor, θ - are the parameters of the target Q network; θ are the parameters of the Q network.
7. The voltage control method of distribution network with reconfigurable soft switches based on joint learning according to claim 1 is characterized in that: The step 4 specifically comprises the following steps: Step 4.1, build R-SOP operation constraints; Step 4.2: For voltage control with R-SOP, the state space S contains node loads and photovoltaic and day-ahead discrete device scheduling results; the action space A consists of active power transmission, reactive power support and feeder selection of R-SOP; the real-time distribution network operation optimization problem is transformed into a Markov decision process, represented by the tuple <S,A,P,R,γ>, where P represents the state transfer function, R is the reward obtained by the agent for executing the action, and γ is the discount factor.
8. The voltage control method of distribution network with reconfigurable soft switches based on joint learning according to claim 1 is characterized in that: In step 4.1, the R-SOP operation constraints are as follows: In equations (19) to (25), equations (19) to (21) are power balance constraints, where: is the active power of the R-SOP DC side connected to node i at time t; is the active power actually transmitted by the R-SOP accessed by node i at time t; is the active power loss of R-SOP connected to node i at time t; is the loss factor of R-SOP; Ω R-SOP It is a set of R-SOP ports; represents the power capacity injected into the node i connected to the R-SOP; is the reactive power injected by the voltage source converter at node i at time t; B i,n represents the switch state of node i on the nth branch connected to the R-SOP; Equations (22) to (25) are the R-SOP capacity constraints, where: It represents the power transmission capacity of the nth branch connected to R-SOP; Table 1 shows the capacity of the R-SOP accessed by node i; B is the reactive power actually transmitted by the R-SOP connected to node i at time t; n Indicates the switch status on the nth branch connected to the R-SOP; For VSC i Reactive power output limit; N s is the total number of branches connected to R-SOP; The environment of the DDPG and DQN joint agent is the distribution system power flow model in Equations (2) to (9), where Equations (5) and (6) are modified as follows to add the active transmission and reactive compensation of R-SOP, as shown in Equations (26) to (27): In formula (26)-formula (27), and is the active power and reactive power emitted by R-SOP on node i at time t.
9. The voltage control method of distribution network with reconfigurable soft switches based on joint learning according to claim 1, characterized in that: The step 4.2 specifically includes the following steps: 1) Construct the state space of the DDPG and DQN joint agent using formula (28): The state space consists of the injected active power and injected reactive power of each node in the distribution network, the active power injected by photovoltaics, and the discrete device actions that constitute the day before, which can be expressed as: In formula (28), and They represent the active load and reactive load of the node at time t respectively, Indicates the active power injected by the photovoltaic power station, tap t represents the tap position of the on-load tapchanger at time t, is the reactive power value of the capacitor bank switched on at time t; 2) Construct the R-SOP action space of the DDPG and DQN joint agent from formula (29): The action space is controlled by DDPG R-SOP (N S -1) Active power transmitted by each port, N S The reactive power provided by the ports and the N controlled by DQN S Feeder selection for each port: In formula (29), a t,DDPG and a t,DQN They represent the power output action and feeder selection action of the DDPG agent and the DQN agent at time t, respectively; and They represent the active transmission and reactive output of the R-SOP port at time t, represents the feeder selection of the R-SOP port at time t, and Ns represents the number of R-SOP ports in the distribution network; 3) Construct the reward function of the DDPG and DQN joint agent from formula (30): Because the deep reinforcement learning algorithm is a strategy designed to explore long-term reward maximization, when designing the reward function of the Markov decision process, the objective function is inverted to obtain: In formula (29), κ1 and κ2 represent the penalty coefficients of power loss and voltage, respectively.
10. The voltage control method of distribution network with reconfigurable soft switches based on joint learning according to claim 1, characterized in that: The step 5 specifically comprises the following steps: First, after the joint agent observes the input state, a deterministic strategy is obtained through the DDPG policy network; noise is added to construct the behavior strategy, and the DQN agent selects the action strategy through the ε-greedy strategy. The two actions together constitute the output strategy of R-SOP, as shown in Equation (31)-Equation (32): In formula (31)-formula (32) is the deterministic policy output by the DDPG policy network, where is the state at time t, θ u =[W,b] are the weights and biases in the DDPG policy network, N(0,σ t ) is a process of adding random quantities. The exploration of action values is completed by adding a noise sample to the network output value. The noise follows a normal distribution with a mean of zero and a standard deviation of σ t , parameter σ t The size of represents the degree of exploration and decreases with the decay rate during training; is the Q value network of DQN, θ = [W, b] is the weight and bias; Combine the actions output by the DDPG and DQN agents and calculate the reward value r of the strategy by interacting with the environment t r and the state at the next moment Will Stored in the experience replay pool; when the samples in the experience replay pool reach the upper limit, a random small batch is sampled from the experience replay pool in each iteration, and the DDPG agent calculates the target Q value through the target value network, as shown in formula (33): In formula (33), y DDPG is the target Q value of the DDPG network; is the state at the next moment; θ μ′ is the target strategy network parameter; θ Q′ is the target value network parameter; Q′ DDPG is the target value network of DDPG; is the strategy of the target value network; r t r represents the reward; γ is the discount factor; The update of the value network is achieved by minimizing the mean square error loss function: its loss function is expressed as minimizing the mean square error, as shown in formula (34): In formula (33), L DDPG (θ) is the loss function of the DDPG value network; Q DDPG is the value network about the state at time t and actions The output, θ Q is the value network parameter; For the policy network, its parameter θ μ The update of is based on the deterministic policy gradient, as shown in formula (35): In the formula, is the gradient after the policy network parameters are updated, s i is the sampling experience state; is the gradient of the value network parameters; is the gradient of the policy network; s r and a r are states and actions respectively, θ Q and θ μ are the parameters of the value network and the strategy network respectively; The target value network and target policy network are updated using a soft update mechanism, as shown in Equation (36)-Equation (37): i Q′ =tθ Q +(1-τ)θ Q′ (36) i μ′ =tθ μ +(1-τ)θ μ′ (37) In formula (36)-(37), τ is a constant, which indicates the update speed of the target network; θ Q′ and θ μ′ are the parameters of the target value network and the target strategy network respectively; θ Q and θ μ are the parameters of the value network and the strategy network; For the update of the DQN network, the target Q value is first calculated, the gap between the Q network output and the target Q value is calculated using the mean square error loss function, and the network parameters are updated using gradient descent, as shown in equations (38)-(39): Formula (38)-In formula (38), L DQN (θ) is the loss function of the DQN network; r t r is the reward value of sampling experience; γ is the discount factor; is the maximum Q value of the next state, where Q′ DQN represents the target Q value network of DQN, To sample the state of the next moment, For sampling experience The action in the state, θ - is the target Q network parameter; is the current Q value, where is the sampled experience state, For sampling experience The action in the state, θ is the Q network parameter; α is the learning rate; is the gradient of the Q network parameters; The parameters of the target Q network of DQN are θ - The soft update mechanism is also used for updating, as shown in formula (40): i - =τθ+(1-τ)θ - (40) Where, θ is the Q network parameter; θ - is the target Q network parameter; τ represents the update speed of the target network.
Citation Information
Cited By
Active power distribution network two-stage cloud edge cooperative scheduling method and device based on deep reinforcement learning
CN120933963A
Compressor energy-saving operation control method and system based on reinforcement learning
CN121165467A