A multi-timescale voltage optimization method for distribution networks based on deep reinforcement learning

Through the multi-time-scale voltage optimization method of deep reinforcement learning and the collaborative decision-making of DQN and DDPG intelligent agents, the voltage optimization problem in the distribution network with a high proportion of distributed photovoltaic access was solved, and the voltage stability and power quality were improved.

CN119482764BActive Publication Date: 2025-10-03STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411627973.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-10-03
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

Traditional voltage optimization methods are difficult to adapt to the dynamic changes of grid structure and operation mode in real time when a high proportion of distributed photovoltaics is connected to the distribution network, resulting in increased difficulty in voltage optimization and serious voltage over-limit problems.

Method used

A multi-time-scale voltage optimization method based on deep reinforcement learning is adopted. The DQN and DDPG agents are used to control the reactive output of long- and short-time-scale equipment respectively. Combined with the constraints of electrical nodes, action decision-making and feedback optimization are performed.

Benefits of technology

It effectively improves the voltage level, prevents voltage from exceeding the limit, optimizes voltage deviation, and improves power quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119482764B_ABST
    Figure CN119482764B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of distribution network voltage optimization and discloses a multi-time-scale voltage optimization method for distribution network based on deep reinforcement learning. The method comprises the following steps: dividing reactive regulation equipment into long-time-scale equipment and short-time-scale equipment, and using a DQN agent to train the action decisions of the long-time-scale equipment and a DDPG agent to train the action decisions of the short-time-scale equipment. Then, the decision-making strategy of the agent is adjusted according to the voltage optimization feedback of the distribution network, and finally the multi-time-scale equipment is collaboratively participated in the voltage optimization of the distribution network. The method realizes the coordinated output and voltage optimization of equipment actions at different time scales. At the same time, on the long time scale, the voltage can be effectively improved to prevent voltage over-limit. On the short time scale, the voltage fluctuation range is also controlled, thereby effectively optimizing the voltage deviation and improving the power quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of distribution network voltage optimization, and in particular to a distribution network multi-timescale voltage optimization method based on deep reinforcement learning. Background Art

[0002] With the high penetration of distributed photovoltaic power generation, the distribution network has shifted from a single-power flow system to a multi-terminal, bidirectional power supply system. The resulting bidirectional power flow and voltage over-limit issues have become increasingly serious. The rational scheduling of diverse power system resources can help mitigate the impact of the randomness and intermittent nature of distributed photovoltaic power generation on the distribution network. By regulating reactive power to assist in regulating voltage levels, it optimizes power flows and effectively reduces energy losses during transmission and distribution.

[0003] Common approaches to optimizing reactive power and voltage in distribution networks include centralized control, local control, and distributed control. However, integrating a high proportion of photovoltaic power into distribution networks significantly increases the difficulty of optimizing traditional methods. The large number of power system resources with varying adjustment timescales exponentially increases the number of optimization constraints. Traditional optimization methods, due to their overreliance on predictions and the limitations of model accuracy, lack real-time optimization capabilities and are difficult to adapt to dynamic changes in grid structure and operation. Summary of the Invention

[0004] In response to the above-mentioned deficiencies in the prior art, the present invention provides a multi-time-scale voltage optimization method for distribution networks based on deep reinforcement learning, which is used to solve the problem of voltage optimization difficulty caused by traditional voltage optimization methods when dealing with a high proportion of distributed photovoltaic access to distribution networks.

[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is:

[0006] A multi-timescale voltage optimization method for distribution network based on deep reinforcement learning, comprising the following steps:

[0007] S1. Based on the device regulation characteristics, the reactive regulation devices involved in the distribution network voltage optimization are divided into long-time scale devices and short-time scale devices. At the same time, DQN agents and DDPG agents are set up to control the reactive output of long-time scale devices and short-time scale devices respectively.

[0008] S2. Establish constraints for each electrical node in the distribution network, and use the DQN agent to obtain the distribution network status of long-time scale devices and the DDPG agent to obtain the distribution network status of short-time scale devices.

[0009] S3. Set the switching decision time for long-time scale devices and short-time scale devices, and obtain the action decision of the DQN agent on the long-time scale devices and the action decision of the DDPG agent on the short-time scale devices at their corresponding switching decision times according to the distribution network status of the long-time scale devices and the distribution network status of the short-time scale devices;

[0010] S4. Based on the voltage over-limit and voltage offset conditions of the distribution network, feedback is provided to the action decisions of the long-time scale devices and the short-time scale devices, and the distribution network feedback of the DQN agent to the long-time scale devices and the distribution network feedback of the DDPG agent to the short-time scale devices are obtained respectively;

[0011] S5. At the time of the long-time-scale device switching decision, the network parameters of the DQN network are updated according to the distribution network status, action decision, and distribution network feedback of the long-time-scale device, thereby obtaining a trained DQN network, i.e., the optimal DQN agent, which is used to output the optimal action decision of the long-time-scale device;

[0012] S6. At the switching decision time of the short-time-scale device, the network parameters of the DDPG network are updated according to the distribution network status, action decision and distribution network feedback of the short-time-scale device, and a trained DDPG network, i.e., the optimal DDPG agent, is obtained to output the optimal action decision of the short-time-scale device;

[0013] S7. Put the optimal DQN agent and the optimal DDPG agent into the distribution network to obtain the optimal action decision of the reactive power regulation equipment.

[0014] Furthermore, the constraints of each electrical node in the distribution network in step S2 are:

[0015]

[0016] Among them, P i , Q i They represent the active power and reactive power of electrical node i, n represents the total number of electrical nodes connected to electrical node i, v i 、v j Represent the voltage of electrical node i and electrical node j respectively, G ij represents the conductance of the line from electrical node i to electrical node j, B ij represents the susceptance of the line from electrical node i to electrical node j, cos represents the cosine function, sin represents the sine function, δ ij The voltage phase difference between the two ends of the line from electrical node i to electrical node j, v i,min 、v i,maxThey represent the lower and upper limits of the voltage of the electrical node i, respectively. v(i,t) represents the voltage of the electrical node i at the tth moment. represents the reactive output of the switched capacitor SCB at electrical node i, Indicates the reactive output corresponding to a single gear of the switching capacitor SCB, Indicates the capacitor level that electrical node i decides to switch at time t, They represent the active power output, reactive power output and capacity of the photovoltaic PV at the electrical node i at the tth moment, It represents the upper limit of the active power output of photovoltaic PV at electrical node i at time t, min represents the minimum value, and max represents the maximum value.

[0017] Furthermore, in step S2, the DQN agent is used to obtain the distribution network state of the long-term equipment:

[0018]

[0019] in, P represents the state of the power distribution network of the long-term equipment obtained by the DQN agent at time t, slack,t , Q slack,t They represent the active power and reactive power of the slack node at time t, respectively. load,t , Q load,t They represent the active power and reactive power of the loads of the electrical nodes except the relaxed nodes at time t, P′ SCB,t , Q′ SCB,t They represent the active output and reactive output of the switched capacitor SCB at time t respectively;

[0020] Among them, the slack node refers to the electrical node connected to the upper-level distribution network.

[0021] Furthermore, in step S2, the distribution network status of the short-time-scale equipment is obtained using the DDPG agent:

[0022]

[0023] in, P′ represents the distribution network state of the short-time scale equipment obtained by the DDPG agent at time t. PV,t , Q′ PV,t They represent the active power output and reactive power output of photovoltaic PV at the t moment respectively, Indicates the maximum active power output of photovoltaic PV.

[0024] Furthermore, in step S3, the DQN agent's action decision for the long-time scale device is:

[0025]

[0026] in, Represents the action decision of the DQN agent on the long-time scale device at time t.

[0027] Furthermore, in step S3, the DDPG agent's action decision for the short-time-scale device is:

[0028]

[0029] in, Represents the action decision of the DDPG agent on the short-time-scale device at time t.

[0030] Furthermore, in step S4, the DQN agent’s feedback to the distribution network of the long-time scale device is:

[0031]

[0032] in, It represents the distribution network feedback of the DQN agent to the long-time scale equipment at time t, and N represents the total number of electrical nodes in the distribution network.

[0033] Furthermore, in step S4, the DDPG agent’s feedback to the distribution network of the short-time-scale device is:

[0034]

[0035] in, represents the distribution network feedback of the DDPG agent to the short-time-scale equipment at time t, v i,base Represents the nominal voltage at electrical node i.

[0036] Furthermore, step S5 specifically includes:

[0037] S51. At the switching decision time of the long-term scale device, the distribution network state, action decision, and distribution network feedback of the long-term scale device are input into the distribution network environment, and the device is transferred to the next state at the same time.

[0038] S52. When updating the network parameters of the DQN network, the distribution network state, action decision, and distribution network feedback of the long-time-scale device are input into the DQN network, and the total loss function of the DQN network is updated using the gradient descent method. When the number of training times reaches the maximum value, a trained DQN network is obtained, i.e., the optimal DQN agent, which is used to output the optimal action decision of the long-time-scale device;

[0039] Among them, the total loss function of the DQN network is:

[0040]

[0041] Where L represents the total loss function of the DQN network, θ represents the network parameters, and θ t represents the network parameters of the DQN network at time t, θ t+1 represents the network parameters of the DQN network at time t+1, E represents the mean calculation, r DQN represents the distribution network feedback of long-time scale devices in the DQN network, γ DQN represents the conversion coefficient of the distribution network feedback of long-time scale devices in the DQN network, Q DQN represents the value evaluation value of the DQN network, a represents the action decision of the long-time scale device, s represents the distribution network state of the long-time scale device, s′ represents the distribution network state after taking the action decision a under the distribution network state s of the long-time scale device, a′ represents the action decision of the distribution network state s′, α represents the learning rate, Represents the gradient.

[0042] Furthermore, step S6 specifically includes:

[0043] S61. At the switching decision time of the short-time-scale device, the distribution network state, action decision, and distribution network feedback of the short-time-scale device are input into the distribution network environment, and the device is transferred to the next state at the same time.

[0044] S62, inputting the action decision of the current state into the value evaluation neural network of the DDPG network, and updating the total loss function of the value evaluation neural network using the gradient descent method to obtain the updated network parameters of the value evaluation neural network;

[0045] S63. Update the action neural network of the DDPG network based on the network parameters of the updated value evaluation neural network. Input the distribution network status, action decision, and distribution network feedback of the short-time-scale device into the action neural network, and use the gradient descent method to update the total loss function of the action neural network. When the number of training times reaches the maximum, the trained action neural network and value evaluation neural network are obtained, that is, the optimal DDPG agent is used to output the optimal action decision of the short-time-scale device.

[0046] Among them, the total loss function of the value evaluation neural network is:

[0047]

[0048] The total loss function of the action neural network is:

[0049]

[0050] Where L′ represents the total loss function of the value assessment neural network, θ Q1 represents the network parameters of the value assessment neural network Q1, represents the network parameters of the value evaluation neural network Q1 at time t-1, represents the network parameters of the value evaluation neural network Q1 at time t, r DDPG represents the distribution network feedback of short-time-scale devices in the DDPG network, γ DDPG represents the reduction factor of the distribution network feedback of short-time-scale devices in the DDPG network, Q Q1 represents the value evaluation value of the value evaluation neural network Q1, s1 represents the distribution network state of the short-time scale device, a1 represents the action decision of the short-time scale device, s1′ represents the distribution network state after the action decision a1 is adopted under the distribution network state s1 of the short-time scale device, a1′ represents the action decision of the distribution network state s1′, α Q1 represents the learning rate of the value evaluation neural network Q1, J represents the total loss function of the action neural network, θ π represents the network parameters of the action neural network π, represents the network parameters of the action neural network π at time t-1, represents the network parameters of the action neural network π at time t, α π represents the learning rate of the action neural network π.

[0051] The present invention has the following beneficial effects:

[0052] The present invention proposes a multi-time-scale voltage optimization method for distribution networks based on deep reinforcement learning. By considering the regulation characteristics of different reactive devices, a two-layer deep learning method of DQN and DDPG is adopted. Through the mutual collaborative decision-making of two intelligent agents, the output and voltage optimization of equipment actions at different time scales are achieved. At the same time, on a long time scale, the voltage can be effectively improved to prevent voltage exceeding the limit. On a short time scale, the voltage fluctuation range is controlled, which effectively optimizes the voltage deviation and improves the power quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is a flow chart of a multi-time-scale voltage optimization method for distribution networks based on deep reinforcement learning proposed in the present invention;

[0054] Figure 2 Schematic diagram of the structure of the improved distribution network system based on the IEEE33 node calculation example in the embodiment;

[0055] Figure 3 This is a schematic diagram of the convergence of intelligent agent training in the embodiment;

[0056] Figure 4 A schematic diagram of an action decision of a discrete capacitor made by a DQN agent in an embodiment;

[0057] Figure 5 A schematic diagram of photovoltaic action decisions made by the DDPG agent in an embodiment;

[0058] Figure 6 Schematic diagram of the box lines before and after the voltage optimization of all nodes on the test day in the embodiment;

[0059] Figure 7 Schematic diagram of the voltage boost effect of the terminal node in the embodiment;

[0060] Figure 8 This is a schematic diagram of the voltage optimization effect at 14:00 on the test day in the embodiment. DETAILED DESCRIPTION

[0061] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0062] like Figure 1 As shown, a multi-time-scale voltage optimization method for distribution network based on deep reinforcement learning includes the following steps S1-S7:

[0063] S1. Based on the equipment regulation characteristics, the reactive regulation equipment involved in the distribution network voltage optimization is divided into long-time-scale equipment and short-time-scale equipment. At the same time, DQN agents and DDPG agents are set up to control the reactive output of long-time-scale equipment and short-time-scale equipment respectively.

[0064] In this embodiment, the reactive power regulation devices involved in the distribution network voltage optimization are classified according to their regulation characteristics. That is, discrete switching capacitors are defined as long-time scale devices, and photovoltaic devices are defined as short-time scale devices. For different types of devices, the gears of discrete switching capacitors are controlled by a DQN agent, and the reactive power output of photovoltaic devices is controlled by a DDPG agent.

[0065] In addition, the distribution network in this embodiment is an improved distribution network system based on the IEEE33 node calculation example, as shown in the following example: Figure 2 As shown, it includes distributed photovoltaics connected to electrical nodes 16, 18, 22, 24, 25, and 31, where each distributed photovoltaic capacity is 0.6MW. It also includes discrete switching capacitors connected to electrical nodes 6, 8, 18, and 25, with a total of four adjustable gears, each with a capacity of 0.2MW, and the gear changes no more than 6 times a day.

[0066] S2. Establish constraints for each electrical node in the distribution network. At the same time, use the DQN agent to obtain the distribution network status of long-time scale equipment and use the DDPG agent to obtain the distribution network status of short-time scale equipment.

[0067] In this example, the agents are switched to training mode, simulating the daily operation of a distribution network to train the action decisions of the two agents and collect state information for each electrical node in the distribution network. To simulate the distribution network's operation, the network must be modeled. By establishing constraints for each electrical node in the distribution network, the DQN agent can obtain the network state for long-term devices, while the DDPG agent can obtain the network state for short-term devices.

[0068] Specifically, the constraints of each electrical node in the distribution network in step S2 are:

[0069]

[0070] Among them, P i , Q i They represent the active power and reactive power of electrical node i, n represents the total number of electrical nodes connected to electrical node i, v i 、v j Represent the voltage of electrical node i and electrical node j respectively, G ij represents the conductance of the line from electrical node i to electrical node j, B ij represents the susceptance of the line from electrical node i to electrical node j, cos represents the cosine function, sin represents the sine function, δ ij The voltage phase difference between the two ends of the line from electrical node i to electrical node j, v i,min 、v i,max They represent the lower and upper limits of the voltage of the electrical node i, respectively. v(i,t) represents the voltage of the electrical node i at the tth moment. represents the reactive output of the switched capacitor SCB at electrical node i, Indicates the reactive output corresponding to a single gear of the switching capacitor SCB, Indicates the capacitor level that electrical node i decides to switch at time t, They represent the active power output, reactive power output and capacity of the photovoltaic PV at the electrical node i at the tth moment, It represents the upper limit of the active power output of photovoltaic PV at electrical node i at time t, min represents the minimum value, and max represents the maximum value.

[0071] Specifically, in step S2, the DQN agent is used to obtain the distribution network status of the long-term equipment:

[0072]

[0073] in, P represents the state of the power distribution network of the long-term equipment obtained by the DQN agent at time t, slack,t , Q slack,t They represent the active power and reactive power of the slack node at time t, respectively. load,t , Q load,t They represent the active power and reactive power of the loads of the electrical nodes except the relaxed nodes at time t, P′ SCB,t , Q′ SCB,t They represent the active output and reactive output of the switched capacitor SCB at time t respectively.

[0074] Among them, the slack node refers to the electrical node connected to the upper-level distribution network.

[0075] Specifically, in step S2, the DDPG agent is used to obtain the distribution network status of the short-time scale equipment:

[0076]

[0077] in, P′ represents the distribution network state of the short-time scale equipment obtained by the DDPG agent at time t. PV,t , Q′ PV,t They represent the active power output and reactive power output of photovoltaic PV at the t moment respectively, Indicates the maximum active power output of photovoltaic PV.

[0078] S3. Set the switching decision time for long-time scale devices and short-time scale devices, and according to the distribution network status of long-time scale devices and short-time scale devices, obtain the action decision of the DQN agent on the long-time scale devices and the action decision of the DDPG agent on the short-time scale devices at their respective corresponding switching decision times.

[0079] In this embodiment, the reactive power regulation equipment that needs to be decided is judged according to time, and the action decision of the reactive power regulation equipment is made by inputting the distribution network status information into the neural network of the corresponding intelligent agent. Among them, the DQN intelligent agent is set to make a switching decision every 6 hours to optimize the voltage limit, and the DDPG intelligent agent makes a photovoltaic reactive power processing decision every 15 minutes to reduce voltage deviation and suppress voltage fluctuation.

[0080] In addition, the switching decision time is set according to the action time of the equipment. That is, the long-time scale equipment is a discrete switching equipment, which is generally set to make a decision every 6 hours; the short-time scale equipment is photovoltaic, which is generally set to make a decision every 15 minutes.

[0081] Specifically, in step S3, the DQN agent's action decision for the long-time scale device is:

[0082]

[0083] in, Represents the action decision of the DQN agent on the long-time scale device at time t.

[0084] Specifically, in step S3, the DDPG agent's action decision for the short-time-scale device is:

[0085]

[0086] in, Represents the action decision of the DDPG agent on the short-time-scale device at time t.

[0087] S4. Based on the voltage over-limit and voltage offset conditions of the distribution network, feedback is provided to the action decisions of the long-time scale devices and the short-time scale devices, and the distribution network feedback of the DQN agent to the long-time scale devices and the distribution network feedback of the DDPG agent to the short-time scale devices are obtained respectively.

[0088] In this embodiment, based on the voltage over-limit and voltage offset conditions of the distribution network, the action reward of the intelligent agent is fed back to the intelligent agent for updating the neural network, that is, updating the DQN network and DDPG network, where the DQN network refers to the deep Q neural network and the DDPG network refers to the action neural network and value evaluation neural network.

[0089] Specifically, in step S4, the DQN agent’s feedback to the distribution network of the long-time scale device is:

[0090]

[0091] in, It represents the distribution network feedback of the DQN agent to the long-time scale equipment at time t, and N represents the total number of electrical nodes in the distribution network.

[0092] In this embodiment, the distribution network feedback of the long-time scale device is the action reward received by the DQN agent.

[0093] Specifically, in step S4, the DDPG agent’s feedback to the distribution network of the short-time-scale device is:

[0094]

[0095] in, represents the distribution network feedback of the DDPG agent to the short-time-scale equipment at time t, v i,base Represents the nominal voltage at electrical node i.

[0096] In this embodiment, the distribution network feedback of the short-time-scale device is the action reward received by the DDPG agent.

[0097] S5. At the time of switching decisions for long-time-scale devices, the network parameters of the DQN network are updated according to the distribution network status, action decisions, and distribution network feedback of the long-time-scale devices to obtain a trained DQN network, i.e., the optimal DQN agent, which is used to output the optimal action decisions for long-time-scale devices.

[0098] In this embodiment, when performing long-time-scale voltage optimization Markov decision-making, the DQN agent inputs the action into the distribution network environment. The environment records the distribution network state, action decision, and distribution network feedback of the long-time-scale equipment under this action, and transmits them to the experience replay pool for storage while transferring to the next state. Each time the neural network parameters are updated, random sampling is performed from the experience replay pool, and stochastic gradient descent is used to update the agent network according to the total loss function, so that it can make better action decisions.

[0099] Specifically, step S5 includes S51-S52:

[0100] S51. At the switching decision time of the long-term scale device, the distribution network state, action decision and distribution network feedback of the long-term scale device are input into the distribution network environment, and the device is transferred to the next state at the same time.

[0101] S52. When updating the network parameters of the DQN network, the distribution network status, action decision and distribution network feedback of the long-time scale device are input into the DQN network, and the total loss function of the DQN network is updated using the gradient descent method. When the number of training times reaches the maximum value, a trained DQN network is obtained, that is, the optimal DQN agent, which is used to output the optimal action decision of the long-time scale device.

[0102] Among them, the total loss function of the DQN network is:

[0103]

[0104] Where L represents the total loss function of the DQN network, θ represents the network parameters, and θ t represents the network parameters of the DQN network at time t, θ t+1 represents the network parameters of the DQN network at time t+1, E represents the mean calculation, r DQN represents the distribution network feedback of long-time scale devices in the DQN network, γ DQN represents the conversion coefficient of the distribution network feedback of long-time scale devices in the DQN network, Q DQNrepresents the value evaluation value of the DQN network, a represents the action decision of the long-time scale device, s represents the distribution network state of the long-time scale device, s′ represents the distribution network state after taking the action decision a under the distribution network state s of the long-time scale device, a′ represents the action decision of the distribution network state s′, α represents the learning rate, Represents the gradient.

[0105] S6. At the switching decision time of the short-time-scale device, the network parameters of the DDPG network are updated according to the distribution network status, action decision and distribution network feedback of the short-time-scale device to obtain a trained DDPG network, that is, the optimal DDPG agent, which is used to output the optimal action decision of the short-time-scale device.

[0106] In this embodiment, when performing short-time-scale voltage optimization Markov decision-making, the DDPG agent maintains two neural networks, namely the action neural network and the value evaluation neural network, which respectively perform action output and action evaluation; when the action neural network outputs the agent action to the environment, the environment records the distribution network state, action decision and distribution network feedback under this action, and transmits it to the experience replay pool for storage, and transfers it to the next state at the same time; at the same time, when the DDPG agent is updated, the value evaluation neural network is updated first according to the data in the experience replay pool, and guides the action neural network to update, so that the DDPG agent outputs better action decisions.

[0107] Specifically, step S6 includes S61-S63:

[0108] S61. At the switching decision time of the short-time-scale device, the distribution network state, action decision and distribution network feedback of the short-time-scale device are input into the distribution network environment, and the device is transferred to the next state at the same time.

[0109] S62. Input the action decision of the current state into the value evaluation neural network of the DDPG network, and use the gradient descent method to update the total loss function of the value evaluation neural network to obtain the updated network parameters of the value evaluation neural network.

[0110] S63. Update the action neural network of the DDPG network based on the network parameters of the updated value evaluation neural network. Input the distribution network status, action decision and distribution network feedback of the short-time-scale equipment into the action neural network, and use the gradient descent method to update the total loss function of the action neural network. When the number of training times reaches the maximum value, the trained action neural network and value evaluation neural network, i.e., the optimal DDPG agent, are obtained, which is used to output the optimal action decision of the short-time-scale equipment.

[0111] Among them, the total loss function of the value evaluation neural network is:

[0112]

[0113] The total loss function of the action neural network is:

[0114]

[0115] Where L′ represents the total loss function of the value assessment neural network, θ Q1 represents the network parameters of the value assessment neural network Q1, represents the network parameters of the value evaluation neural network Q1 at time t-1, represents the network parameters of the value evaluation neural network Q1 at time t, r DDPG represents the distribution network feedback of short-time-scale devices in the DDPG network, γ DDPG represents the reduction factor of the distribution network feedback of short-time-scale devices in the DDPG network, Q Q1 represents the value evaluation value of the value evaluation neural network Q1, s1 represents the distribution network state of the short-time scale device, a1 represents the action decision of the short-time scale device, s1′ represents the distribution network state after the action decision a1 is adopted under the distribution network state s1 of the short-time scale device, a1′ represents the action decision of the distribution network state s1′, α Q1 represents the learning rate of the value evaluation neural network Q1, J represents the total loss function of the action neural network, θ π represents the network parameters of the action neural network π, represents the network parameters of the action neural network π at time t-1, represents the network parameters of the action neural network π at time t, α π represents the learning rate of the action neural network π.

[0116] like Figure 3 As shown, Figure 3 The results demonstrate the convergence of rewards during agent training. During training, the DQN and DDPG agents collaborated to make decisions about the actions of reactive power regulation equipment in the distribution network, with both control objectives including voltage optimization. Consequently, the rewards for both agents converged in the same direction. Furthermore, between rounds 0 and 60, due to the agents' exploration phase, there was no clear convergence trend. However, after 70 rounds of training, the rewards for both agents converged.

[0117] S7. Put the optimal DQN agent and the optimal DDPG agent into the distribution network to obtain the optimal action decision of the reactive power regulation equipment.

[0118] In this embodiment, it is determined whether the distribution network simulation operation time has reached the set training days requirement. If it has, the agent is switched to test mode and put into the distribution network to receive the distribution network status and directly make reactive power regulation equipment action decisions. Specifically, one day is randomly selected to test the two trained agents (i.e., the optimal DQN agent and the optimal DDPG agent). The agent makes reactive power regulation equipment action decisions by receiving the load and voltage conditions of the distribution network; Figure 4-Figure 5 As shown, Figure 4 and Figure 5 The action decisions made by the DQN and DDPG agents for discrete capacitors and photovoltaics. As can be seen from the figure, during the peak load phase, the voltage of each electrical node drops significantly. Therefore, the DQN agent increases the discrete capacitor gear to enhance the reactive power support of the grid. The DDPG agent effectively raises the voltage to prevent over-limit by increasing the reactive power output of the photovoltaics, and effectively reduces the voltage deviation.

[0119] In addition, if Figure 6 As shown, Figure 6 The box plots are of all electrical nodes before and after voltage optimization on the test day. Before optimization, there were voltage over-limit situations at all electrical nodes from 6:00 to 20:00. At 17:00, electrical node 18 experienced serious voltage over-limit, with the lowest voltage being 0.89. However, after voltage optimization was performed using the multi-time-scale voltage optimization method for distribution networks based on deep reinforcement learning proposed in this invention, the distribution network experienced no voltage over-limit situations throughout the day, with the highest voltage being 1.02 and the lowest being 0.96. At the same time, the terminal node 33 was selected to track its voltage changes throughout the day, as shown in the following figure. Figure 7 As shown by Figure 7 It can be seen that this method raises the minimum voltage of the terminal node from 0.95 to 0.98, which improves the power quality. Secondly, the voltage exceeding the limit seriously at 14:00 is selected, and the voltage of 33 nodes before and after optimization is measured, as shown in the figure below. Figure 8 As shown in the figure, the intelligent agent increases the reactive output of the voltage regulating equipment near the out-of-limit nodes, such as nodes 16, 24, and 25, thereby effectively raising the voltage of the out-of-limit nodes, reducing the overall voltage deviation of the power grid, and improving the power quality.

[0120] In summary, the present invention proposes a multi-time-scale voltage optimization method for distribution network based on deep reinforcement learning. By considering the regulation characteristics of different reactive devices, based on deep learning, and through the mutual collaborative decision-making of two intelligent agents, the coordinated output of equipment actions at different time scales is realized, which effectively avoids voltage over-limit and improves the power quality of the terminal node. Specifically: (1) The double-layer deep learning method of DQN and DDPG is used to effectively learn the reactive regulation strategy of the distribution network, and the corresponding voltage optimization goal can be achieved without the need for an accurate physical model of the distribution network; (2) On a long time scale, the voltage can be effectively improved to prevent the voltage from over-limit; (3) On a short time scale, the voltage fluctuation range is also controlled, which effectively optimizes the voltage deviation and improves the power quality.

[0121] Specific embodiments are used in the present invention to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.

[0122] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific descriptions and embodiments. Those skilled in the art can make various other specific variations and combinations based on the technical teachings disclosed in the present invention without departing from the essence of the present invention, and such variations and combinations are still within the scope of protection of the present invention.

Claims

1. A multi-timescale voltage optimization method for distribution network based on deep reinforcement learning, characterized in that: The following steps are involved: S1. Based on the device regulation characteristics, the reactive regulation devices involved in the distribution network voltage optimization are divided into long-time scale devices and short-time scale devices. At the same time, DQN agents and DDPG agents are set up to control the reactive output of long-time scale devices and short-time scale devices respectively. S2. Establish constraints for each electrical node in the distribution network, and use the DQN agent to obtain the distribution network status of long-time scale devices and the DDPG agent to obtain the distribution network status of short-time scale devices. S3. Set the switching decision time for long-time scale devices and short-time scale devices, and obtain the action decision of the DQN agent on the long-time scale devices and the action decision of the DDPG agent on the short-time scale devices at their corresponding switching decision times according to the distribution network status of the long-time scale devices and the distribution network status of the short-time scale devices; S4. Based on the voltage over-limit and voltage offset conditions of the distribution network, feedback is provided to the action decisions of the long-time scale devices and the short-time scale devices, and the distribution network feedback of the DQN agent to the long-time scale devices and the distribution network feedback of the DDPG agent to the short-time scale devices are obtained respectively; S5. At the time of the long-time-scale device switching decision, the network parameters of the DQN network are updated according to the distribution network status, action decision, and distribution network feedback of the long-time-scale device, thereby obtaining a trained DQN network, i.e., the optimal DQN agent, which is used to output the optimal action decision of the long-time-scale device; S6. At the switching decision time of the short-time-scale device, the network parameters of the DDPG network are updated according to the distribution network status, action decision and distribution network feedback of the short-time-scale device, and a trained DDPG network, i.e., the optimal DDPG agent, is obtained to output the optimal action decision of the short-time-scale device; S7. Put the optimal DQN agent and the optimal DDPG agent into the distribution network to obtain the optimal action decision of the reactive power regulation equipment.

2. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 1 is characterized in that: The constraints of each electrical node in the distribution network in step S2 are: Among them, P i , Q i They represent the active power and reactive power of electrical node i, n represents the total number of electrical nodes connected to electrical node i, v i 、v j Represent the voltage of electrical node i and electrical node j respectively, G ij represents the conductance of the line from electrical node i to electrical node j, B ij represents the susceptance of the line from electrical node i to electrical node j, cos represents the cosine function, sin represents the sine function, δ ij The voltage phase difference between the two ends of the line from electrical node i to electrical node j, v i,min 、v i,max They represent the lower and upper limits of the voltage of the electrical node i, respectively. v(i,t) represents the voltage of the electrical node i at the tth moment. represents the reactive output of the switched capacitor SCB at electrical node i, Indicates the reactive output corresponding to a single gear of the switching capacitor SCB, Indicates the capacitor level that electrical node i decides to switch at time t, They represent the active power output, reactive power output and capacity of the photovoltaic PV at the electrical node i at the tth moment, It represents the upper limit of the active power output of photovoltaic PV at electrical node i at time t, min represents the minimum value, and max represents the maximum value.

3. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 2 is characterized in that: In step S2, the DQN agent is used to obtain the distribution network status of the long-term equipment: in, P represents the state of the power distribution network of the long-term equipment obtained by the DQN agent at time t, slack,t , Q slack,t They represent the active power and reactive power of the slack node at time t, respectively. load,t , Q load,t They represent the active power and reactive power of the loads of the electrical nodes except the relaxed nodes at time t, P′ SCB,t , Q′ SCB,t They represent the active output and reactive output of the switched capacitor SCB at time t respectively; Among them, the slack node refers to the electrical node connected to the upper-level distribution network.

4. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 3 is characterized in that: In step S2, the DDPG agent is used to obtain the distribution network status of short-time scale equipment: in, P′ represents the distribution network state of the short-time scale equipment obtained by the DDPG agent at time t. PV,t , Q′ PV,t They represent the active power output and reactive power output of photovoltaic PV at the t moment respectively, Indicates the maximum active power output of photovoltaic PV.

5. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 4 is characterized in that: In step S3, the DQN agent's action decision for the long-time scale device is: in, Represents the action decision of the DQN agent on the long-time scale device at time t.

6. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 5, characterized in that: In step S3, the DDPG agent's action decision for the short-time-scale device is: in, Represents the action decision of the DDPG agent on the short-time-scale device at time t.

7. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 6, characterized in that: In step S4, the DQN agent’s feedback to the distribution network of the long-time scale equipment is: in, It represents the distribution network feedback of the DQN agent to the long-time scale equipment at time t, and N represents the total number of electrical nodes in the distribution network.

8. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 7, characterized in that: In step S4, the DDPG agent's feedback to the distribution network of the short-time scale equipment is: in, represents the distribution network feedback of the DDPG agent to the short-time-scale equipment at time t, v i,base Represents the nominal voltage at electrical node i.

9. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 8, characterized in that: Step S5 specifically includes: S51. At the switching decision time of the long-term scale device, the distribution network state, action decision, and distribution network feedback of the long-term scale device are input into the distribution network environment, and the device is transferred to the next state at the same time. S52. When updating the network parameters of the DQN network, the distribution network state, action decision, and distribution network feedback of the long-time-scale device are input into the DQN network, and the total loss function of the DQN network is updated using the gradient descent method. When the number of training times reaches the maximum value, a trained DQN network is obtained, i.e., the optimal DQN agent, which is used to output the optimal action decision of the long-time-scale device; Among them, the total loss function of the DQN network is: Where L represents the total loss function of the DQN network, θ represents the network parameters, and θ t represents the network parameters of the DQN network at time t, θ t+1 represents the network parameters of the DQN network at time t+1, E represents the mean calculation, r DQN represents the distribution network feedback of long-time scale devices in the DQN network, γ DQN represents the conversion coefficient of the distribution network feedback of long-time scale devices in the DQN network, Q DQN represents the value evaluation value of the DQN network, a represents the action decision of the long-time scale device, s represents the distribution network state of the long-time scale device, s′ represents the distribution network state after taking the action decision a under the distribution network state s of the long-time scale device, a′ represents the action decision of the distribution network state s′, α represents the learning rate, Represents the gradient.

10. The method for multi-time-scale voltage optimization of distribution network based on deep reinforcement learning according to claim 9, characterized in that: Step S6 specifically includes: S61. At the switching decision time of the short-time-scale device, the distribution network state, action decision, and distribution network feedback of the short-time-scale device are input into the distribution network environment, and the device is transferred to the next state at the same time. S62, inputting the action decision of the current state into the value evaluation neural network of the DDPG network, and updating the total loss function of the value evaluation neural network using the gradient descent method to obtain the updated network parameters of the value evaluation neural network; S63. Update the action neural network of the DDPG network based on the network parameters of the updated value evaluation neural network. Input the distribution network status, action decision, and distribution network feedback of the short-time-scale device into the action neural network, and use the gradient descent method to update the total loss function of the action neural network. When the number of training times reaches the maximum, the trained action neural network and value evaluation neural network are obtained, that is, the optimal DDPG agent is used to output the optimal action decision of the short-time-scale device. Among them, the total loss function of the value evaluation neural network is: The total loss function of the action neural network is: Where L′ represents the total loss function of the value assessment neural network, θ Q1 represents the network parameters of the value assessment neural network Q1, represents the network parameters of the value evaluation neural network Q1 at time t-1, represents the network parameters of the value evaluation neural network Q1 at time t, r DDPG represents the distribution network feedback of short-time-scale devices in the DDPG network, γ DDPG represents the reduction factor of the distribution network feedback of short-time-scale devices in the DDPG network, Q Q1 represents the value evaluation value of the value evaluation neural network Q1, s1 represents the distribution network state of the short-time scale device, a1 represents the action decision of the short-time scale device, s1′ represents the distribution network state after the action decision a1 is adopted under the distribution network state s1 of the short-time scale device, a1′ represents the action decision of the distribution network state s1′, α Q1 represents the learning rate of the value evaluation neural network Q1, J represents the total loss function of the action neural network, θ π represents the network parameters of the action neural network π, represents the network parameters of the action neural network π at time t-1, represents the network parameters of the action neural network π at time t, α π represents the learning rate of the action neural network π.

Citation Information

Patent Citations

  • Reactive voltage control method based on multi-time-scale multi-agent deep reinforcement learning

    CN113363997A

  • Comprehensive energy system voltage control method and system based on deep reinforcement learning algorithm

    CN116388280A