Power distribution network voltage regulation method based on deep reinforcement learning of light storage cooperative operation

By using a secure multi-agent algorithm based on deep reinforcement learning to coordinate photovoltaic inverters and energy storage devices, the problem of unstable grid voltage in high-penetration photovoltaic power generation scenarios has been solved. This has enabled efficient, safe, and flexible control of voltage regulation, thereby improving grid stability and photovoltaic absorption capacity.

CN120222494BActive Publication Date: 2026-03-17ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing voltage regulation methods in high-penetration photovoltaic power generation scenarios suffer from problems such as large computational load, high operating cost, reliance on complex accurate physical models, and difficulty in exploring intelligent agents in high-dimensional state spaces, leading to insufficient grid voltage instability and security. In particular, there is little research on voltage regulation in photovoltaic-storage collaborative operation.

Method used

A secure multi-agent algorithm based on deep reinforcement learning is adopted. Through the coordinated control of photovoltaic inverters and energy storage devices, an incremental voltage control model and active power balance rules are established. By utilizing shared state space strategies and deep deterministic strategy gradient algorithms, the actions of photovoltaic inverters and energy storage devices are optimized to achieve voltage regulation and improve system stability.

Benefits of technology

It improves the efficiency of coordinated optimization and control of photovoltaic inverters and energy storage devices, enhances the adaptability and flexibility of the distribution network, reduces the risk of violation of intelligent agent action constraints, improves the security and voltage stability of the power grid, and enhances the photovoltaic absorption capacity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120222494B_ABST
    Figure CN120222494B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on deep reinforcement learning's distribution network light storage collaborative operation voltage regulation method, comprising the following steps: selecting the distribution network to be carried out voltage regulation, establishes the incremental voltage control model of distribution network;Establish the control model of photovoltaic inverter in distribution network, voltage regulation is carried out by security constraint deep reinforcement learning algorithm;Establish the control model of energy storage equipment in distribution network, control the state and output power of the charge or discharge of energy storage equipment;By changing reward function, reduce voltage fluctuation peak;Establish distribution network light storage collaborative operation framework;Distribution network light storage collaborative operation framework is applied to the scene.This method can meet the demand of energy storage equipment according to power flow change dynamic charge and discharge, realize voltage control under distribution network light storage collaborative operation, ensure voltage stability during algorithm deployment and test period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of distribution network voltage risk management, specifically involving a voltage regulation method for the coordinated operation of photovoltaic and energy storage in distribution networks based on deep reinforcement learning. Background Technology

[0002] The global power system is accelerating its transition to sustainable energy, and photovoltaic (PV), as a clean and renewable energy source, has become a crucial pillar of the modern power grid. Energy storage systems (ESS) play a vital role in improving the utilization rate of renewable energy, especially when PV penetration is high. By storing surplus energy during peak power generation and releasing it when power generation is insufficient, energy storage devices can effectively mitigate power fluctuations. However, the inherent intermittency of PV power generation poses a significant challenge to maintaining power system stability. In distributed generation scenarios with high PV penetration, fluctuations in PV power generation can lead to voltage instability in the distribution network, directly impacting grid security and power quality.

[0003] Voltage regulation has always been a significant challenge in power systems. In recent years, the development of advanced communication and measurement technologies has significantly reduced the cost of real-time grid data acquisition, driving the research and application of voltage control methods. Common voltage control strategies include reactive power compensation, droop control, and optimization algorithm-based control methods; in addition, there are rule-based energy storage (ESS) control strategies, which regulate voltage by controlling the charging and discharging of energy storage devices. However, these traditional methods often suffer from high computational costs, high operating costs, reliance on precise physical models, and complex parameter adjustments, limiting their practical application. To address these challenges, reinforcement learning has been gradually applied to voltage regulation. Unlike traditional methods, reinforcement learning does not rely on precise mathematical models but instead makes optimization decisions through continuous interaction with the environment, exhibiting strong adaptability. However, the performance of reinforcement learning in high-dimensional state spaces remains limited, especially in complex power system environments where the agent needs to repeatedly experiment to converge to the optimal strategy. To overcome this problem, deep reinforcement learning (DRL) has emerged. This method utilizes neural networks to handle complex state and action spaces, improving the generalization ability of the strategy. Meanwhile, advancements in artificial intelligence algorithms have also driven the optimization of energy storage device control. For example, DRL-based energy dispatch strategies can improve operational efficiency while considering environmental constraints. Although DRL shows great potential in voltage regulation, it still suffers from insufficient security. In the exploratory phase of reinforcement learning, the voltage stability of the power grid may be affected by the agent's illegal actions. Therefore, safe deep reinforcement learning (safe-DRL) has received widespread attention and has been used for the safe control of renewable energy. However, most existing safe reinforcement learning methods focus on generation-side optimization, with less attention paid to voltage regulation under photovoltaic-storage synergistic operation. Although the energy management applications of energy storage units in the electricity market are relatively mature, their potential in voltage control may still be overlooked. Currently, research on voltage regulation of photovoltaic-storage synergistic operation based on safe-DRL is still relatively limited, and this direction has significant research value and application prospects. Summary of the Invention

[0004] This invention addresses the problems in existing technologies by providing a voltage regulation method for the coordinated operation of photovoltaic and energy storage (ESS) in distribution networks, guided by deep reinforcement learning. This invention mitigates voltage instability by utilizing controllable grid equipment such as photovoltaic inverters to dynamically respond to real-time grid conditions through active voltage regulation. Furthermore, it links the operation of the ESS to the system's active power balance conditions, thereby improving the absorption capacity of renewable energy and enhancing grid resilience.

[0005] The technical solution adopted in this invention is as follows: A voltage regulation method for photovoltaic-storage coordinated operation of distribution networks based on deep reinforcement learning, comprising the following steps:

[0006] Step 1: Select the distribution network to be regulated, which is equipped with photovoltaic inverters and energy storage devices; establish an incremental voltage control model for the distribution network.

[0007] Step 2: Establish a control model for photovoltaic inverters in the distribution network. This control model adopts a multi-agent deep reinforcement learning algorithm with safety constraints. Based on this algorithm, a policy network, a critic network, a target network, and an experience replay pool are established. The network structure of the policy network is changed based on the Russell invariance theorem to reduce the number of voltage limit violations during the training process.

[0008] Step 3: Establish a control model for energy storage devices in the distribution network. This control model adopts the deep deterministic policy gradient algorithm. Based on this algorithm, a policy network, a commentator network, a target network, and an experience replay pool are established.

[0009] Step 4: Embed active power balance rules into the deep deterministic policy gradient algorithm to control the charging or discharging state and power of the energy storage device; by changing the reward function, the energy storage device can play a role in reducing voltage fluctuation spikes.

[0010] Step 5: Adopt a strategy based on shared state space to connect the photovoltaic inverter control model and the energy storage device control model to establish a photovoltaic-storage collaborative operation framework for the distribution network;

[0011] Step 6: Collect power and load data in the distribution network, divide the area according to the distribution network topology, establish bus power flow constraints, and apply the distribution network photovoltaic-storage collaborative operation framework established in Step 5 to the distribution network for voltage regulation.

[0012] Compared with the prior art, the present invention has the following beneficial effects:

[0013] 1) This invention proposes a PV-ESS collaborative operation control framework for distribution networks, employing a shared state-space strategy to achieve coordinated and optimized regulation of photovoltaic inverters and energy storage batteries. This strategy avoids the information silo problem between individual control modules, improves the efficiency of PV-ESS collaborative operation in distribution networks, and enhances the adaptability of control and the flexibility of system operation.

[0014] 2) This invention applies multi-agent deep reinforcement learning (SC-MADRL) based on shared state space to voltage regulation in a photovoltaic-storage collaborative distribution network, and proposes a monotonic strategy network to improve control stability. Compared with the traditional DDPG method, this invention reduces the difficulty of agents exploring in high-dimensional state space, effectively reduces the risk of agent action constraint violations, and enhances the security and controllability of the system.

[0015] 3) This invention designs a Regional Active Power Balance Deep Deterministic Strategy Gradient (RAPB-DDPG) algorithm, embedding power supply and demand balance rules into the DDPG algorithm to guide energy storage devices to optimize scheduling based on photovoltaic output and load demand. Simultaneously, a voltage drop penalty term is introduced to effectively reduce voltage fluctuations and improve voltage stability. Compared to traditional rule-based control methods, this invention can improve photovoltaic absorption capacity while enabling energy storage devices to respond to system state changes in real time, adapt to different scenarios, and improve the voltage regulation capability and operational flexibility of the distribution network. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0017] Figure 1 A framework diagram of a voltage regulation method for photovoltaic-storage coordinated operation of a distribution network based on deep reinforcement learning, provided in an embodiment of the present invention;

[0018] Figure 2 This invention relates to the changes in agent reward value and objective function value during the deployment and training of the optical-storage collaborative operation framework provided in this embodiment.

[0019] Figure 3 These are four typical scenarios obtained through a scenario simplification algorithm, as provided in this embodiment of the invention.

[0020] Figure 4 This refers to the voltage effect without algorithm control.

[0021] Figure 5 A comparison chart of the average bus voltage of the distribution network based on three algorithms;

[0022] Figure 6 This is a comparison chart of the number of times the voltage exceeded the limit during the training process;

[0023] Figure 7 The recovery time required for voltage over-limit based on MADDPG provided in this embodiment of the invention;

[0024] Figure 8 The voltage over-limit recovery time provided by the embodiments of the present invention based on SC-MADDPG;

[0025] Figure 9 A schematic diagram illustrating how energy storage batteries ESU#1 and ESU#2 change during periods of photovoltaic fluctuations due to power supply and demand imbalances, as provided in an embodiment of the present invention.

[0026] Figure 10 A comparison chart showing the changes in SoC of energy storage batteries under rule-based control and RAPB-DDPG control;

[0027] Figure 11 A comparison graph showing voltage changes of energy storage batteries under rule-based control and RAPB-DDPG control;

[0028] Figure 12 The voltage changes in a distribution network with and without energy storage batteries are shown (taking a 14-node example). Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0030] It should be understood that the described embodiments are merely some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of the embodiments of this application.

[0031] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the embodiments of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0032] In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first," "second," "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0033] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0034] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0035] Step 1: Select the basic scenario of the distribution network and establish an incremental voltage control model based on the Distflow optimal power flow model.

[0036] During the execution of step 1, the voltage regulation problem can be modeled as a control task in a discrete dynamic system.

[0037] The power distribution network of this invention is an AC power distribution network with voltage monitoring and regulation capabilities, including an external power source, a main bus, several branch feeders, multiple load nodes, distributed photovoltaic generators, and energy storage devices. The distributed photovoltaic generators use photovoltaic inverters to control their reactive power output. The power distribution network is considered as a system diagram G = {N0, ε}, where N0 = {0, 1, 2, ..., n} represents nodes, and ε represents edges. Each node i has a corresponding active power injection p. i Reactive power injection q i and voltage v i .

[0038] In an exemplary embodiment, an improved IEEE 33 bus is selected as the distribution network to be regulated. This distribution network has a radial structure, containing 33 buses N0 = {0, 1, 2, ..., 33} and edges ε. Photovoltaic generators are deployed on buses 13, 18, 22, 25, 29, and 33, and energy storage batteries are configured on buses 18 and 33. Bus 0 represents the generator bus, which is connected to an external generator. The set of all buses except the generator bus is represented as N = N0 / {0}, and each bus has a corresponding active power injection p. i Reactive power injection q i and voltage v i Based on the Distflow optimal power flow model, the equations for bus voltage and power are established as follows:

[0039]

[0040] v j =v i -2(r ij P ij +x ij Q ij),(i,j)∈ε

[0041] Where j,k∈N,P ij and Q ij The active and reactive power flowing through line (i,j)∈ε represent the power and reactive power respectively; r ij and x ij These are the resistance and reactance on the line (i,j), l ij Represents line current; p j and q j Let represent the active power and reactive power of node j, where is written in matrix-vector form:

[0042] v = Rp + Xq + v01 = Xq + v env

[0043] Where the resistance matrix R = [R ij ] n×n The reactance matrix X = [X ij ] n×n R ij =2×∑ (h,k)∈(i,j) r hk X ij =2×∑ (h,k)∈(i,j) x hk (h,k) represents a segment in the line (i,j); the voltage v is divided into two parts: the controllable part Xq and the uncontrollable part v. env p and q represent the active power and reactive power vectors, respectively, v01 represents the initial voltage v0 multiplied by the unit vector and converted into vector form, and X and R are positive definite n-dimensional matrices;

[0044] The voltage regulation problem can be modeled as a control task in a discrete dynamic system. As time changes, the following closed-loop voltage dynamic control equation can be derived:

[0045] v(t+1)=Xq(t+1)+v env

[0046]

[0047] u i This represents an incremental voltage controller, where ΔT = 3 min is a fixed time window, and q i (t) represents the reactive power injected into node i at time t.

[0048] Step 2: Establish a control model for photovoltaic inverters in the distribution network. This control model adopts a multi-agent deep reinforcement learning algorithm with safety constraints. Based on this algorithm, a policy (Actor) network, a critic (Critic) network, a target network, and an experience replay pool are established. The network structure of the policy network is changed based on the Russell invariance theorem, which significantly reduces the number of voltage overruns during the neural network training process.

[0049] Step 2.1: The photovoltaic inverter is treated as an agent. The safety-constrained multi-agent deep reinforcement learning algorithm (SC-MADDPG) refers to a deep deterministic policy gradient algorithm with safety constraints, which includes a policy network, a critic network, a target network, and an experience replay pool. The critic network is a fully connected neural network, with input being the concatenation of the state space and action space, and outputting a Q-value representing the quality assessment of the state-action pair. A threshold for the experience replay pool is set. The structural parameters of the two target networks are consistent with those of the policy network and the critic network, respectively. The target network uses soft updates. This invention uses the safety-constrained multi-agent deep reinforcement learning (SC-MADDPG) algorithm to plan the optimal voltage regulation problem and build a deep reinforcement learning framework.

[0050] During step 2.1, the constructed Critic network is a fully connected neural network with an initial learning rate of clr_P, a network layer count of K1, and an experience replay pool threshold set to batchsize. The target network soft update coefficient is τ. The parameters of the constructed neural network may differ from those described in this embodiment.

[0051] In one exemplary embodiment, the optimal voltage regulation problem can be formulated as follows:

[0052]

[0053] st.v(t+1)=Xq(t+1)+v env

[0054] q i (t+1)=q i (t)+ΔT·u i (t)

[0055]

[0056] In the SC-MADDPG framework, γ is the discount factor; J(θ) is the objective function, representing the total cost, specifically including voltage deviation cost and control action cost; v i (t) represents the voltage of node i at time t; num = 6 is the number of agents. The structure of an incremental voltage controller is defined, which depends only on the local voltage magnitude.

[0057] In an exemplary embodiment, the Critic network is a fully connected neural network with an initial learning rate clr_P = 2e-4 and the number of layers K1 = 3; the empirical replay pool threshold is set to batchsize = 64. The target network structure and parameters are consistent with the Actor and Critic networks. The target network uses soft updates, and its hyperparameter is τ = 0.1.

[0058] Step 2.2: Modify the network structure of the policy network according to Russell's invariance theorem to establish a monotonically decreasing policy network with security constraints;

[0059] Russell's invariance theorem is used to determine the asymptotic stability of autonomous systems. For a discrete power system x(t+1) = f(x(t)), the theorem can be stated as follows:

[0060] (1) There exists a continuously differentiable Lyapunov function V(x) such that V(x) > 0 for all x, and satisfies V(x) = 0 at the equilibrium point; (2) ΔV(x) = V(f(x)) - V(x) ≤ 0.

[0061] If a Lyapunov function V(x) exists that satisfies both of the above conditions, the entire discrete system will converge to the invariant set S = {x∈R}. n V(f(x)) = V(x)}. The key issue is finding a suitable Lyapunov function and designing a suitable incremental voltage controller u. Theorem 1 provides a sufficient structural condition for designing an incremental voltage controller, thereby ensuring voltage stability:

[0062] (Theorem 1) For node i, f θi (·) is a continuously differentiable function, for Satisfy u i =-f θi (v i ) = 0, and In (-∞, v i ]and On the interval:

[0063]

[0064] at the same time

[0065] Theorem 1 states Needs to be within the range In reality, inverters typically have a sampling frequency in the kHz range, so they generally satisfy the left-hand side of Theorem 1. Diagonal matrix:

[0066]

[0067] Therefore, in order to satisfy the condition on the right side of Theorem 1, it is only necessary to satisfy f θi The Actor network can be monotonically increasing, while it can be monotonically decreasing.

[0068] During step 2.2, a suitable Lyapunov function V(v) is found, and the Russell invariance theorem is introduced as the theoretical support for designing the incremental voltage controller. The policy network is reconstructed into a monotonically decreasing network structure, and the initial learning rate of the Actor network is plr_P. The state space of the agent is defined as S. pv The action space is A pv .

[0069] In an exemplary embodiment, the Lyapunov function under consideration is V(v) = (vf u (v)) Τ X -1 (vf u (v)); In designing the incremental voltage controller, the Actor network is set as a monotonically decreasing network according to Theorem 1. The core expression of the Actor network can be written as y = q·h(x) + z·g(x), where h(x) and g(x) represent the hidden layer outputs, and q and z represent the weight matrices. According to the incremental voltage controller designed in step 3, the Actor network is a monotonically decreasing neural network with an initial learning rate of plcr_P = 1e-4, and the state space of the agent is S. pv ={v i}, v i The voltage of the bus where the photovoltaic generator is located, and the operating space A. pv ={q pv}, q pv This represents the reactive power output of the photovoltaic inverter.

[0070] Step 2.3, design the reward function r of the agent as follows:

[0071] r=-μ1*(u i (t)) 2 -λ1*||max(v node -v max ,0)+min(v node -v min ,0)||2 2

[0072] Where (u i (t)) 2 Represents the cost of agent action control, ||v node -v max ||2 and||v min -v node||2 represents the bus voltage deviation, which is measured by the L2 norm; λ1,μ1 represent the balance coefficients between the terms in the reward function; from a transient perspective, the agent's actions represent the change in reactive power of the photovoltaic inverter. Minimizing the action control cost reduces the reactive power loss of the generator, because the agent's actions represent the change in reactive power ΔQ of the photovoltaic inverter.

[0073] In this embodiment, the agent reward function is as follows:

[0074]

[0075] Where (u i (v i (t))) 2 Represents the cost of agent action control, ||v node -v max ||2 and||v min -v node ||2 represents the bus voltage drop, measured by the L2 norm. If the voltage v is within the safe range [0.95, 1.05], μ1 = 0, λ1 = 0; if the voltage exceeds the limit, then μ1 = 1, λ1 = 100. Step 3: Establish a control model for energy storage devices in the distribution network. This control model uses the Deep Deterministic Policy Gradient Algorithm (DDPG) for control. Based on this algorithm, establish the Actor network, Critic network, and experience replay pool, and set the network parameters.

[0076] In discrete dynamic systems, the control problem of energy storage devices is modeled as a Markov decision process, with the energy storage device acting as an intelligent agent and its state space being S. ESS The action space is A ESS The policy network is a recurrent neural network, and the network input is the state space S. ESS The output is the action space A. ESS This corresponds to the charging and discharging power of the energy storage battery; the critic network is a fully connected neural network, with the input being the concatenation of the state space and action space, and the output being the Q-value, representing the quality evaluation of the state-action pair; an experience replay pool threshold is set, and the structure and parameters of the two target networks are consistent with those of the policy network and the critic network; the target networks use soft updates;

[0077] In this implementation, energy storage devices specifically refer to energy storage batteries. In discrete power systems, the state of charge (SoC) of energy storage batteries changes as follows:

[0078] SoC(t+1) = SoC(t) + p storage ·ΔT / C max

[0079] Where SoC(t) represents the state of charge of the energy storage battery at time t, SoC(t)∈[10%,80%], p storage C represents the charging and discharging power of the intelligent agent. max This represents the maximum capacity of the energy storage device.

[0080] Specifically, during step 3, in the discrete power system, the energy storage device control problem is modeled as a Markov decision process. The Actor network has a learning rate of plr_ESS, a network layer count of K3, and a state space of S. ESS The action space is A ESS The Critic network has a learning rate of clr_ESS and K4 layers. The experience replay pool threshold is set to batchsize. The target network soft update coefficient is τ1.

[0081] In an exemplary embodiment, the Actor network has an initial learning rate of plr_ESS = 1e-4, the number of layers in the neural network is K3 = 3, and the network input is S. ESS =L×P×V×E, the output is the motion space A ESS ={p charge ,p discharge}, p charge p represents the charging power of the energy storage battery. discharge The value represents the discharge power of the energy storage battery; the Critic network has a learning rate of clr_ESS = 2e-4, and the number of layers K4 = 3. Its input is a concatenation of the state space and action space, and its output is the Q-value, representing the quality assessment of the state-action pair. The experience replay pool threshold is set to batchsize = 64. The target network structure and parameters are consistent with the Actor and Critic networks. The target network uses soft updates, and its hyperparameter is τ1 = 0.1.

[0082] Step 4: Embed active power balance rules into the deep deterministic policy gradient algorithm to control the charging or discharging state and power of the energy storage device; by changing the reward function, the energy storage unit can play a role in reducing voltage fluctuation spikes.

[0083] The agent selects the action to be executed in a single instance based on the regional active power balance rule, which refers to the balance between active power supply and demand in the region. Assume P... out P is the total power generation of the regional generators. con It is the total electricity consumption of the region, P bus The sum of active power flowing between all buses within the region, if P out >P con +P bus This indicates that regional supply exceeds demand, resulting in an energy surplus and encouraging smart agents to charge and store energy; conversely, if P...out <P con +P bus This indicates that the regional supply is insufficient to maintain power balance, encouraging the intelligent agent to discharge. A penalty term for voltage fluctuations is added to the energy storage device reward function, r. t+1 as follows:

[0084]

[0085] Where v_loss=|v-1| represents voltage loss, defined as the deviation of the system node voltage from 1; a represents the agent's action, a max The maximum value of the charging power is represented by ΔSoC = (SoC(t+1) - SoC(t)). ΔSoC represents the change in SoC between consecutive time steps. This penalty term is used to limit the fluctuation of SoC in consecutive time steps, because such fluctuations can lead to voltage instability.

[0086] During step 4, the active power balance rule is used as the criterion for the agent's reward; the voltage loss coefficient in the reward function is α, the SoC fluctuation penalty term coefficient is β, and k represents the constant penalty term for the agent choosing an incorrect action; the number of training rounds is m, and the number of interactions between the agent and the environment in a single round is n.

[0087] In an exemplary embodiment, taking region 1 as an example, P out The total power generation P of photovoltaic generators in region 1 pv1 P con Total load P in area 1 load1 P bus P represents the sum of active power flowing into and out of each busbar in region 1. bus1 If P is satisfied pv1 >P load1 +P bus1 If P is greater than demand, it indicates that energy supply in region 1 exceeds demand, encouraging intelligent agents to charge and store energy; conversely, if P is less than demand, it indicates that energy supply in region 1 exceeds demand. pv1 <P load1 +P bus1 This indicates that region 1's energy supply is insufficient to maintain power balance, encouraging the agent to discharge. A penalty term for voltage fluctuations is added to the energy storage unit's reward function. Taking energy surplus as an example, the reward function is designed as follows:

[0088]

[0089] Where v_loss = |v-1| represents voltage loss, defined as the deviation of the system node voltage from 1, α = 1; a represents the agent's action, a maxThe maximum charging power is represented by ΔSoC = (SoC(t+1) - SoC(t)), which represents the change in SoC state between consecutive time steps. This penalty term is used to limit the fluctuation of SoC between consecutive time steps, as such fluctuations can lead to voltage instability. β = 10, k = 1. The number of training rounds is m = 600, and the number of interactions between the agent and the environment in a single round is n = 40.

[0090] Step 5: Connect the photovoltaic inverter control model and the energy storage device control model using a shared state space-based strategy to establish a photovoltaic-energy storage collaborative operation framework for the distribution network.

[0091] The photovoltaic-storage collaborative operation framework for the distribution network includes a control model based on photovoltaic inverters and a control model based on energy storage devices. The control model based on photovoltaic inverters regulates the voltage of the distribution network bus by controlling the reactive power output of photovoltaic inverters to maintain a voltage deviation of no more than ±5%. The control model based on energy storage devices maintains the balance between power supply and demand in the system by selecting the charging and discharging actions of energy storage devices and adjusting the specific active power output, while mitigating voltage fluctuations and improving power quality.

[0092] The control model based on photovoltaic inverters and the control model based on energy storage devices adopt a shared state space strategy to share the same set of state spaces, that is, the states observed by the agents all come from the same environment; the agents in the deep reinforcement learning algorithms involved in the two models select their respective observed states from the shared state space, thereby using unified state variables to drive the execution of the entire framework.

[0093] During step 5, the improved IEEE 33 bus is selected as the research scenario, and the shared state space is S.

[0094] In one exemplary embodiment, the photovoltaic inverter control module is used to control the reactive power output of the photovoltaic inverter, and the energy storage battery control module is used to control the charging and discharging power of the energy storage battery. Figure 1 This is a framework diagram of the photovoltaic-storage coordinated operation control method for distribution networks based on deep reinforcement learning, as illustrated in this invention example. The shared state space is S = L × P × Q × V × E, where L = {(p L ,q L P = {p} represents the set of active and reactive power of the load; pv :p pv ∈(0,∞)} represents the set of active power of photovoltaic generators; Q={q pv :q pv ∈(0,∞)} represents the set of reactive power of photovoltaic generators; V={(v,θ):v∈(0,∞),θ∈[-π,π]} represents the set of bus voltage vectors, including phase angle θ and magnitude v, E={p ESU ,q ESU,SoC represents the set of active power, reactive power, and state of charge of an energy storage unit;

[0095] Step 6: Collect power and load data in the distribution network, divide the area according to the distribution network topology, establish bus power flow constraints, and apply the distribution network photovoltaic-storage collaborative operation framework established in Step 5 to the distribution network for voltage regulation.

[0096] To ensure the stable operation of the distribution network, bus power flow constraints are applied to photovoltaic generators and energy storage:

[0097]

[0098] |P storage (t)|≤P max

[0099] Where N pv This represents the set of busbars connected to the photovoltaic generator, where i represents a busbar; and These are the upper and lower limits of the active power of a photovoltaic generator, P. pv,i This represents the active power of the photovoltaic generator connected to node i; and These are the upper and lower limits of the reactive power of a photovoltaic generator, Q. pv,i P represents the reactive power of the photovoltaic generator connected to node i; max P represents the limit value of charging and discharging power. storage (t) represents the active power of the energy storage device during charging or discharging at time t.

[0100] Step 6.1: Collect photovoltaic power generation active power data;

[0101] During the execution of step 6.1, photovoltaic power generation active power data from any region can be selected, and the number of daily data collections and the sampling frequency ΔT can differ from those described in the embodiments of this invention.

[0102] In one exemplary embodiment, photovoltaic power generation data for a certain location from 2020 to 2022 was collected at a sampling frequency of 15 minutes, with 96 data points collected daily. To ensure that the photovoltaic data remained consistent with the real-time control cycle of the power grid, linear interpolation was used to supplement the collected data, resulting in an interpolated frequency ΔT of 3 minutes.

[0103] Step 6.2: Divide the region according to the minimum path from the branch node terminal to the trunk line;

[0104] During step 6.2, the main line is determined based on the selected distribution network scenario for voltage regulation, and the area is divided according to the minimum path from the branch node terminal to the main line.

[0105] In one exemplary embodiment, nodes 1-6 in the IEEE 33 bus are the backbone nodes. Based on the minimum path from the branch node terminal to the backbone in the network topology, the branch nodes sharing a backbone node are divided into regions. The system is divided into 4 regions: nodes 7-18 are region 1, nodes 19-22 are region 2, nodes 23-25 ​​are region 3, and nodes 26-33 are region 4.

[0106] Step 6.3: Apply the framework proposed in Step 5 to the improved IEEE 33 bus scenario;

[0107] In an exemplary embodiment, the specific training process of the two algorithms under the photovoltaic-storage collaborative operation framework of the distribution network is as follows:

[0108] Step 1: Initialize the Actor, Critic network parameters and the target network parameters; initialize the experience replay pool;

[0109] Step 2: Sample the current observation state and select agent action a through the Actor network. t It interacts with the environment and receives a reward r. t and the next state s t+1 Determine if the single-round training step size n = 40 has been reached. If it has, then d = True; otherwise, it is False.

[0110] Step 3: Storing Experience (s) t ,a t ,r t ,s t+1 d) is added to the experience pool. If the amount of stored experience exceeds a set threshold, batch data (s) is randomly sampled from the experience pool. i ,a i ,r i ,s i+1 ), calculate the target Q value, then calculate the Critic network loss and update the network parameters.

[0111] Step 4: Calculate the policy gradient and update the Actor network; softly update the target network parameters.

[0112] Step 5: Repeat steps 2-4 until the agent learns the optimal policy.

[0113] After performing the above steps, print the reward function and objective function graphs of the agent after training, as shown below. Figure 2 As shown.

[0114] First, the effectiveness of SC-MADDP is verified. In an exemplary embodiment, due to the large number of actual photovoltaic scenarios, a scenario-reduced K-means clustering algorithm is used to extract the four most representative scenarios in 2022 as experimental scenarios, such as... Figure 3 As shown. First, the voltage control effect of applying the SC-MADDPG algorithm was tested in these four scenarios, as follows: Figures 4 to 6 As shown; the voltage control effect is compared with that of no-algorithm control, droop-based control, and standard MADDPG algorithm. The average voltage of 33 buses is calculated, as follows. Figure 4 and Figure 5 As shown, the bus voltage exhibits significant fluctuations without control. Table 1 calculates the system average voltage, voltage fluctuation rate, and peak-to-valley voltage difference based on droop control, MADDPG, and SC-MADDPG algorithms. The results show that SC-MADDPG achieves the lowest fluctuation rate and peak-to-valley voltage difference, effectively stabilizing the voltage, reducing fluctuations, and improving power quality.

[0115] Table 1 - Comparison of Three Algorithms - Voltage Fluctuation Rate and Voltage Peak-to-Voltage Difference

[0116]

[0117] SC-MADDPG can reduce the number of times agents make illegal actions during network training. For example... Figure 6 As shown, SC-MADDPG exhibited only a small number of voltage violations in the first 50 training iterations. Specifically, during training, the probability of voltage violations was 0.21% for the SC-MADDP algorithm, compared to 2.6% for MADDPG. The results demonstrate that SC-MADDPG significantly reduces agent action violations, thereby improving distribution network stability.

[0118] In addition, voltage recovery time is another important indicator for measuring the operational stability of a distribution network. For example... Figure 7 and Figure 8 As shown, photovoltaic output disturbances were artificially introduced at 10:00 and 16:00 (lasting 12 minutes). The comparison results show that SC-MADDPG can quickly stabilize the voltage, while MADDPG leads to frequent voltage violations. Furthermore, the MADDPG algorithm's adjustment process exhibits a delayed response to photovoltaic disturbances. The performance of MADDPG and SC-MADDPG was evaluated in 100 photovoltaic step fluctuation scenarios, and the average duration of voltage violations is shown in Table 2. The results indicate that SC-MADDPG significantly reduces the duration of voltage violations and improves the distribution network's adaptability to external disturbances.

[0119] Table 2 - Performance Comparison of MADDPG and SC-MADDPG in 100 Voltage Violation Scenarios

[0120]

[0121] In an exemplary embodiment, to verify the effectiveness of the RAPB-DDPG algorithm, scenario #2* was selected as the test scenario. A photovoltaic output step disturbance was introduced at 13:00 and 14:00. At 13:00, a decrease in light intensity due to cloud cover was simulated, and at 14:00, a scenario of intense sunlight was simulated. The photovoltaic active power recovered to the original data at 15:00. Figure 9 As shown, during periods of photovoltaic fluctuations, the SoC of the energy storage battery changes with the imbalance between power supply and demand. This demonstrates that the RAPB-DDPG can control the energy storage battery to perform precise charging and discharging operations, thereby maintaining a balance between power supply and demand. The energy storage batteries are designated as ESU#1 and ESU#2, respectively.

[0122] Figure 10 The changes in the SoC of the energy storage battery under rule-based control and RAPB-DDPG control were compared, and the corresponding voltage changes were as follows: Figure 11 As shown in the diagram, the comparison reveals that rule-based control methods cause frequent switching between charging and discharging actions, leading to SoC oscillations, frequent voltage fluctuations, and accelerated battery aging. In contrast, RAPB-DDPG can provide a smoother system response by flexibly adjusting charging and discharging power, thereby improving voltage stability.

[0123] To evaluate the ability of the RAPB-DDPG algorithm to mitigate voltage fluctuations. Figure 12 The voltage variations of bus 14 are shown over 30 consecutive scenarios. The blue area represents the voltage variation, while the black voltage envelope illustrates how energy storage batteries can effectively reduce voltage spikes and improve power quality. Furthermore, energy storage batteries help improve the system's ability to absorb photovoltaic power. We compared the economic losses caused by curtailment of solar power in a distribution network integrating power generation, grid, load, and energy storage with those in a distribution network without energy storage batteries. According to international standards, the unit price for curtailment is US$0.10 per kilowatt-hour. Table 3 shows the daily economic losses and the cumulative losses over a month.

[0124] Table 3 - Economic Losses Caused by Abandoned Solar Power

[0125]

[0126] Table 3 shows that without energy storage batteries, the monthly economic loss due to curtailment of photovoltaic grid connection is US$2031.78, while with energy storage batteries, the loss is reduced to US$7.23, proving that the RAPB-DDPG algorithm can effectively enhance the photovoltaic absorption level of the system.

[0127] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A voltage regulation method for coordinated operation of a power distribution network and a light storage device based on deep reinforcement learning, characterized in that, The method comprises the following steps: Step 1, selecting a power distribution network to be subjected to voltage regulation, wherein a photovoltaic inverter and an energy storage device are arranged in the power distribution network; an incremental voltage control model of the power distribution network is established; Step 2, a control model of the photovoltaic inverter in the power distribution network is established, the control model adopts a safety-constrained multi-agent deep reinforcement learning algorithm, based on which a policy network, a critic network, a target network and an experience replay pool are established, and the network structure of the policy network is changed based on the Russell invariance theorem to reduce the number of voltage out-of-limit occurrences in the training process; Step 3, a control model of the energy storage device in the power distribution network is established, the control model adopts a deep deterministic policy gradient algorithm, based on which a policy network, a critic network, a target network and an experience replay pool are established; Step 4, an active power balance rule is embedded in the deep deterministic policy gradient algorithm to control the state and power of charging or discharging of the energy storage device; by changing the reward function, the energy storage device plays a role in reducing voltage fluctuation peaks; The step 4 is specifically: The agent selects a single-executed action according to a regional active power balance rule, the regional active power balance rule refers to a regional active power supply and demand balance relationship, assuming P out is a total power generation of a regional generator, P con is a total power consumption of the region, P bus represents a sum of active power flowing between each bus in the region, if P out >P con +P bus , it means that the region supply is greater than the demand, energy surplus, and the agent is encouraged to charge the energy storage; on the contrary, if P out <P con +P bus , it means that the region supply is insufficient to maintain power balance, and the agent is encouraged to discharge; a penalty term of voltage fluctuation is added to an energy storage device reward function r t+1 , and the energy storage device reward function r t+1 is as follows: where v_loss = |v - 1| represents the voltage loss, defined as the deviation of the system node voltage from 1; a represents the action of the agent, a max represents the maximum value of the charging power; AS0C = (SoC(t+1) - SoC(t)) represents the amount of change in SoC between consecutive time steps, this penalty term is used to limit the fluctuations of SoC in consecutive time steps, because such fluctuations will lead to voltage instability; a, b represent the penalty term coefficients, k represents the constant term penalty term for the agent selecting the wrong action; Step 5, the photovoltaic inverter control model and the energy storage device control model are connected based on a shared state space to establish a power distribution network photovoltaic and energy storage collaborative operation framework; Step 6, power and load data in the power distribution network are collected, regions are divided according to the power distribution network topology structure, bus flow constraints are established, and the power distribution network photovoltaic and energy storage collaborative operation framework established in step 5 is applied to the power distribution network for voltage regulation; The step 6 is specifically: Collecting photovoltaic active power, node load active power and reactive power in the power distribution network; According to the minimum path of the branch node terminal to the main line, the branch load nodes sharing the main node are divided into several regions, and the regions are optimized and controlled based on the region load and the operating state of the distributed photovoltaic generator; In order to ensure the stable operation of the power distribution network, the photovoltaic generator and the energy storage are subjected to bus flow constraints: |P storage (t)|≤P max where N pv represents the bus set connected with the photovoltaic generator, i represents the bus; and are the upper and lower limits of the active power of the photovoltaic generator, respectively, P pv,i represents the active power of the photovoltaic generator connected with the node i; and are the upper and lower limits of the reactive power of the photovoltaic generator, respectively, Q pv,i represents the reactive power of the photovoltaic generator connected with the node i; P max represents the limit value of the charging and discharging power, P storage (t) represents the active power of the energy storage device charging or discharging at time t.

2. The voltage regulation method for coordinated operation of a power distribution network and an optical storage device based on deep reinforcement learning according to claim 1, characterized in that, The step 1 comprises: The power distribution network is an alternating current power distribution network with voltage monitoring and regulation capability, including an external power source, a main bus, a plurality of branch feeders, a plurality of load nodes, a distributed photovoltaic generator and an energy storage device, wherein the distributed photovoltaic generator adopts a photovoltaic inverter to control the reactive power output thereof; the power distribution network is regarded as a system graph G={N0,ε}, N0={0,1,2,...,n} represents nodes, and ε represents edges; each node i has corresponding active power injection p i , reactive power injection q i and voltage v i ; according to a Distflow optimal power flow model, a matrix vector form of a node voltage equation is established: v = Rp+ Xq+ v01= Xq+ v env where the resistance matrix R = [R ij ] n×n , the reactance matrix X = [X ij ] n×n , R ij = 2 x ∑ (h,k)∈(i,j) r hk , X ij = 2 x ∑ (h,k)∈(i,j) x hk , (h, k) represents a segment in line (i, j), r hk and x hk are the resistance and reactance on line (h, k); the voltage v is divided into two parts, the controllable part Xq and the uncontrollable part v env , p and q represent the active power and reactive power vectors respectively, v01 represents the initial voltage v0 multiplied by the unit vector, which is converted into vector form, X and R are positive definite n-dimensional matrices; The voltage regulation problem is modeled as a control task in a discrete dynamic system that adjusts the reactive power injection of the photovoltaic inverter in response to real-time voltage measurements, thereby regulating the voltage magnitude; the rate of change of the reactive power injection is u i (t) = Δq i (t) / ΔT; zero-order holding of the input is performed, such that the input of the discrete dynamic system remains constant for the sampling time ΔT, and the resulting incremental voltage control model is obtained as follows: v(t + 1) = Xq(t + 1) + v env u i () represents the incremental voltage controller, ΔT is the fixed time window, q i (t) represents the reactive power of the injection node i at time t.

3. The method of claim 1, wherein the method further comprises: The step 2 comprises the following steps: 2.1) the photovoltaic inverter is taken as an agent, the safety-constrained multi-agent deep reinforcement learning algorithm is a safety-constrained deep deterministic policy gradient algorithm, which comprises a policy network, a critic network and a target network and an experience replay pool; the critic network is a fully connected neural network, the input is the splicing of the state space and the action space, the output is the Q value, which represents the quality evaluation of the state action pair; the experience replay pool threshold is set; the structure parameters of the two target networks are consistent with those of the policy network and the critic network; the target network adopts soft update; 2.2) according to the Russell invariance theorem, the network structure of the policy network is changed to establish a safety-constrained monotonically decreasing policy network; 2.3) the reward function r of the agent is designed as follows: r = -μ1*(u i (t)) 2 -λ1*||max(v node -v max ,0)+min(v node -v min ,0)||2 2 where (u i (t)) 2 represents the action control cost of the agent, ||v node -v max ||2and ||v min -v node ||2represents the bus voltage deviation, measured by the two-norm; λ1, μ1 represent the balance coefficient between each term in the reward function; from the transient point of view, the action of the agent represents the change of the reactive power of the photovoltaic inverter, and minimizing the action control cost reduces the loss of the generator reactive power.

4. The method of claim 1, wherein, The step 3 comprises: In discrete dynamic system, the energy storage device control problem is modeled as a Markov decision process, the energy storage device as an agent, the state space S ESS , the action space A ESS ; the policy network is a recurrent neural network, the network input is the state space S ESS , the output is the action space A ESS , corresponding to the charging and discharging power of the energy storage battery; the critic network is a fully connected neural network, the input is the splicing of the state space and the action space, the output is the Q value, representing the quality evaluation of the state-action pair; the experience replay pool threshold is set, the structure and parameters of the two target networks are consistent with the policy network and the critic network; the target network adopts soft update; In a discrete power system, the change of the state of charge SoC of the energy storage device is as follows: SoC(t+1) = SoC(t) + p storage • ΔT / C max Wherein, SoC(t) represents the state of charge of the energy storage device at time t, SoC(t)∈[10%, 80%], p storage represents the charging and discharging power of the intelligent agent, max represents the maximum capacity of the energy storage device.

5. The method of claim 1, wherein, The power distribution network light storage cooperative operation framework in step 5 includes a photovoltaic inverter-based control model and a storage device-based control model; wherein the photovoltaic inverter-based control model adjusts the bus voltage of the power distribution network by controlling the reactive power output of the photovoltaic inverter, and maintains the voltage deviation within ±5%; the storage device-based control model maintains the balance between power supply and demand of the system by selecting the charging and discharging actions of the storage device and adjusting the specific active power output, while slowing down the voltage fluctuation and improving the power quality.

6. The method of claim 1, wherein, In step 5, the photovoltaic inverter control model and the storage device control model are connected by using a shared state space-based strategy, which includes: The photovoltaic inverter-based control model and the storage device-based control model share the same set of state spaces based on the shared state space-based strategy, that is, the states observed by the agents are all from the same environment; the agents involved in the deep reinforcement learning algorithms of the two models select their own observation states from the shared state space, so as to drive the execution of the whole framework by using unified state variables.

Citation Information

Patent Citations

  • Active power distribution network voltage treatment method based on energy storage and photovoltaic inverter cooperative control

    CN118554554A

  • Voltage out-of-limit control system and method based on multi-agent deep reinforcement learning

    CN119401416A