Active power distribution network multi-device cooperation reactive power optimization control method and system
By constructing a multi-device reactive power grid optimization control model with photovoltaics, and using the improved TD3 algorithm optimized by LSTM network, combined with the multi-head self-attention mechanism and the crow search algorithm, the problem of large calculation burden and local optimal state in the reactive power grid optimization control is solved, and accurate and efficient reactive power optimization control effect is achieved.
Patent Information
- Application Number
- CN202411719768.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-11-28
AI Technical Summary
The prior art has a high computational burden in the reactive optimization control of distribution networks, which is prone to falling into a local optimal state, and relies heavily on the accuracy of prediction data, especially when the penetration rate of photovoltaic grid connection is high, it is difficult to achieve the expected control effect.
A multi-device collaborative reactive power optimization control method is proposed for active distribution network. By constructing a multi-device reactive power optimization control model for distribution networks containing photovoltaics, and using the improved TD3 algorithm optimized by LSTM network, combining the multi-head self-attention mechanism and crow search algorithm, it is converted into a Markov decision-making process, and the agent is trained to solve reactive power optimization problems.
It realizes accurate and efficient control of reactive power optimization in the distribution network. By enhancing the capture and utilization of historical state information by the LSTM network, it effectively captures the interdependence between different spatial locations in the real-time state data of the distribution network, improves the reactive power optimization effect, and improves the overall performance and global optimization capabilities.
Smart Images

Figure CN119944717A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power system optimization control, and in particular relates to a method and system for coordinated reactive power optimization control of multiple devices in an active distribution network. Background Art
[0002] With the development of power systems and the widespread application of smart distribution networks, the automation and intelligence levels of power grids are constantly improving. In smart distribution networks, the reasonable control of reactive power is of great significance for ensuring power quality, improving power supply efficiency and reducing line losses. However, due to the wide variety of distribution network equipment, complex operating conditions and high uncertainty, new challenges are brought to the safety and economy of power grid operation. The core work of reactive power optimization is to coordinate the work of various reactive compensation equipment, such as capacitors, static VAR compensators (SVCs) and voltage regulating transformers, to reduce system network losses and achieve reasonable reactive power distribution and voltage regulation. Traditional distribution network reactive power optimization usually establishes a timing optimization model by formulating objective functions and constraints, and determines the output of reactive power compensation devices in each time period by solving the optimization model to achieve the goal of reducing system losses and minimizing voltage deviations.
[0003] After searching the existing technical literature, it was found that the document "Reactive voltage optimization of distribution station area based on target programming method" (Ning Xin, Wang Tongxun, Chen Han, et al. Reactive voltage optimization of distribution station area based on target programming method [J]. Power Capacitors and Reactive Compensation, 2022, 43 (03): 1-7.) proposed a reactive voltage optimization method for distribution station area based on target programming method, designed the priority and secondary goals of voltage optimization, and used particle swarm algorithm to solve it. The document "Frequency reduction model predictive control of MMC under grid voltage unbalance conditions" (Wang Chunlin, Zhao Tao, Xu You, et al. Frequency reduction model predictive control of MMC under grid voltage unbalance conditions [J]. Journal of Electric Power Science and Technology, 2023, 38 (06): 67-75.) established a mathematical model of grid voltage unbalance under different control objectives and proposed a frequency reduction model predictive control strategy. However, the above methods usually have a large computational burden, are prone to fall into local optimal states, and are heavily dependent on the accuracy of prediction data. In addition, due to the uncertainty of PV forecasting affected by terrain, climate and time, it is difficult to accurately quantify the random impact of PV forecasting, which makes it difficult for these methods to achieve the expected control effect when the PV grid penetration rate is high. Deep reinforcement learning (DRL) is a data-driven method that has been shown to achieve online optimization based on partial observation data without relying on forecast data. At present, DRL methods for reactive power optimization in distribution networks mainly include deep Q networks (DQN), deep deterministic policy gradient (DDPG) and proximal policy optimization (PPO). The paper "Deep Reinforcement Learning Based Volt-VAR Optimization in Smart Distribution Systems" (Y. Zhang, X. Wang, J. Wang, et al. Deep Reinforcement Learning Based Volt-VAR Optimization in Smart Distribution Systems [J]. IEEE Transactions on Smart Grid, 2021, 12 (1): 361-371.) transforms the voltage control problem in the distribution network into a DQN framework to obtain the control strategy of the inverter, avoiding the direct solution of a specific optimization model.The paper "Attention Enabled Multi-Agent DRL for Decentralized Volt-VAR Control of Active Distribution System Using PV Inverters and SVCs" (D.Cao, J.Zhao, W.Hu, et al. Attention Enabled Multi-Agent DRL for Decentralized Volt-VAR Control of Active Distribution System Using PV Inverters and SVCs [J]. IEEE Transactions on Sustainable Energy, 2021, 12 (3): 1582-1592.) proposed a reactive power optimization scheme based on DDPG to coordinate the output of multiple photovoltaic inverters to cope with the random impact of photovoltaics. However, the deep reinforcement learning algorithms used in the above papers all use simple linear layers as policy networks, which cannot perform importance analysis on real-time distribution network status information. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for active distribution network multi-device collaborative reactive power optimization control, which can accurately and efficiently solve the reactive power optimization control problem of distribution network, in view of the above-mentioned problems existing in the prior art.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows:
[0006] In a first aspect, the present invention proposes a method for coordinated reactive power optimization control of multiple devices in an active distribution network, comprising:
[0007] S1. Construct a reactive power optimization control model for multiple devices in a distribution network including photovoltaics that takes into account the safety and economy of system operation; construct a TD3 algorithm model optimized by an improved LSTM network, wherein the improved LSTM network integrates a multi-head self-attention mechanism in the LSTM network;
[0008] S2. Convert the above multi-device reactive power optimization problem into a Markov decision process, use the TD3 algorithm optimized by the improved LSTM network to train the intelligent agent and solve the multi-device reactive power optimization problem, so as to obtain the optimal control solution.
[0009] In S1, the improved LSTM network introduces a multi-head self-attention mechanism in each time step, and calculates the similarity between the current time step and the previous time step through the multi-head self-attention mechanism to obtain a weighted vector, which is combined with the input of the current time step as the updated current time step.
[0010] The improved LSTM network uses the crow search algorithm to optimize the number of hidden layer units and the initial learning rate of the LSTM network, including:
[0011] A. Initialize the crow population;
[0012] B. Calculate the fitness function value of each crow at the initial position and update the position of each crow, where the fitness function is:
[0013]
[0014] In the above formula, x is the position of the crow, M is the total number of training rounds, and r is t (x) is the reward function for the tth training;
[0015] The position update formula is:
[0016]
[0017] In the above formula, x a,iter 、fl a,iter are the position and flight step length of crow a at the iterth iteration, r a 、r b is a constant randomly distributed between 0 and 1, m b,iter is the location where crow b hides the food at the iterth iteration, AP a,iter is the perception probability of crow a at the iter-th iteration;
[0018] C. Calculate the fitness function value of the crow at the new position. If the fitness function value of the new position is better than the fitness function value of the hidden food position, update the position of the hidden food, otherwise do not update;
[0019] D. Repeat steps BC until the iteration termination condition is reached, and output the optimal hidden food location at this time.
[0020] The step A comprises:
[0021] A1. Generate the chaotic sequence [x1,x2,···,x k ,···,x N ], where N is the number of crows in the crow population, and
[0022] A2. Map the chaotic sequence to the crow individuals to form the initial crow population:
[0023] x0=x l +(x l -x u )x k+1
[0024] In the above formula, x0 is the initial crow population, x l 、x u are the lower and upper bounds of the search space respectively;
[0025] In the step B, fl a,iter Calculated according to the following formula:
[0026]
[0027] In the above formula, are the minimum and maximum flight steps of crow a, iter is the current iteration number, iter max is the maximum number of iterations.
[0028] In S1, the multi-device reactive power optimization control model of the distribution network including photovoltaics aims to minimize voltage deviation, network loss and equipment operation cost and maximize reactive power margin. Its objective function includes:
[0029]
[0030]
[0031]
[0032] In the above formula, m loss is the network loss cost coefficient, P loss,t , Q margin,t are the network loss and reactive power margin in period t, m v is the voltage deviation cost coefficient, v i,t is the voltage amplitude of node i at time period t, v ref is the rated voltage of the distribution network, m SVC 、m OLTC 、m Q are SVC switching cost coefficient, OLTC regulation cost coefficient, reactive power shortage cost coefficient, N SVC,t 、N OLTC,t are the number of SVC switching times and the number of OLTC tap adjustment times, T is the number of time periods in the control cycle, N is the number of distribution network nodes, and I ij,t is the current of the line between nodes i and j during period t, R ij is the resistance of the line between nodes i and j, is the reactive power upper limit of node i, is the reactive power output by node i during period t;
[0033] The constraints include photovoltaic output constraints, network flow constraints, node voltage constraints, SVC capacity constraints, and on-load tap-changing transformer tap position constraints.
[0034] In S2, the Markov decision process includes a state space, an action space, and a reward function;
[0035] The state space S is:
[0036] S={S V ,S PV ,S SVC ,S L}
[0037] In the above formula, S V is the voltage amplitude of each node in the distribution network, S PV is the output information of the photovoltaic unit, S SVC is the output power of the reactive power compensation equipment, S L is the current electric load collection of the distribution network;
[0038] The action space A is:
[0039] A={A OLTC ,Q PV ,Q SVC}
[0040] In the above formula, A OLTC is the transformer gear, Q PV is the reactive power output of the photovoltaic inverter, Q SVC It is the reactive power output of the reactive power compensation device;
[0041] The reward function r t for:
[0042]
[0043] In a second aspect, the present invention proposes an active distribution network multi-device collaborative reactive power optimization control system, including an optimization model construction module, an optimized TD3 algorithm model construction module, and a model solving module;
[0044] The optimization model building module is used to build a multi-device reactive power optimization control model of a photovoltaic distribution network that takes into account system operation safety and economy;
[0045] The optimized TD3 algorithm model building module is used to build an improved LSTM network optimized TD3 algorithm model, wherein the improved LSTM network integrates a multi-head self-attention mechanism in the LSTM network;
[0046] The model solving module is used to convert the above-mentioned multi-device reactive power optimization problem into a Markov decision process, use the TD3 algorithm optimized by the improved LSTM network to train the intelligent agent and solve the multi-device reactive power optimization problem, so as to obtain the optimal control solution.
[0047] The improved LSTM network uses the crow search algorithm to optimize the number of hidden layer units and the initial learning rate of the LSTM network, including:
[0048] A. Initialize the crow population;
[0049] B. Calculate the fitness function value of each crow at the initial position and update the position of each crow, where the fitness function is:
[0050]
[0051] In the above formula, x is the position of the crow, M is the total number of training rounds, and r is t (x) is the reward function for the tth training;
[0052] The position update formula is:
[0053]
[0054] In the above formula, x a,iter 、fl a,iter are the position and flight step length of crow a at the iterth iteration, r a 、r b is a constant randomly distributed between 0 and 1, m b,iter is the location where crow b hides the food at the iterth iteration, AP a,iter is the perception probability of crow a at the iter-th iteration;
[0055] C. Calculate the fitness function value of the crow at the new position. If the fitness function value of the new position is better than the fitness function value of the hidden food position, update the position of the hidden food, otherwise do not update;
[0056] D. Repeat steps BC until the iteration termination condition is reached, and output the optimal hidden food location at this time.
[0057] The step A comprises:
[0058] A1. Generate the chaotic sequence [x1,x2,···,x k ,···,x N ], where N is the number of crows in the crow population, and
[0059] A2. Map the chaotic sequence to the crow individuals to form the initial crow population:
[0060] x0=x l +(x l -x u )x k+1
[0061] In the above formula, x0 is the initial crow population, x l 、x u are the lower and upper bounds of the search space respectively;
[0062] In the step B, fl a,iter Calculated according to the following formula:
[0063]
[0064] In the above formula, are the minimum and maximum flight steps of crow a, iter is the current iteration number, iter max is the maximum number of iterations.
[0065] The multi-device reactive power optimization control model for the distribution network including photovoltaics aims to minimize voltage deviation, network loss and equipment operation cost and maximize reactive power margin. Its objective function includes:
[0066]
[0067]
[0068]
[0069] In the above formula, m loss is the network loss cost coefficient, P loss,t , Q margin,t are the network loss and reactive power margin in period t, m v is the voltage deviation cost coefficient, v i,t is the voltage amplitude of node i at time period t, v ref is the rated voltage of the distribution network, m SVC 、m OLTC 、m Q are SVC switching cost coefficient, OLTC regulation cost coefficient, reactive power shortage cost coefficient, N SVC,t 、N OLTC,t are the number of SVC switching times and the number of OLTC tap adjustment times, T is the number of time periods in the control cycle, N is the number of distribution network nodes, and I ij,t is the current of the line between nodes i and j during period t, R ij is the resistance of the line between nodes i and j, is the reactive power upper limit of node i, is the reactive power output by node i during period t;
[0070] The constraints include photovoltaic output constraints, network power flow constraints, node voltage constraints, SVC capacity constraints, and on-load tap changer transformer tap position constraints;
[0071] The Markov decision process includes a state space, an action space, and a reward function;
[0072] The state space S is:
[0073] S={S V ,S PV ,S SVC ,S L}
[0074] In the above formula, S V is the voltage amplitude of each node in the distribution network, S PV is the output information of the photovoltaic unit, S SVC is the output power of the reactive power compensation equipment, S L is the current electric load collection of the distribution network;
[0075] The action space A is:
[0076] A={A OLTC ,Q PV ,Q SVC}
[0077] In the above formula, A OLTC is the transformer gear, Q PV is the reactive power output of the photovoltaic inverter, Q SVC It is the reactive power output of the reactive power compensation device;
[0078] The reward function r t for:
[0079]
[0080] Compared with the prior art, the present invention has the following beneficial effects:
[0081] 1. The present invention provides a method for the coordinated reactive power optimization control of multiple devices in an active distribution network. First, a reactive power optimization control model of multiple devices in a distribution network including photovoltaics is constructed, which takes into account the safety and economy of the system operation, and a TD3 algorithm model optimized by an improved LSTM network. Then, the reactive power optimization problem of the multiple devices is converted into a Markov decision process, and the TD3 algorithm optimized by an improved LSTM network is used to train the intelligent agent and solve the reactive power optimization problem of the multiple devices, thereby obtaining the optimal control scheme. The method enhances the capture and utilization of historical state information by the LSTM network by integrating a multi-head self-attention mechanism in the LSTM network, so that the deep neural network can effectively capture the interdependence between different spatial positions in the real-time state data of the distribution network, thereby improving the reactive power optimization effect.
[0082] 2. In the active distribution network multi-device collaborative reactive power optimization control method of the present invention, the improved LSTM network adopts the crow search algorithm to optimize the hyperparameters of the LSTM network, which can effectively improve the overall performance of the reactive power optimization of the distribution network.
[0083] 3. The present invention provides a method for collaborative reactive power optimization control of multiple devices in an active distribution network. On the one hand, in order to solve the problem that the diversity of crow population is poor and affects the optimization accuracy, Tent chaotic mapping is introduced to initialize the population. This process increases the diversity of the crow population, enables it to explore the optimization space more widely, and effectively improves the global optimization capability of the algorithm. On the other hand, the method adopts an adaptive step size strategy to dynamically change the step size decreasing rate according to the optimization requirements of different stages, so that the algorithm can maintain a faster global search in the early stage, quickly traverse the search space, and accelerate the convergence speed of the algorithm in the early stage, and perform refined search with a smaller step size in the later stage to ensure the accuracy of local search. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 This is a flowchart of the method described in Example 1.
[0085] Figure 2 This is the Actor and Critic network structure of the optimized TD3 algorithm model.
[0086] Figure 3 This is a structural diagram of the system described in Example 2.
[0087] Figure 4 The figure is a comparison of the convergence of the algorithm described in the present invention and other algorithms.
[0088] Figure 5 The figure is a comparison of the voltage control effect of the algorithm described in the present invention and other algorithms. DETAILED DESCRIPTION
[0089] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0090] Embodiment 1:
[0091] A multi-device coordinated reactive power optimization control method for active distribution network, such as Figure 1 As shown, the specific steps are as follows:
[0092] 1. Comprehensively consider the system operation safety and economy, and build a reactive power optimization control model for the smart distribution network including photovoltaics.
[0093] Considering the active and reactive regulation characteristics of photovoltaic inverters, the reactive output of reactive compensation equipment (static var compensator, SVC) and controllable devices such as the tap position of on-load tap-changing transformers, the reactive optimization control of the distribution network is realized by minimizing voltage deviation, network loss and equipment operation cost and maximizing reactive margin while meeting the distribution network operation constraints such as power flow constraints and voltage constraints of the distribution system. The objective function is as follows:
[0094]
[0095]
[0096]
[0097] In the above formula, m loss is the network loss cost coefficient, P loss,t , Q margin,t are the network loss and reactive power margin in period t, m v is the voltage deviation cost coefficient, v i,t is the voltage amplitude of node i at time period t, v ref is the rated voltage of the distribution network, m SVC 、m OLTC 、m Q are SVC switching cost coefficient, OLTC regulation cost coefficient, reactive power shortage cost coefficient, N SVC,t 、N OLTC,t are the number of SVC switching times and the number of OLTC tap adjustment times, T is the number of time periods in the control cycle, N is the number of distribution network nodes, and I ij,t is the current of the line between nodes i and j during period t, R ij is the resistance of the line between nodes i and j, is the reactive power upper limit of node i, is the reactive power output by node i during period t;
[0098] The constraints include:
[0099] Photovoltaic output constraints
[0100] The active and reactive power output capabilities of photovoltaic inverters are limited by the equipment capacity, and the power range of the two outputs can be expressed as:
[0101]
[0102]
[0103]
[0104] In the above formula, are the active and reactive power of the photovoltaic unit connected to node i during period t, are the maximum active and reactive power generated by the photovoltaic unit connected to node i, s PV The equipment capacity of the photovoltaic power grid-connected inverter;
[0105] Network flow constraints
[0106]
[0107]
[0108] In the above formula, P i G , are the active power and reactive power injected into the node at bus i, respectively, P i L , are the active load and reactive load at bus i, respectively, v i 、v j are the voltages of busbars i and j respectively, G ij , B ij are the conductance and susceptance between busbars i and j, θ ij is the phase difference between busbars i and j;
[0109] Node Voltage Constraints
[0110] U min ≤U i,t ≤U max
[0111] In the above formula, U i,t is the voltage amplitude of node i at time period t, U max , U min are the upper and lower limits of the node voltage respectively;
[0112] SVC capacity constraints
[0113] Q SVC,min ≤Q SVC,i ≤Q SVC,max
[0114] In the above formula, Q SVC,i is the reactive power output of SVC reactive power compensation device i, Q SVC,max , Q SVC,min They are the maximum and minimum reactive outputs allowed by the SVC reactive compensation equipment respectively;
[0115] On-load tap-changing transformer tap position constraints
[0116] V 1,t =V0±δ OLTC,t ΔV
[0117] In the above formula, V 1,t is the voltage amplitude of the first node in period t, V0 is the voltage amplitude of the first node, δ OLTC,t is the position of the tap during period t, and ΔV is the position of the tap.
[0118] 2. Improve the LSTM network based on the multi-head self-attention mechanism.
[0119] The LSTM network is a specialized recurrent neural network (RNN) architecture that is mainly used to solve the gradient vanishing and gradient exploding problems encountered during long sequence training. Its essence is that it is better at processing longer sequences than ordinary RNNs, while effectively retaining historical features through the gating mechanism. Its structure contains three gating mechanisms and two transmission states, namely: input gate z i 、Forget Gate f , output gate z o , cell state c t and the hidden state h t .
[0120] The LSTM network model constructed in this embodiment includes three different calculation processes: forgetting process, memory selection process and output process. The forgetting stage requires selectively deleting information from the previous node. In the present invention, this involves selectively eliminating irrelevant state information from the previous moment in the process of reactive power optimization of the distribution network. The calculated value z f As a forget gate, it allows control of the previous state c t-1 In the memory selection phase, the input x t Select to memorize. The current input is calculated by the previous process, and the gate signal selection is determined by z i control. This means that the TD3 algorithm ensures the effective capture and storage of historical data related to reactive power optimization. In the output stage, the LSTM network uses historical state information to determine the effective state output. This process is derived from the previous state z o and c o Control and scale using activation functions. These processes can be expressed by the following formulas:
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] This embodiment introduces the multi-head self-attention mechanism MHA mechanism into the LSTM network to improve the utilization of historical state information, so that the deep neural network model can effectively capture and emphasize the interdependence between different spatial positions in the input data. MHA is introduced in each time step of the LSTM network to calculate the similarity between the current time step and the previous time step, thereby obtaining the following weighted vector, which effectively represents the relevant historical state:
[0127]
[0128] In the above formula, Q, K, and V are query matrix, key matrix, and value matrix respectively. K is the dimension of matrix K;
[0129] Combining this vector with the input of the current time step as the updated current time step, MHA allows the model to calculate multiple attention distributions in parallel. The calculation of each attention head is as follows:
[0130] Head i =Attention(QW i q ,KW i k ,VW i ν )
[0131] In the above formula, W i q , W i k , W i ν are the weights of Q, K, and V in the i-th head respectively.
[0132] Finally, the results of all heads are concatenated and linearly transformed to obtain the final weighted vector:
[0133]
[0134] In the above formula, h is the number of heads, W i h is the output transformation matrix.
[0135] The structure of the MHA mechanism includes an encoder, a decoder, and an attention mechanism. The encoder converts the input sequence into a series of historical vectors, while the decoder uses the currently generated vector and the output of the previous time step to predict the subsequent time step. The attention mechanism determines the extent to which the decoder obtains information from the encoder in order to obtain state information about collision avoidance during the prediction process.
[0136] 3. Optimizing the hyperparameters of LSTM network based on CSA
[0137] Improved LSTM network uses the crow search algorithm CSA to optimize the hyperparameters of the LSTM network, including the number of hidden layer units and the initial learning rate. The crow search algorithm is a heuristic optimization algorithm based on swarm intelligence, inspired by the foraging behavior and information transmission process of crow populations. The algorithm simulates the behavior of crow groups in finding food and transmitting information to search for the optimal solution. Its process includes:
[0138] a) Set the number of crows N and the maximum number of iterations iter max , flight length fl and perception probability AP.
[0139] b) Initialize the crow population.
[0140] c) Calculate the fitness function value of each crow at the initial position and update the position of each crow, where:
[0141] The fitness function is:
[0142]
[0143] In the above formula, x is the position of the crow, M is the total number of training rounds, and r is t (x) is the reward function for the tth training;
[0144] By maximizing the reward (i.e., minimizing the negative value of the fitness function), the CSA algorithm can find the optimal LSTM hyperparameters and improve the overall performance of reactive power optimization in distribution networks;
[0145] The position update formula is:
[0146]
[0147] In the above formula, x a,iter 、fl a,iter are the position and flight step length of crow a at the iterth iteration, r a 、r b is a constant randomly distributed between 0 and 1, m b,iter is the location where crow b hides the food at the iterth iteration, AP a,iter is the perception probability of crow a at the iter-th iteration.
[0148] d) Test the feasibility of the new position. If the new position is feasible, the crow will update its position; otherwise, the crow remains in its current position.
[0149] e) Calculate the fitness function value of the new position. If the fitness function value of the new position is better than the fitness function value of the hidden food position, then use the following formula to update the hidden food position, otherwise do not update:
[0150]
[0151] f) Repeat steps c) to e) until the maximum number of iterations iter is reached max , and the optimal hidden food location is taken as the output solution.
[0152] 4. Build a TD3 algorithm model for improved LSTM network optimization.
[0153] Figure 2 It is the Actor and Critic network structure of the optimized TD3 algorithm model.
[0154] The Actor network receives the system state information S as input, processes the input state through the improved LSTM network, captures the dependency of state information at different time steps, further extracts features through the hidden layer of the multi-layer neural network, and finally generates action a at the output layer t , that is, the selected control strategy. At the same time, the Actor network contains a strategy network and a target strategy network, which are used to optimize the control strategy.
[0155] The Critic network receives system state information S and action information A as input. It also enhances the focus on the dependency between state-action pairs by improving the LSTM network. After extracting features through the hidden layer, it generates a Q value in the output layer to evaluate the quality of the current strategy. The Critic network contains two evaluation networks (Q1 and Q2) and their corresponding target evaluation networks. It reduces the deviation of strategy estimation and improves the accuracy of evaluation through the dual Q learning method.
[0156] The above structure enables the improved TD3 algorithm to generate more accurate control strategies in the Actor network and perform more reliable strategy evaluation in the Critic network, thereby realizing the generation of the optimal strategy for coordinated reactive power optimization control of multiple devices in the distribution network.
[0157] In dual Q-learning, the target network plays a key role as a reference for value estimation. This algorithm improves the performance of the current valuation network by identifying and selecting the action with the highest valuation through a greedy strategy. Similarly, the TD3 algorithm learns two Q functions with a common goal while minimizing the mean square error to improve its prediction accuracy. This phenomenon can be greatly amplified by effectively preventing certain areas in the state space. The smaller of the two Q values is used to update the target Q function, which is expressed as:
[0158]
[0159] The use of the target network makes the target update more stable and helps the network adapt to a wider range of training data, while the deep reinforcement learning (DRL) algorithm requires multiple gradient updates to achieve convergence and fit the target. Therefore, the TD3 algorithm updates the Actor network at a lower frequency and updates the Critic network at a higher frequency. The network loss update can be expressed as:
[0160]
[0161] This algorithm incorporates a regularization strategy to alleviate the overfitting problem in the value function and promote a smoother target policy. The principle is based on clipped double Q learning and target policy smoothing, and its expression is:
[0162] α TD3 (s′)=clip(u θ,targ (s′)+clip(ε,-c,c),α low ,α high )
[0163] In the above formula: ε is the noise sampled from the normal distribution, c is the range parameter, α low , α high They are the evaluation networks updated at lower and higher frequencies respectively. Finally, the update of the policy network can be expressed as:
[0164]
[0165] In summary, the proposed algorithm effectively improves the Q-value estimation accuracy of the action-value function in the Critic network and introduces delayed updates of the Actor network, thereby ensuring the stability of training.
[0166] 5. Convert the reactive power optimization problem of multiple devices in the distribution network including photovoltaics into a Markov decision process.
[0167] Deep reinforcement learning combines the decision-making ability of reinforcement learning and the nonlinear function approximation ability of deep neural networks, enabling it to learn effective control strategies in complex environments. In this framework, the reactive power optimization problem of the distribution network is transformed into a Markov decision process (MDP), which consists of three basic elements: state space, action space, and reward function.
[0168] 1) State Space
[0169] The state space S represents the distribution network environment information obtained by the intelligent agent when interacting with the environment. The present invention selects the voltage amplitude of each node in the distribution network environment, the output of each photovoltaic unit, the reactive compensation output of each SVC, and the system electrical load as state variables. The state space is defined as follows:
[0170] S={S V ,S PV ,S SVC ,S L}
[0171] In the above formula, S V is the voltage amplitude of each node in the distribution network, S PV is the output information of the photovoltaic unit, S SVC is the output power of the reactive power compensation equipment, S L is the current electric load collection of the distribution network;
[0172] 2) Action Space
[0173] The action space A is the set of executable operations that the agent takes to achieve reactive power optimization control. In the present invention, the action space is defined as the on-load tap changer (OLTC) gear adjustment, the reactive power regulation of the photovoltaic inverter, and the reactive power output of the SVC. The expression is as follows:
[0174] A={A OLTC ,Q PV ,Q SVC}
[0175] In the above formula, A OLTC is the transformer gear, Q PV is the reactive power output of the photovoltaic inverter, Q SVC It is the reactive power output of the reactive power compensation device;
[0176] 3) Reward Function
[0177] The reward function is a mechanism for evaluating the behavior of the agent to guide the agent to learn the optimal strategy. In the problem of reactive power optimization control of distribution network, the design goal of the reward function is to ensure the voltage stability of the power grid, minimize the power loss and equipment action cost, and maximize the reactive power margin, as shown in the following formula:
[0178]
[0179] 6. Use the intelligent agent in the TD3 algorithm model optimized by the improved LSTM network to solve the multi-device reactive power optimization control model of the distribution network including photovoltaics and obtain the optimal control solution.
[0180] Embodiment 2:
[0181] The steps are the same as those in Example 1, except that:
[0182] This embodiment improves the crow search algorithm based on chaos concept and adaptive step size, specifically:
[0183] (1) Using Tent Chaos Map to Initialize Crow Population
[0184] In step A, the chaotic sequence [x1,x2,···,x k ,···,x N ], where N is the number of crows in the crow population, and
[0185] Then map the chaotic sequence onto individual crows to form the initial crow population:
[0186] x0=x l +(x l -x u )x k+1
[0187] In the above formula, x0 is the initial crow population, x l 、x u are the lower and upper bounds of the search space respectively;
[0188] The chaotic sequence generated by Tent mapping has the characteristics of uniform distribution, which helps the population to traverse the search space more evenly, and has the characteristics of fast iteration, which can improve the global search ability of the algorithm, making the population evenly distributed in the search space, avoiding premature convergence to the local optimal solution, and more effectively finding the optimal hyperparameter combination. In the reactive power optimization problem, the local reactive power demand of the system varies greatly. The piecewise linear structure of Tent mapping enables its chaotic sequence to maintain good randomness in each iteration, increasing the ability of the algorithm to jump out of the local optimum. In addition, the reactive power optimization process involves multiple constraints and boundary conditions. The chaotic sequence of Tent mapping can still maintain stable traversal and randomness under different parameter settings, and has good robustness and strong adaptability.
[0189] (2) Adopting adaptive step size
[0190] The traditional crow search algorithm usually uses a fixed step size when updating the crow position, which causes the algorithm to miss the optimal solution due to the large step size in the early stage of the search, and converge too slowly due to the small step size in the later stage of the search. To make up for this deficiency, this embodiment uses the following nonlinear decreasing flight step size:
[0191]
[0192] In the above formula, are the minimum and maximum flight steps of crow a, iter is the current iteration number, iter max is the maximum number of iterations.
[0193] The above nonlinear decreasing method is more flexible in adjusting the step size, and can dynamically change the step size decreasing rate according to the optimization requirements of different stages, so that the algorithm can maintain a larger step size in the initial stage, maintain a faster global search, and quickly traverse the search space, thereby accelerating the early convergence speed, which is especially important for reactive power optimization problems that require rapid traversal of the search space; in the later stage, a smaller step size can be used for refined search, which helps to find a more superior solution and improve the accuracy of local search.
[0194] Embodiment 3:
[0195] A multi-device coordinated reactive power optimization control system for active distribution network, such as Figure 3 As shown, it includes an optimization model building module, an optimized TD3 algorithm model building module, and a model solving module;
[0196] The optimization model building module is used to build a multi-device reactive power optimization control model of a photovoltaic distribution network that takes into account system operation safety and economy;
[0197] The optimized TD3 algorithm model building module is used to build an improved LSTM network optimized TD3 algorithm model, wherein the improved LSTM network integrates a multi-head self-attention mechanism in the LSTM network;
[0198] The model solving module is used to convert the above-mentioned multi-device reactive power optimization problem into a Markov decision process, use the TD3 algorithm optimized by the improved LSTM network to train the intelligent agent and solve the multi-device reactive power optimization problem, so as to obtain the optimal control solution.
[0199] The improved LSTM network uses the crow search algorithm to optimize the number of hidden layer units and the initial learning rate of the LSTM network, including:
[0200] A. Initialize the crow population;
[0201] B. Calculate the fitness function value of each crow at the initial position and update the position of each crow, where the fitness function is:
[0202]
[0203] In the above formula, x is the position of the crow, M is the total number of training rounds, and r is t (x) is the reward function for the tth training;
[0204] The position update formula is:
[0205]
[0206] In the above formula, x a,iter 、fl a,iter are the position and flight step length of crow a at the iterth iteration, r a 、r b is a constant randomly distributed between 0 and 1, m b,iter is the location where crow b hides the food at the iterth iteration, AP a,iter is the perception probability of crow a at the iter-th iteration;
[0207] C. Calculate the fitness function value of the crow at the new position. If the fitness function value of the new position is better than the fitness function value of the hidden food position, update the position of the hidden food, otherwise do not update;
[0208] D. Repeat steps BC until the iteration termination condition is reached, and output the optimal hidden food location at this time.
[0209] The step A comprises:
[0210] A1. Generate the chaotic sequence [x1,x2,···,x k ,···,x N ], where N is the number of crows in the crow population, and
[0211] A2. Map the chaotic sequence to the crow individuals to form the initial crow population:
[0212] x0=x l +(x l -x u )x k+1
[0213] In the above formula, x0 is the initial crow population, x l 、x u are the lower and upper bounds of the search space respectively;
[0214] In the step B, fl a,iter Calculated according to the following formula:
[0215]
[0216] In the above formula, are the minimum and maximum flight steps of crow a, iter is the current iteration number, iter max is the maximum number of iterations.
[0217] The multi-device reactive power optimization control model for the distribution network including photovoltaics aims to minimize voltage deviation, network loss and equipment operation cost and maximize reactive power margin. Its objective function includes:
[0218]
[0219]
[0220]
[0221] In the above formula, m loss is the network loss cost coefficient, P loss,t , Q margin,t are the network loss and reactive power margin in period t, m v is the voltage deviation cost coefficient, v i,t is the voltage amplitude of node i at time period t, v ref is the rated voltage of the distribution network, m SVC 、m OLTC 、m Q are SVC switching cost coefficient, OLTC regulation cost coefficient, reactive power shortage cost coefficient, N SVC,t 、N OLTC,t are the number of SVC switching times and the number of OLTC tap adjustment times, T is the number of time periods in the control cycle, N is the number of distribution network nodes, and I ij,t is the current of the line between nodes i and j during period t, R ij is the resistance of the line between nodes i and j, is the reactive power upper limit of node i, is the reactive power output by node i during period t;
[0222] The constraints include photovoltaic output constraints, network power flow constraints, node voltage constraints, SVC capacity constraints, and on-load tap changer transformer tap position constraints;
[0223] The Markov decision process includes a state space, an action space, and a reward function;
[0224] The state space S is:
[0225] S={S V ,S PV ,S SVC ,S L}
[0226] In the above formula, S V is the voltage amplitude of each node in the distribution network, S PV is the output information of the photovoltaic unit, S SVC is the output power of the reactive power compensation equipment, S L is the current electric load collection of the distribution network;
[0227] The action space A is:
[0228] A={A OLTC ,Q PV ,Q SVC}
[0229] In the above formula, A OLTC is the transformer gear, Q PV is the reactive power output of the photovoltaic inverter, Q SVC It is the reactive power output of the reactive power compensation device;
[0230] The reward function r t for:
[0231]
[0232] In order to investigate the effectiveness of the method of the present invention, the following performance comparison is performed:
[0233] (1) Convergence performance
[0234] The algorithm described in Example 2 is compared with the convergence of the conventional TD3 algorithm and the DDPG algorithm. The results are as follows: Figure 4 shown.
[0235] It can be seen that the convergence speed of the algorithm of the present invention is significantly better than that of the other two algorithms. The average reward value increases rapidly and tends to be stable with fewer training times, and finally stabilizes near 0. The convergence speed of the TD3 algorithm and the DDPG algorithm is relatively slow, and the average reward value after reaching convergence is also lower than that of the algorithm of the present invention. This result shows that the algorithm of the present invention has a faster convergence speed and better control performance in the coordinated reactive power optimization control of multiple devices in the distribution network.
[0236] (2) Voltage control effect
[0237] The voltage control effect of the method described in Example 2 is compared with that of the TD3 algorithm and the DDPG algorithm. The results are as follows: Figure 5 , as shown in Table 1.
[0238] from Figure 4It can be seen that the system optimized by the method described in the present invention achieves strict control of the system voltage by adjusting the active and reactive power of renewable energy power generation, adjusting the reactive output of reactive compensation equipment, and accurately controlling the on-load tap changer, so as to keep it within a safe range. At the same time, this optimization strategy significantly reduces the active power loss of the system, improves the reactive power margin, and further improves the voltage stability of the system. Since the original system is radially distributed, the node voltage will gradually decrease with the increase of the line length, especially at the end of the line, the voltage drop is particularly significant. Through reactive power optimization, the joint control of the active and reactive power of distributed power sources is achieved, which significantly improves the overall voltage level of the system, reduces the peak-to-valley difference of the voltage, and further enhances the robustness of the system and the voltage control effect.
[0239] As can be seen from Table 1, the algorithm of the present invention is superior to other algorithms in reactive voltage control effect. In terms of average voltage, the algorithm of the present invention controls the system voltage at 1.014pu, which is slightly lower than the 1.0285pu of the DDPG algorithm and the 1.0281pu of the TD3 algorithm, indicating that the algorithm can maintain the system voltage at the target level more accurately. In terms of voltage deviation, the algorithm of the present invention performs best, which is 0, while the voltage deviations of the DDPG algorithm and the TD3 algorithm are 0.071 and 0.0452pu, respectively, indicating that the algorithm of the present invention can better eliminate voltage fluctuations. In terms of average network loss, the algorithm of the present invention also shows the best result, which is 0.0713MW, which is significantly lower than the 0.1005MW of the DDPG algorithm and the 0.0848MW of the TD3 algorithm, showing a significant advantage in optimizing power grid losses. In terms of reactive power margin, the average reactive power margin of the algorithm of the present invention reaches 133.43 kvar, which is higher than 124.83 kvar of the TD3 algorithm and 111.23 kvar of the DDPG algorithm. This means that under the optimization of the algorithm of the present invention, the system has higher reactive power regulation capability and stronger stability.
[0240] Table 1 Reactive power and voltage control effects of different algorithms
[0241]
Claims
1. A method for coordinated reactive power optimization control of multiple devices in an active distribution network, characterized in that: The method comprises: S1. Construct a reactive power optimization control model for multiple devices in a distribution network including photovoltaics that takes into account the safety and economy of system operation; construct a TD3 algorithm model optimized by an improved LSTM network, wherein the improved LSTM network integrates a multi-head self-attention mechanism in the LSTM network; S2. Convert the above multi-device reactive power optimization problem into a Markov decision process, use the TD3 algorithm optimized by the improved LSTM network to train the intelligent agent and solve the multi-device reactive power optimization problem, so as to obtain the optimal control solution.
2. The method for controlling multi-device coordinated reactive power optimization in an active distribution network according to claim 1, characterized in that: In S1, the improved LSTM network introduces a multi-head self-attention mechanism in each time step, and calculates the similarity between the current time step and the previous time step through the multi-head self-attention mechanism to obtain a weighted vector, which is combined with the input of the current time step as the updated current time step.
3. The method for coordinated reactive power optimization control of multiple devices in an active distribution network according to claim 1 or 2, characterized in that: The improved LSTM network uses the crow search algorithm to optimize the number of hidden layer units and the initial learning rate of the LSTM network, including: A. Initialize the crow population; B. Calculate the fitness function value of each crow at the initial position and update the position of each crow, where the fitness function is: In the above formula, x is the position of the crow, M is the total number of training rounds, and r is t (x) is the reward function for the tth training; The position update formula is: In the above formula, x a,iter 、fl a,iter are the position and flight step length of crow a at the iterth iteration, r a 、r b is a constant randomly distributed between 0 and 1, m b,iter is the location where crow b hides the food at the iterth iteration, AP a,iter is the perception probability of crow a at the iter-th iteration; C. Calculate the fitness function value of the crow at the new position. If the fitness function value of the new position is better than the fitness function value of the hidden food position, update the position of the hidden food, otherwise do not update; D. Repeat steps BC until the iteration termination condition is reached, and output the optimal hidden food location at this time.
4. The method for controlling multi-device coordinated reactive power optimization in an active distribution network according to claim 3, characterized in that: The step A comprises: A1. Generate the chaotic sequence [x1,x2,···,x k ,···,x N ], where N is the number of crows in the crow population, and A2. Map the chaotic sequence to the crow individuals to form the initial crow population: x0=x l +(x l -x u )x k+1 In the above formula, x0 is the initial crow population, x l 、x u are the lower and upper bounds of the search space respectively; In the step B, fl a,iter Calculated according to the following formula: In the above formula, are the minimum and maximum flight steps of crow a, iter is the current iteration number, iter max is the maximum number of iterations.
5. The method for coordinated reactive power optimization control of multiple devices in an active distribution network according to claim 1 or 2, characterized in that: In S1, the multi-device reactive power optimization control model of the distribution network including photovoltaics aims to minimize voltage deviation, network loss and equipment operation cost and maximize reactive power margin. Its objective function includes: In the above formula, m loss is the network loss cost coefficient, P loss,t , Q margin,t are the network loss and reactive power margin in period t, m v is the voltage deviation cost coefficient, v i,t is the voltage amplitude of node i at time period t, v ref is the rated voltage of the distribution network, m SVC 、m OLTC 、m Q are SVC switching cost coefficient, OLTC regulation cost coefficient, reactive power shortage cost coefficient, N SVC,t 、N OLTC,t are the number of SVC switching times and the number of OLTC tap adjustment times, T is the number of time periods in the control cycle, N is the number of distribution network nodes, and I ij,t is the current of the line between nodes i and j during period t, R ij is the resistance of the line between nodes i and j, is the reactive power upper limit of node i, is the reactive power output by node i during period t; The constraints include photovoltaic output constraints, network flow constraints, node voltage constraints, SVC capacity constraints, and on-load tap-changing transformer tap position constraints.
6. The method for controlling multi-device coordinated reactive power optimization in an active distribution network according to claim 5, characterized in that: In S2, the Markov decision process includes a state space, an action space, and a reward function; The state space S is: S={S V ,S PV ,S SVC ,S L } In the above formula, S V is the voltage amplitude of each node in the distribution network, S PV is the output information of the photovoltaic unit, S SVC is the output power of the reactive power compensation equipment, S L is the current electric load collection of the distribution network; The action space A is: A={A OLTC ,Q PV ,Q SVC } In the above formula, A OLTC is the transformer gear, Q PV is the reactive power output of the photovoltaic inverter, Q SVC It is the reactive power output of the reactive power compensation device; The reward function r t for:
7. An active distribution network multi-device collaborative reactive power optimization control system, characterized in that: The system includes an optimization model building module, an optimized TD3 algorithm model building module, and a model solving module; The optimization model building module is used to build a multi-device reactive power optimization control model of a photovoltaic distribution network that takes into account system operation safety and economy; The optimized TD3 algorithm model building module is used to build an improved LSTM network optimized TD3 algorithm model, wherein the improved LSTM network integrates a multi-head self-attention mechanism in the LSTM network; The model solving module is used to convert the above-mentioned multi-device reactive power optimization problem into a Markov decision process, use the TD3 algorithm optimized by the improved LSTM network to train the intelligent agent and solve the multi-device reactive power optimization problem, so as to obtain the optimal control solution.
8. The active distribution network multi-device coordinated reactive power optimization control system according to claim 7, characterized in that: The improved LSTM network uses the crow search algorithm to optimize the number of hidden layer units and the initial learning rate of the LSTM network, including: A. Initialize the crow population; B. Calculate the fitness function value of each crow at the initial position and update the position of each crow, where the fitness function is: In the above formula, x is the position of the crow, M is the total number of training rounds, and r is t (x) is the reward function for the tth training; The position update formula is: In the above formula, x a,iter 、fl a,iter are the position and flight step length of crow a at the iterth iteration, r a 、r b is a constant randomly distributed between 0 and 1, m b,iter is the location where crow b hides the food at the iterth iteration, AP a,iter is the perception probability of crow a at the iter-th iteration; C. Calculate the fitness function value of the crow at the new position. If the fitness function value of the new position is better than the fitness function value of the hidden food position, update the position of the hidden food, otherwise do not update; D. Repeat steps BC until the iteration termination condition is reached, and output the optimal hidden food location at this time.
9. The active distribution network multi-device coordinated reactive power optimization control system according to claim 8, characterized in that: The step A comprises: A1. Generate the chaotic sequence [x1,x2,···,x k ,···,x N ], where N is the number of crows in the crow population, and A2. Map the chaotic sequence to the crow individuals to form the initial crow population: x0=x l +(x l -x u )x k+1 In the above formula, x0 is the initial crow population, x l 、x u are the lower and upper bounds of the search space respectively; In the step B, fl a,iter Calculated according to the following formula: In the above formula, are the minimum and maximum flight steps of crow a, iter is the current iteration number, iter max is the maximum number of iterations.
10. An active distribution network multi-device coordinated reactive power optimization control system according to claim 6 or 7, characterized in that: The multi-device reactive power optimization control model for the distribution network including photovoltaics aims to minimize voltage deviation, network loss and equipment operation cost and maximize reactive power margin. Its objective function includes: In the above formula, m loss is the network loss cost coefficient, P loss,t , Q margin,t are the network loss and reactive power margin in period t, m v is the voltage deviation cost coefficient, v i,t is the voltage amplitude of node i at time period t, v ref is the rated voltage of the distribution network, m SVC 、m OLTC 、m Q are SVC switching cost coefficient, OLTC regulation cost coefficient, reactive power shortage cost coefficient, N SVC,t 、N OLTC,t are the number of SVC switching times and the number of OLTC tap adjustment times, T is the number of time periods in the control cycle, N is the number of distribution network nodes, and I ij,t is the current of the line between nodes i and j during period t, R ij is the resistance of the line between nodes i and j, is the reactive power upper limit of node i, is the reactive power output by node i during period t; The constraints include photovoltaic output constraints, network power flow constraints, node voltage constraints, SVC capacity constraints, and on-load tap changer transformer tap position constraints; The Markov decision process includes a state space, an action space, and a reward function; The state space S is: S={S V ,S PV ,S SVC ,S L } In the above formula, S V is the voltage amplitude of each node in the distribution network, S PV is the output information of the photovoltaic unit, S SVC is the output power of the reactive power compensation equipment, S L is the current electric load collection of the distribution network; The action space A is: A={A OLTC ,Q PV ,Q SVC } In the above formula, A OLTC is the transformer gear, Q PV is the reactive power output of the photovoltaic inverter, Q SVC It is the reactive power output of the reactive power compensation device; The reward function r t for:
Citation Information
Patent Citations
Reactive power optimization method based on double-delay depth deterministic strategy gradient
CN116468159A
Power load prediction method and system based on deep learning
CN117455256A
Micro-grid energy optimization scheduling method oriented to source grid load storage
CN118971051A
Neural network feature extractor for actor-critic reinforcement learning models
US20240143975A1
Power system regulation method based on deep reinforcement learning
WO2024092954A1
Cited By
High-permeability photovoltaic power distribution network multi-target voltage optimization method based on genetic algorithm
CN122137034A