Power distribution network security constraint voltage regulation and control method and system based on deep reinforcement learning
Through the improved DDPG algorithm model, the policy network and value network are optimized, and the instability of traditional deep reinforcement learning under the noise interference of power grid is solved, and the rapid and stable regulation of voltage within the safe range is achieved, and the operation reliability and efficiency of the power system are improved.
Patent Information
- Application Number
- CN202510776353.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Traditional deep reinforcement learning algorithms are sensitive to environmental noise and are easily disturbed by noise under dynamic changes in the power grid or external disturbances, affecting control accuracy and system stability, making it difficult to ensure voltage stability in complex scenarios.
Using the improved DDPG algorithm model, through the improvement of policy network and value network, the weight recovery matrix and the bias recovery matrix are introduced, and combined with Relu activation function and layer normalization technology, the control signal generation process is optimized to ensure that the voltage is controlled within a reasonable range.
It improves the stability and accuracy of voltage control, can quickly respond to dynamic changes in the power grid and external disturbances, ensures that the voltage is stable within the range of 0.95~1.05 p.u., reduces system fluctuations, and improves the operating reliability and efficiency of the power system.
Smart Images

Figure CN120300818A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to voltage control technology in a power system, and particularly to a method and system for regulating the voltage of a distribution network with security constraints based on deep reinforcement learning. Background Art
[0002] In a power system, voltage control is a key link to ensure system stability, improve power quality, and equipment safety. With the continuous expansion of the scale and increasing complexity of modern power systems, problems such as voltage fluctuations, overvoltage, and undervoltage have gradually become one of the main factors affecting the stable operation of the power system. The research and application of voltage control technology are of great significance for optimizing grid operation, reducing equipment losses, and improving power transmission efficiency. The purpose of voltage control is to maintain the voltage in the power system within a reasonable range to ensure the normal operation of power equipment and avoid equipment damage and power transmission instability caused by abnormal voltage. Too high or too low voltage may have an adverse impact on the system, affecting the life of equipment, system safety, and overall operation efficiency. Therefore, through effective voltage control methods, the economic, safe, and reliable operation of the power system can be achieved. Traditional voltage control methods rely on static and dynamic devices, such as transformers, reactive power compensation devices, etc. However, with the access of renewable energy, the increase in distributed power sources, and the improvement of the complexity of power loads, voltage control faces new challenges. Modern voltage control not only needs to consider static load changes but also the instantaneous fluctuations and other sudden events caused by the addition of new energy, etc. Moreover, the high-dimensional characteristics of large-capacity power systems make the real-time voltage control problem more complex. Especially when encountering severe but unknown disturbances, how to quickly and optimally implement effective control becomes a major challenge. With the increasing complexity of the grid structure, the traditional method of solving through accurate mathematical models can no longer effectively perform emergency control. Therefore, it is necessary to develop a new real-time implementation control paradigm to effectively and quickly support grid operators to carry out autonomous and effective control actions.
[0003] Traditional voltage control methods, such as centralized control, distributed control, local control, capacitor control, and droop control, are widely used in power systems. However, these methods also have some significant drawbacks. First, traditional control methods often rely on preset models and parameters and have poor adaptability to system dynamic changes. Especially when facing complex situations such as grid load fluctuations and distributed energy access, they cannot flexibly adjust control strategies. Second, these methods usually lack sufficient adaptive capabilities. When the system state changes drastically, it is difficult to make an effective response in a timely manner, resulting in large voltage fluctuations or unstable control. In addition, traditional voltage control methods have poor resistance to environmental noise and interference. Especially when there are uncertainties and external disturbances in the power system, errors are easily generated, affecting the stability of the system.
[0004] However, the traditional DDPG algorithm has a high computational complexity, requires processing a large amount of empirical data and optimizing multiple networks, resulting in high consumption of computing resources and unable to meet the requirements of real-time voltage control. Secondly, the algorithm has a slow convergence speed. Due to the dependence on a large amount of interaction experience and unstable gradient update during the training process, it is difficult to quickly optimize the control strategy in a dynamic environment. Finally, the traditional DDPG algorithm is sensitive to environmental noise. Especially under the dynamic changes of the power grid or external disturbances, it is easily affected by noise interference, which affects the control accuracy and system stability, and it is difficult to ensure voltage stability in complex scenarios. Summary of the Invention
[0005] The technical problem to be solved by the present invention: Aiming at the above problems of the prior art, a method and system for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning are provided. The present invention aims to solve the problems that traditional deep reinforcement learning is sensitive to environmental noise, is easily affected by noise interference under the dynamic changes of the power grid or external disturbances, affects the control accuracy and system stability, and is difficult to ensure voltage stability in complex scenarios.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning, including the following steps: training an improved DDPG algorithm model, and the improved DDPG algorithm model includes improving the policy network of the DDPG algorithm model; the improved policy network is used to generate corresponding actions according to the state of the distribution network, where the state of the distribution network refers to the node voltage state of the distribution network, and the action refers to the control signal for controlling the reactive power injection of the distribution network. The improved policy network generating corresponding actions according to the state of the distribution network includes: S101, calculating and generating a weight recovery matrix and a bias recovery matrix according to the following formula: , , where, is the weight recovery matrix, represents generating an upper triangular matrix, where the elements on and below the diagonal are 0 and the elements above are 1; represents generating an upper triangular matrix, where the elements on and below the diagonal and the two rows below are 0 and the elements above are 1; is a 2-fold identity matrix, that is, the elements on the diagonal are 2 and the other elements are 0; where is the operation of the upper triangular matrix, is the offset from the main diagonal, is the identity matrix, is the bias recovery matrix; S102, applying the weight recovery matrix The forward weight matrix is calculated according to the following formula and the negative weight matrix : ,
[0007] in, is the forward weight matrix, is the negative weight matrix, is the parameter; S103, calculate the forward bias matrix according to the following formula and the negative bias matrix : , , in, is the forward bias matrix, is the negative bias matrix, is the bias reference matrix, and are the upper and lower voltage limits, respectively. is a matrix of all 1s; S104, the forward weight matrix and the distribution network status at time t After multiplication, the Relu activation function is used to activate the result, and the L1 norm of the result is calculated to convert the negative weight matrix and the distribution network status at time t Multiply and add the result of L1 norm calculation as the first control signal; then calculate the difference between the positive bias matrix, negative bias matrix and bias reference matrix respectively, and multiply the two differences by the distribution network state at time t. The second control signal is obtained by multiplying the two control signals, and the two control signals are added to obtain the final control signal for controlling reactive power injection of the distribution network.
[0008] Optionally, the improved DDPG algorithm model includes improving the value network of the DDPG algorithm model, the improved value network is used to generate a corresponding Q value according to the state of the distribution network and the action generated by the strategy network, the improved value network includes a normalization function, a first hidden layer, a second hidden layer and an output layer, and the function expression for generating the corresponding Q value is: , in, Represents the value network according to the distribution network status at time t ,action The generated Q value, is the weight matrix of the output layer; is the bias vector of the output layer, is the output of the second hidden layer, and there is: , , , wherein, is the Leaky ReLU activation function, is the normalization operation, is the weight matrix of the second hidden layer; is the bias vector of the second hidden layer, is the output of the first hidden layer, the bias vector of the first hidden layer, is the weight matrix of the state input, is the weight matrix of the action input, is the normalization result of the distribution network state at time t by the normalization function ; is the normalization result of the action at time t by the normalization function ; is the slope parameter of the negative half axis.
[0009] Optionally, the training of the improved DDPG algorithm model includes: S201, initialize the policy network, value network and their respective target networks, and the parameters of the target network are initialized the same as those of the original network; initialize the experience replay buffer; S202, obtain the current state of the distribution network by interacting with the environment of the distribution network, generate the corresponding action according to the current state of the distribution network through the policy network, obtain the reward and the next state of the action, and put the current state, action, reward and the next state as an experience into the experience replay buffer; wherein the function expression for obtaining the reward of the action is: , wherein, is the reward of node at time t, is the voltage deviation of node at time t, is the control input of node at time t, is the voltage deviation, is the action cost, is the balance coefficient of the voltage deviation, is the balance coefficient of the action cost, is the target voltage value of node and there is: the average value of , , , wherein, is the voltage at time t+1, is the conductance matrix of the power grid, is the reactive power injection at time t+1; is the influence of the external environment on the voltage, and the influence of the external environment on the voltage includes the reactive power injection of the node at time t and time t+1; and are respectively the reactive power injections of node at time t and time t+1, is the time interval, is 's control strategy function, is the policy parameter of the control strategy function of node ; S203, perform experience sampling from the experience replay buffer based on the prioritized experience replay mechanism, including sampling N experiences from the experience replay buffer according to the priority, and the probability of each experience being sampled is: , , , wherein, is the probability of experience i being sampled, is the priority value of experience i, is the priority adjustment factor, is the error value, is the TD error of experience i, is the reward of experience i; is the discount factor, which is used to control the influence degree of future rewards on the current value; is the target value network for the state at time t+1 and the target policy to generate the Q value of the action, is the Q value obtained by the value network based on the distribution network state at time t and the action ; S204, update the parameters of the policy network and the value network respectively by backpropagation by combining the sampled experiences with the loss function; S205, softly update the parameters of the target networks of the policy network and the value network respectively; S206, jump to step S203 until the preset training end condition is reached.
[0010] Optionally, in step S204, the functional expression of the loss function used to update the parameters of the value network is: , , , where, is the loss function used to update the parameters of the value network, is the expected value, is the importance sampling weight of experience i, is the probability that experience i is sampled, is the size of the experience replay buffer, is the Q value obtained by the value network based on the current state and action as well as the parameters ; is the target value of the current experience, is the parameter used to adjust the weight; is the discount factor, which is used to control the influence degree of future rewards on the current action; is the Q value obtained by the target network of the value network according to the state at time t+1 and action ; is the action generated by the policy network according to the state at time t+1 .
[0011] Optionally, in step S204, the functional expression of the loss function used to update the parameters of the policy network is: , , where, is the loss function used to update the parameters of the policy network, is the expected value, is the Q value obtained by the value network based on the distribution network state at time t and action ; is the action obtained by the policy network according to the current state and parameters ; is the weight parameter of entropy regularization, which is used to adjust the trade-off between exploration and exploitation in policy optimization; is the entropy of the policy, is the distribution network state of the policy network at time t The probability distribution of taking action a below.
[0012] Optionally, in step S205, the functional expressions of the parameters of the target networks of the soft update policy network and the value network are: , where, is the parameter of the target network after soft update, is the soft update, is the parameter of the target network corresponding to the policy network or the value network, is the parameter of the target network before soft update; is the step size of the soft update, which controls the update speed of the target network.
[0013] Optionally, when initializing the policy network, the value network and their respective target networks in step S201, it includes adding noise to the output action of the policy network for generating exploratory actions, and its functional expression is: , , where, is the deterministic action in the current state , is the policy network, and are respectively and the noise values at times ; is the mean regression speed, which is used to control the rate at which the noise gradually regresses to the mean; is the random amplitude of the noise; refers to white noise sampled from the standard normal distribution.
[0014] In addition, the present invention also provides a distribution network security-constrained voltage regulation system based on deep reinforcement learning, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the distribution network security-constrained voltage regulation method based on deep reinforcement learning.
[0015] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored, and the computer program or instruction is programmed or configured to execute the distribution network security-constrained voltage regulation method based on deep reinforcement learning through a processor.
[0016] In addition, the present invention also provides a computer program product, including a computer program or instruction, and the computer program or instruction is programmed or configured to execute the distribution network security-constrained voltage regulation method based on deep reinforcement learning through a processor.
[0017] Compared with the prior art, the present invention can mainly achieve the following beneficial effects: The present invention includes training an improved DDPG algorithm model, and the improved DDPG algorithm model includes improving the policy network of the DDPG algorithm model. The improved policy network generates corresponding actions according to the state of the distribution network, including: generating a weight recovery matrix and a bias recovery matrix, calculating the positive and negative weighted matrices and bias matrices, and fusing them with the state of the distribution network at time t to generate two control signals, adding the two control signals to obtain a control signal finally used to control the reactive power injection of the distribution network. By the above method, a structurally controllable weight and bias recovery matrix are introduced into the improved policy network, so as to explicitly regulate the mapping process between the state and the action, and ensure that the generated control instruction can effectively guide the voltage to adjust towards the desired range (such as 0.95~1.05 p.u.). The node voltage state of the distribution network is used as the input of the policy network, processed by the weighted matrix and bias matrix constructed by the upper triangular structure, and a directional weight matrix and a corresponding bias recovery matrix are generated. The above matrices are respectively used to model the control behaviors when the voltage is higher than the upper limit and lower than the lower limit, so as to output a suppression-type reactive power injection signal and an enhancement-type reactive power injection signal respectively. When the node voltage is higher than the upper limit, the positive weight and bias recovery matrix are activated, so that the positive control quantity takes effect and a signal for lowering the voltage is generated; when the node voltage is lower than the lower limit, the negative weight and bias recovery matrix are activated, so that the negative control quantity takes effect and a signal for raising the voltage is generated; if the voltage is in the safe interval, the outputs of the two control signals are both close to 0, and the action tends to be stable or there is no injection. Therefore, the present invention can solve the problems that traditional deep reinforcement learning is sensitive to environmental noise, is easily affected by noise interference under the dynamic changes or external disturbances of the power grid, affects the control accuracy and system stability, and is difficult to ensure voltage stability in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 It is a schematic diagram of the network structure of the improved policy network in the embodiment of the present invention.
[0019] Figure 2 It is a schematic diagram of the network structure of the improved value network in the embodiment of the present invention.
[0020] Figure 3 It is a schematic diagram of the working process of the improved DDPG algorithm model in the embodiment of the present invention.
[0021] Figure 4 It is a schematic diagram of the working principle of the improved DDPG algorithm model in the embodiment of the present invention.
[0022] Figure 5Comparison of the average rewards between the improved DDPG algorithm model and the original DDPG algorithm model in the embodiments of the present invention.
[0023] Figure 6 Voltage control effects of the improved DDPG algorithm model in the embodiments of the present invention and the existing safe-DDPG algorithm model on Node 2.
[0024] Figure 7 Voltage control effects of the improved DDPG algorithm model in the embodiments of the present invention and the existing safe-DDPG algorithm model on Node 7.
[0025] Figure 8 Voltage control effects of the improved DDPG algorithm model in the embodiments of the present invention and the existing safe-DDPG algorithm model on Node 9. Detailed implementation manners
[0026] To enable those skilled in the art of the present technology to better understand the technical solutions of the present invention, the technical solutions of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention. This embodiment provides a method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning, including the following steps: training an improved DDPG algorithm model, and the improved DDPG algorithm model includes improving its policy network for the DDPG algorithm model; the improved policy network is used to generate corresponding actions according to the state of the distribution network, where the state of the distribution network refers to the node voltage state of the distribution network, and the action refers to a control signal for controlling the reactive power injection of the distribution network.
[0027] In this embodiment, the optimization of the structure of the improved policy network mainly focuses on introducing security constraints and enhancing the flexibility of the network to meet the requirements of voltage control, aiming to ensure stability and efficiency during the voltage control process. As Figure 1 shown, the steps for the improved policy network in this embodiment to generate corresponding actions according to the state of the distribution network include: S101, calculating and generating a weight recovery matrix and a bias recovery matrix according to the following formula: , , where, is the weight recovery matrix, represents generating an upper triangular matrix, where the elements on and below the diagonal are 0 and the elements above are 1; represents generating an upper triangular matrix, where the elements on and below the diagonal and the two rows below are 0 and the elements above are 1; is a 2-fold identity matrix, i.e., the elements on the diagonal are 2 and the other elements are 0; where is the operation of the upper triangular matrix, is the offset from the main diagonal, is the identity matrix, is the bias recovery matrix; S102, multiply the weight recovery matrix to calculate the forward weighting matrix and the negative weighting matrix : ,
[0028] where is the forward weighting matrix, is the negative weighting matrix, is a parameter; S103, calculate the forward bias matrix and the negative bias matrix : , , where is the forward bias matrix, is the negative bias matrix, is the bias reference matrix, and are the upper and lower voltage limits respectively, is the all-ones matrix; S104, multiply the forward weighting matrix by the distribution network state at time t , activate the result through the Relu activation function, then calculate the L1 norm of the result, multiply the negative weighting matrix by the distribution network state at time t , and add the result of the multiplication to the calculation result of the L1 norm as the first control signal; then calculate the differences between the forward bias matrix, the negative bias matrix and the bias reference matrix respectively, multiply the two differences by the distribution network state at time t to get the second control signal, and add the two control signals to get the final control signal for controlling the reactive power injection of the distribution network.
[0029] In this embodiment, the improved policy network first introduces the weight recovery matrix and the bias recovery matrix into the distribution network state at time t on the basis of the traditional policy network, and the goal is to restore the monotonicity of the network by adjusting the matrix structure, so as to ensure that voltage control does not produce negative feedback. By adjusting the bias term, the effectiveness and rationality of the control signal are improved. Then, further use and matrices to obtain the positive weighted matrix and the negative weighted matrix , as well as the positive bias matrix and the negative bias matrix . The goal is to control the monotonicity of the network through special matrix construction, so that the output is more stable, ensure that voltage control does not produce negative feedback, and avoid unreasonable control signals by adjusting the bias term. On this basis, the ReLU activation function is used to process the state, and further combined with the weight matrix and bias output. Applied to the weighted result, the negative values are clipped to zero to ensure the non-negative output of the network. The ReLU activation function effectively processes the non-linear characteristics in the network, avoids the vanishing gradient problem, and at the same time enhances the expression ability of the network. The ReLU activation function can be expressed as: , where is the input matrix, that is, a series of previous weight matrices and bias matrices. After passing through the ReLU activation function, the output value is non-negative. The function expression for calculating the L1 norm of the result can be expressed as: , where is the output result, C is the input result, is an operation that restricts the tensor within a given range. It is used in this network to ensure that the input result C does not have negative or overly large values, and is normalized through L1 norm regularization to ensure that the output is within a reasonable range, is a hyperparameter used to scale the bias term, controlling the size of the bias term. is the L1 norm of the input result C. Ensure that the magnitude of the bias term is within a reasonable range after normalization. Through the above operations, the bias terms in the network will remain within a reasonable range during each update, avoiding them being too large or too small.
[0030] As an alternative implementation, the improved DDPG algorithm model in this embodiment includes improving the value network of the DDPG algorithm model. The improved value network is used to generate corresponding Q-values based on the distribution network state and the actions generated by the policy network. Layer normalization combined with residual connections is added to the layers of the original value network. Layer Normalization is a regularization technique that avoids the problems of vanishing gradients or exploding gradients by normalizing the activation values of each layer. It is applied to the output of each layer to ensure that the output of each layer remains within a suitable range during training, thereby accelerating training and stabilizing the learning process. Then, residual connections are stacked on the basis of layer normalization. The main purpose of residual connections is to solve the problem of vanishing gradients in the training of deep networks and to accelerate convergence at the same time. It introduces a "skip connection" to allow the input to flow directly to subsequent layers. As Figure 2 shown, the improved value network in this embodiment includes a normalization function, a first hidden layer, a second hidden layer, and an output layer. The functional expression for generating the corresponding Q-value is: , where represents the Q-value generated by the value network according to the distribution network state at time t, the action ; is the weight matrix of the output layer; is the bias vector of the output layer, is the output of the second hidden layer, and there is: , , , where is the Leaky ReLU activation function, is the normalization operation, is the weight matrix of the second hidden layer; is the bias vector of the second hidden layer, is the output of the first hidden layer, is the bias vector of the first hidden layer, is the weight matrix of the state input, is the weight matrix of the action input, is the normalization result of the normalization function for the distribution network state at time t, is the normalization result of the normalization function for the action at time t, is the slope parameter of the negative half-axis. The purpose of the value network is to evaluate , and first normalize the input state: , .
[0031] Then, in the hidden layer of the value network, a residual connection combined with layer normalization is applied, and on this basis, the Leaky ReLU activation function is used. The Leaky ReLU activation function is used to solve the problem of "dead neurons" that may be caused by the traditional ReLU activation function. The Leaky ReLU activation function avoids the "dead zone" in the negative value region by introducing a small negative slope, enabling the neuron to maintain a certain gradient even when the input is negative: , where is the slope of the negative half-axis (usually a small constant). The Leaky ReLU activation function is applied to each layer of the network to ensure that there is a small slope in the negative value part, avoiding the problem of dead neurons. Finally, the corresponding Q value can be generated through the first hidden layer, the second hidden layer, and the output layer.
[0032] As Figure 3 shown, the training of the improved DDPG algorithm model in this embodiment includes: S201, initialize the policy network, the value network, and their respective target networks. The parameters of the target network are initialized the same as those of the original network; initialize the experience replay buffer; S202, obtain the current state of the distribution network by interacting with the environment of the distribution network, generate the corresponding action according to the current state of the distribution network through the policy network, obtain the reward and the next state of the action, and put the current state, action, reward, and the next state as an experience into the experience replay buffer; S203, perform experience sampling from the experience replay buffer based on the prioritized experience replay mechanism, including sampling N experiences from the experience replay buffer according to the priority, and the probability of each experience being sampled is: , , , where is the probability of experience i being sampled, is the priority value of experience i, is the priority adjustment factor, is the error value, is the TD error of experience i, is the reward of experience i; is the discount factor, which is used to control the influence degree of future rewards on the current value; is the target value network The state at time t+1 and the target policy generate the Q-value of the action, which is the Q-value obtained by the value network based on the state of the distribution network at time t and the action The experience replay buffer improves the sample utilization efficiency by increasing the sampling probability of key experiences, making the policy optimization more efficient; S204. Update the parameters of the policy network and the value network by backpropagation by combining the sampled experiences with the loss function respectively; S205. Soft-update the parameters of the target networks of the policy network and the value network respectively; S206. Jump to step S203 until the preset training end condition is reached.
[0033] As Figure 4 shown, in step S201 of this embodiment, the initialized policy network, value network, and their respective target networks are respectively represented as: policy network , policy target network , value network and value target network , where represents the state, represents the action, the policy network is responsible for generating the deterministic action under the state of the distribution network at time t , the policy network is defined by the parameter , the action directly interacts with the environment to determine the reactive power injection action of the distribution network; the value network is responsible for evaluating the impact (Q-value) of the state of the distribution network at time t and the action on the long-term return of the agent. The output of the value network represents the cumulative reward that can be obtained by starting to execute the action from the state of the distribution network at time t under the current policy. The value network is defined by the parameter ; the policy target network is defined by the parameter ; the value target network is defined by the parameter and adds noise to the output action of the policy network to generate exploratory actions; randomly initialize the parameters of the policy network and assign them to the parameters of the policy target network and the parameters of the value network are assigned to the value target network to ensure the stability of training, which can be expressed as: The policy target network and the value target network are used to calculate the target value , and their weights are the delayed copies (soft updates) of the main network parameters. Through the above settings, the main network is used for policy generation and value evaluation, while the target network provides stable training targets for the main network by slowly changing. Subsequently, the soft update mechanism is gradually adjusted. In this embodiment, in order to improve the sampling efficiency and learning effect, the prioritized replay buffer mechanism is additionally adopted in the DDPG algorithm. Under this mechanism, different experiences are assigned sampling priorities according to their importance (measured by TD error), so that key experiences are utilized more frequently. The storage form of the experience: After each interaction with the environment, a complete five-tuple is stored:
[0034] , where is the state of the distribution network at time t, including voltage and reactive power , is the action generated by the current policy network, such as reactive power adjustment, is the immediate reward calculated according to the cost function, is the state of the distribution network at time t + 1, that is, the next state of the system after the action is executed, is the flag indicating whether it is over. In step S202 of this embodiment, the function expression for obtaining the reward of the action is:
[0035] , where is the reward of node at time t, is the voltage deviation of node at time t, is the control input of node at time t, is the voltage deviation, is the action cost, is the balance coefficient of the voltage deviation, is the balance coefficient of the action cost, is the balance coefficient of the action cost, is the target voltage value of node , The average value, and there is: , , , Wherein, is the voltage at time t + 1, is the conductance matrix of the power grid, is the reactive power injection at time t + 1; is the influence of the external environment on the voltage, and the influence of the external environment on the voltage includes the reactive power injection of the node at time t and time t + 1; and are respectively the reactive power injections of node at time t and time t + 1, is the time interval, is the control strategy function of, is the strategy parameter of the control strategy function of node ; The functional expression of the loss function used to update the parameters of the value network in step S204 of this embodiment is: , , , Wherein, is the loss function used to update the parameters of the value network, is the expected value, is the importance sampling weight of experience i, is the probability that experience i is sampled, is the size of the experience replay buffer, is the Q value obtained by the value network based on the current state and action as well as parameters ; is the target value of the current experience, is the parameter used to adjust the weight; is the discount factor, which is used to control the influence degree of future rewards on the current action; is the Q value obtained by the target network of the value network according to the state at time t + 1 and action ; is the action generated by the policy network according to the state at time t + 1 .
[0036] In step S204 of this embodiment, the functional expression of the loss function used to update the parameters of the policy network is: , , wherein, is the loss function adopted for the parameters of the updated policy network, is the expected value, is the Q value obtained by the value network based on the distribution network state at time t and the action is the action obtained by the policy network according to the current state and the parameters ; is the weight parameter of entropy regularization, which is used to adjust the trade-off between exploration and exploitation in policy optimization; is the entropy of the policy, is the probability distribution of the policy network taking the action a under the distribution network state at time t.
[0037] As Figure 4 shown, in step S205 of this embodiment, the function expressions of the parameters of the target networks of the soft-updated policy network and the value network are: , wherein, is the parameter of the target network after soft update, is the soft update, is the parameter of the target network corresponding to the policy network or the value network, is the parameter of the target network before soft update; is the step size of the soft update, which controls the update speed of the target network.
[0038] The action generated by the policy network is deterministic and lacks randomness. Therefore, in order to enhance the exploration ability of the agent in the action space, the improved algorithm introduces the Ornstein-Uhlenbeck (OU) noise process. The OU noise has temporal correlation and the generated noise is smoother, which helps to achieve more effective exploration in continuous control tasks. As Figure 4 shown, when initializing the policy network, the value network and their respective target networks in step S201 of this embodiment, it includes adding noise to the output action of the policy network to generate exploratory actions, and its function expression is: , , wherein, is the deterministic action under the current state , is the policy network, and respectively and the noise values at the moments, is the mean value of the noise; is the mean reversion speed, which is used to control the rate at which the noise gradually reverts to the mean; is the random amplitude of the noise; refers to white noise sampled from a standard normal distribution. By introducing OU noise, the agent (the improved DDPG algorithm model) can conduct extensive exploration in the initial stage of training, and at the same time, the noise amplitude gradually decreases over time, which helps to focus on high-quality actions in the later stage of training.
[0039] Considering that in the voltage control problem, the design of the reward function has an important impact on the performance of the reinforcement learning algorithm. To verify the effectiveness of the improved DDPG algorithm in voltage control, this embodiment conducts a comparative analysis of the reward functions of the improved DDPG algorithm model and the original DDPG algorithm model. Based on the original DDPG algorithm model, the improved DDPG algorithm model optimizes the design of the reward function for the voltage control scenario, enabling it to better balance voltage stability and control efficiency. Specifically, the reward function of the improved DDPG algorithm model not only considers the penalty for voltage deviation, but also introduces additional constraints such as the smoothness of voltage fluctuations and the economy of control actions, thus more comprehensively reflecting the requirements of actual power grid operation. By comparing the performance of the reward functions of the two algorithms in different voltage disturbance scenarios, the advantages of the improved DDPG algorithm model in the voltage control task can be intuitively evaluated. Figure 5 This is the comparison of the average rewards of the improved DDPG algorithm model and the original DDPG algorithm model in this embodiment, which shows the comparison of the improved DDPG algorithm model in this embodiment and the original DDPG algorithm model (DDPG) in terms of the convergence, stability, and final control effect of the reward function. In the voltage control problem, the DDPG algorithm is usually used to train the agent to effectively control the voltage to reach the target. It is observed that the reward value of the original DDPG (blue curve) performs poorly in the initial stage, but as the training progresses, the reward value gradually rises and tends to be stable. While the optimized DDPG (orange curve) shows better performance in the initial stage of training, and the reward value increases more smoothly and finally tends to a higher level. From Figure 5Several advantages of the improved DDPG algorithm can be reflected as follows: (1) Better initial performance: The improved DDPG algorithm model in this embodiment rapidly increased the reward value around 12 times in the initial stage of training, indicating that through algorithm optimization, the agent can adapt to the environment faster and take appropriate actions in the voltage control task. This is particularly important for real-time control in power systems and can effectively reduce system fluctuations and instability. (2) Improved stability: The improved DDPG algorithm model shows a smoother reward curve. This indicates that the optimized algorithm can learn the voltage control strategy more stably, avoiding the volatility in some training rounds of the original algorithm, thereby improving the reliability of the system. (3) Higher final reward: The reward value of the improved DDPG algorithm model in the later stage of training is significantly higher than that of the original DDPG algorithm model, starting from the first step. This shows that the optimized algorithm can not only achieve good performance in the initial stage but also continuously improve its control strategy, thus obtaining better control effects in the voltage control task. (4) Higher convergence speed: The improved DDPG algorithm model in this embodiment seems to converge faster and reaches a higher reward value in fewer training rounds. This means that through optimization measures, the improved DDPG algorithm model in this embodiment can obtain a more efficient voltage control strategy in a shorter time. It can be seen that in the voltage control problem, the improved DDPG algorithm model can help improve system stability and control efficiency, reduce voltage fluctuations, and respond more quickly to external changes, thereby enhancing the overall performance of the power system.
[0040] To verify the performance advantages of the distribution network security-constrained voltage regulation method based on deep reinforcement learning in this embodiment in voltage control, experiments were carried out on the IEEE 13-bus test feeder in this embodiment, focusing on the voltage control effects of nodes 2, 7, and 9. Figure 6 This is the voltage control effect of the improved DDPG algorithm model in this embodiment and the existing safe-DDPG algorithm model on node 2. Figure 7 This is the voltage control effect of the improved DDPG algorithm model in this embodiment and the existing safe-DDPG algorithm model on node 7. Figure 8 This is the voltage control effect of the improved DDPG algorithm model in this embodiment and the existing safe-DDPG algorithm model on node 9. At Figures 6 to 8In this embodiment, the improved DDPG algorithm model can stabilize the voltage near 1.05 p.u. in a relatively short time and maintain it near this target voltage value. In contrast, the voltage fluctuations of the safe-DDPG algorithm model are relatively large, especially at nodes 2 and 7. The optimized algorithm shows smaller voltage fluctuations. The optimized algorithm is more robust and less likely to let the voltage exceed the set upper limit (1.05 p.u.). In the graph of node 9, the optimized algorithm ensures that the voltage always remains below the specified upper limit (1.05 p.u.), while the safe-DDPG algorithm model shows a certain degree of voltage overrun in the initial stage, especially in the first 200 time steps. This indicates that the optimized algorithm can better control the voltage and reduce the risk of system overrun, which is crucial for voltage control, especially in avoiding equipment damage or system instability. The improved DDPG algorithm model in this embodiment makes the voltage value gradually tend to be stable and remain near the target voltage, indicating that the algorithm is more stable in dynamic control. Although the original algorithm can finally stabilize the voltage, its convergence process is relatively slow and prone to fluctuations. In summary, the improved DDPG algorithm model in this embodiment can converge to the target voltage more quickly in the voltage control task compared with the safe-DDPG algorithm model; it can more stably maintain the voltage within the set range, reducing the overrun risk; it improves the stability and control accuracy of the system, especially when facing multi-node voltage control. It can be seen that the improved DDPG algorithm model in this embodiment can control the voltage within the safe range (i.e., below 1.05 p.u.) more quickly and stably by optimizing the policy network, value network, and the design of the reward function. Under the same high-voltage scenario, the improved DDPG algorithm model in this embodiment can adjust the voltages of nodes 2, 7, and 9 from the state exceeding 1.05 p.u. to the safe range faster than the safe-DDPG algorithm model, and at the same time shows better convergence and stability, demonstrating the significant advantages of the improved DDPG algorithm model in the voltage control task.
[0041] In summary, the method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning in this embodiment provides higher reliability and stability guarantees for voltage control by introducing a security constraint and a stability optimization mechanism. The core of our method is to ensure that the control signal is always within the safe range through an improved policy network design and a dynamic adjustment mechanism, while significantly enhancing the convergence speed and adaptability of the algorithm. Experimental results show that the method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning in this embodiment can not only effectively address the voltage fluctuation problems caused by high-proportion photovoltaic penetration and load fluctuations, but also shorten the response time by more than one-third, while always achieving voltage stability. The method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning in this embodiment provides an efficient and reliable solution for voltage control in power systems and lays a foundation for future research combining other control theory optimization principles.
[0042] In addition, the present invention also provides a secure-constrained voltage regulation system for a distribution network based on deep reinforcement learning, including a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning.
[0043] In addition, the present invention also provides a computer-readable storage medium storing a computer program or instruction, which is programmed or configured to execute the method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning through a processor.
[0044] In addition, the present invention also provides a computer program product including a computer program or instruction, which is programmed or configured to execute the method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning through a processor.
[0045] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present invention can be in the form of a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks. These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0046] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the concept of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as within the protection scope of the present invention.
Claims
1. A method for security-constrained voltage regulation of a distribution network based on deep reinforcement learning, characterized in that It includes the following steps: training an improved DDPG algorithm model, and the improved DDPG algorithm model includes improving the policy network of the DDPG algorithm model. The improved policy network is used to generate corresponding actions according to the distribution network state, where the distribution network state refers to the node voltage state of the distribution network, and the action refers to the control signal for controlling the reactive power injection of the distribution network. The improved policy network generating corresponding actions according to the distribution network state includes: S101, calculating and generating a weight recovery matrix and a bias recovery matrix according to the following formula: , , Among them, is the weight recovery matrix, means to generate an upper triangular matrix, where the elements on and below the diagonal are 0 and the elements above are 1; means to generate an upper triangular matrix, where the elements on and below the diagonal and the two rows below are 0 and the elements above are 1; is a 2-fold identity matrix, that is, the elements on the diagonal are 2 and the other elements are 0; where is the operation of the upper triangular matrix, is the offset from the main diagonal, is the identity matrix, is the bias recovery matrix; S102, restore the weight matrix Calculate the forward weighted matrix according to the following formula and the negative weighted matrix : , Among them, is the positive weighting matrix, is the negative weighting matrix, is the parameter; S103, calculate the forward bias matrix according to the following formula and the reverse bias matrix : , , Among them, is the positive bias matrix, is the negative bias matrix, is the bias reference matrix, and are the upper and lower voltage limits respectively, is the all-ones matrix; S104, multiply the positive weighting matrix by the distribution network state at time t , activate the result through the Relu activation function, then calculate the L1 norm of the result, multiply the negative weighting matrix by the distribution network state at time t , and add the result to the calculation result of the L1 norm as the first control signal; then calculate the differences between the positive bias matrix, the negative bias matrix and the bias reference matrix respectively, multiply the two differences by the distribution network state at time t respectively to obtain the second control signal, and add the two control signals to obtain the final control signal for controlling the reactive power injection of the distribution network.
2. The method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning according to claim 1, wherein The improved DDPG algorithm model includes improving the value network of the DDPG algorithm model. The improved value network is used to generate corresponding Q values according to the distribution network state and the actions generated by the policy network. The improved value network includes a normalization function, a first hidden layer, a second hidden layer, and an output layer. The function expression for generating corresponding Q values is: , Among them, represents the Q value generated by the value network according to the state of the distribution network at time t , action generated, is the weight matrix of the output layer; is the bias vector of the output layer, is the output of the second hidden layer, and there is: , , , Among them, is the Leaky ReLU activation function, is the normalization operation, is the weight matrix of the second hidden layer; is the bias vector of the second hidden layer, is the output of the first hidden layer, the bias vector of the first hidden layer, is the weight matrix of the state input, is the weight matrix of the action input, is the normalization result of the distribution network state at time t by the normalization function , is the normalization result of the action at time t by the normalization function , is the slope parameter of the negative half-axis.
3. The method for voltage regulation of a distribution network with security constraints based on deep reinforcement learning according to claim 1, wherein The training using the improved DDPG algorithm model includes: S201, initializing the policy network, the value network, and their respective target networks. The parameters of the target network are initialized the same as those of the original network; initializing the experience replay buffer; S202, obtaining the current state of the distribution network by interacting with the environment of the distribution network, generating corresponding actions according to the current state of the distribution network through the policy network, obtaining the reward of the action and the next state, and putting the current state, action, reward, and next state as an experience into the experience replay buffer; where the function expression for obtaining the reward of the action is: , wherein, is the reward of the node at time t, is the voltage deviation of the node at time t, is the control input of the node at time t, is the voltage deviation, is the action cost, is the balance coefficient of the voltage deviation, is the balance coefficient of the action cost, is the node the target voltage value of the average value of, and there is: , , , Among them, is the voltage at time t+1, is the conductance matrix of the power grid, is the reactive power injection at time t+1; is the influence of the external environment on the voltage, and the influence of the external environment on the voltage includes the reactive power injection of the node at time t and time t+1; and are respectively the reactive power injections of node at time t and time t+1, is the time interval, is the control strategy function of is the strategy parameter of the control strategy function of node ; S203, performing experience sampling from the experience replay buffer based on the prioritized experience replay mechanism, including sampling N experiences from the experience replay buffer according to the priority, and the probability of each experience being sampled is: , , , where, is the probability that experience i is sampled, is the priority value of experience i, is the priority adjustment factor, is the error value, is the TD error of experience i, is the reward of experience i; is the discount factor, used to control the influence degree of future rewards on the current value; is the target value network for the state at time t + 1 and the Q value of the action generated by the target policy is the Q value obtained by the value network based on the distribution network state at time t and the action and the action is the Q value obtained; S204, updating the parameters of the policy network and the value network respectively through backpropagation by combining the sampled experiences with the loss function; S205, softly updating the parameters of the target networks of the policy network and the value network respectively; S206, jumping to step S203 until the preset training end condition is reached.
4. The method for security-constrained voltage regulation of a distribution network based on deep reinforcement learning according to claim 3, wherein In step S204, the function expression of the loss function used to update the parameters of the value network is: , , , Among them, is the loss function used to update the parameters of the value network, is the expected value, is the importance sampling weight of experience i, is the probability that experience i is sampled, is the size of the experience replay buffer, is the value network based on the current state and action and parameters to obtain the Q value, is the target value of the current experience, is the parameter used to adjust the weight; is the discount factor, which is used to control the influence degree of future rewards on the current action; is the target network of the value network according to the state at time t + 1 and action to obtain the Q value, is the action generated by the policy network according to the state at time t + 1 5. The method for regulating the voltage of a distribution network with security constraints based on deep reinforcement learning according to claim 3, wherein In step S204, the function expression of the loss function used to update the parameters of the policy network is: , , where, is the loss function used to update the parameters of the policy network, is the expected value, is the Q value obtained by the value network based on the distribution network state at time t and the action ; is the action obtained by the policy network according to the current state and the parameters ; is the weight parameter of entropy regularization, which is used to adjust the trade-off between exploration and exploitation in policy optimization; is the entropy of the policy, is the probability distribution of the policy network taking action a under the distribution network state at time t .
6. The method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning according to claim 3, wherein In step S205, the function expression for softly updating the parameters of the target networks of the policy network and the value network respectively is: , Among them, are the parameters of the target network after soft update, is the soft update, are the parameters of the target network corresponding to the policy network or value network, are the parameters of the target network before soft update; is the step size of the soft update, which controls the update speed of the target network.
7. The method for secure-constrained voltage regulation of a distribution network based on deep reinforcement learning according to claim 3, wherein When initializing the policy network, the value network, and their respective target networks in step S201, it includes adding noise to the output action of the policy network to generate exploratory actions, and the function expression is: , , Among them, is the deterministic action in the current state ; is the policy network and are respectively and the noise values at time ; is the mean reversion speed, used to control the rate at which the noise gradually reverts to the mean is the random amplitude of the noise refers to white noise sampled from the standard normal distribution 8. A distribution network security-constrained voltage regulation system based on deep reinforcement learning, comprising a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the distribution network security-constrained voltage regulation method based on deep reinforcement learning according to any one of claims 1 to 7.
9. A computer-readable storage medium storing a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the distribution network security-constrained voltage regulation method based on deep reinforcement learning according to any one of claims 1 to 7 through a processor.
10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or instruction is programmed or configured to execute the distribution network security-constrained voltage regulation method based on deep reinforcement learning according to any one of claims 1 to 7 through a processor.
Citation Information
Patent Citations
Feature filtering defense method for deep reinforcement learning model
CN111600851A
Power distribution network voltage control method and system based on improved DDPG algorithm
CN119834265A