Multi-source energy storage power distribution network voltage regulation and control method based on safety reinforcement learning algorithm
By constructing state space and action space, combining Monte Carlo simulation and PPO algorithm, the intelligent agent is trained to output the optimal control strategy, which solves the problem of insufficient perception of the real-time status of the distribution network in the existing voltage control method, and realizes precise voltage control and safe and stable operation of the system.
Patent Information
- Application Number
- CN202511124774.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing voltage control methods lack the ability to accurately perceive and flexibly respond to the real-time operating status of the distribution network, making it difficult to adjust the voltage quickly and effectively. This causes the node voltage to exceed the qualified range, affecting the power supply quality, and failing to balance system safety and economy.
A voltage control method for a multi-source energy storage distribution network based on a secure reinforcement learning algorithm is adopted to construct a state space and action space. Combined with Monte Carlo simulation and the PPO algorithm, the intelligent agent is trained to output the optimal control strategy. Taking into account the device status, historical information and uncertainty factors, a multi-objective reward function is constructed to optimize the control effect.
It achieves precise control of distribution network voltage, improves power supply quality, reduces equipment operating costs, ensures safe and stable operation of the system, and improves the overall performance of the power system.
Smart Images

Figure CN120638362A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system distribution network, and in particular to a voltage control method for a multi-source energy storage distribution network based on a security reinforcement learning algorithm. Background Art
[0002] Existing technologies for voltage regulation in distribution networks have limitations. Traditional voltage regulation methods mostly rely on fixed rules and preset control strategies, lacking the ability to accurately perceive and flexibly respond to real-time changes in the distribution network's operating status. For example, these traditional methods struggle to quickly and effectively adjust voltage in the face of uncertainties such as random load fluctuations and intermittent output from distributed power sources. This can easily cause node voltages to exceed acceptable ranges, impacting power supply quality.
[0003] Some existing control technologies fail to comprehensively consider the state space when building models. They often focus only on basic parameters such as node voltage amplitude, active power, and reactive power, ignoring important factors such as device status and historical information. This prevents the models from fully capturing the dynamic trends of the system, and control decisions lack accuracy and foresight. Furthermore, existing technologies also fall short in modeling the impact of actions on the system. Operations such as transformer tap adjustment, capacitor switching, and distributed generation reactive power output adjustment fail to comprehensively and accurately analyze their combined impact on the power distribution and voltage levels of the entire distribution network, making it difficult to achieve optimal control results. Existing technologies also fail to effectively balance the safety and economic efficiency of system operation. Some control methods focus solely on voltage control effectiveness, ignoring equipment operating costs and requiring frequent adjustments, resulting in wasted resources. Other methods, while taking into account some safety constraints, are incomplete and cannot fully guarantee safe and stable system operation under various complex operating conditions.
[0004] On this basis, we propose a voltage control method for multi-source energy storage distribution network based on a secure reinforcement learning algorithm. Summary of the Invention
[0005] In view of the above shortcomings in the existing technology, the purpose of the present invention is to propose a multi-source energy storage distribution network voltage control method based on a security reinforcement learning algorithm. The method aims to achieve precise control of the distribution network voltage by building a model, selecting an algorithm and training optimization, thereby ensuring the safe and stable operation of the power system and improving the power supply quality.
[0006] In order to achieve the above objectives, the present invention adopts the following technical solutions:
[0007] A voltage control method for a multi-source energy storage distribution network based on a secure reinforcement learning algorithm includes the following steps:
[0008] S1. Construct the basic framework of the distribution network voltage control model. Each node in the distribution network is considered as a state node in the model. The electrical connection relationship between nodes constitutes the state space. The voltage control operation is used as the action space of the model.
[0009] S2. Establish a state transition probability matrix and solve it using the Newton-Raphson method. Consider the impact of the action on the system and, through Monte Carlo simulation, account for uncertainties such as load fluctuations and distributed power output. After multiple simulations, calculate the state transition probabilities and construct a complete matrix to quantify the probability relationship between "current state-action-next state."
[0010] S3. Define a reward function that integrates multiple objectives, including voltage quality, equipment operating costs, and system safety.
[0011] S4. Reinforcement learning algorithm selection and training: using the constructed model as the environment, inputting preprocessed operating data, using the PPO algorithm to train the intelligent agent, and outputting the optimal control strategy after convergence.
[0012] Furthermore, the state space considers not only node voltage amplitude, active power, and reactive power, but also device status and historical information.
[0013] Furthermore, in addition to the node voltage amplitude, active power, and reactive power, the state space also considers the device status and historical information, and the state vector Expressed as:
[0014] ;
[0015] For the node at time The voltage amplitude; For the node at time Active power; Node at time Reactive power; For the transformer at time The tap position; For the capacitor at time The switch status; It is historical state information, which is formed by splicing the states of multiple time steps in the past and is used to capture the dynamic change trend of the system; Indicates transpose.
[0016] In S1, the action space contains discrete and continuous operations, and the action vector Expressed as:
[0017] ;
[0018] is the adjustment value of the transformer tap, which is , It means downgrading one gear. Indicates that it remains unchanged. Indicates moving up one gear; is the change in the capacitor switching state, which is , Indicates removal, Indicates that it remains unchanged, Indicates investment; is the reactive output adjustment of the distributed generation, which is a continuous value and must meet the upper and lower limit constraints;
[0019] In S2, the method of establishing the state transition probability matrix is:
[0020] S211. Calculate state transition based on power flow equation;
[0021] There is a relationship between the node injection current vector, the node voltage vector, and the node admittance matrix: ; Inject current vector into the node, is the node admittance matrix, is the node voltage vector;
[0022] The active power and reactive power of a node are related to the node voltage amplitude, phase angle, and voltage parameters of adjacent nodes. The expressions are:
[0023] ;
[0024] ;
[0025] in, For nodes Active power; For nodes Reactive power; For nodes The voltage amplitude; For nodes The voltage amplitude of the node is a node adjacent nodes; For nodes The set of adjacent nodes of and Node and The conductance and susceptance between is a node and The voltage phase angle difference between , For nodes The voltage phase angle, For nodes The voltage phase angle;
[0026] The Newton-Raphson method is used to solve the above tidal flow equation, which is expressed as a nonlinear equation system: , is the state variable vector to be solved, where is the state variable vector to be solved, is the number of nodes in the distribution network, Except for the balance node The voltage phase angle of each node, Except for the balance node The voltage amplitude of each node; select a node as the balance node, and its voltage amplitude and voltage phase angle It is known that the iterative formula of the Newton-Raphson method is:
[0027] ; ;
[0028] in, is the Jacobian matrix, whose elements are obtained by taking partial derivatives of the state variables according to the power flow equation; It is The state variables at the iteration vector of It is The iterative process continues until the convergence condition is met. ,in is the preset convergence accuracy; It is The residual of the power flow equation at the iteration;
[0029] S212. Modeling the impact of actions on the system:
[0030] When executing the current action vector When the load is high, it will have different impacts on the operating status of the distribution network;
[0031] Impact of transformer tap adjustment: Changes in tap position will directly affect the transformer's turns ratio;
[0032] The change of the transformation ratio changes the voltage relationship between the nodes on both sides of the transformer and the equivalent impedance of the line, which is finally reflected in the node admittance matrix. On the changes of elements;
[0033] Impact of capacitor switching:
[0034] When Change in the switching state of the capacitor bank When the reactive power injection of the node increases; when the Change in the switching state of the capacitor bank When the reactive power injection is reduced, the change of reactive power will affect the node voltage amplitude and phase angle, and further affect the power distribution and voltage level of the entire distribution network;
[0035] For distributed power sources, the reactive output adjustment directly changes the reactive power injection of the node to which the distributed power source is connected, which will cause the voltage of the node to change. Due to the electrical connection relationship of the distribution network, the change of the node voltage will propagate in the network, affecting the voltage and power distribution of other nodes.
[0036] S213. Considering the uncertainty factors, there are many uncertainties in the actual operation of the distribution network. To accurately reflect the impact of uncertainty on the state transition probability, the Monte Carlo simulation method is used.
[0037] S2131. Determine the probability distribution of uncertainties. When modeling load fluctuations, the active and reactive power at the node follow a normal distribution. For distributed power sources, a Beta distribution is used.
[0038] S2132. Monte Carlo simulation process, for a given current state and current action vector , perform multiple Monte Carlo simulations; in each simulation, randomly generate a set of load and distributed power parameter values according to the probability distribution of the determined uncertainty factors; then, based on the randomly generated parameters, combine the power flow equation to calculate the execution vector of the current action The new state after
[0039] S2133. Calculate the state transition probability and calculate the current action vector from the statistical simulation results. The number of times the state is transferred to each possible state is calculated, and the state transition probability is obtained through a large number of simulations. The probability of transferring to all possible states, the probabilities of all possible states together constitute a row of elements in the state transition probability matrix; repeat the above process for all possible states and action vectors to construct a complete state transition probability matrix.
[0040] Furthermore, the voltage quality reward and punishment are based on node voltage deviation and weight; the equipment operation cost penalty includes the adjustment cost of transformers, capacitors, and distributed power sources; and the system safety reward and punishment are aimed at violations of safety constraints.
[0041] In S3, the steps for constructing the reward function are:
[0042] S321. Define voltage quality reward / penalty items. The voltage quality reward / penalty items are expressed as:
[0043] ;
[0044] in, is the voltage quality reward / penalty item; is the node voltage deviation; is the weight of the node, which is set according to the importance of the node. The weight of important nodes is larger;
[0045] S322. Equipment operating cost penalty item, taking into account the operating costs of transformer tap adjustment, capacitor switching, and distributed generation reactive power output adjustment, the equipment operating cost penalty item is:
[0046] ;
[0047] is the equipment operating cost penalty item; Adjust the number of transformer taps; is the number of capacitor switching times; The number of times the reactive output of distributed generation is adjusted; is the unit cost factor for the number of transformer tap adjustments; is the unit cost coefficient of the capacitor switching times; is the unit cost coefficient of the reactive power output adjustment times of the distributed power source;
[0048] S323. System security reward / penalty items. To ensure the safe operation of the system, a system security indicator is introduced. The system security reward / penalty items are expressed as:
[0049] ;
[0050] Reward / penalty items for system security; is the number of safety constraints; It is The weight of each security constraint is used to adjust the importance of different security constraints; is the safety constraint function, Indicates a violation of the safety constraint; when the system violates the safety constraint, A negative value penalizes the agent and guides it to choose actions that satisfy safety constraints;
[0051] S324. Define a reward function that comprehensively considers the voltage quality, equipment operating costs, and system security of the distribution network. The reward function is expressed as:
[0052] ;
[0053] represents the reward function.
[0054] In S4, the agent and environment interaction training method:
[0055] S411. Initialization: Initialize the parameters of the policy network and the value network, and set the experience replay buffer. The experience replay buffer is used to store the experience data generated by the interaction between the agent and the environment. The role of the experience replay buffer is to break the correlation of the data.
[0056] S412. Training loop;
[0057] S4121. In the initial stage, the agent starts from the initial state and selects actions according to the current strategy network. Greedy strategy, randomly select actions with probability, with The probability of randomly selecting an action is The probability of choosing the action that maximizes the value function; is a hyperparameter;
[0058] The execution selects actions based on the current strategy network. The voltage control model constructed in S1 updates the operating status of the distribution network based on the actions, and feeds back rewards and new status. The agent stores the experience in the experience replay buffer. As the training progresses, the The value of , which enables the agent to transition from random exploration to making decisions based on the learned knowledge;
[0059] S4122. Parameter update phase, when the experience replay buffer accumulates to the set threshold, the parameter update begins, randomly sampling a batch of experience, randomly sampling a batch of experience from the experience replay buffer;
[0060] Update the policy network parameters according to the clipping objective function of the PPO algorithm:
[0061] ;
[0062] represents the loss function; To initialize the policy network; is the minimum function; are the initialization strategy network parameters; are the policy network parameters before updating; is the cropping parameter; Is the advantage function, which represents the action under the current strategy In state degree of advantage; is the policy network parameter before update strategic network; Represents the policy network before update The state sampled and actions Perform expectation calculations; is the truncation function;
[0063] By updating the policy network parameters, the new policy is updated in a more optimal direction while ensuring stability;
[0064] Update the value network parameters by minimizing the loss function of the value network:
[0065] ;
[0066] in, Is the loss function of the value function (ValueFunction), used to optimize the parameters of the value network ; Is a discount factor used to weigh the importance of future rewards, usually with a value between between; are the parameters of the value network; For the value network state Estimated value of The agent is in state Execute an action After that, the reward value of the environment feedback; the value network performs the action The new state reached later Estimated value of Indicates the state obtained by sampling ,award and the new state Perform expectation calculations;
[0067] Adjust the parameters of the value network through the back-propagation algorithm to make it estimate the state value more accurately;
[0068] The soft update method is used to update the policy network and value network parameters to make the training more stable. The update formula is:
[0069] ;
[0070] ;
[0071] is the soft update coefficient, are the updated initialization strategy network parameters; are the parameters of the updated value network;
[0072] The soft update method makes the update of network parameters smoother, avoiding training instability caused by too fast parameter updates;
[0073] S413. Model optimization and convergence. During the training process, continuously monitor the performance indicators of the policy network and value network;
[0074] Average Reward: Calculates the average reward received by the agent during a specific training step. An upward trend in average reward indicates that the agent is learning a more effective control strategy. A higher reward value indicates that the agent's decision-making can achieve better benefits in the current environment.
[0075] The voltage qualification rate measures the percentage of distribution network node voltages within the qualified range within a specific training step. An increase in the voltage qualification rate means that the agent's control actions can better ensure voltage quality and meet user electricity needs.
[0076] Convergence judgment: When the fluctuation of the average reward and voltage qualification rate within continuous specific training steps is less than a certain threshold, the model is considered to have converged; at this time, the trained reinforcement learning voltage control model can adapt to the dynamic changes of the distribution network, accurately perceive the operating status of the distribution network, and output the corresponding optimal control action to ensure voltage quality and system safety.
[0077] In summary, due to the adoption of the above technical solution, the beneficial technical effects of the invention are:
[0078] The control method of this invention incorporates device status and historical information into the state space and uses Monte Carlo simulation to account for uncertainty, enabling the model to accurately perceive real-time operational state changes in the distribution network. During training, the intelligent agent continuously learns the optimal control actions under different states. Compared to traditional methods, it can more quickly and accurately respond to load fluctuations and changes in distributed power supply output, effectively ensuring voltage stability, significantly improving voltage compliance, and meeting user demands for high-quality electricity.
[0079] By introducing a penalty term for equipment operation costs into the reward function, the system comprehensively considers the operational costs of adjusting transformer taps, switching capacitors, and adjusting the reactive output of distributed generation units. When learning the optimal control strategy, the intelligent agent automatically balances control effectiveness with operational costs, avoiding unnecessary equipment operations, thereby reducing the overall cost of distribution network operation and ensuring optimal resource utilization.
[0080] The system's safety reward / penalty settings encourage intelligent agents to strictly adhere to safety constraints during decision-making. By penalizing violations of safety constraints, agents are guided to choose actions that meet safety requirements, effectively avoiding safety issues such as line overloads and voltage overruns. This enhances the reliability and stability of power system operations, ensuring their safe and stable operation.
[0081] By comprehensively considering voltage quality, equipment operating costs, and system safety in constructing a reward function and employing the PPO algorithm for training and optimization, the control model achieves multi-objective balanced optimization. The trained model not only effectively regulates voltage but also reduces operating costs while ensuring system safety. This comprehensively improves the overall performance of distribution network voltage control and provides strong support for the efficient operation of the power system. BRIEF DESCRIPTION OF THE DRAWINGS
[0082] Figure 1 This is a flow chart of a voltage control method for a multi-source energy storage distribution network based on a secure reinforcement learning algorithm;
[0083] Figure 2 Flowchart for establishing state transition probability matrix;
[0084] Figure 3 Flowchart for building a reward function;
[0085] Figure 4 Flowchart of the method for training an intelligent agent using the PPO algorithm. DETAILED DESCRIPTION
[0086] In order to make the purpose, technical solutions and advantages of the invention more clear, the invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described here are only used to explain the invention and are not used to limit the invention.
[0087] like Figure 1 As shown, a voltage control method for a multi-source energy storage distribution network based on a secure reinforcement learning algorithm includes the following steps:
[0088] S1. Construct the basic framework of the distribution network voltage control model. Each node in the distribution network is considered as a state node in the model. The electrical connections between nodes constitute the state space. Voltage control operations (such as adjusting transformer taps and switching capacitors) serve as the model's action space.
[0089] S2. Build a state transition probability matrix and solve it using the Newton-Raphson method. This matrix takes into account the impact of actions on the system (e.g., transformer tap adjustment, which changes the turns ratio). Then, through Monte Carlo simulation, consider uncertainties such as load fluctuations (normal distribution) and distributed power output (beta distribution). After multiple simulations, the state transition probabilities are statistically analyzed to construct a complete matrix that quantifies the probability relationship between "current state-action-next state."
[0090] S3. Define a reward function that integrates multiple objectives, combining voltage quality, equipment operating costs, and system safety. Voltage quality rewards and penalties are based on node voltage deviations and weights; equipment operating cost penalties include the adjustment costs of transformers, capacitors, and distributed generation (DGs); and system safety rewards and penalties target violations of safety constraints.
[0091] S4. Reinforcement learning algorithm selection and training: using the constructed model as the environment, inputting preprocessed operating data, using the PPO algorithm to train the intelligent agent, and outputting the optimal control strategy after convergence.
[0092] In addition to considering node voltage amplitude, active power, and reactive power, the state space also considers device status and historical information. The state space is expressed as:
[0093] In S1, the state space considers not only the node voltage amplitude, active power, and reactive power, but also the device status and historical information. The state vector is expressed as:
[0094] ;
[0095] in:
[0096] ;
[0097] ;
[0098] ;
[0099] ;
[0100] ;
[0101] For the node at time The voltage amplitude; For the node at time Active power; Node at time Reactive power; For the transformer at time The tap position; For the capacitor at time The switch status; It is historical state information, which is formed by splicing the states of multiple time steps in the past and is used to capture the dynamic change trend of the system; For nodes At the moment The voltage amplitude; For nodes At the moment Active power; node At the moment Reactive power; For the Transformer at time The tap position; For the The capacitor group at time The switch status (0 means cut off, 1 means on); is the number of grid nodes; is the number of transformers; is the number of capacitors; It is historical state information, which is formed by splicing the states of multiple time steps in the past and is used to capture the dynamic change trend of the system.
[0102] In S1, the action space contains discrete and continuous operations, and the action vector Expressed as:
[0103] ;
[0104] in:
[0105] ;
[0106] ;
[0107] ;
[0108] in, For the The adjustment value of the transformer tap is , It means downgrading one gear. Indicates that it remains unchanged. Indicates moving up one gear; For the The change in the switching state of the capacitor group is , Indicates removal, Indicates that it remains unchanged, Indicates investment; For the The reactive output adjustment of a distributed power source is a continuous value, which must meet the upper and lower limit constraints, namely:
[0109] ;
[0110] in, For the The lower limit of the reactive output adjustment of a distributed power source, For the The upper limit of the reactive output adjustment of each distributed power source;
[0111] Said S2 includes:
[0112] S211. Calculate state transitions based on the power flow equation. The relationship between the node injection current vector, the node voltage vector, and the node admittance matrix is: ; Inject current vector into the node, is the node admittance matrix, is the node voltage vector;
[0113] The active power and reactive power of a node are related to the node voltage amplitude, phase angle, and voltage parameters of adjacent nodes. The expressions are:
[0114] ;
[0115] ;
[0116] in, For nodes Active power; For nodes Reactive power; For nodes The voltage amplitude; For nodes The voltage amplitude of the node is a node adjacent nodes; For nodes The set of adjacent nodes of and Node and The conductance and susceptance between is a node and The voltage phase angle difference between , For nodes The voltage phase angle, For nodes The voltage phase angle;
[0117] The Newton-Raphson method is used to solve the above tidal flow equation, which is expressed as a nonlinear equation system: ,in is the state variable vector to be solved, is the number of nodes in the distribution network; Except for the balance node The voltage phase angle of each node, Except for the balance node The voltage amplitude of each node; select a node as the balance node, and its voltage amplitude and voltage phase angle It is known that the iterative formula of the Newton-Raphson method is:
[0118] ; ;
[0119] in, is the Jacobian matrix, whose elements are obtained by taking partial derivatives of the state variables according to the power flow equation; It is The state variable vector of the iteration; It is The iterative process continues until the convergence condition is met. ,in is the preset convergence accuracy; It is The residual of the power flow equation at the iteration.
[0120] S212. Modeling the impact of actions on the system:
[0121] When executing action vector When the load is high, it will have different impacts on the operating status of the distribution network;
[0122] Impact of transformer tap adjustment: Changes in tap position will directly affect the transformer's turns ratio;
[0123] The change of the transformation ratio changes the voltage relationship between the nodes on both sides of the transformer and the equivalent impedance of the line, which is finally reflected in the node admittance matrix. For example, for a step-down transformer, if the tap is adjusted up one gear, the secondary voltage will decrease, and the injected current and power distribution of the corresponding node will also change;
[0124] Impact of capacitor switching:
[0125] when (when capacitor is put into operation), the reactive power injection to the node will increase; when When the capacitor is removed, the reactive power injection is reduced accordingly. This change in reactive power will affect the node voltage amplitude and phase angle, and further affect the power distribution and voltage level of the entire distribution network.
[0126] For distributed power sources, the reactive output adjustment directly changes the reactive power injection of the node to which the distributed power source is connected, which will cause the voltage of the node to change. Due to the electrical connection relationship of the distribution network, this change will propagate in the network and affect the voltage and power distribution of other nodes.
[0127] S213. Considering uncertainty factors, there are many uncertainties in actual distribution network operation, such as random load fluctuations and intermittent output of distributed power sources. To accurately reflect the impact of uncertainty on state transition probabilities, a Monte Carlo simulation method is used.
[0128] S2131. Determine the probability distribution of uncertainties. When modeling load fluctuations, the active and reactive power at the node follow a normal distribution. For distributed power sources, a Beta distribution is used.
[0129] S2132. Monte Carlo simulation process, for a given current state and action vector , perform multiple Monte Carlo simulations; in each simulation, randomly generate a set of load and distributed generation parameter values according to the probability distribution of the determined uncertainty factors; then, based on these randomly generated parameters, combine the power flow equation to calculate the current action vector The new state after
[0130] S2133. Calculate the state transition probability and calculate the current action vector from the statistical simulation results. The number of times the state is transferred to each possible state is calculated, and the state transition probability is obtained through a large number of simulations. The probability of transitioning to all possible states, these probabilities together constitute a row of elements in the state transition probability matrix. Repeating the above process for all possible states and action vectors can construct the complete state transition probability matrix.
[0131] In S3, the steps for constructing the reward function are:
[0132] S321. Define voltage quality reward / penalty items. The voltage quality reward / penalty items are expressed as:
[0133] ;
[0134] in, is the voltage quality reward / penalty item; is the node voltage deviation; is the weight of the node, which is set according to the importance of the node. The weight of important nodes is larger;
[0135] S322. Equipment operating cost penalty item, taking into account the operating costs of transformer tap adjustment, capacitor switching, and distributed generation reactive power output adjustment, the equipment operating cost penalty item is:
[0136] ;
[0137] is the equipment operating cost penalty item; Adjust the number of transformer taps; is the number of capacitor switching times; The number of times the reactive output of distributed generation is adjusted; is the unit cost factor for the number of transformer tap adjustments; is the unit cost coefficient of the capacitor switching times; is the unit cost coefficient of the reactive power output adjustment times of the distributed power source;
[0138] S323. System security reward / penalty items. To ensure the safe operation of the system, a system security indicator is introduced. The system security reward / penalty items are expressed as:
[0139] ;
[0140] Reward / penalty items for system security; is the number of safety constraints; It is The weight of each security constraint is used to adjust the importance of different security constraints. When the system violates the security constraint, A negative value penalizes the agent and guides it to choose actions that satisfy safety constraints;
[0141] S324. Define a reward function that comprehensively considers the voltage quality, equipment operating costs, and system security of the distribution network. The reward function is expressed as:
[0142] ;
[0143] represents the reward function.
[0144] In S4, the method of training the agent using the PPO algorithm is as follows:
[0145] S211. Initialization: Initialize the parameters of the policy network and the value network, and set the experience replay buffer. The experience replay buffer is used to store the experience data generated by the interaction between the agent and the environment. The role of the experience replay buffer is to break the correlation of the data.
[0146] S212. Training loop;
[0147] S2121. In the initial stage, the agent starts from the initial state Start by selecting an action based on the current policy network ,use Greedy strategy, randomly select actions with probability, with The probability of randomly selecting an action is The probability of choosing the action that maximizes the value function; It is a hyperparameter, and its initial value can be set larger (such as ), which allows the agent to perform more random exploration in the early stages of training to discover different state-action pairs;
[0148] Execute the action selected according to the current strategy network. The voltage control model constructed in S1 updates the operating status of the distribution network according to the action and feedbacks the reward. and the new state , the agent will experience Stored in the experience replay buffer; gradually decreases as training progresses The value of , which enables the agent to transition from random exploration to making decisions based on the learned knowledge;
[0149] S2122. Parameter update phase: When the experience replay buffer accumulates to the set threshold, parameter update begins, randomly sampling a batch of experience from it, and randomly sampling a batch of experience from the experience replay buffer;
[0150] Update the policy network parameters according to the clipping objective function of the PPO algorithm:
[0151] ;
[0152] represents the loss function; To initialize the policy network; is the minimum function; are the initialization strategy network parameters; are the policy network parameters before updating; is the clipping parameter (usually set to a small value, such as 0.2); Is the advantage function, which represents the action under the current strategy In state degree of advantage; is the policy network parameter before update strategic network; Represents the policy network before update The state sampled and actions Perform expectation calculations; is the truncation function;
[0153] By updating the policy network parameters in this way, the new policy is updated in a more optimal direction while ensuring stability;
[0154] Update the value network parameters by minimizing the loss function of the value network:
[0155] ;
[0156] in, Is the loss function of the value function (ValueFunction), used to optimize the parameters of the value network ; Is a discount factor used to weigh the importance of future rewards, usually with a value between between; are the parameters of the value network; For the value network state Estimated value of The agent is in state Execute an action After that, the reward value of the environment feedback; the value network performs the action The new state reached later Estimated value of Indicates the state obtained by sampling ,award and the new state Perform expectation calculations;
[0157] Adjust the parameters of the value network through the back-propagation algorithm to make it estimate the state value more accurately;
[0158] The soft update method is used to update the policy network and value network parameters to make the training more stable. The update formula is:
[0159] ;
[0160] ;
[0161] is the soft update coefficient (usually a small value, such as ), are the updated initialization strategy network parameters; are the parameters of the updated value network;
[0162] The soft update method makes the update of network parameters smoother, avoiding training instability caused by too fast parameter updates;
[0163] S213. Model optimization and convergence. During training, continuously monitor the performance indicators of the policy network and value network.
[0164] Average Reward: Calculates the average reward received by the agent during a specific training step. An upward trend in average reward indicates that the agent is learning a more effective control strategy. A higher reward value indicates that the agent's decision-making can achieve better benefits in the current environment.
[0165] The voltage qualification rate measures the percentage of distribution network node voltages within the qualified range within a specific training step. An increase in the voltage qualification rate means that the agent's control actions can better ensure voltage quality and meet user electricity needs.
[0166] Convergence judgment: When the fluctuation of the average reward and voltage qualification rate within continuous specific training steps is less than a certain threshold, the model is considered to have converged; at this time, the trained reinforcement learning voltage control model can adapt to the dynamic changes of the distribution network, accurately perceive the operating status of the distribution network, and output the corresponding optimal control action to ensure voltage quality and system safety.
[0167] The above description is a preferred embodiment of the invention and is not intended to limit the invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the invention should be included in the scope of protection of the invention.
Claims
1. A voltage control method for a multi-source energy storage distribution network based on a secure reinforcement learning algorithm, characterized in that: The following steps are involved: S1. Construct the basic framework of the distribution network voltage control model. Each node in the distribution network is considered as a state node in the model. The electrical connection relationship between nodes constitutes the state space. The voltage control operation is used as the action space of the model. S2. Establish a state transition probability matrix and solve it using the Newton-Raphson method. Consider the impact of actions on the system and, through Monte Carlo simulation, account for uncertainties such as load fluctuations and distributed power output. After multiple simulations, calculate the state transition probabilities and construct a complete state transition probability matrix to quantify the probability relationship between "current state-action-next state." S3. Define a reward function that integrates multiple objectives, including voltage quality, equipment operating costs, and system safety. S4. Reinforcement learning algorithm selection and training: using the constructed model as the environment, inputting preprocessed operating data, using the PPO algorithm to train the intelligent agent, and outputting the optimal control strategy after convergence.
2. The voltage control method of a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 1 is characterized in that: The state space considers not only node voltage amplitude, active power, and reactive power, but also device status and historical information.
3. The voltage control method for a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 1 is characterized in that: In addition to node voltage amplitude, active power, and reactive power, the state space also considers device status and historical information. The state vector Expressed as: ; For the node at time The voltage amplitude; For the node at time Active power; Node at time Reactive power; For the transformer at time The tap position; For the capacitor at time The switch status; It is historical state information, which is formed by splicing the states of multiple time steps in the past and is used to capture the dynamic change trend of the system; Indicates transpose.
4. The voltage control method for a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 1 is characterized in that: The action space contains discrete and continuous operations, and the action vector is expressed as: ; is the adjustment value of the transformer tap, which is , It means downgrading one gear. Indicates that it remains unchanged. Indicates moving up one gear; is the change in the capacitor switching state, which is , Indicates removal, Indicates that it remains unchanged, Indicates investment; is the reactive output adjustment of the distributed generation, which is a continuous value and must meet the upper and lower limit constraints.
5. The voltage control method for a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 1 is characterized in that: The method for establishing the state transition probability matrix: S211. Calculate state transition based on power flow equation; There is a relationship between the node injection current vector, the node voltage vector, and the node admittance matrix: ; Inject current vector into the node, is the node admittance matrix, is the node voltage vector; The active power and reactive power of a node are related to the node voltage amplitude, phase angle, and voltage parameters of adjacent nodes. The expressions are: ; ; in, For nodes Active power; For nodes Reactive power; For nodes The voltage amplitude; For nodes The voltage amplitude of the node is a node adjacent nodes; For nodes The set of adjacent nodes of and Node and The conductance and susceptance between is a node and The voltage phase angle difference between , For nodes The voltage phase angle, For nodes The voltage phase angle; The Newton-Raphson method is used to solve the above tidal flow equation, which is expressed as a nonlinear equation system: ,in is the state variable vector to be solved, is the number of nodes in the distribution network; Except for the balance node The voltage phase angle of each node, Except for the balance node The voltage amplitude of each node; select a node as the balance node, and its voltage amplitude and voltage phase angle It is known that the iterative formula of the Newton-Raphson method is: ; ; in, is the Jacobian matrix, whose elements are obtained by taking partial derivatives of the state variables according to the power flow equation; It is The state variables at the iteration vector of It is The iterative process continues until the convergence condition is met. ,in, is the preset convergence accuracy; It is The residual of the power flow equation at the iteration; S212. Modeling the impact of actions on the system: When executing action vector When the load is high, it will have different impacts on the operating status of the distribution network; Impact of transformer tap adjustment: Changes in tap position will directly affect the transformer's turns ratio; The change of the transformation ratio changes the voltage relationship between the nodes on both sides of the transformer and the equivalent impedance of the line, which is finally reflected in the node admittance matrix. On the changes of elements; Impact of capacitor switching: When Change in the switching state of the capacitor bank When the reactive power injection of the node increases; when the Change in the switching state of the capacitor bank When the reactive power injection is reduced, the change of reactive power will affect the node voltage amplitude and phase angle, and further affect the power distribution and voltage level of the entire distribution network; For distributed power sources, the reactive output adjustment directly changes the reactive power injection of the node to which the distributed power source is connected, which will cause the voltage of the node to change. Due to the electrical connection relationship of the distribution network, the change of the node voltage will propagate in the network, affecting the voltage and power distribution of other nodes. S213. Monte Carlo simulation method is used to accurately reflect the impact of uncertainty on state transition probability.
6. The voltage control method for a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 5 is characterized in that: The steps of using the Monte Carlo simulation method to reflect the impact of uncertainty on state transition probability are as follows: S2131. Determine the probability distribution of uncertainties. When modeling load fluctuations, the active and reactive power at the node follow a normal distribution. For distributed power sources, a Beta distribution is used. S2132. Monte Carlo simulation process, for a given current state and current action vector , perform multiple Monte Carlo simulations; In each simulation, a set of load and distributed generation parameter values are randomly generated according to the probability distribution of the determined uncertainty factors; then, based on the randomly generated parameters, the execution vector of the current action is calculated in combination with the power flow equation. The new state after S2133. Calculate the state transition probability and calculate the current action vector from the statistical simulation results. The number of times the state is transferred to each possible state is calculated, and the state transition probability is obtained through a large number of simulations. The probability of transferring to all possible states, the probabilities of all possible states together constitute a row of elements in the state transition probability matrix; repeat the above process for all possible states and action vectors to construct a complete state transition probability matrix.
7. The method for voltage control of a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 1, characterized in that: The voltage quality rewards and penalties are based on node voltage deviation and weight; the equipment operation cost penalty includes the adjustment costs of transformers, capacitors, and distributed power sources; and the system safety rewards and penalties are targeted at violations of safety constraints.
8. The voltage control method for a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 1 is characterized in that: The steps of constructing the reward function are: S321. Define voltage quality reward / penalty items. The voltage quality reward / penalty items are expressed as: ; in, is the voltage quality reward / penalty item; is the node voltage deviation; is the weight of the node, which is set according to the importance of the node. The weight of important nodes is larger; S322. Equipment operating cost penalty item, taking into account the operating costs of transformer tap adjustment, capacitor switching, and distributed generation reactive power output adjustment, the equipment operating cost penalty item is: ; is the equipment operating cost penalty item; Adjust the number of transformer taps; is the number of capacitor switching times; The number of times the reactive output of distributed generation is adjusted; is the unit cost factor for the number of transformer tap adjustments; is the unit cost coefficient of the capacitor switching times; is the unit cost coefficient of the reactive power output adjustment times of the distributed power source; S323. System security reward / penalty items. To ensure the safe operation of the system, a system security indicator is introduced. The system security reward / penalty items are expressed as: ; Reward / penalty items for system security; is the number of safety constraints; It is The weight of each security constraint is used to adjust the importance of different security constraints; is the safety constraint function, Indicates a violation of the safety constraint; when the system violates the safety constraint, A negative value penalizes the agent and guides it to choose actions that satisfy safety constraints; S324. Define a reward function that comprehensively considers the voltage quality, equipment operating costs, and system security of the distribution network. The reward function is expressed as: ; represents the reward function.
9. The method for voltage control of a multi-source energy storage distribution network based on a security reinforcement learning algorithm according to claim 1, characterized in that: The agent and environment interaction training method: S411. Initialization: Initialize the parameters of the policy network and the value network, and set the experience replay buffer. The experience replay buffer is used to store the experience data generated by the interaction between the agent and the environment. The role of the experience replay buffer is to break the correlation of the data. S412. Training loop; S4121. In the initial stage, the agent starts from the initial state and selects actions according to the current strategy network. Greedy strategy, randomly select actions with probability, with The probability of randomly selecting an action is The probability of choosing the action that maximizes the value function; is a hyperparameter; The execution selects actions according to the current strategy network. The voltage control model constructed in S1 updates the operating status of the distribution network according to the action, feedbacks rewards and new status, and the agent stores the experience in the experience playback buffer; as the training progresses, it gradually decreases. The value of , which enables the agent to transition from random exploration to making decisions based on the learned knowledge; S4122. Parameter update phase, when the experience replay buffer accumulates to the set threshold, the parameter update begins, randomly sampling a batch of experience, randomly sampling a batch of experience from the experience replay buffer; Update the policy network parameters according to the clipping objective function of the PPO algorithm: ; represents the loss function; To initialize the policy network; is the minimum function; are the initialization strategy network parameters; are the policy network parameters before updating; is the cropping parameter; Is the advantage function, which represents the action under the current strategy In state degree of advantage; is the policy network parameter before update strategic network; Represents the policy network before update The state sampled and actions Perform expectation calculations; is the truncation function; By updating the policy network parameters, the new policy is updated in a more optimal direction while ensuring stability; Update the value network parameters by minimizing the loss function of the value network: ; in, Is the loss function of the value function, used to optimize the parameters of the value network ; Is a discount factor used to weigh the importance of future rewards, usually with a value between between; are the parameters of the value network; For the value network state Estimated value of The agent is in state Execute an action After that, the reward value of the environment feedback; the value network performs the action The new state reached later Estimated value of Indicates the state obtained by sampling ,award and the new state Perform expectation calculations; Adjust the parameters of the value network through the back-propagation algorithm to make it estimate the state value more accurately; The soft update method is used to update the policy network and value network parameters to make the training more stable. The update formula is: ; ; is the soft update coefficient, are the updated initialization strategy network parameters; are the parameters of the updated value network; The soft update method makes the update of network parameters smoother, avoiding training instability caused by too fast parameter updates; S413. Model optimization and convergence. During the training process, continuously monitor the performance indicators of the policy network and value network; Convergence judgment: When the fluctuation of the average reward and voltage qualification rate within continuous specific training steps is less than a certain threshold, the model is considered to have converged; at this time, the trained reinforcement learning voltage control model can adapt to the dynamic changes of the distribution network, accurately perceive the operating status of the distribution network, and output the corresponding optimal control action to ensure voltage quality and system safety.
10. The method for voltage control of a multi-source energy storage distribution network based on a safety reinforcement learning algorithm according to claim 8, characterized in that: The performance indicators of the policy network and value network include: Average Reward: Calculates the average reward received by the agent during a specific training step. An upward trend in average reward indicates that the agent is learning a more effective control strategy. A higher reward value indicates that the agent's decision-making can achieve better benefits in the current environment. The voltage qualification rate counts the proportion of distribution network node voltages within the qualified range within a specific training step; an increase in the voltage qualification rate means that the intelligent agent's control actions can better ensure voltage quality and meet users' electricity needs.
Citation Information
Cited By
Power distribution network regulation and control method considering distributed power supply and related device
CN120914917A
Active power distribution network operation control method based on safety deep reinforcement learning
CN121485169A