Multi-Time Scale Voltage Regulation Method and System Based on Deep Reinforcement Learning
By introducing a multi-time scale voltage regulation method based on deep reinforcement learning in the active distribution network, using a dual-time scale parallel network and a redundant multi-agent collaborative control system, the problem of the inability to adjust multiple devices at the same time in the prior art is solved, and more efficient and stable voltage control is achieved.
Patent Information
- Application Number
- CN202411696765.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-26
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-11-26
AI Technical Summary
The existing deep reinforcement learning algorithms cannot simultaneously regulate on-load voltage regulators, capacitor banks, inverters, and electrical energy storage systems in the voltage control of active distribution networks, and it is difficult to deal with the joint control needs of multiple time scales and hybrid equipment.
A multi-time scale voltage regulation method based on deep reinforcement learning is proposed. By establishing an intelligent body model, adopting a dual-time scale parallel network structure, combining a redundant multi-agent collaborative control system, and using priority experience playback, Thompson sampling and safety module technology, multi-time scale voltage regulation of the active distribution network is realized.
It significantly improves the learning efficiency and control effect of the agent, can effectively adjust multi-time scales and hybrid devices, reduce the impact of individual agents' wrong decisions on the overall system, and make the system more stable in the face of uncertainty.
Smart Images

Figure CN119209561B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distribution networks, and relates to a voltage control method for an active distribution network based on deep reinforcement learning, specifically to a multi-time scale voltage regulation method and system based on deep reinforcement learning. Background Art
[0002] With the gradual increase in the penetration rate of Renewable Energy Sources (RES) in an Active Distribution Network (ADN), the uncertainty of RES has brought serious operation problems to the ADN, such as voltage fluctuations and violations, high active power losses, etc. Therefore, in order to fully address the problems brought by the uncertainty of RES to the ADN, a real-time online voltage control method needs to be designed for the ADN. Deep Reinforcement Learning (DRL), as a real-time control method, has shown excellent performance in the applications of games, robotics, and power systems.
[0003] Currently, although the voltage regulation method based on DRL has been gradually applied to the ADN, such voltage regulation methods still cannot well solve the requirements of multi-time scale and combined control of hybrid devices existing in voltage regulation. Summary of the Invention
[0004] Aiming at the above problems existing in the existing DRL voltage regulation method, the present invention proposes a multi-time scale voltage regulation method based on deep reinforcement learning for optimizing the voltage control in an active distribution network.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A multi-time scale voltage regulation method based on deep reinforcement learning, comprising the following steps: Step 1. Model an active distribution network with a large amount of renewable energy, and establish an active distribution network voltage optimization model; Step 2. Establish an intelligent agent with continuous and discrete action characteristics including time series characteristics; the intelligent agent includes two networks, one is a slow network for discrete control on the hourly time scale for on-load tap changer and capacitor bank scheduling, and the other is a fast network for continuous control on the minute-level time scale for inverter and electric energy storage system scheduling; the fast network and the slow network of the intelligent agent each have two internal networks, respectively defined as the fast control Q network, the fast control target Q network, the slow control Q network, and the slow control target Q network; Step 3. Use a single-layer Markov to construct a multi-time scale decision-making process to describe the interaction process between the active distribution network voltage optimization model and the intelligent agent, and introduce the distribution network state, the fast network state, and the slow network state; in particular, design three state-coupled state transition functions, and use a time counter to separate the fast time scale and the slow time scale, and use the output of the time counter as an activation signal to control the slow network; Step 4. Propose a redundant multi-agent collaborative control system for multi-time scale voltage regulation of an active distribution network; use multiple intelligent agents in the system to perform the same voltage regulation task, and then adjust the action output of the intelligent agents through a redundant coordination mechanism, and finally apply the coordinated actions to the active distribution network; Step 5. Establish the training process of the intelligent agent; apply the trained intelligent agent, combined with the redundant multi-agent collaborative control system, to the voltage control of the active distribution network to achieve multi-time scale voltage regulation.
[0006] In addition, based on the above multi-time scale voltage regulation method based on deep reinforcement learning, the present invention also proposes a corresponding multi-time scale voltage regulation system based on deep reinforcement learning, which adopts the following technical solutions: The multi-time scale voltage regulation system based on deep reinforcement learning includes the following modules: A model establishment module, which is used to model the active distribution network with a large amount of renewable energy and establish an active distribution network voltage optimization model; An agent construction module, which is used to establish an agent with continuous action and discrete action characteristics including time series characteristics; The agent includes two networks, namely a slow network for discrete control on the time scale of hours for on-load tap changer and capacitor bank scheduling, and a fast network for continuous control on the time scale of minutes for inverter and electrical energy storage system scheduling; The fast network and slow network of the agent also have two internal networks, which are respectively defined as a fast control Q network, a fast control target Q network, a slow control Q network, and a slow control target Q network; An interaction process description module, which uses a single-layer Markov to construct a multi-time scale decision-making process, is used to describe the interaction process between the active distribution network voltage optimization model and the agent, and introduces the distribution network state, the fast network state, and the slow network state; Among them, three state coupling state transition functions are particularly designed, and a time counter is used to separate the fast time scale and the slow time scale, and the output of the time counter is used as an activation signal to control the slow network; A multi-agent collaborative control system construction module, which is used to propose a redundant multi-agent collaborative control system for multi-time scale voltage regulation of an active distribution network based on deep reinforcement learning; Multiple agents are used in the system to execute the same voltage regulation task, and then the action output of the agents is adjusted through a redundant coordination mechanism, and finally the coordinated action is applied to the active distribution network; And a voltage control module for the active distribution network, which is used to apply the trained agent, combined with the redundant multi-agent collaborative control system, to the multi-time scale voltage control of the active distribution network.
[0007] The present invention has the following advantages:
[0008] As described above, the present invention relates to a multi-time scale voltage regulation method and system based on deep reinforcement learning. The present invention makes up for the deficiency of the existing deep reinforcement learning algorithm in the voltage control of active distribution networks, which cannot simultaneously regulate on-load tap changers, capacitor banks, inverters, and electrical energy storage systems. The redundant multi-agent cooperation mechanism proposed by the present invention can significantly reduce the impact of the wrong decisions of individual agents on the overall system, making the system more stable in the face of uncertainties. In addition, the prioritized experience replay and safety module introduced by the present invention can accelerate the convergence speed by using training data more effectively. In addition, the present invention also introduces Thompson sampling, which can provide a flexible exploration mechanism for agents, enabling agents to balance exploration and exploitation when facing unknown environments and avoiding premature convergence to suboptimal solutions. By integrating the prioritized experience replay technology, Thompson sampling technology, and safety module mechanism, the present invention significantly improves the learning efficiency and control effect of agents, and well meets the requirements of multi-time scale and joint control of hybrid devices in voltage regulation. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 It is a flowchart of the multi-time scale voltage regulation method based on deep reinforcement learning in an embodiment of the present invention.
[0010] Figure 2 It is a network structure diagram of the agent built in an embodiment of the present invention.
[0011] Figure 3 It is a network structure diagram of the redundant multi-agent cooperation control system built in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The present invention will be further described in detail below with reference to the drawings and specific embodiments:
[0013] Embodiment 1
[0014] As Figure 1As shown in the figure, this embodiment describes a multi-time-scale voltage regulation method based on deep reinforcement learning. In this method, a detailed voltage optimization model is constructed, and an advanced algorithm strategy is introduced to effectively manage the voltage and the behavior of the equipment. Specifically, in an intelligent agent, a network structure with a double-time-scale parallel described by a single-layer Markov is established. One is a slow control network for controlling on-load tap changers and capacitor banks at the hourly time scale, and the other is a fast control network for controlling inverters and electrical energy storage systems at the minute-level time scale. For the network that controls at the hourly time scale, the present invention designs a control method suitable for on-load tap changers and capacitor banks. For the network that controls at the minute-level time scale, the present invention designs a control method suitable for inverters and electrical energy storage systems. During the control process, the two networks use the same objective function as a guide, and then the obtained scheduling strategies are stacked and uniformly output by the intelligent agent. In addition, the present invention also proposes a redundant multi-agent collaborative decision-making architecture. For the same voltage regulation task, the present invention designs to use multiple intelligent agents, and optimizes the scheduling results of each intelligent agent through a collective decision-making model, and then outputs an optimal scheduling result for the voltage regulation task. This redundant multi-agent collaborative decision-making architecture effectively improves the robustness of the system. On the other hand, the present invention also uses a safety module, a prioritized experience replay technique, and a Thompson sampling technique to improve the control effect of the proposed method.
[0015] The multi-time-scale voltage regulation method based on deep reinforcement learning in this embodiment includes the following steps:
[0016] Step 1. Model the active distribution network with a large amount of renewable energy, and establish an active distribution network voltage optimization model, including the line characteristic model of the active distribution network, the on-load tap changer model, the capacitor bank model, the inverter model installed at the photovoltaic power generation system, the inverter model installed at the wind power generation system, and the electrical energy storage system model.
[0017] This distribution network voltage optimization model is used to calculate the sum of the active power losses of each line within the time range T, the number of distribution network nodes, and the real-time voltage of each node. In particular, a safety module mechanism is used for the electrical energy storage system model.
[0018] Specifically, in order to improve the voltage quality and achieve voltage stability, the present invention realizes voltage regulation by reducing the active power loss. Therefore, the objective function is defined as follows:
[0019] .
[0020] The above formula represents the sum of the active power losses of each node in the distribution network within the time range T.
[0021] Among them, is the number of nodes in the distribution network, represents the th node and the th node in the distribution network, and the line at time The current value at that time, represents the line Impedance.
[0022] The formula of the line characteristic model of the active distribution network is as follows:
[0023] For the line characteristic model of the active distribution network, it is described as the active power and reactive power constraints:
[0024] .
[0025] .
[0026] Among them, and respectively represent the amplitude and phase angle of the voltage of node at moment, and represent The actual injection amounts of active power and reactive power of the node at moment. and are the real part and imaginary part of the admittance element between node and node lines, that is, conductance and susceptance. is node and node The phase angle difference between the lines, formulated as: .
[0027] Actual active injection amount is obtained by connecting to node at The active power output of the photovoltaic power generation device , the active power output of the wind turbine , the active power amount absorbed or released by the energy storage system and the active power demand of the load obtained. The actual reactive injection amount is by connecting to node at The reactive power injection amount of the inverter at the photovoltaic power generation , the reactive power injection amount of the inverter at the wind power generation , the reactive power injection amount of the capacitor bank With the reactive power demand of the load It is obtained.
[0028] In summary, for the nodes accessing the control devices, considering the power flow injection after the device actions is as follows:
[0029] ; .
[0030] The formula expression of the voltage security constraint is as follows:
[0031] .
[0032] Among them, Represents the voltage value of the distribution network node ; Represents the minimum voltage value of each node under the normal operation of the distribution network; Represents the maximum voltage value of each node under the normal operation of the distribution network.
[0033] The formula expression of the on-load tap-changer model is as follows:
[0034] ; ; .
[0035] Among them, Is the primary voltage of the on-load tap-changer at the bus; Is the tap position of the on-load tap-changer; Is the operation times of the on-load tap-changer; Is the maximum operation times of the on-load tap-changer within a day; Is the maximum tap position change range of the on-load tap-changer.
[0036] The formula expression of the capacitor bank model is as follows:
[0037] ; ; .
[0038] Among them, Represents the reactive power injected by the capacitor bank into the distribution network, Represents the reactive power of a group of capacitors; Represents the number of groups of capacitor banks connected to the active distribution network; Is the operation times of the capacitor bank; Is the maximum operation times of the capacitor bank within a day; Is the maximum operation group change range of the capacitor bank.
[0039] The formula expression of the inverter model installed at the photovoltaic power generation system is as follows:
[0040] ; .
[0041] Wherein represents the active power injected by the photovoltaic power generation system into the distribution network, is the maximum active power output of the photovoltaic power generation system; is the available reactive power of the inverter of the photovoltaic power generation system; is the rated power of the inverter of the photovoltaic power generation system. The formula for installing the inverter model at the wind power generation system is as follows:
[0042] ; .
[0043] Wherein, represents the active power injected by the wind power generation system into the distribution network, is the maximum active power output of the wind power generation system; is the available reactive power of the inverter of the wind power generation system; is the rated power of the inverter of the wind power generation system.
[0044] The formula of the electrical energy storage system model is as follows:
[0045] .
[0046] .
[0047] .
[0048] .
[0049] Wherein, , are the minimum and maximum energy storage levels of the energy storage system; , respectively represent at , the electricity content of the energy storage system at node ; represents the charging power of the energy storage system at node at time represents the maximum charging power of the energy storage system; represents the maximum discharging power of the energy storage system; represents the discharging power of the energy storage system at node at time represents node Device efficiency during charging of the energy storage system; Represents a node Device efficiency during discharging of the energy storage system; Represents a time step; Represents a time range, Is the number of nodes in the distribution network.
[0050] Optimize the electrical energy storage system model and use a safety module. The formula is as follows:
[0051] 。
[0052] 。
[0053] 。
[0054] Among them, and respectively represent at 、 The electricity content of the energy storage system at node at the moment.
[0055] Step 2. Propose a Dueling DeepQ-Network algorithm with Hybrid action space, Priority experience replay, and Thompson Sampling characteristics, namely the Hybrid Priority Thompson Sampling Dueling Deep Q-Network, abbreviated as HPTS-Dueling DQN.
[0056] Among them, H represents Hybrid Action Space, P represents Priority Experience Replay, and TS represents Thompson Sampling.
[0057] Use HPTS-Dueling DQN to establish an agent with continuous and discrete action characteristics containing time series characteristics. The agent has two neural network structures. One is a slow network for discrete control on an hourly time scale for on-load tap changer and capacitor bank scheduling. The other is a fast network for continuous control on a minute-level time scale for inverter and electrical energy storage system scheduling.
[0058] The present invention controls all devices using the HPTS-Dueling DQN algorithm, that is, both the fast network and the slow network of the agent are designed based on the HPTS-Dueling DQN algorithm, and the HPTS-DUELING DQN is used to complete real-time multi-time scale voltage regulation within a day. As a deep reinforcement learning algorithm, the HPTS-Dueling DQN can quickly give control strategies and give control according to observations within a day. Especially when controlling OLTC and CB devices, it is not necessary to perform day-ahead control on OLTC and CB according to the predicted results, but directly give control according to the observations of the same day. Other methods either perform day-ahead control on OLTC and CB and intra-day control on PV, WT, and ES, or perform intra-day control on OLTC and CB and day-ahead control on PV, WT, and ES, and they adopt a two-stage or a combination of multiple methods. The present invention uses a method to uniformly schedule OLTC, CB, PV, WT, and ES within a day.
[0059] As Figure 2 shown, each network of the agent (i.e., the fast network and the slow network) has two internal networks, one is the target control network and the other is the control network. These networks are respectively the fast control Q network, the fast control target Q network, the slow control Q network, and the slow control target Q network.
[0060] Step 2.1. Establishment of the time scale control network.
[0061] Based on the Dueling Deep Q-Network (Dueling DQN) algorithm, the Q value is decomposed into two parts: the state value and the advantage function, so as to more effectively estimate the action value under different states.
[0062] For the state value function and the advantage function, a shared feedforward neural network is used for feature extraction.
[0063] Feedforward neural networks are respectively established for the fast control Q network, the fast control target Q network, the slow control Q network, and the slow control target Q network 、 :
[0064] .
[0065] .
[0066] Where is the state input; 、 are weight matrices, 、 is the bias term; 1, 2 represent the number of layers of the neural network, where 1 represents the first layer of the neural network and 2 represents the second layer of the neural network; is the activation function.
[0067] The neural network parameters are represented by That is:
[0068] .
[0069] Then there are the fast network parameters and the slow network parameters .
[0070] Use the state value function to measure the overall value of the current state , indicating the basic return that can be obtained regardless of which action is selected in the current state.
[0071] Note that the fast control Q-network and the fast control target Q-network use the same network structure , and the slow control Q-network and the slow control target Q-network use the same network structure .
[0072] To avoid trouble, in the description of the present invention, the fast network and the slow network are used for description instead.
[0073] The fast control Q-network and the fast control target Q-network have exactly the same structure and parameters when established. The main differences lie in the parameter update method and naming. The slow control Q-network and the slow control target Q-network have exactly the same structure and parameters when established. The main differences lie in the parameter update method and naming.
[0074] Separate state value functions and are established for the fast network and the slow network respectively:
[0075] .
[0076] .
[0077] Among them, , , , represent the network parameters used to calculate the state value network in the fast network. Similarly, , are weight matrices, , are bias terms; 1, 2 represent the number of layers of the neural network, where 1 represents the first layer of the neural network and 2 represents the second layer of the neural network; , , , represent the network parameters for calculating the state value network in a slow network. Similarly, , are weight matrices, , are bias terms; 1, 2 represent the layers of the neural network, i.e., 1 represents the first layer of the neural network and 2 represents the second layer of the neural network.
[0078] Use the advantage function to represent the advantage of performing a specific action in state compared to the average level, i.e., how much better it is to perform a certain action in this state than other actions.
[0079] Establish the advantage functions and respectively for the fast network and the slow network:
[0080] .
[0081] .
[0082] Among them, , , , represent the network parameters for calculating the advantage value network in the fast network, , are weight matrices, , are bias terms; , , , represent the network parameters for calculating the advantage value network in the slow network. Similarly, , are weight matrices, , are bias terms; The subscript 1 and subscript 2 represent the layers of the neural network, i.e., 1 represents the first layer of the neural network and 2 represents the second layer of the neural network.
[0083] The final Q value is calculated by combining the value and the advantage.
[0084] Establish the Q value functions and respectively for the fast network and the slow network:
[0085] .
[0086] .
[0087] Among them, represents the action space of the fast network; represents the action at the next moment of the fast network, and ; represents the action space of the slow network; represents the action at the next moment of the slow network, and . And are the average values of the advantage functions of all actions, used to normalize the advantage values to ensure the stability of Q-value calculation.
[0088] The output of the fast network converts the Q-value into a continuous action value through the tanh function , mapping the action to the range of [-1, 1].
[0089] .
[0090] The output of the slow network selects the optimal action from the discrete action set through the argmax operation. Among them represents the tap position of the on-load tap changer OLTC or the number of capacitor banks CB connected to the ADN:
[0091] .
[0092] Step 2.2. Concatenation of actions.
[0093] In the multi-time frame, the fast network and the slow network alternately generate fast actions and slow actions at different time scales, that is, at the minute-level time scale, the on-load tap changer and capacitor banks do not generate new actions, only the inverter and the electric energy storage system act; while at the hour-level time scale, the on-load tap changer, capacitor, inverter, and electric energy storage system all act; specifically as follows:
[0094] .
[0095] Among them, and are the numbers of continuous actions and discrete actions respectively; represents the new action set of the fast network, represents the old action set of the slow network, represents the new action of the slow network.
[0096] Note that under the action of the time counter, the fast network outputs new action values at each time step, while the slow network needs to wait for the activation signal of the counter. After the slow network receives the activation signal of the counter, it will output new action values. When it does not receive the activation signal of the counter, it outputs the previously updated action, that is, the original old action.
[0097] Represents the content of the new action set of the fast network, Represents the content of the old action set of the slow network, Represents the content of the new action set of the slow network.
[0098] Stack the action sets of the fast network and the slow network. At time, the agent 's action is:
[0099] .
[0100] Step 3. Use a single-layer Markov to construct a multi-time-scale decision-making process to describe the interaction process between the active distribution network voltage optimization model and the agent, and introduce the distribution network state, fast network state, and slow network state.
[0101] Design three state-coupled state transition functions, which are respectively:
[0102] is the state transition function of the environment, i.e., the ADN state, is the state transition function of the internal state of the fast network, is the state transition function of the internal state of the slow network.
[0103] Their coupling form is:
[0104] .
[0105] And use the time counter to separate the fast time scale and the slow time scale, and use the output of the time counter as the activation signal to control the slow network.
[0106] The Markov decision process (MDP) is used to represent the interaction process between the agent and the environment. The usually necessary elements are: . Among them is the state space; is the action space; is the transition probability function; is the reward function, defined according to the specific goals of the system; is the discount factor.
[0107] To address the multi-time-scale problem, a counter is introduced , the different time scales of the two networks are explicitly represented and processed during the decision-making process. A multi-time scale decision-making process is constructed using a single-layer Markov as follows:
[0108] a) State space:
[0109] Define the state .
[0110] Among them, is the current state of the environment, i.e., the ADN ; is the internal state of the fast network; is the internal state of the slow network.
[0111] b) Action space:
[0112] Define the action ;
[0113] Among them, is the action of the fast network at time, is the action of the slow network at time.
[0114] ; Among them represents the control rate of the reactive power of the photovoltaic inverter, represents the control rate of the active power of the inverter of the wind power generation device; represents the active control rate of the energy storage device.
[0115] If then the energy storage system absorbs electrical energy, if then the energy storage system discharges electrical energy.
[0116] For the inverters of photovoltaic and wind turbines, first use and to obtain the available reactive power, and then obtain the reactive power injected into the ADN under the action of the corresponding control rate , ;
[0117] , .
[0118] Among them, represents the available reactive power of the photovoltaic power generation device, represents the available reactive power of the wind turbine, where are the rated capacities of the photovoltaic and wind turbine inverters respectively, and represent the control rates.
[0119] For the energy storage system, the active power actually injected into the ADN is obtained under the influence of the control rate . .
[0120] If then ; If then .
[0121] .
[0122] Among them represents the gear of the tap contact of the OLTC, represents the number of capacitor banks connected to the ADN, that is, by determine the position of the tap of the on-load tap-changer OLTC , and by determine the number of capacitor banks connected to the distribution network .
[0123] c) Transition probability:
[0124] .
[0125] Among them, the transition probability is further decomposed into:
[0126] .
[0127] Among them, describes the transition of the environment, i.e., the state of the ADN, describes the transition of the internal state of the fast network, describes the transition of the internal state of the slow network.
[0128] d) Counter:
[0129] For the slow network, a counter is introduced to track the time step.
[0130] According to the result of the counter, when reaching the hourly time scale of the slow network action, the slow network should perform a new action; at other time steps, it maintains the state of the previous action.
[0131] If the time step reaches the hourly unit then .
[0132] If the time step is in minutes then , then there is an action selection for the slow network:
[0133] .
[0134] Step 4. Propose a redundant multi-agent collaborative control system for multi-time scale voltage regulation of active distribution networks based on deep reinforcement learning. In this embodiment, the redundant multi-agent collaborative control system specifically refers to that after multiple agents execute the same voltage regulation task and give control actions, a multiplication mechanism with lateral interaction is used to coordinate the action outputs of each agent, and then the action of one agent is randomly selected and applied to voltage regulation.
[0135] In the multi-agent collaborative control system, multiple agents are used to execute the same voltage regulation task, and then the action outputs of the agents are adjusted through a redundant coordination mechanism, and finally the coordinated actions are applied to the active distribution network.
[0136] Step 4.1. Establishment of the redundant coordination matrix.
[0137] Suppose there are agents, and the initial redundant coordination matrix is an identity matrix.
[0138] ;
[0139] The above formula means that initially, each agent only refers to its own action and has no influence on each other.
[0140] Step 4.2. Coordination of actions among multiple agents.
[0141] At a certain moment , the action of agent is represented as a vector .
[0142] Concatenate the actions of all agents to form a redundant action matrix with a dimension of . The redundant action matrix is coordinated through the communication matrix to generate the coordinated action :
[0143] .
[0144] Where is the redundant coordination matrix at the th moment, indicating the influence weight of each agent on the actions of other agents.
[0145] Finally, apply the coordinated action .
[0146] Step 4.3. Update of the redundant coordination matrix.
[0147] The redundant coordination matrix is updated according to the system performance after each action execution.
[0148] The performance metric of the system is the reward .
[0149] During each update, noise is added to the matrix , and then the matrix is adjusted according to the performance :
[0150] .
[0151] Among them, the noise is a random matrix subject to a normal distribution:
[0152] .
[0153] Among them, represents the parameter of the normal distribution. To ensure that the weights in the communication matrix are within the preset reasonable range, usually within [0, 1], and the sum of each row is 1, the communication matrix needs to be normalized.
[0154] Assume that after the matrix update in the th step is completed, the new matrix is , and the normalization process is as follows:
[0155] .
[0156] Among them, represents the element in the th row and th column of the matrix , is the normalized matrix.
[0157] Step 5. Establish the training process of the agent based on Steps 1, 2, 3, and 4.
[0158] During the training process, the fast network and the slow network respectively use their own target Q networks , to update the Q value, then the target Q values and are calculated as follows:
[0159] .
[0160] .
[0161] Among them is the immediate reward, represents the discount factor.
[0162] represents the Q-values of all possible actions output by the fast control target Q-network when the input is , and selects the action that maximizes the Q-value from these actions Then this maximized Q-value is used to calculate the target Q-value.
[0163] represents the Q-values of all possible actions output by the slow control target Q-network when the input is , and selects the action that maximizes the Q-value from these actions Then this maximized Q-value is used to calculate the target Q-value.
[0164] To optimize the network parameters of the fast network and the slow network, a loss function is established for the two networks respectively using the target Q-value and the Q-value and , and then the gradient descent method is used to minimize the loss function.
[0165] .
[0166] .
[0167] .
[0168] .
[0169] where is the learning rate, represents the expected value, represents the Q-value obtained when the fast control Q-network takes action under the network parameter in the state ; represents the Q-value obtained when the slow control Q-network takes action under the network parameter in the state ; is the gradient of the loss function of the fast control Q-network with respect to the parameter ; is the gradient of the loss function of the slow control Q-network with respect to the parameter .
[0170] To stabilize the training process, the parameters of the fast control target Q-network and the parameters of the slow control target Q-network are not updated in every training step, but are taken from the parameters of the current fast control Q-network every 100 time steps and Slow Control Q-Network Parameters Copy it over, that is:
[0171] ; .
[0172] This update method reduces the fluctuations during training and ensures more stable calculation of the target value.
[0173] The prioritized experience replay introduced in the Dueling Deep Q-Network algorithm is as follows:
[0174] The experience replay stores experience tuples , and assigns different priorities to different experience tuples through an experience evaluation function, so as to preferentially select the experiences that have a greater impact on the learning process for replay and training. Specifically as follows:
[0175] When storing new experiences in the buffer each time, calculate their errors and assign a priority to each experience according to the error; for the th group of experiences, its experience evaluation function is:
[0176] .
[0177] Among them, represents the input of the target Q-network , outputs the Q-values of all possible actions , selects the action that maximizes the Q-value from these actions Then use this maximized Q-value to calculate the target Q-value.
[0178] The value of and is obtained by taking the average.
[0179] Similarly represents the Q-value obtained when the Q-network takes action in the state of the network parameters .
[0180] The value of and is obtained by averaging.
[0181] The th group of experiences corresponds to the probability of the experience sampling, which is: .
[0182] represents the sum of the priorities of all experiences used to normalize the priorities of each experience. Among them is the total number of experiences in the buffer. An importance sampling weight is adopted when updating the Q-network Correct the sampled samples, and its calculation process is as follows:
[0183] ; is a parameter that controls the smoothness of the weight.
[0184] The Thompson sampling introduced in the Dueling Deep Q-Network algorithm is as follows:
[0185] Thompson sampling combines exploration and exploitation naturally through random sampling. Even if the current average reward of a certain action is high, due to the uncertainty of its distribution, other actions may still be drawn.
[0186] Therefore, the selection of actions will be dynamically adjusted according to their uncertainty, and the specific steps are as follows:
[0187] Let each action have a reward distribution that is a Beta distribution:
[0188] .
[0189] Among them, represents the distribution of action according to the reward (the content of the th group of experience tuples is indexed by ), and the parameters , represent the number of successes and failures respectively, and the initial values are 1.
[0190] Whenever an action needs to be selected, sample from the posterior reward distribution of each action and select the action with the highest reward, that is, in the given state , sample from the reward distribution of each action to generate a sample value for each action, and then select the action with the largest sample value, that is:
[0191] For each action , sample a value from its corresponding Beta distribution: .
[0192] Among them, is the sample value drawn from the Beta distribution, and then select the action with the largest sample value: .
[0193] Update the distribution parameters.
[0194] Reward according to the action to update the corresponding Beta distribution parameters.
[0195] If the reward is positive, increase the number of successful times of this action : .
[0196] If the reward is zero or negative, increase the number of failed times of this action : .
[0197] Step 1 establishes an active distribution network voltage optimization model, step 2 establishes agents for multi-time scale control, step 3 describes the state transition principle of multi-time scales in the active distribution network voltage optimization problem, establishes a multi-time scale decision-making process based on single-layer Markov construction, and step 4 establishes a redundant coordination mechanism for coordination between multiple agents.
[0198] In step 5 of the present invention, the agent is trained. The specific training is divided into a multi-time scale action process, that is, performing a multi-time scale decision-making process based on single-layer Markov construction and a network parameter update process, specifically as follows:
[0199] Multi-time scale decision-making process based on single-layer Markov construction:
[0200] At a certain moment, the active distribution network inputs the state of the distribution network to the agent for multi-time scale control through the active distribution network voltage optimization model. The time counter of the agent outputs a control signal according to the time scale for the activation of the fast network and the slow network. The fast network and the slow network of the agent randomly give control actions according to the state. After integrating the two network actions, the redundant coordination machine is used to coordinate the actions, and then this action is applied to the active distribution network voltage optimization model. The distribution network completes the update of the state according to the state transition principle. At this time, a reward value is obtained using the new state.
[0201] During this process, the state of the distribution network , the action applied to the active distribution network voltage optimization model , the reward value , the new state constitute an experience tuple . Then, an experience evaluation function is used to evaluate the experience tuple to obtain an evaluation result of the experience tuple. According to the quality of the evaluation result, the experience tuple is stored in the experience replay mechanism. Specifically, according to this evaluation result, the experience tuples with better evaluations are preferentially stored in the experience replay mechanism.
[0202] Repeat this process until reaching the upper limit of the experience replay mechanism. During the action sampling process, first establish the reward distribution of actions according to the reward function and actions, and sample actions according to the reward distribution.
[0203] Then perform the agent network parameter update process:
[0204] Extract a mini-batch of experience tuples from the experience replay mechanism. Then establish the reward distribution of actions according to the rewards and actions, and then select the actions for updating parameters based on the reward distribution of actions. Complete the update of the network parameters according to the internal structure of HPTS-DUELING DQN. At the same time, according to the actions and reward values Update the coordination matrix of the redundant coordination mechanism.
[0205] The specific training execution steps of the agent are as follows:
[0206] Step 5.1. Initialization phase.
[0207] Step 5.1.1. Initialize the neural network parameters of the fast network and the slow network.
[0208] Step 5.1.2. Initialize the experience replay buffer, create an experience replay buffer to store the experience tuples generated by the interaction between the agent and the environment.
[0209] Step 5.1.3. Initialize the priority of the experiences, and uniformly set all priorities to 1.
[0210] Step 5.1.4. Initialize the redundant coordination matrix.
[0211] Step 5.1.5. Initialize the Bayesian Q-value distribution.
[0212] Step 5.2. Interaction process.
[0213] Step 5.2.1. Select actions using Thompson sampling.
[0214] Step 5.2.2. Use the redundant multi-agent collaborative control system.
[0215] Step 5.2.3. Execute the Markov process to interact the agent with the active distribution network model.
[0216] Step 5.2.4. Store the experience tuples and assign the initialized priority to each experience tuple.
[0217] Step 5.3. New experience replay training.
[0218] Step 5.3.1. Calculate the loss function and priorities.
[0219] Step 5.3.2. Sample experiences and calculate importance sampling weights.
[0220] Step 5.3.3. Update the Q-network parameters of the fast network and the slow network.
[0221] Step 5.3.4. Update the target Q-network parameters of the fast network and the slow network.
[0222] Step 5.4. Bayesian distribution update.
[0223] After all the training processes are completed, save the agent with the new neural network parameters.
[0224] Apply the agent trained in Step 5 to the voltage control of the active distribution network in combination with the redundant multi-agent collaborative control system. Specifically, after the agent gives a control based on the state of the active distribution network, it acts through the redundant multi-agent collaborative control system, and then the control acts on the active distribution network.
[0225] As an advanced deep reinforcement learning algorithm, the HPTS-DUELING DQN algorithm shows significant advantages in terms of control devices. The core idea of this algorithm is to be able to quickly generate control strategies, enabling the system to make direct decisions based on real-time observation data during intraday scheduling without relying on previous prediction results. This feature is particularly important in the control of key devices such as OLTC (on-load tap changer) and CB (circuit breaker), because traditional methods usually require control planning in advance and rely on the prediction of future load and generation conditions. Compared with other methods, the biggest highlight of the HPTS-DUELING DQN algorithm is that its single method framework can achieve unified scheduling of OLTC, CB, photovoltaic (PV), wind power generation (WT), and energy storage system (ES). This means that during intraday control, the system can respond to changes in real time without having to adopt different control strategies at different time periods. For example, traditional control methods often take the form of two-stage or a combination of multiple methods: one is to perform day-ahead control on OLTC and CB, while performing intraday control on PV, WT, and ES; the other is to perform intraday control on OLTC and CB, while performing day-ahead control on PV, WT, and ES. Although this phased control method is effective in some cases, it often leads to a decrease in response speed and scheduling efficiency. Through the HPTS-DUELING DQN algorithm, the system can achieve higher flexibility and real-time performance during intraday scheduling, and can quickly adapt to the dynamic changes in the power market and the real-time feedback of device status. This not only improves the operating efficiency of the power grid, but also enhances the stability of the system and reduces the risks brought by prediction errors.
[0226] The present invention proposes an active distribution network voltage regulation method with the ability to jointly control hybrid devices and including time series characteristics. Specifically, an intelligent agent capable of outputting discrete actions and continuous actions is first set up, and through a network including time series features, the hourly and minute-level dual time-scale scheduling of on-load tap changers, capacitor banks, and renewable energy inverters is realized. In addition, the present invention also designs a redundant multi-agent collaborative control system, and coordinates the action outputs of each intelligent agent through a multiplication mechanism of horizontal interaction. Moreover, the present invention also integrates the prioritized experience replay technique, the Thompson sampling technique, and the safety module mechanism, thereby improving the learning efficiency and control effect of the intelligent agent.
[0227] Embodiment 2
[0228] This Embodiment 2 describes a multi-time scale voltage regulation system based on deep reinforcement learning, and this system is based on the same inventive concept as the multi-time scale voltage regulation method based on deep reinforcement learning in the above Embodiment 1.
[0229] The multi-time scale voltage regulation system based on deep reinforcement learning includes the following modules:
[0230] A model establishment module for modeling an active distribution network with a large amount of renewable energy and establishing an active distribution network voltage optimization model; an agent construction module for establishing an agent with continuous action and discrete action characteristics including time series characteristics. The agent includes two networks, namely a slow network for discrete control on an hourly time scale for on-load tap changer and capacitor bank scheduling, and a fast network for continuous control on a minute-level time scale for inverter and electric energy storage system scheduling. The fast network and slow network of the agent each have two internal networks, defined as the fast control Q network, the fast control target Q network, the slow control Q network, and the slow control target Q network respectively; an interaction process description module using a single-layer Markov to construct a multi-time scale decision-making process for describing the interaction process between the active distribution network voltage optimization model and the agent, and introducing the distribution network state, the fast network state, and the slow network state. Specifically, three state-coupled state transition functions are designed, and a time counter is used to separate the fast time scale and the slow time scale, and the output of the time counter is used as an activation signal to control the slow network; a multi-agent collaborative control system construction module for proposing a redundant multi-agent collaborative control system for multi-time scale voltage regulation of an active distribution network based on deep reinforcement learning. In the system, multiple agents are used to perform the same voltage regulation task, and then the action outputs of the agents are adjusted through a redundant coordination mechanism, and finally the coordinated actions are applied to the active distribution network; and a voltage control module for the active distribution network for applying the trained agent combined with the redundant multi-agent collaborative control system to the multi-time scale voltage control of the active distribution network.
[0231] It should be noted that in the multi-time scale voltage regulation system based on deep reinforcement learning, the implementation processes of the functions and roles of each functional module are specifically detailed in the corresponding steps of the method in Embodiment 1, and will not be elaborated here.
[0232] Of course, the above description is only a preferred embodiment of the present invention. The present invention is not limited to listing the above embodiments. It should be noted that all equivalent substitutions and obvious deformation forms made by any person skilled in the art under the teaching of this specification fall within the substantial scope of this specification and should be protected by the present invention.
Claims
1. A multi-time scale voltage regulation method based on deep reinforcement learning, characterized in that: The steps include: Step 1. Model the active distribution network with a large amount of renewable energy and establish a voltage optimization model for the active distribution network; Step 2. Establish an intelligent agent with continuous action and discrete action characteristics with time series characteristics; The intelligent agent includes two networks, one is a slow network for discrete control on the hourly time scale for on-load voltage regulators and capacitor bank dispatching, and the other is a fast network for continuous control on the minute-level time scale for inverter and electric energy storage system dispatching; the fast network and slow network of the intelligent agent have two internal networks, which are defined as fast control Q network, fast control target Q network, slow control Q network, and slow control target Q network respectively; The action set of the fast network and the action set of the slow network are stacked to obtain the action of the agent at a certain moment; Step 3. Use a single-layer Markov to construct a multi-time scale decision process to describe the interaction process between the active distribution network voltage optimization model and the intelligent agent, and introduce the distribution network state, fast network state, and slow network state; A state transfer function with three coupled states is specially designed, and a time counter is used to separate the fast time scale from the slow time scale. The output of the time counter is used as an activation signal to control the slow network. definition The transfer function describing the environment, i.e. the ADN state, The transition function describing the internal state of the fast network, A transition function describing the internal state of the slow network; Then the transition probability function The multi-time scale decision process is constructed using a single-layer Markov as follows: a) State space: Define states t =(x t ,y t ,z t ); Among them, x t Is the current state of the environment, i.e. ADN y t is the internal state of the fast network; z t is the internal state of the slow network; in, Indicates the load active power demand, Indicates the reactive power demand of the load, V i,t represents the voltage amplitude of node i at time t, represents the power content of the energy storage system at node i at time t, They represent the active power output of the photovoltaic power generation device and the active power output of the wind turbine connected to the node i at time t respectively; b) Action Space: Defining actions in, is the action of the fast network at time t, is the action of the slow network at time t; in, Indicates the control rate of the reactive power of the photovoltaic inverter, Indicates the control rate of the active power of the inverter of the wind power generation device; Indicates the active power control rate of the energy storage device; For photovoltaic and wind turbine inverters, first use and Get the available reactive power, and then get the reactive power injected into the ADN under the corresponding control rate in, Represents the reactive power available from a photovoltaic power generation device, Represents the reactive power available to the wind turbine; are the rated capacities of photovoltaic and wind turbine inverters respectively, and represents the control rate; For energy storage systems, the control rate The actual active power injected into the ADN is obtained under the influence of if but if but in Indicates the position of the tap contact of the OLTC. Indicates the number of capacitor banks connected to the ADN, that is, Determine the tap of the on-load voltage regulator OLTC t location, through Determine the number of groups of capacitors connected to the distribution network c) Transition probability: Transition probability P(s t+1 |s t ,a t ) is further decomposed into: in, The transfer function describing the environment, i.e. the ADN state, The transition function describing the internal state of the fast network, A transition function describing the internal state of the slow network; d) Counter: For slow networks, introduce a counter c t to keep track of time steps; According to the result of the counter, when the hour time scale of the slow network action is reached, the slow network should take a new action; at other time steps, it maintains the state of the previous action; If the time step reaches the hour unit, then c t =0; If the time step is in minutes, then c t ≠0, then there is a slow network action selection: Step 4. A redundant multi-agent cooperative control system for multi-time-scale voltage regulation based on deep reinforcement learning for active distribution network is proposed; multiple agents are used in the system to perform the same voltage regulation task, and then the action output of the agent is adjusted through a redundant coordination mechanism, and finally the coordinated action is applied to the active distribution network; the redundant coordination mechanism includes establishing a redundant coordination matrix, coordinating the actions between multiple agents, and updating the redundant coordination matrix; Step 5. Establish the training process of the intelligent agent; apply the trained intelligent agent, combined with the redundant multi-agent collaborative control system, to the voltage control of the active distribution network to achieve multi-time scale voltage regulation.
2. The multi-time scale voltage regulation method based on deep reinforcement learning according to claim 1, characterized in that: In step 1, the distribution network voltage optimization model established includes: Line characteristic model of active distribution network, on-load voltage regulator model, capacitor bank model, inverter model installed at photovoltaic power generation system, inverter model installed at wind power generation system and electric energy storage system model; The distribution network voltage optimization model is used to calculate the sum of active power losses of each line within the time range T, the number of distribution network nodes and the real-time voltage value of each node, in which a safety module mechanism is used for the electric energy storage system model.
3. The multi-time-scale voltage regulation method based on deep reinforcement learning according to claim 2 is characterized in that: In the above, the formula of the electric energy storage system model is as follows: Among them, E i min 、E i max are the minimum and maximum energy storage levels of the energy storage system; Respectively represent the power content of the energy storage system at node i at time t and t-1; represents the charging power of the energy storage system at node i at time t; Indicates the maximum charging power of the energy storage system; Indicates the maximum discharge power of the energy storage system; represents the discharge power of the energy storage system at node i at time t; η i c represents the equipment efficiency of the energy storage system when charging at node i; η i d represents the equipment efficiency of the energy storage system at node i when discharging; Δt represents the time step; T represents the time range, and N is the number of distribution network nodes; The electric energy storage system model is optimized using the safety module mechanism. The optimization formula is as follows: in, and They represent the power content of the energy storage system at node i at time t+1 and t respectively.
4. The multi-time scale voltage regulation method based on deep reinforcement learning according to claim 1, characterized in that: The step 2 is specifically as follows: Step 2.
1. Establishment of time scale control network; Designed based on Dueling Deep Q-Network algorithm, the Q value is decomposed into two parts: state value and advantage function. The state value function and advantage function use a shared feedforward neural network for feature extraction. For the fast network and the slow network, a feedforward neural network h is established respectively. fast 、h slow : Among them, s t is the state input; W1 and W2 are weight matrices, b1 and b2 are bias terms; subscripts 1 and 2 represent the number of layers of the neural network, i.e. 1 represents the first layer of the neural network and 2 represents the second layer of the neural network; ReLU is the activation function; The neural network parameters are represented by θ, that is: θ={W1,b1,W2,b2}; Then the fast network parameter θ fast ={W1 fast ,b1 fast ,W2 fast ,b2 fast } and the slow network parameter θ slow ={W1 slow ,b1 slow ,W2 slow ,b2 slow }; Establish state value function V for fast network and slow network respectively fast (s t ) and V slow (s t ): in, represents the network parameters used to calculate the state value network in the fast network, is the weight matrix, is the bias term; Represents the network parameters used to calculate the state value network in the slow network. is the weight matrix, is the bias term; Subscript 1 and subscript 2 represent the number of layers of the neural network, i.e. 1 represents the first layer of the neural network and 2 represents the second layer of the neural network; Establish advantage functions for fast networks and slow networks respectively and in, represents the network parameters used to calculate the advantage value network in the fast network, is the weight matrix, is the bias term; represents the network parameters used to calculate the advantage value network in the slow network, and similarly is the weight matrix, is the bias term; subscript 1 and subscript 2 represent the number of layers of the neural network, i.e. 1 represents the first layer of the neural network and 2 represents the second layer of the neural network; Establish Q value functions for fast networks and slow networks respectively and in, represents the action space of the fast network; Indicates the next action of the fast network, and represents the action space of the slow network; Indicates the next action of the slow network, and and It is the average value of the advantage function of all actions, which is used to normalize the advantage value to ensure that the calculation of Q value is more stable; The output of the fast network is converted into a continuous action value a through the tanh function. fast , map the action to the range [-1,1]; a fast =tanh(Q fast ); The output of the slow network is transformed from the discrete action set into a argmax operation. Select the best action; where 1 , 2,…,N represents the gear position of the on-load voltage regulator OLTC or the number of capacitors CB connected to ADN: Step 2.
2. Action splicing: In multiple time frames, the fast network and the slow network generate fast actions and slow actions alternately at different time scales, respectively; the details are as follows: Among them, n and m are the number of continuous actions and discrete actions respectively; represents the new set of actions for the fast network, represents the old set of actions for the slow network, Indicates new actions for slow networks; Represents the new action set content of the fast network, Represents the old action set content of the slow network, Represents the new action set content of the slow network; Under the action of the time counter, the fast network will output a new action value at each time step, while the slow network needs to wait for the activation signal of the counter; the slow network will output a new action value only after receiving the activation signal of the counter, and will output the last updated action, that is, the original old action, when it does not receive the activation signal of the counter; The action set of the fast network and the action set of the slow network are stacked, and the action a of agent i at time t is i,t for:
5. The multi-timescale voltage regulation method based on deep reinforcement learning according to claim 1, characterized in that: The step 4 is specifically as follows: Step 4.
1. Establishment of redundant coordination matrix; Suppose there are L agents, the initial redundant coordination matrix C0 is an L×L identity matrix; The above formula means that initially, each agent only refers to its own actions and has no influence on each other; Step 4.
2. Coordination of actions between multiple agents; At a certain time t, the action of agent l is represented by a vector a l,t , l∈L; The actions of all agents are concatenated to form a redundant action matrix with a dimension of L×(n+m). The redundant action matrix is connected through the communication matrix C t Coordinate and produce coordinated actions Among them C t is the redundant coordination matrix at the tth moment, which represents the influence weight of each agent on the actions of other agents; n and m are the number of continuous actions and discrete actions, respectively; Step 4.
3. Update of redundant coordination matrix; The redundant coordination matrix is updated according to the system performance after each action is performed; The performance indicator of the system is the reward r t ; Each time the matrix is updated, noise ∈ is added, and then the performance r t Adjust the matrix: C t+1 =C t +∈·r t ; Among them, noise ∈ is a random matrix that obeys the normal distribution: Among them, σ represents the parameter of normal distribution. In order to ensure that the weights in the communication matrix are within the preset reasonable range and keep the sum of each row equal to 1, the communication matrix needs to be normalized; Assume that after the matrix update in step t+1 is completed, the new matrix is C t , the normalization process is as follows: Among them, C t+1 (K,J) represents the matrix C t+1 The element in row K and column J, C t+1 (K,J)' is the normalized matrix.
6. The multi-timescale voltage regulation method based on deep reinforcement learning according to claim 1, characterized in that: In step 5, during the training process, the fast network and the slow network use their respective target Q networks. To update the Q value, the target Q value and The calculation formula is as follows: Where r is the immediate reward and γ represents the discount factor; Indicates the fast control target Q network at input s t+1 Output the Q value of all possible actions, and select the action a that maximizes the Q value from these actions fast ', and then use this maximized Q value to calculate the target Q value; Indicates that the slow control target Q network is at input s t+1 Output the Q value of all possible actions, and select the action a that maximizes the Q value from these actions slow 'Then use this maximized Q value to calculate the target Q value; In order to optimize the network parameters of the fast network and the slow network, the loss function L is established using the target Q value and Q value for the two networks respectively. fast (θ fast ) and L slow (θ slow ), and then use gradient descent to minimize the loss function; Where η is the learning rate, represents the expected value, Denotes the fast control Q network in the network parameter θ fast Next, in state s t Take action when The Q value obtained; Indicates that the slow control Q network has a network parameter θ slow Next, in state s t Take action when The Q value obtained is is the loss function of the fast control Q network with respect to the parameter θ fast The gradient of is the loss function of the slow control Q network with respect to the parameter θ slow The gradient of In order to stabilize the training process, the target Q network is quickly controlled The parameter θ fast ' and slow control target Q network The parameter θ slow 'It is not updated in every training step, but every preset number of time steps, the parameters θ of the current fast control Q network are updated fast and the slow control Q network parameter θ slow Copy it, that is: i fast '←θ fast ; i slow '←θ slow ; The priority experience replay introduced in the Dueling Deep Q-Network algorithm is as follows: Experience replay stores experience tuples (s t ,a t ,r t ,s t+1 ), different experience tuples are given different priorities through the experience evaluation function, so that the experience that has a greater impact on the learning process is preferentially selected for playback and training; Each time a new experience is stored in the buffer, its error is calculated and a priority is assigned to each experience according to the error; for the hth group of experience, its experience evaluation function p h for; p h =|r+γmaxQ target (s t+1 ,a')-Q(s t ,a t )|+∈; Among them, maxQ target (s t+1 ,a') represents the target Q network input s t+1 , output the Q values of all possible actions, select the action a' that maximizes the Q value from these actions, and then use this maximized Q value to calculate the target Q value; maxQ target (s t+1 ,a') is given by and Get the average value; Q(s t ,a t ) indicates that the Q network is in the state of s t Take action a t The Q value obtained; Q(s t ,a t ) is given by and Average is obtained; The probability P(h) of experience sampling corresponding to the hth group of experience is: represents the sum of the priorities of all experiences used to normalize the priority of each experience; u is the total number of experiences in the buffer; an importance sampling weight w is used when updating the Q network i Correct the sampled samples, the calculation process is as follows: β is a parameter that controls the smoothness of the weights; The Thompson sampling introduced in the Dueling Deep Q-Network algorithm is as follows: Thompson sampling naturally combines exploration and exploitation through random sampling. Even if the current average reward of an action is high, other actions may still be sampled due to the uncertainty of its distribution. Therefore, the choice of action is dynamically adjusted according to its uncertainty, and the specific steps are as follows: Suppose that each action a h Rewards h The distribution is Beta distribution: P(r h )~Beta(α h ,b h ); Among them, P(r h ) indicates action a h According to the reward r h The distribution of h , β h Represent the number of successes and failures respectively, with an initial value of 1; whenever an action needs to be selected, sample from the posterior reward distribution of each action and select the action with the highest reward, that is, under a given state s, sample from the reward distribution Beta (α h ,β h ) to generate a sample value δ for each action h Then select the action a with the largest sample value * ,Right now: For each action a h , sample a value from its corresponding Beta distribution: d h ~Beta(α h ,b h ); Among them, δ h is a sample value drawn from a Beta distribution, and then the action with the largest sample value is selected: a * =argmax h d h ; Update the corresponding Beta distribution parameters according to the action reward r; If the reward r is positive, increase the number of successes of the action by α h : a h ←a h +max(0,r); If the reward r is zero or negative, increase the number of failures for that action by β h : b h ←b h +max(0,1-r).
7. The multi-time scale voltage regulation method based on deep reinforcement learning according to claim 1, characterized in that: In step 5, the training is divided into a multi-time scale action process, that is, a multi-time scale decision process and a network parameter update process based on a single-layer Markov model, as follows: The multi-time scale decision-making process based on single-layer Markov is as follows: At a certain moment, the active distribution network inputs the state of the distribution network to the intelligent agent for multi-time scale control through the active distribution network voltage optimization model. The time counter of the intelligent agent outputs a control signal according to the time scale for the activation of the fast network and the slow network. The fast network and the slow network of the intelligent agent randomly give control actions according to the state. After the two network actions are integrated, the redundant coordination machine is used to coordinate the actions, and then the action is applied to the active distribution network voltage optimization model. The distribution network completes the state update according to the state transfer principle; at this time, a reward value is obtained using the new state; In this process, the state of the distribution network s t , Action a acting on the voltage optimization model of the active distribution network t , reward value r t , new state t+1 Composition of experience tuples (s t ,a t ,r t ,s t+1 ), then use the experience evaluation function to evaluate the experience tuple to obtain an evaluation result of the experience tuple, and store the experience tuple into the experience playback mechanism according to the evaluation result; Repeat this process until the upper limit of the experience replay mechanism is reached; Then the agent network parameter update process is carried out, as follows: Extract a small batch of experience tuples from the experience replay mechanism; then establish the reward distribution of the action based on the reward and action, and then select the action for updating the parameters based on the reward distribution of the action; complete the update of the network parameters according to the internal structure of the intelligent agent; at the same time, according to the action a t and the reward value r t Update the coordination matrix of the redundant coordination mechanism.
8. The multi-timescale voltage regulation method based on deep reinforcement learning according to claim 1, characterized in that: In step 5, the trained intelligent agent and the coordination matrix adjusted by the redundant coordination mechanism are obtained. The network parameters and coordination matrix of the intelligent agent remain fixed after training. The active distribution network voltage optimization model inputs the state to the trained intelligent agent, and the intelligent agent gives a control action. The coordination matrix adjusted by the redundant coordination mechanism directly acts on the active distribution network voltage optimization model, that is, the control is completed according to the multi-time scale decision-making process constructed based on the single-layer Markov.
9. A multi-time-scale voltage regulation system based on deep reinforcement learning, used to implement the multi-time-scale voltage regulation method based on deep reinforcement learning as claimed in any one of claims 1 to 8; It is characterized in that The multi-timescale voltage regulation system based on deep reinforcement learning includes the following modules: A model building module is used to model the active distribution network with a large amount of renewable energy and establish a voltage optimization model for the active distribution network; The agent building module is used to build agents with continuous and discrete action characteristics with time series characteristics; The intelligent agent includes two networks, namely a slow network for discrete control on the hourly time scale for on-load voltage regulators and capacitor bank dispatching, and a fast network for continuous control on the minute-level time scale for inverter and electric energy storage system dispatching; the fast network and slow network of the intelligent agent have two internal networks, which are defined as fast control Q network, fast control target Q network, slow control Q network, and slow control target Q network respectively; The interactive process description module uses a single-layer Markov model to construct a multi-time scale decision process to describe the interactive process between the active distribution network voltage optimization model and the intelligent agent, and introduces the distribution network state, fast network state, and slow network state. A state transfer function with three coupled states is specially designed, and a time counter is used to separate the fast time scale from the slow time scale. The output of the time counter is used as an activation signal to control the slow network. Multi-agent cooperative control system building block, used to propose a redundant multi-agent cooperative control system for multi-timescale voltage regulation based on deep reinforcement learning for active distribution networks; Use multiple agents in the system to perform the same voltage regulation task, then adjust the action output of the agents through a redundant coordination mechanism, and finally apply the coordinated actions to the active distribution network; And the voltage control module of the active distribution network is used to apply the trained intelligent agent, combined with the redundant multi-agent collaborative control system, to the multi-time scale voltage control of the active distribution network to achieve multi-time scale voltage regulation.
Citation Information
Patent Citations
Power distribution network multi-time scale reactive voltage control method based on reinforcement learning
CN113489015A
Cited By
Low-voltage distribution network voltage control method and system based on deep reinforcement learning
CN119813234B