A robust control method for delayed adaptive voltage in active power distribution networks
Patent Information
- Application Number
- CN202311730802.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-12-15
AI Technical Summary
[0005]本发明为解决现有方法中未考虑延迟大小未知而导致的电压控制不准确等一系列问题,提供一种主动配电网中延迟自适应电压的鲁棒控制方法,以期能获得具有延迟自适应特性的分散式电压无功控制方案,从而使主动配电网电压稳定,安全运行
[0105] 1. This invention divides the entire active distribution network into different sub-regions and assigns tasks to multiple controllers so that each controller can solve the voltage control problem in the sub-region. This achieves global voltage control with minimal information exchange, thereby effectively reducing communication costs.
Smart Images

Figure CN117713206B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of active distribution network voltage control technology, and particularly relates to a robust control method for delayed adaptive voltage in active distribution networks. Background Technology
[0002] The penetration rate of photovoltaics (PV) in active distribution networks continues to rise, and the back propagation of PV power can easily cause overvoltage hazards, affecting the safe operation of the distribution network. To suppress overvoltage in active distribution network systems, the IEEE 1547 standard, for the first time, allows distributed small-capacity PV inverters to participate in voltage control of the distribution network by outputting reactive power. Compared with traditional voltage control equipment, inverters can respond quickly to voltage fluctuations and reduce network power losses through reactive power compensation. Typically, by considering the system state of the active distribution network (such as load active / reactive power, PV generation, etc.), the controller can design an optimized reactive power scheme for the PV inverter.
[0003] Voltage control in active distribution networks can be categorized into centralized control, local control, and distributed control. When using a centralized control framework, the controller needs to collect all data and make decisions for the entire active distribution network. Therefore, centralized voltage control incurs high computational costs and communication burdens, making it unsuitable for large-scale active distribution networks. The main drawback of local voltage control is the inability to ensure consistent control strategies across different distribution networks. Distributed control enables global voltage control within the active distribution network with minimal information exchange. By dividing the entire active distribution network into different sub-regions and assigning tasks to multiple controllers, each controller can address the voltage control problem within its sub-region.
[0004] However, distributed voltage control involves delays, sometimes reaching up to 30 seconds. These delays include communication delays from system state measurement to controller reception, optimization solution time, communication delays from controller to inverter scheduling commands, and inverter response time. Due to the time-varying nature of these delays, the delay for each future operation time step cannot be precisely known. When the active distribution network experiences severe fluctuations, the system state changes significantly within the delay time, leading to time inconsistencies between the system state at the sampling and control times. This impacts the accuracy of voltage control based on the sampling data. Summary of the Invention
[0005] To address the problems of inaccurate voltage control caused by the lack of consideration for unknown delay magnitude in existing methods, this invention provides a robust control method for delay-adaptive voltage in active distribution networks. The aim is to obtain a distributed voltage and reactive power control scheme with delay-adaptive characteristics, thereby stabilizing the voltage and ensuring safe operation of the active distribution network.
[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0007] The robust control method for delayed adaptive voltage in an active distribution network, as described in this invention, is characterized by the following steps:
[0008] Step 1: Obtain basic parameter information of the active distribution network, including: basic information and parameter information of the active distribution network;
[0009] Step 1.1: Obtain the basic information of the active distribution network:
[0010] The main branches of the active distribution network are determined, and the distribution network is divided into M regions based on the shortest distance from each branch to the main branch. The set of load nodes in the m-th region is denoted as B. m The set of branches is denoted as E. m The set of nodes equipped with photovoltaic inverters is denoted as D. m And obtain the resistance r on the branch between any i-th node and j-th node in the m-th region. ij Reactance x ij Conductivity g ij and susceptance b ij ;
[0011] Step 1.2: Obtain the parameter information of the active distribution network, including: the active power of the load on the i-th node. and reactive power Active power of photovoltaic inverter and reactive power
[0012] Step 2, using the delay value T d To predict the duration, the confidence interval for the parameter information of the active distribution network is predicted:
[0013] When there is a load and photovoltaic power generation equipment on the i-th node, t+T can be predicted using equations (1)-(3). d Active power of load at any time and reactive power and the active power of photovoltaic inverters The confidence interval is 95%;
[0014]
[0015]
[0016]
[0017] In equations (1)-(3), The load active power at time t, Let be the reactive power at time t. Let be the active power of the photovoltaic inverter at time t;
[0018] Step 3: Use equations (4)-(7) to obtain the operational state range sos of the i-th node at time t. i,t ;
[0019]
[0020]
[0021]
[0022]
[0023] In equations (4)-(7), Represents the active power of the load at time t. and reactive power and the active power of photovoltaic inverters The upper bound of the confidence interval; Represents the active power of the load at time t. and reactive power and the active power of photovoltaic inverters The lower bound of the confidence interval; Represents the active power of the load at time t. and reactive power and the active power of photovoltaic inverters The median of the confidence interval; up , . low and. med These represent the upper bound, lower bound, and median of the confidence interval, respectively.
[0024] Step 4: Establish a regional collaborative robust voltage and reactive power control model that minimizes the total node voltage deviation and network loss, including: objective function, variable constraints, power flow constraints, power constraints of photovoltaic inverters, and node voltage constraints;
[0025] Step 4.1: Construct the objective function F of the regional collaborative robust reactive voltage control model at time t using equation (11). t :
[0026]
[0027] In equation (11), ψ t Represents a set of Boolean variables, and This represents the J-th Boolean variable at time t. Indicates the operation status is The objective function value for the m-th region;
[0028] Step 4.2: Construct a function about ψ using equation (12). t Variable constraints;
[0029]
[0030] Step 4.3: Construct the power flow equation constraints for the i-th node at time t using equations (13)-(14);
[0031]
[0032]
[0033] In equations (13)-(14), v j,t Let θ represent the voltage magnitude of the j-th node at time t. i,t and θ j,t This represents the voltage phase angle between the i-th node and the j-th node;
[0034] Step 4.4: Construct the photovoltaic inverter power of the i-th node at time t using equations (15)-(17). Constraints;
[0035]
[0036]
[0037]
[0038] In equations (15)-(17), and Let represent the active power of the photovoltaic inverter at the i-th node at time t. The minimum and maximum values of β, where β represents the capacity factor of the photovoltaic inverter. This represents the apparent power of the photovoltaic inverter at the i-th node;
[0039] Step 4.5: Construct the voltage constraint of the i-th node at time t using equation (18);
[0040] v min ≤v i,t ≤v max (18)
[0041] In equation (18), v min and v max These are the upper and lower bounds of the safe range for node voltages, respectively.
[0042] Step 5: Reconstruct the regional collaborative robust reactive voltage control model using the POMDP model:
[0043] Step 5.1: Define each agent in the POMDP model as managing a partition of an active distribution network, thus obtaining a multi-agent set consisting of M agents.
[0044] Step 5.2: Construct the state space at time t in the robust voltage control POMDP model using equation (19).
[0045]
[0046] In equation (19), This indicates that the operating state range at time t-1 is... The reactive power control command for the photovoltaic inverter at the i-th node is obtained, V i,t ={v i,t ,θ i,t} represents the voltage vector of the i-th node at time t, which includes the voltage magnitude and phase angle;
[0047] Step 5.3: Construct the action space at time t in the POMDP model of regional collaborative robust reactive voltage control using equations (20)-(21).
[0048]
[0049]
[0050] In equations (20)-(21), This represents the set of actions of the m-th agent at time t;
[0051] Step 5.4: Construct the local observation set at time t in the POMDP model of regional collaborative robust reactive voltage control using equations (22)-(23).
[0052]
[0053]
[0054] In equations (22)-(23), This represents the set of local observations of the m-th agent at time t;
[0055] Step 5.5: Construct the reward function r at time t in the POMDP model of regional collaborative robust reactive voltage control using equation (24). t ;
[0056]
[0057] Step 6: Use a multi-agent reinforcement learning algorithm to train the regional collaborative robust reactive voltage control POMDP model offline;
[0058] Step 6.1: Construct the current policy network, denoted as... The target policy network is denoted as μ′ m The current Q-value network is denoted as The target Q-value network is denoted as
[0059] Step 6.2: Construct the training objective function of the multi-agent reinforcement learning algorithm using equation (25).
[0060]
[0061] In equation (25), Represents the current policy network at time t. Network parameters, ρ π Indicates the state The probability of taking a certain action, where γ∈[0,1] represents the discount factor;
[0062] Step 6.3: Obtain a robust voltage control reinforcement learning model through offline training using a multi-agent reinforcement learning algorithm.
[0063] Step 7: By performing delay adaptive processing on the robust voltage control reinforcement learning model, a robust voltage control reinforcement learning model with delay adaptive characteristics is generated, and the robust voltage control reinforcement learning model with delay adaptive characteristics is deployed in each section of the active distribution network.
[0064] The robust control method for delayed adaptive voltage in an active distribution network described in this invention is also characterized in that step 4.1 includes the following steps:
[0065] Step 4.1.1: Use equation (8) to obtain the voltage deviation of the i-th node at time t.
[0066]
[0067] In equation (8), v ref Indicates the reference voltage; v i,t This represents the voltage magnitude of the i-th node at time t;
[0068] Step 4.1.2: Use equation (9) to obtain the network loss between the i-th node and the j-th node at time t.
[0069]
[0070] In equation (9), Re represents the real part, and V i,t and V j,t This represents the voltage vectors of the i-th and j-th nodes at time t, which include voltage magnitude and phase angle, where j represents the node number;
[0071] Step 4.1.3: Construct the objective function f for the m-th region at time t using equation (10). m,t ;
[0072]
[0073] In equation (10), λ1 and λ2 represent the weighting coefficients of the total node voltage deviation and network loss.
[0074] Step 6.3 includes the following steps:
[0075] Step 6.3.1: Input the basic information of the active distribution network;
[0076] Step 6.3.2: Set training hyperparameters, randomly initialize the network parameters of agents in each region of the active distribution network, and initialize the shared experience replay pool;
[0077] Step 6.3.3: Set the maximum number of training rounds to... Let T be the total time of a single round, and set the current training round.
[0078] Step 6.3.4: Initialize t = 0;
[0079] Step 6.3.5: In the active distribution network, the agents in each region obtain their own region's... Local observation set at time t in training rounds
[0080] Step 6.3.6: Based on the power distribution network status in Step 6.3.5, the agents in each region provide the information for the region in the [number]th [phase / stage]. The set of reactive power output of the distributed photovoltaic inverter at time t in the training round And perform the action;
[0081] Step 6.3.7: Calculate the power flow of the m-th region in the current position. Objective function value at time t in training rounds Therefore, calculate the first The total objective function value at time t in the training rounds
[0082] Step 6.3.8, Comparison The size between them, to determine the first The J-th Boolean variable at time t in the training round Therefore, the first equation (11) is used to determine the second equation. Reward at time t in the training round Determine the first using equations (19)-(21) State at time t in the training round and actions
[0083] Step 6.3.9, each area enters the... State at time t+1 in the training round Agents in each region will utilize local experience Stored in the shared experience replay pool middle;
[0084] Step 6.3.10: Agents in each region retrieve data from the shared experience replay pool. Sampling is performed, and the respective network parameters are updated using the back gradient propagation algorithm;
[0085] Step 6.3.11: If t < T, then assign t+1 to t and return to step 6.3.5; otherwise, execute step 6.3.12.
[0086] Step 6.3.12: Use equation (26) to calculate the agent in each region at the 1st rank. Convergence metrics during training rounds
[0087]
[0088] Step 6.3.13, if Then Assign to Then, return to step 6.3.4; otherwise, proceed to step 6.3.14.
[0089] Step 6.3.14: Set the convergence metric to ε. If If the training is successful, it means that all agents have completed training, and the trained current policy network is used as the robust voltage control reinforcement learning model, and step 7 is executed; otherwise, return to step 6.3.2, reset the hyperparameters, and train again.
[0090] Step 7 includes the following steps:
[0091] Step 7.1: Set the maximum number of operation steps to N, and initialize the current number of operation steps to n = 1;
[0092] Step 7.2: Analyze historical latency values and derive the average latency, denoted as . The variance of the delay is denoted as The delay range is denoted as The delay range is divided into N intervals;
[0093] Step 7.3: Randomly select a value from the nth interval and denote it as... Assume the delay follows a Gaussian distribution, denoted as in, express The mean, express The variance;
[0094] Step 7.4, with As the prediction duration, proceed to step 3 to generate the delay. Operating state range of all nodes in the active distribution network at time t Where B represents the set of all nodes in the active distribution network;
[0095] Step 7.5: Execute step 6.3 to obtain the delay. Robust voltage control reinforcement learning model;
[0096] Step 7.6: Use the trained robust voltage control reinforcement learning model to derive the delay as... The reactive power output of the photovoltaic inverter at the node with photovoltaic equipment installed in the active distribution network at time t Where D represents the set of all nodes in the active distribution network that have photovoltaic devices installed;
[0097] Step 7.7: If n < N, then assign n+1 to n and return to step 7.3; otherwise, execute step 7.8.
[0098] Step 7.8: Use equation (27) to obtain the reactive power output of the photovoltaic inverter with delay adaptive characteristics at the j-th node with photovoltaic equipment installed at time t.
[0099]
[0100] In equation (27), Indicates delay The probability of;
[0101] Step 7.9: Deploy the robust voltage control reinforcement learning model with delay adaptive characteristics in each zone of the active distribution network.
[0102] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the robust control method, and the processor is configured to execute the program stored in the memory.
[0103] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to perform the steps of the robust control method.
[0104] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0105] 1. This invention divides the entire active distribution network into different sub-regions and assigns tasks to multiple controllers so that each controller can solve the voltage control problem in the sub-region. This achieves global voltage control with minimal information exchange, thereby effectively reducing communication costs.
[0106] 2. This invention analyzes the probability distribution of system delay, reduces the impact of inconsistent state time characteristics of active distribution network system through confidence interval prediction, and realizes the delay adaptive characteristic in voltage control based on the prediction results of system operating state under different prediction durations.
[0107] 3. This invention remodels the problem using the POMDP model and solves it quickly through a multi-agent reinforcement learning algorithm. Unlike previous works, this method not only robustly performs regional voltage control but also achieves delay-adaptive voltage control even when the delay cannot be accurately known. Furthermore, due to the extremely high computational efficiency of reinforcement learning, the entire voltage control process can be performed in real time, exhibiting strong adaptability and practicality. Attached Figure Description
[0108] Figure 1 This is a topology diagram of the IEEE 33-node power distribution system of the present invention;
[0109] Figure 2 This is a schematic diagram illustrating the principle of voltage control using photovoltaic reactive power in the active distribution network according to the present invention.
[0110] Figure 3 This is a flowchart illustrating the training process of the multi-agent reinforcement learning algorithm of this invention. Detailed Implementation
[0111] The invention will be further described below with reference to the accompanying drawings:
[0112] In this embodiment, the steps of a delayed adaptive voltage robust control method in an active distribution network are as follows:
[0113] Step 1: Obtain basic parameter information of the active distribution network, including: basic information and parameter information of the active distribution network;
[0114] Step 1.1: Obtain basic information about the active distribution network:
[0115] exist Figure 1In the active distribution network shown, to perform distributed voltage control, the active distribution network first needs to be partitioned. The main branches of the active distribution network are determined to be nodes 1 to 6. Based on the shortest distance from each branch to the main branch, the distribution network is divided into 4 regions. The set of load nodes in the m-th region is denoted as B. m The set of branches is denoted as E. m The set of nodes equipped with photovoltaic inverters is denoted as D. m And obtain the resistance r on the branch between any i-th node and j-th node in the m-th region. ij Reactance x ij Conductivity g ij and susceptance b ij ;
[0116] Step 1.2: Since the node voltage fluctuations of the active distribution network are affected by the intermittency of photovoltaic power generation and the randomness of the load, it is necessary to obtain the parameter information of the active distribution network, including: the active power of the load on the i-th node. and reactive power Active power of photovoltaic inverter and reactive power
[0117] Step 2: Due to data transmission delays in the active distribution network, controller calculation time, and actuator response time, the parameter information of the active distribution network at the sampling and control times is mismatched. Therefore, the delay value T is used as the basis for the delay. d To predict the duration, predict the confidence interval of the parameters of the active distribution network:
[0118] When there is a load and photovoltaic power generation equipment on the i-th node, t+T can be predicted using equations (1)-(3). d Active power of load at any time and reactive power and the active power of photovoltaic inverters The confidence interval is 95%;
[0119]
[0120]
[0121]
[0122] In equations (1)-(3), The load active power at time t, Let be the reactive power at time t. Let t be the active power of the photovoltaic inverter at time t.
[0123] Step 3: Since the upper and lower bounds of the predicted confidence interval are the values with the lowest prediction accuracy, while the median is closest to the true value, we classify the predicted values with the same prediction characteristics and use equations (4)-(7) to obtain the operating state range sos of the i-th node at time t. i,t ;
[0124]
[0125]
[0126]
[0127]
[0128] In equations (4)-(7), Let represent the upper bound of the confidence interval for the operation state of the i-th node at time t, where and These represent the upper bounds of the confidence intervals for the load active power, load reactive power, and photovoltaic inverter active power, respectively. Let represent the lower bound of the confidence interval for the operation state of the i-th node, where These represent the lower bounds of the confidence intervals for the load active power, load reactive power, and photovoltaic inverter active power, respectively. Let represent the median of the confidence interval for the operation state of the i-th node at time t, where These represent the median of the confidence intervals for the load active power, load reactive power, and photovoltaic inverter active power, respectively.
[0129] Step 4, according to Figure 2 The voltage control schematic shown is designed to improve the power quality for end users and enable power companies to reduce operating and maintenance costs. It establishes a regional collaborative robust voltage and reactive power control model that minimizes the total node voltage deviation and network loss. The model includes: objective function, variable constraints, power flow constraints, power constraints of photovoltaic inverters, and node voltage constraints.
[0130] Step 4.1: Construct the objective function of the regional collaborative robust reactive voltage control model.
[0131] Step 4.1.1: Use equation (8) to obtain the voltage deviation of the i-th node at time t.
[0132]
[0133] In equation (8), v ref The reference voltage is represented as 1 p.u.; v i,t This represents the voltage magnitude of the i-th node at time t;
[0134] Step 4.1.2: Use equation (9) to obtain the network loss between the i-th node and the j-th node at time t.
[0135]
[0136] In equation (9), Re represents the real part, and V i,t and V j,t Let represent the voltage vectors of the i-th and j-th nodes at time t, which include voltage magnitude and phase angle, where j represents the node number.
[0137] Step 4.1.3: Construct the objective function f for the m-th region at time t using equation (10). m,t ;
[0138]
[0139] In equation (10), λ1 and λ2 represent the weighting coefficients of the total node voltage deviation and network loss, both of which are taken as 0.5;
[0140] Step 4.1.4: Construct the objective function F at time t using equation (11). t ;
[0141]
[0142] In equation (11), ψ t Represents a set of Boolean variables, and Represents the J-th Boolean variable at time t, using The node operation state corresponding to the maximum target value can be selected to ensure robust voltage control; Indicates the operation status is The objective function value for the m-th region.
[0143] Step 4.2: Construct a function about ψ using equation (12). t Variable constraints;
[0144]
[0145] Step 4.3: Construct the power flow equation constraints for the i-th node at time t using equations (13)-(14);
[0146]
[0147]
[0148] In equations (13)-(14), v j,t Let θ represent the voltage magnitude of the j-th node at time t. i,tand θ j,t This represents the voltage phase angle between the i-th node and the j-th node.
[0149] Step 4.4: Construct the photovoltaic inverter power of the i-th node at time t using equations (15)-(17). Constraints;
[0150]
[0151]
[0152]
[0153] In equations (15)-(17), and Let represent the active power of the photovoltaic inverter at the i-th node at time t. The minimum and maximum values are determined based on the minimum and maximum values of photovoltaic power generation data throughout the day. β represents the capacity factor of the photovoltaic inverter, which is set to 0.8. Let represent the apparent power of the photovoltaic inverter at the i-th node, taken as .
[0154] Step 4.5: Construct the voltage constraint of the i-th node at time t using equation (18);
[0155] v min ≤v i,t ≤v max (18)
[0156] In equation (18), v min and v max These are the upper and lower limits of the safe range of the node voltage, respectively, and are set to 0.95 pu and 1.05 pu.
[0157] Step 5: Reconstruct the regional collaborative robust reactive voltage control model using the POMDP model:
[0158] Step 5.1: Define each agent in the POMDP model as managing a partition of an active distribution network, thus obtaining a multi-agent set consisting of M agents.
[0159] Step 5.2: Construct the state space at time t in the robust voltage control POMDP model using equation (19).
[0160]
[0161] In equation (19), This indicates that the operating state range at time t-1 is... The reactive power control command for the photovoltaic inverter at the i-th node is obtained, V i,t ={v i,t ,θ i,t} represents the voltage vector of the i-th node at time t, which includes the voltage magnitude and phase angle.
[0162] Step 5.3: Construct the action space at time t in the POMDP model of regional collaborative robust reactive voltage control using equations (20)-(21).
[0163]
[0164]
[0165] In equations (20)-(21), Let represent the set of actions of the m-th agent at time t.
[0166] Step 5.4: Construct the local observation set at time t in the POMDP model of regional collaborative robust reactive voltage control using equations (22)-(23).
[0167]
[0168]
[0169] In equations (22)-(23), This represents the set of local observations of the m-th agent at time t;
[0170] Step 5.5: Construct the reward function r at time t in the POMDP model of regional collaborative robust reactive voltage control using equation (24). t ;
[0171]
[0172] Step 6, as follows Figure 3 As shown, the POMDP model for regional collaborative robust reactive voltage control is trained offline using a multi-agent reinforcement learning algorithm.
[0173] Step 6.1, Current Policy Network The target policy network is denoted as μ′. m The network consists of three fully connected layers with identical structure, where two hidden layers each have a dimension of 256. The current Q-value network is denoted as... The target Q-value network is denoted as It also has the same structure of 3 fully connected layers, with the two hidden layers having dimensions of respectively. For local observation Dimensions.
[0174] Step 6.2: Construct the training objective function of the multi-agent reinforcement learning algorithm using equation (25).
[0175]
[0176] In equation (25), Represents the current policy network at time t. Network parameters, ρ π Indicates the state The probability of taking a certain action is given by γ∈[0,1], which represents the discount factor and is set to 0.9.
[0177] Step 6.3: Obtain a robust voltage control reinforcement learning model through offline training using a multi-agent reinforcement learning algorithm.
[0178] Step 6.3.1: Input the basic information of the active distribution network;
[0179] Step 6.3.2: Set training hyperparameters, randomly initialize the network parameters of agents in each region of the active distribution network, and initialize the shared experience replay pool;
[0180] Step 6.3.3: Set the maximum number of training rounds. Set the current training round to 4000 and the training time step T for a single round to 1000.
[0181] Step 6.3.4: Initialize the current training time step t = 0;
[0182] Step 6.3.5: In the active distribution network, the agents in each region obtain their own region's... Local observation set at time t in training rounds
[0183] Step 6.3.6: Based on the power distribution network status in Step 6.3.5, the agents in each region provide the information for the region in the [number]th [phase / stage]. The set of reactive power output of the distributed photovoltaic inverter at time t in the training round And perform the action.
[0184] Step 6.3.7: Calculate the power flow of the m-th region in the current position. Objective function value at time t in training rounds Therefore, calculate the first The total objective function value at time t in the training rounds
[0185] Step 6.3.8, Comparison The size between them, to determine the first The J-th Boolean variable at time t in the training round Therefore, the first equation (11) is used to determine the second equation. Reward at time t in the training round Determine the first using equations (19)-(21) State at time t in the training round and actions
[0186] Step 6.3.9, each area enters the... State at time t+1 in the training round Agents in each region will utilize local experience Stored in the shared experience replay pool middle.
[0187] Step 6.3.10: Agents in each region retrieve data from the shared experience replay pool. Sampling is performed, and the respective network parameters are updated using the back gradient propagation algorithm;
[0188] Step 6.3.11: If t < T, then assign t+1 to t and return to step 6.3.5; otherwise, execute step 6.3.12.
[0189] Step 6.3.12: Use equation (26) to calculate the agent in each region at the 1st rank. Convergence metrics during training rounds
[0190]
[0191] Step 6.3.13, if Then Assign to Then, return to step 6.3.4; otherwise, proceed to step 6.3.14.
[0192] Step 6.3.14: Set the convergence metric to ε. If If the training is successful, it means that all agents have completed training, and the trained current policy network is used as the robust voltage control reinforcement learning model, and step 7 is executed; otherwise, return to step 6.3.2, reset the hyperparameters, and train again.
[0193] Step 7: When the magnitude of the delay cannot be precisely known, precise voltage control cannot be made based on a possible delay. Therefore, by performing delay-adaptive processing on the robust voltage control reinforcement learning model, a robust voltage control reinforcement learning model with delay-adaptive characteristics is generated, and the robust voltage control reinforcement learning model with delay-adaptive characteristics is deployed in various sections of the active distribution network.
[0194] Step 7.1: Set the maximum number of operation steps N to 15, and the current number of operation steps n = 1;
[0195] Step 7.2: Analyze historical latency values to obtain the average latency. The time delay is 6 seconds, and the variance of the delay is denoted as . The delay range is 4. The delay range is [1,10]s, and the delay range is divided into 15 equal parts;
[0196] Step 7.3: Randomly select a value from the nth interval and denote it as... Assume the delay follows a Gaussian distribution, denoted as in, express The mean, express The variance.
[0197] Step 7.4, with As the prediction duration, proceed to step 3 to generate the delay. Operating state range of all nodes in the active distribution network at time t Where B represents the set of all nodes in the entire active distribution network;
[0198] Step 7.5: Execute step 6.3 to obtain the delay. A robust voltage control reinforcement learning model.
[0199] Step 7.6: Use the trained robust voltage control reinforcement learning model to derive the delay as... The reactive power output of the photovoltaic inverter at the node with photovoltaic equipment installed in the active distribution network at time t Where D represents the set of all nodes in the entire active distribution network that have photovoltaic devices installed;
[0200] Step 7.7: If n < N, then assign n+1 to n and return to step 7.3; otherwise, execute step 7.8.
[0201] Step 7.8: Use equation (27) to obtain the reactive power output of the photovoltaic inverter with delay adaptive characteristics at the j-th node with photovoltaic equipment installed at time t.
[0202]
[0203] In equation (27), Indicates delay The probability of;
[0204] Step 7.10: Deploy the robust voltage control reinforcement learning model with delay adaptive characteristics in each zone of the active distribution network.
[0205] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0206] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
Claims
1. A robust control method for delayed adaptive voltage in an active distribution network, characterized in that, Includes the following steps: Step 1: Obtain basic parameter information of the active distribution network, including: basic information and parameter information of the active distribution network; Step 1.1: Obtain the basic information of the active distribution network: The main branches of the active distribution network are determined, and the distribution network is divided into M regions based on the shortest distance from each branch to the main branch. The set of load nodes in the m-th region is denoted as B. m The set of branches is denoted as E. m The set of nodes equipped with photovoltaic inverters is denoted as D. m And obtain the resistance r on the branch between any i-th node and j-th node in the m-th region. ij Reactance x ij Conductivity g ij and susceptance b ij ; Step 1.2: Obtain the parameter information of the active distribution network, including: the active power of the load on the i-th node. and reactive power Active power of photovoltaic inverters and reactive power ; Step 2, with delay value To predict the duration, the confidence interval for the parameter information of the active distribution network is predicted: When there is a load and photovoltaic power generation equipment on the i-th node, the prediction is made using equations (1)-(3). Active power of load at any time and reactive power and the active power of photovoltaic inverters The confidence interval is 95%; (1) (2) (3) In equations (1)-(3), for The active power of the load at any moment, for Reactive power at all times for The active power of the photovoltaic inverter at any given time; Step 3: Use equations (4)-(7) to obtain the value of the i-th node. Operating state range at any time ; (4) (5) (6) (7) In equations (4)-(7), express Active power of load at all times and reactive power and the active power of photovoltaic inverters The upper bound of the confidence interval; express Active power of load at all times and reactive power and the active power of photovoltaic inverters The lower bound of the confidence interval; express Active power of load at all times and reactive power and the active power of photovoltaic inverters The median of the confidence interval; , and These represent the upper bound, lower bound, and median of the confidence interval, respectively. Step 4: Establish a regional collaborative robust voltage and reactive power control model that minimizes the total node voltage deviation and network loss, including: objective function, variable constraints, power flow constraints, power constraints of photovoltaic inverters, and node voltage constraints; Step 4.1: Construct a regional collaborative robust reactive voltage control model using equation (11). objective function at time 1 : (11) In equation (11), Represents a set of Boolean variables, and , This represents the J-th Boolean variable at time t. Indicates the operation status is The objective function value for the m-th region; Step 4.1.1: Using equation (8) to obtain Voltage deviation at time i : (8) In equation (8), Indicates the reference voltage; Indicates that the i-th node is in Voltage amplitude at any given moment; Step 4.1.2, using equation (9) to obtain Network loss between node i and node j at time i ; (9) In equation (9), Represents the real part, and This indicates that the voltage amplitude and phase angle are included. The voltage vectors of the i-th node and the j-th node at time i. Indicates the node number; Step 4.1.3: Construct using equation (10) The objective function of the m-th region at time m ; (10) In equation (10), and The weighting coefficients represent the total voltage deviation at the node and the network loss. Step 4.2: Construct about using equation (12) Variable constraints; (12) Step 4.3: Construct using equations (13)-(14) The power flow equation constraints of the i-th node at time i; (13) (14) In equations (13)-(14), Indicates that the j-th node is in Voltage amplitude at time 10:00 and This represents the voltage phase angle between the i-th node and the j-th node; Step 4.4: Construct using equations (15)-(17) Photovoltaic inverter power at the i-th node at time i Constraints; + (15) (16) (17) In equations (15)-(17), and They represent Active power of the photovoltaic inverter at the i-th node at time i The minimum and maximum values, Indicates the capacity factor of a photovoltaic inverter. This represents the apparent power of the photovoltaic inverter at the i-th node; Step 4.5: Construct using equation (18) Voltage constraint of the i-th node at time i; (18) In equation (18), and These are the upper and lower bounds of the safe range for node voltages, respectively. Step 5: Reconstruct the regional collaborative robust reactive voltage control model using the POMDP model: Step 5.1: Define each agent in the POMDP model as managing a partition of an active distribution network, thus obtaining a multi-agent set consisting of M agents. ; Step 5.2: Construct the robust voltage control POMDP model using equation (19). State space at any given moment ; (19) In equation (19), This indicates that the operating state range at time t-1 is... The reactive power control command for the photovoltaic inverter at the i-th node is obtained. This indicates that the voltage amplitude and phase angle are included. The voltage vector of the i-th node at time i; Step 5.3: Construct the POMDP model for regional collaborative robust reactive voltage control using equations (20)-(21). Moment Action Space ; (20) (21) In equations (20)-(21), This represents the set of actions of the m-th agent at time t; Step 5.4: Construct the POMDP model for regional collaborative robust reactive voltage control using equations (22)-(23). Local observation set at time ; (22) (23) In equations (22)-(23), This represents the set of local observations of the m-th agent at time t; Step 5.5: Construct the POMDP model for regional collaborative robust reactive voltage control using equation (24). Time-based reward function ; (24) Step 6: Use a multi-agent reinforcement learning algorithm to train the regional collaborative robust reactive voltage control POMDP model offline; Step 6.1: Construct the current policy network, denoted as... The target policy network is denoted as The current Q-value network is denoted as The target Q-value network is denoted as ; Step 6.2: Construct the training objective function of the multi-agent reinforcement learning algorithm using equation (25). ; (25) In equation (25), Represents the current policy network at time t. Network parameters, Indicates the state The probability of taking a certain action. Indicates the discount factor; Step 6.3: Obtain a robust voltage control reinforcement learning model through offline training using a multi-agent reinforcement learning algorithm. Step 7: By performing delay adaptive processing on the robust voltage control reinforcement learning model, a robust voltage control reinforcement learning model with delay adaptive characteristics is generated, and the robust voltage control reinforcement learning model with delay adaptive characteristics is deployed in each section of the active distribution network; Step 7.1: Set the maximum number of operation steps to N, and initialize the current number of operation steps as follows. ; Step 7.2: Analyze historical latency values and derive the average latency, denoted as . The variance of the delay is denoted as The delay range is denoted as The delay range is divided into N intervals; Step 7.3: Randomly select a value from the nth interval and denote it as... Assume the delay follows a Gaussian distribution, denoted as . ;in, express The mean, express The variance; Step 7.4, with As the prediction duration, proceed to step 3 to generate the delay. Operating state range of all nodes in the active distribution network at time t Where B represents the set of all nodes in the active distribution network; Step 7.5: Execute step 6.3 to obtain the delay. Robust voltage control reinforcement learning model; Step 7.6: Use the trained robust voltage control reinforcement learning model to derive the delay as... The reactive power output of the photovoltaic inverter at the node with photovoltaic equipment installed in the active distribution network at time t Where D represents the set of all nodes in the active distribution network that have photovoltaic devices installed; Step 7.7: If n < N, then assign n+1 to n and return to step 7.3; otherwise, execute step 7.
8. Step 7.8: Use equation (27) to obtain the reactive power output of the photovoltaic inverter with delay adaptive characteristics at the j-th node with photovoltaic equipment installed at time t. : (27) In equation (27), Indicates delay The probability of; Step 7.9: Deploy the robust voltage control reinforcement learning model with delay adaptive characteristics in each zone of the active distribution network.
2. The robust control method for delayed adaptive voltage in an active distribution network according to claim 1, characterized in that, Step 6.3 includes the following steps: Step 6.3.1: Input the basic information of the active distribution network; Step 6.3.2: Set training hyperparameters, randomly initialize the network parameters of agents in each region of the active distribution network, and initialize the shared experience replay pool; Step 6.3.3: Set the maximum number of training rounds to... The total time of a single round is T. Set the current training round. ; Step 6.3.4: Initialize t = 0; Step 6.3.5: In the active distribution network, the agents in each region obtain their own region's... Local observation set at time t in training rounds ; Step 6.3.6: Based on the power distribution network status in Step 6.3.5, the agents in each region provide the information for the region in the [number]th [phase / stage]. The set of reactive power output of the distributed photovoltaic inverter at time t in the training round And perform the action; Step 6.3.7: Calculate the power flow of the m-th region in the current position. Objective function value at time t in training rounds , thus calculating the first The total objective function value at time t in the training rounds ; Step 6.3.8, compare { , The size between} is used to determine the first The J-th Boolean variable at time t in the training round Thus, the first equation (11) is used to determine the second equation. Reward at time t in the training round The first equation is determined using equations (19)-(21). State at time t in the training round and actions ; Step 6.3.9, each area enters the... State at time t+1 in the training round Agents in each region will utilize local experience Stored in the shared experience replay pool middle; Step 6.3.10: Agents in each region retrieve data from the shared experience replay pool. Sampling is performed, and the respective network parameters are updated using the back gradient propagation algorithm; Step 6.3.11: If t < T, then assign t+1 to t and return to step 6.3.5; otherwise, execute step 6.3.
12. Step 6.3.12: Use equation (26) to calculate the agent in each region at the 1st rank. Convergence metrics during training rounds : (26) Step 6.3.13, if < Then +1 is assigned to Then, return to step 6.3.4; otherwise, proceed to step 6.3.
14. Step 6.3.14: Set the convergence metric to... ,like If the training is successful, it means that all agents have completed training, and the trained current policy network is used as the robust voltage control reinforcement learning model, and step 7 is executed; otherwise, return to step 6.3.2, reset the hyperparameters, and train again.
3. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing any of the robust control methods of claims 1-2, the processor being configured to execute the program stored in the memory.
4. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the robust control method according to any one of claims 1-2.
Citation Information
Patent Citations
A robust active and reactive power coordination optimization method for active distribution network based on time series scenario analysis
CN109274134A
Active power distribution network safety scheduling method and device based on reinforcement learning
CN116937586A