A multi-resource coordinated reactive power regulation method for distribution network considering weak communication conditions
Through multi-dimensional weighting and improved multi-agent deep reinforcement learning algorithm, combined with the coordinated operation of multiple resources, the problem of insufficient grid voltage stability under weak communication conditions is solved, and rapid and economical grid voltage optimization and resource scheduling are achieved.
Patent Information
- Application Number
- CN202510608172.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing reactive power regulation methods fail to make full use of the coordination of multiple resources under weak communication conditions, resulting in insufficient grid voltage stability. The missing data repair method relies on a large amount of training data or has high calculation costs, making it difficult to maintain the grid voltage stable when load fluctuates.
Through multi-dimensional weighted missing data repair and improved multi-agent deep reinforcement learning algorithm, data repair is used to utilize the similarity between different station areas and time sections, and combined with the coordinated operation of photovoltaic inverters, energy storage inverters, charging pile inverters and SVG, a new multi-resource collaborative reactive power regulation model for power distribution networks is built to optimize voltage and network losses.
It improves the accuracy and computing speed of data repair, improves the grid voltage stability and operational economy, and realizes fast and flexible multi-resource collaborative reactive power adjustment.
Smart Images

Figure CN120127776B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reactive power regulation, and in particular to a distribution network multi-resource coordinated reactive power regulation method taking weak communication conditions into consideration. Background Art
[0002] With the transformation and upgrading of global energy, promoting green transformation and the construction of new distribution networks in the energy field has provided policy guarantees for the widespread access of distributed energy, and has also put forward higher requirements for improving the reactive power regulation capabilities of distribution networks and optimizing power system operations.
[0003] Reactive power regulation in distribution networks using a variety of resources, including photovoltaic inverters, energy storage inverters, charging station inverters, and static var generators (SVGs), is an emerging voltage regulation method. Leveraging the reactive power regulation capabilities of these resources effectively addresses the issue of voltage violations in new distribution networks. Through a global optimization control strategy, the inverters and SVGs absorb and emit reactive power to maintain a stable system voltage.
[0004] However, existing reactive power regulation methods fail to fully tap the potential of distribution network resources. They typically only consider a single distributed resource, such as photovoltaics or solar-powered storage systems, and fail to integrate multiple resources, including photovoltaics, energy storage, charging station inverters, and SVGs, for coordinated reactive power regulation. In fact, considering the diversity and flexibility of devices like charging station inverters and SVGs, and utilizing multiple resources for coordinated reactive power regulation, will provide a more efficient and flexible solution for voltage regulation in new distribution networks.
[0005] The effectiveness of existing reactive power regulation strategies relies on stable communication conditions. Under these conditions, the control strategy can coordinate the operation of each substation to achieve overall reactive power compensation. However, in practical applications, new distribution networks are characterized by a multitude of devices, and substation intelligent converged terminals upload large amounts of data to the distribution network control center. Furthermore, these substations are often geographically dispersed, making communication susceptible to factors such as inclement weather. This can lead to weak communication conditions, where substation data may be lost during transmission. In contrast, commands sent from the distribution network control center to substations require smaller amounts of data, and the deployment and maintenance of communication equipment are more centralized. Dedicated lines and other communication methods offer greater reliability. To address weak communication conditions, most existing research relies on decentralized control strategies or fixed rules. The core concept is to achieve uniform or proportional distribution of reactive power compensation through adaptive adjustments between local substations, thereby ensuring basic grid operation. Since substations can only make decisions based on local information, they struggle to assess the real-time operational status of the overall system. In emergencies such as load fluctuations, insufficient coordination between substations can lead to uneven reactive power compensation, impacting grid voltage stability.
[0006] Current research on missing data repair falls into three main categories: interpolation, model-based repair, and machine learning. Interpolation estimates missing values using existing adjacent data points, but can introduce significant errors when the relationships between missing data are complex. Both model-based repair and machine learning use existing data to train models for missing data prediction. These methods require large amounts of training data, are computationally expensive, and rely on model accuracy. In reality, operating characteristics across different time sections and substations can be similar. Similarity metrics and weighted evaluations can be used to rapidly repair missing data, thereby improving the coordination of reactive power compensation in weak communication environments.
[0007] In summary, the problems of the prior art are as follows:
[0008] 1) Existing reactive power regulation methods rarely consider the impact of weak communication conditions on reactive power regulation decisions in distribution networks. When communication failures lead to missing data at a substation, existing methods make reactive power decisions based solely on local information, failing to grasp the overall system operating status and lacking coordination. The missing data repair method proposed in this paper, based on multi-dimensional weighting, eliminates the need for extensive training datasets and can quickly and accurately repair missing data, significantly ensuring real-time and coordinated decision-making.
[0009] 2) Most existing reactive power regulation optimization models only perform reactive power regulation for a single distributed resource (such as photovoltaic or photovoltaic storage inverters) and focus only on minimizing network losses and voltage deviations. They fail to fully explore and utilize the potential of coordinated reactive power regulation of multiple resources in new distribution networks, have poor flexibility, and do not comprehensively consider the cost required for the distribution network to maintain voltage stability. Summary of the Invention
[0010] In response to the above problems, the purpose of the present invention is to provide a distribution network multi-resource collaborative reactive power regulation method that takes into account weak communication conditions. It uses the similarities between different substations and different time sections to perform multi-dimensional weighted missing data repair, thereby improving the accuracy of data repair and data processing speed; fully mobilizes the four resources in the distribution network, namely photovoltaic inverters, energy storage inverters, charging pile inverters, and SVG, to operate in coordination, effectively manage voltage exceeding upper and lower limits, maintain the grid voltage at a stable level, and significantly improve the voltage quality and operating economy of the distribution network. The technical solution is as follows:
[0011] A method for coordinated reactive power regulation of a distribution network with multiple resources considering weak communication conditions comprises the following steps:
[0012] Step 1: Obtain data from different areas and time sections of the distribution network;
[0013] Step 2: Use a hash algorithm to calculate the hash value of the data at the intelligent fusion terminal in the substation as the sending end and the distribution network control center as the receiving end, and compare them to detect whether the reactive power regulation demand data is missing during the transmission process;
[0014] Step 3: Based on the missing data, the similarities between different distribution network areas and different time sections are used to perform multi-dimensional comprehensive weighted patching of missing data from the overall interconnection dimension, regional interconnection dimension, overall self-connection dimension, and regional self-connection dimension.
[0015] Step 4: With the goal of minimizing the system's network losses, the voltage offsets of all system nodes, and the cost of the distribution network purchasing active and reactive power from upstream networks, constraints are established, including distribution network constraints, rated capacity constraints, and energy storage system charging and discharging constraints. A distribution network multi-resource collaborative reactive power regulation model based on improved multi-agent deep reinforcement learning is constructed.
[0016] Step 5: Use the multi-agent deep deterministic policy gradient algorithm to solve;
[0017] Step 6: Mobilize the four resources of the distribution network, namely photovoltaic inverters, energy storage inverters, charging pile inverters, and SVG, to operate in coordination to achieve reactive power regulation and voltage optimization.
[0018] The beneficial effects of the present invention are:
[0019] 1) The present invention proposes a missing data repair method based on multi-dimensional weighting, which can perform multi-dimensional weighted repair through the similarity between different stations and different time sections. It not only avoids the dependence on large data sets, but also improves the accuracy and calculation speed of data repair, and greatly ensures the real-time and coordination of reactive power regulation, and avoids the decline in reactive power regulation effect of distribution network due to weak communication.
[0020] 2) This paper proposes a new multi-resource collaborative reactive power regulation model for distribution networks based on an improved Multi-Agent Deep Reinforcement Learning (MADRL) algorithm, and solves the proposed method using the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm. This model not only mobilizes the collaborative work of four resources, namely photovoltaic inverters, energy storage inverters, charging pile inverters, and SVG, greatly improving the flexibility and effectiveness of reactive power regulation and voltage optimization, and better maintaining the voltage stability of the distribution network, but also achieves minimal voltage offset, minimal network loss, and lowest cost by optimizing the objective function, thereby enhancing the operational economy of the distribution network. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a flow chart of the distribution network multi-resource coordinated reactive power regulation method considering weak communication conditions of the present invention. DETAILED DESCRIPTION
[0022] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. The present invention proposes a new distribution network multi-resource collaborative reactive regulation method that takes weak communication conditions into consideration. It uses the similarity between different substations and different time sections to perform multi-dimensional weighted missing data repair, thereby improving the accuracy of data repair and data processing speed; constructs a new distribution network multi-resource collaborative reactive regulation model based on the improved MADRL algorithm, and uses the MADDPG algorithm to solve the proposed method, fully mobilizing the photovoltaic inverter, energy storage inverter, charging pile inverter, and SVG in the distribution network to operate collaboratively, effectively managing the voltage exceeding the upper and lower limits, maintaining the grid voltage at a stable level, and significantly improving the voltage quality and operating economy of the distribution network. The present invention can be used in the software development of the distribution network control center, and can still achieve fast, economical, and flexible distribution network multi-resource collaborative reactive regulation under weak communication conditions.
[0023] The algorithm flow chart of the present invention is as follows Figure 1 The specific process is as follows:
[0024] 1. Data missing assessment based on hash algorithm:
[0025] Data integrity is crucial for the reactive power regulation decision-making of the system, but under weak communication conditions, data loss often occurs during data transmission, thus affecting the reactive power regulation effect. Data loss includes single-item loss and multiple-item loss. For example, the loss of a certain type of data in a single substation is considered a single-item loss, while the loss of multiple types of data in a single substation or the loss of a certain type of data in multiple substations are multiple-item losses. These missing data may be related to each other or independent of each other. In order to accurately evaluate the integrity of the data, the present invention adopts a hash algorithm and uses the SHA-256 algorithm to calculate the hash value of the data at the intelligent fusion terminal (sending end) of the substation and the control center (receiving end) of the distribution network, and compares them to detect whether the reactive power regulation demand data is missing during the transmission process. The data loss rate can be calculated by the following formula:
[0026] (1);
[0027] Where, MIS represents the data missing rate, C Indicates the total number of distribution network areas to be evaluated, B Indicates the total number of attributes of reactive power regulation demand data to be evaluated in a single substation (such as voltage, load reactive power, load active power and other key parameters).N Indicates the number of attributes with empty values. For example, in a distribution network with 30 substations, each substation has eight attributes for reactive power regulation demand data. If a substation has six empty attributes, the data missing rate is 2.5%. A non-zero data missing rate indicates weak communication.
[0028] 2. Missing data filling method based on multi-dimensional weighting:
[0029] For a station M with missing data, its horizontal neighborhood refers to the measurement values from different stations at the same time section, its vertical neighborhood refers to the measurement values from the same station at different time sections, and the local data window refers to the data group containing the horizontal and vertical neighborhood measurement values of M within a given time domain.
[0030] 1) Overall interconnectedness:
[0031] In the overall interconnected dimension, missing data is filled using the inverse distance weighted method, which interpolates missing data based on the neighborhood. Using the cosine similarity between variables as a weighting factor, for a station M with missing data, its horizontal neighborhood data is used to fill the missing data. The formula is as follows:
[0032] (2);
[0033] Where, R em It is the data repair value obtained under the overall interconnected dimension. n 1 is the number of complete data of other stations in the horizontal neighborhood of M, r i is the measurement value of the complete data, l is a weight attenuation factor based on similarity (the higher the similarity, the l The greater the weight). c i Is the vector c (x 11 , x 12 ,…,x 1n ) and the vector b(x 21 , x 22 ,...,x 2n ), the cosine similarity under the same time section, the cosine similarity calculation formula is as follows:
[0034] (3);
[0035] Where, n is the number of attributes contained in the area vector; and are the moduli of vector c and vector b respectively.
[0036] 2) Regional interconnection dimension:
[0037] In the dimension of regional interconnection, based on the user collaborative filtering recommendation method, the missing data are repaired by mining the regional correlation between each substation. According to the measurement values in the local data window, the similarity between the remaining substations and the substation M with missing data is calculated, and the missing data are weighted averaged using this as the weight. The formula is as follows:
[0038] (4);
[0039] Where, R pm It is the data patch value obtained under the regional interconnection dimension. n 2 is the number of complete data points in the lateral neighborhood of M in the local data window. It is the similarity between the vector c of the area with complete data and the vector b of the area M with missing data in the local data window at the same time section. The calculation formula is as follows:
[0040] (5);
[0041] Where, N 1 is the total number of complete time sections of the substation data in the local data window.
[0042] 3) Overall self-connection dimension:
[0043] In the overall self-connected dimension, through the idea of simple exponential smoothing, we use the autocorrelation of a certain area M with missing data, and based on the historical data before the missing data occurs, we use the values of other time sections to fill in the missing data. The formula is as follows:
[0044] (6);
[0045] Where, R es It is the data patch value obtained under the overall self-connection dimension. n 3 is the number of complete data in the vertical neighborhood of station M, r u It is in the time section u The measurement value of the complete data is c` It is a weight attenuation factor based on time series (the shorter the time interval, the smaller the value, and the value range is 0-1). f u It is the difference between the complete data of the station area and the time section where the missing value is located.
[0046] 4) Regional self-connection dimension:
[0047] Data patching in the regional self-connected dimension is similar to that in the regional interconnected dimension. The similarities between different time sections are calculated using the measured values in the local data window, and the missing data are interpolated based on this as the weight. The formula is as follows:
[0048] (7);
[0049] Where, R ps It is the data patch value obtained under the regional self-connection dimension. n 4 is the number of complete data points in the vertical neighborhood of M in the local data window. d u It is the similarity between different time sections of the data missing area in the local data window. The calculation formula is as follows:
[0050] (8);
[0051] Where, N 2 is the local data window, in the time section u and time sections q The number of complete parameters under the same conditions; r uv It is v Parameters in time section u The measured value of r qv It is v Parameters in time section q The measured value.
[0052] 5) Multi-dimensional data patching weighted calculation:
[0053] Integrate the calculation results of the above four dimensions and calculate the data repair value according to the following formula:
[0054] (9);
[0055] Where, 、 、 and are the weights assigned to the four dimensions, y By minimizing the square error between the predicted value and the actual value, each station variable is trained to determine the optimal weight.
[0056] 3. A new distribution network multi-resource coordinated reactive power regulation model based on the improved MADRL algorithm:
[0057] In the new distribution network multi-resource coordinated reactive power regulation, the target decision layer and the optimization control layer will act as independent computing intelligent entities and jointly participate in system optimization.
[0058] 1) Objective function:
[0059] (10);
[0060] Where, N d is the number of decision instructions, 、 and is the adjustment coefficient of each optimization objective, is the network loss of the system, is the voltage offset of all nodes in the system, It is the cost of the distribution network purchasing active and reactive power from the upstream network.
[0061] 2) Constraints:
[0062] ①Distribution network constraints:
[0063] (11);
[0064] (12);
[0065] (13);
[0066] Where, N bus Refers to the number of system buses, and They are t The active power and reactive power obtained from the upstream network at all times, and They are t Time bus j Active power and reactive power generated by photovoltaics; and They are t Time bus j Active power and reactive power generated by the upper energy storage system; and They are t Time bus j The active power and reactive power generated by the charging pile, yes t Time bus j Reactive power generated by the upper SVG; and They are t Time bus j Active load and reactive load; and They are tThe active power loss and reactive power loss of the network at all times, U j It is a busbar j The voltage on.
[0067] When multiple resources are used for coordinated power generation, it is necessary to ensure the supply and demand balance of each busbar, as shown below:
[0068] (14);
[0069] (15);
[0070] Where, H Is connected to the bus j A set of busbars, G jp It is a busbar j To busbar p The conductivity, B jp It is a busbar j To busbar p The electrical susceptance, i t,jp It is a busbar j and busbar p The voltage phase difference between and They are t Time bus j and busbar p voltage.
[0071] ② Rated capacity constraints:
[0072] The reactive power generated by SVG must be less than its rated capacity. The active power and reactive power generated by photovoltaics, energy storage, and charging piles must also be less than their rated capacity, as shown below:
[0073] (16);
[0074] (17);
[0075] (18);
[0076] (19);
[0077] Where, It is a busbar j SVG rated capacity, It is a busbar j The PV rated capacity, It is a busbar j The rated capacity of energy storage, It is a busbarj The rated capacity of the charging pile.
[0078] ③ Energy storage system charging and discharging constraints:
[0079] When the energy storage system is charging and discharging, the state of charge of the energy storage battery on bus j The following conditions must be met:
[0080] (20);
[0081] (twenty one);
[0082] Where, yes t Time busbar j The state of charge of the energy storage system is set to the minimum state of charge considering the service life of the energy storage battery. is 0.3, setting the maximum state of charge is 0.9.
[0083] In order to effectively solve the problem of multi-resource coordinated reactive power regulation in distribution networks, the optimization problem can be constructed as a Markov decision process and solved using the improved MADRL algorithm.
[0084] 1) State space:
[0085] State Space S Used to represent all possible states of the system at each moment, for any state s k ∈ S , which can be expressed as follows:
[0086] (twenty two);
[0087] Where, yes k Busbar in the distribution network at all times j The voltage value, 、 、 They are k Time bus j Active power of photovoltaic inverter, energy storage inverter and charging pile inverter; 、 、 、 They are k Time bus j Reactive power of photovoltaic inverters, energy storage inverters, charging pile inverters, and SVG; yes k Time bus j The state of charge of the upper energy storage battery; is the current network loss of the distribution network; and They are the active power demand information and reactive power demand information of the distribution network, and They are k Always monitor the active and reactive power demands of the upstream network.
[0088] 2) Action Space:
[0089] Action Space A It consists of all possible actions that the agent can take. For any action a j,k ∈ A , which can be expressed as follows:
[0090] (twenty three);
[0091] Where, 、 and yes k Time bus j The active power that the photovoltaic inverter, energy storage inverter, and charging pile inverter need to generate; 、 、 and yes k Time bus j The reactive power required to be generated by the photovoltaic inverter, energy storage inverter, charging pile inverter and SVG; and yes k The active power and reactive power purchased from the upstream network at all times.
[0092] 3) State transfer matrix:
[0093] State transition matrix T It describes the probability of the system state changing over time after performing an action, which can be expressed as:
[0094] (twenty four);
[0095] Where, p ( s k+1 | s k , a j,k ) refers to the original state s k Taking action a j,k After transfer to state sk+1 probability.
[0096] 4) Reward function and discount factor:
[0097] In order to minimize the voltage offset, network loss, and active and reactive power purchase costs, the objective function shown in formula (10) can be transformed into a single-step reward function: R i ,Right now:
[0098] (25);
[0099] Where, e i is a penalty factor used to measure whether the relevant indicator exceeds the constraint value. When the system indicator exceeds the constraint value, the reward function will increase the penalty, and vice versa, it will reduce the penalty. This ensures that the agent provides sufficient safety margin for the operation of the new distribution network while meeting the system constraints.
[0100] In reinforcement learning, the discount factor determines the weight of long-term rewards and reflects the importance of future rewards. When the discount factor is close to 1, it emphasizes future rewards, and when it is close to 0, it emphasizes current rewards. Therefore, in the reactive optimization scenario, the discount factor c It can be calculated using the following formula:
[0101] (26);
[0102] Where, It is the attenuation coefficient of reactive power. The smaller the value, the slower the reactive power decays, which means that long-term rewards have a greater weight in decision-making.
[0103] 5) Agent Strategy:
[0104] The agent strategy defines the actions that the agent takes in a given state. In the multi-resource collaborative reactive power regulation problem proposed in this invention, the agent strategy can be expressed as follows:
[0105] (27);
[0106] Where, F ( s k , a j,k ) indicates that the state s k Take action a j,k The expected return after softmax The function can be converted into a probability distribution.
[0107] 4. MADDPG algorithm solution:
[0108] In solving the reactive power optimization problem, the MADDPG algorithm constructs two neural networks: the actor and the critic. The actor network accepts the current state as input and outputs the corresponding action. The critic network accepts the current state and action as input, calculates the time difference error, and outputs the value corresponding to the state-action. A neural network with three fully connected layers is constructed as the actor network. The input layer dimension corresponds to the state matrix, and the output layer dimension corresponds to the number of reactive power regulation resources in the distribution network. In order to introduce nonlinear relationships, the ReLU function is used as the activation function. The critic network also consists of a three-layer fully connected network. The value under different states can be evaluated using the following formula:
[0109] (28);
[0110] Where, i i are the parameters of the value network, s It is the overall state, a It is an overall action; s i It is i The observation value of an agent, F i ( s , a 1, a 2,..., a n ) is the evaluation function, taking action a i and status s As input value, the action value function is the output value, is the action strategy, which is used to select actions in the process of learning the target strategy. yes right i i The partial differential vector of For parameters i i The performance metric of the corresponding strategy, E Indicates taking the expected value.
[0111] Selecting actions based on probability will result in low convergence efficiency. Therefore, the MADDPG algorithm replaces the random strategy with a deterministic strategy. The gradient of the agent can be expressed as follows:
[0112] (29);
[0113] Where, It is i The strategy of an agent, D (s , a , r , s` ) is the experience back to the buffer area, which is used to record the experience of all agents. is the status s Take action a The visible eigenvectors of s` The state after taking action.
[0114] The loss function of the critic network can be expressed as follows:
[0115] (30);
[0116] Where, Is a delay parameter, s k+1 and a k+1 yes k The state and action corresponding to the +1 moment.
[0117] In summary, the present invention proposes a distribution network multi-resource collaborative reactive power regulation method considering weak communication conditions. By utilizing the similarities between different substations and different time sections of the distribution network, the missing data is comprehensively weighted and repaired from the overall interconnection dimension, regional interconnection dimension, overall self-connection dimension, and regional self-connection dimension. It does not require a large number of data sets as support, and can effectively improve the accuracy of data repair and the speed of calculation. With the minimum voltage offset, the minimum network loss, and the minimum active and reactive power purchase cost as the objective function, a new distribution network multi-resource collaborative reactive power regulation model based on the improved MADRL algorithm is constructed, and the MADDPG algorithm is used to solve the proposed method, which fully mobilizes the four resources of photovoltaic inverters, energy storage inverters, charging pile inverters, and SVG in the distribution network, realizes reactive power regulation and voltage optimization, improves the flexibility and effectiveness of decision-making, and enhances the operation economy of the distribution network.
Claims
1. A distribution network multi-resource coordinated reactive power regulation method considering weak communication conditions, characterized in that: The following steps are involved: Step 1: Obtain data from different areas and time sections of the distribution network; Step 2: Use a hash algorithm to calculate the hash value of the data at the intelligent fusion terminal in the substation as the sending end and the distribution network control center as the receiving end, and compare them to detect whether the reactive power regulation demand data is missing during the transmission process; Step 3: Based on the missing data, the similarities between different distribution network areas and different time sections are used to perform multi-dimensional comprehensive weighted patching of missing data from the overall interconnection dimension, regional interconnection dimension, overall self-connection dimension, and regional self-connection dimension. In the overall interconnectedness dimension, missing data is patched using the inverse distance weighted method, which interpolates missing data based on the neighborhood. Using the cosine similarity between variables as a weighting factor, for a station M with missing data, its lateral neighborhood data is used to fill the gap. In the dimension of regional interconnection, based on the user collaborative filtering recommendation method, the missing data are repaired by mining the regional correlation between each substation. According to the measurement values in the local data window, the similarity between the remaining substations and the substation M with missing data is calculated, and the missing data are weighted averaged based on this similarity. In the overall self-connected dimension, through the idea of simple exponential smoothing, the autocorrelation of a certain area M with missing data is used to fill the missing data with the values of other time sections based on the historical data before the missing data occurs. In the regional self-connected dimension, the similarity between different time sections is calculated using the measured values in the local data window, and the missing data are interpolated based on this weight. Step 4: With the goal of minimizing the system's network losses, the voltage offsets of all system nodes, and the cost of the distribution network purchasing active and reactive power from upstream networks, constraints are established, including distribution network constraints, rated capacity constraints, and energy storage system charging and discharging constraints. A distribution network multi-resource collaborative reactive power regulation model based on improved multi-agent deep reinforcement learning is constructed. Step 5: Use the multi-agent deep deterministic policy gradient algorithm to solve; Step 6: Mobilize the four resources of the distribution network, namely photovoltaic inverters, energy storage inverters, charging pile inverters, and SVG, to operate in coordination to achieve reactive power regulation and voltage optimization.
2. A distribution network multi-resource coordinated reactive power regulation method considering weak communication conditions according to claim 1, characterized in that: In step 2, the specific process of detecting whether the reactive power regulation demand data is missing during the transmission process is as follows: The SHA-256 algorithm is used to calculate the hash value of the data in the intelligent fusion terminal of the substation area and the control center of the distribution network, and the data loss rate is calculated: (1); Where, MIS represents the data missing rate, C Indicates the total number of distribution network areas to be evaluated, B Indicates the total number of attributes of reactive power regulation demand data to be evaluated in a single substation. N Indicates the number of attributes whose values are empty; When the data missing rate is non-zero, it is considered to be a weak communication situation and the reactive power regulation demand data is missing.
3. The method for coordinated reactive power regulation of a distribution network with multiple resources considering weak communication conditions according to claim 1, characterized in that: In step 3, the specific steps of the multi-dimensional comprehensive weighted repair include: Step 3.1: Calculate the data patch value under the overall interconnected dimension: The formula is as follows: (2); Where, R em It is the data repair value obtained under the overall interconnected dimension. n 1 is the number of complete data of other stations in the horizontal neighborhood of station M, r i is the measurement value of the complete data, λ is a weight decay factor based on similarity; c i Is the vector c (x 11 , x 12 , ..., x 1n ) and the vector b(x 21 , x 22 , ..., x 2n ), the cosine similarity between the same time section; the cosine similarity calculation formula is as follows: (3); Where, n is the number of attributes contained in the area vector; and are the moduli of vector c and vector b respectively; and are the elements in vector c and vector b respectively; Step 3.2: Calculate the data patch value under the regional interconnection dimension: The formula is as follows: (4); Where, R pm It is the data patch value obtained under the regional interconnection dimension. n 2 is the number of complete data points in the lateral neighborhood of station M in the local data window; In the local data window, the vector c (x 11 , x 12 , ...,x 1n ) and the vector b(x 21 , x 22 , ..., x 2n ), the similarity between the same time section is calculated as follows: (5); Where, N 1 is the total number of complete time sections of the station area data in the local data window; Step 3.3: Calculate the data patch value under the overall self-connected dimension: The formula is as follows: (6); Where, R es It is the data patch value obtained under the overall self-connection dimension. n 3 is the number of complete data in the vertical neighborhood of station M, r u It is in the time section u The measurement value of the complete data is γ` is a time-based weight decay factor, f u It is the difference between the time section where the complete data of station area M is located and the missing value; Step 3.4: Calculate the data patch value under the regional self-connection dimension: The formula is as follows: (7); Where, R ps This is the result obtained under the regional self-connection dimension. n 4 is the number of complete data points in the longitudinal neighborhood of station M in the local data window; δ u It is the similarity between different time sections of the data missing area in the local data window. The calculation formula is as follows: (8); Where, N 2 is the local data window, in the time section u and time sections q The number of complete parameters under the same conditions; It is v Parameters in time section u The measured value of r qv It is v Parameters in time section q The measured value of Step 3.5: Multi-dimensional data patching weighted calculation: Integrate the calculation results of the above four dimensions and calculate the final data repair value according to the following formula : (9); Where, 、 、 and are the weights assigned to the four dimensions, y is the residual; By minimizing the square error between the predicted value and the actual value, each station variable is trained to determine the optimal weight.
4. A distribution network multi-resource coordinated reactive power regulation method considering weak communication conditions according to claim 1, characterized in that: The specific process of step 4 is as follows: Step 4.1: Determine the objective function F : (10); Where, N d is the number of decision instructions, is the network loss of the system, is the voltage offset of all nodes in the system, It is the cost of the distribution network purchasing active and reactive power from the upstream network; 、 and is the adjustment coefficient of each optimization objective; Step 4.2: Determine the constraints: 1) Grid constraints: (11); (12); (13); Where, N bus Refers to the number of system buses, and They are t The active power and reactive power obtained from the upstream network at all times, and They are t Time bus j Active power and reactive power generated by photovoltaics; and They are t Time bus j Active power and reactive power generated by the upper energy storage system; and They are t Time bus j The active power and reactive power generated by the charging pile, yes t Time bus j Reactive power generated by the upper SVG; and They are t Time bus j Active load and reactive load; and They are t The active power loss and reactive power loss of the network at all times, U j It is a busbar j The voltage on In order to ensure the supply and demand balance of each busbar when multiple resources are coordinated for power generation, the constraints are expressed as follows: (14); (15); Where, H Is connected to the bus j A set of busbars, G jp It is a busbar j To busbar p The conductivity, B jp It is a busbar j To busbar p of susceptance; θ t,jp is busbar j and busbar p The voltage phase difference between and They are t Time bus j and busbar p voltage; 2) Rated capacity constraints: The reactive power generated by SVG is less than its rated capacity. The active power and reactive power generated by photovoltaic, energy storage, and charging piles are also less than their rated capacity, as shown below: (16); (17); (18); (19); Where, It is a busbar j SVG rated capacity, It is a busbar j The PV rated capacity, It is a busbar j The rated capacity of energy storage, It is a busbar j The rated capacity of the charging pile; 3) Energy storage system charging and discharging constraints: When the energy storage system is charging and discharging, the bus j State of charge of the upper energy storage battery The following conditions are met: (20); (21); Where, yes t Time busbar j The state of charge of the energy storage system, taking into account the service life of the energy storage battery; is the minimum state of charge, is the maximum state of charge; Step 4.3: The optimization problem of solving the coordinated reactive power regulation of multiple resources in the distribution network is constructed as a Markov decision process and solved using an improved multi-agent deep reinforcement learning algorithm.
5. A distribution network multi-resource coordinated reactive power regulation method considering weak communication conditions according to claim 4, characterized in that: The solution process of step 4.3 is as follows: Step 4.3.1: Determine the state space: Using state space S Represents all possible states of the system at each moment. For any state s k ∈ S , expressed as follows: (22); Where, yes k The voltage value of bus j in the distribution network at time 、 、 They are k Time bus j Active power of photovoltaic inverter, energy storage inverter and charging pile inverter; 、 、 、 They are k Time bus j Reactive power of photovoltaic inverters, energy storage inverters, charging pile inverters, and SVG; yes k Time bus j The state of charge of the upper energy storage battery; is the current network loss of the distribution network; and They are the active power demand information and reactive power demand information of the distribution network, and They are k Always monitor the active and reactive power demands of the upstream network; Step 4.3.2: Determine the action space: Using action space A Represents all possible actions taken by the agent. For any action a j,k ∈ A , expressed as follows: (23); Where, 、 and yes k Time bus j The active power that the photovoltaic inverter, energy storage inverter, and charging pile inverter need to generate; 、 、 and yes k Time bus j The reactive power required to be generated by the photovoltaic inverter, energy storage inverter, charging pile inverter and SVG; and They are k Active power and reactive power purchased from the upstream network at all times; Step 4.3.3: Determine the state transition matrix: Using the state transition matrix T Describes the probability of the system state changing over time after performing an action, expressed as: (24); Where, p ( s k+1 |s k ,a j,k ) refers to the original state s k Taking action a j,k After transfer to state s k+1 probability; Step 4.3.4: Determine the reward function and discount factor: In order to minimize the voltage offset, network loss, and active and reactive power purchase costs, the objective function shown in formula (10) is transformed into a single-step reward function: R i ,Right now: (25); Where, ε i It is a penalty factor used to measure whether the relevant indicator exceeds the constraint value; In the reactive power optimization scenario, the discount factor γ Calculated using the following formula: (26); Where, is the attenuation coefficient of reactive power; Step 4.3.5: Determine the agent strategy: The agent strategy defines the actions taken by the agent in a given state and is expressed as follows: (27); Where, F ( s k , a j,k ) indicates that the state s k Take action a j,k The expected return after softmax The function can be converted into a probability distribution.
6. A distribution network multi-resource coordinated reactive power regulation method considering weak communication conditions according to claim 5, characterized in that: The specific process of step 5 is as follows: In solving reactive optimization problems, a multi-agent deep deterministic policy gradient algorithm constructs two neural networks: the actor and the critic. The actor neural network accepts the current state as input and outputs the corresponding action. The critic neural network accepts the current state and action as input, calculates the time difference error, and outputs the value corresponding to the state-action. A neural network with three fully connected layers is constructed as the actor neural network. The input layer dimension corresponds to the state matrix, and the output layer dimension corresponds to the number of reactive regulation resources in the distribution network. To introduce nonlinear relationships, the ReLU function is used as the activation function. The critic neural network also includes a three-layer fully connected network and uses the following formula to evaluate the value under different states: (28); Where, θ i are the parameters of the value network, s It is the overall state, a It is an overall action; s i It is i The observation value of an agent, F i ( s , a 1, a 2,..., a n ) is the evaluation function, taking action a i and overall status s As input value, the action-value function is the output value; is the action strategy, which is used to select actions in the process of learning the target strategy; yes right θ i The partial differential vector of For parameters θ i Performance measures of corresponding strategies; E Indicates taking the expected value; Replacing the random strategy with a deterministic strategy, the agent's gradient is expressed as follows: (29); Where, It is i The strategy of each agent; D ( s , a , r , s `) is the experience back to the buffer area, which is used to record the experience of all agents. r For rewards; is the status s Take action a The visible eigenvector of ; s` The state after taking action; The loss function of the critic neural network is expressed as follows: (30); Where, Is a delay parameter, s k+1 and a k+1 yes k The state and action corresponding to the +1 moment.
Citation Information
Patent Citations
Load storage real-time coordination control method and system suitable for regional-level source network
CN119093511A