Adaptive Grid Optimization Method for Enhancing Large-Scale Hydropower Transmission Capacity

By constructing a multi-agent deep reinforcement learning structure, a two-layer transmission network optimization objective function is built, which solves the technical complexity of large-scale hydropower transmission capacity and renewable energy integration, achieves higher accuracy and stability, and optimizes the transmission network expansion planning scheme.

CN119231486BActive Publication Date: 2025-10-31SANXIA JINSHAJIANG YUNCHUAN HYDROPOWER DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411208898.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-10-31
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

Existing grid structure designs struggle to achieve accuracy and stability when dealing with large-scale hydropower transmission capacity. Furthermore, integrating renewable energy presents technical complexity and high costs. Additionally, deep reinforcement learning methods struggle to identify the performance improvement of each indicator by line construction when multiple objective function composites are involved.

Method used

A multi-agent deep reinforcement learning structure is adopted to construct a two-layer transmission network optimization objective function. Independent agents are configured for the optimization objectives of the upper and lower layers, and scheme search is performed through inter-layer constraints. The ε-greedy strategy and convolutional neural network are used for training to optimize the transmission network expansion planning scheme.

Benefits of technology

The adaptive grid optimization method improves the accuracy and stability of large-scale hydropower transmission capacity, enhances the renewable energy absorption level and system reliability of the transmission network, and reduces construction costs and system failure probability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119231486B_ABST
    Figure CN119231486B_ABST
Patent Text Reader

Abstract

This invention provides an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission, comprising: constructing a two-layer transmission network optimization objective function considering the transmission capacity of hydropower stations; wherein the upper-layer optimization objective function is used to evaluate the economic efficiency of the transmission network expansion planning scheme, and the lower-layer optimization objective function is used to evaluate the renewable energy absorption level and system reliability of the transmission network; employing a multi-agent reinforcement learning method to configure an independent agent for each layer of optimization objective to search for schemes and obtain schemes that satisfy the planning model through inter-layer constraints; wherein corresponding reward functions are set according to the structure of the upper and lower layer optimization objective functions. This invention constructs two independent agents, each using the upper and lower layer model optimization objective functions as rewards to search for schemes and independently training its own neural network based on its optimization process data, ultimately obtaining a transmission network expansion planning scheme that satisfies the two-layer model, thereby improving the capacity of large-scale hydropower transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system planning technology, and more specifically, to an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission. Background Technology

[0002] Existing power grid structures typically employ alternating current (AC) and direct current (DC) transmission. AC transmission is widely used due to its mature technology and relatively low construction costs, making it suitable for short- to medium-distance power transmission. However, for ultra-long-distance power transmission, DC transmission demonstrates its superiority, particularly in reducing energy losses during transmission and improving the stability of cross-regional power transmission. DC transmission systems also facilitate connection and management with regional power grids, providing better control capabilities to cope with load changes.

[0003] Furthermore, modern grid designs also attempt to integrate smart grid technologies, including advanced monitoring systems and automation equipment, all aimed at improving the grid's operational efficiency and reliability. By using smart grid technologies, the grid status can be monitored in real time, and power flow can be automatically adjusted to adapt to different power supply and demand conditions, thereby optimizing the performance of the entire power system. However, despite the significant advantages these technologies offer, their implementation and operation still face challenges of technical complexity and high costs. These systems require a high degree of technical expertise in design, construction, and maintenance, which limits their rapid promotion and application. In addition, integration with emerging renewable energy technologies such as wind and solar power presents new technical requirements and challenges, necessitating further technological innovation and system upgrades.

[0004] Deep reinforcement learning (DRL) offers a novel approach to solving extended power system planning models. Traditional DRL-based transmission network extension planning methods require composing multi-objective performance indicators into a single reward function to aid algorithm training. This makes it difficult for the agent to identify the performance improvements of line construction on each indicator, thus hindering the acquisition of the optimal transmission network extension planning scheme. Therefore, researching a DRL structure adapted to solving multi-objective models for transmission network extension planning, in order to improve the performance of network planning structures, is a crucial problem that needs to be solved to promote the application of this method in engineering.

[0005] For adaptive optimization methods of grid structures to enhance large-scale hydropower transmission capacity, higher accuracy, stability, and generalization ability are always required. In terms of accuracy, previous grid construction models have generally achieved high accuracy; however, the complex structure and numerous parameters of these methods may lead to performance fluctuations when faced with different datasets or application scenarios, making the stability of the method particularly important. Summary of the Invention

[0006] The present invention aims to solve at least one of the technical problems existing in the prior art.

[0007] Therefore, this invention provides an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission.

[0008] This invention provides an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission, comprising:

[0009] Considering the power transmission capacity of hydropower stations, a two-layer transmission network optimization objective function is constructed. The upper-layer optimization objective function is used to evaluate the economic efficiency of the transmission network expansion planning scheme, while the lower-layer optimization objective function is used to evaluate the renewable energy consumption level and system reliability of the transmission network.

[0010] A multi-agent reinforcement learning method is adopted to configure an independent agent for each layer of optimization objective to search for solutions and obtain solutions that satisfy the planning model through inter-layer constraints; wherein, corresponding reward functions are set according to the optimization objective function structure of the upper and lower layers respectively.

[0011] The adaptive grid optimization method for improving the capacity of large-scale hydropower transmission according to the above-described technical solution of the present invention may further have the following additional technical features:

[0012] In the above technical solution, the objective function for optimizing the double-layer transmission network structure is:

[0013]

[0014] in, This represents the objective function for optimizing a two-layer transmission network. This represents the upper-level optimization objective function. This represents the lower-level optimization objective function. This indicates the cost of building a power transmission network. Indicates network loss cost, Indicates maintenance costs, This indicates the reduction in hydropower generation. Indicates the load shedding amount. This indicates the electrical dielectric constant of the lines adjacent to the hydroelectric power station.

[0015] In the above technical solution, the constraints of the upper-level optimization objective function include power flow constraints and equipment operation constraints; the power flow constraints include balance constraints constructed based on AC power flow; the equipment operation constraints include voltage amplitude and phase angle constraints, generator output constraints, line transmission capacity constraints, and system load shedding constraints.

[0016] In the above technical solution, the constraints of the upper-level optimization objective function are constructed based on N-1 constraints. Therefore, the constraints of the upper-level optimization objective function under N-1 constraints include:

[0017]

[0018]

[0019]

[0020]

[0021]

[0022]

[0023]

[0024]

[0025] in, This represents the rated active power of the generator set at node j; This represents the active power of wind power discarded by node j; This represents the voltage at node j; This represents the susceptance between node j and node k; This represents the conductivity between node j and node k; This represents the phase angle between node j and node k; To absorb active power for the load at node j; This represents the active power of the load removed from node j; This represents the amount of wind curtailment calculated based on the lower-level model; Indicates the minimum power output ratio; This represents the active power output of the wind farm at node j;

[0026] This represents the reactive power output of the generator at node j; This represents the reactive power input at node j; This represents the reactive power of wind power discarded by node j; This represents the reactive power absorbed by the load at node j; This represents the reactive power of the reactive power compensation device at node j;

[0027] This represents the minimum voltage at node j; This represents the maximum voltage at node j;

[0028] This represents the minimum value of the phase angle at node j; This represents the maximum value of the phase angle at node j;

[0029] This represents the minimum value of the generator's reactive power output; This indicates the maximum value of the generator's reactive power output;

[0030] Indicates the line l The trend; express l ;

[0031] Indicates the minimum load rate; This represents the amount of load to be abandoned calculated based on the lower-level optimization objective function;

[0032] The superscript N-1 indicates the corresponding operating state quantity of the system under the N-1 constraint condition.

[0033] In the above technical solution, a reward function is set according to the comprehensive economic indicators of the power transmission network to optimize the objective function structure of the upper layer. :

[0034]

[0035] in, This represents the baseline of the reward function in the upper-level optimization objective function structure; This indicates the overall economic cost of the power transmission network; This represents the environmental state at time t. , A collection of environmental states;

[0036] The reward function is set according to the renewable energy absorption level and transmission network reliability of the lower-level optimization objective function structure. :

[0037]

[0038] in, This represents the baseline of the reward function for the lower-level optimization objective function structure.

[0039] In the above technical solution, the method of employing multi-agent reinforcement learning to configure an independent agent for each layer of optimization objective to search for solutions and obtain solutions that satisfy the planning model through inter-layer constraints includes:

[0040] The process of searching for solutions using the optimization objective function of each layer as a reward is defined as a layer of intelligent agent, that is, the upper layer optimization objective function corresponds to the upper layer intelligent agent, and the lower layer optimization objective function corresponds to the lower layer intelligent agent;

[0041] The rule stipulates that during the solution search process, the upper-level intelligent agent needs to store three actions with the highest potential value each time and form a Pareto set of candidates;

[0042] The Pareto candidate set is used as the optimization range of the lower-level agent, and the lower-level optimization objective function is used to judge the value of the action. The lower-level agent determines the selected action within the Pareto candidate set.

[0043] When each selected action is executed, Qeval, upper and Qeval, lower are updated, and the parameters are copied to Qtarget, upper and Qtarget, lower at fixed intervals.

[0044] Here, Qeval and Qtarget are value functions with the same initial parameters. Qeval,upper represents the function used by the upper-level agent to select the optimal action, Qeval,lower represents the function used by the lower-level agent to select the optimal action, Qtarget,upper represents the function used by the upper-level agent to predict the value of the selected optimal action, and Qtarget,lower represents the function used by the lower-level agent to predict the value of the selected optimal action.

[0045] In the above technical solution, the upper-layer intelligent agent uses an ε-greedy strategy to search for three routes with the highest potential construction value to form a Pareto solution set. Pareto solution set The rules for its composition include:

[0046]

[0047] Where Prob represents the random probability of the ε-greedy policy. ε Indicates the exploration rate. This represents the process parameters in the Qeval,upper value function. Indicates the best action. This represents the environmental state at time t. , A collection of environmental states; This indicates the action to be selected through random selection. This represents the selected action determined in the Pareto solution set.

[0048] In the above technical solution, the lower-level agent uses an ε-greedy strategy to search the Pareto candidate set to determine the selection action A. t,lower Select action A t,lower The rules for its composition include:

[0049]

[0050]

[0051] in, This represents the process parameters in the Qeval,lower value function.

[0052] In the above technical solution, a multi-agent reinforcement learning method is trained based on a convolutional neural network.

[0053] In the above technical solution, the training of the multi-agent reinforcement learning method based on a deep convolutional network includes:

[0054] Initialize the transmission network expansion planning environment and initialize training parameters

[0055] Based on the current environmental state and agent parameters, multi-agent reinforcement learning is carried out. The upper-layer agent searches to form a Pareto candidate set, and the lower-layer agent determines the selection action within the Pareto candidate set.

[0056] Execute actions and update the state of the reinforcement learning environment, calculate the reward functions of the upper and lower layer agents, form trajectory information and save it;

[0057] Determine if the parameter update interval has been reached. If it is, update the parameters of the convolutional neural network; otherwise, proceed to the next step.

[0058] Determine whether the updated value of the network structure based on the selected action is greater than the value of the network structure before the update. If it is greater, proceed to the next step; otherwise, search again.

[0059] Determine if the maximum number of training sessions has been reached. If it has, proceed to the next step. If it has not, initialize the environment state and start the search again.

[0060] Output the final power grid expansion plan.

[0061] In summary, due to the adoption of the above-mentioned technical features, the beneficial effects of the present invention are:

[0062] This invention proposes an adaptive grid optimization method based on a dual-agent deep reinforcement learning structure to improve the capacity for large-scale hydropower transmission. The invention constructs two independent DDQN agents. Each agent uses the optimization objective function of the upper and lower layer models as reward to search for solutions and independently trains its own neural network based on the optimization process data. Finally, by applying upper and lower layer constraints, a transmission network expansion planning scheme that satisfies the dual-layer model is obtained, thereby improving the capacity for large-scale hydropower transmission.

[0063] Additional aspects and advantages of the invention will become apparent in the following description or may be learned by practice of the invention. Attached Figure Description

[0064] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0065] Figure 1 This is a flowchart of an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission, according to an embodiment of the present invention.

[0066] Figure 2 This is a schematic diagram of the comprehensive economic cost of the PX power grid optimization and transformation network structure obtained using the method of this invention;

[0067] Figure 3 This is a schematic diagram of the comprehensive economic cost of the PX power grid optimization and transformation network structure obtained using the DQN method.

[0068] Figure 4 This is a schematic diagram comparing the line power flow distribution of the PX power grid optimization and transformation network structure obtained using the method of this invention with that of the original network structure.

[0069] Figure 5 This is a schematic diagram comparing the line power flow distribution of the PX power grid optimization and transformation network structure obtained using the DQN method and the PX power grid optimization and transformation network structure obtained using the method of this invention.

[0070] Figure 6 This is a schematic diagram of the action feedback signal of the upper-layer intelligent agent in one embodiment of the present invention;

[0071] Figure 7 This is a schematic diagram of the action feedback signals of the lower-level intelligent agent regarding the amount of wind curtailment and the total amount of load curtailment in one embodiment of the present invention;

[0072] Figure 8 This is a schematic diagram of the action feedback signal of the lower-level intelligent agent regarding the K-order electrical betweenness entropy mean in one embodiment of the present invention;

[0073] Figure 9 This is a schematic diagram comparing the prediction and changes in the value of line construction during the training process of the method of this invention and the DQN method. Detailed Implementation

[0074] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0075] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0076] The following reference Figures 1 to 9 This describes an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission, provided by some embodiments of the present invention.

[0077] Some embodiments of this application provide an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission.

[0078] The first embodiment of this invention proposes an adaptive grid optimization method for improving the capacity of large-scale hydropower transmission, which mainly includes two parts: constructing an optimization objective function and applying a multi-agent reinforcement learning method.

[0079] In constructing the optimization objective function, a two-layer transmission network optimization objective function is constructed considering the transmission capacity of hydropower stations. The upper-layer optimization objective function is used to evaluate the economic efficiency of the transmission network expansion planning scheme, while the lower-layer optimization objective function is used to evaluate the renewable energy consumption level and system reliability of the transmission network.

[0080] Specifically, the objective function for optimizing the dual-layer transmission network structure is:

[0081]

[0082] in, This represents the objective function for optimizing a two-layer transmission network. This represents the upper-level optimization objective function. This represents the lower-level optimization objective function. This indicates the cost of building a power transmission network. Indicates network loss cost, Indicates maintenance costs, This indicates the reduction in hydropower generation. Indicates the load shedding amount. This indicates the electrical dielectric constant of the lines adjacent to the hydroelectric power station.

[0083] A two-layer optimization model was established with reference to different objectives in the objective function of the two-layer transmission network structure optimization.

[0084] In some embodiments, the constraints of the upper-level optimization objective function include power flow constraints and equipment operation constraints; the power flow constraints include balance constraints constructed based on AC power flow; the equipment operation constraints include voltage amplitude and phase angle constraints, generator output constraints, line transmission capacity constraints, and system load shedding constraints.

[0085] In this disclosure, the power flow constraint is a balance constraint constructed based on the alternating current power flow, and its constraint is as follows:

[0086]

[0087]

[0088] in, This represents the rated active power of the generator set at node j; This represents the active power of wind power discarded by node j; This represents the voltage at node j; This represents the susceptance between node j and node k; This represents the conductivity between node j and node k; This represents the phase angle between node j and node k; To absorb active power for the load at node j; This represents the active power of the load removed from node j.

[0089] This represents the reactive power output of the generator at node j; This represents the reactive power input at node j; This represents the reactive power of wind power discarded by node j; This represents the reactive power absorbed by the load at node j; This represents the reactive power of the reactive power compensation device at node j.

[0090] The voltage amplitude and phase angle constraints are as follows:

[0091]

[0092]

[0093] in, This represents the minimum voltage at node j; This represents the maximum voltage at node j; This represents the minimum value of the phase angle at node j; This represents the maximum value of the phase angle at node j.

[0094] Because hydropower plants have strong reactive power regulation capabilities, no special constraints are imposed in this paper. The output constraints for hydropower plants and general generators are as follows:

[0095]

[0096]

[0097]

[0098] in, This represents the minimum value of the generator's active power output; This indicates the maximum value of the generator's active power output; This represents the minimum value of the generator's reactive power output; This indicates the maximum value of the generator's reactive power output; This indicates the active power output of the wind farm; This represents the minimum value of the active power output of a wind farm; This represents the maximum value of the active power output of the wind farm.

[0099] The line transmission capacity constraint is:

[0100]

[0101] in, Indicates the line l The trend; express l .

[0102] The system load shedding constraint is:

[0103]

[0104] in, Indicates the minimum load rate; This represents the amount of load to be abandoned, calculated based on the lower-level optimization objective function.

[0105] In power transmission network planning, the N-1 verification method is the most basic and effective method to ensure the reliability of the power transmission network. Its verification approach involves checking whether the power transmission network can continue to meet stable operating conditions under the new state after any fault occurs in any component of the line. If it does, the N-1 verification passes; otherwise, it fails. In a specific embodiment, the constraints of the upper-level optimization objective function are constructed based on the N-1 constraints. Therefore, the constraints of the upper-level optimization objective function under the N-1 constraints include:

[0106]

[0107]

[0108]

[0109]

[0110]

[0111]

[0112]

[0113]

[0114] The superscript N-1 indicates the operating state variables of the corresponding parameters of the system under the N-1 constraint conditions.

[0115] This represents the amount of wind curtailment calculated based on the lower-level model; Indicates the minimum power output ratio; This represents the active power output of the wind farm at node j.

[0116] In multi-agent reinforcement learning, an independent agent is configured for each layer of optimization objective to search for solutions and obtain solutions that satisfy the planning model through inter-layer constraints; wherein, corresponding reward functions are set according to the optimization objective function structure of the upper and lower layers respectively.

[0117] Specifically, multi-agent reinforcement learning is an algorithmic structural improvement for optimizing each layer of a two- or multi-layer model. It assigns an independent agent to each layer to search for solutions and obtains solutions that satisfy the planning model through inter-layer constraints. Setting up multiple agents in a two-layer model to separate the search tasks for different objective functions can improve the accuracy of judging the construction value of each line, thereby further improving the performance of power grid expansion planning schemes. Since the two agents have different search objectives, corresponding reward functions need to be set according to the optimization objective function structure of the upper and lower layers.

[0118] In some embodiments, a reward function is set based on the comprehensive economic indicators of the power transmission network to define the upper-level optimization objective function structure. :

[0119]

[0120] in, This represents the baseline of the reward function in the upper-level optimization objective function structure; This indicates the overall economic cost of the power transmission network; This represents the environmental state at time t. , A collection of environmental states;

[0121] The reward function is set according to the renewable energy absorption level and transmission network reliability of the lower-level optimization objective function structure. :

[0122]

[0123] in, This represents the baseline of the reward function for the lower-level optimization objective function structure.

[0124] In some embodiments, the method of using multi-agent reinforcement learning to configure an independent agent for each layer of optimization objective to search for solutions and obtain solutions that satisfy the planning model through inter-layer constraints includes the following steps S21-S24.

[0125] S21. The process of searching for a solution using the optimization objective function of each layer as a reward is defined as a layer of intelligent agent, that is, the upper layer optimization objective function corresponds to the upper layer intelligent agent, and the lower layer optimization objective function corresponds to the lower layer intelligent agent.

[0126] S22 stipulates that during the solution search process, the upper-level intelligent agent needs to store three actions with the highest potential value each time and form a Pareto candidate set. It can be understood that "action" refers to an operation performed in the power grid expansion planning scheme, such as adding a new transmission line or reconstructing a transmission line.

[0127] In one specific embodiment, the upper-layer agent uses an ε-greedy strategy to search for three routes with the highest potential construction value to form a Pareto solution set. Specifically, the selection of each path follows an ε-greedy strategy, which introduces an ε threshold to the original greedy strategy. This strategy allows for a probability-wise decision not to choose an action along the path of maximum reward, but rather to select actions randomly. (Pareto solution set) The rules for its composition include:

[0128]

[0129] Where Prob represents the random probability of the ε-greedy strategy, and ε represents the exploration rate. This represents the process parameters in the Qeval,upper value function. Indicates the best action. This represents the environmental state at time t. , A collection of environmental states; This indicates the action to be selected through random selection. This represents the selected action determined in the Pareto solution set.

[0130] S23. The Pareto candidate set is used as the optimization range of the lower-level agent, and the lower-level optimization objective function is used to judge the value of the action. The lower-level agent determines the selected action within the Pareto candidate set.

[0131] Specifically, similar to the upper layer, the lower-layer agent still follows an ε-greedy strategy for searching and uses the lower-layer optimization objective function to judge the value of actions to determine the selected action. The lower-layer agent uses the Pareto candidate set as the optimization range. Action A is selected. t,lower The rules for its composition include:

[0132]

[0133]

[0134] in, This represents the process parameters in the Qeval,lower value function.

[0135] S24. When each selected action is executed, update Qeval,upper and Qeval,lower, and copy the parameters to Qtarget,upper and Qtarget,lower at fixed intervals.

[0136] Here, Qeval and Qtarget are value functions with the same initial parameters. Qeval,upper represents the function used by the upper-level agent to select the optimal action, Qeval,lower represents the function used by the lower-level agent to select the optimal action, Qtarget,upper represents the function used by the upper-level agent to predict the value of the selected optimal action, and Qtarget,lower represents the function used by the lower-level agent to predict the value of the selected optimal action.

[0137] In some embodiments, a multi-agent reinforcement learning method is trained based on a convolutional neural network.

[0138] In one specific embodiment, such as Figure 1 As shown, the training of the multi-agent reinforcement learning method based on deep convolutional networks includes the following steps S101-S107.

[0139] S101. Initialize the transmission network expansion planning environment and initialize the training parameters. Initializing the transmission network expansion planning environment includes setting the initial state corresponding to the initial network structure. Set as current environment state The reward corresponding to the initial network structure is R; initializing training parameters includes setting the maximum number of training sessions. The parameter replication cycle is .

[0140] S102, Based on the current environment status Multi-agent reinforcement learning is performed using the agent parameters. This involves generating two sets of DDQN agents with identical parameters: an upper-layer agent and a lower-layer agent. The upper-layer agent searches to form a Pareto candidate set, and the lower-layer agent determines its selection action from within this set. .

[0141] The upper-layer agent trains a neural network based on a first dataset, which includes data on the state of a given environment. Make a definite choice action The next environmental state formed afterward The corresponding upper-level reward function result Parameters generated during the process , , and The lower-level agent trains the neural network based on a second dataset, which includes data obtained in a given environment. Make a definite choice action The next environmental state formed afterward The corresponding lower-level reward function result Parameters generated during the process , , and .

[0142] S103, Execution Action And update the reinforcement learning environment state, that is, update the grid structure according to the selected action determined in step S102. Calculate the reward functions of the upper and lower layer agents respectively. and The trajectory information is generated and saved; the trajectory information records the data required for neural network training in S102; specifically, the trajectory information is saved in RAM memory, and when training is performed in step S102, the data in RAM can be directly read as the training dataset.

[0143] S104. Determine if the parameter update interval has been reached, i.e., determine if t... step |T step =0, parameter replication period This refers to whether the model's parameters need to be copied or updated during training to maintain consistency; if t step |T step If the value is 0, update the parameters of the convolutional neural network; otherwise, proceed to step S105.

[0144] S105. Determine the updated value V(S) of the network structure based on the selected action. t+1 Is the value of the space frame greater than before the update? (V) base If the value is greater than the specified value, proceed to the next step; otherwise, update the environment state. Set as current environment state Then return to step S102 to perform the search again.

[0145] S106. Determine if the maximum number of training sessions has been reached, i.e., determine the current number of training sessions. Is it greater than the maximum number of training sessions? If the target is reached, proceed to step S107; otherwise, initialize the environment state and start the search again.

[0146] S107, Output the final power grid expansion plan.

[0147] In one specific embodiment, the convolutional neural network may employ a deep residual network to improve the agent performance of the multi-agent DDQN algorithm.

[0148] To better demonstrate the performance of the adaptive grid optimization method proposed in this disclosure, a specific embodiment is illustrated using a power grid planning scheme in the PX region as an example. It is understood that this power grid planning scheme is not an actual power grid plan, but a virtual power grid architecture provided to assist those skilled in the art in understanding it.

[0149] With the large-scale hydropower development and transmission in the PX region, the PX power grid will add four ±800kV UHVDC lines, increasing the UHVDC transmission capacity from the PX region to approximately 34.2 million kilowatts by 2025. Simultaneously, to ensure the safe and stable operation of these multiple UHVDC lines, the PX AC power grid will construct a new BT-SF double-circuit 500kV line and the YCII switching station in 2023, and gradually implement the PX optimization and upgrading project, ultimately achieving the goal of pairwise grouping of the four UHVDC lines (YZ, YC, BZ, and BS). DQN and PSO are used as comparison algorithms, and the parameter comparison of the multi-agent DDQN and DQN is shown in Table 1. Three power grid expansion planning schemes are shown in Tables 2-4. DQN and PSO are both common deep learning methods; DQN is a deep Q-network, and PSO is a particle swarm optimization algorithm. In this embodiment, DDQN represents the multi-agent reinforcement learning method proposed in this disclosure.

[0150] Table 1. Parameters of Multi-Agent DDQN and DQN Algorithms

[0151]

[0152] This disclosure further extends the application scenario to more complex PX power grid systems to further evaluate the performance of the proposed method. We set the loads at nodes 3, 4, 8, 16, 20, 24, 27, and 29 of the power grid system as variable loads, and set lines 30, 32, 33, 34, and 38 as wind farm nodes, increasing the load capacity by 1.10 times, the conventional generator capacity by 1.15 times, and the wind farm capacity by 1.25 times. The planning results of the multi-agent DDQN, DQN, and PSO methods are shown in Tables 2-4.

[0153] Table 2. PX power grid structure and multi-agent DDQN transmission network expansion planning scheme

[0154]

[0155] Table 3. PX power grid structure and multi-agent DDQN transmission network expansion planning scheme

[0156]

[0157] Table 4. PX power grid structure and PSO transmission network expansion planning scheme

[0158]

[0159] Tables 2-4 show that all three methods optimized the stability and economy of the transmission network structure through line construction. The deep reinforcement learning-based method incorporates the order of line construction into its payoff structure, thus its transmission network expansion planning includes this order. However, the PSO method only optimizes the overall transmission network structure. The combined costs of the three methods are 23.10 million, 23.47 million, and 21.64 million, respectively. Among the three methods, the PSO method achieves the lowest construction cost. The DQN method achieves the highest construction cost, while the multi-agent DDQN method's cost falls between the other two. Furthermore, all three methods influence the system structure and power flow distribution through new line construction, resulting in a certain degree of reduction in system network losses.

[0160] The K-order electrical betweenness entropy mean can be used to calculate the uniformity of power flow distribution along lines adjacent to a wind farm, thereby assessing system reliability. The multi-agent DDQN scheme improves the K-order electrical betweenness entropy mean from 110.25MW to 191.50MW, representing the highest optimization magnitude compared to the other three schemes, reflecting a significant reduction in the probability of cascading failures. Furthermore, new power systems need to minimize wind curtailment and load shedding under high uncertainty to improve renewable energy absorption capacity and power supply reliability. Transmission network expansion planning can enhance power support capabilities by constructing new lines and building a more compact transmission network structure. Regarding the modified IEEE RTS24-BUS system (a benchmark system for electricity market and power system operation research) proposed in this disclosure, its significant characteristic is relatively severe load shedding. Both the multi-agent DDQN scheme and the DQN method minimize transmission network load shedding, but the PSO scheme does not achieve good performance in this direction. Meanwhile, the multi-agent DDQN scheme performs best in reducing wind curtailment. Therefore, a comprehensive analysis of the performance of the three schemes reveals that while the PSO scheme has the lowest cost, it neglects the optimization of the K-order electrical betweenness entropy mean and load shedding. This is because multi-objective transmission network expansion is a non-deterministic polynomial (NP) problem, and heuristic learning-based methods are prone to getting trapped in local optima during the transmission network expansion planning process, causing optimization to focus only on improving certain indicators. Deep reinforcement learning-based methods, due to their ε-greedy strategy, allow their agents to potentially escape the optimal line selection determined by the neural network, thus avoiding the influence of local optima to some extent and coordinating the optimization of various model indicators. Although the multi-agent DDQN scheme only slightly outperforms the DQN scheme in terms of wind curtailment and load shedding, its improvement in the K-order electrical betweenness entropy mean is more significant. This confirms that the dual DDQN agent structure improves the transmission network structure optimization capability at each layer, and its judgment of the construction value of each line is more accurate, thus contributing to obtaining a better transmission network expansion planning scheme in the end.

[0161] Figure 2 and Figure 3 The comparison of the overall economic costs of the multi-agent DDQN and DQN schemes illustrates the changes in the overall cost of the solutions obtained by DDQN and DQN. The construction cost of the first three lines in the DQN scheme increases relatively faster than that of the first three lines in the DDQN scheme, requiring higher initial investment at the beginning of the project. Furthermore, based on market price fluctuations and technological advancements, the actual investment required later in the plan may result in greater savings, thus making the capital investment required for the implementation of the DDQN scheme more economical than that of the DQN scheme.

[0162] Figure 4 and Figure 5 The power flow distribution of the multi-agent DDQN scheme and the DQN scheme is shown. The initial power flow of the grid includes two lines with a capacity of over 400MW and three lines with a capacity of over 300MW, indicating poor initial power flow uniformity. Both the multi-agent DDQN and DQN schemes transfer part of the power flow to underloaded lines, improving the utilization rate of underloaded lines and reducing the probability of cascading failures caused by overloaded lines. Further comparison of the power flow distribution shows that for lines with a power flow greater than 200MW, the multi-agent DDQN scheme controls the number of lines with a capacity greater than 250MW to three, while the DQN scheme has seven lines with a capacity greater than 250MW. For lines with a capacity less than 200MW, the multi-agent DDQN and DQN methods show almost identical performance in optimizing the power flow distribution in the range of 150-200MW. For lines with a capacity less than 100MW, the multi-agent DDQN scheme better increases the power flow of underloaded lines. Therefore, the multi-agent DDQN scheme has higher system stability and higher line utilization.

[0163] Compared to the DQN scheme, the multi-agent DDQN scheme achieves better system power flow distribution and lower wind curtailment and load shedding under uncertain conditions with a lower overall economic cost, and the system state transition during construction is smoother. This verifies that the dual-agent structure of the multi-agent DDQN scheme can achieve hierarchical prediction of line construction value. One group searches for lines with high economic efficiency, while the other group searches for lines that can improve renewable energy consumption capacity and system reliability. This structure improves the accuracy of line construction value prediction and helps to form better transmission network expansion planning schemes.

[0164] Figure 6 , Figure 7 and Figure 8The data shows the metrics of approximately 8000 network structures built by the upper and lower layer agents in the Multi-Agent DDQN during 200 training iterations (in order: comprehensive economic cost, total wind curtailment and wind load, and K-order electrical betweenness entropy equilibrium). Before 1000 steps, the comprehensive economic cost metric is concentrated in the poor-performing region, while the lower-layer metrics are more evenly distributed without significant clustering. This is because the Multi-Agent DDQN currently lacks sufficient data for neural network training, resulting in a significant error between its predicted values ​​and the actual environmental benefits. Between 1000 and 2500 steps, the metrics of the planning scheme gradually improve. The neural network has achieved sufficient training, and the prediction error of the route construction values ​​gradually decreases. After 2500 steps, the metrics are evenly distributed between the best and worst. This is because the Multi-Agent DDQN employs an ε-greedy strategy, which allows agents to randomly select route construction results without necessarily choosing the neural network's predictions. This training mechanism can prevent local optima and obtain a more accurate value function.

[0165] Figure 9 The figure shows the change in the sum of line value predictions during the training processes of Multi-Agent DDQN and DQN. Line construction value represents a nonlinear estimate by which agents use neural networks to estimate the improvement in system indicators after the construction of each line. The figure shows that the value prediction of the Multi-Agent DDQN method is significantly lower than that of the DQN method. This is because the Multi-Agent DDQN method uses two neural network structures to form Qeval and Qtarget, making the selection of the optimal line independent of the value prediction of the optimal line. This structure, to some extent, avoids the influence of unexpected overestimation, improves the accuracy of line value prediction, and thus helps solve the task of power grid expansion planning.

[0166] In this specification, the illustrative expressions of the terms used do not necessarily refer to the same embodiments or examples. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0167] Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention shall be included within the scope of protection of this invention.

Claims

1. An adaptive grid optimization method for improving the capacity of large-scale hydropower transmission, characterized in that, include: Considering the power transmission capacity of hydropower stations, a two-layer transmission network optimization objective function is constructed. The upper-layer optimization objective function is used to evaluate the economic efficiency of the transmission network expansion planning scheme, while the lower-layer optimization objective function is used to evaluate the renewable energy consumption level and system reliability of the transmission network. A multi-agent reinforcement learning method is adopted to configure an independent agent for each layer of optimization objective to search for solutions and obtain solutions that satisfy the planning model through inter-layer constraints; wherein, corresponding reward functions are set according to the optimization objective function structure of the upper and lower layers respectively; The method employs a multi-agent reinforcement learning approach, configuring an independent agent for each layer's optimization objective to search for solutions and obtaining solutions that satisfy the planning model through inter-layer constraints. This includes: The process of searching for solutions using the optimization objective function of each layer as a reward is defined as a layer of intelligent agent, that is, the upper layer optimization objective function corresponds to the upper layer intelligent agent, and the lower layer optimization objective function corresponds to the lower layer intelligent agent; The rule stipulates that during the solution search process, the upper-level intelligent agent needs to store three actions with the highest potential value each time and form a Pareto set of candidates; The Pareto candidate set is used as the optimization range of the lower-level agent, and the lower-level optimization objective function is used to judge the value of the action. The lower-level agent determines the selected action within the Pareto candidate set. When each selected action is executed, Qeval, upper and Qeval, lower are updated, and the parameters are copied to Qtarget, upper and Qtarget, lower at fixed intervals. Here, Qeval and Qtarget are value functions with the same initial parameters. Qeval,upper represents the function used by the upper-level agent to select the optimal action, Qeval,lower represents the function used by the lower-level agent to select the optimal action, Qtarget,upper represents the function used by the upper-level agent to predict the value of the selected optimal action, and Qtarget,lower represents the function used by the lower-level agent to predict the value of the selected optimal action.

2. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 1, characterized in that, The objective function for optimizing the dual-layer transmission network structure is: in, This represents the objective function for optimizing a two-layer transmission network. This represents the upper-level optimization objective function. This represents the lower-level optimization objective function. Indicates the cost of power transmission network construction. Indicates network loss cost, Indicates maintenance costs, This indicates the reduction in hydropower generation. Indicates the load shedding amount. This indicates the electrical dielectric constant of the lines adjacent to the hydroelectric power station.

3. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 2, characterized in that, The constraints of the upper-level optimization objective function include power flow constraints and equipment operation constraints; the power flow constraints include balance constraints constructed based on AC power flow; the equipment operation constraints include voltage amplitude and phase angle constraints, generator output constraints, line transmission capacity constraints, and system load shedding constraints.

4. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 3, characterized in that, The constraints of the upper-level optimization objective function are constructed based on the N-1 constraints. Therefore, the constraints of the upper-level optimization objective function under the N-1 constraints include: in, This represents the rated active power of the generator set at node j; This represents the active power of wind power discarded by node j; This represents the voltage at node j; This represents the susceptance between node j and node k; This represents the conductivity between node j and node k; This represents the phase angle between node j and node k; To absorb active power for the load at node j; This represents the active power of the load removed from node j; This represents the amount of wind curtailment calculated based on the lower-level model; Indicates the minimum power output ratio; This represents the active power output of the wind farm at node j; This represents the reactive power output of the generator at node j; This represents the reactive power input at node j; This represents the reactive power of wind power discarded by node j; This represents the reactive power absorbed by the load at node j; This represents the reactive power of the reactive power compensation device at node j; This represents the minimum voltage at node j; This represents the maximum voltage at node j; This represents the minimum value of the phase angle at node j; This represents the maximum value of the phase angle at node j; This represents the minimum value of the generator's reactive power output; This indicates the maximum value of the generator's reactive power output; Indicates the line l The trend; express l ; Indicates the minimum load rate; This represents the amount of load to be abandoned calculated based on the lower-level optimization objective function; The superscript N-1 indicates the corresponding operating state quantity of the system under the N-1 constraint condition.

5. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 4, characterized in that, The reward function is set based on the comprehensive economic indicators of the power transmission network to optimize the objective function structure of the upper-level reward function. : in, This represents the baseline of the reward function in the upper-level optimization objective function structure; This indicates the overall economic cost of the power transmission network; This represents the environmental state at time t. , A collection of environmental states; The reward function is set according to the renewable energy absorption level and transmission network reliability of the lower-level optimization objective function structure. : in, This represents the baseline of the reward function for the lower-level optimization objective function structure.

6. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 1, characterized in that, The upper-level intelligent agent uses an ε-greedy strategy to search for three routes with the highest potential construction value to form a Pareto solution set. Pareto solution set The rules for its composition include: Where Prob represents the random probability of the ε-greedy policy. ε Indicates the exploration rate. This represents the process parameters in the Qeval,upper value function. Indicates the best action. This represents the environmental state at time t. , A collection of environmental states; This indicates the action to be selected through random selection. This represents the selected action determined in the Pareto solution set.

7. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 6, characterized in that, The lower-level agent uses an ε-greedy strategy to search the Pareto candidate set to determine the selection action A. t,lower Select action A t,lower The rules for its composition include: in, This represents the process parameters in the Qeval,lower value function.

8. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 1, characterized in that, Training a multi-agent reinforcement learning method based on a deep convolutional network.

9. The adaptive grid optimization method for improving large-scale hydropower transmission capacity according to claim 8, characterized in that, The training of the multi-agent reinforcement learning method based on deep convolutional networks includes: Initialize the power grid expansion planning environment and initialize training parameters; Based on the current environmental state and agent parameters, multi-agent reinforcement learning is carried out. The upper-layer agent searches to form a Pareto candidate set, and the lower-layer agent determines the selection action within the Pareto candidate set. Execute actions and update the state of the reinforcement learning environment, calculate the reward functions of the upper and lower layer agents, form trajectory information and save it; Determine if the parameter update interval has been reached. If it is, update the parameters of the convolutional neural network; otherwise, proceed to the next step. Determine whether the updated value of the network structure based on the selected action is greater than the value of the network structure before the update. If it is greater, proceed to the next step; otherwise, search again. Determine if the maximum number of training sessions has been reached. If it has, proceed to the next step. If it has not, initialize the environment state and start the search again. Output the final power grid expansion plan.

Citation Information

Patent Citations

  • Power transmission planning under high proportion clean energy access

    CN107679658A

  • Double-layer optimization reconstruction method and system containing distributed wind power distribution network

    CN115411784A