An intelligent 500kV power grid planning method based on source and load balance of extra-high voltage interconnected power grid

CN122512366APending Publication Date: 2026-08-04STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID ZHEJIANG ELECTRIC POWER CO LTD
Filing Date
2026-04-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

数学优化方法(如混合整数线性规划MILP)虽然在理论上能求得最优解,但在处理大规模实际电网时计算复杂度过高,难以在有限时间内收敛;启发式算法(如遗传算法GA、粒子群算法PSO)虽然计算速度较快,但极其依赖参数设置,容易陷入局部最优,且缺乏对电网运行环境的“学习”与“记忆”能力,每次规划都需要从零开始迭代,计算效率低下

Benefits of technology

(1) 本发明所提算法网络训练收敛速度快。传统的电网规划是一个NP-hard的组合优化问题,面对庞大的省级电网节点,潜在的线路组合呈指数级爆炸,极易导致人工智能算法陷入盲目的无效探索。本发明引入了5项基于电力系统物理运行规律的专家经验原则,大大减少了网络训练时的无效探索,加快了DDQN网络的收敛速度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122512366A_ABST
    Figure CN122512366A_ABST
Patent Text Reader

Abstract

A smart planning method for a 500kV power grid based on ultra-high voltage interconnected power grids while considering source-load balance includes: constructing a provincial main power grid model based on the PSD-BPA simulation platform; establishing power grid planning evaluation indicators and constraints; formulating screening principles for the AC / DC planning line set based on engineering experience, developing a screening process according to the principles, and finally embedding the selected planning lines into the action space of reinforcement learning; constructing a power grid planning interactive environment based on the Python programming language, defining the state space, action space, and reward function of reinforcement learning, transforming the multi-objective optimization problem of power grid planning into a Markov Decision Process (MDP), and introducing a dual deep Q-network (DDQN) for network training; using evaluation indicators to perform multi-dimensional analysis and verification of the training results; loading the provincial power grid basic data for the target year to be planned and initializing the planning environment; calling the trained DDQN network to enter the inference mode, performing feature extraction and forward propagation calculation on the current power grid state, and outputting the optimal line construction action; reconstructing the power grid topology based on the output action, and using a power system simulation program to perform safety verification and evaluation indicator analysis on the final generated planning scheme, and outputting a feasibility report.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system operation and planning technology, specifically relating to a provincial 500kV main grid planning method and system that takes into account source-load balance in the context of large-scale access of UHV AC / DC transmission, and is applicable to power grid structure optimization and intelligent planning. Background Technology

[0002] With the deepening implementation of my country's energy transition strategy, provincial power grids are gradually evolving into complex receiving-end grids with a high proportion of UHV AC / DC hybrid connections and renewable energy integration. The dense integration of numerous UHV DC landing points and UHV AC channels has fundamentally altered the power supply structure and power flow distribution characteristics of the provincial 500kV main grid. On the one hand, the receiving-end grid exhibits a new "strong DC, weak AC" pattern, with fault patterns evolving from single AC faults to complex cascading faults involving both AC and DC. On the other hand, the randomness of renewable energy output and the spatial and temporal differences in load distribution have led to increasingly prominent source-load imbalance problems. Against this backdrop, how to construct a robust, flexible, and adaptable 500kV main grid through scientific and rational grid planning has become the primary task for ensuring the safe and stable operation of the power system.

[0003] Power grid planning is a typical nonlinear, nonconvex, multi-constraint NP-hard combinatorial optimization problem. In practical engineering, planners not only need to consider the economic cost of line construction, but also need to take into account multiple physical constraints such as N-1 / N-2 safety criteria, short-circuit current limits, and source-load balance. As the scale of power grid nodes increases, the combinations of candidate lines grow exponentially, resulting in an extremely large solution space for the planning problem (curse of dimensionality). Especially when considering the often conflicting objectives of "limiting short-circuit current" and "improving mutual support capacity," traditional empirical planning methods struggle to quickly find a globally optimal solution that satisfies all safety constraints among massive topology combinations, easily leading to the risk of local power supply blockage or short-circuit capacity exceeding limits in the planning scheme.

[0004] Existing power grid planning methods are mainly divided into two categories: mathematical optimization methods and heuristic algorithms. While mathematical optimization methods (such as Mixed Integer Linear Programming (MILP)) can theoretically find the optimal solution, their computational complexity is too high when dealing with large-scale real-world power grids, making convergence difficult within a finite timeframe. Heuristic algorithms (such as Genetic Algorithm (GA) and Particle Swarm Optimization (PSO)) are computationally fast, but they are highly dependent on parameter settings, easily getting trapped in local optima, and lack the ability to "learn" and "memorize" the power grid's operating environment. Each planning iteration requires starting from scratch, resulting in low computational efficiency. Therefore, there is an urgent need to introduce intelligent planning methods based on Deep Reinforcement Learning (DRL), leveraging its powerful feature extraction and offline learning capabilities to achieve rapid generation and adaptive optimization of power grid planning schemes under complex constraints. Summary of the Invention

[0005] The present invention aims to overcome the above-mentioned shortcomings of the prior art and provide a smart planning method for 500kV power grid based on ultra-high voltage interconnected power grid while taking into account source-load balance.

[0006] This invention proposes a provincial 500kV main grid zoning method that takes into account source-load balance under the background of UHV AC / DC access, providing a theoretically innovative and engineering-practical solution for the planning, design, operation and control of provincial main grid AC / DC under the scenario of high proportion of new energy access.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A smart planning method for a 500kV power grid based on ultra-high voltage interconnected power grids while taking into account source-load balance includes the following steps: S1: Construct a provincial main power grid model based on the PSD-BPA simulation platform, covering ultra-high voltage substations, 500kV substations, large power sources, 500kV main transformers, etc.

[0008] S2: Establish power grid planning evaluation indicators and constraints. The indicator system includes regional unbalance indicators, overall network load rate indicators, and overall network short-circuit current indicators. The constraint space includes node short-circuit current constraints, system minimum inertia constraints, and line thermal stability constraints.

[0009] S3: Based on engineering experience, formulate screening principles for the set of AC / DC planned lines, formulate screening process according to the principles, and finally embed the screened planned lines into the action space of reinforcement learning.

[0010] S4: Construct an interactive environment for power grid planning based on the Python programming language, define the state space, action space and reward function of reinforcement learning, transform the multi-objective optimization problem of power grid planning into a Markov decision process (MDP), and introduce the Double Deep Q-Network (DDQN) algorithm for network training.

[0011] S5: The training results are analyzed from multiple dimensions using evaluation metrics, and the superiority of the proposed method in terms of solution efficiency and solution method is verified by comparing it with traditional planning methods.

[0012] S6: Load the basic data of the provincial power grid for the target year to be planned and initialize the planning environment; call the trained DDQN network to enter the inference mode, perform feature extraction and forward propagation calculation on the current power grid state, and output the optimal line construction action; reconstruct the power grid topology based on the output action, and use the power system simulation program to perform safety verification and evaluation index analysis on the final generated planning scheme, and output a feasibility report.

[0013] In step S1, the specific process of constructing the provincial main power grid model includes: S1-1: Obtain the current status and target year's basic power grid data within the planning area. The basic data includes: UHV substations and their access methods, 500kV substations and busbar structures, 500kV line parameters, main transformer capacity and wiring methods, large power generation capacity and output characteristics, typical load levels and regional load distribution; model the inter-provincial power transmission / receiving channels and external power grids using an equivalent method to form a network dataset with unified per-unit values ​​and unified baseline capacity, and complete node naming, zone identification, and source-load attribution mapping.

[0014] S1-2: In the PSD-BPA platform, establish the provincial main network topology in a node-branch manner, and set the connection relationships between UHV substations, 500kV substations, large power sources and 500kV main transformers; configure electrical parameters and operating constraints for lines, main transformers and power sources, including line R / X / B, rated current and thermal stability limits, main transformer capacity and tap range, power source equivalent reactance and reactive power regulation capability, etc., to obtain a complete simulation model that can be used for power flow and short circuit calculations.

[0015] In step S2, the power grid planning evaluation indicators and constraints are as follows: S2-1: The source load imbalance index is taken as the root mean square error of the source load imbalance in each zone. The load rate index is taken as the top index ranked by load rate. The average load rate of the heavy-load lines is %, and the short-circuit current index is taken as the per-unit value of the average equivalent short-circuit current of the entire network.

[0016] S2-2: Short-circuit current constraint requires that the converted value of the short-circuit current of all nodes in the network after planning does not exceed the limit value; the minimum inertia constraint condition of the system is that the equivalent inertia time constant of the system after planning meets the minimum safety threshold; the thermal stability constraint of the lines is that the actual operating current of all lines after planning must not exceed their rated current carrying capacity. In step S3, the selection principles and steps for AC / DC planning lines based on engineering experience include: S3-1: The selection of access points for AC / DC planned lines should follow these principles: Principle 1: Prioritize sites on 500kV ring networks. This approach leverages the strong topological connectivity and voltage support capabilities of ring networks to maximize the marginal benefits of new lines for global power flow distribution and N-1 fault transfer.

[0017] Principle 2: For AC / direct transmission lines, prioritize large power plants / groups, UHVDC landing points outside the province, and nearby sites. This aims to effectively alleviate power flow congestion at high-capacity source-load injection nodes and improve the transmission and distribution capacity of surplus power.

[0018] Principle 3: The receiving end of cross / straight lines should preferably be close to the load center, i.e., the point where the power flow converges on the transmission line. Selecting such nodes at the receiving end can significantly shorten the power supply radius and reduce voltage drop and network losses caused by long-distance transmission.

[0019] Principle 4: Prefer sites with low short-circuit current for AC line termination. Prioritizing nodes with lower short-circuit current allows for sufficient safety margins for grid densification and strictly prevents the planned short-circuit capacity from exceeding the breaking limit of circuit breakers.

[0020] Principle 5: The distance between candidate AC nodes should be moderate. Too short a distance has limited effect on power flow distribution and is very likely to induce overload risk; too long a distance means excessive construction costs, and high impedance will make it difficult for the line to play a diversion role in system scheduling.

[0021] S3-2: Based on the principles summarized from the above engineering experience, formulate a screening process to remove nodes that do not meet the requirements from the entire network to obtain a set of AC and DC sending and receiving end nodes, and embed the corresponding lines into the action space of reinforcement learning.

[0022] In step S4, the network preparation and training work includes: S4-1: Load the current status and target year power grid basic data in the planning area into PSD-BPA and build a complete simulation model that can be used for power flow and short circuit calculations; S4-2: Constructing an interactive platform for reinforcement learning and PSD-BPA using Python. This involves loading power flow and short-circuit calculation files using Python commands, performing power flow calculations, calculating reward metrics, and saving them to a network-accessible file path. Simultaneously, Python is used to write planning commands to add circuit functionality to the simulation model.

[0023] S4-3: Establish a Markov Decision Process (MDP) model. The state space describes the characteristics of the power grid, the action space corresponds to the set of candidate lines, and the reward space consists of reward values ​​composed of normalized evaluation indicators.

[0024] S4-4: Construct the DDQN algorithm network structure and build two deep neural networks with the same structure but different parameter update frequencies in the Python environment; S4-5: The training process employs an experience playback mechanism, which will generate quadruples through interaction. The data is stored in the experience replay pool and a batch of data is randomly sampled for training.

[0025] In step S5, the algorithm performance verification work includes: S5-1: Based on the DDQN training process, record the changes in the reward curve, combine with evaluation metrics, and compare with traditional heuristic algorithms to find the optimal solution to verify the accuracy of the algorithm.

[0026] S5-2: Based on the trained DDQN network and combined with evaluation metrics, the algorithm's speed is verified by comparing it with traditional heuristic algorithms for finding the optimal solution.

[0027] In step S6, the simulation verification of the planning scheme includes: S6-1: Read the basic data of the provincial power grid to be planned in the target year, including: the predicted load data of each 500kV substation, the installed capacity and access location of the planned new power sources, the electrical parameters (resistance, reactance, susceptance) of the existing grid, and the preset set of candidate lines; map the above data into the reinforcement learning environment to generate an initial state vector; at this time, the state vector contains the initial power flow distribution characteristics, short-circuit current level and grid topology information of the power grid; S6-2: Input the initial state vector into the DDQN network trained in S5, and quickly decide the optimal action based on the experience gained from training; decode the action into a specific planned route to form the target network structure; S6-3: Input the generated final planning scheme into the power system integrated analysis program BPA, perform N-1 / 2 safety check on the system before and after planning, and count the number of line power flow exceeding the limit and the number of power flow non-convergence after the fault. S6-4: Input the generated final planning scheme into the power system integrated analysis program BPA to test the transient stability of the system before and after planning, and record the relative power angle difference and frequency change between generator units before and after the fault.

[0028] This paper addresses the planning problem of AC / DC transmission lines in a provincial 500kV main grid under UHV interconnection. It proposes an expert experience + DDQN intelligent planning method that considers regional source-load balance to solve the problems of the curse of dimensionality in the planning decision space and the coordination of multi-dimensional nonlinear constraints. The main innovations are as follows: (1) Innovation of application scenarios: Existing literature on planning technology that considers regional source-load balance mainly focuses on distribution networks and microgrids. Distribution network planning focuses more on the topology reconfiguration of active distribution networks and the consumption of new energy with high penetration rate. The planning of microgrids and park-level integrated energy systems focuses on the capacity configuration of distributed power sources and source-load coordination. Few literatures study such issues at the 500kV main grid level.

[0029] (2) Innovation in problem-solving approach: Given that flexible DC can adopt virtual synchronous control and has a certain inertia support capability, this invention proposes to use flexible DC lines between provincial zones to promote the balance of source and load in the zones, while avoiding the increase of short-circuit current caused by the construction of new lines.

[0030] (3) Innovation in planning methods: The planning problem of AC / DC transmission lines in the 500kV main grid of a province under UHV interconnection is often solved by traditional methods such as particle swarm optimization and genetic algorithms, which consume a lot of computing power and are difficult to verify the results. This invention designs and trains a DDQN agent in Python for line planning decisions. At the same time, a real-time simulation interaction platform between the agent and BPA is built during training to ensure the reliability of training data to a certain extent. Furthermore, by introducing an expert experience candidate set, the training dimension of the reinforcement learning network is reduced. Finally, it is proved that this method is faster and more accurate than traditional methods.

[0031] The advantages of this invention are as follows: (1) The algorithm proposed in this invention has a fast network training convergence speed. Traditional power grid planning is an NP-hard combinatorial optimization problem. Faced with a large number of provincial power grid nodes, the potential line combinations explode exponentially, which can easily lead artificial intelligence algorithms into blind and ineffective exploration. This invention introduces five expert experience principles based on the physical operation laws of power systems, which greatly reduces ineffective exploration during network training and accelerates the convergence speed of the DDQN network.

[0032] (2) The algorithm proposed in this invention has high planning accuracy. Many artificial intelligence algorithms provide planning schemes that only hold true in mathematical models. Once applied to a real power grid, they often result in fatal errors such as power flow non-convergence or voltage exceeding limits, reducing planning accuracy. This paper constructs a joint interactive platform between the Python algorithm and professional power system simulation software (BPA). All state characteristics (such as short-circuit current, node voltage, and line load rate) and feedback rewards obtained by the agent during training are derived from rigorous physical simulation calculations, rather than simplified linear estimations. The final output planning scheme not only achieves global optimum in source-load balance but also meets the various physical constraints of UHV AC / DC hybrid power grids after simulation verification, demonstrating engineering accuracy.

[0033] (3) The mathematical model proposed in this invention can effectively alleviate the power grid source-load imbalance problem while satisfying physical constraints. Planning flexible direct current lines within the province can achieve precise power transmission to heavily loaded areas and alleviate the source-load imbalance problem between regions. The planning of flexible direct current lines adopts virtual synchronization technology, and its connection will affect the regional inertia and the overall network line load rate. In addition, planning AC lines within a region can alleviate the source-load imbalance problem in some areas within the region and achieve self-sufficiency within the region, but adding new AC lines within the region will inevitably increase the risk of short-circuit current exceeding the standard. The mathematical model proposed in this invention includes the goal of source-load balance between regions, and also incorporates the constraints of regional inertia, short-circuit current and line load rate, ensuring that the final solution algorithm can find the globally optimal network structure scheme that satisfies physical constraints and maximizes the cross-regional power mutual assistance capability after continuous refinement and trial and error. Attached Figure Description

[0034] Figure 1 This is a flowchart of the AC / DC candidate set screening process of the present invention.

[0035] Figure 2 This is a schematic diagram of the interaction framework between the DDQN network and the BPA simulation environment of the present invention.

[0036] Figure 3 This is a convergence comparison chart of the DDQN algorithm of this invention.

[0037] Figure 4 This is a diagram of the regional power grid structure of a certain province.

[0038] Figure 5 This is a comparison chart of the system frequency response before and after planning under fault disturbance.

[0039] Figure 6 This is a comparison chart of the system frequency response before and after planning under load disturbance. Detailed Implementation

[0040] The present invention will be further described below with reference to the embodiments shown in the accompanying drawings.

[0041] A smart planning method for a 500kV power grid based on ultra-high voltage interconnected power grids while taking into account source-load balance includes the following steps: S1: Construct a provincial main grid structure model based on the PSD-BPA simulation platform, covering UHV substations, 500kV substations, large power sources, 500kV main transformers, etc.

[0042] S2: Establish power grid planning evaluation indicators and constraints. The indicator system includes regional unbalance indicators, overall network load rate indicators, and overall network short-circuit current indicators. The constraints include node short-circuit current constraints, system minimum inertia constraints, and line thermal stability constraints.

[0043] S3: Based on engineering experience, formulate screening principles for the set of AC / DC planned lines, formulate screening process according to the principles, and finally embed the screened planned lines into the action space of reinforcement learning.

[0044] S4: Construct an interactive environment for power grid planning based on the Python programming language, define the state space, action space and reward function of reinforcement learning, transform the multi-objective optimization problem of power grid planning into a Markov decision process (MDP), and introduce the Double Deep Q-Network (DDQN) algorithm for network training.

[0045] S5: Based on the trained DDQN network, the final target network structure is generated for the power grid scenario to be planned; the generated planning scheme is evaluated in multiple dimensions using power system simulation tools; and the superiority of the proposed method in terms of solution efficiency and solution method is verified by comparing it with traditional planning methods.

[0046] S6: Based on the simulation model, perform N-1 / 2 safety checks and system transient stability simulation tests to verify the rationality of the planning scheme.

[0047] In step S1, the specific process of constructing the provincial main grid structure power system model includes: S1-1: Obtain the current status and target year's basic power grid data within the planning area. The basic data includes: UHV substations and their access methods, 500kV substations and busbar structures, 500kV line parameters, main transformer capacity and wiring methods, large power generation capacity and output characteristics, typical load levels and regional load distribution. S1-2: In the PSD-BPA platform, establish the provincial main network topology in a node-branch manner, and set the connection relationships between UHV substations, 500kV substations, large power sources and 500kV main transformers; configure the electrical parameters and operating constraints of lines, main transformers and power sources, including the resistance, reactance, admittance, rated current and thermal stability limit of the lines, the capacity and tap range of the main transformers, the equivalent reactance and reactive power regulation capability of the power sources, etc., to obtain a complete simulation model that can be used for power flow and short circuit calculations.

[0048] In step S2, the power grid planning evaluation indicators and constraints are as follows: S2-1: The source-load imbalance index is taken as the root mean square error of the source-load imbalance in each zone. When calculating the power supply for each zone, the output ratio of each type of power source under different operating conditions should be considered. The load rate index is taken as the average load rate of the top 10% of heavily loaded lines ranked by load rate. Considering that the short-circuit current used in actual engineering is related to the rated breaking time of the circuit breaker, the initial effective value of the three-phase short-circuit current at the node is multiplied by a conversion factor to obtain the equivalent short-circuit current of that node. The short-circuit current index is taken as the per-unit value of the average equivalent short-circuit current of the entire network. The specific calculations for the indicators are as follows: 1) Regional source-load imbalance index Considering the different power source types and their responsibilities in the power grid, zones are defined. Power generation Specifically, as shown in equation (1). , , , , , and These are the rated installed capacities of coal-fired power, gas-fired power, pumped storage power, hydropower, nuclear power, wind power, and photovoltaic power, respectively; the coefficient before the rated installed capacity. a , b, c, d, e, f, g This represents the output ratio of various power sources under different operating conditions.

[0049] (1)

[0050] partition The definition of the source load imbalance is shown in equation (2), where Access partition DC power, This is the DC power adjustable coefficient; For partitioning Peak load power within the area, Load output coefficient (2) Source load imbalance in each partition Divide by a uniform base value As shown in equation (3), the dimensionless value is obtained. After standardization, results from different zones and different scales of operation can be directly compared on the same scale.

[0051] (3)

[0052] This invention uses the root mean square error (RMSE) as the statistical measure of partition imbalance.

[0053] The source-load imbalance index between partitions is defined as follows: (4) (5) It reflects the source-load balance relationship between partitions within the system. The smaller the value, the higher the system's source-load self-balancing capability; conversely, the greater the power transmitted across partitions, the more it will exacerbate the risk of systemic cascading failures.

[0054] 2) Load factor

[0055] Line load factor characterizes the stress level of a line when bearing system power flow. This invention uses the system's front... x The average load rate of heavily loaded lines is used as a load rate indicator. As shown in equation (6). The top-ranked systems by load rate x % of the number of heavy-load lines; The load rate of the corresponding line is calculated as shown in equation (7); For the line l The actual operating current, For the line l The rated current limit.

[0056] (6) (7) It reflects the overall line operation status of the system. The smaller the value, the safer the line is in operation; conversely, the larger the value, the higher the line is under high load or overload, and there is a potential operational risk.

[0057] 3) Short-circuit current index

[0058] Considering that the short-circuit current is related to the rated breaking time of the circuit breaker, this invention first defines the circuit breaker switching time node. n equivalent short-circuit current Specifically, as shown in equation (8), where The initial effective value of the three-phase short-circuit current at the node. It is the system reduction factor; then the system short-circuit current index is defined as the per-unit value of the average equivalent short-circuit current of the entire network. As shown in equation (9), where N This represents the total number of nodes in the network. The rated breaking capacity of the circuit breaker; (8) (9) It reflects the short-circuit current level of the entire network. The smaller the value, the smaller the short-circuit impact on the system equipment and the lower the grid construction cost.

[0059] S2-2: Short-circuit current constraint requires that the converted value of the short-circuit current of all nodes in the network after planning does not exceed the limit value; the minimum inertia constraint condition for each region is that the equivalent inertia time constant of the region after planning meets the minimum safety threshold; the line thermal stability constraint is that the actual operating current of all lines after planning must not exceed their rated current carrying capacity. The mathematical expression of the constraints is as follows: 1) Node short-circuit current constraint To ensure that circuit breakers can reliably disconnect fault currents when a power grid fault occurs, the planned grid structure must guarantee that the short-circuit current values ​​of all nodes in the entire grid do not exceed the limit, as shown in equation (10). Based on the current topology z Calculated nodes n The converted value of the short-circuit current; This refers to the rated breaking capacity of the circuit breaker.

[0060] (10)

[0061] 2) Minimum Inertia Constraint for Partitions

[0062] To prevent excessive frequency drops caused by UHVDC blocking or faults in critical interconnection lines, the system must maintain sufficient rotating reserve inertia. Therefore, the equivalent inertial time constant of each zone after planning must meet the minimum safety threshold, as shown in equation (11). For the region k The set of units that are in operation within the facility; and The units g The inertial time constant and rated capacity; The minimum inertia threshold required to maintain system frequency stability. In the planning, it is assumed that the flexible DC transmission uses virtual synchronous control and has a certain inertia support capability, which can improve the inertia of the grid connection.

[0063] (11)

[0064] 3) Line thermal stability constraints

[0065] To ensure the safe operation of the power grid, the actual operating current of all lines must not exceed their rated current carrying capacity, which means that the constraint conditions shown in equation (12) must be met.

[0066] (12)

[0067] In step S3, the selection principles and specific steps for AC / DC line access nodes based on engineering experience include: S3-1: The following selection principles are formulated based on the expert experience candidate route set and engineering experience.

[0068] Principle 1: Prioritize sites on 500kV ring networks. This approach leverages the strong topological connectivity and voltage support capabilities of ring networks to maximize the marginal benefits of new lines for global power flow distribution and N-1 fault transfer. Principle 2: Prioritize large power plants / groups, UHVDC landing points outside the province, and nearby sites for AC / direct transmission lines; this measure aims to effectively alleviate power flow congestion at high-capacity power injection nodes and improve the transmission and distribution capacity of surplus power. Principle 3: The receiving end of cross / straight lines should preferably be close to the load center, i.e., the point where the power flow converges on the transmission line. Selecting such nodes at the receiving end can significantly shorten the power supply radius and reduce voltage drop and network losses caused by long-distance transmission. Principle 4: Prefer sites with low short-circuit current for AC line termination. Prioritizing nodes with lower short-circuit current allows for sufficient safety margins for grid densification and strictly prevents the planned short-circuit capacity from exceeding the breaking capacity of circuit breakers. Principle 5: The distance between candidate AC nodes should be moderate. Too short a distance has limited effect on power flow distribution and is very likely to induce overload risk; too long a distance means excessive construction costs, and high impedance will make it difficult for the line to play a diversion role in system scheduling. It should be noted that principles 4 and 5 above mainly apply to the selection of AC candidate lines. For flexible DC candidate lines, since the converter has flexible active power flow control capabilities and AC / DC fault decoupling characteristics, its connection is not limited by the natural impedance distribution of the AC system, nor does it significantly increase the short-circuit capacity of the receiving-end grid. Therefore, when generating a set of flexible DC candidate lines, only the first three principles need to be followed. In addition, since both flexible DC lines and UHV lines serve as inter-regional power transfer channels in the power grid, the connection point of the former should be appropriately far away from the UHV site; S3-2: Based on the above principles, develop a screening flowchart as follows: Figure 1 As shown. The node net power... This is the difference between the active power flowing out of and injected into a node in the power flow simulation results; For the region k The upper limit of the net power constraint for the mid-node is used to filter out the sending-end nodes that meet the requirements in this area; For the region The lower limit of the net power constraint for the middle node is used to filter the receiving-end nodes that meet the requirements in this area; For the region k The upper limit of the short-circuit current constraint is used to filter out AC sending and receiving end nodes that meet the requirements. Finally, the selected AC and DC sending and receiving end nodes are arranged and combined to obtain the final set of AC and DC candidate lines, which is used as the action decision space for reinforcement learning.

[0069] In step S4, the network preparation and training work includes: S4-1: Load the current status and target year power grid basic data in the planning area into PSD-BPA and build a complete simulation model that can be used for power flow and short circuit calculations; S4-2: As attached Figure 2 As shown, a reinforcement learning and PSD-BPA interactive platform is constructed using Python; power flow and short-circuit calculation files are loaded using Python commands, power flow calculation is performed, reward indicators are calculated and saved to a network-accessible file path; simultaneously, planning commands are written using Python to add line functionality to the simulation model; when calculating the reward indicators, it should be noted that the indicators proposed in S2 are all aimed at minimizing during the optimization process, which contradicts the actual training logic of reinforcement learning. Therefore, the sigmoid function is used to process the indicators into reward values ​​that can be used for reinforcement learning training, and the larger the value, the better; the sigmoid function is a classic function that maps real numbers to (0,1) probability values, as shown in equations (10) and (11); where It is a reference value that maps to real numbers. μ The sensitivity is adjustable as a function; (13) (14) The above indicators are mapped to the range (0,1) using the sigmoid function, and the reference values ​​are taken from the values ​​of each indicator under the initial unplanned power grid. , and ; S4-3: Establish a Markov Decision Process (MDP) model. Define the state vector. Used to describe time The power grid characteristics include topology connection status, active and reactive power injection at each node, and node short-circuit current. Define the action set. For a discrete action space, each action This corresponds to construction within the candidate line set. The three evaluation metrics are converted into scalar reward values, and the constraints are incorporated into the reward function as strong penalty terms; the specific reinforcement learning elements are as follows; 1) State Space

[0070] In order for the network (Agent) to perceive the electrical operating characteristics and topological features of the UHV interconnected power grid, this invention defines the first... The global state of the system at the decision step It is a composite vector containing node operation characteristics and network topology characteristics.

[0071] (15)

[0072] In the formula, and These are the active and reactive power injection vectors for each node, used to reflect the source load distribution. This represents the converted three-phase short-circuit current value at each node, reflecting the electrical strength of the node. For UHV receiving-end power grids, The size of the voltage drop determines the system's ability to withstand voltage drops and whether it meets the short-circuit ratio requirements of multiple DC feeds, and is the core basis for selecting the flexible DC landing point. This is the adjacency matrix of the entire network topology, used to directly describe the physical connections of the power grid. Its expression is: (16) Matrix elements The value can be 0 or 1, where 1 represents a node. With nodes There is a physical connection between them; otherwise, 0 represents no connection. Initially, It only includes the original power grid connections; as the planning progresses, when the network performs the action of adding lines, the element at the corresponding position is directly set to 1.

[0073] 2) Action Space

[0074] To address the combined explosion problem and adapt to the physical characteristics of ultra-high voltage interconnected power grids, this invention employs a discrete action space based on an expert experience candidate set.

[0075] (17)

[0076] in This represents the total number of candidate AC / DC lines. (Action) It is an integer index pointing to the candidate set of expert experience. The execution logic of network actions is as follows: when the output action is... When this condition is met, it indicates that an element of the adjacency matrix in the state space is set to 1, signifying the creation of a new line. If this line is an AC tie line within the region, a new double-circuit AC line will be created in the BPA model; if it is a flexible DC transmission line, the VSC-HVDC model for the corresponding access point will be enabled in the BPA model.

[0077] 3) Reward Space

[0078] The reward function is designed following the principles of objective function guidance and strong constraint penalties. The objective function is denoted as a composite reward value and written into the reward space. The network at time step [time value missing]... Receive a reward after completing the action. A large negative reward value is given when any of the constraints occur during the simulation. And thus, this round of training was terminated ahead of schedule.

[0079] 4) Discount Factor

[0080] Discount factor This indicates the conversion factor for future rewards, which is adjusted based on the characteristics of the task.

[0081] S4-4: Constructing the DDQN algorithm network structure. Build two deep neural networks with identical structures but different parameter update frequencies in a Python environment. Train the network. The parameters are It is used to calculate the Q-value of each action in the current state in real time, guiding action selection. Target network The parameters are This is used to calculate the target Q-value to stabilize the training process; S4-5: Network Training and Policy Update. (Using...) The strategy selects actions based on probability. Explore using random actions, based on probability. Select the action corresponding to the maximum output value of the trained network and utilize it. Use the quadruplets generated by the interaction. The data is stored in the experience replay pool. Once the experience pool reaches a threshold, a batch of data is randomly sampled for training. The mean squared error between the predicted value and the target Q-value is calculated, and the training network parameters are updated using gradient descent. Finally, after a certain number of steps, the parameters of the trained network are synchronized to the target network at a small ratio. The specific process is as follows: 1) Environment Initialization Load the original UHV receiving-end power grid data and initialize DDQN. and Network parameters, clear the experience replay pool D Generate the initial state. The adjacency matrix Only the existing space frame is included.

[0082] 2) Action Selection

[0083] In the decision-making step The network is based on the current state Choose actions based on the expertise of specialists. .

[0084] 3) Simulation interaction and state transition

[0085] Will This is mapped to specific line commissioning instructions, which invoke BPA to perform topology updates, security checks, and indicator extraction. If the constraints are met, the entire network's electrical quantities are extracted and generated. Conversely, if the penalty value is strong, the round is terminated.

[0086] 4) Reward Calculation and Experience Storage

[0087] The overall reward is calculated based on the results from the BPA feedback, and the quadruple for the entire round is calculated. Store in the experience replay pool.

[0088] 5) Network parameter update

[0089] When the amount of data in the playback pool meets the batch requirements, random sampling is performed. Samples, and perform DDQN parameter updates, where parameters Calculated and updated using gradient descent. parameters Update slowly using a soft update method.

[0090] 6) Multi-round training

[0091] Return to step 2) and perform iterative training until the network converges.

[0092] In step S5, the algorithm performance verification work includes: S5-1: Convergence Verification. Table 1 shows the candidate set of expert experience selected for AC / DC planning in this power grid, which is embedded into the action space for network training. The network training convergence process is attached. Figure 3 As shown, the results demonstrate that the planning method proposed in this invention finds a solution very close to the optimal solution given by the enumeration method. Furthermore, after incorporating the expert experience constraints in Table 1, the algorithm's decision-making time is shortened and its robustness is increased, verifying the algorithm's accuracy and speed. In Table 1, the nodes within the thin solid lines correspond to... Figure 4 The green area in the image corresponds to the nodes within the thick solid line box. Figure 4 The nodes in the purple area (within the dashed box) correspond to... Figure 4 The blue area.

[0093] Table 1

[0094] S5-2: Verification of the Accuracy and Speed ​​of the Planning Algorithm. Table 2 shows a comparison of the decision search space. As shown in Table 2, traditional heuristic algorithms require a large amount of computing power, while the DDQN network can reduce the computing power consumption. A provincial power grid pre-plans 3 regional AC lines and 1 inter-regional DC line, using a trained DDQN network for DC line planning. Table 3 shows the results of the planning requirements under different scenarios using the network and traditional algorithms, i.e., a comparison of the planning accuracy of DDQN and enumeration methods. The gray areas in the table represent newly planned lines. As can be seen from the table, the DDQN planning algorithm proposed in this invention not only maintains a high reward value but also greatly shortens the decision time, verifying the accuracy and speed of the algorithm. Table 4 lists the specific values ​​of each indicator and constraint of the planning scheme given by the DDQN network in scenario 1, and gives a comparison of the key performance indicators of the system before and after planning. After the DDQN network coordinates the optimization of the flexible DC landing point and AC tie line, all key performance indicators of the system, except for the short-circuit current indicator, are synergistically improved. System source-load imbalance indicator The value decreased from an initial 0.4973 to 0.3805, an improvement of 23.49%, indicating that the newly added transmission channels effectively broke through the original regional power barriers and significantly reduced the unbalanced power exchange within the system; the average load rate of the top 10% heavily loaded lines in the system... The load factor of critical transmission sections was reduced by 4.50%, significantly enhancing the system's redundancy and robustness against N-1 faults. Notably, the average short-circuit current index in the planning scheme was also reduced. The slight increase is due to the reduced system equivalent impedance resulting from the increased interconnectivity of the network, which objectively enhances the system's voltage support capability and disturbance rejection level. Furthermore, the current values ​​at each node did not exceed limits, meeting the requirements for safe equipment operation. In addition, key constraints remained within limits and had sufficient margins, satisfying the system's safe operation requirements. The statistical results verify the comprehensive superiority of the planning scheme in terms of supply and demand balance, operational safety, and short-circuit stability.

[0095] Table 2

[0096] Table 3

[0097] Table 4

[0098] In step S6, the simulation tests before and after planning include: S6-1: Read the basic data of the provincial power grid to be planned in the target year, including: the predicted load data of each 500kV substation, the installed capacity and access location of the planned new power sources, the electrical parameters (resistance, reactance, susceptance) of the existing grid, and the preset set of candidate line corridors; map the above data into the reinforcement learning environment to generate an initial state vector; at this time, the state vector contains the initial power flow distribution characteristics, short-circuit current level and grid topology information of the power grid; S6-2: Input the initial state vector into the DDQN network trained in S5, set the network's action selection strategy to be fully greedy, calculate the Q-value vectors of all candidate line actions through the forward propagation of the deep neural network, and finally select the action with the largest Q-value; decode the action into a specific planned line to form the target network structure; S6-3: As attached Figure 3 As shown, the network presents the final planned power grid structure. N-1 / 2 safety checks and transient stability analyses are performed on this planned scheme. The N-1 / 2 safety checks of the system before and after planning are shown in Table 5. The results show that the planned scheme significantly improves the static safety level of the system. In the N-1 safety check, the three overloaded lines before planning are completely eliminated, and the system successfully achieves "zero overload" safe operation under N-1 conditions, meeting the core reliability criteria of power grid planning. For the more stringent N-2 extreme conditions, the optimized network structure also exhibits stronger robustness, with four fewer overloaded power flow lines and one fewer non-converging power flow line after the interruption. Table 5

[0099] S6-4: Perform system transient stability analysis on the planning scheme. This invention sets up two disturbance modes. One is that a permanent double-circuit three-phase short-circuit fault occurs on the 500kV line between nodes 10 and 11 at 5s, and is cleared after 0.2s. The frequency difference between the generator sets at nodes 14 and 68 is observed to be as follows: Figure 4 As shown; secondly, by reducing the load in region A by 5000MW at 1s, the frequency difference response between the generator units at node 14 and node 68 is obtained as follows. Figure 5 As shown in the figure. The results indicate that under fault and load disturbances, the frequency difference response oscillations between the generator units after the planning are decayed faster and the steady-state recovery is smoother, resulting in better system transient performance. This is attributed to the fact that the newly planned lines optimize the source-load distribution, reduce regional imbalance and power exchange at key sections, and the use of virtual synchronous control in the flexible DC transmission line increases the system inertia, thereby reducing the relative transient swing amplitude of the frequency between sections and accelerating oscillation decay.

Claims

1. A smart planning method for a 500kV power grid based on ultra-high voltage interconnected power grids while considering source-load balance, characterized in that, Includes the following steps: S1: Construct a provincial main power grid model based on the PSD-BPA simulation platform, covering UHV substations, 500kV substations, large power sources, and 500kV main transformers; S2: Establish power grid planning evaluation indicators and constraints. The indicators include regional unbalance indicators, overall network load rate indicators, and overall network short-circuit current indicators. The constraints include node short-circuit current constraints, system minimum inertia constraints, and line thermal stability constraints. S3: Based on engineering experience, formulate the selection principles for the AC / DC planning line set, formulate the selection process according to the principles, and finally embed the selected planning lines into the action space of reinforcement learning; S4: Based on the Python programming language, an interactive environment for power grid planning is constructed. The state space, action space and reward function of reinforcement learning are defined. The multi-objective optimization problem of power grid planning is transformed into a Markov decision process (MDP). A dual deep Q-network (DDQN) is introduced for network training. S5: Utilize evaluation metrics to conduct multi-dimensional analysis and verification of training results; S6: Load the provincial power grid basic data for the target year to be planned and initialize the planning environment; call the trained DDQN network to enter the inference mode, perform feature extraction and forward propagation calculation on the current power grid status, and output the optimal line construction action; The power grid topology is reconstructed based on the output actions, and the power system simulation program is used to perform safety verification and evaluation index analysis on the final generated planning scheme, and output a feasibility report.

2. The intelligent planning method for a 500kV power grid based on UHV interconnected power grids while considering source-load balance, as described in claim 1, is characterized in that... In step S1, the specific process of constructing the provincial main power grid model includes: S1-1: Obtain the current status and target year power grid basic data within the planning area. The basic data includes: main connection mode of UHV substations, main connection mode of 500kV substations, 500kV line parameters, main transformer parameters, installed capacity and output characteristics of large power sources, typical load levels and regional load distribution, and detailed model parameters of generators and dynamic load elements. S1-2: Establish the provincial main network topology in the PSD-BPA platform, configure electrical parameters for lines, main transformers, generators, dynamic load components, etc., and obtain a complete simulation model that can be used for power flow calculation, short circuit calculation and electromechanical transient analysis.

3. The intelligent planning method for a 500kV power grid based on UHV interconnected power grids while considering source-load balance, as described in claim 1, is characterized in that... In step S2, the power grid planning evaluation indicators and constraints are as follows: S2-1: The source load imbalance index is taken as the root mean square error of the source load imbalance in each zone. The load rate index is taken as the top index ranked by load rate. The average load rate of the heavy-load lines is %, and the short-circuit current index is taken as the per-unit value of the average equivalent short-circuit current of the entire network. S2-2: The short-circuit current constraint requires that the converted value of the short-circuit current of all nodes in the network after planning does not exceed the limit value; the minimum inertia constraint condition of the system is that the equivalent inertia time constant of the system after planning meets the minimum safety threshold. The thermal stability constraint for the lines is that the actual operating current of all lines after planning must not exceed their rated current carrying capacity.

4. The intelligent planning method for a 500kV power grid based on UHV interconnected power grids while considering source-load balance, as described in claim 1, is characterized in that... In step S3, the selection principles and steps for AC / DC planning lines based on engineering experience include: S3-1: The selection of access points for AC / DC planned lines should follow these principles: Principle 1: Prefer sites on 500kV ring networks. This approach can take advantage of the strong topological connectivity and voltage support capabilities of the ring network to maximize the marginal benefits of new lines for global power flow distribution and N-1 fault transfer. Principle 2: The preferred sending end of the AC / straight line is a large power plant / group, an ultra-high voltage DC landing point outside the province, and a nearby site. This measure aims to effectively alleviate the power flow congestion problem at the high-capacity power injection node and improve the transmission and distribution capacity of surplus power. Principle 3: The receiving end of the cross / straight line should preferably be close to the load center, that is, the power flow convergence point of the transmission line. Selecting such a node at the receiving end can significantly shorten the power supply radius and reduce voltage drop and network loss caused by long-distance power transmission. Principle 4: The preferred landing points for AC lines are sites with low short-circuit currents. Prioritizing nodes with low short-circuit currents can provide sufficient safety margins for network densification and strictly prevent the planned short-circuit capacity from exceeding the breaking limit of the circuit breaker. Principle 5: The distance between candidate AC nodes should be moderate. Too short a distance has limited effect on power flow distribution and is very likely to induce overload risk; too long a distance means that the construction cost is too high and the high impedance will make it difficult for the line to play a diversion role in system scheduling. S3-2: Based on the principles summarized from the above engineering experience, formulate a screening process to remove nodes that do not meet the requirements from the entire network to obtain a set of AC and DC sending and receiving end nodes, and embed the corresponding lines into the action space of reinforcement learning.

5. The intelligent planning method for a 500kV power grid based on UHV interconnected power grids while considering source-load balance, as described in claim 1, is characterized in that... In step S4, the preparation and training of the DDQN network includes: S4-1: Load the current status and target year power grid basic data in the planning area into PSD-BPA and build a complete simulation model that can be used for power flow and short circuit calculations; S4-2: Build a reinforcement learning and PSD-BPA interactive platform using Python. Load power flow and short-circuit calculation files, perform power flow calculations, calculate reward indicators and save them to a network-accessible file path through Python commands. At the same time, use Python to write planning commands to add line function reward indicators to the simulation model. S4-3: Establish a Markov Decision Process (MDP) model; the state space is used to describe the characteristics of the power grid, the action space corresponds to the set of candidate lines, and the reward space is the reward value composed of normalized evaluation indicators. S4-4: Construct a DDQN network, building two deep neural networks with the same structure but different parameter update frequencies in the Python environment; S4-5: The training process employs an experience playback mechanism, which will generate quadruples through interaction. The data is stored in the experience replay pool and a batch of data is randomly sampled for training.

6. The intelligent planning method for a 500kV power grid based on UHV interconnected power grids while considering source-load balance, as described in claim 1, is characterized in that... In step S5, the algorithm performance verification work includes: S5-1: Based on the recording of reward curve changes during the training process of the DDQN network, combined with evaluation metrics, the algorithm's accuracy is verified by comparing it with traditional heuristic algorithms for finding the optimal solution. S5-2: Based on the trained DDQN network and combined with evaluation metrics, the algorithm's speed is verified by comparing it with traditional heuristic algorithms to find the optimal solution.

7. The intelligent planning method for a 500kV power grid based on UHV interconnected power grids while considering source-load balance, as described in claim 1, is characterized in that... In step S6, the specific process includes: S6-1: Read the basic data of the provincial power grid to be planned in the target year, including: the predicted load data of each 500kV substation, the installed capacity and access location of the planned new power sources, the electrical parameters of the existing grid, and the preset set of candidate line corridors; map the above data into the reinforcement learning environment to generate an initial state vector; at this time, the state vector contains the initial power flow distribution characteristics, short-circuit current level, and grid topology information of the power grid. S6-2: Input the initial state vector into the DDQN network trained in S5. The network quickly decides the optimal action based on the experience gained from training. The action is then decoded into a specific planned route to form the target network structure. S6-3: Input the generated final planning scheme into the power system integrated analysis program BPA, perform N-1 / 2 safety check on the system before and after planning, and count the number of line power flow exceeding the limit and the number of power flow non-convergence after the fault. S6-4: Input the generated final planning scheme into the power system integrated analysis program BPA to test the transient stability of the system before and after planning, and record the relative power angle difference and frequency change between generator units before and after the fault.