A method and system for arranging transmission line towers based on deep reinforcement learning
Patent Information
- Application Number
- CN202610921273.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2046-06-25
AI Technical Summary
[0007]本发明旨在克服现有技术中排位效率低、难以全局最优以及对复杂地形适应性差的问题,提供一种基于深度强化学习的输电线路杆塔排位方法及系统,实现排位过程的自动化、智能化和成本最小化
[0024]The beneficial effects of this invention are: traditional dynamic programming methods are limited by the discretization of the state space and are prone to getting trapped in local optima under long distances and complex terrain. This invention adopts a policy network based on the Transformer architecture, which utilizes its self-attention mechanism to achieve global perception of the terrain profile, and can simultaneously handle complex constraints such as terrain undulations, obstacles, electrical safety, and mechanical strength.
Smart Images

Figure CN122452382B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power engineering design and artificial intelligence, specifically to a method and system for arranging transmission line towers based on deep reinforcement learning. Background Technology
[0002] The arrangement of transmission line towers is a core aspect of transmission engineering design. Its essence is to seek the global optimal solution between engineering cost and operational safety under complex three-dimensional spatial constraints.
[0003] With the widespread adoption of high-precision LiDAR mapping technology and 3D digital handover, the environmental model for pole and tower placement has evolved from traditional two-dimensional cross-sectional diagrams to high-dimensional, massive-data 3D point cloud profiles. Traditional techniques have the following drawbacks: Computational complexity and search space conflict: Traditional dynamic programming (DP) algorithms heavily rely on the discretization accuracy of the state space. In long-distance, complex terrain scenarios, to avoid the "curse of dimensionality," DP algorithms often have to simplify constraints, making them prone to getting trapped in local optima when dealing with nonlinear constraints.
[0004] The system suffers from weak multi-constraint coupling processing capabilities: Tower placement must simultaneously meet stringent constraints such as electrical insulation distances (to the ground, to buildings, and to vegetation), mechanical strength (angle and span limitations), and terrain and geological conditions. Traditional hard-coded rules struggle to adapt to dynamically changing real-time constraints.
[0005] Lack of strategic flexibility: Existing algorithms struggle to balance the discontinuity of cost functions (such as the cost jumps between different tower types) with the heuristic application of engineering experience, resulting in generated solutions often requiring extensive manual secondary corrections.
[0006] Deep reinforcement learning, through its interactive perception and decision-making mechanism between agents and the environment, can handle high-dimensional state spaces and learn optimal policies in complex environments, providing a new technical approach to solving the high-dimensional combinatorial optimization problem of pole ranking. However, directly applying existing reinforcement learning algorithms to the pole ranking task suffers from problems such as sparse rewards and difficulty in meeting stringent engineering safety constraints, leading to unstable training, slow convergence speed, and ultimately, solutions that fail to meet practical requirements in terms of engineering feasibility and economy. Summary of the Invention
[0007] This invention aims to overcome the problems of low ranking efficiency, difficulty in achieving global optimization, and poor adaptability to complex terrain in the existing technology, and provides a method and system for ranking transmission line towers based on deep reinforcement learning, so as to realize the automation, intelligence and cost minimization of the ranking process.
[0008] The technical solution adopted by this invention to solve its technical problem is a transmission line tower placement method based on deep reinforcement learning, which includes the following steps: S1. Constructing a simulation environment: Based on the digital elevation model and ground feature classification data of the power transmission corridor, construct a simulation environment that supports sequential decision-making and continuous action space to simulate the tower placement process. S2. Design Strategy Network: The strategy network is used to receive the environmental state of the sliding pane and directly output the joint continuous action of the subsequent N towers through the neural network. The joint continuous action includes the continuous action of the span and tower height of each tower to be placed. S3. Simulation Environment Modeling and Reward Mechanism Construction: Design a multi-objective composite reward function to guide the policy network to evaluate the compliance and economy of tower placement during training. Penalize invalid placements with evaluation scores below the threshold, and set specific reward or penalty items for unfavorable terrain and tower type matching. S4. Training using a course-based learning strategy: In the early stages of training, the reward function parameters are adjusted to make the policy network tend to generate compact and conservative tower layouts in order to accumulate effective experience; then the reward parameters are gradually adjusted to guide the policy network to explore optimization schemes with larger spans and lower total costs while meeting compliance constraints, until the strategy converges. S5. Post-processing and scheme evaluation: Post-process the tower placement schemes generated by the trained policy network during the inference phase, collect and statistically analyze the span, tower height, tower type, minimum ground distance, compliance parameters and reward composition for each span, reproduce the tower placement process in an independent evaluation environment, calculate the overall success rate, average compliance parameters and performance indicators of the scheme, and output a structured report.
[0009] Furthermore, in S2, the policy network adopts a Transformer-based architecture, specifically including: S201. State Space Construction: Extract the topographic profile features of the current area to be ranked, including high-order sequences, land cover classification codes, and geospatial coordinates; S202, Feature Encoding: Spatial sequence information is injected into terrain points using location encoding, and global terrain semantic features are extracted through the self-attention mechanism of the Transformer encoder; S203, Multimodal Fusion: The extracted terrain feature embedding vector is concatenated with the real-time scalar state vector and mapped to the action space through a multilayer perceptron; and the policy network is configured to output continuous multidimensional action vectors, which are respectively mapped to the span and tower height parameters of the subsequent N base towers.
[0010] Furthermore, S3 includes: S301. Physical simulation environment construction: Virtually place towers based on the actions output by the policy network; S302, Evaluation Unit Integration: Construct an independent compliance evaluation unit that outputs whether the current tower placement scheme meets the constraints of ground distance, horizontal span LH, and vertical span LV. When the constraints are met... 1. If Then, negative feedback is applied to the policy network through constraint penalty terms; S303. Design a multi-objective coupled reinforcement learning reward function, which includes at least several of the following: progress increment reward, terrain-inspired reward, tower-matching reward, economic penalty, and compliance constraint penalty.
[0011] Furthermore, the progress increment reward is defined as:
[0012] in To advance the weighting factor, The elevation difference sensitivity coefficient, The horizontal distance of the current step forward, This is the absolute value of the height difference in the current gear. The terrain-inspired reward is defined as:
[0013] in The high ground reward coefficient, The valley penalty coefficient, The current elevation. The average elevation within the window. This is a local high point indicator function. Valley indicator function; The tower-type matching reward is defined as follows:
[0014] in Cost sensitivity coefficient The cost of the benchmark tower type, The actual tower construction cost mapped to the current action; The economic penalty is defined as follows:
[0015] in For the height of the tower, For exponential coefficients, This is the cost coefficient for the tower type. and These are the height penalty weight and the tower cost penalty weight, respectively. The compliance constraint penalty is defined as follows:
[0016] in The penalty intensity coefficient, The compliance parameters output by the compliance assessor; The reinforcement learning reward function is:
[0017] in, , , , , For the corresponding weights.
[0018] Furthermore, in S4, the course learning strategies include: S401, Set a dynamic adjustment factor This is used to control the progress weighting factor in the progress increment reward; S402, Initialization Phase: Settings A negative value forces the policy network to learn safely over short intervals. S403, Performance Evolution Phase: When the automatic ranking success rate or training steps reach a preset threshold, Gradually increase from negative values to positive values; S404, Economy-Driven Global Policy Optimization: In the later stages of training, Set to a positive value and increase the weight of the economic penalty.
[0019] Furthermore, S5 includes: S501. In an independent evaluation environment, starting from the first tower of the line, the tower positions are generated one by one in an autoregressive manner. S502. Set an abnormally short span threshold, and classify consecutive tower nodes with adjacent spacing below the threshold into a redundant cluster. Only retain the tower node with the largest surface elevation in the cluster and remove the rest of the nodes. S503. Collect and record span, tower height, conductor-to-ground distance, and compliance parameters output by the compliance assessor for each span, and summarize them into structured records. S504 calculates the automatic ranking success rate, the inference time of a single compliance assessor, and the average decision time per tower, and outputs a structured report.
[0020] A transmission line tower ranking system based on deep reinforcement learning includes: The simulation environment construction module is used to build a simulation environment that supports sequential decision-making and continuous action space based on the digital elevation model and ground feature classification data of the power transmission corridor. The strategy network module is used to receive the environmental state of the sliding pane and determine the joint continuous action of the subsequent N base towers through the output of the neural network. The action includes the continuous action of the span and tower height of each tower to be placed. The reward mechanism construction module is used to design multi-objective composite reward functions, evaluate the compliance and economy of tower placement, impose penalties for ineffective placement, and set specific reward or penalty items for unfavorable terrain and tower type matching. The course learning and training module is used to guide the policy network from a conservative layout to an economically optimal solution by dynamically adjusting the parameters of the reward function. The post-processing and evaluation module is used to post-process the tower placement schemes generated in the inference phase, calculate the overall success rate, average compliance parameters and performance indicators, and output a structured report.
[0021] Furthermore, the policy network module adopts a Transformer-based architecture, including: State space construction unit is used to extract the terrain profile features of the current area to be ranked; The feature encoding unit is used to extract global terrain semantic features using the self-attention mechanism of position encoding and Transformer encoder; A multimodal fusion unit is used to concatenate terrain feature embedding vectors with real-time scalar states; Action space definition unit, used to output a continuous multidimensional action vector mapped to span and tower height parameters.
[0022] Furthermore, the reward mechanism construction module includes: The physical simulation unit is used to virtually place towers and calculate conductor sag based on the actions output by the policy network. The compliance evaluator unit is an independent evaluator used to assess whether the current ranking scheme meets the requirements for ground distance, horizontal clearance, and vertical clearance constraints, and to provide negative feedback when they do not. A multi-objective reward function unit is used to calculate at least one of the following: progress increment reward, terrain-inspired reward, tower matching reward, economic penalty, and compliance constraint penalty.
[0023] Furthermore, the post-processing and evaluation module includes: The abnormally short span merging unit is used to cluster tower nodes with adjacent spacing below a threshold and retain only the node with the highest ground elevation. The span-by-span indicator collection unit is used to record span distance, tower height, distance to ground, compliance parameters, and reward weight; The performance statistics unit is used to calculate success rate, inference time, and decision-making time, and outputs a structured report.
[0024] The beneficial effects of this invention are: traditional dynamic programming methods are limited by the discretization of the state space and are prone to getting trapped in local optima under long distances and complex terrain. This invention adopts a policy network based on the Transformer architecture, which utilizes its self-attention mechanism to achieve global perception of the terrain profile, and can simultaneously handle complex constraints such as terrain undulations, obstacles, electrical safety, and mechanical strength.
[0025] The strategy network of this invention can identify macroscopic terrain features such as ridges and valleys, and adaptively adjust the span and tower height accordingly, generating a ranking scheme that better conforms to actual engineering logic. The strategy network trained by this invention can complete the decision-making for a single tower in a very short time during the inference phase. This enables designers to quickly generate multiple candidate schemes for large-scale transmission corridors and achieve scheme comparison within seconds through parallel computing, greatly improving design efficiency and quality. Attached Figure Description
[0026] Figure 1 This is a flowchart of the present invention; Figure 2 It is a two-dimensional cross-sectional schematic diagram along the centerline of the line; Figure 3 This is a graph showing the relationship between cumulative rewards and the number of training rounds during the training process of the course learning strategy in this embodiment of the invention; Figure 4 This is a graph showing the relationship between the success rate of pole ranking and the number of training rounds during the training process of the course learning strategy in this embodiment of the invention. Figure 5 This is a cross-sectional view of the tower arrangement results according to an embodiment of the present invention; Figure 6 This is the present invention. Figure 5 Detailed results of a local section; Figure 7 This is a CSV report of the results of the pole and tower placement in an embodiment of the present invention. Detailed Implementation
[0027] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0028] Example 1
[0029] This embodiment takes a transmission line project with a total length of approximately 40km and a voltage level of 500kV as an example. According to the "Design Code for 110kV~750kV Overhead Transmission Lines" (GB 50545-2010), the ground distance thresholds are set as follows: 14 meters for residential areas, 11 meters for non-residential areas, 8.5 meters for areas with difficult transportation, 9.0 meters for buildings, 7 meters for forested areas, 14 meters for railways, 6 meters for electrified railway contact networks, 14 meters for highways, 6 meters for crossing power lines, 6.5 meters for cableways, and 11.5 meters for navigable rivers. The tower type and the horizontal span LH, vertical span LV, and kV values are shown in Table 1. Table 1: Conditions for Using Pole Towers
[0030] like Figure 1 The specific implementation steps of the present invention are shown in detail below.
[0031] S1. Constructing the simulation environment
[0032] Digitalization of topographic profiles: See Figure 2 High-precision LiDAR point cloud data of the power transmission corridor was acquired and transformed into a two-dimensional profile sequence along the centerline of the transmission line. A set of two-dimensional profile points with a spacing of 20 meters was extracted. Each cross-section point Includes: Mileage (Unit: m) Elevation (Unit: m) and feature codes (e.g., 0-open space, 1-house, 2-crossing power lines, 3-water system). Feature vectorization is performed, and a sliding pane is used to extract terrain data within a 1500-meter radius before and after the current location to be ranked. Sine wave position coding is used to inject sequence information into each point, generating the input feature matrix. ,in This represents the number of terrain points within the window. For feature dimensions.
[0033] S2, Design Strategy Network
[0034] The policy network employs a Transformer-based architecture to receive the environmental state from the sliding pane and output the joint, continuous actions for the next N towers. In this embodiment, N is set to 2, meaning that each decision simultaneously predicts the span and tower height of the next two towers.
[0035] S201. State Space Construction: Extract the terrain profile features of the current area to be ranked from the simulation environment, including the point cloud height sequence, land feature classification code, and geospatial coordinates within the window.
[0036] S202, Feature Encoding: For each terrain point within the window, its elevation, feature encoding, and relative mileage are concatenated to form an original feature vector, which is then mapped to 256 dimensions through a linear projection layer. After adding sinusoidal position encoding, it is input into a 4-layer Transformer encoder (8 self-attention heads per layer, feedforward network dimension 1024). The encoder outputs a global terrain semantic feature matrix. Then, a fixed-length terrain feature vector is obtained through global average pooling. .
[0037] S203, Multimodal Fusion: Integrating terrain feature vectors With real-time scalar status (current cumulative mileage) The height of the previous tower The altitude of the previous tower Current elevation By splicing together, the fusion features are obtained. Then, input two layers of MLP (128-dimensional hidden layer, 2N-dimensional output layer), and output the action parameters for the subsequent N towers. Output a continuous multi-dimensional action vector, mapped to the span and tower height parameters of the subsequent N towers. For the i-th tower (i=1,2), output... ,in Let be the horizontal span between the i-th tower and the preceding tower. This is the nominal height of the base tower. The network output layer uses the tanh activation function and is then linearly transformed to a specified range. , .
[0038] Policy network parameters: 4 encoder layers, 8 heads, 256 hidden layer dimensions, approximately 1.2 million total parameters. The optimizer used is Adam, with a learning rate of 3e-4.
[0039] S3, Simulation Environment Modeling and Reward Mechanism Construction
[0040] S301. Construction of the physical simulation environment: Based on the actions output by the policy network. .
[0041] S302, Compliance Assessor Integration: The given sag coefficient in this embodiment is: The threshold values for horizontal span LH and vertical span LV are taken as the maximum values of the horizontal spans in the aforementioned tower type library. 860, the maximum value of vertical clearance : 1200; based on the previous two towers , Parameters ( , ) and the current proposed tower location The traverse height (traverse elevation at each cross-section point within the span), the current horizontal span distance LH, vertical span distance LV, and KV value of the tower are calculated using the parabolic equation.
[0042] , Gear spacing: , Gear spacing: , The traverse elevation of each cross-section point within the arch: , Relative The horizontal position of the location.
[0043] Horizontal spacing:
[0044] Vertical distance:
[0045] KV value:
[0046] in, cos , Tower positions Mileage, elevation of topographic points, and altitude.
[0047] Furthermore, the compliance logic is as follows: Compliance of conductor distance to ground: obtaining , Elevation of topographic points within the archive (Surface types include houses, trees, and obstacles; the elevation of topographic points is the sum of the ground elevation and the height of these types of features.) Traverse elevation for each cross-section point. ,statistics If the number of values is less than the ground distance threshold, and the number is 0, then the compliance of tower height, horizontal span LH, and vertical span LV is judged; otherwise, ... 0 Compliance assessment of horizontal clearance (LH) and vertical clearance (LV) values: based on the calculated values. , Determine whether the condition is met. , If all conditions are met, then The value is 1.
[0048] It should be noted that, regarding the tower ,because Not yet generated, in progress A compliance assessment will then be conducted.
[0049] S303. Design a multi-objective coupled reinforcement learning reward function: the total reward is a weighted sum of the components.
[0050] In this embodiment, the weights are set as follows (adjustable during the training phase): , , , , The specific calculation formulas and parameters for each component are as follows: (1) Progress increment reward:
[0051] For example, take , , The current step distance (m). This represents the absolute value (in meters) of the current elevation difference. For example, in gently sloping terrain ( When the span is 300m, .
[0052] (2) Terrain-inspired rewards:
[0053] For example, take , Centered on the current tower location, take a 30m window before and after it, and calculate the average elevation of the terrain points within the window. If the current point elevation If the top-3 highest points within the window are... Otherwise, it is 0. If the current point is located at the bottom of the valley (defined as: the height difference between the highest and lowest points within the window is >20m, and the current point is the lowest point), then Otherwise, it is 0. For example, erecting a tower on a mountain ridge: , ,but Erect a tower in the valley: .
[0054] (3) Tower Matching Rewards:
[0055] For example, take First, determine the horizontal distance of the current gear. Vertical spacing Using the KV value, the tower type usage condition table is consulted to map the appropriate tower type for this node. In this embodiment, the preset tower type library costs include: straight-line tower (cost 100,000 yuan), tension tower (cost 180,000 yuan), and heavy-duty tension tower (cost 250,000 yuan). The benchmark tower type is the straight-line tower, with a cost of... Ten thousand yuan. If the actual tower type currently mapped is a straight tower, then If it is a tension tower, then (Negative reward).
[0056] (4) Economic penalties:
[0057] For example, take , , . The tower height is (m). The cost of the tower (in ten thousand yuan). For example, a straight-line tower with a height of 30m: A heavy-duty tension tower with a height of 48m: .
[0058] (5) Compliance constraints and penalties:
[0059] For example, take If the evaluator outputs... 0, then ;like ,but .
[0060] Example of total reward: Suppose a certain action results in: , , , , The total reward .
[0061] S4. Training using curriculum-based learning strategies
[0062] S401, Parameter Initialization: Set the dynamic adjustment factor Used to control progress increment rewards Value (i.e., the driving weight factor). Initial stage At this time, the actual use That is, the initial .because When the progress reward is negative, it becomes negative, forcing the policy network to use short intervals because the absolute value of the negative reward generated by short intervals is smaller.
[0063] S402, Initialization Phase (Rounds 0-1000): Settings At this stage, the output span of the strategy network is concentrated between 80 and 200 meters, with minimal conductor sag, making it easy to pass compliance assessments. The goal at this stage is to ensure the strategy network adheres to the bottom line of compliance, increasing the success rate of pole and tower placement from the initial 60% to over 80%.
[0064] S403, Performance Evolution Stage (1000-4000 rounds): When the success rate of tower placement reaches 80%, linear interpolation is used to... The value was gradually increased from -0.3 to +0.2, that is... The reward gradually changes from -0.2 to +0.3. It increases by 0.05 every 100,000 steps. At this point, the progress reward becomes positive, and the strategy network begins to try to increase the step distance and actively identify mountain peaks in the terrain to obtain higher rewards.
[0065] S404, Economy-Driven Optimization Phase (After Round 4000): ... Fixed at +0.5 (i.e.) (and economic penalties) weight The value increases from 0.1 to 0.3. At this point, the strategy network learns to trade tower height for terrain: use low-rise towers (15-25m) on ridges and reasonably increase the span (300-500m) on plains to reduce the total number of towers.
[0066] Training algorithm: Proximal policy optimization (PPO) is employed, collecting 2048 experiences per step, and GAE is used to calculate the advantage (…). ), Cutting probability ratio The entropy regularization coefficient is 0.01. The training process consists of 10,000 rounds, with evaluation occurring every 100 rounds.
[0067] S405. Training convergence determination and model output: During training, the system performs a convergence test on the current policy every 100 rounds, and the test metrics include the following five items: (1) Compliance rate: On the independent validation set, the compliance parameters in the statistical strategy generation scheme are used. The proportion of the number of spans in the total number of spans. The convergence threshold is set to ≥98%.
[0068] (2) Cumulative return stability: Calculate the cumulative discounted return over the last 50 full ranked rounds. The moving average and standard deviation. This measures how much the policy network values future rewards, among which... This marks the start of the round (first base tower). The round ends (the end of the route is reached). For the first The immediate reward of the step (i.e. the total reward calculated by S303) has the following convergence criteria: over 50 consecutive rounds, the fluctuation range of the moving average reward is less than ±5% of the average of the most recent 100 rounds, and there is no obvious upward or downward trend.
[0069] (3) Policy update magnitude (KL divergence): In the PPO algorithm, the KL divergence of the action probability distribution before and after each policy update is calculated. The convergence criterion is: if the average KL divergence is <0.01 in 10 consecutive updates, it indicates that the policy parameters have basically stabilized.
[0070] (4) Stability of economic indicators: The standard deviations of the average number of poles per kilometer and the average tower height over the most recent 100 rounds were statistically analyzed. The convergence criteria were: the standard deviation of the number of poles per kilometer < 0.5 poles / km and the standard deviation of the average tower height < 2m.
[0071] (5) Reward Function Value Reference: Although the specific value of the reward function is not directly used as the convergence criterion, the total reward in this embodiment is determined when all four of the above indicators are satisfied. The cumulative reward is usually stable between 400 and 500 (positive value). If all convergence metrics are not met after training reaches the preset maximum number of rounds (10000), the checkpoint with the highest success rate in tower ranking on the validation set is taken as the final model.
[0072] When all five metrics are simultaneously satisfied, the policy network is considered converged, training is terminated, and the current policy network is output as the final model. This model can then be used for subsequent pole ranking inference.
[0073] Training dynamics such as Figure 3 and Figure 4 As shown in the figure, the curves depicting the changes in single-round reward and pole ranking success rate with the number of training rounds are plotted respectively. The two curves show a significant positive correlation: both indicators monotonically increase and eventually converge to a high-performance stable plateau. This dual effect of reward maximization directly translating into increased success rate indicates that it can effectively align the learning objective of the policy network with the expected operational results, thus confirming its practical utility and convergence stability.
[0074] S5. Post-processing and solution evaluation
[0075] S501, Inference Generation: After training, deploy the model to the production environment. See also... Figure 5 Starting from the beginning of the line, each time a terrain window of 1500m before and after the current point is input, the strategy network outputs the span and tower height of the subsequent N towers, updates the current position, and repeats until the entire line is covered. This embodiment generates a total of 96 towers. Figure 6 This is a detailed view of a part of the scene.
[0076] S502, Abnormal Short Gear Counting: Set abnormal short gear count threshold. m. Traverse all adjacent tower spacings. In this embodiment, no cases with a spacing less than 60m were found (the minimum span is 95m), so merging is unnecessary. If a cluster with an elevation below the threshold is found, retain the node with the highest elevation within the cluster and delete the rest.
[0077] S503, Span-by-Span Index Collection: For each span, record the range. Tower height Minimum distance to the ground (Calculated from sag), output of the compliance assessor .
[0078] S504, Performance Evaluation: Automatic positioning success rate: Number of qualified spans (minimum ground distance ≥ 6.5m and) Divide by the total base (96). In this embodiment, all 96 spans are qualified, with a success rate of 100%.
[0079] Table 2 presents the average decision-making time per tower and analyzes its distribution across key components.
[0080] Table 2: Decision-Making Time for Towers
[0081] Output report: See the CSV file for details. Figure 7 This includes tower number, mileage (m), altitude (m), tower height (m), tower type, horizontal span LH, vertical span LV, and KV value, and also generates a cross-sectional view.
[0082] The embodiments described herein are preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Therefore, all equivalent changes made in accordance with the structure, shape, and principle of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for tower ranking of power transmission line based on deep reinforcement learning, characterized in that, Includes the following steps: S1. Constructing a simulation environment: Based on the digital elevation model and ground feature classification data of the power transmission corridor, construct a simulation environment that supports sequential decision-making and continuous action space to simulate the tower placement process. S2. Design Strategy Network: The strategy network is used to receive the environmental state of the sliding pane and directly output the joint continuous action of the subsequent N towers through the neural network. The joint continuous action includes the continuous action of the span and tower height of each tower to be placed. S3. Simulation Environment Modeling and Reward Mechanism Construction: Design a multi-objective composite reward function to guide the policy network to evaluate the compliance and economy of tower placement during training. Penalize invalid placements with evaluation scores below the threshold, and set specific reward or penalty items for unfavorable terrain and tower type matching. S3 includes: S301. Physical simulation environment construction: Virtually place towers based on the actions output by the policy network; S302, Evaluation Unit Integration: Construct an independent compliance evaluation unit that outputs whether the current tower placement scheme meets the constraints of ground distance, horizontal span LH, and vertical span LV. When the constraints are met...
1. If Then, negative feedback is applied to the policy network through constraint penalty terms; S303. Design a multi-objective coupled reinforcement learning reward function, which includes at least several of the following: progress increment reward, terrain-inspired reward, tower matching reward, economic penalty, and compliance constraint penalty. The progress increment reward is defined as: in To advance the weighting factor, The elevation difference sensitivity coefficient, The horizontal distance of the current step forward, This is the absolute value of the height difference in the current gear. The terrain-inspired reward is defined as: in The high ground reward coefficient, The valley penalty coefficient, The current elevation. The average elevation within the window. This is a local high point indicator function. Valley indicator function; The tower-type matching reward is defined as follows: in Cost sensitivity coefficient The cost of the benchmark tower type, The actual tower construction cost mapped to the current action; The economic penalty is defined as follows: in For the height of the tower, For exponential coefficients, This is the cost coefficient for the tower type. and These are the height penalty weight and the tower cost penalty weight, respectively. The compliance constraint penalty is defined as follows: in The penalty intensity coefficient, The compliance parameters output by the compliance assessor; The reinforcement learning reward function is: in, , , , , For the corresponding weight S4. Training using a course-based learning strategy: In the early stages of training, the reward function parameters are adjusted to make the policy network tend to generate compact and conservative tower layouts in order to accumulate effective experience; then the reward parameters are gradually adjusted to guide the policy network to explore optimization schemes with larger spans and lower total costs while meeting compliance constraints, until the strategy converges. S5. Post-processing and scheme evaluation: Post-process the tower placement schemes generated by the trained policy network during the inference phase, collect and statistically analyze the span, tower height, tower type, minimum ground distance, compliance parameters and reward composition for each span, reproduce the tower placement process in an independent evaluation environment, calculate the overall success rate, average compliance parameters and performance indicators of the scheme, and output a structured report.
2. The method for arranging transmission line towers based on deep reinforcement learning as described in claim 1, characterized in that, In S2, the policy network adopts a Transformer-based architecture, specifically including: S201. State Space Construction: Extract the topographic profile features of the current area to be ranked, including high-order sequences, land cover classification codes, and geospatial coordinates; S202, Feature Encoding: Spatial sequence information is injected into terrain points using location encoding, and global terrain semantic features are extracted through the self-attention mechanism of the Transformer encoder; S203, Multimodal Fusion: The extracted terrain feature embedding vector is concatenated with the real-time scalar state vector and mapped to the action space through a multilayer perceptron; and the policy network is configured to output continuous multidimensional action vectors, which are respectively mapped to the span and tower height parameters of the subsequent N base towers.
3. The method for arranging transmission line towers based on deep reinforcement learning according to claim 1, characterized in that, In S4, the course learning strategies include: S401, Set a dynamic adjustment factor This is used to control the progress weighting factor in the progress increment reward; S402, Initialization Phase: Settings A negative value forces the policy network to learn safely over short intervals. S403, Performance Evolution Phase: When the automatic ranking success rate or training steps reach a preset threshold, Gradually increase from negative values to positive values; S404, Economy-Driven Global Policy Optimization: In the later stages of training, Set to a positive value and increase the weight of the economic penalty.
4. The method for arranging transmission line towers based on deep reinforcement learning according to claim 1, characterized in that, S5 includes: S501. In an independent evaluation environment, starting from the first tower of the line, the tower positions are generated one by one in an autoregressive manner. S502. Set an abnormally short span threshold, and classify consecutive tower nodes with adjacent spacing below the threshold into a redundant cluster. Only retain the tower node with the largest surface elevation in the cluster and remove the rest of the nodes. S503. Collect and record span, tower height, conductor-to-ground distance, and compliance parameters output by the compliance assessor for each span, and summarize them into structured records. S504 calculates the automatic ranking success rate, the inference time of a single compliance assessor, and the average decision time per tower, and outputs a structured report.
5. A transmission line tower positioning system based on deep reinforcement learning, characterized in that, A method for arranging transmission line towers based on deep reinforcement learning, used in any one of claims 1-4, comprises: The simulation environment construction module is used to build a simulation environment that supports sequential decision-making and continuous action space based on the digital elevation model and ground feature classification data of the power transmission corridor. The strategy network module is used to receive the environmental state of the sliding pane and determine the joint continuous action of the subsequent N base towers through the output of the neural network. The action includes the continuous action of the span and tower height of each tower to be placed. The reward mechanism construction module is used to design multi-objective composite reward functions, evaluate the compliance and economy of tower placement, impose penalties for ineffective placement, and set specific reward or penalty items for unfavorable terrain and tower type matching. The reward mechanism construction module includes: The physical simulation unit is used to virtually place towers and calculate conductor sag based on the actions output by the policy network. The compliance evaluator unit is an independent evaluator used to assess whether the current ranking scheme meets the requirements for ground distance, horizontal clearance, and vertical clearance constraints, and to provide negative feedback when they do not. A multi-objective reward function unit is used to calculate at least one of the following: progress increment reward, terrain-inspired reward, tower matching reward, economic penalty, and compliance constraint penalty; The course learning and training module is used to guide the policy network from a conservative layout to an economically optimal solution by dynamically adjusting the parameters of the reward function. The post-processing and evaluation module is used to post-process the tower placement schemes generated in the inference phase, calculate the overall success rate, average compliance parameters and performance indicators, and output a structured report.
6. A transmission line tower positioning system based on deep reinforcement learning according to claim 5, characterized in that, The policy network module adopts a Transformer-based architecture, including: State space construction unit is used to extract the terrain profile features of the current area to be ranked; The feature encoding unit is used to extract global terrain semantic features using the self-attention mechanism of position encoding and Transformer encoder; A multimodal fusion unit is used to concatenate terrain feature embedding vectors with real-time scalar states; Action space definition unit, used to output a continuous multidimensional action vector mapped to span and tower height parameters.
7. A transmission line tower positioning system based on deep reinforcement learning according to claim 5, characterized in that, The post-processing and evaluation module includes: The abnormally short span merging unit is used to cluster tower nodes with adjacent spacing below a threshold and retain only the node with the highest ground elevation. The span-by-span indicator collection unit is used to record span distance, tower height, distance to ground, compliance parameters, and reward weight; The performance statistics unit is used to calculate success rate, inference time, and decision-making time, and outputs a structured report.
Citation Information
Patent Citations
Method for intelligently ranking power transmission lines
CN110601065A
Pole tower ranking method, computer equipment and computer readable storage medium
CN119004842A