Power transmission line intelligent planning method based on deep reinforcement learning Ant-D3QN-PER algorithm
By combining ant colony algorithm and deep reinforcement learning algorithm, the Ant-D3QN-PER algorithm with variable resolution raster map and priority experience playback mechanism is used to solve the problem of large-scale maps and multi-constraint processing in transmission line planning in the existing technology, and efficient and low-cost intelligent planning of transmission line is achieved.
Patent Information
- Application Number
- CN202510380904.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
Smart Images

Figure CN120218554A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent planning of transmission lines, and in particular to an intelligent planning method for transmission lines based on the deep reinforcement learning Ant-D3QN-PER algorithm. Background Art
[0002] With the rapid development of technology and economy, the demand for electricity has increased sharply, and the scale of power grid construction has also been growing rapidly. As an important part of power grid construction, the route planning of transmission lines is related to the planning and construction of the entire power grid and has a profound impact on urban construction and residents' lives. Therefore, how to plan transmission lines efficiently and at low cost has become an important research topic in the power industry. Under the background of the digital power grid construction plan proposed by the State Grid, using computer-designed intelligent route selection algorithms has become an inevitable trend and a key means for transmission line planning.
[0003] In the past, traditional manual annotation methods were used for transmission line planning. After field exploration and research by operators in the planning area, combined with transmission line design specifications and construction experience, the routes where wires can be laid were marked on the corresponding map, or marked and designed based on CAD software. This traditional method requires a long construction period, high labor consumption, and is time-consuming and laborious. Using a computer to design an intelligent route selection algorithm based on remote sensing images can effectively solve the above-mentioned drawbacks.
[0004] The reinforcement learning algorithm is an artificial intelligence algorithm that is optimized by continuously interacting with the environment to find better solutions. For the problem of transmission line planning, the reinforcement learning algorithm designs a reward function according to the requirements and costs of transmission line planning to guide the agent to explore in the environment, so as to handle complex decision-making problems that are difficult to solve by traditional algorithms and is suitable for the path planning of transmission lines under multi-constraint scenarios. However, when using general deep reinforcement learning algorithms for transmission line planning, the map is rasterized, which leads to an increase in the training difficulty of the algorithm and a slowdown in the convergence speed as the map and the number of grids increase, and a large amount of computing resources are required. Therefore, how to use deep reinforcement learning to plan a route that meets the actual transmission line planning constraints and has the lowest total cost, and reduce the grid calculation involved in grid map planning, improve the algorithm operation efficiency, and achieve intelligent and efficient planning of transmission lines under multi-constraint conditions is a technical problem. Summary of the Invention
[0005] In order to overcome the defects in the above-mentioned prior art, the present invention provides an intelligent planning method for transmission lines based on the deep reinforcement learning Ant-D3QN-PER algorithm. Considering the situation where multiple constraint conditions coexist and affect each other in the transmission line planning problem, a route that meets the requirements of transmission line planning and has the optimal economic cost is planned, and high operating efficiency can also be achieved in a large-scale planning area.
[0006] An intelligent planning method for transmission lines based on the Ant-D3QN-PER algorithm of deep reinforcement learning, comprising the following steps:
[0007] S1. Screen out the evaluation indicators that affect the transmission line planning, establish an analytic hierarchy structure, combine and determine the weights of the evaluation indicators, calculate the comprehensive weight, and thus construct a cost evaluation model;
[0008] S2. Based on the remote sensing image, segment the rasterized map to obtain a variable-resolution raster map, calculate the cost value of each variable-resolution raster according to the cost evaluation model, and establish a variable-resolution raster map model;
[0009] S3. Improve the algorithm based on the variable-resolution raster map model, combine the ant colony algorithm ACS with the deep reinforcement learning algorithm DQN, and improve it with Double DQN and Dueling Network, and adopt the PER mechanism to obtain the Ant-D3QN-PER algorithm based on the variable-resolution raster map model.
[0010] Preferably, the specific steps of step S1 are as follows:
[0011] S11. Take the planning cost of the transmission line as the goal, analyze the influencing factors of the goal, and further screen out multiple evaluation indicators that affect the planning cost from the influencing factors;
[0012] S12. Build a hierarchical model, where the evaluation indicators correspond to the index layer at the top layer, the influencing factors correspond to the factor layer in the middle layer, and the planning cost corresponds to the goal layer at the bottom layer;
[0013] S13. Solve the comprehensive weight of all evaluation indicators on the planning cost in the hierarchical model based on the fuzzy analytic hierarchy process FAHP.
[0014] Preferably, step S13 specifically includes the following steps:
[0015] S131. Establish a fuzzy complementary matrix A=(a ij ) n×n , i, j = 1, 2,..., n, where n represents the order of the matrix, and the matrix is expressed as follows:
[0016]
[0017] S132. Represent the relative importance degree between two indicators under a certain factor by taking values in the 0.1-0.9 scale method;
[0018] S133. Sum each row in A, and then transform it according to the following formula:
[0019]
[0020] where a i and a j represent the sums of the elements in the i-th row and the j-th row of A respectively, and the transformed fuzzy consistent matrix is obtained
[0021] S134. Calculate the weight of element i corresponding to the upper-layer element h through the following formula:
[0022]
[0023] where represents the sum of the elements in the i-th row of A ~ , and the weights of all elements in the lower layer corresponding to element h are
[0024] S135. Suppose w = [w1, w2,..., w n is the weight of matrix A ~ , and w ij = w i / (w i + w j ), then there is:
[0025] w′ = (w ij ) n×n
[0026] where w' is the eigenmatrix of A ~ ;
[0027] The calculation method of the consistency index is:
[0028]
[0029] If w' and A ~ satisfy I(A ~ , w') ≤ b, then A ~ satisfies the consistency; where b is a constant;
[0030] S136. Set: The weights of transportation conditions and traffic routes in the index layer for the engineering construction factor in the upper-layer factor layer are The weights of temperature, water area, forest land, and hard ground surface in the index layer for the environmental factor in the upper-layer factor layer are The weights of line length and maintenance in the index layer for the conductor cost factor in the upper-layer factor layer are The weights of nature reserves, houses, and industrial land in the index layer for the social impact factor in the upper-layer factor layer are The weights of all factors in the factor layer with respect to the target layer are \(w\). G ;
[0031] The comprehensive weight \(w\) of all evaluation indicators corresponding to the planned cost of the target layer O is calculated in the following manner:
[0032]
[0033] Preferably, step S2 specifically includes the following steps:
[0034] S21, After semantic segmentation of the remote sensing image, a ground object recognition map is obtained. After gray-scale processing of the ground object recognition map, a gray-scale map is obtained, and the gray-scale map is rasterized to obtain an initial raster map;
[0035] S22, Perform quadtree segmentation on each initial raster in the initial raster map in sequence. After an initial raster is segmented, move to the next initial raster to continue the segmentation operation until all initial rasters in the initial raster map have completed quadtree segmentation, forming a variable-resolution raster map;
[0036] S23, Improve the single-node neighborhood structure of the raster into an adaptive node neighborhood structure, and the number of nodes within the raster changes adaptively according to the raster size; if the size of the raster is one or two times the minimum raster size \(S\) min then there is only 1 node in this raster; if the multiple is four or eight times, then there are 4 nodes in this raster; if the size of the raster is the same as the initial raster, then there are 16 nodes in this raster; among them, there are at most 16 nodes within the raster;
[0037] S24, The raster neighborhood structure adopts an adaptive node neighborhood structure. The route cost i from raster node \(u\) j to neighborhood raster node \(b\) is calculated as follows:
[0038]
[0039] where \(v\) u and \(v\) b are the costs of raster \(u\) and raster \(b\), and are the lengths of the line connecting the two points of raster nodes \(u\) i and \(b\) j in rasters \(u\) and \(b\) respectively.
[0040] Preferably, step S22 is specifically as follows:
[0041] Before performing quadtree segmentation on a grid, it is necessary to make judgments according to two segmentation criteria. The segmentation criteria include: (a) First, judge the size of the grid to be segmented. If its size is greater than the pre-set minimum grid size S min , then enter the second segmentation criterion for determination; otherwise, do not perform segmentation judgment on the grid and turn to the next grid to be segmented for determination; (b) Judge the variance of the gray values of the grid to be segmented. If the variance of the gray values of all pixel points in the grid is greater than the set value W f , then perform quadtree segmentation; otherwise, pre-segment the grid into four pre-segmented sub-grids, calculate the average value of the gray variances of the four pre-segmented sub-grids. If the average value of the gray variances of the four pre-segmented sub-grids is greater than the set value W s , then perform actual quadtree segmentation on the grid to be segmented according to the pre-segmentation method; otherwise, no longer perform segmentation judgment on the grid;
[0042] When judging the quadtree segmentation criterion of the grid, it is necessary to calculate the variance σ of the gray values of the grid 2 , and the calculation formula is as follows:
[0043]
[0044] Among them, (r, c) is the coordinate of the pixel point belonging to the grid, Gr(r, c) is the gray value of the pixel point, and the number of pixel points in the grid is denoted as S is the average gray value of the grid;
[0045] Calculating the average value of the gray variances of the four pre-segmented sub-grids is to sum the σ of the four pre-segmented grids 2 and take the average value.
[0046] Preferably, step S3 specifically includes the following steps:
[0047] S31, The node transfer probability p of the improved ACS algorithm based on variable-resolution grids m (u i ,b j ) is:
[0048]
[0049] Among them, p m (u i ,b j ) represents the probability that ant m transfers from node i in grid u to node j in grid b, τ(u, b; φ now ) is the pheromone concentration on the path from grid u to grid b, φ now is the current network parameter, is the shared θ now convolutional network parameter, and They are the parameters of the advantage function network and the state value function network in the Dueling Network, η(u i , b j ) is the unit cost of the route between two grids u and b , the reciprocal of which is the guiding coefficient represents the grid node u i and the grid node b j The included angle between the connection line of the grid nodes and the connection line from u i to the end grid, α, β, are the control coefficients respectively, and Ω m (u) is the other grids that the ant m can reach when it is at grid u except for the grids in the taboo list, and n ~ represents the number of grids that can be selected currently;
[0050] S32, and find the next grid node b according to the following pseudo-random proportion rule j :
[0051]
[0052] Among them, q0 is a constant with a value in [0, 1], q is a random number uniformly distributed in [0, 1], S is the grid node that can be reached by probability selection according to the transition probability, and the ant adds the current grid to the taboo list every time it takes a step;
[0053] S33, using the formulas in steps S31 and S32, the ant starts to search for the route from the starting point to the end point, and for the selected grid node b j , judge whether the connection line of u i -b j satisfies all constraints. If not, exclude bj and repeat the above process to select from other feasible grid nodes; if the constraints are satisfied, the ant moves to the selected grid node b j and takes it as the current position, adding the grid b to the taboo list; until each ant has completed the route search, record the route trajectory of each ant, and calculate Δτ(u, b) obtained by each step of action of the ant according to each section of the route trajectory, where Δτ(u, b) = 1 / η(u i , b j ), record M trajectories u1, b1, Δτ1, u2,..., u n , b n , Δτ n , u n+1 and the route cost of each route;
[0054] S34, based on the PER mechanism, for each trajectory u1, b1, Δτ1, u2,..., u of M antsn , b n , Δτ n , u n+1 Divided into n quadruples (u i , b i , Δτ i , u i+1 ), which are stored in the experience replay pool as experience. Calculate the TD error δ according to the following formula i And the priority pr i :
[0055]
[0056] pr i = |δ i | + 1
[0057] Where γ is the discount factor, taking values in [0, 1], e is the grid that can be selected after the next grid b, and 1 is a positive constant close to zero; Store δ i And pr i Into the experience replay pool together. If the number of experiences in the experience replay pool is full, then according to the order of experience storage, the first M·n quadruples stored are removed, and then the new experience of the ant is stored;
[0058] S35. According to the priorities of all quadruples in the experience replay pool, extract K quadruples with probabilities determined by the priorities; Perform forward propagation on the network and obtain τ(u k , b k ; φ now ) according to the following formula:
[0059]
[0060] Where And Are the output values of the action advantage branch and the state value branch in the Dueling Network respectively, representing the advantages and disadvantages of the ant's action in the current state compared to the average action situation and the ability of the ant to obtain pheromones in the current state. The subscript k represents the kth of the K quadruples extracted;
[0061] S36. Select the optimal action that maximizes the pheromone:
[0062]
[0063] According to the optimal action And the target network, calculate the TD target y k :
[0064]
[0065] Among them, is the current network parameter of the target network in Double DQN. Calculate the TD error δ of each quadruple by the following formula k :
[0066] δ k = τ(u k , b k ; φ now ) - y k
[0067] S37. Obtain the weights ξ of K quadruples according to the following formula k :
[0068]
[0069] Among them, N is the total number of experiences in the experience pool, and λ is a control coefficient with a value in [0, 1]. is the maximum value among K , which plays a role in normalization. Then calculate the weighted loss function according to the following formula:
[0070]
[0071] Update the parameters of the estimation network using gradient descent and θ now :
[0072]
[0073] Thus, update the estimation network parameters to Steps S35 to S37 are looped te times, that is, updating the estimation network parameters te times is regarded as one round of update; then update the target network parameters according to the following formula to
[0074]
[0075] Update the target network parameters once every tg rounds of updating the estimation network parameters;
[0076] S38. Optimize the turning points of the optimal route in this iteration and record it. The grids passed by the optimized route become the trajectory of the optimal route in this iteration u1, b1, Δτ1, u2,..., u n , b n , Δτ n , u n+1 , and divide it into n quadruples (u i , b i , Δτ i , ui+1 ) and randomly extract K quadruples from them, and use φ new and to calculate the loss function and update the network parameters to achieve the effect of updating the pheromone. After every tg iterations, the global optimal route is used for global update; update the guiding factor according to the following formula
[0077]
[0078] wherein represents the control factor required at the (t + 1)-th iteration, and μ is a constant with a value in (0, 1);
[0079] S39. Steps S31 to S38 are one iteration process. If the maximum number of iterations T is reached, the above steps are ended and the global optimal route is output. The global optimal route is obtained by optimizing the corner points, otherwise the iteration process continues.
[0080] Preferably, step S33 is specifically as follows:
[0081] The design and construction of the transmission line shall be carried out in accordance with the relevant transmission line design specifications, and the specifications are introduced as constraints into the route selection algorithm. The specific constraint conditions are as follows:
[0082] (a) When the transmission line crosses an existing line, the tower head distance between the planned transmission line and the existing line is greater than 13 m;
[0083] (b) When the transmission line crosses an existing line, the crossing angle between the planned transmission line and the existing line is greater than 15°, and the crossing angle with the road is greater than 45°;
[0084] (c) The distance between the transmission line and the building is not less than 6 m;
[0085] Add the above constraints to the state transition rule when the ant conducts the search. When the ant determines the next route for each grid search, the above constraints must be satisfied.
[0086] Preferably, step S38 is specifically as follows:
[0087] After tg iterations, select the route with the optimal route cost as the global optimal route. Randomly extract K quadruples from this trajectory, and use φ new to calculate the loss function and update the network parameters again to obtain a new φ new , and complete the global update of this time.
[0088] The present invention also provides a computer program product, which includes a computer program / instructions. When the computer program / instructions are executed by a processor, the intelligent planning method for transmission lines based on the deep reinforcement learning Ant-D3QN-PER algorithm described above is implemented.
[0089] The present invention also provides an electronic device, which is characterized in that it includes a processor, a memory, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the intelligent planning method for transmission lines based on the deep reinforcement learning Ant-D3QN-PER algorithm.
[0090] The advantages of the present invention are as follows:
[0091] (1) The present invention establishes a cost evaluation model for transmission lines, calculates the comprehensive weights of all evaluation indicators in the hierarchical model for the planning cost, and quantifies the evaluation indicators affecting transmission lines by this method, ensuring the rationality and scientificity of weight assignment;
[0092] (2) The present invention proposes a variable-resolution grid map model based on remote sensing images, enabling the grids on the map to accurately fit the ground object contours, reducing redundant grids, and reducing the amount of data that needs to be calculated by subsequent route selection algorithms; the grid neighborhood structure is improved, and an adaptive node neighborhood structure is constructed, increasing the path search direction without reducing the route selection efficiency;
[0093] (3) The Ant-D3QN-PER algorithm based on the resolution grid map proposed by the present invention can plan transmission lines that meet the construction constraint conditions of transmission lines and have low economic costs, and also has high operating efficiency in a large-scale planning area and is not easily trapped in local optimal solutions. Description of the Drawings
[0094] Figure 1 It is a schematic diagram of the overall steps of the present invention.
[0095] Figure 2 It is the evaluation index hierarchical model of the transmission line of the present invention.
[0096] Figure 3 It is the solution process of the weight of the transmission line evaluation index based on FAHP of the present invention.
[0097] Figure 4 It is the rasterization schematic diagram of the remote sensing image of the present invention.
[0098] Figure 5 It is the quadtree segmentation process of the initial grid of the present invention.
[0099] Figure 6 It is the quadtree segmentation process of the present invention.
[0100] Figure 7 It is the comparison diagram of the single-size grid and the variable-resolution grid of the present invention.
[0101] Figure 8 is the adaptive node neighborhood structure of the present invention.
[0102] Figure 9 is the corner point optimization process of the present invention.
[0103] Figure 10 is the overall structure diagram of the Ant-D3QN-PER algorithm in the present invention.
[0104] Figure 11 is the overall flowchart of the Ant-D3QN-PER algorithm in the present invention. Specific embodiments
[0105] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0106] As Figure 1 shown, a transmission line intelligent planning method based on the deep reinforcement learning Ant-D3QN-PER algorithm of the present invention has the following specific implementation steps:
[0107] S1. Screen out the evaluation indicators that affect the transmission line planning, establish an analytic hierarchy structure, combine and determine the weights of the evaluation indicators, calculate the comprehensive weight, and thus construct a cost evaluation model.
[0108] S2. Based on the remote sensing image, segment the rasterized map to obtain a variable-resolution raster map, calculate the cost value of each variable-resolution raster according to the cost evaluation model, and thus establish a variable-resolution raster map model.
[0109] S3. Improve the algorithm based on the variable-resolution raster map model, combine the ant colony algorithm (Ant Colony System, ACS) with the deep reinforcement learning algorithm (Deep Q Network, DQN) to make it networked, introduce Double DQN and the dueling network (Dueling Network) to alleviate the problems of overestimation of values and transmission bias in the algorithm, and enhance the estimation ability of the network to form the Ant-D3QN algorithm. Finally, adopt the prioritized experience replay (Prioritized Experience Replay, PER) mechanism to obtain the Ant-D3QN-PER algorithm based on the variable-resolution raster map model.
[0110] The specific steps of step S1 are as follows:
[0111] S11. Take the planning cost of the transmission line as the goal, divide it into four categories of influencing factors, and then screen out multiple evaluation indicators that have a direct impact on the planning cost from these four categories of influencing factors.
[0112] S12. Construct the evaluation indicators into a hierarchical model as shown in Figure 2 . The evaluation indicators correspond to the index layer in Figure 2 , the four categories of influencing factors correspond to the factor layer, and the planning cost is the goal layer.
[0113] S13. Based on the fuzzy analytic hierarchy process (FAHP), solve the comprehensive weights of all evaluation indicators in the hierarchical model on the planning cost.
[0114] Step S13 specifically includes the following steps:
[0115] S131. Establish a fuzzy complementary matrix A=(a ij ) n×n , (i, j = 1, 2,..., n), where n represents the order of the matrix, and the matrix is expressed as follows
[0116]
[0117] S132. Use the values in the 0.1 - 0.9 scale method to represent the relative importance degree between two indicators under a certain factor, as shown in Table 1:
[0118] Table 1 0.1 - 0.9 scale method and its description
[0119]
[0120] The elements in A satisfy a ij + a ji = 1, and a ii = 0.5. a ij ∈ [0.1, 0.5) indicates that element a ij is more important than a ji . If a ij ∈ (0.5, 0.9], then vice versa.
[0121] S133. Sum each row in A, and then transform it according to the following formula:
[0122]
[0123] where a i and a j respectively represent the sums of the elements in the i-th row and the j-th row in A, and obtain the transformed fuzzy consistent matrix
[0124] S134, calculate the weight of the upper-layer element h corresponding to the element i through the following formula:
[0125]
[0126] where represents the sum of the elements in the i-th row of A ~ and the weights of all elements in the lower layer corresponding to the element h are
[0127] S135, let w = [w1, w2,..., w n be the weight of the matrix A ~ , and w ij = w i / (w i + w j ), (i, j = 1, 2,..., n), then there is:
[0128] w' = (w ij ) n×n
[0129] where w' is the eigenmatrix of A ~ .
[0130] The following is the calculation method of the consistency index:
[0131]
[0132] If w' and A ~ satisfy I(A ~ , w') ≤ b (where b is a constant usually taken as b = 0.1), then A ~ satisfies the consistency.
[0133] S136, set: The weights of transportation conditions and traffic routes in the index layer for the engineering construction factors in the upper-layer factor layer are The weights of temperature, water area, forest land, and hard ground surface in the index layer for the environmental factors in the upper-layer factor layer are The weights of line length and maintenance in the index layer for the conductor cost factors in the upper-layer factor layer are The weights of nature reserves, houses, and industrial land in the index layer for the social impact factors in the upper-layer factor layer are The weights of all factors in the factor layer for the target layer are w G .
[0134] The comprehensive weight w O of all evaluation indicators corresponding to the planning cost of the target layer is calculated in the following way:
[0135]
[0136] The process of solving weights based on FAHP is as follows Figure 3 shown.
[0137] Step S2 specifically includes the following steps:
[0138] S21, After semantic segmentation of the remote sensing image, an object recognition map is obtained. The object recognition map is grayscale processed to obtain a grayscale map, and the grayscale map is rasterized to obtain an initial raster map, specifically as Figure 4 shown.
[0139] S22, Perform quadtree segmentation on each initial raster in the initial raster map in sequence. After an initial raster is segmented, turn to the next initial raster to continue the segmentation operation until all the initial rasters in the initial raster map have completed quadtree segmentation, forming a variable-resolution raster map.
[0140] S23, Improve the single-node neighborhood structure of the raster into an adaptive node neighborhood structure, and the number of nodes in the raster changes adaptively according to the raster size; if the size of the raster is one or two times the minimum raster size S min of, then there is only 1 node in this raster; if the multiple is four or eight times, then there are 4 nodes in this raster; if the size of the raster is the same as the initial raster, then there are 16 nodes in this raster; among them, there are at most 16 nodes in the raster. The adaptive node neighborhood structure is as Figure 8 shown.
[0141] S24, The raster neighborhood structure adopts an adaptive node neighborhood structure. The route cost i from raster node u j to neighborhood raster node b is calculated as follows:
[0142]
[0143] where v u and v b are the costs of raster u and raster b, and are the lengths of the line connecting the two points of raster nodes u i and b j in rasters u and b respectively.
[0144] Step S22 specifically includes the following steps:
[0145] S221, Before performing quadtree segmentation on a raster, it is necessary to judge according to two segmentation criteria. The segmentation criteria include: (a) First, judge the size of the raster to be segmented. If its size is greater than the preset minimum raster size S min, then enter the second segmentation criterion for determination; otherwise, do not perform segmentation judgment on the grid and move to the next grid to be segmented for determination; (b) Judge the variance of the gray values of the grid to be segmented. If the variance of the gray values of all pixel points in the grid is greater than the set value W f , then perform quadtree segmentation; otherwise, pre-segment the grid into four pre-segmented sub-grids, calculate the average value of the gray variances of the four pre-segmented sub-grids. If the average value of the gray variances of the four pre-segmented sub-grids is greater than the set value W s , then perform actual quadtree segmentation on the grid to be segmented according to the pre-segmentation method; otherwise, no longer perform segmentation judgment on the grid.
[0146] When judging the quadtree segmentation criterion for the grid, it is necessary to calculate the variance σ of the gray values of the grid 2 , and the calculation formula is as follows:
[0147]
[0148] Among them, (r, c) is the coordinate of the pixel point belonging to the grid, Gr(r, c) is the gray value of the pixel point, and the number of pixel points in the grid is denoted as S, is the average gray value of the grid.
[0149] Calculating the average value of the gray variances of the four pre-segmented sub-grids is to sum and take the average of the σ of the four pre-segmented grids 2 .
[0150] S222. If the currently processed is the initial grid, then according to the arrangement order of the initial grids in the initial grid map, select one of the initial grids that have not been judged by the segmentation criterion as the grid to be segmented. If the currently processed is a sub-grid, then sequentially select the sub-grid that has not been judged by the segmentation criterion as the grid to be segmented. Specifically as Figure 5 shown.
[0151] S223. Calculate the variance of the gray values in the grid. If the variance is greater than W f , then enter step S224; otherwise, pre-segment the grid into four pre-segmented sub-grids and enter step S225.
[0152] S224. Perform quadtree segmentation on the current grid to generate four new-generation sub-grids. These four new-generation sub-grids become the grids to be segmented, and sequentially serve as the grids to be segmented in the order from left to right and from top to bottom, and enter step S222.
[0153] S225. Calculate the average value of the gray variances of the four pre-segmented sub-grids after pre-segmentation. If it is greater than W s , then enter step S224; otherwise, end the segmentation operation of this grid, return to the previous generation, and enter step S222. The overall steps are as Figure 6As shown, the obtained results are compared as Figure 7 shown below.
[0154] Specifically, step S3 includes the following steps:
[0155] S31. The node transfer probability p m (u i , b j ) of the improved ACS algorithm based on variable-resolution grids is:
[0156]
[0157] where p m (u i , b j ) represents the probability that ant m transfers from node i in grid u to node j in grid b. τ(u, b; φ now ) is the pheromone concentration on the path from grid u to grid b, φ now is the current network parameter, is the shared θ now convolutional network parameter, and are the parameters of the advantage function network and the state value function network in the Dueling Network respectively. η(u i , b j ) is the reciprocal of the unit cost of the route between two grids u and b. The guiding coefficient represents the angle between the connection line of grid nodes u and grid node b i and the connection line from u j to the end grid. α, β, i are the control coefficients respectively. Ω (u) is the other grids that ant m can reach when in grid u excluding the grids in the taboo list. n m represents the number of grids that can be selected currently. ~
[0158] S32. And the next grid node b is obtained according to the following pseudo-random proportional rule j :
[0159]
[0160] where q0 is a constant with a value in [0, 1], q is a random number uniformly distributed in [0, 1], S is the grid node that can be reached by probability selection according to the transfer probability, and the ant adds the current grid to the taboo list every time it takes a step.
[0161] S33, using the formulas of steps S31 and S32, the ant starts searching for a route from the starting point to the end point, and selects the grid node b j , judge u i -b j Does the connection satisfy all constraints? If not, exclude b j Repeat the above process and select from other feasible grid nodes; if the constraints are met, the ant moves to the selected grid node b j As the current position, grid b is added to the taboo table. Until each ant has completed the route search, record the route trajectory of each ant, and calculate the Δτ(u, b) obtained by the ant performing each step according to each segment of the route. Δτ(u, b) = 1 / η(u i ,b j ), record M trajectories u1,b1,Δτ1,u2,…,u n ,b n ,Δτ n ,u n+1 and the route cost for each route.
[0162] S34, based on the PER mechanism, each trajectory u1,b1,Δτ1,u2,…,u n ,b n ,Δτ n ,u n+1 Divide into n quadruplets (u i ,b i ,Δτ i ,u i+1 ) is stored in the experience replay pool as experience, and the TD error δ is calculated according to the following formula i With priority pr i :
[0163]
[0164] pr i =|δ i |+1
[0165] Where γ is the discount coefficient, which takes values in [0,1], e is the grid that can be selected after the next grid b, and 1 is a positive constant close to zero. i With pr i They are stored together in the experience replay pool. If the experience in the experience replay pool is full, the first M n quadruplets stored are removed according to the order in which the experience is stored, and then the new experience of the ant is stored.
[0166] S35, according to the priority of all the four-tuples in the experience playback pool, K four-tuples are extracted according to the probability determined by the priority; forward propagation is performed on the network, and τ(uk , b k ; φ now ):
[0167]
[0168] where and are the output values of the action advantage branch and the state value branch in the Dueling Network, respectively, representing the degree of superiority or inferiority of the ant's action in the current state compared to the average action situation and the ability of the ant to obtain pheromones in the current state. The subscript k represents the k-th of the K quadruples extracted.
[0169] S36, select the optimal action that maximizes the pheromone:
[0170]
[0171] According to the optimal action and the target network, calculate the TD target y k :
[0172]
[0173] where is the current network parameter of the target network in Double DQN. Calculate the TD error δ of each quadruple by the following formula k :
[0174] δ k = τ(u k , b k ; φ now ) - y k
[0175] S37, obtain the weights ξ of the K quadruples according to the following formula k :
[0176]
[0177] where N is the total number of experiences in the experience pool, λ is a control coefficient with a value in [0, 1], is the maximum value among the K , which plays a role in normalization. Then, calculate the weighted loss function according to the following formula:
[0178]
[0179] Update the parameters of the estimation network and θ now :
[0180]
[0181] Thereby update the estimated network parameters For Steps S35 to S37 are looped te times, that is, the update of the estimated network parameters is performed te times as one round of update; then update the target network parameters according to the following formula For
[0182]
[0183] Update the target network parameters once after every tg rounds of updating the estimated network parameters.
[0184] S38. Optimize the corner points of the iterative optimal route in this iteration and record it. The grids passed by the optimized route become the iterative optimal route trajectory u1,b1,Δτ1,u2,…,u n ,b n ,Δτ n ,u n+1 , and divide it into n quadruples (u i ,b i ,Δτ i ,u i+1 ). Randomly extract K quadruples from them, and use φ new and to find the loss function and update the network parameters so as to achieve the effect of updating the pheromone. After every tg iterations, perform global update using the globally optimal route. Update the guiding factor according to the following formula
[0185]
[0186] where represents the control factor required in the (t + 1)-th iteration, and μ is a constant with a value in (0,1);
[0187] S39. Steps S31 to S38 are one iteration process. If the maximum number of iterations T is reached, end the above steps and output the globally optimal route, which is obtained by optimizing the corner points, otherwise continue the iteration process. The Ant-D3QN-PER algorithm structure is as Figure 10 shown Figure 11 showing the overall algorithm flow.
[0188] Step S33 is specifically as follows:[[]]
[0189] The design and construction of the transmission line shall be carried out in accordance with the relevant transmission line design specifications, and the specifications are introduced as constraints into the route selection algorithm. The specific constraint conditions are as follows:[[]]
[0190] (a) When the transmission line crosses an existing line, the tower head distance between the planned transmission line and the existing line is greater than 13 m;
[0191] (b) When the transmission line crosses an existing line, the crossing angle between the planned transmission line and the existing line is greater than 15°, and the crossing angle with the road is greater than 45°;
[0192] (c) The distance between the transmission line and the building is not less than 6 m;
[0193] Add the above constraints to the state transition rule when the ant conducts a search. When the ant determines the next route for each grid search, the above constraints must be satisfied.
[0194] Step S38 specifically includes the following steps:
[0195] S381. The corner point optimization process is as Figure 9 shown. Figure 9 In it, the gray grids represent the non-crossable areas in the line planning. Set the corner point weight as cr. Starting from the starting point of the route, each time judge whether three corner points can be optimized, and advance towards the end point in turn until the end point is judged. First, judge whether the line connected by points z1, z2, and z3 can be optimized. Points z1, z2, and z3 are collinear and no corner is generated. However, the line z1 - z3 has one less node than the line z1 - z2 - z3, and one less tower pole is required. So the line z1 - z3 has a lower cost, and remove node z2. Then judge whether the line z1 - z3 - z4 can be optimized. Node z3 is a corner point, and the cost of the line z1 - z3 - z4 is the sum of z1 - z3 and z3 - z4 plus cr. If the cost of the line z1 - z4 is lower than that of the line z1 - z3 - z4, then remove node z3 and connect nodes z1 and z4. When judging the line z1 - z7 - z8, node z7 is a corner point, but the line z1 - z8 crosses the non-crossable area, so node z7 cannot be removed. Then judge the line z7 - z8 - z9 until the end point z 10 .
[0196] S382. After tg iterations, select the route with the optimal route cost as the global optimal route. Randomly extract K quadruples from this trajectory, and use φ new to calculate the loss function and update the network parameters again to obtain φ new , and complete the global update for this time.
[0197] The above are only the preferred embodiments of the present invention, and are not used to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A transmission line intelligent planning method based on deep reinforcement learning Ant-D3QN-PER algorithm, characterized in that: The following steps are involved: S1, screen out the evaluation indicators that have an impact on transmission line planning, establish a hierarchical analysis structure, combine and confirm the weights of the evaluation indicators, calculate the comprehensive weights, and construct a cost evaluation model; S2, based on the remote sensing image, the rasterized map is segmented to obtain a variable resolution raster map, the cost value of each variable resolution raster is calculated according to the cost evaluation model, and a variable resolution raster map model is established; S3, based on the variable resolution grid map model, the algorithm is improved by combining the ant colony algorithm ACS with the deep reinforcement learning algorithm DQN, and improved by Double DQN and Dueling Network. The PER mechanism is adopted to obtain the Ant-D3QN-PER algorithm based on the variable resolution grid map model.
2. According to claim 1, a transmission line intelligent planning method based on deep reinforcement learning Ant-D3QN-PER algorithm is characterized in that: The specific steps of step S1 are as follows: S11, taking the planning cost of the transmission line as the target, analyzing the influencing factors of the target, and then selecting multiple evaluation indicators that have an impact on the planning cost from the influencing factors; The transportation conditions and traffic routes in the indicator layer correspond to the engineering construction factors in the upper factor layer; Temperature, water area, forest land and hard surface in the indicator layer correspond to environmental factors in the previous factor layer; Line length and maintenance in the indicator layer correspond to the conductor cost factor in the previous factor layer; The nature reserves, houses and industrial land in the indicator layer correspond to the social impact factors in the previous factor layer; S12, a hierarchical model is built, where the evaluation index corresponds to the top index layer, the influencing factors correspond to the middle factor layer, and the planning cost corresponds to the bottom target layer; S13, based on the fuzzy analytic hierarchy process FAHP, the comprehensive weights of all evaluation indicators in the hierarchical model for planning costs are solved.
3. According to claim 2, a transmission line intelligent planning method based on deep reinforcement learning Ant-D3QN-PER algorithm is characterized in that: Step S13 specifically includes the following steps: S131, by comparing the relative importance of a certain element in the hierarchical model to each element in the next layer, a fuzzy complementary matrix A is established. ij ) n×n , i, j = 1, 2, ..., n, where n represents the order of the matrix, and the matrix is represented as follows: S132, the relative importance between two indicators under a certain factor is expressed by the value of 0.1 to 0.9 scaling method; S133, sum each row in A, and then transform it according to the following formula: Among them, a i and a j Respectively represent the sum of the elements in the i-th row and the j-th row in A, and obtain the transformed fuzzy consistent matrix S134, calculate the weight of element i corresponding to element h in the previous layer by the following formula: in, Indicates A ~ The sum of the elements in the i-th row, the weights of all elements in the next layer corresponding to element h are S135, let w=[w1,w2,...,w n ] is the matrix A ~ The weight of w ij =w i / (w i +w j ), then: in′=(in ij ) n×n Among them, w' is A ~ The feature matrix of The calculation method of consistency index is: If w' and A ~ Satisfy I(A ~ ,w')≤b, then A ~ Satisfy consistency; where b is a constant; S136, set: the weight of transportation conditions and traffic routes in the indicator layer to the engineering construction factors in the upper factor layer is The weights of temperature, water area, forest land and hard surface in the index layer to the environmental factors in the upper factor layer are: The weights of line length and maintenance in the indicator layer on the wire cost factor in the previous factor layer are: The weights of nature reserves, houses and industrial land in the indicator layer to the social factors in the upper factor layer are: The weight of all factors in the factor layer to the target layer is wG; All evaluation indicators correspond to the comprehensive weight w of the planning cost of the target layer O It is calculated by:
4. According to claim 1, a transmission line intelligent planning method based on deep reinforcement learning Ant-D3QN-PER algorithm is characterized in that: Step S2 specifically includes the following steps: S21, obtaining a land object recognition map after semantic segmentation of the remote sensing image, grayscale processing of the land object recognition map to obtain a grayscale map, and rasterizing the grayscale map to obtain an initial raster map; S22, performing quadtree segmentation on each initial grid in the initial grid map in turn, and when the segmentation of one initial grid is completed, turning to the next initial grid to continue the segmentation operation until all the initial grids in the initial grid map have completed the quadtree segmentation, thereby forming a variable resolution grid map; S23, the single-node neighborhood structure of the grid is improved into an adaptive node neighborhood structure, and the number of nodes in the grid changes adaptively according to the grid size; if the grid size is the minimum grid size S min If the multiple is one or two times, there is only one node in the grid; if the multiple is four or eight times, there are four nodes in the grid; if the size of the grid is the same as the initial grid, there are 16 nodes in the grid; there are at most 16 nodes in the grid; S24, the grid neighborhood structure adopts an adaptive node neighborhood structure, from the grid node u i To the neighboring grid node b j Route cost The calculation is as follows: Among them, v u and v b is the cost of grid u and grid b, and is the grid node u i and b j The length of the line connecting the two points in grids u and b respectively.
5. According to claim 4, a method for intelligent planning of power transmission lines based on deep reinforcement learning Ant-D3QN-PER algorithm is characterized in that: Step S22 is specifically as follows: Before quadtree segmentation of a grid, two segmentation criteria must be used to determine the size of the grid to be segmented. The segmentation criteria include: (a) first determine the size of the grid to be segmented. If its size is larger than the preset minimum grid size S, min , then enter the second segmentation criterion for judgment, otherwise the grid is not segmented and the next grid to be segmented is judged; (b) the gray value variance of the grid to be segmented is judged, if the gray value variance of all pixels in the grid is greater than the set value W f , then quadtree segmentation is performed, otherwise the grid is pre-divided into four pre-divided sub-grids, and the grayscale variance mean of the four pre-divided sub-grids is calculated. If the grayscale variance mean of the four pre-divided sub-grids is greater than the set value W s , then the actual quadtree segmentation is performed on the segmented grid according to the pre-segmentation method, otherwise the grid is no longer segmented; When judging the quadtree segmentation criterion for a grid, it is necessary to calculate the gray value variance σ of the grid 2 , the calculation formula is as follows: Among them, (r, c) is the coordinate of the pixel point in the grid, Gr(r, c) is the gray value of the pixel point, and the number of pixels in the grid is recorded as S. is the mean gray value of the grid; Calculating the mean grayscale variance of the four pre-divided sub-grids is to divide the σ 2 Sum and average.
6. The method for intelligent planning of power transmission lines based on deep reinforcement learning Ant-D3QN-PER algorithm according to claim 1, characterized in that: Step S3 specifically includes the following steps: S31, node transfer probability p of the improved ACS algorithm based on variable resolution grid m (u i ,b j )for: Among them, p m (u i ,b j ) represents the probability that ant m moves from node i in grid u to node j in grid b, τ(u,b;φ now ) is the pheromone concentration on the path from grid u to grid b, φ now are the current network parameters, is the shared θ now Convolutional network parameters, and They are the parameters of the advantage function network and the state value function network in DuelingNetwork, η(u i ,b j ) is the unit cost of the route between two grids u and b The reciprocal of Represents grid node u i and grid node b j Connect with u i The angle between the grid line and the end point, α, β, They are the control coefficient, Ω m (u) is the other grids that ant m can reach after removing the grids in the taboo table when it is at grid u, n ~ Indicates the number of grids that can be selected currently; S32, and find the next grid node b according to the following pseudo-random ratio rule j : Where q0 is a constant with a value in [0,1], q is a random number uniformly distributed in [0,1], S is the probability of selecting the reachable grid node according to the transition probability, and the ant adds the current grid to the taboo table every time it takes a step; S33, using the formulas of steps S31 and S32, the ant starts searching for a route from the starting point to the end point, and selects the grid node b j , judge u i -b j Does the connection satisfy all constraints? If not, exclude b j Repeat the above process and select from other feasible grid nodes; if the constraints are met, the ant moves to the selected grid node b j and as the current position, add grid b to the taboo table; until each ant has completed the route search, record the route trajectory of each ant, and calculate the Δτ(u,b) obtained by the ant performing each step according to each segment of the route trajectory, where Δτ(u,b)=1 / η(u i ,b j ), record M trajectories u1,b1,Δτ1,u2,…,u n ,b n ,Δτ n ,u n+1 and the route cost for each route; S34, based on the PER mechanism, each trajectory u1,b1,Δτ1,u2,…,u n ,b n ,Δτ n ,u n+1 Divide into n quadruplets (u i ,b i ,Δτ i ,u i+1 ) is stored in the experience replay pool as experience, and the TD error δ is calculated according to the following formula i With priority pr i : pr i =|δ i |+I Where γ is the discount coefficient, which takes values in [0,1], e is the grid that can be selected after the next grid b, and I is a positive constant close to zero; i With pr i The ants will store their experiences in the experience replay pool. If the experience in the experience replay pool is full, the first M n quadruplets will be removed according to the order in which the experiences were stored, and then the new experiences of the ants will be stored. S35, according to the priority of all the four-tuples in the experience playback pool, K four-tuples are extracted according to the probability determined by the priority; forward propagation is performed on the network, and τ(u k ,b k ; φ now ): in, and are the output values of the action advantage branch and the state value branch in the Dueling Network, respectively, indicating the superiority of the ant's action in the current state compared to the average action and the ability of the ant to obtain pheromones in the current state. The subscript k indicates the kth of the K extracted quadruple; S36, select the optimal action that maximizes pheromones: According to the best action and the target network calculates the TD target y k : in, is the current network parameter of the target network in Double DQN, and the TD error δ of each quadruple is calculated by the following formula k : d k =τ(u k ,b k ;f now )-y k S37, the weights ξ of the K quadruple groups are obtained according to the following formula k : Where N is the total amount of experience in the experience pool, λ is the control coefficient with a value of [0,1], There are K The maximum value in plays a normalization role, and then the weighted loss function is calculated according to the following formula: Update the parameters of the estimated network using gradient descent and θ now : Thus, the estimated network parameters are updated for Steps S35 to S37 are executed te times in a loop, i.e., te times of estimated network parameter updates constitute one round of updates; then the target network parameters are updated according to the following formula: for Update the target network parameters once every tg rounds after updating the estimated network parameters; S38, optimizing the corner points of the iterative optimal route in this iteration and recording it, and the grids passed by the optimized route become the iterative optimal route trajectory u1,b1,Δτ1,u2,…,u n ,b n ,Δτ n ,u n+1 , which is divided into n quadruples (u i ,b i ,Δτ i ,u i+1 ), randomly extract K quadruples from them, and use φ new and Calculate the loss function and update the network parameters to achieve the effect of updating pheromones. After each tg iterations, the global optimal route is used for global update. Update the guidance factor according to the following formula in, represents the control factor required for the t+1th iteration, μ is a constant with a value of (0,1); S39, steps S31 to S38 are an iterative process. If the maximum number of iterations T is reached, the above steps are terminated and the global optimal route is output. The global optimal route is obtained by optimizing the corner points. Otherwise, the iterative process continues.
7. The method for intelligent planning of power transmission lines based on deep reinforcement learning Ant-D3QN-PER algorithm according to claim 6, characterized in that: Step S33 is specifically as follows: The design and construction of transmission lines must meet the relevant transmission line design specifications. The specifications are introduced into the line selection algorithm as constraints. The specific constraints are as follows: (a) When the transmission line crosses over an existing line, the distance between the planned transmission line and the tower head of the existing line is greater than 13m; (b) When a transmission line crosses an existing line, the intersection angle between the planned transmission line and the existing line is greater than 15°, and the intersection angle with the road is greater than 45°; (c) The distance between the transmission line and the building is not less than 6m; The above constraints are added to the state transition rules when the ants are searching. Each time the ants search the grid to determine the next route, the above constraints must be met.
8. The method for intelligent planning of power transmission lines based on deep reinforcement learning Ant-D3QN-PER algorithm according to claim 6, characterized in that: Step S38 is specifically as follows: After tg iterations, the route with the best route cost is selected as the global optimal route. K quadruplets are randomly extracted from this trajectory and φ is used new Calculate the loss function and update the network parameters again to get the new φ new , complete this global update.
9. A computer program product, characterized in that It includes a computer program / instruction, which, when executed by a processor, implements a transmission line intelligent planning method based on deep reinforcement learning Ant-D3QN-PER algorithm as described in any one of claims 1 to 8.
10. An electronic device, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements a transmission line intelligent planning method based on deep reinforcement learning Ant-D3QN-PER algorithm as described in any one of claims 1 to 8.