A GIS and InSAR-based underground expressway reinforcement learning intelligent route selection method and computer system
Through the reinforcement learning method combining GIS and InSAR, the problem of insufficient comprehensive consideration of multi-source data in underground expressway route selection was solved, the standardization and precision of line design was achieved, the design efficiency and accuracy were improved, and various optimization requirements were met.
Patent Information
- Application Number
- CN202411585328.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-07
- Publication Date
- 2025-10-14
- Estimated Expiration
- 2044-11-07
AI Technical Summary
Existing technologies lack comprehensive consideration of multi-source data in urban underground expressway line selection, have a low level of quantification, are limited by traditional heuristic optimization algorithms, and reinforcement learning models are unable to meet the requirements of underground road line design, resulting in insufficient design efficiency and accuracy.
A reinforcement learning method based on GIS and InSAR is used to establish a rasterized observation space, combine remote sensing information, geological information and settlement data, construct a continuous line space, use the reinforcement learning model to optimize the line parameters, and combine the multi-index evaluation model to carry out line design.
It has achieved standardization and precision in the design of underground express routes, reduced interference from human factors, comprehensively considered multiple requirements such as economy and safety, and improved design efficiency and accuracy.
Smart Images

Figure CN119537502B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to an application technology of artificial intelligence in road alignment design, in particular to a GIS and InSAR-based underground expressway reinforcement learning intelligent alignment method. BACKGROUND
[0002] With the continuous superposition of regional integration and urban functions in China, the contradictions of urban population, space and resources are increasingly prominent, and the development and utilization of underground space resources have become a trend and necessity. In major cities, as an extension and supplement of the ground road system, underground expressways are developing rapidly. However, there are many indicators to be considered in the alignment of newly-built underground roads, and it is difficult to quantitatively evaluate. In addition to the line parameters, obstacle avoidance, traffic function improvement and other factors required for conventional road alignment, the construction is also affected by geological conditions and other factors. Improper construction will lead to deformation of the surrounding environment and cracking of houses. Therefore, how to comprehensively consider various factors and optimize the line has important challenges.
[0003] For the traditional intelligent alignment method of road alignment design, it is mostly based on the exploration of the known line intersection set in the simplified profile group or grid space, and a large number of possible line schemes cannot be detected, which has great limitations. The current reinforcement learning method is mostly for the pathfinding problem in the grid space, and the generated line is mostly a grid or polyline line, which does not have the line elements required by the specification and cannot be applied to the alignment of underground expressways.
[0004] The current problem is that:
[0005] 1) In the alignment stage of urban underground expressways, there are problems such as insufficient multi-source data acquisition and quantitative application in the current planning and design method. It is mostly based on the ground surface and ground object information in the GIS or CAD drawing, and the experience or AHP method is used to determine the plane alignment, vertical alignment, plane-vertical combination and other parameters. The comprehensive consideration of the above-ground and underground multi-source data and the settlement of the alignment area is lacking, and the quantitative level is low.
[0006] 2) The current heuristic optimization algorithm for road alignment design is limited, and the reinforcement learning model based on the pathfinding problem cannot meet the basic requirements of underground road alignment design in the action space design and alignment space design. This is the place that needs to be improved in the present application. SUMMARY
[0007] The technical problem to be solved by the application is to provide a GIS and InSAR-based underground expressway reinforcement learning intelligent alignment method, which is beneficial to improve the efficiency and accuracy of underground road design and reduce the interference of human factors.
[0008] In order to solve the above technical problems, the application provides a GIS and InSAR-based underground expressway reinforcement learning intelligent route selection method, comprising the following steps:
[0009] Step S1, a grid observation space is established based on the urban environment multi-source data provided by the geographic information system (GIS) and the interferometric synthetic aperture radar (InSAR), and an initial linear continuous route space is established according to the specification parameterization of the plane line of the underground expressway, and the two together build a route selection space of the reinforcement learning model;
[0010] Further, the route selection space of the reinforcement learning model comprises the following steps:
[0011] Step S11, a GIS grid data set is established according to the remote sensing information, geological information and hydrological information provided by the GIS, the historical settlement data calculated by the InSAR is extracted, the average rate, average value and extreme value of the past settlement are calculated by a statistical method to form an InSAR settlement grid data set, and the two are unified into grid data of the same precision to form an observation space, thereby providing state information for the subsequent reinforcement learning model;
[0012] Step S12, three types of plane basic line types required in the road specification are parameterized, the line type is divided into a main line type and a secondary line type according to the combination characteristics required by the design specification, a default unit line type is formed by the main line type and the secondary line type, the parameterized unit line type is represented by three parameters of radius, angle and spiral line parameter, and finally the underground expressway design route formed by a plurality of line types is represented by a unique sequence formed by the parameters of each group and forms a route space, thereby providing accurate route information of the final optimized route;
[0013] Further, step S12 comprises: constructing the mathematical equations of straight lines, easement curves and circular curves by mathematical methods, wherein the straight lines and the circular curves are represented by a radius R and an angle θ as the main line type, the easement curve is represented by a spiral line parameter A as the secondary line type, a unit line type is formed by splicing a main line type with a secondary line type, and is determined by a group of parameters (R i ,θ i ,A i ), and finally the route formed by a plurality of unit line types is represented by a parameter sequence
[0014] [(R0,θ0,A0),(R1,θ1,A1),(R2,θ2,A2),…,(R n ,θ n ,A n )].
[0015] Step S2, the strategy network of the reinforcement learning model receives the parameter information of the route selection space and outputs the line type parameters;
[0016] Furthermore, the reinforcement learning model includes: the reinforcement learning network includes a value network, a target value network, a policy network and a target policy network, wherein the policy network and the target policy network output line type parameters (R i ,θ i ,A i ).
[0017] Step S3: The line type builder modifies the line type parameters according to the specification requirements and updates the line selection space;
[0018] Furthermore, the line type builder is based on the output line type parameters (R i ,θ i ,A i ) Update line space;
[0019] The line type builder is constructed based on the line type parameters (R i ,θ i ,A i ) Determine the types of primary and secondary line types, calculate and generate the line type parameter equation based on the coordinates and direction of the line end, and modify the line type parameters according to the specification requirements. The specific parameter equation is as follows:
[0020] straight line: Circular curve: Cloak line:
[0021] Where: α represents the direction of the line end, (x0, y0) is the position coordinate of the line end, s, s' are the parameters of the clothoid line, where s = A 2 / R i ,s'=A 2 / R i-1 ;
[0022] The sign of the clothoid line depends on the line type parameters (R i ,θ i ,A i ) and (R i-1 ,θ i-1 ,A i-1 )Sure.
[0023] Step S4: The multi-index evaluation model outputs reward and penalty values based on the parameter information of the new line selection space. The value network of the reinforcement learning model receives the reward and penalty values, calculates the gradient, and updates the strategy model by passing the gradient.
[0024] Furthermore, the multi-index evaluation model and reinforcement model training include:
[0025] In step S41, a multi-index evaluation model comprehensively evaluates the currently generated route based on multiple factors according to the updated route selection space. The multi-index evaluation model considers factors including the design parameters of the current route, feature classification, obstacle information, and target information from GIS remote sensing information, and settlement information provided by InSAR. The multi-index evaluation model considers evaluation aspects including target reachability, obstacle avoidance, economy, and safety. The specific reward and penalty values are calculated as follows:
[0026] Reward=(w1r1+w2r2+w3r3+w4r4+w5r5)×β (4);
[0027] Reward for approaching the target: r = (d i -d i-1 ) / D(5);
[0028] Rewards for course correction:
[0029]
[0030] Penalty for pointing to an obstacle: r3 = sense_r / obstacle ≤20° (7);
[0031] Line length penalty: r4 = (1-(d i -d i-1 ) / l) (8);
[0032] Penalty for subsidence risk: r5 = (0.01n i -m i ) / l (9);
[0033] Where, d i The distance between the route generated in step i and the target; sense_r is the range of obtaining status information; obstacle ≤20° is the distance to the nearest obstacle within 20° of the line direction; l is the length of the currently generated set of unit lines; n i All blocks that the current set of unit lines passes through; m i The high-risk blocks that the current set of lines passes through; reward_shaping is the modification of the reward value, which increases as the line approaches the end point. When the line reaches the end point, a fixed reward value of 100 is obtained, and if it hits an obstacle, a fixed penalty value of -100 is obtained;
[0034] In step S42, the reinforcement learning value network and the target value network calculate the gradient based on the obtained reward value to update the value network, and transfer the gradient to the value policy network to update it at the same time. The target network partially copies the parameters of the policy network and the value network for soft update. The gradient is calculated as follows:
[0035] GradQ=Q+Reward-Q' (10);
[0036] Where GradQ represents the updated gradient of the value network, Q represents the maximum future reward value estimated by the value network, Reward represents the reward and punishment value returned by the evaluation model, and Q' represents the maximum future reward value estimated by the target value network after the update.
[0037] Step S5, looping steps S2-S4 until a round of line selection is completed and the accumulated reward value is calculated;
[0038] Furthermore, completing a round of line selection includes: looping steps S2-S4 until a round of line selection is completed, wherein the conditions for completing a round of line selection are reaching the end point, hitting an obstacle, and exceeding the limited number of cycles; after a round of line selection is completed, the total value of the entire reward value is counted.
[0039] Step S6: When the accumulated average reward value does not reach the optimization target, loop through steps S1-S5 until the optimization target is reached and the optimal design is found.
[0040] Furthermore, achieving the optimization goal includes: looping steps S1-S5 until the total reward value reaches the optimization goal. The condition for achieving the optimization goal is that the total reward value continues to rise and reaches a stable state. At this time, the converged strategy network generates a path that is the optimal path.
[0041] The beneficial effects of the present invention are:
[0042] 1) Designing the line from the perspective of linear parameters can achieve standardized underground expressway design in one step;
[0043] Traditional heuristic algorithms often require a pathfinding algorithm or a traversal algorithm to initially obtain a gridded corridor for line selection, and then use a heuristic algorithm to optimize the line. However, this application designs lines from the perspective of line parameters, which can effectively break down the specifications into each step of line generation, ensuring line standardization while achieving line generation in one go.
[0044] 2) The line selection space is set to a combination of linear continuous space and raster data space, which can help the model fully explore the environment, achieve precise line design, and explore the global optimal solution;
[0045] Traditional heuristic algorithms are mostly optimized based on the intersection set of an initial route. Since the number of initial intersection sets is fixed and their positions can only be moved on the section or grid nodes, in actual design, the number and distribution of intersection sets are not constrained. This leads to the neglect of a large number of possible route combinations when designing routes based on the distribution of intersection sets, which has huge limitations and often fails to obtain the global optimal solution. However, the routes in this application are drawn in a continuous space, which can ensure the accurate design of the routes and can be freely drawn within the specification range.
[0046] 3) The evaluation model combines GIS, InSAR, and line parameter information to consider multiple requirements, including economy and safety, making it more comprehensive;
[0047] 4) Traditional pathfinding often only considers factors such as target arrival, obstacle avoidance, and earth excavation. The data combined is often limited to elevation maps, and the factors considered are relatively narrow. However, this application considers remote sensing information, geological information, and hydrological information provided by GIS to establish a GIS raster dataset, as well as historical settlement data calculated by InSAR, and route design information. Data can be obtained from the route selection space to achieve rapid evaluation and achieve comprehensive optimization goals. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are intended to explain the present invention and do not constitute an undue limitation of the present invention. In the accompanying drawings:
[0049] Figure 1 is a flow chart of a specific embodiment of the present invention;
[0050] Figure 2 is a flow chart of reinforcement learning according to a specific embodiment of the present invention;
[0051] Figure 3 1 is a schematic diagram of the process of constructing a line type by a standardized line type builder according to an embodiment of the present invention;
[0052] Figure 4 1 is a flow chart of multi-factor evaluation for underground road route selection according to an embodiment of the present invention;
[0053] Figure 5 It is a schematic diagram of a final generated circuit according to an embodiment of the present invention. DETAILED DESCRIPTION
[0054] The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0055] like Figure 1 As shown, the present invention provides an intelligent line selection method for underground expressways based on reinforcement learning using GIS and InSAR, comprising the following steps:
[0056] Step S1: establishing a gridded observation space based on multi-source urban environment data provided by a geographic information system (GIS) and an interferometric synthetic aperture radar (InSAR), and establishing a linear continuous line space based on a standardized parameterized underground rapid plane line shape. The two together constitute the line selection space of the reinforcement learning model.
[0057] Specifically, S1 includes the following steps S11 to S12:
[0058] Step S11: GIS is used to obtain relevant data within the selected area, including hydrological information, surface remote sensing images, elevation, slope, aspect, and feature classification data. InSAR statistics are used to calculate the subsidence rate, extreme subsidence value, and mean subsidence value within the area. Aligning these data to grids of equal size completes the establishment of the observation space, providing state information for the subsequent reinforcement learning model.
[0059] Step S12: The route parameter space is a linear continuous space. A set of initial route endpoints, directions, and line type parameters are required. The endpoint positions can be determined based on the coordinates of the starting point in the observation space. The initial line type parameters are in the form of (R0, θ0, A0), which are set based on the line type before the starting point of the route. The route space is then formed according to the route formation sequence generated by the reinforcement learning model, which is in the form of [(R0, θ0, A0), (R1, θ1, A1), (R2, θ2, A2), ..., (R n ,θ n ,A n )], providing accurate route information of the final optimized route.
[0060] Step S2: The strategy network of the strong chemistry model receives parameter information of the line selection space and outputs line type parameters;
[0061] See also Figure 2 The reinforcement learning model includes a value network, a target value network, a policy network, and a target policy network. The target network and the corresponding network have the same network structure. Among them, the value network receives the state information and line type parameters (R i ,θ i ,A i ) predicts the maximum reward value to be obtained in the future, and the policy network receives the state information and outputs the linear parameter (R i ,θ i ,A i );
[0062] Specifically, according to the different state information, the value network and the policy network adopt a multi-layer linear connection network or a convolution network, and the output value is within the required range through an activation function. In addition, in order to make the model explore more lines, an exploration rate conforming to a normal distribution is added in the output linear parameter (R i ,θ i ,A i ) in the initial stage, and the exploration rate gradually decreases with the gradual stabilization of the model training strategy, so as to prevent interference with the convergence of the model.
[0063] S3, the linear constructor corrects the linear parameter according to the specification requirements, and updates the alignment space;
[0064] Referring to Figure 3 , specifically, the normalized linear constructor divides the linear into a main linear and a secondary linear, wherein the main linear includes a straight line and a circular curve, and the secondary linear is a transition curve or a straight line. The normalized linear constructor reads the linear parameter (R i ,θ i ,A i ), determines whether the main linear is a straight line or a circular curve according to the size of the radius R i , wherein the threshold is set to 10000m according to the specification, that is, a straight line is set when it exceeds 10000m, and the line is drawn according to formula (1), and the rest is a circular curve, which is drawn according to formula (3). Then, the type of the transition curve is determined in combination with the previous linear parameter (R i-1 ,θ i-1 ,A i-1 ), and the transition curve can adopt a clothoid, and in some cases, a straight line can also be used. Here, the type is determined according to the length of different transition curves, and the clothoid is shorter, and the line is drawn according to formula (3). The specification correction aims to correct the size of the generated line according to the specification requirements. When the generated line radius is less than the minimum available radius 200m, or the clothoid parameter A does not meet the specification requirements, the constructor corrects the input linear parameter according to the specification, so as to ensure that the linear combination is consistent with the specification. Finally, the alignment space is updated according to the corrected linear parameter.
[0065] S4, the multi-index evaluation model outputs the reward and punishment value according to the parameter information of the new alignment space, the value network of the reinforcement learning model receives the reward and punishment value to calculate the gradient for updating, and the gradient updates the policy model;
[0066] Referring to Figure 4The line selection space consists of a linear continuous space and a grid space. The evaluation model obtains the design parameters of the line from the linear continuous space and projects the generated line type into the grid space for unified evaluation using formulas (1)-(3) provided in the line type builder. The data set in the grid space should include feature classification, obstacle classification, target information, etc. The multi-factor comprehensive evaluation can be combined with the grid data to calculate the reward and penalty values according to formulas (4)-(9).
[0067] Specifically, the reinforcement learning model training involves obtaining the maximum future reward value predicted by the value network through each of the four networks, then calculating the gradient according to formula (10). The value model is then updated through backpropagation based on this gradient, and the policy network is updated based on the gradient of the transferred linear parameters. To ensure the stability of the model during training, a soft update method is used when updating the target network, copying the corresponding network parameters at a low rate for update. In addition, an experience pool is constructed in the model, storing state information, linear parameters, and reward and penalty values within the same cycle as a group to improve data training efficiency.
[0068] S5, loop S2-S4 until a round of line selection is completed and the accumulated reward value is calculated;
[0069] Specifically, the completion of a round of route selection is determined by a multi-factor evaluation model. This model determines basic information such as whether the route currently intersects environmental obstacles, whether it has reached its destination, and whether the number of loops has been exceeded. When the conditions for completing a round of route selection are met, the current loop ends, and the environment is reset for the next round of route selection. If the end conditions are not met, the loop continues until they are met.
[0070] S6, when the accumulated average reward value does not reach the optimization target, loop S1-S5 until the optimization target is reached and the optimal design is found.
[0071] See also Figure 5 Specifically, the model calculates the total reward value for each round of training. During training, the total reward value for each round gradually increases, eventually reaching a stable convergence value. At this point, the optimization of the target is complete, generating the final optimized path.
[0072] Through continuous iteration and optimization, the reinforcement learning model can gradually learn the optimal line design solution.
[0073] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A reinforcement learning intelligent route selection method for underground expressways based on GIS and InSAR, comprising the following steps: Step S1: establishing a gridded observation space based on multi-source urban environment data provided by Geographic Information System (GIS) and Interferometric Synthetic Aperture Radar (InSAR), and establishing an initial linear continuous line space based on a standardized parameterized underground rapid plane line shape. The two together construct a line selection space for a reinforcement learning model. Step S2: The policy network of the reinforcement learning model receives parameter information of the line selection space and outputs line type parameters; Step S3: The line type builder modifies the line type parameters according to the specification requirements and updates the line selection space; Step S4: The multi-index evaluation model outputs reward and penalty values based on the parameter information of the new line selection space. The value network of the reinforcement learning model receives the reward and penalty values, calculates the gradient, updates it, and transmits the gradient to update the strategy model. This includes: In step S41, a multi-index evaluation model comprehensively evaluates the currently generated route based on multiple factors according to the updated route selection space. The multi-index evaluation model considers factors including the design parameters of the current route, feature classification, obstacle information, and target information from GIS remote sensing information, and settlement information provided by InSAR. The multi-index evaluation model considers evaluation aspects including target reachability, obstacle avoidance, economy, and safety. The specific reward and penalty values are calculated as follows: (4); in: Correction for reward value; Rewards for approaching the target: (5); Where: D is the distance from the starting point to the destination point of the route; Rewards for course correction: (6); Penalty for pointing to obstacles: (7); Line length penalty: (8); Penalty for Subsidence Risk: (9); Where, Refers to the angle between the tangent line at the end of the line and the line connecting the endpoint to the target point when the line is generated to the i-th step; The distance between the route generated for step i and the target; The scope of obtaining status information; is the distance to the nearest obstacle within 20° of the line direction; l is the length of the currently generated set of unit lines; All the blocks that the current set of unit lines passes through; A set of high-risk blocks that are currently generated by the line type; The reward value is modified, which increases as the line approaches the end point. When the line reaches the end point, it gets a fixed reward value of 100. If it hits an obstacle, it gets a fixed penalty value of -100. In step S42, the reinforcement learning value network and the target value network calculate the gradient based on the obtained reward value to update the value network, and transfer the gradient to the value policy network to update it at the same time. The target network partially copies the parameters of the policy network and the value network for soft update. The gradient is calculated as follows: (10); Where, represents the updated gradient of the value network, represents the maximum future reward value estimated by the value network, Indicates the reward and punishment value returned by the evaluation model. Represents the maximum future reward value estimated by the updated target value network; Step S5, looping steps S2-S4 until a round of line selection is completed and the accumulated reward value is calculated; Step S6: When the accumulated average reward value does not reach the optimization target, loop steps S1-S5 until the optimization target is reached and the optimal design is found.
2. The method for intelligent route selection for underground expressways based on reinforcement learning using GIS and InSAR according to claim 1 is characterized by: The step S1 includes the following steps: Step S11: A GIS raster dataset is established based on the remote sensing information, geological information, and hydrological information provided by GIS. Historical settlement data calculated by InSAR are extracted. The average rate, mean value, and extreme value of past settlement are calculated using statistical methods to form an InSAR settlement raster dataset. The two are unified into raster data of the same accuracy to form an observation space, providing state information for the subsequent reinforcement learning model. Step S12, parameterize the three types of basic plane line types required by the road specifications. The line types are divided into primary line types and secondary line types based on the combination characteristics required by the design specifications. The default set of unit line types consists of a primary line type and a secondary line type. The parameterized set of unit line types is represented by three parameters: radius, turning angle, and spiral line parameters. Finally, the underground expressway design line composed of multiple sets of line types is represented by a unique sequence of parameters of each set and forms a line space, providing accurate line information for the final optimized line.
3. The intelligent route selection method for underground expressways based on reinforcement learning using GIS and InSAR according to claim 2 is characterized by: The step S12 includes: constructing mathematical equations of straight lines, transition curves, and circular curves by mathematical methods, wherein straight lines and circular curves are represented as main line types by radius R and rotation angle θ, transition curves are represented as secondary line types by clothoid parameter A, a set of unit line types is composed of secondary line types spliced together with a main line type, and a set of parameters is used to represent the main line type. Determine the parameter sequence of the line composed of multiple groups of unit line types express; in: is the radius of the circular curve generated in step i, is the turning angle of the circular curve generated in step i, The clothoid parameters corresponding to the transition curve generated in step i.
4. The method for intelligent route selection for underground expressways based on reinforcement learning using GIS and InSAR according to claim 1 is characterized by: The step S2 comprises: The reinforcement learning network includes a value network, a target value network, a policy network, and a target policy network. The policy network and the target policy network output line parameters by receiving the state information of the line selection space. .
5. The method for intelligent route selection for underground expressways based on reinforcement learning using GIS and InSAR according to claim 1 is characterized by: The step S3 includes: the line type builder outputs the line type parameters according to the reinforcement learning strategy network. Update line space; The line type builder is constructed based on the line type parameters Determine the types of primary and secondary line types, calculate and generate the line type parameter equation based on the coordinates and direction of the line end, and modify the line type parameters according to the specification requirements. The specific parameter equation is as follows: straight line: (1); Circular curve: (2); Cloak line: (3); Where: Indicates the direction of the line end, is the position coordinate of the end of the line, is the clothoid parameter, where , A is the clothoid parameter; The sign of the clothoid line depends on the line type parameters of the current and previous steps. and Sure.
6. The method for intelligent route selection for underground expressways based on reinforcement learning using GIS and InSAR according to claim 1, characterized in that: The step S5 comprises: Repeat steps S2-S4 until a round of line selection is completed, where the conditions for completing a round of line selection are reaching the end point, hitting an obstacle, and exceeding the limited number of cycles; after a round of line selection is completed, the total value of the entire reward value is calculated.
7. The method for intelligent route selection for underground expressways based on reinforcement learning using GIS and InSAR according to claim 1 is characterized by: The step S6 comprises: Repeat steps S1-S5 until the total reward value reaches the optimization target. The condition for achieving the optimization target is that the total reward value continues to rise and reaches a stable state. At this time, the path generated by the converged policy network is the optimal path.
8. A computer system comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the underground expressway reinforcement learning intelligent line selection method based on GIS and InSAR as described in any one of claims 1-7.
Citation Information
Patent Citations
Method for establishing quantitative model of road performance evaluation indexes
CN114880744A
Underground road reinforcement learning intelligent line selection method based on multi-source data
CN118095595A