An intelligent route selection method for underground roads based on multi-source data using reinforcement learning

By integrating GIS and InSAR data in urban underground expressway selection, combining evaluation models and reinforcement learning algorithms, intelligent line selection of underground roads is realized, solving the problem of insufficient comprehensive consideration of multi-source data in the existing technology, and improving design efficiency and accuracy.

CN118095595BActive Publication Date: 2025-05-30TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410224388.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-29
Publication Date
2025-05-30
Estimated Expiration
2044-02-29

AI Technical Summary

Technical Problem

During the selection stage of urban underground expressway line, the existing planning and design methods lack comprehensive consideration for multi-source data, the quantitative level is low, making it difficult to effectively optimize the route.

Method used

The intelligent line selection method of underground road reinforcement learning based on multi-source data is adopted. By integrating GIS and InSAR data, combining evaluation models and reinforcement learning algorithms, we comprehensively use digital information such as geology, geology, and surface settlement to dynamically decide the route direction to achieve optimal line design.

Benefits of technology

It improves the efficiency and accuracy of underground road design, realizes the automated design of underground road line selection, reduces interference from human factors, and provides technical support for underground road construction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118095595B_ABST
    Figure CN118095595B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent route selection method for underground roads based on multi-source data using reinforcement learning. The intelligent route selection method for underground roads using reinforcement learning includes: establishing a route selection space by fusing multi-source data of GIS and InSAR and adopting a parameterization method; after generating a section of parameterized route each time, inputting relevant evaluation parameters in the route selection space into an evaluation model for evaluation; inputting relevant parameters of the route selection space into a reinforcement learning model for reinforcement learning; inputting the line type parameters of the newly generated next section of route into a route construction network to generate a new route selection space according to the input parameters; repeating steps S2 to S4 until an optimal design is found in the route selection space. The intelligent route selection method for underground roads of the present invention can realize the automated design of underground road route selection, thereby being conducive to improving the efficiency and accuracy of underground road design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an application technology of artificial intelligence in road route selection design, and particularly to an intelligent route selection method for underground roads based on multi-source data using reinforcement learning. Background Art

[0002] With the continuous superposition of regional integration and urban functions in China, contradictions such as urban population, space, and resources have become increasingly prominent, and the development and utilization of underground space resources have become a trend and necessity. In major cities, as an extension and supplement to the ground road system, underground expressways have developed rapidly. However, there are many indicators to be considered in the route selection of newly built underground roads, and it is difficult to conduct quantitative evaluation. In addition to considering elements such as linear parameters, obstacle avoidance, and traffic function improvement required for conventional road route selection, it is also affected by factors such as geological conditions. Improper construction will cause deformation of the surrounding environment and cracking of houses. Therefore, how to comprehensively consider various factors and optimize the route is an important challenge.

[0003] The current problems are as follows:

[0004] In the route selection stage of urban underground expressways, considering various influencing factors comprehensively, there are problems such as insufficient acquisition and quantitative application of multi-source data in the current planning and design methods. Mostly relying on surface feature information in GIS or CAD drawings, using means such as experience or analytic hierarchy process to determine parameters such as horizontal alignment, vertical alignment, and horizontal-vertical combination, lacking comprehensive consideration of multi-source data such as above-ground, underground, and settlement in the route selection area, and having a low level of quantification. Summary of the Invention

[0005] The purpose of the present invention is to provide an intelligent route selection method for underground roads based on multi-source data using reinforcement learning. Through GIS (Geographic Information System) and InSAR (Interferometric Synthetic Aperture Radar), combined with artificial intelligence algorithms, comprehensively utilize digital information such as surface features, geology, and surface settlement and parameterize them. Using the reinforcement learning method, convert the above information into reward and punishment values and make dynamic decisions on the route direction, conduct optimal route design, and achieve automated design of underground road route selection. The invention is beneficial to improving the efficiency and accuracy of underground road design.

[0006] In order to achieve the above technical purpose, the present invention adopts the following technical solutions:

[0007] An intelligent route selection method for underground roads based on multi-source data using reinforcement learning, the intelligent route selection method for underground roads using reinforcement learning includes steps S1 to S5;

[0008] S1, establish a route selection space by fusing multi-source data of GIS, InSAR, and an information model including regional hydrogeology, underground buildings and obstacles, and using a parameterization method;

[0009] S2. After generating a parametric route segment each time, input the relevant evaluation parameters in the route selection space into the evaluation model for evaluation;

[0010] S3. Input the relevant parameters of the route selection space into the reinforcement learning model for reinforcement learning;

[0011] S4. Input the line type parameters of the newly generated next route segment into the route construction network, and generate a new route selection space according to the input parameters;

[0012] S5. Repeat steps S2 to S4 until the optimal design is found in the route selection space.

[0013] Furthermore, step S1 includes:

[0014] S11. Parametrize the basic line type according to the specification requirements. The parameters mainly involved in the parametrization include radius and turning angle;

[0015] S12. When constructing the route selection space, the multi-source information parametrization process is based on GIS data, considering the surrounding environmental factors. For the visible ground object categories, machine vision technology is used to automatically identify and label the ground objects in a vast area.

[0016] Furthermore, step S12 also includes: introducing InSAR data, including ground settlement rate and time series offset value information, and constructing a digital spatial distribution matrix in the three-dimensional grid space.

[0017] Furthermore, step S2 also includes: the evaluation model comprehensively evaluates the route according to the technical, safety, sustainability, and economic evaluation indicators, and returns the corresponding evaluation level. At the same time, according to the evaluation results, corresponding reward and punishment values are generated to reflect the quality of the route design, rewarding good designs and punishing poor designs;

[0018] The reward and punishment values mainly consider the goal achievement situation and the evaluation level of the evaluation model, and are calculated in the form of a design reward function. The goal achievement situation mainly reflects the distance between the end of the currently generated route and the target location, and the evaluation level reflects the comprehensive evaluation result of the currently generated route; the design principle of the reward function is to reward the routes that are beneficial to achieving the goal and improving the evaluation level of the evaluation model, and vice versa.

[0019] Furthermore, the evaluation model consists of evaluation indicators and evaluation methods, and is used to evaluate the generated routes;

[0020] The evaluation indicators include technical indicators, safety indicators, sustainability indicators, and economic indicators, forming a multi-level indicator system;

[0021] The evaluation method uses the method of an index matrix to complete the grading of the line.

[0022] Further, step S3 further includes: during the reinforcement learning process, the route selection space and the reinforcement learning model perform interactive learning through the reward and punishment value; the reinforcement learning model continuously optimizes the parameters of the policy network to find the optimal line design scheme;

[0023] The reinforcement learning model predicts the future line type parameters based on the characteristics of the current route selection space and the reward and punishment value, and selects the optimal line type parameters as the parameter values of the next section of the line.

[0024] Further, the update and optimization of the reinforcement learning model parameters include:

[0025] The line construction network constructs the spatial distribution of the tunnel in the route selection space according to the generated line type parameters, so as to generate the next route selection space;

[0026] The policy network receives the current route selection space, outputs the current line type parameters, and then sends them into the line construction network;

[0027] The value network receives the reward and punishment value (represented by r) returned by the evaluation model, and calculates the future maximum total reward (represented by Q) according to the current route selection space and action; at the same time, the value network also calculates the future maximum total reward (represented by Q') according to the updated route selection space and action; the gradient of the value network is calculated through formula 1 for update, and at the same time, the gradient of the action is also returned to help the policy network for update;

[0028] GradQ = Q + r - Q' (1)

[0029] In the formula,

[0030] GradQ represents the update gradient of the value network,

[0031] Q represents the future maximum total reward calculated according to the current route selection space and action,

[0032] r represents the reward and punishment value returned by the evaluation model,

[0033] Q' represents the future maximum total reward calculated according to the updated route selection space and action.

[0034] Further, step S5 further includes: the reinforcement learning is updated by continuously receiving the reward and punishment value feedback from the environment, and continuously optimizing the parameters of the policy network until reaching the end point and the reward value reaches the expectation or converges. The line generated by the final policy network is the optimal line.

[0035] In the intelligent route selection method for underground roads based on reinforcement learning of the present invention, through steps such as integrating multi-source data of GIS and InSAR, an evaluation model, reinforcement learning, and route construction, the linear parameters are gradually optimized, and finally the optimal road design plan that meets the specification requirements is obtained, realizing the automated design of underground road route selection. This is conducive to improving the efficiency and accuracy of underground road design, reducing the interference of human factors, and providing strong technical support for underground road construction.

[0036] Based on the reinforcement learning model, the automatic generation of linear parameters and the automatic update of the policy network can be realized, effectively improving the route selection efficiency of designers.

[0037] Through the intelligent evaluation system for the route selection of urban underground expressways and combined with intelligent algorithms, the route can be evaluated in real time, and the results are fed back into machine learning in the form of reward and punishment values, providing a basis for subsequent route generation and achieving the optimization of route selection. Brief Description of the Drawings

[0038] Figure 1 It is a schematic flow chart of an intelligent route selection method for underground roads based on multi-source data of the present invention;

[0039] Figure 2 It is a schematic flow chart of the reinforcement learning involved in the intelligent route selection method for underground roads of the present invention. Detailed Embodiments

[0040] The following uses specific embodiments to further illustrate the present invention:

[0041] This embodiment provides an intelligent route selection method for underground roads based on multi-source data. This intelligent route selection method for underground roads can realize the automated design of underground road route selection, which is conducive to improving the efficiency and accuracy of underground road design.

[0042] See Figure 1 , the intelligent route selection method for underground roads in this embodiment specifically includes the following steps S1 to S5.

[0043] S1. By integrating multi-source data of GIS (Geographic Information System), InSAR (Interferometric Synthetic Aperture Radar), and an information model including regional hydrogeology, underground buildings, and obstacles, and using a parameterization method to establish a route selection space. In this process, the parameterized line type is constrained according to the specifications to ensure that the route design meets the standards and regulations. At the same time, machine vision algorithms are used to annotate and obtain some GIS data for subsequent route design.

[0044] Specifically, step S1 includes the following steps S11 to S12:

[0045] S11, Parametrize common basic line types according to specification requirements, such as circular curves, clothoid curves, straight lines, etc. The parameters mainly involved in parametrization include radius and deflection angle.

[0046] S12, When constructing the route selection space, the multi-source information parametrization process is based on GIS data, considering surrounding environmental factors, such as various buildings (structures), pile foundations, underground pipelines, and tunnels. For visible ground feature categories, machine vision technology is used to automatically identify and label ground features in a vast area. Additionally, InSAR data is introduced, including information such as ground settlement rate and temporal offset value, and a digital spatial distribution matrix is constructed in the three-dimensional grid space.

[0047] In the actual implementation process of the method, a three-dimensional space for route selection is established based on the currently determined starting and ending points of the line. The three-dimensional space is discretized into a grid-like network space, a space composed of cubes with side length L in the dimension of X×Y×Z. Among them, a two-dimensional planar space is constructed according to the top view. Based on the grid data obtained from GIS, grids occupied by land, grids occupied by pile foundations, pipelines, etc. are generated in the grid route selection space, and the subsequently generated tunnel routes will also occupy corresponding positions in the route selection space.

[0048] S2, After each generation of a parametrized line segment, relevant evaluation parameters in the route selection space are input into the evaluation model for assessment. The evaluation model will comprehensively evaluate the line according to evaluation indicators such as technicality, safety, sustainability, and economy, and return the corresponding evaluation grade. At the same time, corresponding reward and punishment values are generated according to the evaluation results to reflect the quality of the line design, rewarding good designs and punishing poor designs.

[0049] The reward and punishment values are mainly calculated in the form of a designed reward function considering the goal achievement situation and the evaluation grade of the evaluation model. The goal achievement situation mainly reflects the distance between the end of the currently generated line and the target location, and the evaluation grade reflects the comprehensive evaluation result of the currently generated line. The design principle of the reward function is to reward lines that are beneficial to achieving the goal and improving the evaluation grade of the evaluation model, and vice versa.

[0050] The evaluation model consists of evaluation indicators and evaluation methods and is used to evaluate the generated lines. The evaluation indicators mainly include technical indicators, safety indicators, sustainability indicators, economic indicators, etc., forming a multi-level indicator system. The evaluation method uses the method of indicator matrix to complete the grade division of the line.

[0051] S3, Take the relevant parameters of the route selection space as input values and input them into the reinforcement learning model for reinforcement learning.

[0052] During the reinforcement learning process, the route selection space and the reinforcement learning model interact and learn through reward and punishment values. The reinforcement learning model continuously optimizes the parameters of the policy network to find the optimal route design solution. Specifically, based on the characteristics of the current route selection space and the reward and punishment values, it predicts the future line type parameters and selects the optimal line type parameters as the parameter values for the next section of the route.

[0053] That is to say, during the reinforcement learning process, the model adjusts the parameters of the value network according to the previously generated reward and punishment values. The value network is used to evaluate the quality of various behaviors. At the same time, the policy network outputs the line type parameters of the next section of the route based on the feedback of the value network, gradually optimizing the route design.

[0054] Specifically,

[0055] See Figure 2 The reinforcement learning model mainly consists of three parts: a policy network, a value network, and a route construction network. These networks work together to update and optimize the parameters of the model to achieve better performance.

[0056] Among them,

[0057] The role of the policy network is to receive the current route selection space, output the line type parameters of the route according to the characteristics of the route selection space, and receive the gradient of the action parameters returned by the value network for updating. The output line type parameters are then sent into the route construction network.

[0058] The role of the route construction network is to update the route selection space according to the route selection space and the line type parameters output by the policy network, and then input it into the evaluation network to obtain the returned reward and punishment values.

[0059] The role of the value network is to output the maximum rewards that can be obtained currently and in the future respectively according to the current route selection space and the updated route selection space, calculate the gradient update network in combination with the reward and punishment values, and return the gradient of the action parameters.

[0060] For the reinforcement learning model, the process of parameter update and optimization is roughly as follows:

[0061] 1) The route construction network constructs the spatial distribution of the tunnel within the route selection space according to the generated line type parameters, thereby generating the next section of the route selection space.

[0062] 2) The policy network receives the current route selection space, outputs the current line type parameters, and then sends them into the route construction network.

[0063] 3) The value network receives the reward and punishment value returned by the evaluation model (denoted by r), and calculates the future maximum total reward (denoted by Q) based on the current route selection space and actions. At the same time, the value network also calculates the future maximum total reward (denoted by Q') based on the updated route selection space and actions. The gradient of the value network is calculated by Formula 1 for update, and the gradient of the actions is also returned to help update the policy network;

[0064] GradQ = Q + r - Q’ (1)

[0065] In the formula,

[0066] GradQ represents the update gradient of the value network,

[0067] Q represents the future maximum total reward calculated based on the current route selection space and actions,

[0068] r represents the reward and punishment value returned by the evaluation model,

[0069] Q’ represents the future maximum total reward calculated based on the updated route selection space and actions.

[0070] Throughout the process, the reinforcement learning model continuously learns, updates, and optimizes parameters based on the reward and punishment values feedback by the evaluation network, improving the quality and performance of the generated routes, and ultimately achieving a more efficient tunnel route selection design.

[0071] S4. Input the line type parameters of the newly generated next section of the line into the line construction network, and generate a new route selection space based on these parameters. If the newly generated line "does not reach the end of the line within the specified number of steps" or "does not reach the target required total reward value when reaching the end of the line", then return to the starting point to generate a new line.

[0072] S5. Repeat steps S2 to S4 until the optimization target value is reached, that is, find the optimal design scheme in the route selection space.

[0073] Reinforcement learning is updated by continuously receiving the reward and punishment values feedback by the environment, continuously optimizing the parameters of the policy network until reaching the end and the reward value reaches the expectation or converges. At this time, the line generated by the final policy network is the optimal line.

[0074] Through continuous iteration and optimization, the reinforcement learning model can gradually learn the best line design scheme.

[0075] The intelligent route selection method for underground roads in this embodiment can achieve the automated design of underground road route selection. By integrating steps such as GIS and InSAR multi-source data, evaluation models, reinforcement learning, and route construction, the linear parameters are gradually optimized, and finally the optimal road design plan that meets the specification requirements is obtained. This is beneficial to improving the efficiency and accuracy of underground road design, reducing the interference of human factors, and providing strong technical support for underground road construction. Moreover, based on the reinforcement learning model, the automatic generation of linear parameters and the automatic update of the policy network can be realized, effectively improving the route selection efficiency of designers. Through the intelligent evaluation system for the route selection of urban underground expressways and combined with intelligent algorithms, the route can be evaluated in real time, and the results are fed back to machine learning in the form of reward and punishment values, providing a basis for subsequent route generation and achieving the optimization of route selection.

[0076] The above are only the preferred embodiments of the present invention and are not used to limit the protection scope of the present invention. Therefore, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent line selection method for underground roads based on multi-source data through reinforcement learning, characterized in that: The underground road reinforcement learning intelligent line selection method comprises steps S1 to S5; S1, by integrating multi-source data of GIS, InSAR and information models including regional hydrogeology, underground structures and obstacles, and using parameterized methods to establish the line selection space; S2, after each parameterized line is generated, the relevant evaluation parameters in the line selection space are input into the evaluation model for evaluation; The evaluation model comprehensively evaluates the route according to technical, safety, sustainability and economic evaluation indicators, returns the corresponding evaluation level, and generates corresponding reward and punishment values ​​according to the evaluation results; S3, inputting the relevant parameters of the line selection space into the reinforcement learning model to perform reinforcement learning; The reinforcement learning model consists of a strategy network, a value network and a circuit construction network; The value network receives the reward and punishment values ​​returned by the evaluation model, and calculates the maximum total reward in the future based on the current line selection space and action, and calculates the maximum total reward in the future based on the updated line selection space and action, and updates by calculating the gradient of the value network, and returns the gradient of the action to help the policy network update; The strategy network receives the current line selection space, outputs the current line type parameters, and then sends them to the line construction network; The line construction network constructs the spatial distribution of tunnels in the line selection space according to the generated line type parameters, thereby generating the line selection space for the next section; S4, inputting the line type parameters of the newly generated next section of the line into the line construction network, and generating a new line selection space according to the input parameters; S5, repeating steps S2 to S4 until the optimal design is found in the line selection space.

2. According to claim 1, a method for intelligent line selection of underground roads based on multi-source data through reinforcement learning, characterized in that: Step S1 includes: S11, according to the requirements of the specification, the basic line type is parameterized, and the parameters mainly involved in the parameterization include radius and angle; S12, when constructing the line selection space, the multi-source information parameterization process is based on GIS data, taking into account the surrounding environmental factors. For the types of objects that can be seen intuitively, machine vision technology is used to automatically identify and label the objects in a wide area.

3. According to claim 2, a method for intelligent line selection of underground roads based on multi-source data through reinforcement learning, characterized in that: Step S12 also includes: introducing InSAR data, including ground subsidence rate and time series offset value information, and constructing a digital spatial distribution matrix in the three-dimensional grid space.

4. According to claim 1, a method for intelligent line selection of underground roads based on multi-source data through reinforcement learning, characterized in that: Step S2 also includes: the evaluation model comprehensively evaluates the line according to the technical, safety, sustainability and economic evaluation indicators, and returns the corresponding evaluation level. At the same time, the corresponding reward and punishment values ​​are generated according to the evaluation results to reflect the quality of the line design, reward good designs and punish poor designs; The reward and punishment values ​​are mainly calculated by considering the goal achievement and the evaluation level of the evaluation model through the design of a reward function. The goal achievement mainly reflects the distance between the end of the currently generated route and the target location, and the evaluation level reflects the comprehensive evaluation result of the currently generated route. The design principle of the reward function is to reward routes that are conducive to achieving goals and improving the evaluation level of the evaluation model, and vice versa.

5. According to claim 4, a method for intelligent line selection of underground roads based on multi-source data through reinforcement learning is characterized in that: The evaluation model consists of evaluation indicators and evaluation methods, and is used to evaluate the generated routes; The evaluation indicators include technical indicators, safety indicators, sustainability indicators and economic indicators, forming a multi-level indicator system; The evaluation method adopts an indicator matrix to complete the level division of the lines.

6. According to claim 1, a method for intelligent line selection of underground roads based on multi-source data through reinforcement learning, characterized in that: The step S3 also includes: in the reinforcement learning process, the line selection space and the reinforcement learning model interact with each other through the reward and punishment values; the reinforcement learning model continuously optimizes the parameters of the strategy network to find the optimal line design solution; The reinforcement learning model predicts the future line parameters and selects the optimal line parameters as the parameter values ​​of the next line segment based on the characteristics and reward and penalty values ​​of the current line selection space.

7. According to claim 6, a method for intelligent line selection of underground roads based on multi-source data through reinforcement learning, characterized in that: Reinforcement learning model parameter updating and optimization include: The line construction network constructs the spatial distribution of tunnels in the line selection space according to the generated line type parameters, thereby generating the line selection space for the next section; The strategy network receives the current line selection space, outputs the current line type parameters, and then sends them to the line construction network; The value network receives the reward and punishment value returned by the evaluation model (represented by r), and calculates the maximum total reward in the future (represented by Q) based on the current line selection space and action. At the same time, the value network also calculates the maximum total reward in the future (represented by Q') based on the updated line selection space and action. The gradient of the value network is calculated and updated through formula 1, and the gradient of the action is also returned to help the policy network update. GradQ=Q+r-Q' (1) In the formula, GradQ represents the updated gradient of the value network. Q represents the maximum total reward in the future calculated based on the current line selection space and actions. r represents the reward and punishment value returned by the evaluation model, Q' represents the maximum total reward in the future calculated based on the updated line selection space and actions.

8. According to claim 1, the method for intelligent line selection of underground roads based on multi-source data through reinforcement learning is characterized by: Step S5 also includes: reinforcement learning is updated by continuously accepting the reward and punishment values ​​of environmental feedback, and continuously optimizing the parameters of the strategy network until the end point is reached and the reward value reaches the expectation or converges, and the route generated by the final strategy network is the optimal route.

Citation Information

Patent Citations

  • Path planning method based on heuristic deep reinforcement learning

    CN112325897A

  • Urban underground space development and construction platform based on geological big data

    CN116432270A