Parameterized three-dimensional modeling method for shield tunnel based on reinforcement learning double-layer optimization

By using a two-layer optimization method based on reinforcement learning, the design parameters and layout of shield tunnel segments are optimized simultaneously, solving the problem of resource waste in shield tunnel design and achieving efficient and low-error design.

CN120671264BActive Publication Date: 2026-03-03TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510934479.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2026-03-03
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The existing shield tunnel segment design and layout design are disconnected, resulting in significant limitations, making it impossible to achieve dual-layer optimization, and causing serious waste of resources.

Method used

A two-layer optimization method based on reinforcement learning is adopted to optimize the design parameters and layout of shield tunnel segments through a global policy network and a sequential decision network, thereby achieving synchronous optimization of segment design parameters and layout.

Benefits of technology

It improved the efficiency of shield tunnel segment design and layout, reduced track errors, and achieved rational utilization of resources and efficient design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671264B_ABST
    Figure CN120671264B_ABST
Patent Text Reader

Abstract

The application discloses a shield tunnel parameterized three-dimensional modeling method based on reinforcement learning double-layer optimization, which comprises the following steps: inputting planning line parameters through a global policy network, and outputting segment design parameters; constructing a layout design space according to the segment design parameters and the planning line; outputting the position of the next ring segment based on the layout space state by a sequential decision network; determining a reward and punishment value by calculating the deviation value of the segment axis and the design line, and training the sequential decision layer network; completing line layout and calculating the cumulative reward and punishment value through an iterative cycle; training the global parameter layer network according to the cumulative reward and punishment value and updating the global parameters, repeating the above process until the convergence condition is reached, and finally obtaining the optimized segment parameters and layout, and constructing a three-dimensional model. The application realizes double-layer optimization of segment design and layout, improves the design efficiency, and reduces the line deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of shield tunnel parameter modeling, and particularly relates to a parametric 3D modeling method for shield tunnels based on reinforcement learning-based two-layer optimization. Background Technology

[0002] With the expansion and development of urbanization, urban surface space resources can no longer meet the needs of further development, making the development of underground space a hot topic. As an indispensable part of underground space development, the construction of urban underground shield tunnels has seen rapid development in recent years. The design and layout of tunnel segments is a crucial aspect of shield tunnel construction. Reasonable design and layout of segments according to the planned route can effectively avoid deviations in the shield tunnel line and waste of segment resources.

[0003] Currently, the design and layout of tunnel segments for shield tunnels are typically separated into two separate processes. When the wedge shape does not meet the tunnel curvature requirements, it is often necessary to design segments with new wedge shapes, resulting in a waste of resources and money. Existing optimization methods can only optimize the layout under given segment parameters, which are closely related to the design of the layout, thus limiting the optimization of the layout.

[0004] Therefore, current shield tunnel segment design and layout design are disconnected. Parameters such as segment wedge shape, average segment width, and number of longitudinal bolts are all closely related to segment layout, and optimization of layout alone has limitations. Furthermore, the design parameters for segments are often determined by designers based on their own experience, which frequently leads to situations where the wedge shape is insufficient and specially made segments are required during actual layout.

[0005] Current methods for optimizing shield tunnel segments and layout design have significant limitations, making it impossible to achieve dual-layer optimization of shield tunnel segments and layout. Summary of the Invention

[0006] The purpose of this invention is to provide a shield tunnel segment design and layout method based on reinforcement learning-based two-layer optimization. Through a two-layer reinforcement learning optimization algorithm, the method simultaneously optimizes the shield tunnel segment parameters and layout, fully considering the influence of the shield tunnel segment design parameters on the layout design. Furthermore, it can rapidly generate the segment design parameters and layout based on the given planned route curve. This improves the efficiency of shield tunnel segment design and layout, and by optimizing the segment design parameters and layout, minimizes the error between the shield tunnel route and the planned design route.

[0007] To achieve the above technical objectives, this invention provides a parametric 3D modeling method for shield tunnels based on reinforcement learning-based two-layer optimization, comprising:

[0008] Input the parameters of the shield tunnel planning route into the global strategy network, and output the segment design parameters of the shield tunnel.

[0009] Construct a segment layout design space based on the segment design parameters and planned route;

[0010] Based on the sequential decision network, the position of the next assembled segment is output according to the state of the segment layout design space;

[0011] Update the position of the next ring segment, and calculate the reward and penalty value based on the deviation between the segment axis and the designed route. Use the reward and penalty value to train the sequential decision layer network.

[0012] The iteration loop continues until the layout of the complete route is completed, and the cumulative reward and penalty value for one round is calculated.

[0013] The global parameter layer network is trained and the global parameters are updated based on the accumulated reward and punishment values. If the convergence condition is not met, the parameters of the shield tunnel planning route are re-input into the global strategy network, and the segment design parameters of the shield tunnel are obtained. Then, a subsequent round of layout is performed until the convergence condition is met, and the final optimized segment parameters and layout are obtained.

[0014] Based on the final optimized segment parameters and layout, a 3D model is constructed to obtain the 3D modeling results.

[0015] Preferably, the process of inputting the parameters of the planned shield tunnel route into the global strategy network and outputting the segment design parameters of the shield tunnel includes:

[0016] Randomly generate the planned route of the shield tunnel and obtain the design parameter sequence of the planned route; among which, the basic design parameters include the three-dimensional curvature of the design route and the rate of change of curvature of the transition curve;

[0017] The basic design parameters are arranged according to the linear arrangement order of the planning curve to obtain the design parameter sequence of the planning curve;

[0018] The global policy network is a neural network composed of fully connected layers. The design parameter sequence is input into the global policy network to calculate the corresponding segment design parameters.

[0019] The segments use universal wedge rings, and the segment design parameters include: the average width of the lining ring, the amount of wedges, and the number of longitudinal bolts.

[0020] Preferably, the process of constructing the segment layout design space based on the segment design parameters and planned route includes:

[0021] The layout lines are parameterized and represented as a sequence composed of axis vectors, circumferential vectors, and corresponding node coordinates.

[0022] The planned route is divided into planar combinations and vertical combinations. The planar combinations are further divided into basic line types of straight lines, circular curves, and transition curves. The vertical combinations are divided into basic line types of circular curves and straight lines. Parametric equations about the length of the planar line types are constructed.

[0023] Preferably, based on the sequential decision network, during the process of outputting the position of the next assembled segment according to the state of the segment layout design space, the state parameters of the segment layout design space include: segment design parameters, the relative position of the current position of the lining ring to the planned route, the axial vector and circumferential vector of the lining ring at the current position, and the direction vector of the corresponding planned route segment;

[0024] The sequential decision network is a neural network composed of fully connected layers. It takes the state parameters of the layout space as input and outputs the splicing position of the lower ring lining ring.

[0025] Preferably, the position of the next ring segment is updated, and a reward or penalty value is calculated based on the deviation between the segment axis and the designed route. The process includes:

[0026] Calculate the axis vector at different splicing positions based on the average width of the lining ring, the amount of wedge, and the number of longitudinal bolts, and obtain the layout axis vector transfer table;

[0027] The nodes and direction vectors of the next lining ring are calculated using the splicing position and the layout axis vector table, and the layout line sequence of the segment layout design space is updated.

[0028] Calculate the deviation of the axis based on the parametric equation of the planned route and the equation of the current axis;

[0029] The reward or penalty value is calculated based on the deviation value. A deviation from the planned route axis is penalized, while reducing the deviation value or maintaining the axis below a preset deviation value threshold results in a reward value.

[0030] Preferably, the sequential decision-making layer network includes a sequential decision-making network, a value network, a target sequential decision-making network, and a target value network;

[0031] The process of training a sequential decision-making network using reward and punishment values ​​includes:

[0032] The sequential decision network, value network, target sequential decision network, and target value network are updated sequentially based on the reward and penalty values.

[0033] Preferably, the process of iterating until the complete route layout is completed and calculating the cumulative reward / penalty value for one round includes:

[0034] When the cumulative reward value of the current layout path reaches stable convergence, the global parameter layer network is trained and the global parameters are updated based on the cumulative reward and penalty value.

[0035] If convergence is not achieved, update the order decision layer and reset the spatial layout space.

[0036] Preferably, the global parameter layer network includes a global parameter network, a global value network, a target global parameter network, and a target global value network;

[0037] The process of training the global parameter layer network and updating the global parameters based on the accumulated reward and penalty values ​​includes:

[0038] Each network is updated based on the optimized average lining ring deviation value for the entire line.

[0039] When the optimized average lining ring deviation value exceeds the set threshold, the global parameter layer is updated, and the parameters of the shield tunnel planning route are input into the global strategy network, and the segment design parameters of the shield tunnel are output.

[0040] When the optimized average lining ring deviation value reaches stable convergence, the segment parameters and layout are constructed based on any input planned route.

[0041] Preferably, the process of constructing a 3D model based on the final optimized segment parameters and layout includes:

[0042] Based on the segment design parameters, the basic lining ring is obtained by constructing cylinders to divide the segment.

[0043] Based on the layout design parameters, the lining ring is assembled in space by translation and rotation, and finally a visual 3D model of the tunnel is constructed.

[0044] Preferably, the segment design parameters include the wedge shape, width, thickness, diameter, and segmentation position of each ring segment of the lining ring;

[0045] The layout design parameters include the center position coordinates of the lining ring, the axis vector, and the corner vector.

[0046] Compared with the prior art, the present invention has the following advantages and technical effects:

[0047] This invention simultaneously optimizes the segment design parameters and the lining ring layout parameters, allowing for adjustments to segment parameters during the layout process to achieve dual-layer optimization, thereby enabling a more reasonable segment and layout design.

[0048] Traditionally, segment design and lining ring layout are usually done separately. In actual layout, inappropriate segment wedge amounts often necessitate the custom-made segments with additional wedge amounts, leading to resource waste. However, simultaneously optimizing segment parameters and lining layout allows for adjustments to segment parameters during the layout process, achieving dual-layer optimization.

[0049] This invention employs a two-layer optimization method to randomly initialize planned routes, allowing the global parameter layer to learn the characteristics of different planned routes. After sufficient optimization learning, the two-layer reinforcement learning model can output optimized parameters for any planned route. The two-layer reinforcement learning process of this invention can achieve different types of route optimization for different types of planned routes. Attached Figure Description

[0050] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0051] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention;

[0052] Figure 2 This is a schematic diagram of the reinforcement learning two-layer optimization process according to an embodiment of the present invention;

[0053] Figure 3 This is a schematic diagram illustrating the learning and updating process of the reinforcement learning model according to an embodiment of the present invention.

[0054] Figure 4 This is a schematic diagram of the typesetting design parameters according to an embodiment of the present invention. Detailed Implementation

[0055] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0056] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0057] like Figure 1-4 As shown in the figure, this embodiment provides a parametric 3D modeling method for shield tunnels based on reinforcement learning-based two-layer optimization. This method can realize the automatic optimization design of shield tunnel segments and layout, improve the efficiency of segment design and layout, and achieve low-error, high-efficiency optimization design of shield tunnels.

[0058] like Figure 1 As shown, the shield tunnel segment design and layout method based on reinforcement learning dual-layer optimization in this embodiment specifically includes the following steps S1 to S7.

[0059] S1, Global Strategy Network: Input the parameters of the currently randomly planned route, and output the shield tunnel segment design parameters;

[0060] S2. Construct a segment layout design space based on segment design parameters and planned routes;

[0061] S3. The sequential decision network outputs the position of the next assembled segment based on the state of the layout space.

[0062] S4. The layout design space updates the position of the next ring segment, and calculates the reward and penalty value based on the deviation value between the segment axis and the design line.

[0063] S5. The sequential decision-making layer network is trained based on reward and punishment values;

[0064] S6. Repeat S2-S5 until the layout of the complete route is completed, and calculate the cumulative reward and penalty value for one round.

[0065] S7. The global parameter layer network is trained and updated based on the accumulated reward and punishment values. If the convergence condition is not met, it returns to step S1 and performs a subsequent round of layout. After the convergence condition is met, the final optimized segment parameters and layout can be obtained.

[0066] S8. Based on the parameters optimized from S7, construct a 3D model using common modeling methods.

[0067] The dual-layer parameter reinforcement learning optimization method described in this embodiment includes a segment parameter optimization layer and a layout design optimization layer. Specifically, the shield tunnel segment design parameters are set as global parameters, and the shield tunnel segment layout parameters are set as sequential optimization layer parameters. The deviation between the axis and the planned design route is used as the optimization objective to achieve dual optimization of the shield tunnel segment and layout design. The dual-layer parameter reinforcement learning optimization model includes two layers of networks: a global parameter layer network and a sequential decision layer network. The main optimization steps include: the global policy network takes the parameters of the currently planned route as input and outputs the shield tunnel segment design parameters; a segment layout design space is constructed based on the segment design parameters and the planned route; the sequential decision network outputs the position of the next ring of assembled segments based on the state of the layout space; the layout design space updates the position of the next ring of segments and calculates reward and penalty values ​​based on the deviation between the segment axis and the designed route; the sequential decision network is trained based on the reward and penalty values; the above steps are repeated until the layout of the complete route is completed, and the cumulative reward and penalty values ​​for one round are calculated; the global policy network is trained based on the cumulative reward and penalty values ​​and updates the global parameters for the next round of layout; finally, after the model converges, it can optimize the segment design and layout for any given design route and construct a visualized 3D model.

[0068] Furthermore, S1, the global policy network, takes the parameters of the currently randomly planned route as input and outputs the shield tunnel segment design parameters, including:

[0069] A random shield tunnel planning route is generated, and a sequence of design parameters for the planned route is obtained. The basic design parameters include the three-dimensional curvature of the design route and the rate of change of curvature of the transition curve. The basic design parameters are arranged according to the linear arrangement of the planned curve to form the design parameter sequence for that planned curve.

[0070] The global policy network is a neural network composed of fully connected layers. The design parameter sequence is input into the global policy network to calculate the corresponding segment design parameters. The segments adopt a general wedge ring, and the segment design parameters mainly include: the average width of the lining ring, the wedge amount, and the number of longitudinal bolts.

[0071] Specifically, the parameters and global policy network for the randomized route planning described in S1 include:

[0072] The planned route is divided into combinations of basic alignments based on horizontal and vertical alignments. Then, the three-dimensional radius of curvature, length, and other parameters of the corresponding curves are calculated. The radius of curvature of transition curves and the radius of curvature of varying horizontal and vertical combinations are averaged over length. The radius of curvature of straight lines is 10,000 meters. The radii of curvature are arranged in sequence in the form of [(r1,l1),(r2,l2),(r3,l3),…,(r n ,l n The sequence of )] represents the characteristic values ​​of the currently planned route.

[0073] like Figure 2 As shown, the global policy network is constructed as a fully connected neural network. The input layer takes the feature values ​​of the planned route as input. To ensure the network's adaptability to planned routes with different configurations, the input layer should have a large amount of input data. For shorter planned routes, straight segments can be used to supplement the input data to ensure consistent size. The output layer contains the segment design parameters for the corresponding planned route. The lining ring uses a universal double-sided wedge ring, and the main parameters are the wedge shape of the segment, the average width, and the number of longitudinal bolts.

[0074] Furthermore, S2, based on the segment design parameters and planned route, constructs a segment layout design space, including:

[0075] The layout lines are parameterized and represented as a sequence composed of axis vectors, circumferential vectors, and corresponding node coordinates.

[0076] The planned route is divided into two parts: a horizontal combination and a vertical combination. The horizontal combination is further divided into basic line types such as straight lines, circular curves, and transition curves, while the vertical combination is divided into basic line types such as circular curves and straight lines. Parametric equations about the length of the horizontal line types are constructed using mathematical methods.

[0077] Specifically, the typesetting design space described in S2 includes the parametric typesetting sequence of the lining ring and the parametric equations of the planned route:

[0078] Parametric layout sequence of lining rings: such as Figure 4 As shown, the vector through the lining ring axis and circumferential vector The spatial orientation of the lining ring is represented by the coordinates of the center point of the splicing surface of the lining ring, which represent the absolute position (x, y, z) of the lining ring in space.

[0079] Based on the above three main parameters, the spatial information of a single lining ring can be represented, and a complete lining ring layout can be represented as follows: sequence.

[0080] Constructing parametric equations for the planned route: The planned route is divided into planar combinations and vertical combinations. The planar combinations consist of straight lines, circular curves, and transition curves, while the vertical combinations consist of straight lines and circular curves. Parametric equations for the length of the planar route can be constructed using mathematical methods.

[0081] Furthermore, S3, the sequential decision network, outputs the position of the next assembled segment based on the state of the layout space, including:

[0082] The status parameters of the layout space include: segment design parameters, the relative position of the lining ring to the planned route, the axial vector and circumferential vector of the lining ring at the current position, and the direction vector of the corresponding planned route segment.

[0083] The sequential decision network is a neural network composed of fully connected layers. It takes the state parameters of the layout space as input and outputs the splicing position of the lower ring lining ring.

[0084] Specifically, the typesetting space state and order decision network described in S3 includes:

[0085] Furthermore, the typesetting space state described in S3 includes:

[0086] The segment design parameters include the relative position of the lining ring to the planned route, the axial vector and circumferential vector of the lining ring at the current position, and the direction vector of the corresponding planned route segment. The relative position of the lining ring to the planned route can be expressed as a distance and a unit perpendicular vector. The direction vector of the corresponding planned route segment can be obtained by calculating the derivative of the parametric equation of the planned route at the perpendicular foot position.

[0087] Furthermore, the sequential decision network described in S3 includes:

[0088] like Figure 2As shown, the sequential decision network is constructed as a fully connected neural network, where the input parameters are the layout space state parameters and the output parameters are the splicing positions of the next lining ring. The output parameters are the number of longitudinal bolts in the global parameters. Based on the preset range of variation of longitudinal bolts, multiple output heads are pre-set for the sequential decision layer to meet the needs of different splicing conditions.

[0089] Furthermore, S4 updates the position of the next ring segment in the layout design space and calculates reward and penalty values ​​based on the deviation values ​​between the segment axis and the designed route, including:

[0090] Based on the average width, wedge amount, and number of longitudinal bolts, the axis vectors at different splicing positions are calculated, forming a layout axis vector transfer table. The layout space is updated by calculating the node and direction vectors of the next lining ring using the splicing position and the layout axis vector table, thus updating the layout line sequence.

[0091] The deviation of the axis is calculated based on the parametric equation of the planned route and the equation of the current axis. Rewards and penalties are calculated based on the deviation; excessive deviation from the planned route will result in a penalty, while reducing the deviation or maintaining a small deviation will result in a reward.

[0092] Specifically, S4's layout design space updates and reward / penalty value calculations include:

[0093] Calculate the axial vector of the next lining ring by referring to formulas (1) and (2), and derive the position vector of the next ring based on the position vector of the current ring and the splicing position. The specific formulas are as follows:

[0094]

[0095] In the formula: Let represent the axial vector of the i-th ring, the circumferential vector of the lining, and the normal vector of the splicing surface, respectively. θ is the included angle of the wedge surface, and its relationship with the wedge quantity is as follows: α is the angle of rotation of the lining ring. When the angle is 180 degrees, the splicing axis is a straight line.

[0096] The offset distance of the splicing surface center from the planned route and the angle between the splicing lining ring axis and the planned route direction are calculated after the splicing of the lining ring are completed. The greater the offset distance or the larger the angle, the greater the penalty value obtained by the sequential decision layer network, and vice versa.

[0097] S5. The sequential decision-making layer network is trained based on reward and punishment values;

[0098] Specifically, the training of the sequential decision layer network includes:

[0099] like Figure 3As shown, the State Value Network updates its value by calculating the deviation value through reward and penalty values ​​to more accurately predict the current lining ring splicing state value. The Sequential Decision Network updates its value in reverse based on the value estimated by the State Value Network to optimize the strategy and increase the reward value obtained by subsequent strategies.

[0100] Furthermore, S6, and repeat S2-S5 until the complete layout of the route is completed, and calculate the cumulative reward and penalty value for one round, including:

[0101] When the cumulative reward value of the current layout path reaches stable convergence, proceed to the next step S7.

[0102] If convergence is not achieved, the sequential decision layer is updated, returning to step S2 and resetting the spatial layout space.

[0103] Specifically, completing the route layout and calculating cumulative rewards and penalties includes:

[0104] There are two types of termination conditions for layout: one is completing the layout of the entire route within the deviation requirement along the planned route; the other is exceeding the critical value of the allowable deviation for layout. When layout is terminated due to the second type of condition, the environment needs to be reset for the next round of layout. When layout is terminated due to the first type of condition, and the cumulative reward and penalty value has reached stability, it is considered that the layout has reached the optimal level under the design parameters of that group of tunnel segments.

[0105] S7. The global parameter layer network is trained and updated based on the accumulated reward and punishment values. If the convergence condition is not met, it returns to step S1 and performs a subsequent round of layout. After the convergence condition is met, the final optimized segment parameters and layout can be obtained.

[0106] Specifically, the update and convergence conditions for the global parameter layer network include:

[0107] like Figure 3 As shown, the state value network is updated by the average reward and penalty value of the lining ring in the sequential decision layer to better predict the state value based on the characteristics of the planned line and the segment design parameters. The global parameter strategy network is updated by the state value predicted by the state value network to optimize the output segment design parameters to obtain a higher average reward and penalty value of the lining ring.

[0108] The convergence condition of the global decision layer is that the average reward and penalty value of the lining ring of the model gradually reaches a stable value, indicating that the design parameters of the tunnel segment given by the global parameter layer network have reached the optimal value.

[0109] Furthermore, the update of the reinforcement learning two-layer optimization network includes updating the sequential decision layer (S5) and the global parameter layer (S7). The sequential decision layer includes a sequential decision network, a value network, a target sequential decision network, and a target value network, which are updated sequentially based on reward / penalty values. The global parameter layer includes a global parameter network, a global value network, a target global parameter network, and a target global value network, which are updated based on the average lining ring deviation value for the entire line optimization.

[0110] When the optimized average lining ring deviation value exceeds the set threshold, the global parameter layer is updated and returns to step S1. When the optimized average lining ring deviation value reaches stable convergence, the update of the two-layer optimization network is achieved, which can construct suitable segment parameters and layout based on any input planning route.

[0111] S8. Based on the parameters optimized from S7, construct a 3D model using common modeling methods.

[0112] Furthermore, the parameters used in the modeling optimization output include:

[0113] Segment design parameters: wedge shape, width, thickness, diameter of the lining ring, and the segmentation position of each ring segment.

[0114] Layout design parameters: coordinates of the center position of the lining ring, axis vector, and corner vector.

[0115] The model can be constructed using common modeling techniques based on the above basic parameters.

[0116] Based on the aforementioned segment design parameters, the basic lining ring can be constructed by dividing the tunnel into cylindrical sections. Following the layout design parameters, the lining ring can be assembled in space through translation and rotation. Finally, a visualized 3D model of the tunnel is constructed.

[0117] Taking the modeling method provided by the trimesh library in Python as an example, firstly, based on the segment design parameters, the lining ring is constructed using the `cylinder` method and Boolean operations in `trimesh.creation`. Then, based on the segment angles, the `box` method in `trimesh.creation` is used to construct a half-plane to segment the lining ring, thus achieving segmentation of the lining ring and construction of the wedge beams. Next, the spatial rotation matrix is ​​calculated based on the spatial attitude parameters, and the `apply_transform` method provided by trimesh adjusts the lining ring to the correct attitude. Finally, based on the spatial coordinates, the `apply_translation` method is used to assemble the lining ring into the correct position. By constructing a Python script, the above process can be repeated to achieve rapid parametric modeling of shield tunnels.

[0118] This embodiment of the shield tunnel segment design and layout method based on reinforcement learning and two-layer optimization enables automatic design and layout of shield tunnel segments. This improves the efficiency of shield tunnel segment and layout design and reduces layout deviations, providing an efficient and convenient design method for shield tunnel construction. Furthermore, based on two-layer reinforcement learning, it can simultaneously optimize parameters at different levels. During the reinforcement learning process, by continuously learning the layouts of different planned routes, it reduces the final average lining ring deviation value, thereby achieving segment design and layout optimization for different planned routes and enabling rapid segment design and layout optimization for any given planned route.

[0119] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for parameterized three-dimensional modeling of a shield tunnel based on reinforcement learning double-layer optimization, characterized in that, The method comprises the following steps: inputting parameters of a shield tunnel planning route into a global strategy network to obtain segment design parameters of the shield tunnel; constructing a segment layout design space according to the segment design parameters and the planning route; outputting a position of a next ring segment according to a state of the segment layout design space based on a sequential decision network; updating the position of the next ring segment and calculating a reward and punishment value according to a deviation value of a segment axis and a design route, and training the sequential decision network through the reward and punishment value; iterating until the layout of the complete route is completed, and calculating a cumulative reward and punishment value; training the global parameter network according to the cumulative reward and punishment value and updating the global parameters, inputting the parameters of the shield tunnel planning route into the global strategy network again to obtain the segment design parameters of the shield tunnel when a convergence condition is not reached, and performing subsequent layout, until the final optimized segment parameters and layout are obtained after the convergence condition is reached; constructing a three-dimensional model according to the final optimized segment parameters and layout to obtain a three-dimensional modeling result; The process of inputting the parameters of the shield tunnel planning route into the global strategy network to obtain the segment design parameters of the shield tunnel comprises the following steps: randomly generating a shield tunnel planning route and obtaining a design parameter sequence of the shield tunnel planning route; wherein the basic design parameters include a three-dimensional curvature of the design route and a curvature change rate of a transition curve; arranging the basic design parameters according to a linear arrangement order of the planning curve to obtain the design parameter sequence of the planning curve; wherein the global strategy network is a neural network composed of full connection layers, and the design parameter sequence is input into the global strategy network to calculate corresponding segment design parameters; the segment adopts a general wedge-shaped ring, and the segment design parameters include an average width of a lining ring, a wedge-shaped amount, and a longitudinal bolt quantity; The process of constructing a segment layout design space according to the segment design parameters and the planning route comprises the following steps: parameterizing the layout route and representing it as a sequence composed of an axis vector, a ring vector, and a corresponding node coordinate combination; splitting the planning route into a plane combination and a vertical combination, wherein the plane combination is further divided into basic line types of a straight line, a circular curve, and a transition curve, the vertical combination is divided into basic line types of a circular curve and a straight line, and a parameter equation about a plane line length is constructed; The process of updating the segment layout design space and calculating the reward and punishment value comprises the following steps: calculating an axis vector of a next lining ring, and deriving a position vector of the next ring according to a position vector of a current ring and a splicing position according to formulas (1) and (2); The expressions of formulas (1) and (2) are as follows: In the formula: respectively represent the first ring axis vector, lining ring vector and splicing surface normal vector, is the included angle of the wedge surface, and the relationship with the wedge amount is , is the angle of lining ring rotation, and the splicing axis is a straight line when the angle is 180 degrees.

2. The method according to claim 1, wherein In the process of outputting the position of the next ring segment according to the state of the segment layout design space based on the sequential decision network, the state parameters of the segment layout design space include segment design parameters, a relative position of a current position of a lining ring and the planning route, an axis vector and a ring vector of the current position of the lining ring, and a direction vector corresponding to a segment of the planning route; the sequential decision network is a neural network composed of full connection layers, and the state parameters of the layout space are input to output the splicing position of the next ring lining ring.

3. The method of claim 1, wherein, the process of updating the position of the next segment and calculating the reward and penalty value according to the deviation of the segment axis and the design line comprises: calculating the axis vector at different splicing positions according to the average width, wedge amount and longitudinal bolt quantity of the lining ring to obtain a layout axis vector transfer table; calculating the node and direction vector of the next lining ring through the splicing position and the layout axis vector table, and updating the layout line sequence of the segment layout design space; calculating the deviation value of the axis according to the parametric equation of the planning line and the equation of the current axis; calculating the reward and penalty value according to the deviation value, wherein the axis deviating from the planning line is punished, and the axis with a deviation value less than a preset deviation value threshold is rewarded.

4. The method of claim 1, wherein, the sequential decision layer network comprises a sequential decision network, a value network, a target sequential decision network and a target value network; the process of training the sequential decision layer network through the reward and penalty value comprises: updating the sequential decision network, the value network, the target sequential decision network and the target value network in turn according to the reward and penalty value.

5. The method of claim 1, wherein, the process of iterating until the layout of the complete line is completed and calculating the accumulated reward and penalty value of a round comprises: when the accumulated reward value of the current layout line reaches stable convergence, training the global parameter layer network according to the accumulated reward and penalty value and updating the global parameters; when the convergence is not reached, updating the sequential decision layer and resetting the spatial layout space.

6. The method of claim 1, wherein, the global parameter layer network comprises a global parameter network, a global value network, a target global parameter network and a target global value network; the process of training the global parameter layer network according to the accumulated reward and penalty value and updating the global parameters comprises: updating each network according to the optimized average lining ring deviation value of the whole line; when the optimized average lining ring deviation value exceeds the set threshold, updating the global parameter layer, inputting the parameters of the shield tunnel planning line into the global strategy network and outputting the segment design parameters of the shield tunnel; when the optimized average lining ring deviation value reaches stable convergence, constructing the segment parameters and layout according to any input planning line.

7. The method of claim 1, wherein, the process of constructing a three-dimensional model according to the finally optimized segment parameters and layout comprises: constructing a basic lining ring by segmenting a cylinder according to the segment design parameters; obtaining the assembly of the lining ring in space by translation and rotation according to the layout design parameters, and finally constructing a visual three-dimensional model of the tunnel.

8. The method of claim 7, wherein, the segment design parameters comprise the wedge amount, width, thickness, diameter of the lining ring and the segmentation position of each ring segment; the layout design parameters comprise the center position coordinates of the lining ring, the axis vector and the angle vector.

Citation Information

Patent Citations

  • Three-dimensional parameterized simulation design method for intelligent bed structure

    CN119272539A

  • Underground expressway reinforcement learning intelligent line selection method based on GIS and InSAR and computer system

    CN119537502A