Shield tunnel parameterization three-dimensional modeling method based on reinforcement learning double-layer optimization

Through the reinforcement learning double-layer optimization method, the design parameters and layout of shield tunnel segments are optimized simultaneously, which solves the problem of resource waste in shield tunnel design and achieves efficient and low-error design and modeling.

CN120671264AActive Publication Date: 2025-09-19TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510934479.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-19
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The existing shield tunnel segment design and layout design are separated from each other, resulting in great limitations, inability to achieve double-layer optimization, and serious waste of resources.

Method used

A two-layer optimization method based on reinforcement learning is adopted to optimize the design parameters and layout of shield tunnel segments through the global strategy network and sequential decision network, thereby achieving synchronous optimization of the design parameters and layout of shield tunnel segments.

Benefits of technology

It improves the efficiency of shield tunnel segment design and layout, reduces line errors, and achieves rational utilization of resources and efficient design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671264A_ABST
    Figure CN120671264A_ABST
Patent Text Reader

Abstract

The invention discloses a shield tunnel parameterization three-dimensional modeling method based on reinforcement learning double-layer optimization. The method comprises the steps that planning line parameters are input through a global strategy network, and segment design parameters are output; constructing a typesetting design space according to the segment design parameters and the planned route; the sequential decision network outputs the position of the next segment based on the typesetting space state; a reward and punishment value is determined by calculating a deviation value between a segment axis and a design line, and a sequence decision-making layer network is trained; iteratively and circularly finishing line typesetting and calculating an accumulated reward and punishment value; and training a global parameter layer network and updating global parameters according to the accumulated reward and punishment values, repeating the above process when convergence is not carried out until a convergence condition is reached, finally obtaining optimized segment parameters and typesetting, and constructing a three-dimensional model. According to the method, double-layer optimization of segment design and typesetting is realized, the design efficiency is improved, and the line deviation is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of shield tunnel parameter modeling, and in particular relates to a shield tunnel parameterized three-dimensional modeling method based on reinforcement learning double-layer optimization. Background Art

[0002] With the expansion and development of urbanization, urban surface space resources are unable to meet the needs of further development, and the development of underground space has become a hot topic. The construction of urban underground shield tunnels, as an indispensable component of underground space development, has experienced rapid development in recent years. Segment design and layout are crucial aspects of shield tunnel construction. Proper segment design and layout based on the planned route can effectively avoid shield route deviation and waste of segment resources.

[0003] Currently, segment design and layout for shield tunnels are typically split into two separate processes. When the wedge shape doesn't meet tunnel curvature requirements, a new wedge shape is often required, resulting in a waste of resources and money. Existing optimization methods can only optimize the layout within given segment parameters, which are closely related to the layout design, thus significantly limiting their application.

[0004] Therefore, current shield segment design and layout design are separated. Parameters such as segment wedge, average segment width, and the number of longitudinal bolts are all closely related to segment layout, and single-pronged layout optimization has limitations. Segment design parameters are often determined by designers based on their own experience, and in actual layout, insufficient wedge is often encountered, necessitating custom segments.

[0005] The current optimization of shield tunnel segments and layout design has great limitations, and it is impossible to achieve double-layer optimization of shield tunnel segments and layout. Summary of the Invention

[0006] The present invention aims to provide a shield tunnel segment design and layout method based on reinforcement learning dual-layer optimization. This method utilizes a dual-layer reinforcement learning optimization algorithm to simultaneously optimize the parameters and layout of shield tunnel segments, fully considering the impact of shield tunnel segment design parameters on the layout design. Furthermore, it can rapidly generate segment design parameters and layout based on a given planned route curve. This method improves the efficiency of shield tunnel segment design and layout, and by optimizing segment design parameters and layout, minimizes the error between the shield tunnel route and the planned design route.

[0007] To achieve the above technical objectives, the present invention provides a parametric 3D modeling method for shield tunnels based on reinforcement learning and double-layer optimization, comprising:

[0008] Input the parameters of the shield tunnel planning route into the global strategy network, and output the shield tunnel segment design parameters;

[0009] Constructing a segment layout design space according to the segment design parameters and the planned route;

[0010] Based on the sequential decision network, the position of the next ring of assembled segments is output according to the state of the segment layout design space;

[0011] Update the position of the next ring segment and calculate the reward and penalty value based on the deviation between the segment axis and the design line. Use the reward and penalty value to train the sequential decision layer network.

[0012] Iterate the loop until the layout of the entire route is completed, and calculate the accumulated reward and punishment values ​​of one round;

[0013] The global parameter layer network is trained and updated based on the accumulated reward and penalty values. If convergence conditions are not met, the parameters of the shield tunnel planning route are re-entered into the global strategy network, and the shield tunnel segment design parameters are output. A subsequent round of layout is performed until the convergence conditions are met and the final optimized segment parameters and layout are obtained.

[0014] Based on the final optimized segment parameters and layout, a 3D model is constructed to obtain the 3D modeling results.

[0015] Preferably, the process of inputting the parameters of the shield tunnel planning route into the global strategy network and outputting the obtained shield tunnel segment design parameters includes:

[0016] Randomly generate a shield tunnel planning route and obtain a sequence of shield tunnel planning route design parameters; the basic design parameters include the three-dimensional curvature of the design route and the curvature change rate of the transition curve;

[0017] Arranging the basic design parameters according to the linear arrangement order of the planning curve to obtain a design parameter sequence of the planning curve;

[0018] The global strategy network is a neural network composed of fully connected layers, and the design parameter sequence is input into the global strategy network to calculate the corresponding segment design parameters;

[0019] The segments adopt universal wedge rings, and the segment design parameters include: average width of the lining ring, wedge amount, and number of longitudinal bolts.

[0020] Preferably, the process of constructing the segment layout design space according to the segment design parameters and the planned route includes:

[0021] The layout line is parameterized and expressed as a sequence consisting of axis vector, annular vector and corresponding node coordinates.

[0022] The planned route is split into a plane combination and a vertical combination, wherein the plane combination is further divided into basic line shapes of straight lines, circular curves and transition curves, and the longitudinal line shape of the vertical combination is divided into basic line shapes of circular curves and straight lines, and a parametric equation about the length of the plane line shape is constructed.

[0023] Preferably, in the process of outputting the position of the next ring of assembled segments according to the state of the segment layout design space based on the sequential decision network, the state parameters of the segment layout design space include: segment design parameters, the relative position of the current position of the lining ring and the planned route, the axis vector and the annular vector of the lining ring at the current position, and the direction vector corresponding to the planned route segment;

[0024] The sequential decision network is a neural network composed of fully connected layers, which inputs the state parameters of the typesetting space and outputs the splicing position of the lower ring lining ring.

[0025] Preferably, the position of the next ring segment is updated, and the reward and penalty values ​​are calculated based on the deviation between the segment axis and the design line. The process includes:

[0026] The axis vectors at different splicing positions are calculated based on the average width of the lining ring, the amount of wedge and the number of longitudinal bolts, and the typesetting axis vector transfer table is obtained;

[0027] Calculating the nodes and direction vectors of the next lining ring by using the splicing position and the layout axis vector table, and updating the layout line sequence of the segment layout design space;

[0028] Calculate the axis deviation value based on the parameter equation of the planned route and the equation of the current axis;

[0029] Reward and penalty values ​​are calculated based on the deviation value, wherein an axis that deviates from the planned route is penalized, and an axis that reduces the deviation value or maintains a deviation value less than a preset threshold value is rewarded.

[0030] Preferably, the sequential decision layer network includes a sequential decision network, a value network, a target sequential decision network and a target value network;

[0031] The process of training the sequential decision layer network through reward and punishment values ​​includes:

[0032] The sequential decision network, the value network, the target sequential decision network and the target value network are updated in sequence according to the reward and punishment values.

[0033] Preferably, the process of iterating and looping until the layout of the entire line is completed and calculating the accumulated reward and penalty values ​​for one round includes:

[0034] When the cumulative reward value of the current typesetting line reaches stable convergence, the global parameter layer network is trained and the global parameters are updated according to the accumulated reward and penalty values;

[0035] When convergence is not reached, the sequential decision layer is updated and the spatial layout space is reset.

[0036] Preferably, the global parameter layer network includes a global parameter network, a global value network, a target global parameter network and a target global value network;

[0037] The process of training the global parameter layer network and updating the global parameters based on the accumulated reward and penalty values ​​includes:

[0038] Update each network based on the average lining ring deviation value optimized for the entire line;

[0039] When the optimized average lining ring deviation value exceeds the set threshold, the global parameter layer is updated, and the parameters of the shield tunnel planning route are input into the global strategy network, and the shield tunnel segment design parameters are output;

[0040] When the optimized average lining ring deviation value reaches stable convergence, the segment parameters and layout are constructed according to the arbitrarily input planning line.

[0041] Preferably, the process of constructing a 3D model based on the final optimized segment parameters and layout includes:

[0042] According to the segment design parameters, the basic lining ring is constructed by constructing cylinders for segmentation;

[0043] According to the layout design parameters, the lining rings are assembled in space through translation and rotation, and finally a visual 3D model of the tunnel is constructed.

[0044] Preferably, the segment design parameters include the wedge shape, width, thickness, diameter of the lining ring, and the segment division position of each ring segment;

[0045] The layout design parameters include the lining ring center position coordinates, axis vectors, and rotation angle vectors.

[0046] Compared with the prior art, the present invention has the following advantages and technical effects:

[0047] The present invention optimizes the segment design parameters and the lining ring layout parameters simultaneously, and can adjust the segment parameters during the layout process to achieve double-layer optimization, thereby achieving a more reasonable segment and layout design.

[0048] Traditionally, segment design and lining layout are performed separately. In practice, inappropriate segment wedges often require custom-made segments with different wedges, resulting in wasted resources. Simultaneous optimization of segment parameters and lining layout allows for adjustment of segment parameters during the layout process, achieving dual-level optimization.

[0049] The present invention uses a two-layer optimization method to randomly initialize planned routes, allowing the global parameter layer to learn the characteristics of different planned routes. After sufficient optimization and learning, the two-layer reinforcement learning model can output optimized parameters for any planned route. The two-layer reinforcement learning process of the present invention can achieve different types of planned route optimization for different types of planned routes. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:

[0051] Figure 1 Schematic diagram of a method flow in an embodiment of the present invention;

[0052] Figure 2 A schematic diagram of the process of reinforcement learning dual-layer optimization according to an embodiment of the present invention;

[0053] Figure 3 A schematic diagram of the learning and updating process of the reinforcement learning model according to an embodiment of the present invention;

[0054] Figure 4 Schematic diagram of typesetting design parameters according to an embodiment of the present invention. DETAILED DESCRIPTION

[0055] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0056] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0057] like Figure 1-4 As shown, this embodiment provides a parametric three-dimensional modeling method for shield tunnels based on reinforcement learning double-layer optimization. This method can realize the automatic optimization design of shield tunnel segments and layouts. The invention can improve the efficiency of segment design and layout, and realize low-error, high-efficiency optimization design of shield tunnels.

[0058] like Figure 1 As shown, the shield tunnel segment design and layout method based on reinforcement learning double-layer optimization of this embodiment specifically includes the following steps S1 to S7.

[0059] S1, the global strategy network inputs the parameters of the current random planning line and outputs the design parameters of the shield tunnel segment;

[0060] S2. Construct segment layout design space based on segment design parameters and planned routes;

[0061] S3, the sequential decision network outputs the position of the next ring assembly segment based on the state of the layout space;

[0062] S4. The layout design space updates the position of the next ring segment and calculates the reward and penalty value based on the deviation between the segment axis and the design line;

[0063] S5, the sequential decision layer network is trained based on the reward and punishment values;

[0064] S6. Repeat S2-S5 until the layout of the entire route is completed, and calculate the accumulated reward and punishment value of one round;

[0065] S7. The global parameter layer network is trained and the global parameters are updated based on the accumulated reward and penalty values. If the convergence condition is not met, it returns to step S1 and performs a subsequent round of layout. After the convergence condition is met, the final optimized segment parameters and layout can be obtained.

[0066] S8. Based on the parameters optimized and outputted in S7, a three-dimensional model is constructed using common modeling methods.

[0067] The dual-layer parameter reinforcement learning optimization method described in this embodiment includes a segment parameter optimization layer and a layout design optimization layer. Specifically, shield tunnel segment design parameters are set as global parameters, and shield segment layout parameters are set as parameters in a sequential optimization layer. The deviation between the axis and the planned design line is used as the optimization target, achieving dual optimization of shield tunnel segments and layout design. The dual-layer parameter reinforcement learning optimization model includes two layers of networks: a global parameter layer network and a sequential decision layer network. The main optimization steps include: the global strategy network inputs the parameters of the current planned route and outputs the design parameters of the shield tunnel segments; the segment layout design space is constructed based on the segment design parameters and the planned route; the sequential decision network outputs the position of the next ring of assembled segments based on the state of the layout space; the layout design space updates the position of the next ring of segments, and calculates the reward and penalty values ​​based on the deviation value between the segment axis and the design route; the sequential decision network is trained based on the reward and penalty values; the above steps are repeated until the layout of the complete route is completed, and the accumulated reward and penalty values ​​of one round are calculated; the global strategy network is trained based on the accumulated reward and penalty values ​​and updates the global parameters for the subsequent round of layout; finally, after the model reaches convergence, it can realize the optimization of segment design and layout for any given design route, and construct a visual three-dimensional model.

[0068] Furthermore, S1, the global strategy network inputs the parameters of the current random planning line and outputs the shield tunnel segment design parameters, including:

[0069] A shield tunnel planning route is randomly generated, and a design parameter sequence for the shield tunnel planning route is obtained. The basic design parameters include the three-dimensional curvature of the design route and the curvature change rate of the transition curve. These basic design parameters are arranged according to the linear arrangement order of the planning curve to form the design parameter sequence for the planning curve.

[0070] The global strategy network is a neural network composed of fully connected layers. The design parameter sequence is input into the global strategy network to calculate the corresponding segment design parameters. The segments use universal wedge rings. The segment design parameters mainly include the average width of the lining ring, the amount of wedge, and the number of longitudinal bolts.

[0071] Specifically, the parameters of the random planning circuit and the global policy network described in S1 include:

[0072] The planned line is split into a combination of basic linear shapes according to the horizontal and vertical directions, and then the three-dimensional curvature radius, length and other parameters of the corresponding curves are calculated. The curvature radius of the transition curve and the curvature radius of the changing horizontal and vertical combination are averaged according to the length. The curvature radius of the straight line is 10,000 meters. The curvature radius is sequentially formed in the form of [(r1,l1),(r2,l2),(r3,l3),…,(r n ,l n )], representing the characteristic value of the current planned route.

[0073] like Figure 2 As shown, the global strategy network is constructed as a fully connected neural network. The input layer receives the eigenvalues ​​of the planned route. To ensure the network's adaptability to planned routes of varying configurations, a large amount of input data is recommended. For shorter planned routes, straight line segments can be used to ensure consistent input data size. The output layer contains the segment design parameters for the corresponding planned route. The lining ring adopts a universal double-sided wedge ring. The main parameters are the segment wedge shape, average width, and number of longitudinal bolts.

[0074] Furthermore, S2 constructs a segment layout design space based on the segment design parameters and the planned route, including:

[0075] The layout line is parameterized and represented as a sequence consisting of an axis vector, a circular vector, and corresponding node coordinates.

[0076] The planned route is divided into two parts: plane combination and vertical combination. The plane combination is further divided into basic line types such as straight lines, circular curves and transition curves, and the vertical line type is divided into basic line types such as circular curves and straight lines. The parametric equation about the length of the plane line type is constructed through mathematical methods.

[0077] Specifically, the typesetting design space described in S2 includes the parametric typesetting sequence of the lining ring and the parametric equations of the planning line:

[0078] Lining ring parameterized typesetting sequence: Figure 4 As shown, the vector through the lining ring axis and the circumferential vector Indicates the spatial posture of the lining ring, and the coordinates of the center point of the lining ring splicing surface represent the absolute position (x, y, z) of the lining ring in space.

[0079] The spatial information of a single lining ring can be expressed based on the above three main parameters. A complete lining ring layout can be expressed as sequence.

[0080] Construct parametric equations for the planned routes: Split the planned routes into plane combinations and vertical combinations. The plane combinations consist of straight lines, circular curves, and transition curves, while the vertical combinations consist of straight lines and circular curves. Mathematical methods can be used to construct parametric equations for the length of the plane routes for each line type.

[0081] Furthermore, S3, the sequential decision network outputs the position of the next ring assembly segment based on the state of the layout space, including:

[0082] The status parameters of the layout space include: segment design parameters, the relative position of the lining ring between the current position and the planned route, the axis vector and annular vector of the lining ring at the current position, and the direction vector corresponding to the planned route segment.

[0083] The sequential decision network is a neural network composed of fully connected layers, which inputs the state parameters of the typesetting space and outputs the splicing position of the lower ring lining ring.

[0084] Specifically, the typesetting space state and sequence decision network described in S3 includes:

[0085] Furthermore, the typesetting space state described in S3 includes:

[0086] Segment design parameters, the relative position of the lining ring to the planned route, the axial and circumferential vectors of the lining ring at the current position, and the direction vector corresponding to the planned route segment. The relative position of the lining ring to the planned route can be expressed as a distance and a perpendicular unit vector. The direction vector corresponding to the planned route segment can be obtained by calculating the derivative of the planned route parameter equation at the foot of the perpendicular.

[0087] Furthermore, the sequential decision network described in S3 includes:

[0088] like Figure 2As shown in the figure, the sequential decision network is constructed as a fully connected neural network, in which the input parameter is the layout space state parameter, and the output parameter is the splicing position of the next lining ring. The parameter quantity at the output end is the number of longitudinal bolts in the global parameter. According to the preset range of variation of the longitudinal bolts, multiple output heads are pre-set for the sequential decision layer to meet the needs of different splicing conditions.

[0089] Furthermore, S4, the layout design space updates the position of the next ring segment and calculates the reward and penalty value based on the deviation value between the segment axis and the design line, including:

[0090] The axis vectors at different splicing positions are calculated based on the average width, wedge amount, and number of longitudinal bolts, and a layout axis vector transfer table is constructed. To update the layout space, the node and direction vectors of the next lining ring are calculated based on the splicing position and the layout axis vector table, and the layout line sequence is updated.

[0091] The deviation value of an axis is calculated based on the parametric equation of the planned route and the equation of the current axis. Rewards and penalties are calculated based on the deviation value. Axis that deviate too much from the planned route are penalized, while axes that reduce the deviation value or maintain a small deviation value are rewarded.

[0092] Specifically, S4's layout design space update and reward and penalty value calculation include:

[0093] Calculate the axis vector of the next lining ring. Refer to formulas (1) and (2) to deduce the position vector of the next ring based on the position vector of the current ring and the splicing position. The specific formula is as follows:

[0094]

[0095] Where: They represent the axis vector of the i-th ring, the lining hoop vector and the normal vector of the splicing surface respectively, θ is the angle of the wedge surface, and its relationship with the wedge amount is: α is the rotation angle of the lining ring. When the angle is 180 degrees, the splicing axis is a straight line.

[0096] After the lining ring is spliced, the offset distance of the center of the splicing surface from the planned route and the angle between the axis of the splicing lining ring and the direction of the planned route are calculated. The longer the offset distance or the larger the angle, the greater the penalty value obtained by the sequential decision layer network. Conversely, the greater the reward value obtained.

[0097] S5, the sequential decision layer network is trained based on the reward and punishment values;

[0098] Specifically, the training of the sequential decision layer network includes:

[0099] like Figure 3As shown in the figure, the state value network is updated by calculating the deviation value through the reward and punishment value to more accurately predict the current lining ring splicing state value. The sequential decision network is updated in reverse based on the value estimated by the state value network to optimize the strategy to increase the reward value obtained by the subsequent strategy.

[0100] Furthermore, S6, S2-S5 are repeated until the layout of the entire route is completed, and the accumulated reward and penalty values ​​of one round are calculated, including:

[0101] When the cumulative reward value of the current typesetting line reaches stable convergence, the process proceeds to the next step S7.

[0102] When convergence is not reached, the sequential decision layer is updated, returning to step S2 and resetting the spatial layout space.

[0103] Specifically, completing route layout and calculating accumulated rewards and penalties includes:

[0104] There are two types of layout termination conditions: one is completing the layout of the entire route along the planned route within the required deviation value, and the other is exceeding the critical value of the layout's allowable deviation. When layout is terminated due to the second type of condition, the environment must be reset for the next round of layout. When layout is terminated due to the first type of condition and the accumulated reward and penalty values ​​have reached a stable level, the layout is considered to have reached optimization for that set of segment design parameters.

[0105] S7. The global parameter layer network is trained and the global parameters are updated based on the accumulated reward and penalty values. If the convergence condition is not met, it returns to step S1 and performs a subsequent round of layout. After the convergence condition is met, the final optimized segment parameters and layout can be obtained.

[0106] Specifically, the global parameter layer network update and convergence conditions include:

[0107] like Figure 3 As shown in the figure, the state value network is updated through the average reward and penalty value of the lining ring of the sequential decision layer to better predict the state value based on the planned line characteristics and segment design parameters. The global parameter strategy network is updated through the state value predicted by the state value network to optimize the output segment design parameters to obtain a higher average reward and penalty value of the lining ring.

[0108] The convergence condition of the global decision layer is that the average reward and penalty value of the lining ring of the model gradually reaches a stable value, indicating that the segment design parameters given by the global parameter layer network have reached the optimal value.

[0109] Furthermore, the reinforcement learning two-layer optimization network updates include updates to the sequential decision layer (S5) and the global parameter layer (S7). The sequential decision layer includes a sequential decision network, a value network, a target sequential decision network, and a target value network, and each network is updated sequentially based on the reward and penalty values. The global parameter layer includes a global parameter network, a global value network, a target global parameter network, and a target global value network, and each network is updated based on the average lining ring deviation value optimized for the entire line section.

[0110] When the optimized average lining ring deviation exceeds the set threshold, the global parameter layer is updated and the process returns to step S1. When the optimized average lining ring deviation reaches stable convergence, the two-layer optimization network has been updated and can construct appropriate segment parameters and layout based on any input planned route.

[0111] S8. Based on the parameters optimized and outputted in S7, a three-dimensional model is constructed using common modeling methods.

[0112] Furthermore, the optimized output parameters used in modeling include:

[0113] Segment design parameters: wedge shape, width, thickness, diameter of the lining ring, and the division position of each ring segment.

[0114] Layout design parameters: lining ring center position coordinates, axis vector, angle vector.

[0115] The model can be constructed using common modeling methods based on the above basic parameters.

[0116] Based on the segment design parameters, the basic lining rings can be constructed by constructing cylinders for segmentation. Based on the layout design parameters, the lining rings can be assembled in space through translation and rotation. Finally, a visual 3D model of the tunnel is constructed.

[0117] Taking the modeling method provided by the Python trimesh library as an example, the lining ring is first constructed based on the segment design parameters using the cylinder method in trimesh.creation and Boolean operations. Furthermore, based on the segment angles, the box method in trimesh.creation constructs half-planes to segment the lining ring and construct the wedge-shaped beams. A spatial rotation matrix is ​​then calculated based on the spatial pose parameters. The apply_transform method in trimesh adjusts the lining ring to the correct pose. Finally, the apply_translation method in trimesh reassembles the lining ring to the correct position based on the spatial position coordinates. By constructing a Python script, this process can be repeated to achieve rapid parametric modeling of shield tunnels.

[0118] The shield tunnel segment design and layout method based on reinforcement learning and dual-layer optimization in this embodiment can achieve automatic design and layout of shield tunnel segments, which is beneficial for improving the efficiency of shield tunnel segment and layout design and reducing segment layout deviation, providing an efficient and convenient design method for shield tunnel construction. Furthermore, based on dual-layer reinforcement learning, parameters at different levels can be optimized simultaneously. During the reinforcement learning process, by continuously learning the layout of different planned routes, the final average lining ring deviation value is reduced, thereby achieving segment design and layout optimization for different planned routes, and achieving rapid segment design and layout optimization for any given planned route.

[0119] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A parametric 3D modeling method for shield tunnels based on reinforcement learning double-layer optimization, characterized in that: include: Input the parameters of the shield tunnel planning route into the global strategy network, and output the shield tunnel segment design parameters; Constructing a segment layout design space according to the segment design parameters and the planned route; Based on the sequential decision network, the position of the next ring of assembled segments is output according to the state of the segment layout design space; Update the position of the next ring segment and calculate the reward and penalty value based on the deviation between the segment axis and the design line. Use the reward and penalty value to train the sequential decision layer network. Iterate the loop until the layout of the entire route is completed, and calculate the accumulated reward and punishment values ​​of one round; The global parameter layer network is trained and updated based on the accumulated reward and penalty values. If convergence conditions are not met, the parameters of the shield tunnel planning route are re-entered into the global strategy network, and the shield tunnel segment design parameters are output. A subsequent round of layout is performed until the convergence conditions are met and the final optimized segment parameters and layout are obtained. Based on the final optimized segment parameters and layout, a 3D model is constructed to obtain the 3D modeling results.

2. The method according to claim 1, characterized in that The process of inputting the parameters of the shield tunnel planning route into the global strategy network and outputting the shield tunnel segment design parameters includes: Randomly generate a shield tunnel planning route and obtain a sequence of shield tunnel planning route design parameters; the basic design parameters include the three-dimensional curvature of the design route and the curvature change rate of the transition curve; Arranging the basic design parameters according to the linear arrangement order of the planning curve to obtain a design parameter sequence of the planning curve; The global strategy network is a neural network composed of fully connected layers, and the design parameter sequence is input into the global strategy network to calculate the corresponding segment design parameters; The segments adopt universal wedge rings, and the segment design parameters include: average width of the lining ring, wedge amount, and number of longitudinal bolts.

3. The method according to claim 1, characterized in that The process of constructing the segment layout design space based on the segment design parameters and planned routes includes: The layout line is parameterized and expressed as a sequence consisting of axis vector, annular vector and corresponding node coordinates. The planned route is split into a plane combination and a vertical combination, wherein the plane combination is further divided into basic line shapes of straight lines, circular curves and transition curves, and the longitudinal line shape of the vertical combination is divided into basic line shapes of circular curves and straight lines, and a parametric equation about the length of the plane line shape is constructed.

4. The method according to claim 1, wherein In the process of outputting the position of the next ring of assembled segments according to the state of the segment layout design space based on the sequential decision network, the state parameters of the segment layout design space include: segment design parameters, the relative position of the current lining ring position and the planned route, the axis vector and the annular vector of the current lining ring position, and the direction vector corresponding to the planned route segment; The sequential decision network is a neural network composed of fully connected layers, which inputs the state parameters of the typesetting space and outputs the splicing position of the lower ring lining ring.

5. The method according to claim 1, wherein The process of updating the position of the next ring segment and calculating the bonus and penalty value based on the deviation between the segment axis and the design line includes: The axis vectors at different splicing positions are calculated based on the average width of the lining ring, the amount of wedge and the number of longitudinal bolts, and the typesetting axis vector transfer table is obtained; Calculating the nodes and direction vectors of the next lining ring by using the splicing position and the layout axis vector table, and updating the layout line sequence of the segment layout design space; Calculate the axis deviation value based on the parameter equation of the planned route and the equation of the current axis; Reward and penalty values ​​are calculated based on the deviation value, wherein an axis that deviates from the planned route is penalized, and an axis that reduces the deviation value or maintains a deviation value less than a preset threshold value is rewarded.

6. The method according to claim 1, characterized in that The sequential decision layer network includes a sequential decision network, a value network, a target sequential decision network and a target value network; The process of training the sequential decision layer network through reward and punishment values ​​includes: The sequential decision network, the value network, the target sequential decision network and the target value network are updated in sequence according to the reward and punishment values.

7. The method according to claim 1, characterized in that The process of iterating until the layout of the entire route is completed and the accumulated reward and penalty values ​​for one round are calculated includes: When the cumulative reward value of the current typesetting line reaches stable convergence, the global parameter layer network is trained and the global parameters are updated according to the accumulated reward and penalty values; When convergence is not reached, the sequential decision layer is updated and the spatial layout space is reset.

8. The method according to claim 1, characterized in that The global parameter layer network includes a global parameter network, a global value network, a target global parameter network and a target global value network; The process of training the global parameter layer network and updating the global parameters based on the accumulated reward and penalty values ​​includes: Update each network based on the average lining ring deviation value optimized for the entire line; When the optimized average lining ring deviation value exceeds the set threshold, the global parameter layer is updated, and the parameters of the shield tunnel planning route are input into the global strategy network, and the shield tunnel segment design parameters are output; When the optimized average lining ring deviation value reaches stable convergence, the segment parameters and layout are constructed according to the arbitrarily input planning line.

9. The method according to claim 1, characterized in that Based on the final optimized segment parameters and layout, the process of building a 3D model includes: According to the segment design parameters, the basic lining ring is constructed by constructing cylinders for segmentation; According to the layout design parameters, the lining rings are assembled in space through translation and rotation, and finally a visual 3D model of the tunnel is constructed.

10. The method according to claim 9, characterized in that The segment design parameters include the wedge shape, width, thickness, diameter of the lining ring, and the segment division position of each ring segment; The layout design parameters include the lining ring center position coordinates, axis vectors, and rotation angle vectors.

Citation Information

Patent Citations

  • Three-dimensional parameterized simulation design method for intelligent bed structure

    CN119272539A

  • Underground expressway reinforcement learning intelligent line selection method based on GIS and InSAR and computer system

    CN119537502A