High-precision shield segment splicing pose recognition method

By using multi-source sensor data processing and robust optimization methods, the stability and robustness issues of high-precision shield segment assembly in shield tunnel construction were solved, enabling high-precision pose recognition and assembly under complex working conditions, and improving the system's reliability and automation level.

CN121921375APending Publication Date: 2026-04-24BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2026-02-06
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve high-precision shield segment assembly under complex conditions during shield tunnel construction. In particular, when there are obstructions, missing features, and a high proportion of external points, sensor installation posture and redundancy configuration are limited, and multi-source sensing and high-precision attitude estimation suffer from insufficient stability and robustness.

Method used

Depth data is acquired using multiple sensors, preprocessed to generate clean point clouds, structured geometric features and sparse matching features are extracted, virtual completion is performed using a CAD model, and the initial pose is calculated through multi-factor robust optimization, triggering a restricted six-degree-of-freedom global search to compensate for the pose, and finally outputting the pose matrix and quality indicators.

Benefits of technology

It achieves high-precision and robust pose estimation in environments with occlusion and dust, significantly improving the pose reconstruction success rate, reducing the misfit rate and the frequency of manual intervention, and enhancing system availability and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121921375A_ABST
    Figure CN121921375A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision shield segment splicing pose recognition method, and belongs to the technical field of shield tunnel segment splicing automation, three-dimensional vision measurement and robot pose estimation. The method comprises the steps of realizing high-precision identification of the pose of the segment through multi-source single-frame data access and preprocessing, self-adaptive point cloud cleaning, double-domain feature extraction and CAD virtual completion, multi-factor robust optimization and limited six-degree-of-freedom global search compensation, stably controlling a translation error within + / -2mm and an attitude angle error within + / -0.5 degree, and realizing high-precision identification of the pose of the segment. The high success rate is still kept when the external point proportion is high or the features are lost, and automatic assembly control of the shield tunneling machine is supported by outputting the pose matrix and the quality index. According to the method, high robustness and precision are still kept in industrial sites with high proportions of sheltered and missing features and exterior points, manual intervention and assembly errors are remarkably reduced, and intelligent assembly efficiency and safety are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of automated shield tunnel segment assembly, three-dimensional vision measurement and robot pose estimation technology, and particularly relates to a high-precision shield tunnel segment assembly pose recognition method. Background Technology

[0002] Shield tunneling, as the most widely used fully mechanized method in cut-and-cover tunnel construction, has been widely adopted in urban rail transit, railways, highways, and municipal integrated pipeline projects. The tunneling process relies on the cutterhead breaking up the strata, the muck removal system clearing debris, and the combined support of the machine shell and installed concrete segments to form a ring-shaped lining structure. The accuracy of segment assembly directly affects the closure of the ring joints, waterproofing performance, internal contour dimensional control, and subsequent operational safety. Currently, on-site positioning still generally relies on manual visual inspection, laser indication, or single 3D point cloud registration (such as single ICP) for assistance. Under complex working conditions, it is difficult to consistently achieve millimeter-level translation and sub-angle-level attitude requirements, and its consistency, efficiency, and robustness are insufficient. The on-site environment has typical adverse characteristics: narrow space, dense equipment, and limited passageways, which lead to obstruction or dynamic changes in the field of view; dust, high humidity, mud mist, and oil stains easily adhere to optical / laser windows, causing enhanced scattering, signal loss, or increased ranging noise; uneven lighting, reflection, and flicker lead to unstable visual characteristics; propulsion, hoisting, and hydraulic actions introduce low- to mid-frequency vibrations and micro-displacements, causing slow drift in the internal alignment of the sensor due to external interference; and mutual interference from multiple sources limits the sensor's installation posture and redundancy configuration.

[0003] Under these conditions, multi-source sensing and high-precision attitude estimation face several key technical challenges: Single-frame data dominance: To ensure timing, the applicability of multi-frame temporal fusion and long-term stable filtering is reduced; Insufficient view overlap: Large differences in camera / laser coverage render traditional point cloud denoising strategies relying on multi-view consistency or repeated observation statistics ineffective; Diverse noise and outlier types: Random scattered points, reflection pseudo-points, moving components, mud-attached debris, etc., coexist, and fixed thresholds or single-stage filtering are prone to over-reduction or residual contamination; Missing and incomplete features: Holes, arc edges, and seam lines are occluded or contaminated, and insufficient complete geometric constraints lead to unstable initial values ​​or constraint degradation; Risk of local extrema: When initial pose deviation or outlier ratio is high, classic ICP / point-to-plane registration is prone to getting trapped in local optima and failing to detect failure; Prior exploitation... Shortcomings: Existing CAD models lack systematic point cloud filtering, missing data completion, and search space shrinkage for geometric priors such as hole array layout, arc radius, and seam direction; Parameters are not adaptive: Depth noise increases with distance, and fixed plane / distance thresholds lead to false detections at long distances or missed detections at short distances; Lack of weighted and uncertain outputs: Existing methods mostly output only a single transformation result, lacking residual structure, feature contribution, and attitude confidence, limiting subsequent control and error compensation; Insufficient degradation resistance and fallback mechanisms: When sensors fail intermittently, mirror contamination occurs, or local features are missing, there is a lack of restricted global search and multi-level optimization compensation chains; Conflict between real-time performance and reliability: There is a contradiction between high-precision models (multi-factor fitting, global search) and limited embedded / industrial computing resources, lacking on-demand triggering and hierarchical degradation strategies.

[0004] Furthermore, the industry has set higher comprehensive indicators for automated assembly (positioning error ≤ ±2 mm, attitude angle error ≤ ±0.2°, shorter single-stage time frame, significantly reduced misassembly rate, and high long-term availability), imposing stringent requirements on the stable output of algorithms under environments with multiple disturbances, contamination, and feature degradation. Existing publicly available technologies typically focus on local aspects (such as simple filtering, single-class geometric fitting, or single ICP registration), and lack a systematic solution for the comprehensive constraints of "multiple sources in a single frame + incomplete features + multi-morphological noise + the need for rollback-compatible global compensation." Based on this, it is necessary to propose a high-precision shield tunnel segment pose recognition and assembly method that can: perform depth and density adaptive multi-stage cleaning of point clouds, fuse structured geometry and sparse descriptor features, introduce CAD geometric priors and virtual completion, combine robust multi-factor joint optimization and support restricted six-degree-of-freedom global compensation search, and output attitude uncertainty and quality indicators to meet the needs of practical engineering for accuracy, stability, and traceability in parallel. Summary of the Invention

[0005] To address the aforementioned technical issues, this invention proposes a high-precision shield tunnel segment splicing pose recognition method. This method maintains high robustness and accuracy even in industrial settings with high levels of occlusion, missing features, and external points, significantly reducing manual intervention and assembly errors, and improving intelligent assembly efficiency and safety.

[0006] To achieve the above objectives, the present invention provides a high-precision shield tunnel segment splicing pose recognition method, comprising: Acquire single-frame depth data from multiple source sensors and preprocess the depth data to generate a clean point cloud; Structured geometric features and sparse matching features are extracted based on the pure point cloud, and virtual completion is performed using a CAD model when features are missing. Using the structured geometric features and sparse matching features, the initial pose of the tube segment is calculated through multi-factor robust optimization, and a restricted six-degree-of-freedom global search is triggered to compensate for the pose when the pose quality does not meet the standard. Output the final pose matrix and quality index.

[0007] Optionally, the preprocessing of depth data includes: For each depth pixel, a validity weight factor is calculated based on distance constraints, local gradients, and incident angles to generate a validity mask; Guided bilateral filtering is performed on the effective depth data for edge conformal filtering, where the spatial kernel and range kernel are adaptively determined based on the point distance and depth noise model; Each depth of view data is mapped to a unified global coordinate system through intrinsic inverse projection and extrinsic matrix to form a primary point set; Perform reprojection consistency detection on points in multi-view overlapping regions to reduce the weight of points with residuals exceeding the threshold; Point-level multi-factor confidence scores are constructed based on the number of observations, incident angle, measurement variance, plane fitting residual, and reprojection residual, and point weights are calculated using a monotonic compression function. The voxel distance field is updated using weighted voxel fusion based on a truncated signed distance function, where point weights are used for weighted updates. A multi-stage adaptive denoising strategy, employing plane segmentation, soft-label noise scoring, density clustering, statistical filtering, and radius filtering, preserves structural details and suppresses outliers.

[0008] Optionally, the process of extracting structured geometric features and sparse matching features includes: Principal plane features are extracted from the pure point cloud, plane parameters are refined by normal clustering and least squares, and the variance of the fitting residuals is estimated. Candidate arc segments are identified from the edge points of the plane based on the normal gradient and curvature, and the center, radius and normal of the circle are obtained by circle fitting. In polar coordinates, the feature boundary of the hole is extracted by the grid holes in the plane parameterized coordinates, and the elliptical hole center position and hole diameter are fitted. ISS key point detection is performed on the downsampled point cloud, and key points are selected based on the ratio of neighborhood covariance eigenvalues. Calculate FPFH descriptors for keypoints, construct histograms, and perform matching to obtain initial transformations.

[0009] Optionally, the process of virtual completion using a CAD model includes: When the number of supporting point clouds of real features is lower than the threshold, the corresponding features in the CAD model are projected as virtual features according to the current coarse pose. The virtual features are given amplified covariance, with the amplification factor ranging from 5 to 10; During the iterative optimization process, the amplification factor is gradually reduced based on the number of real points falling within the support radius of the virtual feature, so that the virtual feature converges to the level of the real constraint.

[0010] Optionally, the process of calculating the initial pose of the tube segment through multi-factor robust optimization includes: A unified multi-factor robust optimization problem is constructed, incorporating planar features, arc features, hole features, matching point pairs, and virtual features into the loss function, where a robust kernel is used to suppress outliers in the loss function; The pose is parameterized using Lie algebras and solved iteratively using the Levenberg-Marquardt algorithm. Constraints are enabled in stages: first, plane and hole features are used, then arc and high-confidence matching are introduced, and finally all matching is enabled and refined. Factors with abnormal standardized residuals and low weights are removed after each round of optimization.

[0011] Optionally, the process of triggering a restricted six-DOF global search to compensate for pose includes: Candidate rotation sets are generated by discretizing in the rotation space with coarse-grained step size, and the consistency cost with the design prior direction is calculated to filter candidates. For each candidate rotation, the candidate translation is quickly estimated by minimizing the translation cost function using the distance field; The candidate rotation and translation combinations are sorted according to their comprehensive scores, and a complete multi-factor nonlinear optimization is performed on the top K groups. If the pose improvement exceeds the threshold, the update is accepted; otherwise, the angle step size is reduced and the search is restarted until the timeout occurs.

[0012] Optionally, the process of outputting the final pose matrix and quality metrics includes: Output a 4x4 homogeneous pose matrix; Output attitude quality metrics, including residual statistics, contribution rate of each feature category, and outlier rate; Output the attitude covariance matrix, which is calculated using an approximate Hessian method and includes the translation standard deviation and the attitude angle standard deviation. When the pose is invalid, a failure flag is returned, which can be used by the control layer to trigger manual review or re-sampling strategy.

[0013] Optionally, the process for determining if the pose quality is substandard includes: Calculate the condition number of the pose covariance matrix. If the condition number exceeds the threshold, it is determined that the standard has not been met. Check the number of measured support point clouds for key structural features; if it is less than the minimum cardinality, it is determined that the standard has not been met. Analyze the residual distribution; if the residual distribution has a long tail, it is determined that the standard has not been met. Check the correlation between pose drift direction and virtual feature weights; if the correlation is high, it is determined that the standard is not met.

[0014] Technical advantages of this invention: This invention discloses a high-precision shield tunnel segment splicing pose recognition method. In shield tunnel segment assembly pose recognition, through multi-source single-frame data access and preprocessing, adaptive point cloud cleaning, dual-domain feature extraction and CAD virtual completion, multi-factor robust optimization and restricted six-degree-of-freedom global search compensation, high-precision and robust pose estimation is achieved. It can stably control translation error and attitude angle error, and can still maintain usable initial values ​​and converge through compensation even under occlusion, dust or feature loss. Compared with the prior art, it significantly improves the pose reconstruction success rate and reduces the local extreme value failure rate. At the same time, by triggering global search on demand, it ensures the balance between real-time performance and computing resources, thereby reducing the misassembly rate and the frequency of manual intervention, and improving the system availability and consistency. Attached Figure Description

[0015] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart illustrating a high-precision shield tunnel segment splicing pose recognition method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the data preprocessing, fusion, and multi-stage adaptive denoising structure in an embodiment of the present invention; Figure 3 This is a schematic diagram of the dual-domain feature extraction and virtual completion structure according to an embodiment of the present invention; Figure 4 This is a schematic diagram of the pose robust optimization, quality assessment, and restricted six-degree-of-freedom global compensation structure in an embodiment of the present invention; Figure 5 This is a structural diagram of the segment pose stitching and pose recognition system according to an embodiment of the present invention. Detailed Implementation

[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0017] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0018] like Figure 1 As shown, this embodiment provides a high-precision shield tunnel segment splicing pose recognition method, including: acquiring single-frame depth data from multi-source sensors, and preprocessing the depth data to generate a clean point cloud; Structured geometric features and sparse matching features are extracted based on the pure point cloud, and virtual completion is performed using a CAD model when features are missing. Using the structured geometric features and sparse matching features, the initial pose of the tube segment is calculated through multi-factor robust optimization, and a restricted six-degree-of-freedom global search is triggered to compensate for the pose when the pose quality does not meet the standard. Output the final pose matrix and quality index.

[0019] Furthermore, the preprocessing of depth data includes: For each depth pixel, a validity weight factor is calculated based on distance constraints, local gradients, and incident angles to generate a validity mask; Guided bilateral filtering is performed on the effective depth data for edge conformal filtering, where the spatial kernel and range kernel are adaptively determined based on the point distance and depth noise model; Each depth of view data is mapped to a unified global coordinate system through intrinsic inverse projection and extrinsic matrix to form a primary point set; Perform reprojection consistency detection on points in multi-view overlapping regions to reduce the weight of points with residuals exceeding the threshold; Point-level multi-factor confidence scores are constructed based on the number of observations, incident angle, measurement variance, plane fitting residual, and reprojection residual, and point weights are calculated using a monotonic compression function. The voxel distance field is updated using weighted voxel fusion based on a truncated signed distance function, where point weights are used for weighted updates. A multi-stage adaptive denoising strategy, employing plane segmentation, soft-label noise scoring, density clustering, statistical filtering, and radius filtering, preserves structural details and suppresses outliers.

[0020] Furthermore, the process of extracting structured geometric features and sparse matching features includes: Principal plane features are extracted from the pure point cloud, plane parameters are refined by normal clustering and least squares, and the variance of the fitting residuals is estimated. Candidate arc segments are identified from the edge points of the plane based on the normal gradient and curvature, and the center, radius and normal of the circle are obtained by circle fitting. In polar coordinates, the feature boundary of the hole is extracted by the grid holes in the plane parameterized coordinates, and the elliptical hole center position and hole diameter are fitted. ISS (Intrinsic Shape Signatures) key point detection is performed on the downsampled point cloud, and key points are selected based on the ratio of neighborhood covariance eigenvalues. Calculate FPFH (Fast Point Feature Histograms) descriptors for keypoints, construct histograms, and perform matching to obtain the initial transformation.

[0021] Furthermore, the process of virtual completion using CAD models includes: When the number of supporting point clouds of real features is lower than the threshold, the corresponding features in the CAD model are projected as virtual features according to the current coarse pose. The virtual features are given amplified covariance, with the amplification factor ranging from 5 to 10; During the iterative optimization process, the amplification factor is gradually reduced based on the number of real points falling within the support radius of the virtual feature, so that the virtual feature converges to the level of the real constraint.

[0022] Furthermore, the process of calculating the initial pose of the tube segment through multi-factor robust optimization includes: A unified multi-factor robust optimization problem is constructed, incorporating planar features, arc features, hole features, matching point pairs, and virtual features into the loss function, where a robust kernel is used to suppress outliers in the loss function; The pose is parameterized using Lie algebras and solved iteratively using the Levenberg-Marquardt algorithm. Constraints are enabled in stages: first, plane and hole features are used, then arc and high-confidence matching are introduced, and finally all matching is enabled and refined. Factors with abnormal standardized residuals and low weights are removed after each round of optimization.

[0023] Furthermore, the process of triggering a restricted six-DOF global search to compensate for pose includes: Candidate rotation sets are generated by discretizing in the rotation space with coarse-grained step size, and the consistency cost with the design prior direction is calculated to filter candidates. For each candidate rotation, the candidate translation is quickly estimated by minimizing the translation cost function using the distance field; The candidate rotation and translation combinations are sorted according to their comprehensive scores, and a complete multi-factor nonlinear optimization is performed on the top K groups. If the pose improvement exceeds the threshold, the update is accepted; otherwise, the angle step size is reduced and the search is restarted until the timeout occurs.

[0024] Furthermore, the process of outputting the final pose matrix and quality indicators includes: Output a 4x4 homogeneous pose matrix; Output attitude quality metrics, including residual statistics, contribution rate of each feature category, and outlier rate; Output the attitude covariance matrix, which is calculated using the Hessian (information matrix) obtained by the Gauss-Newton approximation, and give the standard deviations of the translation component and the attitude angle component.

[0025] When the pose is invalid, a failure flag is returned, which can be used by the control layer to trigger manual review or re-sampling strategy.

[0026] Furthermore, the process for determining if the pose quality fails to meet the standards includes: Calculate the condition number of the pose covariance matrix. If the condition number exceeds the threshold, it is determined that the standard has not been met. Check the number of measured support point clouds for key structural features; if it is less than the minimum cardinality, it is determined that the standard has not been met. Analyze the residual distribution; if the residual distribution has a long tail, it is determined that the standard has not been met. Check the correlation between pose drift direction and virtual feature weights; if the correlation is high, it is determined that the standard is not met.

[0027] Specifically, the implementation process of this embodiment includes: like Figure 1 As shown, this embodiment provides a high-precision shield tunnel segment splicing pose recognition method, which mainly includes, from top to bottom: a multi-source single-frame data access and calibration association layer, a data preprocessing and fusion adaptive denoising layer, a dual-domain feature construction and virtual completion layer, a pose robust solution and constrained global compensation layer, and an output and quality feedback closed-loop layer. Addressing common problems in actual construction such as incomplete overlap of multi-source equipment fields of view, increased depth noise and outlier ratio due to dust and humidity, missing structures like holes and arcs, and susceptibility to local extrema in single-step ICP (Iterative Closest Point) algorithms, this implementation method achieves stable millimeter-level translation and sub-angle-level pose solutions through point-level confidence-driven multi-stage cleaning, complementary modeling of structured geometry and sparse matching, multi-factor optimization with switchable constraints, and on-demand triggered hierarchical global search collaboration. During system operation, each stage achieves a low-coupling interface through standardized data objects (such as weighted point clouds, feature factor sets, and covariance information), supporting on-demand pruning in computationally limited scenarios.

[0028] During the device access phase, multiple LiDAR, structured light, or depth cameras register through a unified device access service and are associated with their intrinsic and extrinsic parameters, time synchronization, and health status. To mitigate micro-pose drift caused by mechanical vibration, this implementation employs a soft synchronization strategy: recording the timestamp of each device within the trigger time window Δt. Linear extrinsic parameter interpolation correction is performed for time offsets exceeding a threshold. All raw data is stored as point clouds and their attributes, and each device has a separate thread controlling data acquisition. If a device is missing data for a given frame but its health indicators are in a degraded state (e.g., window contamination), the system will return an acquisition failure and wait for new instructions.

[0029] like Figure 2 and Figure 3 As shown, the data preprocessing and fusion stage first performs a depth map validity mask. For each depth pixel... According to distance constraints Local gradient threshold Angle of incidence Together with the original confidence level, they construct the validity weight factor. The local gradient can be estimated using Sobel, and the incident angle is calculated from the normal n of the nearest 3×3 depth-fitting plane and the line-of-sight direction l. Subsequently, edge conformal filtering is performed to suppress high-frequency depth noise while preserving the geometric discontinuities between the surface and the hole edge. This implementation uses guided bilateral filtering: ; in This is the original depth value. The center pixel to be updated. The set of pixels within the neighborhood of the center pixel. Weights. Spatial kernel. The adaptive value is taken as a multiple of the average point distance, and the value range kernel is... According to the depth noise model ( The depth difference between a pixel and its neighboring pixels. This is used for range and kernel variance estimation. After filtering, each depth of view is inversely projected using intrinsic parameters and then analyzed using the extrinsic parameter matrix. Mapping to a unified global coordinate system forms a primary set of points. Fast reprojection consistency checks are performed on point pairs in multi-view overlapping regions. Points with residuals greater than a threshold have their weights reduced rather than being immediately removed, preventing early weakening of areas such as hole edges and thin walls.

[0030] Preferably, this invention constructs a multi-factor confidence score for each sampling point. The multi-factors include at least one or more of the following: the number of observations, the incident angle factor, the inverse of the measurement variance, the plane fitting residual penalty, and the reprojection residual penalty. The multi-factors are then linearly or nonlinearly combined and input into a monotonic compression function (e.g., logistic) to obtain the point weights. This point weight is used for the weighted update of voxel distance field fusion and the initial weight allocation for subsequent multi-factor residual optimization. The weighted voxel fusion stage uses the truncated signed distance function (TSDF) for updating: ; in The truncated normalized distance of a point along the line of sight to the center of the voxel This fusion buffers random noise and provides a fast range field scoring structure for subsequent global compensation searches.

[0031] In the multi-stage adaptive denoising pipeline, the principal plane set is first extracted using RANSAC through plane segmentation. Record the residual from each point to the nearest plane. and its structural labels; then, using multiple features to construct soft-label noise probabilities: ; in These are the bias coefficient and the influence factor. This is an estimate of the reciprocal of the density; The variance of the normal vector is represented. The density clustering stage incorporates adaptive parameters. Using DBSCAN (Density-Based Spatial Clustering of Applications with Noise) clustering, isolated low-density and high-density clusters are identified. The region is labeled as a set of candidate outliers; when using a combined statistical and radius-based dual-filter decision-making process, the statistical filter is based on... Deemed abnormal. According to Scaling is applied; the radius filtering criterion is neighborhood counting. To avoid accidentally deleting actual structures, CAD residual retracement was used to re-scan the deleted point set. Calculate the residual of the design model ,like If the number of true point clouds in a given area is insufficient, the points are restored by reducing their weights and labeled as "re-scanned points". The output is a clean, uniform point cloud and its corresponding confidence level.

[0032] like Figure 3 and Figure 4As shown, the feature extraction stage constructs dual-domain features of structured geometry and sparse matching on the clean point cloud. Specifically, the structured geometry branch first performs normal clustering and parameter refinement on the main planes, updates the plane normal n and distance d using least squares, and estimates the variance of the fitting residuals. As a constraint on covariance, for plane edge points, the normal gradient threshold and curvature are used. Identify the set of potential edges, where The eigenvalues ​​are the local neighborhood covariance matrix (PCA) of the point. Candidate arc segments are segmented using topology tracking and orientation continuity detection, and the center c, radius R, and plane normal are obtained using Taubin circle fitting. The effective angular domain of the arc segment is constrained by the tangential variation at the endpoints. Hole feature detection is performed in polar coordinates using plane parameterized coordinates. The process involves extracting the boundaries of missing internal regions through grid holes, and then fitting ellipses to the boundary points. The covariance between the hole center position and the hole diameter is derived from the fitting residuals and the point weight distribution. Since the arc and hole of the shield tunnel segment have prior design dimensions, a penalty term can be introduced during the fitting process. To avoid radius shift caused by partial occlusion.

[0033] The sparse matching branch performs ISS keypoint detection after downsampling. For each candidate point, the neighborhood covariance eigenvalue is calculated. ,like and Furthermore, points with responses greater than the largest neighboring point are retained. Feature description uses the FPFH descriptor: each point and its kNN neighborhood are calculated using Darboux boxes. (Rotation angle, tilt angle, fold angle), construct a 33-dimensional histogram and then perform weighted aggregation; in the matching stage, perform nearest and second nearest distance ratio tests, cross-checks, and geometric consistency tests on the descriptors of the two datasets (point cloud and CAD template) to obtain the initial transformation using RANSAC. To improve the robustness of pose estimation, this embodiment dynamically estimates the uncertainty of points within the matching context based on the similarity of the matching descriptors and the local density of the point cloud. Specifically, let the descriptor distance between the matching points be... The local density is Then its variance is defined as: ; in The standard variance is... This is an adjustment coefficient. This form can adaptively reduce the impact of similarity and sparse region matching on the final pose optimization, thereby improving overall robustness and accuracy.

[0034] When structural features are missing (e.g., handhole occlusion > 50% or edge curvature is damaged), the system will attempt to use a CAD template for virtual completion. This will be based on the current coarse pose. The design model will include elements not supported by real observations (number of point clouds). The center of the hole, arc segment, or auxiliary positioning plane is projected onto the point cloud coordinate system to generate virtual features. For each virtual feature, its covariance is assigned as follows: ,in The baseline covariance of this feature under the design model. The amplification factor is preferably set between 5 and 10. During subsequent iterative optimization, the system periodically adjusts the value based on the support radius of the real point falling within this virtual feature. The number within gradually converges iteratively, eventually making This gradually transforms the virtual feature into a level that approximates the real constraint.

[0035] like Figure 4 As shown, the pose solving and compensation stage first constructs a unified multi-factor robust optimization problem. The pose to be estimated is parameterized as follows: Using Lie algebra minimum perturbation Iterative updates (rotation and translation). Example of the overall objective function form: ; in The Huber robust kernel is used to suppress residual outliers. Derived from the aforementioned combination of point-level confidence and feature fitting weights, This represents the normalized representation of the aforementioned features. For a set of planar constraints, For the set of arc constraints, A set of constraints for holes or grooves. For the corresponding set of key points, For the remaining possible set of constraints, Represents the range field consistency term, within the summation term For each point, the distance field query function R is the rotation matrix, and t is the translation vector. The distance scale can be represented as the variance of the distance field. This refers to the weight hyperparameter. The Jacobian matrix is ​​analytically derived from the derivative of the error from a point to a plane or a point to a circle with respect to ξ. Specifically, this example uses the Levenberg-Marquardt solution. To avoid being drawn in by erroneous key points early on, the constraints are divided into three stages: the first stage uses only the plane and holes; the second stage introduces arcs and high-confidence matching; and the third stage enables full matching and refinement. Standardized residuals are removed after each stage. And factors with low weights.

[0036] After completing the above solution process, fine-tuning of the ICP is performed. A voxel pyramid (step size (l, 2l, 4l)) is constructed for the target point cloud, and the point-to-plane and point-to-plane residual models are run sequentially from coarse to fine: ; If the mean of the finest scale residuals and 95th percentile satisfy and , and The threshold parameter is used, and the pose covariance is calculated on the main diagonal (translation sub-block). Rotating sub-blocks If all values ​​are less than their respective thresholds, then the solution is considered valid.

[0037] The results were then evaluated for quality using an approximation of Hessian. Calculate attitude covariance Output translation standard deviation with attitude angle standard deviation Simultaneously, the contribution rate of each feature category was calculated. This is used to record constraint coverage completeness in the log. The trigger logic will detect if: 1) condition number ;2) The measured support for key structures (holes / arcs) is less than the minimum cardinality;3) The residual distribution exhibits a long tail ( ); 4) If the pose drift direction height is related to the virtual feature weight, then the restricted six degrees of freedom hierarchical search logic is invoked.

[0038] On the rotation space SO(3) with the current coarse-grained step size Construct a set of candidate rotations (For example, uniform discretization by axis angle or uniform projection of a quaternion hypercube). For each candidate, compare the consistency between the observed structural feature directions (plane normal, hole axis, etc.) and the design prior directions, and calculate the directional cost: ; in It is the normal vector of the observation. This is the model normal vector, used to remove rotations with orientation deviations exceeding a threshold. For the retained... We use the constructed TSDF to perform a fast coarse search on the translation, minimizing: ; in It is the truncated square penalty function. The cutoff function threshold is obtained by gradient descent or multi-start cube scan to obtain several candidate thresholds. .

[0039] All Sort by comprehensive score, select the top K groups to perform full multi-factor nonlinear optimization; if a group is better than the current solution and exceeds the improvement threshold in both residual statistics and covariance (main diagonal standard deviation), then it is accepted for update.

[0040] If no improvement is made in this round, reduce the angle step size (e.g.) Enter the iteration, and set the maximum processing time during the iteration process; if a better solution is not obtained after the time is exceeded, return the original solution and output an error for the upper layer to return according to the preset strategy.

[0041] This implementation also introduces a parameter adaptive feedback mechanism: during operation, the true residual variance will be... Structural support rate Out-of-point rate Write to the statistics buffer and periodically adjust the plane threshold. Soft label weight coefficient Updates are made based on criteria such as matching and rejection thresholds. For example, if N consecutive frames... If the result is significantly lower than the prediction of the design noise model, the slope of the soft-label noise probability mapping can be appropriately tightened to obtain a more "decisive" filtering. Conversely, if feature loss frequently triggers virtual completion and the global search trigger rate increases, the initial filtering threshold and the plane segmentation distance limit can be relaxed to improve the structural candidate retention rate.

[0042] The output stage system provides the final 4x4 homogeneous pose matrix. Or failure markers, attitude quality metric set Q (including residual statistics and category contribution) Out-of-point rate Covariance matrix The control layer can determine the tunnel boring machine's travel based on the returned results. For abnormal frames or scenarios, manual review or a strategy of adding a new frame can be implemented. All intermediate data (filtering parameters, feature fitting results, reasons for elimination, and candidate search scores) are structured and recorded, forming a traceable chain that supports subsequent quality analysis and model parameter retraining.

[0043] like Figure 5An embodiment of a tunnel segment pose stitching and recognition system, as shown, consists of four parts: a front-end display layer, an application development layer, an application foundation layer, and a system perception layer. The graphical user interface (GUI) comprises two parts for operators to interact with the system for actual control. The GUI visualizes the stitching results; the tunnel boring machine (TBM) interface is used for sending commands and receiving results. The application development layer consists of two parts: data persistence and data processing, used for actual business logic processing. Data persistence saves the collected sensor data and the results returned from the process; data processing performs algorithmic calculations on the tunnel segment data and returns the final results. The application foundation layer provides the underlying implementation for sensor hardware control, request communication, and the system algorithm's basic runtime library to ensure system operation. Specifically, in this embodiment, the point cloud runtime library is a data processing runtime library based primarily on PCL. The persistent database uses MySQL as the persistent storage method and contains two databases: `dataSets` for storing the original collected data storage location and `computeResult` for storing the final calculation results. Machine learning runtime library: Anaconda is used for building and maintaining the virtual environment, and PyTorch is used as the algorithm development framework. A network communication library is used to implement request interactions; specifically, in this embodiment, the system as a whole uses WebSocket as the communication protocol and JSON as the data transmission format for information exchange. The system's perception layer directly interacts with the sensing device for actual data acquisition functions.

[0044] In a specific engineering example, the typical single-frame input points number approximately 2–4 million. After validity masking and edge conformal filtering, 70–85% are retained. After multi-stage denoising, the number of points is reduced to 40–55%, and structural features are extracted to include 2–4 principal planes, 6–12 holes, 1–3 arc segments, and 300–600 high-response keypoints. Initial multi-factor optimization iterations of 5–15 times can converge to a translation root mean square error of <3mm. After coordinating multi-resolution ICP and triggering a small amount of global compensation, the convergence is reduced to ≤2mm. In the optimal working condition, when the holes and arcs are intact, the accuracy reaches within 1.5mm and 0.3°, respectively. Compared to single ICP, with an external point rate of 30% and hole occlusion of 50%, the success rate (meeting the error threshold) of this invention is increased from approximately 62% to over 90%, and 90% of the failed cases can be downgraded to usable poses through hierarchical search.

[0045] It should be noted that the weighting function, threshold adaptive strategy, cost function form, sampling density, kernel function type, etc., mentioned above can all be adjusted according to different tunnel segment specifications, sensor arrangement, and construction environment pollution characteristics; any sub-module (e.g., TSDF fusion or arc fitting) can be replaced with an equivalent function without affecting the overall idea of ​​this invention. For scenarios without configured CAD templates, the virtual completion module can degenerate into statistical morphological template generation (based on historical frame aggregation), and the weight covariance amplification strategy is still applicable. The method described in this invention can also be extended to other assembly and inspection tasks that require high-precision attitude estimation under single-frame conditions, such as prefabricated component assembly, circumferential lining inspection, and cylindrical structure calibration, which is a reasonable extension within the scope of protection of the claims of this invention. In summary, through multi-source single-frame adaptive cleaning, dual-domain feature redundancy modeling, robust multi-factor optimization, and constrained global compensation collaboration, this implementation maintains high accuracy and high stability in shield tunnel segment pose recognition under high noise and feature loss conditions.

[0046] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A high-precision method for recognizing the pose of tunnel segment splicing, characterized in that, include: Acquire single-frame depth data from multiple source sensors and preprocess the depth data to generate a clean point cloud; Structured geometric features and sparse matching features are extracted based on the pure point cloud, and virtual completion is performed using a CAD model when features are missing. Using the structured geometric features and sparse matching features, the initial pose of the tube segment is calculated through multi-factor robust optimization, and a restricted six-degree-of-freedom global search is triggered to compensate for the pose when the pose quality does not meet the standard. Output the final pose matrix and quality index.

2. The high-precision shield tunnel segment splicing pose recognition method as described in claim 1, characterized in that, The preprocessing of depth data includes: For each depth pixel, a validity weight factor is calculated based on distance constraints, local gradients, and incident angles to generate a validity mask; Guided bilateral filtering is performed on the effective depth data for edge conformal filtering, where the spatial kernel and range kernel are adaptively determined based on the point distance and depth noise model; Each depth of view data is mapped to a unified global coordinate system through intrinsic inverse projection and extrinsic matrix to form a primary point set; Perform reprojection consistency detection on points in multi-view overlapping regions to reduce the weight of points with residuals exceeding the threshold; Point-level multi-factor confidence scores are constructed based on the number of observations, incident angle, measurement variance, plane fitting residual, and reprojection residual, and point weights are calculated using a monotonic compression function. The voxel distance field is updated using weighted voxel fusion based on a truncated signed distance function, where point weights are used for weighted updates. A multi-stage adaptive denoising strategy, employing plane segmentation, soft-label noise scoring, density clustering, statistical filtering, and radius filtering, preserves structural details and suppresses outliers.

3. The high-precision shield tunnel segment splicing pose recognition method as described in claim 1, characterized in that, The process of extracting structured geometric features and sparse matching features includes: Principal plane features are extracted from the pure point cloud, plane parameters are refined by normal clustering and least squares, and the variance of the fitting residuals is estimated. Candidate arc segments are identified from the edge points of the plane based on the normal gradient and curvature, and the center, radius and normal of the circle are obtained by circle fitting. In polar coordinates, the feature boundary of the hole is extracted by the grid holes in the plane parameterized coordinates, and the elliptical hole center position and hole diameter are fitted. ISS key point detection is performed on the downsampled point cloud, and key points are selected based on the ratio of neighborhood covariance eigenvalues. Calculate FPFH descriptors for keypoints, construct histograms, and perform matching to obtain initial transformations.

4. The high-precision shield tunnel segment splicing pose recognition method as described in claim 1, characterized in that, The process of virtual completion using a CAD model includes: When the number of supporting point clouds of real features is lower than the threshold, the corresponding features in the CAD model are projected as virtual features according to the current coarse pose. The virtual features are given amplified covariance, with the amplification factor ranging from 5 to 10; During the iterative optimization process, the amplification factor is gradually reduced based on the number of real points falling within the support radius of the virtual feature, so that the virtual feature converges to the level of the real constraint.

5. The high-precision shield tunnel segment splicing pose recognition method as described in claim 1, characterized in that, The process of calculating the initial pose of the tube segment through multi-factor robust optimization includes: A unified multi-factor robust optimization problem is constructed, incorporating planar features, arc features, hole features, matching point pairs, and virtual features into the loss function, where a robust kernel is used to suppress outliers in the loss function; The pose is parameterized using Lie algebras and solved iteratively using the Levenberg-Marquardt algorithm. Constraints are enabled in stages: first, plane and hole features are used, then arc and high-confidence matching are introduced, and finally all matching is enabled and refined. Factors with abnormal standardized residuals and low weights are removed after each round of optimization.

6. The high-precision shield tunnel segment splicing pose recognition method as described in claim 1, characterized in that, The process of triggering a restricted six-degrees-of-freedom global search to compensate for pose includes: Candidate rotation sets are generated by discretizing in the rotation space with coarse-grained step size, and the consistency cost with the design prior direction is calculated to filter candidates. For each candidate rotation, the candidate translation is quickly estimated by minimizing the translation cost function using the distance field; The candidate rotation and translation combinations are sorted according to their comprehensive scores, and a complete multi-factor nonlinear optimization is performed on the top K groups. If the pose improvement exceeds the threshold, the update is accepted; otherwise, the angle step size is reduced and the search is restarted until the timeout occurs.

7. The high-precision shield tunnel segment splicing pose recognition method as described in claim 1, characterized in that, The process of outputting the final pose matrix and quality indicators includes: Output a 4x4 homogeneous pose matrix; Output attitude quality metrics, including residual statistics, contribution rate of each feature category, and outlier rate; Output the attitude covariance matrix, which is calculated using an approximate Hessian method and includes the translation standard deviation and the attitude angle standard deviation. When the pose is invalid, a failure flag is returned, which can be used by the control layer to trigger manual review or re-sampling strategy.

8. The high-precision shield tunnel segment splicing pose recognition method as described in claim 1, characterized in that, The process for determining whether a person's pose quality is substandard includes: Calculate the condition number of the pose covariance matrix. If the condition number exceeds the threshold, it is determined that the standard has not been met. Check the number of measured support point clouds for key structural features; if it is less than the minimum cardinality, it is determined that the standard has not been met. Analyze the residual distribution; if the residual distribution has a long tail, it is determined that the standard has not been met. Check the correlation between pose drift direction and virtual feature weights; if the correlation is high, it is determined that the standard is not met.