Multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP (Inductively Coupled Plasma)
The adaptive layer-wise attention and efficient ICP method for multi-modal point cloud registration addresses dynamic environment challenges by improving feature fusion and alignment precision, offering robustness and efficiency in point cloud alignment.
Patent Information
- Application Number
- CN202510407717.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-15
AI Technical Summary
The existing point cloud registration methods have insufficient robustness in dynamic environment and cross-modal feature alignment, limited generalization capabilities, and multimodal point cloud registration fails to fully explore the spatial structure consistency between multimodal features, resulting in limited registration accuracy.
The multimodal point cloud dynamic registration method is adopted with adaptive hierarchical attention and efficient ICP. Through multimodal feature fusion and spatial hierarchical adaptive attention mechanism, including coarse registration and precise registration stages, combined with point cloud geometric coding and BEV semantic coding, the attention range is dynamically adjusted, and point cloud registration is optimized using an improved ICP method.
It significantly improves the accuracy and efficiency of point cloud registration, and can perform point cloud registration more accurately in complex dynamic environments, reducing errors and instability, and enhancing robustness in dynamic objects and scene changes.
Smart Images

Figure CN120318281A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of xx, specifically a multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP. Background Art
[0002] Existing point cloud registration methods mainly include learning-based point cloud registration, Transformer-based registration methods, and multi-modal point cloud registration methods.
[0003] Learning-based point cloud registration methods have become the main research direction for point cloud alignment. Such methods usually adopt a coarse-to-fine strategy to improve the robustness and accuracy of registration by relaxing the strict point-to-point matching constraints. Point-level methods (such as Minkowski and KPConv) extract dense point features through deep learning, while Patch-based methods (such as PPFNet, D3Feat) use local region features to enhance matching stability.
[0004] Transformer-based point cloud registration. The global attention mechanism introduced by the Transformer architecture has been widely applied in the point cloud registration task. CoFiNet uses a standard self-attention mechanism to construct global geometric consistency, while GeoTrans combines geometric structure information to enhance feature robustness. PEAL reduces the interference of non-overlapping regions by combining a unidirectional attention mechanism with an overlap prior. In addition, DCATr improves the matching accuracy by limiting the attention range to focus on specific geometric features.
[0005] Multi-modal point cloud registration. The application of multi-modal information fusion in point cloud registration has gradually increased. RPVNet uses range images to assist point cloud segmentation to improve registration accuracy. Methods such as BYOC, PCRC-G, and IGReg fuse image features and point cloud information to improve matching accuracy. In addition, PEAL2D infers the overlapping regions of point clouds using 2D images, and BEV features, due to their global edge information advantages, have been widely used in 3D object detection and scene recognition and are gradually applied to the point cloud registration task to improve the matching accuracy in the coarse registration stage.
[0006] However, existing point cloud registration methods still have many deficiencies in dynamic environments and cross-modal feature alignment. Learning-based point cloud registration methods highly depend on the distribution of training data, resulting in limited generalization ability. The application of Transformer in point cloud registration still has limitations in large-scale or dynamic environments, which may introduce a large amount of information about irrelevant points, causing matching ambiguity. In multi-modal point cloud registration, most methods use simple splicing or directly introduce additional supervision, without fully exploring the spatial structure consistency between multi-modal features, resulting in insufficient matching robustness and still limited registration accuracy.
[0007] Therefore, a new solution to the above problems is needed. Summary of the Invention
[0008] The purpose of the present invention is to provide a multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP, including two main stages: coarse registration and fine registration. High-precision and high-efficiency point cloud registration is achieved through multi-modal feature fusion and spatial hierarchical adaptive attention mechanism, aiming to solve the problem of lidar point cloud registration.
[0009] To achieve the above purpose, the present invention provides the following technical solutions: A multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP includes at least the following steps:
[0010] S1: Input, receive the source measurement point cloud P and the source BEV view P ~ , and the target measurement point cloud Q and the target BEV view Q ~ ;
[0011] S2: Feature extraction is performed through a feature extraction module. Point-level feature encoding is respectively performed on the source point cloud P and the target point cloud Q, and then point-level features F p and F q are extracted; The source BEV view P ~ and the target BEV view Q~ are encoded to obtain BEV features and
[0012] S3: Build a feature fusion module, and respectively fuse the feature F p of the source point cloud P with the feature of the source BEV to obtain Fuse the feature F q of the target point cloud Q with the feature of the target BEV to obtain
[0013] S4: Build a spatial hierarchical adaptive attention module. The spatial hierarchical adaptive attention module dynamically adjusts the attention range according to the distance between each point and other points in the global context. The spatial hierarchical adaptive attention module includes three levels: local module attention, regional expansion attention, and global coupling attention;
[0014] S5: Build a superpoint matching module, and map the features to a shared feature space through the superpoint matching function and ;
[0015] S6: Build an Iterative Closest Point (ICP) optimization module and implement fine registration using an improved ICP method. The improved ICP method combines dynamic corresponding point adjustment, adaptive weight optimization, and a hierarchical convergence strategy to further optimize the rigid body transformation between point clouds, improve the convergence speed and the final registration accuracy, and then output the final rigid body transformation matrix {R, T}.
[0016] Furthermore, the feature extraction module is a dual-modal joint encoding architecture, consisting of a Point-Encoder and a Bird's Eye View Encoder (BEV-Encoder). The Point-Encoder performs point cloud geometric encoding, and the BEV-Encoder performs BEV semantic encoding;
[0017] The feature extraction module enhances the feature expression ability through the complementary fusion of point cloud geometric encoding and BEV semantic encoding;
[0018] The Point-Encoder adopts a four-level KPConv convolutional network to perform spatial downsampling and feature extraction layer by layer. After four levels of processing, it outputs a superpoint set with 1 / 32 downsampling. Each superpoint carries a 512-dimensional feature vector, significantly reducing the computational complexity while retaining key geometric structure information;
[0019] The BEV feature encoding constructs a four-stage feature pyramid based on the improved ResNet-18 and introduces a channel attention module in the fourth stage to adaptively enhance the semantic feature weights of key regions, which include but are not limited to road boundaries and building outlines.
[0020] Furthermore, the feature fusion derivation of the feature fusion module includes at least the following steps:
[0021] First, perform point cloud to BEV projection, that is, map the 3D point cloud to the 2D bird's eye view (BEV). The BEV projection provides a global top-down view of the scene, which helps to extract structured features. Given a point cloud P, where (x i , y i , z i ) represents the coordinates of each point, project it into the BEV image I ∈ RH×W through formula (1), and calculate the pixel position (ui, vi) of the point;
[0022]
[0023] where, x max , x min , y max and y min represent the boundary range of the point cloud in the XY plane, that is, the maximum and minimum boundaries of the point cloud in X and Y; W×H represents the resolution size of the BEV image; represents rounding down;
[0024] Then, the weight calculation of the feature weights is performed. To improve the quality of the features extracted from the BEV image, semantic weights need to be assigned to it. For each point p in the point cloud i , by calculating the features within the corresponding BEV image region, the fusion of the feature weights is obtained. Specifically, the BEV image feature F IP (u, v) and the interpolation kernel function G(p i , u, v) are combined and expressed by formula (2):
[0025]
[0026] Through this formula, the points are weighted and aggregated using the BEV image features. The interpolation kernel function G(p i , u, v) incorporates the neighborhood information by considering the contribution of each pixel in the BEV image to the point cloud features, thereby achieving feature fusion;
[0027] The bilinear interpolation algorithm is used to aggregate the BEV features within the neighborhood, and the BEV features are mapped to the point cloud feature space through a learnable weight matrix. The feature fusion strategy performs channel concatenation on the point cloud geometric features and the projected BEV features.
[0028] Furthermore, in the initial layer of the spatial hierarchical adaptive attention module, i.e., L = 1, local module attention is adopted. The spatial hierarchical adaptive attention module refines the local correlation by focusing on nearby superpoints, thereby effectively reducing the ambiguity that may be caused by the initial global attention;
[0029] In the subsequent layers, i.e., L > 1, regional expansion attention is adopted, and the attention range gradually expands, enabling the establishment of a multi-scale spatial representation of the superpoints. This step-by-step approach enhances the robustness of the superpoint features. A similar method also applies to Q ~ ;
[0030] In the calculation process of the at the L = k layer, global coupling attention is adopted. To achieve step-by-step self-attention, the masked attention score is expressed as:
[0031]
[0032] Among them, is a layer-specific mask used to filter the attention scores according to the Euclidean distance between points;
[0033] The mask is calculated by comparing the superpoint with all other superpoints The maximum distance between them is divided into S segments. For a given layer L = k, a distance threshold is defined as follows:
[0034]
[0035] Where: represents the maximum Euclidean distance between superpoints; is a factor used to adjust the attention range of each layer;
[0036] Boolean mask matrix:
[0037]
[0038] indicates that when the Euclidean distance between two superpoints is less than or equal to the threshold then the attention score between them will be retained, i.e., the mask value is 1; otherwise, the attention score is filtered out, i.e., the mask value is 0;
[0039] Masked attention calculation:
[0040] Z (k) = Softmax(M (k) ⊙ E) · XW V (6)
[0041] Where M (k) is the mask matrix of the L = k layer, E represents the attention score matrix, and XW V represents the value projection of the point cloud features; the role of the attention score matrix E here is to calculate an attention weight for each pair of points, and the mask matrix M (k) determines which attention scores will be retained. After Softmax normalization, the final output is calculated through weighted superpoint features.
[0042] Furthermore, in the feature normalization stage, the hyperpoint matching module maps the encoded and to the unit spherical space to obtain the normalized features and Subsequently, the Gaussian similarity matrix is expressed as follows:
[0043]
[0044] Where γ is a hyperparameter used to control the similarity decay rate; is the Euclidean distance between the normalized features and ;
[0045] By performing on Perform row and column double normalization to eliminate distribution bias:
[0046]
[0047] denotes the normalization factor of the i-th row, representing the sum of similarities between point i and all other points;
[0048] denotes the normalization factor of the j-th column, representing the sum of similarities between point j and all other points
[0049] Construct a matching set by selecting the superpoint pair with the highest correlation score As the initial correspondence for subsequent fine registration.
[0050] Furthermore, the improved ICP method at least includes the following steps:
[0051] In the superpoint matching stage, obtain a set of initial matching point pairs with high confidence As the input for ICP fine registration;
[0052] Since reliable matching point pairs have been obtained, directly calculate the rigid body transformation (R0, T0), avoiding the traditional ICP's dependence on nearest neighbor search for initialization; the transformation is solved by weighted least squares, using SVD to decompose the covariance matrix to obtain the rotation matrix R0, and calculate the translation vector T0; this initialization method can significantly accelerate the convergence speed and improve the accuracy;
[0053] In the subsequent ICP iteration process, dynamically optimize the matching point pairs, transform the source point cloud, and update the point cloud position;
[0054] Local nearest neighbor search, if the error of the current matching point exceeds the set threshold, update it to the local optimal matching, and perform matching screening in combination with feature similarity to avoid false matching interfering with ICP optimization;
[0055] Then perform adaptive weight optimization. In order to improve the reliability of the matching point pairs, the contribution of each point pair is controlled by the weight w i The calculation is based on feature similarity, making the weights of points with similar features higher; geometric consistency ensures the consistency of matching points in the local structure;
[0056] To optimize the calculation efficiency and registration accuracy, adopt hierarchical ICP optimization:
[0057] Superpoint-level ICP, only use the superpoint matching point pairs for rough registration to obtain the global transformation;
[0058] Local area ICP, perform ICP in the superpoint neighborhood to correct the local error;
[0059] Global fine ICP performs the final optimization based on the complete point cloud to ensure overall accuracy.
[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0061] 1. The present invention proposes a novel point cloud registration method, which combines an adaptive hierarchical attention mechanism with an efficient ICP optimization strategy. The adaptive hierarchical attention mechanism enhances the multi-scale spatial representation of superpoint features by gradually expanding the attention range, enabling the model to perform point cloud registration more accurately in complex dynamic environments. This method significantly improves the registration accuracy and reduces the errors and instabilities in traditional methods.
[0062] 2. The present invention proposes a multi-modal feature fusion strategy, which adopts the joint encoding of point cloud geometric features and BEV (Bird's Eye View) semantic features. By feature fusion, the representation ability of the point cloud is enhanced, and through an efficient superpoint matching module, the ability to map point cloud features to a shared feature space is further improved, thereby enhancing the accurate recognition and removal ability of dynamic objects. This method effectively reduces the registration errors caused by feature loss;
[0063] 3. The present invention proposes a spatial hierarchical adaptive attention module. Through the spatial hierarchical adaptive attention module, the attention range is dynamically adjusted according to the Euclidean distance between points, realizing multi-scale feature learning from local to global. This module can effectively reduce the ambiguity that may be brought by the initial global attention and enhance the robustness of superpoint features by gradually expanding the attention range layer by layer. This technology has significant advantages in dealing with dynamic objects and scene changes in complex dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for describing the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0065] Figure 1 It is a schematic diagram of the whole of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.
[0067] The present invention has strong generalization ability in the rough registration stage, can effectively match cross-modal point clouds in complex environments, and improve the robustness of registration. At the same time, in the fine registration stage, it combines an optimized ICP algorithm to reduce the computational cost while ensuring the accuracy, so as to balance the registration accuracy and computational efficiency, and is applicable to the efficient point cloud alignment task of lidar SLAM in dynamic environments.
[0068] Please refer to Figure 1 , a multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP, at least including the following steps:
[0069] S1: Input, receive the source measurement point cloud P and the source BEV view P ~ , and the target measurement point cloud Q and the target BEV view Q ~ ;
[0070] S2: Perform feature extraction through the feature extraction module, perform point-level feature encoding on the source point cloud P and the target point cloud Q respectively, and then extract the point-level features F p and F q ; Encode the source BEV view P ~ and the target BEV view Q ~ to obtain the BEV features and
[0071] S3: Build a feature fusion module, and fuse the feature F p of the source point cloud P with the feature of the source BEV to obtain Fuse the feature F q of the target point cloud Q with the feature of the target BEV to obtain
[0072] S4: Build a spatial hierarchical adaptive attention module. The spatial hierarchical adaptive attention module dynamically adjusts the attention range according to the distance between each point and other points in the global context. The spatial hierarchical adaptive attention module includes three levels: local module attention, regional expansion attention, and global coupling attention;
[0073] S5: Build a superpoint matching module, and map the features to the shared feature space through the superpoint matching functions and ;
[0074] S6: Build an Iterative Closest Point (ICP) optimization module and implement fine registration using an improved ICP method. The improved ICP method combines dynamic corresponding point adjustment, adaptive weight optimization, and a hierarchical convergence strategy to further optimize the rigid body transformation between point clouds, improve the convergence speed and the final registration accuracy, and then output the final rigid body transformation matrix {R, T}.
[0075] The feature extraction module has a dual-modal joint encoding architecture, consisting of a Point-Encoder and a Bird's Eye View (BEV)-Encoder. The Point-Encoder performs geometric encoding on the point cloud, and the BEV-Encoder performs semantic encoding on the BEV;
[0076] The feature extraction module enhances the feature expression ability through the complementary fusion of point cloud geometric encoding and BEV semantic encoding;
[0077] The Point-Encoder adopts a four-level KPConv convolutional network to perform spatial downsampling and feature extraction layer by layer. After four-level processing, it outputs a superpoint set with 1 / 32 downsampling. Each superpoint carries a 512-dimensional feature vector, significantly reducing the computational complexity while retaining key geometric structure information;
[0078] The BEV feature encoding constructs a four-stage feature pyramid based on the improved ResNet-18 and introduces a channel attention module in the fourth stage to adaptively enhance the semantic feature weights of key regions, including but not limited to road boundaries and building contours.
[0079] The feature fusion derivation of the feature fusion module includes at least the following steps:
[0080] First, perform point cloud to BEV projection, that is, map the 3D point cloud to the 2D bird's eye view (BEV). The BEV projection provides a global top-down view of the scene, which helps to extract structured features. Given a point cloud P, where (x i , y i , z i ) represents the coordinates of each point, project it into the BEV image I ∈ RH×W through formula (1) to calculate the pixel position (ui, vi) of the point;
[0081]
[0082] where, x max , x min , y max and y min represent the boundary range of the point cloud in the XY plane, that is, the maximum and minimum boundaries of the point cloud in X and Y; W×H represents the resolution size of the BEV image; represents rounding down;
[0083] Then, the weight calculation of feature weights is performed. To improve the quality of the features extracted from the BEV image, semantic weights need to be assigned to it. For each point p in the point cloud i , by calculating the features within the corresponding BEV image region, the fusion of feature weights is obtained. Specifically, the BEV image feature F IP (u, v) and the interpolation kernel function G(p i , u, v) are combined and expressed by formula (2):
[0084]
[0085] Through this formula, the points are weighted and aggregated using the BEV image features. The interpolation kernel function G(p i , u, v) incorporates neighborhood information by considering the contribution of each pixel in the BEV image to the point cloud features, thus achieving feature fusion;
[0086] The bilinear interpolation algorithm is used to aggregate the BEV features within the neighborhood. The BEV features are mapped to the point cloud feature space through a learnable weight matrix, and the feature fusion strategy concatenates the point cloud geometric features and the projected BEV features in the channel dimension.
[0087] In the initial layer of the spatial hierarchical adaptive attention module, i.e., when L = 1, local module attention is adopted. The spatial hierarchical adaptive attention module refines local correlations by focusing on nearby superpoints, thus effectively reducing the ambiguity that may be caused by the initial global attention;
[0088] In the subsequent layers, i.e., when L > 1, regional expansion attention is adopted, and the attention range gradually expands, enabling the establishment of multi-scale spatial representations of superpoints. This step-by-step approach enhances the robustness of the superpoint features. A similar method also applies to Q ~ ;
[0089] During the calculation at the L = k layer , global coupling attention is adopted. To achieve step-by-step self-attention, the masked attention score is expressed as:
[0090]
[0091] where is a layer-specific mask used to filter the attention scores based on the Euclidean distance between points;
[0092] The mask is calculated by dividing the maximum distance between the superpoint and all other superpoints into S segments. For a given layer L = k, the distance threshold is defined as:
[0093]
[0094] Wherein: represents the maximum Euclidean distance between superpoints, is a factor used to adjust the attention range of each layer;
[0095] Boolean mask matrix:
[0096]
[0097] indicates that when the Euclidean distance between two superpoints is less than or equal to the threshold then the attention score between them will be retained, i.e., the mask value is 1; otherwise, the attention score is filtered out, i.e., the mask value is 0;
[0098] Masked attention calculation:
[0099] Z (k) = Softmax(M (k) ⊙ E)·XW V (6)
[0100] Wherein, M (k) is the mask matrix of the L = k-th layer, E represents the attention score matrix, and XW V represents the value projection of the point cloud feature; the role of the attention score matrix E here is to calculate an attention weight for each pair of points, and the mask matrix M (k) determines which attention scores will be retained. After Softmax normalization, the final output is calculated through weighted superpoint features.
[0101] In the feature normalization stage, the superpoint matching module maps the encoded and to the unit spherical space to obtain the normalized features and Subsequently, the Gaussian similarity matrix is represented as follows:
[0102]
[0103] Wherein, γ is a hyperparameter used to control the similarity decay rate; is the Euclidean distance between the normalized features and ;
[0104] By performing row and column double normalization on the distribution bias is eliminated:
[0105]
[0106] It represents the normalization factor of the i-th row and the sum of the similarities between point i and all other points;
[0107] It represents the normalization factor of the j-th column and the sum of the similarities between point j and all other points
[0108] The matching set is formed by selecting the superpoint pair with the highest correlation score As the initial correspondence for subsequent fine registration.
[0109] The improved ICP method at least includes the following steps:
[0110] In the superpoint matching stage, a set of initial matching point pairs with high confidence is obtained As the input for ICP fine registration;
[0111] Since reliable matching point pairs have been obtained, the rigid body transformation (R0, T0) is directly calculated, avoiding the traditional ICP's dependence on nearest neighbor search for initialization; the transformation is solved by weighted least squares, the rotation matrix R0 is obtained by SVD decomposing the covariance matrix, and the translation vector T0 is calculated; this initialization method can significantly accelerate the convergence speed and improve the accuracy;
[0112] In the subsequent ICP iteration process, the matching point pairs are dynamically optimized, the source point cloud is transformed, and the point cloud position is updated;
[0113] Local nearest neighbor search, if the error of the current matching point exceeds the set threshold, it is updated to the local optimal matching, and the matching is screened by combining feature similarity to avoid the interference of false matching on ICP optimization;
[0114] Then, adaptive weight optimization is performed. In order to improve the reliability of the matching point pairs, the contribution of each point pair is controlled by the weight w i Its calculation is based on feature similarity, making the points with higher feature similarity have higher weights; geometric consistency (such as normal vector similarity) to ensure the consistency of the matching points in the local structure;
[0115] To optimize the calculation efficiency and registration accuracy, hierarchical ICP optimization is adopted:
[0116] Superpoint-level ICP, only using the superpoint matching point pairs for rough registration to obtain the global transformation;
[0117] Local region ICP, performing ICP in the superpoint neighborhood to correct local errors;
[0118] Global fine ICP, performing the final optimization based on the complete point cloud to ensure the overall accuracy.
[0119] In summary, by combining the adaptive hierarchical attention mechanism with the efficient ICP optimization strategy, the present invention exhibits the following remarkable advantages in the multi-modal point cloud registration task:
[0120] 1. Improve registration accuracy: Through the dual-modal joint encoding architecture, the geometric features of the point cloud and the BEV semantic features are fused to enhance the feature expression ability and reduce the loss of feature information during the matching process. At the same time, the spatial hierarchical adaptive attention mechanism is adopted to gradually expand the attention calculation range from local association to global coupling, enhancing the robustness of feature matching. Compared with traditional ICP or registration methods based on single-modal features, the proposed solution of the present invention has higher matching accuracy in complex environments, can effectively reduce the registration error, and improve the accuracy of point cloud alignment;
[0121] 2. Enhance computational efficiency: Traditional ICP methods perform nearest neighbor search in the global range, with a high computational complexity and difficulty in meeting the requirements of real-time applications. The superpoint matching strategy proposed in the present invention conducts similarity matching by constructing a high-dimensional feature space, significantly reducing the search space. In addition, an adaptive weight adjustment mechanism is combined in the improved ICP optimization process, so that the contribution of the matching point pairs is constrained by feature similarity and geometric consistency, avoiding the interference of false matches, and thus accelerating the convergence speed. Experiments show that compared with the standard ICP, this method optimizes the computational efficiency while ensuring the registration accuracy, achieving faster point cloud alignment.
[0122] 3. Enhance adaptability to dynamic environments: Due to the movement or occlusion of target objects in dynamic environments, traditional point cloud registration methods are easily affected, resulting in misregistration or drift. The present invention adopts a multi-modal feature fusion strategy, combines BEV information to provide more stable global structure constraints, and at the same time, through dynamically adjusting the attention range and the multi-level ICP optimization strategy, realizes robust processing of dynamic objects. This method is particularly suitable for autonomous driving, robot navigation, and dynamic environment reconstruction tasks, and can maintain stable registration effects in different scenarios.
[0123] In summary, the method of the present invention is superior to the prior art in terms of registration accuracy, computational efficiency, and adaptability to dynamic environments, providing an efficient and robust solution for point cloud registration and having broad application prospects.
[0124] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights.
Claims
1. A multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP, characterized in that: At least include the following steps: S1: Input, receiving the source measurement point cloud P and the source BEV view P ~ , as well as the target measurement point cloud Q and the target BEV view Q ~ ; S2: Perform feature extraction through the feature extraction module, perform point-level feature encoding on the source point cloud P and the target point cloud Q respectively, and then extract the point-level features F p and F q ; Encode the source BEV view P ~ and the target BEV view Q ~ to obtain the BEV features and S3: Build a feature fusion module, and respectively fuse the feature F of the source point cloud P p with the feature of the source BEV to obtain Fuse the feature F of the target point cloud Q q with the feature of the target BEV to obtain S4: Build a spatial hierarchical adaptive attention module. The spatial hierarchical adaptive attention module dynamically adjusts the attention range according to the distance between each point and other points in the global context. The spatial hierarchical adaptive attention module includes three levels: local module attention, regional expansion attention, and global coupling attention; S5: Build a superpoint matching module to map features and to a shared feature space through superpoint matching; S6: Build an iterative closest point optimization module and implement fine registration using an improved ICP method. The improved ICP method combines dynamic corresponding point adjustment, adaptive weight optimization, and hierarchical convergence strategy to further optimize the rigid body transformation between point clouds, so as to improve the convergence speed and the final registration accuracy, and then output the final rigid body transformation matrix.
2. The multimodal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP according to claim 1, characterized in that: The feature extraction module is a bimodal joint encoding architecture, consisting of a point encoder and a bird's-eye view encoder. The point encoder encodes through point cloud geometry, and the bird's-eye view encoder encodes through BEV semantics; The feature extraction module enhances the feature expression ability through the complementary fusion of point cloud geometry encoding and BEV semantic encoding; The point encoder adopts a four-level KPConv convolutional network, performs spatial downsampling and feature extraction layer by layer. After four-level processing, it outputs a superpoint set with 1 / 32 downsampling. Each superpoint carries a 512-dimensional feature vector, significantly reducing the computational complexity while retaining key geometric structure information; The BEV feature encoding constructs a four-stage feature pyramid based on the improved ResNet-18, and introduces a channel attention module in the fourth stage, thereby adaptively enhancing the semantic feature weights of key regions. The key regions include but are not limited to road boundaries and building contours.
3. The multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP according to claim 1, wherein: The feature fusion derivation of the feature fusion module at least includes the following steps: First, perform point cloud to BEV projection, that is, map the 3D point cloud to the 2D bird's-eye view (BEV). The BEV projection provides a global top-down view of the scene, which helps to extract structured features. Given a point cloud P, where (x i , y i , z i ) represents the coordinates of each point, project it into the BEV image I ∈ RH×W through formula (1), and calculate the pixel position (ui, vi) of the point; Among them, x max , x min , y max and y min represent the boundary range of the point cloud in the XY plane, that is, the maximum and minimum boundaries of the point cloud in X and Y; W×H represents the resolution size of the BEV image; represents rounding down; Then, the weight calculation of the feature weights is performed. To improve the quality of the features extracted from the BEV image, semantic weights need to be assigned to it. For each point p in the point cloud i , by calculating the features within the corresponding BEV image region, the fusion of the feature weights is obtained. Specifically, the BEV image feature F IP (u, v) and the interpolation kernel function G(p i , u, v) are combined and expressed by formula (2): Through this formula, the points are weighted and aggregated using BEV image features, and the interpolation kernel function G(p i , u, v) integrates neighborhood information by considering the contribution of each pixel in the BEV image to the point cloud features, thereby achieving feature fusion; Adopt the bilinear interpolation algorithm to aggregate the BEV features in the neighborhood, map the BEV features to the point cloud feature space through a learnable weight matrix, and the feature fusion strategy performs channel splicing on the point cloud geometric features and the projected BEV features.
4. The multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP according to claim 1, characterized in that: In the initial layer of the spatial hierarchical adaptive attention module, that is, when L = 1, local module attention is adopted. The spatial hierarchical adaptive attention module realizes the refinement of local association by focusing on nearby superpoints, thus effectively reducing the ambiguity that may be caused by the initial global attention; In the subsequent layers, that is, when L>1, regional expansion attention is adopted, and the attention range gradually expands, so as to be able to establish a multi-scale spatial representation of superpoints. This step-by-step method enhances the robustness of superpoint features; Calculation at the L = k layer The process of using global coupled attention to achieve stepwise self-attention, the masked attention score is expressed as: Among them, is a layer-specific mask used to filter attention scores based on the Euclidean distance between points; Mask is calculated by dividing the maximum distance between the superpoint and all other superpoints into S segments. For a given layer L = k, the distance threshold is defined as: Wherein: represents the maximum Euclidean distance between superpoints; is a factor used to adjust the attention range of each layer; Boolean mask matrix: Indicates that when the Euclidean distance between two superpoints is less than or equal to the threshold then the attention score between them is retained, i.e., the mask value is 1; otherwise, the attention score is filtered out, i.e., the mask value is 0; Mask attention calculation: Z (k) = Softmax(M (k) ⊙E)·XW V (6) Among them, M (k) is the mask matrix of the L = k-th layer, E represents the attention score matrix, and XW V represents the value projection of the point cloud feature; the role of the attention score matrix E here is to calculate an attention weight for each pair of points, while the mask matrix M (k) determines which attention scores will be retained. After Softmax normalization, the final output is calculated through weighted superpoint features.
5. The multi-modal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP according to claim 1, characterized in that: In the feature normalization stage, the superpoint matching module maps the encoded and to the unit spherical space to obtain the normalized features and Subsequently, the Gaussian similarity matrix is expressed as follows: Among them, γ is a hyperparameter used to control the similarity decay rate; is the normalized feature and is the Euclidean distance between; By performing double normalization of rows and columns to eliminate distribution bias: represents the normalization factor for the i-th row, representing the sum of the similarities between point i and all other points; represents the normalization factor for the j-th column, representing the sum of the similarities between point j and all other points Form a matching set by selecting the superpoint pair with the highest correlation score as the initial correspondence for subsequent fine registration.
6. The multimodal point cloud dynamic registration method based on adaptive hierarchical attention and efficient ICP according to claim 5, characterized in that: The improved ICP method at least includes the following steps: In the superpoint matching stage, a set of initial matching point pairs with high confidence is obtained as the input for ICP fine registration; Since reliable matching point pairs have been obtained, directly calculate the rigid body transformation (R0, T0), avoiding the initialization of traditional ICP relying on nearest neighbor search; the transformation is solved by weighted least squares, the rotation matrix R0 is obtained by SVD decomposing the covariance matrix, and the translation vector T0 is calculated; this initialization method can significantly accelerate the convergence speed and improve the accuracy; In the subsequent ICP iteration process, dynamically optimize the matching point pairs, transform the source point cloud, and update the point cloud position. Local nearest neighbor search. If the error of the current matching point exceeds the set threshold, it is updated to the local optimal match, and the matching is screened by combining the feature similarity to avoid the interference of false matches on ICP optimization; Then, adaptive weight optimization is performed. To improve the reliability of matching point pairs, the contribution of each point pair is controlled by a weight w i which is calculated based on feature similarity, making points with more similar features have higher weights; and geometric consistency, ensuring the consistency of matching points in the local structure To optimize the calculation efficiency and registration accuracy, hierarchical ICP optimization is adopted: Superpoint-level ICP, only using the superpoint matching point pairs for rough registration to obtain the global transformation; Local region ICP, performing ICP in the superpoint neighborhood to correct the local error; Global fine ICP, performing the final optimization based on the complete point cloud to ensure the overall accuracy.
Citation Information
Cited By
Multi-modal point cloud tower inclination detection method based on Lie group structure self-attention
CN121259071A