Robust point cloud registration method and system based on global spatial perception and multistage filtering

By introducing global spatial perception and multi-level filtering technology into point cloud registration, combined with transformer network for feature processing and filtering, the problem of information loss and error correspondence in point cloud registration is solved, and the accuracy and robustness of registration are improved.

CN119941807APending Publication Date: 2025-05-06YANTAI UNIV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411797641.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

During point cloud registration, downsampling leads to loss of global and local information, and there are inevitably errors in the correspondence generated after feature matching, affecting the registration accuracy and robustness.

Method used

A robust point cloud registration method based on global spatial perception and multi-level filtering is adopted. By acquiring the initial characteristics of point clouds and global spatial structural characteristics, feature aggregation and information interaction are combined with transformer networks, and then upsampling and multi-level filtering are performed to obtain the transformation matrix.

Benefits of technology

Effectively capture the rich spatial relationships between point clouds, supplement global context information, improve network robustness and registration accuracy, and reduce the impact of errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941807A_ABST
    Figure CN119941807A_ABST
Patent Text Reader

Abstract

The invention discloses a robust point cloud registration method and system based on global spatial perception and multistage filtering. The method comprises the following steps: respectively obtaining point cloud initial features and spatial structure features of global spatial perception based on point cloud information; inputting the point cloud initial features and the spatial structure features into a transform network for feature enhancement, and obtaining enhanced features; and performing up-sampling processing and multi-stage filtering on the enhanced features to obtain a transformation matrix, and completing point cloud registration. According to the method, global spatial perception and multi-stage filtering are introduced, so that abnormal values can be effectively eliminated while feature matching is completed finally, and the precision and efficiency of point cloud registration are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of point cloud registration, and in particular to a robust point cloud registration method and system based on global space perception and multi-level filtering. Background Art

[0002] Point cloud registration is a key technology for 3D reconstruction. Point cloud registration is the problem of estimating the transformation matrix required to align two overlapping point clouds. Point cloud registration is widely used in traditional fields such as 3D reconstruction, localization, and pose estimation, and has also been applied in emerging fields such as autonomous driving, robotics, and virtual reality. With the rise of deep learning methods, the application of deep learning to point cloud registration has made sufficient progress, but improving registration accuracy and high robustness in the point cloud registration process is still a challenging task.

[0003] Regardless of the method used for point cloud registration, the point cloud must first be downsampled. This downsampling will lead to the loss of global and local information in the point cloud, and this loss of information seems to be inevitable. After completing the matching through features, the generated correspondences will inevitably have erroneous correspondences that interfere with the registration results. How to remove erroneous correspondences as much as possible is also a difficult problem. Summary of the invention

[0004] In order to solve the above technical problems, the present invention proposes a robust point cloud registration method and system based on global space perception and multi-level filtering to provide a solution to achieve a balance between accuracy and robustness.

[0005] On the one hand, to achieve the above object, the present invention provides a robust point cloud registration method based on global space perception and multi-level filtering, comprising:

[0006] Based on the point cloud information, the initial features of the point cloud and the spatial structure features of global space perception are obtained respectively;

[0007] Inputting the initial features of the point cloud and the spatial structure features into a transformer network for feature aggregation and information interaction to obtain enhanced features;

[0008] The enhanced features are subjected to upsampling and multi-level filtering to obtain a transformation matrix and complete point cloud registration.

[0009] Preferably, the process of obtaining the initial features of the point cloud is: performing a downsampling operation and a feature extraction operation on the point cloud information, wherein the point cloud information includes a source point cloud and a target point cloud;

[0010] Wherein, the feature extraction operation is:

[0011] Using kernel point convolution KPConv, through multi-layer downsampling, the original point cloud is regarded as a dense point cloud, and the dense point cloud is sampled into a sparse point cloud, while the initial features of the point cloud are extracted.

[0012] Preferably, the spatial structural features of the global spatial perception are obtained as follows:

[0013] Each point in the sparse point cloud is taken as the core, and a local patch is constructed with several points closest to the core point cloud. The point pair features between each point in the local patch and the core point are calculated, and the point pair features are aggregated with the global context position information through an aggregation network to obtain the spatial structure features of the global space perception.

[0014] Preferably, the aggregation network is composed of several one-dimensional convolutions, maximum pooling layers and multi-layer perceptrons, wherein each one-dimensional convolution includes an activation function and a normalization layer, and the outputs of all one-dimensional convolutions and the outputs of the maximum pooling layer are linked and sent to the multi-layer perceptron together.

[0015] Preferably, the transformer network is constructed by stacking a three-layer structure of a first graph attention module, a cross attention module, and a second graph attention module.

[0016] Preferably, the initial point cloud features and the spatial structure features are input into a transformer network for feature aggregation and information interaction to obtain enhanced features, specifically:

[0017] A mixed feature is obtained by bitwise addition of the spatial structure feature and the initial feature of the point cloud, and the mixed feature and the spatial structure feature are input into the first graph attention module for processing; wherein, in the first graph attention module, the spatial structure feature is subjected to a weight score obtained by a Sigmoid() mapping function, and the weight score is weightedly multiplied with the graph attention feature to obtain an enhanced feature; the operation in the second graph attention module is consistent with that in the first graph attention module.

[0018] Preferably, the enhanced features are subjected to upsampling and multi-stage filtering to obtain a transformation matrix:

[0019] Upsampling the enhanced features to obtain dense point cloud features, and performing linear projection and similarity distribution calculation on the dense point cloud features to obtain overlapping probability distribution and significant probability distribution respectively;

[0020] The overlapping probability distribution and the significance probability distribution are multiplied to obtain a comprehensive score, and a point cloud with a high score is selected from the dense point cloud according to the comprehensive score as a reliable point cloud for spatial confidence filtering to obtain the transformation matrix.

[0021] On the other hand, to achieve the above object, the present invention also provides a robust point cloud registration system based on global space perception and multi-level filtering, comprising:

[0022] Point cloud feature acquisition module: used to obtain the initial features of the point cloud and the spatial structure features of global space perception based on the point cloud information;

[0023] Transformer processing module: used for inputting the initial features of the point cloud and the spatial structure features into the transformer network for feature aggregation and information interaction to obtain enhanced features;

[0024] Point cloud registration module: used to perform upsampling and multi-level filtering on the enhanced features, obtain the transformation matrix, and complete point cloud registration.

[0025] Preferably, the point cloud feature acquisition module includes:

[0026] A first point cloud feature acquisition unit: used to downsample the point cloud information and extract the initial features of the point cloud;

[0027] The second point cloud feature acquisition unit is used to acquire spatial structure features with global space perception.

[0028] Preferably, the point cloud registration module comprises:

[0029] An upsampling unit is used to upsample the enhanced features to obtain dense point cloud features, and obtain overlapping probability distribution and saliency probability distribution of the dense point cloud features through linear projection and similarity distribution calculation respectively;

[0030] Multi-stage filtering unit: used for multiplying the overlapping probability distribution and the significance probability distribution to obtain a comprehensive score, selecting reliable point clouds from the dense point cloud according to the comprehensive score to perform spatial confidence filtering, and obtaining the transformation matrix.

[0031] Compared with the prior art, the present invention has the following advantages and technical effects:

[0032] (1) The present invention provides a robust point cloud registration method based on global spatial perception and multi-level filtering. By combining point cloud features and spatial structural features with global spatial perception, it can better capture the rich spatial relationships between point clouds, supplement global context information, and better adapt to the distribution of different point clouds. This enables the network to understand the local spatial features and global structural features of the point cloud at the same time, and obtain richer information to improve the robustness of the network. By introducing global spatial perception and multi-level filtering, outliers can be effectively excluded while completing feature matching at the end, thereby improving the accuracy and efficiency of point cloud registration.

[0033] (2) The spatial structural features obtained by the global spatial perception method in the present invention can effectively capture spatial information that is easily overlooked in point cloud registration, and are used to supplement the missing structural information in downsampling, add global context information, and ensure the accuracy of the predicted transformation. The spatial structural features are also introduced into the transformer's graph attention and participate in the calculation as attention weight parameters. This innovation can effectively enhance the model's ability to focus on key registration areas, and at the same time can better capture the common features between two point clouds in cross-attention, enhance the correlation between data, and thus improve the registration success rate, bringing new ideas and possibilities to research and application in related fields.

[0034] (3) The present invention also proposes a multi-level filtering method, which is composed of a reliable point selection module and an optimal transmission-guided spatial confidence filtering method. It can ensure that the point cloud is screened first before feature matching, removes a large number of interfering point clouds in non-overlapping areas, and greatly improves the registration speed. The optimal transmission-guided spatial confidence filtering method uses the optimal transmission guidance to construct a spatial consistency matrix. Through the characteristics of spatial consistency, it can effectively remove erroneous matches in the correspondence and improve the registration accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0036] Figure 1 A flowchart of a robust point cloud registration method based on global space perception and multi-stage filtering according to an embodiment of the present invention;

[0037] Figure 2 It is a schematic diagram of the main framework of an embodiment of the present invention;

[0038] Figure 3 Schematic diagram of global spatial perception of an embodiment of the present invention: (a) is a schematic diagram of obtaining spatial structure features using global spatial perception, and (b) is a schematic diagram of using spatial structure features in a transformer structure;

[0039] Figure 4 A schematic diagram of the use of spatial structure features in graph attention according to an embodiment of the present invention;

[0040] Figure 5 A schematic diagram of a spatial confidence filtering module for optimal transmission guidance of multi-stage filtering according to an embodiment of the present invention. DETAILED DESCRIPTION

[0041] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0042] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0043] First, the technical terms in this embodiment are explained:

[0044] 1. Global spatial perception:

[0045] Global spatial perception in point cloud data processing refers to the comprehensive spatial understanding and analysis of the entire scene through point cloud data, including information such as the scene's geometric structure, object positions, and spatial relationships.

[0046] Global spatial perception methods for point cloud data include:

[0047] Point cloud registration: By aligning multiple point cloud data sets to form a complete 3D model, we can understand the global space. The registration process needs to consider factors such as the overlapping area between point clouds and the transformation matrix.

[0048] 3D reconstruction: Use point cloud data to reconstruct the 3D model of an object. This requires point cloud preprocessing, feature extraction, surface reconstruction and other steps to ultimately form a complete 3D model.

[0049] Semantic segmentation: Classify point cloud data into different objects or areas, such as ground, buildings, vegetation, etc. This requires labeling of each point in the point cloud, usually combined with methods such as deep learning.

[0050] In 3D reconstruction, global spatial awareness is achieved through the following steps:

[0051] Data collection: Use laser scanning, depth cameras, binocular cameras and other methods to obtain point cloud data.

[0052] Preprocessing: including denoising, filtering and other operations to improve the quality of point cloud data.

[0053] Feature extraction: Extract geometric features, texture features, etc. from the point cloud for subsequent processing and analysis.

[0054] Surface reconstruction: Based on the extracted features, the surface of the object is reconstructed using voxelization, meshing and other methods.

[0055] 2. Multi-stage filtration:

[0056] In point cloud data processing, multi-level filtering is a technology that filters data step by step through multiple filtering stages, aiming to improve filtering accuracy and reduce the rate of false positives. Multi-level filtering technology is based on traditional filtering technology, which introduces multiple filtering stages to filter data step by step, ultimately retaining the most valuable data. Each stage is usually based on different filtering dimensions, such as rule-based, statistical, or machine learning methods, and ultimately makes a comprehensive decision based on the output of filters at each stage.

[0057] Multi-level filtering technology can filter data at multiple levels, thereby improving filtering accuracy and reducing the rate of false positives. It is also applicable to various complex data processing scenarios and can handle different types of data noise and outliers.

[0058] 3. Point cloud data:

[0059] Point cloud data refers to a set of vectors in a three-dimensional coordinate system, usually expressed in the form of X, Y, and Z three-dimensional coordinates, which is used to represent the outer surface shape of an object. In addition to geometric position information, point cloud data can also contain other attributes, such as RGB color, grayscale value, depth, and segmentation results. Point cloud data acquisition methods include laser radar, stereo vision, and structured light technologies.

[0060] Method for obtaining point cloud data:

[0061] LiDAR: LiDAR acquires point cloud data of the target object’s surface by emitting a laser beam and measuring the time and intensity of the returning laser. This method can collect high-precision three-dimensional data and is widely used in fields such as robot navigation and three-dimensional map construction.

[0062] Stereo vision: Two cameras are used to simultaneously shoot the same object, and the displacement and angle between the two cameras are calculated to determine the three-dimensional coordinates of each point on the surface of the object. Stereo vision is used in high-precision point cloud data collection, but is limited by environmental and camera factors.

[0063] Structured light: By projecting a structured light source such as a grating or a coded light strip to illuminate the surface of the target object, the camera uses information such as the reflection of the grating pattern or the deformation of the light strip to calculate the three-dimensional coordinates. Structured light requires that the surface of the object to be measured has smooth reflective properties.

[0064] This paper proposes a robust point cloud registration method based on global spatial perception and multi-level filtering, such as Figure 1 ,include:

[0065] Based on the point cloud information, the initial features of the point cloud and the spatial structure features of global space perception are obtained respectively;

[0066] Inputting the initial features of the point cloud and the spatial structure features into the transformer network for feature aggregation and information interaction to obtain enhanced features;

[0067] The enhanced features are upsampled and multi-level filtered to obtain the transformation matrix and complete the point cloud registration.

[0068] This embodiment can better capture the rich spatial relationships between point clouds, supplement global context information, and better adapt to the distribution of different point clouds by combining point cloud features with spatial structural features with global spatial perception. This enables the transformer network to understand the local spatial features and global structural features of the point cloud at the same time, and obtain richer information to improve the robustness of the network. By introducing global spatial perception and multi-level filtering, outliers can be effectively excluded while completing feature matching at the end, which significantly improves the accuracy and efficiency of point cloud registration.

[0069] Furthermore, the point cloud information includes: a source point cloud and a target point cloud.

[0070] The initial features of the point cloud are obtained by performing downsampling operations and feature extraction operations on the point cloud information to obtain the initial features of the point cloud with point cloud information and the spatial structure features with global spatial perception.

[0071] Specifically, the feature extraction operation is:

[0072] Using kernel point convolution KPConv, through multi-layer downsampling, the original point cloud is regarded as a dense point cloud, and the dense point cloud is sampled into a sparse point cloud, while the initial features of the point cloud are extracted.

[0073] The method for obtaining spatial structural features of global spatial perception includes: taking each point in the sparse point cloud as the core, constructing a local patch based on several points closest to the core point cloud, calculating the point pair features between each point in the local patch and the core point, and aggregating the obtained point pair features with global context position information as features through an aggregation network.

[0074] Furthermore, the aggregation network is composed of several one-dimensional convolutions, maximum pooling layers and multi-layer perceptrons, wherein each one-dimensional convolution includes a ReLu() activation function and a normalization layer, and the outputs of all one-dimensional convolutions and the outputs of the maximum pooling layer are linked through the concat function () and sent to the multi-layer perceptron together.

[0075] In this embodiment, the aggregation network includes: three one-dimensional convolutions, a maximum pooling layer, and a multi-layer perceptron. After each one-dimensional convolution, a ReLu() activation function and a normalization layer are used, and the outputs of the three one-dimensional convolutions and the output of the maximum pooling layer are linked together through the concat function () and sent to the multi-layer perceptron.

[0076] Furthermore, the transformer network diagram is constructed by stacking three layers of structures: first-image attention module-cross-attention module-second-image attention module.

[0077] Furthermore, the initial features of the point cloud and the spatial structure features are input into the transformer network for feature aggregation and information interaction to obtain enhanced features, specifically:

[0078] A mixed feature is obtained by bitwise addition of the spatial structure feature and the initial feature of the point cloud, and the mixed feature and the spatial structure feature are input into the first graph attention module for processing; wherein, in the first graph attention module, the spatial structure feature is subjected to a weight score obtained by a Sigmoid() mapping function, and the weight score is weightedly multiplied with the graph attention feature to obtain an enhanced feature; the operation in the second graph attention module is consistent with that in the first graph attention module.

[0079] In this embodiment, the method of using the spatial structure feature input transformer includes: the spatial structure feature will be bitwise added with the point cloud feature before entering the graph attention module; when completing the graph attention module, the spatial structure feature will first be processed using the Sigmoid() mapping function and used as a weight to weight the final graph attention feature.

[0080] Furthermore, the enhanced features are upsampled and multi-level filtered to obtain the transformation matrix:

[0081] The enhanced features are upsampled to obtain dense point cloud features, and the dense point cloud features are respectively used to obtain overlapping probability distribution and significant probability distribution through linear projection and similarity distribution calculation;

[0082] The overlapping probability distribution and the significance probability distribution are multiplied to obtain a comprehensive score, and according to the comprehensive score, a point cloud with a high score is selected from the dense point cloud as a reliable point cloud for spatial confidence filtering to obtain the transformation matrix.

[0083] Specifically, the point cloud features obtained by the transformer network are upsampled to obtain dense point cloud features, and these features are used to obtain overlapping probability distribution and significance probability distribution through linear projection and similarity distribution calculation.

[0084] The overlapping probability distribution and the significance probability distribution are multiplied to obtain a comprehensive score, and reliable point clouds are selected from the dense point cloud for spatial confidence filtering according to the comprehensive score.

[0085] Spatial confidence filtering includes: constructing an optimal transmission-guided confidence matrix through a spatial consistency matrix, selecting reliable seed points in the confidence matrix, estimating the transformation matrix for each seed point, selecting the best one for optimization and outputting the final transformation matrix.

[0086] This embodiment also provides a robust point cloud registration system based on global space perception and multi-level filtering, including:

[0087] Point cloud feature acquisition module: used to obtain the initial features of the point cloud and the spatial structure features of global space perception based on the point cloud information;

[0088] Transformer processing module: used to input the initial features of the point cloud and the spatial structure features into the transformer network for feature aggregation and information interaction to obtain enhanced features;

[0089] Point cloud registration module: used to perform upsampling and multi-level filtering on the enhanced features, obtain the transformation matrix, and complete point cloud registration.

[0090] Furthermore, the point cloud feature acquisition module includes:

[0091] The first point cloud feature acquisition unit is used to downsample the point cloud information and extract the initial features of the point cloud;

[0092] The second point cloud feature acquisition unit is used to acquire spatial structure features with global space perception.

[0093] Furthermore, the point cloud registration module includes:

[0094] Upsampling unit: used to upsample the enhanced features to obtain dense point cloud features, and obtain overlapping probability distribution and saliency probability distribution of the dense point cloud features through linear projection and similarity distribution calculation respectively;

[0095] Multi-level filtering unit: used to multiply the overlapping probability distribution and the significance probability distribution to obtain a comprehensive score, select reliable point clouds from the dense point cloud according to the comprehensive score, perform spatial confidence filtering, and obtain a transformation matrix.

[0096] This embodiment proposes a global spatial perception method. The spatial structural features obtained by the global spatial perception method can effectively capture spatial information that is easily overlooked in point cloud registration, and are used to supplement the missing structural information in downsampling, add global context information, and ensure the accuracy of the predicted transformation. Spatial structural features are also introduced into the transformer's graph attention and participate in the calculation as attention weight parameters. This innovation can effectively enhance the model's ability to focus on key registration areas, and at the same time can better capture the common features between two point clouds in cross-attention, enhance the correlation between data, and thus improve the registration success rate, bringing new ideas and possibilities for research and application in related fields.

[0097] The multi-level filtering method proposed in this embodiment is composed of a reliable point selection module and an optimal transmission-guided spatial confidence filtering method. It can ensure that the point cloud is screened first before feature matching, removes a large number of interfering point clouds in non-overlapping areas, and greatly improves the registration speed. The optimal transmission-guided spatial confidence filtering method uses the optimal transmission guidance to construct a spatial consistency matrix. Through the characteristics of spatial consistency, it can effectively remove erroneous matches in the correspondence and improve the registration accuracy.

[0098] In order to more clearly express the technical solution of the present invention, the following specific embodiments are provided to introduce the solution:

[0099] Combining this embodiment and Figure 2 , and explain in detail how the present invention solves technical problems in actual work.

[0100] First, the source point cloud and target point cloud required for point cloud registration are obtained. This embodiment uses the mainstream standard datasets in the field of point cloud registration. Specifically, you can download and obtain them by visiting the official websites of commonly used public datasets such as 3DMatch, 3DLoMatch, and KITTI odometry. These datasets are widely used in this field and can provide important support for the development and evaluation of point cloud registration methods.

[0101] Based on the convolutional neural network, an encoding layer is constructed to downsample and extract the initial feature information of the point cloud.

[0102] Using kernel point convolution KPConv, through multi-layer downsampling, the dense point cloud is sampled into a sparse point cloud, and the initial features of the point cloud are extracted at the same time.

[0103] Each point in the sparse point cloud is used as the core to construct a local patch, and the point pair features between each point in the local patch and the core point are calculated to complete the structural encoding. The obtained point pair features are aggregated through the aggregation network to aggregate the global context position information as the spatial structural features. Among them, the aggregation network consists of three one-dimensional convolutions, a maximum pooling layer, and a multi-layer perceptron. After each one-dimensional convolution, the ReLu() activation function and the normalization layer are used, and the outputs of the three one-dimensional convolutions and the output of the maximum pooling layer are passed through. The process of obtaining spatial structural features is as follows Figure 3 (a) shown.

[0104] The initial features and spatial structure features of the point cloud are sent to the transformer. The transformer consists of a three-layer architecture: the first image attention module-the cross attention module-the second image attention module.

[0105] First, the spatial structure features will be bitwise added to the initial features of the point cloud before being input into the first image attention module to supplement the point cloud structure information; they will also be sent to the first image attention module as weights to weight the final image attention features. Figure 3 (b) as shown.

[0106] In the first image attention module, the spatial structure features are first processed using the mapping function σ() and used as weights.

[0107] The mixed features between the initial features of the point cloud and the spatial structure features before the first attention module are linearly mapped to W Q It is also a layer of feature F 1 , linearly map the point cloud position into W K ,W V .W Q With W K At the same time, it is sent to the graph neural network to extract the second-layer feature F 2 , the obtained F 2 Then with W V At the same time, it is sent to the graph neural network to extract the three-layer feature F 3 . The three-layer feature F 1 ,F 2 ,F 3 Through the link layer, they are linked to form a comprehensive feature F MLP , and cross-multiply with the above weights, we can get a new feature Output with spatial structure feature weights.

[0108] The specific summary is:

[0109] F MLP =MLP(Concat(F 1 ,F 2 ,F3 ));

[0110] Output=σ(W P )*F MLP ;

[0111] Where MLP is a multi-layer perceptron, Concat is the concat() function, and W P It is the spatial structure feature.

[0112] Specific design such as Figure 4 shown.

[0113] The sparse point cloud features obtained through the transformer network are propagated to the entire point cloud, and the evaluation scores are obtained: overlap probability and significance probability.

[0114] The overlapping probability distribution and the significance probability distribution are multiplied to obtain a comprehensive score. The comprehensive score is used as the probability according to the size of the score. Reliable point features are selected from the entire point cloud and correspondence is generated through feature matching. This is used as the spatial confidence filtering for the next step of optimal transmission guidance.

[0115] The optimal transmission guided spatial confidence filtering constructs a feature similarity matrix C as the transmission cost by matching the generated features, and finds the cost p at the i-th and j-th positions on C. i and q j , the transmission amount (weight) of T at the i-th and j-th positions. Use the optimal transmission algorithm to iteratively optimize the similarity matrix to obtain the optimal transmission confidence matrix T, which is:

[0116]

[0117] Using the soft-nms method, the highest score is selected from the optimal transmission confidence matrix as the seed for the subsequent construction of the spatial consistency matrix. Then the optimal transmission confidence matrix and the spatial consistency matrix are sent together for transformation matrix rotation and optimization, and finally the transformation matrix {R, t} is obtained to complete the point cloud registration. Figure 5 shown.

[0118] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A robust point cloud registration method based on global spatial perception and multi-level filtering, characterized in that: include: Based on the point cloud information, the initial features of the point cloud and the spatial structure features of global space perception are obtained respectively; Inputting the initial features of the point cloud and the spatial structure features into a transformer network for feature aggregation and information interaction to obtain enhanced features; The enhanced features are subjected to upsampling and multi-level filtering to obtain a transformation matrix and complete point cloud registration.

2. The method according to claim 1, characterized in that The process of obtaining the initial features of the point cloud is: performing a downsampling operation and a feature extraction operation on the point cloud information, wherein the point cloud information includes a source point cloud and a target point cloud; Wherein, the feature extraction operation is: Using kernel point convolution KPConv, through multi-layer downsampling, the original point cloud is regarded as a dense point cloud, and the dense point cloud is sampled into a sparse point cloud, while the initial features of the point cloud are extracted.

3. The method according to claim 2, characterized in that The spatial structural features of the global spatial perception are obtained as follows: Each point in the sparse point cloud is taken as the core, and a local patch is constructed based on several points closest to the core point cloud. The point pair features between each point in the local patch and the core point are calculated, and the point pair features are aggregated with the global context position information through an aggregation network to obtain the spatial structure features of the global space perception.

4. The method according to claim 3, characterized in that The aggregation network is composed of several one-dimensional convolutions, maximum pooling layers and multi-layer perceptrons, wherein each one-dimensional convolution includes an activation function and a normalization layer, and the outputs of all one-dimensional convolutions and the outputs of the maximum pooling layer are linked and sent to the multi-layer perceptron together.

5. The method according to claim 1, characterized in that The transformer network is composed of a three-layer structure consisting of a first image attention module, a cross attention module, and a second image attention module.

6. The method according to claim 5, characterized in that The initial features of the point cloud and the spatial structure features are input into the transformer network for feature aggregation and information interaction to obtain enhanced features, specifically: A mixed feature is obtained by bitwise addition of the spatial structure feature and the initial feature of the point cloud, and the mixed feature and the spatial structure feature are input into the first graph attention module for processing; wherein, in the first graph attention module, the spatial structure feature is subjected to a weight score obtained by a Sigmoid() mapping function, and the weight score is weightedly multiplied with the graph attention feature to obtain an enhanced feature; the operation in the second graph attention module is consistent with that in the first graph attention module.

7. The method according to claim 6, characterized in that The enhanced features are subjected to upsampling and multi-stage filtering to obtain a transformation matrix: Upsampling the enhanced features to obtain dense point cloud features, and performing linear projection and similarity distribution calculation on the dense point cloud features to obtain overlapping probability distribution and significant probability distribution respectively; The overlapping probability distribution and the significance probability distribution are multiplied to obtain a comprehensive score, and a point cloud with a high score is selected from the dense point cloud according to the comprehensive score as a reliable point cloud for spatial confidence filtering to obtain the transformation matrix.

8. A robust point cloud registration system based on global spatial perception and multi-level filtering, applied to the method according to any one of claims 1 to 7, characterized in that: include: Point cloud feature acquisition module: used to obtain the initial features of the point cloud and the spatial structure features of global space perception based on the point cloud information; Transformer processing module: used for inputting the initial features of the point cloud and the spatial structure features into the transformer network for feature aggregation and information interaction to obtain enhanced features; Point cloud registration module: used to perform upsampling and multi-level filtering on the enhanced features, obtain the transformation matrix, and complete point cloud registration.

9. The system according to claim 8, characterized in that The point cloud feature acquisition module includes: A first point cloud feature acquisition unit: used to downsample the point cloud information and extract the initial features of the point cloud; The second point cloud feature acquisition unit is used to acquire spatial structure features with global space perception.

10. The system according to claim 8, characterized in that The point cloud registration module comprises: An upsampling unit is used to upsample the enhanced features to obtain dense point cloud features, and obtain overlapping probability distribution and saliency probability distribution of the dense point cloud features through linear projection and similarity distribution calculation respectively; Multi-stage filtering unit: used for multiplying the overlapping probability distribution and the significance probability distribution to obtain a comprehensive score, selecting reliable point clouds from the dense point cloud according to the comprehensive score to perform spatial confidence filtering, and obtaining the transformation matrix.

Citation Information

Cited By

  • Robust point cloud registration method and device based on common-view query embedding

    CN120543608A

  • A three-dimensional environment perception method and device based on cross-modal bidirectional prior guidance

    CN122506545A

  • A three-dimensional environment perception method and device based on cross-modal bidirectional prior guidance

    CN122506545B