Three-dimensional point cloud registration model design method for global and local attention
Through a three-dimensional point cloud registration model of global and local attention, the problem of feature extraction and matching in point cloud registration is solved, and efficient and accurate point cloud registration is achieved, which is suitable for fields such as intelligent transportation and autonomous driving.
Patent Information
- Application Number
- CN202510194691.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art is difficult to effectively extract features and improve the accuracy and efficiency of point pair matching and transformation matrix calculations in three-dimensional point cloud registration, which affects the accuracy and efficiency of registration.
A three-dimensional point cloud registration model of global and local attention was designed. By collecting multi-view point cloud data, a backbone feature extraction network of the deep network model is constructed, and feature matching is performed by combining global and local attention mechanisms. The rigid body transformation matrix is calculated through iterative optimization to generate a complete point cloud.
It improves the accuracy and efficiency of point cloud registration, can adaptively calculate the transformation matrix, simplify the registration process, and ensure the reliability and flexibility of the results.
Smart Images

Figure CN120259378A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a method for designing a three-dimensional point cloud registration model with global and local attention. Background Art
[0002] Three-dimensional point cloud registration refers to calculating a spatial transformation matrix through rotation and translation, and converting point cloud data sets under two completely different coordinate systems to the same coordinate system through the spatial transformation matrix, so that the overlapping parts of two point cloud data sets with different postures are aligned, realizing a more accurate description of the same target or scene, and being widely used in many industries and fields such as robot navigation, three-dimensional scene reconstruction, and optical measurement. At present, most deep learning registration methods are based on correspondence matching. The neural network finds key points and their related feature descriptors in the point cloud, then extracts fine correspondences, and inputs them into RANSAC and other robust estimators to recover the transformation matrix. In the field of point cloud registration, the attention mechanism has been proven to be an effective way to strengthen the deep feature relationship for matching the point pair relationship between the source point cloud and the target point cloud during point cloud registration. However, there are still some challenges in point cloud registration, and it is necessary to solve how to effectively extract point cloud features, improve the matching between point pairs, and the calculation between transformation matrices to improve the configuration accuracy and efficiency.
[0003] The present invention proposes a method for designing a three-dimensional point cloud registration model with global and local attention, aiming to efficiently extract the feature relationship between the source point cloud and the target point cloud, and provide a new, more flexible and accurate point cloud registration method. By fully considering uncertainty factors, combining historical data and simulation technology, the present invention aims to provide an effective solution for the development of intelligent transportation systems and autonomous driving technologies. Summary of the Invention
[0004] To achieve the above and other related purposes, the present invention discloses a method for designing a three-dimensional point cloud registration model with global and local attention, including: Step S1: Collect discrete sampling data of the scene surface from different perspectives to form a source point cloud and a target point cloud, and the point cloud data is a set of three-dimensional coordinate points; Step S2: Construct a backbone feature extraction network of a deep network model, and extract multi-level features of the source point cloud and the target point cloud respectively. The backbone feature extraction network includes a kernel point convolution layer and a feature fusion module; Step S3: Based on the extracted multi-level features, perform rough matching of feature points through the global and local attention mechanism, and combine point-level fine matching in the point cloud intersection area to generate point pair correspondence; Step S4: Based on the point pair correspondence, calculate the rigid body transformation matrix from the source point cloud to the target point cloud through iterative optimization; Step S5: Transform the source point cloud to the target point cloud coordinate system using the rigid body transformation matrix, complete the registration, and generate the complete point cloud.
[0005] Further, the step S1 includes: Collect multi-viewpoint cloud data of the scene through a three-dimensional sensing device, and the device includes a lidar, a depth camera, or a three-dimensional scanner; Perform noise reduction, downsampling, and semantic annotation processing on the original point cloud data; Define the storage formats of the source point cloud data, the target point cloud data, and the corresponding true transformation matrix.
[0006] Further, the calculation of the kernel point convolution layer satisfies the following formula: ; where, is the input feature set, is the convolution kernel parameter, is the convolution center point, is the point set within the domain of, is a certain point within the domain, is the corresponding feature vector, indicating a weighted sum of all points within the neighborhood of the convolution center point, and the weights are determined by the convolution kernel parameters, so as to obtain a new output feature.
[0007] Further, in step S3, the global and local attention mechanism includes: Calculate the similarity scores between feature points through the query weight matrix, the key weight matrix, and the value weight matrix; Based on the similarity scores, perform weighted fusion on the feature vectors to generate the dependencies within and across the point clouds; Generate a matching score matrix through the dot product operation of the normalized feature matrix, and screen the feature point pairs corresponding to the maximum matching score.
[0008] Further, step S4 includes: Randomly select three pairs of non-collinear matching points to calculate the initial rigid body transformation matrix; Screen the inliers through a distance threshold, and iteratively optimize to obtain the optimal transformation matrix corresponding to the maximum inlier set.
[0009] Further, the feature fusion module realizes the fusion of global information and local information by upsampling the deep features and splicing them with the shallow features.
[0010] Further, the point-level fine matching includes: Based on the calculation of the feature vector similarity, use the nearest neighbor matching method to determine the specific point pairs in the point cloud intersection area.
[0011] Further, the complete point cloud is generated by merging the coordinate spaces of the transformed source point cloud and the target point cloud.
[0012] On the other hand, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above method is implemented.
[0013] By adopting the above technical solutions, the present invention designs a point cloud registration neural network applicable to multiple scenarios. The model can adaptively calculate the transformation matrix between the source point cloud and the target point cloud, thereby simplifying the registration process. The present invention has a strong uncertainty processing ability. This method can effectively handle the point matching problem between the source point cloud and the target point cloud, ensuring the reliability of the registration result. The coupling degree between the modules of the present invention is relatively low, which is convenient for future expansion and optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1 is a flowchart of the present invention; Figure 2 is a three-dimensional display diagram of the source point cloud of the present invention; Figure 3 is a three-dimensional display diagram of the target point cloud of the present invention; Figure 4 is a display diagram of the registration result of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0016] Refer to Figure 1 , the embodiments of the present invention provide a method for designing a three-dimensional point cloud registration model with global and local attention, including the following steps: Step S1: Collect discrete sampling data of the scene surface from different perspectives to form a source point cloud and a target point cloud. Refer to Figure 2 and Figure 3 , the point cloud data is a set of three-dimensional coordinate points; Specifically, S1 includes: Point cloud data collection: Collect point cloud data of the scene at different perspective coordinates through devices such as lidar, cameras, depth sensors, 3D scanners, laser rangefinders, structured light scanners, etc.; Data collection should cover different types of scenes (including indoor and outdoor scenes) to ensure the diversity and representativeness of the data; Use point cloud generation software to convert the RGB-D images of the scene collected by the acquisition device into point cloud data.
[0017] Data cleaning and processing: Remove noise and use filtering algorithms (such as median filtering, Gaussian filtering) to remove outliers and noise; Data downsampling (including uniform sampling, random sampling, etc.) to reduce the quantity of point cloud data and improve the speed of subsequent processing; Semantic annotation, assigning semantic labels to each point.
[0018] Data format: Define the source point cloud data as src.npy, representing the point cloud data to be registered; Define the target point cloud data as ref.npy, representing the point cloud data that the source point cloud wants to be registered to the target coordinate system; Define the true transformation matrix as gt.npy, representing the rotation angle and translation position that the source point cloud needs to be registered to the target point cloud.
[0019] Step S2: Construct the backbone feature extraction network of the depth network model, and extract multi-level features of the source point cloud and the target point cloud respectively. The backbone feature extraction network includes a kernel point convolution layer and a feature fusion module.
[0020] Specifically, S2 includes: First, read the input source point cloud and target point cloud data in matrix form. For example, if there are n point clouds, the size of the point cloud matrix is (n, 3), which respectively record the x, y, and z coordinate values of the point cloud data in its own coordinate system.
[0021] For the input point cloud matrix, perform feature extraction through a feature extraction network. The feature extraction network is constructed by multiple kernel point convolution layers. The convolution principle of KPConv is mainly defined based on kernel points and related functions. The convolution formula of KPConv can be expressed as: ; Among them, is the input feature set, is the convolution kernel parameter, is the convolution center point, is the point set within the domain of is a certain point within the domain, is the corresponding feature vector, which represents the weighted sum of all points within the neighborhood of the convolution center point, and the weights are determined by the convolution kernel parameters, thereby obtaining a new output feature.
[0022] After each convolution, the output features are abstracted to another dimensional space through a fully connected layer neural network, enabling these features to be better linearly separated.
[0023] The feature fusion module performs feature fusion on the point cloud features extracted at different depths. By upsampling the deep features and fusing them with the shallow features in a concatenated manner, the shallow features are made to have rich global information; Output the multi-level features after fusion; Step S3: Based on the extracted multi-level features, perform coarse matching of feature points through a global and local attention mechanism, and combine with point-level fine matching in the point cloud intersection area to generate point pair correspondence relationships.
[0024] Specifically, it includes: Assign the point cloud points under the point cloud to the nearest feature points extracted by the backbone network to obtain a subset of point cloud points, representing the point cloud region features represented by the feature points.
[0025] Calculate the similarity scores between feature points by querying the weight matrix, key weight matrix, and value weight matrix.
[0026] Based on the similarity scores, perform weighted fusion on the feature vectors to generate dependencies within the point cloud and across point clouds.
[0027] Generate a matching score matrix through the dot product operation of the normalized feature matrix, and screen the feature point pairs corresponding to the maximum matching scores.
[0028] Based on the calculation of feature vector similarity, use the nearest neighbor matching method to determine the specific point pairs in the point cloud intersection area.
[0029] In this embodiment, the above steps are refined as follows: Assign the point cloud points under the point cloud to the nearest feature points extracted by the backbone network to obtain a subset of point cloud points, representing the point cloud region features represented by the feature points; Introduce a global and local attention mechanism to learn the relationships within the point cloud and between point clouds, and obtain the corresponding relationships between feature points. The specific steps are as follows: Generate Q, K, and V matrices, corresponding to the query weight, key weight, and value weight respectively, enabling the model to capture the dependencies between any two elements within the input sequence; Calculate the similarity between Q and K to obtain the similarity score; Use the obtained similarity scores to weight V, so as to weight each vector in the feature matrix; Match the extracted key feature points through the key point matching module to obtain the corresponding relationship of key points. The specific steps are as follows: Perform a dot product on the two normalized feature matrices x and y to obtain the matrix xy; Calculate the matching score matrix between feature points in the two feature matrices. The formula is as follows: Extract the target feature score matrix ref and the source feature score matrix src from the matching score matrix by dimension; Perform a dot product on the target feature score matrix and the source feature score matrix to obtain the final matching score matrix; Traverse the matrix to obtain the largest n matching scores and the index coordinates of their corresponding points; Match the point clouds of the regions represented by the two feature points through the point matching module to obtain the corresponding matching relationship between the point clouds. The specific steps are as follows: Filter out the target feature key points and source feature key points in the regions with overlapping parts through the matching scores; Obtain the point clouds corresponding to the matched target feature points and source feature points; Calculate the similarity between the feature vectors extracted from the point clouds through the distance calculation formula; Find the point most similar to the target point cloud through the nearest neighbor matching method; Step S4: Based on the corresponding relationship of the point pairs, calculate the rigid body transformation matrix from the source point cloud to the target point cloud through iterative optimization.
[0030] Specifically include: Step S41: Perform geometric verification. By randomly selecting 3 non-collinear matching point pairs, calculate an initial rigid body transformation matrix; Step S42: Use this transformation matrix to transform all points within the key point domain of the source point cloud to the space of the target point cloud, and calculate the distance between the transformed points and the nearest points in the target point cloud. If the distance is less than the set threshold, these point pairs are considered inliers; Step S43: Repeat Step S41 and Step S42, and select the transformation matrix with the largest number of inliers as the final result to obtain the final conversion relationship between the two coordinate systems.
[0031] Step S5: Use the rigid body transformation matrix to transform the source point cloud to the target point cloud coordinate system, complete the registration and generate a complete point cloud.
[0032] Specifically include: By multiplying the point cloud data in the source point cloud data by the calculated rigid body transformation matrix, the coordinates of each point are transformed into the target point cloud coordinate system; Refer to Figure 4 , and merge the transformed source point cloud data with the target point cloud data to obtain the finally registered point cloud data.
[0033] On the other hand, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above method is implemented.
[0034] Those skilled in the art of this technology can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined.
[0035] For the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0036] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0037] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A design method for a 3D point cloud registration model with global and local attention, characterized in that Including: Step S1: Collect discrete sampling data of the scene surface from different perspectives to form a source point cloud and a target point cloud, where the point cloud data is a set of three-dimensional coordinate points; Step S2: Construct a backbone feature extraction network of a deep network model, and extract multi-level features of the source point cloud and the target point cloud respectively. The backbone feature extraction network includes a kernel point convolution layer and a feature fusion module; Step S3: Based on the extracted multi-level features, perform coarse matching of feature points through a global and local attention mechanism, and combine point-level fine matching in the point cloud intersection area to generate point pair correspondence; Step S4: Based on the point pair correspondence, calculate the rigid body transformation matrix from the source point cloud to the target point cloud through iterative optimization; Step S5: Use the rigid body transformation matrix to transform the source point cloud into the target point cloud coordinate system, complete registration and generate a complete point cloud.
2. The method according to claim 1, wherein The step S1 includes: Collect multi-perspective point cloud data of the scene through a three-dimensional sensing device, and the device includes a lidar, a depth camera or a three-dimensional scanner; Perform noise reduction, downsampling and semantic annotation processing on the original point cloud data; Define the storage formats of the source point cloud data, the target point cloud data and the corresponding true transformation matrix.
3. The method according to claim 1, wherein The calculation of the kernel point convolution layer satisfies the following formula: ; Among them, is the input feature set, is the convolutional kernel parameter, is the convolutional center point, is a point set within the domain of, is a certain point within the domain, is the corresponding feature vector, which represents the weighted sum of all points within the neighborhood of the convolutional center point, and the weights are determined by the convolutional kernel parameter, thereby obtaining a new output feature.
4. The method according to claim 1, wherein In step S3, the global and local attention mechanism includes: Calculate the similarity score between feature points through a query weight matrix, a key weight matrix and a value weight matrix; Based on the similarity score, perform weighted fusion on the feature vectors to generate dependencies within the point cloud and across point clouds; Generate a matching score matrix through the dot product operation of the normalized feature matrix, and screen the feature point pairs corresponding to the maximum matching score.
5. The method according to claim 1, wherein Step S4 includes: Randomly select three non-collinear pairs of matching points to calculate the initial rigid body transformation matrix; Filter inliers through a distance threshold, and iteratively optimize to obtain the optimal transformation matrix corresponding to the maximum inlier set.
6. The method according to claim 1, characterized in that The feature fusion module realizes the fusion of global information and local information by upsampling deep features and splicing them with shallow features.
7. The method according to claim 1, wherein Point-level fine matching includes: Based on the calculation of feature vector similarity, use the nearest neighbor matching method to determine the specific point pairs in the point cloud intersection area.
8. The method according to claim 7, characterized in that The complete point cloud is generated by merging the coordinate spaces of the transformed source point cloud and the target point cloud.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method described in any one of claims 1-8.
Citation Information
Cited By
Arch dam point cloud registration method based on deep learning and multi-target feature input
CN120431139A