A multi-modal three-dimensional point cloud registration method, system, terminal and storage medium

Through the multimodal three-dimensional point cloud registration method, feature extraction, fusion and key point recognition technology are used to solve the problem of accuracy deviation of point cloud data at different perspectives, and high-precision point cloud alignment is achieved.

CN119863498BActive Publication Date: 2025-07-22GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510354703.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-22
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

In the prior art, there are large accuracy deviations in three-dimensional point cloud data when rotating and translating estimating at different perspectives, and accurate alignment cannot be achieved.

Method used

The multimodal three-dimensional point cloud registration method is adopted to achieve the alignment of point cloud data through point cloud feature extraction, feature enhancement and fusion, significance score calculation, key point recognition and rigid transformation parameter calculation.

Benefits of technology

It significantly improves the accuracy and robustness of point cloud registration, and can quickly achieve accurate alignment of point cloud data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119863498B_ABST
    Figure CN119863498B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal three-dimensional point cloud registration method, system, terminal and storage medium. The method includes: obtaining initial point cloud data, and performing point cloud feature extraction processing and point cloud image feature extraction processing to obtain point cloud features and point cloud image features; performing feature enhancement processing and feature fusion processing on the point cloud features and point cloud image features to obtain hybrid features; performing feature stitching processing and significance score calculation according to the hybrid features to obtain a significance score, and performing feature selection processing according to the significance score to obtain key points and key point features; performing matching matrix construction processing and related weight calculation according to the key points and key point features to obtain target rigid transformation parameters, and aligning the initial point cloud data according to the target rigid transformation parameters. The present invention can quickly realize the accurate alignment of point cloud data by calculating the optimal rigid transformation parameters and registering the point cloud data according to the optimal rigid transformation parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of three-dimensional point cloud, and particularly to a multi-modal three-dimensional point cloud registration method, system, terminal and computer-readable storage medium. Background Art

[0002] Three-dimensional point cloud registration is a key technology in the field of computer vision. Its core goal is to accurately align point cloud data from different positions and poses, so as to realize data display, comparison and fusion in a unified coordinate system. With the rapid development of high-precision sensor technology, point cloud has gradually become the mainstream form of three-dimensional data representation, and point cloud registration technology has been widely applied in fields such as autonomous driving, augmented reality, and three-dimensional printing.

[0003] However, point cloud data naturally has characteristics such as partiality and sparsity, which lead to large accuracy deviations in the estimation of rotation and translation of point clouds from different perspectives, and it is impossible to achieve precise alignment of point clouds.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of the present invention is to provide a multi-modal three-dimensional point cloud registration method, system, terminal and computer-readable storage medium, aiming to solve the problem that point cloud data naturally has characteristics such as partiality and sparsity, which lead to large accuracy deviations in the estimation of rotation and translation of point clouds from different perspectives, and it is impossible to achieve precise alignment of point clouds.

[0006] To achieve the above purpose, the present invention provides a multi-modal three-dimensional point cloud registration method, and the multi-modal three-dimensional point cloud registration method includes the following steps:

[0007] Obtain initial point cloud data, and perform point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features;

[0008] Perform feature enhancement processing and feature fusion processing on the point cloud features and the point cloud image features to obtain mixed features;

[0009] Perform feature stitching processing and significance score calculation according to the mixed features to obtain a significance score, and perform feature selection processing according to the significance score to obtain key points and key point features;

[0010] Perform matching matrix construction processing and relevant weight calculation according to the key points and the key point features to obtain target rigid transformation parameters, and align the initial point cloud data according to the target rigid transformation parameters.

[0011] Optionally, in the multi-modal three-dimensional point cloud registration method, the initial point cloud data includes source point cloud data and target point cloud data; the point cloud features include source point cloud features and target point cloud features;

[0012] The obtaining of the initial point cloud data and the performing of point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features specifically includes:

[0013] Obtain the source point cloud data and the target point cloud data, perform point cloud feature extraction processing on the source point cloud data and the target point cloud data to obtain the source point cloud features and the target point cloud features;

[0014] Perform point cloud image feature extraction processing on the source point cloud data and the target point cloud data to obtain point cloud image features.

[0015] Optionally, in the multi-modal three-dimensional point cloud registration method, the performing of point cloud image feature extraction processing on the source point cloud data and the target point cloud data to obtain point cloud image features specifically includes:

[0016] Perform multi-view projection processing on the source point cloud data and the target point cloud data to obtain a plurality of two-dimensional image sequences;

[0017] Use a CNN network to perform feature extraction on the plurality of two-dimensional image sequences to obtain a plurality of projection features;

[0018] Use an aggregation function to perform fusion processing on the plurality of projection features to obtain the point cloud image features.

[0019] Optionally, in the multi-modal three-dimensional point cloud registration method, after the obtaining of the initial point cloud data and the performing of point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features, it further includes:

[0020] Obtain the overlapping point features and non-overlapping point features between the source point cloud data and the target point cloud data, construct a feature pair set according to the overlapping point features and the non-overlapping point features, and perform overlapping contrast learning on the source point cloud data and the target point cloud data according to the feature pair set;

[0021] Perform feature space projection processing and similarity maximization calculation on the source point cloud data and the target point cloud data to obtain a similarity result, and perform multi-modal contrast learning on the source point cloud data and the target point cloud data according to the similarity result.

[0022] Optionally, in the multimodal three-dimensional point cloud registration method, the hybrid features include source point cloud hybrid features and target point cloud hybrid features;

[0023] Performing feature enhancement processing and feature fusion processing on the point cloud features and the point cloud image features to obtain hybrid features specifically includes:

[0024] Using an attention mechanism to perform feature interaction learning processing on the source point cloud features and the target point cloud features to obtain source point cloud enhanced features and target point cloud enhanced features;

[0025] Inputting the source point cloud enhanced features and the point cloud image features into a multi-layer perceptron to obtain a first output dimension, and performing attention weight calculation and feature fusion processing according to the first output dimension to obtain source point cloud hybrid features;

[0026] Inputting the target point cloud enhanced features and the point cloud image features into the multi-layer perceptron to obtain a second output dimension, and performing attention weight calculation and feature fusion processing according to the second output dimension to obtain target point cloud hybrid features.

[0027] Optionally, in the multimodal three-dimensional point cloud registration method, performing feature splicing processing and significance score calculation according to the hybrid features to obtain a significance score, and performing feature selection processing according to the significance score to obtain key points and key point features specifically includes:

[0028] Performing max pooling processing on the source point cloud hybrid features and the source point cloud data to obtain a global feature vector;

[0029] Performing splicing processing on the global feature vector and the target point cloud hybrid features to obtain a spliced feature;

[0030] Calculating the significance score corresponding to each feature in the spliced feature through a one-dimensional convolutional network, and selecting in descending order of the significance score to obtain a preset number of key points and key point features.

[0031] Optionally, in the multimodal three-dimensional point cloud registration method, performing matching matrix construction processing and related weight calculation according to the key points and the key point features to obtain target rigid transformation parameters, and aligning the initial point cloud data according to the target rigid transformation parameters specifically includes:

[0032] Constructing a spatial coordinate combination and a hybrid feature combination according to the key points and the key point features, and performing two-way parallel convolution processing on the spatial coordinate combination and the hybrid feature combination to obtain a spatial coordinate matching matrix and a feature matching matrix;

[0033] If the results of the spatial coordinate matching matrix and the feature matching matrix are inconsistent, the spatial coordinate matching matrix and the feature matching matrix are subjected to splicing processing, aggregation processing, and convolution processing to obtain a matching score;

[0034] Calculate the relevant weights corresponding to the matching score, and use the weighted singular value decomposition method to solve the relevant weights to obtain the target rigid transformation parameters;

[0035] Align the source point cloud data and the target point cloud data according to the target rigid transformation parameters.

[0036] In addition, to achieve the above object, the present invention also provides a multi-modal three-dimensional point cloud registration system, wherein the multi-modal three-dimensional point cloud registration system includes:

[0037] A point cloud feature extraction module, configured to obtain initial point cloud data, and perform point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features;

[0038] A hybrid feature generation module, configured to perform feature enhancement processing and feature fusion processing on the point cloud features and the point cloud image features to obtain hybrid features;

[0039] A feature selection processing module, configured to perform feature splicing processing and significance score calculation according to the hybrid features to obtain a significance score, and perform feature selection processing according to the significance score to obtain key points and key point features;

[0040] A target rigid transformation parameter generation module, configured to perform matching matrix construction processing and relevant weight calculation according to the key points and the key point features to obtain target rigid transformation parameters, and align the initial point cloud data according to the target rigid transformation parameters.

[0041] In the present invention, initial point cloud data is acquired, and point cloud feature extraction processing and point cloud image feature extraction processing are performed on the initial point cloud data to obtain point cloud features and point cloud image features; feature enhancement processing and feature fusion processing are performed on the point cloud features and the point cloud image features to obtain hybrid features; feature stitching processing and significance score calculation are performed according to the hybrid features to obtain a significance score, and feature selection processing is performed according to the significance score to obtain key points and key point features; a matching matrix construction process and related weight calculation are performed according to the key points and the key point features to obtain target rigid transformation parameters, and the initial point cloud data is aligned according to the target rigid transformation parameters. By performing feature extraction on the initial point cloud data, the present invention obtains point cloud features and point cloud image features, constructs hybrid features based on the point cloud features and the point cloud image features, then determines key points and key point features according to the hybrid features, calculates the optimal rigid transformation parameters of the point cloud data according to the key points and the key point features, and finally can align the point cloud data according to the optimal rigid transformation parameters, significantly improving the accuracy and robustness of point cloud registration. Description of the Drawings

[0042] Figure 1 is a flowchart of a preferred embodiment of the multi-modal three-dimensional point cloud registration method of the present invention;

[0043] Figure 2 is a schematic diagram of the overall structure implementation process of a preferred embodiment of the multi-modal three-dimensional point cloud registration method of the present invention;

[0044] Figure 3 is a schematic diagram of the structure of the correspondence search module of a preferred embodiment of the multi-modal three-dimensional point cloud registration method of the present invention;

[0045] Figure 4 is a schematic diagram of the overall architecture of a preferred embodiment of the multi-modal three-dimensional point cloud registration method of the present invention;

[0046] Figure 5 is a structural diagram of a preferred embodiment of the multi-modal three-dimensional point cloud registration system of the present invention;

[0047] Figure 6 is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Embodiments

[0048] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] 3D point cloud registration is a key technology in the field of computer vision. Its core goal is to accurately align point cloud data from different positions and postures, so as to achieve data display, comparison and fusion in a unified coordinate system. With the rapid development of high-precision sensor technology, point cloud has gradually become the mainstream form of 3D data representation, and point cloud registration technology has been widely used in fields such as autonomous driving, augmented reality and 3D printing. However, point cloud data naturally has characteristics such as partiality and sparsity, which often leads to large accuracy deviations when estimating the rotation and translation of point clouds from different perspectives. Therefore, how to quickly and accurately achieve the registration of partially overlapping point clouds has always been an important technical problem that needs to be solved in this field.

[0050] With the breakthrough progress of deep learning in the field of computer vision, the three-dimensional point cloud registration method based on neural network has gradually become a research hotspot. The processing flow of deep point cloud registration mainly consists of three core links: feature extraction, matching search and registration optimization. Specifically, the neural network first extracts high-dimensional feature representations from the input source point cloud and target point cloud. These features need to have rotation invariance and local description capabilities; then the correspondence between point clouds is established based on the extracted features, and feature matching is usually achieved by using methods such as nearest neighbor search or attention mechanism; finally, the optimal rigid transformation parameters are obtained by optimizing the preset loss function (such as point-to-point distance, feature similarity, etc.), thereby completing the precise alignment of point clouds. Compared with traditional registration algorithms, deep learning methods can more comprehensively capture the geometric shape, spatial structure and semantic features of point clouds with their powerful feature learning capabilities, and have stronger environmental adaptability and robustness. This end-to-end learning framework not only simplifies the registration process, but also continuously improves model performance in a data-driven manner, providing a more reliable solution for practical applications.

[0051] To solve the above problems, the present invention proposes a multimodal 3D point cloud registration method, which achieves higher precision results by perceiving the global shape. At the same time, two contrastive learning strategies are introduced, where overlapping contrastive learning is used to emphasize overlapping point features, and cross-modal contrastive learning is used to construct the correspondence between 2D and 3D. In addition, a method for predicting point cloud masks is designed to extract key points and reduce computing resource consumption.

[0052] The multimodal three-dimensional point cloud registration method described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the multimodal three-dimensional point cloud registration method includes the following steps:

[0053] Step S10: Obtain the initial point cloud data, and perform point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features. The initial point cloud data includes source point cloud data and target point cloud data; the point cloud features include source point cloud features and target point cloud features.

[0054] As Figure 2 shown, the present invention proposes a three-dimensional point cloud registration method based on multi-modal contrast learning, which mainly consists of four parts, namely feature extraction, multiple contrast learning, attention fusion and mask prediction, and correspondence search.

[0055] Specifically, obtain the source point cloud data and the target point cloud data, and perform point cloud feature extraction processing on the source point cloud data and the target point cloud data to obtain the source point cloud features and the target point cloud features.

[0056] The specific process of feature extraction is as follows: The input is two sets of point clouds, including source point cloud data (where is the first source point cloud, is the th source point cloud, is the th source point cloud) and target point cloud data (where is the first target point cloud, is the th target point cloud, is the th target point cloud). The goal of the present invention is to calculate the optimal rigid transformation parameters, and the rigid transformation parameters include the rotation matrix and the translation vector . According to the optimal rigid transformation parameters, align the two sets of point clouds (the alignment method is: by finding the optimal rotation matrix and the translation vector , align the source point cloud to the target point cloud through rigid body transformation). Among them, is the three-dimensional special orthogonal group (Special Orthogonal Group), which describes the rotation transformation in three-dimensional space, is the three-dimensional real number space. It should be particularly noted that the present invention does not require a strict one-to-one correspondence between the source point cloud and the target point cloud, so it can handle the case where the number of points is different ( , where is the total number of source point clouds, is the total number of target point clouds).

[0057] The feature extraction process of the present invention is divided into two parallel branches of point cloud and image. In the point cloud branch, the point cloud And Regarding each point of as the vertex of the graph, point-wise features are extracted through the EdgeConv operation (EdgeConv extracts point-wise features by dynamically constructing a local graph and aggregating edge features. The process includes: 1. Constructing a dynamic K-nearest neighbor graph; 2. Calculating edge features; 3. Aggregating edge features), and the KNN algorithm is used to dynamically construct the graph structure at each layer to expand the receptive field, and finally the point cloud features are obtained And . Among them, the process of constructing the graph structure is as follows: for each point, find the K nearest neighbor points in terms of Euclidean distance to form a K-nearest neighbor graph. Usually, by calculating the distance matrix between points, and then selecting the K points corresponding to the smallest K distances in each row as neighbor points

[0058] Perform multi-view projection processing on the source point cloud data and the target point cloud data to obtain multiple two-dimensional image sequences; use a CNN network (Convolutional Neural Network) to extract features from the multiple two-dimensional image sequences to obtain multiple projection features; use an aggregation function to fuse the multiple projection features to obtain the point cloud image features

[0059] In the image branch, by projecting the three-dimensional point cloud onto different views, corresponding two-dimensional image sequences are generated, and then image features are extracted , the image features The expression of is: ; where is the replication operation, is the aggregation function, represents the projected picture, is one of the different views

[0060] Among them, the extraction process of the image features is as follows: by projecting the three-dimensional point cloud (including the source point cloud and the target point cloud) from V different views, multiple two-dimensional images are obtained, then a CNN network is used to extract features from these projected images, and finally the features of all views are fused through an aggregation function to obtain the final image features . Essentially, it is to convert 3D information into 2D features through multi-view projection + CNN and then integrate them

[0061] Further, obtain the overlapping point features and non-overlapping point features between the source point cloud data and the target point cloud data, construct a feature pair set according to the overlapping point features and the non-overlapping point features, and perform overlapping contrast learning on the source point cloud data and the target point cloud data according to the feature pair set.

[0062] Multi-modal contrast learning consists of two parts, overlapping contrast learning and multi-modal contrast learning. Among them, overlapping contrast learning aims to enhance the feature expression of the overlapping region while suppressing the interference of the non-overlapping region. The specific implementation process is as follows: First, transform the source point cloud through the ground truth rigid transformation (referring to the rotation matrix and the translation vector ), and then determine the overlapping points based on the minimum distance threshold with the target point cloud (where the minimum distance is calculated as the Euclidean space distance of the three-dimensional coordinates. When the calculated Euclidean space distance is less than the preset threshold, the point cloud position at this time is determined as the overlapping point); Secondly, extract the overlapping point features and non-overlapping point features of the two groups of point clouds respectively through the overlapping selection mechanism (the extraction process is: determine the indexes of the overlapping points and non-overlapping points respectively through the threshold, and then make a selection); Then, construct a feature pair set (that is, the subsequent positive pair set + negative pair set 1 + negative pair set 2, used for overlapping contrast learning): Define the overlapping point feature pairs of the two groups of point clouds as the positive pair set, and the overlapping point features of the point cloud and the non-overlapping point features of the point cloud constitute the negative pair set 1, and the overlapping point features of the point cloud and the non-overlapping point features of the point cloud constitute the negative pair set 2.

[0063] Perform feature space projection processing and similarity maximization calculation on the source point cloud data and the target point cloud data to obtain a similarity result, and perform multi-modal contrast learning on the source point cloud data and the target point cloud data according to the similarity result.

[0064] It can be understood that the cross-modal contrast learning module is responsible for constructing the corresponding mapping relationship between the two-dimensional and three-dimensional features. The process includes: First, project the point cloud features ( and )and the image feature into a unified feature space ( represents the dimension of the point feature) respectively through the max pooling operation to obtain the projection vectors , and . Subsequently, calculate the mean value of and to obtain the unified projection representation of the point cloud modality. In the invariant space, the goal is to maximize and the similarity between them (for cross-modal contrast learning) because they both correspond to the same object.

[0065] Step S20: Perform feature enhancement processing and feature fusion processing on the point cloud features and the point cloud image features to obtain hybrid features. Among them, the hybrid features include source point cloud hybrid features and target point cloud hybrid features.

[0066] Specifically, use the attention mechanism to perform feature interaction learning processing on the source point cloud features and the target point cloud features to obtain source point cloud enhanced features and target point cloud enhanced features; input the source point cloud enhanced features and the point cloud image features into a multi-layer perceptron to obtain a first output dimension, and perform attention weight calculation and feature fusion processing according to the first output dimension to obtain source point cloud hybrid features; input the target point cloud enhanced features and the point cloud image features into the multi-layer perceptron to obtain a second output dimension, and perform attention weight calculation and feature fusion processing according to the second output dimension to obtain target point cloud hybrid features.

[0067] The present invention uses a double-layer attention fusion structure to enhance the extraction of context information of features. Among them, the first layer of attention mechanism focuses on the information interaction between and to generate enhanced point cloud features through feature interaction and , and this process automatically highlights the key region information in the point cloud.

[0068] Among them, the specific implementation process of generating enhanced point cloud features and is as follows: Interact the original point cloud features and through the attention mechanism to achieve information exchange and complementarity between features, thereby generating enhanced features and . Generally speaking, it is to let the two sets of features "pay attention" to the important information of each other, learn from and complement each other, and finally obtain a richer feature representation.

[0069] The second layer of attention mechanism aims to enhance the discriminability of point features by integrating global shape and texture information. Taking the source point cloud as an example, first input the point cloud features and the image features into the MLP (Multi-Layer Perceptron) for processing, and use the output of as the query array , where is the output dimension of the MLP. The output of serves as key-value pairs (including , K and V , both of which belong to key-value pairs). By calculating the attention weights to capture global shape and texture information (where is the matrix transpose), the source point cloud hybrid feature that fuses multi-modal information is finally generated: ; Similarly, the target point cloud hybrid feature can be obtained in the same way.

[0070] Step S30: Perform feature stitching processing and significance score calculation based on the hybrid feature to obtain a significance score, and perform feature selection processing based on the significance score to obtain key points and key point features.

[0071] Specifically, perform max pooling processing on the source point cloud hybrid feature and the source point cloud data to obtain a global feature vector; perform stitching processing on the global feature vector and the target point cloud hybrid feature to obtain a stitched feature; calculate the significance score corresponding to each feature in the stitched feature through a one-dimensional convolutional network, and select in descending order of the significance score to obtain a preset number of key points and key point features.

[0072] In the present invention, the mask prediction module is set to optimize the feature expression by screening discriminative features and suppressing non-discriminative features. The specific process of mask prediction is as follows: First, perform max pooling operations on the source point cloud hybrid feature and the source point cloud data , copy the globally pooled feature vector and stitch it with the target point cloud hybrid feature ; Subsequently, calculate the significance score for each feature through a one-dimensional convolutional network. The higher the score, the stronger the discriminability of the feature, which is more conducive to subsequent matching; finally, select the top-K points according to the significance score, set their mask values to 1, and the rest to 0, and accordingly screen out the key point coordinates and key point features of K key points, where the key point coordinates and key point features are used to guide the correspondence search.

[0073] Step S40: Perform matching matrix construction processing and related weight calculation based on the key points and the key point features to obtain target rigid transformation parameters, and align the initial point cloud data according to the target rigid transformation parameters.

[0074] The present invention provides a corresponding relationship search module. This module adopts a cascaded framework of coarse registration and fine registration, and independently searches for corresponding points using hybrid features and spatial coordinates respectively, as Figure 3 shown. By reasonably integrating the information in the feature space and the geometric space at different registration stages, the registration accuracy is effectively improved. This cascaded registration strategy can not only make full use of information at different scales, but also gradually refine the registration result to avoid local optimal solutions.

[0075] Specifically, a spatial coordinate combination and a hybrid feature combination are constructed according to the key points and the key point features, and two-way parallel convolution processing is performed on the spatial coordinate combination and the hybrid feature combination to obtain a spatial coordinate matching matrix and a feature matching matrix; if the results of the spatial coordinate matching matrix and the feature matching matrix are inconsistent, then splicing processing, aggregation processing and convolution processing are performed on the spatial coordinate matching matrix and the feature matching matrix to obtain a matching score; calculate the relevant weights corresponding to the matching score, and use the weighted singular value decomposition method to solve the relevant weights to obtain target rigid transformation parameters; align the source point cloud data and the target point cloud data according to the target rigid transformation parameters.

[0076] According to the coordinates of the selected key points and their corresponding key point features , a spatial coordinate combination and a hybrid feature combination can be constructed. Among them, the spatial coordinate combination contains the geometric position relationship between point pairs (i.e., between the source point cloud and the target point cloud), while the hybrid feature combination encodes the high-level semantic information between point pairs. These two combinations describe the similarity between point pairs from different perspectives and provide complementary information for subsequent matching.

[0077] As Figure 3 shown ( Figure 3 in which K is the number of key points, represents the dimension of the point feature, M is the total number of the target point cloud, is the initial rotation matrix, is the initial translation vector, is the rotation matrix after the nth iteration, is the translation vector after the nth iteration. The present invention processes the spatial coordinate combination and the hybrid feature combination through two-way parallel convolution operations respectively, compresses the high-dimensional features into a one-dimensional matching matrix, including a spatial coordinate matching matrix and a feature matching matrix . Among them, the spatial coordinate matching matrix Reflect the geometric similarity between point pairs, feature matching matrix Reflect the semantic similarity between point pairs. The larger the matrix element value, the higher the matching possibility of the corresponding point pair. This dual-channel feature extraction strategy can capture both the local geometric structure and the global semantic information of the point cloud.

[0078] The generation of the final matching matrix is based on the spatial coordinate matching matrix and the feature matching matrix fusion. When the results of the two matching matrices (i.e., the spatial coordinate matching matrix and the feature matching matrix ) are consistent, it can be regarded as a high-confidence match; when there is inconsistency, it is necessary to evaluate the degree of conflict to determine the matching credibility. Specifically, after splicing the two matching matrices, use the convolution operation to map them to a high-dimensional feature space, and then obtain the matching score of the point through the maximum aggregation and convolution operations. Therefore, the weight of the th point pair is defined as: ; where, is the weight definition of the th point pair, represents the indicator function, represents the median of the matching scores of all points, that is, the middle value is taken after sorting all the scores. is the matching score of the th point, which is calculated by fusing the spatial coordinate matching matrix and the feature matching matrix . After obtaining the relevant weights, use the weighted singular value decomposition (SVD) to solve the final rigid transformation: ; where, is the rotation matrix; is the translation vector; is used to calculate the minimum and values; is the coordinate of the th point in the source point cloud; is the corresponding point coordinate found in the target point cloud according to ; The double vertical bars represent the solution of the Euclidean distance.

[0079] In summary, the present invention introduces the point cloud projection image as an auxiliary information source, and designs a dual contrast learning strategy: the overlapping contrast learning focuses on the feature extraction of the overlapping area, and the cross-modal contrast learning establishes the mapping relationship between the two-dimensional and three-dimensional features. At the same time, a mask prediction module is innovatively introduced for key point recognition, and finally the rigid transformation is solved through the correspondence search. The overall architecture diagram of the present invention is as shown in Figure 4as shown in Figure 4 in is the source point cloud feature, is the target point cloud feature, is the image feature, is the source point cloud mixed feature, is the target point cloud mixed feature, and SVD is the weighted singular value decomposition).

[0080] Advantages of the present invention:

[0081] The present invention proposes a three-dimensional point cloud registration method based on multi-modal contrast learning. By introducing dual strategies of overlapping contrast learning and cross-modal contrast learning to enhance feature expression, combining a double-layer attention mechanism to achieve efficient feature fusion, designing a mask prediction module to automatically identify key points, and finally adopting a cascaded registration framework to jointly optimize the feature space and geometric space information, the accuracy and robustness of point cloud registration are significantly improved, providing a more reliable solution for practical applications.

[0082] The three-dimensional point cloud registration method proposed by the present invention mainly includes four core innovation points: 1. A feature learning strategy that combines feature extraction with multiple contrast learning, where overlapping contrast learning enhances the discriminability of overlapping region features, and cross-modal contrast learning realizes the fusion of two-dimensional and three-dimensional features; 2. A feature fusion method based on a double-layer attention mechanism to enhance the expression of context information; 3. An innovative mask prediction module for key point recognition; 4. A cascaded correspondence search framework that combines feature space and geometric space information.

[0083] In addition, on the premise of keeping the overall framework unchanged, each module of the present invention can be replaced by other technical solutions: Feature extraction can use deep learning architectures such as Transformer, contrast learning can introduce other loss functions such as triplet loss, mask prediction can use traditional methods based on geometric features such as ISS (Intrinsic Shape Signatures), Harris3D (Harris 3D Feature Detector), etc., and correspondence search can also select traditional matching algorithms such as RAN, SAC, etc. as alternative solutions.

[0084] Furthermore, as Figure 5 shown, based on the above multi-modal three-dimensional point cloud registration method, the present invention also correspondingly provides a multi-modal three-dimensional point cloud registration system, wherein, the multi-modal three-dimensional point cloud registration system includes:

[0085] A point cloud feature extraction module 51, configured to obtain initial point cloud data, and perform point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features;

[0086] The hybrid feature generation module 52 is configured to perform feature enhancement processing and feature fusion processing on the point cloud feature and the point cloud image feature to obtain a hybrid feature;

[0087] The feature selection processing module 53 is configured to perform feature splicing processing and significance score calculation according to the hybrid feature to obtain a significance score, and perform feature selection processing according to the significance score to obtain key points and key point features;

[0088] The target rigid transformation parameter generation module 54 is configured to perform matching matrix construction processing and related weight calculation according to the key points and the key point features to obtain target rigid transformation parameters, and align the initial point cloud data according to the target rigid transformation parameters.

[0089] Further, as Figure 6 shown, based on the above multi-modal three-dimensional point cloud registration method and system, the present invention also correspondingly provides a terminal, and the terminal includes a processor 10, a memory 20, and a display 30. Figure 6 Only some components of the terminal are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.

[0090] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal. Further, the memory 20 may also include both the internal storage unit and the external storage device of the terminal. The memory 20 is used to store application software installed on the terminal and various types of data, such as program codes installed on the terminal. The memory 20 may also be used to temporarily store data that has been output or will be output. In one embodiment, a multi-modal three-dimensional point cloud registration program 40 is stored on the memory 20, and the multi-modal three-dimensional point cloud registration program 40 can be executed by the processor 10, so as to implement the multi-modal three-dimensional point cloud registration method in the present application.

[0091] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chips, and is used to run program codes stored in the memory 20 or process data, such as executing the multi-modal three-dimensional point cloud registration method, etc.

[0092] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0093] In one embodiment, when the processor 10 executes the multi-modal three-dimensional point cloud registration program 40 in the memory 20, the steps of the multi-modal three-dimensional point cloud registration method are implemented.

[0094] In summary, the present invention provides a multi-modal three-dimensional point cloud registration method, system and terminal. The method includes: obtaining initial point cloud data, and performing point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features; performing feature enhancement processing and feature fusion processing on the point cloud features and the point cloud image features to obtain mixed features; performing feature stitching processing and significance score calculation according to the mixed features to obtain a significance score, and performing feature selection processing according to the significance score to obtain key points and key point features; performing matching matrix construction processing and related weight calculation according to the key points and the key point features to obtain target rigid transformation parameters, and aligning the initial point cloud data according to the target rigid transformation parameters. By extracting features from the initial point cloud data, the present invention obtains point cloud features and point cloud image features, constructs mixed features according to the point cloud features and point cloud image features, then determines key points and key point features according to the mixed features, calculates the optimal rigid transformation parameters of the point cloud data according to the key points and key point features, and finally can align the point cloud data according to the optimal rigid transformation parameters, significantly improving the accuracy and robustness of point cloud registration.

[0095] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the element.

[0096] Of course, those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The program can be stored in a computer-readable storage medium readable by a computer. When the program is executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a memory, a magnetic disk, an optical disk, etc.

[0097] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description. All such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A multi-modal three-dimensional point cloud registration method, characterized in that The multi-modal three-dimensional point cloud registration method includes: Obtain initial point cloud data, and perform point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features; The initial point cloud data includes source point cloud data and target point cloud data; the point cloud features include source point cloud features and target point cloud features; The obtaining of the initial point cloud data, and performing point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features specifically includes: Obtain the source point cloud data and the target point cloud data, and perform point cloud feature extraction processing on the source point cloud data and the target point cloud data to obtain the source point cloud features and the target point cloud features; Perform point cloud image feature extraction processing on the source point cloud data and the target point cloud data to obtain point cloud image features; Adopt an attention mechanism to perform feature interaction learning processing on the source point cloud features and the target point cloud features to obtain source point cloud enhanced features and target point cloud enhanced features; Input the source point cloud enhanced features and the point cloud image features into a multi-layer perceptron to obtain a first output dimension, and perform attention weight calculation and feature fusion processing according to the first output dimension to obtain source point cloud mixed features; Input the target point cloud enhanced features and the point cloud image features into the multi-layer perceptron to obtain a second output dimension, and perform attention weight calculation and feature fusion processing according to the second output dimension to obtain target point cloud mixed features; Perform max pooling processing on the source point cloud mixed features and the source point cloud data to obtain a global feature vector; Perform splicing processing on the global feature vector and the target point cloud mixed features to obtain spliced features; Calculate the significance score corresponding to each feature in the spliced features through a one-dimensional convolutional network, and select them in descending order of the significance score to obtain a preset number of key points and key point features; Perform matching matrix construction processing and related weight calculation according to the key points and the key point features to obtain target rigid transformation parameters, and align the initial point cloud data according to the target rigid transformation parameters.

2. The multimodal three-dimensional point cloud registration method according to claim 1, wherein The performing of point cloud image feature extraction processing on the source point cloud data and the target point cloud data to obtain point cloud image features specifically includes: Perform multi-view projection processing on the source point cloud data and the target point cloud data to obtain a plurality of two-dimensional image sequences; Adopt a CNN network to perform feature extraction on the plurality of two-dimensional image sequences to obtain a plurality of projection features; Adopt an aggregation function to perform fusion processing on the plurality of projection features to obtain the point cloud image features.

3. The multimodal three-dimensional point cloud registration method according to claim 1, characterized in that After obtaining the initial point cloud data, and performing point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features, it further includes: Obtain the overlapping point features and non-overlapping point features between the source point cloud data and the target point cloud data, construct a feature pair set according to the overlapping point features and the non-overlapping point features, and perform overlapping contrast learning on the source point cloud data and the target point cloud data according to the feature pair set; Perform feature space projection processing and similarity maximization calculation on the source point cloud data and the target point cloud data to obtain a similarity result, and perform multi-modal contrast learning on the source point cloud data and the target point cloud data according to the similarity result.

4. The multimodal three-dimensional point cloud registration method according to claim 1, wherein The process of constructing a matching matrix and calculating relevant weights according to the key points and the key point features to obtain target rigid transformation parameters, and aligning the initial point cloud data according to the target rigid transformation parameters specifically includes: Construct a spatial coordinate combination and a mixed feature combination according to the key points and the key point features, and perform two-way parallel convolution processing on the spatial coordinate combination and the mixed feature combination to obtain a spatial coordinate matching matrix and a feature matching matrix; If the results of the spatial coordinate matching matrix and the feature matching matrix are inconsistent, perform splicing processing, aggregation processing, and convolution processing on the spatial coordinate matching matrix and the feature matching matrix to obtain a matching score; Calculate the relevant weights corresponding to the matching score, and use the weight singular value decomposition method to solve the relevant weights to obtain target rigid transformation parameters; Align the source point cloud data and the target point cloud data according to the target rigid transformation parameters.

5. A multi-modal three-dimensional point cloud registration system that implements the multi-modal three-dimensional point cloud registration method according to any one of claims 1-4, characterized in that, The multi-modal three-dimensional point cloud registration system includes: A point cloud feature extraction module, configured to obtain initial point cloud data, and perform point cloud feature extraction processing and point cloud image feature extraction processing on the initial point cloud data to obtain point cloud features and point cloud image features; A mixed feature generation module, configured to perform feature enhancement processing and feature fusion processing on the point cloud features and the point cloud image features to obtain mixed features; A feature selection processing module, configured to perform feature splicing processing and significance score calculation according to the mixed features to obtain a significance score, and perform feature selection processing according to the significance score to obtain key points and key point features; A target rigid transformation parameter generation module, configured to perform matching matrix construction processing and relevant weight calculation according to the key points and the key point features to obtain target rigid transformation parameters, and align the initial point cloud data according to the target rigid transformation parameters.

6. A terminal, characterized in that, The terminal includes: a memory, a processor, and a multi-modal three-dimensional point cloud registration program stored on the memory and executable on the processor. When the multi-modal three-dimensional point cloud registration program is executed by the processor, the steps of the multi-modal three-dimensional point cloud registration method according to any one of claims 1-4 are implemented.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a multi-modal three-dimensional point cloud registration program. When the multi-modal three-dimensional point cloud registration program is executed by a processor, the steps of the multi-modal three-dimensional point cloud registration method according to any one of claims 1-4 are implemented.

Citation Information

Patent Citations

  • Multi-modal point cloud registration method based on image and geometric information guidance

    CN117095033A

  • Semantic component attitude estimation method based on deep learning

    CN117218343A