Image point cloud registration method based on multi-modal uncertainty modeling and modal alignment
Through the method of multimodal uncertainty modeling and modal alignment, the importance of image patches is quantified, noise interference is reduced, and the attention to key areas is improved, and the problem of insufficient accuracy and success rate of image and point cloud registration in complex scenarios in the prior art is solved, achieving high-precision and robust registration effect.
Patent Information
- Application Number
- CN202510359605.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-04
AI Technical Summary
Existing image and point cloud registration methods are susceptible to excessive attention from noise areas in complex scenarios, ignoring key areas, resulting in a decrease in matching accuracy and success rate.
The multimodal uncertainty modeling and modal alignment method are used to quantify the importance of image patches through the uncertainty modeling mechanism, reduce noise interference, and increase the attention of key areas through the layered matching module and the adversarial modal alignment module, and deal with scale differences caused by perspective scaling in combination with dynamic adjustment.
It significantly improves the registration success rate in occlusion scenes, improves matching accuracy and robustness, reduces the computational complexity, and adapts to image and point cloud registration in complex scenes.
Smart Images

Figure CN120259384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to an image point cloud registration method based on multi-modal uncertainty modeling and modal alignment. Background Art
[0002] Image and point cloud registration aims to determine the rigid body transformation of the point cloud to the camera coordinate system, including the calculation of the rotation matrix and the translation matrix. The process includes cross-modal matching of the image and the point cloud, and the final pose transformation is completed through a pose estimator. This technology has important application value in tasks such as 3D reconstruction, SLAM (Simultaneous Localization and Mapping), and visual localization. However, the image is dense two-dimensional grid data, while the point cloud is sparse and irregular three-dimensional data. This difference in data representation poses a great challenge to the cross-modal matching of the image and the point cloud.
[0003] In the prior art, two main methods for solving image and point cloud registration have been proposed: one is the method of detection followed by matching. In this type of method, key points are first detected separately in the image and the point cloud, and matching is performed based on the semantic features of the key points. Since image key points depend on texture and color, while point cloud key points depend on geometric structure, this modal difference makes the detection of key points more difficult. In addition, image and point cloud descriptors encode different visual information, and it is difficult to extract consistent descriptors for effective matching; the other is the direct registration method without detection. This type of method does not rely on key point detection, but directly performs cross-modal matching on the image and the point cloud. For example, the 2D3D-MATR method, through a coarse-to-fine registration pipeline, first establishes patch-level matching of image and point cloud features, and then further refines it into dense pixel-to-point matching. This type of method significantly improves the inlier ratio of registration by introducing context information and multiple receptive fields.
[0004] Although the method without detection overcomes the deficiencies of detection followed by matching to a certain extent, the existing methods assign equal attention to all regions in the image, which may lead to excessive attention to noise regions while ignoring key regions, limiting the performance of image and point cloud registration technology in complex scenarios. Therefore, the present invention proposes an image point cloud registration method based on multi-modal uncertainty modeling and modal alignment. Summary of the Invention
[0005] The purpose of the present invention is to provide an image point cloud registration method based on multi-modal uncertainty modeling and modal alignment. Through the uncertainty modeling mechanism, the importance of image patches is quantified, enabling the network to pay more attention to key regions, reducing noise interference, and thus improving the accuracy and success rate of matching.
[0006] To achieve the above object, the present invention provides the following technical solutions: An image point cloud registration method based on multi-modal uncertainty modeling and modal alignment, which is applied to an indoor scene and includes the following steps:
[0007] Receive image data and point cloud data, and pair them into groups;
[0008] Use backbone networks for feature extraction of different modalities to extract image features and point cloud features from the image data and point cloud data, and use a multi-layer self-attention module and a cross-attention module to perform interactive processing on the image features and point cloud features;
[0009] Construct a hierarchical matching module based on uncertainty modeling to achieve hierarchical matching and interaction of image features and point cloud features, and use the uncertainty modeling method to assign a large variance to unaligned image patches during the matching process;
[0010] Construct an adversarial modal alignment module to reduce the difference in cross-modal distributions of image features and point cloud features based on adversarial learning, and achieve alignment of the domains of image features and point cloud features;
[0011] Construct a loss function to train the overall multi-modal feature extraction model until convergence, and obtain an overall multi-modal feature extraction model that meets the standards.
[0012] Further, the multi-modal feature extraction backbone network extracts image features by using a ResNet network and a feature pyramid network respectively for the received image data and point cloud data and extracts point cloud features by using a KPFCNN network;
[0013] The image features and point cloud features are respectively represented as and at low resolution for rough matching; and are respectively represented as and at high resolution for fine matching.
[0014] Further, construct a hierarchical matching module based on uncertainty modeling to achieve hierarchical matching and interaction of image features and point cloud features, and use the uncertainty modeling method to assign a large variance to unaligned image patches during the matching process, specifically as follows:
[0015] (31) Multi-scale feature extraction:
[0016] Divide the image data into h×w patches, and extract features of different scales through a lightweight three-stage CNN, denoted as [i1, f i1 , [i2, f i2 , [i3, f i3, the point cloud data is partitioned into point cloud patches through the nearest neighbor grid partitioning method, denoted as [p1, f p1 ;
[0017] (32) Uncertainty Modeling:
[0018] (32.1) In the uncertainty estimation layer, the features of image patches at different scales are reconstructed, and the input feature f ix , x ∈ {1, 2, 3}, is fed into three uncertainty estimation layers;
[0019] Model the single-image feature distribution as a Gaussian distribution, and obtain the feature f for the x-th layer through uncertainty modeling ixu , the feature f ixu is reconstructed by a Gaussian distribution parameterized by the mean vector μ (x) and the covariance matrix Σ (x) ;
[0020] (32.2) Define a term q representing entropy, where the variance is positively correlated with the entropy, and the entropy is defined as follows:
[0021]
[0022] Σ represents the variance values of different image patches. The larger the variance, the larger the entropy;
[0023] (32.3) Through calculation, the total uncertainty loss formula is obtained by combining the losses of all layers, as follows:
[0024]
[0025] In the formula, γ is the threshold of the sum of the total uncertainties, where q (x) , x ∈ {1, 2, 3} represents the entropy of each uncertainty estimation layer;
[0026] Sample ò from the standard Gaussian distribution, and calculate the sample ò ~ N(0, I);
[0027] (32.4) The rough matching loss L coarse and the fine matching loss L fine use the circle loss. Given an anchor descriptor di, the descriptors of its positive and negative sample pairs are respectively and
[0028]
[0029] In the formula, is the l2 feature distance, is the weight of the positive and negative sample pairs, and are the scaling factors;
[0030] Since the fine-level matching is derived from the coarse-level matching, the uncertainty modeling affects the coarse matching loss L coarse , a larger variance Σ is assigned to the unaligned image patches to reduce the interference of noisy image patches;
[0031] (33) Hierarchical matching and interaction:
[0032] After patch-level matching, the image patches and the point cloud patches interact through the cross-attention mechanism. Taking the first-level matching as an example, the image and point cloud features are projected as:
[0033] Q = f p1 W q , K = f i1u W k , V = f i1u W v
[0034] where are the projection weights of the query, key, and value;
[0035] The attention features of the anchor set are calculated as follows:
[0036]
[0037] The initial score map is calculated through cosine similarity, and the score map is refined to finally form dense matching pairs; after feature matching, the rigid body transformation [R, t] from the point cloud to the camera coordinate system is estimated by the PnP-RANSAC algorithm; the initial transformation is obtained through coarse matching, and then combined with the refined matching to optimize the pose estimation result, improving the registration accuracy.
[0038] Furthermore, an adversarial modality alignment module is constructed to reduce the cross-modal distribution difference between the image features and the point cloud features based on adversarial learning, realizing the alignment of the image feature and point cloud feature domains, specifically as follows:
[0039] (41) Distinguish the image feature f i I and the point cloud feature from d', distinguish the image domain with label 1 and the point cloud domain with label 0, and send the image feature f n i I n d into the domain classifier, which consists of three connected layers and is used to predict the label d n ;
[0040] (42) Calculate the domain alignment loss L d through the cross-entropy loss according to the difference between the predicted label and the true domain label, and the domain alignment loss L dThe definitions are as follows:
[0041]
[0042] Where N is the total number of samples, d' n is the true label of the nth sample, and d n is the predicted value of the nth sample;
[0043] The gradient of the domain alignment loss is backpropagated through the gradient reversal layer and is reversed to before being propagated back to the multi-modal feature extraction backbone network. The reversed gradient represents the feature space difference between the image and point cloud domains learned by the classifier. By reversing the gradient, the difference between the image and point cloud features is reduced, achieving the alignment of the image and point cloud feature domains.
[0044] Furthermore, the overall multi-modal feature extraction model is an overall model composed of a feature extraction backbone network, a hierarchical matching module based on uncertainty modeling, and an adversarial modal alignment module;
[0045] The total loss function of the overall multi-modal feature extraction model includes a rough matching loss L coarse , a fine matching loss L fine , an uncertainty loss L sig and a modal alignment loss L d , which are specifically as follows:
[0046] The rough matching loss L coarse and the fine matching loss L fine :
[0047] Given an anchor descriptor whose positive and negative sample pair descriptors are respectively and then the loss function is defined as follows:
[0048]
[0049] Where is the L2 distance, and are the individual weights of the positive and negative sample pairs respectively, and are scaling factors;
[0050] The uncertainty loss L sig :
[0051] The uncertainty constraint makes the sum of the variance values of all image patches a fixed value;
[0052] The modal alignment loss L d :
[0053]
[0054] Among them, N is the total number of samples, and d' n is the true label of the nth sample, and d n is the predicted value of the nth sample;
[0055] The final total loss function is defined as:
[0056] L = L coarse + L fine + L sig + L d .
[0057] According to the second aspect of the present invention, the present invention provides an image point cloud registration system based on multi-modal uncertainty modeling and modal alignment, which is used to implement the above-mentioned image point cloud registration method based on multi-modal uncertainty modeling and modal alignment, including:
[0058] A receiving unit, which is used to receive image data and point cloud data and perform preprocessing;
[0059] A feature extraction unit, which is used to use a pre-trained multi-modal feature extraction backbone network to extract features from the image data and point cloud data to obtain image features and point cloud features, and use a multi-layer self-attention module and a cross-attention module to perform interactive processing on the image features and point cloud features;
[0060] A hierarchical matching unit, which is used to construct a hierarchical matching module based on uncertainty modeling to realize hierarchical matching and interaction between image features and point cloud features, and use an uncertainty modeling method to assign a large variance to unaligned image blocks during the matching process;
[0061] A construction unit, which is used to construct an adversarial modal alignment module, and based on adversarial learning, reduce the difference in cross-modal distributions between image features and point cloud features to achieve alignment of the domains of image features and point cloud features;
[0062] A training output unit, which is used to construct a loss function to train the overall multi-modal feature extraction model until convergence to obtain an overall multi-modal feature extraction model that meets the standards.
[0063] According to the third aspect of the present invention, the present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, the above-mentioned image point cloud registration method based on multi-modal uncertainty modeling and modal alignment is adopted.
[0064] According to the fourth aspect of the present invention, the present invention provides a storage medium containing computer-executable instructions, and when the computer-executable instructions are executed by a computer processor, they are used to execute the above-mentioned image point cloud registration method based on multi-modal uncertainty modeling and modal alignment.
[0065] The present invention has at least the following beneficial effects:
[0066] 1. Through the uncertainty modeling mechanism, the present invention quantifies the importance of image patches, enables the network to pay more attention to key regions, reduces noise interference, and significantly improves the registration success rate in occlusion scenarios.
[0067] 2. By extracting multi-scale features through the hierarchical matching module and combining dynamic adjustment, the present invention effectively addresses the scale differences caused by perspective scaling, and the matching accuracy is significantly better than existing methods on public datasets.
[0068] 3. By designing the feature alignment module, the present invention uses the gradient reversal layer to narrow the feature domain differences between the image and the point cloud, realizes the consistent representation of the image and point cloud features, reduces the cross-modal matching error, and further improves the efficiency, robustness and task generalization ability of registration.
[0069] 4. The overall framework of the present invention combines lightweight design, significantly reduces the computational complexity, improves the registration efficiency, adapts to complex scenarios, and demonstrates excellent robustness and promotion ability.
[0070] Of course, it is not necessary for any product implementing the present invention to simultaneously achieve all the above-mentioned advantages. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 is a schematic flow chart of the method of the present invention;
[0072] Figure 2 is a schematic diagram of the framework principle of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0074] Please refer to Figure 1 - Figure 2 , the present invention provides a technical solution: an image point cloud registration method based on multi-modal uncertainty modeling and modal alignment, including the following steps:
[0075] S1. Receive image data and point cloud data, and pair them into groups. It should be noted that in an indoor scene, image data and point cloud data of the same scene are collected by a camera and a radar respectively;
[0076] S2. Use backbone networks for feature extraction of different modalities to extract image features and point cloud features from the image data and point cloud data, and use multi-layer self-attention modules and cross-attention modules to perform interactive processing on the image features and point cloud features;
[0077] Specifically, for the input image and point cloud First, use ResNet and Feature Pyramid Network (FPN) to extract image features respectively, and use the KPFCNN network to extract point cloud features. The features are respectively represented as and for coarse matching; they are respectively represented as and for fine matching. To improve the consistency and expression ability of multi-modal features, use multi-layer self-attention and cross-attention modules to perform interactive processing on image and point cloud features. Through the feature representation generated by this backbone network, the image feature and point cloud feature are respectively denoted as f i ' and f′ p ;
[0078] S3. Construct a hierarchical matching module based on uncertainty modeling to achieve hierarchical matching and interaction between image features and point cloud features. In the matching process, use the uncertainty modeling method to assign a large variance to the unaligned image patches. To achieve efficient cross-modal matching, the hierarchical matching module enhances the robustness and accuracy of matching through multi-scale feature extraction and uncertainty modeling;
[0079] S31. Multi-scale feature extraction:
[0080] Divide the image data into h×w patches, and extract features of different scales through a lightweight three-stage CNN, denoted as [i1, f i1 , [i2, f i2 , [i3, f i3 . Divide the point cloud data into point cloud patches through the nearest neighbor grid division method, denoted as [p1, f p1 . Multi-scale feature extraction can effectively address the scale inconsistency problem caused by perspective scaling in images;
[0081] S32. Uncertainty modeling:
[0082] In the uncertainty estimation layer, in this embodiment, the features of image patches at different scales are reconstructed, and the input feature f ix , x ∈ {1, 2, 3}, is sent to three uncertainty estimation layers. The feature distribution of a single image is modeled as a Gaussian distribution, and the feature f ixu obtained for the x-th layer through uncertainty modeling is reconstructed as a Gaussian distribution parameterized by the mean vector μ (x) and the covariance matrix Σ (x) , and this process is as shown in Figure 2 ;
[0083] A term q representing entropy is utilized, where the variance is positively correlated with the entropy. It is mainly used to prevent the variance of the flat solution from tending to zero. The term q effectively helps maintain the level of uncertainty during the training process. The entropy of a sample is defined as follows:
[0084]
[0085] Σ represents the variance values of different image patches. The larger the variance, the larger the entropy;
[0086] Through calculation, the total uncertainty loss formula is obtained by combining the losses of all layers, as follows:
[0087]
[0088] In the formula, γ is the threshold of the sum of total uncertainties, where q (x) , x belongs to {1, 2, 3}, represents the entropy of each layer;
[0089] Obviously, using L sig , the model aims to maintain the overall level of the variance of training samples. Here, this embodiment emphasizes the use of the reparameterization technique in sampling. Direct sampling will generate random feature samples, and these samples do not allow the gradient to propagate back to the previous layers. To solve this problem, in this embodiment, samples ò are first drawn from the standard Gaussian distribution, and then we calculate the samples ò ∼ N(0, I), rather than directly drawing samples from N(μ, Σ). This method separates the sampling process from parameter training and allows the effective backpropagation of gradients. Through L sig , the overall model for multi-modal feature extraction in this embodiment obtains two key capabilities: 1. Assigns a smaller variance to correctly matched image patches and a larger variance to incorrectly matched image patches; 2. In image matching, image patches with a larger variance have less influence on model training, reducing the misleading of the network that may be caused by incorrect matches;
[0090] Modeling the uncertainty of images and point cloud patches is computationally expensive and difficult to converge. Therefore, in this embodiment, only the uncertainty of images is modeled. As a two-dimensional structure, images are arranged in a regular grid and are easier to model than unordered, sparse, and irregular three-dimensional point clouds. Multi-level image patches naturally integrate multi-scale information and effectively convey contextual details to obtain more accurate and robust matches.
[0091] The specific experimental verification is as follows:
[0092] The experimental results show that the method of this embodiment outperforms all comparison methods on the RGB-D Scenes V2 dataset. In terms of the inlier ratio, the method of this embodiment achieves the highest score in all scenarios, with an average value of 35.1, significantly outperforming the sub-optimal method 2D3D-MATR (32.4); in terms of the feature matching recall rate, the method of this embodiment also performs outstandingly, with an average value as high as 94.4, and Scene.11 and Scene.12 reaching 100.0 and 99.0 respectively, far exceeding other methods; in contrast, although 2D3D-MATR performs better, its average value is only 90.8, and the average performance of other methods is lower than 60%; in addition, in terms of the registration recall rate, the method of this embodiment still remains leading, with an average value of 63.4, and is particularly outstanding in Scene.13 (74.2). In contrast, the sub-optimal method 2D3D-MATR only reaches 56.4, and the average values of the remaining methods are between 30% and 40%;
[0093] Generally speaking, the method of this embodiment achieves the best performance in all evaluation metrics. Especially in terms of the feature matching recall rate and the registration recall rate, it demonstrates strong robustness and matching ability, significantly improving the accuracy and stability of RGB-D scene matching.
[0094] It should be further noted that the reasons why the multi-modal feature extraction backbone network of this embodiment has the above two capabilities are as follows:
[0095] The rough matching loss L coarse and the fine matching loss L fine use the general circle loss;
[0096] Given an anchor descriptor di, the descriptors of its positive and negative sample pairs are respectively and
[0097]
[0098] In the formula, is the l2 feature distance, is the weight of the positive and negative sample pairs, and is the scaling factor;
[0099] Since the fine-level matching is derived from the coarse-level matching, the uncertainty modeling mainly affects L coarse , based on this understanding, let's discuss two capabilities of the model:
[0100] 1. Why are larger variances assigned to mis-matched image patches? L coarse Constrained by the circle loss, where mis-matching leads to an increase in Δ n while a decrease in Δ p , resulting in a higher L coarse , image patches with larger variances are less misleading to the network during training. However, controlling the total variance cannot meet the requirement of L coarse by reducing all variances. Then, which image patches should be assigned larger variances? Even with smaller variances, the mis-matched image patches still have a relatively large L coarse , but reducing the variance of correctly matched image patches can directly reduce L coarse , so the model assigns larger variances to unaligned image patches;
[0101] 2. Why do image patches with larger variances have less impact on model training?
[0102] If a larger uncertainty is assigned to an image patch, the samples of that patch will have a larger variance Σ, causing the result Z to deviate from the original image patch μ. Therefore, when collecting the feature space vectors of point cloud patches into different Z (1) , Z (2) ,..., Z (n) and taking their average, their gradients may cancel each other out, as shown in Figure 2 . Conversely, when the variance of an image patch is small, it will generate consistent gradients after entering the network, enhancing its importance. Therefore, the uncertainty modeling in this embodiment can determine which image patches are more or less important, thereby changing their impact on model training, making it more focused on key information patches and reducing the impact of noise patches;
[0103] S33. Hierarchical Matching and Interaction:
[0104] After patch-level matching, the image and the point cloud patch interact through a cross-attention mechanism. Taking the first layer as an example, these two sets of features are first projected as:
[0105] Q = f p1 W q , K = f i1u W k , V = f i1uW v
[0106] where are the projection weights of the query, key, and value;
[0107] The attention feature of the anchor set is calculated as follows:
[0108]
[0109] Use this to perform cross-attention hierarchical update of the point cloud features, enabling the full interaction between the point cloud features and the image features, further enhancing the feature expression ability. Calculate the initial score map (Score Map) through cosine similarity. Extract the k values with higher scores in the score map as the image-point cloud pairs of the rough matching pairs. Subsequently, introduce the point-level features in the block for similarity calculation and select the highest one. Finally, form the dense matching pairs. After completing the feature matching, estimate the rigid body transformation [R, t] from the point cloud to the camera coordinate system through the PnP-RANSAC algorithm. Obtain the initial transformation through rough matching, and then combine the refined matching to optimize the pose estimation result, improving the registration accuracy;
[0110] S4. Construct an adversarial modality alignment module to reduce the difference in the cross-modal distribution between the image features and the point cloud features based on adversarial learning, and achieve the alignment of the image feature domain and the point cloud feature domain, specifically as follows:
[0111] Design an adversarial modality alignment module for the domain difference between the image and point cloud modalities, and reduce the difference in the cross-modal feature distribution based on adversarial learning;
[0112] In this embodiment, the feature f extracted from the image i I and the feature extracted from the point cloud are distinguished from d' n The image domain (labeled as label 1) and the point cloud domain (labeled as label 0) are distinguished. The image feature f i I and the point cloud feature are fed into the domain classifier, which consists of three connected layers and is used to predict the label d n . Calculate the domain alignment loss L d through the cross-entropy loss according to the difference between the predicted label and the true domain label; d The domain alignment loss L
[0113]
[0114] where N is the total number of samples, d' n is the true label of the nth sample, and d n is the predicted value of the nth sample;
[0115] Gradient of domain alignment loss Backpropagated through the Gradient Reversal Layer (GRL) and reversed before being propagated back to the feature extraction backbone network The reversed gradient represents the feature space difference between the images and point cloud domains learned by the classifier. By reversing these gradients, the difference is alleviated, making the extracted features more consistent. Then, the adjusted features are classified to ensure they maintain their intrinsic properties despite the modality alignment. Through this adversarial method, the feature extractor is optimized and the classification is blurred, ultimately achieving the alignment of the image and point cloud feature domains;
[0116] S5. Construct a loss function to train the overall multi-modal feature extraction model until convergence, obtaining an overall multi-modal feature extraction model that meets the standards;
[0117] The overall multi-modal feature extraction model is an overall model composed of a feature extraction backbone network, a hierarchical matching module based on uncertainty modeling, and an adversarial modality alignment module;
[0118] The total loss function of the overall multi-modal feature extraction model includes a rough matching loss L coarse , a fine matching loss L fine , an uncertainty loss L sig and a modality alignment loss L d , specifically as follows:
[0119] Rough matching loss L coarse and fine matching loss L fine :
[0120] Given an anchor descriptor whose positive and negative sample pair descriptors are respectively and then the loss function is defined as follows:
[0121]
[0122] where is the L2 distance, and are the individual weights of the positive and negative sample pairs respectively, and are scaling factors;
[0123] Uncertainty loss L sig :
[0124] The uncertainty constraint makes the sum of the variance values of all image patches a fixed value;
[0125] Modality alignment loss L d :
[0126]
[0127] Among them, N is the total number of samples, d' n is the true label of the nth sample, and d n is the predicted value of the nth sample;
[0128] The final total loss function is defined as:
[0129] L = L coarse + L fine + L sig + L d .
[0130] In summary, through technical means such as multi-modal feature extraction, uncertainty modeling, and modal alignment, the present invention significantly improves the accuracy, robustness, and efficiency of image and point cloud registration, and is widely applicable to scenarios such as 3D reconstruction, SLAM, and visual positioning.
[0131] Embodiment 2:
[0132] This embodiment provides an image-point cloud registration system based on multi-modal uncertainty modeling and modal alignment, which is used to implement the above-mentioned image-point cloud registration method based on multi-modal uncertainty modeling and modal alignment, and includes:
[0133] A receiving unit, which is used to receive image data and point cloud data and perform preprocessing;
[0134] A feature extraction unit, which is used to use a pre-trained multi-modal feature extraction backbone network to extract features from the image data and point cloud data to obtain image features and point cloud features, and use a multi-layer self-attention module and a cross-attention module to perform interactive processing on the image features and point cloud features;
[0135] A hierarchical matching unit, which is used to construct a hierarchical matching module based on uncertainty modeling to realize hierarchical matching and interaction between image features and point cloud features, and adopt an uncertainty modeling method to assign a large variance to unaligned image blocks during the matching process;
[0136] A construction unit, which is used to construct an adversarial modal alignment module to reduce the difference in cross-modal distributions between image features and point cloud features based on adversarial learning, and realize the alignment of the domains of image features and point cloud features;
[0137] A training and output unit, which is used to construct a loss function to train the overall multi-modal feature extraction model until convergence to obtain an overall multi-modal feature extraction model that meets the standards.
[0138] Specifically, the above receiving unit, feature extraction unit, hierarchical matching unit, construction unit, and training output unit can be embedded in a computer processing system. The computer, based on the above-provided image point cloud registration method based on multi-modal uncertainty modeling and modal alignment, calls the above units to complete the task of registering the image point cloud; the above receiving unit, feature extraction unit, hierarchical matching unit, construction unit, and training output unit can perform operations according to the specific steps given by the above image point cloud registration method based on multi-modal uncertainty modeling and modal alignment.
[0139] It should be noted that it should be understood that the division of each unit of the above system is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these units can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; it is also possible that some units are implemented in the form of software called by processing elements, and some units are implemented in the form of hardware. For example, the receiving unit can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above signal processing unit. The implementation of other units is similar. In addition, all or part of these units can be integrated together or can be independently implemented. The processing element mentioned here can be an integrated circuit with the ability to process signals. In the implementation process, each step of the above method or each of the above units can be completed through the hardware integrated logic circuit or software-form instructions in the processor element.
[0140] For example, the above units can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain unit above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these units can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0141] Embodiment 3:
[0142] The present invention provides a terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The computer program stored in the memory can run on the processor. When the processor loads and executes the computer program, the above-mentioned image point cloud registration method based on multi-modal uncertainty modeling and modal alignment is adopted.
[0143] It should be noted that the terminal device can be a computer device such as a desktop computer, a laptop computer, or a cloud server. And the terminal device includes but is not limited to a processor and a memory. For example, the terminal device may further include input / output devices, network access devices, and a bus, etc.
[0144] Furthermore, the processor can adopt a central processing unit (CPU). Of course, according to the actual usage situation, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc. This application does not make any restrictions in this regard.
[0145] Embodiment 4:
[0146] The present invention provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the above-mentioned image point cloud registration method based on multi-modal uncertainty modeling and modal alignment when executed by a computer processor.
[0147] Among them, the computer program can be stored in a computer-readable medium. The computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some middleware form, etc. The computer-readable medium includes any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer-readable medium includes but is not limited to the above components.
[0148] It should be noted that, in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0149] For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. When an element is referred to as being "assembled on", "mounted on", "fixed to" or "disposed on" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "upper", "lower", "left", "right" and similar expressions used herein are for illustrative purposes only and do not represent the only embodiments.
[0150] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
[0151] In the description of this specification, the description with reference to terms such as "an embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
Claims
1. An image point cloud registration method based on multi-modal uncertainty modeling and modal alignment, which is applied to indoor scenes, is characterized in that It includes the following steps: Receive image data and point cloud data, and pair them into groups; Use the feature extraction backbone network of different modalities to extract features from the image data and point cloud data, obtain image features and point cloud features, and use the multi-layer self-attention module and cross-attention module to perform interactive processing on the image features and point cloud features; Construct a hierarchical matching module based on uncertainty modeling to achieve hierarchical matching and interaction between image features and point cloud features, and use the uncertainty modeling method to assign a large variance to the unaligned image patches during the matching process; Construct an adversarial modality alignment module to reduce the difference in cross-modal distributions between image features and point cloud features based on adversarial learning, and achieve the alignment of the domains of image features and point cloud features; Construct a loss function to train the overall multi-modal feature extraction model until convergence, and obtain an overall multi-modal feature extraction model that meets the standards.
2. The method for image point cloud registration based on multi-modal uncertainty modeling and modal alignment according to claim 1, characterized in that: The multi-modal feature extraction backbone network is for the received image data and point cloud data respectively uses the ResNet network and the feature pyramid network to extract image features, and uses the KPFCNN network to extract point cloud features; The image features and point cloud features are respectively represented as and for rough matching; they are respectively represented as and for fine matching.
3. The method for image-point cloud registration based on multi-modal uncertainty modeling and modal alignment according to claim 1, characterized in that, Construct a hierarchical matching module based on uncertainty modeling to achieve hierarchical matching and interaction between image features and point cloud features, and use the uncertainty modeling method to assign a large variance to the unaligned image patches during the matching process, specifically as follows: (31) Multi-scale feature extraction: The image data is segmented into h×w patches, and features at different scales are extracted through a lightweight three-stage CNN, denoted as [i1, f i1 , [i2, f i2 , [i3, f i3 . The point cloud data is partitioned into point cloud patches through the nearest neighbor grid partitioning method, denoted as [p1, f p1 ; (32) Uncertainty modeling: (32.1) In the uncertainty estimation layer, the features of image patches at different scales are reconstructed, and the input feature f ix , x ∈ {1, 2, 3}, are fed into three uncertainty estimation layers; Model the distribution of individual image features as a Gaussian distribution, and obtain the feature f for the x-th layer through uncertainty modeling ixu , the feature f ixu is reconstructed by a Gaussian distribution parameterized by the mean vector μ (x) and the covariance matrix Σ (x) . (32.2) Define a term q representing entropy, where the variance is positively correlated with the entropy, and the entropy is defined as follows: Σ represents the variance values of different image patches. The larger the variance, the larger the entropy; (32.3) Through calculation, the total uncertainty loss formula is obtained by combining the losses of all layers, as follows: where γ is the threshold of the total sum of uncertainties, where q (x) , and x ∈ {1, 2, 3} represents the entropy of each uncertainty estimation layer; Draw a sample ò from the standard Gaussian distribution, and calculate the sample ò~N(0,I); (32.4) Coarse matching loss L coarse and fine matching loss L fine The circle loss is used. Given an anchor descriptor di, the descriptors of its positive and negative sample pairs are respectively and In the formula, is the l2 feature distance, is the weight of the positive and negative sample pairs, and are scaling factors; Since the fine-level matching is derived from the coarse-level matching, the uncertainty modeling affects the coarse matching loss L coarse , a larger variance Σ is assigned to the unaligned image patches to reduce the interference of noisy image patches; (33) Hierarchical matching and interaction: After patch-level matching, the image patches and point cloud patches interact through the cross-attention mechanism. Taking the first-layer matching as an example, the image and point cloud features are projected as: Q = f p1 W q ,K = f i1u W k ,V = f i1u W v In the formula are the projection weights of the query, key, and value; The attention feature of the anchor set is calculated as follows: Calculate the initial score map through cosine similarity, refine the score map, and finally form dense matching pairs; after completing feature matching, estimate the rigid body transformation [R,t] from the point cloud to the camera coordinate system through the PnP-RANSAC algorithm; obtain the initial transformation through coarse matching, and then combine the refined matching to optimize the pose estimation result, improving the registration accuracy.
4. The method for image point cloud registration based on multi-modal uncertainty modeling and modal alignment according to claim 3, wherein Construct an adversarial modality alignment module to reduce the difference in cross-modal distributions between image features and point cloud features based on adversarial learning, and achieve the alignment of the domains of image features and point cloud features, specifically as follows: (41) Distinguish the image feature f i I and the point cloud feature from d', distinguish the image domain with label 1 and the point cloud domain with label 0, and send the image feature f n i I and the point cloud feature into the domain classifier, which consists of three connected layers and is used to predict the label d n ; (42) Calculate the domain alignment loss \(L\) based on the difference between the predicted label and the true domain label through cross - entropy loss d , the domain alignment loss \(L\) d is defined as follows: where N is the total number of samples, d' n is the true label of the nth sample, and d n is the predicted value of the nth sample; Gradient of domain alignment loss Backpropagated through the gradient reversal layer and reversed before propagating back to the multi-modal feature extraction backbone network The reversed gradient represents the feature space difference between the images and point cloud domains learned by the classifier. By reversing the gradient, the difference between image and point cloud features is reduced, achieving alignment of the image and point cloud feature domains.
5. The method for image point cloud registration based on multi-modal uncertainty modeling and modal alignment according to claim 4, characterized in that: The overall multi-modal feature extraction model is an overall model composed of a feature extraction backbone network, a hierarchical matching module based on uncertainty modeling, and an adversarial modality alignment module; The total loss function of the multi-modal feature extraction overall model includes a rough matching loss L coarse , a fine matching loss L fine , an uncertainty loss L sig and a modality alignment loss L d , which are specifically as follows: Coarse matching loss L coarse and fine matching loss L fine : Given an anchor descriptor The descriptors of its positive and negative sample pairs are respectively and The loss function is defined as follows: Among them, is the L2 distance, and are the individual weights of the positive and negative sample pairs respectively, and are scaling factors; Uncertainty loss L sig : The uncertainty constraint makes the sum of the variance values of all image patches a fixed value; Modal alignment loss L d : where N is the total number of samples, d' n is the true label of the n-th sample, and d n is the predicted value of the n-th sample; The final total loss function is defined as: L = L coarse + L fine + L sig + L d 。 6. An image point cloud registration system based on multimodal uncertainty modeling and modality alignment, which is used to implement the multimodal uncertainty modeling and modality alignment-based image point cloud registration method described in any one of claims 1 to 5, characterized in that, It includes: A receiving unit for receiving image data and point cloud data and performing preprocessing; A feature extraction unit for using the pre-trained multi-modal feature extraction backbone network to extract features from the image data and point cloud data, obtain image features and point cloud features, and use the multi-layer self-attention module and cross-attention module to perform interactive processing on the image features and point cloud features; A hierarchical matching unit, which is used to construct a hierarchical matching module based on uncertainty modeling, realize the hierarchical matching and interaction between image features and point cloud features, and adopt an uncertainty modeling method to assign a large variance to unaligned image patches during the matching process; A construction unit, which is used to construct an adversarial modality alignment module, reduce the difference in cross-modal distributions between image features and point cloud features based on adversarial learning, and realize the alignment of the domains of image features and point cloud features; A training and output unit, which is used to construct a loss function to train the overall multi-modal feature extraction model until convergence, and obtain an overall multi-modal feature extraction model that meets the standards.
7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that The memory stores a computer program that can run on a processor. When the processor loads and executes the computer program, the method for image-point cloud registration based on multi-modal uncertainty modeling and modality alignment described in any one of claims 1 to 5 is adopted.
8. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions are used to execute the method for image-point cloud registration based on multi-modal uncertainty modeling and modality alignment described in any one of claims 1 to 5 when executed by a computer processor.