Image point cloud registration method and system based on implicit correspondence learning
Through an implicit correspondence learning method, combined with geometric priors and multi-layer perceptron network, the error matching problems caused by inaccurate detection of overlapping areas in image-point cloud registration are solved, and the image-point cloud registration effect with high accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510380596.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-17
AI Technical Summary
The existing image-point cloud registration methods have problems with inaccurate detection of overlapping area boundaries, direct mismatch of cross-modal data, and failure to directly optimize camera poses.
Image-point cloud registration is realized through geometric prior guidance overlaid region detection, implicit corresponding learning module and pose regression module. The specific steps include obtaining scene point cloud and image data, extracting 3D and 2D features, using a multi-layer perceptron network to learn probability and poses within the view cone, iteratively apply pixels and point attention, obtaining 3D and 2D key points, and performing pose regression through a multi-layer perceptron network.
It significantly improves the accuracy and robustness of image-point cloud registration, avoids the problem of wrong alignment in traditional methods, and directly optimizes the camera position, improving the accuracy of registration results.
Smart Images

Figure CN120163852A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and specifically to an image-point cloud registration method and system based on implicit correspondence learning. Background Art
[0002] Image-point cloud registration is a fundamental task in the field of computer vision, aiming to estimate the camera pose of a given image in a 3D scene point cloud. It plays an important role in many computer vision tasks, including SLAM (Simultaneous Localization and Mapping), 3D reconstruction, augmented / virtual reality, and visual localization. Although there have been advancements in the cross-modal registration task of image-point cloud in recent years, due to the inherent difference that images capture two-dimensional grid data while point clouds represent structures in an unordered format, the image-point cloud registration task remains challenging.
[0003] Currently, the methods for image-point cloud registration can be divided into two major categories: including non-matching methods and matching-based methods. Non-matching methods utilize the camera imaging principle as a prior for predicting the camera pose. For example, DeepI2P converts the registration problem into a classification problem and designs a binary classification network to determine whether each scene point is inside or outside the camera frustum. Then, the classification result is input into a camera back-projection solver to obtain the camera pose. However, these methods only determine a rough correspondence between the image and the points. For example, whether a 3D point is inside the image frustum, which is not sufficient for accurate camera pose estimation. Different from non-matching methods, matching-based methods focus on establishing fine-grained 2D-3D correspondences and use the PnP-RANSAC algorithm to estimate the camera pose, achieving better performance. These methods generally include three important steps: (1) Design a detector to identify the overlapping regions between the image and the point cloud. (2) Establish 2D-3D correspondences within the identified overlapping regions. (3) Employ the PnP-RANSAC algorithm as a post-processing module to obtain the camera pose.
[0004] Despite the success of matching-based methods, there are still three significant problems: (1) Most matching-based methods use a point-by-point classification method to determine whether each point is within the overlapping region, ignoring the global projection relationship between the image and the point cloud. Therefore, existing methods often have inaccurate boundary detection of the overlapping region. (2) Matching-based methods essentially establish 2D-3D correspondences by directly calculating the feature similarity between the image and the point cloud. These methods use metric learning techniques to force the corresponding 2D and 3D features to align. However, due to the inherent data differences between the image and the point cloud, directly processing cross-modal data results in a large number of false matches. (3) After establishing the 2D-3D correspondences, most existing methods use the PnP-RANSAC algorithm as a post-processing module to estimate the camera pose. These methods mainly take the 2D-3D correspondences as the optimization objective, rather than directly optimizing the camera pose. Summary of the Invention
[0005] To address the deficiencies mentioned in the above background art, the purpose of the present invention is to provide an image-point cloud registration method and system based on implicit correspondence learning.
[0006] In a first aspect, the purpose of the present invention can be achieved through the following technical solutions: An image-point cloud registration method based on implicit correspondence learning, the method comprising the following steps:
[0007] Obtain scene point cloud data and image data, extract 3D features and 2D features from the scene point cloud data and the image data, and process the 3D features and the 2D features based on an overlapping region detection strategy guided by geometric priors to obtain a point cloud mask;
[0008] Filter out the points predicted to be outside the overlapping region based on the point cloud mask, so that the randomly initialized learnable correspondence query features only focus on the potential overlapping region, iteratively apply pixel attention and point attention to the learnable correspondence query features, and after multiple iterations, obtain a 3D key point detector and a 2D key point detector, and further interact with the 3D features and the 2D features to obtain 3D key points and 2D key points, and process the 3D key points and the 2D key points based on two pre-established parallel multi-layer perceptron networks to obtain the point cloud registration result.
[0009] In combination with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of extracting 3D features and 2D features from the scene point cloud data and the image data:
[0010] For the input scene point cloud and the image from the same scene respectively use the KPFCNN network to extract 3D features Extract 2D features using ResNet and Feature Pyramid Network (FPN) Represent the image pixel coordinate matrix as
[0011] Combined with the first aspect, in some implementations of the first aspect, the method further includes: for the extracted 2D feature F I Perform average pooling (Avgpool) to obtain the 2D global feature, and concatenate the feature channels with the 3D feature F P Use a multi-layer perceptron (MLP) to learn the probability that each point is within the camera frustum The probability O within the entire frustum P The learning is expressed as follows:
[0012] O P = MLP([F P , Tile N (Avgpool(F I ))]),
[0013] where MLP(·) represents the multi-layer perceptron, [·,·] represents feature channel concatenation, Tile N (·) represents replicating the feature N times, and Avgpool(·) represents the average pooling operation;
[0014] Using the probability O within the frustum P and the coordinate information of the point cloud P, use another multi-layer perceptron (MLP) to regress the frustum pose T f ∈ SE(3), and the frustum pose T f consists of a 3D rotation matrix R f ∈ SO(3) and a 3D translation vector where two parallel multi-layer perceptrons are respectively applied to regress R f and t f :
[0015] R f , t f = ρ(MLP([O P , P])), MLP([O P , P]),
[0016] Select to use the 6D representation of rotation as the regression target of the first MLP, and ρ(·) represents the transformation from the 6D representation to the 3×3 rotation matrix R f
[0017] Combined with the first aspect, in some implementations of the first aspect, the method further includes: Define the projection function from point to pixel according to the given camera intrinsic matrix as follows:
[0018] [x i ,y i ,z i T =K·(R f ·p i +t f ),
[0019] u i =[u xi ,u yi T =[x i / z i ,y i / z i T ,
[0020] f(p i ;K,R f ,t f )=u i ,
[0021] Wherein, p i ∈P represents a 3D point, and u i represents a 2D pixel, and a point cloud mask is obtained
[0022]
[0023] Wherein, represents that the 3D point p i is projected inside the image I, represents that the 3D point p i is projected outside the image I.
[0024] Combined with the first aspect, in some implementations of the first aspect, the method further includes: after the point cloud mask M P is obtained:
[0025] Initialize N q corresponding query features with a set of learnable parameters By aggregating information from the 2D feature F I and the 3D feature F P , convert the corresponding query feature F q into 2D and 3D key point detectors, iteratively apply pixel attention and point attention, and for each iteration k≥0, use as the input and interact with the 2D feature F I and the 3D feature F P :
[0026]
[0027] Among them, Attention_Pixel represents pixel attention, and Attention_Point represents point attention. And perform it L times alternately to obtain the updated corresponding query feature
[0028] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of iteratively applying pixel attention and point attention to the learnable corresponding query feature, and obtaining the 3D key point detector and the 2D key point detector after multiple iterations:
[0029] Pixel attention:
[0030] For the k-th iteration, use the output from the previous iteration Aggregate the information from the 2D feature F I to obtain
[0031]
[0032] Among them, represents a linear mapping, and use the multi-head attention mechanism to update
[0033]
[0034] The output attention feature is then passed through a multi-layer perceptron to obtain
[0035] Point attention:
[0036] For the k-th iteration, use the output of the pixel attention Aggregate the information from the 3D feature F P to obtain
[0037]
[0038] The output attention feature is then passed through a multi-layer perceptron to obtain
[0039] Combined with the first aspect, in some implementation manners of the first aspect, the method further includes: the process of obtaining the 3D key points and the 2D key points:
[0040] After L iterations, obtain the updated corresponding query feature The corresponding query feature is adaptively transformed according to the content of the input point cloud and the image, captures the same key point pattern, and then uses As inputs, a 3D keypoint detector D is obtained using point attention and pixel attention respectively P and a 2D keypoint detector D I . A 3D keypoint heatmap is obtained by calculating the inner product between the 3D keypoint detector and the 3D features and then performing a Softmax operation Then, the predicted point cloud mask M P is used to suppress irrelevant points:
[0041]
[0042] Furthermore, a 2D keypoint heatmap
[0043] H I = Softmax · D I Flatten(F P ) T ).
[0044] where Flatten(·) represents a flattening operation that converts the first two feature dimensions into one feature dimension; then the 3D keypoints and 2D keypoints
[0045] K P = H P P, K I = H I E I .
[0046] Combined with the first aspect, in some implementations of the first aspect, the method further includes: the process of processing the 3D keypoints and 2D keypoints based on two pre-established parallel multi-layer perceptron networks to obtain a point cloud registration result is as follows:
[0047] Based on the frustum pose T f regarded as an initial camera pose to be refined, and corresponding transformation is performed on the 3D keypoints K P , where the difference between R f , t f and the true camera pose R gt , t gt is:
[0048] ΔR gt = R gt · (R f ) -1 , Δt gt = t gt - ΔR gt · t f .
[0049] The multi - layer perceptron is applied to elevate the dimension of the correspondence features, and the features are paired through feature splicing to obtain fused features
[0050] F f = MLP([D P , MLP(K P ), D I , MLP(K I )]).
[0051] Global information is injected into the point - pixel correspondence by embedding the average feature of F f :
[0052]
[0053] The pose - sensitive feature f pose is learned from all correspondence information through average pooling operation:
[0054] f pose = Avgpool(MLP(F s )).
[0055] Two parallel multi - layer perceptrons are applied to respectively regress ΔR and Δt as the point cloud registration result:
[0056] ΔR, Δt = ρ(MLP(f pose ), MLP(f pose
[0057] In a second aspect, to achieve the above object, the present invention discloses an image - point cloud registration system based on implicit correspondence learning, including:
[0058] A data processing module, configured to obtain scene point cloud data and image data, extract features from the scene point cloud data and the image data to obtain 3D features and 2D features, and process the 3D features and the 2D features based on a geometric - prior - guided overlapping region detection strategy to obtain a point cloud mask;
[0059] A point cloud registration module, configured to filter out points predicted outside the overlapping region based on the point cloud mask, enabling the randomly initialized learnable correspondence query features to only focus on potential overlapping regions, iteratively applying pixel attention and point attention to the learnable correspondence query features, obtaining a 3D key - point detector and a 2D key - point detector after multiple iterations, further interacting with the 3D features and the 2D features to obtain 3D key points and 2D key points, and processing the 3D key points and the 2D key points based on two pre - established parallel multi - layer perceptron networks to obtain the point cloud registration result.
[0060] In another aspect of the present invention, in order to achieve the above object, a terminal device is disclosed, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. The memory stores a computer program capable of running on the processor. When the processor loads and executes the computer program, a method for image-point cloud registration based on implicit correspondence learning as described above is adopted.
[0061] Advantages of the present invention:
[0062] The present invention has three specific designs for image-point cloud registration. First, the proposed geometric prior can guide the model to achieve high accuracy in overlapping region detection. Second, the proposed implicit correspondence learning module can implicitly obtain accurate 2D-3D correspondences, avoiding the problem of forced alignment of image and point cloud features in traditional methods. Finally, a pose regression module is proposed to handle incorrect 2D-3D correspondences and generate accurate camera poses. Brief Description of the Drawings
[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Figure 1 It is a schematic flowchart of the method of the present invention;
[0065] Figure 2 It is a schematic diagram of the image-point cloud registration model framework based on implicit correspondence learning of the present invention;
[0066] Figure 3 It is a schematic diagram of the system structure of the present invention. Detailed Embodiments
[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0068] Embodiment 1:
[0069] As Figure 1 shown, a method for image-point cloud registration based on implicit correspondence learning, the method includes the following steps:
[0070] S101: Obtain the scene point cloud data and image data, extract features from the scene point cloud data and image data to obtain 3D features and 2D features, and process the 3D features and 2D features based on the overlapping region detection strategy guided by geometric priors to obtain a point cloud mask;
[0071] For the input scene point cloud and the image from the same scene We first use the KPFCNN network to extract 3D features respectively Use ResNet and Feature Pyramid Network (FPN) to extract 2D features At the same time, we represent the image pixel coordinate matrix as
[0072] To accurately detect the overlapping region of the point cloud, inspired by the fact that the image in the point cloud must correspond to a frustum, we design a geometric prior-guided overlapping region detection strategy to obtain more accurate overlapping region detection results. Specifically, we first perform average pooling (Avgpool) on the extracted 2D feature F I to obtain the 2D global feature, and perform feature channel concatenation with the 3D feature F P Then we use a multi-layer perceptron (MLP) to learn the probability that each point is located within the camera frustum Generally speaking, the learning of the overall probability O P within the frustum is expressed as follows:
[0073] O P = MLP([F P , Tile N (Avgpool(F I ))]),
[0074] where MLP(·) represents the multi-layer perceptron, [·,·] represents feature channel concatenation, Tile N (·) represents copying the feature N times, and Avgpool(·) represents the average pooling operation.
[0075] Since the probability O P within the frustum is not accurate enough for overlapping region detection, we propose to estimate the frustum pose to obtain more accurate overlapping region detection results. Using the probability O P within the frustum and the coordinate information of the point cloud P, we use another multi-layer perceptron (MLP) to regress the frustum pose T f ∈ SE(3). Among them, the frustum pose T f consists of a 3D rotation matrix R f ∈ SO(3) and a 3D translation vector Composition. Specifically, we apply two parallel multi-layer perceptrons respectively to regress and obtain R f and t f :
[0076] R f , t f = ρ(MLP([O P , P])), MLP([O P , P]),
[0077] where we choose to use the rotated 6D representation as the regression target of the first MLP, and ρ(·) represents the transformation from the 6D representation to the 3×3 rotation matrix R f . According to the given camera intrinsic matrix we define the point-to-pixel projection function as follows:
[0078] [x i , y i , z i T = K·(R f ·p i + t f ),
[0079] u i = [u xi , u yi T = [x i / z i , y i / z i T ,
[0080] f(p i ; K, R f , t f ) = u i ,
[0081] where p i ∈P represents a 3D point and u i represents a 2D pixel. We obtain the point cloud mask
[0082]
[0083] where if the 3D point p i is projected inside the image I, if the 3D point p i is projected outside the image I.
[0084] S102: Filter out the points predicted to be outside the overlapping region based on the point cloud mask, so that the randomly initialized learnable correspondence query features only focus on the potential overlapping region. Apply pixel attention and point attention to the iteration of the learnable correspondence query features. After multiple iterations, obtain the 3D key point detector and the 2D key point detector, and further interact with the 3D features and 2D features to obtain the 3D key points and 2D key points. Process the 3D key points and 2D key points based on two pre-established parallel multi-layer perceptron networks to obtain the point cloud registration result.
[0085] After obtaining the point cloud mask M P we implicitly obtain a set of robust correspondences between the image and the point cloud. Specifically, we first initialize N q learnable correspondence query features with a set of learnable parameters Then, we transform the correspondence query feature F I and 3D feature F P into 2D and 3D key point detectors by aggregating the information from the 2D feature F q . To this end, we iteratively apply pixel attention followed by point attention. For each iteration k≥0, we take as the input and interact with the 2D feature F I and 3D feature F P :
[0086]
[0087] where Attention_Pixel represents pixel attention and Attention_Point represents point attention. These two steps are alternately executed L times to obtain the updated correspondence query feature
[0088] Pixel attention:
[0089] For the k-th iteration, we use output from the previous iteration to aggregate the information from the 2D feature F I to obtain
[0090]
[0091] where represents a linear mapping. Then we use the multi-head attention mechanism to update
[0092]
[0093] The output attention features are then passed through a multi-layer perceptron to obtain
[0094] Point attention:
[0095] For the k-th iteration, we use the output of the pixel attention to aggregate the information from the 3D feature F P to obtain
[0096]
[0097]
[0098] Here, we use the point cloud mask M P to focus the information interaction on meaningful regions. The output attention features are then passed through a multi-layer perceptron to obtain
[0099] After L iterations, we obtain the updated corresponding query features These corresponding query features can adaptively transform according to the content of the input point cloud and image, capturing the same key point patterns. Next, we use as the input, and use point attention and pixel attention respectively to obtain the 3D key point detector D P and the 2D key point detector D I . By calculating the inner product between the 3D key point detector and the 3D feature and then performing the Softmax operation, we can obtain the 3D key point heat map Meanwhile, we use the predicted point cloud mask M P to suppress irrelevant points:
[0100]
[0101] Similarly, we can obtain the 2D key point heat map
[0102] H I = Softmax(D I Flatten(F P )) T .
[0103] where Flatten(() represents the flattening operation, that is, converting the first two feature dimensions into one feature dimension. Finally, we can obtain the 3D key points and the 2D key points
[0104] K P = H P P, KI = H I E I .
[0105] After obtaining the 2D-3D correspondence, we use a learnable network to regress the camera pose. Since we estimate the frustum pose T in the geometric prior-guided overlap region detector f , and this is a rough camera pose estimate, we can regard T f as an initial camera pose to be refined, and perform corresponding transformations on the 3D keypoints K P . In this case, we have R f , t f and the difference between the true camera pose R gt , t gt :
[0106] ΔR gt = R gt · (R f ) -1 , Δt gt = t gt - ΔR gt · t f .
[0107] We first apply a multi-layer perceptron to elevate the dimensionality of the correspondence features and pair the correspondences through feature concatenation to obtain the fused feature
[0108] F f = MLP([D P , MLP(K P ), D I , MLP(K I )]).
[0109] Then, we inject the global information into the point-pixel correspondence through the average feature of the embedded F f to enhance the correspondence information:
[0110]
[0111] The pose-sensitive feature f pose can be learned from all the correspondence information through an average pooling operation:
[0112] f pose = Avgpool(MLP(F s )).
[0113] Finally, we apply two parallel multi-layer perceptrons to regress ΔR and Δt respectively:
[0114] ΔR, Δt = ρ(MLP(f pose )), MLP(f pose ).
[0115] Loss function design.
[0116] The loss function of this method consists of the following parts:
[0117] Classification loss Constraining the probability O that the predicted points in the overlapping region detector guided by geometric priors are located within the camera frustum P to be as accurate as possible, and calculating the binary cross-entropy loss by comparing with the true classification label O gt as follows:
[0118]
[0119] Frustum pose prediction loss Constraining the predicted frustum pose R f , t f in the overlapping region detector guided by geometric priors to be consistent with the true camera pose R gt , t gt as follows:
[0120]
[0121] Correspondence loss Constraining the predicted 2D and 3D key points to be consistent:
[0122]
[0123] Diversity loss Constraining the predicted 2D and 3D key points not to cluster in the same area, increasing the difference between key points:
[0124]
[0125]
[0126] Camera pose prediction loss Constraining the predicted camera pose to be consistent with the true value:
[0127]
[0128] The final total loss function is defined as:
[0129]
[0130] Through technical means such as geometric prior-guided overlapping region detection, implicit correspondence learning, and pose regression, the present invention significantly improves the accuracy and robustness of image-point cloud registration and can be widely applied to scenarios such as 3D reconstruction, SLAM, and visual positioning.
[0131] Specifically, the solution of the present invention will be further elaborated through the following embodiments:
[0132] The present invention can be widely applied to systems in the fields of 3D reconstruction, SLAM (Simultaneous Localization and Mapping), visual positioning, etc., and is particularly suitable for scenarios that require image and point cloud data registration and pose estimation. In specific implementations, the present invention can be integrated into various front-end devices, such as drones, mobile robots, and autonomous driving vehicles, to process real-time image and point cloud data. These devices operate in dynamic environments and need to quickly and accurately complete pose estimation tasks to achieve autonomous navigation and decision-making. With the method of the present invention, the front-end device can utilize the image and point cloud data collected by the camera and LiDAR sensor, and through geometric prior-guided overlapping region detection, accurately identify the corresponding regions between the image and the point cloud. This method not only improves the accuracy of overlapping region detection but also reduces the false matches caused by factors such as noise and occlusion. On this basis, the implicit correspondence learning module avoids the problem of directly aligning image and point cloud features by learning query vectors, and can robustly find the correct 2D-3D correspondence relationship under the different backgrounds of the image and the point cloud. The pose regression module of the present invention directly inputs the 2D-3D key point information into a multi-layer perceptron (MLP) through end-to-end optimization and can directly regress the pose of the camera. In addition to the real-time application of the front-end device, the present invention can also be deployed in the background server to support large-scale and multi-scene image and point cloud data processing tasks. In application scenarios such as urban 3D reconstruction and environmental modeling, by batch processing image and point cloud data, the background server can provide high-precision pose estimation and registration results for multiple scenes. This is of great significance for supporting multi-modal analysis and 3D reconstruction of large datasets, especially in applications such as autonomous driving and robot perception, and can provide accurate environmental understanding and decision-making basis for the system.
[0133] Embodiment 2: As Figure 3 shown, to achieve the above object, the present invention discloses an image-point cloud registration system based on implicit correspondence learning, including:
[0134] A data processing module, configured to obtain scene point cloud data and image data, extract 3D features and 2D features from the scene point cloud data and the image data, and process the 3D features and the 2D features based on a geometric prior-guided overlapping region detection strategy to obtain a point cloud mask;
[0135] The point cloud registration module is used to filter out points predicted outside the overlapping region based on the point cloud mask, enabling the randomly initialized learnable correspondence query features to only focus on potential overlapping regions. Pixel attention and point attention are iteratively applied to the learnable correspondence query features. After multiple iterations, a 3D key point detector and a 2D key point detector are obtained, and further interact with the 3D features and 2D features to obtain 3D key points and 2D key points. Based on two pre-established parallel multi-layer perceptron networks, the 3D key points and 2D key points are processed to obtain the point cloud registration result.
[0136] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0137] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electro-magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.
[0138] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0139] The above has shown and described the basic principles, main features, and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.
Claims
1. An image point cloud registration method based on implicit correspondence learning, characterized in that, The method includes the following steps: Obtain scene point cloud data and image data, extract features from the scene point cloud data and image data to obtain 3D features and 2D features, and process the 3D features and 2D features based on a geometric prior-guided overlapping region detection strategy to obtain a point cloud mask; Filter out the points predicted to be outside the overlapping region based on the point cloud mask, so that the randomly initialized learnable correspondence query features only focus on the potential overlapping regions, iteratively apply pixel attention and point attention to the learnable correspondence query features, and after multiple iterations, obtain a 3D key point detector and a 2D key point detector, further interact with the 3D features and 2D features to obtain 3D key points and 2D key points, and process the 3D key points and 2D key points based on two pre-established parallel multi-layer perceptron networks to obtain a point cloud registration result.
2. The image point cloud registration method based on implicit correspondence learning according to claim 1, characterized in that, The process of extracting features from the scene point cloud data and image data to obtain 3D features and 2D features: For the input scene point cloud and the image from the same scene Use the KPFCNN network to extract 3D features respectively Use ResNet and the Feature Pyramid Network (FPN) to extract 2D features Represent the image pixel coordinate matrix as 3. The image point cloud registration method based on implicit correspondence learning according to claim 2, characterized in that, For the extracted 2D feature F I Perform average pooling Avgpool to obtain the 2D global feature, and concatenate it with the 3D feature F P Perform feature channel concatenation, and use a multi-layer perceptron MLP to learn the probability that each point is within the camera frustum The probability O within the entire frustum P The learning representation is as follows: O P = MLP([F P , Tile N (Avgpool(F I ))]), Among them, MLP(·) represents a multi-layer perceptron, [·, ·] represents feature channel concatenation, and Tile N (·) represents copying the feature N times, and Avgpool(·) represents an average pooling operation; Utilize the probability O within the frustum P and the coordinate information of the point cloud P, and use another multi-layer perceptron MLP to regress the frustum pose T f ∈ SE(3), where the frustum pose T f is composed of the 3D rotation matrix R f ∈ SO(3) and the 3D translation vector respectively, and two parallel multi-layer perceptrons are applied to regress R f and t f : R f ,t f = ρ(MLP([O P ,P])), MLP([O P ,P]), Select to use the rotated 6D representation as the regression target of the first MLP, and ρ(·) represents the transformation from the 6D representation to the 3×3 rotation matrix R f transformations.
4. The image point cloud registration method based on implicit correspondence learning according to claim 3, characterized in that, According to the given camera intrinsic matrix Define the projection function from points to pixels as follows: [x i ,y i ,z i T = K·(R f ·p i +t f ), u i = [u xi , u yi T = [x i / z i , y i / z i T , f(p i ; K,R f ,t f ) = u i , where p i ∈ P represents a 3D point, and u i represents a 2D pixel, and the point cloud mask Among them, represents that the 3D point p i is projected inside the image I, represents that the 3D point p i is projected outside the image I.
5. The image point cloud registration method based on implicit correspondence learning according to claim 4, characterized in that, The point cloud mask M P After obtaining: Initialize N with a set of learnable parameters q corresponding query features By aggregating information from 2D feature F I and 3D feature F P transform the corresponding query feature F q into 2D and 3D keypoint detectors, iteratively apply pixel attention and point attention, and for each iteration k≥0, take as the input and interact with 2D feature F I and 3D feature F P as follows: Among them, Attention_Pixel represents pixel attention, and Attention_Point represents point attention. And it is executed L times alternately to obtain the updated corresponding query feature 6. The image point cloud registration method based on implicit correspondence learning according to claim 1, characterized in that, The process of iteratively applying pixel attention and point attention to the learnable correspondence query features and obtaining a 3D key point detector and a 2D key point detector after multiple iterations: Pixel attention: For the k-th iteration, use the output from the previous iteration Aggregate the information from the 2D feature F I to obtain Among them, represents a linear mapping and is updated using the multi-head attention mechanism The output attention features are then obtained through a multi-layer perceptron Point attention: For the k-th iteration, using the output of the pixel attention aggregate the information from the 3D feature F P to obtain The output attention features are then passed through a multi-layer perceptron to obtain 7. A method for image point cloud registration based on implicit correspondence learning according to claim 1, characterized in that The process of obtaining 3D key points and 2D key points: After L iterations, the updated corresponding query features are obtained The corresponding query features are adaptively transformed according to the content of the input point cloud and image, capturing the same key point pattern, and then using as the input, the 3D key point detector D is obtained using point attention and pixel attention respectively P and the 2D key point detector D I . The 3D key point heat map is obtained by calculating the inner product between the 3D key point detector and the 3D features and then performing the Softmax operation Then, the predicted point cloud mask M P is used to suppress irrelevant points: H P = Softmax(D P F P T + Tile N q ((M P - 1)·+∞)). Furthermore, a 2D key point heat map is obtained. H I = Softmax(D I Flatten(F P ) T ). Among them, Flatten(·) represents the flattening operation, which converts the first two feature dimensions into one feature dimension; then 3D key points are obtained and 2D key points K P = H P P, K I = H I E I 。 8. A method for image point cloud registration based on implicit correspondence learning according to claim 1, characterized in that The process of processing the 3D key points and 2D key points based on two pre-established parallel multi-layer perceptron networks to obtain a point cloud registration result is as follows: Based on the pose T of the frustum f is regarded as an initial camera pose to be refined, and the 3D key point K P is transformed accordingly, where R f , t f and the difference between the true camera pose R gt , t gt is as follows: ΔR gt = R gt ·(R f ) -1 , Δt gt = t gt -ΔR gt ·t f . Apply a multi-layer perceptron to perform dimensionality elevation on the correspondence features, and perform pairing through corresponding feature splicing to obtain fused features F f = MLP([D P , MLP(K P ), D I , MLP(K I )]). Inject global information into point-pixel correspondence through the average features of the embedded F f : Pose-sensitive feature f pose Learned from all corresponding information through average pooling operation: f pose = Avgpool(MLP(F s )). Apply two parallel multi-layer perceptrons to respectively regress to obtain ΔR and Δt as the point cloud registration result: ΔR, Δt = ρ(MLP(f pose )), MLP(f pose )。 9. An image point cloud registration system based on implicit correspondence learning, characterized in that It includes: A data processing module, configured to obtain scene point cloud data and image data, extract features from the scene point cloud data and image data to obtain 3D features and 2D features, and process the 3D features and 2D features based on a geometric prior-guided overlapping region detection strategy to obtain a point cloud mask; A point cloud registration module, configured to filter out the points predicted to be outside the overlapping region based on the point cloud mask, so that the randomly initialized learnable correspondence query features only focus on the potential overlapping regions, iteratively apply pixel attention and point attention to the learnable correspondence query features, and after multiple iterations, obtain a 3D key point detector and a 2D key point detector, further interact with the 3D features and 2D features to obtain 3D key points and 2D key points, and process the 3D key points and 2D key points based on two pre-established parallel multi-layer perceptron networks to obtain a point cloud registration result.
10. A terminal device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that A computer program capable of running on a processor is stored in the memory. When the processor loads and executes the computer program, a method for image point cloud registration based on implicit correspondence learning according to any one of claims 1 to 8 is adopted.