Multi-scale image point cloud cross-modal registration method and system fusing depth information
Patent Information
- Application Number
- CN202610779219.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-06-02
AI Technical Summary
虽然这种方式取得了一定进展,但结构化二维图像域与无序三维几何域之间的固有模态差异很容易会导致模糊或错误匹配,限制了配准的准确性和鲁棒性
1、通过从待测目标的RGB图像中恢复出深度图,提取两个并行的特征流,确保了RGB图像的外观信息和深度图的几何信息被独立处理,从而保留了各自的独特特性。通过将几何上下文深度融合到图像表示中,大大缩小了深度感知的二维图像特征与三维点云特征之间的固有模态差距,结合交错排列的自注意力层与交叉注意力层来增强图像特征与点云特征,并采用将初始匹配关系逐步向下一级的大尺度进行映射的方式进行特征点匹配,提高了图像与点云之间配准的准确性和鲁棒性。
Smart Images

Figure CN122335923B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and system for cross-modal registration of multi-scale image point clouds that integrates depth information. Background Technology
[0002] Geometric registration is a fundamental task in computer vision and robotics, with the core objective of finding the optimal geometric transformation to unify spatial data from different coordinate systems into a common reference frame. This precise data alignment is a crucial prerequisite for downstream visual computing applications, including 3D reconstruction, augmented / virtual reality, and visual localization. Specifically, it enables the fusion of data acquired from multiple devices, providing a more complete and comprehensive representation of the scene. Based on data type, geometric registration can be broadly categorized into three types: image-to-image, point cloud-to-point cloud, and image-to-point cloud registration. The first two are homomodal tasks because they require aligning data from the same domain, a field with extensive research and methods ranging from traditional feature descriptors to modern deep learning-based approaches. In contrast, image-to-point cloud registration requires aligning structured 2D images from a camera with unordered 3D point clouds. The differences in their data structures and dimensions create significant modal gaps, significantly increasing the difficulty of accurate registration.
[0003] In the field of image-to-point cloud registration, the mainstream paradigm adopts a correspondence-based approach. This involves first establishing a reliable cross-modal correspondence, and then solving for the rigid transformation. While this approach has made some progress, the inherent modal differences between the structured 2D image domain and the unordered 3D geometric domain can easily lead to fuzziness or incorrect matching, limiting the accuracy and robustness of the registration. Summary of the Invention
[0004] In view of this, the present invention proposes a method and system for cross-modal registration of multi-scale image point clouds that integrates depth information.
[0005] The technical solution of this invention is implemented as follows: The first aspect of this invention provides a multi-scale image point cloud cross-modal registration method that fuses depth information, comprising: Feature extraction and feature fusion are performed on the RGB image and depth map of the target to be tested to obtain two-dimensional image fusion features; The self-attention module is used to enhance the 2D image fusion features and the 3D point cloud features of the target object, respectively, to obtain corresponding first and second enhanced features. One of the first and second enhanced features is then used alternately as the cross-attention layer. Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features are obtained after processing by the cross-attention module; The top-k algorithm is used to obtain the first correspondence between the image enhancement features and the point cloud enhancement features, and the feature pairs corresponding to the first correspondence are mapped to a preset scale to obtain the corresponding first target features and second target features. Neighborhood point features within a preset range are obtained with the first target features and the second target features as the center, and matching is performed based on the neighborhood point features to obtain the second correspondence. The second correspondence is used to complete the registration of the RGB image and the 3D point cloud of the target under test.
[0006] Based on the above technical solutions, preferably, the step of extracting and fusing features from the RGB image and depth map of the target to obtain two-dimensional image fusion features includes: The RGB image of the target under test is processed using a depth estimation neural network to obtain the corresponding depth map; Feature extraction is performed on the RGB image and the depth map using a 2D backbone network to obtain the corresponding first image features and first depth features. The first image features and the first depth features are fused to obtain two-dimensional image fusion features at different scales.
[0007] Based on the above technical solutions, preferably, the feature fusion of the first image feature and the first depth feature to obtain two-dimensional image fusion features at different scales includes: The first image feature and the first depth feature are flattened, and the first image feature and the first depth feature are spliced and self-attention enhanced based on the length of the flattened feature to obtain the initial image fusion feature; The initial image fusion features are summed using the feature length, and the summation result is concatenated with the first image features and the first depth features to obtain two-dimensional image fusion features.
[0008] Based on the above technical solutions, preferably, the alternating use of one of the first enhancement feature and the second enhancement feature as the cross-attention layer... Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features processed by the cross-attention module are obtained, including: Using the first enhanced feature as the cross-attention layer The second enhanced feature serves as the cross-attention layer. and The image enhancement features are obtained after processing by the cross-attention module; Using the second enhanced feature as the cross-attention layer The first enhanced feature serves as the cross-attention layer. and The point cloud enhanced features are obtained after processing by the cross-attention module.
[0009] Based on the above technical solutions, preferably, the step of mapping the feature pairs corresponding to the first correspondence to a preset scale to obtain the corresponding first target feature and second target feature, obtaining neighborhood point features within a preset range centered on the first target feature and the second target feature respectively, and performing matching based on the neighborhood point features to obtain the second correspondence includes: Map the feature pairs corresponding to the first correspondence to the first preset scale to obtain the corresponding third target feature and fourth target feature; The first neighborhood point features within a preset range are obtained, centered on the third target feature and the fourth target feature respectively, and matching is performed based on the first neighborhood point features to obtain an initial correspondence. The feature pairs corresponding to the initial correspondence are mapped to a second preset scale to obtain the corresponding first target feature and second target feature. Then, the second neighborhood point features within a preset range are obtained with the first target feature and the second target feature as the center, and the matching is performed based on the second neighborhood point features to obtain the second correspondence. The second preset scale is larger than the first preset scale.
[0010] Based on the above technical solutions, preferably, the step of matching based on the features of the neighboring points to obtain the second correspondence includes: Using a spatial mask, the features of neighboring points within a two-dimensional search window centered on the first target feature are matched with the features of neighboring points in the corresponding three-dimensional neighborhood to obtain a second correspondence.
[0011] Based on the above technical solutions, preferably, the registration of the RGB image and the 3D point cloud of the target under test using the second correspondence includes: The second correspondence is solved using the PnP-RANSAC algorithm to obtain the transformation matrix, and the RGB image and 3D point cloud of the target under test are registered using the transformation matrix.
[0012] More preferably, a second aspect of the present invention provides a multi-scale image point cloud cross-modal registration system that integrates depth information, comprising: a feature processing module, a feature enhancement module, a feature matching module, and a data registration module; wherein, The feature processing module is configured to extract and fuse features from the RGB image and depth map of the target to obtain two-dimensional image fusion features. The feature enhancement module is configured to use a self-attention module to enhance the two-dimensional image fusion features and the three-dimensional point cloud features of the target object respectively, to obtain corresponding first enhanced features and second enhanced features, and to alternately use one of the first enhanced features and the second enhanced features as the cross-attention layer. Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features are obtained after processing by the cross-attention module; The feature matching module is configured to use the top-k algorithm to obtain a first correspondence between the image enhancement features and the point cloud enhancement features, and to map the feature pairs corresponding to the first correspondence to a preset scale to obtain the corresponding first target features and second target features. It then obtains neighborhood point features within a preset range centered on the first target features and the second target features, and performs matching based on the neighborhood point features to obtain a second correspondence. The data registration module is configured to use the second correspondence to complete the registration of the RGB image and the three-dimensional point cloud of the target under test.
[0013] More preferably, a third aspect of the present invention provides an electronic device, including a processor and a memory; the memory has a computer program stored thereon, wherein the computer program, when executed by the processor, implements the multi-scale image point cloud cross-modal registration method with fused depth information as described in the first aspect.
[0014] More preferably, a fourth aspect of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the multi-scale image point cloud cross-modal registration method with fused depth information as described in the first aspect.
[0015] The multi-scale image point cloud cross-modal registration method and system that integrates depth information of the present invention has the following advantages over the prior art: 1. By recovering the depth map from the RGB image of the target object and extracting two parallel feature streams, the appearance information of the RGB image and the geometric information of the depth map are processed independently, thus preserving their unique characteristics. By fusing the geometric context depth into the image representation, the inherent modal gap between the depth-aware 2D image features and 3D point cloud features is greatly reduced. Alternating self-attention layers and cross-attention layers are used to enhance image and point cloud features. Feature point matching is performed by progressively mapping the initial matching relationship to the next larger scale, improving the accuracy and robustness of image-point cloud registration.
[0016] 2. By mapping the initial matching relationship to the next level and defining the local search neighborhood with higher resolution, the entire matching process is strictly limited to regions with relatively higher matching probabilities. This phased process, advancing layer by layer, allows the density and accuracy of the corresponding point set to gradually increase as the feature scale increases, ultimately resulting in a dense set of fine-grained matching relationships. This improves both registration efficiency and registration accuracy.
[0017] 3. By calculating the similarity between image enhancement features and point cloud enhancement features, the top-k algorithm is used to select the most likely matching feature pairs, eliminating a large number of low-confidence matches and reducing the cardinality of false matches. Coarse-matching feature pairs are mapped to a preset scale, eliminating the scale inconsistency problem caused by modal differences and providing standardized input for subsequent fine-matching. Using the coarse-matching feature pairs as the center, neighborhood point features are searched within a preset range. Secondary verification is performed using local geometric consistency, effectively filtering isolated false matches caused by modal differences. Furthermore, spatial context information is introduced through neighborhood constraints, enhancing the robustness of the matching. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating a multi-scale image point cloud cross-modal registration method that integrates depth information, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a multi-scale image point cloud cross-modal registration system that fuses depth information, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0020] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0021] In some embodiments, such as Figure 1 As shown, Figure 1This is a flowchart illustrating a multi-scale image point cloud cross-modal registration method that fuses depth information, provided by an embodiment of the present invention. The multi-scale image point cloud cross-modal registration method that fuses depth information provided by the present invention includes: S110 performs feature extraction and feature fusion on the RGB image and depth map of the target to be tested, and obtains two-dimensional image fusion features.
[0022] S120: The self-attention module is used to enhance the fusion features of the two-dimensional image and the three-dimensional point cloud features of the target object, respectively, to obtain the corresponding first enhanced feature and second enhanced feature. One of the first enhanced feature and the second enhanced feature is used alternately as the cross-attention layer. Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features are obtained after processing by the cross-attention module.
[0023] S130, the top-k algorithm is used to obtain the first correspondence between image enhancement features and point cloud enhancement features, and the feature pairs corresponding to the first correspondence are mapped to a preset scale to obtain the corresponding first target features and second target features. The neighborhood point features within a preset range are obtained with the first target features and the second target features as the center, and the matching is performed based on the neighborhood point features to obtain the second correspondence.
[0024] S140, using the second correspondence to complete the registration of the RGB image and 3D point cloud of the target under test.
[0025] In this embodiment, the RGB image contains pixel grid and texture information of the target object, but lacks depth information; the depth map can provide geometric distance but has low resolution; the 3D point cloud is disordered and sparse, but contains complete geometric structure. Using self-attention on the fusion features of the 2D image can capture long-range dependencies in the image, such as global texture consistency, thereby enhancing the expression of structured information. Using self-attention on the 3D point cloud features can mine the contextual associations of local geometric structures in the point cloud, improving data organization. By alternately using image enhancement features or point cloud enhancement features as queries, and another feature as a key and value, dynamic alignment of features between modalities is achieved. For example, image enhancement features can be used to guide point cloud enhancement features to focus on geometric regions that match the texture, thereby reducing modal bias during direct matching.
[0026] The top-k algorithm is used to calculate the similarity between image enhancement features and point cloud enhancement features, filtering out the most likely matching feature pairs to eliminate a large number of low-confidence matches. Based on this, neighborhood point features are searched within a preset range centered on the coarsely matched feature pairs, and secondary verification is performed using local geometric consistency to filter out isolated false matches caused by modal differences. By selecting precise feature point pairs, the registration accuracy and reliability between different modalities are ensured.
[0027] In some embodiments, feature extraction and feature fusion are performed on the RGB image and depth map of the target to be tested to obtain two-dimensional image fusion features, including: A depth estimation neural network is used to process the RGB image of the target to be tested, and the corresponding depth map is obtained. Feature extraction is performed on the RGB image and depth map using a 2D backbone network to obtain the corresponding first image features and first depth features. Feature fusion is performed on the first image features and the first depth features to obtain two-dimensional image fusion features at different scales.
[0028] In this embodiment, by using clues such as texture, shadows, and relative positions between objects in the RGB image, a depth estimation neural network, such as MiDaS (Monocular Depth Estimation via a Multi-Scale Deep Network) or DPT (Dense Prediction Transformer), is used to learn the mapping relationship from pixel values to depth values, thereby obtaining the corresponding depth map. The 2D backbone network may include an RGB branch and a depth branch, which are used to extract features from the RGB image and the depth map respectively.
[0029] In some embodiments, feature fusion is performed on the first image features and the first depth features to obtain two-dimensional image fusion features at different scales, including: The first image features and the first depth features are flattened, and the first image features and the first depth features are spliced and self-attention enhanced based on the length of the flattened features to obtain the initial image fusion features; The initial image fusion features are summed using the feature length, and the summation result is concatenated with the first image features and the first depth features to obtain the two-dimensional image fusion features.
[0030] In this embodiment, the first image features are obtained through a 2D backbone network. and first depth features (Where i=1, 2, 3, 4, the feature scale gradually increases from i=1 to i=4) Then, the feature fusion module is used to perform same-modal feature fusion. and Flattening makes ,in, Represents the feature length after flattening. Represents the number of channels. Along the feature length Will and Feature concatenation is performed, and the concatenated features are input into a multi-head self-attention module to obtain initial image fusion features. The initial image fusion features are then processed along the feature length... Perform summation and compare with the first image features. and first depth features Along the passage By stitching the images together, we obtain the 2D image fusion features. :
[0031] ; in, Represents the summation operation. Representative feature splicing operation, Multi-head self-attention module.
[0032] In some embodiments, one of the first enhancement feature and the second enhancement feature is used alternately as the cross-attention layer. Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features processed by the cross-attention module are obtained, including: Using the first enhanced feature as the cross-attention layer The second enhancement feature serves as the cross-attention layer. and The image enhancement features are obtained after processing by the cross-attention module; Using the second enhanced feature as the cross-attention layer The first enhanced feature serves as the cross-attention layer. and The point cloud enhanced features are obtained after processing by the cross-attention module.
[0033] In this embodiment, to reduce the inherent differences between the first and second enhancement feature modalities, an attention mechanism is used to fuse them across modalities to obtain the image enhancement features. and point cloud augmentation features : ; in, This represents an attention-based fusion module, implemented through alternating self-attention layers and cross-attention layers. Specifically, it involves fusion features of two-dimensional images. and 3D point cloud features The features are first obtained through a self-attention layer. and characteristics ,Will As a cross-attention layer , As a cross-attention layer and Image enhancement features are obtained. Similarly, As a cross-attention layer , As a cross-attention layer and Point cloud augmentation features were obtained. .
[0034] In some embodiments, feature pairs corresponding to the first correspondence are mapped to a preset scale to obtain corresponding first target features and second target features. Neighborhood point features within a preset range are obtained, centered on the first target features and the second target features, respectively. Matching is performed based on the neighborhood point features to obtain the second correspondence, including: Map the feature pairs corresponding to the first correspondence to the first preset scale to obtain the corresponding third target features and fourth target features; The first neighborhood point features within a preset range are obtained, centered on the third target feature and the fourth target feature respectively, and the initial correspondence is obtained by matching based on the first neighborhood point features; The feature pairs corresponding to the initial correspondence are mapped to the second preset scale to obtain the corresponding first target feature and second target feature. The second neighborhood point features within a preset range are obtained with the first target feature and the second target feature as the center, and the matching is performed based on the second neighborhood point features to obtain the second correspondence. The second preset scale is greater than the first preset scale.
[0035] In this embodiment, the top-k algorithm is used to obtain the first correspondence between image enhancement features and point cloud enhancement features. : ; Based on all Local search windows are defined on both the image and the point cloud, and feature matching is performed within each local search window. For the image, all coarse matching relationships are... Mapping back to larger-scale features The corresponding mapped coordinates are obtained from the above. :
[0036] ; in, This represents a two-dimensional mapping function. Set the radius to the center. A local search window.
[0037] In terms of point clouds, all coarse matching relationships are... Mapping back to larger-scale features The corresponding mapping region is obtained above. : ; in, This represents a three-dimensional mapping function.
[0038] Map the point features within the local search window to the mapped region. Matching the point features within the image yields a more refined correspondence between the point clouds. : ; Features at higher scales and ,as well as and Performing the same operation yields the most detailed image point cloud correspondence. .
[0039] In some embodiments, matching is performed based on neighborhood point features to obtain a second correspondence, including: Using spatial masks, the features of neighboring points within a two-dimensional search window centered on the first target feature are matched with the features of neighboring points in the corresponding three-dimensional neighborhood to obtain a second correspondence.
[0040] Points that are spatially close to each other are more geometrically or semantically related. By limiting the search range through spatial masks, interference from irrelevant features can be reduced, thus lowering computational complexity. When matching 2D image features with 3D point cloud features, spatial masks can force matching to occur within locally corresponding regions, avoiding mismatches caused by global searches.
[0041] In some embodiments, the registration of the RGB image and the 3D point cloud of the target under test is performed using the second correspondence, including: The second correspondence is solved using the PnP-RANSAC algorithm to obtain the transformation matrix, and the RGB image and 3D point cloud of the target under test are registered using the transformation matrix.
[0042] The Perspective-n-Point (PnP) algorithm estimates the camera pose using known 3D-2D point correspondences. Mathematically, it solves for the transformation matrix in the camera projection model, minimizing the error between the transformed 3D point's projected position on the image plane and the actual 2D point's position. The Random Sample Consensus (RANSAC) algorithm iteratively verifies the number of inliers by randomly sampling a minimum set of points, ultimately selecting the transformation matrix corresponding to the model with the most inliers. This approach combines the advantages of both algorithms, using the transformation matrix to transform the 3D point cloud from its own coordinate system to the camera coordinate system, or the 2D image from its own coordinate system to the 3D coordinate system, achieving spatial alignment and eliminating misalignment problems caused by modal differences.
[0043] In some embodiments, please refer to Figure 2 , Figure 2 This is a schematic diagram of the structure of a multi-scale image point cloud cross-modal registration system that fuses depth information, provided in an embodiment of the present invention. The present invention provides a multi-scale image point cloud cross-modal registration system 200 that fuses depth information, comprising: a feature processing module 210, a feature enhancement module 220, a feature matching module 230, and a data registration module 240; wherein,
[0044] The feature processing module 210 is configured to extract and fuse features from the RGB image and depth map of the target to be tested, and obtain two-dimensional image fusion features. Feature enhancement module 220 is configured to use a self-attention module to enhance the fusion features of the two-dimensional image and the three-dimensional point cloud features of the target object, respectively, to obtain corresponding first enhanced features and second enhanced features, and to alternately use one of the first enhanced features and the second enhanced features as the cross-attention layer. Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features are obtained after processing by the cross-attention module; The feature matching module 230 is configured to use the top-k algorithm to obtain the first correspondence between image enhancement features and point cloud enhancement features, and map the feature pairs corresponding to the first correspondence to a preset scale to obtain the corresponding first target feature and second target feature. It then obtains neighborhood point features within a preset range centered on the first target feature and the second target feature, and performs matching based on the neighborhood point features to obtain the second correspondence. The data registration module 240 is configured to use the second correspondence to complete the registration of the RGB image and the 3D point cloud of the target under test.
[0045] In some embodiments, the feature processing module 210 is specifically configured as follows: A depth estimation neural network is used to process the RGB image of the target to be tested, and the corresponding depth map is obtained. Feature extraction is performed on the RGB image and depth map using a 2D backbone network to obtain the corresponding first image features and first depth features. Feature fusion is performed on the first image features and the first depth features to obtain two-dimensional image fusion features at different scales.
[0046] In some embodiments, the feature processing module 210 is specifically configured as follows: The first image features and the first depth features are flattened, and the first image features and the first depth features are spliced and self-attention enhanced based on the length of the flattened features to obtain the initial image fusion features; The initial image fusion features are summed using the feature length, and the summation result is concatenated with the first image features and the first depth features to obtain the two-dimensional image fusion features.
[0047] In some embodiments, the feature enhancement module 220 is specifically configured as follows: Using the first enhanced feature as the cross-attention layer The second enhancement feature serves as the cross-attention layer. and The image enhancement features are obtained after processing by the cross-attention module; Using the second enhanced feature as the cross-attention layer The first enhanced feature serves as the cross-attention layer. and The point cloud enhanced features are obtained after processing by the cross-attention module.
[0048] In some embodiments, the feature matching module 230 is specifically configured as follows: Map the feature pairs corresponding to the first correspondence to the first preset scale to obtain the corresponding third target features and fourth target features; The first neighborhood point features within a preset range are obtained, centered on the third target feature and the fourth target feature respectively, and the initial correspondence is obtained by matching based on the first neighborhood point features; The feature pairs corresponding to the initial correspondence are mapped to the second preset scale to obtain the corresponding first target feature and second target feature. The second neighborhood point features within a preset range are obtained with the first target feature and the second target feature as the center, and the matching is performed based on the second neighborhood point features to obtain the second correspondence. The second preset scale is greater than the first preset scale.
[0049] In some embodiments, the feature matching module 230 is specifically configured as follows: Using spatial masks, the features of neighboring points within a two-dimensional search window centered on the first target feature are matched with the features of neighboring points in the corresponding three-dimensional neighborhood to obtain a second correspondence.
[0050] In some embodiments, the data registration module 240 is specifically configured as follows: The second correspondence is solved using the PnP-RANSAC algorithm to obtain the transformation matrix, and the RGB image and 3D point cloud of the target under test are registered using the transformation matrix.
[0051] It should be noted that the multi-scale image point cloud cross-modal registration system with fused depth information provided in this application embodiment and the multi-scale image point cloud cross-modal registration method with fused depth information provided in this application embodiment are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned multi-scale image point cloud cross-modal registration method with fused depth information, and the repeated parts will not be described again.
[0052] In some embodiments, please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device 300 provided in this embodiment includes a processor 310 and a memory 320; the memory 320 stores a computer program, wherein the computer program, when executed by the processor, implements the aforementioned multi-scale image point cloud cross-modal registration method that fuses depth information.
[0053] Specifically, processor 310 may include, for example, a general-purpose microprocessor, an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. Processor 310 may also include onboard memory for caching purposes. Processor 310 may be a single processing unit or multiple processing units for performing different actions of the method flow according to embodiments of this application.
[0054] The memory 320 may be any medium capable of containing, storing, transmitting, propagating, or transmitting instructions. For example, the memory 320 may include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, apparatuses, or propagation media. Specific examples of the memory 320 include: magnetic storage devices such as magnetic tape or hard disk drives (HDDs); optical storage devices such as optical discs (CD-ROMs); and may also be random access memory (RAM) or flash memory; and / or wired / wireless communication links.
[0055] This application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, this program implements the aforementioned multi-scale image point cloud cross-modal registration method with fused depth information. This computer-readable medium may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into that device / apparatus / system. The aforementioned computer-readable medium carries one or more programs, which, when executed, implement the method according to the embodiments of this application.
[0056] According to embodiments of this application, a computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wired, optical fiber, radio frequency signals, etc., or any suitable combination thereof.
[0057] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in this application. In particular, the features described in the various embodiments and / or claims of this application can be combined and / or combined in various ways without departing from the spirit and teachings of this application. All such combinations and / or combinations fall within the scope of this application. Therefore, the scope of this application should not be limited to the above embodiments, but should be defined not only by the appended claims, but also by their equivalents. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this invention should be included within the protection scope of this invention.
Claims
1. A method for multi-scale image point cloud cross-modality registration with fused depth information, characterized in that, include: Feature extraction and feature fusion are performed on the RGB image and depth map of the target to be tested to obtain two-dimensional image fusion features; The self-attention module is used to enhance the 2D image fusion features and the 3D point cloud features of the target object, respectively, to obtain corresponding first and second enhanced features. One of the first and second enhanced features is then used alternately as the cross-attention layer. Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features are obtained after processing by the cross-attention module; The method involves: using the top-k algorithm to obtain a first correspondence between the image enhancement features and the point cloud enhancement features; mapping the feature pairs corresponding to the first correspondence to a preset scale to obtain corresponding first target features and second target features; obtaining neighborhood point features within a preset range centered on the first target features and the second target features; and performing matching based on the neighborhood point features to obtain a second correspondence. This includes: mapping the feature pairs corresponding to the first correspondence to a first preset scale to obtain corresponding third target features and fourth target features; obtaining first neighborhood point features within a preset range centered on the third target features and the fourth target features; performing matching based on the first neighborhood point features to obtain an initial correspondence; mapping the feature pairs corresponding to the initial correspondence to a second preset scale to obtain corresponding first target features and second target features; and using spatial masks to match the neighborhood point features within a two-dimensional search window centered on the first target feature and the neighborhood point features within a three-dimensional neighborhood centered on the second target feature to obtain a second correspondence. The second preset scale is larger than the first preset scale. The second correspondence is used to complete the registration of the RGB image and the 3D point cloud of the target under test.
2. The multi-scale image point cloud cross-modal registration method with fused depth information as described in claim 1, characterized in that, The RGB image and depth map of the target under test are subjected to feature extraction and feature fusion to obtain two-dimensional image fusion features, including: The RGB image of the target under test is processed using a depth estimation neural network to obtain the corresponding depth map; Feature extraction is performed on the RGB image and the depth map using a 2D backbone network to obtain the corresponding first image features and first depth features. The first image features and the first depth features are fused to obtain two-dimensional image fusion features at different scales.
3. The multi-scale image point cloud cross-modal registration method with fused depth information as described in claim 2, characterized in that, The feature fusion of the first image features and the first depth features to obtain two-dimensional image fusion features at different scales includes: The first image feature and the first depth feature are flattened, and the first image feature and the first depth feature are spliced and self-attention enhanced based on the length of the flattened feature to obtain the initial image fusion feature; The initial image fusion features are summed using the feature length, and the summation result is concatenated with the first image features and the first depth features to obtain two-dimensional image fusion features.
4. The multi-scale image point cloud cross-modality registration method fusing depth information as claimed in claim 1, wherein, The alternating use of one of the first enhancement feature and the second enhancement feature as the cross-attention layer Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features processed by the cross-attention module are obtained, including: using the first enhanced feature as an input of a cross-attention layer , using the second enhanced feature as an input of a cross-attention layer and to obtain an image enhanced feature processed by the cross-attention module Using the second enhanced feature as the cross-attention layer The first enhanced feature serves as the cross-attention layer. and The point cloud enhanced features are obtained after processing by the cross-attention module.
5. The multi-scale image point cloud cross-modal registration method with fused depth information as described in claim 1, characterized in that, The registration of the RGB image and 3D point cloud of the target under test using the second correspondence includes: The second correspondence is solved using the PnP-RANSAC algorithm to obtain the transformation matrix, and the RGB image and 3D point cloud of the target under test are registered using the transformation matrix.
6. A system for multi-modal registration of multi-scale image point clouds fused with depth information, the system comprising: include: The module comprises a feature processing module, a feature enhancement module, a feature matching module, and a data registration module; among which, The feature processing module is configured to extract and fuse features from the RGB image and depth map of the target to obtain two-dimensional image fusion features. The feature enhancement module is configured to use a self-attention module to enhance the two-dimensional image fusion features and the three-dimensional point cloud features of the target object respectively, to obtain corresponding first enhanced features and second enhanced features, and to alternately use one of the first enhanced features and the second enhanced features as the cross-attention layer. Another feature serves as the cross-attention layer. and The image enhancement features and point cloud enhancement features are obtained after processing by the cross-attention module; The feature matching module is configured to use a top-k algorithm to obtain a first correspondence between the image enhancement features and the point cloud enhancement features, and map the feature pairs corresponding to the first correspondence to a preset scale to obtain corresponding first target features and second target features. It then obtains neighborhood point features within a preset range centered on the first target features and the second target features, and performs matching based on the neighborhood point features to obtain a second correspondence. This includes: mapping the feature pairs corresponding to the first correspondence to a first preset scale to obtain corresponding third target features and fourth target features; obtaining first neighborhood point features within a preset range centered on the third target features and the fourth target features, and performing matching based on the first neighborhood point features to obtain an initial correspondence; mapping the feature pairs corresponding to the initial correspondence to a second preset scale to obtain corresponding first target features and second target features, and using spatial masks to match neighborhood point features within a two-dimensional search window centered on the first target feature and neighborhood point features within a three-dimensional neighborhood centered on the second target feature to obtain a second correspondence; the second preset scale is larger than the first preset scale. The data registration module is configured to use the second correspondence to complete the registration of the RGB image and the three-dimensional point cloud of the target under test.
7. An electronic device comprising a processor and a memory; the memory storing a computer program, wherein, When the computer program is executed by the processor, it implements the multi-scale image point cloud cross-modal registration method that fuses depth information as described in any one of claims 1 to 5.
8. A non-transitory computer-readable storage medium, comprising: It stores a computer program, wherein the computer program, when executed by a processor, implements the multi-scale image point cloud cross-modal registration method that fuses depth information as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Image point cloud registration method based on self-calibration attention mechanism
CN117934570A
Progressive optimization method for multi-modal point cloud registration
CN121564049A