Cross-domain repositioning method and device based on camera and laser radar
By fusing RGB images and LiDAR point cloud features using perspective projection and attention mechanisms in robot localization and perception technology, the problems of limited single-modal localization and difficulty in cross-modal matching are solved, achieving efficient and accurate cross-domain relocalization.
Patent Information
- Application Number
- CN202511352926.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-22
- Publication Date
- 2025-11-25
AI Technical Summary
In existing technologies, single-sensor relocation methods have limitations in complex environments, multimodal fusion technology has high computational costs and is difficult to achieve real-time processing, and the feature differences and data alignment problems between different sensor modes have not been effectively solved.
By acquiring RGB images and LiDAR point cloud data, the point cloud data is converted into a projected depth map in the camera coordinate system using perspective projection transformation. Features are extracted using a feature extraction network, and feature fusion is performed using an attention mechanism. Cross-modal relocalization is achieved by matching using cosine similarity and kd-tree index.
It achieves efficient and accurate relocation between camera and LiDAR data in complex environments, reducing computational complexity and improving robustness and real-time processing capabilities.
Smart Images

Figure CN121010748A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of robot localization and perception technology, and in particular to a cross-domain relocalization method and apparatus based on camera and lidar. Background Technology
[0002] In the fields of autonomous driving and robot navigation, achieving accurate relocalization and location recognition is one of the key tasks.
[0003] Traditional methods primarily rely on a single sensor, such as a camera or LiDAR. However, single-modal methods have limitations in complex environments. Camera images are prone to recognition difficulties under changes in lighting and weather conditions, while LiDAR point clouds, although robust to environmental changes, are expensive and computationally complex.
[0004] In recent years, multimodal fusion technology has attracted attention, but existing methods often rely on complex operations such as depth estimation or semantic segmentation, which are computationally expensive and difficult to implement in real time. Furthermore, feature differences and data alignment issues between different sensor modalities are also major challenges in cross-modal location recognition. Therefore, a precise and efficient cross-domain relocalization method is currently lacking. Summary of the Invention
[0005] This invention provides a cross-domain relocation method and apparatus based on camera and lidar, which relates to cross-modal position recognition technology based on camera and lidar (LiDAR), and achieves robust and accurate global positioning by fusing different sensor modalities.
[0006] In a first aspect, the present invention provides a cross-domain relocalization method based on a camera and lidar, comprising:
[0007] S1. Acquire single-frame RGB images and LiDAR point cloud data;
[0008] S2. Convert the lidar point cloud data into a projected depth map in the camera coordinate system through perspective projection transformation;
[0009] S3. The first feature extraction network is used to extract image features from the RGB image data, and the second feature extraction network is used to extract point cloud features from the lidar projection depth map.
[0010] S4. A feature fusion method based on attention mechanism, aligning feature data of RGB image and LiDAR point cloud;
[0011] S5. Based on the aligned feature data, find the target point cloud feature that is most similar to the RGB image feature from the feature data of the global lidar point cloud, and perform relocation based on the position data corresponding to the target point cloud feature.
[0012] Optionally, S2 specifically includes:
[0013] The coordinates of the lidar are transformed to the camera coordinate system using an extrinsic parameter matrix;
[0014] The laser point cloud data in the camera coordinate system is subjected to perspective projection using the camera intrinsic parameter matrix to obtain a two-dimensional projection depth map of the laser point cloud data.
[0015] Optionally, the first feature extraction network in S3 is ResNet-50.
[0016] Optionally, the second feature extraction network in S3 is the SalsaNext network.
[0017] Optionally, S4 specifically includes:
[0018] Global relationships within RGB image features and point cloud features are extracted based on a self-attention mechanism.
[0019] The global relationship is used as input to a cross-modal attention mechanism to extract fusion features from RGB images and LiDAR point clouds.
[0020] Optionally, S5 specifically includes:
[0021] Cosine similarity is used to calculate the matching degree between RGB image features and point cloud features;
[0022] The kd-tree indexing method is used to find the target point cloud features that are most similar to the RGB image features from the feature data of the LiDAR point cloud;
[0023] Obtain the location information of the target point cloud and relocate it based on the location information.
[0024] Optionally, the loss function of the method includes: intra-class consistency loss function, inter-class consistency loss function, and cross-modal contrast loss function.
[0025] Secondly, the present invention provides a cross-domain relocation device based on a camera and lidar, comprising:
[0026] The data acquisition module is used to acquire single-frame RGB images and LiDAR point cloud data;
[0027] The coordinate transformation module is used to convert LiDAR point cloud data into an image-like projection depth map in the camera coordinate system through perspective projection transformation.
[0028] The feature extraction module is used to extract image features from RGB image data using a first feature extraction network and to extract point cloud features from the LiDAR projection depth map using a second feature extraction network.
[0029] The feature fusion module is used for attention-based feature fusion methods to align feature data of RGB images with LiDAR point clouds.
[0030] The matching and positioning module is used to find the target point cloud feature data that is most similar to the RGB image features from the feature data of the lidar point cloud based on the aligned feature data, and to perform repositioning based on the target point cloud feature data.
[0031] Optionally, the coordinate transformation module is specifically used for:
[0032] The coordinates of the lidar are transformed to the camera coordinate system using an extrinsic parameter matrix;
[0033] The laser point cloud data in the camera coordinate system is projected through the camera intrinsic parameter matrix to obtain the projection depth map of the laser point cloud data.
[0034] Optionally, the feature fusion module is specifically used for:
[0035] Global relationships within RGB image features and point cloud features are extracted based on a self-attention mechanism.
[0036] The global relationship is used as input to a cross-modal attention mechanism to extract fusion features from RGB images and LiDAR point clouds.
[0037] This invention utilizes perspective projection transformation to convert LiDAR point cloud data into a projected depth map in the camera coordinate system. Different feature extraction networks are used to extract feature maps from both the RGB image data and the LiDAR projected depth map. Then, an attention-based feature fusion method is employed to align the feature data of the RGB image and the LiDAR point cloud. Finally, the target point cloud feature most similar to the RGB image feature is found from the LiDAR point cloud feature data. Relocalization is then performed based on the position data corresponding to the target point cloud feature. This method leverages perspective projection transformation, deep learning feature extraction, scene attention mechanisms, and efficient matching strategies to achieve efficient relocalization between camera and LiDAR data, solving problems such as limited single-modal localization, difficulty in cross-modal matching, and high computational complexity. Attached Figure Description
[0038] Figure 1 A flowchart illustrating a cross-domain relocation method based on a camera and lidar, provided as an embodiment of the present invention;
[0039] Figure 2 This is a flowchart of a scene attention mechanism provided in an embodiment of the present invention;
[0040] Figure 3 This is a schematic diagram of constructing a cross-modal inter-class feature consistency descriptor provided in an embodiment of the present invention;
[0041] Figure 4 This is a structural diagram of a cross-domain relocation device based on a camera and lidar provided in an embodiment of the present invention. Detailed Implementation
[0042] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0043] Example
[0044] Figure 1 This is a flowchart illustrating a cross-domain relocalization method based on camera and LiDAR, provided in an embodiment of the present invention. The aim is to address the limitations of single-modal localization, the difficulty of cross-modal matching, and high computational complexity. This method achieves efficient relocalization between camera and LiDAR data by utilizing perspective projection transformation, deep learning feature extraction, scene attention mechanisms, and efficient matching strategies.
[0045] Specifically, the steps include the following:
[0046] S1. Acquire single-frame RGB images and LiDAR point cloud data.
[0047] A monocular camera can be used to acquire RGB images, with an image size of H×W. Let the RGB image be I. t ∈R H×W×3 The image is used as query data input to the cross-modal matching system.
[0048] Point cloud data P is acquired using lidar, with each frame containing N 3D points:
[0049] P={(x i ,y i ,z i ,r i )}, i=1,2,…,N
[0050] Where, r i The reflection intensity is used. Pre-stored point cloud data is used to build a global map database for subsequent matching.
[0051] S2. Convert the LiDAR point cloud data into a projected depth map in the camera coordinate system through perspective projection transformation.
[0052] To align RGB and LiDAR data within the same reference frame, this invention employs perspective projection transformation, with the specific steps as follows:
[0053] (1) Coordinate transformation: through the external parameter matrix TL2C Point cloud data P in LiDAR coordinates L The transformation formula for the camera coordinate system is as follows:
[0054] P c =T L2C P L
[0055] Among them, P c For point cloud data in camera coordinates, T L2C This is the LiDAR->camera extrinsic parameter matrix.
[0056] (2) Perspective projection: Using the camera intrinsic parameter K, the lidar point cloud data P in the camera coordinate system is projected. c Perform perspective projection to obtain the projection depth map of the laser point cloud data:
[0057] p i =K·P c
[0058] Where, p i = (u,v,d) represents pixel coordinates, and d represents the depth value.
[0059] The depth map D is obtained after perspective projection. t ∈R H×W Crop point cloud data that exceeds the camera's field of view to ensure that the RGB data is aligned with the LiDAR image.
[0060] This step ensures modal alignment between the RGB image and LiDAR data, enabling them to be matched in the same coordinate system.
[0061] S3. The first feature extraction network is used to extract image features from the RGB image data, and the second feature extraction network is used to extract point cloud features from the lidar projection depth map.
[0062] The first feature extraction network is ResNet-50, and the second feature extraction network is SalsaNext network.
[0063] Specifically, ResNet-50 is used to extract visual features F from RGB image data I. C extract:
[0064] F C =ResNet(I)
[0065] The SalsaNext network is used to extract the feature F of the LiDAR projected depth map D. L :
[0066] F L =SalsaNext(D)
[0067] Furthermore, Batch Normalization is used to normalize the RGB and LiDAR features:
[0068]
[0069] By employing a lightweight network architecture (ResNet-50 + SalsaNext), the computational overhead of the network is reduced. Furthermore, the prototype of this invention uses CUDA for parallel computation, improving the feature extraction speed.
[0070] S4. A feature fusion method based on attention mechanism, aligning feature data of RGB images and LiDAR point clouds.
[0071] Since RGB images and LiDAR point clouds belong to different modalities, relying solely on independent feature extraction networks cannot guarantee effective matching between them within the same feature space. Therefore, this embodiment employs a feature fusion method based on an attention mechanism to achieve feature alignment between RGB images and LiDAR point clouds.
[0072] See Figure 2 First, the global relationships within the RGB image features and point cloud features are extracted based on the self-attention mechanism. The expression for the self-attention mechanism is as follows:
[0073]
[0074] Then, the global relationship is used as input to a cross-modal attention mechanism to extract fusion features from the RGB image and the LiDAR point cloud. The expression for the cross-modal attention mechanism is:
[0075]
[0076] S5. Based on the aligned feature data, find the target point cloud feature that is most similar to the RGB image feature from the feature data of the lidar point cloud, and perform relocation based on the position data corresponding to the target point cloud feature.
[0077] This embodiment employs a cosine similarity-based matching method. By calculating the similarity between RGB image features and all point cloud features in the LiDAR database, the most similar matching result is found. Specifically, it includes the following steps:
[0078] (1) Similarity Calculation: Cosine similarity is used to calculate the matching degree between RGB image features and point cloud features. The similarity calculation formula is as follows:
[0079]
[0080] (2) Database retrieval: The kd-tree indexing method is used to find the target point cloud features most similar to the RGB image features from the feature data of the LiDAR point cloud, with a search complexity of O(logN). This retrieval method has low retrieval complexity, which can optimize the database retrieval speed. Furthermore, CUDA is used to compute similarity in parallel during the retrieval process to improve matching efficiency.
[0081] (3) Obtain the location information of the target point cloud and relocate it based on the location information.
[0082] Based on the above embodiments, the embodiments of the present invention consider multiple loss schemes, including cross-modal contrast, intra-modal consistency, and matched supervision loss, when constructing the loss function of the network, in order to increase the robustness of cross-modal learning.
[0083] See details Figure 3 The loss function in this embodiment of the invention includes:
[0084] Intra-class consistency loss function:
[0085]
[0086] Where M is the batch size. This represents the corresponding positive sample embedded in the same modality. This loss forces the shared attention mechanism to generate stable and consistent embeddings for modal-specific inputs, thereby improving the robustness of downstream cross-modal matching.
[0087] Inter-class consistency loss function:
[0088]
[0089] Among them, τ is the cosine similarity, and τ is the temperature scale. This loss causes the model to generate highly similar embeddings for positive sample pairs (the same scene) and different embeddings for negative sample pairs.
[0090] Cross-modal contrastive loss function:
[0091]
[0092] Given normalized image features and point cloud features Where sim() calculates cosine similarity, τ is the temperature scale, and B is the training epoch size. A multi-loss scheme combining cross-modal comparison, intra-modal consistency, and matching supervision loss improves representation robustness and retrieval accuracy.
[0093] See Figure 4 This invention also provides a cross-domain relocation device based on a camera and lidar, comprising:
[0094] The data acquisition module is used to acquire single-frame RGB images and LiDAR point cloud data;
[0095] The coordinate transformation module is used to convert LiDAR point cloud data into an image-like projection depth map in the camera coordinate system through perspective projection transformation.
[0096] The feature extraction module is used to extract image features from RGB image data using a first feature extraction network and to extract point cloud features from the LiDAR projection depth map using a second feature extraction network.
[0097] The feature fusion module is used for attention-based feature fusion methods to align feature data of RGB images with LiDAR point clouds.
[0098] The matching and positioning module is used to find the target point cloud feature data that is most similar to the RGB image features from the feature data of the global LiDAR point cloud based on the aligned feature data, and to perform repositioning based on the position data corresponding to the target point cloud features.
[0099] Specifically, the coordinate transformation module is used for:
[0100] The coordinates of the lidar are transformed to the camera coordinate system using an extrinsic parameter matrix;
[0101] The laser point cloud data in the camera coordinate system is subjected to perspective projection using the camera intrinsic parameter matrix to obtain a two-dimensional projection depth map of the laser point cloud data.
[0102] The feature fusion module is specifically used for:
[0103] Global relationships within RGB image features and point cloud features are extracted based on a self-attention mechanism.
[0104] The global relationship is used as input to a cross-modal attention mechanism to extract fusion features from RGB images and LiDAR point clouds.
[0105] The matching and positioning module is specifically used for:
[0106] Cosine similarity is used to calculate the matching degree between RGB image features and point cloud features;
[0107] The kd-tree indexing method is used to find the target point cloud features that are most similar to the RGB image features from the feature data of the LiDAR point cloud;
[0108] Obtain the location information of the target point cloud and relocate it based on the location information.
[0109] The cross-domain relocation device based on camera and lidar provided in this embodiment of the invention can execute the cross-domain relocation method based on camera and lidar provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method, which will not be described in detail here.
[0110] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A cross-domain relocalization method based on camera and lidar, characterized in that, include: S1. Acquire single-frame RGB images and LiDAR point cloud data; S2. Convert the lidar point cloud data into a projected depth map in the camera coordinate system through perspective projection transformation; S3. The first feature extraction network is used to extract image features from the RGB image data, and the second feature extraction network is used to extract point cloud features from the lidar projection depth map. S4. A feature fusion method based on attention mechanism, aligning feature data of RGB image and LiDAR point cloud; S5. Based on the aligned feature data, find the target point cloud feature that is most similar to the RGB image feature from the feature data of the global lidar point cloud, and perform relocation based on the position data corresponding to the target point cloud feature.
2. The method according to claim 1, characterized in that, S2 specifically includes: The coordinates of the lidar are transformed to the camera coordinate system using an extrinsic parameter matrix; The laser point cloud data in the camera coordinate system is subjected to perspective projection using the camera intrinsic parameter matrix to obtain a two-dimensional projection depth map of the laser point cloud data.
3. The method according to claim 1, characterized in that, The first feature extraction network in S3 is ResNet-50.
4. The method according to claim 1, characterized in that, The second feature extraction network in S3 is the SalsaNext network.
5. The method according to claim 1, characterized in that, S4 specifically includes: Global relationships within RGB image features and point cloud features are extracted based on a self-attention mechanism. The global relationship is used as input to a cross-modal attention mechanism to extract fusion features from RGB images and LiDAR point clouds.
6. The method according to claim 1, characterized in that, S5 specifically includes: Cosine similarity is used to calculate the matching degree between RGB image features and point cloud features; The kd-tree indexing method is used to find the target point cloud features that are most similar to the RGB image features from the feature data of the LiDAR point cloud; Obtain the location information of the target point cloud and relocate it based on the location information.
7. The method according to claim 1, characterized in that, The loss functions of the method include: intra-class consistency loss function, inter-class consistency loss function, and cross-modal contrast loss function.
8. A cross-domain relocation device based on a camera and lidar, characterized in that, include: The data acquisition module is used to acquire single-frame RGB images and LiDAR point cloud data; The coordinate transformation module is used to convert LiDAR point cloud data into an image-like projection depth map in the camera coordinate system through perspective projection transformation. The feature extraction module is used to extract image features from RGB image data using a first feature extraction network and to extract point cloud features from the LiDAR projection depth map using a second feature extraction network. The feature fusion module is used for attention-based feature fusion methods to align feature data of RGB images with LiDAR point clouds. The matching and positioning module is used to find the target point cloud feature data that is most similar to the RGB image features from the feature data of the global LiDAR point cloud based on the aligned feature data, and to perform repositioning based on the position data corresponding to the target point cloud feature data.
9. The apparatus according to claim 8, characterized in that, The coordinate transformation module is specifically used for: The coordinates of the lidar are transformed to the camera coordinate system using an extrinsic parameter matrix; The laser point cloud data in the camera coordinate system is subjected to perspective projection using the camera intrinsic parameter matrix to obtain a two-dimensional projection depth map of the laser point cloud data.
10. The apparatus according to claim 8, characterized in that, The feature fusion module is specifically used for: Global relationships within RGB image features and point cloud features are extracted based on a self-attention mechanism. The global relationship is used as input to a cross-modal attention mechanism to extract fusion features from RGB images and LiDAR point clouds.