A method and system for joint calibration of lidar and camera with spherical spatial alignment

By using a spherical alignment method, the problem of difficult feature alignment between LiDAR and camera was solved, achieving high-precision cross-domain feature fusion and improving the environmental understanding capability in autonomous driving.

CN120088339BActive Publication Date: 2025-11-14WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510197030.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-11-14
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In existing technologies, feature alignment between lidar and cameras is difficult, resulting in poor performance of multi-sensor fusion methods in semantic segmentation and target detection, especially in terms of stability and accuracy under different lighting conditions.

Method used

By employing a spherical space alignment method, multiple spherical spaces are constructed to extract key features from LiDAR point cloud data and camera image data. Linear calibration and feature enhancement are then performed, and combined with geometric alignment and semantic weight alignment, cross-domain feature fusion is achieved.

Benefits of technology

It significantly improves the alignment accuracy between different sensor modalities, enhances the fusion effect of multimodal information, and improves the ability to understand the environment in autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088339B_ABST
    Figure CN120088339B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for joint calibration of LiDAR and camera using spherical space alignment, comprising: acquiring point cloud data from a vehicle-mounted LiDAR and image data from a camera; constructing spherical domains formed by multiple spherical spaces, and extracting key features from the LiDAR point cloud data and camera image data; constructing a perception spherical space from the multiple spherical domains, and accessing different modal features in the multiple spherical domains through linear calibration to determine cross-domain fusion information of the perception spherical space; performing feature enhancement on the cross-domain fusion information of the perception spherical space to obtain enhanced cross-domain perception features; and sequentially performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fused feature information. This invention, by introducing the concept of a perception spherical space, utilizes deep learning alignment of geometric features from the radar and semantic features from the camera in multimodal features, effectively overcoming the difficulty in eliminating the misalignment between radar and camera information in physical methods, and significantly improving the alignment accuracy between different sensor modalities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method and system for joint calibration of lidar and camera with spherical spatial alignment. Background Technology

[0002] Semantic scene understanding is a fundamental task for many applications, such as autonomous driving and robotics. Specifically, in autonomous driving scenarios, it provides fine-grained environmental information for advanced motion planning and enhances the safety of autonomous vehicles. A key task in semantic scene understanding is semantic segmentation, which assigns a class label to each data point in the input data to help autonomous vehicles better understand their environment. Recent research on semantic segmentation methods can be categorized into three types based on the sensors used: pure camera methods, LiDAR-only methods, and multi-sensor fusion methods. Camera-only methods have made significant progress using a large number of open datasets. Because images captured by cameras contain rich appearance information (such as texture and color), camera-only methods can provide fine and accurate semantic segmentation results. However, as passive sensors, cameras are susceptible to changes in lighting conditions and are therefore unreliable.

[0003] To address this issue, researchers performed semantic segmentation of point clouds using LiDAR. Compared to methods using only cameras, LiDAR-only methods are more robust under different lighting conditions; LiDAR provides reliable and accurate spatial depth information of the physical environment. Unfortunately, semantic segmentation using LiDAR alone remains challenging due to the sparse and irregular distribution of point clouds. Furthermore, the lack of texture and color information in point clouds leads to severe classification errors in fine-grained segmentation tasks using pure LiDAR methods. To overcome these shortcomings of camera-only and LiDAR-only methods, a simple solution is to integrate multimodal data from both sensors, i.e., multi-sensor fusion methods. However, data from RGB cameras and LiDAR exist in completely different ways: for example, cameras capture data in perspective views, while LiDAR captures data in 3D views. This significant domain gap between RGB cameras and LiDAR complicates the fusion process. In the field of object detection, effectively representing and processing objects at different scales remains one of the core challenges.

[0004] Given the significant advances in 2D perception, one approach is to project LiDAR point clouds onto 2D images and then process the RGB-D data using a CNN. However, this method of projecting LiDAR data onto a 2D plane introduces severe geometric distortions. Consequently, it performs poorly for geometry-oriented tasks such as 3D point cloud segmentation. Another approach involves point-to-point fusion methods. These methods augment LiDAR point clouds with semantic labels and visual features, then apply existing LiDAR-based segmentation networks to predict the semantic label for each 3D point. However, these point-level fusion methods face challenges in effectively performing semantically oriented tasks. This is due to the inevitable semantic loss that occurs during camera-to-LiDAR projection. In a typical 32-beam LiDAR scanner, only 5% of the camera features are aligned with the LiDAR points, thus most features are ignored. As LiDAR systems have fewer beams, this difference in data density becomes more pronounced, further complicating network learning.

[0005] Therefore, given the limitations of existing technologies, a method is needed that can precisely align camera and LiDAR features. Summary of the Invention

[0006] This invention provides a method and system for joint calibration of lidar and camera with spherical spatial alignment, in order to overcome the deficiencies existing in the prior art.

[0007] In a first aspect, the present invention provides a method for joint calibration of a lidar and camera with spherical spatial alignment, comprising:

[0008] Acquire point cloud data from the vehicle's LiDAR and image data from the camera;

[0009] Construct a spherical domain formed by multiple spherical spaces, and extract key features from the lidar point cloud data and camera image data based on the spherical domain;

[0010] A perceptual sphere space is constructed from multiple spheres. By accessing different modal features in multiple spheres through linear calibration, cross-domain fusion information of the perceptual sphere space is determined.

[0011] The feature enhancement module is used to enhance the cross-domain fusion information of the perception sphere space to obtain enhanced cross-domain perception features.

[0012] The enhanced cross-domain perception features are sequentially geometrically aligned and semantically weighted to obtain fused feature information.

[0013] According to the spherical space-aligned joint calibration method for lidar and camera provided by the present invention, after acquiring the lidar point cloud data and camera image data, the method further includes:

[0014] Based on the transformation equation, the lidar point cloud data is transformed into a 2D-range map containing depth information, and point cloud features are extracted by convolution.

[0015] The camera image data is convolved using a ResNet front-end convolutional neural network to obtain color image features. The P2R equation system is then used to transform the image into a 2D-range graph containing depth information. The formula for transforming the Cartesian coordinate system p3D k = (xk, yk, zk) to the polar coordinate system psph k = (rk, θk, φk) is as follows:

[0016]

[0017]

[0018]

[0019] The formula for transforming from spherical polar coordinates to a 2D-range plot is as follows:

[0020]

[0021]

[0022]

[0023] in, and These are the upper and lower limits of the vertical field of view. and These are the width and height of the final generated image.

[0024] According to the present invention, a joint calibration method for lidar and camera with spherical spatial alignment is provided, which constructs spherical domains formed by multiple spherical spaces, and extracts key features of lidar point cloud data and camera image data based on the spherical domains, including:

[0025] The 2D-range map containing depth information generated from point cloud features and color features is spherically segmented, and the spherical range of each different radius is determined as a spherical domain. The features of the corresponding depth are then projected onto the spherical domain plane.

[0026] The color image features contain semantic information, and the point cloud features contain spatial information.

[0027] According to the present invention, a joint calibration method for lidar and camera with spherical spatial alignment is provided, which enhances the cross-domain fusion information of the perceived spherical space through a feature enhancement module to obtain enhanced cross-domain perception features, including:

[0028] Point cloud features are mapped to a range view through P2R operations to extract features, and a 2D feature extractor is used for downsampling.

[0029] The 2D-range map features are mapped back to the corresponding point cloud using the R2P operation.

[0030] According to the present invention, a joint calibration method for lidar and camera with spherical spatial alignment is provided, wherein the geometric alignment includes:

[0031] A camera calibration matrix is ​​used to establish the correlation between LiDAR point cloud data and camera image data, for those located in... Each point in The corresponding pixel coordinates (u,v) are calculated using the following transformation formula:

[0032]

[0033] in, This represents the camera extrinsic parameter matrix, including rotation and translation components. This represents the camera intrinsic parameter matrix.

[0034] According to the present invention, a joint calibration method for lidar and camera with spherical spatial alignment is provided, wherein the semantic weight alignment includes:

[0035] 3D point depth is estimated using camera features. The 3D point depth is then compared with LiDAR point depth data. The difference analysis is used to evaluate the effect of camera feature assistance, and features that are more similar are assigned higher weights. The fusion process formula is as follows:

[0036]

[0037]

[0038] in, It is lidar scanning data. It is the existing feature corresponding to the current 3D point. It is the activation function in a multilayer perceptron. and These are two weight matrices used to perform further linear transformations on the features. These are features resulting from a single fusion step. These are the features resulting from the second step of fusion.

[0039] Secondly, the present invention also provides a spherically aligned lidar and camera joint calibration system, comprising:

[0040] The acquisition module is used to acquire point cloud data from the vehicle's LiDAR and image data from the camera.

[0041] The extraction module is used to construct a spherical domain formed by multiple spherical spaces, and extract key features from the lidar point cloud data and camera image data based on the spherical domain;

[0042] The determination module is used to construct a perception sphere space from multiple spheres, access different modal features in multiple spheres through linear calibration, and determine the cross-domain fusion information of the perception sphere space.

[0043] The enhancement module is used to enhance the cross-domain fusion information of the perceptual sphere space through the feature enhancement module to obtain enhanced cross-domain perceptual features;

[0044] The fusion module is used to sequentially perform geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fused feature information.

[0045] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the spherical space alignment method for joint calibration of a lidar and a camera as described above.

[0046] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the spherical space-aligned lidar and camera joint calibration method as described above.

[0047] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the spherical space alignment method for joint calibration of lidar and camera as described above.

[0048] The present invention provides a spherical space aligned LiDAR and camera joint calibration method and system. By introducing the concept of perceptual spherical space, it utilizes the geometric features from radar and the semantic features from camera in multimodal features for deep learning alignment. This can effectively overcome the misalignment between radar information and camera information that is difficult to eliminate in physical methods, significantly improve the alignment accuracy between different sensor modalities, and effectively facilitate the application of multimodal information in the autonomous driving industry. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0050] Figure 1This is one of the flowcharts illustrating the spherical space alignment method for joint calibration of lidar and camera provided by the present invention.

[0051] Figure 2 This is the second flowchart illustrating the spherical space alignment method for joint calibration of lidar and camera provided by the present invention.

[0052] Figure 3 This is a schematic diagram of the spherical space alignment lidar and camera joint calibration system provided by the present invention;

[0053] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0055] Figure 1 This is one of the flowcharts illustrating the spherical space-aligned joint calibration method for lidar and camera provided in this embodiment of the invention, such as... Figure 1 As shown, it includes:

[0056] Step 100: Acquire point cloud data from the vehicle's LiDAR and image data from the camera;

[0057] Step 200: Construct a spherical domain formed by multiple spherical spaces, and extract key features from the lidar point cloud data and camera image data based on the spherical domain;

[0058] Step 300: Construct a perceptual sphere space from multiple spheres, access different modal features in multiple spheres through linear calibration, and determine the cross-domain fusion information of the perceptual sphere space;

[0059] Step 400: Enhance the cross-domain fusion information of the perceptual sphere space using the feature enhancement module to obtain enhanced cross-domain perceptual features;

[0060] Step 500: Perform geometric alignment and semantic weight alignment on the enhanced cross-domain perception features in sequence to obtain fused feature information.

[0061] Specifically, such as Figure 2 As shown, it includes the following steps:

[0062] Step 1: Acquire point cloud data from the onboard LiDAR and color image data from the onboard camera.

[0063] Step 2, Multi Spatial Feature Extraction (MFE), involves creating multiple spherical spaces called spheres and extracting key features from each sphere, such as semantic information carried in color photographs and spatial geometric information carried in radar point clouds.

[0064] Step 3 introduces the concept of Perception Sphere Space (PSS), which consists of multiple spherical domains, each preserving the unique data characteristics of its corresponding modality. The PSS calibration design allows points to access features from different modalities through linear calibration and supports parallel access to features of points in adjacent spherical domains.

[0065] Step 4: Feature enhancement is performed using the feature enhancement module. This involves extracting point cloud features and downsampling them within the extent view. It maps point cloud features to the extent view via a P2R operation, extracts features, and downsamples them using a 2D feature extractor. Then, an R2P operation maps the extent view features back to the corresponding point cloud.

[0066] Step 5: Through domain-aware feature alignment, the point cloud data and color image data are corrected geometrically and semantically to generate a unified feature representation that can be applied to downstream tasks.

[0067] Furthermore, the point cloud data obtained in step 1 is transformed into a 2D-range map containing depth information through a set of transformation equations, and then convolution is performed to extract point cloud features for subsequent processing.

[0068] The color image data obtained in step 1 is processed by a pre-processing ResNet convolutional neural network to obtain a feature map, which is then transformed into a 2D-range image containing depth information using a set of transformation equations (P2R). The formula for transforming from Cartesian coordinate system p3D k = (xk, yk, zk) to spherical polar coordinate system psph k = (rk, θk, φk) is as follows:

[0069]

[0070]

[0071]

[0072] The formula for transforming from spherical polar coordinates to a 2D-range map is as follows (discarding the original depth information). )

[0073]

[0074] and These are the upper and lower limits of the vertical field of view. ).

[0075] and These are the width and height of the final generated image.

[0076] Furthermore, in step 2, the 2D-range map containing depth information generated from the point cloud features and color image features is spherically segmented, and each spherical range with a different radius is called a spherical domain. The features of the corresponding depth are then projected onto the spherical domain plane, where the color image features carry semantic information and the point cloud features carry spatial information.

[0077] Furthermore, the Perceptual Sphere Space (PSS) introduced in step 3 facilitates unified multimodal feature representation and fusion by integrating multiple sphere domains. This approach aligns with the physical modeling of sensors and significantly improves the efficiency of feature indexing. The sphere domains unify the modeling methods of different sensors by mapping information from various sensors, thereby achieving cross-domain fusion through PSS.

[0078] Furthermore, the specific implementation process of the feature enhancement module in step 4 is as follows: extracting point cloud features and downsampling in the 2D-range map. It maps the point cloud features to the range view through a P2R operation, extracts the features, and performs downsampling using a 2D feature extractor. Then, it maps the 2D-range map features back to the corresponding point cloud through an R2P operation.

[0079] Furthermore, the geometric alignment in step 5 is specifically implemented as follows: In a multi-sensor system, the encoding is performed in... and The data provides complementary information. To effectively utilize this synergy, we employ a camera calibration matrix to establish the relationship between LiDAR points and image pixels. Specifically, for each point... lie in In the diagram, the corresponding pixel coordinates (u,v) can be calculated using the following transformation formula:

[0080]

[0081] in, This represents the extrinsic parameter matrix of the camera, which includes rotation and translation components. This represents the intrinsic parameter matrix of the camera.

[0082] Furthermore, the semantic weight alignment in step 5 is implemented as follows: First, we use camera features to estimate the depth of 3D points, then compare them with LiDAR depth data, and evaluate the camera feature assistance effect through difference analysis. Features that are closer in size are assigned higher weights. The mathematical formula for the fusion process is as follows:

[0083]

[0084]

[0085] It is scanning data from LiDAR, that is, feature information provided by LiDAR, which is good at expressing geometric depth information (such as the distance and outline of an object). It refers to the existing features corresponding to the current point (a 3D point), such as features previously extracted using other methods or perceptual models. These are features resulting from a single fusion step. These are the features after the second fusion step. A Multilayer Perceptron (MLP) is a simple neural network module that extracts new and more useful information from the input (the concatenated features). The role of the MLP is to "process" and "refine" features. σ is an activation function, typically used to make the output of the data non-linear, helping the model better represent complex features. and These are two weight matrices used to perform further linear transformations on the features.

[0086] This process can extract deeper patterns or features. The multiplication operation refers to residual linking. This avoids information loss after multiple processing steps and enhances the model's expressive power.

[0087] Finally, these weighted strategies are used to integrate the geometric features of LiDAR with semantic information from the camera to generate comprehensive hybrid features for use in downstream tasks, such as scene perception in autonomous driving.

[0088] The spherical space-aligned lidar and camera joint calibration system provided by the present invention will be described below. The spherical space-aligned lidar and camera joint calibration system described below can be referred to in correspondence with the spherical space-aligned lidar and camera joint calibration method described above.

[0089] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include a processor 410, a communication interface 420, a memory 430, and a communication bus 440. The processor 410, communication interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a spherical space-aligned joint calibration method for LiDAR and camera. This method includes: acquiring vehicle-mounted LiDAR point cloud data and camera image data; constructing spherical domains formed by multiple spherical spaces; extracting key features from the LiDAR point cloud data and camera image data based on the spherical domains; constructing a perception spherical space from the multiple spherical domains; accessing different modal features in the multiple spherical domains through linear calibration to determine cross-domain fusion information of the perception spherical space; performing feature enhancement on the cross-domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross-domain perception features; and sequentially performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fused feature information.

[0090] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0091] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the spherical space alignment method for joint calibration of LiDAR and camera provided by the above methods. The method includes: acquiring vehicle-mounted LiDAR point cloud data and camera image data; constructing spherical domains formed by multiple spherical spaces, and extracting key features of the LiDAR point cloud data and camera image data based on the spherical domains; constructing a perception spherical space from the multiple spherical domains, accessing different modal features in the multiple spherical domains through linear calibration, and determining cross-domain fusion information of the perception spherical space; performing feature enhancement on the cross-domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross-domain perception features; and sequentially performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fused feature information.

[0092] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a method for joint calibration of LiDAR and camera with spherical space alignment provided by the methods described above. This method includes: acquiring point cloud data of vehicle-mounted LiDAR and image data of a camera; constructing spherical domains formed by multiple spherical spaces, and extracting key features of the LiDAR point cloud data and camera image data based on the spherical domains; constructing a perception spherical space from the multiple spherical domains, accessing different modal features in the multiple spherical domains through linear calibration, and determining cross-domain fusion information of the perception spherical space; performing feature enhancement on the cross-domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross-domain perception features; and sequentially performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fused feature information.

[0093] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0094] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for joint calibration of lidar and camera with spherical spatial alignment, characterized in that, include: Acquire point cloud data from the vehicle's LiDAR and image data from the camera; Construct a spherical domain formed by multiple spherical spaces, and extract key features from the lidar point cloud data and camera image data based on the spherical domain; A perceptual sphere space is constructed from multiple spheres. By accessing different modal features in multiple spheres through linear calibration, cross-domain fusion information of the perceptual sphere space is determined. The feature enhancement module is used to enhance the cross-domain fusion information of the perception sphere space to obtain enhanced cross-domain perception features. The enhanced cross-domain perception features are sequentially geometrically aligned and semantically weighted to obtain fused feature information; After acquiring the vehicle-mounted LiDAR point cloud data and camera image data, the following is also included: Based on the transformation equation, the lidar point cloud data is transformed into a 2D-range map containing depth information, and point cloud features are extracted by convolution. The camera image data is convolved using a ResNet front-end convolutional neural network to obtain color image features. The P2R equation system is then used to transform the image into a 2D-range map containing depth information. The formula for transforming the Cartesian coordinate system p3D k = (xk, yk, zk) to the polar coordinate system psph k = (rk, θk, φk) is as follows: The formula for transforming from spherical polar coordinates to a 2D-range plot is as follows: in, and These are the upper and lower limits of the vertical field of view. and These are the width and height of the final generated image; The geometric alignment includes: A camera calibration matrix is ​​used to establish the correlation between LiDAR point cloud data and camera image data, for those located in... Each point in The corresponding pixel coordinates (u,v) are calculated using the following transformation formula: in, This represents the camera extrinsic parameter matrix, including rotation and translation components. Represents the camera intrinsic parameter matrix; The semantic weight alignment includes: 3D point depth is estimated using camera features. The 3D point depth is then compared with LiDAR point depth data. The difference analysis is used to evaluate the effect of camera feature assistance, and features that are more similar are assigned higher weights. The fusion process formula is as follows: in, It is lidar scanning data. It is the existing feature corresponding to the current 3D point. It is the activation function in a multilayer perceptron. and These are two weight matrices used to perform further linear transformations on the features. These are features resulting from a single fusion step. These are the features resulting from the second step of fusion.

2. The method for joint calibration of lidar and camera with spherical spatial alignment according to claim 1, characterized in that, Constructing spherical domains formed by multiple spherical spaces, and extracting key features from the lidar point cloud data and camera image data based on the spherical domains, including: The 2D-range map containing depth information generated from point cloud features and color features is spherically segmented, and the spherical range of each different radius is determined as a spherical domain. The features of the corresponding depth are then projected onto the spherical domain plane. The color image features contain semantic information, and the point cloud features contain spatial information.

3. The method for joint calibration of lidar and camera with spherical spatial alignment according to claim 1, characterized in that, The feature enhancement module enhances the cross-domain fusion information of the perceptron space to obtain enhanced cross-domain perceptual features, including: Point cloud features are mapped to a range view through P2R operations to extract features, and a 2D feature extractor is used for downsampling. The 2D-range map features are mapped back to the corresponding point cloud using the R2P operation.

4. A spherically aligned lidar and camera joint calibration system, based on the spherically aligned lidar and camera joint calibration method according to any one of claims 1 to 3, characterized in that, include: The acquisition module is used to acquire point cloud data from the vehicle's LiDAR and image data from the camera. The extraction module is used to construct a spherical domain formed by multiple spherical spaces, and extract key features from the lidar point cloud data and camera image data based on the spherical domain; The determination module is used to construct a perception sphere space from multiple spheres, access different modal features in multiple spheres through linear calibration, and determine the cross-domain fusion information of the perception sphere space. The enhancement module is used to enhance the cross-domain fusion information of the perceptual sphere space through the feature enhancement module to obtain enhanced cross-domain perceptual features; The fusion module is used to sequentially perform geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fused feature information.

5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the spherical space alignment method for joint calibration of lidar and camera as described in any one of claims 1 to 3.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the spherical space alignment method for joint calibration of lidar and camera as described in any one of claims 1 to 3.

7. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the spherical space alignment method for joint calibration of lidar and camera as described in any one of claims 1 to 3.

Citation Information

Patent Citations

  • Large-scale point cloud semantic segmentation method based on superpoint graph

    CN108319957A

  • Space alignment method and system for vehicle-road cooperation

    CN116229713A