Spherical space alignment laser radar and camera combined calibration method and system

Through the joint calibration method of spherical space alignment lidar and camera, the problem of low alignment accuracy of camera and LiDAR feature in the prior art is solved, and more efficient multimodal data fusion is achieved, and the application effect in the autonomous driving industry is improved.

CN120088339AActive Publication Date: 2025-06-03WUHAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510197030.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-03
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art is difficult to finely align camera and LiDAR features, resulting in problems such as semantic loss and low alignment accuracy during multimodal data fusion.

Method used

Using a joint calibration method of spherical space alignment, a spherical domain formed by multiple spherical spaces is constructed by acquiring the on-board lidar point cloud data and camera image data, key features are extracted, and cross-domain fusion is carried out through the concept of perceived spherical space. The feature enhancement module is used to enhance the fusion information, and combine geometric alignment and semantic weight alignment to generate fusion feature information.

Benefits of technology

It significantly improves the alignment accuracy between different sensor modes, effectively overcomes the difficult misalignment of radar information and camera information in physical methods, and helps the application of multimodal information in the autonomous driving industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088339A_ABST
    Figure CN120088339A_ABST
Patent Text Reader

Abstract

The invention provides a laser radar and camera combined calibration method and system for spherical space alignment. The method comprises the following steps: acquiring vehicle-mounted laser radar point cloud data and camera image data; constructing a sphere domain formed by a plurality of spherical spaces, and extracting key features of laser radar point cloud data and camera image data; constructing a sensing ball space by the plurality of ball domains, accessing different modal features in the plurality of ball domains through linear calibration, and determining cross-domain fusion information of the sensing ball space; carrying out feature enhancement on the spatial cross-domain fusion information of the sensing ball to obtain enhanced cross-domain sensing features; geometric alignment and semantic weight alignment are carried out on the enhanced cross-domain perception features in sequence, and fusion feature information is obtained. According to the method, the perception sphere space concept is introduced, deep learning alignment is carried out by using the geometric features from the radar and the semantic features from the camera in the multi-modal features, the dislocation that radar information and camera information are difficult to eliminate in a physical method is effectively overcome, and the alignment precision between different sensor modals is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for jointly calibrating a lidar and a camera with spherical space alignment. Background Art

[0002] Semantic scene understanding is a fundamental task for many applications such as autonomous driving and robotics. Specifically, in the autonomous driving scenario, it provides fine-grained environmental information for high-level motion planning and enhances the safety of autonomous vehicles. A key task in semantic scene understanding is semantic segmentation, which assigns a class label to each data point in the input data to help autonomous vehicles better understand their environment. According to the sensors used in semantic segmentation methods, recent research can be classified into three categories: pure camera methods, LiDAR-only methods, and multi-sensor fusion methods. Using a large number of open datasets, pure camera methods have made significant progress. Since the images captured by cameras have rich appearance information (such as texture and color), pure camera methods can provide fine and accurate semantic segmentation results. However, as passive sensors, cameras are vulnerable to changes in lighting conditions and thus unreliable.

[0003] To solve this problem, researchers perform semantic segmentation on point clouds from LiDAR. Compared with methods that only use cameras, LiDAR-only methods are more robust under different lighting conditions; LiDAR provides reliable and accurate spatial depth information of the physical environment. Unfortunately, due to the sparse and irregular distribution of point clouds, semantic segmentation using only LiDAR is still challenging. In addition, the lack of texture and color information in point clouds leads to serious classification errors in the fine-grained segmentation tasks of pure LiDAR methods. To address these drawbacks of pure camera and LiDAR-only methods, a simple solution is to integrate multi-modal data from the two sensors, i.e., the multi-sensor fusion method. However, data from RGB cameras and LiDAR exist in completely different ways: for example, cameras capture data in a perspective view, while LiDAR captures data in a three-dimensional view. This significant domain gap between RGB cameras and LiDAR complicates the fusion process. In the field of object detection, effectively representing and processing objects at different scales has always been one of the core challenges.

[0004] Given the significant progress in 2D perception, it is considered to project LiDAR point clouds onto 2D images and then use CNN to process RGB-D data. However, this method of projecting LiDAR data onto a 2D plane introduces serious geometric deformations. Therefore, it performs poorly for geometry-oriented tasks such as 3D point cloud segmentation. Another approach involves point-to-point fusion methods. These methods enhance LiDAR point clouds through semantic labels and visual features, and then apply existing LiDAR-based segmentation networks to predict the semantic labels of each 3D point. However, these point-level fusion methods face challenges in effectively performing semantic-oriented tasks. This is due to the inevitable semantic loss during the camera-to-LiDAR projection process. In a typical 32-beam LiDAR scanner, only 5% of the camera features are aligned with LiDAR points, so most features are ignored. As LiDAR systems have fewer lines, this difference in data density becomes more obvious, further complicating network learning.

[0005] Therefore, in view of the limitations of the prior art, a method capable of finely aligning camera and LiDAR features is needed. Summary of the Invention

[0006] The present invention provides a method and system for joint calibration of a LiDAR and a camera with spherical space alignment to solve the defects existing in the prior art.

[0007] In a first aspect, the present invention provides a method for joint calibration of a LiDAR and a camera with spherical space alignment, including: Obtaining LiDAR point cloud data and camera image data of a vehicle; Constructing a spherical domain formed by multiple spherical spaces, and extracting key features of the LiDAR point cloud data and camera image data based on the spherical domain; Constructing a perception spherical space from multiple spherical domains, accessing different modality features in multiple spherical domains through linear calibration, and determining cross-domain fusion information of the perception spherical space; Enhancing the cross-domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross-domain perception features; Performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features in sequence to obtain fusion feature information.

[0008] According to the method for joint calibration of a LiDAR and a camera with spherical space alignment provided by the present invention, after obtaining the LiDAR point cloud data and camera image data of a vehicle, it further includes: Converting the LiDAR point cloud data into a 2D-range map containing depth information based on a transformation equation, and extracting point cloud features through convolution; The pre - installed convolutional neural network ResNet is used to perform convolutional processing on the camera image data to obtain color image features, and perform P2R conversion on the equations to a 2D - range map containing depth information. The formula for converting from the rectangular coordinate system p3D k = (xk, yk, zk) to the spherical polar coordinate system psph k = (rk, θk, φk) is as follows:

[0009]

[0010]

[0011] The formula for converting from the spherical polar coordinate system to the 2D - range map is as follows:

[0012]

[0013]

[0014] Among them, and are the upper and lower limits of the vertical field of view, and are the width and height of the finally generated image.

[0015] According to a method for jointly calibrating a lidar and a camera with spherical space alignment provided by the present invention, a spherical domain formed by multiple spherical spaces is constructed, and key features of the lidar point cloud data and the camera image data are extracted based on the spherical domain, including: Performing spherical segmentation on the 2D - range map containing depth information generated from the point cloud feature and the color feature, determining the spherical range of each different radius as the spherical domain, and projecting the features of the corresponding depth onto the spherical domain plane; Among them, the color image feature contains semantic information, and the point cloud feature contains spatial information.

[0016] According to a method for jointly calibrating a lidar and a camera with spherical space alignment provided by the present invention, feature enhancement is performed on the cross - domain fusion information of the perceptual spherical space through a feature enhancement module to obtain enhanced cross - domain perceptual features, including: Mapping the point cloud feature to the range view to extract features through P2R operation, and performing downsampling using a 2D feature extractor; Mapping the 2D - range map feature back to the corresponding point cloud through R2P operation.

[0017] According to a method for jointly calibrating a lidar and a camera with spherical space alignment provided by the present invention, the geometric alignment includes: The association between lidar point cloud data and camera image data is established using the camera calibration matrix. For each point within , the corresponding pixel coordinates (u, v) are calculated through the following transformation formula:

[0018] where represents the camera extrinsic parameter matrix, including rotation and displacement components, and represents the camera intrinsic parameter matrix.

[0019] According to a lidar and camera joint calibration method with spherical space alignment provided by the present invention, the semantic weight alignment includes: Estimate the depth of 3D points using camera features, compare the depth of 3D points with lidar point depth data, evaluate the auxiliary effect of camera features through difference analysis, and assign higher weights to closer features. The fusion process formula is as follows:

[0020]

[0021] where is the lidar scan data, is the existing feature corresponding to the current 3D point, is the activation function in the multi-layer perceptron, and are two weight matrices used for further linear transformation of the features, is the feature after one-step fusion, is the feature after the second-step fusion.

[0022] In a second aspect, the present invention also provides a lidar and camera joint calibration system with spherical space alignment, including: An acquisition module for acquiring lidar point cloud data and camera image data of a vehicle; An extraction module for constructing a spherical domain formed by multiple spherical spaces and extracting key features of the lidar point cloud data and camera image data based on the spherical domain; A determination module for constructing a perception spherical space from multiple spherical domains, accessing different modality features in multiple spherical domains through linear calibration, and determining the cross-domain fusion information of the perception spherical space; An enhancement module for enhancing the cross-domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross-domain perception features; A fusion module for sequentially performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fusion feature information.

[0023] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method for jointly calibrating a lidar and a camera with spherical space alignment as described in any one of the above.

[0024] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for jointly calibrating a lidar and a camera with spherical space alignment as described in any one of the above.

[0025] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for jointly calibrating a lidar and a camera with spherical space alignment as described in any one of the above.

[0026] The method and system for jointly calibrating a lidar and a camera with spherical space alignment provided by the present invention introduce the concept of a perception spherical space, and use the geometric features from the lidar and the semantic features from the camera in multi-modal features for deep learning alignment. It can effectively overcome the misalignment between lidar information and camera information that is difficult to eliminate in physical methods, significantly improve the alignment accuracy between different sensor modalities, and effectively facilitate the application of multi-modal information in the field of autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0028] Figure 1 is one of the flow diagrams of the method for jointly calibrating a lidar and a camera with spherical space alignment provided by the present invention; Figure 2 is another flow diagram of the method for jointly calibrating a lidar and a camera with spherical space alignment provided by the present invention; Figure 3 is the structural diagram of the system for jointly calibrating a lidar and a camera with spherical space alignment provided by the present invention; Figure 4 is the structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] To make the objectives, technical solutions and advantages of the present invention more clear, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0030] Figure 1 is one of the flow schematic diagrams of the method for joint calibration of a lidar and a camera with spherical space alignment provided by an embodiment of the present invention. As Figure 1 shown, it includes: Step 100: Obtain the lidar point cloud data and camera image data of the vehicle. Step 200: Construct a spherical domain formed by multiple spherical spaces, and extract key features of the lidar point cloud data and camera image data based on the spherical domain. Step 300: Construct a perception sphere space from multiple spherical domains, access different modality features in multiple spherical domains through linear calibration, and determine the cross-domain fusion information of the perception sphere space. Step 400: Perform feature enhancement on the cross-domain fusion information of the perception sphere space through a feature enhancement module to obtain enhanced cross-domain perception features. Step 500: Perform geometric alignment and semantic weight alignment on the enhanced cross-domain perception features in sequence to obtain fusion feature information.

[0031] Specifically, as Figure 2 shown, it includes the following steps: Step 1, obtain the point cloud data of the lidar carried on the vehicle and the color image data of the vehicle-mounted camera.

[0032] Step 2, Multi-spatial Feature Extraction (MFE), create multiple spherical spaces called spherical domains, and extract key features in the spherical domains respectively, such as the semantic information carried in the color photo and the spatial geometric information carried in the lidar point cloud.

[0033] Step 3, introduce the concept of Perception Sphere Space (PSS), which is composed of multiple spherical domains, and each spherical domain retains the unique data characteristics of its corresponding modality. The PSS calibration design allows points to access features from different modalities through linear calibration and supports parallel access to the features of points in adjacent spherical domains. Step 4: Perform feature enhancement through the feature enhancement module. Extract point cloud features and downsample in the range view. It maps the point cloud features to the range view through P2R operation, extracts features, and downsamples using a 2D feature extractor. Then it maps the range view features to the corresponding point cloud through R2P operation.

[0034] Step 5: Through domain-aware feature alignment, geometrically and semantically correct the point cloud data and color image data twice to generate a unified feature representation that can be applied to downstream tasks.

[0035] Furthermore, for the point cloud data obtained in Step 1, it is transformed into a 2D-range map with depth information through a transformation equation system, and then convolution is performed to extract point cloud features for subsequent processing.

[0036] For the color image data obtained in Step 1, after being processed by the preposed convolutional neural network ResNet convolution, a feature map is obtained, and then it is transformed into a 2D-range map with depth information through a transformation equation system (P2R). The formula for converting from the rectangular coordinate system p3D k = (xk, yk, zk) to the spherical coordinate system psph k = (rk, θk, φk) is as follows:

[0037]

[0038]

[0039] The formula for further converting from the spherical coordinate system to the 2D-range map is as follows (discarding the original depth information )

[0040]

[0041] and are the upper and lower limits of the vertical field of view ( ).

[0042] and are the final generated image width and height.

[0043] Furthermore, for the 2D-range map with depth information generated from the point cloud features and color image features in Step 2, spherical segmentation is performed. Each spherical range with a different radius is called a spherical domain, and the features with corresponding depths are projected onto the spherical domain plane. Among them, the color image features carry semantic information, and the point cloud features carry spatial information.

[0044] Furthermore, the Perceptual Sphere Space (PSS) introduced in step 3 promotes unified multi-modal feature representation and fusion by integrating multiple spherical domains. This approach is consistent with the physical modeling of sensors and significantly improves the efficiency of feature indexing. The spherical domains unify the modeling methods of different sensors by mapping the information from various sensors into them, thus achieving cross-domain fusion through PSS.

[0045] Furthermore, the specific implementation process of the feature enhancement module in step 4 is as follows: Extract point cloud features and downsample in the 2D-range map. It maps the point cloud features to the range view through P2R operation, extracts features, and downsamples using a 2D feature extractor. Then, it maps the 2D-range map features back to the corresponding point cloud through R2P operation.

[0046] Furthermore, for the geometric alignment in step 5, the specific implementation process is as follows: In the multi-sensing system, the data encoded in and provides complementary information. To effectively utilize this synergy, we use the camera calibration matrix to establish the connection between LiDAR points and image pixels. Specifically, for each point located in , its corresponding pixel coordinates (u, v) can be calculated through the following transformation formula:

[0047] where represents the external parameter matrix of the camera, which includes rotation and displacement components, and represents the internal parameter matrix of the camera.

[0048] Furthermore, for the semantic weight alignment in step 5, the specific implementation process is as follows: First, we use camera features to estimate the depth of 3D points, then compare them with the LiDAR depth data, and evaluate the auxiliary effect of camera features through difference analysis. Higher weights are assigned to closer features. The mathematical formula for the fusion process is as follows:

[0049]

[0050] is the scan data from LiDAR, that is, the feature information provided by the lidar, which is good at expressing geometric depth information (such as the distance and contour of an object). is the existing feature corresponding to the current point (a 3D point), such as the feature extracted by other methods or perception models before, is the feature after one-step fusion, These are the features after the second - step fusion. The multi - layer perceptron (MLP) is a simple neural network module that extracts new and more useful information based on the input (concatenated features). The role of the MLP is to "process" and "refine" the features. σ is an activation function, which is usually used to make the output of the data non - linear, helping the model better represent complex features. and are two weight matrices used to perform further linear transformations on the features.

[0051] This process can extract deeper - level patterns or characteristics. The multiplication operation refers to the residual connection, which can avoid information loss after multiple processes and enhance the expressive power of the model.

[0052] Finally, these weighting strategies are used to integrate the geometric features of LiDAR with the semantic information from the camera to generate comprehensive hybrid features for downstream tasks such as scene perception in autonomous driving.

[0053] The following describes the ball - space alignment LiDAR - camera joint calibration system provided by the present invention. The ball - space alignment LiDAR - camera joint calibration system described below can be correspondingly referred to the ball - space alignment LiDAR - camera joint calibration method described above.

[0054] Figure 4 An example of the physical structure diagram of an electronic device is shown as Figure 4 shown. The electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the ball - space alignment LiDAR - camera joint calibration method, which includes: obtaining the LiDAR point cloud data and camera image data of the vehicle; constructing a spherical domain formed by multiple spherical spaces, and extracting key features of the LiDAR point cloud data and camera image data based on the spherical domain; constructing a perception spherical space from multiple spherical domains, accessing different - modality features in multiple spherical domains through linear calibration, and determining the cross - domain fusion information of the perception spherical space; performing feature enhancement on the cross - domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross - domain perception features; sequentially performing geometric alignment and semantic weight alignment on the enhanced cross - domain perception features to obtain fusion feature information.

[0055] In addition, when the logical instructions in the above-mentioned memory 430 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0056] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for jointly calibrating a lidar and a camera with spherical space alignment provided by the above-mentioned various methods. The method includes: obtaining lidar point cloud data and camera image data of a vehicle; constructing a spherical domain formed by multiple spherical spaces, and extracting key features of the lidar point cloud data and camera image data based on the spherical domain; constructing a perception spherical space from multiple spherical domains, accessing different modality features in multiple spherical domains through linear calibration, and determining cross-domain fusion information of the perception spherical space; performing feature enhancement on the cross-domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross-domain perception features; sequentially performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fusion feature information.

[0057] On another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the method for jointly calibrating a lidar and a camera with spherical space alignment provided by the above-mentioned various methods. The method includes: obtaining lidar point cloud data and camera image data of a vehicle; constructing a spherical domain formed by multiple spherical spaces, and extracting key features of the lidar point cloud data and camera image data based on the spherical domain; constructing a perception spherical space from multiple spherical domains, accessing different modality features in multiple spherical domains through linear calibration, and determining cross-domain fusion information of the perception spherical space; performing feature enhancement on the cross-domain fusion information of the perception spherical space through a feature enhancement module to obtain enhanced cross-domain perception features; sequentially performing geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fusion feature information.

[0058] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.

[0059] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A laser radar and camera joint calibration method for spherical space alignment, characterized in that: include: Obtain vehicle-mounted LiDAR point cloud data and camera image data; Constructing a spherical domain formed by a plurality of spherical spaces, and extracting key features of the laser radar point cloud data and the camera image data based on the spherical domain; The perception sphere space is constructed from multiple sphere domains, and the different modal features in multiple sphere domains are accessed through linear calibration to determine the cross-domain fusion information of the perception sphere space; The feature enhancement module is used to enhance the cross-domain fusion information of the perception ball space to obtain enhanced cross-domain perception features; The enhanced cross-domain perception features are sequentially subjected to geometric alignment and semantic weight alignment to obtain fused feature information.

2. The spherical space aligned laser radar and camera joint calibration method according to claim 1, characterized in that: After obtaining the on-board LiDAR point cloud data and camera image data, it also includes: The laser radar point cloud data is converted into a 2D-range map containing depth information based on a transformation equation, and point cloud features are extracted by convolution; The camera image data is convolved with the pre-convolutional neural network ResNet to obtain color image features, and the equation group P2R is converted into a 2D-range map containing depth information. The formula for converting from the rectangular coordinate system p3D k = (xk, yk, zk) to the spherical polar coordinate system psph k = (rk, θk, φk) is as follows: The formula for converting from the spherical polar coordinate system to the 2D-range graph is as follows: in, and are the upper and lower limits of the vertical field of view, and is the final image width and height.

3. The spherical space aligned laser radar and camera joint calibration method according to claim 1, characterized in that: Constructing a spherical domain formed by multiple spherical spaces, and extracting key features of the laser radar point cloud data and camera image data based on the spherical domain, including: Perform spherical segmentation on the 2D-range map containing depth information generated by point cloud features and color features, determine each spherical range with different radius as a spherical domain, and project the features of the corresponding depth onto the spherical domain plane; The color image features contain semantic information, and the point cloud features contain spatial information.

4. The spherical space aligned laser radar and camera joint calibration method according to claim 1, characterized in that: The feature enhancement module is used to enhance the cross-domain fusion information of the perception ball space to obtain enhanced cross-domain perception features, including: Map point cloud features to range view extracted features through P2R operation and use 2D feature extractor for downsampling; The 2D-range map features are mapped back to the corresponding point cloud through the R2P operation.

5. The spherical space aligned laser radar and camera joint calibration method according to claim 1, characterized in that: The geometric alignment includes: The camera calibration matrix is ​​used to establish the association between the lidar point cloud data and the camera image data. Each point in , the corresponding pixel coordinates (u, v) are calculated by the following transformation formula: in, represents the camera extrinsic matrix, including rotation and displacement components, Represents the camera intrinsic parameter matrix.

6. The spherical space aligned laser radar and camera joint calibration method according to claim 1, characterized in that: The semantic weight alignment includes: The camera features are used to estimate the 3D point depth, and the 3D point depth is compared with the lidar point depth data. The auxiliary effect of the camera features is evaluated through difference analysis, and higher weights are given to features that are closer. The fusion process formula is as follows: in, is the LiDAR scan data, is the existing feature corresponding to the current 3D point, is the activation function in the multi-layer perceptron, and Are two weight matrices used to further linearly transform the features. is the feature after one step of fusion. It is the feature after the second step fusion.

7. A laser radar and camera joint calibration system for spherical space alignment, characterized in that: include: The acquisition module is used to obtain the vehicle-mounted lidar point cloud data and camera image data; An extraction module, used to construct a spherical domain formed by multiple spherical spaces, and extract key features of the laser radar point cloud data and the camera image data based on the spherical domain; A determination module is used to construct a perception sphere space from multiple sphere domains, access different modal features in multiple sphere domains through linear calibration, and determine cross-domain fusion information of the perception sphere space; An enhancement module, used for performing feature enhancement on the cross-domain fusion information of the perception ball space through a feature enhancement module to obtain enhanced cross-domain perception features; The fusion module is used to sequentially perform geometric alignment and semantic weight alignment on the enhanced cross-domain perception features to obtain fused feature information.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the laser radar and camera joint calibration method for spherical space alignment as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the laser radar and camera joint calibration method for spherical space alignment as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the laser radar and camera joint calibration method for spherical space alignment as described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Large-scale point cloud semantic segmentation method based on superpoint graph

    CN108319957A

  • Space alignment method and system for vehicle-road cooperation

    CN116229713A

  • Detection method, system, and device based on fusion of image and point cloud information, and storage medium

    WO2022156175A1