ROS-based Fusion Calibration System and Method for LiDAR and Camera

By performing dynamic object removal and multi-scale feature extraction on the ROS platform, combining RANSAC algorithm and global weighted least squares optimization, the accuracy and adaptability problems of traditional calibration methods in dynamic environments are solved, and high-precision and highly robust lidar and camera fusion calibration is achieved.

CN120065187BActive Publication Date: 2025-07-25SHENZHEN YONGTAI PHOTOELECTRIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510543552.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-25
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The traditional lidar and camera calibration methods have inaccurate accuracy in dynamic environments, severe mismatch of position estimation, large calculation volume, low efficiency, and lack optimization strategies for specific scenarios.

Method used

Data is obtained through the ROS platform, dynamic object culling and multi-scale feature extraction are performed, combined with RANSAC algorithm and global weighted least squares optimization, feature matching and pose estimation are performed, and multi-angle calibration parameter evaluation and scene-specific optimization are performed.

Benefits of technology

The robustness and accuracy of calibration are improved in a dynamic environment, the adaptability and stability of the algorithm are enhanced, and the fusion calibration is achieved with high precision, high robustness and adaptive.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120065187B_ABST
    Figure CN120065187B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of the fusion calibration of lidar and cameras, and particularly to a fusion calibration system and method for lidar and cameras based on ROS. The method includes the following steps: obtaining the original lidar point cloud, the original camera image, and the odometer information through the ROS platform; performing dynamic object detection on the original lidar point cloud and the original camera image, and removing the dynamic objects to obtain the static environment point cloud and the static environment image; performing multi-scale environmental feature extraction on the static environment point cloud and the static environment image, and performing feature fusion to obtain a fusion feature set; performing adaptive feature enhancement processing on the fusion feature set to obtain a multi-modal feature map; performing feature matching on the multi-modal feature map and performing motion compensation to obtain a correspondence set. The present invention realizes a high-precision, highly robust and adaptive fusion calibration method through the fusion calibration technology of lidar and cameras.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the fusion calibration of lidar and camera, and particularly to a lidar and camera fusion calibration system and method based on ROS. Background Art

[0002] Lidar and camera are two commonly used sensors in autonomous driving systems. Lidar provides accurate three-dimensional point cloud data, while the camera provides rich texture information. Fusing the data of these two sensors can complement each other's advantages and improve the accuracy and robustness of environmental perception. However, to achieve effective fusion, accurate calibration must be performed first to determine the spatial transformation relationship between the lidar and the camera. Traditional calibration methods mainly rely on specific calibration boards or manual markings and are carried out in a static environment. These methods perform well in a laboratory environment but have limitations in practical applications:

[0003] Challenges in dynamic environments: Traditional methods assume that the calibration environment is static and ignore the influence of dynamic objects. This is unrealistic in actual driving scenarios because the road is full of moving vehicles, pedestrians, and other objects. These dynamic objects will be misidentified as calibration features, resulting in inaccurate calibration results. In a dynamic environment, moving objects will be misidentified as static features, leading to incorrect feature matching and ultimately affecting the calibration accuracy. For example, the corner points on a moving car may be mis-matched to the corner points of a background building.

[0004] Challenges in pose estimation: There are incorrect matches during the feature matching process, and these incorrect matches will seriously affect the accuracy of pose estimation. Although global optimization methods can reduce the influence of incorrect matches, they have a large amount of computation and low efficiency. In addition, the lack of optimization strategies for specific scenarios makes it difficult to achieve the optimal calibration accuracy. Summary of the Invention

[0005] Based on this, it is necessary to provide a lidar and camera fusion calibration system and method based on ROS to solve at least one of the above technical problems.

[0006] A lidar and camera fusion calibration method based on ROS includes the following steps:

[0007] Step S1: Obtain the original lidar point cloud, the original camera image, and the odometer information through the ROS platform; perform dynamic object detection on the original lidar point cloud and the original camera image, and remove the dynamic objects to obtain the static environment point cloud and the static environment image; perform multi-scale environmental feature extraction on the static environment point cloud and the static environment image, and perform feature fusion to obtain a fusion feature set; perform adaptive feature enhancement processing on the fusion feature set to obtain a multi-modal feature map;

[0008] Step S2: Perform feature matching on the multi-modal feature map and perform motion compensation to obtain a set of correspondence relations;

[0009] Step S3: Perform pose estimation on the set of correspondence relations to obtain rough pose estimation data; perform global pose optimization based on the rough pose estimation data, and perform scene-based local optimization to obtain a locally optimized pose; perform iterative optimization based on the locally optimized pose to obtain calibration parameters;

[0010] Step S4: Evaluate the calibration parameters from multiple angles according to the calibration parameters to obtain a verification report;

[0011] Step S5: Monitor the environment of the original lidar point cloud and the original camera image, and perform current scene classification to obtain the current scene category; adjust the multi-modal feature map and calibration parameters according to the verification report, calibration parameters, and current scene category, and perform scene-specific optimization to obtain optimized calibration parameters.

[0012] Through dynamic object culling and multi-scale feature extraction, the present invention effectively reduces the impact of dynamic environments and targets of different scales on the calibration results. Through feature fusion and adaptive enhancement, a more robust and discriminative multi-modal feature map is constructed, laying a foundation for subsequent precise matching and pose estimation. Motion compensation based on feature matching of the multi-modal feature map and odometer information effectively improves the accuracy and reliability of feature matching. Especially in dynamic environments, a more precise set of corresponding relationships can be obtained, providing reliable data for subsequent pose estimation. The RANSAC algorithm is used for rough pose estimation to effectively eliminate the influence of mismatches. Global weighted least squares optimization and scene-based local optimization further improve the accuracy of pose estimation and enhance the adaptability of the algorithm to different scenarios. Finally, more precise calibration parameters are obtained through iterative optimization. Multi-angle calibration parameter evaluation, including reprojection error analysis, scene analysis, key point alignment verification, and robustness evaluation, can comprehensively and quantitatively evaluate the accuracy, applicability, and stability of the calibration results, providing a reference for algorithm improvement and practical applications. Through environmental monitoring, scene classification, and feature weight adjustment, the adaptive ability of the calibration algorithm is realized, enabling the calibration system to dynamically adjust parameters according to changes in the environment and scene, thereby maintaining a high calibration accuracy and robustness in complex and changing practical application scenarios. Therefore, the present invention provides a fusion calibration method for lidar and camera based on ROS, which effectively solves the challenges brought by dynamic environments by performing dynamic object detection and culling in the data preprocessing stage, improving the robustness of calibration. In terms of pose estimation, the present invention combines the RANSAC algorithm, global weighted least squares optimization, and scene-based local optimization strategies, taking into account both accuracy and efficiency and enhancing the adaptability to different scenarios. Through adaptive feature weight adjustment and iterative optimization, the stability and reliability of the calibration results are further enhanced, ultimately realizing a high-precision, highly robust, and adaptive fusion calibration scheme in complex dynamic environments.

[0013] Preferably, step S1 includes the following steps:

[0014] Step S11: Obtain the original lidar point cloud, the original camera image, and odometer information; perform point cloud motion compensation on the original lidar point cloud according to the odometer information to obtain the compensated point cloud;

[0015] Step S12: Perform dynamic object detection on the compensated point cloud and the original camera image respectively, and perform dynamic object culling to obtain the static environment point cloud and the static environment image;

[0016] Step S13: Perform multi-scale plane feature extraction on the static environment point cloud to obtain multi-scale plane information;

[0017] Step S14: Extract multi-scale edge and corner features from the static environment point cloud to obtain multi-scale geometric features;

[0018] Step S15: Extract image features from the static environment image to obtain image feature information;

[0019] Step S16: Project the multi-scale plane information and multi-scale geometric features into the image feature information to obtain a fused feature set;

[0020] Step S17: Perform adaptive feature enhancement on the fused feature set to obtain a weighted fused feature set;

[0021] Step S18: Generate a multi-modal feature map from the weighted fused feature set to obtain a multi-modal feature map.

[0022] In the present invention, the odometer information is used to perform motion compensation on the lidar point cloud, and the point cloud data collected at different times is transformed into the same coordinate system, effectively eliminating the point cloud distortion caused by vehicle motion, and providing a more accurate data basis for subsequent feature extraction and matching. By separately performing dynamic object detection and removal on the point cloud and the image, the interference of moving objects on the calibration result is effectively removed, the robustness and accuracy of the calibration algorithm are improved, and the calibration process is more focused on the static environment features. Extracting plane features from point clouds with different resolutions can more comprehensively capture the plane information in the scene, improve the perception ability of plane objects with different sizes and distances, and thus enhance the adaptability of the calibration algorithm. Similarly, the extraction of multi-scale edge and corner features can more finely capture the geometric features in the scene, improve the perception ability of geometric objects with different scales, and further enrich the feature information for calibration. Extracting ORB features from static environment images provides rich texture information, which is complementary to the point cloud geometric features, enhancing the reliability and accuracy of feature matching. Projecting the multi-scale plane information and geometric features onto the image plane and associating them with the image features realizes the effective fusion of point cloud data and image data, providing multi-modal feature correspondence for subsequent pose estimation. By performing adaptive feature enhancement on the fused feature set and assigning different weights according to the confidence of the features, environmental stability, and scene type, the contribution of reliable features is improved, and the influence of noise and unstable features is reduced, thereby improving the accuracy of pose estimation. Organizing the weighted fused feature set into a structured multi-modal feature map facilitates efficient access and use of feature information in subsequent steps, providing a unified data structure for feature matching and pose estimation.

[0023] Preferably, step S12 includes the following steps:

[0024] Step S121: Perform point cloud motion estimation and segmentation on the compensated point cloud according to the odometer information to obtain the dynamic point cloud and the static point cloud segmentation result;

[0025] Step S122: Perform image background modeling on the original camera image, and perform foreground segmentation to obtain a foreground image;

[0026] Step S123: Perform optical flow calculation on the original camera image, and perform dynamic area detection to obtain a dynamic area mask;

[0027] Step S124: Project the dynamic point cloud onto the foreground image, and determine dynamic objects according to the static point cloud segmentation result and the dynamic area mask to obtain the final dynamic objects;

[0028] Step S125: Remove dynamic objects from the compensated point cloud and the original camera image according to the final dynamic objects to obtain a static environment point cloud and a static environment image.

[0029] The present invention estimates the movement of the point cloud through odometer information, and combines the density clustering algorithm to segment the point cloud into dynamic and static parts, realizing the preliminary identification of moving objects in the point cloud, laying a foundation for subsequent accurate dynamic object removal. Use the Gaussian mixture model to perform background modeling on the image, and extract the potentially moving regions in the image through foreground segmentation, providing prior information at the image level for subsequent confirmation of dynamic objects in combination with point cloud information. Obtain the motion information of image pixels through optical flow calculation and generate a dynamic area mask, further refining the identification of moving regions in the image and effectively compensating for the possible deficiencies of the background modeling method. Project the dynamic point cloud onto the image plane, and combine the foreground image and the dynamic area mask to realize the fusion of point cloud and image information, thereby more accurately determining dynamic objects and reducing the probability of false detection and missed detection. According to the finally determined dynamic objects, the corresponding parts are removed from the compensated point cloud and the original camera image, obtaining point cloud and image data that only contain static environment information, providing purer data for subsequent feature extraction and matching, and improving the calibration accuracy.

[0030] Preferably, step S17 includes the following steps:

[0031] Step S171: Evaluate the feature confidence of the fusion feature set to obtain the feature confidence;

[0032] Step S172: Calculate the voxel displacement change within the point cloud motion time window according to the fusion feature set to obtain voxel displacement change data; calculate the image background update frequency according to the fusion feature set; perform environmental stability evaluation according to the voxel displacement change data and the image background update frequency to obtain an environmental stability evaluation value;

[0033] Step S173: Perform an initial judgment on the scene type according to the fusion feature set to obtain initial scene category judgment data;

[0034] Step S174: Feature-weight the fused feature set according to the feature confidence, the environmental stability evaluation value, and the initial scene category judgment data to obtain a weighted fused feature set.

[0035] In the present invention, by evaluating the reprojection error, geometric characteristics, and descriptor quality of the fused features, the confidence score of each feature is obtained, providing a basis for subsequent feature weighting, enabling more reliable features to play a greater role in the calibration process. By analyzing the voxel displacement change of the point cloud and the image background update frequency, the stability of the current environment is evaluated, providing environmental information for subsequent feature weighting, enabling more stable features to obtain higher weights in a dynamic environment. According to the three-dimensional point distribution in the fused feature set, the scene type is initially judged, providing prior information for subsequent scene-specific optimization, enabling the calibration algorithm to adopt more appropriate optimization strategies according to different scene types. Considering the feature confidence, the environmental stability evaluation value, and the initial scene category judgment data comprehensively, the fused feature set is weighted, highlighting the contribution of reliable features and reducing the influence of noise and unstable features, thereby improving the robustness and accuracy of the calibration algorithm.

[0036] Preferably, step S2 includes the following steps:

[0037] Step S21: Perform a preliminary match on the multi-modal feature map based on geometric constraints to obtain a candidate match set;

[0038] Step S22: Perform a fine match on the candidate match set based on descriptors to obtain a preliminary match set;

[0039] Step S23: Perform motion compensation on the preliminary match set to obtain a compensated match set;

[0040] Step S24: Perform outlier rejection and optimization on the compensated match set according to the multi-modal feature map to obtain a correspondence set.

[0041] In the present invention, by setting distance and angle thresholds, a preliminary match is performed on the multi-modal feature map, quickly screening out candidate match pairs that meet geometric constraints, effectively reducing the computational amount of subsequent fine matching and improving the matching efficiency. A fine match based on ORB descriptors is performed on the candidate match set, and by calculating the Hamming distance, match pairs with similar feature descriptors are further screened out, improving the matching accuracy. Using odometer information to perform motion compensation on the preliminary match set, converting feature points at different times to the same coordinate system, eliminating the influence of vehicle motion on the matching result, and improving the matching reliability, especially in a dynamic environment. Using the RANSAC algorithm to perform outlier rejection on the compensated match set, further removing incorrect matches and optimizing the final correspondence set, providing a more accurate matching result for subsequent pose estimation.

[0042] Preferably, step S3 includes the following steps:

[0043] Step S31: Perform a rough pose estimation based on RANSAC on the correspondence set to obtain rough pose estimation data and an inlier set;

[0044] Step S32: Perform global weighted least squares optimization according to the multi-modal feature map, the rough pose estimation data, and the inlier set to obtain a globally optimized pose;

[0045] Step S33: Perform scene-based local optimization according to the multi-modal feature map, the globally optimized pose, and the inlier set to obtain a locally optimized pose;

[0046] Step S34: Perform iterative optimization according to the locally optimized pose, the correspondence set, and the multi-modal feature map to obtain calibration parameters.

[0047] In the present invention, by using the RANSAC algorithm to estimate a rough pose transformation from the correspondence set, the influence of mismatches is effectively eliminated, and a preliminary pose estimation result is obtained, providing a good initial value for subsequent optimization. Based on the rough pose estimation and the inlier set, a global weighted least squares optimization problem is constructed, and the feature confidence is used as the weight to globally optimize the pose, improving the accuracy of pose estimation and effectively utilizing the reliability information of the features. According to the scene type, using scene-specific geometric constraints (such as the planarity of roads, the verticality of buildings, the cylindrical shape of tunnels), the globally optimized pose is locally optimized, further improving the accuracy of pose estimation and enhancing the adaptability of the algorithm to different scenes. By iteratively executing the rough pose estimation, global optimization, and local optimization steps, the pose estimation result is continuously improved until convergence, and finally more accurate calibration parameters are obtained, improving the overall accuracy and stability of calibration.

[0048] Preferably, step S33 includes the following steps:

[0049] Step S331: Extract scene local data from the multi-modal feature map to obtain scene local data, where the scene local data includes scene odometry data, scene image data, and scene point cloud data;

[0050] Step S332: Extract odometry features from the scene odometry data to obtain odometry features;

[0051] Step S333: Detect road markings in the scene image data and detect curbs to obtain road image features;

[0052] Step S334: Calculate the road surface flatness for the scene point cloud data to obtain the road surface flatness; extract the building features from the scene point cloud data to obtain the building features; extract the point cloud density features from the scene point cloud data to obtain the point cloud density features; generate the road point cloud features based on the road surface flatness, building features, and point cloud density features to obtain the road point cloud features;

[0053] Step S335: Construct a scene recognition model based on a support vector machine according to the odometer features, road image features, and road point cloud features to obtain the scene recognition model; use the scene recognition model to perform scene model recognition to obtain the model recognition scene categories, where the model recognition scene categories include highway scenes, urban road scenes, and tunnel scenes;

[0054] Step S336: Construct a constraint optimization problem based on road markings for the highway scene to obtain the highway scene optimization problem; construct a constraint optimization problem based on building corners and curbs for the urban road scene to obtain the urban road scene optimization problem; construct a constraint optimization problem based on tunnel walls for the tunnel scene to obtain the tunnel scene optimization problem;

[0055] Step S337: Perform non-linear optimization on the highway scene optimization problem, urban road scene optimization problem, and tunnel scene optimization problem according to the global optimized pose and the inlier set to obtain the local optimized pose.

[0056] The present invention extracts local scene data from a multi-modal feature map, including odometer data, image data, and point cloud data, providing a necessary data basis for subsequent scene recognition and local optimization. By extracting odometer features, such as the average speed and acceleration of a vehicle, the motion information of the scene is provided, which helps to distinguish different scene types, such as highways, urban roads, etc. By detecting road markings and curbs and extracting road image features, the structural information of the scene is provided, which helps to identify road scenes and provides support for optimization based on road geometric constraints. By calculating the road surface flatness, extracting building features and point cloud density features, and combining them into road point cloud features, the three-dimensional structural information of the scene is provided, which helps to identify road scenes and building scenes and provides support for optimization based on three-dimensional geometric constraints. Using the extracted odometer features, road image features, and road point cloud features, a scene recognition model based on a support vector machine is trained, and the model is used to identify the current scene category (highway, urban road, tunnel), providing a basis for subsequent scene-specific optimization. According to the identified scene category, different optimization problems are constructed, and different geometric constraint conditions, such as the parallelism of road markings, the perpendicularity of building corners, the cylindrical shape of tunnel walls, etc., are designed for highway, urban road, and tunnel scenes respectively, providing more accurate constraints for subsequent local optimization. Using the global optimization pose and the inlier set, the constructed scene-specific optimization problem is nonlinearly optimized to obtain the local optimization pose, further improving the accuracy of pose estimation and enhancing the adaptability of the algorithm to different scenes.

[0057] Preferably, step S4 includes the following steps:

[0058] Step S41: Perform reprojection error analysis based on calibration parameters, the original lidar point cloud, and the original camera image to obtain reprojection error statistics;

[0059] Step S42: Perform scene analysis based on calibration parameters, the original lidar point cloud, the original camera image, and odometer information to obtain scene performance evaluation;

[0060] Step S43: Perform key point alignment verification based on calibration parameters, the original lidar point cloud, and the original camera image to obtain key point alignment results;

[0061] Step S44: Perform robustness evaluation based on calibration parameters, the original lidar point cloud, the original camera image, and odometer information to obtain a robustness evaluation result;

[0062] Step S45: Generate a verification report for the reprojection error statistics, scene performance evaluation, key point alignment results, and robustness evaluation results to obtain a verification report.

[0063] By calculating the reprojection error of lidar point cloud projected onto the image and conducting statistical analysis, the present invention can quantitatively evaluate the accuracy of calibration parameters, providing a direct numerical index for judging the quality of calibration results. Performance evaluation is carried out separately for different scenarios (such as highways, urban roads, tunnels), which can understand the applicability of the calibration algorithm in different scenarios and provide a direction for subsequent algorithm improvement. By manually selecting key points on the image and comparing them with the lidar point cloud projection, the accuracy of calibration parameters can be visually verified, especially in some areas with obvious geometric features. By testing under different environmental conditions (such as different lighting, weather, vehicle speeds), the robustness of calibration parameters can be evaluated, and the stability and reliability of the algorithm in complex environments can be verified. Integrating the evaluation results into the verification report can comprehensively and systematically evaluate the calibration results, providing clear performance indicators and analysis conclusions for users.

[0064] Preferably, step S5 includes the following steps:

[0065] Step S51: Monitor the environment for the original lidar point cloud, the original camera image, and the odometer information to obtain environmental change information;

[0066] Step S52: Classify the current scene according to the environmental change information to obtain the current scene category;

[0067] Step S53: Adjust the feature weights according to the verification report, calibration parameters, environmental change information, and the current scene category to obtain the adjusted multi-modal feature map;

[0068] Step S54: Adjust the calibration parameters according to the adjusted multi-modal feature map to obtain the adjusted calibration parameters;

[0069] Step S55: Conduct scene-specific optimization according to the adjusted multi-modal feature map, the adjusted calibration parameters, and the current scene category to obtain the optimized calibration parameters.

[0070] By monitoring information such as ambient light intensity, point cloud density, vehicle speed, and image clarity in real time, the present invention can obtain environmental change information, providing a basis for adaptive calibration and enabling the calibration algorithm to be dynamically adjusted according to environmental changes. According to the environmental change information and the extracted features, the current scene can be classified to determine the current scene type, such as highway, urban road, tunnel, etc., providing guidance for subsequent scene-specific optimization. According to the verification report, calibration parameters, environmental change information, and the current scene category, the feature weights in the multi-modal feature map can be dynamically adjusted to improve the adaptability of the calibration algorithm in different environments and scenes. For example, in an environment with poor lighting, the weight of image features can be reduced and the weight of point cloud features can be increased. According to the adjusted multi-modal feature map, pose estimation and optimization are performed again to obtain calibration parameters that are more suitable for the current environment and scene, improving the calibration accuracy. According to the adjusted calibration parameters and the current scene category, scene-specific optimization can be carried out to further improve the calibration accuracy. For example, in a highway scene, optimization can be performed using the parallelism constraint of road markings. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is a schematic flowchart of the steps of a fusion calibration method for a lidar and a camera based on ROS;

[0072] Figure 2 It is a schematic flowchart of the detailed implementation steps of step S1 in the present invention.

[0073] The realization of the object, functional features, and advantages of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0075] In addition, the drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.

[0076] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.

[0077] To achieve the above object, please refer to Figures 1 to 2 , a fusion calibration method for lidar and camera based on ROS, comprising the following steps:

[0078] Step S1: Obtain the original lidar point cloud, the original camera image, and the odometer information through the ROS platform; perform dynamic object detection on the original lidar point cloud and the original camera image, and remove the dynamic objects to obtain the static environment point cloud and the static environment image; perform multi-scale environmental feature extraction on the static environment point cloud and the static environment image, and perform feature fusion to obtain a fusion feature set; perform adaptive feature enhancement processing on the fusion feature set to obtain a multi-modal feature map;

[0079] Step S2: Perform feature matching on the multi-modal feature map and perform motion compensation to obtain a correspondence set;

[0080] Step S3: Perform pose estimation on the correspondence set to obtain rough pose estimation data; perform global pose optimization based on the rough pose estimation data, and perform scene-based local optimization to obtain a locally optimized pose; perform iterative optimization based on the locally optimized pose to obtain calibration parameters;

[0081] Step S4: Perform multi-angle calibration parameter evaluation based on the calibration parameters to obtain a verification report;

[0082] Step S5: Monitor the environment of the original lidar point cloud and the original camera image, and perform current scene classification to obtain the current scene category; adjust the multi-modal feature map and the calibration parameters according to the verification report, the calibration parameters, and the current scene category, and perform scene-specific optimization to obtain optimized calibration parameters.

[0083] In the embodiments of the present invention, referring to Figure 1 shown, is a schematic flowchart of the steps of the fusion calibration method for lidar and camera based on ROS of the present invention. In this example, the fusion calibration method for lidar and camera based on ROS includes the following steps:

[0084] Step S1: Obtain the original lidar point cloud, the original camera image, and the odometer information through the ROS platform; perform dynamic object detection on the original lidar point cloud and the original camera image, and remove the dynamic objects to obtain the static environment point cloud and the static environment image; perform multi-scale environmental feature extraction on the static environment point cloud and the static environment image, and perform feature fusion to obtain a fused feature set; perform adaptive feature enhancement processing on the fused feature set to obtain a multi-modal feature map;

[0085] In the embodiment of the present invention, the original lidar point cloud, the original camera image, and the odometer information are received from the ROS platform. Then, through point cloud motion compensation and dynamic object detection and removal, the static environment information is separated. Next, multi-scale feature extraction is performed on the static environment point cloud and image, and features of different modalities are fused to construct a fused feature set. Finally, according to the feature confidence, environmental stability, and scene type, the fused feature set is weighted and enhanced to generate a multi-modal feature map, providing input for subsequent feature matching and pose estimation.

[0086] Step S2: Perform feature matching on the multi-modal feature map, and perform motion compensation to obtain a set of corresponding relationships;

[0087] In the embodiment of the present invention, geometric constraints are used to perform preliminary matching on the features in the multi-modal feature map to screen candidate matching pairs. Then, the similarity of the descriptors is calculated to perform fine matching on the candidate matching set to obtain a preliminary matching set. Next, the odometer information is used to perform motion compensation on the preliminary matching set to eliminate the influence of vehicle motion. Finally, through geometric information and feature confidence, outliers are removed and optimized to obtain the final set of corresponding relationships, providing a basis for subsequent pose estimation.

[0088] Step S3: Perform pose estimation on the set of corresponding relationships to obtain rough pose estimation data; perform global pose optimization based on the rough pose estimation data, and perform scene-based local optimization to obtain a locally optimized pose; perform iterative optimization based on the locally optimized pose to obtain calibration parameters;

[0089] In the embodiment of the present invention, the RANSAC algorithm is used to perform rough pose estimation on the set of corresponding relationships and remove outliers to obtain rough pose estimation data and an inlier set. Then, based on the inlier set and the multi-modal feature map, global weighted least squares optimization is performed to improve the accuracy of pose estimation. Next, according to the scene category, scene-based local optimization is performed, and scene-specific constraints are used to further improve the calibration accuracy. Finally, iterative optimization is performed to continuously update the pose estimation until the result converges to obtain the final calibration parameters.

[0090] Step S4: Perform multi-angle calibration parameter evaluation according to the calibration parameters to obtain a verification report;

[0091] In the embodiment of the present invention, according to the calibration parameters, the reprojection error is calculated to evaluate the accuracy of the calibration result. Then, for different scenarios, the performance of the calibration result is evaluated. Then, the alignment degree of key points in the LiDAR and camera coordinate systems is verified. In addition, under different environmental conditions, the calibration process is repeated to evaluate the robustness of the calibration result. Finally, a verification report is generated to summarize the evaluation information and comprehensively evaluate the accuracy and stability of the calibration result.

[0092] Step S5: Monitor the environment of the original LiDAR point cloud and the original camera image, classify the current scenario to obtain the current scenario category; adjust the multi-modal feature map and calibration parameters according to the verification report, calibration parameters and the current scenario category, and perform scenario-specific optimization to obtain optimized calibration parameters.

[0093] In the embodiment of the present invention, the environmental changes are continuously monitored, including illumination, dynamic objects, road types and weather, etc. Then, according to the environmental change information, the current scenario is classified. Then, combining the verification report, calibration parameters, environmental change information and the current scenario category, the weights of the features in the multi-modal feature map are adjusted, and the calibration parameters are fine-tuned to adapt to environmental changes and scenario changes. Finally, for different scenarios, specific optimization strategies are used to further improve the accuracy of the calibration result to obtain optimized calibration parameters.

[0094] Preferably, step S1 includes the following steps:

[0095] Step S11: Obtain the original LiDAR point cloud, the original camera image and the odometer information; perform point cloud motion compensation on the original LiDAR point cloud according to the odometer information to obtain the compensated point cloud;

[0096] Step S12: Detect dynamic objects in the compensated point cloud and the original camera image respectively, and remove the dynamic objects to obtain the static environment point cloud and the static environment image;

[0097] Step S13: Extract multi-scale plane features from the static environment point cloud to obtain multi-scale plane information;

[0098] Step S14: Extract multi-scale edge and corner features from the static environment point cloud to obtain multi-scale geometric features;

[0099] Step S15: Extract image features from the static environment image to obtain image feature information;

[0100] Step S16: Project the multi-scale plane information and the multi-scale geometric features into the image feature information to obtain a fused feature set;

[0101] Step S17: Perform adaptive feature enhancement on the fused feature set to obtain a weighted fused feature set;

[0102] Step S18: Generate a multi-modal feature map from the weighted fused feature set to obtain a multi-modal feature map.

[0103] As an example of the present invention, refer to Figure 2 As shown, in this example, the said Step S1 includes:

[0104] Step S11: Obtain the original lidar point cloud, the original camera image, and the odometer information; perform point cloud motion compensation on the original lidar point cloud according to the odometer information to obtain the compensated point cloud;

[0105] In the embodiments of the present invention, the original data from the lidar, camera, and odometer are received through the ROS (Robot Operating System) platform. The lidar data is published in the form of a point cloud, and its ROS message type is sensor_msgs / PointCloud2, which contains the three-dimensional coordinates (x, y, z) of each point and other attributes, such as the intensity value. The original image data of the camera is published in the form of the ROS message type sensor_msgs / Image, which contains the pixel data, width, height, and encoding information of the image. The odometer information is usually published through the nav_msgs / Odometry message type, which contains the pose information (position and orientation) and speed information of the vehicle. The pose information of the vehicle, including the position (x, y, z) and orientation (quaternion), is extracted from the nav_msgs / Odometry message. Using the method of linear interpolation, according to the scan timestamp of the point cloud and the timestamp of the odometer, the pose of the vehicle corresponding to each point at the scanning moment is calculated. Assuming that the scanning period of the lidar is T, for any point in the point cloud, its scanning timestamp t is less than the current timestamp, and the pose of the point at the moment t is calculated by linear interpolation. Then, each point is transformed from its scanning moment coordinate system to the current moment coordinate system. This involves coordinate transformation. First, the point cloud coordinate system is transformed to the vehicle coordinate system, and then, according to the vehicle pose information, the point cloud coordinates are transformed to the world coordinate system or the global coordinate system. This process can be implemented using the tf library of ROS. Using the coordinate transformation function provided by the tf library, the conversion between coordinate systems can be conveniently performed. Finally, the point cloud after motion compensation is obtained, eliminating the influence of vehicle motion on the point cloud and preparing for subsequent static environment analysis.

[0106] Step S12: Perform dynamic object detection on the compensated point cloud and the original camera image respectively, and perform dynamic object removal to obtain a static environment point cloud and a static environment image;

[0107] In the embodiments of the present invention, the motion of the point cloud is detected by comparing the displacements between two consecutive frames of the point cloud. Specifically, the ICP (Iterative Closest Point) algorithm can be used to estimate the rigid transformation between two adjacent frames of the point cloud. The point cloud is segmented into multiple voxels, and the displacement of each voxel between adjacent frames is calculated. If the displacement of the voxel exceeds a predefined threshold, it is considered that the voxel may belong to a dynamic object. A clustering-based motion segmentation method can also be used. For example, based on Euclidean distance clustering, the point cloud is segmented into multiple clusters, and the motion vector of each cluster is calculated. If the motion vector of the cluster exceeds the threshold, it is considered that the cluster is a dynamic object. A background modeling method based on time series is adopted, such as the Gaussian Mixture Model (GMM). For each pixel, a Gaussian mixture model is established to describe the color distribution of the pixel at different time periods. When a new frame of the image arrives, the color of the pixel is matched with the GMM. If the color matches one of the Gaussian distributions in the GMM, it is considered that the pixel belongs to the background. Otherwise, it is considered that the pixel belongs to the foreground. The optical flow algorithm, such as the Lucas-Kanade algorithm, is used to calculate the pixel motion vector between two adjacent frames of the image. According to the distribution of the optical flow field, the dynamic regions in the image are detected. For example, if the optical flow vectors in a certain region change violently, it is considered that there may be a dynamic object in this region. The dynamic point cloud obtained in step S121 is projected onto the image plane. Combining the foreground segmentation result obtained in step S122 and the dynamic region mask obtained in step S123, a more accurate determination of the dynamic object is made. If a point cloud point is marked as a dynamic object and its projected point on the image plane is located in the foreground region or the dynamic region, it is considered that the point cloud point belongs to the dynamic object. The point cloud and image pixels corresponding to the dynamic object determined in step S124 are removed. The static environment point cloud and the static environment image are obtained, providing clean data for subsequent feature extraction and matching.

[0108] Step S13: Extract multi-scale plane features from the static environment point cloud to obtain multi-scale plane information;

[0109] In the embodiments of the present invention, the point cloud space is segmented into voxel grids of different sizes. The size of the voxels can be adjusted according to the actual application scenario. By using voxels of different sizes, the extraction of multi-scale plane features can be achieved. For each voxel, its normal vector is calculated. The PCA (Principal Component Analysis) algorithm can be used to estimate the normal vector of the voxel. For all the points in the voxel, its covariance matrix is calculated, and the covariance matrix is subjected to eigenvalue decomposition. The eigenvector corresponding to the minimum eigenvalue is the normal vector of the voxel. Planarity indicates whether the distribution of the point cloud in the voxel is close to a plane. The eigenvalue can be used to calculate the planarity. If the ratio of the minimum eigenvalue to the middle eigenvalue and the maximum eigenvalue is close to 0, it is considered that the voxel is a plane. According to the normal vector and planarity of the voxel, multi-scale plane information is constructed. The normal vectors and planarities of voxels of different scales are stored in a data structure. For example, a multi-scale pyramid structure can be used to organize this information. For each scale, the normal vector, planarity of the voxel, and the center coordinates of the voxel are stored.

[0110] Step S14: Extract multi-scale edge and corner features from the static environment point cloud to obtain multi-scale geometric features;

[0111] In the embodiments of the present invention, the static environment point cloud is voxelized. Similar to step S13, voxel grids of different sizes are used to achieve multi-scale feature extraction. For each voxel, its point cloud density is calculated. The point cloud density can reflect the local structural features of this area. Calculate the number of points inside the voxel, and divide the number of points by the volume of the voxel to obtain the point cloud density. Surface curvature can be used to determine whether this area is a plane, an edge, or a corner. The normal vector of the point cloud can be used to calculate the surface curvature. For each point in the voxel, calculate the normal vectors of its neighboring points, and calculate the variance of the neighboring normal vectors. The larger the variance, the greater the surface curvature. According to the point cloud density and surface curvature, it is determined whether the voxel is an edge or a corner. For example, if a voxel has a large point cloud density and a large surface curvature, it is considered that the voxel may be a corner. If a voxel has a small point cloud density and a large surface curvature, it is considered that the voxel may be an edge. The edge, corner information, point cloud density, and surface curvature of voxels of different scales are stored in a data structure. For example, a multi-scale pyramid structure can be used to organize this information. For each scale, the edge, corner flags, point cloud density, surface curvature of the voxel, and the center coordinates of the voxel are stored.

[0112] Step S15: Extract image features from the static environment image to obtain image feature information;

[0113] In an embodiment of the present invention, an image feature extraction algorithm is selected to detect key points in an image. Descriptors of the key points are calculated. Image feature information is constructed. The positions, descriptors, and other relevant information of the key points are stored in a data structure. For example, a key point list can be used to store the information of all key points. The information of each key point includes its coordinates, descriptor, scale, and direction in the image, etc.

[0114] Step S16: Project the multi-scale plane information and multi-scale geometric features into the image feature information to obtain a fused feature set;

[0115] In an embodiment of the present invention, the initial values of the internal parameter matrix and external parameter matrix of the lidar and camera are obtained. The internal parameter matrix describes the internal geometric characteristics of the camera, such as focal length and principal point. The external parameter matrix describes the relative position and attitude between the lidar coordinate system and the camera coordinate system. These parameters can be obtained through manual measurement or a simple calibration process and stored in the ROS parameter server. For each key image feature, through the initial external parameter matrix, the plane information and geometric features in the lidar point cloud are projected onto the image plane. During the projection process, first, the plane information (normal vector, flatness) of the lidar, and edge and corner features (point cloud density, surface curvature, edge, corner flag) are transformed into the camera coordinate system. For each feature point of the lidar, using the internal parameter matrix, it is projected onto the image plane. After projection, the image features are associated with the projected lidar features. For example, if the distance between an image feature point and a projected lidar plane feature point on the image plane is less than a predefined threshold, it is considered that the image feature point corresponds to the lidar plane feature point. Similarly, the image features can also be associated with the projected lidar edge and corner features. A fused feature set is constructed. The fused feature set contains image features, projected lidar features, and the association relationships between them. A data structure, such as a hash table, can be used to store the fused feature set. The key of the hash table is the ID of the image feature, and the value is a list containing the lidar features associated with the image feature. The fused feature set provides multi-modal feature information for subsequent steps.

[0116] Step S17: Perform adaptive feature enhancement on the fused feature set to obtain a weighted fused feature set;

[0117] In the embodiments of the present invention, the confidence of the fused feature set is evaluated. The confidence of a feature reflects the reliability of the feature. For image features, the confidence can be calculated based on the response value, scale, etc. of the feature. For lidar features, the confidence can be calculated based on its flatness, point cloud density, surface curvature, etc. It is also possible to combine image features and lidar features. For example, if an image feature corresponds to a lidar plane feature and the lidar plane feature has a high flatness, it can be considered that the confidence of the image feature is also high. The environmental stability is evaluated. The environmental stability describes the dynamic degree of the current environment. The environmental stability can be evaluated by calculating the voxel displacement change within the point cloud motion time window. The greater the displacement change, the more unstable the environment. It is also possible to calculate the update frequency of the image background. The higher the update frequency, the more unstable the environment. Considering the voxel displacement change and the image background update frequency comprehensively, the environmental stability evaluation value is obtained. The initial judgment of the scene category is performed. The scene category can affect the weight of the features. For example, in a tunnel scene, the weight of the image features should be reduced, while the weight of the lidar features should be increased. The initial judgment of the scene category can be performed based on the odometer information and the image information at the current moment. For example, if the odometer information shows that the vehicle is inside the tunnel, it can be determined that the current scene is a tunnel scene. The features of the fused feature set are weighted according to the feature confidence, the environmental stability evaluation value, and the initial judgment data of the scene category. For each feature, a weight is calculated based on its confidence, the environmental stability evaluation value, and the scene category. The weight calculation can adopt the method of weighted average. For example, if a feature has a high confidence, high environmental stability, and the current scene is a static scene, the weight of this feature should be relatively high. The calculated weight is applied to the feature descriptor in the fused feature set. The weighted fused feature set is obtained, where each feature has its corresponding weight, preparing for subsequent matching and optimization.

[0118] Step S18: Generate a multi-modal feature map from the weighted fused feature set to obtain a multi-modal feature map;

[0119] In an embodiment of the present invention, a data structure of a multimodal feature map is created. The data structure can be a map in which weighted fusion feature sets at different times are stored. The map can use the nav_msgs / OccupancyGrid message type of ROS, or a custom message type. The weighted fusion feature set is added to the multimodal feature map. For each frame of data, the information of all features in the weighted fusion feature set, including the feature ID, projection coordinates under different sensor perspectives, descriptors, confidence, weights, etc., is stored in the multimodal feature map. When storing, the feature ID can be used as the key and the feature information as the value. A spatial index structure, such as a KD tree, can also be used to organize the multimodal feature map for quick search and access to features. The multimodal feature map is updated and maintained. As data is continuously input, the multimodal feature map needs to be continuously updated, for example, by adding new frame data and deleting old frame data. Maintaining the multimodal feature map can ensure the integrity and consistency of its data. The multimodal feature map contains the fused geometric features, image features, feature descriptors, confidence, and the projection coordinates of each feature under different sensor perspectives.

[0120] Preferably, step S12 comprises the following steps:

[0121] Step S121: performing point cloud motion estimation and segmentation on the compensated point cloud according to the odometer information to obtain dynamic point cloud and static point cloud segmentation results;

[0122] Step S122: performing image background modeling on the original camera image and performing foreground segmentation to obtain a foreground image;

[0123] Step S123: performing optical flow calculation on the original camera image and performing dynamic area detection to obtain a dynamic area mask;

[0124] Step S124: projecting the dynamic point cloud into the foreground image, determining the dynamic object according to the static point cloud segmentation result and the dynamic area mask, and obtaining the final dynamic object;

[0125] Step S125: According to the final dynamic object, dynamic objects are eliminated from the compensated point cloud and the original camera image to obtain a static environment point cloud and a static environment image.

[0126] In the embodiments of the present invention, the pose transformation between the current frame and the previous frame of point cloud is calculated. The ICP (Iterative Closest Point) algorithm is adopted, with the current frame of point cloud as the target point cloud and the previous frame of point cloud after pose transformation as the source point cloud, and the rigid transformation matrix between the two frames of point cloud is iteratively calculated. At the same time, according to the odometer information, the pose change of the vehicle between the two frames of point cloud is obtained. The rigid transformation matrix obtained by the ICP algorithm is compared with the pose change provided by the odometer. If the difference between the two is too large, it indicates that there are dynamic objects in the environment, and the pose estimation result obtained by the ICP algorithm is unreliable. Point cloud motion segmentation is performed. The point cloud is divided into multiple voxels. The ICP algorithm is used to calculate the displacement of each voxel between adjacent frames. If the displacement of a voxel exceeds a predefined threshold, the voxel is considered dynamic. The selection of the threshold needs to be adjusted according to the actual application scenario. It is also possible to use the pose change provided by the odometer to transform the previous frame of point cloud into the coordinate system of the current frame, and then calculate the distance between the corresponding voxels of the two frames of point cloud. If the distance is greater than the threshold, the voxel is considered dynamic. It is also possible to combine the intensity information of the point cloud. If there is an obvious change in the point cloud intensity, it may also indicate that the point cloud point belongs to a dynamic object. According to the judgment result of the voxel dynamics, the point cloud is segmented into a dynamic point cloud and a static point cloud. The point cloud points in the voxels marked as dynamic are divided into the dynamic point cloud. The point cloud points in the voxels not marked as dynamic are divided into the static point cloud. The static point cloud segmentation result and the dynamic point cloud are obtained, providing a basis for subsequent processing.

[0127] Background modeling is performed on the original camera image, and foreground segmentation is performed to identify dynamic objects in the image. A background modeling method based on time series is adopted, for example, the Gaussian Mixture Model (GMM). For each pixel, a GMM is established to describe the color distribution of the pixel at different time periods. The GMM consists of multiple Gaussian distributions, and each Gaussian distribution has a mean, variance, and weight. When the first frame of image arrives, the colors of all pixels are initialized as the background. For each subsequent frame of image, the matching degree between the color of each pixel and the GMM is calculated. If the color of the pixel matches one of the Gaussian distributions in the GMM, the pixel is considered to belong to the background, and the parameters of the GMM, including the mean, variance, and weight, are updated. If the color of the pixel does not match any of the Gaussian distributions in the GMM, the pixel is considered to belong to the foreground, and a new Gaussian distribution is created, or the Gaussian distribution with the smallest weight in the GMM is updated with the color of the pixel. Foreground segmentation can be achieved by setting a threshold. For each pixel, if the color difference between it and the background exceeds a threshold, the pixel is considered to belong to the foreground. The selection of the threshold needs to be adjusted according to the actual application scenario. The foreground image is obtained. In the foreground image, foreground pixels are usually marked as white and background pixels are marked as black.

[0128] The optical flow algorithm, such as the Lucas-Kanade algorithm, is used to calculate the pixel motion vectors between two adjacent frames of images. The Lucas-Kanade algorithm is based on the following assumptions: the gray value of the same pixel in adjacent frames remains unchanged; the motion vectors of adjacent pixels are the same. For each pixel in the image, calculate its motion vector in adjacent frames, that is, the optical flow. Calculating the optical flow requires solving a system of linear equations, which is composed of the gray value changes and gradient information of the pixels. An iterative method can be used to solve this system of linear equations. After obtaining the optical flow field, perform dynamic region detection. Analyze the motion vectors of the optical flow field, and calculate the amplitude and direction of the optical flow vectors. If the amplitude of the optical flow vectors in a certain region is large, or the direction of the optical flow vectors changes violently, it is considered that there may be dynamic objects in this region. Set a threshold. If the amplitude of the optical flow vector exceeds this threshold, or the direction change of the optical flow vector exceeds this threshold, it is considered that this pixel belongs to the dynamic region. Generate a dynamic region mask. The dynamic region mask is a binary image with the same size as the original image. If a pixel belongs to the dynamic region, set the value of this pixel to 1 (white) in the dynamic region mask, otherwise set it to 0 (black).

[0129] In the embodiment of the present invention, project the dynamic point cloud obtained in step S121 onto the image plane. Projection requires the use of the camera intrinsic matrix and extrinsic matrix. For each point in the dynamic point cloud, convert it from the point cloud coordinate system to the camera coordinate system, and then use the intrinsic matrix to project it onto the image plane. Combine the static point cloud segmentation result, the foreground image, and the dynamic region mask to determine the dynamic object. For each point projected onto the image plane, determine whether it belongs to a dynamic object. If a point cloud point is marked as a dynamic object, and its projected point on the image plane is located in the foreground region of the foreground image and within the dynamic region of the dynamic region mask, then it is considered that this point cloud point belongs to the final dynamic object. If this point cloud point only meets one or two of the above conditions, it needs to be judged according to the actual situation. For example, a confidence threshold can be set. If the confidence exceeds the threshold, it is considered that this point cloud point belongs to the dynamic object. Obtain the final dynamic object. The final dynamic object includes point cloud points and image pixels. These point cloud points and image pixels all belong to the dynamic object and need to be removed in subsequent steps.

[0130] Perform dynamic object removal on the compensated point cloud. For each point in the compensated point cloud, determine whether it belongs to the final dynamic object. This can be achieved by judging whether the point cloud point is within the range of the final dynamic object point cloud, or whether the projected point of the point cloud point on the image plane is within the range of the final dynamic object image pixels. If the point cloud point belongs to the final dynamic object, remove it from the compensated point cloud. Perform dynamic object removal on the original camera image. For each pixel in the original camera image, determine whether it belongs to the final dynamic object. This can be achieved by judging whether the pixel is within the range of the final dynamic object image pixels. If the pixel belongs to the final dynamic object, remove it from the original camera image. A smoother removal method can also be adopted. For example, use the color values of neighboring pixels to fill the removed pixels. Finally, obtain the static environment point cloud and the static environment image. The static environment point cloud only contains the point cloud points in the static environment, excluding the dynamic objects. The static environment image only contains the image pixels in the static environment, excluding the dynamic objects. The static environment point cloud and the static environment image provide clean data for subsequent feature extraction and matching.

[0131] Preferably, step S17 includes the following steps:

[0132] Step S171: Perform feature confidence evaluation on the fused feature set to obtain the feature confidence;

[0133] Step S172: Based on the fused feature set, calculate the voxel displacement change within the point cloud motion time window to obtain the voxel displacement change data; calculate the image background update frequency based on the fused feature set; perform environmental stability evaluation based on the voxel displacement change data and the image background update frequency to obtain the environmental stability evaluation value;

[0134] Step S173: Perform an initial judgment on the scene type based on the fused feature set to obtain the initial scene category judgment data;

[0135] Step S174: Perform feature weighting on the fused feature set according to the feature confidence, the environmental stability evaluation value, and the initial scene category judgment data to obtain the weighted fused feature set.

[0136] In the embodiments of the present invention, for image features, such as key points extracted by algorithms such as SIFT and ORB, the confidence is calculated using their inherent characteristics. For SIFT features, the confidence can be evaluated according to the response value of the feature. The higher the response value, the more prominent the feature point in the image and the higher the confidence. A linear function can be set to map the response value to a confidence value between 0 and 1. For ORB features, the confidence can be evaluated according to the corner response value of the feature. The higher the corner response value, the higher the confidence. The scale information of the feature can also be combined. The larger the scale, the larger the coverage area of the feature point in the image and the relatively lower the confidence. For lidar features, such as plane features, the confidence is calculated according to their planarity. The higher the planarity, the closer the area is to a plane and the higher the confidence. The planarity value can be used as the confidence value, or a linear function can be set to map the planarity value to a confidence value between 0 and 1. For edge and corner features, the confidence is evaluated according to their point cloud density and surface curvature. The higher the point cloud density and the larger the surface curvature, the more likely the area is to be an edge or a corner and the higher the confidence. The point cloud density and surface curvature can be weighted and summed to obtain a comprehensive confidence value. For fused features, the association between image features and lidar features is considered. If an image feature corresponds to a lidar plane feature and the planarity of the lidar plane feature is high, then the confidence of the image feature can be considered to be high. The confidence of the image feature and the confidence of the lidar plane feature can be weighted and fused to obtain a comprehensive confidence value. Finally, each feature is associated with a confidence value, providing a basis for subsequent weighted processing.

[0137] Calculate the voxel displacement change within the motion time window of the point cloud. Select a time window, for example, the past 5 frames of point cloud data. Divide the point cloud into multiple voxels. For each voxel, calculate its displacement change within the time window. The ICP algorithm can be used to estimate the rigid transformation between two adjacent frames of point cloud, and then calculate the voxel displacement according to the rigid transformation. The odometry information can also be used to calculate the displacement of the voxel between different frames. Calculate the standard deviation of the displacement of each voxel within the time window. The larger the standard deviation, the more unstable the motion of the voxel. Average the standard deviations of the displacements of all voxels to obtain an average voxel displacement change value. Calculate the image background update frequency. Use the GMM background model constructed in step S122. Calculate the background update frequency of each pixel. The background update frequency refers to the frequency at which the color of the pixel does not match the existing Gaussian distribution in the GMM and a new Gaussian distribution needs to be created or the existing Gaussian distribution needs to be updated. Count the number of background updates of each pixel over a period of time, divide the number of updates by the length of the time window to obtain the background update frequency. Average the background update frequencies of all pixels to obtain an average image background update frequency. Conduct environmental stability assessment. Considering the average voxel displacement change value and the average image background update frequency comprehensively, obtain the environmental stability assessment value. The weighted average method can be adopted. For example, a weight can be set, multiply the average voxel displacement change value by this weight, multiply the average image background update frequency by another weight, and then add the two results to obtain the environmental stability assessment value. A threshold can also be set. If the average voxel displacement change value or the average image background update frequency exceeds this threshold, the environment is considered unstable and the environmental stability assessment value is low. Obtain the environmental stability assessment value, which reflects the dynamic degree of the current environment.

[0138] Extract information from the fused feature set, such as odometer information, image information, point cloud information, etc. According to the odometer information, it can be judged whether the vehicle is in scenarios such as tunnels, highways, urban roads, etc. According to the image information, features such as road markings, traffic signs, traffic lights, etc. can be extracted, and these features can be used to judge the current scenario. According to the point cloud information, features such as road surface flatness and building features can be extracted, and these features can also be used to judge the current scenario. Construct scene type judgment rules. The scene type judgment rule is an if-else logical judgment. According to the extracted feature information, the type of the current scene is judged. For example, if the odometer information shows that the vehicle speed is high and there are road markings in the image, it can be judged that the current scene is a highway scene. If the odometer information shows that the vehicle is inside a tunnel, it can be judged that the current scene is a tunnel scene. If the odometer information shows that the vehicle speed is low and there are buildings in the image, it can be judged that the current scene is an urban road scene. Finally, obtain the initial judgment data of the scene category. The initial judgment data of the scene category can be a category label, such as highway, urban road, tunnel, etc. It can also be a probability distribution, indicating the probability that the current scene belongs to different categories.

[0139] For each feature, obtain its feature confidence, environmental stability evaluation value, and the initial judgment data of the scene category. According to this information, calculate the weight of the feature. The weight calculation can adopt the method of weighted average. For example, different weights can be set, corresponding to the feature confidence, environmental stability evaluation value, and the initial judgment data of the scene category respectively. For an image feature, if its confidence is high, the environmental stability is high, and the current scene is an urban road scene, then the weight of this image feature should be high. For a lidar feature, if its confidence is high, the environmental stability is high, and the current scene is a tunnel scene, then the weight of this lidar feature should be high. Different weights can be adjusted according to the actual application scenario. Apply the calculated weights to the feature descriptors in the fused feature set. Multiply the feature descriptor by its corresponding weight. Obtain the weighted fused feature set, where each feature is accompanied by its corresponding weight. The weighted fused feature set prepares for subsequent matching and optimization.

[0140] Preferably, step S2 includes the following steps:

[0141] Step S21: Perform a preliminary match on the multi-modal feature map based on geometric constraints to obtain a candidate match set;

[0142] Step S22: Perform a fine match on the candidate match set based on descriptors to obtain a preliminary match set;

[0143] Step S23: Perform motion compensation on the preliminary match set to obtain a compensated match set;

[0144] Step S24: Remove outliers and optimize the compensation matching set according to the multi-modal feature map to obtain a correspondence set.

[0145] In the embodiment of the present invention, feature information is extracted from the multi-modal feature map. For example, the positions of image features, the positions of lidar features, etc. For each image feature, search for the lidar feature that may correspond to it in the multi-modal feature map. Use geometric constraints to perform preliminary matching on the features. Geometric constraints can be based on the distance, angle, etc. between the features. For an image feature and a lidar feature, if the distance between them in space is less than a predefined threshold, they are considered potential matching pairs. The direction information of the features can also be used. For example, the gradient direction of the image feature and the normal vector direction of the lidar plane feature. If the angle between the two directions is less than a threshold, they are considered potential matching pairs. Multi-scale information can also be combined. For example, if the scale of the image feature matches the voxel size of the lidar feature, they are considered potential matching pairs. Construct a candidate matching set.

[0146] Extract feature descriptors from the multi-modal feature map. For each matching pair in the candidate matching set, obtain the descriptor of its image feature and the descriptor of the lidar feature. The descriptor is a numerical representation of the local features of the feature, used to describe the appearance of the feature. Calculate the similarity of the descriptors. Different measurement methods can be used for the similarity of the descriptors, such as Euclidean distance, Hamming distance, etc. If Euclidean distance is used, calculate the Euclidean distance between the two descriptors. If Hamming distance is used, calculate the Hamming distance between the two descriptors. The smaller the similarity value, the more similar the two descriptors are, and the higher the matching possibility. Perform descriptor matching. Set a similarity threshold. If the similarity of two descriptors is less than this threshold, this matching pair is considered a correct match. The ratio test can also be used to calculate the similarity ratio between the best match and the second-best match. If this ratio is less than a threshold, this matching pair is considered a correct match. Construct a preliminary matching set.

[0147] Obtain odometry information. Extract the pose information of the vehicle from the nav_msgs / Odometry message of ROS, including position and orientation. Obtain the pose information of the current frame and the previous frame. Calculate the pose transformation between two frames of point clouds. Use the odometry information to calculate the pose change of the current frame relative to the previous frame. Perform motion compensation. For each matching pair in the preliminary matching set, project it back to the coordinate system of the previous frame. If an image feature and a lidar feature match in the current frame, use the pose transformation to project this image feature and lidar feature back to the coordinate system of the previous frame. In this way, the influence of vehicle motion on the matching can be eliminated. Construct a compensation matching set.

[0148] Using the geometric information in the multi-modal feature map to eliminate outliers. For each matching pair in the compensated matching set, calculate its reprojection error in the image plane. The reprojection error refers to the distance between the lidar feature projected onto the image plane and the position of the image feature in the image plane. If the reprojection error is greater than a threshold, then this matching pair is considered an outlier and is removed from the compensated matching set. Using the feature confidence in the multi-modal feature map to weight the matching. For each matching pair in the compensated matching set, obtain the confidence of its image feature and lidar feature. The confidence of the matching can be weighted according to the confidence of the image feature and lidar feature. Matching pairs with high confidence should have higher weights. Perform matching optimization. Adopt an optimization algorithm, for example, weighted least squares method, to optimize the matching. The optimization goal is to minimize the reprojection error and consider the confidence of the features. Construct a correspondence set.

[0149] Preferably, step S3 includes the following steps:

[0150] Step S31: Perform a rough pose estimation based on RANSAC on the correspondence set to obtain rough pose estimation data and an inlier set;

[0151] Step S32: Perform global weighted least squares optimization according to the multi-modal feature map, the rough pose estimation data, and the inlier set to obtain the globally optimized pose;

[0152] Step S33: Perform scene-based local optimization according to the multi-modal feature map, the globally optimized pose, and the inlier set to obtain the locally optimized pose;

[0153] Step S34: Perform iterative optimization according to the locally optimized pose, the correspondence set, and the multi-modal feature map to obtain the calibration parameters.

[0154] In an embodiment of the present invention, a set of matching pairs is randomly selected from the correspondence set. For example, 3 pairs of matching pairs are selected. Each matching pair contains the ID of the image feature and the ID of the lidar feature. According to the selected matching pairs, the pose transformation between the lidar and the camera is calculated. Since there are only 3 pairs of matching pairs, a unique pose transformation can be determined. The pose transformation includes a rotation matrix and a translation vector. A linear method can be used to solve the pose transformation. The calculation method is to project the lidar feature onto the image plane, and then calculate the residual between the projected point and the corresponding image feature. According to the principle of minimizing the residual, the pose transformation is solved. Using the calculated pose transformation, all the matching pairs in the correspondence set are verified. The lidar feature is projected onto the image plane, and the reprojection error between the projected point and the corresponding image feature is calculated. If the reprojection error is less than a predefined threshold, the matching pair is considered an inlier. Otherwise, the matching pair is considered an outlier. Repeat the above process, randomly select multiple sets of matching pairs, calculate the pose transformation, and verify all the matching pairs. Select the pose transformation with the largest number of inliers as the best pose transformation. Construct the rough pose estimation data and the inlier set.

[0155] Feature information is extracted from the multi-modal feature map, such as the position of the image feature, the position of the lidar feature, etc. The matching pair information is obtained from the inlier set. Each matching pair contains the ID of the image feature and the ID of the lidar feature. An optimization objective function is constructed. The optimization objective function is the weighted sum of the reprojection errors. The reprojection error refers to the distance between the lidar feature projected onto the image plane and the position of the image feature on the image plane. Weighting means weighting the reprojection error according to the confidence of the feature. Features with higher confidence should have higher weights. The rough pose estimation data is used as the initial value of the optimization algorithm. The optimization algorithm can adopt the Gauss-Newton method or the Levenberg-Marquardt algorithm. The pose transformation is iteratively optimized to minimize the optimization objective function. In each iteration, the reprojection error is calculated, the gradient and the Hessian matrix are calculated, and the pose transformation is updated. Construct the globally optimized pose.

[0156] Extract local scene data from the multi-modal feature map. The local scene data includes scene odometer data, scene image data, and scene point cloud data, which correspond to specific scenes, such as highways, urban roads, tunnels, etc. Perform scene recognition. Use the scene recognition model to recognize the current scene. The scene recognition model can be based on machine learning algorithms, such as support vector machines (SVM). The training data of the scene recognition model comes from the multi-modal feature map. The training data contains the feature information of different scenes and the corresponding scene category labels. The scene recognition model can identify the category of the current scene, such as highways, urban roads, tunnels, etc. According to the scene category, construct a local optimization problem. For highway scenes, road markings can be used as constraints. Extract the road marking features in the image, such as lane lines, edge lines, etc. Associate the road marking features with the plane features in the lidar point cloud. The optimization objective function can include the reprojection error and the parallelism constraint of the road markings. For urban road scenes, building corners and curbs can be used as constraints. Extract the building corner and curb features in the image, and associate the building corner and curb features with the corner and edge features in the lidar point cloud. The optimization objective function can include the reprojection error and the perpendicularity constraint of the building corners and curbs. For tunnel scenes, the constraints of the tunnel wall can be used. Extract the tunnel wall features in the lidar point cloud. The optimization objective function can include the reprojection error and the parallelism constraint of the tunnel wall. Use the globally optimized pose to solve the local optimization problem. The optimization algorithm can use the Gauss-Newton method or the Levenberg-Marquardt algorithm. Iteratively optimize the pose transformation to minimize the objective function of the local optimization problem. Construct the local optimized pose.

[0157] Use the locally optimized pose as the initial pose. Use the locally optimized pose to calculate the reprojection error again. Use the reprojection error to update the inlier set. Set a reprojection error threshold. If the reprojection error is less than this threshold, then this matching pair is considered an inlier. Use the inlier set to recalculate the pose transformation. Use the multi-modal feature map and the newly calculated pose transformation to perform global weighted least squares optimization to obtain a new globally optimized pose. Use the new globally optimized pose to perform scene-based local optimization to obtain a new locally optimized pose. Repeat the above process for iterative optimization. The iteration stop condition can be that the change in the reprojection error is less than a threshold, or the number of iterations reaches a maximum value. Obtain the constructed calibration parameters.

[0158] Preferably, step S33 includes the following steps:

[0159] Step S331: Extract local scene data from the multi-modal feature map to obtain local scene data, where the local scene data includes scene odometer data, scene image data, and scene point cloud data;

[0160] Step S332: Extract odometer features from the scene odometer data to obtain odometer features;

[0161] Step S333: Detect road markings in the scene image data and detect curbs to obtain road image features;

[0162] Step S334: Calculate the road surface flatness from the scene point cloud data to obtain the road surface flatness; extract building features from the scene point cloud data to obtain building features; extract point cloud density features from the scene point cloud data to obtain point cloud density features; generate road point cloud features based on the road surface flatness, building features, and point cloud density features to obtain road point cloud features;

[0163] Step S335: Construct a scene recognition model based on a support vector machine according to the odometer features, road image features, and road point cloud features to obtain a scene recognition model; use the scene recognition model to perform scene model recognition to obtain the model recognition scene categories, where the model recognition scene categories include highway scenes, urban road scenes, and tunnel scenes;

[0164] Step S336: Construct a constraint optimization problem based on road markings for the highway scene to obtain a highway scene optimization problem; construct a constraint optimization problem based on building corners and curbs for the urban road scene to obtain an urban road scene optimization problem; construct a constraint optimization problem based on tunnel walls for the tunnel scene to obtain a tunnel scene optimization problem;

[0165] Step S337: Perform non-linear optimization on the highway scene optimization problem, urban road scene optimization problem, and tunnel scene optimization problem according to the global optimization pose and the inlier set to obtain the local optimization pose.

[0166] In an embodiment of the present invention, data is obtained from a multi-modal feature map according to the current time. The obtained data includes odometry information, image information, and point cloud information of the current frame. A time window is constructed. The length of the time window can be adjusted according to the actual application scenario. For example, it can be 5 seconds. Odometry data, image data, and point cloud data within the time window are obtained from the multi-modal feature map. The odometry data within the time window includes pose information and speed information. The image data within the time window includes static environment images. The point cloud data within the time window includes static environment point clouds. Scene local data is obtained. The scene local data contains the odometry data, image data, and point cloud data within the time window, and these data correspond to a specific scene, providing a basis for subsequent scene recognition and local optimization.

[0167] The pose information within the time window is obtained from the scene odometry data. The pose information includes the position and attitude of the vehicle. The driving speed and acceleration of the vehicle are calculated. The speed of the vehicle can be obtained by calculating the pose change between adjacent moments. The acceleration of the vehicle can be obtained by calculating the speed change between adjacent moments. The turning rate of the vehicle is calculated. The turning rate of the vehicle can be obtained by calculating the change rate of the vehicle's attitude. An odometry feature is constructed. The odometry feature includes the average speed, acceleration, turning rate, etc. of the vehicle.

[0168] The scene image data is preprocessed. The preprocessing includes operations such as image denoising and image enhancement. Gaussian filtering can be used for image denoising. Histogram equalization can be used for image enhancement. Road marking detection is performed. A method based on edge detection is used to detect road markings in the image. The Canny operator can be used for edge detection. After extracting the edges, the Hough transform is used to detect lines, and the detected lines are used as road markings. A deep learning-based road marking detection method can also be used. For example, a semantic segmentation network is used to segment the road markings in the image. Curb detection is performed. A pattern matching method is used to detect curbs in the image. The template of the curb can be defined first. Then, the area matching the template is searched in the image. A deep learning-based curb detection method can also be used. For example, an object detection network is used to detect curbs in the image. Road image features are constructed. The road image features include the number, direction, curvature, etc. of the road markings. The road image features also include the position, size, shape, etc. of the curbs.

[0169] Perform road flatness calculation. Segment the point cloud into multiple grids. For each grid, calculate its normal vector. The PCA algorithm can be used to calculate the normal vector. Calculate the variance of the normal vectors of the point cloud within the grid. The smaller the variance, the flatter the area. Perform building feature extraction. The method based on point cloud segmentation can be used to segment the point cloud into multiple clusters. For each cluster, calculate its geometric features, such as length, width, height, volume, etc. The building detection method based on deep learning can also be used, such as using a point cloud segmentation network to segment the buildings in the point cloud. Perform point cloud density feature extraction. Segment the point cloud into multiple grids. Calculate the number of points in each grid. Divide the number of points by the volume of the grid to obtain the point cloud density. Construct road point cloud features.

[0170] Prepare training data. The training data includes feature data of different scenarios and corresponding scenario category labels. The scenario category labels include highway scenarios, urban road scenarios, and tunnel scenarios. The feature data includes odometer features, road image features, and road point cloud features. Construct a scenario recognition model. Adopt the Support Vector Machine (SVM) algorithm. The SVM algorithm is a supervised learning algorithm used for classification and regression. Input the training data into the SVM algorithm for model training. During the training process, the SVM algorithm learns the boundaries between the features of different scenarios and constructs a classifier. Different kernel functions can be selected, such as linear kernel function, radial basis kernel function, etc. Obtain the scenario recognition model. Use the scenario recognition model for scenario model recognition. Input the odometer features, road image features, and road point cloud features extracted in steps S332, S333, and S334 into the scenario recognition model. The scenario recognition model will predict the category of the current scenario. Obtain the model recognition scenario category.

[0171] For highway scenarios, construct a constraint optimization problem based on road markings. Extract road marking features in the image, such as lane lines, edge lines, etc. Associate the road marking features with the plane features in the lidar point cloud. The optimization objective function can include reprojection error and the parallelism constraint of road markings. The parallelism constraint means that road markings should be parallel on the image plane. For urban road scenarios, construct a constraint optimization problem based on building corners and curbs. Extract building corner and curb features in the image. Associate the building corner and curb features with the corner and edge features in the lidar point cloud. The optimization objective function can include reprojection error and the perpendicularity constraint of building corners and curbs. The perpendicularity constraint means that building corners and curbs should be perpendicular on the image plane. For tunnel scenarios, construct a constraint optimization problem based on tunnel walls. Extract tunnel wall features in the lidar point cloud. The optimization objective function can include reprojection error and the parallelism constraint of tunnel walls. The parallelism constraint means that tunnel walls should be parallel in the point cloud space.

[0172] Obtain the globally optimized pose. The globally optimized pose is the pose obtained in step S32. The locally optimized pose is a fine-tuning relative to the globally optimized pose. Obtain the inlier set. The inlier set is the set of inliers obtained in step S31, and the inlier set contains reliable matching points. Perform non-linear optimization. Use the Gauss-Newton method or the Levenberg-Marquardt algorithm to solve the local optimization problem. The optimization objective is to minimize the objective function of the local optimization problem. For example, for a highway scenario, the optimization objective is to minimize the reprojection error and the parallelism constraint of road markings. The optimization variable is the pose transformation between the lidar and the camera. Construct the locally optimized pose.

[0173] Preferably, step S4 includes the following steps:

[0174] Step S41: Perform reprojection error analysis based on the calibration parameters, the original lidar point cloud, and the original camera image to obtain reprojection error statistics;

[0175] Step S42: Perform scene analysis based on the calibration parameters, the original lidar point cloud, the original camera image, and the odometer information to obtain a scene performance evaluation;

[0176] Step S43: Perform key point alignment verification based on the calibration parameters, the original lidar point cloud, and the original camera image to obtain a key point alignment result;

[0177] Step S44: Perform robustness evaluation based on the calibration parameters, the original lidar point cloud, the original camera image, and the odometer information to obtain a robustness evaluation result;

[0178] Step S45: Generate a verification report for the reprojection error statistics, the scene performance evaluation, the key point alignment result, and the robustness evaluation result to obtain a verification report.

[0179] In the embodiments of the present invention, calibration parameters are obtained. The calibration parameters include the rotation matrix and translation vector of the lidar relative to the camera. These parameters are obtained in step S3. The original lidar point cloud and the original camera image are obtained. The original lidar point cloud data is obtained from the ROS platform, and its ROS message type is sensor_msgs / PointCloud2. The original camera image data is obtained, and its ROS message type is sensor_msgs / Image. The lidar point cloud is projected onto the image plane. For each point in the lidar point cloud, using the calibration parameters, it is transformed into the camera coordinate system. Then, using the intrinsic matrix of the camera, the point is projected onto the image plane. The intrinsic matrix describes the internal parameters of the camera, such as the focal length and the principal point. The reprojection error is calculated. The reprojection error is the distance between the projected point and the image feature. The image feature can be manually marked or automatically extracted. If manually marked, feature points corresponding to the lidar point cloud need to be marked on the image. If automatically extracted, features such as corners and edges in the image can be extracted and matched with the projected points. The Euclidean distance between the projected point and the corresponding image feature is calculated as the reprojection error. The reprojection errors are statistically analyzed. The mean, standard deviation, maximum value, and minimum value of all reprojection errors are calculated. A histogram of the reprojection errors is plotted. The reprojection error statistics are obtained.

[0180] Calibration parameters, the original lidar point cloud, the original camera image, and odometry information are obtained. The calibration parameters include the rotation matrix and translation vector of the lidar relative to the camera. The original lidar point cloud and the original camera image are obtained from the ROS platform. The odometry information is obtained from the ROS nav_msgs / Odometry message and includes the position and attitude of the vehicle. Scene recognition is performed. The scene recognition model in step S335 can be used to recognize the current scene. The scene recognition model can determine whether the current scene is a highway, an urban road, a tunnel, etc. For different scenes, performance evaluation is carried out. For the highway scene, the alignment degree of the road markings can be evaluated. The road surface points in the lidar point cloud are projected onto the image plane. The road markings in the image are detected. The distance between the projected points and the road markings is calculated. The smaller the distance, the better the calibration result. For the urban road scene, the alignment degree of the building edges can be evaluated. The building points in the lidar point cloud are projected onto the image plane. The building edges in the image are detected. The distance between the projected points and the building edges is calculated. The smaller the distance, the better the calibration result. For the tunnel scene, the alignment degree of the tunnel wall can be evaluated. The tunnel wall points in the lidar point cloud are projected onto the image plane. The tunnel wall can be marked on the image by manual marking. The distance between the projected points and the tunnel wall is calculated. The smaller the distance, the better the calibration result.

[0181] Obtain calibration parameters, the original lidar point cloud, and the original camera image. The calibration parameters include the rotation matrix and translation vector of the lidar relative to the camera. The original lidar point cloud and the original camera image are obtained from the ROS platform. Select key points in the image. The key points can be road markings, curbs, traffic signs, etc. The key points need to have obvious features so as to find the corresponding points in the lidar point cloud. Find the points in the lidar point cloud corresponding to the key points in the image. Project the key points in the image into the lidar coordinate system. For each key point in the image, use the calibration parameters to transform it into the lidar coordinate system. Then, search for the point closest to this point in the lidar point cloud. The alignment error refers to the distance between the coordinates of the key point in the image in the lidar coordinate system and the coordinates of the corresponding point in the lidar point cloud. Calculate the Euclidean distance as the alignment error. Construct the key point alignment result. The key point alignment result includes the position of the key point, the coordinates in the lidar coordinate system and the camera coordinate system, and the alignment error.

[0182] Obtain calibration parameters, the original lidar point cloud, the original camera image, and odometry information. The calibration parameters include the rotation matrix and translation vector of the lidar relative to the camera. The original lidar point cloud and the original camera image are obtained from the ROS platform. The odometry information is obtained from the ROS's nav_msgs / Odometry message, including the position and attitude of the vehicle. Repeat the calibration process under different environmental conditions. Different environmental conditions include: Light change: Repeat the calibration process under different lighting conditions, such as daytime, night, cloudy, etc. Dynamic background: Repeat the calibration process in the presence of dynamic objects, such as vehicles, pedestrians, etc. Driving speed: Repeat the calibration process at different driving speeds, such as low speed, high speed, etc. For each environmental condition, perform the following steps: Use the methods in steps S1 to S3 for calibration. Use the methods in steps S41 to S43 to evaluate the accuracy of the calibration result. Calculate the robustness evaluation index. The robustness evaluation index can include the reprojection error, the change in the key point alignment error, etc. Obtain the robustness evaluation result. The robustness evaluation result includes the accuracy index of the calibration result under different environmental conditions.

[0183] Obtain the reprojection error statistics, the scene performance evaluation, the key point alignment result, and the robustness evaluation result. These results are obtained in steps S41 to S44. Organize this evaluation information. Summarize the reprojection error, the scene performance evaluation index, the key point alignment error under different scenarios, and the robustness evaluation result under different environmental conditions. Generate a verification report.

[0184] Preferably, step S5 includes the following steps:

[0185] Step S51: Monitor the environment using the raw lidar point cloud, the raw camera image, and the odometry information to obtain environmental change information;

[0186] Step S52: Classify the current scene based on the environmental change information to obtain the current scene category;

[0187] Step S53: Adjust the feature weights according to the verification report, the calibration parameters, the environmental change information, and the current scene category to obtain an adjusted multi-modal feature map;

[0188] Step S54: Adjust the calibration parameters according to the adjusted multi-modal feature map to obtain adjusted calibration parameters;

[0189] Step S55: Perform scene-specific optimization according to the adjusted multi-modal feature map, the adjusted calibration parameters, and the current scene category to obtain optimized calibration parameters.

[0190] In the embodiment of the present invention, the raw lidar point cloud, the raw camera image, and the odometry information are obtained. The raw lidar point cloud is obtained from the ROS sensor_msgs / PointCloud2 message. The raw camera image is obtained from the ROS sensor_msgs / Image message. The odometry information is obtained from the ROS nav_msgs / Odometry message and includes the position and attitude of the vehicle. Then, environmental change monitoring is performed. The light change is monitored. The average gray value or brightness value of the image is calculated. The average gray value of the current frame image is compared with that of the historical frame image. If the change exceeds a threshold, it indicates that the light has changed. The histogram of the image can also be used to observe the change of the histogram to judge the light change. The appearance of dynamic objects is monitored. The dynamic object detection method in step S12 is used to detect the dynamic objects in the image and the point cloud. The number and movement speed of the dynamic objects are counted. If the number of dynamic objects increases or the movement speed accelerates, it indicates that there is a change in the dynamic objects. The road type change is monitored. The odometry information is used to judge whether the vehicle is in a highway, an urban road, a tunnel or other scenes. The road marking features in the image can also be used to judge the road type. The weather change is monitored. A weather sensor can be used to obtain weather information, such as sunny, rainy, snowy, etc. The image information, such as the clarity and contrast of the image, can also be used to judge the weather change. The environmental change information is constructed.

[0191] Obtain environmental change information. The environmental change information is obtained in step S51 and includes information such as light change, appearance of dynamic objects, road type change, and weather change. Construct scene classification rules. The scene classification rules are an if-else logical judgment. Based on the environmental change information, the category of the current scene is judged. For example: If the light changes drastically, the number of dynamic objects is small, and the road type is highway, then the current scene is judged to be the highway daytime scene. If the light changes drastically, the number of dynamic objects is small, the road type is highway, and the weather is rainy, then the current scene is judged to be the highway rainy scene. If the light change is small, the number of dynamic objects is large, and the road type is urban road, then the current scene is judged to be the urban road scene. If the light change is small, the number of dynamic objects is small, and the road type is tunnel, then the current scene is judged to be the tunnel scene. Obtain the current scene category.

[0192] Obtain the verification report, calibration parameters, environmental change information, and the current scene category. The verification report is generated in step S45 and includes reprojection error, scene performance evaluation, key point alignment result, and robustness evaluation result. The calibration parameters are obtained in step S34 and include the rotation matrix and translation vector. The environmental change information is obtained in step S51 and includes information such as light change, appearance of dynamic objects, road type change, and weather change. The current scene category is obtained in step S52. Then, based on this information, adjust the weights of the features. If the verification report shows a high reprojection error, it indicates that the accuracy of the calibration result is low. The weight of the feature can be reduced, or the confidence of the feature can be increased. If the environment changes drastically, for example, the light changes drastically, the weight of the image feature can be reduced, and the weight of the lidar feature can be increased. If the current scene is the tunnel scene, the weight of the image feature can be reduced, and the weight of the lidar feature can be increased. If the current scene is the highway scene, the weight of the road marking feature can be increased. If the current scene is the urban road scene, the weight of the building feature can be increased. Construct the adjusted multi-modal feature map.

[0193] Obtain the adjusted multi-modal feature map. The adjusted multi-modal feature map is obtained in step S53. Use the adjusted multi-modal feature map to recalculate the pose. The pose estimation and optimization method in step S3 can be used. Using the calibration parameters obtained in step S34 as the initial value, perform iterative optimization. Obtain the adjusted calibration parameters.

[0194] Obtain the adjusted multi-modal feature map, the adjusted calibration parameters, and the current scene category. The adjusted multi-modal feature map is obtained in step S53. The adjusted calibration parameters are obtained in step S54. The current scene category is obtained in step S52. If the current scene is a highway scene, the road markings can be used as a constraint. Extract the road marking features in the image, such as lane lines, edge lines, etc. Associate the road marking features with the plane features in the lidar point cloud. Construct an optimization objective function, which can include the reprojection error and the parallelism constraint of the road markings. Use the adjusted calibration parameters as the initial value for iterative optimization. If the current scene is an urban road scene, the building corners and curbs can be used as constraints. Extract the building corner and curb features in the image, and associate the building corner and curb features with the corner and edge features in the lidar point cloud. Construct an optimization objective function, which can include the reprojection error and the perpendicularity constraint of the building corners and curbs. Use the adjusted calibration parameters as the initial value for iterative optimization. If the current scene is a tunnel scene, the tunnel wall surface can be used as a constraint. Extract the tunnel wall surface features in the lidar point cloud. Construct an optimization objective function, which can include the reprojection error and the parallelism constraint of the tunnel wall surface. Use the adjusted calibration parameters as the initial value for iterative optimization. Obtain the optimized calibration parameters.

[0195] Preferably, the present invention also provides a ROS-based lidar and camera fusion calibration system for performing the ROS-based lidar and camera fusion calibration method described above. The ROS-based lidar and camera fusion calibration system includes:

[0196] An environmental perception module for obtaining the original lidar point cloud, the original camera image, and the odometer information through the ROS platform; performing dynamic object detection on the original lidar point cloud and the original camera image, and removing the dynamic objects to obtain the static environment point cloud and the static environment image; performing multi-scale environmental feature extraction on the static environment point cloud and the static environment image, and performing feature fusion to obtain a fusion feature set; performing adaptive feature enhancement processing on the fusion feature set to obtain a multi-modal feature map;

[0197] A motion feature matching module for performing feature matching on the multi-modal feature map and performing motion compensation to obtain a correspondence set;

[0198] A pose estimation and optimization module for performing pose estimation on the correspondence set to obtain rough pose estimation data; performing global pose optimization based on the rough pose estimation data, and performing scene-based local optimization to obtain a locally optimized pose; performing iterative optimization based on the locally optimized pose to obtain calibration parameters;

[0199] The calibration result evaluation module is used to evaluate the calibration parameters from multiple angles according to the calibration parameters and obtain a verification report;

[0200] The adaptive calibration module is used to monitor the environment of the original point cloud of the lidar and the original image of the camera, classify the current scene, and obtain the current scene category; according to the verification report, calibration parameters and the current scene category, adjust the multi-modal feature map and calibration parameters, and perform scene-specific optimization to obtain optimized calibration parameters.

[0201] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the application documents are intended to be encompassed within the present invention.

[0202] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features invented herein.

Claims

1. A calibration method for the fusion of lidar and camera based on ROS, characterized in that, It includes the following steps: Step S1 includes the following steps: Step S11: Obtain the original lidar point cloud, the original camera image, and the odometer information; perform point cloud motion compensation on the original lidar point cloud according to the odometer information to obtain the compensated point cloud; Step S12: Perform dynamic object detection on the compensated point cloud and the original camera image respectively, and perform dynamic object removal to obtain the static environment point cloud and the static environment image; Step S13: Extract multi-scale plane features from the static environment point cloud to obtain multi-scale plane information; Step S14: Extract multi-scale edge and corner features from the static environment point cloud to obtain multi-scale geometric features; Step S15: Extract image features from the static environment image to obtain image feature information; Step S16: Project the multi-scale plane information and the multi-scale geometric features into the image feature information to obtain a fused feature set; Step S17: Perform adaptive feature enhancement on the fused feature set to obtain a weighted fused feature set, where Step S17 includes the following steps: Step S171: Evaluate the feature confidence of the fused feature set to obtain the feature confidence; Step S172: Calculate the voxel displacement change based on the fused feature set within the point cloud motion time window to obtain the voxel displacement change data; calculate the image background update frequency based on the fused feature set to obtain the image background update frequency; perform environmental stability evaluation based on the voxel displacement change data and the image background update frequency to obtain the environmental stability evaluation value; Step S173: Perform an initial judgment on the scene type based on the fused feature set to obtain the initial scene category judgment data; Step S174: Perform feature weighting on the fused feature set according to the feature confidence, the environmental stability evaluation value, and the initial scene category judgment data to obtain a weighted fused feature set; Step S18: Generate a multi-modal feature map from the weighted fused feature set to obtain a multi-modal feature map; Step S2: Perform feature matching on the multi-modal feature map and perform motion compensation to obtain a correspondence set; Step S3: Perform pose estimation on the correspondence set to obtain rough pose estimation data; perform global pose optimization based on the rough pose estimation data, and perform scene-based local optimization to obtain the locally optimized pose; perform iterative optimization based on the locally optimized pose to obtain the calibration parameters; Step S4: Evaluate the calibration parameters from multiple angles according to the calibration parameters to obtain a verification report; Step S5: Monitor the environment of the original lidar point cloud and the original camera image, and perform current scene classification to obtain the current scene category; adjust the multi-modal feature map and the calibration parameters according to the verification report, the calibration parameters, and the current scene category, and perform scene-specific optimization to obtain the optimized calibration parameters.

2. The method for fusing and calibrating a lidar and a camera based on ROS according to claim 1, wherein, Step S12 includes the following steps: Step S121: Perform point cloud motion estimation and segmentation on the compensated point cloud according to the odometer information to obtain the dynamic point cloud and the static point cloud segmentation result; Step S122: Perform image background modeling on the original camera image and perform foreground segmentation to obtain the foreground image; Step S123: Calculate the optical flow of the original camera image and perform dynamic region detection to obtain a dynamic region mask; Step S124: Project the dynamic point cloud onto the foreground image, and determine the dynamic object based on the static point cloud segmentation result and the dynamic region mask to obtain the final dynamic object; Step S125: Remove the dynamic object from the compensated point cloud and the original camera image according to the final dynamic object to obtain the static environment point cloud and the static environment image.

3. The method for fusing and calibrating a lidar and a camera based on ROS according to claim 1, wherein, Step S2 includes the following steps: Step S21: Perform preliminary matching based on geometric constraints on the multi-modal feature map to obtain a candidate matching set; Step S22: Perform fine matching based on descriptors on the candidate matching set to obtain a preliminary matching set; Step S23: Perform motion compensation on the preliminary matching set to obtain a compensated matching set; Step S24: Remove outliers and optimize the compensated matching set according to the multi-modal feature map to obtain a correspondence set.

4. The method for fusing and calibrating a lidar and a camera based on ROS according to claim 1, wherein, Step S3 includes the following steps: Step S31: Perform rough pose estimation based on RANSAC on the correspondence set to obtain rough pose estimation data and an inlier set; Step S32: Perform global weighted least squares optimization according to the multi-modal feature map, the rough pose estimation data, and the inlier set to obtain the globally optimized pose; Step S33: Perform scene-based local optimization according to the multi-modal feature map, the globally optimized pose, and the inlier set to obtain the locally optimized pose; Step S34: Perform iterative optimization according to the locally optimized pose, the correspondence set, and the multi-modal feature map to obtain calibration parameters.

5. The method for fusing and calibrating a lidar and a camera based on ROS according to claim 4, wherein Step S33 includes the following steps: Step S331: Extract scene local data from the multi-modal feature map to obtain scene local data, where the scene local data includes scene odometer data, scene image data, and scene point cloud data; Step S332: Extract odometer features from the scene odometer data to obtain odometer features; Step S333: Detect road markings in the scene image data and detect curbs to obtain road image features; Step S334: Calculate the road surface flatness of the scene point cloud data to obtain the road surface flatness; extract building features from the scene point cloud data to obtain building features; extract point cloud density features from the scene point cloud data to obtain point cloud density features; generate road point cloud features according to the road surface flatness, the building features, and the point cloud density features to obtain road point cloud features; Step S335: Construct a scene recognition model based on a support vector machine according to the odometer features, the road image features, and the road point cloud features to obtain a scene recognition model; use the scene recognition model to perform scene model recognition to obtain the model recognition scene category, where the model recognition scene category includes highway scenes, urban road scenes, and tunnel scenes; Step S336: Construct the constraint optimization problem based on road markings for the highway scenario to obtain the highway scenario optimization problem; construct the constraint optimization problem based on building corners and curbs for the urban road scenario to obtain the urban road scenario optimization problem; construct the constraint optimization problem based on tunnel walls for the tunnel scenario to obtain the tunnel scenario optimization problem. Step S337: Perform non-linear optimization on the highway scenario optimization problem, urban road scenario optimization problem, and tunnel scenario optimization problem according to the global optimization pose and the inlier set to obtain the local optimization pose.

6. The method for fusing and calibrating a lidar and a camera based on ROS according to claim 1, wherein Step S4 includes the following steps: Step S41: Conduct reprojection error analysis based on calibration parameters, the original lidar point cloud, and the original camera image to obtain reprojection error statistics. Step S42: Conduct scene analysis based on calibration parameters, the original lidar point cloud, the original camera image, and odometer information to obtain scene performance evaluation. Step S43: Conduct key point alignment verification based on calibration parameters, the original lidar point cloud, and the original camera image to obtain the key point alignment result. Step S44: Conduct robustness evaluation based on calibration parameters, the original lidar point cloud, the original camera image, and odometer information to obtain the robustness evaluation result. Step S45: Generate a verification report based on the reprojection error statistics, scene performance evaluation, key point alignment result, and robustness evaluation result to obtain the verification report.

7. The method for fusing and calibrating a lidar and a camera based on ROS according to claim 1, characterized in that, Step S5 includes the following steps: Step S51: Conduct environmental monitoring on the original lidar point cloud, the original camera image, and odometer information to obtain environmental change information. Step S52: Classify the current scene according to the environmental change information to obtain the current scene category. Step S53: Adjust the feature weights according to the verification report, calibration parameters, environmental change information, and the current scene category to obtain the adjusted multi-modal feature map. Step S54: Adjust the calibration parameters according to the adjusted multi-modal feature map to obtain the adjusted calibration parameters. Step S55: Conduct scene-specific optimization according to the adjusted multi-modal feature map, the adjusted calibration parameters, and the current scene category to obtain the optimized calibration parameters.

8. A fusion calibration system for lidar and camera based on ROS, characterized in that, A ROS-based lidar and camera fusion calibration system for performing the ROS-based lidar and camera fusion calibration method as claimed in claim 1, the system comprising: An environmental perception module for obtaining the original lidar point cloud, the original camera image, and odometer information through the ROS platform; performing dynamic object detection on the original lidar point cloud and the original camera image, and removing the dynamic objects to obtain the static environment point cloud and the static environment image; performing multi-scale environmental feature extraction on the static environment point cloud and the static environment image, and performing feature fusion to obtain a fusion feature set; performing adaptive feature enhancement processing on the fusion feature set to obtain a multi-modal feature map. A motion feature matching module for performing feature matching on the multi-modal feature map and performing motion compensation to obtain a correspondence set. A pose estimation and optimization module is used to perform pose estimation on the correspondence set to obtain rough pose estimation data; perform global pose optimization based on the rough pose estimation data, and perform scene-based local optimization to obtain a locally optimized pose; perform iterative optimization based on the locally optimized pose to obtain calibration parameters; A calibration result evaluation module is used to perform multi-angle calibration parameter evaluation based on the calibration parameters to obtain a verification report; An adaptive calibration module is used to monitor the environment of the original lidar point cloud and the original camera image, and perform current scene classification to obtain the current scene category; perform multi-modal feature map and calibration parameter adjustment based on the verification report, calibration parameters, and the current scene category, and perform scene-specific optimization to obtain optimized calibration parameters.

Citation Information

Patent Citations

  • 3D target detection method and device based on fusion of laser radar and binocular camera

    CN117011388A

  • Vision and radar fused target positioning method and device

    CN118244281A