Multi-modal sensor external parameter failure sensing and online calibration method

By using a spatial attention model and an unsupervised scene flow training framework to correct point cloud distortion and estimate the velocity of moving targets, and combining this with a factor graph optimization model, the problem of multi-sensor extrinsic parameter drift under the vehicle platform was solved, achieving higher accuracy and robustness in online calibration.

CN121304801APending Publication Date: 2026-01-09JIMEI UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511369368.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

On vehicle platforms, the extrinsic parameters of multi-sensor systems are easily affected by complex road conditions, leading to drift. Existing online calibration methods are difficult to effectively detect extrinsic parameter inaccuracies and perform real-time corrections. In particular, feature extraction and matching are difficult in dynamic environments, and existing methods fail to effectively utilize data streams for optimization.

Method used

A spatial attention model and unsupervised scene flow training framework are used to correct point cloud distortion and estimate the velocity of moving targets. Semantic features are extracted from attention heatmap regions, and multi-sensor extrinsic parameter estimation and failure detection are performed by combining a factor graph optimization model, thus constructing a tightly coupled optimization mechanism.

Benefits of technology

It improves the calibration accuracy and stability of multi-sensor systems in dynamic environments, enhances sparse feature matching in small visible areas, and improves the robustness and accuracy of online calibration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304801A_ABST
    Figure CN121304801A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal sensor external parameter failure perception and online calibration method. The method comprises the steps of performing speed estimation and point cloud distortion correction on a three-dimensional scene moving target based on a space attention model and an unsupervised scene flow training framework; on the basis of distortion correction and data of a significant target estimation heat map, semantic feature extraction of a spatial moving target is performed through an attention heat map region, and multi-modal data feature matching of a laser radar and a camera is performed through consistency estimation between semantic information and a significant three-dimensional target pose; according to streaming data of an online scene, time-space and attention heat map information are fused, a factor graph optimization model is constructed by using multi-sensor external parameters, online tight coupling optimization is carried out, the external parameters are further accurately estimated, failure perception is carried out by using consistency discrimination of optimization convergence, and a closed loop is formed on a result. According to the method, aiming at an online calibration environment and based on a space attention mechanism, a multi-sensor high-precision fusion external parameter drift problem existing in large-scale road surveying and mapping is optimized, and a better solution is provided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of calibration of lidar and camera, and particularly refers to a multi-modal sensor external parameter failure perception and online calibration method. BACKGROUND

[0002] In recent years, with the continuous development of the field of assisted driving and unmanned driving, through the optimized fusion of lidar, camera and modern sensors such as inertial measurement unit and global positioning system, supplemented by effective software algorithms, efficient road scene perception and reconstruction can be achieved; low-cost, high-precision large-scale road mapping becomes possible: using the accurate depth measurement point cloud of lidar, the data of camera, GPS and inertial navigation are fused to perform real-time reconstruction of the road scene “crowdsourcing”, which can further provide important data basis for vehicle networking and high-precision traffic map reconstruction. Lidar and camera have become the mainstream configuration of intelligent vehicles. For a vehicle-mounted platform containing multiple sensors, the calibration between sensors is an important link, accurate sensor external parameters are the core prerequisite for consistent perception, and are also the necessary condition for the correct operation of the fusion algorithm. Traditionally, sensors are calibrated in an ideal factory environment, regular target objects (such as chessboard, etc.) with clear geometric relationship are laid out within the field of view (FOV) of the sensor, and feature / distance prior is used for extraction and matching optimization to calculate accurate external parameters. However, after the vehicle-mounted platform is put into operation, the external parameters of the vehicle-mounted sensors are prone to drift due to complex road conditions and other factors, thereby affecting the accuracy of the external parameters between sensors and greatly affecting the subsequent perception and road mapping. Therefore, being able to effectively detect the misalignment of external parameters and correct the external parameters in real time is the core basis for multi-sensor fusion of unmanned driving and large-scale road high-precision mapping. Currently, there are many challenges in online calibration of multiple sensors. First, the online calibration scene is generally in the process of vehicle-mounted platform operation, and the road scene has many dynamic targets. The dynamic distortion, noise and other factors of the lidar sensor itself need to be corrected in advance. Second, compared with the ideal environment in the factory, the light and scene of the natural environment on the road change unpredictably, and the feature extraction and matching are more difficult. Especially in the case of large external parameter error, the strategy based on nearest neighbor matching is prone to failure, especially for the case of small common viewing area, more robust feature extraction, matching and optimization mechanism are needed. Finally, tracking and failure perception of the accuracy of external parameters are the trigger and exit premise of online calibration, and are also an important data loop of the online calibration process.

[0003] Calibration between sensors is generally divided into target calibration and non-target calibration by whether there is a target involved. Through real-time or non-real-time, calibration is divided into online calibration and offline calibration. For laser radar-camera calibration algorithm, the current mainstream and mature scheme is a target feature point matching scheme, and offline calibration based on natural scene usually only low-level feature extraction and matching (such as point and plane features). There are two optimization problems in this method: (1) the first is the feature extraction problem: the above algorithm depends on ideal and uniformly distributed significant features. For the calibration environment of road scene, this feature extraction scheme is usually prone to degeneration. (2) This feature matching method is usually sensitive to initial value. In the online calibration scene, facing small common view area, the initial error is large and cannot effectively converge.

[0004] In addition, in the online calibration scene, there are the following problems: (1) online calibration in the vehicle dynamic environment, the original laser radar point cloud usually contains dynamic target distortion and self-motion distortion, and the fusion of data will bring great challenges; (2) online calibration usually contains a large amount of data stream, and the existing calibration method does not effectively utilize this characteristic to constrain and optimize the result; (3) in the online environment, the sensor external parameter may slowly drift or suddenly misalign, and there is no effective method for detecting misalignment. For the three core problems of online calibration, there is no mature research result in the world. SUMMARY

[0005] In order to overcome the deficiencies of the prior art, the present application provides a modal sensor external parameter failure perception and online calibration method, which optimizes the multi-sensor high-precision fusion external parameter drift problem existing in large-scale road mapping based on a spatial attention mechanism in the online calibration environment, and proposes a more optimal solution.

[0006] The present application provides a multi-modal sensor external parameter failure perception and online calibration method, which comprises:

[0007] Based on the spatial attention model and the unsupervised scene flow training framework, the speed of the three-dimensional scene moving target is estimated and the point cloud distortion is corrected;

[0008] Based on the data of distortion correction and saliency target estimation heat map, the semantic feature extraction of spatial moving target is performed through the attention heat map area, and the multi-modal data feature matching of laser radar and camera is performed through the consistency estimation between semantic information and salient three-dimensional target pose;

[0009] According to the streaming data of the online scene, the space-time and attention heat map information are fused, a factor graph optimization model is constructed by using the multi-sensor external parameter, online tight coupling optimization is performed, accurate estimation of the external parameter is further performed, and failure perception is performed by using the consistency discrimination of optimization convergence, and a closed loop is formed on the result.

[0010] Further, according to the multi-modal sensor extrinsic parameter failure perception and online calibration method provided in the application, the point cloud distortion correction comprises:

[0011] Under the known target motion speed, the equation for motion interpolation and distortion correction of the distorted point cloud cluster is formula 1:

[0012]

[0013] wherein, p represents the coordinate of p i in the laser radar coordinate system A, and the transformation of the coordinate system i to A is represented as then the transformation of each point in a frame of point cloud to the coordinate system of the first point is respectively The above formula converts the corresponding point coordinates to the coordinate system A of the first point.

[0014] Further, according to the multi-modal sensor extrinsic parameter failure perception and online calibration method provided in the application, in the formula 1, also represents the distance, the sampling time Δt is known, and the speed S in the Δt time period is assumed by the short-time uniform motion assumption, and then the formula 2 is obtained by

[0015]

[0016] The point cloud distortion correction is converted to the estimation of the three-dimensional target speed by the formula 2; the three-dimensional target comprises a three-dimensional environmental structure target and a motion target, and the three-dimensional structure of the motion target is further restored by estimating the speed of the motion target.

[0017] Further, according to the multi-modal sensor extrinsic parameter failure perception and online calibration method provided in the application, the speed estimation of the three-dimensional scene motion target based on the spatial attention model comprises:

[0018] The target motion speed is estimated by a two-stage model, and a motion heat map of the three-dimensional scene is generated based on the spatial attention model, comprising formula 3:

[0019]

[0020] The spatial attention model divides the three-dimensional grid scene into a motion area and a static area by KNN loss, and the KNN loss is defined as formula 3; wherein NN is a grid neighbor search algorithm, p i is a static area point cloud, and f i is a motion area point cloud.

[0021] ​By employing a small velocity estimation network to estimate the velocity of the moving region, the problem of full-scene velocity estimation is transformed into a problem of velocity estimation of a small area grid of a heatmap.

[0022] Furthermore, according to the multimodal sensor extrinsic parameter failure perception and online calibration method provided in this application, the unsupervised scene flow training framework directly uses raw, unlabeled data input for training, and directly outputs the scene flow estimation velocity vector end-to-end, including:

[0023] The motion T of the target is estimated using a model M. i The target point cloud predicted by M from the previous frame and the current frame is then processed by T. i Projecting to the next frame, the distance d between the projected moving target position and the corresponding moving target position in the actual next frame is 0, which is transformed into the target equation as follows:

[0024] T i =M(p t p t+1 );

[0025]

[0026] Where NN is the nearest neighbor search function, e i This is the ultimate loss.

[0027] Furthermore, based on the multimodal sensor extrinsic parameter failure detection and online calibration method provided in this application, the consistency constraint construction for the relative pose of a salient three-dimensional target includes:

[0028] The salient moving targets and their velocities in the current scene are obtained by using a heatmap of moving targets, and then optimization equations are constructed using the synchronous isomorphism hypothesis to iteratively estimate accurate extrinsic parameters.

[0029] Furthermore, based on the multimodal sensor extrinsic parameter failure detection and online calibration method provided in this application, the optimization equation is constructed based on the spatial attitude consistency assumption theory as follows:

[0030] When the position and attitude of the salient target in sensor A are projected onto sensor B through extrinsic parameters, the target distance error and attitude error are 0; the resulting optimization equation is:

[0031]

[0032] Where π represents the pinhole projection model, and f represents the image lens distortion projection model. For point cloud semantic feature points, q i These are two-dimensional candidate regions for semantic image.

[0033] Furthermore, according to the multimodal sensor extrinsic parameter failure detection and online calibration method provided in this application, the online tight coupling optimization is a tight coupling optimization of multi-sensor calibration based on attention feature semantic constraints, including:

[0034] Constructing an information matrix based on attention-feature semantic factors;

[0035] The information matrix, used as weights in the factor graph, is imported into the objective equation for global optimization.

[0036]

[0037] in, For external parameters to be optimized, O i For the current frame observation information, For common-view target observation within the sliding window, S i O is the constraint factor between sensors. i , All of these are semantic attention constraint factors, and ∑α, ∑δ, ∑∈ represent the covariance matrices of different factors, respectively;

[0038] The convergence state of the optimization equation is tracked to assess whether the extrinsic parameters have failed.

[0039] Furthermore, according to the multimodal sensor extrinsic parameter failure detection and online calibration method provided in this application, the evaluation of extrinsic parameter calibration, which involves accurately estimating the extrinsic parameters, is divided into quantitative evaluation of accuracy and quantitative evaluation of precision.

[0040] The accuracy quantitative assessment employs multiple calibrations, and the overall standard deviation is used to evaluate the calibration results.

[0041] Furthermore, according to the multimodal sensor extrinsic parameter failure detection and online calibration method provided in this application, the accuracy quantitative assessment is evaluated by designing an experiment using pixel distance error, including:

[0042] The accuracy of the calibration is quantitatively and objectively evaluated by measuring the pixel-to-target centerline error distribution of the two-dimensional projection of the target point cloud at different locations within a common field of view.

[0043] The beneficial effects of this invention are as follows: The multimodal sensor extrinsic parameter failure perception and online calibration method provided by this invention is based on a spatial attention model and an unsupervised training framework based on cycle consistency. It generates velocity and semantic heatmaps for moving targets in the scene context, and performs more accurate velocity estimation through the scene flow of the heatmap, thereby correcting motion distortion and providing an accurate data foundation and attention heatmap feature data flow for subsequent online calibration.

[0044] Based on the heatmap feature information of the attention mechanism, multimodal data feature matching is performed through salient scene target extraction and target pose estimation. This improves the online matching and calibration of sparse features in small visible areas, enhancing the robustness and accuracy of the results.

[0045] Then, failure detection and online optimization are performed by using attention feature heatmaps and spatiotemporal consistency error constraints. We explore the use of more semantic observations to optimize direct extrinsic parameter relationships, and use attention heatmap features to construct factor graph optimization to further improve calibration accuracy. Attached Figure Description

[0046] The technical solution and other beneficial effects of this application will become apparent from the following detailed description of specific embodiments in conjunction with the accompanying drawings.

[0047] Figure 1 This is a schematic diagram of the multimodal sensor extrinsic parameter failure detection and online calibration method provided in this embodiment.

[0048] Figure 2 This is a schematic diagram of the process of generating distorted point clouds.

[0049] Figure 3 The images show a comparison of the three sets of distorted point clouds and the corrected point clouds provided in this embodiment.

[0050] Figure 4 This is the scene flow estimation framework based on the spatial attention model in this embodiment.

[0051] Figure 5 This diagram illustrates the calculation of the result loss through the bijective transformation consistency of the domain, representing the cyclic consistency loss.

[0052] Figure 6 A schematic diagram of the construction process for saliency-based 3D target semantics-PNP.

[0053] Figure 7 A schematic diagram of a sliding window factor graph optimization model that combines pose, observational landmark information (attention heatmap target), and sensor extrinsic context constraints.

[0054] Figure 8 This is a schematic diagram of center point pixel error estimation used for radar camera calibration accuracy evaluation. Detailed Implementation

[0055] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0056] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0057] The following disclosure provides many different embodiments or examples for implementing different structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed. In addition, various specific examples of processes and materials are provided in this application, but those skilled in the art will recognize the application of other processes and / or the use of other materials.

[0058] The embodiments of this application will now be further described in conjunction with the accompanying drawings and specific implementation details.

[0059] Figure 1 This is a schematic diagram of the multimodal sensor extrinsic parameter failure detection and online calibration method provided in this embodiment.

[0060] like Figure 1 As shown, in order to address the problems of online calibration of multiple sensors in large-scale road mapping scenarios, the method provided in this embodiment optimizes the problem of extrinsic parameter drift in high-precision fusion of multiple sensors in large-scale road mapping in three aspects: point cloud distortion correction, online calibration of multimodal sensors, and calibration fusion failure perception and tight coupling optimization.

[0061] Specifically, the method includes:

[0062] Velocity estimation and point cloud distortion correction for moving targets in 3D scenes are performed based on a spatial attention model and an unsupervised scene flow training framework.

[0063] Based on the data from distortion correction and salient target estimation heatmaps, semantic features of spatially moving targets are extracted through attention heatmap regions, and multimodal data feature matching of lidar and camera is performed through consistency estimation between semantic information and salient 3D target poses.

[0064] Based on streaming data from online scenarios, spatiotemporal and attention heatmap information is fused, and a factor graph optimization model is constructed using multi-sensor extrinsic parameters. Online tightly coupled optimization is then performed, and the extrinsic parameters are further accurately estimated. Finally, failure detection is achieved using consistency judgment of optimization convergence, thus forming a closed loop in the results.

[0065] Specifically, in the online calibration process of vehicle-mounted mapping, the vehicle itself is usually in motion, and the road environment often contains complex three-dimensional moving targets. Since LiDAR data is typically non-instantaneous measurement, its non-measurement noise usually manifests as self-motion distortion at high speeds and blurring and ghosting of moving targets, with varying degrees of distortion at different speeds. For current mainstream automotive-grade reciprocating LiDARs, these problems are often more severe. During sensor calibration, it is essential to effectively correct this distortion noise to meet the accuracy and stability requirements of online calibration.

[0066] Therefore, in this embodiment, a spatial attention model and an unsupervised training framework based on cycle consistency are used to generate velocity and semantic heatmaps for moving targets in the scene context. The scene flow of these heatmaps is then used to perform more accurate velocity estimation, thereby correcting motion distortion and providing an accurate data foundation and attention heatmap feature data flow for subsequent online calibration.

[0067] In point cloud distortion correction based on a spatial attention model, the first step is to analyze the theory of point cloud distortion correction. The essence of motion distortion in 3D point clouds is the deviation in the measurement coordinate system caused by different time deviations in target measurement. Due to the motion characteristics of the target, its relative coordinates are constantly changing, so the data measured at different times cannot be unified into a single spatial coordinate system.

[0068] Figure 2 The process of distortion is demonstrated. Because the lidar measures the target at different times, the target will produce different degrees of motion blur depending on its relative motion state, which is the so-called "distortion".

[0069] Figure 3 This embodiment demonstrates the effects of three sets of distorted point clouds and corrected point clouds, such as... Figure 3As shown in the three examples, three different moving targets are used to demonstrate the effects before and after motion distortion correction. Different colored point clouds represent data collected at different times. It can be seen that the point cloud after accurate velocity correction is clearer and better restores the three-dimensional structure of the target itself.

[0070] Therefore, a core aspect of distortion correction is the estimation of the moving target's velocity. By estimating the velocity and using a motion interpolation algorithm at time t, the original three-dimensional structure of the moving target can be reconstructed.

[0071] As shown in Equation (1), the equations for motion interpolation and distortion correction of distorted point cloud clusters under known target velocity are presented:

[0072]

[0073] In this embodiment, with p i The coordinates in the lidar coordinate system A, the transformation from coordinate system i to A is expressed as: The transformations from the coordinate system of each point in a point cloud frame to the coordinate system of the first point are as follows: The coordinates of the corresponding point are transferred to the coordinate system A of the first point using the above formula (1).

[0074] In Formula 1, This also represents the distance traveled. The sampling time Δt is known. Using the assumption of short-term uniform motion, we assume the velocity within the time interval Δt is S. Then, through... Formula 2 can be obtained:

[0075]

[0076] The point cloud distortion correction is converted into an estimate of the velocity of a three-dimensional target using Formula 2. The three-dimensional target includes a three-dimensional environmental structure target and a moving target. The three-dimensional structure of the moving target is further restored by estimating the velocity of the moving target.

[0077] Specifically, formula (2) transforms the problem of point cloud distortion correction into the problem of estimating the velocity of three-dimensional targets. The three-dimensional targets can be divided into two categories. One category is three-dimensional environmental structural targets, which are usually static targets. The reason for the blurring caused by this type is relative motion, that is, it is necessary to accurately estimate its own motion in order to calculate the motion distortion of the static three-dimensional environment in reverse, and then correct it. In the vehicle scenario, GPS+RTK combined navigation can generally be used to obtain relative motion, and the LIO radar inertial navigation odometer algorithm can be fused to make accurate trajectory T and attitude R estimation, and self-motion distortion correction can be performed through R and T. The other category is moving targets, where it is necessary to accurately estimate the velocity of the moving targets in order to further restore the three-dimensional structure of the moving targets. The difficulty of point cloud distortion correction is the distortion correction of moving targets.

[0078] Then, based on the spatial attention model, region segmentation and velocity estimation are performed on moving targets in the 3D scene to correct motion distortion. Current mainstream scene-flow-based models typically take the entire 3D scene or a meshed 3D scene as input, which introduces two problems: first, the input data is a 3D point cloud covering the entire scene, leading to increased model complexity; second, the entire scene data usually contains complex moving targets, which significantly impacts the accuracy and precision of the network's motion target feature extraction and velocity estimation. Inspired by the spatial attention mechanism, the method provided in this embodiment uses a two-stage model for target motion velocity estimation:

[0079] like Figure 4 As shown, firstly, a motion heatmap is generated for the 3D scene based on a spatial attention model, including public...

[0080] Formula 3:

[0081]

[0082] This model divides the 3D mesh scene into moving and stationary regions using KNN loss. KNN loss is defined by formula (3), where NN is the mesh nearest neighbor search algorithm, p i For a point cloud in a static region, f i A point cloud is generated for the motion region. Then, a small velocity estimation network is used to estimate the velocity of the motion region. This transforms the full-scene velocity estimation problem into a small-region grid velocity estimation problem based on a heatmap, thus reducing the scope and scale of the model and lowering computational complexity. Simultaneously, the two-stage small network ensures the accuracy of velocity regression. Furthermore, this framework possesses inherent parallelism, improving computational efficiency.

[0083] Finally, a context-based unsupervised scene flow training framework is used to train and optimize the two-stage model. The biggest challenge currently facing scene flow algorithms is obtaining the ground truth. Training with manually labeled datasets (KITTI, FlyThings 3D, etc.) introduces two problems: first, the transfer problem between data from different scanning modes; and second, the long-tail problem caused by insufficient data volume and the high cost of manual annotation for increasing data volume. These are risks that will be faced in practical applications.

[0084] Therefore, the method provided in this embodiment proposes an unsupervised two-stage network training framework for scene flow, which directly uses raw, unlabeled data input for training and outputs the scene flow estimation velocity vector end-to-end.

[0085] Figure 5 The example demonstrates the cycle consistency loss, calculating the resulting loss through the bijective transformation consistency of the domain, such as... Figure 5 As shown, the core of the unsupervised training framework proposed in this embodiment is frame consistency estimation. Specifically, it is assumed that the motion T of the target can be estimated using a model M. i Then, the target point cloud predicted by M from the previous frame and the current frame is used by T. i Theoretically, when projected onto the next frame, the distance d between the projected position of the moving target and the corresponding position of the moving target in the actual next frame is 0, which can be transformed into the target equation as follows:

[0086] T i =M(p t p t+1 );

[0087]

[0088] Where NN is the nearest neighbor search function, e i This is the ultimate loss.

[0089] Furthermore, in the method provided in this embodiment, a cycle-consistency loss from Generative Adversarial Networks (GANs) is introduced to further constrain the network's convergence. The cycle-consistency loss is used to achieve spatial transformation between two domains in the absence of paired data. Assume a map X... domin Image x translated to Y domin Get image y, then translate it back to X domin Given G(F(x)), and similarly, images y and G(F(y)), then x and G(F(x)), and y and G(F(y)), should be exactly the same. The differences between them can then serve as a monitoring signal.

[0090]

[0091] For quantitative evaluation of this algorithm, datasets such as KITTI and Fly Things 3D can be used. Furthermore, for cross-validation on our self-developed dataset, this method uses the Crispness Score to evaluate the distortion-corrected point cloud. The Crispness Score is defined as follows:

[0092]

[0093] Where T is the point cloud frame of the tracked target, and n i This represents the number of point clouds in the i-th frame. for The nearest neighbors of the point cloud are represented by ω, which is the weight. This formula, from the perspective of information entropy, structurally represents the "clarity" of the point cloud and can indirectly evaluate the accuracy of velocity estimation and distortion correction.

[0094] In LiDAR-camera multimodal semantic feature matching and online calibration, unlike offline static indoor calibration, the online calibration scenario of LiDAR cameras based on road environment usually involves high noise and several long-tail problems. Specifically, there are two problems that need to be solved: (1) In the deployment of sensors in vehicle scenarios, in order to obtain a high coverage of the entire field of view, the common field of view between sensors is usually compressed, and the initial values ​​of the extrinsic parameters are poor. The traditional calibration method using the nearest neighbor feature matching strategy is prone to failure, and it is necessary to explore a more robust high-dimensional feature matching method. (2) In the online calibration scenario of LiDAR and camera calibration, it is usually not possible to use the offline accumulation of multiple frames to extract dense features, so it is necessary to solve the problem of robust sparse feature extraction, matching and posterior optimization of multimodal data of three-dimensional sparse point cloud and two-dimensional image. For the large initial error problem and multimodal matching problem in small overlapping field of view scenarios, in this method, based on the heatmap feature information of the first part of the attention mechanism, multimodal data feature matching is performed by extracting salient scene targets and estimating target pose. This improves the online matching and calibration of sparse features in small visible areas, enhancing the robustness and accuracy of the results.

[0095] Specifically, in the feature matching of lidar-camera multimodal data for salient target pose estimation, based on the data of distortion correction and salient 3D target heatmap, this method uses the consistency constraint of the relative pose of salient 3D targets to construct the initial extrinsic parameter estimate, and then uses a multi-sensor tightly coupled graph optimization model for precise fine-tuning.

[0096] First, the consistency constraints for the relative poses of salient 3D targets are constructed. This includes obtaining salient moving targets and their velocities in the current scene through a moving target heatmap, followed by the synchronous isomorphism assumption: theoretically, when the position and orientation of a salient target in sensor A are projected onto sensor B through extrinsic parameters, its target distance error and orientation error are zero. This is transformed into the following optimization equation:

[0097]

[0098] Where π represents the pinhole projection model, and f represents the image lens distortion projection model. For point cloud semantic feature points, q i These are two-dimensional candidate regions for semantic image.

[0099] Based on this spatial pose consistency assumption, optimization equations can be constructed within the context to iteratively estimate accurate extrinsic parameters. Since images are two-dimensional targets and point cloud data are three-dimensional targets, calibration matching based solely on data dimensions is cross-dimensional matching. When the initial error of the extrinsic parameters is too large, it is prone to getting trapped in local minima, leading to calibration failure. The innovation of this method lies in using an attention mechanism to project heatmap feature data, narrowing the feature matching range, and elevating both the two-dimensional data of the image and the three-dimensional data of the point cloud to a high-dimensional semantic level (x, y, z, pose, speed, semantic). Nearest neighbor retrieval is then performed in the high-dimensional space to construct "semantic PNP" matching. Figure 6 It illustrates the specific process of semantic PNP.

[0100] In consistency-tightly coupled optimization and failure detection based on spatial attention heatmaps, failure detection and continuous optimization of multi-sensor extrinsic parameters are crucial issues in automotive environments. During multi-sensor operation, the extrinsic parameter status of sensors needs to be tracked in real time. Once sensor extrinsic parameters become inaccurate, the algorithm should quickly and accurately detect and identify the failure, then initiate calibration. Once calibration is complete, a successful calibration status should be identified, and the calibration optimization calculation should be exited. Furthermore, in multi-sensor environments, pairwise sensor calibration strategies can lead to cumulative errors and consistency errors. More observations are needed to improve the consistency of the final results. This method uses attention feature heatmaps and spatiotemporal consistency error constraints for failure detection and online optimization. It explores using more semantic observations to optimize direct extrinsic parameter relationships and utilizes attention heatmap features to construct factor graph optimization to further improve calibration accuracy.

[0101] Specifically, semantic PNP can be used to obtain initial values ​​of extrinsic parameters for a single frame. In vehicular road scenarios, there are often high noise levels, complex multi-moving targets, and occlusion. These challenging scenarios typically introduce unpredictable risks. However, in online vehicular calibration scenarios, data stream information from consecutive frames and calibration information between sensors can usually be obtained. Multi-frame and multi-sensor optimization can be further introduced to improve the accuracy and stability of the results.

[0102] In this method, multi-sensor calibration is tightly coupled and optimized based on attention-feature semantic constraints. Traditionally, multi-frame calibration results, i.e., six-degree-of-freedom extrinsic parameter data, are used as observations to construct a factor map for optimization. This can lead to estimation errors when the scene is simple or features are degenerate. In this method, an information matrix is ​​constructed based on attention-feature semantic factors, M =<speed,semantic> The information matrix, used as weights in the factor graph, is imported into the objective equation for global optimization.

[0103]

[0104] in, For external parameters to be optimized, O i For the current frame observation information, For common-view target observation within the sliding window, S i O is the constraint factor between sensors. i , All are semantic attention constraint factors, where ∑α, ∑δ, and ∑∈ represent the covariance matrices of different factors. Simultaneously, the convergence state of the optimization equation is tracked to evaluate the failure of the extrinsic parameters.

[0105] Figure 7 A sliding window factor graph optimization model that combines pose, observed landmark information (attention heatmap target), and sensor extrinsic context constraints.

[0106] The evaluation of extrinsic parameter calibration, which involves accurately estimating the external parameters, is divided into quantitative accuracy evaluation and quantitative precision evaluation. The quantitative accuracy evaluation employs multiple calibrations, and the overall standard deviation is used to assess the calibration results. Quantitative precision evaluation is one of the challenges in sensor calibration, as it is typically difficult to directly obtain accurate attitude parameters (true values) between sensors through measurement. In this method, however, the quantitative precision evaluation is conducted using pixel distance error through experimental design. Figure 8 The following are examples:

[0107] The accuracy of the calibration is quantitatively and objectively evaluated by measuring the error distribution from the pixels (red) of the two-dimensional projection of the target point cloud at different locations within the common field of view to the center line of the image target.

[0108] In summary, large-scale road mapping inevitably faces scenarios such as "high dynamics" and "feature sparsity". These scenarios can lead to a series of serious problems such as data distortion, multi-sensor extrinsic parameter drift and feature degradation. The method provided in this embodiment starts from the attention mechanism and uses attention heatmap features to penetrate the whole system to perform distortion correction, calibration, and tight coupling failure perception optimization.

[0109] This method addresses the data distortion problem caused by the "high dynamic range" in online road mapping calibration. It uses a scene flow model based on spatial attention to estimate the velocity of moving targets and correct distortion. This model is trained using an unsupervised machine learning framework with frame-to-frame consistency constraints, eliminating the need for expensive manual labeling and providing an effective solution to sensor transfer adaptation and long-tail problems. Furthermore, this framework can simultaneously output the velocity of the moving target, background, and target, aiming to correct data distortion and provide a precise data foundation while offering semantic heatmap information for downstream perception / calibration tasks in autonomous driving. It is expected to effectively improve the accuracy and stability of calibration from the data source in online, high-dynamic environments.

[0110] Online calibration of dynamic scenes in road environments is a core foundation for large-scale road mapping applications using multi-sensor fusion. Currently, multi-sensor calibration is mainly performed offline, and online calibration of road environments primarily focuses on nearest-neighbor feature detection and matching. For dynamic scenes with large initial extrinsic parameters, effective matching and convergence are difficult. Furthermore, current algorithms mainly rely on single-frame data and rarely utilize spatiotemporal constraints for noise suppression and optimization. This method proposes a salient target pose estimation framework, utilizing an initial attention heatmap for feature extraction and matching of LiDAR-camera multimodal data. It aims to solve the dynamic registration and calibration problem under large initial error values. Based on this, it explores dynamic failure perception and tightly coupled extrinsic parameter optimization based on contextual constraints, providing effective theoretical and technical support for robust online calibration in large-scale road mapping scenarios.

[0111] Therefore, the multimodal sensor extrinsic parameter failure perception and online calibration method provided by this invention is based on a spatial attention model and an unsupervised training framework based on cycle consistency. It generates velocity and semantic heatmaps for moving targets in the scene context, and performs more accurate velocity estimation through the scene flow of the heatmap, thereby correcting motion distortion and providing an accurate data foundation and attention heatmap feature data flow for subsequent online calibration.

[0112] Based on the heatmap feature information of the attention mechanism, multimodal data feature matching is performed through salient scene target extraction and target pose estimation. This improves the online matching and calibration of sparse features in small visible areas, enhancing the robustness and accuracy of the results.

[0113] Then, failure detection and online optimization are performed by using attention feature heatmaps and spatiotemporal consistency error constraints. We explore the use of more semantic observations to optimize direct extrinsic parameter relationships, and use attention heatmap features to construct factor graph optimization to further improve calibration accuracy.

[0114] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the present invention. Finally, it should be noted that in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0115] The above provides a detailed description of a modal sensor extrinsic parameter failure detection and online calibration method provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the technical solutions and core ideas of this application. Those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for detecting and calibrating extrinsic parameters of a multimodal sensor online, characterized in that, The method includes: Velocity estimation and point cloud distortion correction for moving targets in 3D scenes are performed based on a spatial attention model and an unsupervised scene flow training framework. Based on the data from distortion correction and salient target estimation heatmaps, semantic features of spatially moving targets are extracted through attention heatmap regions, and multimodal data feature matching of lidar and camera is performed through consistency estimation between semantic information and salient 3D target poses. Based on streaming data from online scenarios, spatiotemporal and attention heatmap information is fused, and a factor graph optimization model is constructed using multi-sensor extrinsic parameters. Online tightly coupled optimization is then performed, and the extrinsic parameters are further accurately estimated. Finally, failure detection is achieved using consistency judgment of optimization convergence, thus forming a closed loop in the results.

2. The method for detecting and calibrating extrinsic parameters of a multimodal sensor online according to claim 1, characterized in that, The point cloud distortion correction includes: Given the target's velocity, the equations for motion interpolation and distortion correction of the distorted point cloud clusters are given by Formula 1: Among them, with p i The coordinates in the lidar coordinate system A, the transformation from coordinate system i to A is expressed as: The transformations from the coordinate system of each point in a point cloud frame to the coordinate system of the first point are as follows: The above formula transforms the coordinates of the corresponding point into the coordinate system A of the first point.

3. The method for detecting and calibrating the extrinsic parameters of a multimodal sensor online according to claim 2, characterized in that, In Formula 1, This also represents the distance traveled. The sampling time Δt is known. Using the assumption of short-term uniform motion, we assume the velocity within the time interval Δt is S. Then, through... Formula 2 can be obtained: The point cloud distortion correction is converted into an estimate of the velocity of a three-dimensional target using Formula 2. The three-dimensional target includes a three-dimensional environmental structure target and a moving target. The three-dimensional structure of the moving target is further restored by estimating the velocity of the moving target.

4. The method for detecting and calibrating the extrinsic parameters of a multimodal sensor online according to claim 3, characterized in that, The velocity estimation of moving targets in a 3D scene based on the spatial attention model includes: The target's motion velocity is estimated using a two-stage model, and a motion heatmap is generated for the 3D scene based on a spatial attention model, including Equation 3: The spatial attention model divides the 3D mesh scene into moving and stationary regions using KNN loss, which is defined in Equation 3; where NN is the mesh nearest neighbor search algorithm, p i For a point cloud in a static region, f i Point cloud for the motion region; By employing a small velocity estimation network to estimate the velocity of the moving region, the problem of full-scene velocity estimation is transformed into a problem of velocity estimation of a small area grid of a heatmap.

5. The method for detecting and calibrating the extrinsic parameters of a multimodal sensor online according to claim 4, characterized in that, The unsupervised scene flow training framework directly uses raw, unlabeled data input for training and outputs the scene flow estimation velocity vector end-to-end, including: The motion T of the target is estimated using a model M. i The target point cloud predicted by M from the previous frame and the current frame is then processed by T. i Projecting to the next frame, the distance d between the projected moving target position and the corresponding moving target position in the actual next frame is 0, which is transformed into the target equation as follows: T i =M(p t ,p t+1 ); Where NN is the nearest neighbor search function, e i This is the ultimate loss.

6. The method for detecting and calibrating the extrinsic parameters of a multimodal sensor online according to claim 1, characterized in that, The construction of consistency constraints for the relative pose of salient 3D targets includes: The salient moving targets and their velocities in the current scene are obtained by using a heatmap of moving targets, and then optimization equations are constructed using the synchronous isomorphism hypothesis to iteratively estimate accurate extrinsic parameters.

7. The method for detecting and calibrating the extrinsic parameters of a multimodal sensor online according to claim 6, characterized in that, Based on the assumption of spatial attitude consistency, the optimization equation is constructed as follows: When the position and attitude of the salient target in sensor A are projected onto sensor B through extrinsic parameters, the target distance error and attitude error are 0; the resulting optimization equation is: Where π represents the pinhole projection model, and f represents the image lens distortion projection model. For point cloud semantic feature points, q i These are two-dimensional candidate regions for semantic image.

8. The method for detecting and calibrating the extrinsic parameters of a multimodal sensor online according to claim 1, characterized in that, The online tight coupling optimization is a tight coupling optimization of multi-sensor calibration based on attention feature semantic constraints, including: Constructing an information matrix based on attention-feature semantic factors; The information matrix, used as weights in the factor graph, is imported into the objective equation for global optimization. in, For external parameters to be optimized, O i For the current frame observation information, For common-view target observation within the sliding window, S i O is the constraint factor between sensors. i , All of these are semantic attention constraint factors, and ∑α, ∑δ, ∑∈ represent the covariance matrices of different factors, respectively; The convergence state of the optimization equation is tracked to assess whether the extrinsic parameters have failed.

9. The method for detecting and calibrating extrinsic parameters of a multimodal sensor online according to claim 8, characterized in that, The assessment of external parameter calibration, which involves making accurate estimates of external parameters, is divided into quantitative assessment of accuracy and quantitative assessment of precision. The accuracy quantitative assessment employs multiple calibrations, and the overall standard deviation is used to evaluate the calibration results.

10. The method for detecting and calibrating the extrinsic parameters of a multimodal sensor online according to claim 9, characterized in that, The accuracy quantitative assessment is performed through an experiment using pixel distance error, including: The accuracy of the calibration is quantitatively and objectively evaluated by measuring the pixel-to-target centerline error distribution of the two-dimensional projection of the target point cloud at different locations within a common field of view.

Citation Information

Cited By

  • Online self-calibration and SLAM method and system for multi-depth vision unmanned aerial vehicle

    CN121837385A