A Fusion Calibration Method for 4D Millimeter-Wave Radar and Visible Light Camera Applicable to Environmental Perception

By employing a three-stage optimization strategy and a deep learning calibration network, the robustness and real-time performance issues of 4D millimeter-wave radar and visible light cameras in real road environments were resolved, achieving high-precision dynamic calibration suitable for vehicle-mounted environmental perception in complex dynamic scenarios.

CN122131258APending Publication Date: 2026-06-02ZHEJIANG SCI-TECH UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG SCI-TECH UNIV
Filing Date
2026-05-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In real-world road environments, existing calibration methods for 4D millimeter-wave radar and visible light cameras suffer from insufficient robustness and difficulty in achieving real-time performance. In particular, errors accumulate significantly in complex dynamic scenarios, making it difficult to meet the response requirements of vehicle systems.

Method used

A three-stage optimization strategy is adopted, including extrinsic parameter optimization, intrinsic parameter distortion optimization, and overall optimization. By combining the region matching loss function and the deep learning calibration network, and constructing 3D-2D corresponding data pairs, the radar-camera extrinsic parameters, camera intrinsic parameters, and distortion parameters are optimized in stages to achieve high-precision dynamic calibration.

Benefits of technology

It significantly improves the accuracy and stability of calibration results, reduces computational complexity, and enhances calibration accuracy and real-time performance in complex road scenarios, meeting the response requirements of vehicle systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131258A_ABST
    Figure CN122131258A_ABST
Patent Text Reader

Abstract

This invention relates to a fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception. It uses a triangular radar corner reflector to collect radar point cloud and camera image calibration data at different distances and angles. After preprocessing, 3D-2D corresponding data pairs are constructed, and a region-constrained optimization framework is built. Using a region matching loss function as the error metric, the radar-camera extrinsic parameters, camera intrinsic parameters, and distortion parameters are optimized in stages to obtain a static benchmark calibration parameter set. Temporal synchronous data of different types of environmental features are collected, and supervisory labels are generated through spatiotemporal alignment and frame-by-frame parameter correction to train the calibration network model. The model is deployed to a computing platform, and dynamic calibration parameters are output in real time for parameter updates. This invention significantly improves the accuracy and stability of calibration results, achieving the dual effects of reducing computational power and increasing efficiency while reducing noise and maintaining accuracy, and greatly enhancing the model's generalization ability in real-world complex road scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of general image data processing or generation, and in particular to a fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception. Background Technology

[0002] To meet the demands of in-vehicle environmental perception in real-world road environments, multi-sensor fusion perception has become a core approach to improve the robustness and reliability of environmental perception. 4D millimeter-wave radar, with its all-weather capability and precise radial velocity sensing, can adapt to the complex conditions of real roads, while visible light cameras excel in spatial resolution and environmental semantic understanding. The two technologies complement each other effectively, forming a core sensor combination suitable for real-world road perception. Achieving high-precision fusion of these two sensors in real-world road environments hinges on accurately calibrating their spatial extrinsic parameters, namely the rotation matrix R and the translation vector t.

[0003] Currently, the online calibration technology using 4D millimeter-wave radar and visible light cameras still faces two prominent bottlenecks in its practical application in real road environments:

[0004] On the one hand, traditional geometric optimization methods, such as the PnP algorithm, often simplify radar corner reflectors to single-point centroid matching when processing them, ignoring their inherent triangular region geometry. In real road environments, these methods are easily affected by image segmentation noise and outliers in radar point clouds, resulting in insufficient calibration robustness. At the same time, these methods usually fix the initial intrinsic parameters of the camera and do not correct for the systematic projection bias caused by intrinsic parameter errors. This leads to a significant lack of calibration stability in various real road scenarios such as urban roads, rural roads, and highways.

[0005] On the other hand, while existing deep learning calibration methods, such as DeepFocal and PoseNet, have the potential for real-time calibration, their training labels are highly dependent on the output of traditional methods, which easily leads to error accumulation and propagation. Moreover, such errors are further amplified in the complex and dynamic scenarios of real roads. At the same time, the model architecture is designed for dense LiDAR point clouds, which is difficult to adapt to the characteristics of sparse 4D millimeter-wave radar point clouds and drastic fluctuations in radar cross section (RCS). The iterative inference structure leads to high inference latency, which cannot meet the stringent requirements of millisecond-level response of vehicle systems in real road environments.

[0006] Therefore, there is an urgent need in this field for a new online calibration method that deeply integrates the physical characteristics of sensors. This method needs to be adapted to the complex dynamic scenarios of real roads, and while taking into account calibration accuracy and real-time performance, it should also have strong environmental adaptability. This would break through the application limitations of existing technologies in real road environments and ensure the accuracy and stability of multi-sensor fusion perception in real roads. Summary of the Invention

[0007] This invention solves the problems existing in the prior art and provides a fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception, especially suitable for high-precision dynamic calibration in complex scenarios such as real roads.

[0008] The technical solution adopted in this invention is a fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception, comprising the following steps:

[0009] S1 constructs a data acquisition scenario using a triangular radar corner reflector, a visible light camera, and a 4D millimeter-wave radar. The triangular radar corner reflector is used to acquire radar point cloud and camera image calibration data at different distances and angles.

[0010] S2 preprocesses the acquired radar point cloud and camera images to construct 3D-2D corresponding data pairs;

[0011] S3 Based on the characteristics of the acquisition scenario, a regional constraint optimization framework is constructed. The regional matching loss function is used as the error metric for parameter optimization. Combined with the optimization algorithm, the radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters are optimized in stages to obtain the optimal radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters.

[0012] S4 collects sample data of different types of environmental characteristics, and simultaneously collects image data output by the visible light camera and point cloud data output by the 4D millimeter-wave radar to obtain time-series synchronized data.

[0013] S5 uses time-synchronized data as training data and a complete set of calibration parameters consisting of optimal radar-camera extrinsic parameters, camera intrinsic parameters, and camera distortion parameters as a static benchmark calibration parameter set. Through spatiotemporal alignment verification and frame-by-frame parameter adaptive correction, frame-by-frame supervision labels of dynamic road data are generated to train and optimize the calibration network model.

[0014] S6 deploys the trained calibration network model to the computing platform, processes the synchronously acquired 4D millimeter-wave radar point cloud and visible light camera images in real time, updates the time delay with preset calibration parameters, and outputs the radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters at the current moment based on the calibration network model, thus completing the update of calibration parameters.

[0015] Preferably, the visible light camera image is processed using a preset image processing algorithm to perform edge detection and contour filtering, and extract the two-dimensional image region corresponding to each radar corner reflector;

[0016] For 4D millimeter-wave radar point cloud data, candidate points are selected based on the preset features of radar corner reflectors, and spatial clustering is performed on the candidate points. Each cluster is assigned to a radar corner reflector, and its three-dimensional spatial coordinate point set is determined.

[0017] The two-dimensional image region corresponding to each radar corner reflector is associated with a set of three-dimensional spatial coordinate points to obtain a 3D-2D corresponding data pair with an associated reference.

[0018] Preferably, in S3, the phased targeted optimization includes sequentially executed extrinsic parameter optimization, intrinsic parameter distortion optimization, and overall optimization. Each phase uses the region matching loss function as the error metric and combines optimization algorithms to achieve parameter optimization.

[0019] Preferably, in the extrinsic parameter optimization, the initial intrinsic parameters of the camera are fixed, and the initial extrinsic parameters are used as initial values. The corresponding region matching loss function and optimization algorithm are used to optimize only the radar-camera extrinsic parameters to obtain refined extrinsic parameters. Based on the refined extrinsic parameters, the original 3D-2D corresponding data pairs are projected for verification to identify and remove projection outliers. The initial intrinsic parameters of the camera are continuously fixed, and the extrinsic parameters of the cleaned data are optimized again using the corresponding region matching loss function and optimization algorithm to obtain further refined extrinsic parameters.

[0020] Preferably, in the intrinsic parameter distortion optimization, the refined extrinsic parameters and the cleaned 3D-2D corresponding data pairs are used as inputs. The extrinsic parameters are fixed, and the corresponding region matching loss function and optimization algorithm are used to optimize only the camera intrinsic parameters and camera distortion parameters to obtain the optimized camera intrinsic parameters and camera distortion parameters.

[0021] Preferably, in the overall optimization, the refined extrinsic parameters, optimized camera intrinsic parameters, camera distortion parameters, and cleaned 3D-2D corresponding data pairs are used as inputs. The corresponding region matching loss function and optimization algorithm are adopted to simultaneously optimize the radar-camera extrinsic parameters, camera intrinsic parameters, and camera distortion parameters, and finally output high-precision radar-camera extrinsic parameters, camera intrinsic parameters, and distortion parameters.

[0022] Preferably, the region matching loss function includes error calculation rules that differentiate between designs within and outside the region;

[0023] Construct a single-point projection error function:

[0024]

[0025] in, For the i-th radar point, project it onto the camera image as a pixel. The Euclidean distance to the midpoint C of the triangle region of the radar corner reflector. The average distance from the midpoint to the edge of the triangle. This represents the loss corresponding to the i-th set of data. This represents the triangular region of the radar corner reflector;

[0026] Obtain the region matching loss function.

[0027]

[0028] Where N represents the total number of radar points; Let the coordinates of the i-th radar point be expressed as the three-dimensional coordinates in the radar coordinate system. Let θ = {R, t, A, η} be the set of variables to be optimized, where R is the radar-camera rotation matrix, t is the translation vector, A is the camera intrinsic parameter, and η is the camera distortion parameter. This is the radar-camera pixel projection mapping function.

[0029] Preferably, in S5, the optimal parameters obtained in S3 are used as the initial reference values ​​for each frame tag; the preprocessing process of step S2 is repeated for each frame of synchronized data in S4 to construct 3D-2D corresponding data pairs.

[0030] Calculate the average projection error based on the reference parameters. If the error is less than the preset first threshold, the reference parameters are directly used as the frame label. Otherwise, the reference parameters are used as the initial values, and the radar-camera extrinsic parameters and / or camera intrinsic parameters are adjusted by a preset amplitude. The correction parameters are obtained by minimizing the region matching loss function through the optimization algorithm and used as the frame label.

[0031] Abnormal frame labels that have an average projection error greater than a preset second threshold after correction are removed, and the final frame-by-frame supervision label set is composed of the baseline parameters and / or correction parameters of all valid frames.

[0032] Preferably, the calibration network model is an improved CalibNet model;

[0033] The model's input layer includes a feature reconstruction module, which is used to concatenate 3D-2D corresponding data pairs in the channel dimension and the time difference information between the current frame and adjacent frames as input;

[0034] A lightweight encoder and a decoder are sequentially arranged after the input layer. The lightweight encoder uses depth-separable convolutional blocks and compressed excitation modules connected in series. The decoder has three parallel regression branches at the end, which output new radar-camera extrinsic parameters, camera intrinsic parameters and distortion parameters respectively.

[0035] Preferably, in S6, the rapid dynamic calibration adopts a differentiated parameter update strategy, wherein the radar-camera extrinsic parameters are updated in full real time, and the camera intrinsic parameters and distortion parameters are updated using a reference value locking and online incremental correction method.

[0036] This invention relates to a fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception. It uses a triangular radar corner reflector to collect radar point cloud and camera image calibration data at different distances and angles. After preprocessing, 3D-2D corresponding data pairs are constructed, and a regional constraint optimization framework is built. Using a regional matching loss function as the error metric, the radar-camera extrinsic parameters, camera intrinsic parameters, and distortion parameters are optimized in stages to obtain a static benchmark calibration parameter set. Temporal synchronous data of different types of environmental features are collected, and supervisory labels are generated through spatiotemporal alignment and frame-by-frame parameter correction to train a calibration network model. The model is then deployed to a computing platform to output dynamic calibration parameters in real time, completing parameter updates.

[0037] The beneficial effects of this invention are as follows:

[0038] (1) Unlike the traditional PnP method, which only uses the center point of the corner reflector for matching and ignores the inherent geometric characteristics of its triangular region, this invention proposes a region matching loss function based on region constraints, which extends the matching tolerance from a single centroid point to the triangular region. It constructs matching constraints based on the original geometric features of the calibration object, avoiding the propagation of calibration error caused by centroid extraction deviation in single-point matching.

[0039] (2) Design a differentiated error calculation rule for the region / outside in the region matching loss function. This not only preserves the calibration value of the effective point cloud in the region, but also imposes a strong penalty on abnormal projection points outside the region. This achieves the integration of data cleaning and error measurement, and significantly improves the accuracy and stability of the calibration results.

[0040] (3) Design a progressive three-stage parameter optimization strategy of external parameter vector optimization, internal parameter distortion optimization and overall optimization. By locking the optimization variables in layers, the computational dimension and computing power consumption of each stage are greatly reduced, solving the problem that existing technologies cannot balance real-time performance and calibration accuracy. A refined data cleaning mechanism is embedded in the three-stage optimization. In the first stage, after the initial external parameter refinement, outliers are removed by projection verification. The second external parameter refinement is completed based on the cleaned effective data, which effectively avoids noise from participating in subsequent parameter fine-tuning and achieves the dual effect of reducing computing power and improving efficiency and removing noise and maintaining accuracy.

[0041] (4) Construct a closed-loop system that combines physical optimization and deep learning. Use the high-precision three-stage physical optimization results as the training supervision label for the calibration network model. Completely get rid of the dependence on traditional calibration methods, avoid error accumulation and transmission from the source, and greatly improve the generalization ability of the model in real complex road scenarios. Attached Figure Description

[0042] Figure 1 This is a flowchart of the method of the present invention;

[0043] Figure 2 This is a flowchart illustrating the implementation of the method of the present invention;

[0044] Figure 3 This is a diagram of the improved CalibNet network model architecture in this invention;

[0045] Figure 4 This is a schematic diagram of the device configuration in this invention. The arrows in (a) and (b) indicate the positions of the corner reflectors in the pixel coordinate system and the radar coordinate system, respectively. Detailed Implementation

[0046] The present invention will be further described in detail below with reference to embodiments, but the scope of protection of the present invention is not limited thereto.

[0047] like Figure 1 , Figure 2 As shown, the present invention relates to a fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception. The method is described below with reference to specific embodiments.

[0048] S1 constructs a data acquisition scenario using a triangular radar corner reflector, a visible light camera, and a 4D millimeter-wave radar. The triangular radar corner reflector is used to acquire radar point cloud and camera image calibration data at different distances and angles.

[0049] In this invention, triangular radar corner reflectors are arranged within a specified calibration distance range of the target calibration scene, such as... Figure 4 As shown, the corner reflector, as the core reference for calibration, has a radar cross section (RCS) value that is significantly higher than that of surrounding objects, which can provide clear physical characteristics for subsequent screening and clustering of radar point clouds.

[0050] In this embodiment, the corner reflector is fixed to the horizontal calibration bracket to keep its spatial attitude stationary throughout the process, so as to avoid additional calibration errors caused by the displacement of the calibration object. Then, the 4D millimeter-wave radar and the visible light camera are integrated and fixed to the calibration gimbal using a rigid connection method. Through the fastening connection structure, it is ensured that there is no obvious bending deformation of the two during the subsequent movement and angle adjustment process, and the calibration deviation caused by the relative displacement between the sensors is completely eliminated.

[0051] Subsequently, while keeping the triangular radar corner reflector stationary, multiple sets of radar point cloud and camera image calibration data were simultaneously acquired at different distances and angles by adjusting the horizontal rotation and pitch angles of the calibration pan-tilt unit. During the acquisition process, the data synchronization between the radar and the camera was strictly ensured, and the timestamp error was controlled within the millisecond level. The acquired calibration data sets will serve as the raw dataset for subsequent sensor data preprocessing and parameter optimization.

[0052] S2 preprocesses the acquired radar point cloud and camera images to construct 3D-2D corresponding data pairs;

[0053] After preprocessing the radar point cloud and camera images acquired by S1, the two-dimensional image features and three-dimensional point cloud features corresponding to the triangular radar corner reflector are extracted respectively. After completing the one-to-one matching of features, 3D-2D corresponding data pairs are constructed to provide basic data support for subsequent phased parameter optimization.

[0054] (2-1) Extraction of two-dimensional triangular regions from visible light camera images

[0055] Visible light camera images are processed using preset image processing algorithms, such as OpenCV, to perform edge detection and contour filtering, and extract the two-dimensional image region corresponding to each radar corner reflector;

[0056] Specifically, image denoising is first performed using an image filtering algorithm to eliminate the interference of environmental noise on image features. Then, an edge detection algorithm is used to extract edge features from the image. Finally, contour filtering is performed based on features such as contour area and contour shape to accurately extract the two-dimensional triangular region Ω corresponding to each radar corner reflector. i The pixel coordinates of the three vertices of each two-dimensional triangular region are recorded to realize the feature localization of the corner reflector in the image domain; thus, the image feature map is obtained.

[0057] (2-2) Determination of 3D coordinate point set of 4D millimeter-wave radar point cloud

[0058] For 4D millimeter-wave radar point cloud data, candidate points are selected based on the preset features of radar corner reflectors, and spatial clustering is performed on the candidate points. Each cluster is assigned to a radar corner reflector, and its three-dimensional spatial coordinate point set is determined.

[0059] Specifically, the collected point cloud data is processed based on the preset features of the radar corner reflector; candidate points are selected by specifying feature thresholds, and environmental noise and outliers with feature values ​​below the threshold are removed, while valid point cloud data related to the corner reflector are retained. Then, the density clustering algorithm DBSCAN is used to spatially cluster the candidate points. By setting a reasonable neighborhood radius and minimum number of cluster points, each cluster is uniquely associated with a radar corner reflector. Isolated points after clustering are removed, and finally, the three-dimensional spatial coordinate point set of each corner reflector in the radar coordinate system is determined, resulting in a sparse feature map.

[0060] (2-3) Construction of 3D-2D Correspondence Data Pairs

[0061] The two-dimensional image region corresponding to each radar corner reflector is associated with a set of three-dimensional spatial coordinate points to obtain a 3D-2D corresponding data pair with an associated reference.

[0062] Specifically, based on the consistency between the timestamp and the calibration viewpoint, the two-dimensional triangular region Ω of each radar corner reflector is obtained. iBy combining the obtained three-dimensional spatial coordinate point set of the corresponding radar corner reflector, feature matching between the image domain and the point cloud domain is completed, forming a complete 3D-2D corresponding data pair, which provides standardized input data for subsequent phased parameter optimization.

[0063] S3 Based on the characteristics of the acquisition scenario, a regional constraint optimization framework is constructed. The regional matching loss function is used as the error metric for parameter optimization. Combined with the optimization algorithm, the radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters are optimized in stages. The optimal radar-camera extrinsic parameters R*,t*, camera intrinsic parameters A* and camera distortion parameters η* are obtained by solving the problem.

[0064] The "features of the acquisition scene" here actually refers to the geometric features of the triangular region of the radar corner reflector.

[0065] In this embodiment, the initial values ​​of each parameter to be optimized are obtained through existing conventional calibration methods. The complete calibration parameter set includes three categories, totaling 16 parameters to be optimized. The initial intrinsic parameter A0 and the initial distortion parameter η0 of the camera pre-calibration are solved using Zhang Zhengyou's checkerboard calibration method. By shooting the checkerboard calibration board at different angles and combining it with the camera imaging model, the initial values ​​of the parameters including the focal length f are obtained. x ,f y Main point c x ,c y The initial intrinsic parameters A0 and pixel tilt coefficients s, and the initial distortion parameters η0 including radial distortion k1, k2, k3 and tangential distortion p1, p2; the initial extrinsic parameters R0, t0 of the radar-camera system are obtained by coarse calibration using the RANSAC-PnP algorithm, based on the triangular radar corner reflector array with known spatial positions arranged in S1, combined with the two-dimensional triangular region features extracted in the camera image domain and the three-dimensional spatial coordinate point set in the radar coordinate system. The accuracy of the obtained initial extrinsic parameters meets the basic requirements for subsequent fine optimization.

[0066] In S3, the phased targeted optimization includes sequentially executed extrinsic parameter optimization, intrinsic parameter distortion optimization, and overall optimization. Each phase uses the region matching loss function as the error metric and combines the optimization algorithm Trust-Region Reflective (TRR) to achieve parameter optimization.

[0067] Specifically, in the extrinsic parameter optimization, the initial intrinsic parameter A0 of the camera is fixed, and the initial extrinsic parameters R0,t0 are used as initial values, and the corresponding region matching loss function is adopted. The optimization algorithm TRR only optimizes the radar-camera extrinsic parameters to obtain refined extrinsic parameters R1,t1. Based on the refined extrinsic parameters R1,t1, the original 3D-2D corresponding data pairs are projected for verification. In this embodiment, the reprojection error threshold is set to 4 pixels. That is, for each radar point, the Euclidean distance between its projected position on the image plane and the corresponding triangular region is calculated. If the distance is greater than 4 pixels, it is judged as an outlier and removed. Identifying and removing projection outliers completes the cleaning of basic data and removes the interference of invalid data on parameter optimization. The initial intrinsic parameters A0 of the camera are continuously fixed, and the extrinsic parameters of the cleaned data are optimized again using the corresponding region matching loss function and the optimization algorithm TRR to obtain further refined extrinsic parameters R2,t2, which provides an accurate extrinsic parameter basis for subsequent intrinsic parameter distortion optimization.

[0068] In the intrinsic parameter distortion optimization, the refined extrinsic parameters R2, t2 and the cleaned 3D-2D corresponding data pairs are used as inputs. The extrinsic parameters are fixed (and do not participate in the optimization), and the corresponding region matching loss function is adopted. The optimization algorithm TRR only optimizes the camera intrinsic parameters and camera distortion parameters. The camera distortion parameters to be optimized include radial distortion and tangential distortion. By optimizing the parameters, the projection deviation caused by camera lens distortion is corrected, the systematic calibration error caused by lens distortion is eliminated, and the optimized camera intrinsic parameters and camera distortion parameters η1 are obtained.

[0069] In the overall optimization, the refined extrinsic parameters R2,t2, the optimized camera intrinsic parameters, the camera distortion parameter η1, and the cleaned 3D-2D corresponding data pairs are used as inputs, and the corresponding region matching loss function is adopted. Along with the optimization algorithm TRR, the radar-camera extrinsic parameters, camera intrinsic parameters, and camera distortion parameters are optimized simultaneously, ultimately outputting high-precision radar-camera extrinsic parameters, camera intrinsic parameters, and distortion parameters. During the optimization process, the camera intrinsic parameters are specifically corrected to compensate for the systematic projection deviation caused by the initial intrinsic parameter A0 calibration error, thus making up for the error defects of the pre-calibrated intrinsic parameters. After the algorithm iterative convergence, a high-precision complete calibration parameter set is finally output. This parameter set includes the optimal radar-camera extrinsic parameters R* and t*, the optimal camera intrinsic parameter A*, and the optimal camera distortion parameter η*, which can meet the accuracy requirements of fusion calibration of 4D millimeter-wave radar and visible light camera.

[0070] The region matching loss function includes error calculation rules that differentiate between the region and the outside. In this embodiment, the region matching loss function has three specific forms for different optimization stages. Their core structure is the same, and only the dimensions of the optimization variables are different.

[0071] Construct a single-point projection error function:

[0072]

[0073] in, For the i-th radar point, project it onto the camera image as a pixel. The Euclidean distance to the midpoint C of the triangle region of the radar corner reflector. The superscript 'c' in this context refers to a camera. The average distance from the midpoint to the edge of the triangle. This represents the loss corresponding to the i-th set of data. This represents the triangular region of the radar corner reflector;

[0074] Where C is obtained by averaging the pixel coordinates V1, V2, and V3 of the three vertices of the triangle region of the radar corner reflector, i.e. ; , Represents the two-dimensional Euclidean norm;

[0075] Obtain the region matching loss function (general form).

[0076]

[0077] Where N represents the total number of radar points; Let the coordinates of the i-th radar point be expressed as the three-dimensional coordinates in the radar coordinate system. The superscript 'r' stands for radar. Let θ = {R, t, A, η} be the set of variables to be optimized, where R is a (3×3) radar-camera rotation matrix, t is a (3×1) translation vector, A is the camera intrinsic parameter, and η is the camera distortion parameter. This is the radar-camera pixel projection mapping function.

[0078] Accordingly, the three sets of targeted region matching loss functions are as follows:

[0079]

[0080] Set of variables to be optimized ={R,t} For extrinsic parameter vector optimization, the camera intrinsic parameter A0 and distortion parameter η0 are fixed, and only the radar-camera extrinsic parameters R and t are optimized.

[0081]

[0082] Set of variables to be optimized ={ }, R2,t2 are the optimal extrinsic parameters obtained after optimizing the extrinsic parameter vector; this function is used for intrinsic parameter distortion optimization, fixing the radar-camera extrinsic parameters R2 and t2, and only optimizing the camera intrinsic parameter A and distortion parameter η;

[0083]

[0084] Set of variables to be optimized ={ This function is used for overall optimization, simultaneously optimizing all calibration parameters, and outputting the final high-precision radar-camera extrinsic parameters R* and t*, camera intrinsic parameters A*, and distortion parameter η*. The superscript * indicates the optimal solution.

[0085] S4 collects sample data of different types of environmental characteristics, and simultaneously collects image data output by the visible light camera and point cloud data output by the 4D millimeter-wave radar to obtain time-series synchronized data.

[0086] Specifically, sensor data is collected for different types of real road scenarios and various weather and lighting conditions. The real road scenarios include, but are not limited to, urban roads, rural roads, and highways. The various weather and lighting conditions include, but are not limited to, sunny days, rainy days, foggy days, and strong backlight at night.

[0087] S5 uses time-synchronized data as training data and a complete set of calibration parameters consisting of optimal radar-camera extrinsic parameters, camera intrinsic parameters, and camera distortion parameters as a static benchmark calibration parameter set. Through spatiotemporal alignment verification and frame-by-frame parameter adaptive correction, frame-by-frame supervision labels of dynamic road data are generated to train and optimize the calibration network model.

[0088] When generating frame-by-frame supervision labels, the validity of each frame of data is first screened: the number of radar points belonging to corner reflectors in the frame is counted. If it is less than a preset threshold (e.g., 4), or the average reprojection error calculated based on the reference parameters exceeds a preset second threshold (e.g., 2 pixels), the frame data is determined to be unreliable and is directly discarded without generating a label; supervision labels are generated for the remaining valid frames according to the reference parameters or correction parameters.

[0089] In this invention, the mapping logic from static calibration results to dynamic scenes is as follows: since the 4D millimeter-wave radar and the visible light camera are rigidly connected and have no relative displacement or deformation, the static reference calibration parameters obtained in S3 are the core basis for dynamic scene calibration. Only a small-scale correction is needed for the small system errors of the sensor caused by different road scenes and weather lighting conditions, without the need for complete recalibration.

[0090] In S5, the optimal parameters obtained in S3 are used as the initial reference values ​​for each frame tag; the preprocessing process of step S2 is repeated for each frame of synchronization data in S4 to construct 3D-2D corresponding data pairs.

[0091] Calculate the average projection error based on the reference parameters. If the error is less than the preset first threshold, the reference parameters are directly used as the frame label. Otherwise, the reference parameters are used as the initial values, and the radar-camera extrinsic parameters and / or camera intrinsic parameters are adjusted by a preset amplitude. The correction parameters are obtained by minimizing the region matching loss function through the optimization algorithm TRR, and used as the frame label.

[0092] Abnormal frame labels that have an average projection error greater than a preset second threshold after correction are removed, and the final frame-by-frame supervision label set is composed of the baseline parameters and / or correction parameters of all valid frames.

[0093] Specifically, calculating the average projection error based on the reference parameters means that for each frame of data, using the current reference parameters (i.e., the optimal extrinsic, intrinsic, and distortion parameters obtained in S3), all radar points belonging to the corner reflector are projected onto the image plane, the Euclidean distance between each projected point and the corresponding triangular region is calculated, and the average distance of all points is taken as the average projection error of the frame. If the average projection error is less than a preset first threshold (in this embodiment, the first threshold is set to 4 pixels), then the reference parameters are directly used as the label of the frame; otherwise, the reference parameters are used as the initial value, and the radar-camera extrinsic and / or camera intrinsic parameters are adjusted by a preset amplitude. The correction parameters are obtained by minimizing the region matching loss function through an optimization algorithm and used as the label of the frame. The preset amplitude adjustment is specifically set as follows: the rotation angle increment of the extrinsic parameters is set to ±0.1°, the translation vector increment is set to ±0.01m, the focal length increment of the intrinsic parameters is set to ±0.1%, the principal point offset increment is set to ±1 pixel, and the distortion parameter increment is set to ±0.001.

[0094] The optimization algorithm uses Trust-Region Reflective (TRR), with a maximum number of iterations set to 100 and a convergence tolerance of 10. -5 If the average projection error of the frame after correction is still greater than the preset second threshold (the second threshold is set to 2 in this embodiment), then the frame is removed and not included in the training label set.

[0095] In this invention, the calibration network model is an improved CalibNet model;

[0096] like Figure 3 As shown, the input layer of the model includes a feature reconstruction module, which is used to stitch together 3D-2D corresponding data pairs (sparse feature maps generated by 4D millimeter-wave radar point cloud preprocessing and visible light camera image feature maps) and the time difference information between the current frame and adjacent frames as input in the channel dimension;

[0097] A lightweight encoder and a decoder are sequentially arranged after the input layer. The lightweight encoder uses depth-separable convolutional blocks and compressed excitation modules in series to reduce the number of parameters and computational load. The decoder has three parallel regression branches at the end, which output the latest radar-camera extrinsic parameters, camera intrinsic parameters and distortion parameters respectively.

[0098] The model is jointly trained using uncertainty-weighted loss and introduces temporal consistency constraints. Label smoothing loss across multiple consecutive frames is used to encourage continuous changes in the parameter sequence. The model uses a static benchmark calibration parameter set and frame-by-frame supervised labels as training objectives. It employs mean squared error and region matching loss for joint supervision. After training, it can be deployed to a computing platform for rapid dynamic calibration.

[0099] S6 deploys the trained calibration network model to the computing platform, processes the synchronously acquired 4D millimeter-wave radar point cloud and visible light camera images in real time, updates the time delay with preset calibration parameters, and outputs the radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters at the current moment based on the calibration network model, thus completing the update of calibration parameters.

[0100] In this invention, it should be noted that the update delay of the calibration parameters must be less than the data acquisition cycle of the environmental perception sensor, so as to achieve dynamic adaptive adjustment of the calibration parameters.

[0101] In S6, rapid dynamic calibration employs a differentiated parameter update strategy. Radar-camera extrinsic parameters are updated in full real-time, while camera intrinsic parameters and distortion parameters are updated using a base value locking and online incremental correction. Online incremental fine-tuning of camera intrinsic distortion parameters is reasonable because temperature changes, light fluctuations, and slight lens contamination in real-world road environments can cause minor shifts in these parameters. Fixing these parameters completely would easily introduce systematic calibration errors. In practice, online incremental fine-tuning uses the base parameters obtained from static calibration as anchor points, making small-scale corrections only when projection errors exceed the limit for multiple consecutive frames. This effectively corrects the actual deviations of sensor parameters while avoiding computational overload and parameter oscillation issues caused by full updates, thus balancing the real-time performance of dynamic calibration with system stability.

[0102] In application, the trained model is deployed to a computing platform, including but not limited to NVIDIA Jetson AGX Orin / Xavier series, embedded industrial control computer with GPU accelerator card, or FPGA accelerator; the computing platform needs to be pre-installed with a deep learning inference framework, such as TensorRT, ONNX Runtime, etc., and ensure hard real-time synchronization with the sensor data acquisition cycle; the calibration parameter update latency is preset to 10ms~50ms according to the platform's computing power, and must be less than the sensor data acquisition cycle to ensure the continuity and stability of dynamic calibration.

[0103] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0104] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0105] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0106] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0107] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0108] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception, characterized in that: Includes the following steps: S1 constructs a data acquisition scenario using a triangular radar corner reflector, a visible light camera, and a 4D millimeter-wave radar. The triangular radar corner reflector is used to acquire radar point cloud and camera image calibration data at different distances and angles. S2 preprocesses the acquired radar point cloud and camera images to construct 3D-2D corresponding data pairs; S3 Based on the characteristics of the acquisition scenario, a regional constraint optimization framework is constructed. The regional matching loss function is used as the error metric for parameter optimization. Combined with the optimization algorithm, the radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters are optimized in stages to obtain the optimal radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters. S4 collects sample data of different types of environmental characteristics, and simultaneously collects image data output by the visible light camera and point cloud data output by the 4D millimeter-wave radar to obtain time-series synchronized data. S5 uses time-synchronized data as training data and a complete set of calibration parameters consisting of optimal radar-camera extrinsic parameters, camera intrinsic parameters, and camera distortion parameters as a static benchmark calibration parameter set. Through spatiotemporal alignment verification and frame-by-frame parameter adaptive correction, frame-by-frame supervision labels of dynamic road data are generated to train and optimize the calibration network model. S6 deploys the trained calibration network model to the computing platform, processes the synchronously acquired 4D millimeter-wave radar point cloud and visible light camera images in real time, updates the time delay with preset calibration parameters, and outputs the radar-camera extrinsic parameters, camera intrinsic parameters and camera distortion parameters at the current moment based on the calibration network model, thus completing the update of calibration parameters.

2. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 1, characterized in that: In S2, a preset image processing algorithm is used to perform edge detection and contour filtering on the visible light camera image, and extract the two-dimensional image region corresponding to each radar corner reflector. For 4D millimeter-wave radar point cloud data, candidate points are selected based on the preset features of radar corner reflectors, and spatial clustering is performed on the candidate points. Each cluster is assigned to a radar corner reflector, and its three-dimensional spatial coordinate point set is determined. The two-dimensional image region corresponding to each radar corner reflector is associated with a set of three-dimensional spatial coordinate points to obtain a 3D-2D corresponding data pair with an associated reference.

3. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 1, characterized in that: In S3, the phased targeted optimization includes sequentially executed extrinsic parameter optimization, intrinsic parameter distortion optimization, and overall optimization. Each phase uses the region matching loss function as the error metric and combines optimization algorithms to achieve parameter optimization.

4. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 3, characterized in that: In the extrinsic parameter optimization, the initial intrinsic parameters of the camera are fixed, and the initial extrinsic parameters are used as the initial values. The corresponding region matching loss function and optimization algorithm are used to optimize only the radar-camera extrinsic parameters to obtain the refined extrinsic parameters. Based on the refined extrinsic parameters, the original 3D-2D corresponding data pairs are projected for verification to identify and remove projection outliers. The initial intrinsic parameters of the camera are continuously fixed, and the extrinsic parameters are optimized again on the cleaned data using the corresponding region matching loss function and optimization algorithm to obtain the further refined extrinsic parameters.

5. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 4, characterized in that: In the intrinsic distortion optimization, the refined extrinsic parameters and the cleaned 3D-2D corresponding data pairs are used as inputs. The extrinsic parameters are fixed, and the corresponding region matching loss function and optimization algorithm are used to optimize only the camera intrinsic parameters and camera distortion parameters to obtain the optimized camera intrinsic parameters and camera distortion parameters.

6. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 5, characterized in that: In the overall optimization, the refined extrinsic parameters, optimized camera intrinsic parameters, camera distortion parameters, and cleaned 3D-2D corresponding data pairs are used as inputs. The corresponding region matching loss function and optimization algorithm are adopted to simultaneously optimize the radar-camera extrinsic parameters, camera intrinsic parameters, and camera distortion parameters, and finally output high-precision radar-camera extrinsic parameters, camera intrinsic parameters, and distortion parameters.

7. A fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception, as described in any one of claims 3 to 6, characterized in that: The region matching loss function includes error calculation rules that differentiate between designs within and outside the region; Construct a single-point projection error function: , in, For the i-th radar point, project it onto the camera image as a pixel. The Euclidean distance to the midpoint C of the triangle region of the radar corner reflector. The average distance from the midpoint to the edge of the triangle. This represents the loss corresponding to the i-th set of data. This represents the triangular region of the radar corner reflector; Obtain the region matching loss function. , Where N represents the total number of radar points; Let the coordinates of the i-th radar point be expressed as the three-dimensional coordinates in the radar coordinate system. Let θ = {R, t, A, η} be the set of variables to be optimized, where R is the radar-camera rotation matrix, t is the translation vector, A is the camera intrinsic parameter, and η is the camera distortion parameter. This is the radar-camera pixel projection mapping function.

8. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 1, characterized in that: In S5, the optimal parameters obtained in S3 are used as the initial reference values ​​for each frame tag; the preprocessing process of step S2 is repeated for each frame of synchronization data in S4 to construct 3D-2D corresponding data pairs. Calculate the average projection error based on the reference parameters. If the error is less than the preset first threshold, the reference parameters are directly used as the frame label. Otherwise, the reference parameters are used as the initial values, and the radar-camera extrinsic parameters and / or camera intrinsic parameters are adjusted by a preset amplitude. The correction parameters are obtained by minimizing the region matching loss function through the optimization algorithm and used as the frame label. Abnormal frame labels that have an average projection error greater than a preset second threshold after correction are removed, and the final frame-by-frame supervision label set is composed of the baseline parameters and / or correction parameters of all valid frames.

9. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 1, characterized in that: The calibration network model is an improved CalibNet model; The model's input layer includes a feature reconstruction module, which is used to concatenate 3D-2D corresponding data pairs in the channel dimension and the time difference information between the current frame and adjacent frames as input; A lightweight encoder and a decoder are sequentially arranged after the input layer. The lightweight encoder uses depth-separable convolutional blocks and compressed excitation modules connected in series. The decoder has three parallel regression branches at the end, which output new radar-camera extrinsic parameters, camera intrinsic parameters and distortion parameters respectively.

10. The fusion calibration method for 4D millimeter-wave radar and visible light camera suitable for environmental perception according to claim 1, characterized in that: In S6, the rapid dynamic calibration adopts a differentiated parameter update strategy, in which the radar-camera extrinsic parameters are updated in full real time, while the camera intrinsic parameters and distortion parameters are updated by locking the reference value and online incremental correction.