4D imaging millimeter wave radar-camera external parameter calibration method and device
Through the improved PointNet++ network and residual network block extraction, multi-scale features are combined with the Transformer module to fusion across modal features, the problem of 4D imaging millimeter wave radar and camera external parameters are easily affected by the environment, and fast, real-time and accurate external parameters estimation is achieved.
Patent Information
- Application Number
- CN202311830646.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-12-28
AI Technical Summary
The prior art is susceptible to environmental influences during the online calibration of 4D imaging mmWave radar and camera external parameters, resulting in low immediacy and accuracy.
The improved PointNet++ network and multiple residual network blocks are used to extract point cloud and image features at multiple scales, and cross-modal features interaction and fusion are combined with the Transformer module, and the external parameters are estimated using iterative refinement from coarse to fine.
It realizes fast, real-time and accurate online calibration of 4D imaging mmWave radar and camera external parameters without being affected by the environment, improving the accuracy and operating efficiency of external parameters.
Smart Images

Figure CN120235953A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of sensor calibration, and in particular to a method and device for calibrating the extrinsic parameters of a 4D imaging millimeter-wave radar and a camera. Background Art
[0002] Recently, autonomous driving technology has received extensive attention. To make autonomous vehicles safer and more reliable, autonomous driving systems are equipped with various perception sensors. 4D imaging millimeter-wave radars and cameras are two of the most commonly used perception sensors. To accurately fuse the data of 4D imaging millimeter-wave radars and cameras, accurate calibration of the extrinsic parameters of 4D imaging millimeter-wave radars and cameras is an essential prerequisite. Most existing technologies use special calibration targets, such as polygon calibration plates and infrared images, for extrinsic parameter calibration. Although these methods have achieved good results, they can only be used for initial calibration before product delivery and cannot be used for online calibration during product use because they require specific calibration environments and calibration tools. Existing online calibration methods mostly use methods based on traditional feature matching and optimization, but such methods perform poorly in environments with strongly textured surfaces and shadows. With the development of deep learning technology, methods using deep learning for extrinsic parameter calibration have emerged. Such methods show the robustness of the algorithm in the face of different environments. However, in order to accurately estimate the calibration matrix in a larger range, an iterative network will be applied, which will increase the scale of the network and reduce the operating efficiency of the online calibration method. Therefore, how to quickly, real-time, and accurately perform online calibration of the extrinsic parameters of 4D imaging millimeter-wave radars and cameras without being affected by the environment has become a problem to be solved in this field. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and device for calibrating the extrinsic parameters of a 4D imaging millimeter-wave radar and a camera, which can overcome the defects of the existing technologies that are vulnerable to environmental influences and result in low instantaneity and accuracy of online calibration of the extrinsic parameters of 4D imaging millimeter-wave radars and cameras.
[0004] The purpose of the present invention can be achieved by the following technical solutions:
[0005] According to the first aspect of the present invention, a method for calibrating the extrinsic parameters of a 4D imaging millimeter-wave radar and a camera is provided, including the following steps:
[0006] S1, acquiring and preprocessing the point cloud of the 4D imaging millimeter-wave radar and the camera image, and initializing the initial extrinsic parameter matrix and the camera intrinsic parameter matrix of the 4D imaging millimeter-wave radar and the camera;
[0007] S2. Based on the preprocessed point cloud and camera image, use the improved PointNet++ network to extract point cloud features at multiple scales, and use multiple residual network blocks to extract image features at multiple scales. The improved PointNet++ network synchronously encodes the spatial, velocity, and intensity information of the preprocessed point cloud;
[0008] S3. Based on the point cloud features, the image features, the initial extrinsic matrix, and the camera intrinsic matrix, use the Transformer module to obtain point cloud-image fusion features at multiple scales;
[0009] S4. Based on the point cloud-image fusion features at multiple scales, estimate the extrinsic parameters of the 4D imaging millimeter-wave radar-camera in an iterative refinement manner from coarse to fine.
[0010] As a preferred technical solution, the process of using the improved PointNet++ network to extract point cloud features at multiple scales specifically includes:
[0011] S201. Sample the preprocessed point cloud using the farthest point sampling method to obtain sampling points;
[0012] S202. Obtain a point set with the sampling points as the center of the sphere to get the grouped point cloud. The number of 4D radar points in the point set is the same as the preset number of search points;
[0013] S203. Use the improved PointNet++ network to perform feature aggregation on the grouped point cloud to obtain point cloud features at a single scale;
[0014] S204. Based on the preset number of scales, repeat S201 - S203 to obtain point cloud features at multiple scales.
[0015] As a preferred technical solution, the process of using multiple residual network blocks to extract image features at multiple scales specifically includes:
[0016] S211. Based on the preprocessed camera image, use the first 2D convolutional neural network and the second 2D convolutional neural network to obtain the first image data and the residual data. The stride of the first 2D convolutional neural network is 2, and the stride of the second 2D convolutional neural network is 1;
[0017] S212. Based on the preprocessed camera image, use the third 2D convolutional neural network to obtain the second image data, and add the second image data to the residual data to obtain the third image data. The stride of the third 2D convolutional neural network is 2;
[0018] S213. Based on the third image data, use the activation function to obtain image features at a single scale;
[0019] S214. Based on the preset number of residual network blocks, repeat S211 - S213 to obtain image features at multiple scales. Each of the residual network blocks includes a first 2D convolutional neural network, a second 2D convolutional neural network, and a third 2D convolutional neural network.
[0020] As a preferred technical solution, S3 specifically includes:
[0021] S301. Obtain sampling points of the point cloud features, project the sampling points onto the image features by using the initial extrinsic parameter matrix and the camera intrinsic parameter matrix, and obtain projection points and first image features. The first image features are the image features corresponding to the projection points;
[0022] S302. Connect the first point cloud features and the first image features to obtain connection features. The first point cloud features are the point cloud features corresponding to the sampling points;
[0023] S303. Use the Transformer module to perform cross-modal feature interaction among the first point cloud features, the first image features, and the connection features to obtain cross-modal interaction features;
[0024] S304. Aggregate the first point cloud features, the first image features, and the cross-modal interaction features to obtain point cloud-image fusion features.
[0025] As a preferred technical solution, the way of iterative refinement from coarse to fine specifically includes:
[0026] S401. According to the point cloud-image fusion features at the largest scale, use a combination layer of 2D convolution - activation function - max pooling function and a combination layer of 2D convolution - ReLU activation function - average pooling function in sequence to estimate the initial extrinsic parameter values;
[0027] S402. According to the point cloud-image fusion features at other scales, use a fully connected layer to iteratively optimize the initial extrinsic parameter values to obtain the final extrinsic parameter estimation values.
[0028] According to the second aspect of the present invention, there is provided a 4D imaging millimeter-wave radar-camera extrinsic parameter calibration device, including a data acquisition module, a feature extraction module, a feature fusion module, and an extrinsic parameter estimation module that are signal-connected.
[0029] The data acquisition module is used to acquire and preprocess 4D imaging millimeter-wave radar point clouds and camera images, initialize the initial extrinsic parameter matrix and the camera intrinsic parameter matrix of the 4D imaging millimeter-wave radar-camera, send the preprocessed point clouds and camera images to the feature extraction module, and send the initial extrinsic parameter matrix and the camera intrinsic parameter matrix to the feature fusion module;
[0030] The feature extraction module is used to receive the preprocessed point cloud and camera image, extract point cloud features at multiple scales using an improved PointNet++ network, extract image features at multiple scales using multiple residual network blocks, and send the point cloud features and the image features to the feature fusion module. The improved PointNet++ network synchronously encodes the spatial, velocity, and intensity information of the preprocessed point cloud;
[0031] The feature fusion module is used to receive the point cloud features, the image features, the initial extrinsic parameter matrix, and the camera intrinsic parameter matrix, obtain point cloud-image fusion features at multiple scales using a Transformer module, and send the point cloud-image fusion features at multiple scales to the extrinsic parameter estimation module;
[0032] The extrinsic parameter estimation module is used to receive the point cloud-image fusion features at multiple scales and estimate the extrinsic parameters of the 4D imaging millimeter-wave radar-camera in a coarse-to-fine iterative refinement manner.
[0033] As a preferred technical solution, the feature extraction module includes a point cloud sampling unit, a point cloud grouping unit, and a point cloud feature aggregation unit that are signal-connected,
[0034] The point cloud sampling unit is used to sample the preprocessed point cloud using the farthest point sampling method, obtain the sampled points, and send them to the point cloud grouping unit;
[0035] The point cloud grouping unit is used to receive the sampled points, obtain a point set with the sampled points as the center of the sphere, obtain the grouped point cloud, and send it to the point cloud feature aggregation unit. The number of 4D radar points in the point set is the same as the preset number of search points;
[0036] The point cloud feature aggregation unit is used to receive the grouped point cloud and perform feature aggregation on the grouped point cloud using an improved PointNet++ network to obtain point cloud features at a single scale.
[0037] As a preferred technical solution, the feature extraction module further includes a first image and residual acquisition unit, a third image acquisition unit, and an image feature acquisition unit that are signal-connected,
[0038] The first image and residual acquisition unit is used to obtain first image data and residual data based on the preprocessed camera image using a first 2D convolutional neural network and a second 2D convolutional neural network and send them to the third image acquisition unit. The stride of the first 2D convolutional neural network is 2, and the stride of the second 2D convolutional neural network is 1;
[0039] The third image acquisition unit is configured to receive the first image data and the residual data, obtain second image data by using a third 2D convolutional neural network based on the preprocessed camera image, add the second image data to the residual data to obtain third image data, and send the third image data to the image feature acquisition unit, where the stride of the third 2D convolutional neural network is 2;
[0040] The image feature acquisition unit is configured to receive the third image data and obtain image features at a single scale by using an activation function.
[0041] As a preferred technical solution, the feature fusion module includes a first image feature acquisition unit, a connection feature acquisition unit, a cross-modal interaction feature acquisition unit, and a feature aggregation unit that are connected in series,
[0042] The first image feature acquisition unit is configured to obtain sampling points of the point cloud feature, project the sampling points onto the image feature by using the initial extrinsic parameter matrix and the camera intrinsic parameter matrix, obtain projection points and first image features, and send the projection points and the first image features to the connection feature acquisition unit, where the first image feature is the image feature corresponding to the projection points;
[0043] The connection feature acquisition unit is configured to receive the first image feature, connect the first point cloud feature and the first image feature to obtain a connection feature, and send the connection feature to the cross-modal interaction feature acquisition unit, where the first point cloud feature is the point cloud feature corresponding to the sampling points;
[0044] The cross-modal interaction feature acquisition unit is configured to receive the connection feature and perform cross-modal feature interaction among the first point cloud feature, the first image feature, and the connection feature by using a Transformer module to obtain cross-modal interaction features, and send the cross-modal interaction features to the feature aggregation unit;
[0045] The feature aggregation unit is configured to receive the cross-modal interaction features and aggregate the first point cloud feature, the first image feature, and the cross-modal interaction features to obtain point cloud-image fusion features.
[0046] As a preferred technical solution, the extrinsic parameter estimation module includes an initial extrinsic parameter value estimation unit and a final extrinsic parameter value estimation unit that are connected in series,
[0047] The initial extrinsic parameter value estimation unit is configured to estimate an initial extrinsic parameter value according to the point cloud-image fusion features at the maximum scale by sequentially using a combination layer of 2D convolution-activation function-maximum pooling function and a combination layer of 2D convolution-ReLU activation function-average pooling function, and send the initial extrinsic parameter value to the final extrinsic parameter value estimation unit;
[0048] The final extrinsic parameter value estimation unit is configured to receive the initial extrinsic parameter value, and iteratively optimize the initial extrinsic parameter value by using a fully connected layer according to the point cloud-image fusion features at other scales, so as to obtain the final extrinsic parameter estimation value.
[0049] Compared with the prior art, the present invention has the following beneficial effects:
[0050] 1. The present invention corrects the extrinsic parameters for a 4D imaging millimeter-wave radar and an image. In view of the rich feature of the 4D radar point cloud, the traditional Pointnet++ network is improved, and the spatial, velocity, and intensity information of the preprocessed point cloud is synchronously encoded, so that it can make full use of the rich point cloud information of the 4D radar point cloud, and estimates the extrinsic parameters by a method of gradually refining from coarse to fine. The extrinsic parameters are estimated jointly by using the image features and the 4D imaging millimeter-wave radar point cloud features at different scales, which can be immune to the influence of the external environment and improve the accuracy of the extrinsic parameter estimation.
[0051] 2. In the process of using multiple residual network blocks to extract image features at multiple scales, the present invention can extract multi-scale image features by using only a small number of convolutional neural networks, and fully considers three modalities: point modality, image modality, and connection modality. The Transformer module is used for feature interaction and feature fusion between modalities, which can not only make full use of the 3D geometric information contained in the 4D imaging millimeter-wave radar points and the rich semantic information of the image to achieve the effect of full fusion of the 4D imaging millimeter-wave radar point features and the image features, but also avoid a large increase in the scale of the network, improve the operation efficiency of the extrinsic parameter estimation, and further improve the immediacy of the online calibration of the extrinsic parameters.
[0052] 3. The present invention adopts the method of specifying the number of neighborhood points, presets the specified number of search points, and obtains the grouped point cloud according to the preset number, avoiding the situation of not finding neighborhood points, which can improve the efficiency of point cloud feature extraction and further improve the immediacy of the online calibration of the extrinsic parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is a schematic flowchart of the method in an embodiment of the present invention;
[0054] Figure 2 is a schematic principle diagram of point cloud-image fusion in an embodiment of the present invention;
[0055] Figure 3 is a schematic principle diagram of iterative extrinsic parameter estimation in an embodiment of the present invention;
[0056] Figure 4 is a schematic structural diagram of the device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0058] Embodiment
[0059] As Figure 1 shown, this embodiment provides a 4D imaging millimeter-wave radar-camera extrinsic parameter calibration method, and the specific implementation process is as follows:
[0060] Implement step S1, obtain and preprocess the 4D imaging millimeter-wave radar point cloud and the camera image, and initialize the initial extrinsic parameter matrix and the camera intrinsic parameter matrix of the 4D imaging millimeter-wave radar-camera. Specifically, downsample the 4D imaging millimeter-wave radar point cloud to 512 points by random sampling, and scale the camera image to 1 / 2 of the original size.
[0061] Implement step S2, based on the preprocessed point cloud and camera image, use the improved PointNet++ network to extract point cloud features at multiple scales, and use multiple residual network blocks to extract image features at multiple scales. The improved PointNet++ network synchronously encodes the spatial, velocity, and intensity information of the preprocessed point cloud.
[0062] Input the input point cloud and camera image into the point cloud feature extraction network and the image feature extraction network respectively. Among them, the point cloud feature extraction process is as follows:
[0063] Step S201, sample the preprocessed point cloud using the farthest point sampling method to obtain sampling points. That is, use the farthest point sampling method to sample the input preprocessed point cloud, and select a part of the points to represent the entire point cloud;
[0064] Step S202, obtain the point set with the sampling point as the center of the sphere to get the grouped point cloud, and the number of 4D radar points in the point set is the same as the preset number of search points. That is, by setting the number of search points, obtain the point set with each sampling point as the center of the sphere and search for the points that meet the number of search points to group the input preprocessed point cloud;
[0065] Step S203: Use the improved PointNet++ network to perform feature aggregation on the grouped point clouds to obtain point cloud features at a single scale. The original PointNet++ network only encodes the spatial information of the point clouds. 4D radar points contain not only the spatial information of the points but also the motion information and intensity information of the points. To make full use of the 4D radar point cloud information, the PointNet++ block is improved so that it can encode the spatial, motion, and intensity information of the point clouds simultaneously. Use structures such as multi-layer perceptrons to perform feature learning on the five-dimensional information of the spatial information, velocity information, and intensity information of the grouped point clouds and use the max pooling operation for feature aggregation;
[0066] Step S204: Based on the preset number of scales, repeat steps S201 - S203 to obtain point cloud features at multiple scales. Specifically, repeat the above steps to obtain the 4D imaging millimeter-wave radar point cloud features at three scales. The number of 4D imaging millimeter-wave radar points extracted at each scale is 256, 128, and 64 respectively. The coordinate sets of the 4D imaging millimeter-wave radar points at different scales are denoted as P0, P1, and P2 respectively, and the point cloud features extracted at different scales are denoted as RF0, RF1, and RF2 respectively;
[0067] While steps S201 - S204 are being executed, the image feature extraction process is as follows:
[0068] Step S211: Based on the preprocessed camera image, use the first 2D convolutional neural network with a stride of 2 and the second 2D convolutional neural network with a stride of 1 to obtain the first image data and the residual data;
[0069] Step S212: Based on the preprocessed camera image, use the third 2D convolutional neural network with a stride of 2 to obtain the second image data, and add the second image data to the residual data to obtain the third image data;
[0070] Step S213: Based on the third image data, use the activation function to obtain the image features at a single scale.
[0071] Steps S211 - S213, that is, take the preprocessed camera image as the input. First, pass it through a 2D convolutional neural network with a stride of 2 and a convolution kernel of 3×3 and a 2D convolutional neural network with a stride of 1 and a convolution kernel of 3×3 to obtain the residual between the input and the output first image data; then, the input image data passes through another 2D convolutional neural network with a stride of 2 and a convolution kernel of 3×3 and is added to the previously obtained residual; finally, after passing through the LeakyReLU activation function, the image features at a single scale are obtained.
[0072] Step S214, based on the preset number of residual network blocks, repeat Step S211 to Step S213 to obtain image features at multiple scales. Each residual network block includes a first 2D convolutional neural network, a second 2D convolutional neural network, and a third 2D convolutional neural network, that is, each residual network block is composed of two 2D convolutional neural networks with a stride of 2 and a convolutional kernel of 3×3 and a 2D convolutional neural network with a stride of 1 and a convolutional kernel of 3×3. In this embodiment, a total of three consecutive residual network blocks are set, and image features at three scales can be obtained. The sizes of the image features extracted at each scale are 1 / 4, 1 / 8, and 1 / 16 of the original image size, respectively. The features of the image at different scales are denoted as IF0, IF1, and IF2.
[0073] Implement Step S3. Based on the point cloud features, image features, initial extrinsic matrix, and camera intrinsic matrix, use the Transformer module to obtain point cloud-image fusion features at multiple scales. As Figure 2 shown, the specific process is as follows:
[0074] Step S301, using the initial extrinsic matrix of the 4D imaging millimeter-wave radar-camera and the camera intrinsic matrix K, a 4D imaging millimeter-wave radar sampling point with coordinates P and related point cloud feature RF (i.e., the first point cloud feature) in the 4D imaging millimeter-wave radar coordinate system can be projected onto the feature map of the RGB image to obtain a projected point p with coordinates (u, v) and the corresponding image feature IF p (i.e., the first image feature). The projection process is as follows:
[0075]
[0076] In the formula, Z represents the depth value of the 4D imaging millimeter-wave radar point in the camera coordinate system, P represents the homogeneous coordinate of the 4D imaging millimeter-wave radar point. After multiplying T IR and P, take the first three coordinates so that it can be multiplied by K. If the coordinates (u, v) are not integers, use bilinear interpolation for retrieval;
[0077] Step S302, connect the retrieved image feature IF p with the input point cloud feature RF to obtain a connection feature CF;
[0078] Step S303, use the Transformer module to achieve cross-modal feature interaction and multi-modal feature fusion among the 4D radar point feature RF, image feature IF p and connection feature CF. Specifically: project the 4D radar point feature RF onto the query feature Q P = RF·W Q and the key feature K P= RF·W K Similarly, the image features are projected onto Q I and K I For the connection feature CF, first use two linear transformations to map it to different feature spaces to obtain and After that, project and onto the value features and The above W Q 、W K 、 and are learnable linear mappings. Therefore, the context information of the image modality is obtained through the attention weight W P←I = Softmax(Q I (K P ) T ), and then interact with the connection modality through to obtain the connection feature guided by the image modality. Similarly, the context information of the point modality can be obtained through and encoded into the connection feature guided by the point modality.
[0079] Step S304, aggregate the original 4D radar feature RF, the image feature IF p and the features F P←I and F I←P with cross-modal interaction to obtain the point cloud-image fusion feature, that is as the fusion feature.
[0080] Implement step S4. Based on the point cloud-image fusion feature at multiple scales, use the method of iterative refinement from coarse to fine to estimate the extrinsic parameters of the 4D imaging millimeter-wave radar-camera. As Figure 3 shown, the specific process is as follows:
[0081] Step S401, according to the point cloud-image fusion feature at the largest scale, successively use the combination layer of 2D convolution-activation function-maximum pooling function and the combination layer of 2D convolution-ReLU activation function-average pooling function to estimate the initial extrinsic parameter value. Specifically, apply two combination layers of 2D convolution-ReLU activation function-maximum pooling function and one combination layer of 2D convolution-ReLU activation function-average pooling function to the fusion feature at the largest scale in sequence to obtain the feature matching information of the 4D imaging millimeter-wave radar and the camera;
[0082] Step S402, according to the point cloud-image fusion features at other scales, use the fully connected layer to iteratively optimize the initial extrinsic parameter value to obtain the final extrinsic parameter estimation value. Specifically, it includes:
[0083] Step S4021: According to the feature matching information of the 4D imaging millimeter-wave radar and the camera, two independent fully connected layers are used to obtain the initial translation estimate and the initial rotation estimate respectively, where the initial translation is represented by a translation vector and the initial rotation is represented by a quaternion. The extrinsic matrix at the second scale can be obtained based on the initial translation estimate and the initial rotation estimate.
[0084] Step S4022: Use the extrinsic matrix at the second scale to perform coordinate transformation on the 4D imaging millimeter-wave radar point set at the first scale: According to the transformed coordinate P1 ′ , through the fusion step S3 and the estimation steps S401 and S4021, the extrinsic matrix at the first scale can be obtained.
[0085] Step S4023: Use the extrinsic matrix at the first scale to perform coordinate transformation on the 4D imaging millimeter-wave radar point set at the zero scale: According to the transformed coordinate P0 ′ , through the fusion step S3 and the estimation steps S401 and S4021, the extrinsic matrix at the zero scale can be obtained.
[0086] Step S4024: Finally, the extrinsic parameter estimation between the 4D imaging millimeter-wave radar and the camera is obtained by multiplying .
[0087] Furthermore, this embodiment also provides a 4D imaging millimeter-wave radar-camera extrinsic parameter calibration device, as Figure 4As shown. The device includes a data acquisition module, a feature extraction module, a feature fusion module, and an extrinsic parameter estimation module that are signal-connected. The data acquisition module is used to acquire and preprocess 4D imaging millimeter-wave radar point clouds and camera images, initialize the initial extrinsic parameter matrix and camera intrinsic parameter matrix of the 4D imaging millimeter-wave radar-camera, send the preprocessed point clouds and camera images to the feature extraction module, and send the initial extrinsic parameter matrix and camera intrinsic parameter matrix to the feature fusion module; the feature extraction module is used to receive the preprocessed point clouds and camera images, extract point cloud features at multiple scales using an improved PointNet++ network, and extract image features at multiple scales using multiple residual network blocks, send the point cloud features and image features to the feature fusion module, and the improved PointNet++ network synchronously encodes the spatial, velocity, and intensity information of the preprocessed point clouds; the feature fusion module is used to receive the point cloud features, image features, initial extrinsic parameter matrix, and camera intrinsic parameter matrix, and use a Transformer module to obtain point cloud-image fusion features at multiple scales, and send the point cloud-image fusion features at multiple scales to the extrinsic parameter estimation module; the extrinsic parameter estimation module is used to receive the point cloud-image fusion features at multiple scales, and estimate the extrinsic parameters of the 4D imaging millimeter-wave radar-camera in a coarse-to-fine iterative refinement manner. Each module can implement all the foregoing method steps, and the specific implementation process is as described above and will not be elaborated here.
[0088] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A 4D imaging millimeter-wave radar-camera extrinsic parameter calibration method, characterized in that, Including the following steps: S1. Obtain and preprocess the 4D imaging millimeter-wave radar point cloud and camera image, and initialize the initial extrinsic parameter matrix and camera intrinsic parameter matrix of the 4D imaging millimeter-wave radar-camera; S2. Based on the preprocessed point cloud and camera image, use the improved PointNet++ network to extract point cloud features at multiple scales, and use multiple residual network blocks to extract image features at multiple scales. The improved PointNet++ network synchronously encodes the spatial, velocity, and intensity information of the preprocessed point cloud; S3. Based on the point cloud features, the image features, the initial extrinsic parameter matrix, and the camera intrinsic parameter matrix, use the Transformer module to obtain point cloud-image fusion features at multiple scales; S4. Based on the point cloud-image fusion features at multiple scales, estimate the extrinsic parameters of the 4D imaging millimeter-wave radar-camera in a coarse-to-fine iterative refinement manner.
2. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration method according to claim 1, characterized in that, The process of using the improved PointNet++ network to extract point cloud features at multiple scales specifically includes: S201. Sample the preprocessed point cloud using the farthest point sampling method to obtain sampling points; S202. Obtain a point set with the sampling point as the center of the sphere to obtain the grouped point cloud. The number of 4D radar points in the point set is the same as the preset number of search points; S203. Use the improved PointNet++ network to perform feature aggregation on the grouped point cloud to obtain point cloud features at a single scale; S204. Based on the preset number of scales, repeat S201-S203 to obtain point cloud features at multiple scales.
3. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration method according to claim 1, wherein, The process of using multiple residual network blocks to extract image features at multiple scales specifically includes: S211. Based on the preprocessed camera image, use the first 2D convolutional neural network and the second 2D convolutional neural network to obtain the first image data and residual data. The stride of the first 2D convolutional neural network is 2, and the stride of the second 2D convolutional neural network is 1; S212. Based on the preprocessed camera image, use the third 2D convolutional neural network to obtain the second image data, and add the second image data to the residual data to obtain the third image data. The stride of the third 2D convolutional neural network is 2; S213. Based on the third image data, use the activation function to obtain image features at a single scale; S214. Based on the preset number of residual network blocks, repeat S211-S213 to obtain image features at multiple scales. Each residual network block includes the first 2D convolutional neural network, the second 2D convolutional neural network, and the third 2D convolutional neural network.
4. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration method according to claim 1, characterized in that The S3 specifically includes: S301. Obtain the sampling points of the point cloud features, project the sampling points onto the image features using the initial extrinsic parameter matrix and the camera intrinsic parameter matrix to obtain projection points and the first image features. The first image features are the image features corresponding to the projection points; S302. Connect the first point cloud features and the first image features to obtain connection features. The first point cloud features are the point cloud features corresponding to the sampling points; S303. Use the Transformer module to perform cross-modal feature interaction among the first point cloud feature, the first image feature, and the connection feature to obtain a cross-modal interaction feature. S304. Aggregate the first point cloud feature, the first image feature, and the cross-modal interaction feature to obtain a point cloud-image fusion feature.
5. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration method according to claim 1, wherein The way of iterative refinement from coarse to fine specifically includes: S401. According to the point cloud-image fusion feature at the largest scale, successively use a combination layer of 2D convolution-activation function-max pooling function and a combination layer of 2D convolution-ReLU activation function-average pooling function to estimate the initial extrinsic parameter value. S402. According to the point cloud-image fusion feature at other scales, use a fully connected layer to iteratively optimize the initial extrinsic parameter value to obtain the final extrinsic parameter estimation value.
6. A 4D imaging millimeter-wave radar-camera extrinsic parameter calibration device, characterized in that, It includes a data acquisition module, a feature extraction module, a feature fusion module, and an extrinsic parameter estimation module with signal connections. The data acquisition module is used to acquire and preprocess the 4D imaging millimeter-wave radar point cloud and the camera image, initialize the initial extrinsic parameter matrix and the camera intrinsic parameter matrix of the 4D imaging millimeter-wave radar-camera, send the preprocessed point cloud and camera image to the feature extraction module, and send the initial extrinsic parameter matrix and the camera intrinsic parameter matrix to the feature fusion module. The feature extraction module is used to receive the preprocessed point cloud and camera image, extract the point cloud feature at multiple scales using an improved PointNet++ network, extract the image feature at multiple scales using multiple residual network blocks, and send the point cloud feature and the image feature to the feature fusion module. The improved PointNet++ network synchronously encodes the spatial, velocity, and intensity information of the preprocessed point cloud. The feature fusion module is used to receive the point cloud feature, the image feature, the initial extrinsic parameter matrix, and the camera intrinsic parameter matrix, and use the Transformer module to obtain the point cloud-image fusion feature at multiple scales, and send the point cloud-image fusion feature at multiple scales to the extrinsic parameter estimation module. The extrinsic parameter estimation module is used to receive the point cloud-image fusion feature at multiple scales and estimate the extrinsic parameters of the 4D imaging millimeter-wave radar-camera in an iterative refinement manner from coarse to fine.
7. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration device according to claim 6, characterized in that, The feature extraction module includes a point cloud sampling unit, a point cloud grouping unit, and a point cloud feature aggregation unit with signal connections. The point cloud sampling unit is used to sample the preprocessed point cloud using the farthest point sampling method to obtain sampling points and send them to the point cloud grouping unit. The point cloud grouping unit is used to receive the sampling points, obtain a point set with the sampling points as the center of the sphere, obtain the grouped point cloud and send it to the point cloud feature aggregation unit. The number of 4D radar points in the point set is the same as the preset number of search points. The point cloud feature aggregation unit is used to receive the grouped point cloud and use the improved PointNet++ network to perform feature aggregation on the grouped point cloud to obtain the point cloud feature at a single scale.
8. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration device according to claim 6, characterized in that, The feature extraction module further includes a first image and residual acquisition unit, a third image acquisition unit, and an image feature acquisition unit that are signal-connected. The first image and residual acquisition unit is configured to obtain first image data and residual data based on the preprocessed camera image by using a first 2D convolutional neural network and a second 2D convolutional neural network, and send them to the third image acquisition unit. The stride of the first 2D convolutional neural network is 2, and the stride of the second 2D convolutional neural network is 1. The third image acquisition unit is configured to receive the first image data and the residual data, obtain second image data based on the preprocessed camera image by using a third 2D convolutional neural network, add the second image data to the residual data to obtain third image data, and send it to the image feature acquisition unit. The stride of the third 2D convolutional neural network is 2. The image feature acquisition unit is configured to receive the third image data and obtain image features at a single scale by using an activation function.
9. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration device according to claim 6, characterized in that, The feature fusion module includes a first image feature acquisition unit, a connection feature acquisition unit, a cross-modal interaction feature acquisition unit, and a feature aggregation unit that are signal-connected. The first image feature acquisition unit is configured to obtain sampling points of the point cloud feature, project the sampling points onto the image feature by using the initial extrinsic parameter matrix and the camera intrinsic parameter matrix, obtain projection points and a first image feature, and send them to the connection feature acquisition unit. The first image feature is the image feature corresponding to the projection points. The connection feature acquisition unit is configured to receive the first image feature, connect the first point cloud feature and the first image feature to obtain a connection feature, and send it to the cross-modal interaction feature acquisition unit. The first point cloud feature is the point cloud feature corresponding to the sampling points. The cross-modal interaction feature acquisition unit is configured to receive the connection feature, perform cross-modal feature interaction among the first point cloud feature, the first image feature, and the connection feature by using a Transformer module, obtain cross-modal interaction features, and send them to the feature aggregation unit. The feature aggregation unit is configured to receive the cross-modal interaction features, and aggregate the first point cloud feature, the first image feature, and the cross-modal interaction features to obtain point cloud-image fusion features.
10. The 4D imaging millimeter-wave radar-camera extrinsic parameter calibration device according to claim 6, wherein The extrinsic parameter estimation module includes an initial extrinsic parameter value estimation unit and a final extrinsic parameter value estimation unit that are signal-connected. The initial extrinsic parameter value estimation unit is configured to estimate an initial extrinsic parameter value based on the point cloud-image fusion features at the maximum scale by using a combination layer of 2D convolution-activation function-maximum pooling function and a combination layer of 2D convolution-ReLU activation function-average pooling function in sequence, and send it to the final extrinsic parameter value estimation unit. The final extrinsic parameter value estimation unit is configured to receive the initial extrinsic parameter value, and perform iterative optimization on the initial extrinsic parameter value by using a fully connected layer based on the point cloud-image fusion features at other scales to obtain a final extrinsic parameter estimation value.
Citation Information
Patent Citations
Laser radar and camera external parameter online calibration method and device and storage medium
CN115393448A
Real-time target detection method based on laser radar and camera data fusion
CN115546594A
Camera and millimeter wave radar front fusion road surface target detection method
CN116259024A
Three-dimensional target detection method based on multi-modal fusion and deep attention mechanism
CN116612468A
End-to-end camera and laser radar calibration method for automatic driving
CN117152262A
Cited By
4D millimeter wave radar and visible light camera fusion calibration method suitable for environmental perception
CN122131258A
A 4D millimeter wave radar and visible light camera fusion calibration method suitable for environmental perception
CN122131258B