Point cloud density enhancement method based on near-infrared image and millimeter wave radar point cloud fusion
By fusing near-infrared images with millimeter-wave radar point cloud data, point cloud density is enhanced using depth estimation algorithm and registration algorithm, solving the problem of sparseness and low resolution of millimeter-wave radar point clouds, and achieving more refined environmental perception and all-weather perception capabilities.
Patent Information
- Application Number
- CN202510005113.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-02
AI Technical Summary
Point cloud data generated by millimeter-wave radars are usually sparse and have low resolution, making it difficult to provide detailed environmental details.
By fusing near-infrared images with millimeter-wave radar point cloud data, absolute depth estimation is used to perform absolute depth estimation, convert it into three-dimensional point cloud data, and register and fuse it through ICP algorithm and K nearest neighbor algorithm to enhance point cloud density.
It significantly improves the resolution and detailed performance of the millimeter-wave radar point cloud, provides more refined spatial details, achieves more accurate perception of the environment, and forms an all-weather and full-environment perception system.
Smart Images

Figure CN120088145A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to technologies such as artificial intelligence depth estimation algorithms, monocular vision, and millimeter-wave radar sensors, and specifically relates to a point cloud enhancement method based on the fusion of near-infrared images and millimeter-wave radar point clouds. Background Art
[0002] With the rapid development of fields such as autonomous driving, intelligent transportation, and industrial automation, environmental perception technology has become a key link. As an important sensor, millimeter-wave radar is widely used in tasks such as vehicle detection, obstacle recognition, and distance measurement due to its excellent performance in various harsh weather conditions (such as rain, snow, and fog). However, the point cloud data generated by millimeter-wave radar is often sparse and has low resolution, making it difficult to provide detailed environmental details.
[0003] Near-infrared cameras have their unique advantages in environmental perception. Near-infrared light is not easily affected by visible light conditions, can still provide clear images in low-light environments, and has a certain ability to resist fog and dust. Combining near-infrared cameras with millimeter-wave radar is expected to make up for the limitations of single sensors and enhance the richness and accuracy of point cloud data.
[0004] The present invention fuses millimeter-wave radar point cloud data and near-infrared image data, and designs a point cloud enhancement method based on the fusion of near-infrared images and millimeter-wave radar point clouds. By fusing the perception data of near-infrared cameras and millimeter-wave radars, the resolution and detail performance of the point cloud can be significantly improved, making up for the shortage of sparse millimeter-wave radar point clouds, providing more refined spatial details for the point cloud data, and achieving more accurate perception of the environment. The reliability of millimeter-wave radar under harsh weather conditions, combined with the superior performance of near-infrared cameras in low-light environments, can form an all-weather and all-environment perception system. This multi-sensor fusion technology can provide consistent and high-quality perception data under different environmental conditions, improving the robustness and reliability of the system. Summary of the Invention
[0005] The purpose of the present invention is to propose a method for enhancing the point cloud density by fusing near-infrared images and millimeter-wave radar point clouds, so as to solve the problem that the point cloud data generated by millimeter-wave radar is usually sparse and has low resolution, making it difficult to provide detailed environmental details.
[0006] The solution to achieve the purpose of the present invention is: a method for enhancing the point cloud density based on the fusion of near-infrared images and millimeter-wave radar point clouds, including the following steps:
[0007] Step 1, preprocess the original image data captured by the near-infrared camera, and the preprocessing operations include image scaling and pixel data normalization operations;
[0008] Step 2: Use a monocular depth estimation algorithm to perform absolute depth estimation on the preprocessed near-infrared image data, and obtain the depth information of each pixel point in the preprocessed near-infrared image;
[0009] Step 3: Utilize the depth information of each pixel point in the preprocessed near-infrared image to convert the two-dimensional preprocessed near-infrared image data into three-dimensional near-infrared image point cloud data, and transform it into the radar coordinate system;
[0010] Step 4: Use a pass-through filter and a ground point removal algorithm to denoise the three-dimensional near-infrared image point cloud data in the radar coordinate system;
[0011] Step 5: Use the ICP algorithm to register the original radar point cloud data and the denoised near-infrared image point cloud data to obtain the corrected near-infrared image point cloud;
[0012] Step 6: Use the K-nearest neighbor algorithm to match the corrected near-infrared image with the original radar point cloud and perform fusion to enhance the density of the original radar point cloud.
[0013] Furthermore, in Step 1, preprocess the original image data captured by the near-infrared camera. The preprocessing operations include image scaling and pixel data normalization. The specific method is as follows:
[0014] Assume the original image resolution is W×H, and the size of the image data after preprocessing is 1064×616. Perform proportional scaling on the original image data, calculate the scaling factor scale, and use gray bars to fill the remaining pixels to make its resolution 1064×616. The method for calculating the scaling factor scale is as follows:
[0015]
[0016] Perform normalization on the scaled near-infrared image. According to the statistically obtained mean μ and standard deviation σ of the pixel values, for each pixel value I c (i, j), where c represents the channel, and i and j represent the row and column of the pixel respectively. The normalized pixel value I' c (i, j) is expressed as:
[0017]
[0018] Furthermore, in Step 2, use a monocular depth estimation algorithm to perform absolute depth estimation on the preprocessed near-infrared image data, and obtain the depth information of each pixel point in the preprocessed near-infrared image. The specific method is as follows:
[0019] (A) Network architecture design
[0020] The monocular depth estimation algorithm uses a Vision Transformer structure to predict absolute depth, including an encoder module and a decoder module. The near-infrared image is input into the encoder for encoding, and four different-sized features are extracted. In the decoder module, the four different-sized features are used as inputs for feature decoding, and the final depth map is output.
[0021] (c) Encoder module
[0022] The encoder module includes sub-modules of Patch Embedding, Position Embedding, Transformer Blocks, and Layer Nomalization, and the sub-modules are executed serially.
[0023] The Patch Embedding module uses a 14×14 convolution with a stride of 14 to divide the input image into several non-overlapping patches and embeds each patch into a vector.
[0024] The Position Embedding module adds position information to each patch embedding to preserve the positional relationship of the patches in the image and adds a special token [CLS] for classification.
[0025] 12 identical Transformer Block structures are connected in series. In the Transformer Block, the input passes through Layer Normalization 1, the multi-head self-attention mechanism layer, Layer Scale 1, Layer Normalization 2, the feed-forward neural network, and Layer Scale 2 in sequence. Among them, the output of the multi-head self-attention mechanism layer is connected with the output of Layer Scale 1 through a residual connection as the input of Layer Normalization 2, and the output of the feed-forward neural network is connected with the output of Layer Scale 2 through a residual connection as the output of the Transformer Block.
[0026] Layer Nomalization normalizes the output of the Transformer Blocks to obtain four different-sized features.
[0027] (d) Decoder module
[0028] The decoder module includes sub - modules such as Token2feature, DecoderFeature, Depth Regressor, NormalPredictor, ContextFeatureEncoder, and GRU Update Block. The output of the encoder passes through the Token2feature and DecoderFeature sub - modules in sequence, and then enters two branches. After passing through the Depth Regressor and Normal Predictor sub - modules respectively, they are merged into the ContextFeatureEncoder sub - module, and finally enter the GRUUpdate Block sub - module to obtain the output;
[0029] Token2feature converts the output of the encoder into a multi - scale feature map, facilitating subsequent depth and normal prediction;
[0030] DecoderFeature decodes the multi - scale features into initial depth and normal feature maps, including 3 FuseBlock sub - modules for feature fusion and up - sampling operations. The FusionBlock consists of two 3×3 convolutional layers with a stride of 1 and an up - sampling layer. Through 3 times of feature fusion and up - sampling, high - resolution depth and normal feature maps are gradually generated;
[0031] Depth Regressor is used to convert the high - resolution depth feature map into a depth probability distribution and calculate the depth expectation value, consisting of two 3×3 convolutional layers with a stride of 1 and a Relu activation function;
[0032] Normal Predictor is used to predict the normal from the high - resolution normal feature map, consisting of two 3×3 convolutional layers with a stride of 1, four 1×1 convolutional layers with a stride of 1, and two Relu activation functions;
[0033] ContextFeatureEncoder encodes the outputs of Depth Regressor and Normal Predictor as input features to generate the initial hidden state and context features required for GRU Update. ContextFeatureEncoder contains the ResidualBlock sub - module and convolutional operations, responsible for converting the input features into context features. The ResidualBlock contains two 3×3 convolutional layers with a stride of 1, two normalization layers, and two Relu activation functions. Through the residual connection operation, the input of the ResidualBlock is added to the output features processed by the ResidualBlock to generate the output;
[0034] The GRU Update takes the output of the BlockContextFeatureEncoder as input and gradually optimizes depth and normal prediction. The GRU Update Block contains ConvGRU and FlowHead sub-modules. The ConvGRU contains three 3×3 convolutional layers with a stride of 1 (convz, convr, convq) for updating the hidden state of the GRU. The FlowHead contains four 3×3 convolutional layers with a stride of 1 and an activation function for generating incremental updates of depth and normals. The generated depth information is the required depth map.
[0035] Further, in step 3, using the depth information of each pixel in the preprocessed near-infrared image, the two-dimensional preprocessed near-infrared image data is converted into three-dimensional near-infrared image point cloud data and transformed into the radar coordinate system. The specific method is as follows:
[0036] Let the extrinsic parameters from the near-infrared camera to the millimeter-wave radar be T, and the intrinsic parameters K of the near-infrared camera are as follows, where fx and fy are the focal lengths of the camera in the horizontal and vertical directions, and cx and cy are the principal point coordinates of the camera in the horizontal and vertical directions;
[0037]
[0038] Let the coordinates of any pixel in the near-infrared image in the image plane be (u, v), and its corresponding depth value be d(u, v). Converting this pixel coordinate into the three-dimensional point coordinates (x, y, z) of the near-infrared image data in the camera coordinate system, there is the following equation relationship:
[0039] d(u, v)·[u, v, 1] T =K·[x, y, z] T
[0040] Further derivation gives the calculation formula for the three-dimensional point coordinates of the near-infrared image data as:
[0041]
[0042] Let a point P in the camera coordinate system c be transformed to point P r in the radar coordinate system. The method is as follows:
[0043] [P r , 1] T =[P c , 1] T ·T
[0044] Through two coordinate system transformations, the near-infrared image data is converted from the image plane coordinate system to the radar coordinate system. The above operation is repeated for all pixel data on the entire near-infrared image to obtain three-dimensional near-infrared image point cloud data.
[0045] Further, in step 4, a pass-through filter and a ground point removal algorithm are used to denoise the three-dimensional near-infrared image point cloud data in the radar coordinate system. The specific method is as follows:
[0046] Step 4-1: Use a pass-through filter to denoise the three-dimensional near-infrared image point cloud data in the radar coordinate system, filtering out points outside the effective distance and effective height of the near-infrared camera. Let the effective distance of the near-infrared camera be (DisL, DisH), and the effective height of detection be (HeightL, HeightH). The pass-through filter method is as follows:
[0047] Let the set of three-dimensional near-infrared image point cloud data in the radar coordinate system be P = {p i | 1 ≤ i ≤ N}, where i represents the i-th point and N represents the total number of points in the image point cloud. The three-dimensional coordinates of each point can be expressed as p i = (p ix , p iy , p iz ). The image point cloud set after the pass-through filter is expressed as:
[0048] P' = {p i | 1 ≤ i ≤ N, DisL < p ix < DisH, HeightH < p iz < HeightH}
[0049] Step 4-2: Use a ground point removal algorithm to remove invalid ground points;
[0050] First, create a pair of two-dimensional grid maps Map1 and Map2, which are used to record the maximum and minimum values of the point cloud height in each grid respectively;
[0051] Let be the image point cloud set, where p i represents the i-th point, and (x i , y i , z i ) is its coordinate. The grid index idx of each point is:
[0052]
[0053] where Δx and Δy are the resolutions of the grid in the x and y directions respectively, and x min and y minis the minimum coordinate of the grid map. Project each point in the image point cloud into the corresponding grid and update the height information in the grid map;
[0054]
[0055] For each grid, calculate the difference between the corresponding elements in the grid map recording the maximum height and the grid map recording the minimum height to obtain the height difference grid map MapDiff:
[0056] MapDiff idx = Map1 idx - Map2 idx
[0057] Traverse each point and check whether the height difference of the grid where the height difference grid map is located is greater than the threshold threshold. If it is greater than the threshold, the point in the grid is considered not a ground point and is retained in the non-ground point cloud set p′. If it is less than or equal to the threshold, it is considered a ground point and is removed. The method is as follows:
[0058] p′ = {p|MapDiff idx > threshold}.
[0059] Furthermore, in step 5, the ICP algorithm is used to register the original radar point cloud data and the denoised near-infrared image point cloud data to obtain the corrected near-infrared image point cloud. The specific method is as follows:
[0060] Let the original radar point cloud set The denoised near-infrared image point cloud set where The ICP algorithm finds the transformation matrix T such that p c after transformation is as aligned as possible with p r ;
[0061] The transformation matrix T is decomposed into a rotation matrix R and a translation vector t, that is:
[0062]
[0063] The goal of the ICP algorithm is to minimize the following error metric:
[0064]
[0065] where represents the distance to the nearest radar point, w i is the distance-based weight:
[0066]
[0067] where ε is a very small integer to avoid the coincidence of two points;
[0068] After multiple iterations, a transformation matrix is obtained, that is, the corrected extrinsic matrix T′ from the camera to the radar, where the rotation matrix is R′ and the translation vector is t′. The corrected near-infrared image point cloud is obtained using the corrected extrinsic matrix. The correction method is as follows:
[0069]
[0070] Further, in step 6, the K-nearest neighbor algorithm is used to match the corrected near-infrared image with the original radar point cloud and fuse them to enhance the density of the original radar point cloud. The specific method is as follows:
[0071] Let the original radar point cloud be Use the K-nearest neighbor algorithm to find its K nearest neighbor image points in the corrected near-infrared image point cloud respectively The corresponding Euclidean distance The weight of the image point is expressed as:
[0072]
[0073] where σ is the attenuation factor, the maximum confidence of the image point is 0.5, and the weight of the radar point is w r = 1 - w c , and the fused point cloud set is expressed as The fusion formula is:
[0074]
[0075] The enhanced point cloud set P enhance is the superposition of the original radar point cloud set and the fused point cloud set, expressed as: P enhance = M ∩ P r
[0076] A point cloud enhancement system based on the fusion of near-infrared images and millimeter-wave radar point clouds implements the point cloud enhancement method based on the fusion of near-infrared images and millimeter-wave radar point clouds to achieve point cloud enhancement based on the fusion of near-infrared images and millimeter-wave radar point clouds.
[0077] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the point cloud enhancement method based on the fusion of near-infrared images and millimeter-wave radar point clouds to achieve point cloud enhancement based on the fusion of near-infrared images and millimeter-wave radar point clouds.
[0078] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the described method for enhancing point cloud based on the fusion of near-infrared image and millimeter-wave radar point cloud is implemented, achieving the enhancement of point cloud based on the fusion of near-infrared image and millimeter-wave radar point cloud.
[0079] Compared with the prior art, the significant advantages of the present invention include: (1) By fusing the high-resolution image data of the near-infrared camera, the resolution and detail performance of the millimeter-wave radar point cloud are significantly improved. The near-infrared image provides rich detail information, which can effectively complement the sparsity of the millimeter-wave radar point cloud, making the overall point cloud data more refined and accurate. (2) The combination of the near-infrared camera and the millimeter-wave radar enables the perception system to provide consistent and high-quality perception data under various environmental conditions, such as bad weather, low-light environments, and complex lighting conditions. The millimeter-wave radar performs excellently under bad weather conditions such as rain, snow, and fog, while the near-infrared camera has excellent performance at night or in low-light environments. The combination of the two can form an all-weather and all-environment perception system. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 is the algorithm framework diagram of the present invention.
[0081] Figure 2 is the absolute depth estimation encoder framework diagram.
[0082] Figure 3 is the absolute depth estimation decoder framework diagram.
[0083] Figure 4 is the effect diagram of depth estimation and point cloud density enhancement.
[0084] Figure 5 is the comparison effect diagram of model object detection before and after point cloud enhancement. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0085] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0086] As Figure 1 shown, a method for enhancing point cloud based on the fusion of near-infrared image and millimeter-wave radar point cloud according to the present invention includes the following steps:
[0087] Step 1, preprocess the original image data captured by the near-infrared camera. The preprocessing operations include image scaling and pixel data normalization operations;
[0088] Assume the original image resolution is W×H, and the size of the image data after preprocessing is 1064×616. Scale the original image data proportionally, calculate the scaling factor scale, and fill the remaining pixels with gray bars to make its resolution 1064×616. The method for calculating the scaling factor scale is as follows:
[0089]
[0090] Perform normalization on the scaled near-infrared image. According to statistics, obtain the mean μ and standard deviation σ of the pixel values. For each pixel value I c (i,j), where c represents the channel, and i and j represent the row and column of the pixel respectively. The normalized pixel value I' c (i,j) is expressed as:
[0091]
[0092] Step 2: Use a monocular depth estimation algorithm to perform absolute depth estimation on the preprocessed near-infrared image data, and obtain the depth information of each pixel point in the preprocessed near-infrared image;
[0093] (A) Network architecture design
[0094] The monocular depth estimation algorithm uses a Vision Transformer structure to predict absolute depth, including an encoder module and a decoder module. Input the near-infrared image into the encoder for encoding to extract four different-sized features; in the decoder module, use the four different-sized features as inputs for decoding to output the final depth map;
[0095] (e) Encoder module
[0096] The encoder module contains sub-modules such as Patch Embedding, Position Embedding, Transformer Blocks, and Layer Nomalization, and the sub-modules are executed serially;
[0097] The Patch Embedding module uses a 14×14 convolution with a stride of 14 to divide the input image into several non-overlapping patches and embed each patch into a vector;
[0098] The Position Embedding module adds position information to each patch embedding to preserve the position relationship of the patches in the image and adds a special token [CLS] for classification;
[0099] Twelve identical Transformer Block structures are connected in series. In the Transformer Block, the input passes through Layer Normalization 1, the multi-head self-attention mechanism layer, Layer Scale 1, Layer Normalization 2, the feed-forward neural network, and Layer Scale 2 in sequence. Among them, the output of the multi-head self-attention mechanism layer is connected with the output of Layer Scale 1 through a residual connection as the input of Layer Normalization 2, and the output of the feed-forward neural network is connected with the output of Layer Scale 2 through a residual connection as the output of the Transformer Block;
[0100] Layer Nomalization normalizes the output of the Transformer Blocks to obtain four different-sized features;
[0101] (f) Decoder module
[0102] The decoder module includes sub-modules such as Token2feature, DecoderFeature, Depth Regressor, NormalPredictor, ContextFeatureEncoder, and GRU Update Block. The output of the encoder passes through the Token2feature and DecoderFeature sub-modules in sequence, and then enters two branches, passing through the Depth Regressor and Normal Predictor sub-modules respectively and then merging into the ContextFeatureEncoder sub-module, and finally entering the GRUUpdate Block sub-module to obtain the output;
[0103] Token2feature converts the output of the encoder into a multi-scale feature map to facilitate subsequent depth and normal prediction;
[0104] DecoderFeature decodes the multi-scale features into the initial depth and normal feature maps, including 3 FuseBlock sub-modules for feature fusion and upsampling operations. The FusionBlock consists of two 3×3 convolutional layers with a stride of 1 and an upsampling layer. Through 3 times of feature fusion and upsampling, high-resolution depth and normal feature maps are gradually generated;
[0105] Depth Regressor is used to convert the high-resolution depth feature map into a depth probability distribution and calculate the depth expectation value, consisting of two 3×3 convolutional layers with a stride of 1 and a Relu activation function;
[0106] The Normal Predictor is used to predict normals from a high-resolution normal feature map. It consists of two 3×3 convolutional layers with a stride of 1, four 1×1 convolutional layers with a stride of 1, and two Relu activation functions.
[0107] The ContextFeatureEncoder encodes the outputs of the Depth Regressor and the Normal Predictor as input features to generate the initial hidden state and context features required for GRU Update. The ContextFeatureEncoder contains the ResidualBlock sub-module and convolutional operations, and is responsible for converting the input features into context features. The ResidualBlock contains two 3×3 convolutional layers with a stride of 1, two normalization layers, and two Relu activation functions. Through the residual connection operation, the input of the ResidualBlock is added to the output features processed by the ResidualBlock to generate the output.
[0108] The GRU Update takes the output of the BlockContextFeatureEncoder as input and gradually optimizes the depth and normal predictions. The GRU Update Block contains the ConvGRU and FlowHead sub-modules. The ConvGRU contains three 3×3 convolutional layers (convz, convr, convq) for updating the hidden state of the GRU. The FlowHead contains four 3×3 convolutional layers and one activation function for generating the incremental updates of the depth and normals. The generated depth information is the required depth map.
[0109] Step 3: Using the depth information of each pixel in the preprocessed near-infrared image, convert the two-dimensional preprocessed near-infrared image data into three-dimensional near-infrared image point cloud data and transform it into the radar coordinate system.
[0110] Let the extrinsic parameters from the near-infrared camera to the millimeter-wave radar be T, and the intrinsic parameters K of the near-infrared camera are as follows, where fx and fy are the focal lengths of the camera in the horizontal and vertical directions, and cx and cy are the principal point coordinates of the camera in the horizontal and vertical directions.
[0111]
[0112] Let the coordinates of any pixel in the near-infrared image in the image plane be (u, v), and its corresponding depth value be d(u, v). Convert this pixel coordinate into the three-dimensional point coordinates (x, y, z) of the near-infrared image data in the camera coordinate system, and there is the following equation relationship:
[0113] d(u, v)·[u, v, 1] T = K·[x, y, z] T
[0114] Further derivation yields the calculation formula for the three-dimensional point coordinates of the near-infrared image data as follows:
[0115]
[0116] Let a point P in the camera coordinate system c be transformed to point P in the radar coordinate system r as follows:
[0117] [P r , 1] T = [P c , 1] T ·T
[0118] Through two coordinate system transformations, the near-infrared image data is converted from the image plane coordinate system to the radar coordinate system. Repeating the above operations for all pixel data on the entire near-infrared image yields the three-dimensional near-infrared image point cloud data.
[0119] Step 4: Use the pass-through filter and ground point removal algorithm to denoise the three-dimensional near-infrared image point cloud data in the radar coordinate system;
[0120] Step 4-1: Use the pass-through filter to denoise the three-dimensional near-infrared image point cloud data in the radar coordinate system, filtering out points outside the effective distance and effective height of the near-infrared camera. Let the effective distance of the near-infrared camera be (DisL, DisH), and the detected effective height be (HeightL, HeightH). The pass-through filter method is as follows:
[0121] Let the set of three-dimensional near-infrared image point cloud data in the radar coordinate system be P = {p i | 1 ≤ i ≤ N}, where i represents the i-th point and N represents the total number of points in the image point cloud. The three-dimensional coordinates of each point can be expressed as p i = (p ix , p iy , p iz ). The image point cloud set after the pass-through filter is expressed as:
[0122] P' = {p i | 1 ≤ i ≤ N, DisL < p ix < DisH, HeightH < p iz < HeightH}
[0123] Step 4-2: Use the ground point removal algorithm to remove invalid ground points;
[0124] First, create a pair of two-dimensional grid maps Map1 and Map2, which are used to record the maximum and minimum values of the point cloud height in each grid respectively;
[0125] Let be the image point cloud set, where p i represents the i-th point, and (x i , y i , z i ) are its coordinates. The grid index idx of each point is:
[0126]
[0127] where Δx and Δy are the resolutions of the grid in the x and y directions respectively, and x min and y min are the minimum coordinates of the grid map. Project each point in the image point cloud into the corresponding grid and update the height information in the grid map;
[0128]
[0129] For each grid, calculate the difference between the corresponding elements in the grid map recording the maximum height and the grid map recording the minimum height to obtain the height difference grid map MapDiff:
[0130] MapDiff idx = Map1 idx - Map2 idx
[0131] Traverse each point and check whether the height difference in the grid where the height difference grid map is located is greater than the threshold threshold. If it is greater than the threshold, it is considered that the points in this grid are not ground points and are retained in the non-ground point cloud set p′. If it is less than or equal to the threshold, it is considered a ground point and is excluded. The method is as follows:
[0132] p′ = {p|MapDiff idx > threshold}.
[0133] Step 5, use the ICP algorithm to register the original radar point cloud data and the denoised near-infrared image point cloud data to obtain the corrected near-infrared image point cloud;
[0134] Let the original radar point cloud set The denoised near-infrared image point cloud set where The ICP algorithm finds the transformation matrix T such that p c after transformation is as aligned as possible with p r ;
[0135] The transformation matrix T is decomposed into a rotation matrix R and a translation vector t, i.e.:
[0136]
[0137] The goal of the ICP algorithm is to minimize the following error metric:
[0138]
[0139] where p r i represents the radar point closest to p c i and w i is the distance-based weight:
[0140]
[0141] where ε is a very small integer to avoid the situation of two points coinciding;
[0142] After multiple iterations, a transformation matrix is obtained, i.e., the corrected extrinsic matrix T' from the camera to the radar, where the rotation matrix is R' and the translation vector is t'. The corrected near-infrared image point cloud is obtained using the corrected extrinsic matrix The correction method is as follows:
[0143]
[0144] Step 6, use the K-nearest neighbor algorithm to match the corrected near-infrared image with the original radar point cloud and fuse them to enhance the density of the original radar point cloud;
[0145] Let the original radar point cloud be Use the K-nearest neighbor algorithm to find its K nearest neighbor image points in the corrected near-infrared image point cloud respectively The corresponding Euclidean distance The image point weight is expressed as:
[0146]
[0147] where σ is the attenuation factor, the maximum image point confidence is 0.5, and the radar point weight is w r = 1 - w c , and the fused point cloud set is denoted as The fusion formula is:
[0148]
[0149] The enhanced point cloud set P enhanceThe superposition of the original radar point cloud set and the fused point cloud set is expressed as:
[0150] P enhance = M ∩ P r
[0151] The present invention also proposes a point cloud enhancement system based on the fusion of near-infrared images and millimeter-wave radar point clouds, implements the point cloud enhancement method based on the fusion of near-infrared images and millimeter-wave radar point clouds, and realizes the point cloud enhancement based on the fusion of near-infrared images and millimeter-wave radar point clouds.
[0152] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the point cloud enhancement method based on the fusion of near-infrared images and millimeter-wave radar point clouds, and realizes the point cloud enhancement based on the fusion of near-infrared images and millimeter-wave radar point clouds.
[0153] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the point cloud enhancement method based on the fusion of near-infrared images and millimeter-wave radar point clouds, and realizes the point cloud enhancement based on the fusion of near-infrared images and millimeter-wave radar point clouds.
[0154] Embodiment
[0155] In order to verify the effectiveness of the solution of the present invention, the following experiments are carried out.
[0156] The mainstream point cloud-based object detection model PointPillar is selected. The model is trained on the original millimeter-wave radar point cloud and the point cloud with enhanced point cloud density using the present invention respectively, and its effects in day and night scenes are detected. The depth estimation effect of the near-infrared image and the enhancement effect of the point cloud density enhancement method under night conditions are as Figure 4 shown. Among them Figure 4 -a is the original image of the near-infrared image, Figure 4 -b and 4-c respectively represent the image point cloud map and the image point cloud shown, Figure 4 -e and 4-f respectively represent the original millimeter-wave radar point cloud and the enhanced point cloud after point cloud density enhancement. The performance of PointPillar on the original point cloud and the enhanced point cloud is as Figure 5 shown, where Figure 5 -a and 5-b are night scenes, while 5-c and 5-d are day scenes. It can be seen that the false detection rate and the missed detection rate of the PointPillar model trained on the enhanced point cloud are greatly reduced, showing more excellent accuracy and robustness.
[0157] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A point cloud density enhancement method based on the fusion of near infrared image and millimeter wave radar point cloud, characterized in that: The following steps are involved: Step 1, preprocessing the raw image data taken by the near-infrared camera, the preprocessing operation includes image scaling and pixel data normalization operations; Step 2: Use a monocular depth estimation algorithm to perform absolute depth estimation on the preprocessed near-infrared image data to obtain the depth information of each pixel of the preprocessed near-infrared image; Step 3, using the depth information of each pixel of the preprocessed near-infrared image, convert the two-dimensional preprocessed near-infrared image data into three-dimensional near-infrared image point cloud data, and convert it to the radar coordinate system; Step 4, using a straight-through filter and ground point removal algorithm to denoise the three-dimensional near-infrared image point cloud data in the radar coordinate system; Step 5, using the ICP algorithm to register the radar original point cloud data and the denoised near-infrared image point cloud data to obtain a corrected near-infrared image point cloud; Step 6: Use the K nearest neighbor algorithm to match the corrected near-infrared image and fuse it with the original radar point cloud to enhance the density of the original radar point cloud.
2. The point cloud density enhancement method based on the fusion of near infrared image and millimeter wave radar point cloud according to claim 1 is characterized in that: Step 1: preprocess the raw image data taken by the near-infrared camera. The preprocessing operation includes image scaling and pixel data normalization. The specific method is as follows: Assume that the original image resolution is W×H, and the size of the image data after preprocessing is 1064×616. Scale the original image data proportionally, calculate the scaling factor scale, and use gray bars to fill the remaining pixels to make its resolution 1064×616. The method for calculating the scaling factor scale is as follows: The scaled near-infrared image is normalized, and the mean μ and standard deviation σ of the pixel values are obtained according to statistics. For each pixel value I c (i, j), where c represents the channel, i and j represent the row and column of the pixel respectively, and the normalized pixel value I' c (i,j) is expressed as:
3. The point cloud density enhancement method based on the fusion of near infrared image and millimeter wave radar point cloud according to claim 1 is characterized in that: Step 2: Use a monocular depth estimation algorithm to perform absolute depth estimation on the preprocessed near-infrared image data to obtain the depth information of each pixel of the preprocessed near-infrared image. The specific method is as follows: (A) Network architecture design The monocular depth estimation algorithm uses the Vision Transformer structure to predict the absolute depth. It includes an encoder module and a decoder module. The near infrared image is input into the encoder for encoding and four features of different sizes are extracted. In the decoder module, features of four different sizes are taken as input, feature decoding is performed, and the final depth map is output; (a) Encoder module The encoder module includes Patch Embedding, Position Embedding, Transformer Blocks, and LayerNomalization sub-modules, and the sub-modules are executed serially; The Patch Embedding module uses a 14×14 convolution with a stride of 14 to divide the input image into several non-overlapping patches and embed each patch into a vector; The Position Embedding module adds position information to each patch embedding to preserve the positional relationship of the patches in the image and adds a special marker symbol [CLS] for classification; Twelve identical Transformer Block structures are connected in series. In the Transformer Block, the input passes through layer normalization 1, multi-head self-attention mechanism layer, Layer Scale 1, layer normalization 2, feedforward neural network, and LayerScale 2 in sequence. The output of the multi-head self-attention mechanism layer is residually connected with the output of Layer Scale 1 as the input of layer normalization 2, and the output of the feedforward neural network is residually connected with the output of Layer Scale 2 as the output of the Transformer Block. Layer Nomalization normalizes the output of Transformer Blocks to obtain features of four different sizes; (b) Decoder module The decoder module includes Token2feature, DecoderFeature, Depth Regressor and NormalPredictor, ContextFeatureEncoder, and GRU Update Block submodules. The output of the encoder passes through the Token2feature and DecoderFeature submodules in turn, and then enters two branches, passes through the Depth Regressor and Normal Predictor submodules respectively, and then merges into the ContextFeatureEncoder submodule, and finally enters the GRUUpdate Block submodule to get the output; Token2feature converts the encoder output into a multi-scale feature map to facilitate subsequent depth and normal prediction; DecoderFeature decodes multi-scale features into initial depth and normal feature maps, including three FuseBlock submodules for feature fusion and upsampling operations. FusionBlock consists of two 3×3 convolutional layers with a step size of 1 and an upsampling layer. Through three times of feature fusion and upsampling, high-resolution depth and normal feature maps are gradually generated; Depth Regressor is used to convert high-resolution depth feature maps into depth probability distribution and calculate the depth expectation value. It consists of two 3×3 convolutional layers with a step size of 1 and a Relu activation function; Normal Predictor is used to predict normals from high-resolution normal feature maps. It consists of two 3×3 convolutional layers with a stride of 1, four 1×1 convolutional layers with a stride of 1, and two ReLU activation functions. ContextFeatureEncoder encodes the output of Depth Regressor and Normal Predictor as input features to generate the initial hidden state and context features required by GRU Update. ContextFeatureEncoder contains the ResidualBlock submodule and convolution operation, which is responsible for converting the input features into context features. ResidualBlock contains two 3×3 convolution layers with a step size of 1, two normalization layers and two Relu activation functions. Through the residual connection operation, the input of ResidualBlock is added to the output features processed by ResidualBlock to generate the output; GRU Update takes the output of BlockContextFeatureEncoder as input and gradually optimizes the depth and normal predictions. GRU Update Block contains ConvGRU and FlowHead submodules. ConvGRU contains three 3×3 convolutional layers (convz, convr, convq) with a step size of 1, which are used to update the hidden state of GRU. FlowHead contains four 3×3 convolutional layers with a step size of 1 and an activation function to generate incremental updates of depth and normals. The generated depth information is the required depth map.
4. The point cloud density enhancement method based on the fusion of near infrared image and millimeter wave radar point cloud according to claim 1 is characterized in that: Step 3, using the depth information of each pixel of the preprocessed near-infrared image, convert the two-dimensional preprocessed near-infrared image data into three-dimensional near-infrared image point cloud data, and convert it to the radar coordinate system. The specific method is: Assume that the external parameter from the near-infrared camera to the millimeter-wave radar is T, and the internal parameter K of the near-infrared camera is as follows, where fx, fy are the focal lengths of the camera in the horizontal and vertical directions, and cx, cy are the principal point coordinates of the camera in the horizontal and vertical directions; Assume that the coordinates of any pixel in the near-infrared image on the image plane are (u, v), and its corresponding depth value is d(u, v). Convert this pixel coordinate to the three-dimensional point coordinates (x, y, z) of the near-infrared image data in the camera coordinate system, and there is the following equation: d(u,v)·[u,v,1] T =K·[x,y,z] T The three-dimensional point coordinate calculation formula of near-infrared image data is further deduced as follows: Let the camera coordinate system be a point P c Convert to radar coordinate system point P r Next, the method is as follows: [P r ,1] T =[P c ,1] T ·T Through two coordinate system transformations, the near-infrared image data is converted from the image plane coordinate system to the radar coordinate system. The above operation is repeated for all pixel data on the entire near-infrared image to obtain the three-dimensional near-infrared image point cloud data.
5. The point cloud density enhancement method based on the fusion of near infrared image and millimeter wave radar point cloud according to claim 1 is characterized in that: Step 4: Use the straight-through filtering and ground point removal algorithm to denoise the 3D near-infrared image point cloud data in the radar coordinate system. The specific method is as follows: Step 4-1: Use straight-through filtering to denoise the 3D near-infrared image point cloud data in the radar coordinate system, and filter out all points outside the effective distance and effective height of the near-infrared camera. Suppose the effective distance of the near-infrared camera is (DisL, DisH), and the effective height of detection is (HeightL, HeightH). The straight-through filtering method is as follows: Assume that the three-dimensional near-infrared image point cloud data set P in the radar coordinate system is i |1≤i≤N}, where i represents the i-th point, N represents the total number of points in the image point cloud, and the three-dimensional coordinates of each point can be expressed as p i =(p ix ,p iy ,p iz ), the image point cloud set after direct filtering is expressed as: P'={p i |1≤i≤N,DisL<p ix <DisH,HeightH<p iz <HeightH} Step 4-2, use the ground point elimination algorithm to eliminate invalid ground points; First, create a pair of two-dimensional grid maps Map1 and Map2, which are used to record the maximum and minimum values of the point cloud height in each grid; set up is a set of image point clouds, where p i represents the i-th point, (x i ,y i ,z i ) is its coordinate, and the grid index idx of each point is: Among them, Δx and Δy are the resolution of the grid in the x and y directions respectively. min and min is the minimum coordinate of the grid map. Project each point in the image point cloud to the corresponding grid and update the height information in the grid map; For each grid, calculate the difference between the corresponding elements in the grid map that records the maximum height and the grid map that records the minimum height, and get the height difference grid map MapDiff: MapDiff idx =Map1 idx -Map2 idx Traverse each point and check whether the height difference of the grid where its height difference grid map is located is greater than the threshold threshold. If it is greater than the threshold, the point in the grid is considered not to be a ground point and is retained in the non-ground point cloud set p'. If it is less than or equal to the threshold, it is considered to be a ground point and is removed. The method is as follows: p′={p|MapDiff idx >threshold}。 6. The point cloud density enhancement method based on the fusion of near infrared image and millimeter wave radar point cloud according to claim 1 is characterized in that: Step 5: Use the ICP algorithm to register the radar original point cloud data and the denoised near-infrared image point cloud data to obtain the corrected near-infrared image point cloud. The specific method is as follows: Assume the radar original point cloud set Denoised near infrared image point cloud collection in The ICP algorithm seeks a transformation matrix T such that p c After transformation and p r Align as much as possible; The transformation matrix T is decomposed into a rotation matrix R and a translation vector t, namely: The goal of the ICP algorithm is to minimize the following error metric: in Indicates distance The nearest radar point, w i is the distance-based weight: Among them, ε is a very small integer to avoid the situation where two points coincide with each other; After multiple iterations, a transformation matrix is obtained, that is, the modified camera to radar external parameter matrix T′, in which the rotation matrix is R′ and the translation vector is t′. The modified near-infrared image point cloud is obtained using the modified external parameter matrix The correction method is as follows:
7. The point cloud enhancement method based on the fusion of near infrared image and millimeter wave radar point cloud according to claim 1 is characterized in that: Step 6: Use the K nearest neighbor algorithm to match the corrected near infrared image and the original radar point cloud, and fuse them to enhance the density of the original radar point cloud. The specific method is as follows: Assume the original radar point cloud is Use the K nearest neighbor algorithm to find the corrected near infrared image point cloud. The K nearest neighbor image points in Its corresponding Euclidean distance The image point weight is expressed as: Where σ is the attenuation factor, the maximum confidence of the image point is 0.5, and the radar point weight is w r =1-w c , the fused point cloud set is represented as The fusion formula is: Enhanced point cloud set P enhance It is the superposition of the original radar point cloud set and the fused point cloud set, expressed as: P enhance =M∩P r .
8. A point cloud enhancement system based on the fusion of near infrared image and millimeter wave radar point cloud, characterized in that: Implement the point cloud enhancement method based on the fusion of near-infrared image and millimeter-wave radar point cloud as described in any one of claims 1-7 to achieve point cloud enhancement based on the fusion of near-infrared image and millimeter-wave radar point cloud.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the point cloud enhancement method based on the fusion of near-infrared image and millimeter-wave radar point cloud according to any one of claims 1 to 7 is implemented to achieve point cloud enhancement based on the fusion of near-infrared image and millimeter-wave radar point cloud.
10. A computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the point cloud enhancement method based on the fusion of near-infrared image and millimeter-wave radar point cloud according to any one of claims 1 to 7 is implemented to achieve point cloud enhancement based on the fusion of near-infrared image and millimeter-wave radar point cloud.
Citation Information
Patent Citations
Millimeter wave radar and vision fused three-dimensional target detection method based on attention mechanism
CN114708585A
Target tracking method based on visible light, infrared and laser radar data fusion
CN116258744A
Depth completion method based on millimeter wave radar and camera fusion
CN117808689A
Multiband image fusion coloring method suitable for underground space point cloud map
CN118334153A
Laser radar point cloud filtering and enhancing method based on image priori knowledge under severe weather condition
CN118446932A
Cited By
Radar point cloud and vision fusion target detection method and system based on frustum mapping
CN122695394A