Depth completion method based on infrared feature mining and cross-modal fusion

By adopting the depth completion method based on infrared feature mining and cross-modal fusion in the autonomous driving system, the problem of poor completion of sparse 3D point cloud depth maps is solved, and high-precision point cloud completion in highlight and dark scenes is achieved, which enhances the robustness of the system.

CN119205583BActive Publication Date: 2025-05-09NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411724123.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-05-09
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

In the existing autonomous driving systems, the depth map completion effect of sparse 3D point clouds is not good, especially in highlight and dark scenes, and the imaging quality of visible images is greatly affected by changes in lighting conditions and lacks robustness.

Method used

The depth completion method based on infrared feature mining and cross-modal fusion is adopted, and the depth map is completed through infrared adaptive edge enhancement and feature extraction module and cross-modal feature stereoscopic diffusion module, combined with infrared images and sparse depth maps.

Benefits of technology

It improves the accuracy of object edge completion in sparse depth images, and can provide high-precision point cloud completion under all-weather conditions, especially in low-illumination environments, and enhances the robustness of depth completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205583B_ABST
    Figure CN119205583B_ABST
Patent Text Reader

Abstract

The present invention discloses a depth completion method based on infrared feature mining and cross-modal fusion, comprising the following steps: collecting infrared images and sparse depth maps, and adding infrared adaptive edge enhancement and feature extraction modules and cross-modal feature stereo diffusion modules on the basis of the existing KBnet network. The depth completion method based on infrared feature mining and cross-modal fusion of the present invention proposes an infrared adaptive edge enhancement and feature extraction module in view of the low contrast and low texture characteristics of infrared images, which provides more features for the completion network and contributes to the accuracy of object edge completion in sparse depth images. In view of the problem of guiding high-precision point cloud completion with low-resolution infrared images, a cross-modal feature stereo fusion module is proposed, which enhances the cross-modal feature fusion performance and can smoothly fill in the missing depth area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a depth completion method based on infrared feature mining and cross-modal fusion, and belongs to the technical field of depth completion. Background Art

[0002] In autonomous driving systems, the acquisition of scene depth information is crucial. Currently, most systems use lidar to project laser beams into the scene and collect echoes to obtain 3D point clouds. However, even the 3D point clouds collected by high-end lidars are still highly sparse, and point clouds cannot be obtained at all in areas of reflection, projection, and beyond the working distance range. Therefore, completing the depth map obtained by projecting the sparse 3D point cloud in two-dimensional space is an important part of the work of the autonomous driving system.

[0003] Since sparse point clouds lack too much scene information, it is necessary to use information from other modalities to guide depth completion. The commonly used method is to use visible light images to guide depth completion. However, due to the limitations of imaging principles, visible light cameras cannot take into account both highlight and dark scene information. Infrared cameras capture thermal radiation emitted by objects rather than visible light for imaging. Therefore, they can better preserve the structural information of highlights and dark areas in scenes with strong lighting contrast, such as meeting vehicles without a barrier at night, driving against the light at dawn or sunset, and entering and exiting tunnels.

[0004] In addition, the image quality of visible light cameras is greatly affected by changes in lighting conditions. The structural information of visible light images is submerged in a large amount of noise, which may affect the depth completion effect guided by visible light and lack good robustness. Infrared cameras have good all-weather performance and can guarantee image quality even in low illumination, providing sufficient scene information.

[0005] Therefore, a depth completion method based on infrared feature mining and cross-modal fusion is needed to solve the above problems. Summary of the invention

[0006] Purpose of the invention: In view of the problems existing in the prior art, the present invention provides a depth completion method based on infrared feature mining and cross-modal fusion.

[0007] A depth completion method based on infrared feature mining and cross-modal fusion includes the following steps:

[0008] Step 1: Use an infrared camera to collect infrared images, and use a laser radar to collect sparse depth maps;

[0009] Step 2: construct a depth completion network, which includes an infrared adaptive edge enhancement and feature extraction module, an infrared branch encoder, a depth branch encoder, a decoder, and a cross-modal feature stereo diffusion module.

[0010] The infrared adaptive edge enhancement and feature extraction module includes an edge enhancement module and a feature extraction module. The edge enhancement module includes a noise reduction module, a Roberts operator and a Canny operator. The noise reduction module is used to filter the noise in the infrared image to obtain a noise-reduced infrared image. The Roberts operator and the Canny operator respectively detect the image edge of the noise-reduced infrared image to obtain a Roberts image edge and a Canny image edge.

[0011] The infrared image, the edge of the Roberts image and the edge of the Canny image are connected and input into the feature extraction module to obtain an edge-enhanced infrared image, wherein the feature extraction module is composed of a convolution layer and four residual blocks;

[0012] The sparse depth map is input into the depth branch encoder to obtain a depth map code;

[0013] The infrared adaptive edge enhancement and feature extraction module is used to process the infrared image to obtain the deep features of the infrared image, and the deep features of the infrared image are input into the infrared branch encoder to obtain the infrared image code;

[0014] The depth map code and the infrared image code are connected and input into the decoder to obtain a pre-completed depth map;

[0015] The pre-completed depth map and the deep features of the infrared image together form a stereo convolution layer that is input into the cross-modal feature stereo diffusion module;

[0016] The cross-modal feature stereo diffusion module is used to fuse the pre-completed depth map and the deep features of the infrared image to obtain a dense depth map;

[0017] Step 3: Output dense depth map.

[0018] Furthermore, the noise reduction module in step 2 is a Gaussian filter.

[0019] Furthermore, the edge enhancement module in step 2 further includes a Sobel operator, and the Sobel operator is used to determine the high and low thresholds in the Canny operator.

[0020] Furthermore, the Sobel operator is used to determine the high and low thresholds in the Canny operator, including the following steps:

[0021] Step 21: Calculate the horizontal and vertical gradients at pixel i using the Sobel operator. and For each pixel, the gradient magnitude is calculated according to the following formula based on its gradient components in the x and y directions:

[0022] ,

[0023] In the formula, and is the gradient in the horizontal and vertical directions at pixel i;

[0024] Step 22: After the gradient amplitude of all pixels is obtained, the high threshold and the low threshold are calculated according to the following formula:

[0025] ,

[0026] Where high and low are the high threshold and low threshold, respectively, m and n are the number of pixels in the x direction and y direction, respectively. is a constant.

[0027] Furthermore, the horizontal and vertical operators dx and dy of the Roberts operator in step 2 are expressed by the following formulas:

[0028] .

[0029] Furthermore, in step 2, the Roberts operator detects the image edge of the denoised infrared image, specifically: the horizontal and vertical operators of the Roberts operator are used as two convolution kernels, the denoised infrared image is convolved, and the image edge is obtained after the average value is calculated.

[0030] Furthermore, in step 2, the detection of the image edge of the de-noised infrared image by the Roberts operator further includes a filtering step: filtering the obtained image edge by using a Butterworth low-pass filter.

[0031] Furthermore, the output of the residual block in step 2 is consistent with the input size.

[0032] Furthermore, the cross-modal feature stereo diffusion module in step 2 includes a feature diffusion module, and the feature diffusion module is a convolutional spatial propagation network CSPN.

[0033] Furthermore, the stereo convolution layer includes deep features of infrared images, depth map Z, image X and image Y connected in sequence. The depth map Z is obtained by projecting the pre-completed depth map to the camera perspective into a two-dimensional grayscale image. The grayscale value is used to represent the distance of a point in the three-dimensional space corresponding to the pixel from the camera. The length and width of the deep features of the infrared image, the depth map Z, image X and image Y are the same. Image X and image Y correspond to the depth map Z and are expressed by the following formula:

[0034] ,

[0035] In the formula, , , and is the camera’s internal parameter, To pre-complete the coordinates of the points in the depth map, Yes The corresponding coordinates on the camera's pixel plane.

[0036] Beneficial effects: The depth completion method based on infrared feature mining and cross-modal fusion of the present invention proposes an infrared adaptive edge enhancement and feature extraction module for the low contrast and low texture characteristics of infrared images, which provides more features for the completion network and helps to improve the accuracy of object edge completion in sparse depth images. In view of the problem of guiding high-precision point cloud completion with low-resolution infrared images, a cross-modal feature stereo fusion module is proposed to enhance the cross-modal feature fusion performance and smoothly fill in the missing depth area. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Schematic diagram of the structure of the deep completion network of the present invention;

[0038] Figure 2 The effect diagram of edge detection of infrared images by different operators;

[0039] Figure 3 A schematic diagram of projecting a 3D point cloud to a camera's perspective;

[0040] Figure 4 It is a schematic diagram of the structure of the three-dimensional convolution layer;

[0041] Figure 5 It is a schematic diagram of the structure of a multi-modal camera system;

[0042] Figure 6 Infrared, visible light and depth images are fused in pairs;

[0043] Figure 7 This is a qualitative comparison chart of different modules in the 771 dataset;

[0044] Figure 8 This is a qualitative comparison chart of other different algorithms under 771 datasets. DETAILED DESCRIPTION

[0045] The preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings to more clearly and completely illustrate the technical solutions of the present invention.

[0046] See also Figure 1As shown, the present invention proposes an infrared adaptive edge enhancement and feature extraction module (IR AEE) and a cross-modal feature stereo diffusion module (3D CMFD). The former enables the network to deeply mine the edge features of objects in infrared images by enhancing the blurred edges of objects in infrared images, and guides the accurate diffusion of edge pixels in depth images. The latter restores and strengthens the spatial position information of points from 2D depth images, and fuses them with infrared images, enhancing the cross-modal feature fusion performance, so that the network can smoothly fill in the missing depth area while retaining the original accurate depth value.

[0047] 1. Infrared Image Guided Depth Completion

[0048] Inspired by the KBnet network in the prior art, the present invention constructs it into a new depth completion module. Figure 1 As shown in the figure, the entire network includes a depth completion module (KBnet network), an infrared adaptive edge enhancement and feature extraction module, and a cross-modal feature stereo diffusion module. Among them, the KBnet network uses the unsupervised depth completion of the calibrated back-projection layer, and uses the calibration matrix and deep feature descriptor to back-project each pixel in the two-dimensional image into the three-dimensional space, achieving good results without supervision.

[0049] The infrared image is input into the infrared adaptive edge enhancement and feature extraction module, and the deep features of the infrared image are extracted and entered into the convolution layer of the infrared branch. The sparse depth map is input into the convolution layer of the depth branch. Through the external parameter matrix between the two cameras, each pixel in the infrared image is reversely projected into the three-dimensional space to generate a 3D position code. The output of each layer is used as the input of the next layer. After being decoded by the decoder, the pre-completed depth map is output. The pre-completed depth map and the infrared image together form a stereo convolution layer, which is input into the cross-modal feature stereo diffusion module to better integrate the common features between the two, and finally output an accurate dense depth map.

[0050] 1. Infrared adaptive edge enhancement and feature extraction module (IR AEE)

[0051] In the infrared guided depth completion task, the present invention hopes to provide better suppression for the edge noise of objects in the depth image by enhancing the edge information of objects in the infrared image. The edge density of infrared images in different scenes will be slightly different, and excessive edge enhancement will be noise to the network. Therefore, the present invention designs an infrared adaptive edge enhancement and feature extraction module, which can adaptively determine the edge density that needs to be enhanced and fully mine the information of infrared images through the network to meet the needs of various scenes.

[0052] Common edge detection operators include Sobel operator, Prewitt operator, Laplacian operator, Roberts operator and Canny operator. Among them, Prewitt operator has poor noise resistance, and the original infrared image contains a lot of thermal noise, so the edge detection effect of Prewitt operator on infrared image is not ideal. Laplacian operator is more sensitive to image changes and is suitable for detecting more subtle edges. Considering that infrared images have fewer detailed features, Laplacian operator is not suitable for the task of the present invention. Figure 2 It can also be proved that these two operators have poor edge detection effects on infrared images. The Roberts operator has a fast calculation speed and detects image edges from the diagonal direction. The Canny operator has a good anti-noise effect and has better edge refinement ability. Therefore, the present invention uses the Roberts operator and the Canny operator to jointly detect image edges, and uses the Sobel operator to adaptively determine the high and low thresholds of the Canny operator.

[0053] Specifically, the present invention first uses a Gaussian filter to filter out the noise in part of the infrared image, and then sends the image to the Roberts and Canny operator detection channels respectively. For an image of size m*n, where the coordinates of any pixel i are (x, y), the high and low thresholds in the Canny operator are adaptively determined with the help of the Sobel operator. The gradients in the horizontal and vertical directions at the pixel i are calculated by the Sobel operator. and For each pixel, the gradient magnitude is calculated using the Euclidean distance formula according to its gradient components in the x and y directions:

[0054] ,

[0055] After obtaining the gradient amplitude of all pixels, take the average value and the coefficient It can be fine-tuned according to different batches of data, and the low threshold is usually set to half of the high threshold to connect edge pixels to form a complete edge.

[0056] ,

[0057] Roberts operator horizontal and vertical operators and As shown below, it is not difficult to find from its form that the Roberts operator has a better effect on edge detection of positive and negative 45 degrees:

[0058] .

[0059] The horizontal and vertical operators are used as the convolution kernels of two convolutions, and the image is convolved. The edge of the image is obtained by taking the average value. The Roberts operator is less resistant to noise than the Canny operator, and some high-frequency edge noise is prone to appear after detection. Therefore, a Butterworth low-pass filter is used to filter it. A filter H with the same size as the image is obtained according to the following formula:

[0060] ,

[0061] Where D0 is the cutoff frequency and D(u,v) is the distance between the point and the center in the frequency domain, which is defined as follows:

[0062] ,

[0063] If the detected edge is directly merged with the original image, an abrupt highlight edge may be generated, which affects the feature extraction network from learning the correct edge information. Therefore, the present invention dilutes the edge and merges it with the original image, so that the grayscale value near the edge transitions smoothly, and strives to enhance the edge without making it abrupt. Finally, the low-contrast infrared image is grayscale stretched to make full use of the data bit width and obtain a larger dynamic range.

[0064] The infrared image after adaptive edge enhancement is sent to the feature extraction network. The feature extraction network is mainly composed of a convolution layer and four basic residual blocks. The output of each residual block is consistent with the input size, that is, the infrared image is not downsampled, so that its features can be learned globally and matched with the depth map. Since the infrared image only has grayscale information, the residual blocks are all single-channel. While ensuring the feature mining effect, the number of residual blocks is appropriately reduced, which reduces the pressure on the feature extraction network. After adaptive edge enhancement and feature extraction, the infrared image enters the infrared branch convolution layer, that is, the infrared branch coding layer. The feature extraction module can well mine infrared features, so that the depth image can better maintain the accuracy of depth pixels at the edges of the object boundary.

[0065] 2. Cross-modal feature stereo diffusion module (3D CMFD)

[0066] In infrared images, the internal temperature of the same object is relatively close, which is reflected in the grayscale of the image; there are temperature differences between different objects, which are reflected in the grayscale values. In depth images, the depth values ​​are generally relatively close inside the object, but there are differences between different objects. Therefore, there are certain similarities between the two modalities. In theory, the fusion of infrared modal information can play a positive role in the diffusion of depth values. However, the resolution of infrared images is low, while the depth value accuracy is high, so it is not a simple matter to fuse the two.

[0067] See also Figure 3 As shown, in order to make it easier for computers to understand and learn three-dimensional space information, and to convert 3D point clouds and 2D infrared images to spaces of the same dimensionality, the present invention generally projects the 3D point cloud to the camera's viewing angle to become a 2D grayscale image, and uses the grayscale value to represent the distance of a point in the three-dimensional space corresponding to the pixel from the camera. However, this process weakens the x-axis and y-axis coordinate information of a point in the three-dimensional space. Therefore, the present invention proposes a stereo convolutional layer to strengthen the x-axis and y-axis information, and encodes the infrared image and the depth image at the same time. Figure 3 As shown, P (X, Y, Z) is a point in three-dimensional space, that is, a point in the 3D point cloud collected by the sensor, and P' (u, v) is the corresponding point on the pixel plane imaged by the camera system, that is, a pixel in the depth image. There is a transformation from the camera coordinate system to the pixel coordinate system:

[0068] ,

[0069] Where (u, v) is the coordinate of the pixel, , , and is the intrinsic parameter of the camera. A simple transformation of it yields:

[0070] ,

[0071] See also Figure 4 As shown in the figure, through the depth map Z, the corresponding X and Y are obtained and added to the convolution channel together with the infrared image, so that the information of the point in the three-dimensional space and the infrared image is represented in one convolution layer. The infrared and pre-completed depth maps are better encoded to make the feature fusion more efficient and effective.

[0072] Excluding the interference of noise, the depth information directly obtained by depth sensors such as LiDAR is sparse but accurate. When completing the depth map, it is necessary to ensure that these accurate depth values ​​will not be changed at will. In addition, filling similar depth areas should be soft and smooth to avoid the jump of depth values ​​and the re-introduction of noise. The present invention uses the convolutional spatial propagation network CSPN as the basic feature diffusion module. The convolutional spatial propagation network is a simple and efficient linear propagation model that learns the affinity between adjacent pixels through cyclic convolution operations. Specifically, the convolutional spatial propagation network CSPN has proved its effectiveness in various depth estimation tasks.

[0073] In order to reduce the computational pressure of the network, the present invention accelerates CSPN and converts the translation calculation when calculating the affinity map into a convolution operation, so that the originally sequentially executed process can be calculated in parallel, greatly improving the efficiency. Specifically, if D0 is used to represent the pre-completed depth map, the image will be obtained after CSPN iterates n times to obtain a dense depth map D n , for pixel i, the information of its domain N(i) will diffuse to it during the iteration, written as:

[0074] ,

[0075] Among them, a ji is the affinity between pixel i and pixel j. If the domain size is m*m, then i has m 2 m adjacent pixels need to be processed 2 The convolution operation reduces the number of calculations to 1. x , x) represents a translation operator, which moves an affinity graph Ax along the -x direction:

[0076] ,

[0077] This converts sequential calculations into parallel calculations, making full use of the GPU's computing power for acceleration. Subsequent experiments have shown that the infrared feature diffusion module can make full use of the feature-assisted network of infrared images to fill in the missing depth area while retaining the original accurate depth value, and the network can run on a single GPU with 24G video memory.

[0078] 2. Experiment

[0079] 1. Multimodal camera system construction

[0080] In the current field of depth completion, there are relatively abundant paired datasets of visible light images and point clouds, such as the KITTI dataset mainly for road scenes, the NVU V2 dataset and Matterport3D dataset for indoor scenes, etc., while there are relatively few paired datasets of infrared images and point clouds.

[0081] See also Figure 5As shown, the present invention has produced a set of multi-modal co-optical axis camera systems to collect the required data. The system consists of a visible light camera, an infrared camera and a laser radar. Among them, the infrared camera and the visible light camera are placed at 90 degrees, and there is a dichroic mirror between the two cameras. The dichroic mirror and the optical axes of the two cameras are placed at 45 degrees. The dichroic mirror completely transmits the light in the visible light band and almost completely reflects the light in the infrared band. Therefore, the visible light camera and the infrared camera are co-optical axis, and the viewing angles of the two cameras are almost exactly the same after alignment. The laser radar is installed directly above the visible light camera, and the optical axis is parallel to the visible light camera, and the distance does not exceed 50mm. Therefore, the difference in viewing angle is extremely small, so the system can collect paired three-modal images. At the same time, the whole system is small and compact, and it is easy to be placed on various platforms.

[0082] To calibrate an infrared camera, you need to heat the calibration plate to create a temperature difference between it and the background so that it can be clearly identified in the infrared image. LiDAR works at a longer distance and requires a larger calibration plate to be clearly identified in the point cloud. Therefore, if you calibrate the infrared camera and LiDAR directly, you need to heat a larger calibration plate each time, which is obviously unrealistic.

[0083] In order to solve the cross-modal calibration problem, the present invention introduces a visible light camera. Since the visible light camera and the infrared camera share the same optical axis and the working distance is much closer than the laser radar, a small calibration plate can be used for rapid calibration. At the same time, the spectral characteristics of the visible light camera are close to those of the human eye, and it is also easier to calibrate with the laser radar. Using a visible light camera as a bridge, infrared images and point clouds can be easily registered. In addition, the data collected by the system can also make a fair comparison of the depth completion effects guided by visible light and infrared images respectively.

[0084] Figure 6 Pairwise fused infrared, visible light, and depth images collected for the system.

[0085] 2. Experimental details

[0086] Dataset: The 771 dataset was collected and produced using the above-mentioned multimodal co-optical axis camera system. This dataset is a small dataset of complex lighting road scenes, including scenes such as night, backlighting of vehicles, and high contrast between dawn and dusk. The dataset contains a total of 771 paired infrared and depth images of different road conditions with a resolution of 1477*758. Due to the use of the livoxavia rotating mirror lidar, according to its working characteristics, within a certain range, the longer the acquisition time at a fixed position, the higher the density of the point cloud. This feature is used to produce the GT of the depth image. The sparsity of the GT is about 25%, and the sparsity of the original depth map is about 3%, both of which are close to the KITTI dataset. The present invention uses 700 images for training and 71 images for testing.

[0087] Evaluation method: In the evaluation, this paper only considers the valid pixels in GT. As for the evaluation indicators, the standard indicators used in previous depth completion work are used, including the mean absolute error (MAE) and root mean square error (RMSE) in millimeters. RMSE is selected as the measurement indicator to judge the quality of the current model, and MAE is for reference only.

[0088] Experimental parameter setting: The network of the present invention uses the original infrared and sparse depth images as input with a resolution of 1400*750. The learning rates of the six stages are set to 5e-5, 1e-4, 15e-5, 1e-4, 5e-5, 2e-5 respectively. The batchsize is set to 1. The present invention trains the model on an RTX 3090.

[0089] 3. Ablation experiment

[0090] Since the dataset of the present invention has not been used in other works, in order to verify whether infrared images have greater advantages than visible light images in the depth completion task under the same scenario, the present invention uses a multimodal co-optical axis system to collect about 10,000 paired trimodal road datasets, of which day and night scenes each account for about 50%. The dataset was tested using the most basic completion module, and the results are as follows.

[0091] Table 1. Comparison of the effects of KBnet network on RGB images and infrared images

[0092]

[0093] The value before the ground truth refers to the ratio of the number of valid pixels in the input sparse depth map to the ground truth, which is used to simulate the point clouds collected by different types of lidars. It is not difficult to find that infrared depth completion has certain advantages over visible light depth completion at different sparsities. This is determined by the characteristics of the infrared modality itself, proving that the invention is reasonable and effective in introducing infrared images into the depth completion task.

[0094] In order to better evaluate the effect of each module, the present invention designs ablation experiments to verify the effects of the infrared adaptive edge enhancement and feature extraction module and the cross-modal feature stereo diffusion module. To speed up the experiment, the present invention uses a smaller 771 data set, which is different from the above experiment, but still follows the control variable method. The experimental results are summarized in Table 2.

[0095] Table 2. Comparison of the effects of different modules

[0096]

[0097] Figure 7The figures are qualitative comparisons of different module combinations under the 771 dataset, where (a) is the actual image; (b) is the infrared image; (c) is the KBnet network processing diagram; (d) is the KBnet+feature extraction processing diagram; (e) is the KBnet+feature diffusion processing diagram; (f) is KBnet+feature extraction+feature diffusion (the present invention); the most significantly improved areas are highlighted with boxes.

[0098] Infrared adaptive edge enhancement and feature extraction module: This paper adds an adaptive edge enhancement and feature extraction module based on KBnet, and uses simple additive fusion on cross-modal feature fusion to obtain the output depth map. According to the first two rows of Table 2, compared with the results of KBnet direct completion, the addition of the infrared feature extraction module has an improvement of 1.4% in RMSE. In actual effect, the completed depth map significantly adds more detailed features from the infrared image, and the attached Figure 7 In the first row of images, the edge of the signboard is restored more clearly, but the edge position of the pillar in the third row with the background in the baseline is not well restored, and the method of the present invention can clearly complete it. This is due to the addition of the adaptive edge enhancement and feature extraction module, which increases the information of the edge of the object in the infrared mode.

[0099] Cross-modal feature stereo fusion module: This paper performs simple convolution extraction on the infrared image, feeds the high-dimensional features into the cross-modal feature stereo fusion module, and obtains the output depth map. According to the first and third rows of Table 2, compared with the results of KBnet, the addition of the cross-modal feature stereo fusion module has an improvement of 1.1% in RMSE. Figure 7 For the car in the 5th row, due to the presence of the car window, the original depth map has a block area with missing depth. The baseline completion effect does not smoothly complete this part, resulting in more noise points. However, the method of the present invention can accurately and smoothly complete the car window to the same depth as the car body, with less noise points. For the sign in the 6th row, it can be found from GT that there is a clear depth boundary line. After baseline completion, the boundary line becomes blurred, and the cross-modal feature stereo fusion module ensures that the sparse depth points of the input are accurately retained, so the depth boundary line can be completely retained.

[0100] The method of the present invention achieves the best results after adding both modules. Compared with KBnet, the RMSE is improved by 6.0%. From the actual results, the method of the present invention can fill in the missing depth to the greatest extent, and at the same time, there is a clear boundary between the edges of objects at different depths. Information from infrared images can also be correctly fused.

[0101] The present invention also tests the influence of the number of layers of the residual block of the feature extraction network in the infrared adaptive edge enhancement and feature extraction module on the result on the same RTX TITAN, and the results are as follows:

[0102] Table 3. The influence of the number of layers of the feature extraction network residual block on the results

[0103]

[0104] The results in Table 3 show that increasing the number of network layers can indeed improve the indicators, but it also occupies more GPU resources and greatly increases the training time. Therefore, the infrared image features extracted by the 18-layer network structure are sufficient for deep completion, and the marginal effect of increasing the number of network layers is very obvious. Therefore, in the complete completion algorithm, the present invention still uses Resnet18 to strike a balance between the completion effect and computing power requirements.

[0105] 4. Comparison with other algorithms

[0106] Since the present invention uses a self-made data set, it also replaces the private data set when testing other methods, and the training parameters remain the same. From Table 4, it can be seen that the method of the present invention has a certain improvement over other methods. The adaptive edge enhancement and feature extraction module and the cross-modal feature stereo fusion module designed for the infrared modality of the present invention can adapt well to the task and achieve better results.

[0107] Table 4. Comparison of the effects of the present invention and different algorithms

[0108]

[0109] Figure 8 It is a qualitative comparison diagram of other different algorithms under 771 data sets, among which, part (a) is the actual image; part (b) is the infrared image; part (c) is the ENet network processing image; part (d) is the PENet network processing image; part (e) is the processing image of the present invention; the area with the most significant improvement is highlighted with a box.

[0110] In summary, the present invention proposes a depth completion method based on infrared feature mining and cross-modal fusion. According to the task requirements of depth completion, a special infrared adaptive edge enhancement and feature extraction module is designed to solve the problem of blurred edges of infrared images. Based on the similarity between infrared and point cloud modalities, a cross-modal feature stereo fusion module is designed, which enables infrared images to provide good information for completion, completes the completion task excellently, and greatly improves the accuracy of edge completion. Ablation experiments have proved the effectiveness of the two modules proposed in the present invention. The method of the present invention achieves excellent results in complex lighting scenes, which is an important scene for depth completion, and has excellent all-weather depth completion characteristics.

Claims

1. A depth completion method based on infrared feature mining and cross-modal fusion, characterized in that: The following steps are involved: Step 1: Use an infrared camera to collect infrared images, and use a laser radar to collect sparse depth maps; Step 2: construct a depth completion network. The depth completion network uses the unsupervised depth completion of the calibrated back-projection layer to back-project each pixel in the two-dimensional image into the three-dimensional space using the calibration matrix and the deep feature descriptor. The depth completion network includes an infrared adaptive edge enhancement and feature extraction module, an infrared branch encoder, a depth branch encoder, a decoder, and a cross-modal feature stereo diffusion module. The infrared adaptive edge enhancement and feature extraction module includes an edge enhancement module and a feature extraction module. The edge enhancement module includes a noise reduction module, a Roberts operator and a Canny operator. The noise reduction module is used to filter the noise in the infrared image to obtain a noise-reduced infrared image. The Roberts operator and the Canny operator respectively detect the image edge of the noise-reduced infrared image to obtain a Roberts image edge and a Canny image edge. The infrared image, the edge of the Roberts image and the edge of the Canny image are connected and input into the feature extraction module to obtain an edge-enhanced infrared image, wherein the feature extraction module is composed of a convolution layer and four residual blocks; The sparse depth map is input into the depth branch encoder to obtain a depth map code; The infrared adaptive edge enhancement and feature extraction module is used to process the infrared image to obtain the deep features of the infrared image, and the deep features of the infrared image are input into the infrared branch encoder to obtain the infrared image code; The depth map code and the infrared image code are connected and input into the decoder to obtain a pre-completed depth map; The pre-completed depth map and the deep features of the infrared image together form a stereo convolution layer that is input into the cross-modal feature stereo diffusion module; The cross-modal feature stereo diffusion module is used to fuse the pre-completed depth map and the deep features of the infrared image to obtain a dense depth map; Step 3: Output dense depth map.

2. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 1, characterized in that: The noise reduction module in step 2 is a Gaussian filter.

3. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 1, characterized in that: The edge enhancement module in step 2 also includes a Sobel operator, which is used to determine the high and low thresholds in the Canny operator.

4. The depth completion method based on infrared feature mining and cross-modal fusion as claimed in claim 3, characterized in that: The Sobel operator is used to determine the high and low thresholds in the Canny operator, including the following steps: Step 21: Calculate the horizontal and vertical gradients at pixel i using the Sobel operator. and For each pixel, the gradient magnitude is calculated according to the following formula based on its gradient components in the x and y directions: , In the formula, and is the horizontal and vertical gradient at pixel i, is the gradient amplitude, is the gradient direction; Step 22: After the gradient amplitude of all pixels is obtained, the high threshold and the low threshold are calculated according to the following formula: , Where high and low are the high threshold and low threshold, respectively, m and n are the number of pixels in the x direction and y direction, respectively. is a constant.

5. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 1, characterized in that: The horizontal and vertical operators of the Roberts operator in step 2 and It is expressed by the following formula: 。 6. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 1, characterized in that: In step 2, the Roberts operator detects the image edge of the denoised infrared image, specifically: the horizontal and vertical operators of the Roberts operator are used as two convolution kernels, the denoised infrared image is convolved, and the image edge is obtained after the average value is calculated.

7. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 6, characterized in that: In step 2, the Roberts operator detecting the image edge of the de-noised infrared image further includes a filtering step: filtering the obtained image edge using a Butterworth low-pass filter.

8. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 1, characterized in that: The output of the residual block in step 2 is consistent with the input size.

9. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 1, characterized in that: The cross-modal feature stereo diffusion module in step 2 includes a feature diffusion module, and the feature diffusion module is a convolutional spatial propagation network CSPN.

10. The depth completion method based on infrared feature mining and cross-modal fusion according to claim 1, characterized in that: The stereo convolution layer includes deep features of infrared images, depth map Z, image X and image Y connected in sequence. The depth map Z is obtained by projecting the pre-completed depth map to the camera perspective into a two-dimensional grayscale image. The grayscale value is used to represent the distance of a point in the three-dimensional space corresponding to the pixel from the camera. The length and width of the deep features of the infrared image, the depth map Z, image X and image Y are the same. Image X and image Y correspond to the depth map Z and are expressed by the following formula: , In the formula, , , and is the camera’s internal parameter, To pre-complete the coordinates of the points in the depth map, Yes The corresponding coordinates on the camera's pixel plane.

Citation Information

Patent Citations

  • Multi-modal image semantic segmentation method based on cross-modal feature enhancement and interaction

    CN115546489A

  • Device and method for assisting laparoscopic surgery - directing and maneuvering articulating tool

    US20220395159A1