A depth estimation method combining visible light camera and infrared camera
By introducing basic CCF and CRF modules into the depth estimation network, the feature information of visible light and infrared cameras is effectively fused, and the problem of insufficient feature fusion in the prior art is solved, and high-precision depth estimation is achieved.
Patent Information
- Application Number
- CN202310197043.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-03-03
AI Technical Summary
The prior art is difficult to effectively integrate the feature information of visible light cameras and infrared cameras, resulting in the inability to fully utilize the information of the two modal features in the depth estimation task.
A deep estimation network VINet is proposed, and features are fusion through the CCF basic module and the CRF basic module. The CCF basic module combines features of different scales through channel shuffling and splicing operations, while the CRF basic module selectively emphasizes relevant features through compression-excitation operations and attention mechanisms.
The accuracy of depth estimation results is improved, and high accuracy can be achieved while ensuring speed, making full use of visible light and infrared features.
Smart Images

Figure CN116563360B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision depth estimation, and in particular to a depth estimation method combining a visible light camera and an infrared camera. Background Art
[0002] Nowadays, with the development of computer vision applications, the demand for all-weather depth information acquisition has increased significantly. Depth information refers to the actual vertical distance from an object to a measuring device. All-weather depth information acquisition requires that the depth estimation method can effectively obtain the distance information of the object both during the day and at night. To this end, it is necessary to use visible light cameras and infrared cameras to complete image acquisition together to ensure that effective image data can be collected both during the day and at night. The current difficulty lies in how to fuse the information of visible light cameras and infrared cameras, that is, how to perform effective multimodal feature fusion. In recent years, convolutional neural networks (CNNs) have made remarkable achievements in computer vision tasks and have become the main method for depth estimation. Features at different levels in CNNs contain different information. Shallow features are rich in image details, while deep features have stronger semantic information. Recently, many studies related to visual recognition have proved that fully fusing features at different levels can effectively enhance the representation ability of multimodal features. The existing method "Li Yuqi, Zhao Haitao. Scene depth estimation based on step-by-step adaptive fusion of infrared and visible light images. Applied Optics, 41(1):24-32, 2020" fuses the visible light and infrared feature representations at different stages and then uses them for depth estimation. However, even with this feature fusion strategy, the features of the two modalities cannot be fully utilized. How to fuse these two features in a more effective way is very important for the depth estimation task of combining visible light cameras and infrared cameras.
[0003] The current status of research on depth estimation of joint visible light camera and infrared camera is as follows: It is generally believed that in this task, the visible light camera and infrared camera have been calibrated, and the visible light image and the infrared image represent the same scene. Referring to other similar fields, such as the depth estimation method of joint visible light camera and lidar, there are two types of methods that can fuse these two data. First, such as "Xinjing Cheng, Peng Wang and Ruigang Yang. Depth Estimation via Affinity Learned with Convolutional Spatial Propagation Network. ECCV 2018.", the two images are directly fused in a splicing manner, and then CNNs are used to extract features. This method is simple and direct, but it has been proved that its fusion effect is poor in the depth estimation method of joint visible light camera and lidar. Therefore, no similar public method has been found in the depth estimation of joint visible light camera and infrared camera. Second, CNNs are used to extract the features of visible light image and infrared image respectively, and then feature fusion is performed. This method is more thorough in fusion and can obtain better depth estimation results. According to the second design method, “Yuqi Li, Haitao Zhao. Scene depth estimation based on step-by-step adaptive fusion of infrared and visible light images. Applied Optics, 41(1):24-32, 2020” designed a pixel-level adaptive fusion method of visible light images and infrared images for depth estimation. Its disadvantage is that the cross-fusion of the two modal features is not deep, and the different information in different features cannot be fully utilized. “Yuqi Li, Haitao Zhao, Zhengwei Hu, Qianqian Wang, Yuru Chen. IVFuseNet: Fusion of infrared and visible light images for depth prediction. Information Fusion, 58, 2020.” proposed common feature fusion, full feature fusion and high-resolution reconstruction network to complete the fusion of the two modal features. Its disadvantage is that it does not explore the inherent characteristics of the two modalities and cannot give play to the complementary characteristics of the two cameras. Summary of the invention
[0004] The purpose of the present invention is to provide a depth estimation method combining a visible light camera and an infrared camera, and propose a depth estimation network to improve the accuracy of the depth estimation result through visible light and infrared features at different stages.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A depth estimation method that combines visible light cameras and infrared cameras is proposed. A depth estimation network VINet is proposed based on CCF basic modules and CRF basic modules. The CCF basic module can fuse the information of the two modal features while extracting visible light features and infrared features of different scales. It can not only cross the two modal information, but also retain rich information and avoid errors caused by the crossing of erroneous information. The CRF basic module completes the fusion of visible light features and infrared features according to the sensor characteristics, selectively emphasizes the features of more relevant modalities during fusion, and can mine the intrinsic properties of each modal feature.
[0007] The specific steps are as follows:
[0008] Step 1) Construct the CCF basic module; the CCF basic module includes the channel shuffle operation module, the splicing operation module and the 1×1 convolution layer;
[0009] The input feature is the visible light image feature map f of the same scale in the i-th stage i V And infrared image feature map f i I , set the size of both tensors to h×w×c; first use the Channel Shuffle operation to recombine the c channels of each feature in the two feature maps to form a new feature; the front of the visible light image feature map and the infrared image feature map The channels form a new visible light feature. channels form a new infrared feature; the combination of each channel in the new visible light feature is The combination of each channel in the new infrared feature is Then, the concatenation operation is used to fuse the input features and the new features along the channel dimension to obtain fused features of size h×w×2c; the visible light image feature map f i V Combined with the new visible light feature map, infrared image feature map f i I It is concatenated and fused with the new infrared features; finally, a 1×1 convolutional layer is used to compress the number of channels of the fused features to c and output the visible light features. and infrared characteristics Tensors of size h×w×c respectively;
[0010] Step 2) construct a CRF basic module; the CRF basic module includes a compression-excitation operation module, a multiplication operation module, an addition operation, a normalization operation module and a splicing operation; the compression-excitation operation includes a global flat pooling layer, a fully connected layer and a Sigmoid activation function;
[0011] Input the visible light features of the same scale output by the CCF basic module in the i-th stage and infrared characteristics First, the Squeeze-and-Excitation module is used to obtain the visible light characteristics. and infrared characteristics The global statistics of and infrared feature attention vector The size of the two attention vectors is 1×c; then, the attention vector is multiplied with its corresponding feature to obtain the activation feature map in the visible light modality. and infrared modality activation feature map The two attention vectors are added and normalized along the channel dimension using the Softmax function to obtain the cross-modal attention vector. Cross-modal attention vector Respectively with visible light characteristics and infrared characteristics Perform multiplication operation to obtain the visible light cross-modal activation feature map and infrared cross-modal activation feature maps Visible light modality activation feature map Cross-modal activation feature map with visible light Addition operation obtains visible light cross-modal activation feature map Infrared modality activation feature map and infrared cross-modal activation feature maps Addition operation to obtain infrared cross-modal activation feature map Finally, the concatenation operation is used to fuse the visible light cross-modal activation feature maps along the channel dimension. and infrared cross-modal activation feature maps Get the concatenated features, use a 1×1 convolution layer to compress the number of channels of the concatenated features to c, and then use a 3×3 convolution layer with a step size of 1 to obtain the final fusion feature map
[0012] Step 3) Building a VINet network through the CCF basic module and the CRF basic module constructed in step 1) and step 2); the entire VINet network is divided into two parts, an encoder and a decoder, and the encoder structure of the visible light image and the infrared image is the same;
[0013] The encoder consists of a convolution operation network layer and three Resnet network layers; the Resnet network layer is implemented by "He Kaiming, Zhang Xiangyu, Ren Shaoqing, Sun Jian. Deep residual learning for image recognition. CVPR 2016"; the convolution operation includes a 3×3 convolution layer with a step size of 2 and a 3×3 convolution layer with a step size of 1; the feature scales obtained by each network layer are as well as right as well as The visible light features and infrared features of the characteristic scale are respectively input into the decoder after executing the CCF basic module;
[0014] The decoder includes a 4-layer network structure, which is used to gradually restore the resolution of visible light features and infrared features to their original resolution; The network structure of feature scale first uses the CRF basic module to complete the multi-modal feature fusion, and then uses the deconvolution layer to complete the fusion feature map of multi-modal feature fusion. Upsample by 2 times to get Decoded features at feature scale; obtained by upsampling The decoded features of the feature scale and the CRF fusion The feature scale feature is concatenated along the channel dimension and then compressed through a 1×1 convolutional layer to reduce the number of channels. The number of feature channels of the feature scale is consistent; repeat the above operation until you get Decoding features of feature scales; The decoded features of the feature scale are passed through a deconvolution layer to obtain the features of the input resolution, which are used as the depth map for the final depth estimation;
[0015] Step 4) According to the feature map obtained in step 3), a 3×3 convolution layer with a step size of 1 and a ReLU activation function are used to obtain the depth estimation result; the loss function is constructed using the mean absolute error to train the network, and the loss is: Where V represents all pixels, D p represents the depth prediction at pixel position p, Represents the true depth value at pixel position p.
[0016] Beneficial effects of the present invention: A depth estimation method for a combined visible light camera and infrared camera of the present invention constructs a neural network based on two lightweight modules (CCF basic module and CRF basic module), and builds VINet through the CCF basic module and the CRF basic module. The CCF basic module strengthens the information exchange between the two modal features in the feature encoding stage, and the CRF basic module explores the characteristics of each modal feature and selectively emphasizes more relevant features during fusion. Both the CCF basic module and the CRF basic module have the characteristics of lightweight and cross-modal intersection, and are therefore suitable for deployment in the depth estimation of a combined visible light camera and infrared camera. While ensuring speed, they can also achieve high accuracy, providing a new solution for the depth estimation task of a combined visible light camera and infrared camera. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1(a) is the overall structure diagram of the depth estimation network VINet;
[0018] Figure 1(b) is a structural diagram of the basic module of CCF;
[0019] Figure 1(c) is the structural diagram of the basic module of CRF.
[0020] In the figure: is the addition operation; For splicing operation; is a normalization operation; ⊙ is a multiplication operation. DETAILED DESCRIPTION
[0021] In order to make the objectives, technical solutions and advantages of the present invention more clear, the specific implementation modes of the present invention are further described in detail below.
[0022] As shown in Figure 1(a), the VINet network is built by using the CCF basic module and the CRF basic module. The entire VINet network is divided into two parts: the encoder and the decoder. The encoder structure of visible light images and infrared images is the same. The encoder is a convolution operation network layer and three Resnet network layers. The convolution operation includes a 3×3 convolution layer with a step size of 2 and a 3×3 convolution layer with a step size of 1. The feature scales obtained by each network layer are as well as right as well as The visible light features and infrared features of the characteristic scale are respectively input into the decoder after executing the CCF basic module;
[0023] The decoder includes a 4-layer network structure, which is used to gradually restore the resolution of visible light features and infrared features to their original resolution; The feature scale features first use the CRF basic module to complete the multi-modal feature fusion, and then use the deconvolution layer to complete the fusion feature map of multi-modal feature fusion Upsample by 2 times to get Decoded features at feature scale; obtained by upsampling The decoded features of the feature scale and the CRF fusion The scale feature is concatenated along the channel dimension and then compressed through a 1×1 convolution layer to reduce the number of channels. The number of feature channels of the feature scale is consistent; repeat the above operation until you get Decoding features of feature scales; The decoded features of the feature scale are passed through a deconvolution layer to obtain the features of the input resolution, which are used as the depth map for the final depth estimation.
[0024] As shown in Figure 1(b), the CCF basic module includes a channel shuffle operation module, a splicing operation module and a 1×1 convolutional layer; in "Liu, Lina, Xibin Song, Xiaoyang Lyu, Junwei Diao, Mengmeng Wang, Yong Liu, and Liangjun Zhang. Fcfr-net: Feature fusion based coarse-to-fine residual learning for depth completion. AAAI 2021.", there is a structure similar to CCF, but only the channel shuffle operation is used, which may introduce erroneous information crossover. The difference between the present invention and the channel shuffle operation is that after using the channel shuffle to complete the information crossover, the original feature is reused to retain the original modal information to alleviate potential errors.
[0025] The input feature is the visible light image feature map f of the same scale in the i-th stage i V And infrared image feature map f i I , set the size of both tensors to h×w×c; first use the Channel Shuffle operation to recombine the c channels of each feature in the two feature maps to form a new feature; the front of the visible light image feature map and the infrared image feature map The channels form a new visible light feature. channels form a new infrared feature; the combination of each channel in the new visible light feature is The combination of each channel in the new infrared feature is Then, the concatenation operation is used to fuse the input features and the new features along the channel dimension to obtain fused features of size h×w×2c; the visible light image feature map f i V Combined with the new visible light feature map, infrared image feature map f i I It is concatenated and fused with the new infrared features; finally, a 1×1 convolutional layer is used to compress the number of channels of the fused features to c and output the visible light features. and infrared characteristics are tensors of size h×w×c respectively.
[0026] As shown in Figure 1(c), the basic module of CRF includes a compression-excitation operation module, a multiplication operation module, an addition operation, a normalization operation module and a splicing operation; the compression-excitation operation includes a global flat pooling layer, a fully connected layer and a Sigmoid activation function; in "Wei Ji, Jingjing Li, Shuang Yu, Miao Zhang, Yongri Piao, Shunyu Yao, Qi Bi, Kai Ma, Yefeng Zheng, Huchuan Lu and Li Cheng. Calibrated RGB-D Salient Object Detection. CVPR 2021.", there is a structure similar to CRF. The main difference is that the Max function is used in the reference to filter the attention vectors of the two modalities, and the present invention uses a normalization function to complete the fusion of the attention vectors, so that the features can be comprehensively integrated instead of focusing only on obvious feature areas.
[0027] Input the visible light features of the same scale output by the CCF basic module in the i-th stage and infrared characteristics First, the Squeeze-and-Excitation module is used to obtain the visible light characteristics. and infrared characteristics The global statistics of and infrared feature attention vector The size of the two attention vectors is 1×c; then, the attention vector is multiplied with its corresponding feature to obtain the activation feature map in the visible light modality. and infrared modality activation feature map The two attention vectors are added and normalized along the channel dimension using the Softmax function to obtain the cross-modal attention vector. Cross-modal attention vector Respectively with visible light characteristics and infrared characteristics Perform multiplication operation to obtain the visible light cross-modal activation feature map and infrared cross-modal activation feature maps Visible light modality activation feature map Cross-modal activation feature map with visible light Addition operation obtains visible light cross-modal activation feature map Infrared modality activation feature map and infrared cross-modal activation feature maps Addition operation to obtain infrared cross-modal activation feature map Finally, the concatenation operation is used to fuse the visible light cross-modal activation feature maps along the channel dimension. and infrared cross-modal activation feature maps Get the concatenated features, use a 1×1 convolution layer to compress the number of channels of the concatenated features to c, and then use a 3×3 convolution layer with a step size of 1 to obtain the final fusion feature map
[0028] The proposed VINet is implemented on NVIDIA 2080Ti GPU and Intel(R) Xeon(R) CPU E5-2678 v3 using CUDA 10.2 and cuDNN 7.3.1 backend. The Adam optimizer is used with parameter β 1 =0.9,β 2 =0.999. The initial learning rate is set to 1e-3. The batch size is 8. Horizontal flipping, color jittering, and random cropping are used to increase the dataset. In the comparative experiment, the baseline model does not use the CCF basic module, and the CRF basic module is replaced by a normal splicing operation and a convolution layer. When comparing, the CCF basic module and the CRF basic module are added to the baseline model in turn, and the effectiveness of each module is proved by comparing the results of the model after adding the modules with the baseline model. Finally, both modules are added to verify the performance of the proposed VINet model. Finally, it is found that the model with both basic modules added can achieve high accuracy while ensuring speed.
[0029] The above are specific embodiments of the present invention and the technical principles used. If the changes made according to the concept of the present invention do not exceed the spirit covered by the description and drawings, they should still fall within the scope of protection of the present invention.
Claims
1. A depth estimation method combining visible light camera and infrared camera, It is characterized in that Based on the CCF basic module and the CRF basic module, a depth estimation network VINet is proposed for depth estimation of the combined visible light camera and infrared camera. The specific steps are as follows: Step 1), build the CCF basic module; the CCF basic module includes the channel shuffle operation module, the splicing operation module and the 1×1 convolution layer; The input feature is the visible light image feature map f of the same scale in the i-th stage i V And infrared image feature map f i I , set the size of both tensors to h×w×c; first use the ChannelShuffle operation to recombine the c channels of each feature in the two feature maps to form a new feature; the front of the visible light image feature map and the infrared image feature map The channels form a new visible light feature. channels form a new infrared feature; the combination of each channel in the new visible light feature is The combination of each channel in the new infrared feature is Then, the concatenation operation is used to fuse the input features and the new features along the channel dimension to obtain fused features of size h×w×2c; the visible light image feature map f i V Combined with the new visible light feature map, infrared image feature map f i I Stitching and fusion with new infrared features; Finally, use a 1×1 convolutional layer to compress the number of channels of the fused features to c, and output the visible light features and infrared features which are tensors of size h×w×c respectively; Step 2) construct the CRF basic module; The basic modules of CRF include compression-excitation operation module, multiplication operation module, addition operation, normalization operation module and concatenation operation; the compression-excitation operation includes global flat pooling layer, fully connected layer and Sigmoid activation function; Input the visible light features of the same scale output by the CCF basic module in the i-th stage and infrared characteristics First, the Squeeze-and-Excitation module is used to obtain the visible light characteristics. and infrared characteristics The global statistics of and infrared feature attention vector The size of the two attention vectors is 1×c; then, the attention vector is multiplied with its corresponding feature to obtain the activation feature map in the visible light modality. and infrared modality activation feature map The two attention vectors are added and normalized along the channel dimension using the Softmax function to obtain the cross-modal attention vector. Cross-modal attention vector Respectively with visible light characteristics and infrared characteristics Perform multiplication operation to obtain the visible light cross-modal activation feature map and infrared cross-modal activation feature maps Visible light modality activation feature map Cross-modal activation feature map with visible light Addition operation obtains visible light cross-modal activation feature map Infrared modality activation feature map and infrared cross-modal activation feature maps Addition operation to obtain infrared cross-modal activation feature map Finally, the concatenation operation is used to fuse the visible light cross-modal activation feature maps along the channel dimension. and infrared cross-modal activation feature maps Get the concatenated features, use a 1×1 convolution layer to compress the number of channels of the concatenated features to c, and then use a 3×3 convolution layer with a step size of 1 to obtain the final fusion feature map Step 3) Building a VINet network through the CCF basic module and the CRF basic module constructed in step 1) and step 2); the entire VINet network is divided into two parts, an encoder and a decoder, and the encoder structure of the visible light image and the infrared image is the same; The encoder consists of a convolution operation network layer and three Resnet network layers. The convolution operation includes a 3×3 convolution layer with a step size of 2 and a 3×3 convolution layer with a step size of 1. The feature scales obtained by each network layer are as well as The characteristic scale is as well as The visible light features and infrared features are respectively input into the decoder after executing the CCF basic module; The decoder consists of a 4-layer network structure, which is used to gradually restore the resolution of visible light features and infrared features to their original resolution; Input to the decoder The features of the feature scale first complete the multi-modal feature fusion using the basic CRF module, and then use the deconvolution layer to perform upsampling by a factor of 2 to obtain the decoded features of the scale; Upsampling obtained The decoded features of the scale and the CRF fusion The feature scale is concatenated along the channel dimension and then compressed by a 1×1 convolutional layer. The number of feature channels of the feature scale is consistent; repeat the above operation until you get Decoding features of feature scales; The decoded features of the feature scale are passed through a deconvolution layer to obtain the features of the input resolution, which are used as the depth map for the final depth estimation; Step 4) According to the feature map obtained in step 3), a 3×3 convolution layer with a step size of 1 and a ReLU activation function is used to obtain the depth estimation result; The mean absolute error is used to construct the loss function to train the network, and the loss is: Where V represents all pixels, D p represents the depth prediction at pixel position p, Represents the true depth value at pixel position p.
Citation Information
Patent Citations
Foresight scene depth estimation method based on self-supervised learning
CN113313732A
3D information sensing method and system based on RGB and infrared images
CN114782541A