Deep Image Super-Resolution Method and System Based on Uncertainty-Aware Feature Transmission

Through the uncertainty-aware feature transmission network and iterative up-down sampling pipeline, the problem of resolution gaps and cross-modal gaps between depth maps and RGB maps is solved, and high-quality depth map super-resolution reconstruction is achieved, reducing noise and blurring, and improving the reconstruction effect.

CN115511708BActive Publication Date: 2025-07-08WUHAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211135383.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-07-08
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

In the existing depth map super-resolution methods, the resolution gap and cross-modal gap between the depth map and RGB guide images lead to texture replication artifacts and depth bleeding, making it difficult to accurately reconstruct high-resolution depth maps.

Method used

Using an uncertain-aware feature transmission method, the features of low-resolution depth maps and high-resolution RGB maps are extracted through the uncertain-aware feature transmission network, symmetric uncertainty maps are used to reduce texture mismatch, and an iterative up-down sampling pipeline is constructed to reduce noise amplification and blurring.

Benefits of technology

Effectively reduces texture replication artifacts in reconstruction results, improves the reconstruction quality of high-resolution depth maps, reduces redundancy consumption of computing resources, and is easy to integrate into existing methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115511708B_ABST
    Figure CN115511708B_ABST
Patent Text Reader

Abstract

The present invention discloses a depth map super-resolution method and system based on uncertainty-aware feature transmission. By constructing a pipeline of iterative upsampling and downsampling during feature transmission to replace the pre-interpolation upsampling in the existing method, the resolution gap between the depth map and the RGB guidance image is eliminated while avoiding side effects such as noise amplification. The present invention proposes a symmetric uncertainty scheme that can model the uncertainty of depth features during feature transmission. Then, the generated uncertainty map is used to weight the RGB features to remove the RGB features that do not match the texture of the depth image, alleviating the texture replication phenomenon caused by the cross-modal gap between the two images. In each iteration, the present invention can obtain the uncertainty map by only propagating forward once, reducing the redundant consumption of computing resources. At the same time, the present invention is easy to integrate into the existing color-guided depth image super-resolution model and further effectively improves the performance of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image reconstruction, and relates to a depth map super-resolution method and system, specifically to a depth map super-resolution method and system based on uncertainty-aware feature transfer. Background Art

[0002] Depth images are an important complement to the RGB modality and can provide depth information for humans or computer vision systems to better understand scenes. Better scene understanding is beneficial to research in many fields of computer vision, such as scene recognition, autonomous navigation, 3D reconstruction, etc. These tasks usually rely on high-quality depth information. However, the depth maps obtained by existing commercial depth sensors usually have low resolution and are difficult to be used for various computer vision tasks. Therefore, depth map super-resolution is a practical and valuable technology that elevates depth maps from a low-resolution space to a high-resolution space.

[0003] Some existing depth map super-resolution methods usually utilize registered high-resolution RGB images under the same scene to guide the reconstruction of depth maps (Literatures 1 and 2). Such methods are called color-guided depth map super-resolution. Currently, color-guided depth map super-resolution methods mainly face two major problems:

[0004] 1. Resolution gap: The inconsistent resolution sizes of the depth map and the RGB guidance image result in the inability to directly fuse the features of the two modalities;

[0005] 2. Cross-modal gap: The textures of the depth map and the RGB guidance image do not completely match. This will cause texture replication artifacts and depth bleeding phenomena in the reconstructed high-resolution depth map.

[0006] The basic training and testing steps of conventional color-guided depth map super-resolution methods are as follows:

[0007] 1. Prepare an RGB-depth image pair dataset and divide the dataset into a training set and a testing set;

[0008] 2. Input the data in the training set into a neural network for training, including steps such as the construction of a basic network, feature extraction of RGB images and depth maps, feature fusion, loss optimization, etc.;

[0009] 3. Save the optimal model during the training process, and finally use this model to test the data in the testing set to obtain the model performance results.

[0010] Regarding the resolution gap between the depth map and the RGB guidance image, current methods usually use pre-interpolation upsampling to increase the resolution of the depth map to be the same as that of the RGB guidance image. However, this approach brings some side effects, such as noise amplification and blurring. In addition, existing methods usually have two branches or sub-networks, one for extracting features of the low-resolution depth map and the other for extracting features of the corresponding high-resolution RGB image. They transfer the high-frequency features extracted from the RGB image to the depth map branch or sub-network to better restore the edge details in the depth map. However, this approach ignores the cross-modal gap between the two images. Not all of the high-frequency information in the RGB image is required for depth map reconstruction.

[0011] In summary, how to make the depth features and RGB features consistent in spatial size while avoiding the above side effects and accurately estimating and removing the RGB features that do not match the depth image texture during the feature transmission process, so that the model can accurately reconstruct high-resolution depth maps is an urgent problem to be solved.

[0012] [Reference 1] He, Linzhi, et al. "Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021.

[0013] [Reference 2] Tang, Qi, et al. "BridgeNet: A Joint Learning Network of Depth Map Super-Resolution and Monocular Depth Estimation." Proceedings of the 29th ACM International Conference on Multimedia. 2021. Summary of the Invention

[0014] Aiming at the problems existing in the prior art, the present invention provides a Symmetric Uncertainty-aware Feature Transmission (SUFT) technology to reduce the resolution gap and cross-modal gap between the depth map and the RGB guidance image and improve the performance of the depth super-resolution method.

[0015] The technical solution adopted by the method of the present invention is: a depth map super-resolution method based on uncertainty-aware feature transmission, comprising the following steps:

[0016] Step 1: For the input image, extract the features of the low-resolution depth image and the high-resolution RGB guidance image through the RGB branch and the Depth branch of the uncertainty-aware feature transmission network;

[0017] Input the features of the low-resolution depth image and the high-resolution RGB guidance image into the SUFT module of the uncertainty-aware feature transmission network. The SUFT module first copies and horizontally flips the input depth features in the spatial dimension, and then projects these two horizontally mirrored depth features into the high-resolution domain:

[0018]

[0019]

[0020] where is the feature extracted from the low-resolution depth map, is the high-resolution depth feature obtained by upsampling, is the high-resolution depth feature after flipping, HFlip(·) and (·)↑ s represent the horizontal flipping operation and the up-projection operation with a scaling factor of s, respectively;

[0021] The uncertainty-aware feature transmission network is composed of an RGB branch, a Depth branch, and an SUFT module as a whole;

[0022] The RGB branch is composed of a first 3×3 convolutional layer, a first residual block, a second residual block, and a third residual block connected in sequence. Input the high-resolution RGB image, and after passing through the RGB branch, extract the features of the high-resolution RGB image and input them into the corresponding SUFT module;

[0023] The Depth branch is composed of a second 3×3 convolutional layer, a first residual group, a second residual group, a third residual group, a fourth residual group, an up-projection unit, a fifth residual group, a sixth residual group, a third 3×3 convolutional layer, and a bicubic linear interpolation module. Input the low-resolution depth map, and after passing through the Depth branch, extract the features of the low-resolution depth map and input them into the corresponding SUFT module. Finally, add the high-frequency components of the high-resolution depth map extracted by the network and the low-frequency components of the high-resolution depth map obtained by the bicubic linear interpolation module element by element to output the reconstructed high-resolution depth map;

[0024] The first residual block, the second residual block, and the third residual block are composed of two 3×3 convolutional layers and one rectified linear unit layer; the first residual group, the second residual group, the third residual group, and the fourth residual group are composed of eight convolutional layers, four rectified linear unit layers, and four channel attention modules; the fifth residual group and the sixth residual group are composed of sixteen convolutional layers, eight rectified linear unit layers, and eight channel attention modules; the up-projection unit is composed of two convolutional layers with adaptive kernel sizes, two deconvolutional layers with adaptive kernel sizes, and four rectified linear unit layers; the bicubic linear interpolation module upsamples the input low-resolution depth map to obtain a blurred high-resolution depth map;

[0025] Step 2: Calculate the spatial distribution of symmetric uncertainty using the two horizontally mirrored high-resolution depth features obtained in Step 1 to obtain an uncertainty map

[0026] Step 3: Multiply the uncertainty map obtained in Step 2 by the high-resolution RGB guidance image features extracted in Step 1, and then concatenate it with the upsampled high-resolution depth features along the channel axis:

[0027]

[0028] where are the features extracted from the high-resolution RGB guidance image, are the fused features, and [·;·] represents the concatenation operation along the channel axis;

[0029] Step 4: Map the fused features back to the low-resolution spatial domain through the down-projection unit:

[0030]

[0031] where (·)↓ s represents the down-projection operation with a scale factor of s.

[0032] The technical solution adopted by the system of the present invention is: A depth map super-resolution system based on uncertainty-aware feature transmission, including the following modules:

[0033] Module 1: For the input image, extract the features of the low-resolution depth image and the high-resolution RGB guidance image through the RGB branch and the Depth branch of the uncertainty-aware feature transmission network;

[0034] The features of both the low-resolution depth image and the high-resolution RGB guidance image are input into the SUFT module based on the uncertainty-aware feature transfer network. The SUFT module first replicates and horizontally flips the input depth features in the spatial dimension, and then projects these two horizontally mirrored depth features into the high-resolution domain:

[0035]

[0036]

[0037] where is the feature extracted from the low-resolution depth map, is the high-resolution depth feature obtained by upsampling, is the high-resolution depth feature after flipping, and HFlip(·) and (·)↑ s represent the horizontal flipping operation and the up-projection operation with a scaling factor of s, respectively;

[0038] The uncertainty-aware feature transfer network is composed of an RGB branch, a Depth branch, and an SUFT module as a whole;

[0039] The RGB branch is composed of a first 3×3 convolutional layer, a first residual block, a second residual block, and a third residual block connected in sequence. The high-resolution RGB image is input, and after passing through the RGB branch, the features of the high-resolution RGB image are extracted and passed into the corresponding SUFT module;

[0040] The Depth branch is composed of a second 3×3 convolutional layer, a first residual group, a second residual group, a third residual group, a fourth residual group, an up-projection unit, a fifth residual group, a sixth residual group, a third 3×3 convolutional layer, and a bicubic linear interpolation module. The low-resolution depth map is input, and after passing through the Depth branch, the features of the low-resolution depth map are extracted and passed into the corresponding SUFT module. Finally, the high-frequency components of the high-resolution depth map extracted by the network and the low-frequency components of the high-resolution depth map obtained through the bicubic linear interpolation module are added element by element to output the reconstructed high-resolution depth map;

[0041] The first residual block, the second residual block, and the third residual block are composed of two 3×3 convolutional layers and one rectified linear unit layer; the first residual group, the second residual group, the third residual group, and the fourth residual group are composed of eight convolutional layers, four rectified linear unit layers, and four channel attention modules; the fifth residual group and the sixth residual group are composed of sixteen convolutional layers, eight rectified linear unit layers, and eight channel attention modules; the up-projection unit is composed of two convolutional layers with self-adaptive kernel sizes, two deconvolutional layers with self-adaptive kernel sizes, and four rectified linear unit layers; the bicubic bilinear interpolation module upsamples the input low-resolution depth map to obtain a blurred high-resolution depth map;

[0042] Module 2: Calculate the spatial distribution of symmetric uncertainty using the two horizontally mirrored high-resolution depth features obtained in Module 1 to obtain an uncertainty map

[0043] Module 3: Multiply the uncertainty map obtained in Module 2 by the high-resolution RGB guidance image features extracted in Module 1, and then concatenate it with the upsampled high-resolution depth features along the channel axis:

[0044]

[0045] where are the features extracted from the high-resolution RGB guidance image, are the fused features, and [·;·] represents the concatenation operation along the channel axis;

[0046] Module 4: Map the fused features back to the low-resolution spatial domain through the down-projection unit:

[0047]

[0048] where (·)↓ s represents the down-projection operation with a scale factor of s.

[0049] The present invention has the following advantages:

[0050] (1) In the feature transmission, the present invention constructs an iterative upsampling and downsampling pipeline to replace the commonly used pre-interpolation upsampling, which can eliminate the resolution difference and provide an error feedback mechanism for the projection error in each feature fusion stage to reduce noise amplification and blur.

[0051] (2) The symmetric uncertainty scheme proposed by the present invention can accurately select the effective information in the RGB guidance image during the feature transmission process, reducing the false textures generated in the reconstruction result. And based on this scheme, the network can obtain the uncertainty map by only propagating forward once in each iteration, reducing the redundant consumption of computing resources.

[0052] (3) The method proposed by the present invention can be added to the architecture of the existing color-guided depth map super-resolution method as an independent module, with the characteristics of easy integration. Description of the Drawings

[0053] Figure 1 It is the structural diagram of the uncertainty-aware feature transmission network according to the embodiment of the present invention;

[0054] Figure 2 It is the structural diagram of the SUFT module according to the embodiment of the present invention;

[0055] Figure 3 It is the calculation flow chart of the symmetric uncertainty according to the embodiment of the present invention. Detailed Embodiment

[0056] For the convenience of those of ordinary skill in the art to understand and implement the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0057] A depth map super-resolution method based on uncertainty-aware feature transmission provided by the present invention includes the following steps:

[0058] Step 1: For the input image, extract the features of the low-resolution depth image and the high-resolution RGB guidance image through the RGB branch and the Depth branch of the uncertainty-aware feature transmission network;

[0059] Input the features of the low-resolution depth image and the high-resolution RGB guidance image into the SUFT module of the uncertainty-aware feature transmission network. The SUFT module in this embodiment first copies and horizontally flips the input depth features in the spatial dimension, and then projects these two horizontally mirrored depth features into the high-resolution domain:

[0060]

[0061]

[0062] Where is the feature extracted from the low-resolution depth map, is the high-resolution depth feature obtained by upsampling, is the flipped high-resolution deep feature, HFlip(·) and (·)↑ s They represent the horizontal flipping operation and the upward projection operation with a scaling factor of s respectively;

[0063] Please see Figure 1 ,The uncertainty-aware feature transmission network of this embodiment is composed of an RGB branch, a Depth branch and a SUFT module;

[0064] The RGB branch of this embodiment is composed of a first 3×3 convolutional layer, a first residual block, a second residual block, and a third residual block connected in sequence. A high-resolution RGB image is input, and the features of the high-resolution RGB image are extracted through the RGB branch to be passed into the corresponding SUFT module.

[0065] The Depth branch of this embodiment is composed of a second 3×3 convolutional layer, a first residual group, a second residual group, a third residual group, a fourth residual group, an upper projection unit, a fifth residual group, a sixth residual group, a third 3×3 convolutional layer, and a bicubic linear interpolation module. A low-resolution depth map is input, and the features of the low-resolution depth map are extracted through the Depth branch to be passed into the corresponding SUFT module. Finally, the high-frequency components of the high-resolution depth map extracted by the network and the low-frequency components of the high-resolution depth map obtained through the bicubic linear interpolation module are added element by element, and a reconstructed high-resolution depth map is output;

[0066] The first residual block, the second residual block, and the third residual block of this embodiment are composed of two 3×3 convolutional layers and one rectified linear unit layer; the first residual group, the second residual group, the third residual group, and the fourth residual group of this embodiment are composed of eight convolutional layers, four rectified linear unit layers, and four channel attention modules; the fifth residual group and the sixth residual group of this embodiment are composed of sixteen convolutional layers, eight rectified linear unit layers, and eight channel attention modules; the upper projection unit of this embodiment is composed of two kernel size adaptive convolutional layers, two kernel size adaptive deconvolutional layers, and four rectified linear unit layers; the bicubic linear interpolation module of this embodiment upsamples the input low-resolution depth map to obtain a blurred high-resolution depth map;

[0067] Please see Figure 2 ,The SUFT module of this embodiment is composed of a first upper projection unit, a second upper projection unit, an uncertainty module, and a lower projection unit, which inputs features extracted from a high-resolution RGB image and a low-resolution depth map, removes texture mismatch information in the RGB image features, and outputs fusion features of the high-resolution RGB image and the low-resolution depth map into the Depth branch;

[0068] The uncertainty module of this embodiment consists of a convolutional layer and a normalization layer. The depth map features of the input mirror image are subjected to element-wise subtraction and absolute value operations. The resulting difference map performs maximum and mean operations along the channel axis respectively, and then is concatenated along the channel axis. After that, it passes through the convolutional layer and the normalization layer to output the symmetric uncertainty map;

[0069] The first up-projection unit, the second up-projection unit and the down-projection unit of this embodiment consist of two convolutional layers with self-adaptive kernel sizes, two deconvolutional layers with self-adaptive kernel sizes and four rectified linear unit layers;

[0070] Step 2: Calculate the spatial distribution of symmetric uncertainty using the high-resolution depth features of the two horizontal mirrors obtained in Step 1 to obtain the uncertainty map

[0071] Please refer to Figure 3 , and the specific implementation of Step 2 in this embodiment includes the following sub-steps:

[0072] Step 2.1: Horizontally flip again to align it spatially with ; then perform element-wise subtraction operation and take the absolute value to preliminarily calculate the uncertainty map:

[0073]

[0074] where represents the absolute difference between the two depth features;

[0075] Step 2.2: Perform average pooling and max pooling operations on along the channel axis to summarize its channel information and generate two two-dimensional information maps; then concatenate these two two-dimensional information maps along the channel axis, and then perform convolution on them by a standard convolutional layer to generate a two-dimensional symmetric uncertainty map; the values of the symmetric uncertainty map are finally normalized to the range of [0,1]:

[0076]

[0077] where represents the normalized symmetric uncertainty map, AvgPool(·) represents the average pooling operation, MaxPool(·) represents the max pooling operation, Conv(·) represents the convolution operation, and the specific operation of the normalization Norm(·) is expressed as:

[0078]

[0079] where ∈ is a small value to avoid division by zero during the calculation process, and the default value is 1e -12 , X normis the result of normalization, where max and min represent the maximum and minimum values of the input data X respectively.

[0080] Step 3: Multiply the uncertainty map obtained in Step 2 by the high-resolution RGB guidance image features extracted in Step 1, so that the RGB features corresponding to the regions with greater uncertainty in the depth map obtain higher weights, and vice versa; then concatenate it with the upsampled high-resolution depth features along the channel axis:

[0081]

[0082] where are the features extracted from the high-resolution RGB guidance image, are the fused features, and [·;·] represents the concatenation operation along the channel axis;

[0083] Step 4: Map the fused features back to the low-resolution spatial domain through the down-projection unit:

[0084]

[0085] where (·)↓ s represents the down-projection operation with a scale factor of s, which ensures that the output of the SUFT module has the same spatial size as the input, so that multi-level feature fusion can be performed. By embedding the SUFT module into the multi-stage fusion network, the resolution difference can be eliminated while providing an error feedback mechanism for the projection error in each feature fusion stage to reduce noise amplification and blurring.

[0086] The following further elaborates on the solution of this embodiment through experiments.

[0087] The deep learning framework adopted in this embodiment is Pytorch, version 1.9.0, and the CUDA version is 11.3. The hardware environment for the experiment is an NVIDIA GeForce RTX 3090 graphics card, and the processor is an Intel(R) Xeon(R) Gold 6240C. The specific implementation process of the depth map super-resolution method based on uncertainty-aware feature transmission is as follows:

[0088] The uncertainty-aware feature transmission network in this embodiment can be added as an independent module to the existing color-guided depth map super-resolution network architecture. Only need to remove the pre-interpolation upsampling, and after setting the input of the uncertainty-aware feature transmission module to the RGB features and depth features extracted by the CNN, connect it to the existing deep neural network architecture. In the experiment, the uncertainty-aware feature transmission module was embedded into a simple multi-stage fusion model for implementation.

[0089] The uncertainty-aware feature transmission network of this embodiment is a trained network. Its training process includes the following steps:

[0090] (1) Data preparation: Prepare low-resolution depth maps, corresponding RGB guidance images, and high-resolution depth maps as training and test data.

[0091] The present invention uses the NYU v2, Middlebury, and RGB-D-D datasets. The NYUv2 dataset contains 1449 RGB-depth image pairs, of which 1000 pairs are used as training data and 449 pairs are used as test data. The Middlebury dataset includes a total of 30 RGB-depth image pairs from the Middlebury 2001, 2005, and 2006 datasets. There are 2215 RGB-depth image pairs in the RGB-D-D dataset for training and 405 RGB-depth image pairs for testing. Among them, the values of the depth maps in the NYUv2 dataset and the RGB-D-D dataset represent 16-bit absolute depth in millimeters, and the values of the depth maps in Middlebury represent 8-bit relative depth. In addition, the present invention also evaluates the proposed method under the real-world manner setting of the RGB-D-D dataset. Under this setting, there are 2215 pairs of RGB-depth images for training and 405 pairs for testing. Among them, the low-resolution depth maps are captured by a mobile phone ToF camera and have a size of 192×144; the high-resolution depth maps are captured by an industrial ToF camera and have a size of 512×384. The degradation of the low-resolution depth maps under this setting is more complex, so it is more challenging for the depth map super-resolution method.

[0092] The above three datasets all have three scaling factors: ×4, ×8, and ×16. The low-resolution depth images are obtained from the high-resolution depth images through bicubic interpolation (except for the real-world manner). During training, the original high-resolution depth maps and high-resolution RGB images are cropped into fixed-size blocks of 256×256, which can accelerate the training speed without weakening the network performance. When the selected scale factors are ×4, ×8, and ×16, the corresponding low-resolution depth maps are respectively divided into blocks of size 64×64, 32×32, and 16×16. The present invention is trained on the NYUv2 dataset and tested on the test set of NYUv2, the Middlebury dataset, and the test set of RGB-D-D to verify the performance and generalization ability of the present invention. In addition, the present invention is tested using the model trained under the ×4 scaling factor condition on the NYUv2 dataset in the real-world manner setting of the RGB-D-D dataset to analyze the effectiveness of the present invention in real scenarios. For the NYUv2 dataset and the RGB-D-D dataset, the RMSE is measured in centimeters; for the Middlebury dataset, the RMSE is measured on the original scale of the provided differences.

[0093] During training, the batch size is set to 1, and the model is optimized using the Adam optimizer, where β1 = 0.9, β2 = 0.999, and ∈ = 1e -8 . The initial learning rate of the network is set to 1e -4 , and the learning rate is reduced by a factor of 0.1 every 100 epochs.

[0094] (2) Feed the training image pairs into the uncertainty-aware feature transfer network for training.

[0095] (3) Network optimization and parameter update.

[0096] The update includes two parts: forward propagation and backward propagation. Forward propagation calculates the output and loss function through the network. To make a fair comparison with existing methods, the present invention uses the same loss function as the existing methods when training the network, that is, the L1 loss function, which has been proven to have better performance and convergence than the L2 loss in the depth map super-resolution task. Given a training set that contains N low-resolution depth maps and the corresponding high-resolution RGB guiding images as inputs, and the target depth map as the ground truth:

[0097]

[0098] where, represents the i-th low-resolution depth map in the dataset, represents the \(i\)-th high-resolution RGB guidance image in the dataset, represents the \(i\)-th high-resolution target depth map in the dataset, represents the color-guided depth image super-resolution model, and \(\theta\) represents the learned set of parameters. Then, during the backpropagation process, the gradient of the backpropagated loss is used to update the network through the optimization strategy of stochastic gradient descent.

[0099] In the training stage, RGB-depth image pairs are input in batches. Forward propagation is used to calculate various losses, and backward propagation is used to update the network parameters. After multiple iterations, the final network model is obtained.

[0100] (4) Network testing.

[0101] In the testing stage, the network is not trained or its parameters updated. The trained model is used to process the test RGB-depth image pairs to reconstruct the high-resolution depth map. The quality of the reconstruction result is measured by calculating the root mean square error (RMSE) between the high-resolution depth map reconstructed by the network and the true high-resolution depth map. The lower the RMSE value, the better the quality of the reconstruction result. The calculation method is as follows:

[0102]

[0103] where \(i\) and \(j\) represent the horizontal and vertical coordinates of the pixel point, respectively, and \(D\) ij represents the pixel value at position \((i, j)\) in the true high-resolution image, and \(D'\) ij represents the pixel value at position \((i, j)\) in the high-resolution image reconstructed by the network. \(H\) and \(W\) represent the height and width of \(D\), respectively.

[0104] To verify the effectiveness of the present invention, the present invention is compared with existing depth map super-resolution methods. The existing depth map super-resolution methods mainly include:

[0105] (1) DJF: Yijun Li, Jia-Bin Huang, Narendra Ahuja, and Ming-Hsuan Yang. 2016. Deep joint image filtering. In ECCV. 154–169.

[0106] (2) SVLRM: Jinshan Pan, Jiangxin Dong, Jimmy S Ren, Liang Lin, Jinhui Tang, and Ming Hsuan Yang. 2019. Spatially variant linear representation models for joint filtering. In CVPR. 1702–1711.

[0107] (3) DJFR: Yijun Li, Jia-Bin Huang, Narendra Ahuja, and Ming-Hsuan Yang. 2019. Joint image filtering with deep convolutional networks. IEEE TPAMI (2019), 1909–1923.

[0108] (4) FDKN, DKN: Beomjun Kim, Jean Ponce, and Bumsub Ham. 2019. Deformable kernel networks for guided depth map upsampling. IJCV. 579–600.

[0109] (5) FDSR: Lingzhi He, Hongguang Zhu, Feng Li, Huihui Bai, Runmin Cong, Chunjie Zhang, Chunyu Lin, Meiqin Liu, and Yao Zhao. 2021. Towards fast and accurate real-world depth super-resolution: Benchmark dataset and baseline. In CVPR. 9229–9238.

[0110] (6) JIIF: Jiaxiang Tang, Xiaokang Chen, and Gang Zeng. 2021. Joint implicit image function for guided depth super-resolution. In ACM MM. 4390–4399.

[0111] (7) CTKT: Baoli Sun, Xinchen Ye, Baopu Li, Haojie Li, Zhihui Wang, and Rui Xu. 2021. Learning scene structure guidance via cross-task knowledge transfer for single depth super-resolution. In CVPR. 7792–7801.

[0112] (8) BridgeNet: Qi Tang, Runmin Cong, Ronghui Sheng, Lingzhi He, Dan Zhang, Yao Zhao, and Sam Kwong. 2021. BridgeNet: A Joint Learning Network of Depth Map Super-Resolution and Monocular Depth Estimation. In ACM MM. 2148–2157.

[0113] The tests were conducted on the NYU v2 dataset, and the results are shown in Table 1:

[0114] Table 1

[0115]

[0116] The tests were conducted on the Middlebury dataset and the RGB-D-D dataset, and the results are shown in Table 2:

[0117] Table 2

[0118]

[0119] It can be seen from Table 1 and Table 2 that: compared with the depth map super-resolution models in recent years, the RMSE result obtained by the present invention is higher than that of the existing methods, and the image reconstruction quality is significantly improved. It can be seen from the results under the real-world setting in Table 2 that: the present invention has a better reconstruction effect than the existing methods when facing more complex degradations in the real world, which proves the robustness of the present invention and its potential in processing actual depth map super-resolution tasks in real-world scenarios. There are mainly two reasons for this: 1. The iterative upsampling used in the present invention causes less noise amplification and blurring than the pre-interpolation upsampling commonly used in the existing methods. 2. The symmetric uncertainty scheme proposed in the present invention can effectively reduce the cross-modal gap between the two modal images, thereby reducing the texture replication artifacts in the reconstruction results.

[0120] The present invention constructs an iterative upsampling and downsampling pipeline in feature transmission to replace the pre-interpolation upsampling commonly used in existing methods, reducing side effects such as noise amplification and blurring while eliminating the resolution gap. Specifically, the present invention upsamples the depth features before each feature fusion to make them consistent with the RGB features in spatial size, and projects the high-resolution features back to the low-resolution spatial domain after each fusion for subsequent operations. The present invention also proposes a symmetric uncertainty scheme to narrow the cross-modal gap between the two modal images. It calculates the uncertainty of the features through a simple and effective flipping operation to estimate the regions in the RGB features that do not match the depth texture, and assigns low weights to these mismatched parts to avoid misguiding the depth map recovery. The technical solution adopted by the present invention can be added as an independent module to the architecture of existing color-guided depth image super-resolution methods.

[0121] It should be understood that the above description of the preferred embodiment is relatively detailed, and thus it should not be considered as limiting the protection scope of the present invention. Under the inspiration of the present invention, those of ordinary skill in the art can still make substitutions or deformations without departing from the scope protected by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of protection claimed by the present invention shall be subject to the appended claims.

Claims

1. A deep graph super-resolution method based on uncertainty-aware feature transmission, characterized in that The following steps are involved: Step 1: For the input image, extract the features of the low-resolution depth image and the high-resolution RGB guidance image through the RGB branch and Depth branch based on the uncertainty-aware feature transfer network; The features of both the low-resolution depth image and the high-resolution RGB guidance image are input into the SUFT module based on the uncertainty-aware feature transfer network, which first copies and horizontally flips the input depth features in the spatial dimension, and then projects the two horizontally mirrored depth features to the high-resolution domain: Among them are features extracted from the low-resolution depth map, are high-resolution depth features obtained by upsampling, are high-resolution depth features after flipping, where HFlip(·) and (·)↑ s represent the horizontal flipping operation and the up-projection operation with a scaling factor of s, respectively; The uncertainty perception feature transmission network is composed of an RGB branch, a Depth branch and a SUFT module. The RGB branch is composed of a first 3×3 convolutional layer, a first residual block, a second residual block, and a third residual block connected in sequence. A high-resolution RGB image is input, and the features of the high-resolution RGB image are extracted through the RGB branch to be passed into the corresponding SUFT module. The Depth branch consists of a second 3×3 convolutional layer, a first residual group, a second residual group, a third residual group, a fourth residual group, an upper projection unit, a fifth residual group, a sixth residual group, a third 3×3 convolutional layer and a bicubic linear interpolation module. A low-resolution depth map is input, and the features of the low-resolution depth map are extracted through the Depth branch to be passed into the corresponding SUFT module. Finally, the high-frequency components of the high-resolution depth map extracted by the network and the low-frequency components of the high-resolution depth map obtained through the bicubic linear interpolation module are added element by element, and a reconstructed high-resolution depth map is output; The first residual block, the second residual block and the third residual block are composed of two 3×3 convolutional layers and one rectified linear unit layer; the first residual group, the second residual group, the third residual group and the fourth residual group are composed of eight convolutional layers, four rectified linear unit layers and four channel attention modules; the fifth residual group and the sixth residual group are composed of sixteen convolutional layers, eight rectified linear unit layers and eight channel attention modules; the upper projection unit is composed of two kernel size adaptive convolutional layers, two kernel size adaptive deconvolutional layers and four rectified linear unit layers; the bicubic linear interpolation module upsamples the input low-resolution depth map to obtain a blurred high-resolution depth map; Step 2: Calculate the spatial distribution of symmetric uncertainty using the two horizontally mirrored high-resolution depth features obtained in Step 1 to obtain an uncertainty map Step 3: Multiply the uncertainty map obtained in Step 2 by the high-resolution RGB guidance image features extracted in Step 1, and then concatenate with the upsampled high-resolution depth features along the channel axis: wherein are features extracted from a high-resolution RGB guidance image, are the fused features; [·;·] represents the concatenation operation along the channel axis; Step 4: Map the fused features back to the low-resolution spatial domain through the down-projection unit: where (·)↓ s represents the down-projection operation with a scale factor of s.

2. The depth map super-resolution method based on uncertainty-aware feature transmission according to claim 1, characterized in that: In step 1, the SUFT module, which consists of a first upper projection unit, a second upper projection unit, an uncertainty module, and a lower projection unit, inputs features extracted from a high-resolution RGB image and a low-resolution depth map, removes texture mismatch information in the RGB image features, and outputs fusion features of the high-resolution RGB image and the low-resolution depth map into the Depth branch; The uncertainty module consists of a convolution layer and a normalization layer. The depth map features of the input image are subjected to element-by-element subtraction and absolute value operations. The obtained difference map is subjected to maximum and mean operations along the channel axis and then spliced ​​along the channel axis. Then, it passes through the convolution layer and the normalization layer to output a symmetric uncertainty map. The first upper projection unit, the second upper projection unit, and the lower projection unit are composed of two convolution layers with self-adaptive kernel sizes, two deconvolution layers with self-adaptive kernel sizes, and four rectified linear unit layers.

3. The depth map super-resolution method based on uncertainty-aware feature transmission according to claim 1, wherein: The specific implementation of Step 2 includes the following sub-steps: Step 2.1: Horizontally flip again to align it with spatially; then take the absolute value after element-wise subtraction operation to preliminarily calculate the uncertainty map: wherein represents the absolute difference between two depth features; Step 2.2: For perform average pooling and max pooling operations along the channel axis to summarize its channel information and generate two two-dimensional information maps; then concatenate these two two-dimensional information maps along the channel axis, and then perform convolution on them by a standard convolutional layer to generate a two-dimensional symmetric uncertainty map; the values of the symmetric uncertainty map are finally normalized to the range of [0, 1]: Among them represents the normalized symmetric uncertainty graph, AvgPool(·) represents the average pooling operation, MaxPool(·) represents the maximum pooling operation, Conv(·) represents the convolution operation, and the specific operation of the normalization Norm(·) is expressed as: where ∈ is a small value to avoid division by zero during the calculation, and X norm is the result of normalization, and max and min represent the maximum and minimum values of the input data X, respectively.

4. The depth map super-resolution method based on uncertainty-aware feature transmission according to any one of claims 1-3, characterized in that: The uncertainty-aware feature transfer network is a trained uncertainty-aware feature transfer network; During the training process, the update includes two parts: forward propagation and backward propagation; forward propagation calculates the output and the L1 loss function through the network, given the training set It takes N low-resolution depth maps and the corresponding high-resolution RGB guidance images as inputs, with the target depth map as the ground truth: Among them, represents the i-th low-resolution depth map in the dataset, represents the i-th high-resolution RGB guidance image in the dataset, represents the i-th high-resolution target depth map in the dataset, represents the color-guided depth image super-resolution model, and θ represents the learned set of parameters; then, during the backpropagation process, the gradient of the backpropagated loss is used to update the network through the optimization strategy of stochastic gradient descent; The quality of the reconstruction result is measured by calculating the root mean square error (RMSE) between the high-resolution depth map reconstructed by the network and the real high-resolution depth map; the lower the RMSE value, the better the quality of the reconstruction result. where i and j represent the horizontal and vertical coordinates of the pixel point, respectively, D ij represents the pixel value at the position (i, j) in the true high-resolution image, and D' ij represents the pixel value at the position (i, j) in the high-resolution image reconstructed by the network, and H and W represent the height and width of D, respectively.

5. A depth map super-resolution system based on uncertainty-aware feature transmission, characterized in that, It includes the following modules: Module 1: For the input image, extract the low-resolution depth image and the high-resolution RGB guidance image through the RGB branch and the Depth branch of the uncertainty-aware feature transfer network; Input the features of the low-resolution depth image and the high-resolution RGB guidance image into the SUFT module of the uncertainty-aware feature transfer network. The SUFT module first copies and horizontally flips the input depth features in the spatial dimension, and then projects these two horizontally mirrored depth features into the high-resolution domain: where are the features extracted from the low-resolution depth map, are the high-resolution depth features obtained by upsampling, are the high-resolution depth features after flipping, where HFlip(·) and (·)↑ s represent the horizontal flipping operation and the up-projection operation with a scaling factor of s, respectively; The uncertainty-aware feature transfer network is composed of an RGB branch, a Depth branch, and an SUFT module as a whole; The RGB branch is sequentially composed of a first 3×3 convolution layer, a first residual block, a second residual block, and a third residual block. Input the high-resolution RGB image, and after passing through the RGB branch, extract the features of the high-resolution RGB image and pass them into the corresponding SUFT module; The Depth branch is composed of a second 3×3 convolution layer, a first residual group, a second residual group, a third residual group, a fourth residual group, an upper projection unit, a fifth residual group, a sixth residual group, a third 3×3 convolution layer, and a bicubic linear interpolation module. Input the low-resolution depth map, and after passing through the Depth branch, extract the features of the low-resolution depth map and pass them into the corresponding SUFT module. Finally, element-wise add the high-frequency components of the high-resolution depth map extracted by the network and the low-frequency components of the high-resolution depth map obtained through the bicubic linear interpolation module to output the reconstructed high-resolution depth map; The first residual block, the second residual block, and the third residual block are composed of two 3×3 convolution layers and one rectified linear unit layer; the first residual group, the second residual group, the third residual group, and the fourth residual group are composed of eight convolution layers, four rectified linear unit layers, and four channel attention modules; the fifth residual group and the sixth residual group are composed of sixteen convolution layers, eight rectified linear unit layers, and eight channel attention modules; the upper projection unit is composed of two convolution layers with self-adaptive kernel sizes, two deconvolution layers with self-adaptive kernel sizes, and four rectified linear unit layers; the bicubic linear interpolation module upsamples the input low-resolution depth map to obtain a blurred high-resolution depth map; Module 2: Calculate the spatial distribution of symmetric uncertainty using the two horizontally mirrored high-resolution depth features obtained in Module 1 to obtain an uncertainty map Module 3: Multiply the uncertainty map obtained in Module 2 by the high-resolution RGB guidance image features extracted in Module 1, and then concatenate with the upsampled high-resolution depth features along the channel axis: where are the features extracted from the high-resolution RGB guidance image, are the fused features, and [·;·] represents the concatenation operation along the channel axis; Module 4: Map the fused features back to the low-resolution spatial domain through the downward projection unit: Among them, (·)↓ s represents the down-projection operation with a scale factor of s.

Citation Information

Patent Citations

  • Depth map super-resolution reconstruction method based on convolutional neural networks

    CN107358576A

  • Depth image super-resolution reconstruction method based on hierarchical feature feedback fusion

    CN111882485A