Depth image arbitrary resolution reconstruction method guided by color image
By designing a neural network model containing multiple deep learning modules, the problem of difficulty in realizing super-resolution reconstruction of arbitrary resolution depth maps in the prior art is solved, and efficient and flexible super-resolution reconstruction of depth maps is achieved.
Patent Information
- Application Number
- CN202510052786.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to realize the super-resolution reconstruction of depth maps for any resolution, and the deep neural network-based methods have shortcomings in efficiency and cost in the training and inference process.
A neural network model based on deep learning is designed, including step-by-step downsampling module, fusion module, output sampling module, cross-aggregation module and attention projection module. Through the combination of these modules, the depth map super-resolution reconstruction of arbitrary resolution targets is achieved.
This method can perform super-resolution inference reconstruction for depth map targets with different resolutions in one training, reducing the memory resource requirements of graphics card required for training and improving the efficiency and flexibility of the model.
Smart Images

Figure CN119941510A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of computer vision, digital image processing and digital signal processing, and relates to super-resolution reconstruction of depth images, and in particular to a method for reconstructing a depth map of a target with arbitrary resolution under the guidance of a color image. Background Art
[0002] Depth images are particularly important for scenes that need to perceive distance information and reflect three-dimensional structures. They are widely used in 3D reconstruction, semantic scene understanding, and autonomous driving. Depth image acquisition can be divided into passive and active methods based on the principle. The passive method relies on texture for stereo matching, and the calculation results are often wrong and missing. The active method is limited by the sensor capabilities and has problems such as low sampling resolution or missing noise. The above problems have hindered the efficient acquisition of high-resolution depth images and their application to tasks with high precision requirements.
[0003] Existing depth acquisition equipment is often equipped with a color image sensor to capture color images of the same scene. Thanks to more mature technology, the resolution of the acquired color images is often higher. Because of their similarity in scene structure with depth images, they can be used to guide the super-resolution reconstruction of depth maps. Methods for this task are also one of the research hotspots in the field of computer vision.
[0004] The current color image guided deep image super-resolution reconstruction methods can be roughly divided into two categories: traditional algorithms and learning-based methods. Traditional methods can be divided into two main research directions: filtering-based and optimization-based. Learning-based methods mainly use deep neural network modeling and directly learn complex nonlinear mapping relationships from data. The effect is more prominent and has become the current mainstream method. However, the training and reasoning involved in the current method model are basically only performed for fixed super-resolution multiples (such as 4 times), and super-resolution scenes with different multiples can only be trained separately multiple times. Due to the shortcomings of neural networks such as large number of parameters and high resource requirements for training and storage, the method has shortcomings in efficiency and cost in practical applications. Therefore, it is of great value to propose a color image-guided deep image super-resolution reconstruction method that can effectively target any target resolution. Summary of the invention
[0005] The purpose of the present invention is to fill the deficiencies in the prior art and provide an effective solution for super-resolution reconstruction of depth maps of arbitrary resolution. A neural network model based on deep learning and a training method thereof are proposed, which can perform arbitrary reconstruction according to a given coordinate position, thereby completing super-resolution reconstruction of depth maps of arbitrary resolution.
[0006] The reconstruction model designed by the present invention comprises: a step-by-step downsampling module, a fusion module, an output sampling module, a cross aggregation module and an attention projection module, and achieves deep super-resolution reconstruction at a given high-resolution sampling target position through the following steps:
[0007] Step 1: For the same scene, obtain high-resolution color images r that match each other Hi and the low-resolution depth image d Lo , and for the depth image d Lo Upsample to r Hi For the same resolution, we get d Hi The mutual matching refers to the corresponding relationship between the color image and the depth image collected by the same depth / color image sensor with the same viewing angle or a corresponding relationship, or multiple depth / color image sensors aligned with external parameters, while ensuring that the two have the same projection.
[0008] Step 2: r Hi and d Hi Input to the step-by-step downsampling module to obtain shallow features of color images and depth images at multiple resolutions; first pass through the linear layer U r and U d After being embedded into the feature maps r1 and d1 of the specified dimensions, they are then downsampled L-1 times in succession, recorded as and Get the shallow feature map r at each resolution l and d l , where l∈2,...,L; the above downsampling operation and It consists of a layer of convolution kernel size k and a step size of The convolutional layer of the multi-layer convolutional network is executed, and k is generally greater than or equal to 3. and At the same time, the feature dimension is doubled to ensure that information is not lost as the resolution decreases. k and the starting features r1 and d1 dimensions w can be used as pre-configured parameters and set according to the expected model size and required hardware.
[0009] Step 3: Use the fusion module to fusion the shallow feature maps of color images and depth images at various resolutions l and d l Fusion is performed to obtain a fused feature map. For the lth resolution, l∈1,...,L, the fusion operator is applied The operations completed are as follows: First, the shallow feature r is concatenated l and d l Stacked in the feature dimension and passed through the feature encoding network E in turn l, Self-Attention I l and the MLP layer F l , get the fusion feature map f l . Feature Encoding Network E l It can be any deep convolutional neural network, such as EDSR network. Self-attention operator I l It can be any Transformer-like network module with kernel integral, such as the Galerkin-style self-attention module.
[0010] Step 4: Use the output sampling module to obtain the sampling features representing different scales for the given coordinate x of the sampling target point at each resolution; for the lth resolution, l∈1,...,L, use the sampling operation based on the learnable method For the fusion feature map f l Sampling: First, the fusion features of the four nearest corner points of the sampling target point are collected and weighted according to the proportion of the diagonal area to the surrounding area. After being cascaded with the relative distance information of the sampling target point and the four corner points, they are input into the MLP layer G. l , to autonomously learn the combination between them and finally obtain the sampled feature map z l The coordinate x of the sampling target point is based on the normalized range of 0.0 to 1.0, and can be any number and any position within the range, thus giving the model the ability to reconstruct at any resolution other than the input guidance color image resolution. Output sampling module based on a learnable approach The specific operation method is as follows:
[0011]
[0012] in, are the coordinates of the four nearest corner points of the sampling target point, δ(x,x i )=xx i Get relative distance information, is a weighted operation using the diagonal area as the weight. For x i The diagonal coordinates, φ(x,x i ) calculates the area of the rectangle formed by two points, and Concat is a cascade operation. The surrounding area refers to the rectangular area surrounded by the four nearest corner points of the sampling target point. The diagonal area refers to the area between the diagonal point of the rectangle with the nearest corner point in the surrounding area and the sampling target point, which is calculated by multiplying the horizontal coordinate difference and the vertical coordinate difference of the two points.
[0013] Step 5: Use the cross aggregation module to start from the resolution with the smallest label l, and use multiple aggregation operators to guide the enhancement and aggregation operations of the resolution sampling feature map step by step to obtain the aggregated feature map; where, for the lth aggregation operator, l∈2,...,L, the operation is completed as follows: Use the aggregated feature map m l-1 As a guide, the sampled feature map z l Input to the cross attention C l-1 Further enhance the features and combine the obtained feature map with m l-1 Perform stacking and cascading on the feature dimension to obtain the updated aggregate feature map m l , In particular, the starting fusion feature map m1 is z1. Aggregation operator The cross attention C used l-1 Specifically, it is a Galerkin-style cross-attention module, which operates as follows:
[0014] C l-1 (z l ,m l-1 )=P(O( <Ln(z l W Q ), <Ln(m l-1 W K ) T ,Ln(m l-1 W V )>> / N)+z l )+z l
[0015] Among them, W Q , W K , W V Linear operations project features to the same feature dimension, Ln(·) is layer normalization, <·,·> represents inner product operation, N is the number of feature points involved in the operation, O(·) is the MLP layer including the activation layer, and the dimension is converted to the same as z l Similarly, P(·) is a linear layer.
[0016] Step 6: Use the attention projection module to the final aggregate feature map m L Apply global attention and finally project the target depth image, that is, the super-resolved depth image; among them, for the final aggregated feature map m L , apply the attention projection layer Complete the projection operation: m L First, the self-attention operator S is used to integrate global information in the feature dimension, and then the depth value of the sampled target position is projected through the MLP layer D. At the same time, for the input low-resolution depth map d LoUse traditional sampling, such as bilinear, to process the jump connection branch, and add the output to the output of the main part of the model to make the main network of the model learn the residual, which can reduce the difficulty of model training and promote convergence.
[0017] In a complete training of the model network involved in the present invention for all training data, each training sample uses a different, random, and continuous distribution range of resolutions during data pre-processing (further expressed as a resolution of 100% of the input low-resolution depth image d Lo The target depth images (distributed in the resolution range corresponding to magnifications of 1.0 to 16.0) are used as true value supervision, so that the method can perform super-resolution reasoning and reconstruction of depth map targets with different resolutions after a single training.
[0018] At the same time, for a single target depth image, an arbitrary number of points (further expressed as far less than the total number of pixels in the actual target high-resolution depth image) are randomly used as true value supervision for training, thereby reducing the network's computing power requirements and graphics card memory requirements, allowing the network to be trained with limited resources.
[0019] Compared with the prior art, the method of the present invention has at least the following beneficial effects:
[0020] (1) The model of this method only needs to be trained once, and it can reconstruct deep image super-resolution targets of any resolution.
[0021] (2) During the training process, the model of this method only needs to select a small number of depth image points as supervision, which greatly reduces the graphics card memory resources required for training, thereby reducing the requirements for model training.
[0022] (3) This model method uses the self-attention mechanism to effectively integrate global information to facilitate better reconstruction.
[0023] (4) This method uses multi-resolution feature extraction and aggregated feature-guided refinement design to help the model better cope with scenes with larger super-resolution spans. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is a schematic diagram of the overall module division and the structural flow of the model network of an embodiment of the present invention;
[0025] Figure 2 is the fusion operator of the lth ('∈1, ..., L) resolution in the fusion module of the embodiment of the present invention The structural flow diagram of the model network;
[0026] Figure 3The sampling operation based on a learnable manner at the lth (l∈1, ..., L) resolution in the output sampling module of the embodiment of the present invention is The structural flow diagram of the model network;
[0027] Figure 4 is the lth (l∈1, ..., L-1) aggregation operator in the cross aggregation module of the embodiment of the present invention The structural flow diagram of the model network;
[0028] Figure 5 The attention projection layer of the attention projection module of the embodiment of the present invention The structural flow diagram of the model network;
[0029] Figure 6 The aggregation operator of the embodiment of the present invention The Galerkin-style cross-attention module C used to guide the refinement l The structural flow diagram of the model network;
[0030] Figure 7 An example of super-resolution reconstruction of a depth map of arbitrary resolution under the guidance of a color image is provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] The examples are as follows:
[0033] The present invention discloses a method for reconstructing a depth image with arbitrary resolution guided by a color image. This embodiment takes the training and testing on the Middlebury public depth image dataset as an example to illustrate the implementation method of the present invention with specific scene data. The steps are as follows:
[0034] Step 1: Use the Middlebury dataset training set to train high-resolution color images r that match each other Hi and a high-resolution depth image d gt As a training data pair; use a randomly obtained downsampling factor in the range of ×1.0 to ×16.0 magnification to perform bilinear downsampling to obtain a low-resolution depth image d Lo , and r Hi Together as input to train the model network, where the depth image d is pre- Lo Upsample to r Hi The same resolution size gets d Hi . Figure 1The overall structure and process of the model network of the present invention are shown. Specifically, in one forward pass of the model, a small number of points at a certain position will be super-reconstructed. The various links are as follows:
[0035] S1-1. Hi and d Hi Input to the step-by-step downsampling module to obtain the shallow features of color images and depth images at multiple resolutions. It first passes through the linear layer U r and U d After being embedded into the feature maps r1 and d1 of the specified dimensions, continuous downsampling operations are performed L-1 times respectively. and Get the shallow feature map r at each resolution l and d l , where L∈2,...,L. In this embodiment, L is 3, and the subsequent steps are the same.
[0036] S1-2. Use the fusion module to fuse the color image and depth image at each resolution obtained in S1-1 to obtain a fused feature map. Apply the fusion operator to the lth (l∈1,...,L) resolution Complete the following operations: First, the shallow feature map r is concatenated l and d l Stacked in the feature dimension and passed through the feature encoding network E in turn l , Self-Attention I l and the MLP layer F l , get the fusion feature map f l . Figure 2 Demonstrates the fusion operator The overall structure and process, including E l Using an EDSR network with 16 blocks, self-attention I l Use the Galerkin form of self-attention module.
[0037] S1-3. Randomly select a small number of sampling positions x, and use the output sampling module to obtain the sampling features for x at each resolution, that is, use a sampling operation based on a learnable method The fusion feature f of the lth resolution output in S1-2 l Sampling: First, the fusion features of the four nearest corner points of the sampling target point are collected and weighted according to the proportion of the diagonal area to the surrounding area. After being cascaded with the relative distance information of the sampling target point and the four corner points, they are input into the MLP layer G. l , autonomous learning combination to obtain output sampling feature z l . Figure 3 Demonstrates sampling operations in a learnable manner Overall structure and process.
[0038] S1-4. Use the cross aggregation module to guide the aggregation of the sampling features of each resolution obtained in S1-3 step by step starting from the resolution with the smaller label, so as to obtain the aggregated features: For the lth (here only for l∈2,...,L) aggregation operator Complete the following operations: Sample feature z l Input to the cross attention C l-1 , using the aggregate feature m l-1 Further guide the refinement and compare the obtained features with m l-1 Cascade and stack in the feature dimension to obtain the updated aggregate feature m l , in particular, the starting fusion feature m1 is z1. Figure 4 Demonstrates sampling operations in a learnable manner The overall structure and process. The cross attention C l-1 Using Galerkin-style cross-attention modules, Figure 6 Its overall structure and process are demonstrated.
[0039] S1-5. Use the attention projection module to the final aggregate feature m in S1-4 L Apply global attention and finally project the super-resolved depth image of the target The attention projection module mainly uses the attention projection layer Complete the projection operation: m L First, the self-attention operator S is used to integrate global information in the feature dimension, and then the depth residual value of the target sampling position is projected through the MLP layer D At the same time, for the input low-resolution depth map d Lo Use bilinear sampling as the skip connection branch and compare its output with Add together to get the final target depth value d t . Figure 4 Demonstrates sampling operations in a learnable manner Overall structure and process.
[0040] Step 2: Use the sampling position x obtained in S1-3 to calculate d gt The values at the corresponding positions are collected, and MAE is used as the loss function to supervise the update of model parameters in back propagation. The Adam optimizer used in back propagation corrects the gradient according to the learning rate, and uses the learning rate adjustment strategy of warmup+cosine annealing.
[0041] Step 3: Alternate forward and backward propagation from step 1 to step 2 to train the model for 600 rounds to complete model training.
[0042] Step 4: Generate gridded position coordinates according to the target resolution, use the model trained in step 3, refer to the forward process of step 1 for prediction, and finally obtain the super-resolution result of the depth map of the target resolution. Figure 6 The super-resolution reconstruction effect of depth map for targets with different magnifications on an image in the Middlebury test set is demonstrated.
Claims
1. A color image guided depth image arbitrary resolution reconstruction method, comprising: A low-resolution depth image and a relatively high-resolution color image of the same scene are input into a pre-trained color image guided depth map super-resolution reconstruction model, and the additional information extracted from the high-resolution color image is used as a guide for depth map reconstruction to achieve a more accurate low-resolution depth map super-resolution goal. The method is characterized in that the reconstruction model includes: a step-by-step downsampling module, a fusion module, an output sampling module, a cross aggregation module, and an attention projection module; This reconstruction method includes the following steps: Step 1: For the same scene, obtain high-resolution color images r that match each other Hi and the low-resolution depth image d Lo , and for the depth image d Lo Upsample to r Hi For the same resolution, we get d Hi ; The mutual matching refers to the corresponding relationship between the color image and the depth image collected by the same depth / color image sensor with the same viewing angle or a corresponding relationship or multiple depth / color image sensors aligned with external parameters, while ensuring that the two have the same projection; Step 2: r Hi and d Hi Input to the step-by-step downsampling module to obtain shallow features of color images and depth images at multiple resolutions; first pass through the linear layer U r and U d After being embedded into the feature maps r1 and d1 of the specified dimensions, they are then downsampled L-1 times in succession, recorded as and Get the shallow feature map r at each resolution l and d l , where l∈2,...,L; Step 3: Use the fusion module to fusion the shallow feature maps of color images and depth images at various resolutions l and d l Fusion is performed to obtain a fused feature map; for the lth resolution, l∈1,...,L, the fusion operator is applied The completed operations are as follows: First, the shallow feature map r is concatenated l and d l Stacked in the feature dimension and passed through the feature encoding network E in turn l , Self-Attention I l and the MLP layer F l , get the fusion feature map f l ; Step 4: Use the output sampling module to obtain the sampling features representing different scales for the given coordinate x of the sampling target point at each resolution; for the lth resolution, l∈1,...,L, use the sampling operation based on the learnable method For the fusion feature map f l Sampling: First, the fusion features of the four nearest corner points of the sampling target point are collected and weighted according to the proportion of the diagonal area to the surrounding area. After being cascaded with the relative distance information of the sampling target point and the four corner points, they are input into the MLP layer G. l , in order to autonomously learn the combination between them and finally obtain the sampling feature map zl; the surrounding area refers to the rectangular area surrounded by the four nearest corner points of the sampling target point; the diagonal area refers to the area between the diagonal point of the rectangle with the nearest corner point in the surrounding area and the sampling target point, which is calculated by multiplying the horizontal coordinate difference and the vertical coordinate difference of the two points; Step 5: Use the cross aggregation module to start from the resolution with the smallest label l, and use multiple aggregation operators to guide the enhancement and aggregation operations of the resolution sampling feature map step by step to obtain the aggregated feature map; among them, for the lth aggregation operator l∈2,...,L, the operation is completed as follows: using the aggregated feature map m l-1 As a guide, the sampled feature map z l Input to the cross attention C l-1 Further enhance the features and combine the obtained feature map with m l-1 Perform stacking and cascading on the feature dimension to obtain the updated aggregate feature map m l ,In particular, the starting fusion feature map m1 is z1; Step 6: Use the attention projection module to the final aggregated feature map m L Apply global attention and finally project the target depth image, that is, the super-resolved depth image; for the final aggregate feature map mL, apply the attention projection layer Complete the projection operation: m L First, the global information is integrated in the feature dimension through the self-attention operator S, and then the depth value of the sampled target position is projected through the MLP layer D.
2. The method for reconstructing a depth image with arbitrary resolution guided by a color image according to claim 1, characterized in that: Downsampling in step 2 and It consists of a layer of convolution kernel size k and a step size of The convolutional layer of the multi-layer convolutional network is executed, and k is generally greater than or equal to 3. and At the same time, the feature dimension is doubled to ensure that information is not lost as the resolution decreases. k and the starting features r1 and d1 dimensions w can be used as pre-configured parameters and set according to the expected model size and required hardware.
3. The method for color image guided depth image arbitrary resolution reconstruction as claimed in claim 1, characterized in that: In step 3, the feature encoding network E l It can be any deep convolutional neural network, such as EDSR network.
4. The method for reconstructing a depth image with arbitrary resolution guided by a color image as claimed in claim 1, characterized in that: Self-attention operator I in step 3 or step 6 l and S can be any Transformer-like network module containing kernel integrals, such as a Galerkin-style self-attention module.
5. The method for color image guided depth image arbitrary resolution reconstruction as claimed in claim 1, characterized in that: The coordinate x of the sampled target point in step 4 is based on the normalized range of 0.0 to 1.0, and can be any number and any position within the range, thus giving the model the ability to reconstruct at any resolution different from the input guided color image resolution.
6. The method for color image guided depth image arbitrary resolution reconstruction as claimed in claim 1, characterized in that: Output sampling module based on learnable method in step 4 The specific operation method is as follows: in, are the coordinates of the four nearest corner points of the sampling target point, δ(x, x i )=xx i Get relative distance information, is a weighted operation using the diagonal area as the weight. For x i The diagonal coordinates, φ(x, x i ) calculates the area of the rectangle formed by two points, and Concat is a cascade operation.
7. The method for reconstructing a depth image with arbitrary resolution guided by a color image as claimed in claim 1, characterized in that: Aggregation operator in step 5 The cross attention C used l-1 Specifically, it is a Galerkin-style cross-attention module, which operates as follows: C l-1 (z l ,m l-1 )=P(O(<Ln(z l W Q ),<Ln(m l-1 W K ) T ,Ln(m l-1 W V )>> / N)+z l )+z l Among them, W Q , W K , W V Linear operations project features to the same feature dimension, Ln(·) is layer normalization, <·, ·> represents inner product operations, N is the number of feature points involved in the operation, O(·) is the MLP layer including the activation layer, and the dimension is converted to the same as z l Similarly, P(·) is a linear layer.
8. The method for reconstructing a depth image with arbitrary resolution guided by a color image as claimed in claim 1, characterized in that: The main part of the model learns the residual, that is, for the input low-resolution depth map d Lo Use traditional sampling, such as bilinear interpolation, to process the branches as skip connections, and add the output to the output of the main part of the model to reduce the difficulty of model training and promote convergence.
9. The method according to any one of claims 1 to 8, characterized in that: In a complete training of the model network of the method for all training data, each training sample uses a different, random, and continuous distribution range of resolutions in the data pre-processing (further manifested as a resolution of d compared to the input low-resolution depth image d Lo The target depth images (distributed in the resolution range corresponding to magnifications of 1.0 to 16.0) are used as true value supervision, so that the method can perform super-resolution reasoning and reconstruction of depth map targets with different resolutions after a single training.
10. The method according to any one of claims 1 to 9, characterized in that: When training the model network of the method, for a single target depth image, the random sampling operation uses an arbitrary number of points (further manifested as far less than the total number of pixels in the actual target high-resolution depth image) as true value supervision for training, thereby reducing the network's computing power requirements and graphics card memory requirements, allowing the network to be trained with limited resources.