A deep image restoration system and method based on gated cyclic feature fusion
By adopting the gated cyclic feature fusion module and the spatial propagation module in the depth image repair system, the problem of neglected correlation between features in the prior art is solved, and a higher quality dense depth image repair is achieved.
Patent Information
- Application Number
- CN202210170142.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-23
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-02-23
AI Technical Summary
Existing depth image repair methods ignore the correlation between features when calculating the affinity matrix, resulting in low quality of dense depth image repair.
A deep image repair system based on gated loop feature fusion is adopted, and multi-scale feature encoding, decoding and iterative diffusion are realized through shallow feature extraction module, gated loop feature fusion module and spatial propagation module to improve repair performance.
By building a dual network structure with rough repair and fine repair, the system can learn complex mapping relationships more effectively, significantly improving the repair quality of dense depth images.
Smart Images

Figure CN114529793B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a deep image restoration system and method based on gated cyclic feature fusion. Background Art
[0002] In recent years, deep learning frameworks have been widely used in the field of deep image restoration. Some methods incorporate surface normal information into the deep restoration network, some methods stack sparse depth images and color image features of the same scene and pass them into the self-supervised network, and use photometric consistency loss to supervise the restoration process; some methods combine depth and color information in the normalized network to complete deep restoration. Among these methods, multi-level feature fusion or multimodal feature fusion is often completed by simple pixel-by-pixel addition or feature stacking.
[0003] In addition, some of the latest deep image restoration methods use a coarse-fine network architecture, that is, a coarse restoration network combined with a fine restoration network. Among them, in the fine restoration network, some researchers use the convolutional spatial propagation network model (CSPN), which iteratively diffuses adjacent points under the guidance of the affinity matrix to correct the depth result. Subsequently, these researchers proposed CSPN++, which improves the restoration performance by adaptively learning the convolution kernel size and the number of diffusion iterations. Some researchers proposed a non-local spatial propagation network model (NLSPN), which uses the affinity matrix between non-local neighborhood points to guide the depth correction during the iterative diffusion process. The affinity matrix determines the speed and direction of spatial propagation, and its accuracy will greatly affect the depth correction performance of the fine restoration network. However, these methods currently only use a simple convolution layer to calculate the affinity matrix, ignoring the study of the correlation between features, which reduces the restoration quality of dense deep images. Summary of the invention
[0004] The purpose of the present invention is to provide a deep image restoration system and method based on gated cyclic feature fusion, so as to achieve the technical effect of improving the quality of deep image restoration.
[0005] In a first aspect, the present invention provides a deep image restoration system based on gated cyclic feature fusion, comprising: a shallow feature extraction module, a gated cyclic feature fusion module and a spatial propagation module;
[0006] The shallow feature extraction module is used to extract shallow features from the input color image and the sparse depth image, and stack the extracted shallow features into a unified shallow feature;
[0007] The gated cyclic feature fusion module includes an encoder and a decoder; the encoder includes S scaled encoding units connected in sequence; the encoding unit includes R residual blocks connected in sequence; the decoder includes S decoding units connected in sequence and symmetrically arranged with the encoding unit; except that the first decoding unit corresponding to the first encoding unit includes a gated cyclic unit and a convolution layer connected to the corresponding gated cyclic unit, the remaining decoding units all include a gated cyclic unit and an upsampling layer connected to the corresponding gated cyclic unit; wherein S and R are both integers greater than 1;
[0008] The encoder is used to encode at multiple scales according to the unified shallow features to obtain low-level features required for feature fusion in each decoding unit; the decoder is used to decode in sequence from the Sth decoding unit through the acquired initial high-level features to obtain a roughly restored first dense depth image, and at the same time output the high-level features processed by the gated recurrent unit in the first decoding unit;
[0009] The spatial propagation module is used to perform depth image correction by iterative updating according to the sparse depth image, the first dense depth image and the high-level features to obtain a finely repaired second dense depth image.
[0010] Furthermore, the last residual block of the first S-1 coding units in the encoder is down-sampled.
[0011] Furthermore, the spatial propagation module includes a dimension-by-dimension attention module, a convolution layer and a spatial propagation network; the dimension-by-dimension attention module includes a feature channel attention unit, a feature height attention unit, a feature width attention unit and a Concat layer; the feature channel attention unit is used to analyze the channel attention weight of the high-level feature, and multiply the channel attention weight with the high-level feature and output it; the feature height attention unit is used to analyze the height attention weight of the high-level feature, and multiply the height attention weight with the high-level feature and output it; the feature width attention unit is used to analyze the width attention weight of the high-level feature, and multiply the width attention weight with the high-level feature and output it; the Concat layer in the dimension-by-dimension attention module is used to stack the output results of the three attention units into a unified feature; the convolution layer in the spatial propagation module obtains the corresponding affinity matrix according to the unified feature analysis; the spatial propagation network takes the sparse depth image and the first dense depth image as input, and guides the iterative diffusion and update between neighborhood pixels through the affinity matrix to obtain the second dense depth image.
[0012] Furthermore, the feature channel attention unit includes a global pooling layer, a "1×1 convolution layer-ReLU layer-1×1 convolution layer-Sigmoid layer" combination structure and a multiplier; the feature height attention unit and the above-mentioned feature width attention unit both include a global pooling layer, a "Resize layer-1×1 convolution layer-ReLU layer-1×1 convolution layer-Sigmoid layer-Resize layer" combination structure and a multiplier; the high-level features first obtain corresponding one-dimensional statistical signals through the global pooling layers in the feature channel attention unit, the feature height attention unit and the feature width attention unit respectively; secondly, the corresponding attention weights are obtained through the corresponding combination structure processing; then, the corresponding attention weights are multiplied pixel by pixel with the high-level features through the corresponding multipliers; finally, the outputs of the three attention units are stacked into a unified feature through the Concat layer.
[0013] Furthermore, the shallow feature extraction module includes 2 n×n convolutional layers and one Concat layer; one n×n convolutional layer is used to extract shallow color features from the input color image, and one n×n convolutional layer is used to extract shallow sparse depth features from the input sparse depth image; the Concat layer is used to stack the shallow color features and shallow sparse depth features into a unified shallow feature.
[0014] In a second aspect, the present invention provides a deep image restoration method based on gated cyclic feature fusion, which is applied to the above-mentioned deep image restoration system based on gated cyclic feature fusion, comprising:
[0015] S1. Get the deep image restoration training set {I i , X i , Y i gt}, where i represents a variable, and 1≤i≤N, N represents the number of images of each type; X represents a sparse depth image; I represents a color image of the same scene; Y gt represents the corresponding real dense depth image;
[0016] S2. extract shallow features from the input color image and the sparse depth image through a shallow feature extraction module, and stack the extracted shallow features into a unified shallow feature;
[0017] S3. Processing the unified shallow features through the gated cyclic feature fusion module to obtain a roughly restored first dense depth image, and outputting the high-level features processed by the gated cyclic unit in the first decoding unit;
[0018] S4. Performing depth image correction by iteratively updating the spatial propagation module according to the sparse depth image, the first dense depth image and the high-level features to obtain a finely repaired second dense depth image.
[0019] Further, the method further includes: S5. using the average L2 error between the N finely restored second dense depth images and the corresponding true dense depth images as a loss function to optimize the parameters of the depth image restoration system, wherein the loss function is:
[0020]
[0021] In the above formula, Θ represents the parameters of the whole system; i represents the variable, and 1≤i≤N, N represents the number of each type of image; Ⅱ(·) is the marker function; Y gt represents the corresponding real dense depth image; Y represents the finely restored second dense depth image; ⊙ represents pixel-by-pixel multiplication.
[0022] The beneficial effects that can be achieved by the present invention are: the deep image restoration system and method based on gated cyclic feature fusion provided by the present invention constitute a dual network structure of coarse restoration and fine restoration through a gated cyclic feature fusion module. Compared with the existing technology, it has a stronger ability to learn complex mapping relationships and can restore higher quality dense depth images. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments of the present invention are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 A schematic diagram of the topological structure of a deep image restoration system based on gated cyclic feature fusion provided by an embodiment of the present invention;
[0025] Figure 2 A schematic diagram of the topological structure of a gated cyclic feature fusion module provided in an embodiment of the present invention;
[0026] Figure 3 A schematic diagram of a gated recurrent unit provided in an embodiment of the present invention;
[0027] Figure 4 A schematic diagram of the topological structure of a spatial propagation module provided in an embodiment of the present invention;
[0028] Figure 5A schematic diagram of the topological structure of a dimension-by-dimension attention module provided in an embodiment of the present invention;
[0029] Figure 6 A schematic flow chart of a deep image restoration method based on gated cyclic feature fusion provided in an embodiment of the present invention.
[0030] Icon: 10-deep image restoration system; 100-shallow feature extraction module; 200-gated recurrent feature fusion module; 210-encoder; 220-decoder; 221-gated recurrent unit; 300-spatial propagation module; 310-dimension-by-dimension attention module. DETAILED DESCRIPTION
[0031] The technical solutions in the embodiments of the present invention will be described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0032] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0033] Please see Figure 1 , Figure 2 and Figure 3 , Figure 1 A schematic diagram of the topological structure of a deep image restoration system based on gated cyclic feature fusion provided by an embodiment of the present invention; Figure 2 A schematic diagram of the topological structure of a gated cyclic feature fusion module provided in an embodiment of the present invention; Figure 3 A schematic diagram of a gated recurrent unit according to an embodiment of the present invention.
[0034] In one implementation, an embodiment of the present invention provides a deep image restoration system 10 based on gated cyclic feature fusion, and the deep image restoration system 10 includes: a shallow feature extraction module 100, a gated cyclic feature fusion module 200 and a spatial propagation module 300; the shallow feature extraction module 100 is used to extract shallow features from the input color image and the sparse depth image, and stack the extracted shallow features into a unified shallow feature F0; the gated cyclic feature fusion module 200 includes an encoder 210 and a decoder 220; the encoder 210 includes S scale encoding units connected in sequence; the encoding unit includes R residual blocks connected in sequence; the decoder 220 includes S decoding units connected in sequence symmetrically arranged with the encoding unit; except that the first decoding unit corresponding to the first encoding unit includes a gated recurrent unit 221 (gated recurrent unit, GRU) and a convolution layer (CONV layer) connected to the corresponding gated recurrent unit 221, the remaining decoding units all include a gated recurrent unit and an upsampling layer connected to the corresponding gated recurrent unit; wherein S and R are both integers greater than 1; the encoder 210 is used to encode at multiple scales according to the unified shallow feature F0, and obtain the low-level features required for the fusion of the gated recurrent unit features in each decoding unit; the decoder 220 is used to decode in sequence from the Sth decoding unit through the acquired initial high-level features, to obtain a roughly repaired first dense depth image Y0, and at the same time output the high-level feature Q1 obtained by processing the gated recurrent unit in the first decoding unit; the spatial propagation module 300 is used to perform depth image correction by iterative updating according to the sparse depth image X, the first dense depth image Y0 and the high-level feature Q1, to obtain a finely repaired second dense depth image Y.
[0035] Specifically, Figure 2 As shown, the encoder includes S scale coding units from left to right, each coding unit includes R sequentially connected residual blocks, and the unified shallow feature F0 is encoded from the first coding unit through the S scale coding units in sequence; the decoder includes S sequentially connected decoding units symmetrically arranged with the coding units; except that the first decoding unit corresponding to the first coding unit includes a gated recurrent unit 221 (gated recurrent unit, GRU) and a convolution layer (CONV layer) connected to the corresponding gated recurrent unit 221, the remaining decoding units (i.e., the second to S decoding units) all include a gated recurrent unit and an upsampling layer (CONV layer) connected to the corresponding gated recurrent unit Figure 2 The UPSAMPLE layer in is the upsampling layer).
[0036] In the above implementation process, the shallow feature extraction module 100 first extracts shallow features from the input color image and sparse depth image, and stacks the extracted shallow features into a unified shallow feature; then, the U network composed of the encoder 210 and the decoder 220 in the gated cyclic feature fusion module 200 performs multi-scale encoding and decoding according to the unified shallow feature, and obtains a roughly repaired first dense depth image and a high-level feature processed by the gated cyclic unit in the first decoding unit; finally, the spatial propagation module 300 corrects the depth image by iterative updating according to the sparse depth image, the first dense depth image and the high-level feature, and obtains a finely repaired second dense depth image. The gated cyclic feature fusion module 200 forms a dual network structure of rough repair and fine repair, which has a stronger ability to learn complex mapping relationships compared to the prior art and can repair higher quality dense depth images.
[0037] Specifically, the processing flow of the encoder 210 is as follows: the unified shallow feature F0 is passed into the encoder 210, and passes through S encoding scales in sequence; wherein each scale is sequentially residually learned by R residual blocks, and the Rth residual block also needs to downsample the feature size to expand the perception domain. The low-level feature extracted by the rth residual block (1≤r≤R) in the sth scale (1≤s≤S) of the encoder 210 is represented as F s,r ; Then the output of the Rth residual block is F s,R , F s,R It can be expressed as:
[0038] F s,R =↓f s,R (f s,R-1 (…f s,1 (F s,0 )))
[0039] In the above formula, F s,0 =F s-1,R Corresponds to the output of the encoder at the s-1th scale; f s,r is the residual learning function of the rth residual block of the sth scale of the encoder; ↓ represents the downsampling operation.
[0040] Specifically, each stage of the gated recurrent unit 221 contains three convolutional layers, two Sigmoid (σ) layers, one tanh layer, three pixel-by-pixel multipliers (⊙) and one pixel-by-pixel adder (⊕), which together constitute a reset gate and an update gate; the reset gate determines which information of the previous hidden state will be stored and which will be forgotten in the current stage; the update gate determines which new information will be added to the current hidden state.
[0041] The processing flow of decoder 220 is as follows: each scale is subjected to multi-level feature fusion by the corresponding gated recurrent unit; the first S-1 scales are subjected to the upsampling layer ( Figure 2 The UPSAMPLE layer in the UPSAMPLE layer is used to upsample the feature size, and the decoding unit corresponding to the encoding unit of the first scale is sampled by the convolution layer ( Figure 2 The CONV layer in the decoder 220 reconstructs a roughly repaired dense depth image Y0. Taking the decoder 220 scale s as an example, the multi-level features include the initial high-level features Q transmitted from the decoder 220 at the s+1th scale s+1,↑ ( Figure 2 is None) and the low-level features F passed by the encoder scale s s,0 , F s,1 , ..., F s,R-1 ; Then the output of decoder 220 scale s is:
[0042] Q s,↑ =↑Q s =↑f GRFB (F s,0 ,F s,1 ,…,F s,R-1 ,Q s+1,↑ )
[0043] In the above formula, f GRFB represents the function of the gated recurrent unit; ↑ represents the upsampling function of the upsampling layer; Q s,↑ Represents the high-level features of the decoder output at the s-th scale.
[0044] The gated recurrent unit in the decoder scale s (i.e., the gated recurrent unit S) can be expanded into R stages, corresponding to the R hidden states h r , the high-level feature Q passed by the decoder at the s+1th scale s+1,↑ (None) as the initial hidden state h0, and the R low-level features (i.e., F s,0 , F s,1 , ..., F s,R-1 ) is passed to each stage in turn as the input of each stage, and the hidden state is updated stage by stage. Taking the rth stage as an example, its processing flow includes: resetting the gate, updating the gate, candidate hidden state calculation and hidden state calculation. r-1 and the input F of the current stage s,R-r After stacking, the incoming weight is W x The convolutional layer and Sigmoid (σ) layer are used to get the reset gate output x r ; The previous hidden state h r-1 and the input F of the current stage s,R-r After stacking, the incoming weight is W cThe convolution layer and Sigmoid (σ) layer are used to obtain the update gate output z r The expressions of reset gate and update gate are:
[0045] x r =σ(W x *[h r-1 ,F s,R-r ]),
[0046] z r =σ(W z *[h r-1 ,F s,R-r ]).
[0047] Then, x r With the previous hidden state h r-1 Multiply pixel by pixel to determine which information of the previous hidden state will be stored and which information will be forgotten. Then add it to the input feature F of the current stage s,R-r Stacking, and passing in weight W h The convolutional layer and tanh layer are used to obtain the candidate hidden state The expression is:
[0048]
[0049] Finally, the update gate outputs zr in the previous hidden state h r-1 and candidate hidden states Adaptively select the current hidden state h r , the expression is:
[0050]
[0051] In the above way, the gated recurrent unit can achieve effective fusion of multi-level features through stage-by-stage update of the hidden state.
[0052] In one implementation, the last residual block of the first S-1 coding units in the encoder is downsampled. In this way, the perception domain can be expanded.
[0053] In one embodiment, if Figure 1 As shown, the shallow feature extraction module 100 includes 2 n×n convolutional layers ( Figure 1 COMV layer in) and a Concat layer ( Figure 1 CAT layer in ); one n×n convolutional layer is used to extract shallow color features from the input color image, and one n×n convolutional layer is used to extract shallow sparse depth features from the input sparse depth image; the Concat layer is used to stack the shallow color features and the shallow sparse depth features into a unified shallow feature.
[0054] Please see Figure 4 and Figure 5 , Figure 4 A schematic diagram of the topological structure of a spatial propagation module provided in an embodiment of the present invention; Figure 5 A schematic diagram of the topological structure of the dimension-by-dimension attention module provided in an embodiment of the present invention.
[0055] In one embodiment, the spatial propagation module 300 includes a dimension-by-dimension attention module 310, a convolution layer and a spatial propagation network; the dimension-by-dimension attention module 310 includes a feature channel attention unit, a feature height attention unit, a feature width attention unit and a Concat layer; the feature channel attention unit is used to analyze the channel attention weight of the high-level feature, and multiply the channel attention weight with the high-level feature and output it; the feature height attention unit is used to analyze the height attention weight of the high-level feature, and multiply the height attention weight with the high-level feature and output it; the feature width attention unit is used to analyze the width attention weight of the high-level feature, and multiply the width attention weight with the high-level feature and output it; the Concat layer in the dimension-by-dimension attention module 310 is used to stack the output results of the three attention units into a unified feature; the convolution layer in the spatial propagation module 300 obtains the corresponding affinity matrix based on the unified feature analysis; the spatial propagation network takes the sparse depth image and the first dense depth image as input, and guides the iterative diffusion and update between neighborhood pixels through the affinity matrix to obtain the second dense depth image.
[0056] In one embodiment, the feature channel attention unit includes a global pooling layer, a "1×1 convolution layer-ReLU layer-1×1 convolution layer-Sigmoid layer" combination structure and a multiplier; the feature height attention unit and the above-mentioned feature width attention unit both include a global pooling layer, a "Resize layer-1×1 convolution layer-ReLU layer-1×1 convolution layer-Sigmoid layer-Resize layer" combination structure and a multiplier; the high-level features first obtain the corresponding one-dimensional statistical signals through the global pooling layers in the feature channel attention unit, the feature height attention unit and the feature width attention unit respectively; secondly, the corresponding attention weights are obtained through the corresponding combination structure processing; then, the corresponding attention weights are multiplied pixel by pixel with the high-level features through the corresponding multipliers; finally, the outputs of the three attention units are stacked into a unified feature through the Concat layer. In the above implementation process, the height or width of the one-dimensional statistical signal can be scaled to a fixed value through the first Resize layer, and the attention weight size can be adjusted to be consistent with the height and width of the feature Q through the second Resize layer.
[0057] Specifically, the processing flow of the spatial propagation module 300 is as follows: the high-level feature Q output by the gated recurrent feature fusion module 200 is passed to the dimension-by-dimension attention module 310, the dependency of the features in each dimension is learned, and attention weights are generated based on these relationships, and multiplied with the dimension-by-dimension weights to achieve adaptive adjustment of Q; the adjusted Q is passed to the CONV layer to calculate the affinity matrix w; the affinity matrix w, the sparse depth image X and the roughly repaired first dense depth image Y0 are passed to the spatial propagation network, and the affinity matrix guides the iterative diffusion and update between adjacent pixels in Y0, thereby obtaining a finely repaired second dense depth image Y. In an embodiment of the present invention, Figure 2 Q1 in is the Q in the above process.
[0058] The specific processing flow of the space propagation network is: Let Y0 = (y m,n )∈R H×W ,y m,n Represents the pixel value at position (m, n) in Y0, y m,n At the tth iteration, the affinity matrix can be used to represent the neighborhood set N m,n Updated to:
[0059]
[0060] Where (m, n) and (i, j) represent the positions of the reference point and the neighboring point respectively. The affinity value between (m, n) and (i, j) is It is used as a weight to control the speed of propagation of the depth value on the neighborhood (i, j) to the point (m, n). In order to ensure the stability of propagation, the affinity values in the neighborhood set need to be normalized in absolute value in advance. The weight of the reference point is:
[0061]
[0062] In addition, the spatial propagation network also needs to take a permutation operation at each iteration to retain the valid pixels in the sparse depth image X. The permutation operation can be expressed as:
[0063]
[0064] If X m,n is a valid pixel, then Replace with X m,n After T iterations, the depth image correction function is completed and a finely repaired second dense depth image Y is obtained.
[0065] Please see Figure 6 , Figure 6 A schematic flow chart of a deep image restoration method based on gated cyclic feature fusion provided in an embodiment of the present invention.
[0066] In one implementation, the embodiment of the present invention further provides a deep image restoration method based on gated cyclic feature fusion applied to the above-mentioned deep image restoration system 10, the specific contents of which are described as follows.
[0067] S1. Get the deep image restoration training set {I i , X i , Y i gt}, where i represents a variable, and 1≤i≤N, N represents the number of images of each type; X represents a sparse depth image; I represents a color image of the same scene; Y gt represents the corresponding real dense depth image.
[0068] S2. Extract shallow features from the input color image and sparse depth image through the shallow feature extraction module, and stack the extracted shallow features into a unified shallow feature.
[0069] Specifically, the expression is as follows:
[0070] F0=f SF (X,I)
[0071] Among them, F0 represents the unified shallow feature formed by stacking shallow color features and shallow sparse depth features, and f SF Represents the performance function of the shallow feature extraction module 100.
[0072] S3. Processing is performed according to the unified shallow features through the gated cyclic feature fusion module to obtain a roughly restored first dense depth image, and at the same time, high-level features processed by the gated cyclic unit in the first decoding unit are output.
[0073] Specifically, the expression is as follows:
[0074] (Y0,Q1)=f U (F0)
[0075] Among them, f U represents the performance function of the gated cyclic feature fusion module 200, Q1 represents the high-level features, and Y0 represents the roughly restored first dense depth image.
[0076] S4. The spatial propagation module performs depth image correction by iteratively updating the sparse depth image, the first dense depth image and the high-level features to obtain a finely repaired second dense depth image. Specifically, the expression is as follows:
[0077] Y=f CSPN (X,Y0,Q1)
[0078] Among them, f CSPNrepresents the performance function of the spatial propagation module 300, and Y represents the finely restored second dense depth image.
[0079] In one embodiment, the method further includes: S5. using the average L2 error between the N finely restored second dense depth images and the corresponding true dense depth images as a loss function to optimize the parameters of the depth image restoration system 10, wherein the loss function is:
[0080]
[0081] In the above formula, Θ represents the parameters of the entire network; i represents the variable, and 1≤i≤N, N represents the number of each type of image; Ⅱ(·) is the marker function; Y gt represents the corresponding real dense depth image; Y represents the finely restored second dense depth image; ⊙ represents pixel-by-pixel multiplication.
[0082] The parameters of the system are optimized by setting the loss function to further improve the dense depth image.
[0083] In order to better illustrate the effectiveness of the present invention, the embodiment of the present invention also uses a comparative experiment to demonstrate the deep image restoration effect, the specific content of which is as follows.
[0084] Dataset: The present invention uses the KITTI training set and the NYUv2 training set. KITTI is currently the world's largest computer vision algorithm evaluation dataset for autonomous driving scenarios. Its training set contains 85,898 depth images and corresponding color images. The present invention's test uses the KITTI validation set and the NYUv2 test set.
[0085] Evaluation indicators: For the KITTI dataset, the root mean square error (RMSE), mean absolute error (MAE), inverse depth root mean square error (iRMSE) and inverse depth mean absolute error (iMAE) are used to evaluate the model performance; for the NYUv2 dataset, the root mean square error (RMSE), the average absolute value of the relative error (REL) and δ i To evaluate the model performance, δ i Indicates that the relative error is less than a given threshold i(i∈{1.2 5 ,1.25 2 ,1.25 3}) in pixels.
[0086] The present invention uses the KITTI validation set and the NYUv2 test set to compare the model performance. The comparative experiment selects 12 representative deep image restoration methods to compare with the experimental results of the present invention. The experimental results are shown in Table 1 and Table 2. The 12 representative deep image restoration methods include:
[0087] Method 1 (SparseConvs): The method proposed by Uhrig et al., reference "J. Uhrig, N. Schneider, L. Schneider, U. Franke, T. Brox, and A. Geiger, Sparsity invariant cnns, in: Proc. Int. Conf. 3D Vis., 2017, pp. 11-20.".
[0088] Method 2 (Sparse2Dense): The method proposed by Ma et al., reference "F. Ma, GV Cavalheiro, and S. Karaman, Self-supervised sparse-to-dense: Self-supervised depth completion from lidar and monocular camera, in: Proc. IEEE Int. Conf. Robot. Autom., 2019, pp. 3288-3295.".
[0089] Method 3 (PwP): The method proposed by Xu et al., reference “Y.Xu, X.Zhu, J.Shi, G.Zhang, H.Bao, and H.Li, Depth completion from sparse LiDAR data with depth-normal constraints, in: Proc.IEEE Int.Conf.Comput.Vis., Oct.2019, pp.2811-2820.”.
[0090] Method 4 (NConv-CNN): The method proposed by Eldesokey et al., reference “A. Eldesokey, M. Felsberg, and F. S. Khan, Confidence Propagation through CNNs for Guided Sparse Depth Regression, IEEE Trans. Pattern Anal. Mach. Intell. 42 (10) (2020) 2423-2436.”.
[0091] Method 5 (MSG-CHN): The method proposed by Li et al., reference “A. Li, Z. Yuan, Y. Ling, W. Chi, and C. Zhang, A multi-scale guided cascade hourglass network for depth completion, in: Proc. IEEE Winter Conf. Appl. Comput. Vis., 2020, pp. 32-40.”
[0092] Method 6 (NLSPN): The method proposed by Park et al., reference "J.Park, K.Joo, Z.Hu, C.-K.Liu, and I.So Kweon, Non-local spatial propagation network for depth completion, in: Proc.European Conf. on Comput.Vis., 2020, pp.120-136.".
[0093] Method 7 (HMS-Net): The method proposed by Huang et al., reference "Z. Huang, J. Fan, S. Cheng, S. Yi, X. Wang, and H. Li, Hms-net: Hierarchical multi-scale sparsity-invariant network for sparse depth completion, IEEE Trans. on Image Process. 29 (2019) 3429-3441.".
[0094] Method 8 (GuideNet): The method proposed by Tang et al., reference "J. Tang, F. P. Tian, W. Feng, J. Li, and P. Tan, Learning guided convolutional network for depth completion, IEEE Trans. Image Process. 30 (2020) 1116-1129.".
[0095] Method 9 (ACMNet): The method proposed by Zhao et al., reference "S. Zhao, M. Gong, H. Fu, and D. Tao, Adaptive context-aware multi-modal network for depth completion, IEEE Trans. Image Process. 30 (2021) 5264-5276.".
[0096] Method 10 (S2D): The method proposed by Ma et al., reference “F. Ma and S. Karaman, Sparse-to-dense: Depth prediction from sparse depth samples and a single image, in: Proc. IEEE Int. Conf. Robot. Autom., May 2018, pp. 4796-4803.”
[0097] Method 11 (CSPN): The method proposed by Cheng et al., reference “X. Cheng, P. Wang, and R. Yang, Depth estimation via affinity learned with convolutional spatial propagation network, in: Proc. European Conf. on Comput. Vis., 2018, pp. 108-125.”.
[0098] Method 12 (DeepLiDAR): The method proposed by Qiu et al., reference “J. Qiu, Z. Cui, Y. Zhang, X. Zhang, S. Liu, B. Zeng, and M. Pollefeys, DeepLiDAR: Deep surface normal guided depth prediction for outdoor scene from sparse LiDAR data and single color image, in: Proc. IEEE Conf. Comput. Vis. Pattern Recognit., Jun. 2019, pp. 3313-3322.”.
[0099] It can be seen from Tables 1 and 2 (the best value and the second best value are indicated by black bold and underline respectively) that in most cases, the objective evaluation index value of the method provided by the present invention is optimal, and the restoration performance is significantly better than some currently representative deep image restoration methods.
[0100] Table 1 Comparison of objective evaluation indicators on the KITTI dataset
[0101]
[0102] Table 2 Comparison of objective evaluation indicators on the NYUv2 dataset (the number of effective pixels of sparse depth images is 200 and 500 respectively)
[0103]
[0104] To sum up, the embodiments of the present invention provide a deep image restoration system and method based on gated cyclic feature fusion, which forms a dual network structure of coarse restoration and fine restoration through the gated cyclic feature fusion module. Compared with the existing technology, it has stronger ability to learn complex mapping relationships and can restore higher quality dense depth images.
[0105] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A deep image restoration system based on gated cyclic feature fusion, characterized in that: include: Shallow feature extraction module, gated recurrent feature fusion module, and spatial propagation module; The shallow feature extraction module is used to extract shallow features from the input color image and the sparse depth image, and stack the extracted shallow features into a unified shallow feature; The gated cyclic feature fusion module includes an encoder and a decoder; the encoder includes S scaled encoding units connected in sequence; the encoding unit includes R residual blocks connected in sequence; the decoder includes S decoding units connected in sequence and symmetrically arranged with the encoding unit; except that the first decoding unit corresponding to the first encoding unit includes a gated cyclic unit and a convolution layer connected to the corresponding gated cyclic unit, the remaining decoding units all include a gated cyclic unit and an upsampling layer connected to the corresponding gated cyclic unit; wherein S and R are both integers greater than 1; The encoder is used to encode at multiple scales according to the unified shallow features to obtain low-level features required for fusion of gated recurrent unit features in each decoding unit; the decoder is used to decode sequentially from the Sth decoding unit through the acquired initial high-level features to obtain a roughly restored first dense depth image, and at the same time output the high-level features processed by the gated recurrent unit in the first decoding unit; The spatial propagation module is used to perform depth image correction by iterative updating according to the sparse depth image, the first dense depth image and the high-level features to obtain a finely repaired second dense depth image.
2. The deep image restoration system based on gated cycle feature fusion according to claim 1 is characterized in that: The last residual block of the first S-1 coding units in the encoder is down-sampled.
3. The deep image restoration system based on gated cycle feature fusion according to claim 1, characterized in that: The spatial propagation module includes a dimension-by-dimension attention module, a convolution layer and a spatial propagation network; the dimension-by-dimension attention module includes a feature channel attention unit, a feature height attention unit, a feature width attention unit and a Concat layer; the feature channel attention unit is used to analyze the channel attention weight of the high-level feature, and multiply the channel attention weight with the high-level feature and output it; the feature height attention unit is used to analyze the height attention weight of the high-level feature, and multiply the height attention weight with the high-level feature and output it; the feature width attention unit is used to analyze the width attention weight of the high-level feature, and multiply the width attention weight with the high-level feature and output it; the Concat layer in the dimension-by-dimension attention module is used to stack the output results of the three attention units into a unified feature; The convolution layer in the spatial propagation module obtains a corresponding affinity matrix according to the unified feature analysis; the spatial propagation network takes the sparse depth image and the first dense depth image as input, and guides iterative diffusion and updating between neighborhood pixels through the affinity matrix to obtain the second dense depth image.
4. The deep image restoration system based on gated cycle feature fusion according to claim 3 is characterized in that: The feature channel attention unit includes a global pooling layer, a "1×1 convolution layer-ReLU layer-1×1 convolution layer-Sigmoid layer" combination structure and a multiplier; the feature height attention unit and the feature width attention unit both include a global pooling layer, a "Resize layer-1×1 convolution layer-ReLU layer-1×1 convolution layer-Sigmoid layer-Resize layer" combination structure and a multiplier; the high-level features first obtain the corresponding one-dimensional statistical signals through the global pooling layers in the feature channel attention unit, the feature height attention unit and the feature width attention unit respectively; secondly, the corresponding attention weights are obtained through the corresponding combination structure processing; then, the corresponding attention weights are multiplied pixel by pixel with the high-level features through the corresponding multipliers; finally, the outputs of the three attention units are stacked into a unified feature through the Concat layer.
5. The deep image restoration system based on gated cycle feature fusion according to claim 1, characterized in that: The shallow feature extraction module includes two n×n convolutional layers and one Concat layer; one n×n convolutional layer is used to extract shallow color features from an input color image, and one n×n convolutional layer is used to extract shallow sparse depth features from an input sparse depth image; The Concat layer is used to stack the shallow color features and the shallow sparse depth features into a unified shallow feature.
6. A deep image restoration method based on gated cyclic feature fusion, applied to the deep image restoration system based on gated cyclic feature fusion according to any one of claims 1 to 5, characterized in that: include: S1. Get the deep image restoration training set {I i , X i , Y i gt }, where i represents a variable, and 1≤i≤N, N represents the number of images of each type; X represents a sparse depth image; I represents a color image of the same scene; Y gt represents the corresponding real dense depth image; S2. extract shallow features from the input color image and the sparse depth image through a shallow feature extraction module, and stack the extracted shallow features into a unified shallow feature; S3. Processing the unified shallow features through the gated cyclic feature fusion module to obtain a roughly restored first dense depth image, and outputting the high-level features processed by the gated cyclic unit in the first decoding unit; S4. Performing depth image correction by iteratively updating the spatial propagation module according to the sparse depth image, the first dense depth image and the high-level features to obtain a finely repaired second dense depth image.
7. The method according to claim 6, characterized in that The method further comprises: S5. Use the average L2 error between the N finely restored second dense depth images and the corresponding true dense depth images as the loss function to optimize the parameters of the deep image restoration system, where the loss function is: In the above formula, Θ represents the parameters of the whole system; i represents the variable, and 1≤i≤N, N represents the number of each type of image; Ⅱ(·) is the marker function; Y gt represents the corresponding real dense depth image; Y represents the finely restored second dense depth image; ⊙ represents pixel-by-pixel multiplication.