Depth map super-resolution reconstruction system and method for multi-scale feature fusion correction

Through a depth map super-resolution reconstruction system with multi-scale feature fusion correction, multi-scale feature extraction, fusion and enhancement modules are used, combined with bidirectional cross-scale feature collaboration, the problem of low reconstruction quality in the existing technology is solved, and high-quality high-resolution depth maps are generated.

CN120543382APending Publication Date: 2025-08-26CHINA GRAPHICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717710.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing depth map super-resolution reconstruction method fails to fully explore color feature information, resulting in low reconstruction quality and failure to effectively utilize the two-way interaction between multi-scale features, limiting the ability of cross-scale features to represent the synergistic representation of cross-scale features, resulting in artifacts and blur problems in image reconstruction.

Method used

A depth map super-resolution reconstruction system with multi-scale feature fusion correction is adopted, including multi-scale feature extraction, fusion, enhancement and bidirectional cross-scale feature collaboration modules. The multi-scale feature fusion module is used to initially integrate depth and color features, and multi-scale features are enhanced by color guide features, and high-resolution depth maps are generated by combining residual reconstruction technology.

Benefits of technology

Improve image reconstruction quality, reduce edge blur and artifacts, and generate high-resolution depth maps with clear edges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543382A_ABST
    Figure CN120543382A_ABST
Patent Text Reader

Abstract

The invention discloses a depth map super-resolution reconstruction system and method for multi-scale feature fusion correction, and belongs to the technical field of image processing. According to the method, multi-scale features are extracted from a high-resolution color image and a low-resolution depth image, and the multi-scale depth features and multi-color guide features are preliminarily fused through a multi-scale feature fusion module to obtain fused features; feature enhancement is carried out in combination with the fusion features and color guide features, and the fusion features are enhanced and corrected by using high-frequency color information, so that the problems of edge blur and artifacts are effectively relieved; and finally, information contained in different scale features is transmitted through bidirectional interaction by a bidirectional cross-scale feature cooperation module, a corresponding high-resolution depth map is reconstructed in combination with a residual error, and a depth map of a target scale is obtained through step-by-step interpolation. Therefore, the ability of the system to reconstruct a high-resolution depth map with clearer edges and fewer artifacts is improved, and the problem of low image reconstruction quality in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a depth map super-resolution reconstruction system and method with multi-scale feature fusion correction. Background Art

[0002] Depth map super-resolution reconstruction is an important and challenging task in the field of image processing. Its goal is to generate depth images with higher resolution and richer details by using low-resolution depth images. In recent years, it has become a research hotspot in the fields of image processing and computer vision.

[0003] Many existing depth map super-resolution reconstruction methods tend to use priors such as high-resolution color images to guide reconstruction to improve the network's expressive ability and obtain higher evaluation indicators. However, existing methods only consider using color features at a single scale for guidance, resulting in the failure to fully exploit rich color feature information; simple superposition or attention fusion makes it difficult to accurately align features and easily introduces reconstruction distortions such as artifacts and blur; the one-way information transmission in the reconstruction stage ignores the two-way interaction between multi-scale features, especially the guiding role of large-scale features on small-scale features, thereby limiting the collaborative representation ability of cross-scale features. Therefore, the results of detail feature processing in super-resolution reconstruction need to be improved, and there is a problem of low image reconstruction quality. Summary of the Invention

[0004] In response to the above-mentioned deficiencies in the prior art, the present invention provides a multi-scale feature fusion and correction depth map super-resolution reconstruction system and method to solve the problem of low image reconstruction quality in the prior art.

[0005] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a depth map super-resolution reconstruction system with multi-scale feature fusion correction, comprising: a first multi-scale feature extraction module, a second multi-scale feature extraction module, a multi-scale feature fusion module, a first multi-scale feature enhancement module, a second multi-scale feature enhancement module, a third multi-scale feature enhancement module and a bidirectional cross-scale feature collaboration module; The first multi-scale feature extraction module is used to extract multi-scale features from the high-resolution color image to obtain a first color guidance feature, a second color guidance feature, and a third color guidance feature; The second multi-scale feature extraction module is used to extract multi-scale features from the low-resolution depth map to obtain first-scale depth features, second-scale depth features and third-scale depth features; The multi-scale feature fusion module is used to fuse the first color guide feature, the second color guide feature, the third color guide feature, the first scale depth feature, the second scale depth feature and the third scale depth feature to obtain a first fused feature, a second fused feature and a third fused feature; The first multi-scale feature enhancement module is used to enhance the first fusion feature and the first color guide feature to obtain a first color fusion enhanced feature; The second multi-scale feature enhancement module is used to enhance the second fusion feature and the second color guide feature to obtain a second color fusion enhanced feature; The third multi-scale feature enhancement module is used to enhance the third fusion feature and the third color guide feature to obtain a third color fusion enhanced feature; The bidirectional cross-scale feature collaboration module is used to process the low-resolution depth map, the first color fusion enhancement feature, the second color fusion enhancement feature and the third color fusion enhancement feature to obtain the first target scale depth map, the second target scale depth map and the third target scale depth map.

[0006] Furthermore, the first multi-scale feature extraction module includes a stacked convolution layer and two 1×1 Stride Convs. The stacked convolution layer is used to extract the third color-guided feature from the high-resolution color image. One Stride Conv is used to perform feature extraction on the third color-guided feature to obtain the second color-guided feature. The other Stride Conv is used to perform feature extraction on the second color-guided feature to obtain the first color-guided feature. Among them, Stride Conv is a strided convolution layer.

[0007] Furthermore, the second multi-scale feature extraction module includes a stacked convolution layer and 3 Ups, the stacked convolution layer is used to extract initial-scale depth features from the low-resolution depth map, the first Up is used to extract first-scale depth features from the initial-scale depth features, the second Up is used to extract second-scale depth features from the first-scale depth features, and the third Up is used to extract third-scale depth features from the second-scale depth features, wherein Up is a sub-pixel convolution layer.

[0008] Furthermore, the expression of the multi-scale feature fusion module is: in, The upsampling scale after recombinant fusion is The deep features of The upsampling scale after recombinant fusion is The deep features of The upsampling scale after recombinant fusion is The deep features of is the first scale depth feature, is the second scale depth feature, is the third scale depth feature, The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The color characteristics of is the first color guide feature, is the second color guide feature, is the third color guide feature, C is the channel connection layer, is bilinear interpolation, R is the residual block Res, i The values ​​are 1, 2, and 3 respectively. The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The deep features of The convolution kernel is Convolutional layers of size is the fusion feature output by the multi-scale feature fusion module, i When 1, is the first fusion feature, i When it is 2, is the second fusion feature, i It is 3 o'clock, It is the third fusion feature.

[0009] Furthermore, the first multi-scale feature enhancement module, the second multi-scale feature enhancement module and the third multi-scale feature enhancement module have the same structure, and all include: a multi-scale feature refinement submodule, a shallow detail enhancement submodule, a deep feature extraction submodule, a semantic correction submodule, a channel connection layer and a 1×1 Conv; The multi-scale feature refinement submodule is used to extract detail features from the fused features using receptive fields of different sizes, fuse the detail features, unify the channel dimensions of the fused detail features, and obtain the output features of the multi-scale feature refinement submodule; The deep feature extraction submodule is used to perform deep feature extraction on the color guide feature to obtain the output features of the deep feature extraction submodule; The shallow detail enhancement submodule is used to enhance the details of the output features of the color guidance feature and multi-scale feature refinement submodule to obtain the output features of the shallow detail enhancement submodule; The semantic correction submodule is used to perform semantic correction on the output features of the deep feature extraction submodule and the output features of the multi-scale feature refinement submodule to obtain the output features of the semantic correction submodule; The channel connection layer is used to connect the output features of the semantic correction submodule and the output features of the shallow detail enhancement submodule through channels to obtain the output features of the channel connection layer; The 1×1 Conv is used to fuse the output features of the channel connection layer to obtain color fusion enhanced features.

[0010] Furthermore, the expression of the multi-scale feature refinement submodule is: in, is the output feature of the multi-scale feature refinement submodule, C is the channel connection layer, For the first dilated convolution, For the second dilated convolution, For the third dilated convolution, is the output of the first dilated convolution, is the output of the second dilated convolution DilatedConv, is the output of the third dilated convolution, The convolution kernel is Convolutional layers of size To fusion features, i The values ​​are 1, 2, and 3 respectively. is the first fusion feature, is the second fusion feature, It is the third fusion feature.

[0011] Furthermore, the expression of the shallow detail enhancement submodule is: in, is the output feature of the shallow detail enhancement submodule, is the inverse residual mask, is point-wise convolution, is the depthwise convolution, is the Sigmoid activation function, The convolution kernel is Convolutional layers of size is the output feature of the multi-scale feature refinement submodule, For element-wise multiplication, It is a color guide feature.

[0012] Furthermore, the deep feature extraction submodule includes: 4 sequentially connected stacked convolutional layers.

[0013] Furthermore, the expression of the semantic correction submodule is: in, is the output feature of the semantic correction submodule, is the spatial attention weight, C is the channel connection layer, is the average pooled GMP, is the maximum pooling GAP, is the output feature of the deep feature extraction submodule, For element-wise multiplication, is the color feature after weighting, is the output feature of the multi-scale feature refinement submodule, The convolution kernel is Convolutional layers of size The convolution kernel is Convolutional layers of size is the generated semantic mask.

[0014] Furthermore, the expression of the bidirectional cross-scale feature collaboration module is: in, is the intermediate feature of the LTH submodule, and is the output feature of the LTH submodule, is the intermediate feature of the HTL submodule, and is the output feature of the HTL submodule, Up is the sub-pixel convolution layer, is the inverse sub-pixel convolution layer, C is the channel connection layer, is a low-resolution depth map, The convolution kernel is Convolutional layers of size The convolution kernel is Convolutional layers of size RCAB is the residual channel attention, is the bilinear interpolation function, is the output fusion feature of the HTL submodule and the LTH submodule, 、 and is the target scale depth map, and To enhance the features of color fusion, and is the output feature of the RRM submodule.

[0015] A depth map super-resolution reconstruction method based on multi-scale feature fusion correction includes the following steps: S1. Extract multi-scale features from the high-resolution color image to obtain a first color guidance feature, a second color guidance feature, and a third color guidance feature; S2. Extract multi-scale features from the low-resolution depth map to obtain first-scale depth features, second-scale depth features, and third-scale depth features; S3, fusing the first color guidance feature, the second color guidance feature, the third color guidance feature, the first scale depth feature, the second scale depth feature, and the third scale depth feature to obtain a first fused feature, a second fused feature, and a third fused feature; S4. Perform feature enhancement on the first fusion feature and the first color guidance feature to obtain a first color fusion enhancement feature; S5. Enhance the second fusion feature and the second color guide feature to obtain a second color fusion enhanced feature; S6. Enhance the third fusion feature and the third color guidance feature to obtain a third color fusion enhancement feature; S7. Reconstruct a first target scale depth map, a second target scale depth map, and a third target scale depth map according to the low-resolution depth map, the first color fusion enhancement feature, the second color fusion enhancement feature, and the third color fusion enhancement feature.

[0016] The beneficial effects of the present invention are as follows: the present invention extracts multi-scale features from high-resolution color images and low-resolution depth images respectively, and preliminarily fuses the multi-scale depth features with the multi-color guide features through a multi-scale feature fusion module to obtain fused features, thereby enhancing the feature interaction between different upsampling factors; then, the fused features and the color guide features are combined for feature enhancement, and the fused features are enhanced and corrected using high-frequency color information, thereby effectively alleviating edge blur and artifact problems; finally, the information contained in the features of different scales is transmitted through a bidirectional cross-scale feature collaboration module through bidirectional interaction, and the corresponding high-resolution depth map is reconstructed in combination with the residual, and then the depth map of the target scale is obtained by step-by-step interpolation, thereby improving the system's ability to reconstruct a high-resolution depth map with clearer edges and fewer artifacts, thereby solving the problem of low image reconstruction quality in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a system block diagram of a depth map super-resolution reconstruction system with multi-scale feature fusion correction; Figure 2 Schematic diagram of the structure of the first multi-scale feature extraction module; Figure 3 Schematic diagram of the structure of the second multi-scale feature extraction module; Figure 4 It is a structural diagram of the multi-scale feature fusion module; Figure 5 It is a structural diagram of the residual block Res; Figure 6 Schematic diagram of the structures of the first multi-scale feature enhancement module, the second multi-scale feature enhancement module, and the third multi-scale feature enhancement module; Figure 7 Schematic diagram of the structure of the multi-scale feature refinement submodule; Figure 8 This is a schematic diagram of the structure of the deep feature extraction submodule; Figure 9 Schematic diagram of the structure of the bidirectional cross-scale feature collaboration module; Figure 10 It is a structural diagram of the LTH submodule; Figure 11 It is a structural diagram of the HTL submodule; Figure 12 This is a structural diagram of the ResBlock submodule; Figure 13 Schematic diagram of the structure of the RRM submodule. DETAILED DESCRIPTION

[0018] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0019] Example 1, as Figure 1 As shown, a depth map super-resolution reconstruction system with multi-scale feature fusion correction includes: a first multi-scale feature extraction module, a second multi-scale feature extraction module, a multi-scale feature fusion module, a first multi-scale feature enhancement module, a second multi-scale feature enhancement module, a third multi-scale feature enhancement module and a bidirectional cross-scale feature collaboration module; The first multi-scale feature extraction module is used to extract high-resolution color images Extracting multi-scale features to obtain a first color guidance feature, a second color guidance feature, and a third color guidance feature; The second multi-scale feature extraction module is used to extract the low-resolution depth map Extracting multi-scale features to obtain first-scale depth features, second-scale depth features, and third-scale depth features; The multi-scale feature fusion module is used to fuse the first color guide feature, the second color guide feature, the third color guide feature, the first scale depth feature, the second scale depth feature and the third scale depth feature to obtain a first fused feature, a second fused feature and a third fused feature; The first multi-scale feature enhancement module is used to enhance the first fusion feature and the first color guide feature to obtain a first color fusion enhanced feature; The second multi-scale feature enhancement module is used to enhance the second fusion feature and the second color guide feature to obtain a second color fusion enhanced feature; The third multi-scale feature enhancement module is used to enhance the third fusion feature and the third color guide feature to obtain a third color fusion enhanced feature; The bidirectional cross-scale feature collaboration module is used to process the low-resolution depth map, the first color fusion enhancement feature, the second color fusion enhancement feature and the third color fusion enhancement feature to obtain the first target scale depth map, the second target scale depth map and the third target scale depth map.

[0020] The multi-scale feature fusion module is combined with the residual block to achieve the preliminary fusion of scale depth features and color guidance features through the cross-modal feature interaction mechanism, thereby enhancing the feature interaction between different upsampling factors.

[0021] The three multi-scale feature enhancement modules use color-guided features and multi-scale feature refinement to enhance and correct the fused features, and use the high-frequency information of color features to guide the correction of the initially fused features to reduce edge blur and artifacts.

[0022] The bidirectional cross-scale feature collaboration module transfers the information contained in features of different scales through bidirectional interaction, and combines with the RRM sub-module to reconstruct the corresponding high-resolution depth map.

[0023] In this embodiment, if Figure 2 As shown, the first multi-scale feature extraction module includes a stacked convolution layer and two 1×1 Stride Convs. The stacked convolution layer is used to extract the third color-guided feature from the high-resolution color image. One Stride Conv is used to perform feature extraction on the third color-guided feature to obtain the second color-guided feature. The other Stride Conv is used to perform feature extraction on the second color-guided feature to obtain the first color-guided feature. Among them, Stride Conv is a strided convolution layer.

[0024] The stacked convolutional layer includes two 3×3 Conv layers and one 1×1 Conv layer.

[0025] In this embodiment, if Figure 3 As shown, the second multi-scale feature extraction module includes a stacked convolution layer and 3 Ups. The stacked convolution layer is used to extract the initial scale depth features of the low-resolution depth map. The first Up is used to extract the first scale depth features from the initial scale depth features. The second Up is used to extract the second scale depth features from the first scale depth features. The third Up is used to extract the third scale depth features from the second scale depth features, where Up is a sub-pixel convolution layer.

[0026] In the two multi-scale feature extraction modules, stacked convolution layers are used to extract shallow features, stride convolution layers are used for downsampling, and sub-pixel convolution layers are used for upsampling.

[0027] High-resolution color images and low-resolution depth maps Input the multi-scale feature extraction module to obtain a multi-scale feature map, specifically: high-resolution color map and low-resolution depth maps Shallow feature extraction is performed by stacking convolution layers consisting of two 3×3 convolution layers and one 1×1 convolution layer, and then feature maps of different scales are obtained through strided convolution layers and sub-pixel convolution layers. and The process can be expressed as: in, and is the color guidance feature, and is the scale depth feature, For the stacked convolutional layers in the first multi-scale feature extraction module, is the strided convolution layer for downsampling, For the stacked convolutional layers in the second multi-scale feature extraction module, is the output feature of the stacked convolution layer in the second multi-scale feature extraction module, is the sub-pixel convolution layer for upsampling.

[0028] like Figure 4 As shown, the multi-color guide feature , multi-scale deep features The output of the residual block Res of the same scale is fused through the channel connection layer and the 1×1 convolution layer to unify the channel dimension and obtain the preliminary fusion feature. .

[0029] The expression of the multi-scale feature fusion module is: in, The upsampling scale after recombinant fusion is The deep features of The upsampling scale after recombinant fusion is The deep features of The upsampling scale after recombinant fusion is The deep features of is the first scale depth feature, is the second scale depth feature, is the third scale depth feature, The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The color characteristics of is the first color guide feature, is the second color guide feature, is the third color guide feature, C is the channel connection layer, is bilinear interpolation, R is the residual block Res, i The values ​​are 1, 2, and 3 respectively. The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The deep features of The convolution kernel is Convolutional layers of size is the fusion feature output by the multi-scale feature fusion module, i When 1, is the first fusion feature, i When it is 2, is the second fusion feature, i It is 3 o'clock, It is the third fusion feature.

[0030] like Figure 5 As shown, the residual block Res includes: 2 3×3 Convs, 1 1×1 Conv and an adder.

[0031] In this embodiment, if Figure 6 As shown, the first multi-scale feature enhancement module, the second multi-scale feature enhancement module and the third multi-scale feature enhancement module have the same structure, and all include: a multi-scale feature refinement submodule, a shallow detail enhancement submodule, a deep feature extraction submodule, a semantic correction submodule, a channel connection layer and a 1×1 Conv; The multi-scale feature refinement submodule is used to extract detail features from the fused features using receptive fields of different sizes, fuse the detail features, unify the channel dimensions of the fused detail features, and obtain the output features of the multi-scale feature refinement submodule; The deep feature extraction submodule is used to perform deep feature extraction on the color guide feature to obtain the output features of the deep feature extraction submodule; The shallow detail enhancement submodule is used to enhance the details of the output features of the color guidance feature and multi-scale feature refinement submodule to obtain the output features of the shallow detail enhancement submodule; The semantic correction submodule is used to perform semantic correction on the output features of the deep feature extraction submodule and the output features of the multi-scale feature refinement submodule to obtain the output features of the semantic correction submodule; The channel connection layer is used to connect the output features of the semantic correction submodule and the output features of the shallow detail enhancement submodule through channels to obtain the output features of the channel connection layer; The 1×1 Conv is used to fuse the output features of the channel connection layer to obtain color fusion enhanced features.

[0032] like Figure 7 As shown, the initial fusion features Input into the multi-scale feature refinement submodule to obtain the output features of the multi-scale feature refinement submodule Specifically: through three dilated convolutions with different dilation rates, we use receptive fields of different sizes to capture more details and restore detailed features of different scales as much as possible, thereby improving the initial fusion features. To enhance, the outputs of three different dilated convolutions are fused through a 1×1 convolution layer and the channel dimensions are unified and output .

[0033] The expression of the multi-scale feature refinement submodule is: in, is the output feature of the multi-scale feature refinement submodule, C is the channel connection layer, For the first dilated convolution, For the second dilated convolution, For the third dilated convolution, is the output of the first dilated convolution, is the output of the second dilated convolution DilatedConv, is the output of the third dilated convolution, The convolution kernel is Convolutional layers of size To fusion features, i The values ​​are 1, 2, and 3 respectively. is the first fusion feature, is the second fusion feature, It is the third fusion feature.

[0034] like Figure 8 As shown in the figure, the deep feature extraction submodule includes: 4 sequentially connected stacked convolutional layers, using 4 stacked convolutional layers consisting of two 3×3 convolutional layers and one 1×1 convolutional layer to extract the color guide features. Perform deep feature extraction to obtain .

[0035] like Figure 6 As shown, in the shallow detail enhancement submodule, the color guide features are And the multi-scale feature refinement submodule output features Perform point-by-point convolution and depth-wise convolution to achieve the goal of reducing the amount of calculation while maintaining sensitivity to multi-scale features and capturing multi-scale information. Then, the Sigmoid function is used to obtain the residual mask of the area with large differences in the RGB-D modality, and the reverse residual mask is used. Improve the weight of favorable information and suppress interference information to achieve color guidance features Redistribution of information weights, outputting feature maps with enhanced details .

[0036] The expression of the shallow detail enhancement submodule is: in, is the output feature of the shallow detail enhancement submodule, is the inverse residual mask, is point-wise convolution, is the depthwise convolution, is the Sigmoid activation function, The convolution kernel is Convolutional layers of size is the output feature of the multi-scale feature refinement submodule, For element-wise multiplication, It is a color guide feature.

[0037] like Figure 6 As shown, the semantic correction submodule uses average pooling GMP and maximum pooling GAP, a 1×1 convolution layer and a Sigmoid function to obtain the spatial attention weight. , for deep color features The information weight is redistributed to obtain , thereby improving the weight of favorable information and suppressing interference information; then a 1×1 convolution layer and a 3×3 convolution layer are used to and The features obtained after channel splicing generate semantic masks , thus Perform semantic enhancement and correction, and output semantic correction features .

[0038] The expression of the semantic correction submodule is: in, is the output feature of the semantic correction submodule, is the spatial attention weight, C is the channel connection layer, is the average pooled GMP, is the maximum pooling GAP, is the output feature of the deep feature extraction submodule, For element-wise multiplication, is the color feature after weighting, is the output feature of the multi-scale feature refinement submodule, The convolution kernel is Convolutional layers of size The convolution kernel is Convolutional layers of size is the generated semantic mask.

[0039] like Figure 6 As shown in the figure, the outputs of the shallow detail enhancement submodule and the semantic correction submodule are fused through channel connection and a 1×1 convolution layer to obtain the color fusion enhancement features. , the specific expression is: .

[0040] In this embodiment, if Figure 9 As shown, the bidirectional cross-scale feature collaboration module includes multiple LTH sub-modules, multiple HTL sub-modules and multiple RRM sub-modules.

[0041] The color fusion enhancement feature of the multi-scale feature enhancement module and low-resolution depth maps They are input into the bidirectional cross-scale feature collaboration module to obtain the final reconstruction result. The bidirectional cross-scale feature collaboration module first processes each input feature in parallel, where the LTH submodule forwardly transfers the information contained in the small-scale feature map to the large-scale, and the HTL submodule reversely transfers the large-scale feature information. Since there is only one input when the maximum-scale input feature information is reversely transferred, the RRM submodule is directly used to refine and transfer it, that is, ; Then, the RRM submodule is used to further effectively fuse and refine the features obtained after processing by the LTH and HTL modules to obtain , in preparation for the subsequent generation of the final reconstruction image; finally, all Interpolate step by step to obtain the depth map of the final required scale .

[0042] The expression of the bidirectional cross-scale feature collaboration module is: in, is the intermediate feature of the LTH submodule, and is the output feature of the LTH submodule, is the intermediate feature of the HTL submodule, and is the output feature of the HTL submodule, Up is the sub-pixel convolution layer, is the inverse sub-pixel convolution layer, C is the channel connection layer, is a low-resolution depth map, The convolution kernel is Convolutional layers of size The convolution kernel is Convolutional layers of size RCAB is the residual channel attention, is the bilinear interpolation function, is the output fusion feature of the HTL submodule and the LTH submodule, 、 and is the target scale depth map, and To enhance the features of color fusion, and is the output feature of the RRM submodule.

[0043] i The values ​​are 1, 2, and 3 respectively. is the first color fusion enhancement feature, is the second color fusion enhancement feature, For the third color fusion enhancement feature, is the first target scale depth map, is the second target scale depth map, is the third target scale depth map, is the output feature of the first RRM submodule, is the output feature of the second RRM submodule, is the output feature of the third RRM submodule.

[0044] like Figure 10 As shown in Figure 1, the LTH submodule includes: 1 sub-pixel convolution layer Up, a channel connection layer and a ResBlock submodule.

[0045] like Figure 11 As shown in the figure, the HTL submodule includes: inverse sub-pixel convolution layer Down, channel connection layer and ResBlock submodule.

[0046] The inverse sub-pixel convolution layer Down is used for downsampling.

[0047] like Figure 12 As shown in the figure, the ResBlock submodule includes: a 1×1 convolutional layer, two residual channel attention RCABs and a 3×3 convolutional layer. The ResBlock submodule is used to further refine and enhance the input features.

[0048] like Figure 13 As shown in the figure, the RRM submodule includes: one 1×1 convolutional layer and two 3×3 convolutional layers, which are combined with local residual connections to effectively fuse and refine the input features.

[0049] Example 2, a depth map super-resolution reconstruction method using multi-scale feature fusion correction, comprising the following steps: S1. Extract multi-scale features from the high-resolution color image to obtain a first color guidance feature, a second color guidance feature, and a third color guidance feature; S2. Extract multi-scale features from the low-resolution depth map to obtain first-scale depth features, second-scale depth features, and third-scale depth features; S3, fusing the first color guidance feature, the second color guidance feature, the third color guidance feature, the first scale depth feature, the second scale depth feature, and the third scale depth feature to obtain a first fused feature, a second fused feature, and a third fused feature; S4. Perform feature enhancement on the first fusion feature and the first color guidance feature to obtain a first color fusion enhancement feature; S5. Enhance the second fusion feature and the second color guide feature to obtain a second color fusion enhanced feature; S6. Enhance the third fusion feature and the third color guidance feature to obtain a third color fusion enhancement feature; S7. Reconstruct a first target scale depth map, a second target scale depth map, and a third target scale depth map according to the low-resolution depth map, the first color fusion enhancement feature, the second color fusion enhancement feature, and the third color fusion enhancement feature.

[0050] The specific implementation process of Example 2 is the same as that of Example 1.

[0051] S1 is specifically as follows: a stacked convolution layer consisting of two 3×3 convolution layers and one 1×1 convolution layer is used to extract shallow features of the input high-resolution color image, and then two strided convolutions are used to operate on the extracted shallow features to obtain color guide features of different scales. .

[0052] S2 is specifically: extract shallow features of the input low-resolution depth map through a stacked convolution layer consisting of two 3×3 convolution layers and one 1×1 convolution layer, and operate on the extracted shallow depth features through three sub-pixel convolution layers Up to obtain multi-scale depth features .

[0053] S3 is specifically: the residual block Res is used to further refine and enhance the features obtained by sampling at different scales; several 1×1 convolutions are used to fuse the outputs of the residual blocks Res at the same scale and unify the channel dimensions to obtain preliminary fusion features at all levels. .

[0054] S4, S5, and S7 are as follows: through the multi-scale feature refinement submodule, we use the receptive fields of different scales to capture more details and restore the detailed features of different scales as much as possible, thereby improving the initial fusion features. Enhancement is performed by using the shallow detail enhancement submodule to learn the parts of the color image that are irrelevant to the depth information and remove them, thereby suppressing the interference information in the color features, strengthening the spatial position that best matches the depth features, and realizing the adaptive guidance of the color features to the depth features, thereby obtaining the detail enhancement features. Through the deep feature extraction submodule and the semantic correction submodule, the global semantic information contained in the deep color feature is used to perform semantic correction on the features obtained after being enhanced by the multi-scale feature refinement submodule to obtain the semantic correction feature. ; By using a 1×1 convolution layer to combine the outputs of the shallow detail enhancement submodule and the semantic correction submodule and Fusion is performed to obtain each color fusion enhancement feature .

[0055] The specific process of the multi-scale feature refinement submodule includes: through three dilated convolutions with different dilation rates, using receptive fields of different sizes to capture more details and restore detailed features of different scales as much as possible, thereby refining the initial fusion features. Enhancement; fuse the outputs of three different dilated convolutions through a 1×1 convolution layer and unify the channel dimensions and output .

[0056] The specific process of the shallow detail enhancement submodule includes: And the multi-scale feature refinement submodule output features Perform point-by-point convolution and depth-wise convolution to achieve the goal of reducing the amount of calculation while maintaining sensitivity to multi-scale features and capturing multi-scale information; use the Sigmoid function to obtain the residual mask of the area with large differences in the RGB-D modality, and use the inverse residual mask to enhance the weight of favorable information and suppress interference information to achieve color guidance features. Redistribution of information weights, outputting feature maps with enhanced details .

[0057] The deep feature extraction submodule uses four stacked convolutional layers consisting of two 3×3 convolutional layers and one 1×1 convolutional layer to extract the color guidance features. Perform deep feature extraction to obtain deep color features .

[0058] The specific process of the semantic correction submodule includes: using average pooling GMP, maximum pooling GAP, a 1×1 convolution layer and Sigmoid function to obtain spatial attention weights , for deep color features The information weight is redistributed to enhance the weight of favorable information and suppress interference information; then a 1×1 convolution layer and a 3×3 convolution layer are used to and The features obtained after channel splicing generate semantic masks , thus Perform semantic enhancement and correction, and output semantic correction features .

[0059] Finally, the outputs of the shallow detail enhancement module and the semantic correction submodule are fused through channel connection and a 1×1 convolution layer to obtain the color fusion enhancement features. .

[0060] S7 is specifically: the information contained in the small-scale feature map is converted through an LTH submodule consisting of a sub-pixel convolution layer Upsample for upsampling and a ResBlock submodule The information contained in the large-scale feature map is forwarded to the large-scale feature map through the HTL submodule consisting of a downsampling inverse sub-pixel convolution layer Down and a ResBlock submodule. The forward transmission is to the small-scale feature map. Since there is only one input when the maximum-scale input feature information is transmitted backward, the RRM submodule is directly used to refine and transmit it. The RRM submodule uses a 1×1 convolution layer and two 3×3 convolution layers combined with local residual connections to and The features obtained after element-by-element addition Further integration and refinement yields , and finally all Interpolate step by step to obtain the depth map of the final required scale and obtain the final reconstruction result .

[0061] The loss function of the depth map super-resolution reconstruction system is the mean absolute error loss function, and the specific expression is: , L 1 is the loss value, N is the number of pixels, is the target scale depth map, are the real depth maps of different scales, To perform the norm operation on the difference between two images, j Is a positive integer.

[0062] Specifically, the present invention uses real and synthetic datasets to verify the effectiveness of the proposed scheme. Bicubic interpolation is used to downsample the real depth maps to obtain LR depth maps of different scales. In the real dataset, to reduce training time while not affecting network performance, the first 1000 RGB-D image pairs from the NYU v2 dataset are used as the training dataset and then sliced ​​into 256×256 pixel sub-images. The remaining 449 RGB-D image pairs are used as the test dataset. Furthermore, the model trained on the NYU v2 dataset is directly tested for generalization on the Lu RGB-D dataset (6 pairs) and the Middlebury RGB-D dataset (30 pairs) provided by Zhao et al. Furthermore, for training and testing on the synthetic datasets, a total of 82 RGB-D image pairs are selected from the Middlebury and MPI Sintel datasets as the training set, and 6 RGB-D image pairs are selected from the Middlebury (2005) dataset as the test set. As with the NYU v2 dataset, the RGB-D image pairs are cropped to 256×256 pixels. All training data are normalized. In order to fully utilize the information contained in the image, the original image is rotated. And flip (horizontally, vertically) processing to enhance the data.

[0063] After the training set is built, the model can be trained and tested on the pytorch framework. The Adam optimization algorithm is used to optimize the network parameters, where the parameter settings of the Adam optimizer are 0.9, is 0.999, and the learning rate is . Root mean square error (RMSE) was used to evaluate model performance.

[0064] The present invention uses the NYU dataset, the Middlebury dataset, and the Lu dataset to verify the effectiveness of the proposed model. The comparative methods based on learning in the comparative experiment 111, including: JIIF, DKN, FDKN, FDSR, AHMF, DCTNet, DAGF, GeoDSR, DASNet, TM-GAN, and DSR_Diff, are compared and tested. The experimental results are shown in Table 1. The comparative experimental depth map super-resolution reconstruction methods include: Tang JX, Chen XK, Zeng G. Joint implicit image function for guideddepth super-resolution[C] / / ACM International Conference on Multimedia, NewYork: ACM, 2021: 4390–4399. Kim B, Ponce J, Ham B. Deformable kernel networks for joint imagefiltering[J]. International Journal of Computer Vision, 2021, 129(2): 579-600. He L, Zhu H, Li F, et al. Towards fast and accurate real-world depthsuper-resolution: Benchmark dataset and baseline[C] / / IEEE / CVF Conference on Computer Vision and Pattern Recognition. Nashville: IEEE, 2021: 9225-9234. Zhong Z W, Liu X M, Jiang J J, et al. High-resolution depth mapsimaging via attention-based hierarchical multi-modal fusion[J]. IEEETransactions on Image Processing, 2022, 31: 648-663. Zhong Z W, Liu X M, Jiang J J, et al. Deep attentional guided imagefiltering[J]. IEEE Transactions on Neural Networks and Learning Systems,2024, 35(9): 12236-12250. Wang X H, Chen X H, Ni B B, et al. Learning continuous depthrepresentation via geometric spatial aggregator[C] / / AAAI Conference onArtificial Intelligence. Washington: AAAI, 2023: 2698-2706. Hou Y, Nie L, Lin C Y, et al. Digging into depth-adaptive structurefor guided depth super-resolution[J]. Displays, 2024, 84: 102752. Zhu J, Koh V K Z, Lin Z P, et al. TM-GAN: A transformer-based multi-modal generative adversarial network for guided depth image super-resolution[J]. IEEE Journal on Emerging and Selected Topics in Circuits and Systems,2024, 14(2): 261-274. Shi Y, Cao H, Xia B, et al. DSR-Diff: Depth map super-resolution withdiffusion model[J]. Pattern Recognition Letters, 2024, 184: 225-231. Table 1 Comparison of RMSE values ​​on three test sets.

[0065] Table 1 RMSE statistics of NYU dataset, Middlebury dataset and Lu dataset As can be seen from Table 1 (the optimal value is highlighted in black), in most cases, the RSME of the present invention is the lowest, and the reconstruction effect is significantly better than other reconstruction methods.

[0066] The present invention fuses the extracted multi-scale depth features and color guidance features through a multi-scale feature fusion module. Subsequently, the multi-scale feature enhancement module uses multi-scale feature refinement and layered color feature guidance to enhance the details of the initially fused features, and uses the high-frequency guidance information of the color features to reduce edge blur and artifacts. Finally, the bidirectional cross-scale feature collaboration module further utilizes the multi-scale feature information to promote the mutual collaboration of fused features at different scales, thereby improving the reconstruction quality of the depth map and effectively solving common phenomena such as texture transfer.

[0067] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A depth map super-resolution reconstruction system with multi-scale feature fusion correction, characterized by: include: A first multi-scale feature extraction module, a second multi-scale feature extraction module, a multi-scale feature fusion module, a first multi-scale feature enhancement module, a second multi-scale feature enhancement module, a third multi-scale feature enhancement module and a bidirectional cross-scale feature collaboration module; The first multi-scale feature extraction module is used to extract multi-scale features from the high-resolution color image to obtain a first color guidance feature, a second color guidance feature, and a third color guidance feature; The second multi-scale feature extraction module is used to extract multi-scale features from the low-resolution depth map to obtain first-scale depth features, second-scale depth features, and third-scale depth features; The multi-scale feature fusion module is used to fuse the first color guidance feature, the second color guidance feature, the third color guidance feature, the first scale depth feature, the second scale depth feature and the third scale depth feature to obtain a first fused feature, a second fused feature and a third fused feature; The first multi-scale feature enhancement module is used to enhance the first fusion feature and the first color guide feature to obtain a first color fusion enhanced feature; The second multi-scale feature enhancement module is used to enhance the second fusion feature and the second color guide feature to obtain a second color fusion enhanced feature; The third multi-scale feature enhancement module is used to enhance the third fusion feature and the third color guidance feature to obtain a third color fusion enhancement feature; The bidirectional cross-scale feature collaboration module is used to process the low-resolution depth map, the first color fusion enhancement feature, the second color fusion enhancement feature and the third color fusion enhancement feature to obtain a first target scale depth map, a second target scale depth map and a third target scale depth map.

2. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 1 is characterized in that: The first multi-scale feature extraction module includes a stacked convolution layer and two 1×1 Stride Convs. The stacked convolution layer is used to extract the third color guidance feature from the high-resolution color image. One Stride Conv is used to perform feature extraction on the third color guidance feature to obtain the second color guidance feature. The other Stride Conv is used to perform feature extraction on the second color guidance feature to obtain the first color guidance feature, wherein the Stride Conv is a strided convolution layer.

3. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 1 is characterized in that: The second multi-scale feature extraction module includes a stacked convolution layer and three Ups, the stacked convolution layer is used to extract initial-scale depth features from the low-resolution depth map, the first Up is used to extract first-scale depth features from the initial-scale depth features, the second Up is used to extract second-scale depth features from the first-scale depth features, and the third Up is used to extract third-scale depth features from the second-scale depth features, wherein Up is a sub-pixel convolution layer.

4. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 1, characterized in that: The expression of the multi-scale feature fusion module is: in, The upsampling scale after recombinant fusion is The deep features of The upsampling scale after recombinant fusion is The deep features of The upsampling scale after recombinant fusion is The deep features of is the first scale depth feature, is the second scale depth feature, is the third scale depth feature, The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The color characteristics of is the first color guide feature, is the second color guide feature, is the third color guide feature, C is the channel connection layer, is bilinear interpolation, R is the residual block Res, i The values ​​are 1, 2, and 3 respectively. The upsampling scale after recombinant fusion is The color characteristics of The upsampling scale after recombinant fusion is The deep features of The convolution kernel is Convolutional layers of size is the fusion feature output by the multi-scale feature fusion module, i When 1, is the first fusion feature, i When it is 2, is the second fusion feature, i It is 3 o'clock, It is the third fusion feature.

5. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 1, characterized in that: The first multi-scale feature enhancement module, the second multi-scale feature enhancement module and the third multi-scale feature enhancement module have the same structure and all include: a multi-scale feature refinement submodule, a shallow detail enhancement submodule, a deep feature extraction submodule, a semantic correction submodule, a channel connection layer and a 1×1 Conv; The multi-scale feature refinement submodule is used to extract detail features from the fused features using receptive fields of different sizes, fuse the detail features, unify the channel dimensions of the fused detail features, and obtain the output features of the multi-scale feature refinement submodule; The deep feature extraction submodule is used to perform deep feature extraction on the color guide feature to obtain output features of the deep feature extraction submodule; The shallow detail enhancement submodule is used to enhance the details of the output features of the color guidance feature and the multi-scale feature refinement submodule to obtain the output features of the shallow detail enhancement submodule; The semantic correction submodule is used to perform semantic correction on the output features of the deep feature extraction submodule and the output features of the multi-scale feature refinement submodule to obtain the output features of the semantic correction submodule; The channel connection layer is used to perform channel connection on the output features of the semantic correction submodule and the output features of the shallow detail enhancement submodule to obtain the output features of the channel connection layer; The 1×1 Conv is used to fuse the output features of the channel connection layer to obtain color fusion enhancement features.

6. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 5, characterized in that: The expression of the multi-scale feature refinement submodule is: in, is the output feature of the multi-scale feature refinement submodule, C is the channel connection layer, For the first dilated convolution, For the second dilated convolution, For the third dilated convolution DilatedConv, is the output of the first dilated convolution, is the output of the second dilated convolution, is the output of the third dilated convolution, The convolution kernel is Convolutional layers of size To fusion features, i The values ​​are 1, 2, and 3 respectively. is the first fusion feature, is the second fusion feature, It is the third fusion feature.

7. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 5, characterized in that: The expression of the shallow detail enhancement submodule is: in, is the output feature of the shallow detail enhancement submodule, is the inverse residual mask, is point-wise convolution, is the depthwise convolution, is the Sigmoid activation function, The convolution kernel is Convolutional layers of size is the output feature of the multi-scale feature refinement submodule, For element-wise multiplication, It is a color guide feature.

8. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 5, characterized in that: The expression of the semantic correction submodule is: in, is the output feature of the semantic correction submodule, is the spatial attention weight, C is the channel connection layer, is the average pooled GMP, is the maximum pooling GAP, is the output feature of the deep feature extraction submodule, For element-wise multiplication, is the color feature after weighting, is the output feature of the multi-scale feature refinement submodule, The convolution kernel is Convolutional layers of size The convolution kernel is Convolutional layers of size is the generated semantic mask.

9. The depth map super-resolution reconstruction system with multi-scale feature fusion correction according to claim 1, characterized in that: The expression of the bidirectional cross-scale feature collaboration module is: in, is the intermediate feature of the LTH submodule, and is the output feature of the LTH submodule, is the intermediate feature of the HTL submodule, and is the output feature of the HTL submodule, Up is the sub-pixel convolution layer, is the inverse sub-pixel convolution layer, C is the channel connection layer, is a low-resolution depth map, The convolution kernel is Convolutional layers of size The convolution kernel is Convolutional layers of size RCAB is the residual channel attention, is the bilinear interpolation function, is the output fusion feature of the HTL submodule and the LTH submodule, 、 and is the target scale depth map, and To enhance the features of color fusion, and is the output feature of the RRM submodule.

10. A method for super-resolution reconstruction of a depth map using multi-scale feature fusion correction, implemented based on the system for super-resolution reconstruction of a depth map using multi-scale feature fusion correction according to any one of claims 1 to 9, characterized in that: The following steps are involved: S1. Extract multi-scale features from the high-resolution color image to obtain a first color guidance feature, a second color guidance feature, and a third color guidance feature; S2. Extract multi-scale features from the low-resolution depth map to obtain first-scale depth features, second-scale depth features, and third-scale depth features; S3, fusing the first color guidance feature, the second color guidance feature, the third color guidance feature, the first scale depth feature, the second scale depth feature, and the third scale depth feature to obtain a first fused feature, a second fused feature, and a third fused feature; S4. Perform feature enhancement on the first fusion feature and the first color guidance feature to obtain a first color fusion enhancement feature; S5. Enhance the second fusion feature and the second color guide feature to obtain a second color fusion enhanced feature; S6. Enhance the third fusion feature and the third color guidance feature to obtain a third color fusion enhancement feature; S7. Reconstruct a first target scale depth map, a second target scale depth map, and a third target scale depth map according to the low-resolution depth map, the first color fusion enhancement feature, the second color fusion enhancement feature, and the third color fusion enhancement feature.