Monocular depth estimation method based on attention feature fusion interaction

By constructing a multi-scale feature interaction fusion module, the attention feature fusion interaction technology is used to solve the problem of insufficient feature correlation in monocular depth estimation, and the accuracy and efficiency of depth estimation are improved.

CN120339356APending Publication Date: 2025-07-18CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510504608.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When the feature extraction is insufficient and the spatial structure information of the features is lost, the existing monocular depth estimation method leads to insufficient correlation between global features and local features, making it difficult to generate an accurate depth map.

Method used

A multi-scale feature interaction fusion module is built, attention weighting and feature fusion are carried out through the AT module, ATDS module and ATLG module, feature representation ability is enhanced, combined with local and global feature interaction, and multi-head attention mechanism and position coding are used to improve the accuracy of depth estimation.

Benefits of technology

Through multi-scale feature fusion and attention mechanism, the accuracy and efficiency of depth estimation are enhanced, the risk of overfitting is reduced, and the training stability and output quality of the model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339356A_ABST
    Figure CN120339356A_ABST
Patent Text Reader

Abstract

The invention provides a monocular depth estimation method based on attention feature fusion interaction, and relates to the field of three-dimensional imaging and artificial intelligence. The method comprises the following steps: carrying out data enhancement preprocessing on input data; extracting a multi-scale feature map after data enhancement by using a depth encoder network; the AT module performs attention weighting on different channels of the highest layer of features in the feature map to highlight important features, and the ATDS module fuses and optimizes feature images of different levels to ensure that each layer of processed feature image still keeps important information, the attention of the model to the important information in the feature map is enhanced, and the accuracy of the model is improved. The ATLG module processes the features of each level to realize interaction of local and global features, so that the representation capability of a feature graph is enhanced, and a multi-scale feature interaction fusion module composed of an AT module, an ATDS module and an ATLG module is constructed; and outputting the depth image through the decoder network. According to the method, features from different scales are enhanced, the correlation between global and local features is improved, and improvement of the overall precision of a depth estimation task is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of three-dimensional imaging and artificial intelligence, and particularly relates to a monocular depth estimation method based on attention feature fusion interaction. Background Art

[0002] Monocular depth estimation (MDE) refers to estimating the three-dimensional structure of a scene from a single image and representing it in the form of a depth map. As a fundamental task in three-dimensional vision, MDE plays a crucial role in application scenarios such as autonomous driving and three-dimensional imaging, and thus has attracted much attention. Given the uncertainty of depth maps and the nature of dense prediction, traditional methods are difficult to generate satisfactory depth maps. With the development of deep learning, deep neural networks have conducted extensive research on the MDE task and achieved excellent results.

[0003] Currently, the mainstream monocular depth estimation methods are divided into unsupervised learning methods and supervised learning methods. Unsupervised learning methods do not require collecting real depth labels. During training, a stereo image pair composed of the original image and the target image is used. First, the encoder predicts the depth map of the original image, and then the decoder combines the target image and the predicted depth map to reconstruct the original image, and the loss is calculated by comparing the reconstructed image with the original image. The supervised learning method is one of the most popular methods currently. Usually, depth labels are collected using a depth camera or lidar, and image depth estimation is processed as a regression task or a classification task.

[0004] Due to insufficient feature extraction and the loss of spatial structure information in the feature extraction stage in the encoding network of most monocular depth estimation models, and due to convolution operations and downsampling in the network, the correlation between global features and local features is insufficient, and there is still a problem of loss of global appearance structure information in the feature learning process, making it difficult to integrate global features and local features, resulting in scale ambiguity and distortion of the predicted depth map and lack of fine-grainedness. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] Aiming at the deficiencies of the prior art, the present invention provides a monocular depth estimation method based on attention feature fusion interaction. The constructed multi-scale feature interaction and fusion module can achieve multi-scale feature fusion, make full use of feature information at different levels, and combine local feature extraction and global feature interaction, while enhancing the feature representation ability, retaining spatial information, and improving the accuracy and efficiency of the depth estimation task.

[0007] (2) Technical Solutions

[0008] The present invention specifically adopts the following technical solutions to achieve the above object:

[0009] A monocular depth estimation method based on attention feature fusion interaction, comprising the following steps:

[0010] S1: Perform preprocessing of data augmentation on the input data;

[0011] S2: Use the depth encoder module to extract multi-scale feature maps after data augmentation;

[0012] S3: Use the AT module to perform attention weighting between different channels on the highest-level feature in the extracted multi-scale feature maps to highlight important features;

[0013] S4: Use the ATDS module to fuse and optimize feature images at different levels to ensure that each processed feature image still retains important information and enhance the model's attention to important information in the feature maps;

[0014] S5: Use the ATLG module to process features at each level to achieve the interaction between local and global features, thereby enhancing the representation ability of the feature maps;

[0015] S6: Construct a multi-scale feature interaction fusion module, which consists of the AT module, the ATDS module, and the ATLG module;

[0016] S7: Send the feature maps processed by the multi-scale feature interaction fusion module into the depth decoder module to output the depth image.

[0017] Furthermore, the step S1 includes:

[0018] S11: Obtain pseudo-label data from the data without real depth labels through a pre-trained depth network, and add Gaussian blur and spatial distortion to it;

[0019] S12: Use the real label data and the pseudo-label data for joint training.

[0020] Furthermore, the step S2 includes: The encoder gradually extracts feature maps of different scales through an initial convolutional layer, a max pooling layer, and multiple residual blocks.

[0021] Furthermore, the step S3 includes:

[0022] S31: Calculate an attention map with the same spatial size as the input feature map, and then multiply this attention map element-wise with the original input feature map to highlight the important parts in the input image and weaken the unimportant parts;

[0023] S32: Use a 3x3 convolutional kernel to perform further feature extraction on the adjusted feature map;

[0024] S33: Use group normalization to normalize the convolutional feature map;

[0025] S34: Use the ReLU activation function to add non-linearity to the model.

[0026] Furthermore, the step S4 includes:

[0027] S41: Generate a spatial attention map and a channel attention map, and multiply them element-wise with the input feature map respectively to enhance or suppress the information in certain regions.

[0028] S42: Perform the first convolution operation with a 3x3 convolutional kernel, and determine whether to perform the first downsampling according to the resolution of the data;

[0029] S43: Perform normalization and non-linear transformation through group normalization and the ReLU activation function.

[0030] S44: Perform the second convolution operation with a 3x3 convolutional kernel, with a stride of 2, for downsampling.

[0031] S45: Perform normalization and non-linear transformation again through group normalization and the ReLU activation function.

[0032] Furthermore, the step S5 includes:

[0033] S51: Reshape the input feature map x from the shape (B, C, H, W) to (B, H*W, C) for the calculation of the multi-head attention mechanism;

[0034] S52: Add relative position encoding to capture spatial information;

[0035] S53: Perform layer normalization on the input features, then perform feature interaction through the multi-head self-attention mechanism, scale the output of the self-attention mechanism using a learnable scaling parameter, and add it back to the input features;

[0036] S54: Reshape and transpose the features from (B, H*W, C) to (B, C, H, W);

[0037] S55: Use point convolution to expand the number of feature channels, introduce non-linearity with the GeLU activation function, use depthwise separable convolution to keep the number of channels unchanged, and apply the GeLU activation function again;

[0038] S56: Use point convolution again to compress the number of feature channels, and then scale the features using a learnable scaling parameter;

[0039] S57: Perform a residual connection between the processed features and the original input, and perform random depth regularization to randomly discard paths to prevent overfitting, and then return the feature map.

[0040] Further, step S6 includes: The module receives feature maps of multiple scales at the encoder. The highest-level feature map x_4 is sent to the AT module for processing. The processed feature map passes through the ATLG module to obtain the output of x_4, denoted as x_c4. The x_c4 feature map is processed by the ATDS module and then concatenated with the x_3 feature map and sent to a convolutional layer for processing. The processed feature map then passes through the ATLG module, and then the output of x_3 is obtained, denoted as x_c3. The feature maps of x_2 and x_1 are processed in this way successively to obtain the outputs of x_2 and x_1, denoted as x_c2 and x_c1. Finally, x_c4, x_c3, x_c2, and x_c1 are returned.

[0041] Further, step S7 includes: By combining spatial and channel attention mechanisms in the decoder, the model's ability to represent input data and output quality can be significantly improved while maintaining computational efficiency.

[0042] Further, the loss function of the overall network is expressed by the following formula: where d i and are the target and predicted depth values respectively, E i and are the edge intensities of the target and predicted depth maps respectively, N is the number of valid pixels, M is the number of valid edge pixels, λ is a hyperparameter used to control the influence of the variance and mean of the logarithmic difference, θ is the threshold, and α is the weight of the edge protection term.

[0043] (III) Beneficial Effects

[0044] Compared with the prior art, the present invention provides a monocular depth estimation method based on attention feature fusion interaction, having the following beneficial effects:

[0045] By combining feature maps of different levels, details and global structures in the image can be better captured, making full use of feature information at different levels and improving the accuracy of depth estimation; using multiple attention mechanisms can adaptively weight different features, highlight important feature regions, suppress unimportant regions, and enhance the representation ability of key information; through the multi-head attention mechanism and position encoding, the interaction between local and global features is realized and their correlation is enhanced, enabling better processing of texture-rich regions, edges, and details; through residual connections and layer normalization, it helps to improve the training stability of the model, reduce the risk of overfitting, and accelerate the convergence of the model; this method can be integrated with most fully convolutional backbone networks to implement an end-to-end depth training model. Description of the Drawings

[0046] Figure 1Schematic flow diagram of a monocular depth estimation method based on attention feature fusion interaction according to an embodiment of the present invention;

[0047] Figure 2 Schematic flow diagram of the AT module according to an embodiment of the present invention;

[0048] Figure 3 Schematic flow diagram of the ATDS module according to an embodiment of the present invention;

[0049] Figure 4 Schematic flow diagram of the ATLG module according to an embodiment of the present invention;

[0050] Figure 5 Schematic flow diagram of the multi-scale feature interaction and fusion module according to an embodiment of the present invention;

[0051] Figure 6 Depth map predicted according to an embodiment of the present invention. Detailed implementation manners

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] Embodiment

[0054] As Figure 1 shown, a schematic flow diagram of a monocular depth estimation method based on attention feature fusion interaction provided by the present invention, and its specific implementation method includes the following steps:

[0055] S1: Perform preprocessing of data augmentation on the input data;

[0056] S2: Use the depth encoder module to extract multi-scale feature maps after data augmentation;

[0057] S3: Use the AT module to perform attention weighting between different channels on the highest-level features in the extracted multi-scale feature maps to highlight important features;

[0058] S4: Use the ATDS module to fuse and optimize feature images at different levels to ensure that each processed feature image still retains important information and enhance the model's attention to important information in the feature maps;

[0059] S5: Use the ATLG module to process the features at each level to achieve the interaction between local and global features, thereby enhancing the representation ability of the feature maps;

[0060] S6: Construct a multi-scale feature interaction and fusion module, which consists of an AT module, an ATDS module, and an ATLG module;

[0061] S7: Send the feature map processed by the multi-scale feature interaction and fusion module into the depth decoder module to output the depth image.

[0062] Furthermore, the step S1 includes:

[0063] S11: Obtain pseudo-label data from the data without real depth labels through a pre-trained depth network, and add Gaussian blur and spatial distortion to it;

[0064] S12: Use the real label data and the pseudo-label data for joint training.

[0065] Furthermore, the step S2 includes: The encoder gradually extracts feature maps of different scales through an initial convolutional layer, a max pooling layer, and multiple residual blocks.

[0066] As Figure 2 shown, it is a schematic flowchart of the AT module described in the embodiment of the present invention. Furthermore, the step S3 includes:

[0067] S31: Calculate an attention map with the same spatial size as the input feature map, and then multiply this attention map element-wise with the original input feature map to highlight the important parts in the input image and weaken the unimportant parts;

[0068] S32: Use a 3x3 convolutional kernel to perform further feature extraction on the adjusted feature map;

[0069] S33: Use group normalization to normalize the convolved feature map;

[0070] S34: Use the ReLU activation function to add non-linearity to the model.

[0071] As Figure 3 shown, it is a schematic flowchart of the ATDS module described in the embodiment of the present invention. Furthermore, the step S4 includes:

[0072] S41: Generate a spatial attention map and a channel attention map, and multiply them element-wise with the input feature map respectively to enhance or suppress the information in certain regions.

[0073] S42: Perform the first convolution operation through a 3x3 convolutional kernel, and decide whether to perform the first downsampling according to the resolution of the data;

[0074] S43: Perform normalization and non-linear transformation through group normalization and the ReLU activation function.

[0075] S44: Perform the second convolution operation with a 3x3 convolution kernel, a stride of 2, for downsampling.

[0076] S45: Normalize and perform non-linear transformation again through group normalization and ReLU activation function.

[0077] As Figure 4 shown, it is a schematic flowchart of the ATLG module described in the embodiment of the present invention. Further, the step S5 includes:

[0078] S51: Reshape the input feature map x from the shape (B, C, H, W) to (B, H*W, C) for the calculation of the multi-head attention mechanism;

[0079] S52: Add relative position encoding to capture spatial information;

[0080] S53: Perform layer normalization on the input features, then perform feature interaction through the multi-head self-attention mechanism, scale the output of the self-attention mechanism using a learnable scaling parameter, and add it back to the input features;

[0081] S54: Reshape and transpose the features from (B, H*W, C) to (B, C, H, W);

[0082] S55: Use point convolution to expand the number of feature channels, introduce non-linearity with the GeLU activation function, use depthwise separable convolution to keep the number of channels unchanged, and apply the GeLU activation function again;

[0083] S56: Use point convolution again to compress the number of feature channels, and then scale the features using a learnable scaling parameter;

[0084] S57: Perform residual connection on the processed features and the original input, and perform random depth regularization to randomly discard paths to prevent overfitting, and then return the feature map.

[0085] As Figure 5 shown, it is a schematic flowchart of the multi-scale feature interaction and fusion module described in the embodiment of the present invention. Further, step S6 includes: The module receives feature maps of multiple scales at the encoder, sends the highest-level feature map x_4 to the AT module for processing, and passes the processed feature map through the ATLG module to obtain the output of x_4, denoted as x_c4; After the x_c4 feature map is processed by the ATDS module, it is concatenated with the x_3 feature map and sent to the convolutional layer for processing, and the processed feature map is then passed through the ATLG module, and then the output of x_3 is obtained, denoted as x_c3; The feature maps of x_2 and x_1 are processed in this way in turn to obtain the outputs of x_2 and x_1, denoted as x_c2 and x_c1, and finally x_c4, x_c3, x_c2, and x_c1 are returned.

[0086] Furthermore, step S7 includes: by combining spatial and channel attention mechanisms in the decoder, it is possible to significantly improve the model's ability to express input data and output quality while maintaining computational efficiency.

[0087] Furthermore, the loss function of the overall network is represented by the following formula: where d i and are the target and predicted depth values respectively, E i and are the edge intensities of the target and predicted depth maps respectively, N is the number of valid pixels, M is the number of valid edge pixels, λ is a hyperparameter used to control the influence of the variance and mean of the logarithmic difference, θ is the threshold, and α is the weight of the edge protection term.

[0088] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A monocular depth estimation method based on attention feature fusion interaction, characterized in that It includes the following steps: S1: Perform preprocessing of data augmentation on the input data; S2: Use a deep encoder module to extract multi-scale feature maps after data augmentation; S3: Use the AT module to perform attention weighting between different channels on the highest-level features in the extracted multi-scale feature maps to highlight important features; S4: Use the ATDS module to fuse and optimize feature images at different levels to ensure that each processed feature image still retains important information and enhance the model's attention to important information in the feature maps; S5: Use the ATLG module to process the features at each level to achieve the interaction between local and global features, thereby enhancing the representation ability of the feature maps; S6: Construct a multi-scale feature interaction and fusion module, which consists of the AT module, the ATDS module, and the ATLG module; S7: Send the feature maps processed by the multi-scale feature interaction and fusion module into the deep decoder module to output the depth image.

2. The monocular depth estimation method based on attention feature fusion interaction according to claim 1, wherein The step S1 includes: jointly training with real labels and pseudo-label data, where the pseudo-label data is the data predicted by a pre-trained deep network for data without real depth labels, and adding strong perturbations to it.

3. The monocular depth estimation method based on attention feature fusion interaction according to claim 1, wherein, The step S3 includes: S31: Generate a spatial attention image and multiply it element-wise with the input feature map; S32: Use a 3x3 convolutional kernel to perform further feature extraction on the adjusted feature map; S33: Use group normalization to normalize the convolutional feature map; S34: Use the ReLU activation function to add non-linearity to the model.

4. A monocular depth estimation method based on attention feature fusion interaction according to claim 1, characterized in that The step S4 includes: S41: Generate a spatial attention map and a channel attention map, and multiply them element-wise with the input feature map respectively to enhance or suppress the information in certain regions; S42: Perform the first convolution operation through a 3x3 convolutional kernel, and decide whether to perform the first downsampling according to the resolution of the data; S43: Perform normalization and non-linear transformation through group normalization and the ReLU activation function; S44: Perform the second convolution operation through a 3x3 convolutional kernel with a stride of 2 for downsampling; S45: Perform normalization and non-linear transformation again through group normalization and the ReLU activation function.

5. A monocular depth estimation method based on attention feature fusion interaction according to claim 1, characterized in that The step S5 includes: S51: Reshape the input feature map x from shape (B, C, H, W) to (B, H*W, C) for the calculation of the multi-head attention mechanism; S52: Add relative position encoding to capture spatial information; S53: Perform layer normalization on the input features, then perform feature interaction through the multi-head self-attention mechanism, scale the output of the self-attention mechanism using a learnable scaling parameter, and add it back to the input features; S54: Reshape and transpose the features from (B, H*W, C) to (B, C, H, W); S55: Use point convolution to expand the number of feature channels, introduce non-linearity with the GeLU activation function, use depthwise separable convolution to keep the number of channels unchanged, and apply the GeLU activation function again; S56: Use point convolution again to compress the number of feature channels, and then scale the features using a learnable scaling parameter; S57: Perform a residual connection between the processed features and the original input, and through stochastic depth regularization, randomly discard paths to prevent overfitting, and then return the feature map.

6. The monocular depth estimation method based on attention feature fusion interaction according to claim 1, wherein The step S6 described above includes: The module receives feature maps of multiple scales at the encoder. The highest-level feature map x_4 is sent to the AT module for processing. The processed feature map passes through the ATLG module to obtain the output of x_4, denoted as x_c4; the x_c4 feature map is processed through the ATDS module and then concatenated with the x_3 feature map and sent to the convolutional layer for processing. The processed feature map then passes through the ATLG module, and then the output of x_3 is obtained, denoted as x_c3; the feature maps of x_2 and x_1 are processed in this way successively to obtain the outputs of x_2 and x_1, denoted as x_c2 and x_c1. Finally, x_c4, x_c3, x_c2, and x_c1 are returned.

7. A monocular depth estimation method based on attention feature fusion interaction according to claim 1, characterized in that, The loss function of the overall network is expressed by the following formula: where d i and are the target and predicted depth values respectively, E i and are the edge intensities of the target and predicted depth maps respectively, N is the number of valid pixels, M is the number of valid edge pixels, λ is a hyperparameter used to control the influence of the variance and mean of the logarithmic difference, θ is the threshold, and α is the weight of the edge protection term.