A monocular depth estimation method, device, storage medium and computer program product
By introducing a cross-dimensional attention mechanism into the monocular depth estimation model for dynamic convolution processing, the problem of high computational complexity is solved, achieving efficient depth estimation that is suitable for edge devices and real-time applications.
Patent Information
- Application Number
- CN202511203944.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Existing monocular depth estimation algorithms consume excessive computational resources and have high computational complexity, resulting in decreased prediction efficiency, especially in edge devices and real-time applications.
A cross-dimensional attention mechanism is adopted, including first convolution kernel attention, input channel attention, and output channel attention. The image features to be predicted are processed by dynamic convolution, which reduces computational complexity and improves prediction efficiency.
It significantly reduces computational complexity, improves prediction efficiency, and is suitable for resource-constrained devices and real-time application scenarios, thereby enhancing the accuracy and computational efficiency of depth estimation.
Smart Images

Figure CN120747186B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image depth estimation technology, and in particular to a monocular depth estimation method, apparatus, storage medium, and computer program product. Background Technology
[0002] Current monocular depth estimation algorithms include Metric3Dv2 and Depth-Anything. Metric3Dv2, by introducing a unified camera space transformation module, enables the model to generalize well across a large number of different camera datasets, but its computational complexity remains high, making it difficult to use in real-time applications. Depth-Anything, on the other hand, uses a large-scale unlabeled dataset and trains its depth estimation algorithm through self-supervised learning. Its main advantages lie in the diversity of data and improved generalization ability. However, Depth-Anything relies on a large dataset for training, resulting in a massive computational burden, and its applicability in certain specific scenarios (such as edge devices) still needs optimization.
[0003] In summary, current monocular depth estimation algorithms still suffer from excessive computational resource consumption and high computational complexity, leading to decreased prediction efficiency. Summary of the Invention
[0004] This application provides a monocular depth estimation method, apparatus, storage medium, and computer program product that can reduce computational complexity and improve prediction efficiency.
[0005] The technical solution of this application embodiment is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a monocular depth estimation method, the method comprising:
[0007] Obtain the image to be predicted;
[0008] The image to be predicted is input into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted.
[0009] The first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. The cross-dimensional attention mechanism includes a first convolution kernel attention, input channel attention, and output channel attention.
[0010] Secondly, embodiments of this application provide a monocular depth estimation device, which includes: an acquisition unit and an input unit; wherein,
[0011] The acquisition unit is used to acquire the image to be predicted;
[0012] The input unit is used to input the image to be predicted into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, the first module is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism including a first convolution kernel attention, input channel attention and output channel attention.
[0013] Thirdly, embodiments of this application provide a monocular depth estimation device, which includes: a processor and a memory; wherein,
[0014] The memory is used to store computer programs that can run on the processor;
[0015] The processor is configured to execute the monocular depth estimation method as described above when running the computer program.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program code, which, when executed by a computer, implements the monocular depth estimation method as described above.
[0017] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the monocular depth estimation method as described above.
[0018] This application provides a monocular depth estimation method, apparatus, storage medium, and computer program product. The method includes: acquiring an image to be predicted; inputting the image to be predicted into a first model for monocular depth estimation to obtain a target depth image corresponding to the image to be predicted; wherein the first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism including first convolution kernel attention, input channel attention, and output channel attention. Therefore, this application embodiment can input the acquired image to be predicted into the first model for monocular depth estimation, and the first model includes a first module that can perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. That is, this application embodiment introduces an attention mechanism in three dimensions—convolution kernel size, input channel, and output channel—so that dynamic convolution can be performed on the first image features corresponding to the image to be predicted based on these three dimensions of attention mechanism, thereby greatly reducing computational complexity and improving the prediction efficiency of the first model. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the monocular depth estimation method proposed in the embodiments of this application. Figure 1 ;
[0020] Figure 2 This is a schematic diagram of the composition structure of the initial first model proposed in the embodiments of this application;
[0021] Figure 3 This is a schematic diagram of the modified cross-attention feature fusion mechanism proposed in the embodiments of this application;
[0022] Figure 4 This is a schematic diagram illustrating the logic of dynamic convolution implementation proposed in the embodiments of this application;
[0023] Figure 5 This is a schematic diagram illustrating the implementation logic of wavelet attention transform proposed in an embodiment of this application;
[0024] Figure 6 This is a schematic diagram of the monocular depth estimation method proposed in the embodiments of this application. Figure 2 ;
[0025] Figure 7 This is a schematic diagram of the composition and structure of the monocular depth estimation device proposed in the embodiments of this application. Figure 1 ;
[0026] Figure 8 This is a schematic diagram of the composition and structure of the monocular depth estimation device proposed in the embodiments of this application. Figure 2 . Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for explaining the relevant application and not for limiting the application. Furthermore, it should be noted that, for ease of description, only the parts related to the relevant application are shown in the accompanying drawings.
[0028] Current monocular depth estimation algorithms include Metric3Dv2 and Depth-Anything. Metric3Dv2, by introducing a unified camera space transformation module, enables the model to generalize well across a large number of different camera datasets, effectively improving depth estimation accuracy. However, its computational complexity remains high, making it difficult to use in real-time applications. Depth-Anything uses a large-scale unlabeled dataset and trains its depth estimation algorithm through self-supervised learning. Its main advantages lie in data diversity and improved generalization ability. However, Depth-Anything relies on large-scale datasets for training, resulting in a massive computational burden, and its applicability in certain specific scenarios (such as edge devices) still needs optimization.
[0029] Depth estimation obtains depth information by determining the spatial distance from the visual sensor to the observed scene. This technology has been widely applied in various fields such as 3D reconstruction, visual navigation, and obstacle detection. Although binocular depth estimation methods can achieve good accuracy, they often require high-end hardware. Furthermore, processing data simultaneously from two cameras not only presents the problem of positional discrepancies but also demands significant computational resources.
[0030] In summary, current depth estimation algorithms still suffer from excessive computational resource consumption and high computational complexity, leading to decreased prediction efficiency.
[0031] To address the high computational complexity of current depth estimation algorithms, this application provides a monocular depth estimation method, apparatus, storage medium, and computer program product. The method includes: acquiring an image to be predicted; inputting the image to be predicted into a first model for monocular depth estimation to obtain a target depth image corresponding to the image to be predicted; wherein the first model includes a first module, which is at least used for dynamic convolution processing of first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, including first convolution kernel attention, input channel attention, and output channel attention. Therefore, this application embodiment can input the acquired image to be predicted into the first model for monocular depth estimation. The first model includes a first module that can perform dynamic convolution processing of the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. That is, this application embodiment introduces an attention mechanism in three dimensions—convolution kernel size, input channel, and output channel—to dynamically convolve the first image features corresponding to the image to be predicted based on these three dimensions of attention, thereby significantly reducing computational complexity and improving the prediction efficiency of the first model.
[0032] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0033] This application provides a monocular depth estimation method. Figure 1 This is a schematic diagram of the monocular depth estimation method proposed in the embodiments of this application. Figure 1 ,like Figure 1 As shown, the monocular depth estimation method may include the following steps:
[0034] Step 101: Obtain the image to be predicted.
[0035] In the embodiments of this application, the monocular depth estimation device can acquire the image to be predicted.
[0036] It should be noted that, in the embodiments of this application, the image to be predicted can be a single red-green-blue (RGB) image, and this application does not specifically limit the type of image to be predicted.
[0037] Step 102: Input the image to be predicted into the first model to perform monocular depth estimation to obtain the target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism including first convolution kernel attention, input channel attention and output channel attention.
[0038] In the embodiments of this application, after acquiring the image to be predicted, the monocular depth estimation device can input the image to be predicted into the first model to perform monocular depth estimation, thereby obtaining the target depth image corresponding to the image to be predicted.
[0039] It should be noted that, in the embodiments of this application, the first model can be obtained by training an initial first model based on a first loss function and a training dataset; wherein, the training dataset includes L RGB images and the corresponding real depth images of the RGB images, where L is a positive integer.
[0040] It should be noted that, in the embodiments of this application, Figure 2 This is a schematic diagram of the composition structure of the initial first model proposed in the embodiments of this application, as shown below. Figure 2 As shown, the initial first model may include an initial encoder, an initial decoder, an initial first module, and an initial second module. This application does not specifically limit the number and type of modules included in the initial first model.
[0041] It should be noted that, in the embodiments of this application, the initial first module may be an initial local multidimensional convolutional attention module, which can be used to perform dynamic convolution processing on RGB images based on a cross-dimensional attention mechanism.
[0042] It should be noted that, in the embodiments of this application, the initial second module may be an initial wavelet attention transformation module, which can be used to perform dimensionality reduction processing on the image features of RGB images. This application does not specifically limit the module type of the initial second module.
[0043] It should be noted that in the embodiments of this application, the first loss function is associated with the predicted depth image corresponding to the RGB image and the real depth image corresponding to the RGB image, and the first loss function L can be as shown in the following formula (1).
[0044] (1)
[0045] in, =0.85, =10, This represents the true depth image corresponding to the RGB image. This represents the predicted depth image corresponding to the RGB image, and N represents the number of samples included in the training dataset.
[0046] In other words, in the embodiments of this application, the monocular depth estimation device can first train an initial first model based on a first loss function and a training dataset to obtain a trained first model, and then perform prediction processing on the target depth image corresponding to the image to be predicted based on the trained first model.
[0047] It should be noted that, in the embodiments of this application, the first model can be used to perform monocular depth estimation on the image to be predicted. The first model may include a first module. This application does not specifically limit the number and type of modules included in the first model.
[0048] It should be noted that, in the embodiments of this application, the first module may be a local multidimensional convolutional attention module. The first module is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, which can greatly reduce the computational complexity. This application does not specifically limit the type of the first module.
[0049] It should be noted that, in the embodiments of this application, the first image feature may include image features after the decoder integrates the features extracted by the encoder at different levels.
[0050] It should be noted that, in the embodiments of this application, the first model may include an encoder and a decoder in addition to the first module. This application does not specifically limit the number and type of modules included in the first model.
[0051] Optionally, in the embodiments of this application, when the monocular depth estimation device inputs the image to be predicted into the first model to perform monocular depth estimation and obtains the target depth image corresponding to the image to be predicted, it can first input the image to be predicted into the encoder for feature extraction to obtain N second image features; where N is a positive integer; then, it can perform fusion processing on the N second image features based on the deformable cross-attention mechanism in the decoder to obtain the first image features; and then, it can generate the target depth image based on the first image features.
[0052] For example, in the embodiments of this application, when the monocular depth estimation device inputs the image to be predicted to the encoder for feature extraction, it may use a multi-layer convolutional neural network for feature extraction. This application does not specifically limit the structure of the encoder.
[0053] Furthermore, in the embodiments of this application, after the monocular depth estimation device extracts features from the image to be predicted through the encoder to obtain N second image features, it can perform fusion processing on the N second image features based on the deformable cross-attention mechanism in the decoder, thereby obtaining the first image features.
[0054] Optionally, in embodiments of this application, Figure 3 This is a schematic diagram of the modified cross-attention feature fusion mechanism proposed in the embodiments of this application, as shown below. Figure 3 As shown, when a monocular depth estimation device fuses N second image features based on a deformable cross-attention mechanism in the decoder, it is assumed that N is of size 2, and one of the second image features is... Figure 3 The query feature, another second image feature is Figure 3 The Value feature in the algorithm involves convolving and encoding the Query feature to obtain convolved and encoded image features. These convolved and encoded features are then concatenated and fed into two linear layers. The offset output from one of these linear layers is combined with the convolved Value feature, and this combined feature is then fused with the output of the other linear layer. This fused feature is then fed into a linear layer to obtain the first image feature. These represent the height, width, and channels of an image feature, respectively. For example, a color image has three channels: red, green, and blue. D represents the depth of the image feature.
[0055] In other words, in the embodiments of this application, when the monocular depth estimation device performs fusion processing on different second image features through the decoder, it can selectively integrate features from different levels of the encoder by using a layer-by-layer attention mechanism, thereby improving the depth estimation capability of the model. Furthermore, the deformable cross-attention mechanism can focus only on some key information in the second image features, thereby significantly reducing computational complexity.
[0056] For example, in an embodiment of this application, the monocular depth estimation device can perform fusion processing on different second image features using the following formula (2) to obtain the first image features. .
[0057] (2)
[0058] in, Represents the number of attention heads, This represents the number of sampling points (much smaller than the number of pixels in the feature map, HW). , , For trainable parameters, For attention weights, This represents the image features after the Query feature convolution process. This represents the image features after the Query features have been encoded. Indicates query features, Represents the Value feature. This represents the image features after convolution of the Query features. This represents the first image feature.
[0059] It should be noted that, in the embodiments of this application, after the monocular depth estimation device performs fusion processing on N second image features based on the deformable cross-attention mechanism in the decoder to obtain the first image features, it can generate a target depth image based on the first image features.
[0060] It should be noted that, in the embodiments of this application, the first model may also include a second module, which may be a wavelet attention transform module. This application does not specifically limit the structural type of the second module.
[0061] Optionally, in an embodiment of this application, when the monocular depth estimation device generates a target depth image based on the first image features, it can input the first image features into the first module for correction processing to obtain the processed second image features, and input the first image features into the second module for dimensionality reduction processing to obtain the processed third image features; then the processed second image features and the processed third image features can be combined to obtain the target depth image.
[0062] Optionally, in an embodiment of this application, when the monocular depth estimation device inputs the first image features into the first module for correction processing to obtain the processed second image features, it can perform fusion processing on the first convolutional kernel attention, input channel attention, and output channel attention based on the first image features to obtain the target attention weight; then the target attention weight can be combined with the first image features to obtain the processed second image features.
[0063] Optionally, in the embodiments of this application, when the monocular depth estimation device performs fusion processing on the first convolution kernel attention, input channel attention, and output channel attention based on the first image features to obtain the target attention weight, it can perform convolution processing on the first image features to obtain the convolved fourth image features; then the first attention weight can be determined based on the fourth image features and the input channel attention.
[0064] Furthermore, the second attention weight can be determined based on the first attention weight and the second convolutional kernel dimension attention; wherein, the second convolutional kernel dimension attention is determined based on the first convolutional kernel attention and the first parameter, the first parameter being used to learn information about the first image features; finally, the target attention weight can be determined based on the second attention weight and the output channel attention.
[0065] For example, in an embodiment of this application, the monocular depth estimation device can calculate the processed second image features using the following formula (3). .
[0066] (3)
[0067] in, Indicates the first image feature, This indicates attention along the convolution kernel dimension. Indicates input channel attention. Indicates output channel attention. Indicates learnable parameters, This represents the processed second image feature. This represents the dot product operator.
[0068] Exemplary, in an embodiment of this application, Figure 4 This is a schematic diagram illustrating the logic of dynamic convolution implementation proposed in the embodiments of this application, such as... Figure 4 As shown, the first image features ( The convolution process is performed to obtain the fourth image feature, which can then be combined with the input channel attention (...). The first attention weight is obtained by performing a dot product operation on the first convolutional kernel dimension attention. Then, the first attention weight is convolved with the second convolutional kernel dimension attention to obtain the second attention weight. The second convolutional kernel dimension attention is based on the first convolutional kernel dimension attention. ), first parameter ( The second attention weight is obtained by performing a dot product operation, and then the second attention weight can be combined with the output channel attention ( The dot product is performed to obtain the target attention weights, which can then be combined with the first image features to obtain the processed second image features. ).
[0069] In other words, in the embodiments of this application, an attention mechanism can be introduced in three dimensions: convolutional kernel size, input channel, and output channel. This allows the weights of the convolutional kernels to be dynamically adjusted by learning the features of the input image, thereby improving the network's adaptability to different scenarios. Furthermore, when the first module corrects the first image features, the first convolutional kernel attention, input channel attention, and output channel attention can be fused to obtain the fused target attention weights. These target attention weights can then be combined with the first image features, thereby enhancing the feature extraction capability of the first model without significantly increasing the number of parameters.
[0070] It should be noted that, in the embodiments of this application, the monocular depth estimation device can input the first image features into the first module for correction processing to obtain the processed second image features, and can also input the first image features into the second module for dimensionality reduction processing to obtain the processed third image features.
[0071] Optionally, in an embodiment of this application, when the monocular depth estimation device inputs the first image features into the second module for dimensionality reduction processing to obtain the processed third image features, it can perform discrete wavelet transform to separate the first image features to obtain M first image sub-features; where M is a positive integer; then, based on multi-head self-attention and the M first image sub-features, it can determine M second image sub-features; where the second image sub-features contain local feature information of the first image sub-features; then, the first image features can be subjected to preset operation processing with the M second image sub-features respectively to obtain the processed third image features.
[0072] For example, in the embodiments of this application, the first image sub-feature may include high-frequency detail information obtained by separating the first image features through discrete wavelet transform, or it may include other information. This application does not specifically limit the type of feature information included in the first image sub-feature.
[0073] Exemplary, in an embodiment of this application, Figure 5 This is a schematic diagram illustrating the implementation logic of wavelet attention transform proposed in an embodiment of this application, as shown below. Figure 5 As shown, the monocular depth estimation device can perform image estimation on the first image features ( The separation process is performed to obtain the first image sub-features after separation. Then, the first image sub-features can be passed through two linear layers and a multi-head attention module to obtain the second image sub-features. That is, the second image sub-feature can focus on the local detail information of the first image sub-feature, and then the image features input to the linear layer ( ) respectively with the second image sub-features ( The features are multiplied to obtain the fused image features. The fused feature vector can then be connected with the feature vector obtained by integrating the first image sub-features and processed through a linear layer to obtain the final third image features.
[0074] In other words, in the embodiments of this application, the monocular depth estimation device can be based on lossless downsampling of wavelet transform. That is, the embodiments of this application can use discrete wavelet transform (DWT) to replace the traditional pooling method, which can reduce the computational complexity while retaining high-frequency detail information. In addition, during the wavelet attention transformation process, global context information (i.e., the second image sub-feature) and local detail information (i.e., the image features of the input linear layer) can be effectively fused, thereby improving the accuracy of pixel-level classification.
[0075] For example, in an embodiment of this application, the calculation formula of the second module (i.e., the wavelet attention transformation module) is as shown in the following formula (4).
[0076] (4)
[0077] in, This represents the feature map recovered after inverse wavelet transform. represents the trainable parameters, multi-head attention is used to capture long-range dependencies, and X represents the first image feature.
[0078] Furthermore, in the embodiments of this application, after the monocular depth estimation device inputs the first image features into the first module for correction processing to obtain the processed second image features, and inputs the first image features into the second module for dimensionality reduction processing to obtain the processed third image features, the processed second image features and the processed third image features can be combined to obtain the target depth image.
[0079] Optionally, in an embodiment of this application, the monocular depth estimation device can combine the processed second image features and the processed third image features using the following formula (5).
[0080] (5)
[0081] in, This represents the center value of the i-th depth interval (i.e., the processed second image feature). This represents the corresponding probability value (i.e., the processed third image feature). This represents the depth image of the target.
[0082] It should be noted that in the embodiments of this application, the first model performs monocular depth estimation on the image to be predicted, significantly improving the computational efficiency and prediction accuracy of monocular depth estimation. First, the deformable cross-attention decoder can reduce redundant computation, enabling the model to run efficiently and making it suitable for embedded devices or real-time applications. Second, experiments show that on the NYU and KITT datasets, this application reduces the absolute relative error (AbsRel) by 11.7% and 10.3% respectively compared to related depth estimation methods, and requires fewer parameters. Furthermore, the introduction of wavelet attention transform allows the model to better preserve high-frequency information, reducing depth prediction error and thus improving overall accuracy. Therefore, the first model proposed in this application has a lighter design, can run on devices with limited computing resources, and maintains high performance.
[0083] In summary, the monocular depth estimation device can first train an initial first model based on a first loss function and a training dataset to obtain a trained first model. Then, the image to be predicted can be input into the encoder in the first model for feature extraction to obtain N second image features, where N is a positive integer. Then, the N second image features can be fused based on the deformable cross-attention mechanism in the decoder to obtain the first image features, thereby improving the depth estimation capability of the model. Moreover, the deformable cross-attention mechanism can focus only on some key information in the second image features, thus significantly reducing computational complexity. Then, the first image features can be input into the first module for correction processing to obtain the processed second image features, and the first image features can be input into the second module for dimensionality reduction processing to obtain the processed third image features. Finally, the processed second image features and the processed third image features can be combined to obtain the target depth image. That is, the embodiments of this application, through the deformable cross-attention decoder, the first module, and the second module in the first model, can achieve high-precision depth estimation while maintaining high computational efficiency, and can be applied to real-time scenarios such as autonomous driving and robot navigation.
[0084] This application provides a monocular depth estimation method, which includes: acquiring an image to be predicted; inputting the image to be predicted into a first model for monocular depth estimation to obtain a target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism including first convolution kernel attention, input channel attention, and output channel attention. Therefore, this application embodiment can input the acquired image to be predicted into the first model for monocular depth estimation, and the first model includes a first module that can perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. That is, this application embodiment introduces an attention mechanism in three dimensions—convolution kernel size, input channel, and output channel—so that dynamic convolution can be performed on the first image features corresponding to the image to be predicted based on these three dimensions of attention mechanism, thereby greatly reducing computational complexity and improving the prediction efficiency of the first model.
[0085] Based on the above embodiments, another embodiment of this application provides a monocular depth estimation method. This method combines deformable attention, local multidimensional convolutional attention, and wavelet attention mechanisms to improve depth prediction accuracy while effectively reducing computational complexity, making it more suitable for resource-constrained devices and real-time application scenarios. This application proposes a monocular depth estimation model, SimMDE (i.e., the first model), to achieve a balance between computational complexity and prediction accuracy. This model employs a DCF decoder, an LMC module (i.e., the first module), and a WAT module (i.e., the second module), which can improve depth estimation accuracy while significantly reducing computational complexity.
[0086] It should be noted that, in the embodiments of this application, Figure 6 This is a schematic diagram of the monocular depth estimation method proposed in the embodiments of this application. Figure 2 ,like Figure 6As shown, SimMDE (i.e., the first model) adopts an encoder-decoder structure and models the depth estimation task as an ordinal regression problem. The first model mainly includes the following key components: Encoder: responsible for extracting multi-scale features (i.e., N second image features) from the input RGB image, using a multi-layer convolutional neural network for feature extraction; Decoder: adopts a deformable cross-attention feature fusion (DCF) mechanism to reduce computational complexity while enhancing multi-scale feature fusion capability; Probability Prediction Branch: adopts a Local Multidimensional Convolutional Attention (LMC) module (i.e., the first module) to improve local feature extraction capability and make depth prediction more accurate; Pixel Classification Branch: adopts a Wavelet Attention Transformer (WAT) module (i.e., the second module) to achieve global information capture without increasing computational complexity, thereby improving pixel-level classification capability. The final depth image can be obtained by combining the outputs of the probability prediction branch and the pixel classification branch, and then displayed in a user interface (UI).
[0087] It should be noted that, in the embodiments of this application, most current decoders rely on bilinear upsampling and convolution operations to recover depth information. However, this method is difficult to effectively reconstruct details in complex scenes. This application proposes a deformable cross-attention feature fusion (DCF) mechanism; for specific implementation details, see [link to implementation details]. Figure 3 ,like Figure 3 As shown, when the monocular depth estimation device fuses N second image features based on the deformable cross-attention mechanism, it is assumed that the size of N is 2, and one of the second image features is... Figure 3 The query feature, another second image feature is Figure 3 The Value feature can be processed by convolution and encoding of the Query feature to obtain convolutional and encoded image features. Then, the convolutional and encoded image features can be concatenated and input into two linear layers. The offset output from one of the linear layers can be combined with the convolutional feature of the Value feature, and the combined feature can be fused with the image features output from the other linear layer. Finally, the fused image features can be input into a linear layer to obtain the first image feature.
[0088] In other words, in the embodiments of this application, when the monocular depth estimation device performs fusion processing on different second image features through the decoder, it can selectively integrate features from different levels of the encoder by using a layer-by-layer attention mechanism, thereby improving the depth estimation capability of the model. Furthermore, the deformable cross-attention mechanism can focus only on some key information in the second image features, thereby significantly reducing computational complexity.
[0089] It should be noted that, in the embodiments of this application, the deformable attention mechanism allows the network to focus on only a small subset of key points, thereby significantly reducing the quadratic complexity problem of the Transformer. Furthermore, during multi-scale feature fusion, DCF can selectively integrate features from different levels of the encoder through a layer-by-layer attention mechanism, improving the model's depth estimation capability.
[0090] It should be noted that in the embodiments of this application, the parameters of the convolutional kernel in traditional convolutional neural networks (CNNs) are fixed and cannot adapt to the features of different input images. Therefore, this application proposes an LMC module (i.e., the first module), the core idea of which is to introduce an attention mechanism in three dimensions: kernel size, input channels, and output channels, to achieve dynamic convolution. For specific implementation details, see [link to implementation details]. Figure 4 The first image features can be convolved to obtain the fourth image features. Then, the fourth image features can be combined with the input channel attention (…). The first attention weight is obtained by performing a dot product operation on the first convolutional kernel dimension attention. Then, the first attention weight is convolved with the second convolutional kernel dimension attention to obtain the second attention weight. The second convolutional kernel dimension attention is based on the first convolutional kernel dimension attention. ), first parameter ( The second attention weight is obtained by performing a dot product operation, and then the second attention weight can be combined with the output channel attention ( The target attention weights are obtained by performing dot product processing. Then, the target attention weights can be combined with the first image features to obtain the processed second image features.
[0091] It should be noted that in the embodiments of this application, an adaptive convolution kernel is proposed, which can dynamically adjust the weight of the convolution kernel by learning the features of the input image, thereby improving the adaptability of the network to different scenarios. Multi-dimensional attention is also proposed. LMC can perform weighted processing in the channel dimension, spatial dimension and feature dimension to improve the effectiveness of convolution operation. The mathematical expression of the LMC module is shown in the above formula (3).
[0092] In other words, in the embodiments of this application, an attention mechanism can be introduced in three dimensions: convolutional kernel size, input channel, and output channel. This allows the weights of the convolutional kernels to be dynamically adjusted by learning the features of the input image, thereby improving the network's adaptability to different scenarios. Furthermore, when the first module corrects the first image features, the first convolutional kernel attention, input channel attention, and output channel attention can be fused to obtain the fused target attention weights. These target attention weights can then be combined with the first image features, thereby enhancing the feature extraction capability of the first model without significantly increasing the number of parameters.
[0093] It should be noted that in the embodiments of this application, current Transformer architectures generally use pooling for dimensionality reduction to reduce computational complexity, but this often leads to information loss. Therefore, the embodiments of this application propose a WAT module (i.e., the second module), the specific implementation details of which can be found in [link to implementation details]. Figure 5 ,like Figure 5 As shown, the monocular depth estimation device can perform image estimation on the first image features ( The separation process is performed to obtain the first image sub-features after separation. Then, the first image sub-features can be passed through two linear layers and a multi-head attention module to obtain the second image sub-features. That is, the second image sub-feature can focus on the local detail information of the first image sub-feature, and then the image features input to the linear layer ( ) respectively with the second image sub-features ( The features are multiplied to obtain the fused image features. The fused feature vector can then be connected with the feature vector obtained by integrating the first image sub-features and processed through a linear layer to obtain the final third image features.
[0094] In other words, in the embodiments of this application, the monocular depth estimation device can be based on lossless downsampling of wavelet transform. That is, the embodiments of this application can use discrete wavelet transform to replace the traditional pooling method, which can reduce the computational complexity while retaining high-frequency detail information. In the wavelet attention transformation process, global context information (i.e., the second image sub-feature) and local detail information (i.e. the image features of the input linear layer) can be effectively fused, thereby improving the accuracy of pixel-level classification.
[0095] For example, in an embodiment of this application, the calculation formula for the second module is as shown in formula (4) above.
[0096] Furthermore, in the embodiments of this application, the final depth estimate (i.e. the target depth image) can be obtained by combining the outputs of the probability prediction branch and the pixel classification branch. That is, the processed second image features and the processed third image features can be combined to obtain the target depth image, which can be calculated by the above formula (5).
[0097] It should be noted that in the embodiments of this application, the loss function (i.e. the first loss function) can be scale-invariant loss. The first loss function is associated with the predicted depth image corresponding to the RGB image and the real depth image corresponding to the RGB image. The first loss function can be as shown in the above formula (1). That is, the first model proposed in the embodiments of this application can be obtained by training the initial first model based on the first loss function and the training dataset. The training dataset includes L RGB images and the real depth images corresponding to the RGB images.
[0098] It should be noted that, in the embodiments of this application, the DCF decoder, LMC and WAT mechanism in the SimMDE model (i.e. the first model) proposed in this application can achieve high-precision depth estimation while maintaining high computational efficiency, and are suitable for real-time scenarios such as autonomous driving and robot navigation.
[0099] It should be noted that in the embodiments of this application, the first model performs monocular depth estimation on the image to be predicted, significantly improving the computational efficiency and prediction accuracy of monocular depth estimation. First, the deformable cross-attention decoder can reduce redundant computation, enabling the model to run efficiently and making it suitable for embedded devices or real-time applications. Second, experiments show that on the NYU and KITT datasets, this application reduces the absolute relative error (AbsRel) by 11.7% and 10.3% respectively compared to related depth estimation methods, and requires fewer parameters. Furthermore, the introduction of wavelet attention transform allows the model to better preserve high-frequency information, reducing depth prediction error and thus improving overall accuracy. Therefore, the first model proposed in this application has a lighter design, can run on devices with limited computing resources, and maintains high performance.
[0100] In summary, the monocular depth estimation device can first train an initial first model based on a first loss function and a training dataset to obtain a trained first model. Then, the image to be predicted can be input into the encoder in the first model for feature extraction to obtain N second image features, where N is a positive integer. Then, the N second image features can be fused based on the deformable cross-attention mechanism in the decoder to obtain the first image features, thereby improving the depth estimation capability of the model. Moreover, the deformable cross-attention mechanism can focus only on some key information in the second image features, thus significantly reducing computational complexity. Then, the first image features can be input into the first module for correction processing to obtain the processed second image features, and the first image features can be input into the second module for dimensionality reduction processing to obtain the processed third image features. Finally, the processed second image features and the processed third image features can be combined to obtain the target depth image. That is, the embodiments of this application, through the deformable cross-attention decoder, the first module, and the second module in the first model, can achieve high-precision depth estimation while maintaining high computational efficiency, and can be applied to real-time scenarios such as autonomous driving and robot navigation.
[0101] This application provides a monocular depth estimation method, which includes: acquiring an image to be predicted; inputting the image to be predicted into a first model for monocular depth estimation to obtain a target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism including first convolution kernel attention, input channel attention, and output channel attention. Therefore, this application embodiment can input the acquired image to be predicted into the first model for monocular depth estimation, and the first model includes a first module that can perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. That is, this application embodiment introduces an attention mechanism in three dimensions—convolution kernel size, input channel, and output channel—so that dynamic convolution can be performed on the first image features corresponding to the image to be predicted based on these three dimensions of attention mechanism, thereby greatly reducing computational complexity and improving the prediction efficiency of the first model.
[0102] Based on the above embodiments, this application provides a monocular depth estimation device. Figure 7 Schematic diagram of the composition of a monocular depth estimation device Figure 1 ,like Figure 7 As shown, the device 10 includes: an acquisition unit 11 and an input unit 12; wherein,
[0103] The acquisition unit 11 is used to acquire the image to be predicted;
[0104] The input unit 12 is used to input the image to be predicted into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, the first module is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism includes a first convolution kernel attention, input channel attention and output channel attention.
[0105] In the embodiments of this application, further, Figure 8 Schematic diagram of the composition of a monocular depth estimation device Figure 2 ,like Figure 8 As shown, the monocular depth estimation device 10 proposed in this application embodiment may further include a processor 13, a memory 14 storing instructions executable by the processor 13, and further, the device 10 may also include a communication interface 15 and a bus 16 for connecting the processor 13, the memory 14 and the communication interface 15.
[0106] In the embodiments of this application, the processor 13 can be at least one of the following: Application-Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that for different devices, the electronic device used to implement the above-mentioned processor function can also be other types, and this application embodiment does not specifically limit the specific types. The device 10 may also include a memory 14, which can be connected to the processor 13. The memory 14 is used to store executable program code, which includes computer operation instructions. The memory 14 may include high-speed RAM memory and may also include non-volatile memory, such as at least two disk drives.
[0107] In embodiments of this application, bus 16 is used to connect communication interface 15, processor 13 and memory 14 and the mutual communication between these devices.
[0108] In embodiments of this application, memory 14 is used to store instructions and data.
[0109] Furthermore, in the embodiments of this application, the processor 13 is used to acquire the image to be predicted;
[0110] The image to be predicted is input into a first model for monocular depth estimation to obtain a target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism including a first convolution kernel attention, input channel attention and output channel attention.
[0111] In practical applications, the aforementioned memory 14 can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 13.
[0112] This application provides a monocular depth estimation device. The device acquires an image to be predicted; it inputs the image to be predicted into a first model for monocular depth estimation to obtain a target depth image corresponding to the image to be predicted. The first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. This cross-dimensional attention mechanism includes first convolution kernel attention, input channel attention, and output channel attention. Therefore, this application allows the acquired image to be predicted to be input into a first model for monocular depth estimation. The first model includes a first module that performs dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. Specifically, this application introduces an attention mechanism in three dimensions: convolution kernel size, input channels, and output channels. This allows for dynamic convolution of the first image features corresponding to the image to be predicted based on these three dimensions, significantly reducing computational complexity and improving the prediction efficiency of the first model.
[0113] This application provides a computer-readable storage medium storing a program that, when executed by a processor, implements the monocular depth estimation method described above.
[0114] Specifically, the program instructions corresponding to a monocular depth estimation method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the program instructions corresponding to a monocular depth estimation method in the storage media are read or executed by an electronic device, the following steps are included:
[0115] Obtain the image to be predicted;
[0116] The image to be predicted is input into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted.
[0117] The first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism. The cross-dimensional attention mechanism includes a first convolution kernel attention, input channel attention, and output channel attention.
[0118] This application also provides a computer program product, including a computer program that can be executed by the processor 13 of the monocular depth estimation device 10 to perform the steps described in any of the foregoing methods.
[0119] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0120] This application is described with reference to schematic and / or block diagrams of implementations of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the schematic and / or block diagrams can be implemented by computer program instructions, and combinations of blocks in the schematic and / or block diagrams can be implemented. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the schematic and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0121] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the implementation flow diagram. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0122] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0123] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.
Claims
1. A monocular depth estimation method, characterized in that, The method includes: Obtain the image to be predicted; The image to be predicted is input into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, which is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism including a first convolution kernel attention, input channel attention and output channel attention, and the first model also includes an encoder and a decoder; The step of inputting the image to be predicted into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted includes: The image to be predicted is input into the encoder for feature extraction to obtain N second image features; where N is a positive integer; The N second image features are fused based on the deformable cross-attention mechanism in the decoder to obtain the first image features; Generate the target depth image based on the first image features; The first model further includes a second module; the step of generating the target depth image based on the first image features includes: The first image features are input into the first module for correction processing to obtain the processed second image features, and the first image features are input into the second module for dimensionality reduction processing to obtain the processed third image features. The processed second image features and the processed third image features are combined to obtain the target depth image; The step of inputting the first image features into the second module for dimensionality reduction processing to obtain the processed third image features includes: The first image features are separated by discrete wavelet transform to obtain M first image sub-features; where M is a positive integer. Based on multi-head self-attention, M second image sub-features are determined from the M first image sub-features; wherein, the second image sub-features contain local feature information of the first image sub-features; The first image feature is processed by a preset operation with each of the M second image sub-features to obtain the processed third image feature.
2. The method according to claim 1, characterized in that, The step of inputting the first image features into the first module for correction processing to obtain the processed second image features includes: Based on the first image features, the first convolutional kernel attention, the input channel attention, and the output channel attention are fused to obtain the target attention weights; The target attention weights are combined with the first image features to obtain the processed second image features.
3. The method according to claim 2, characterized in that, The step of fusing the first convolutional kernel attention, the input channel attention, and the output channel attention based on the first image features to obtain the target attention weights includes: The first image feature is convolved to obtain the fourth image feature after convolution. The first attention weight is determined based on the fourth image feature and the input channel attention. The second attention weight is determined based on the first attention weight and the second convolutional kernel dimension attention; wherein, the second convolutional kernel dimension attention is determined based on the first convolutional kernel attention and the first parameter, and the first parameter is used to learn information about the first image features; The target attention weight is determined based on the second attention weight and the output channel attention.
4. The method according to claim 1, characterized in that, The first model is obtained by training an initial first model based on a first loss function and a training dataset; wherein, the training dataset includes L red-green-blue images and the corresponding real depth images of the red-green-blue images, and L is a positive integer; The initial first model includes an initial encoder, an initial decoder, an initial first module, and an initial second module. The first loss function is associated with the predicted depth image corresponding to the red-green-blue image and the real depth image corresponding to the red-green-blue image.
5. A monocular depth estimation device, characterized in that, The monocular depth estimation device includes: an acquisition unit and an input unit; wherein... The acquisition unit is used to acquire the image to be predicted; The input unit is used to input the image to be predicted into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted; wherein, the first model includes a first module, the first module is at least used to perform dynamic convolution processing on the first image features corresponding to the image to be predicted based on a cross-dimensional attention mechanism, the cross-dimensional attention mechanism includes a first convolution kernel attention, input channel attention and output channel attention, and the first model also includes an encoder and a decoder. The step of inputting the image to be predicted into the first model for monocular depth estimation to obtain the target depth image corresponding to the image to be predicted includes: The image to be predicted is input into the encoder for feature extraction to obtain N second image features; where N is a positive integer; The N second image features are fused based on the deformable cross-attention mechanism in the decoder to obtain the first image features; Generate the target depth image based on the first image features; The first model further includes a second module; the step of generating the target depth image based on the first image features includes: The first image features are input into the first module for correction processing to obtain the processed second image features, and the first image features are input into the second module for dimensionality reduction processing to obtain the processed third image features. The processed second image features and the processed third image features are combined to obtain the target depth image; The step of inputting the first image features into the second module for dimensionality reduction processing to obtain the processed third image features includes: The first image features are separated by discrete wavelet transform to obtain M first image sub-features; where M is a positive integer. Based on multi-head self-attention, M second image sub-features are determined from the M first image sub-features; wherein, the second image sub-features contain local feature information of the first image sub-features; The first image feature is processed by a preset operation with each of the M second image sub-features to obtain the processed third image feature.
6. A monocular depth estimation device, characterized in that, The monocular depth estimation device includes: a processor and a memory; wherein... The memory is used to store computer programs that can run on the processor; The processor is configured to perform the method as described in any one of claims 1-4 when running the computer program.
7. A computer-readable storage medium, characterized in that, The storage medium stores computer program code, which, when executed by a computer, performs the method described in any one of claims 1-4.
8. A computer program product, comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method according to any one of claims 1-4.
Citation Information
Patent Citations
Monocular depth estimation method based on attention feature fusion and multistage correction
CN116883476A
Multi-dimensional attention for dynamic convolution kernels
CN119072689A