Monocular depth estimation method and system based on text guidance and multi-scale fusion
By introducing text information and multi-scale fusion technology in monocular depth estimation, the problem of existing methods relying on image features and ignoring text information is solved, achieving more accurate depth prediction and stronger robustness.
Patent Information
- Application Number
- CN202510111292.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-23
AI Technical Summary
In the prior art, monocular depth estimation methods often rely solely on image features, ignoring the potential value of text information, resulting in limitations in processing complex scenes, diverse objects, and capturing global and local semantic information.
A monocular depth estimation method based on text guidance and multi-scale fusion is adopted. By acquiring the visual features of the original image and the text features of the original text, the text features are mapped to the same dimension as the image features, and introduced into the image features, performing bidirectional cross-guiding, combining multi-scale Laplace residual calculation, feature fusion is performed to generate the estimated depth image.
It realizes more accurate depth prediction, captures the edge characteristics of the image on different scales, improves the model's expressive ability and robustness of local information, and enables the model to understand the local information and global scene layout of details at the same time, adapt to complex scenes, and improves the recovery effect of depth boundaries and details.
Smart Images

Figure CN119941816A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image recognition, and relates to a monocular depth estimation method and system based on text guidance and multi-scale fusion. Background Art
[0002] Depth estimation from monocular images has long been a key task in many practical applications and is widely used in many fields such as robotics, autonomous driving, virtual reality, etc. Due to these rich possibilities, many researchers have devoted a lot of effort to solving the monocular depth estimation problem.
[0003] Since the advent of deep learning, deep neural networks have been successfully applied to various visual tasks and achieved remarkable results. Many excellent works based on deep neural networks have also appeared in the field of monocular depth estimation, demonstrating its great potential in solving such visual tasks. Although deep neural network-based methods have demonstrated a strong ability to reveal depth layout without the need for domain knowledge, they still face many challenges, including scene complexity, diversity, and variability in object shapes and sizes. Specifically, most existing methods use visual features that mainly rely on images for depth prediction, which are simply upsampled back to the original size through a decoding process in a symmetric architecture and finally converted into a depth map. Although progress has been made to some extent, there are still limitations in dealing with complex scenes and subtle details.
[0004] In recent years, multimodal learning methods have shown great potential in improving depth estimation performance. For example, VPD introduces the fusion of visual and textual information, and achieves more accurate depth prediction by jointly learning visual features and textual descriptions. WorDepth uses text descriptions as semantic priors to improve the accuracy of depth estimation through conditional generation models. In addition, CLIPDepth innovatively combines the semantic knowledge of the CLIP model with the depth estimation task, converting the depth value regression problem into a distance classification problem. This method achieves zero-shot depth estimation, surpassing existing unsupervised methods without the need for additional training data. These studies show that fusing information from multiple modalities can significantly improve the performance of depth estimation tasks, especially when dealing with complex and diverse scenes.
[0005] With the rapid development of deep learning and computer vision, monocular depth estimation has made significant progress in the research of various encoder-decoder architectures. However, traditional depth estimation methods often rely only on image features and often ignore the potential value of text information. This leads to limitations in dealing with complex scenes, diverse objects, and capturing global and local semantic information, making monocular depth estimation an inherently ill-posed problem. Summary of the invention
[0006] The purpose of the present invention is to solve the problem that traditional depth estimation methods in the prior art often rely only on image features and often ignore the potential value of text information. This leads to limitations in processing complex scenes, diverse objects, and capturing global and local semantic information. A monocular depth estimation method and system based on text guidance and multi-scale fusion is provided.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A monocular depth estimation method based on text guidance and multi-scale fusion includes the following steps:
[0009] Obtaining visual features of the original image and text features of the original text;
[0010] Mapping the text features to the same dimension as the visual features of the original image to obtain processed text features, introducing the processed text features into the visual features of the original image to obtain image features with semantic guidance, and enhancing the text features based on the visual features to obtain enhanced text features;
[0011] By introducing enhanced text features into image features with semantic guidance through bidirectional cross-guidance, the final enhanced image features are obtained to obtain an enhanced image;
[0012] Multi-scale Laplace residual calculation is performed on the original image and the enhanced image to obtain residual information between the original image and the enhanced image, and a fusion operation is performed based on the residual information and the enhanced image features to obtain an estimated depth image.
[0013] A further improvement of the present invention is:
[0014] The text features are mapped to the same dimension as the image information, and the processed text features ti are obtained by the following formula:
[0015] t i =Proj i ·t
[0016] Among them, Proj i represents the projection matrix corresponding to scale i, whose dimension is 1024×Ci; Ci is the number of channels of the i-th layer image feature;
[0017] The method of introducing the processed text features into the visual features of the original image to obtain image features with semantic guidance, and enhancing the text features based on the visual features of the original image to obtain enhanced text features includes:
[0018] Use text features ti as queries and image features fi as keys and values, and enhance them through the cross-attention mechanism to generate semantically guided image features;
[0019] The query, key, and value are transformed through multiple different linear transformations, and then the attention of multiple heads is calculated respectively, and the output results of multiple heads are spliced to obtain enhanced image features.
[0020] The enhanced image features It is expressed by the following formula:
[0021]
[0022] Where H represents the number of heads, W o Represents the output linear transformation matrix; Ai represents the output result of the head;
[0023] The whole process is defined as cross attention:
[0024]
[0025] Among them, T i Represents the text features after projection; Represents the flattening of image features.
[0026] The text features are enhanced based on the visual features of the original image to obtain enhanced text features, including:
[0027]
[0028] Among them, T i represents the text features after projection; Represents the flattening of image features.
[0029] The enhanced text features are introduced into the image features with semantic guidance through bidirectional cross-guidance to obtain the final enhanced image features, including:
[0030] Enhanced text features As query input, the enhanced image features As keys and values, further enhanced image features are obtained
[0031]
[0032] in, represents enhanced image features; Represents enhanced text features;
[0033] The enhanced image features are concatenated with the original image features, and then a 2D convolution operation is performed;
[0034] The feature representation is further optimized using weighted residual connections to obtain the final enhanced features.
[0035] The obtaining of residual information between the original image and the enhanced image includes:
[0036] Calculate the Laplace residual L of the original image k :
[0037] L k =I k -Up(I k+1 )
[0038] The multi-scale Laplace residual is calculated for the enhanced image to obtain the enhanced Laplace residual corresponding to the original image:
[0039]
[0040] Among them, en anced_image k represents the enhanced image;
[0041] Compute the difference residual between the original image and the enhanced image:
[0042]
[0043] Among them, L k Represents the Laplace residual of the kth layer of the original image; represents the Laplace residual of the kth layer of the enhanced image; ΔL k Represents the difference between the Laplace residual of the enhanced image and the original image at the kth level.
[0044] Based on the residual information and enhanced image features, a fusion operation is performed to obtain an estimated depth image, including:
[0045] For each scale k, these residual features are weighted through the channel attention mechanism and the spatial attention mechanism to obtain the weighted residual features;
[0046] The weighted residual features are concatenated and the fused features are obtained through convolutional layer fusion to obtain the fused residual feature representation;
[0047] The fused residual feature representation and enhanced image features are input into the decoder to generate an estimated depth image.
[0048] A monocular depth estimation system based on text guidance and multi-scale fusion, comprising:
[0049] An initial feature extraction module, used to obtain the visual features of the original image and the text features of the original text;
[0050] A feature enhancement module is used to map text features to the same dimension as the visual features of the original image to obtain processed text features, introduce the processed text features into the visual features of the original image to obtain image features with semantic guidance, enhance the text features based on the visual features to obtain enhanced text features, introduce the enhanced text features into the image features with semantic guidance through bidirectional cross-guidance to obtain final enhanced image features, and obtain an enhanced image;
[0051] The residual calculation module is used to perform multi-scale Laplace residual calculation on the original image and the enhanced image, obtain the residual information between the original image and the enhanced image, and perform a fusion operation based on the residual information and the enhanced image features to obtain an estimated depth image.
[0052] A terminal device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any method described in the present invention when executing the computer program.
[0053] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any method described in the present invention.
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] The present invention discloses a monocular depth estimation method based on text guidance and multi-scale fusion. By dynamically fusing text features to enhance image features, richer semantic information is obtained, and more accurate depth prediction is achieved. By capturing edge features of images at different scales and performing multi-modal alignment with enhanced image features, the model's expression ability and robustness for local information are further improved, so that the model can understand detailed local information and global scene layout at the same time. This method can not only adapt to complex scenes, but also improve the restoration effect of image depth boundaries and details. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0057] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0058] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0059] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0060] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.
[0061] In the description of the embodiments of the present invention, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. indicate an orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.
[0062] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0063] In the description of the embodiments of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0064] The present invention is further described in detail below in conjunction with the accompanying drawings:
[0065] See also Figure 1 , the present invention discloses a monocular depth estimation method based on text guidance and multi-scale fusion, aiming to achieve high-precision monocular depth estimation. A new model framework, TFDepth (text-guided multi-scale fusion network for monocular depth estimation), is proposed. TFDepth uses text features to guide the representation and fusion of image features, thereby obtaining richer semantic information and helping the model to better understand the details and global structure in the image. The framework makes full use of the complementary advantages of multimodal features, thereby enhancing the accuracy and robustness of monocular depth estimation.
[0066] The image and text features are weightedly fused using the cross-semantic attention module (CSAM) and the multi-scale residual fusion module (MSRFM) to accurately capture the image details and global structure. After being encoded, the image and the corresponding text description are integrated into the model as prior knowledge. The weights of the image and text features are dynamically adjusted through the cross-semantic attention mechanism, so that the model can effectively utilize these semantic information in different depth estimation tasks. In order to improve the depth estimation accuracy of the model, this embodiment also uses a multi-scale residual fusion module to optimize the restoration of depth boundaries and details by capturing the fine-grained changes of the image at different resolutions. In addition, the DINOv2 encoder used in the decoder further enhances the expressiveness of image features and ensures the accuracy of the depth estimation results.
[0067] The DINOv2 encoder refers to the encoder disclosed in Oquab M, Darcet T, Moutakanni T, et al. Dinov2: Learning robust visual features without supervision [J]. arXiv preprint arXiv: 2304.07193, 2023.
[0068] The model structure of this method includes a pre-trained encoder, a cross-semantic attention module (CSAM), a multi-scale residual fusion module (MSRFM) and an image decoder.
[0069] Among them, the encoder part is used to extract rich feature information from the input image and text. In terms of image feature extraction, inspired by Depth_Anything (from the document Yang L, Kang B, Huang Z, et al. Depth anything: Unleashing the power of large-scale unlabeled data [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024: 10371-10381), this embodiment migrates the powerful semantic capabilities of the pre-trained DINOv2 model to the depth estimation model of this embodiment, using frozen and non-frozen DINOv2 encoders to make full use of its powerful semantic capabilities to ensure efficient extraction of image features. For the processing of text information, this embodiment uses the text encoder of the pre-trained visual-language model CLIP. CLIP provides a shared potential space that can effectively integrate language priors into image features. Specifically, given a text description t = {t1, t2, ...}, this embodiment converts it into a feature vector of size b × 1024 through the text encoder of CLIP, where b is the batch size. These text features are then dynamically weighted fused with image features through the Cross-Semantic Attention Module (CSAM) to ensure that the information of the two modalities can effectively complement each other at different levels.
[0070] CSAM adopts a multi-head cross-attention mechanism, which uses the interaction between image features and text features to dynamically adjust the weighted proportion of the two, thereby enhancing the semantic information of the image at different scales. It establishes effective connections between image and text features at different levels to ensure the deep complementarity of information, thereby improving the accuracy of depth estimation.
[0071] The decoder effectively integrates features and detail information at multiple scales through the Multi-Scale Residual Fusion Module (MSRFM), and uses the Laplacian residual of the image to help restore the details and depth boundaries in the image. The key point of this module is to effectively fuse the detail information of the original image and the enhanced image by calculating the residual information at different scales, thereby achieving accurate depth estimation in the decoder.
[0072] Finally, after being processed by the multi-scale residual fusion module, all fused multi-scale features are input into the LapDepth (from the literature Song M, Lim S, Kim W. Monocular depth estimation using laplacian pyramid-based depth residuals [J]. IEEE transactions on circuits and systems for video technology, 2021, 31 (11): 4381-4393.) decoder, and the final depth map is gradually restored by recursively synthesizing the depth residual. Generate the final depth map. In the following section, this embodiment will introduce in detail the functions of each module and its role in the entire network, as well as how to improve the accuracy of depth estimation by training and optimizing the loss function.
[0073] The specific steps include:
[0074] Step 1: The RGB image is first processed by data augmentation to improve the generalization ability of the model; then, the data augmented image and the corresponding text description are respectively passed through their respective encoders for feature extraction. Specifically, the image is passed through the pre-trained DINOv2 encoder to extract multi-level visual features, while the text description is passed through the CLIP text encoder to obtain semantic features.
[0075] Specifically include:
[0076] Step 1.1, Text Encoder
[0077] In order to incorporate language priors into monocular depth estimation, this example uses the text encoder of the pre-trained vision-language model CLIP. CLIP provides a shared latent space for visual and text embedding by default. Specifically, given a text description t = {t 1 ,t 2 ...}, this embodiment encodes it through the CLIP text encoder to obtain a feature vector of size b×1024, where b is the batch size. The generated text embedding captures rich semantic information and acts as a bridge to align text descriptions with visual features, facilitating a more context-aware depth estimation process.
[0078] Step 1.2: Frozen and non-frozen encoders extract image features
[0079] Inspired by Depth_Anything, this embodiment believes that introducing high-level semantic information into the depth estimation model can bring significant improvements. Based on this idea, this embodiment migrates the powerful semantic capabilities of the pre-trained DINOv2 model to the depth estimation model of this embodiment, and uses frozen and non-frozen DINOv2 encoders for image feature extraction. In addition, this embodiment also adopts the feature alignment loss proposed in Depth_Anything.
[0080]
[0081] In this loss function, cos represents the cosine similarity between two feature vectors, f is the feature extracted by the deep model S, and f′ is the feature from the frozen DINOv2 encoder.
[0082] Step 2: The features of the image and text are weighted and fused in the cross-semantic attention module to generate a joint feature representation. This fusion module uses a text-image feature fusion module based on the cross-attention mechanism to enhance image features by dynamically fusing text features to achieve more accurate depth prediction. The cross-attention mechanism dynamically weights image features and text features at different scales, thereby effectively combining information from the two modalities and improving the model's understanding of scene semantics and the accuracy of depth prediction.
[0083] The specific steps include:
[0084] In the depth estimation task, multimodal features such as images and text have been shown to significantly improve model performance. In order to effectively fuse these multimodal features and improve the accuracy of depth estimation, this study proposes a cross-semantic attention module based on cross-attention. This module uses the cross-semantic attention module to enhance image and text features at multiple scales, and focuses on key information through a multi-scale residual fusion module, and finally further optimizes the feature representation through weighted residual connections.
[0085] Step 2.1 Text feature projection:
[0086] In order to effectively integrate text information into image features, this embodiment maps the text features extracted by the text encoder to the same dimension as the image features. Given that the dimension of the text features is B×1024, where B is the batch size and 1024 is the dimension of each text feature. In order to enable the text features to be fused with the image features, this embodiment projects the text features onto the number of channels Ci at each scale through a multi-layer linear mapping. Specifically, this embodiment uses the projection matrix Proj i To map the text feature t to the feature space of each scale, we can get the text feature t i :
[0087] t i =Proj i t (2)
[0088] Among them, t i ∈R B×Ci It is the text feature after projection. i is the projection matrix corresponding to scale i, and its dimension is 1024×C i ; C i is the number of channels of the i-th layer image feature. Through this mapping, the dimension of the text feature is consistent with the dimension of the image feature, which is convenient for subsequent feature fusion.
[0089] Step 2.2 Cross-attention operation:
[0090] At each scale, this embodiment uses a multi-head cross-attention mechanism to interact text features with image features, thereby introducing the semantic information of the text into the image features. Specifically, this embodiment uses the text feature t i As the query, the image feature f i As the key and value, the cross-attention mechanism is used for enhancement to generate semantically guided image features Ienhanced. In order to further enhance the relationship between image and text features, this embodiment adopts a multi-head attention mechanism. In each head, the query, key, and value are mapped to different spaces through different linear transformations, and the attention value of each head is calculated. For the hth head, the output is:
[0091]
[0092] in, is the query, which represents the representation of text features after linear transformation;
[0093] is the key, which indicates the representation of image features after linear transformation; It is a value, which represents the representation of image features after linear transformation.
[0094] in, A are the linear transformation matrices of the query, key, and value of each head, with dimensions Ci×Ci, which map the input features to the space adapted to the attention mechanism. h is the output of the hth head. Finally, the final enhanced image features are obtained by concatenating the outputs of all heads and performing linear transformation
[0095]
[0096] Where H is the number of heads, W ois the linear transformation matrix of the output. This embodiment defines the whole process as cross attention:
[0097]
[0098] Among them, T i Represents the projected text feature, with a shape of B×1×C i ; Indicates the flattening of image features; Represents the image features enhanced by the cross-attention mechanism;
[0099] Through the cross-attention operation, the image features are enhanced and the semantic information in the text is integrated. In order to better integrate the image and text features, this embodiment proposes a two-way cross-guidance mechanism. The original image features are also involved in the process of enhancing the text features, forming a two-way cross-guidance:
[0100]
[0101] Finally, this embodiment will enhance the text features As query input, the enhanced image features As keys and values, we further enhance the image features and obtain
[0102]
[0103] This process is performed at each scale to better capture the semantic relationship between the image and the text. At the same time, in order to further improve the accuracy of feature expression, the model balances the information between the original image features and the enhanced image features. At each scale, this embodiment fuses the enhanced image features with the original image features. Specifically, this embodiment fuses the enhanced image features. With the original image feature I i After concatenation, the feature representation is further optimized using weighted residual connections, and the enhanced features are finally obtained:
[0104]
[0105] where w 1 ,w 2 =Softmax(weights) (9)
[0106] Among them, weights is a learnable parameter vector, which is initially a random value and is learned during the training process. This module performs similar operations at multiple scales, gradually enhancing the image and text features of each scale, which are then used in downstream tasks. Through this multi-scale feature fusion strategy, this embodiment can make full use of image and text information at different scales, thereby improving the accuracy and robustness of depth estimation.
[0107] This module performs similar operations at multiple scales, gradually enhancing the image and text features at each scale. Through this multi-scale feature fusion strategy, this embodiment can make full use of image and text information at different scales, thereby improving the accuracy and robustness of depth estimation.
[0108] By dynamically fusing text features to enhance image features, more accurate depth prediction is achieved.
[0109] A multi-scale residual fusion module is designed to further improve the model's expressiveness and robustness of local information by capturing edge features of images at different scales and performing multi-modal alignment with enhanced image features.
[0110] Step 3: The multi-scale residual fusion module further improves the accuracy of depth estimation by capturing subtle changes in the image. Inspired by LapDepth, this module uses Laplacian pyramid residuals to perform multi-scale processing on images and text-enhanced images to generate rich detail information. This detail information is then input into the decoder. The decoder uses a layer-by-layer reconstruction strategy similar to LapDepth to gradually synthesize depth residuals through recursive relationships to restore the depth information of the image. Specifically, the Laplacian pyramid residual captures high-frequency information and local structural changes in the image, helping the decoder to better recover the details in the depth map.
[0111] The following steps are involved:
[0112] The original image I K Downsample to different scales k, k∈{1,2,3,4} to obtain multi-scale image features. At each scale, this embodiment calculates the Laplace residual L k , which represents the changes in image features at different resolutions. The formula is as follows:
[0113] L k =I k -Up(I k+1 ) (10)
[0114] Where k represents the level index of the Laplacian pyramid; I K This is done by downsampling the original input image to 1 / 2 k -1Up(·) represents the upsampling operation (using bilinear interpolation).
[0115] This embodiment also applies the enhanced image (processed by the image enhancement network) to this method to obtain richer detail information. The enhanced image is calculated through a similar multi-scale Laplace residual to obtain an enhanced Laplace residual corresponding to the original image:
[0116]
[0117] In order to effectively combine the detail differences between the original image and the enhanced image, the difference residual between the original image and the enhanced image is calculated:
[0118]
[0119] Where, ΔL k The difference residual captures the changes introduced during the image enhancement process and provides key information for subsequent feature fusion.
[0120] In order to further improve the accuracy of depth estimation, this embodiment proposes a multi-scale feature fusion method to further restore the local details of the image by combining image features at multiple scales. k , Laplace residual of enhanced image and the difference between the two residuals ΔL k The definition is as follows:
[0121] f1 k Represents the Laplace residual L of the kth layer of the original image k ;
[0122] f2 k Represents the Laplace residual of the kth layer of the enhanced image
[0123] f3 k : represents the difference residual ΔL between the two at the kth layer k .
[0124] For each scale k, these features are weighted by the channel attention mechanism (CA) and the spatial attention mechanism (SA):
[0125] Fi k = fi k ·CA(fi k )·SA(fi k ·CA(fi k )) i∈{1,2,3},k∈{1,2,3,4} (13)
[0126] CA(·) calculates the weight of each channel and weights the features by multiplying them channel by channel to focus on important channel information. The spatial attention mechanism SA(·) generates a spatial weight map by calculating the average and maximum values of each position and applies it to the spatial dimension of the feature map, so that the model can pay more attention to important areas in the image.
[0127] After applying the channel and spatial attention mechanisms, this embodiment concatenates the three weighted residual features of each scale and fuses them through a convolutional layer to obtain a fuse. k The purpose of this operation is to integrate feature information from different sources and extract a more compact high-dimensional representation through the convolution layer. On this basis, in order to further improve the accuracy of multi-scale feature fusion, this embodiment also performs weighted averaging on the original image features, enhanced image features, and difference residual features of each scale. The weighted averaging operation can ensure that the three features are reasonably balanced during the fusion process, thereby avoiding a certain feature dominating the fusion process. After obtaining the weighted average feature, this embodiment processes it through 1x1 convolution to obtain a more compact and refined feature representation f res k , to reduce the dimension of the feature and extract a more compact representation. Finally, this embodiment combines the above two features by weighted fusion to obtain the final fusion feature f at each scale. final k :
[0128] fuse k =Conv2d(concat(F1 k ,F2 k ,F3 k )) (14)
[0129] f res k ==Conv2d((f1 k +f2 k +f3 k ) / 3) (15)
[0130] f final k =αfuse k +β·f res k (16)
[0131] Among them, α and β are learnable fusion weights, and the initial values are random values, which are used to balance the contribution of fusion features and residual features. Through this weighted fusion strategy, this embodiment can effectively fuse original image features, enhanced image features, and difference residual features at multiple scales, thereby improving the accuracy and robustness of depth estimation. The effect of the module has been verified in the ablation experiment. The experimental results show that the use of multi-scale residual fusion strategy significantly improves the model's adaptability to complex scenes, especially in the recovery of depth boundaries and the retention of details.
[0132] Step 4: The decoder effectively restores the overall depth structure and detail information of the image by synthesizing the depth residual layer by layer. This process not only improves the accuracy of depth estimation, but also enhances the robustness of the model in complex and diverse scenes. By integrating the multimodal information of images and texts, combining the cross-attention mechanism and detail extraction module, the monocular depth estimation method based on text guidance and multi-scale fusion has achieved significant performance improvement in the monocular depth estimation task, demonstrating its extensive potential and superiority in practical applications.
[0133] This embodiment also discloses the training process of the model:
[0134] Scale invariant loss: This example uses the true value To minimize a supervised loss function. In order to improve the stability of training in diverse scenarios, this embodiment adopts a scale-invariant depth loss, which promotes the scale invariance of the prediction results by calculating the logarithmic difference of the depth image.
[0135]
[0136] in Ω represents the image space, N e is the number of valid pixels, y is the predicted depth, and γ is a scaling factor that controls the sensitivity of the loss.
[0137] Feature alignment loss: According to this embodiment, the feature alignment loss proposed in Depth_Anything is also introduced to migrate powerful semantic capabilities to the depth estimation model of this embodiment:
[0138]
[0139] In this loss function, cos represents the cosine similarity between two feature vectors, f is the feature extracted by the deep model S, and f′ is the feature from the frozen DINOv2 encoder. This alignment loss encourages the depth estimation model to maintain the semantic consistency of the pre-trained DINOv2 features, enhancing the model's ability to capture meaningful representations for accurate depth prediction.
[0140] Final loss function: The final loss function consists of a weighted combination of scale invariance loss and feature alignment loss. The specific form is:
[0141] L=αL SI +βL FA (19)
[0142] Among them, α and β are weight coefficients for balancing L1 loss and L2 loss, respectively. After a lot of experiments, their values are set to 10 and 1, respectively.
[0143] The method also discloses an embodiment
[0144] Step 1, process the dataset (NYU, KITTI)
[0145] This embodiment evaluates the method of this embodiment on two different datasets: NYU Depth V2 for indoor scenes and KITTI for outdoor environments.
[0146] The NYU Depth V2 dataset contains 120K pairs of RGB images and depth images, which were taken in 464 indoor scenes using a Microsoft Kinect sensor with a resolution of 640×480 pixels. This embodiment follows the dataset partitioning method proposed in the existing literature (Lee J H, Han MK, Ko DW, et al. From big to small: Multi-scale local planar guidance for monocular depth estimation [J]. arXiv preprint arXiv: 1907.10326, 2019.), which contains 24,231 training images and 654 test images.
[0147] The existing document is David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from a single image using a multi-scale deep network. Advances inneural information processing systems, 27, 2014. 4, 5, 6, 1. The KITTI dataset contains various road environment images collected from autonomous driving scenes, with a resolution of 1242×375 pixels. In order to compare the performance, this embodiment adopts the partitioning strategy proposed by Eigen et al. in the document David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from asingle image using a multi-scale deep network. Advances inneural information processing systems, 27, 2014. 4, 5, 6, 1. The test set includes 697 images selected from 29 scenes, and the training set consists of 23,488 images from the remaining 32 scenes. According to the methods of existing literature (Shariq Farooq Bhat, Ibraheem Alhashim, and Peter Wonka. Adabins: Depth estimation using adaptive bins. In Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition, pages 4009–4018, 2021. 2, 3, 5, 6, 7, 8, 1. and Weihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu, and PingTan. Neural window fully-connected crfs for monocular depth estimation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 3916-3925, 2022. 2, 5, 6, 8), after clearing invalid ground truth samples, this embodiment finally obtains 652 valid test images.During the testing phase, according to the guidelines of the KITTI dataset, this embodiment limits the maximum value of the prediction output to the order of 80 meters.
[0148] Step 2, Experimental Setup (Evaluation Metrics, Hyperparameters)
[0149] Step 2.1 Training
[0150] The proposed method is implemented based on the PyTorch framework. In the experiment, the text encoder part adopted the same settings as in the existing literature (Zeng Z, Wang D, Yang F, et al. Wordepth: Variational language prior for monocular depth estimation [C] / / Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2024: 9708-9719). Specifically, this embodiment uses the ResNet-50 version of CLIP as a text encoder, which is responsible for extracting text features. In order to improve computational efficiency, the title generation part of the image uses ExpansionNet-v2. This architecture effectively extracts text information and generates accurate descriptions for the image, which helps the depth estimation task of this embodiment.
[0151] Hyperparameters: During the optimization process, this embodiment uses the AdamW optimizer and does not use weight decay. This embodiment sets the same learning rate of 1×10-4 for the encoder and decoder, and the weight decay is set to 0. The eps parameter of the optimizer is set to 1×10-3 to ensure the stability of numerical calculations. Under this scheduler, the model is trained on the KITTI and NYU Depth V2 datasets for 50 cycles. The weight α of the scale-invariant loss is set to 10, and the weight of the feature loss is set to 1. In the training phase, in order to reduce overfitting and improve the generalization ability of the model, this embodiment uses online data enhancement technology in the training phase, and the experimental setting follows the practice in WorDepth. Specifically, this embodiment implements a variety of random image transformations on the NYU Depth V2 and KITTI datasets, including adjusting the brightness, gamma value, and color intensity of the image. At the same time, this embodiment randomly flips and rotates the image horizontally to enhance the diversity of the image. These data enhancement methods can effectively increase the diversity of the training set and improve the adaptability of the model to different scenarios.
[0152] Step 2.2 Evaluation indicators
[0153] This embodiment adopts the evaluation indicators proposed by Eigen et al. (specifically from the literature D. Eigen, C. Puhrsch, and R. Fergus, "Depth map prediction from a single image using a multi-scale deep network," in Proc. Adv. Neural Inf. Process. Syst., Dec. 2014, pp. 2366–2374.), which are widely used in the performance evaluation of monocular depth estimation. Specifically, this embodiment uses the mean absolute relative error (Abs Rel), the root mean square error (RMSE), the absolute error in the logarithmic space (log10), the logarithmic root mean square error (RMSElog), and the threshold accuracy (δi) to quantitatively evaluate the method of this study.
[0154] Step 2.3 Quantitative Results - Visualization
[0155] The experimental results of this embodiment on the NYU Depth V2 dataset show that compared with existing depth estimation techniques, the model of this embodiment shows significant improvement in all evaluation indicators. In particular, the TFDepth method performs outstandingly on the threshold accuracy δ<1.25, which measures the deviation ratio between the predicted result and the true value within a specific range. Specifically, when using vit-l as the backbone, the δ<1.25 accuracy of the TFDepth method reaches 0.969, which is significantly better than other existing state-of-the-art methods. Even when using vit-s as the backbone, the performance of the TFDepth method exceeds many current existing methods, further verifying the superiority of the method of this embodiment. The experimental results of this embodiment show that the TFDepth method has been improved in all evaluation indicators, demonstrating its significant progress in depth estimation accuracy, especially in dealing with complex scale estimation tasks. By using text modality to enhance image features, this embodiment can capture scene details more accurately, so that the depth estimation results are closer to the true value. This improvement further verifies the advantages of the model of this embodiment in various depth estimation tasks.
[0156] The experimental results of the method of this embodiment on the KITTI dataset use Eigen Split (from the literature David Eigen, Christian Puhrsch, and Rob Fergus. Depth map prediction from asingle image using a multi-scale deep network. Advances in neural information processing systems, 27, 2014. 4, 5, 6, 1) division. Compared with existing depth estimation techniques, the TFDepth method also achieves state-of-the-art performance. Similar to the NYU Depth V2 dataset, the TFDepth method improves on the threshold accuracy of δ<1.25, but the relative performance improvement on this indicator is not as significant as on NYU Depth V2. This difference may be due to the wider range of sizes and shapes of objects in outdoor scenes, and objects of the same category may have different sizes and shapes. For example, the word "car" may refer to a sedan, a coupe, or a hatchback - they differ in size (such as a coupe is usually smaller than a sedan) and shape (a hatchback has a higher roof and a connected trunk). Although text descriptions as priors provide flexibility between generality and specificity, over-reliance on text features may be counterproductive when the description is vague. The TFDepth method effectively alleviates these problems by introducing a cross-attention mechanism and enhancing image features with text features. Specifically, text features dynamically adjust the weights of image features through the CSAM module, so that the model can more accurately identify and distinguish objects of different forms. This multimodal feature enhancement strategy not only improves the accuracy of depth estimation, but also enables the model to maintain a high level of performance in complex outdoor scenes. Experimental results show that when the TFDepth method deals with outdoor objects with high diversity and complexity, the positive effects of prior information significantly outweigh the potential negative effects, further verifying the advantages of the model of this embodiment in various depth estimation tasks.
[0157] Step 2.4 Ablation
[0158] In order to evaluate the contribution of the fusion module and the detail extraction module to the model performance, this embodiment conducted an ablation experiment on the NYU Depth V2 dataset, using VIT-S as the backbone. Specifically, this embodiment set the following three experimental configurations: (1) removing the cross-semantic attention module and the multi-scale residual fusion module; (2) removing the multi-scale residual fusion module; (3) using the complete model. The experimental results show that the complete model performs best on all evaluation indicators, significantly better than the configuration with some modules removed. This shows that the cross-semantic attention module effectively combines text and image features through the cross-attention mechanism, significantly improving the accuracy of depth estimation. At the same time, the multi-scale residual fusion module plays a key role in capturing subtle changes in the image, and its existence further enhances the model's adaptability to complex scenes. Removing the cross-semantic attention module and the multi-scale residual fusion module resulted in a significant drop in performance, indicating that these two modules play an indispensable role in the overall model. Although only removing the multi-scale residual fusion module also resulted in a drop in performance, its impact was slightly weaker than when all modules were removed. This further verifies the importance of the multi-scale residual fusion module in improving the accuracy of depth estimation. In summary, the synergy of the cross-semantic attention module and the multi-scale residual fusion module significantly improves the depth estimation performance of the model in complex and diverse scenarios, verifying the effectiveness and superiority of the model design of this embodiment.
[0159] The present invention also discloses a monocular depth estimation system based on text guidance and multi-scale fusion, comprising:
[0160] An initial feature extraction module, used to obtain the visual features of the original image and the text features of the original text;
[0161] A feature enhancement module is used to map text features to the same dimension as the visual features of the original image to obtain processed text features, introduce the processed text features into the visual features of the original image to obtain image features with semantic guidance, enhance the text features based on the visual features to obtain enhanced text features, introduce the enhanced text features into the image features with semantic guidance through bidirectional cross-guidance to obtain final enhanced image features, and obtain an enhanced image;
[0162] The residual calculation module is used to perform multi-scale Laplace residual calculation on the original image and the enhanced image, obtain the residual information between the original image and the enhanced image, and perform a fusion operation based on the residual information and the enhanced image features to obtain an estimated depth image.
[0163] A schematic diagram of a terminal device provided in an embodiment of the present invention. The terminal device of this embodiment includes: a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are implemented.
[0164] The computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to accomplish the present invention.
[0165] The terminal device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0166] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0167] The memory may be used to store the computer programs and / or modules, and the processor implements various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory.
[0168] If the module / unit integrated in the terminal device is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0169] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A monocular depth estimation method based on text guidance and multi-scale fusion, characterized in that: The following steps are involved: Obtaining visual features of the original image and text features of the original text; Mapping the text features to the same dimension as the visual features of the original image to obtain processed text features, introducing the processed text features into the visual features of the original image to obtain image features with semantic guidance, and enhancing the text features based on the visual features to obtain enhanced text features; By introducing enhanced text features into image features with semantic guidance through bidirectional cross-guidance, the final enhanced image features are obtained to obtain an enhanced image; Multi-scale Laplace residual calculation is performed on the original image and the enhanced image to obtain residual information between the original image and the enhanced image, and a fusion operation is performed based on the residual information and the enhanced image features to obtain an estimated depth image.
2. The monocular depth estimation method based on text guidance and multi-scale fusion according to claim 1, characterized in that: The text features are mapped to the same dimension as the image information, and the processed text features ti are obtained by the following formula: t i =Proj i ·t Among them, Proj i represents the projection matrix corresponding to scale i, and its dimension is 1024×Ci; Ci is the number of channels of the i-th layer image feature; The method of introducing the processed text features into the visual features of the original image to obtain image features with semantic guidance, and enhancing the text features based on the visual features of the original image to obtain enhanced text features includes: Use text features ti as queries and image features fi as keys and values, and enhance them through the cross-attention mechanism to generate semantically guided image features; The query, key, and value are transformed through multiple different linear transformations, and then the attention of multiple heads is calculated respectively, and the output results of multiple heads are spliced to obtain enhanced image features.
3. The monocular depth estimation method based on text guidance and multi-scale fusion according to claim 2, characterized in that: The enhanced image features It is expressed by the following formula: Where H represents the number of heads, W o Represents the linear transformation matrix of the output; A i Indicates the output result of the header; The whole process is defined as cross attention: Among them, T i Represents the text features after projection; Represents the flattening of image features.
4. The monocular depth estimation method based on text guidance and multi-scale fusion according to claim 2, characterized in that: The text features are enhanced based on the visual features of the original image to obtain enhanced text features, including: Among them, T i represents the text features after projection; Represents the flattening of image features.
5. The monocular depth estimation method based on text guidance and multi-scale fusion according to claim 4, characterized in that: The enhanced text features are introduced into the image features with semantic guidance through bidirectional cross-guidance to obtain the final enhanced image features, including: Enhanced text features As query input, the enhanced image features As keys and values, further enhanced image features are obtained in, represents enhanced image features; Represents enhanced text features; The enhanced image features are concatenated with the original image features, and then a 2D convolution operation is performed; The feature representation is further optimized using weighted residual connections to obtain the final enhanced features.
6. The monocular depth estimation method based on text guidance and multi-scale fusion according to claim 1, characterized in that: The obtaining of residual information between the original image and the enhanced image includes: Calculate the Laplace residual L of the original image k : L k =I k -Up(I k+1 ) The multi-scale Laplace residual is calculated for the enhanced image to obtain the enhanced Laplace residual corresponding to the original image: Among them, en anced_image k represents the enhanced image; Compute the difference residual between the original image and the enhanced image: Among them, L k Represents the Laplace residual of the kth layer of the original image; represents the Laplace residual of the kth layer of the enhanced image; ΔL k Represents the difference between the Laplace residual of the enhanced image and the original image at the kth level.
7. The monocular depth estimation method based on text guidance and multi-scale fusion according to claim 6, characterized in that: Based on the residual information and enhanced image features, a fusion operation is performed to obtain an estimated depth image, including: For each scale k, these residual features are weighted through the channel attention mechanism and the spatial attention mechanism to obtain the weighted residual features; The weighted residual features are concatenated and the fused features are obtained through convolutional layer fusion to obtain the fused residual feature representation; The fused residual feature representation and enhanced image features are input into the decoder to generate an estimated depth image.
8. A monocular depth estimation system based on text guidance and multi-scale fusion, characterized in that: include: An initial feature extraction module, used to obtain the visual features of the original image and the text features of the original text; A feature enhancement module is used to map text features to the same dimension as the visual features of the original image to obtain processed text features, introduce the processed text features into the visual features of the original image to obtain image features with semantic guidance, enhance the text features based on the visual features to obtain enhanced text features, introduce the enhanced text features into the image features with semantic guidance through bidirectional cross-guidance to obtain final enhanced image features, and obtain an enhanced image; The residual calculation module is used to perform multi-scale Laplace residual calculation on the original image and the enhanced image, obtain the residual information between the original image and the enhanced image, and perform a fusion operation based on the residual information and the enhanced image features to obtain an estimated depth image.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Monocular image depth estimation method based on multi-scale residual pyramid attention network model
CN112001960A
Monocular depth prediction algorithm based on multi-scale progressive interaction and aggregation cross attention features
CN116485860A
Visual saliency prediction method and system under text guidance, terminal and medium
CN117351231A
Efficient scene text image super-resolution method with semantic guidance
CN118608385A
Scene adaptive video compression method and system based on natural language guidance
CN118972590A
Cited By
Self-supervised monocular depth estimation method based on spatial-semantic prior feature enhancement
CN120580273A