A Depth Estimation Method, Device, Terminal and Medium Combining Cross-Modal Information
By combining the depth estimation method of cross-modal information, the visual language pre-trained model is used to extract images and scene features, and the depth text feature set and codebook are obtained, which solves the problem of long training time and poor accuracy of traditional methods, and achieves efficient depth estimation.
Patent Information
- Application Number
- CN202410359264.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-27
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-03-27
AI Technical Summary
The monocular depth estimation method based on traditional machine learning and deep learning in the prior art requires a lot of training time and computing resources, but the depth estimation method based on pre-trained large models has poor accuracy.
Using a depth estimation method combining cross-modal information, the target image is input into the trained visual language pre-trained model, image features and scene features are extracted, and the depth text feature set, scene-like feature set and depth codebook are obtained. Based on these features, the depth weight matrix and the target depth box are determined to perform depth estimation.
The training steps are reduced, the accuracy and generalization ability of depth estimation are improved, and the depth range calculation is adapted to different scenarios.
Smart Images

Figure CN118212282B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to a depth estimation method, device, terminal and medium that combines cross-modal information. Background Art
[0002] Large-scale vision-language pre-training models (VLMs) are very effective in image classification tasks because they combine visual and language information. Monocular depth estimation is a downstream task of computer vision and has important applications in the fields of autonomous driving, virtual reality, and robotics. However, common traditional methods based on machine learning or deep learning require a large amount of data, are computationally expensive, and consume a long time.
[0003] Some methods propose directly using vision-language pre-training models for monocular depth estimation. Although this method can avoid the cumbersome training process, it has limitations. That is, the scene depth ranges corresponding to different images are not the same, and existing methods estimate images with different scene depth ranges according to the same standard, resulting in poor accuracy of the depth estimation results.
[0004] Therefore, there are defects in the prior art and it needs to be improved and developed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a depth estimation method, device, terminal and medium that combines cross-modal information in view of the above-mentioned defects of the prior art, aiming to solve the problem that the methods based on traditional machine learning and deep learning in the prior art require a large amount of training time and computing power resources, but the depth estimation method based on pre-trained large models has poor accuracy.
[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0007] A depth estimation method that combines cross-modal information, wherein the method includes:
[0008] Input a target image into a trained vision-language pre-training model to extract the image features and scene features of the target image;
[0009] Obtain a pre-stored depth text feature set, a class scene feature set, and a depth codebook. The depth text feature set is used to reflect several depth categories. The class scene feature set includes scene features corresponding to different scenes. The depth codebook includes several scenes and corresponding depth bins, and each depth bin includes depth values corresponding to each depth category;
[0010] Determine a depth weight matrix based on the image features and the depth text feature set, and determine a target depth bin based on the scene features, the class scene feature set, and the depth codebook;
[0011] Obtain a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix.
[0012] In one implementation, extracting the image features and scene features of the target image includes:
[0013] Divide the target image into a plurality of image patches, extract features from the plurality of image patches respectively and then splice them to obtain image features;
[0014] Encode the entire image of the target image to generate scene features.
[0015] In one implementation, the depth text feature set is updated during training based on preset depth category description information, and the class scene feature set is obtained by extracting scene features corresponding to different scenes from a pre-constructed training sample set; the depth codebook is updated during training based on a preset initial depth codebook.
[0016] In one implementation, determining a depth weight matrix based on the image features and the depth text feature set includes:
[0017] Calculate a first cosine similarity between the image features and the depth text feature set to estimate the depth category of each image patch;
[0018] Normalize the first cosine similarity to obtain a depth weight matrix;
[0019] Determining a target depth bin based on the scene features, the class scene feature set, and the depth codebook includes:
[0020] Calculate a second cosine similarity between the scene features and the class scene feature set;
[0021] Determine a target scene in the class scene feature set that matches the target image according to the second cosine similarity;
[0022] Search the depth codebook according to the target scene to obtain the target depth bin corresponding to the target image.
[0023] In one implementation, obtaining a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix includes:
[0024] Based on the depth weights corresponding to each depth category in the depth weight matrix, perform a weighted sum of the depth values corresponding to each depth category in the target depth bin to obtain a depth estimation result, where the depth estimation result includes the depth of each pixel in the target image.
[0025] In one implementation, the training steps of the vision-language pre-training model include:
[0026] Construct a training sample set, where the training sample set includes a number of training images corresponding to different scenarios, and depth value labels corresponding to each training image;
[0027] Input the current training image into the initial vision-language pre-training model to extract the training image features and training scenario features of the current training image;
[0028] Obtain a depth text feature set, a class scenario feature set, and a preset initial depth codebook. The depth text feature set is obtained based on preset depth category description information. The class scenario feature set is obtained by extracting scenario features corresponding to different scenarios from a pre-constructed training sample set. The initial depth codebook includes a number of scenarios and corresponding depth bins;
[0029] Determine a training depth weight matrix based on the training image features and the depth text feature set, and determine a target training depth bin based on the training scenario features, the class scenario feature set, and the initial depth codebook;
[0030] Obtain a depth estimation result according to the depth values corresponding to each depth category in the target training depth bin and the depth weight matrix, and train the initial vision-language pre-training model based on the depth value label of the current training image;
[0031] After completing the training of the initial vision-language pre-training model using all the training images in the training sample set, obtain a trained vision-language pre-training model.
[0032] In one implementation, obtaining the depth text feature set includes:
[0033] Obtain a number of preset depth category description information, where the depth category description information includes depth category description words and context information;
[0034] Perform word embedding on the depth category description words to obtain depth category vectors, and convert the context information into context vectors. Use the vector pair formed by the depth category vectors and context vectors as the input text of the depth category description information;
[0035] After inputting the input text into the initial vision-language pre-training model, extract depth text features to form a depth text feature set;
[0036] Among them, both the context vector and the initial depth codebook are learned and updated during training, and the context vector and the depth codebook after completion of the update are obtained after the training is completed.
[0037] The present invention also discloses a depth estimation device combining cross-modal information, wherein the device includes:
[0038] An extraction module, configured to input a target image into a pre-trained vision-language pre-trained model to extract the image feature and the scene feature of the target image;
[0039] An acquisition module, configured to acquire a pre-stored depth text feature set, a class scene feature set, and a depth codebook, where the depth text feature set is used to reflect a plurality of depth categories, the class scene feature set includes scene features corresponding to different scenes, and the depth codebook includes a plurality of scenes and corresponding depth bins, and each depth bin includes depth values corresponding to each depth category;
[0040] A determination module, configured to determine a depth weight matrix based on the image feature and the depth text feature set, and determine a target depth bin based on the scene feature, the class scene feature set, and the depth codebook;
[0041] An estimation module, configured to obtain a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix.
[0042] The present invention also discloses a terminal, which includes: a memory, a processor, and a depth estimation program combining cross-modal information stored on the memory and executable on the processor, and when the depth estimation program combining cross-modal information is executed by the processor, the steps of the depth estimation method combining cross-modal information as described above are implemented.
[0043] The present invention also discloses a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the depth estimation method combining cross-modal information as described above.
[0044] A depth estimation method, device, terminal and medium combining cross-modal information provided by the present invention, the method comprising: inputting a target image into a trained vision-language pre-trained model to extract the image features and scene features of the target image; obtaining a pre-stored depth text feature set, class scene feature set and depth codebook, the depth text feature set being used to reflect a number of depth categories, the class scene feature set including scene features corresponding to different scenes, and the depth codebook including a number of scenes and corresponding depth bins, each depth bin including depth values corresponding to each depth category; determining a depth weight matrix based on the image features and the depth text feature set, and determining a target depth bin based on the scene features, the class scene feature set and the depth codebook; and obtaining a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix. By using the vision-language pre-trained model, the present invention avoids a large number of training steps, obtains a depth text feature set, a class scene feature set and a depth codebook, compares the features of the target image with them, and then determines the current scene, realizing the calculation of images with different depth ranges in different scenes according to different depth bins, and improving the accuracy of depth estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flowchart of a preferred embodiment of the depth estimation method combining cross-modal information in the present invention;
[0046] Figure 2 is a logical schematic block diagram of a preferred embodiment of the depth estimation method combining cross-modal information in the present invention;
[0047] Figure 3 is a comparison of task performance with other methods on the NYU V2 dataset in the depth estimation method combining cross-modal information in the present invention;
[0048] Figure 4 is a functional principle block diagram of a preferred embodiment of the depth estimation device combining cross-modal information in the present invention;
[0049] Figure 5 is a functional principle block diagram of a preferred embodiment of the terminal in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0051] Compared with existing depth estimation methods, the method based on the vision-language pre-training model has higher accuracy and generality, and can be fine-tuned to adapt to different tasks, so it has high practical value. The combination of the depth estimation problem and the large model in the present invention can incorporate text information to help achieve more accurate depth estimation, and at the same time can adapt to different scenarios, showing better generalization performance.
[0052] Please refer to Figure 1 , Figure 1 which is the flowchart of the depth estimation method combining cross-modal information in the present invention. As Figure 1 shown, the depth estimation method combining cross-modal information described in the embodiments of the present invention includes:
[0053] Step S100: Input the target image into the trained vision-language pre-training model to extract the image features and scene features of the target image.
[0054] The vision-language pre-training model has strong expressiveness in some tasks due to the combination of vision and language information. Among them, the contrastive language-image pre-training model (CLIP) in the vision-language pre-training model has excellent image-text matching ability. Depth estimation is to obtain the distance information from each point in the scene of the image to the camera, and monocular means that only one image is provided. The monocular depth estimation task mainly uses image information. In the embodiments of the present application, the information of an additional modality of language is added to achieve the combination of cross-modal information. The combination of cross-modal information in the embodiments of the present application is realized through the vision-language pre-training model. The contrastive language-image pre-training model (CLIP model) can be preferentially selected. Because the model has excellent image-text matching ability, a model for solving the problem can be obtained only through a small amount of sample training. Furthermore, the combination of cross-modal information is realized through few-shot learning, so as to complete the monocular depth estimation task.
[0055] The CLIP model is commonly used in multi-modal tasks of image-text matching. By using the contrastive learning method, the CLIP model uses a training data set containing 400 million text-image pairs to train a powerful text encoder and image encoder. For a pre-trained CLIP model like this, when facing new data (new input images and a series of texts to be matched), there is no need to retrain. The text encoder and image encoder are directly used to obtain text and image features, and the similarity of the images is obtained by calculating the dot product of the two types of features, so as to accurately select the text most similar to the input image.
[0056] To solve the downstream task of depth estimation, the embodiments of the present application can use the method of few-shot learning to extract information from limited data samples and perform model fine-tuning. When training the model, from the perspective of text, the present invention can better extract text features, enabling the text features to assist the image task to the greatest extent; from the perspective of images, it can extract local and global information of the image and combine it with the text for application to the downstream task of depth estimation.
[0057] In the embodiments of the present application, "extracting the image features and scene features of the target image" in step S100 specifically includes: dividing the target image into a plurality of image patches, extracting features from the plurality of image patches respectively and then splicing them to obtain image features; encoding the entire image of the target image to generate scene features.
[0058] The vision-language pre-training model includes an image encoder and a text encoder. The embodiments of the present application use the image encoder to extract the image features and scene features of the target image. Specifically, the image features F are generated by splicing each image patch of the target image. v The scene features F are generated by encoding the entire image of the target image. s This is to facilitate identifying the scene where the target image is located.
[0059] As Figure 1 shown, the depth estimation method combining cross-modal information described in this embodiment further includes:
[0060] Step S200, obtaining a pre-stored depth text feature set, a class scene feature set, and a depth codebook. The depth text feature set is used to reflect several depth categories. The class scene feature set includes scene features corresponding to different scenes. The depth codebook includes several scenes and corresponding depth bins, and each depth bin includes depth values corresponding to each depth category.
[0061] In the embodiments of the present application, the depth text feature set is updated during training based on preset depth category description information. The class scene feature set is obtained by extracting scene features corresponding to different scenes from a pre-constructed training sample set. The depth codebook is updated during training based on a preset initial depth codebook.
[0062] Since images from different scenarios have different depth distributions, the embodiments of the present application design learnable prompts and a learnable depth codebook to effectively adapt a vision-language pre-training model (such as a CLIP model) to different scenarios. Specifically, both the depth text feature set and the class-scenario feature set are used as learnable prompts. Before training the model, the training samples of all scenarios are extracted into the image encoder to obtain a set of class-specific image features, called scenario features, forming the class-scenario feature set. At the same time, all depth category description information is input into the text encoder to obtain a set of depth text features, so as to facilitate subsequent calculation of the similarity degree with the target image based on the depth text feature set and the class-scenario feature set, and then determine the scenario where the target image is located, realizing different processing of different scenarios according to different criteria, and improving the accuracy of the depth estimation method based on the pre-trained large model.
[0063] As Figure 1 shown, the depth estimation method combining cross-modal information described in this embodiment further includes:
[0064] Step S300: Determine a depth weight matrix based on the image features and the depth text feature set, and determine a target depth bin based on the scenario features, the class-scenario feature set, and the depth codebook.
[0065] In the embodiments of the present application, "determine a depth weight matrix based on the image features and the depth text feature set" in step S300 specifically includes: calculating a first cosine similarity between the image features and the depth text feature set to estimate the depth category of each image patch; normalizing the first cosine similarity to obtain a depth weight matrix. By calculating the depth weight matrix in the embodiments of the present application, a basis is provided for calculating the depth estimation result.
[0066] In an embodiment of the present application, "determine a target depth bin based on the scenario features, the class-scenario feature set, and the depth codebook" in step S300 specifically includes: calculating a second cosine similarity between the scenario features and the class-scenario feature set; determining a target scenario in the class-scenario feature set that matches the target image according to the second cosine similarity; and looking up the depth codebook according to the target scenario to obtain the target depth bin corresponding to the target image.
[0067] Specifically, by calculating the second cosine similarity between the scenario features and the class-scenario feature set, the scenario where the target image is currently located can be identified, and the one-hot encoding method is used to select the target depth bin corresponding to this scenario from the depth codebook, realizing obtaining different depth bins for different scenarios and improving the accuracy of the depth estimation method based on the pre-trained large model.
[0068] As Figure 1As shown in the figure, the depth estimation method combining cross-modal information described in this embodiment further includes:
[0069] Step S400: Obtain a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix.
[0070] In the embodiment of the present application, step S400 specifically includes: Based on the depth weights corresponding to each depth category in the depth weight matrix, perform weighted summation on the depth values corresponding to each depth category in the target depth bin to obtain a depth estimation result, and the depth estimation result includes the depth of each pixel in the target image. Specifically, perform weighted summation on the depth estimation weights and the depth values in the depth bin to obtain the final depth value. The depth of each pixel in each image patch is equal to the depth of the image patch.
[0071] In order to seek a balance between high training costs and method effects, the present invention uses a vision-language pre-training model for monocular depth estimation, which can reduce training costs, and thus can achieve training with a very small amount of data, that is, few-shot learning, and calculate images with different depth ranges in different scenes according to different depth bins, ensuring the accuracy and generalization of depth estimation.
[0072] In an embodiment of the present application, the training steps of the vision-language pre-training model include:
[0073] Construct a training sample set, where the training sample set includes a number of training images corresponding to different scenes, and depth value labels corresponding to each training image;
[0074] Input the current training image into the initial vision-language pre-training model to extract the training image features and training scene features of the current training image;
[0075] Obtain a depth text feature set, a class scene feature set, and a preset initial depth codebook. The depth text feature set is obtained based on preset depth category description information, the class scene feature set is obtained by extracting scene features corresponding to different scenes from the pre-constructed training sample set, and the initial depth codebook includes a number of scenes and corresponding depth bins;
[0076] Determine a training depth weight matrix based on the training image features and the depth text feature set, and determine a target training depth bin based on the training scene features, the class scene feature set, and the initial depth codebook;
[0077] Obtain a depth estimation result according to the depth values corresponding to each depth category in the target training depth bin and the depth weight matrix, and train the initial vision-language pre-training model based on the depth value label of the current training image;
[0078] After training the initial vision - language pre - trained model with all the training images in the training sample set, a trained vision - language pre - trained model is obtained.
[0079] Specifically, when depth estimation needs to be performed on a series of images, this series of images can be used as a data set, and one or several pictures corresponding to each scene in the data set are selected as training images to form a training sample set. It can be understood that the trained model has better estimation accuracy for other images in the data set, that is, the target image can also belong to the images in the data set.
[0080] That is to say, when applying the vision - language pre - trained model for depth estimation, only a small number of samples are selected from the data set to be estimated for annotation and training, and the depth estimation of other images in the data set by the trained vision - language pre - trained model will be more accurate.
[0081] Before training, the training samples of all scenes are extracted into the scene encoder to obtain a set of class - specific image features, forming a class - scene feature set. At the same time, all depth category description information is input into the text encoder to obtain a set of depth text features.
[0082] In the embodiment of the present application, obtaining the depth text feature set includes:
[0083] Obtain a plurality of preset depth category description information, where the depth category description information includes depth category description words and context information;
[0084] Perform word embedding on the depth category description words to obtain depth category vectors, and convert the context information into context vectors. The vector pair formed by the depth category vectors and context vectors is used as the input text of the depth category description information;
[0085] After inputting the input text into the initial vision - language pre - trained model, extract depth text features to form a depth text feature set.
[0086] Among them, the context vector and the initial depth encoding book are both learned and updated during training, and the context vector and depth encoding book that are updated are obtained after training.
[0087] Specifically, when designing the text input, the form of “This object is [Depth CLASS]” (This object is very close / far from us) is adopted. For example Figure 2As shown, in the embodiments of the present application, the continuous variable of depth is artificially classified and mapped to discrete values. Specifically, the depth categories include "huge", "very close", "close", "not far", "a bit far", "far", and "invisible", corresponding to different numerical values when the depth bins are different. For example, "invisible" may represent a distance of 5 meters.
[0088] Although the artificially designed text prompts conform to English semantics and grammar rules, they are often not intuitive for the model. Therefore, in the embodiments of the present application, training is carried out through few-shot learning. After converting the context words other than the category into vectors, learning and updating are performed. The text input is in the form of "vector + category". The context vector is V = [V1, V2,..., V n , and the category is the vector m of the word after word embedding i . After simply connecting the two, it becomes the text input T i = [V m i , which is input into the text encoder Φ t to obtain a series of text features
[0089] The depth encoding book Θ is the key to converting the discrete values of this classification result into depth values. The samples in the embodiments of the present application are composed of taking one image from each scene category. Therefore, in the depth encoding book, each scene has a depth bin. For the scene image S j used for few-shot learning, that is, the training image, first input them into the scene encoder Φ s to obtain class-specific scene features The scene encoder Φ here s can directly come from the image encoder in the CLIP model pre-trained based on ResNet-50.
[0090] When the input image is I, I obtains the input scene feature F s = Φ s (I).
[0091] Next, calculate the similarity between the scene feature of the input image and the class-specific scene feature, so as to identify the most suitable depth bin in the depth encoding book Θ through the one-shot encoding technology. The similarity j between the input image I and the scene image S is calculated as follows:
[0092]
[0093] Among them, The scene features in the scene feature set of the representation class, and T is the transpose symbol. The scene with the maximum similarity is selected according to a series of similarities, so as to select the corresponding depth bin θ in the deep codebook * for subsequent calculations.
[0094] For the input image I, it is also input into an image encoder without a pooling layer to obtain the image feature F v = Φ v (I).
[0095] Then, calculate and F v to obtain the cosine similarity between them, and get its similarity score:
[0096]
[0097] Perform Softmax function processing on the similarity score vector to obtain the depth weight d = Softmax(s v ).
[0098] The depth estimation result obtained by weighted summation of the depth weight and the depth bin represents the depth of each pixel in the image block. The calculation formula of the depth estimation result is:
[0099]
[0100] where θ represents the depth value in the depth bin.
[0101] The present invention finds a balance between the time-consuming and laborious traditional methods and the new methods with insufficient generalization and low accuracy through few-shot learning, and explores the monocular depth estimation task by adding cross-modal information in a new few-shot learning setting, aiming to improve the depth estimation performance of the pre-trained large model by using a small group of annotated samples; since images from different scenes have different depth distributions, to solve this problem, the present invention designs learnable prompts and a learnable depth codebook to effectively adapt the cross-modal model to different scenes.
[0102] As Figure 3 shown, the method proposed by the present invention is evaluated on the NYU V2 dataset. The method of the present invention is higher than various unsupervised methods in terms of accuracy. Specifically, in each scene of this application embodiment, one image is selected to train the network, so that the method of the present invention is 10.6% higher than the previous state-of-the-art unsupervised method in terms of the mean absolute relative error (MARE), and can even be compared with fully supervised methods.
[0103] In one embodiment, as Figure 4As shown above, based on the above-mentioned depth estimation method that combines cross-modal information, the present invention also correspondingly provides a depth estimation device that combines cross-modal information, including:
[0104] An extraction module 100, configured to input a target image into a pre-trained vision-language pre-trained model to extract the image features and scene features of the target image;
[0105] An acquisition module 200, configured to acquire a pre-stored depth text feature set, a class scene feature set, and a depth codebook. The depth text feature set is used to reflect several depth categories. The class scene feature set includes scene features corresponding to different scenes. The depth codebook includes several scenes and corresponding depth bins, and each depth bin includes depth values corresponding to each depth category;
[0106] A determination module 300, configured to determine a depth weight matrix based on the image features and the depth text feature set, and determine a target depth bin based on the scene features, the class scene feature set, and the depth codebook;
[0107] An estimation module 400, configured to obtain a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix.
[0108] Figure 5 It is a schematic structural diagram of a terminal provided in an embodiment of the present application. The terminal may include:
[0109] A memory 501, a processor 502, and a computer program stored on the memory 501 and executable on the processor 502.
[0110] When the processor 502 executes the program, it implements the depth estimation method that combines cross-modal information provided in the above embodiment.
[0111] Further, the terminal further includes:
[0112] A communication interface 503, configured to communicate between the memory 501 and the processor 502.
[0113] The memory 501 is used to store a computer program executable on the processor 502.
[0114] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0115] If the memory 501, the processor 502, and the communication interface 503 are implemented independently, the communication interface 503, the memory 501, and the processor 502 can be interconnected via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0116] Optionally, in a specific implementation, if the memory 501, the processor 502, and the communication interface 503 are integrated on a single chip, the memory 501, the processor 502, and the communication interface 503 can communicate with each other via an internal interface.
[0117] The processor 502 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0118] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the depth estimation method for combining cross-modal information as described above.
[0119] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0120] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0121] Any process or method description represented in a flowchart or described otherwise herein may be understood to represent a module, segment, or portion of code including one or N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the technical field to which the embodiments of the present application pertain.
[0122] The logic and / or steps represented in a flowchart or described otherwise herein, for example, may be considered as a sequenced list of executable instructions for implementing a logical function and may be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can read instructions from and execute instructions by the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" may be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium may even be paper or other suitable medium on which a program can be printed, as the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0123] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0124] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. When this program is executed, it includes one or a combination of the steps of the method embodiments.
[0125] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0126] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
[0127] In summary, a depth estimation method, device, terminal, and medium combining cross-modal information disclosed by the present invention, the method comprising: inputting a target image into a pre-trained vision-language pre-trained model to extract image features and scene features of the target image; obtaining a pre-stored depth text feature set, a class scene feature set, and a depth codebook, the depth text feature set being used to reflect a plurality of depth categories, the class scene feature set including scene features corresponding to different scenes, and the depth codebook including a plurality of scenes and corresponding depth bins, each depth bin including depth values corresponding to each depth category; determining a depth weight matrix based on the image features and the depth text feature set, and determining a target depth bin based on the scene features, the class scene feature set, and the depth codebook; and obtaining a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix. By using the vision-language pre-trained model, the present invention avoids a large number of training steps, obtains the depth text feature set, the class scene feature set, and the depth codebook, compares the features of the target image with them, and then determines the current scene, realizing the calculation of images with different depth ranges in different scenes according to different depth bins, and improving the accuracy of depth estimation.
[0128] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A depth estimation method combining cross-modal information, characterized in that, The method includes: Input the target image into a pre-trained vision-language pre-training model to extract the image features and scene features of the target image; Obtain a pre-stored depth text feature set, a class scene feature set, and a depth codebook. The depth text feature set is used to reflect several depth categories. The class scene feature set includes scene features corresponding to different scenes. The depth codebook includes several scenes and corresponding depth bins, and each depth bin includes depth values corresponding to each depth category; Determine a depth weight matrix based on the image features and the depth text feature set, and determine a target depth bin based on the scene features, the class scene feature set, and the depth codebook; Obtain a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix; The depth text feature set is updated during training based on preset depth category description information. The class scene feature set is obtained by extracting scene features corresponding to different scenes from a pre-constructed training sample set. The depth codebook is updated during training based on a preset initial depth codebook; Determining a depth weight matrix based on the image features and the depth text feature set includes: Calculate a first cosine similarity between the image features and the depth text feature set to estimate the depth category of each image patch; Normalize the first cosine similarity to obtain a depth weight matrix; Determining a target depth bin based on the scene features, the class scene feature set, and the depth codebook includes: Calculate a second cosine similarity between the scene features and the class scene feature set; Determine a target scene in the class scene feature set that matches the target image according to the second cosine similarity; Search the depth codebook according to the target scene to obtain the target depth bin corresponding to the target image; The training steps of the vision-language pre-training model include: Construct a training sample set, which includes several training images corresponding to different scenes and depth value labels corresponding to each training image; Input the current training image into an initial vision-language pre-training model to extract the training image features and training scene features of the current training image; Obtain a depth text feature set, a class scene feature set, and a preset initial depth codebook. The depth text feature set is obtained based on preset depth category description information. The class scene feature set is obtained by extracting scene features corresponding to different scenes from a pre-constructed training sample set. The initial depth codebook includes several scenes and corresponding depth bins; Determine a training depth weight matrix based on the training image features and the depth text feature set, and determine a target training depth bin based on the training scene features, the class scene feature set, and the initial depth codebook; Obtain a depth estimation result according to the depth values corresponding to each depth category in the target training depth bin and the depth weight matrix, and train the initial vision-language pre-training model based on the depth value label of the current training image; After training the initial vision-language pre-trained model with all the training images in the training sample set, a trained vision-language pre-trained model is obtained; Obtain a depth text feature set, including: Obtain a plurality of preset depth category description information, where the depth category description information includes depth category description words and context information; Perform word embedding on the depth category description words to obtain depth category vectors, and convert the context information into context vectors. Use the vector pair formed by the depth category vectors and context vectors as the input text of the depth category description information; After inputting the input text into the initial vision-language pre-trained model, extract depth text features to form a depth text feature set; Among them, both the context vector and the initial depth codebook are learned and updated during training, and the updated context vector and depth codebook are obtained after training; When performing depth estimation on a plurality of images, use the plurality of images as a data set, and select one or several pictures corresponding to each scene in the data set as training images to form a training sample set; the target image is an image in the data set.
2. The depth estimation method for combining cross-modal information according to claim 1, wherein Extract the image features and scene features of the target image, including: Slice the target image into a plurality of image patches, extract features from the plurality of image patches respectively and then splice them to obtain image features; Encode the entire image of the target image to generate scene features.
3. The depth estimation method for combining cross-modal information according to claim 1, wherein Obtain a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix, including: Based on the depth weights corresponding to each depth category in the depth weight matrix, perform weighted summation on the depth values corresponding to each depth category in the target depth bin to obtain a depth estimation result, where the depth estimation result includes the depth of each pixel in the target image.
4. A depth estimation device that combines cross-modal information, characterized in that, The device includes: An extraction module, configured to input a target image into the trained vision-language pre-trained model to extract the image features and scene features of the target image; An acquisition module, configured to acquire a pre-stored depth text feature set, a class scene feature set, and a depth codebook. The depth text feature set is used to reflect a plurality of depth categories. The class scene feature set includes scene features corresponding to different scenes. The depth codebook includes a plurality of scenes and corresponding depth bins, and each depth bin includes depth values corresponding to each depth category; A determination module, configured to determine a depth weight matrix based on the image features and the depth text feature set, and determine a target depth bin based on the scene features, the class scene feature set, and the depth codebook; An estimation module, configured to obtain a depth estimation result according to the depth values corresponding to each depth category in the target depth bin and the depth weight matrix; The depth text feature set is updated during training based on preset depth category description information. The class scene feature set is obtained by extracting scene features corresponding to different scenes from a pre-constructed training sample set. The depth codebook is updated during training based on a preset initial depth codebook; Determining a depth weight matrix based on the image features and the depth text feature set includes: Calculating a first cosine similarity between the image features and the depth text feature set to estimate the depth category of each image patch; Normalizing the first cosine similarity to obtain a depth weight matrix; Determining a target depth bin based on the scene features, the class scene feature set, and the depth codebook, includes: Calculating a second cosine similarity between the scene features and the class scene feature set; Determining a target scene in the class scene feature set that matches the target image according to the second cosine similarity; Searching the depth codebook according to the target scene to obtain the target depth bin corresponding to the target image; The training steps of the visual language pre-training model include: Constructing a training sample set, which includes a number of training images corresponding to different scenes, and depth value labels corresponding to each training image; Inputting the current training image into the initial visual language pre-training model to extract the training image features and training scene features of the current training image; Obtaining a depth text feature set, a class scene feature set, and a preset initial depth codebook. The depth text feature set is obtained based on preset depth category description information. The class scene feature set is obtained by extracting scene features corresponding to different scenes from a pre-constructed training sample set. The initial depth codebook includes a number of scenes and corresponding depth bins; Determining a training depth weight matrix based on the training image features and the depth text feature set, and determining a target training depth bin based on the training scene features, the class scene feature set, and the initial depth codebook; Obtaining a depth estimation result according to the depth values corresponding to each depth category in the target training depth bin and the depth weight matrix, and training the initial visual language pre-training model based on the depth value label of the current training image; After completing the training of the initial visual language pre-training model using all the training images in the training sample set, obtaining a trained visual language pre-training model; Obtaining a depth text feature set, includes: Obtaining a number of preset depth category description information, where the depth category description information includes depth category description words and context information; Performing word embedding on the depth category description words to obtain depth category vectors, and converting the context information into context vectors, and using the vector pair formed by the depth category vectors and context vectors as the input text of the depth category description information; After inputting the input text into the initial visual language pre-training model, extracting depth text features to form a depth text feature set; Among them, both the context vector and the initial depth codebook are learned and updated during training, and a completed updated context vector and depth codebook are obtained after training; When performing depth estimation on a number of images, using the number of images as a data set, and selecting one or several pictures corresponding to each scene in the data set as training images to form a training sample set; the target image is the image in the data set.
5. A terminal, characterized in that, Includes: A memory, a processor, and a depth estimation program combined with cross-modal information stored on the memory and executable on the processor, the steps of the depth estimation method combined with cross-modal information as described in any one of claims 1 to 3 are implemented when the depth estimation program combined with cross-modal information is executed by the processor.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program can be executed to implement the steps of the depth estimation method combined with cross-modal information as described in any one of claims 1 to 3.