Monocular image depth estimation method, device and equipment based on diffusion model and target prompt
By using a method combining diffusion model with target cues in monocular depth estimation, the object detection network is used to extract object information and generate fine-grained prompt information, the problem of lack of targetedness and accuracy of prompt information in the prior art is solved, and the accuracy of depth estimation and the universality of the model are improved.
Patent Information
- Application Number
- CN202510226151.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
With the diversity of visual scenes, the existing monocular depth estimation method based on diffusion model cannot fully capture the scene details characteristics based on scene category information alone, resulting in the lack of targetedness and accuracy of the generated prompt information, which limits the model's understanding of complex scenes and the accuracy of depth estimation.
The monocular image depth estimation method based on diffusion model and target prompt is adopted, and the image is detected through the pre-trained target detection network, object information is extracted, and the target detection result is used to process the target detection results to generate target prompt information, and a richer and diverse detailed information is generated based on the image content.
The depth estimation model's understanding of complex scenarios and the accuracy of depth estimation is improved, the template-based prompts are avoided monotonous and redundant, and the limitations of high requirements for scene labeling data are overcome, and it has stronger versatility and adaptability.
Smart Images

Figure CN120147390A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a monocular image depth estimation method, device and equipment based on a diffusion model and target prompts. Background Art
[0002] Depth estimation refers to inferring the distance from each pixel point in a scene to the camera imaging plane, that is, the depth value, by analyzing a single or multiple images. In the field of computer vision, depth estimation is considered to be a very important research direction and is widely applied in many frontier technology fields such as autonomous driving, three-dimensional scene reconstruction, and augmented reality. Accurate depth information can not only help the system better understand and analyze the surrounding environment, but also provide support for achieving more complex visual tasks. Especially in scenarios involving spatial perception and interaction, it plays a crucial role.
[0003] According to the number of images and viewpoints, depth estimation methods can be divided into monocular depth estimation and multi-view depth estimation. In multi-view depth estimation, it is a common technique to infer depth information by calculating the disparity between different viewpoints (i.e., the position difference of the same object in different images). For example, humans judge the distance of an object by comparing the disparities of the images formed by the left and right eyes. Similarly, binocular cameras also use this principle for depth estimation. However, multi-view cameras usually have a high cost, and have disadvantages such as high cost, strong sensitivity to the detection environment, and poor universality. In contrast, monocular cameras not only have a low cost, are small in size and easy to use, but also monocular images contain rich scene information. Therefore, how to extract depth information from monocular images has become an important research topic and has important academic and application values.
[0004] Early researchers mainly relied on geometric theories or manually designed features (such as focus and defocus, image vanishing points, and shadows, etc.) to extract depth information from monocular images. However, it is difficult to make breakthroughs in these methods at the theoretical level. With the in-depth research, scholars have gradually turned to using deep learning methods to solve the problem of monocular depth estimation. The introduction of deep learning enables researchers to automatically extract complex features from a large amount of data through an end-to-end learning method, which not only significantly improves the performance of monocular depth estimation, but also enhances the adaptability of the model in different environments and conditions. In recent years, some latest methods have tried to use diffusion models to achieve the monocular depth estimation task. Thanks to training with large-scale datasets, these models have shown strong generalization ability and can significantly improve the accuracy of the predicted depth map.
[0005] However, in the current monocular depth estimation method based on the diffusion model, the diffusion model is usually prompted only by the scene category information. Due to the diversity of visual scenes, the scene category information alone cannot fully capture the detailed features of the scene, resulting in the lack of pertinence and accuracy of the generated prompt information. For example, in an autonomous driving scenario, relying solely on the category information of "urban roads" cannot accurately describe specific traffic conditions, weather conditions, or road structures and other details, thereby limiting the model's understanding of complex scenes and the accuracy of depth estimation. In addition, since the scene category information is used as a text description of the image, the scene label or description information corresponding to each image must be obtained in advance during the training process. This requirement limits the application of the model in the absence of scene annotation data or incomplete scene information, resulting in higher requirements for the training data set. Therefore, how to improve the accuracy of monocular depth estimation based on the diffusion model is a technical problem to be solved in the present invention. Summary of the invention
[0006] Based on the above technical problems, the present invention provides a method, device and equipment for monocular image depth estimation based on a diffusion model and target prompts, aiming to improve the accuracy of monocular depth estimation based on the diffusion model.
[0007] A first aspect of the present invention provides a monocular image depth estimation method based on a diffusion model and target prompts, the method comprising: Inputting the sample detection image into a depth estimation model to be trained, wherein the depth estimation model to be trained includes at least: a pre-trained image encoder, a pre-trained target detection network, a target prompt module to be trained, a denoising network to be trained, and an interactive decoder to be trained; Processing the sample detection image by using the pre-trained image encoder to obtain a shallow spatial representation corresponding to the sample detection image; Processing the sample detection image through the pre-trained target detection network to extract the target detection result of the sample detection image; Processing the target detection result by the target prompt module to be trained to obtain target prompt information corresponding to the sample detection image; Inputting the shallow spatial representation and the target prompt information into the denoising network to be trained to extract multi-scale features; Inputting the multi-scale features into the interactive decoder to be trained to obtain a sample depth map; Based on the label depth map corresponding to the sample detection image and the sample depth map, the target prompt module to be trained, the denoising network to be trained and the interactive decoder to be trained are trained until a trained depth estimation model is obtained; Input the image to be detected into the trained depth estimation model to obtain the depth map corresponding to the image to be detected.
[0008] In a second aspect of the present invention, a monocular image depth estimation device based on a diffusion model and target prompts is provided. The device includes: An image input module for inputting a sample detection image into a depth estimation model to be trained, where the depth estimation model to be trained at least includes: a pre-trained image encoder, a pre-trained target detection network, a target prompt module to be trained, a denoising network to be trained, and an interactive decoder to be trained; An image encoding module for processing the sample detection image through the pre-trained image encoder to obtain a shallow spatial representation corresponding to the sample detection image; A target detection module for processing the sample detection image through the pre-trained target detection network to extract the target detection result of the sample detection image; A prompt acquisition module for processing the target detection result through the target prompt module to be trained to obtain the target prompt information corresponding to the sample detection image; A multi-scale feature extraction module for inputting the shallow spatial representation and the target prompt information into the denoising network to be trained to extract multi-scale features; A depth map output module for inputting the multi-scale features into the interactive decoder to be trained to obtain a sample depth map; A model training module for training the target prompt module to be trained, the denoising network to be trained, and the interactive decoder to be trained based on the labeled depth map corresponding to the sample detection image and the sample depth map until a trained depth estimation model is obtained; A depth estimation module for inputting the image to be detected into the trained depth estimation model to obtain the depth map corresponding to the image to be detected.
[0009] In a third aspect of the present invention, an electronic device is provided. The electronic device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the method for monocular image depth estimation based on a diffusion model and target prompts in the first aspect of the embodiments of the present invention.
[0010] In the monocular image depth estimation method based on diffusion model and target prompt provided by the present invention, it is proposed to use the objects in the scene as prompt information to guide the diffusion model to perform the monocular depth estimation task: perform object detection on the sample detection image through a pre-trained object detection network, extract the objects (i.e., targets) existing in the sample detection image to obtain the object detection result, and process the object detection result based on the to-be-trained target prompt module proposed by the present invention to generate the corresponding target prompt information. The target prompt information generated by this method is closely combined with the shallow spatial representation (image content) corresponding to the sample detection image extracted by the pre-trained image encoder, providing richer and more diverse details, enhancing the depth estimation model's understanding ability of complex scenes and the accuracy of depth estimation, and at the same time avoiding the monotony and redundancy of template-based prompts. In addition, since the target prompt information generated by the present invention only depends on the input sample detection image itself and does not require additional scene annotation data, it overcomes the limitation that the model cannot be applied in the case of lack of scene annotation data or incomplete scene information, and avoids the defect of high requirements for the training data set. Based on this, the monocular image depth estimation method based on diffusion model and target prompt proposed by the present invention has stronger generality and adaptability and can be widely applied to various common monocular depth estimation data sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0012] Figure 1 is the average attention map between the prompt word and the image feature during the diffusion process shown in an embodiment of the present invention; Figure 2 is the step flowchart of a monocular image depth estimation method based on diffusion model and target prompt shown in an embodiment of the present invention; Figure 3 is the overall schematic diagram of training a to-be-trained depth estimation model shown in an embodiment of the present invention; Figure 4 is the structural schematic diagram of a target prompt module shown in an embodiment of the present invention; Figure 5 is the structural schematic diagram of an interactive decoder shown in an embodiment of the present invention; Figure 6 is the flowchart of a monocular image depth estimation method based on diffusion model and target prompt shown in an embodiment of the present invention; Figure 7It is the training and inference flowchart of a monocular depth estimation system shown in an embodiment of the present invention; Figure 8 It is the training and inference flowchart of an autonomous driving obstacle detection system shown in an embodiment of the present invention; Figure 9 It is the visualization diagram of the prediction results of different methods shown in an embodiment of the present invention; Figure 10 It is the visualization diagram of the influence of the detected target on the depth prediction of the result shown in an embodiment of the present invention; Figure 11 It is the structural block diagram of a monocular image depth estimation device based on a diffusion model and target prompts provided in an embodiment of the present invention. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0014] Aiming at various disadvantages brought by using scene information as the text description of an image in the current monocular depth estimation method based on a diffusion model, the present invention can find by observing the average attention map of the denoising process of the diffusion model that the diffusion model has an obvious perception of the target words in the text prompt and can generate a corresponding attention map. As Figure 1 shown, Figure 1 is the average attention map between the prompt word and the image feature during the diffusion process shown in an embodiment of the present invention. It can be seen from Figure 1 that the average attention maps corresponding to the two prompt words "cat" and "burger" present their respective corresponding appearance forms. Based on this, the present invention directly extracts prompt information from the input image to provide fine-grained prompt information, and proposes to use the objects in the scene as prompt information to guide the diffusion model to perform the monocular depth estimation task. The prompts generated by this method are closely combined with the image content, providing richer and more diverse details, thereby improving the understanding ability of the depth estimation model for complex scenes and the accuracy of depth estimation, and at the same time avoiding the monotony and redundancy of template-based prompts. And since the prompt information generated by the present invention only depends on the input image itself and does not require additional scene annotation data, it has stronger versatility and adaptability and can be widely applied to various common monocular depth estimation datasets.
[0015] Please refer to Figure 2 , Figure 2It is a flowchart of the steps of a monocular image depth estimation method based on a diffusion model and target prompts shown in an embodiment of the present invention. As Figure 2 shown, the monocular image depth estimation method based on a diffusion model and target prompts provided in this embodiment at least includes the following steps: Step S11: Input the sample detection image into the depth estimation model to be trained. The depth estimation model to be trained at least includes: a pre-trained image encoder, a pre-trained target detection network, a target prompt module to be trained, a denoising network to be trained, and an interactive decoder to be trained.
[0016] In this embodiment, first input the sample detection image into the depth estimation model to be trained. Among them, the depth estimation model to be trained at least includes: a pre-trained image encoder, a pre-trained target detection network, a target prompt module to be trained, a denoising network to be trained, and an interactive decoder to be trained. This embodiment prepares a sample data set for training the depth estimation model. The sample data set includes: sample detection images and label depth maps corresponding to the sample detection images. The label depth map is the label corresponding to the sample detection image and is the true depth map corresponding to the sample detection image. Among them, the sample detection image is an image sample to be depth-estimated for training the depth estimation model.
[0017] Step S12: Process the sample detection image through the pre-trained image encoder to obtain the shallow spatial representation corresponding to the sample detection image.
[0018] In this embodiment, after the sample detection image is input into the depth estimation model to be trained, the sample detection image can be input into the pre-trained image encoder and the pre-trained target detection network respectively, and the sample detection image is processed accordingly through the pre-trained image encoder and the pre-trained target detection network respectively. Among them, the sample detection image can be processed in parallel by the image encoder and the target detection network. After the sample detection image is input into the pre-trained image encoder, the sample detection image can be encoded through the pre-trained image encoder to obtain the shallow spatial representation corresponding to the sample detection image.
[0019] Among them, the latent space is represented as the feature representation corresponding to the shallow space. The shallow space (Latent Space) refers to the low-dimensional representation space of data after being encoded by a neural network. In this space, the high-dimensional features of the data are compressed into lower-dimensional abstract representations, which are convenient for processing and generation tasks. The pre-trained image encoder in this embodiment is a pre-trained image encoder, and during the training process of the depth estimation model, the model parameters of the pre-trained image encoder are fixed. In an alternative embodiment, the image encoder can be a Variational Autoencoder (VAE), such as: Vector Quantized Variational Autoencoder (VQ-VAE), etc. Among them, the variational autoencoder is a generative model. By introducing the idea of variational inference, the model can learn the latent distribution of the data and generate new samples. The VAE maps the data to the latent space through the encoder and then reconstructs the data from the latent space through the decoder.
[0020] Step S13: Process the sample detection image through the pre-trained object detection network to extract the object detection result of the sample detection image.
[0021] In this embodiment, after the sample detection image is input into the pre-trained object detection network, the pre-trained object detection network can perform object detection on the sample detection image, that is, detect each object in the image, and extract the object detection result of the sample detection image. The objects in this embodiment can be understood as various objects in the image, such as tables, benches, computers, etc., which are not limited thereto. The pre-trained object detection network in this embodiment is a pre-trained object detection network, and during the training process of the depth estimation model, the model parameters of the pre-trained object detection network are fixed.
[0022] Since the object prompt information in this embodiment is used as information to assist the depth estimation model in depth prediction, it does not depend on a specific detector, and various types of detectors can be used to extract the object detection result (object information in the scene) in the sample detection image. In an alternative embodiment, the object detection network can be a detector of the DETR or YOLO series, which is not limited thereto.
[0023] Step S14: Process the object detection result through the object prompt module to be trained to obtain the object prompt information corresponding to the sample detection image.
[0024] In this embodiment, after using the target detection network to extract the objects existing in the sample detection image, that is, after the target detection network extracts the target detection result of the sample detection image, the target prompt module to be trained proposed in this embodiment processes the target detection result, encodes the target detection result (i.e., the object information in the scene corresponding to the sample detection image), and obtains the target prompt information corresponding to the sample detection image, so as to be used as the prompt of the diffusion model.
[0025] In this embodiment, the target prompt information output by the target prompt module includes: the relative spatial relationship between the objects in the sample detection image. Among them, according to the perspective principle, the size of the same type of objects in the image is negatively correlated with their corresponding depth. For example, in the real world, for two objects with similar sizes, if they are closer to the camera, they will appear larger in the image, while the farther objects will appear smaller. Based on this, by learning the relative spatial relationship between the objects in the scene in this embodiment, it is helpful to generate a more accurate depth map.
[0026] Step S15: Input the shallow spatial representation and the target prompt information into the denoising network to be trained, and extract multi-scale features.
[0027] The diffusion model used in this embodiment only includes: the denoising network in the diffusion model, and does not involve the noise addition part of the diffusion model. After obtaining the shallow spatial representation and the target prompt information corresponding to the sample detection image in this embodiment, the shallow spatial representation and the target prompt information can be input into the denoising network to be trained, and the multi-scale features corresponding to the sample detection image are extracted through the denoising network to be trained.
[0028] Step S16: Input the multi-scale features into the interactive decoder to be trained, and obtain the sample depth map.
[0029] In this embodiment, after obtaining the multi-scale features, the multi-scale features are input into the interactive decoder to be trained, and the interactive decoder is used to guide the fusion between different scale features in the multi-scale features, and the sample depth map output by the interactive decoder to be trained is obtained. The sample depth map is the depth map obtained by the depth estimation model to be trained for depth estimation of the sample detection image during the model training process.
[0030] Step S17: Based on the label depth map corresponding to the sample detection image and the sample depth map, train the target prompt module to be trained, the denoising network to be trained, and the interactive decoder to be trained until a trained depth estimation model is obtained.
[0031] In this embodiment, the target prompt module to be trained, the denoising network to be trained, and the interactive decoder to be trained can be trained based on the labeled depth map corresponding to the sample detection image and the sample depth map corresponding to the sample detection image until the trained target prompt module, the trained denoising network, and the trained interactive decoder are obtained, so as to obtain a trained depth estimation model composed of a pre-trained image encoder, a pre-trained target detection network, the trained target prompt module, the trained denoising network, and the trained interactive decoder.
[0032] In an alternative embodiment, the loss between the labeled depth map and the sample depth map can be calculated (such as calculating the scale-invariant loss, etc.), and the model parameters of the target prompt module to be trained, the denoising network to be trained, and the interactive decoder to be trained can be updated based on the calculated loss until a trained depth estimation model is obtained.
[0033] Step S18: Input the image to be detected into the trained depth estimation model to obtain the depth map corresponding to the image to be detected.
[0034] In this embodiment, the trained depth estimation model is used for monocular image depth estimation. During the application process of the trained depth estimation model, the image to be detected can be input into the trained depth estimation model to obtain the depth map corresponding to the image to be detected output by the trained depth estimation model. Among them, the image to be detected is an image that needs to perform depth estimation, the depth map is the result of depth estimation for the image to be detected, and the depth map includes the depth value corresponding to each pixel in the image to be detected.
[0035] In this embodiment, a method is proposed to use the objects in the scene as prompt information to guide the diffusion model for monocular depth estimation tasks: perform object detection on the sample detection image through a pre-trained object detection network, extract the objects (i.e., targets) existing in the sample detection image to obtain the object detection result, and process the object detection result based on the object prompt module to be trained proposed by the present invention to generate corresponding object prompt information. The object prompt information generated by this method is closely combined with the shallow spatial representation (image content) corresponding to the sample detection image extracted by the pre-trained image encoder, providing richer and more diverse details, enhancing the depth estimation model's understanding ability of complex scenes and the accuracy of depth estimation, and at the same time avoiding the monotony and redundancy of template-based prompts. In addition, since the object prompt information generated by the present invention only depends on the input sample detection image itself and does not require additional scene annotation data, it overcomes the limitation that the model cannot be applied in the case of lack of scene annotation data or incomplete scene information, and avoids the defect of high requirements for the training data set. Based on this, the monocular image depth estimation method based on the diffusion model and object prompt proposed by the present invention has stronger generality and adaptability and can be widely applied to various common monocular depth estimation data sets.
[0036] In one embodiment, as Figure 3 shown, Figure 3 is an overall schematic diagram showing the model training of a depth estimation model to be trained in an embodiment of the present invention. In Figure 3 , the depth estimation model to be trained at least includes: a pre-trained image encoder (VQVAE), a pre-trained object detection network (Detector), an object prompt module to be trained (OPM, Object Prompt Module), a denoising network to be trained (Denoising UNet), and an interactive decoder to be trained (GLID). Among them, the network parameters of VQVAE and Detector are fixed, and the network parameters of OPM, Denoising UNet, and GLID can be updated.
[0037] In Figure 3In it, the input image x (Input) is first compressed by the vectorized variational autoencoder (VQ-VAE) to 1 / 8 of its original size, obtaining the corresponding latent space representation (Latent Space Features) Z. Meanwhile, the detector is used to extract the objects appearing in the input image, and the object prompt module (OPM) is used to convert the object detection results into object prompt information (Prompts) C. Then, a single-step denoising process is performed in the denoising network, and the latent space representation Z interacts with the object prompt information C through the cross-attention mechanism, and multi-scale features (Multi-Scale Features) F are extracted therefrom. Finally, the global-local interaction decoder (GLID) proposed in this embodiment is used to guide the fusion of features at different scales and generate the final depth map (Pred). Then, based on the final depth map (Pred) and the labeled depth map (GT), the scale-invariant loss is calculated, thereby training the depth estimation model being trained.
[0038] Combined with the above embodiments, in one implementation manner, the present invention further provides a monocular image depth estimation method based on a diffusion model and object prompts. In this method, the object detection results include: class label information and bounding box information of the object; the above step S14 may specifically include steps S21 to S23: Step S21: In the object prompt module to be trained, process the class label information to obtain the label prompt information corresponding to the sample detection image.
[0039] In this embodiment, in the object prompt module to be trained, the object detection results are divided into two parts, namely the label and the bounding box, to encode the detection results corresponding to these two parts into prompt information. Among them, the class label information in the object detection results can be processed to obtain the label prompt information corresponding to the sample detection image.
[0040] Step S22: Process the bounding box information to obtain the position prompt information and size prompt information corresponding to the sample detection image.
[0041] In this embodiment, in the object prompt module to be trained, the bounding box information in the object detection results can also be processed to obtain the position prompt information and size prompt information corresponding to the sample detection image.
[0042] Step S23: Connect the label prompt information, the position prompt information, and the size prompt information, and obtain the object prompt information through convolutional fusion.
[0043] In this embodiment, in the target prompt module to be trained, after obtaining the label prompt information, position prompt information, and size prompt information, the label prompt information, position prompt information, and size prompt information can be concatenated, and the target prompt information is obtained through convolutional fusion and output.
[0044] Combined with the above embodiments, in one implementation, the present invention further provides a monocular image depth estimation method based on a diffusion model and target prompts. In this method, the "processing the category label information to obtain the label prompt information corresponding to the sample detection image" in step S21 above specifically includes step S31 and step S32: Step S31: Based on the category label information, retrieve the target label feature vector corresponding to the category label information from the label prompt set.
[0045] In this embodiment, after obtaining the category label information of the target detection result, the target label feature vector corresponding to the category label information can be retrieved from the label prompt set based on the category label information. Among them, the label prompt set includes: a plurality of label feature vectors, each label feature vector corresponding to different object label information; the label feature vector is a text feature encoded based on the object label information in the training dataset corresponding to the target detection network.
[0046] Since labels play a crucial role in guiding the model to recognize and understand objects in the scene, in this embodiment, a pre-trained text encoder (CLIP, Contrastive Language-Image Pretraining) can be used in advance to separately encode each label (i.e., each object label information) in the training dataset corresponding to the target detection network. Thus, each label will be converted into its corresponding text feature , so as to obtain the label feature vector corresponding to each object label information. After all object label information is converted into label feature vectors, these label feature vectors can be combined together to obtain the label prompt set . It can be expressed by the following formula:
[0047] Among them, represents the i-th label in the training dataset, and n represents the total number of labels.
[0048] CLIP is a model trained by a contrastive learning method that can map images and text to the same latent space. It jointly trains image and text data, enabling the model to understand the relationship between text and images, and is widely used in tasks such as image generation, search, and classification.
[0049] During the model training process (similarly for model inference), based on the class label information provided by the object detection network , retrieve the target label feature vector corresponding to this class label information from the label hint set .
[0050] Step S32: Based on the target label feature vector, obtain the label hint information.
[0051] In this embodiment, the label hint information can be obtained based on the retrieved target label feature vector. Specifically, since there may be multiple valid targets (multiple objects) in the image to be detected, therefore, based on multiple class label information, the corresponding target label feature vectors can be retrieved from the label hint set respectively, and then these multiple retrieved target label feature vectors are concatenated to obtain the label hint information corresponding to the image to be detected .
[0052] In this embodiment, the label hint information is learnable, and its weight will be continuously updated during the entire training process, so that the label hint becomes more refined and has context information as the model is trained.
[0053] Combined with the above embodiments, in one implementation, the present invention also provides a monocular image depth estimation method based on a diffusion model and target hints. In this embodiment, the bounding box information includes: the position information and size information corresponding to the object; and, the above step S22 can specifically include steps S41 to S44: Step S41: Determine the target cell index based on the position information, and retrieve the target position feature vector corresponding to the position information from the position hint set based on the target cell index.
[0054] In this embodiment, the role of the bounding box information is to provide the position and size information of each object for the depth estimation model, so as to achieve a rough construction of the scene. In this embodiment, the target cell index can be determined first based on the position information in the bounding box information, and then the target position feature vector corresponding to this position information can be retrieved from the position hint set based on the target cell index. Among them, the position hint set includes: multiple position feature vectors, each position feature vector corresponds to a different cell index; the position feature vector is the position feature corresponding to a cell after dividing the sample detection image into multiple cells.
[0055] In this embodiment, the image to be detected can be divided into multiple cells in advance, such as divided into an n×n network, and a position hint set of the same size is constructed , where each cell corresponds to a position feature vector, and each cell corresponds to a cell index respectively.
[0056] In an optional example, the target center point corresponding to the position information can be calculated based on the position information, then the target cell index where the target center point is located can be determined based on the target center point, and finally, based on the target cell index, retrieve the target position feature vector corresponding to the position information from the position hint set for the position information.
[0057] Step S42: Obtain the position hint information based on the target position feature vector.
[0058] In this embodiment, the position hint information can be obtained based on the retrieved target position feature vector, and the position hint information is used to pass the relative position information of the target in the image to be detected. Specifically, since there may be multiple valid targets (multiple objects) in the image to be detected, the target cell index can be determined based on multiple position information respectively, and the target position feature vectors corresponding to each can be retrieved from the position hint set, and then the multiple retrieved target position feature vectors are concatenated to obtain the position hint information corresponding to the image to be detected .
[0059] In this embodiment, in order to be consistent with the label hint information the position hint information also uses the way of learnable features, that is, the position hint information is learnable, and its weight will be continuously updated during the entire training process, so that the position hint becomes more refined and has context information as the model is trained.
[0060] In an optional example, the bounding box information obtained from the target detection network contains position information represented in the form of xyxy, which respectively represent the coordinates of the upper left corner and the lower right corner of the object. It should be noted that the coordinates of the position information used in this embodiment are in the normalized form, that is, the coordinate values are within the range of (0, 1).
[0061] Step S43: Determine the target size index based on the size information, and retrieve the target size feature vector corresponding to the size information from the size hint set based on the target size index.
[0062] In this embodiment, the target size index can be determined based on the size information in the bounding box information, and then the target size feature vector corresponding to the size information can be retrieved from the size hint set. The size hint set includes: a plurality of size feature vectors, each size feature vector corresponding to a different size index; the size feature vector is the size feature corresponding to one level after dividing the height and width of the object into multiple levels.
[0063] In this embodiment, similar to the position hint, the height and width of the object (i.e., the target) in the image to be detected can be divided into multiple levels in advance, for example, divided into n levels, and a size hint set with a shape of n×n can be constructed. Each level corresponds to a size feature vector. Among them, each size feature vector corresponds to a different size index.
[0064] In an optional example, the target size index corresponding to the size information can be calculated based on the height information and width information in the size information, and finally, based on the target size index, from the size hint set the target size feature vector corresponding to the size information can be retrieved.
[0065] Step S44: Obtain the size hint information based on the target size feature vector.
[0066] In this embodiment, the size hint information can be obtained based on the retrieved target size feature vector. Specifically, since there may be multiple valid targets (multiple objects) in the image to be detected, the target size index can be determined based on multiple size information respectively, and the corresponding target size feature vectors can be retrieved from the size hint set, and then the retrieved multiple target size feature vectors can be concatenated to obtain the size hint information corresponding to the image to be detected. 。
[0067] In this embodiment, in order to be consistent with the label hint information the size hint information also uses the way of learnable features, that is, the size hint information is learnable, and its weight will be continuously updated during the entire training process, so that the size hint becomes more refined and has context information as the model is trained.
[0068] In an optional implementation manner, according to the following formula, the height information and width information of the target can be used to calculate the corresponding target size index . Subsequently, the corresponding target size feature vector can be retrieved from the size hint set : ; Among them, n represents the level corresponding to the height and / or width.
[0069] In an optional embodiment, after obtaining the label hint information , position hint information and size hint information , these three hint information can be connected through the following formula, and the target hint information C is obtained through convolutional fusion: .
[0070] In an embodiment, as Figure 4 shown, Figure 4 is a schematic structural diagram of a target hint module shown in an embodiment of the present invention. In Figure 4 , using the target detection network (Detector) to perform inference on the input image (image to be detected), the label information of the objects appearing in the image to be detected and the bounding box information can be obtained. After that, the target hint module (OPM) is divided into two parts: label and bounding box, to encode the corresponding detection results into hint information.
[0071] This method first uses a text encoder (pre-trained CLIP) to separately encode each label (AllLabels) in the sample dataset, and each label will be converted into its corresponding text feature . After all the labels are converted into text features, this method combines these features together to form an initial label hint set . This process corresponds to the initialization step in Figure 4 .
[0072] During the training and inference process, the target hint module retrieves the corresponding feature vectors from the label hint set according to the label information provided by the target detection network (Detector). This process corresponds to the retrieval step in Figure 4 . Since there may be multiple valid targets in the input image, the retrieved feature vectors are concatenated together to form the label hint information of the current input image.
[0073] In Figure 4 , for the position hint, the image to be detected is first divided into an n×n grid, and an object position hint set of the same size is constructed , where each cell corresponds to a position feature. Next, the cell index where the object center point is located is calculated using the position information (x, y) in the bounding box information provided by the target detection network (Detector). Then, the corresponding position cue is retrieved from to provide the relative position information of the object in the image. Since there may be multiple valid objects in the input image, the retrieved position cues are concatenated together to form the position cue information of the current input image .
[0074] In Figure 4 , similar to the position cue, for the size cue, first divide the height and width of the object into n levels, and create a set of object size cues with a shape of n×n . Then, the index corresponding to the size cue can be calculated using the size information (the height and width information of the object) in the bounding box information provided by the target detection network (Detector). Subsequently, the corresponding size cue can be retrieved from the size cue set . Since there may be multiple valid objects in the input image, the retrieved size cues are concatenated together to form the size cue information of the current input image .
[0075] Finally, these three cues are concatenated and fused through Concat-Conv to form the final object cue information (Prompts).
[0076] Combined with the above embodiments, in one implementation, the present invention also provides a monocular image depth estimation method based on a diffusion model and object cues. In this embodiment, step S15 above may specifically include step S51 and step S52: Step S51: In the denoising network to be trained, perform single-step denoising on the shallow spatial representation and the object cue information.
[0077] Step S52: Interact the shallow spatial representation and the object cue information through a cross-attention mechanism to extract the multi-scale features.
[0078] In this embodiment, after the shallow spatial representation and the object cue information are input into the denoising network to be trained, in this denoising network to be trained, a single-step denoising process is performed on the shallow spatial representation and the object cue information, and the shallow spatial representation and the object cue information are interacted through a cross-attention mechanism to extract the multi-scale features, thereby avoiding the randomness generated by the conventional monocular depth estimation method based on a diffusion model that usually generates a depth image by gradually denoising from random noise.
[0079] In this embodiment, compared with gradually restoring from noise to a depth map in a traditional diffusion model, on the one hand, the prediction efficiency of depth estimation is improved by performing single-step denoising. On the other hand, the object bounding box and label are extracted through the object detection network of this embodiment, and object prompts are generated through the object prompt module. Compared with generating an image based on a caption to describe the image in the related art, its model parameters are smaller and the generation efficiency of prompt information is higher, thereby further improving the prediction efficiency of depth estimation, overcoming the problem of the high time cost in the inference process due to the inherently iterative denoising characteristics of the diffusion model itself in the related art, and improving the stability and inference efficiency of the monocular depth estimation method based on the diffusion model.
[0080] Currently, related methods have deficiencies in multi-scale feature fusion and depth information restoration, resulting in low feature utilization: Specifically, when dealing with multi-scale features, related methods usually directly splice features of different scales. This simple feature fusion method cannot effectively capture the complementarity between multi-scale features, resulting in limited prediction accuracy. The differences in spatial resolution and semantic information of features at each scale are not considered in the related art.
[0081] To solve this problem, in combination with the above embodiments, in one implementation manner, the present invention also provides a monocular image depth estimation method based on a diffusion model and object prompts. In this method, the multi-scale features at least include: a first resolution feature, a second resolution feature, a third resolution feature, and a fourth resolution feature with increasing resolution from small to large; and, the above step S16 can specifically include steps S61 to S63: Step S61: Use the first resolution feature as the global feature and the second resolution feature as the local feature, and perform interaction between the global feature and the local feature to obtain a first interaction fusion feature.
[0082] In this embodiment, in order to make full use of the information between features of different scales, a global-local interaction decoder (i.e., interaction decoder) is constructed, aiming to promote the interaction between multi-scale features and effectively fuse them. Among them, the global feature in this embodiment is a feature with a smaller resolution and more semantic information; the local feature is a feature with a larger resolution and rich detail information.
[0083] First, use the first resolution feature in the multi-scale features as the global feature and the second resolution feature in the multi-scale features as the local feature, and perform interaction between the global feature and the local feature to obtain a first interaction fusion feature.
[0084] Step S62: Use the first interaction fusion feature as the global feature and the third resolution feature as the local feature, and perform an interaction between the global feature and the local feature to obtain a second interaction fusion feature.
[0085] In this embodiment, after obtaining the first interaction fusion feature, use the first interaction fusion feature as the global feature and the third resolution feature in the multi-scale features as the local feature, and perform an interaction between the global feature and the local feature again to obtain a second interaction fusion feature.
[0086] Step S63: Use the second interaction fusion feature as the global feature and the fourth resolution feature as the local feature, and perform an interaction between the global feature and the local feature to obtain the sample depth map.
[0087] In this embodiment, after obtaining the second interaction fusion feature, use the second interaction fusion feature as the global feature and the fourth resolution feature in the multi-scale features as the local feature, and perform an interaction between the global feature and the local feature again to obtain the sample depth map output by the interaction decoder.
[0088] In some other alternative embodiments, the multi-scale features may include more or fewer size features. For example, they may include size features of 2 different resolutions, or more size features such as 6, 8, etc. At this time, it is similar to the steps including size features of 4 different resolutions: first, perform an interaction between the two with the smallest resolutions through the global feature and the local feature to obtain a fusion feature, then use the fusion feature as the global feature and the feature with the third smallest resolution as the local feature to continue performing an interaction between the global feature and the local feature to obtain a fusion feature, and so on, until the feature with the largest resolution is used as the local feature, and perform an interaction between the global feature and the local feature with the final fusion feature obtained from all previous resolution features as the global feature to obtain the output sample depth map. Among them, the operation of performing an interaction between the global feature and the local feature each time is the same.
[0089] In this embodiment, through the proposed interaction decoder, utilize the complementary characteristics of the global feature and the local feature to promote the fusion of features of different types, while retaining detailed information, and finally enhance the model's overall understanding ability of the scene.
[0090] Combined with the above embodiments, in one implementation, the present invention also provides a monocular image depth estimation method based on a diffusion model and target prompts. In this method, the "performing an interaction between the global feature and the local feature" in the above step S61, step S62, or step S63 may specifically include steps S71 to S75: Step S71: Extract a channel attention map from the global features, and perform element-wise multiplication of the channel attention map and the local features to obtain a first fused feature.
[0091] In this embodiment, after determining the local features and the global features, first extract a channel attention map from the global features, and perform element-wise multiplication of the obtained channel attention map and the local features to obtain a first fused feature, thereby realizing extracting a channel attention map from the global features to provide semantic guidance for the local features.
[0092] Step S72: Extract a spatial attention map from the local features, and perform element-wise multiplication of the spatial attention map and the global features to obtain a second fused feature.
[0093] In this embodiment, after determining the local features and the global features, a spatial attention map can also be extracted from the local features, and perform element-wise multiplication of the obtained spatial attention map and the global features to obtain a second fused feature, thereby realizing extracting a spatial attention map from the local features to provide detail guidance for the global features.
[0094] It should be noted that this embodiment does not limit the execution order between step S71 and step S72. For example, step S71 can be executed first and then step S72, or step S72 can be executed first and then step S71, or step S71 and step S72 can be executed simultaneously.
[0095] Step S73: Perform element-wise addition of the first fused feature and the local features to obtain a third fused feature.
[0096] In this embodiment, after obtaining the first fused feature, the first fused feature and the local features can be added element-wise to obtain a third fused feature.
[0097] Step S74: Perform element-wise addition of the second fused feature and the global features to obtain a fourth fused feature.
[0098] In this embodiment, after obtaining the second fused feature, the second fused feature and the global features can be added element-wise to obtain a fourth fused feature.
[0099] Step S75: Perform splicing fusion on the fourth fused feature and the third fused feature.
[0100] In this embodiment, after obtaining the third fused feature and the fourth fused feature, the fourth fused feature and the third fused feature can be concatenated and fused. For example, the third fused feature and the fourth fused feature can be concatenated (Concat) first and then convolutionally fused (Conv-BN-Norm) to obtain an interactive fused feature, such as the first interactive fused feature, the second interactive fused feature, or the sample depth map.
[0101] In this embodiment, in view of the characteristics of the monocular depth estimation task, a dedicated decoder module (interactive decoder) is designed. By using different types of attention mechanisms, this module fully considers the characteristics of different feature types and generates customized attention maps for each feature, so as to achieve efficient fusion of multi-scale features and accurate reconstruction of depth information.
[0102] In a specific example, the interactive fused feature can be obtained through the following formula: ; where is the global feature, is the channel attention map, is the spatial attention map, is the local feature, Concat is concatenation, Conv is convolutional fusion, is the fourth fused feature, is the third fused feature.
[0103] In an alternative embodiment, before calculating the channel attention map, a 2-fold upsampling operation is first performed on the global feature to align the sizes of the two features.
[0104] In an alternative embodiment, the calculation method of the channel attention map is as follows: ; where is the global feature, is average pooling, Conv is convolution, is the activation function, where before performing average pooling, is first upsampled by 2 times.
[0105] In another alternative embodiment, the calculation method of the spatial attention map is as follows: ; Among them, in order to expand the receptive field, this embodiment uses the large kernel attention mechanism (LKA) to extract the spatial attention map. LKA is a lightweight attention mechanism that decomposes a convolution kernel with a large number of parameters into a depthwise convolution (DW-Conv), a depthwise dilated convolution (DW-D-Conv), and a 1×1 convolution (1×1 Conv). While maintaining the receptive field, it can reduce the overall number of parameters. is the local feature, is the activation function.
[0106] In one embodiment, as Figure 5 shown, Figure 5 is a schematic structural diagram of an interactive decoder shown in an embodiment of the present invention. In Figure 5 , specifically, first, the channel attention map is extracted from the global feature to provide semantic guidance for the local feature , and at the same time, the spatial attention map is extracted from the local feature to provide detailed guidance for the global feature . Finally, the interacted features are concatenated (Concat) and then convolutionally fused (Conv-BN-Norm) to obtain the interactive fusion feature .
[0107] In one embodiment, as Figure 6 shown, Figure 6 is a flowchart of a monocular image depth estimation method based on a diffusion model and target prompts shown in an embodiment of the present invention. In Figure 6 , the input image is loaded and input into the trained depth estimation model. In the trained depth estimation model, the object category and detection box information in the image are extracted by the target detector, and then the category and detection box are encoded as text prompts (i.e., target prompt information) based on the target prompt module, and, the input image is mapped to the shallow space by the image encoder to obtain the shallow space representation; then, the shallow space representation and the text prompt are denoised through the denoising network and multi-scale features are extracted. Finally, the interactive decoder restores the depth information from the multi-scale features and predicts the depth map.
[0108] Combined with the above embodiments, in one implementation manner, the present invention further provides a monocular image depth estimation method based on a diffusion model and a target prompt. In this method, the image to be detected is an image captured in real time by a vehicle-mounted camera, and the depth map includes the depth values of each pixel in the image; and, after the above step S18, steps S81 and S82 may further be included: Step S81: Based on the depth map, compare the size of the first region composed of pixels with depth values less than a preset safety distance with a preset region.
[0109] In this embodiment, after obtaining the depth map corresponding to the image to be detected, which includes the depth values of each pixel in the image, the size of the first region composed of pixels with depth values less than a preset safety distance (such as 1 m, which can be freely set) can be compared with a preset region based on the depth map.
[0110] Step S82: When the first region is larger than the preset region, mark the first region as an obstacle for warning and display.
[0111] In this embodiment, when it is determined that the first region is larger than or equal to the preset region, it can be determined that the first region is a potential obstacle, and the first region is marked as an obstacle for warning and display to implement obstacle detection for autonomous driving and achieve warning prompts during the autonomous driving process.
[0112] In an alternative embodiment, an online monocular depth estimation system is provided. The online monocular depth estimation system is designed to perform depth prediction on the pictures provided by users and return visual results to help users better understand the spatial structure and object distribution in the images. The specific flowchart of the system is as Figure 7 shown Figure 7 is the training and inference flowchart of a monocular depth estimation system shown in an embodiment of the present invention. In Figure 7 it, the system specifically includes two main processes: model training and real-time inference. Since the depth distribution ranges of indoor and outdoor scenes are quite different, it is difficult for a single model to take both scenes into account. Therefore, the system trains and infers them separately. The system hardware configuration and software version are shown in Table 1: Table 1 System Hardware Configuration and Software Version Information Table
[0113] Among them, the training stage: During the training phase, the indoor scenes are trained using the NYU Depth V2 dataset, which contains 464 indoor scenes, a total of more than 24,000 RGB images and their corresponding depth maps, with each image having a resolution of 640×480; for outdoor scenes, the KITTI dataset is used for training, which is commonly used in the field of autonomous driving. It contains RGB images, LiDAR data, and camera calibration information taken from different perspectives, with a total of approximately 23,000 images sized 1238×374. To enhance data diversity and improve the generalization ability of the model, this system uses methods such as random cropping, horizontal flipping, contrast adjustment, and hue adjustment to augment the training data. Among them, the cropping size for the NYU Depth V2 dataset is 576×448, and the cropping size for the KITTI dataset is 768×352, and the triggering probability of the remaining augmentation methods is set to 50%.
[0114] In the part of model construction, this system uses Pytorch 2.0.0 to implement the network structure and uses the scale-invariant loss L to train the model. The specific calculation method is as follows: ; where, and represent the depth map values corresponding to the i-th pixel in the real depth map and the model-predicted depth map respectively, and N is the total number of image pixels. A one-cycle learning rate strategy with an Adam optimizer is used. The learning rate increases from to in the first half, and then decreases from to in the second half. To prevent overfitting, the weight decay of the diffusion model part is set to 0.01. The scaling factor set in Stable Diffusion is used to scale the output of the latent encoder so that the rescaled latent features have unit standard deviation. 4 V100-32G cards are used for training, with a batch size of 3 for each card, and a total of 25 rounds of training are conducted. After each round of training, the model's performance on the test set is evaluated, and the corresponding weights are saved.
[0115] Inference phase: During the inference phase, the system receives the user-input image and loads it into memory. Then, it preprocesses the image, padding it to make its size a multiple of 64. After that, it loads the model with the best performance during the training phase to perform inference on the input image and obtains the predicted depth map. Since the depth map obtained by model inference is a 16-bit single-channel image, it needs to be converted into a three-channel RGB visualization image for the user to view. Finally, the depth image and the converted visualization image are output and displayed.
[0116] In addition, in another embodiment, an autonomous driving obstacle detection system is proposed. This autonomous driving obstacle detection system aims to detect obstacles in front of the vehicle in real time through monocular depth estimation technology and provide distance prompts to assist the driver or the autonomous driving system in making safe decisions. The specific flowchart of the system is as Figure 8 shown Figure 8 is the training and inference flowchart of an autonomous driving obstacle detection system shown in an embodiment of the present invention. To adapt to the actual road environment, the system uses a self-collected dataset for training and testing to ensure that the model can adapt to diverse road scenarios. Before training and inference, it is necessary to calibrate the in-vehicle camera to obtain the internal and external parameters of the camera to correct the image and eliminate the influence of lens distortion.
[0117] The process and parameter settings in the training phase are the same as those in the previous embodiment, and the only difference between the two lies in the dataset used. To construct a high-quality training dataset, this system uses in-vehicle cameras and LiDAR devices to collect data in the actual road environment. The collected data includes: ① RGB images: with a resolution of 1920×1080, covering various lighting and weather conditions such as day, night, rain, and fog. ② Depth information: obtaining high-precision depth maps through LiDAR devices for supervising the training of the model. The total number of collected datasets is about 5000 images, covering various scenarios such as urban roads, highways, and rural roads.
[0118] During the inference phase, the system first loads the weights of the model with the best performance during the training phase. Subsequently, it receives the images captured in real time by the in-vehicle camera, loads them into memory for inference, and obtains the predicted depth map. Based on the depth map, the system further performs obstacle detection. It calculates the distance information of each area in front of the vehicle according to the depth map. Through threshold segmentation and connected component analysis, it identifies the areas in the depth map where the distance is less than the preset safety distance, marks them as potential obstacles, and issues a warning prompt to the driver or the autonomous driving system.
[0119] In one embodiment, the present invention has a more accurate depth prediction effect compared with other methods. Table 2 shows the performance comparison results of the present invention with other methods on the KITTI dataset.
[0120] Table 2 Quantitative comparison results on the KITTI dataset
[0121] Compared with other monocular depth estimation methods, the present invention shows obvious advantages on the KITTI dataset and achieves leading results in terms of accuracy and precision. Specifically, the root mean square error (RMSE) of the present invention (Ours) is 1.967, and the relative error value (REL) is 0.047, both of which are the best among all methods, and it also ranks first in terms of the accuracy rate index. The calculation formulas for specific indicators are as follows: ; Among them, in the above formula, represents the predicted depth of the th pixel point, represents the true depth of the th pixel point, represents the total number of pixel points in the image. The calculation method of the accuracy rate is specifically as follows: For each pixel point in the image, calculate the maximum value between and , and count the percentage of pixel points whose statistical results fall within the given threshold interval in the total pixel points. During the evaluation process, the threshold threshold is often taken as 1.25, and n is usually one of the three values 1, 2, and 3.
[0122] In addition, the present invention is significantly superior to other methods in terms of the accuracy and continuity of the prediction results. For example, traffic signs in Figure 9 and objects such as plants and computers in Figure 10 have clearer contours, which is attributed to the use of object information as a hint by this method. Among them, Figure 9 is a visualization diagram of the prediction results of different methods shown in an embodiment of the present invention. Figure 10 is a visualization diagram of the influence of the detected target on the depth prediction of the results shown in an embodiment of the present invention.
[0123] The present invention demonstrates excellent generalization ability by leveraging the rich prior knowledge provided by the diffusion model and is superior to other methods in various performance indicators. The specific comparison results are shown in Table 3. All methods are trained on the NYU Depth V2 indoor dataset and then evaluated for performance on the unseen SUN RGB-D dataset.
[0124] Table 3 Comparison of the generalization ability of different methods
[0125] It should be noted that, for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the described action sequences, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0126] Based on the same inventive concept, an embodiment of the present invention provides a monocular image depth estimation device based on a diffusion model and a target prompt. Refer to Figure 11 , Figure 11 which is a structural block diagram of a monocular image depth estimation device based on a diffusion model and a target prompt provided by an embodiment of the present invention. As Figure 11 shown, the monocular image depth estimation device based on a diffusion model and a target prompt in this embodiment may include: An image input module, configured to input a sample detection image into a depth estimation model to be trained, where the depth estimation model to be trained at least includes: a pre-trained image encoder, a pre-trained target detection network, a target prompt module to be trained, a denoising network to be trained, and an interactive decoder to be trained; An image encoding module, configured to process the sample detection image through the pre-trained image encoder to obtain a shallow spatial representation corresponding to the sample detection image; A target detection module, configured to process the sample detection image through the pre-trained target detection network to extract a target detection result of the sample detection image; A prompt acquisition module, configured to process the target detection result through the target prompt module to be trained to obtain target prompt information corresponding to the sample detection image; A multi-scale feature extraction module, configured to input the shallow spatial representation and the target prompt information into the denoising network to be trained to extract multi-scale features; A depth map output module, configured to input the multi-scale features into the interactive decoder to be trained to obtain a sample depth map; A model training module, configured to train the target prompt module to be trained, the denoising network to be trained, and the interactive decoder to be trained based on the label depth map corresponding to the sample detection image and the sample depth map until a trained depth estimation model is obtained; A depth estimation module, configured to input an image to be detected into the trained depth estimation model to obtain a depth map corresponding to the image to be detected.
[0127] Optionally, the object detection result includes: class label information and bounding box information of an object; the hint acquisition module includes: The label hint module is configured to process the class label information in the target hint module to be trained to obtain label hint information corresponding to the sample detection image; The bounding box hint module is configured to process the bounding box information to obtain position hint information and size hint information corresponding to the sample detection image; The hint generation module is configured to connect the label hint information, the position hint information, and the size hint information, and obtain the target hint information through convolutional fusion.
[0128] Optionally, the label hint module includes: The first retrieval module is configured to retrieve a target label feature vector corresponding to the class label information from a label hint set based on the class label information; The first generation module is configured to obtain the label hint information based on the target label feature vector; Wherein, the label hint set includes: a plurality of label feature vectors, each label feature vector corresponding to different object label information; the label feature vector is a text feature encoded based on the object label information in the training dataset corresponding to the object detection network.
[0129] Optionally, the bounding box information includes: position information and size information corresponding to an object; the bounding box hint module includes: The second retrieval module is configured to determine a target cell index based on the position information, and retrieve a target position feature vector corresponding to the position information from a position hint set based on the target cell index; The second generation module is configured to obtain the position hint information based on the target position feature vector; The third retrieval module is configured to determine a target size index based on the size information, and retrieve a target size feature vector corresponding to the size information from a size hint set based on the target size index; The third generation module is configured to obtain the size hint information based on the target size feature vector; Wherein, the position hint set includes: a plurality of position feature vectors, each position feature vector corresponding to a different cell index; the position feature vector is the position feature corresponding to a cell after dividing the sample detection image into a plurality of cells; The size hint set includes: a plurality of size feature vectors, each size feature vector corresponding to a different size index; the size feature vector is the size feature corresponding to one level after dividing the height and width of the object into multiple levels.
[0130] Optionally, the multi-scale feature extraction module includes: A single-step denoising module for performing single-step denoising on the shallow spatial representation and the target hint information in the denoising network to be trained; A feature interaction module for interacting the shallow spatial representation and the target hint information through a cross-attention mechanism to extract the multi-scale features.
[0131] Optionally, the multi-scale features at least include: a first resolution feature, a second resolution feature, a third resolution feature, and a fourth resolution feature in ascending order of resolution; The depth map output module includes: A first interaction module for using the first resolution feature as the global feature and the second resolution feature as the local feature to interact between the global feature and the local feature to obtain a first interaction fusion feature; A second interaction module for using the first interaction fusion feature as the global feature and the third resolution feature as the local feature to interact between the global feature and the local feature to obtain a second interaction fusion feature; A third interaction module for using the second interaction fusion feature as the global feature and the fourth resolution feature as the local feature to interact between the global feature and the local feature to obtain the sample depth map.
[0132] Optionally, the first interaction module, the second interaction module, or the third interaction module includes: A first processing module for extracting a channel attention map from the global feature, multiplying the channel attention map element-wise with the local feature to obtain a first fusion feature; A second processing module for extracting a spatial attention map from the local feature, multiplying the spatial attention map element-wise with the global feature to obtain a second fusion feature; A third processing module for adding the first fusion feature and the local feature element-wise to obtain a third fusion feature; A fourth processing module for adding the second fusion feature and the global feature element-wise to obtain a fourth fusion feature; A fifth processing module for splicing and fusing the fourth fusion feature and the third fusion feature.
[0133] Optionally, the image to be detected is an image captured in real time by a vehicle-mounted camera, and the depth map includes the depth values of each pixel in the image; the apparatus further includes: A depth comparison module, configured to, after obtaining the depth map corresponding to the image to be detected, based on the depth map, compare the size of a first region composed of pixels with depth values less than a preset safety distance with a preset region; An early warning module, configured to, when the first region is larger than the preset region, mark the first region as an obstacle for early warning and display.
[0134] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes, it implements the steps in the monocular image depth estimation method based on a diffusion model and a target prompt described in any one of the above embodiments of the present invention.
[0135] For the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0136] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0137] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0138] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, and the combination of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1a device for the functions specified in one or more boxes.
[0139] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 a box or more boxes.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process. Therefore, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 a box or more boxes.
[0141] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0142] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal device including the said element.
[0143] The above has introduced in detail a monocular image depth estimation method, device and equipment based on a diffusion model and target prompts. In this article, specific examples are used to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A monocular image depth estimation method based on diffusion model and target hint, characterized in that: The method comprises: Inputting the sample detection image into a depth estimation model to be trained, wherein the depth estimation model to be trained includes at least: a pre-trained image encoder, a pre-trained target detection network, a target prompt module to be trained, a denoising network to be trained, and an interactive decoder to be trained; Processing the sample detection image by using the pre-trained image encoder to obtain a shallow spatial representation corresponding to the sample detection image; Processing the sample detection image through the pre-trained target detection network to extract the target detection result of the sample detection image; Processing the target detection result by the target prompt module to be trained to obtain target prompt information corresponding to the sample detection image; Inputting the shallow spatial representation and the target prompt information into the denoising network to be trained to extract multi-scale features; Inputting the multi-scale features into the interactive decoder to be trained to obtain a sample depth map; Based on the label depth map corresponding to the sample detection image and the sample depth map, the target prompt module to be trained, the denoising network to be trained and the interactive decoder to be trained are trained until a trained depth estimation model is obtained; The image to be detected is input into the trained depth estimation model to obtain a depth map corresponding to the image to be detected.
2. The monocular image depth estimation method based on diffusion model and target hint according to claim 1, characterized in that: The target detection result includes: category label information and bounding box information of the object; the target detection result is processed by the target prompt module to be trained to obtain target prompt information corresponding to the sample detection image, including: In the target prompt module to be trained, the category label information is processed to obtain label prompt information corresponding to the sample detection image; Processing the bounding box information to obtain position prompt information and size prompt information corresponding to the sample detection image; The label prompt information, the position prompt information and the size prompt information are connected, and the target prompt information is obtained through convolution fusion.
3. The monocular image depth estimation method based on diffusion model and target hint according to claim 2, characterized in that: Processing the category label information to obtain label prompt information corresponding to the sample detection image includes: Based on the category label information, a target label feature vector corresponding to the category label information is retrieved from a label prompt set; Based on the target label feature vector, obtaining the label prompt information; The label prompt set includes: a plurality of label feature vectors, each of which corresponds to different object label information; the label feature vector is a text feature obtained by encoding the object label information in the training data set corresponding to the target detection network.
4. The monocular image depth estimation method based on diffusion model and target hint according to claim 2, characterized in that: The bounding box information includes: position information and size information corresponding to the object; processing the bounding box information to obtain position prompt information and size prompt information corresponding to the sample detection image includes: Determine a target cell index based on the position information, and retrieve a target position feature vector corresponding to the position information from a position prompt set based on the target cell index; Based on the target position feature vector, obtaining the position prompt information; Determine a target size index based on the size information, and retrieve a target size feature vector corresponding to the size information from a size prompt set based on the target size index; Based on the target size feature vector, obtaining the size prompt information; The position prompt set includes: a plurality of position feature vectors, each of which corresponds to a different cell index; the position feature vector is a position feature corresponding to a cell after the sample detection image is divided into a plurality of cells; The size prompt set includes: a plurality of size feature vectors, each of which corresponds to a different size index; the size feature vector is a size feature corresponding to a level after the height and width of the object are divided into multiple levels.
5. The monocular image depth estimation method based on diffusion model and target hint according to claim 1, characterized in that: Inputting the shallow spatial representation and the target prompt information into the denoising network to be trained, and extracting multi-scale features, including: In the denoising network to be trained, performing single-step denoising on the shallow spatial representation and the target prompt information; The shallow spatial representation and the target prompt information are interacted through a cross-attention mechanism to extract the multi-scale features.
6. The monocular image depth estimation method based on diffusion model and target hint according to claim 1, characterized in that: The multi-scale features include at least: a first resolution feature, a second resolution feature, a third resolution feature and a fourth resolution feature in increasing order of resolution; Inputting the multi-scale features into the interactive decoder to be trained to obtain a sample depth map, including: Taking the first resolution feature as a global feature and the second resolution feature as a local feature, performing interaction between the global feature and the local feature to obtain a first interactive fusion feature; Taking the first interactive fusion feature as the global feature, taking the third resolution feature as the local feature, and performing interaction between the global feature and the local feature to obtain a second interactive fusion feature; The second interactive fusion feature is used as the global feature, the fourth resolution feature is used as the local feature, and the global feature and the local feature are interacted to obtain the sample depth map.
7. The monocular image depth estimation method based on diffusion model and target hint according to claim 6, characterized in that: Interacting the global feature with the local feature includes: Extract a channel attention map from the global feature, and multiply the channel attention map by the local feature element by element to obtain a first fusion feature; Extracting a spatial attention map from the local features, and multiplying the spatial attention map by the global features element by element to obtain a second fused feature; Adding the first fused feature and the local feature element by element to obtain a third fused feature; Adding the second fused feature to the global feature element by element to obtain a fourth fused feature; The fourth fusion feature and the third fusion feature are spliced and fused.
8. The monocular image depth estimation method based on diffusion model and target hint according to any one of claims 1 to 7, characterized in that: The image to be detected is an image captured in real time by a vehicle-mounted camera, and the depth map includes the depth value of each pixel in the image; After obtaining the depth map corresponding to the image to be detected, the method further includes: Based on the depth map, a first area consisting of pixels having depth values less than a preset safety distance is compared in size with a preset area; When the first area is larger than the preset area, the first area is marked as an obstacle for early warning and display.
9. A monocular image depth estimation device based on diffusion model and target hint, characterized in that: The device comprises: An image input module, used for inputting a sample detection image into a depth estimation model to be trained, wherein the depth estimation model to be trained comprises at least: a pre-trained image encoder, a pre-trained target detection network, a target prompt module to be trained, a denoising network to be trained, and an interactive decoder to be trained; An image encoding module, used for processing the sample detection image through the pre-trained image encoder to obtain a shallow spatial representation corresponding to the sample detection image; An object detection module is used to process the sample detection image through the pre-trained object detection network to extract the object detection result of the sample detection image; A prompt acquisition module, used for processing the target detection result through the target prompt module to be trained to obtain target prompt information corresponding to the sample detection image; A multi-scale feature extraction module, used for inputting the shallow spatial representation and the target prompt information into the denoising network to be trained to extract multi-scale features; A depth map output module, used for inputting the multi-scale features into the interactive decoder to be trained to obtain a sample depth map; A model training module, used for training the target prompt module to be trained, the denoising network to be trained and the interactive decoder to be trained based on the label depth map corresponding to the sample detection image and the sample depth map, until a trained depth estimation model is obtained; The depth estimation module is used to input the image to be detected into the trained depth estimation model to obtain a depth map corresponding to the image to be detected.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is executed by the processor, the monocular image depth estimation method based on diffusion model and target prompt is implemented as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Transparent object single-view multi-scale depth estimation method and system
CN113139999A
Generating shadows for placement objects in depth estimation scene of two-dimensional image
CN117830473A
Modifying digital images via depth-aware object movement
CN117853613A
Monocular depth estimation and visual odometer method based on diffusion model
CN117911477A
Semantic discrete monocular depth estimation method based on adaptive pixel clustering
CN118570271A