Absolute depth estimation method and depth image imaging device
By inputting low-quality depth maps and 2D maps into trained absolute depth estimation models, using multi-layer networks to extract and fusion features, the problem of low absolute depth estimation accuracy is solved, and high-precision absolute depth estimation is achieved.
Patent Information
- Application Number
- CN202510077475.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art has low accuracy in absolute depth estimation and is difficult to accurately recover absolute depth values that match the actual physical world, especially in applications requiring precise spatial positioning.
By obtaining the low-quality depth map and 2D map of the target scene, it is input into the trained absolute depth estimation model. The model includes a third network for extracting 2D image features, the first network for extracting and processing features of the low-quality depth map so that its resolution is the same as the 2D image features, and the second network outputs the absolute depth map based on both.
High-precision absolute depth estimation of the target scene is achieved, the accuracy of absolute depth estimation of the image is improved, and the problem of low accuracy in the prior art is solved.
Smart Images

Figure CN120031941A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing, and in particular to an absolute depth estimation method and a depth image imaging device. Background Art
[0002] With the development of three-dimensional (3D) imaging technology and applications such as augmented reality (AR) and virtual reality (VR), the demand for accurately measuring the distance of objects in the environment is growing. Currently, there are many methods to obtain environmental depth information, among which the more common ones include time of flight (ToF) technology and structured light technology. These technologies achieve high-precision depth mapping by emitting a specific pattern of light and calculating the distance of the target based on the time or phase change of the reflected light. However, in order to achieve high-quality depth measurement effects, such systems usually require the use of expensive professional hardware equipment, which limits their popularity in a wider range of application scenarios.
[0003] At the same time, deep learning-based depth estimation methods have also developed rapidly in recent years. These methods can use single or multiple images as input and predict the relative depth of each pixel in the image relative to other points by training models. Although this method has been able to provide satisfactory relative depth information in many cases, it still has shortcomings in absolute depth estimation. Due to the lack of direct physical measurement methods, learning-based methods often find it difficult to accurately restore absolute depth values that are consistent with the actual physical world, which poses a challenge to some applications that require precise spatial positioning.
[0004] Therefore, developing a new method that can effectively improve the accuracy of absolute depth estimation is of great significance to promoting the advancement of related technologies. Summary of the invention
[0005] Based on this, it is necessary to provide an absolute depth estimation method and a depth image imaging device that can solve the above-mentioned technical problems.
[0006] In a first aspect, the present application provides an absolute depth estimation method, the method comprising:
[0007] Obtain low-quality depth maps and 2D images of the target scene;
[0008] Input the low-quality depth map and the 2D map into a trained and complete absolute depth estimation model to obtain the absolute depth map of the target scene. Among them, the trained and complete absolute depth estimation model includes a first network, a second network, and a third network. The third network is used to extract the image features of the 2D image. The first network is used to extract the image features of the low-quality depth map and process the image features of the low-quality depth map into the same resolution as the image features of the 2D image extracted by the third network. The second network is used to output the absolute depth map according to the image features of the 2D image and the processed image features of the low-quality depth map;
[0009] Perform absolute depth estimation on the target scene according to the absolute depth map.
[0010] In some embodiments, the first network includes a shallow convolutional network, the second network includes a DPT network, and the third network includes a ViT network; inputting the low-quality depth map and the 2D map into a trained and complete absolute depth estimation model to obtain the absolute depth map of the target scene includes:
[0011] Extract the first image features of the 2D map through the ViT network;
[0012] Extract the second image features of the low-quality depth map through the shallow convolutional network;
[0013] Process the resolution of the second image features into the same as the first image features through the shallow convolutional network;
[0014] Input the second image features processed by the shallow convolutional network and the first image features into the DPT network to obtain the absolute depth map.
[0015] In some embodiments, the shallow convolutional network is further used to denoise the second image features.
[0016] In some embodiments, the training method of the trained and complete absolute depth estimation model includes:
[0017] Obtain the image data of the target scene, where the image data includes: 2D images, low-quality depth images, and 3D data;
[0018] Perform three-dimensional reconstruction on the 2D image to obtain a pseudo-ground truth depth image;
[0019] Synthesize the pseudo-ground truth depth image and the 3D data to obtain the ground truth depth image;
[0020] Constructing an image training pair by combining the 2D image and the low-quality depth image under the same target scene with the true depth image;
[0021] The image training pair is input into the absolute depth estimation model to be trained, and the training is performed until convergence to obtain a fully trained absolute depth estimation model.
[0022] In some of the embodiments, acquiring image data of the target scene includes:
[0023] Acquire a 2D image and a low-quality depth image of the target scene based on a preset shooting device;
[0024] The target scene is three-dimensionally scanned to obtain 3D data of the target scene, wherein the 3D data includes one of the following: grid data, point cloud data, and depth image data.
[0025] In some embodiments, synthesizing the pseudo true depth image with the 3D data to obtain the true depth image includes:
[0026] Acquire an image gradient of the pseudo true value depth image and an image depth of the 3D data;
[0027] The image gradient and the image depth are input into a preset neural network model to obtain the true value depth image.
[0028] In some embodiments, the training method of the fully trained absolute depth estimation model includes:
[0029] Acquire three-dimensional data of the target scene;
[0030] Performing a first preset processing on the three-dimensional data to obtain a 2D image and a true depth image;
[0031] Performing a second preset processing on the true depth image to obtain a low-quality depth image;
[0032] Constructing an image training pair by combining the 2D image and the low-quality depth image under the same target scene with the true depth image;
[0033] The image training pair is input into the absolute depth estimation model to be trained, and the training is performed until convergence to obtain a fully trained absolute depth estimation model.
[0034] In some embodiments, the first preset processing includes: projection processing and rendering processing.
[0035] In some embodiments, the second preset processing includes: blurring processing, noise adding processing, downsampling processing and sparse anchor frame interpolation processing.
[0036] In some of these embodiments, the 2D image includes: an RGB image and / or a grayscale image.
[0037] In a second aspect, the present application provides a depth image imaging device, the device including:
[0038] An image acquisition unit, configured to acquire a low-quality depth map and a 2D map of a target scene;
[0039] A deep learning inference unit, configured to input the low-quality depth map and the 2D map into a trained complete absolute depth estimation model to obtain an absolute depth map of the target scene, where the trained complete absolute depth estimation model includes a first network, a second network, and a third network, the third network is configured to extract image features of the 2D image, the first network is configured to extract image features of the low-quality depth map and process the image features of the low-quality depth map into the same resolution as the image features of the 2D image extracted by the third network, and the second network is configured to output an absolute depth map according to the image features of the 2D image and the processed image features of the low-quality depth map;
[0040] A processing unit, configured to perform absolute depth estimation on the target scene according to the absolute depth map.
[0041] In some of these embodiments, the first network includes a shallow convolutional network, and the shallow convolutional network is further configured to perform denoising processing on the image features of the low-quality depth map.
[0042] In some of these embodiments, the depth image device further includes a camera unit, the camera unit is configured to acquire a low-quality depth map and a 2D image of a target scene and send them to the image acquisition unit, where the camera unit is wirelessly connected to the image acquisition unit, and / or, the camera unit is wiredly connected to the image acquisition unit.
[0043] The above-mentioned absolute depth estimation method and depth image imaging device obtain a low-quality depth map and a 2D map of the target scene; input the low-quality depth map and the 2D map into a fully trained absolute depth estimation model to obtain an absolute depth map of the target scene, wherein the fully trained absolute depth estimation model includes a first network, a second network and a third network, the third network is used to extract image features of the 2D image, the first network is used to extract image features of the low-quality depth map and process the image features of the low-quality depth map to the same resolution as the image features of the 2D image extracted by the third network, and the second network is used to output an absolute depth map according to the image features of the 2D image and the image features of the processed low-quality depth map; according to the absolute depth map, the absolute depth estimation of the target scene is performed, thereby realizing the absolute depth estimation of the captured target scene, solving the problem of low accuracy of absolute depth estimation of images in the prior art, and improving the accuracy of absolute depth estimation of images. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0045] Figure 1 is a schematic flow chart of an absolute depth estimation method in one embodiment;
[0046] Figure 2 is a flowchart of a training method for an absolute depth estimation model in one embodiment;
[0047] Figure 3 is a flowchart of a training method for an absolute depth estimation model in another embodiment;
[0048] Figure 4 is a flow chart of an absolute depth estimation method according to another embodiment;
[0049] Figure 5 4 is a structural block diagram of a depth image imaging device in one embodiment. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0051] In an exemplary embodiment, Figure 1As shown, an absolute depth estimation method is provided, comprising the following steps S102 to S110. Among them:
[0052] Step S102: Obtain a low-quality depth map and a 2D map of the target scene.
[0053] It should be noted that a low-quality depth image generally refers to a depth image with certain defects or deficiencies, and these defects may include but are not limited to: noise, low resolution, etc.
[0054] In this step, a 2D image can be acquired by a digital camera / smartphone camera.
[0055] Step S104, inputting the low-quality depth map and the 2D map into a fully trained absolute depth estimation model to obtain an absolute depth map of the target scene, wherein the fully trained absolute depth estimation model includes a first network, a second network, and a third network, the third network is used to extract image features of the 2D image, the first network is used to extract image features of the low-quality depth map and process the image features of the low-quality depth map to the same resolution as the image features of the 2D image extracted by the third network, and the second network is used to output an absolute depth map according to the image features of the 2D image and the image features of the processed low-quality depth map.
[0056] In this step, the trained absolute depth estimation model processes the image including the following steps:
[0057] Step S1041, obtaining a low-quality depth map and a 2D map.
[0058] Step S1042: segment both the low-quality depth map and the 2D image into image blocks of fixed sizes.
[0059] In step S1043, the image blocks obtained after the low-quality depth map and the 2D map are segmented are linearly embedded into vectors.
[0060] Step S1044, input the image blocks after the 2D image is embedded as vectors into the third network, and input the image blocks after the low-quality depth image is embedded as vectors into the first network. The third network performs deep processing on the image block features of the image blocks after the 2D image is embedded as vectors, and outputs the feature vectors of each image block of the 2D image; the first network performs deep processing on the image block features of the image blocks after the low-quality depth image is embedded as vectors, and outputs the low-resolution feature vectors of each image block.
[0061] Step S1045, the feature vector output by the third network is resampled using a cross-attention mechanism or pixel by pixel.
[0062] In step S1046, the first network recombines the low-resolution feature vector into a high-resolution feature vector with the same resolution as the feature vector of the image block of the 2D image, and outputs the high-resolution feature vector, which is then used by the second network to predict the absolute depth map of the target scene based on the high-resolution feature vector and the feature vector of the image block of the 2D image.
[0063] It should be noted that the absolute depth map of the target scene is obtained through the trained absolute depth estimation model, and the absolute depth map is a high-resolution and denoised image.
[0064] Step S106: Estimating the absolute depth of the target scene according to the absolute depth map.
[0065] Based on the above steps S102 to S106, in this embodiment, the low-quality depth map and 2D map of the target scene are used as the input of the trained absolute depth estimation model, and then the image features are extracted based on the third network and the first network processes the image features extracted by the third network into image features with the same resolution. Finally, the image features with the same resolution are processed based on the second network to obtain an absolute depth map, and depth estimation is performed. This method can realize image depth estimation of low-quality depth maps and 2D maps, solves the problem of low accuracy of image absolute depth estimation in the prior art, and improves the accuracy of image absolute depth estimation.
[0066] It should be noted that the absolute depth map refers to a depth image that does not have certain defects or deficiencies of the low-quality depth image described in the above embodiments, that is, a high-resolution, denoised depth image.
[0067] In some embodiments, the first network includes a shallow convolutional network, the second network includes a DPT (DensePrediction Transformer) network, and the third network includes a ViT (Vision Transformer) network; Figure 4 As shown, the low-quality depth map and the 2D map are input into a fully trained absolute depth estimation model to obtain the absolute depth map of the target scene, including: extracting the first image feature of the 2D map through the ViT network; extracting the second image feature of the low-quality depth map through a shallow convolutional network; processing the resolution of the second image feature to be the same as that of the first image feature through the shallow convolutional network; inputting the second image feature and the first image feature processed by the shallow convolutional network into the DPT network to obtain the absolute depth map.
[0068] In this embodiment, the ViT network is used to extract features from the input 2D image. ViT can capture the global information in the image and provide rich semantic features for subsequent processing. For a given low-quality depth map, a shallow convolutional network (CNN) is used to extract its features. The shallow network can quickly capture basic spatial structure information. Since the image features extracted by the ViT network and the image features extracted by the shallow convolutional network may have different resolutions, the shallow convolutional network described above needs to be used again to convert the image features of the low-quality depth map into a resolution that matches the image features of the 2D image. This step ensures that features from two different sources can be effectively fused in the next step. The processed second image features are input into the DPT network together with the first image features. The DPT network is a network architecture designed specifically for depth prediction. It can combine multi-source features for more accurate depth estimation, thereby outputting a high-quality absolute depth map.
[0069] It should be noted that the above-mentioned shallow convolutional network may also include two sub-networks, one sub-network extracts the features of the low-quality depth map, another sub-network converts the resolution of the image features, and another sub-network can also denoise the image features at the same time.
[0070] In some embodiments, the shallow convolutional network is also used to perform denoising on the second image feature, which can reduce the noise in the image feature.
[0071] In some of the embodiments, Figure 2 As shown, the training method of the fully trained absolute depth estimation model includes:
[0072] Step S202: Acquire image data of the target scene, wherein the image data includes: a 2D image, a low-quality depth image, and 3D data.
[0073] In this step, 2D images can be acquired through digital cameras / smartphone cameras, low-quality depth images can be collected through structured light sensors, time-of-flight (ToF) cameras, or binocular vision systems, and 3D data can be constructed through laser scanners, RGB-D cameras, or stereo matching algorithms.
[0074] Step S204: perform three-dimensional reconstruction on the 2D image to obtain a pseudo-true value depth image.
[0075] In this step, the pseudo-true depth image refers to an image that is generated in some way in the fields of 3D reconstruction, computer vision, etc., which is close to the real depth information. This image is not the real depth data directly measured by a depth sensor (such as a lidar or a structured light camera), but the result estimated by an algorithm. In this embodiment, the estimation is mainly performed by a 3D reconstruction algorithm, and the main methods of 3D reconstruction include but are not limited to the following methods:
[0076] Method 1: Method based on multi-view geometry: This method relies on photos of the same scene taken from multiple perspectives. By calculating the correspondence between these photos (such as feature point matching), the three-dimensional position of objects in the scene can be estimated using the triangulation principle. SfM (Structure from Motion) and MVS (Multi-View Stereo) are specific implementations of this type of method.
[0077] Method 2: Deep learning-based methods: In recent years, with the development of deep learning technology, it has become possible to use neural networks to directly predict depth information from a single or multiple images. This method does not require explicit feature matching, but relies on a large amount of training data to allow the model to learn how to infer depth information from image content.
[0078] It should be noted that 2D images can be images from iPhone videos and DSLR fisheye videos among other devices.
[0079] Step S206: synthesize the pseudo true value depth image with the 3D data to obtain a true value depth image.
[0080] In this step, a preset algorithm can be used to fuse the pseudo-true value depth image and 3D data. This process can select different preset algorithms according to different applications, such as weighted average algorithm, nearest neighbor interpolation algorithm, Kalman filter algorithm and other algorithms that can implement the solution of this application, with the purpose of using the high precision of 3D data to correct the error in the pseudo-true value depth image, thereby improving the accuracy of the overall depth estimation.
[0081] It should be noted that the synthesized true depth image may require further post-processing, such as removing noise, filling missing values, smoothing boundaries, etc., to generate the final true depth image. Finally, the synthesized true depth image can also be verified to check its accuracy and reliability. The verification method can be performed by comparing with known true depth values or by visual inspection.
[0082] Step S208 , constructing an image training pair by combining the 2D image and the low-quality depth image under the same target scene with the true depth image.
[0083] Step S210: input the image training pair into the absolute depth estimation model to be trained, and train until convergence to obtain a fully trained absolute depth estimation model.
[0084] It should be noted that, in this embodiment, the true depth image is the training label of the absolute depth estimation model to be trained, which provides a learning target for the model so that the model can adjust its internal parameters to minimize the difference between the predicted depth map and the true depth image.
[0085] Based on the above steps 202 to 210, by acquiring image data of the target scene, wherein the image data includes: a 2D image, a low-quality depth image and 3D data; performing three-dimensional reconstruction on the 2D image to obtain a pseudo-true value depth image; synthesizing the pseudo-true value depth image with the 3D data to obtain a true value depth image; constructing an image training pair using the 2D image and the low-quality depth image and the true value depth image under the same target scene; inputting the image training pair into the absolute depth estimation model to be trained, and training until convergence to obtain a fully trained absolute depth estimation model, the training of the absolute depth estimation model is realized, and the low-quality depth map is incorporated into the absolute depth estimation model as a hint, so as to realize absolute depth estimation of the image, solve the problem of low accuracy of absolute depth estimation of images in the prior art, and improve the accuracy of absolute depth estimation of images.
[0086] In one of the embodiments, acquiring image data of a target scene includes: acquiring a 2D image and a low-quality depth image of the target scene based on a preset shooting device; performing a three-dimensional scan on the target scene to obtain 3D data of the target scene, wherein the 3D data includes one of the following: mesh data, point cloud data, and depth image data.
[0087] In this embodiment, the preset shooting device may include but is not limited to: a digital camera, a smart phone. The three-dimensional scanning method may be implemented by the following devices: a laser scanner, an RGB-D camera, and a stereo matching algorithm.
[0088] In one of the embodiments, synthesizing a pseudo true-value depth image with 3D data to obtain a true-value depth image includes: obtaining an image gradient of the pseudo true-value depth image and an image depth of the 3D data; and inputting the image gradient and the image depth into a preset neural network model to obtain a true-value depth image.
[0089] In this embodiment, the advantages of the image gradient of the pseudo true value depth image and the image depth of the 3D data are combined, and a preset neural network model synthesizes the image gradient of the pseudo true value depth image and the image depth of the 3D data to obtain a simulated true value depth image.
[0090] In an exemplary embodiment, Figure 3As shown, another absolute depth estimation model training method is provided, including the following steps S302 to S310.
[0091] Step S302: Acquire three-dimensional data of the target scene.
[0092] In this step, the method of acquiring three-dimensional data may be, but is not limited to, 3D scanning, stereo vision, and photogrammetry.
[0093] Step S304: performing a first preset processing on the three-dimensional data to obtain a 2D image and a true depth image.
[0094] In this step, the first preset processing includes but is not limited to: projection processing and rendering processing.
[0095] Step S306: performing a second preset process on the true depth image to obtain a low-quality depth image.
[0096] In this step, the second preset processing includes but is not limited to: blurring processing, noise adding processing, downsampling processing and sparse anchor frame interpolation processing.
[0097] Step S308: constructing an image training pair using the 2D image and the low-quality depth image under the same target scene and the true depth image.
[0098] Step S310: input the image training pair into the absolute depth estimation model to be trained, and train until convergence to obtain a fully trained absolute depth estimation model.
[0099] It should be noted that, in this embodiment, the true depth image is the training label of the absolute depth estimation model to be trained, which provides a learning target for the model so that the model can adjust its internal parameters to minimize the difference between the predicted depth map and the true depth image.
[0100] Based on the above steps 302 to 310, by acquiring image data of the target scene, wherein the image data includes: acquiring three-dimensional data of the target scene; performing a first preset processing on the three-dimensional data to obtain a 2D image and a true depth image; performing a second preset processing on the true depth image to obtain a low-quality depth image; constructing an image training pair using the 2D image and the low-quality depth image and the true depth image under the same target scene; inputting the image training pair into the absolute depth estimation model to be trained, and training until convergence to obtain a fully trained absolute depth estimation model, the training of the absolute depth estimation model is realized, and the low-quality depth map is incorporated into the absolute depth estimation model as a hint, so as to realize absolute depth estimation of the image, solve the problem of low accuracy of absolute depth estimation of images in the prior art, and improve the accuracy of absolute depth estimation of images.
[0101] In one embodiment, the 2D image includes: an RGB image and / or a grayscale image.
[0102] It should be noted that RGB images, RGB (Red, Green, Blue) images are color images composed of three color channels: red, green, and blue. Each pixel is represented by a triplet, such as (R, G, B), where R, G, and B represent the intensity values of the three colors red, green, and blue, respectively. Grayscale images are single-channel images with only brightness information but no color information. Each pixel has only one value, which is used to represent the brightness level of the position, usually ranging from 0 (black) to 255 (white), or sometimes other ranges, depending on the number of bits used.
[0103] In this embodiment, the training of the absolute depth estimation model based on the RGB image and the grayscale image can be implemented, so that the trained absolute depth estimation model can perform image depth estimation on the RGB image and the grayscale image.
[0104] Based on the same inventive concept, the embodiment of the present application also provides an image imaging device for implementing the above-mentioned image depth estimation method. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above-mentioned method, so the specific limitations in one or more image imaging device embodiments provided below can refer to the limitations of the image depth estimation method above, and will not be repeated here.
[0105] In an exemplary embodiment, Figure 5 As shown, a depth image imaging device is provided, the device comprising:
[0106] An image acquisition unit 51 is used to acquire a low-quality depth map and a 2D map of a target scene;
[0107] A deep learning reasoning unit 52, configured to input the low-quality depth map and the 2D map into a fully trained absolute depth estimation model to obtain an absolute depth map of the target scene, wherein the fully trained absolute depth estimation model includes a first network, a second network, and a third network, the third network is configured to extract image features of the 2D image, the first network is configured to extract image features of the low-quality depth map and process the image features of the low-quality depth map to have the same resolution as the image features of the 2D image extracted by the third network, and the second network is configured to output an absolute depth map according to the image features of the 2D image and the image features of the processed low-quality depth map;
[0108] The processing unit 53 is used to estimate the absolute depth of the target scene according to the absolute depth map.
[0109] In one of the embodiments, the first network includes a shallow convolutional network, which is also used to denoise image features of the low-quality depth map.
[0110] In one of the embodiments, the depth image device also includes a camera unit, which is used to obtain a low-quality depth map and a 2D image of the target scene and send them to the image acquisition unit, wherein the camera unit and the image acquisition unit are wirelessly connected, and / or the camera unit and the image acquisition unit are wired connected.
[0111] In some cases, a hybrid mode may be used, that is, some of the camera units are connected to the image acquisition unit 51 by wires, while others are connected wirelessly. This solution can flexibly balance cost, performance, and convenience according to actual needs.
[0112] Each module in the above-mentioned image forming apparatus can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute operations corresponding to each module.
[0113] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0114] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. An absolute depth estimation method, characterized in that: The method comprises: Obtain low-quality depth maps and 2D images of the target scene; Inputting the low-quality depth map and the 2D map into a fully trained absolute depth estimation model to obtain an absolute depth map of the target scene, wherein the fully trained absolute depth estimation model includes a first network, a second network, and a third network, the third network is used to extract image features of the 2D image, the first network is used to extract image features of the low-quality depth map and process the image features of the low-quality depth map to the same resolution as the image features of the 2D image extracted by the third network, and the second network is used to output an absolute depth map according to the image features of the 2D image and the image features of the processed low-quality depth map; An absolute depth estimation is performed on the target scene according to the absolute depth map.
2. The absolute depth estimation method according to claim 1, characterized in that: The first network includes a shallow convolutional network, the second network includes a DPT network, and the third network includes a ViT network; Inputting the low-quality depth map and the 2D map into a fully trained absolute depth estimation model to obtain the absolute depth map of the target scene includes: Extracting a first image feature of the 2D image through the ViT network; Extracting a second image feature of the low-quality depth map through the shallow convolutional network; Processing the resolution of the second image feature to be the same as that of the first image feature through the shallow convolutional network; The second image features and the first image features processed by the shallow convolutional network are input into the DPT network to obtain an absolute depth map.
3. The absolute depth estimation method according to claim 2, characterized in that: The shallow convolutional network is also used to perform denoising on the second image feature.
4. The absolute depth estimation method according to claim 1, characterized in that: The training method of the fully trained absolute depth estimation model includes: Acquire image data of a target scene, wherein the image data includes: a 2D image, a low-quality depth image, and 3D data; Performing three-dimensional reconstruction on the 2D image to obtain a pseudo-true value depth image; Synthesizing the pseudo true value depth image with the 3D data to obtain a true value depth image; Constructing an image training pair by combining the 2D image and the low-quality depth image under the same target scene with the true depth image; The image training pair is input into the absolute depth estimation model to be trained, and the training is performed until convergence to obtain a fully trained absolute depth estimation model.
5. The absolute depth estimation method according to claim 4, characterized in that: Acquiring image data of the target scene includes: Acquire a 2D image and a low-quality depth image of the target scene based on a preset shooting device; The target scene is three-dimensionally scanned to obtain 3D data of the target scene, wherein the 3D data includes one of the following: grid data, point cloud data, and depth image data.
6. The absolute depth estimation method according to claim 4, characterized in that: The pseudo true value depth image is synthesized with the 3D data to obtain a true value depth image, comprising: Acquire an image gradient of the pseudo true value depth image and an image depth of the 3D data; The image gradient and the image depth are input into a preset neural network model to obtain the true value depth image.
7. The absolute depth estimation method according to claim 1, characterized in that: The training method of the fully trained absolute depth estimation model includes: Acquire three-dimensional data of the target scene; Performing a first preset processing on the three-dimensional data to obtain a 2D image and a true depth image; Performing a second preset processing on the true depth image to obtain a low-quality depth image; Constructing an image training pair by combining the 2D image and the low-quality depth image under the same target scene with the true depth image; The image training pair is input into the absolute depth estimation model to be trained, and the training is performed until convergence to obtain a fully trained absolute depth estimation model.
8. The absolute depth estimation method according to claim 7, characterized in that: The first preset processing includes: projection processing and rendering processing.
9. The absolute depth estimation method according to claim 7, characterized in that: The second preset processing includes: blurring processing, noise adding processing, downsampling processing and sparse anchor frame interpolation processing.
10. The absolute depth estimation method according to claim 4 or 7, characterized in that: The 2D image includes: an RGB image and / or a grayscale image.
11. A depth image imaging device, characterized in that: The device comprises: An image acquisition unit, used to acquire a low-quality depth map and a 2D map of a target scene; A deep learning reasoning unit, configured to input the low-quality depth map and the 2D map into a fully trained absolute depth estimation model to obtain an absolute depth map of the target scene, wherein the fully trained absolute depth estimation model comprises a first network, a second network, and a third network, wherein the third network is configured to extract image features of the 2D image, the first network is configured to extract image features of the low-quality depth map and process the image features of the low-quality depth map to have the same resolution as the image features of the 2D image extracted by the third network, and the second network is configured to output an absolute depth map according to the image features of the 2D image and the image features of the processed low-quality depth map; A processing unit is configured to perform absolute depth estimation on the target scene according to the absolute depth map.
12. The depth image imaging device according to claim 11, characterized in that: The first network includes a shallow convolutional network, and the shallow convolutional network is also used to denoise the image features of the low-quality depth map.
13. The depth image imaging device according to claim 11, characterized in that: The depth image device also includes a camera unit, which is used to obtain a low-quality depth map and a 2D image of the target scene and send them to the image acquisition unit, wherein the camera unit is wirelessly connected to the image acquisition unit, and / or the camera unit is wiredly connected to the image acquisition unit.
Citation Information
Patent Citations
Depth estimation method and device for automatic driving scene and autonomous vehicle
CN111680554A
Real-time depth completion method based on pseudo depth map guidance
CN112861729A
Indoor scene monocular image depth estimation method based on deep learning
CN114638870A
Scene depth estimation model training method, scene depth estimation method and device
CN118247323A
Depth estimation method based on multi-view self-supervised learning
CN118552596A