Depth Map Enhancement Method and Device

By extracting features from the original visual data and performing multi-scale processing, the dependence of the depth map enhancement method on geometric symmetry in the prior art is solved, and the shape missing area completion and overall density of the depth map is achieved, and the quality and accuracy of data acquisition are improved.

CN114820344BActive Publication Date: 2025-06-10TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210295510.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-06-10
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

In the prior art, the enhancement method of depth map requires the target object to meet specific geometric symmetry, which makes it difficult to use a general enhancement model to complete and dense geometric features during the depth acquisition process, resulting in the problems of information defects and low accuracy.

Method used

By obtaining the initial depth map from the original visual data, multi-scale feature extraction of the alternating convolution and deconvolution modules, a feature view with the depth map three-dimensional coordinates as the three-channel, and scale compression and convolution are performed to obtain the feature vectors of the two stages. Then, the initial depth map is spatially transformed and feature extraction, and the depth map features are strengthened based on the two-stage feature vectors to restore the depth structure, and finally the final depth map is obtained through multi-layer perceptron mapping.

Benefits of technology

While maintaining the original depth geometry, it can complete the missing shapes in the original data and dense the overall depth map to adapt to shape defects and sparseness of different degrees and angles, and improve the output accuracy and data acquisition quality of low-precision acquisition equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820344B_ABST
    Figure CN114820344B_ABST
Patent Text Reader

Abstract

The present application discloses a depth map enhancement method and apparatus. The method includes: obtaining an initial depth map from original visual data, performing multi-scale feature extraction on the initial depth map using an alternating convolution and deconvolution module to obtain a feature view, performing scale compression and convolution to sequentially obtain two-stage feature vectors; performing spatial transformation and feature extraction on the initial depth map to obtain depth map features, enhancing the depth map features based on the two-stage feature vectors, restoring the depth structure to generate a feature map, and then obtaining the final depth map through mapping by a multi-layer perceptron. Thus, the technical problem in the related art is solved, where due to the need for the target object to satisfy a certain specific geometric symmetry for feature extraction and prediction, it is difficult to use a general enhancement model to complete and densify geometric features during depth acquisition, resulting in incomplete information and low accuracy during acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to a method and device for depth map enhancement. Background Art

[0002] With the progress of technology, the acquisition and obtaining of three-dimensional object models have become increasingly easy. In particular, depth maps have attracted more and more research attention due to their applications in various fields such as autonomous driving and robotics.

[0003] Due to many factors such as the low resolution of devices and occlusion, the data quality of the depth directly obtained by depth scanning devices such as lidar and depth cameras is usually very poor. The low-quality data is mainly manifested in aspects such as the incomplete depth space structure and information redundancy. Therefore, the enhancement methods and systems for depth maps have become increasingly important in practical engineering applications.

[0004] In related technologies, the geometric feature-based completion and densification methods adopted require the target object to satisfy a certain specific geometric symmetry, and then use the obtained images for feature extraction and prediction. However, the geometric differences between real-world objects are relatively large, and it is difficult to use a general enhancement model for geometric feature completion and densification during depth acquisition, resulting in information incompleteness and low accuracy during acquisition, which urgently needs to be improved. Summary of the Invention

[0005] This application provides a method and device for depth map enhancement to solve the technical problem that in related technologies, due to the need for the target object to satisfy a certain specific geometric symmetry and then perform feature extraction and prediction, it is difficult to use a general enhancement model for geometric feature completion and densification during depth acquisition, resulting in information incompleteness and low accuracy during acquisition.

[0006] The first aspect embodiment of this application provides a method for depth map enhancement, including the following steps: obtaining an initial depth map from the original visual data; performing multi-scale feature extraction on the initial depth map by an alternating convolution and deconvolution module, and gradually restoring it by the deconvolution module to obtain a feature view with depth Figure 3 dimensional coordinates as three channels; performing scale compression and convolution on the feature view to obtain two-stage feature vectors in sequence; performing spatial transformation and feature extraction on the initial depth map to obtain depth map features, and strengthening the depth map features based on the two-stage feature vectors and restoring the depth structure to generate a first three-channel feature map; and obtaining a second three-channel feature map by mapping the feature view through a multi-layer perceptron, and fusing the first three-channel feature map and the second three-channel feature map, and mapping through a multi-layer perceptron to obtain the final depth map.

[0007] Optionally, in an embodiment of the present application, obtaining the initial depth map from the original visual data includes: obtaining a low-quality depth map that meets a preset condition from the original visual data; performing preprocessing and feature extraction on the low-quality depth map and the original infrared and depth map data to generate the initial depth map.

[0008] Optionally, in an embodiment of the present application, obtaining the low-quality depth map that meets a preset condition from the original visual data includes: collecting the original visual data based on a preset acquisition perspective threshold to obtain the low-quality depth map.

[0009] Optionally, in an embodiment of the present application, performing spatial transformation and feature extraction on the initial depth map to obtain depth map features, strengthening the depth map features based on the two-stage feature vectors, and restoring the depth structure to generate the first three-channel feature map includes: performing a three-dimensional spatial transformation on the initial depth map to obtain the depth after the three-dimensional transformation; extracting a first feature based on the depth after the three-dimensional transformation, and fusing the first feature with the view feature vector of the feature view to obtain a first-stage fusion feature; extracting a second feature based on the first-stage fusion feature, and fusing the second feature with the view feature vector to obtain a second-stage fusion feature; and obtaining the first three-channel feature map by mapping the second-stage fusion feature through a multi-layer perceptron.

[0010] An embodiment of the second aspect of the present application provides a depth map enhancement device, including: an acquisition module for obtaining an initial depth map from the original visual data; a feature extraction module for performing multi-scale feature extraction of an alternating convolution and deconvolution module on the initial depth map, and gradually restoring the depth by the deconvolution module to obtain a feature view with depth Figure 3 dimensional coordinates as three channels; a calculation module for performing scale compression and convolution on the feature view to sequentially obtain two-stage feature vectors; a strengthening module for performing spatial transformation and feature extraction on the initial depth map to obtain depth map features, strengthening the depth map features based on the two-stage feature vectors, and restoring the depth structure to generate the first three-channel feature map; and a fusion module for obtaining a second three-channel feature map by mapping the feature view through a multi-layer perceptron, and fusing the first three-channel feature map and the second three-channel feature map, and mapping through a multi-layer perceptron to obtain the final depth map.

[0011] Optionally, in an embodiment of the present application, the acquisition module includes: an acquisition unit for obtaining a low-quality depth map that meets a preset condition from the original visual data; a preprocessing unit for performing preprocessing and feature extraction on the low-quality depth map and the original infrared and depth map data to generate the initial depth map.

[0012] Optionally, in one embodiment of the present application, the acquisition unit is further used to collect the original visual data based on a preset collection viewing angle threshold to obtain the low-quality depth map.

[0013] Optionally, in one embodiment of the present application, the enhancement module includes: a three-dimensional transformation unit, used to perform a spatial three-dimensional transformation on the initial depth map to obtain a depth after the three-dimensional transformation; a first fusion unit, used to extract a first feature based on the depth after the three-dimensional transformation, and fuse the first feature with the view feature vector of the feature view to obtain a first-stage fused feature; a second fusion unit, used to extract a second feature based on the first-stage fused feature, and fuse the second feature with the view feature vector to obtain a second-stage fused feature; a mapping unit, used to map the second-stage fused feature multi-layer perceptron to obtain the first three-channel feature map.

[0014] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the depth map enhancement method as described in the above embodiment.

[0015] The fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the depth map enhancement method as described in any one of claims 1 to 4.

[0016] The embodiment of the present application can obtain an initial depth map from the original visual data, extract and enhance the two-stage feature vector and depth map features from the initial depth map, restore the depth structure, and obtain the final depth map through multi-layer perceptron mapping. While maintaining the original depth geometric structure, the shape missing areas in the original data can be completed and the overall depth map can be densified. It has the ability to adapt to different degrees and angles of shape incompleteness and sparsity, and can effectively improve the effective output accuracy of low-precision acquisition equipment and improve the quality of data acquisition. Therefore, it solves the technical problem in the related art that the target object needs to meet a certain specific geometric symmetry before feature extraction and prediction, which makes it difficult to use a general enhancement model to complete and densify geometric features during depth acquisition, resulting in incomplete information and low accuracy in acquisition.

[0017] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The above and / or additional aspects and advantages of the present application will become apparent and readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:

[0019] Figure 1 FIG. is a flowchart of a depth map enhancement method provided according to an embodiment of the present application;

[0020] Figure 2 FIG. is a schematic diagram of the principle of a depth map enhancement method according to an embodiment of the present application;

[0021] Figure 3 FIG. is a schematic diagram of the principle of an enhancement network of a depth map enhancement method according to an embodiment of the present application;

[0022] Figure 4 FIG. is a schematic diagram of the principle of a generator of a depth map enhancement method according to an embodiment of the present application;

[0023] Figure 5 FIG. is a flowchart of a depth map enhancement method according to a specific embodiment of the present application;

[0024] Figure 6 FIG. is a schematic structural diagram of a depth map enhancement device provided according to an embodiment of the present application;

[0025] Figure 7 FIG. is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. Detailed Embodiments

[0026] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.

[0027] The depth map enhancement method and apparatus according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the technical problem in the related art mentioned in the above background art that due to the need for a target object to satisfy a certain specific geometric symmetry, and then feature extraction and prediction are performed, it is difficult to use a general enhancement model to complete and densify geometric features during depth acquisition, resulting in incomplete information and low accuracy during acquisition, the present application provides a depth map enhancement method. In this method, an initial depth map can be obtained from the original visual data. From the initial depth map, two-stage feature vectors and depth map features are extracted and enhanced, the depth structure is restored, and the final depth map is obtained through a multi-layer perceptron mapping. While maintaining the original depth geometric structure, it can complete the missing shape areas in the original data and densify the overall depth map, and has the ability to adapt to shape defects and sparsity of different degrees and angles, which can effectively improve the effective output accuracy of low-precision acquisition devices and improve the quality of data acquisition. Thus, the technical problem in the related art that due to the need for a target object to satisfy a certain specific geometric symmetry, and then feature extraction and prediction are performed, it is difficult to use a general enhancement model to complete and densify geometric features during depth acquisition, resulting in incomplete information and low accuracy during acquisition is solved.

[0028] Specifically, Figure 1 FIG. is a schematic flowchart of a depth map enhancement method provided by an embodiment of the present application.

[0029] As Figure 1 shown, the depth map enhancement method includes the following steps:

[0030] In step S101, an initial depth map is obtained from the original visual data.

[0031] In actual implementation, the embodiments of the present application can use a depth camera and a camera for supporting acquisition to obtain the original visual data. After obtaining the view and a sparse and incomplete low-quality depth map from the original visual data, an initial depth map is obtained after data processing including but not limited to this.

[0032] Optionally, in an embodiment of the present application, obtaining an initial depth map from the original visual data includes: obtaining a low-quality depth map that meets a preset condition from the original visual data; performing preprocessing and feature extraction on the low-quality depth map and the infrared and depth map original data to generate an initial depth map.

[0033] As a possible implementation manner, the embodiments of the present application can obtain a low-quality depth map that meets preset conditions from the original visual data, and perform preprocessing and feature extraction on the view, infrared, and depth map original data obtained from the original data to generate an initial depth map for the processing of subsequent steps. The embodiments of the present application use infrared data as an auxiliary to enhance the data output of low-power and low-performance depth cameras and lidars, which will be described in detail below.

[0034] It should be noted that the preset conditions for the low-quality depth map can be set by those skilled in the art according to the actual situation, and no specific limitation is made here.

[0035] Optionally, in an embodiment of the present application, obtaining a low-quality depth map that meets preset conditions from the original visual data includes: collecting the original visual data based on a preset acquisition view angle threshold to obtain a low-quality depth map.

[0036] Specifically, the embodiments of the present application can use a depth camera and a camera for supporting collection, and collect the original visual data based on a preset acquisition view angle threshold to obtain a low-quality depth map. The embodiments of the present application collect the original visual data based on the preset acquisition view angle threshold, which is beneficial to the subsequent strengthening and densification of the low-quality depth map and lays a foundation for generating a high-quality image.

[0037] It should be noted that the preset view angle threshold will change accordingly according to different acquisition targets, and the view angle threshold can be set by those skilled in the art according to the actual situation, and no specific limitation is made here.

[0038] In step S102, multi-scale feature extraction of the initial depth map is performed by an alternating convolution and deconvolution module, and the deconvolution module is used for gradual restoration to obtain a feature view with depth Figure 3 dimensional coordinates as three channels.

[0039] Furthermore, the embodiments of the present application can perform multi-scale feature extraction of the initial depth map obtained through the above steps by an alternating convolution and deconvolution module, and the deconvolution module is used for gradual restoration to obtain a feature view with depth Figure 3 dimensional coordinates as three channels. The embodiments of the present application can extract multi-modal features by gradually extracting geometric features at different levels, which is convenient for subsequent feature fusion and restoration to generate a high-precision image.

[0040] In step S103, scale compression and convolution are performed on the feature view to obtain two-stage feature vectors in sequence.

[0041] In the actual implementation process, the embodiment of the present application can perform scale compression and convolution on the feature view and infrared features, so as to obtain a two-stage feature vector, realize the extraction of multimodal features, and facilitate the subsequent feature fusion and restoration to generate a high-precision image.

[0042] In step S104, the initial depth map is spatially transformed and feature extracted to obtain depth map features, and the depth map features are enhanced based on the feature vectors of the two stages, and the depth structure is restored to generate a first three-channel feature map.

[0043] Specifically, the embodiment of the present application can perform spatial transformation and feature extraction on the initial depth map obtained in step S101, and input the two-stage feature vector obtained in step S103, strengthen the depth map features and restore the depth structure, and then generate a first three-channel feature map. The embodiment of the present application can strengthen the depth map features and restore the depth structure, which is beneficial for completing the shape missing areas in the original data while maintaining the original depth geometry structure, and densifying the overall depth map. It has the ability to adapt to different degrees and angles of shape incompleteness and sparsity, and can effectively improve the effective output accuracy of low-precision acquisition equipment and improve the quality of data acquisition.

[0044] Optionally, in one embodiment of the present application, the initial depth map is spatially transformed and feature extracted to obtain depth map features, and the depth map features are enhanced based on the feature vectors of the two stages, and the depth structure is restored to generate a first three-channel feature map, including: performing a spatial three-dimensional transformation on the initial depth map to obtain a depth after the three-dimensional transformation; extracting a first feature based on the depth after the three-dimensional transformation, and fusing the first feature with a view feature vector of a feature view to obtain a first-stage fused feature; extracting a second feature based on the first-stage fused feature, and fusing the second feature with the view feature vector to obtain a second-stage fused feature; and mapping the second-stage fused feature multilayer perceptron to obtain the first three-channel feature map.

[0045] For example, in the actual implementation process, the steps of performing deep feature enhancement in the embodiment of the present application are as follows:

[0046] 1. Perform a spatial three-dimensional transformation on the initial depth map to obtain the depth after the three-dimensional transformation;

[0047] 2. Extract the first feature according to the depth after the three-dimensional transformation, and fuse the first feature with the view feature vector of the feature view to obtain the first stage fusion feature;

[0048] 3. Extract the second feature through the first stage fusion feature, and fuse the second feature with the view feature vector to obtain the second stage fusion feature;

[0049] 4. The first three-channel feature map is obtained by mapping the second-stage fusion features through a multi-layer perceptron.

[0050] For example, as Figure 2 shown, the method for enhancing depth features is as follows:

[0051] The point branch extracts features from points and generates enhanced point features using an attention mask generated from the image features in the image branch;

[0052] Then, the enhanced point features are forwarded to a fully connected layer to reconstruct a point set representing the global geometry, which contributes to another subset for finally predicting depth.

[0053] Specifically, in the embodiments of the present application, the original N input points in the 3D representation in the Euclidean space (N×3) can be converted to a fixed C dimension in the feature space (N×C), and the EdgeConv learning module proposed in DGCNN is used to extract point features.

[0054] In the spatial transformation layer, in the embodiments of the present application, an estimated 3×3 matrix can be used to align the input point set to the canonical space. To estimate the 3×3 matrix, in the embodiments of the present application, a tensor can be used to concatenate the coordinates of each point and the coordinate differences of its k adjacent points, where the image feature point feature ARC global feature CA×R is Figure 2 the feature enhancement module in the point branch.

[0055] It can be understood that the feature enhancement module fuses the global features from the view modality and the geometric local features F extracted from the point modality through the attention mechanism p . Specifically, the K-dimensional feature vector (as shown in the top row of Figure 2 ) from the first enhanced unit view branch is concatenated with the repeated points N times and then compressed into an Nx1-dimensional vector using an MLP. This vector is normalized to the range of [0,1] through the sigmoid function, and an attention mask m a (Nx1) can be obtained. Thereafter, by performing the element-wise multiplication of m a and the point feature F′, the enhanced local point feature F p is realized.

[0056] Furthermore, in the embodiments of the present application, the local features enhanced by the global image features can be obtained:

[0057]

[0058] The final enhanced point feature F e (N×2×C) can be realized by concatenating the local enhanced point feature F′ with N times the repeated global point features obtained by average pooling of F p .

[0059] In step S105, the second three-channel feature map is obtained by mapping with the feature view multi-layer perceptron, and the first three-channel feature map and the second three-channel feature map are fused, and the final depth map is obtained by mapping with the multi-layer perceptron.

[0060] As a possible implementation manner, the embodiments of the present application can perform multi-layer perceptron mapping on the compressed feature vectors obtained from the above steps to obtain the second three-channel feature map, and fuse the first three-channel feature map and the second three-channel feature map obtained in the above steps, and then perform multi-layer perceptron mapping to obtain the restored dense and complete depth map. The embodiments of the present application can complement the missing shape areas in the original data while maintaining the original depth geometric structure, and densify the overall depth map, and have the ability to adapt to shape defects and sparsity of different degrees and different angles, and can effectively improve the effective output accuracy of low-precision acquisition devices and improve the quality of data acquisition.

[0061] The following Figures 2 to 5 As shown, a detailed description of the embodiments of the present application will be given with an example.

[0062] The embodiments of the present application include the following steps:

[0063] Step S501: Acquisition of original visual data. In the actual execution process, the embodiments of the present application can use a depth camera and a camera, and perform supporting acquisition on the original visual data based on a preset acquisition view threshold to obtain the original visual data, and obtain views and sparse and incomplete low-quality depth maps from the original visual data.

[0064] It should be noted that the preset view threshold will change accordingly according to different acquisition targets, and the view threshold can be set by those skilled in the art according to the actual situation, and no specific limitation is made here.

[0065] Step S502: Preprocessing and feature extraction of the original data. As a possible implementation manner, the embodiments of the present application can obtain a low-quality depth map that meets the preset conditions from the original visual data, and perform preprocessing and feature extraction on the view, infrared, and depth map original data obtained from the original data to generate an initial depth map for subsequent processing steps.

[0066] It should be noted that the preset conditions for the low-quality depth map can be set by those skilled in the art according to the actual situation, and no specific limitation is made here.

[0067] Step S503: Obtain the depth Figure 3Furthermore, the embodiment of the present application can perform multi-scale feature extraction of the initial depth map obtained by the above steps by alternating convolution and deconvolution modules, and gradually restore it by the deconvolution module to obtain a depth map. Figure 3 The dimensional coordinates are used as the feature views of the three channels. The embodiments of the present application can realize the extraction of multimodal features by gradually extracting geometric features at different levels, which is convenient for subsequent feature fusion and restoration to generate high-precision images.

[0068] Step S504: Obtaining two-stage feature vectors. In the actual implementation process, the embodiment of the present application can perform scale compression and convolution on the feature view and infrared features to obtain two-stage feature vectors, realize multimodal feature extraction, and facilitate subsequent feature fusion and restoration to generate high-precision images.

[0069] Step S505: Enhance the depth map features and restore the depth structure. Specifically, the embodiment of the present application can perform spatial transformation and feature extraction on the initial depth map obtained in step S101, and input the two-stage feature vector obtained in step S103, enhance the depth map features and restore the depth structure, and then generate a first three-channel feature map. By enhancing the depth map features and restoring the depth structure, the embodiment of the present application is beneficial for completing the shape missing areas in the original data while maintaining the original depth geometry, and densifying the overall depth map. It has the ability to adapt to different degrees and angles of shape incompleteness and sparsity, and can effectively improve the effective output accuracy of low-precision acquisition equipment and improve the quality of data acquisition.

[0070] In the actual implementation process, the steps of performing deep feature enhancement in the embodiment of the present application are as follows:

[0071] 1. Perform a spatial three-dimensional transformation on the initial depth map to obtain the depth after the three-dimensional transformation;

[0072] 2. Extract the first feature according to the depth after the three-dimensional transformation, and fuse the first feature with the view feature vector of the feature view to obtain the first stage fusion feature;

[0073] 3. Extract the second feature through the first stage fusion feature, and fuse the second feature with the view feature vector to obtain the second stage fusion feature;

[0074] 4. The first three-channel feature map is obtained by mapping the second-stage fusion feature multi-layer perceptron.

[0075] For example, Figure 2 As shown, the method of deep feature enhancement is:

[0076] The point branch extracts features from points and generates enhanced point features using an attention mask generated from the image features in the image branch;

[0077] Then, the enhanced point features are forwarded to a fully connected layer to reconstruct a point set representing the global geometry, which contributes to another subset for the final depth prediction.

[0078] Specifically, the embodiments of the present application can convert the original N input points of the 3D representation in the Euclidean space (N×3) into a fixed C dimension in the feature space (N×C), and use the EdgeConv learning module proposed in DGCNN to extract point features.

[0079] In the spatial transformation layer, the embodiments of the present application can align the input point set to the canonical space using the estimated 3×3 matrix. To estimate the 3×3 matrix, the embodiments of the present application can use a tensor to concatenate the coordinates of each point and the coordinate differences of its k adjacent points, where the image feature point feature ARC global feature CA×R is Figure 2 The feature enhancement module in the point branch.

[0080] It can be understood that the feature enhancement module fuses the global features from the view modality and the geometric local features F extracted from the point modality through the attention mechanism p . Specifically, the K-dimensional feature vector from the first enhanced unit view branch (as shown in the top row of Figure 2 ) is concatenated with the N-fold repeated points and then compressed into an Nx1-dimensional vector using an MLP. By normalizing this vector to the range of [0,1] through the sigmoid function, an attention mask m a (Nx1) can be obtained. Thereafter, by performing the element-wise multiplication of m a with the point feature F ′ , the enhanced local point feature F p is achieved.

[0081] Furthermore, the embodiments of the present application can obtain the local features enhanced by the global image features:

[0082]

[0083] The final enhanced point feature F e (N×2×C) can be achieved by concatenating the local enhanced point feature F′ with N times the global point feature obtained by average pooling of F p .

[0084] Step S506: Fuse the three-channel feature maps, and through multi-layer perceptron mapping, obtain the restored dense and complete depth map. As a possible implementation manner, the embodiment of the present application can perform multi-layer perceptron mapping on the compressed feature vectors obtained from the above steps to obtain a second three-channel feature map, and fuse the first three-channel feature map and the second three-channel feature map obtained in the above steps, and then through multi-layer perceptron mapping, obtain the restored dense and complete depth map. The embodiment of the present application can, while maintaining the original depth geometry, complete the missing shape regions in the original data, and densify the overall depth map, and has the ability to adapt to shape defects and sparsity of different degrees and different angles, and can effectively improve the effective output accuracy of low-precision acquisition devices and improve the quality of data acquisition.

[0085] Further, the depth map enhancement method of the embodiment of the present application mainly includes the following aspects:

[0086] 1. Multi-modal feature extraction, fusion and restoration generation. The depth enhancement network is an adversarial architecture, including a generator and a discriminator. As Figure 3 shown, the enhancement network is an adversarial architecture, including a generator and a discriminator. As Figure 4 shown, the generator is a cascaded structure of two enhancement units with similar structures, taking the low-quality depth X and the single-view reference image I as inputs. Each enhancement unit consists of three parallel functional branches: the view branch, the point branch, and the fusion branch, which predict point sets by processing view, infrared features, depth features, and multi-modal image point fusion features respectively.

[0087] As Figure 4 shown, the depth generator has two cascaded enhancement units. The first unit directly takes the original image I and the depth map X as inputs, and transfers them into the latent embedding space, and then forwards the features to the next unit; the second unit is functionally similar to the first unit, but the output is a three-dimensional point set X. From the perspective of the data flow, the generator includes three parallel branches, which predict depth from view, depth, and multi-modal image point features respectively, and the output depth is the union of the predicted point sets of the three branches.

[0088] Among them, the view branch takes the image and the infrared image as inputs, and gradually extracts geometric features of different levels.

[0089] The depth branch, the point branch extracts features from the points, and uses the attention mask generated from the image features in the image branch to generate enhanced point features, so as to forward the enhanced point features to the fully connected layer to reconstruct the point set representing the global geometry, which contributes to another subset for the final depth prediction.

[0090] Fusion branch,The fusion branch mainly fuses two feature streams.

[0091] 2. Generate a deep discriminator. Adversarial training has promoted the development of a series of representation and generation applications in image representation, but there is little work on applying this architecture to depth completion, so the embodiment of the present application can add a discriminator to perform adversarial training in the framework, with the goal of identifying true and false depth. The embodiment of the present application can use a joint loss function that considers completeness and distribution uniformity to train the network to obtain a "coarse" complete depth, and then fine-tune the network with adversarial loss to achieve a "fine" complete depth, and use PointNet as a binary classification network for the discriminator to distinguish whether the prediction results are from the generated set X or the real set Y. The embodiment of the present application can be trained using adversarial loss in an improved Wasserstein GAN:

[0092]

[0093] Where D represents the set of 1-Lipschitz functions, y~PY and are the point set samples extracted from the generated data and the actual data, respectively, and the polynomial in the second term is the gradient penalty.

[0094] According to the depth map enhancement method proposed in the embodiment of the present application, an initial depth map can be obtained from the original visual data, and two-stage feature vectors and depth map features can be extracted and enhanced from the initial depth map to restore the depth structure, and the final depth map can be obtained through multi-layer perceptron mapping. While maintaining the original depth geometry, the shape missing areas in the original data can be completed, and the overall depth map can be densified, which has the ability to adapt to different degrees and angles of shape incompleteness and sparsity, and can effectively improve the effective output accuracy of low-precision acquisition equipment and improve the quality of data acquisition. Thus, the technical problem in the related art that when performing data acquisition, due to the large geometric differences of objects, it is difficult to complete features through geometric symmetry, so that the collected data has incomplete information and low precision, and it is difficult to achieve high-precision image output is solved.

[0095] Next, the depth map enhancement device proposed according to the embodiment of the present application is described with reference to the accompanying drawings.

[0096] Figure 6 It is a block diagram of a depth map enhancement device according to an embodiment of the present application.

[0097] like Figure 6 As shown, the depth map enhancement device 10 includes: an acquisition module 100, a feature extraction module 200, a calculation module 300, an enhancement module 400 and a fusion module 500.

[0098] Specifically, an acquisition module 100 is configured to acquire an initial depth map from the original visual data.

[0099] A feature extraction module 200 is configured to perform multi-scale feature extraction on the initial depth map through an alternating convolution and deconvolution module, and gradually recover the depth through the deconvolution module to obtain a feature view with depth Figure 3 dimension coordinates as a three-channel feature view.

[0100] A calculation module 300 is configured to perform scale compression and convolution on the feature view to sequentially obtain two-stage feature vectors.

[0101] A strengthening module 400 is configured to perform spatial transformation and feature extraction on the initial depth map to obtain depth map features, strengthen the depth map features based on the two-stage feature vectors, and recover the depth structure to generate a first three-channel feature map.

[0102] A fusion module 500 is configured to obtain a second three-channel feature map through multi-layer perceptron mapping of the feature view, fuse the first three-channel feature map and the second three-channel feature map, and perform multi-layer perceptron mapping to obtain the final depth map.

[0103] Optionally, in an embodiment of the present application, the acquisition module 100 includes: an acquisition unit and a preprocessing unit.

[0104] Wherein, the acquisition unit is configured to obtain a low-quality depth map that meets preset conditions from the original visual data.

[0105] The preprocessing unit is configured to perform preprocessing and feature extraction on the low-quality depth map and the original infrared and depth map data to generate an initial depth map.

[0106] Optionally, in an embodiment of the present application, the acquisition unit is further configured to collect the original visual data based on a preset acquisition perspective threshold to obtain a low-quality depth map.

[0107] Optionally, in an embodiment of the present application, the strengthening module 400 includes: a three-dimensional transformation unit, a first fusion unit, a second fusion unit, and a mapping unit.

[0108] Wherein, the three-dimensional transformation unit is configured to perform a spatial three-dimensional transformation on the initial depth map to obtain the depth after the three-dimensional transformation.

[0109] The first fusion unit is configured to extract a first feature based on the depth after the three-dimensional transformation, and fuse the first feature with the view feature vector of the feature view to obtain a first-stage fusion feature.

[0110] The second fusion unit is configured to extract a second feature based on the first-stage fusion feature, and fuse the second feature with the view feature vector to obtain a second-stage fusion feature.

[0111] The mapping unit is used to map the second-stage fusion feature multi-layer perceptron to obtain the first three-channel feature map.

[0112] It should be noted that the aforementioned explanation of the depth map enhancement method embodiment is also applicable to the depth map enhancement device of this embodiment, and will not be repeated here.

[0113] According to the depth map enhancement device proposed in the embodiment of the present application, an initial depth map can be obtained from the original visual data, and two-stage feature vectors and depth map features can be extracted and enhanced from the initial depth map to restore the depth structure, and the final depth map can be obtained through multi-layer perceptron mapping. While maintaining the original depth geometry, the shape missing areas in the original data can be completed, and the overall depth map can be densified. It has the ability to adapt to different degrees and angles of shape incompleteness and sparsity, and can effectively improve the effective output accuracy of low-precision acquisition equipment and improve the quality of data acquisition. Thus, the technical problem in the related art that when performing data acquisition, due to the large geometric differences of objects, it is difficult to complete features through geometric symmetry, so that the collected data has incomplete information and low precision, and it is difficult to achieve high-precision image output is solved.

[0114] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:

[0115] A memory 701 , a processor 702 , and a computer program stored in the memory 701 and executable on the processor 702 .

[0116] When the processor 702 executes the program, the depth map enhancement method provided in the above embodiment is implemented.

[0117] Furthermore, the electronic device further comprises:

[0118] The communication interface 703 is used for communication between the memory 701 and the processor 702 .

[0119] The memory 701 is used to store computer programs that can be executed on the processor 702 .

[0120] The memory 701 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0121] If the memory 701, the processor 702, and the communication interface 703 are implemented independently, the communication interface 703, the memory 701, and the processor 702 can be interconnected via a bus to complete communication with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 only a thick line is used to represent it in Figure 7 , but it does not mean that there is only one bus or one type of bus.

[0122] Optionally, in a specific implementation, if the memory 701, the processor 702, and the communication interface 703 are integrated on a single chip, the memory 701, the processor 702, and the communication interface 703 can complete communication with each other through an internal interface.

[0123] The processor 702 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0124] This embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-described depth map enhancement method is implemented.

[0125] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0126] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0127] Any process or method description represented in a flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or more N executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present application pertain.

[0128] The logic and / or steps represented in a flowchart or described otherwise herein, for example, can be considered as an ordered list of executable instructions for implementing a logical function and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more N wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0129] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0130] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0131] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0132] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, or the like. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A method for depth map enhancement, characterized in that, it includes the following steps: Obtain an initial depth map from the original visual data; Perform multi-scale feature extraction on the initial depth map using an alternating convolution and deconvolution module, and gradually restore it by the deconvolution module to obtain a feature view with the three-dimensional coordinates of the depth map as three channels; Perform scale compression and convolution on the feature view to obtain two-stage feature vectors in sequence; Perform spatial transformation and feature extraction on the initial depth map to obtain depth map features, strengthen the depth map features based on the two-stage feature vectors, and restore the depth structure to generate a first three-channel feature map; and Perform multi-layer perceptron mapping on the two-stage feature vectors to obtain a second three-channel feature map, fuse the first three-channel feature map and the second three-channel feature map, and perform multi-layer perceptron mapping after fusing the first three-channel feature map and the second three-channel feature map to obtain the final depth map; The obtaining the initial depth map from the original visual data includes: Obtain a low-quality depth map that meets the preset conditions from the original visual data; Perform preprocessing and feature extraction on the low-quality depth map and the original infrared depth map data to generate the initial depth map; The performing spatial transformation and feature extraction on the initial depth map to obtain depth map features, strengthening the depth map features based on the two-stage feature vectors, and restoring the depth structure to generate a first three-channel feature map includes: Perform three-dimensional spatial transformation on the initial depth map to obtain the depth after three-dimensional transformation; Extract the first feature based on the depth after three-dimensional transformation, and fuse the first feature with the view feature vector of the feature view to obtain a first-stage fusion feature; Extract the second feature based on the first-stage fusion feature, and fuse the second feature with the view feature vector to obtain a second-stage fusion feature; Perform multi-layer perceptron mapping on the second-stage fusion feature to obtain the first three-channel feature map.

2. The method according to claim 1, characterized in that, the obtaining the low-quality depth map that meets the preset conditions from the original visual data includes: Collect the original visual data based on a preset acquisition perspective threshold to obtain the low-quality depth map.

3. A depth map enhancement device, characterized in that, it includes: An acquisition module for obtaining an initial depth map from the original visual data; A feature extraction module for performing multi-scale feature extraction on the initial depth map using an alternating convolution and deconvolution module, and gradually restoring it by the deconvolution module to obtain a feature view with the three-dimensional coordinates of the depth map as three channels; A calculation module for performing scale compression and convolution on the feature view to obtain two-stage feature vectors in sequence; A strengthening module for performing spatial transformation and feature extraction on the initial depth map to obtain depth map features, strengthening the depth map features based on the two-stage feature vectors, and restoring the depth structure to generate a first three-channel feature map; and A fusion module, which is used to perform multi-layer perceptron mapping on the two-stage feature vectors to obtain a second three-channel feature map, fuse the first three-channel feature map and the second three-channel feature map, and perform multi-layer perceptron mapping after fusing the first three-channel feature map and the second three-channel feature map to obtain a final depth map; The acquisition module includes: an acquisition unit and a preprocessing unit; The acquisition unit is used to obtain a low-quality depth map that meets preset conditions from the original visual data; The preprocessing unit is used to preprocess and extract features from the low-quality depth map and the original infrared depth map data to generate the initial depth map; The enhancement module includes: A three-dimensional transformation unit, which is used to perform spatial three-dimensional transformation on the initial depth map to obtain the depth after three-dimensional transformation; A first fusion unit, which is used to extract a first feature based on the depth after three-dimensional transformation, and fuse the first feature with the view feature vector of the feature view to obtain a first-stage fusion feature; A second fusion unit, which is used to extract a second feature based on the first-stage fusion feature, and fuse the second feature with the view feature vector to obtain a second-stage fusion feature; A mapping unit, which is used to perform multi-layer perceptron mapping on the second-stage fusion feature to obtain the first three-channel feature map.

4. The apparatus according to claim 3, wherein, The acquisition unit is further configured to collect the original visual data based on a preset acquisition view angle threshold to obtain the low-quality depth map.

5. An electronic device, wherein, including: A memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the program to implement the depth map enhancement method according to any one of claims 1-2.

6. A computer-readable storage medium, on which a computer program is stored, wherein, The program is executed by a processor to be used to implement the depth map enhancement method according to any one of claims 1-2.

Citation Information

Patent Citations

  • Depth map enhancement method based on deep convolutional neural network

    CN111080688A

  • Image background estimation method based on depth map segmentation

    CN113436220A