Time-of-Flight Depth Enhancement

Through the end-to-end neural network combined with multimodal information for upsampling and depth injection, the problem of limited resolution of time-of-flight sensors under strong light is solved, and a high-quality time-of-flight depth map is generated, meeting the needs of 2D and 3D computer vision applications.

CN114174854BActive Publication Date: 2025-08-12HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980098894.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-08-07
Publication Date
2025-08-12
Estimated Expiration
2039-08-07

AI Technical Summary

Technical Problem

Existing time-of-flight sensors have multipath reflection problems under strong ambient light, and the spatial resolution of related images is limited, resulting in low-quality depth maps, which makes it difficult to meet the needs of 2D and 3D computer vision applications.

Method used

Through an end-to-end trainable neural network, combined with multimodal information of RAW-related signals, color images and depth images, guided upsampling and depth injection are performed to generate high-resolution improved time-of-flight depth maps.

Benefits of technology

Improved resolution, accuracy and accuracy of the time-of-flight depth map, solved the problems of multipath reflection and low resolution, recovered lost data and improved depth discontinuity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114174854B_ABST
    Figure CN114174854B_ABST
Patent Text Reader

Abstract

An image processing system is provided for receiving an input time-of-flight depth map, the input time-of-flight depth map representing the distance between an object in an image and a camera at multiple pixel positions in the corresponding image, and for generating an improved time-of-flight depth map of the image based on the map, the input time-of-flight depth map being generated from at least one related image, the related image representing the overlap between the transmitted light signal and the reflected light signal at the multiple pixel positions under a given phase shift, and the system is provided for generating the improved time-of-flight depth map from the input time-of-flight depth map based on the color representation of the corresponding image and at least one related image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to image formation in digital photography, and in particular to generating enhanced time-of-flight depth maps of images. Background Art

[0002] Time-of-flight (ToF) sensors are used to measure the distance between an object captured by a camera and the sensor plane. Figure 1 As shown in (a) of the figure, a scene is illuminated with a pulsed light source. The light is reflected by objects in the scene, and the round-trip time of the light is measured. Using this measured round-trip time and knowing the speed of light, the distance of the reflecting object from the camera can be estimated. This calculation can be performed for each pixel in the image. The distance stored at a specific pixel location is called the depth of that pixel, and the 2D image encoding these values for each pixel is called a depth map.

[0003] In ToF imaging, the depth map is obtained by RAW measurement by correlating the reflected light with the phase-shifted input pulse, e.g. Figure 1 The resulting image is called a correlation image. While this sensing approach works well in low-light scenarios and is computationally faster than other depth estimation methods such as stereo vision, current time-of-flight sensors suffer from drawbacks such as multipath reflections and have problems in strong ambient light. Furthermore, the spatial resolution of the correlation image is very limited. This ultimately hinders the use of ToF depth in 2D and 3D computer vision applications, where high resolution, accuracy, and precision are key aspects of high-quality data and a satisfying user experience.

[0004] Figure 2 (a) shows a photographic image, Figure 2 (b) shows the aligned ToF depth map corresponding to the image. The photographic image is formed by an RGB camera. In the ToF depth map, the distance is encoded in grayscale, with brighter areas being farther away.

[0005] In the dashed circle at 201, small objects have been over-smoothed or eliminated. For example, the distance to the thin black cable is not measured correctly. However, at 202, the gradient is well recovered even for this visually challenging part of the image. Therefore, the ToF sensor provides correct gradient measurements in this reflective area. Classical depth estimation methods have difficulty with these textureless areas. In the solid circles shown in 204 and 205, it can be seen that dark objects as well as distant objects are not captured correctly. In addition, the low depth image resolution (240×180 pixels in this example) will have additional information loss after image alignment, which can be seen in the following examples: Figure 2 This can be seen in the lower right corner 203 of (b), which further constrains the usable portion of the ToF depth map.

[0006] Several attempts have been made to overcome the drawbacks of low-quality ToF data by enriching the data with another input source or leveraging the power of machine learning through data preprocessing.

[0007] ToF methods using deep learning include "Deep end-to-end time-of-flight imaging" by Su, Shuochen et al., in Proceedings of the 2018 IEEE Conference on Computer Vision and Pattern Recognition. This method proposes an end-to-end learning pipeline that maps raw correlation signals to depth maps. The network is trained on a synthetic dataset. This approach can generalize to real data to some extent.

[0008] In another approach, US 9,760,837 B1 describes a method for depth estimation using time of flight. This method utilizes the RAW correlation signal and produces a depth output of the same resolution.

[0009] In Agresti, Gianluca et al., “Deep learning for confidence information in stereo and ToF data fusion,” Proceedings of the IEEE International Conference on Computer Vision, 2017, and Agresti, Gianluca and Pietro Zanuttigh, “Deep learning for multi-path error removal in ToF sensors,” Proceedings of the European Conference on Computer Vision, 2018, classic stereo is fused with time-of-flight sensing to improve the resolution and accuracy of the two synthetically created data modalities. The RGB data input pipeline is not learned, so RGB is only used indirectly by leveraging the predicted stereo depth map. The ToF data is reprojected separately and upsampled using a bilateral filter.

[0010] US 8,134,637 B2 proposes a method for super-resolving depth images from a ToF sensor using RGB images without learning. This method is not a single-step method, so the multiple individual modules of the method propagate errors, and each error accumulates through the pipeline.

[0011] It is desirable to develop a method for generating enhanced ToF depth maps of images. Summary of the Invention

[0012] According to a first aspect, an image processing system is provided, which is used to receive an input time-of-flight depth map, the input time-of-flight depth map representing the distance between an object in an image and a camera at multiple pixel positions in the corresponding image, and to generate an improved time-of-flight depth map of the image based on the map, the input time-of-flight depth map being generated from at least one related image, the related image representing the overlap between the emitted light signal and the reflected light signal at the multiple pixel positions under a given phase shift, and the system being used to generate the improved time-of-flight depth map from the input time-of-flight depth map based on the color representation of the corresponding image and at least one related image.

[0013] Therefore, the input ToF depth map can be enriched with features from the RAW correlation signal and processed using co-modal guidance from the aligned color image. Thus, the system exploits cross-modality advantages. ToF depth errors can also be corrected. Missing data can be recovered, and multipath ambiguity can be resolved using RGB guidance.

[0014] The resolution of the color representation of the corresponding image may be higher than the resolution of the input time-of-flight depth map and / or the at least one related image.This may increase the resolution of the improved time-of-flight depth map.

[0015] The system can be used to generate improved time-of-flight depth maps through a trained artificial intelligence model. The trained artificial intelligence model can be an end-to-end trainable neural network. Because the pipeline is trainable end-to-end, accessing all three different modalities (color, depth, and RAW correlation) simultaneously can improve the overall recovered depth map. This can improve the resolution, accuracy, and precision of the ToF depth map.

[0016] The model is trained using at least one of: an input time-of-flight depth map, a correlation image, and a color representation of the image.

[0017] The system is configured to combine the input time-of-flight depth map with the at least one correlation image to form a correlation-enriched time-of-flight depth map. Enriching the input ToF depth map with encoded features of the low-resolution RAW correlation signal can help reduce depth errors.

[0018] The system is configured to generate the improved time-of-flight depth map by hierarchically upsampling the correlation-rich time-of-flight depth map based on the color representation of the corresponding image. This can help improve and ameliorate depth discontinuities.

[0019] The improved time-of-flight depth map may have a higher resolution than the input time-of-flight depth map. This may result in improvements when rendering images captured by the camera.

[0020] The color representation of the corresponding image may be a color-separated representation. The color representation may be an RGB representation. This may be a convenient color representation for use when processing depth maps.

[0021] According to a second aspect, a method is provided for generating an improved time-of-flight depth map of an image based on an input time-of-flight depth map, wherein the input time-of-flight depth map represents the distance between an object in the image and a camera at multiple pixel positions in the corresponding image, the input time-of-flight depth map is generated from at least one related image, and the related image represents the overlap between the emitted light signal and the reflected light signal at the multiple pixel positions under a given phase shift, the method comprising generating the improved time-of-flight depth map from the input time-of-flight depth map based on the color representation of the corresponding image and at least one related image.

[0022] Therefore, the input ToF depth map can be enriched with features from the RAW correlation signal and processed using co-modal guidance from the aligned color image. Thus, the proposed method exploits cross-modality advantages. ToF depth errors can also be corrected. Missing data can be recovered, and multipath ambiguity can be resolved using RGB guidance.

[0023] The resolution of the color representation of the corresponding image may be higher than the resolution of the input time-of-flight depth map and / or the at least one related image.This may increase the resolution of the improved time-of-flight depth map.

[0024] The method may include generating an improved time-of-flight depth map through a trained artificial intelligence model. The trained artificial intelligence model may be an end-to-end trainable neural network. Because the pipeline is trainable end-to-end, accessing all three different modalities (color, depth, and RAW correlation) simultaneously can improve the overall recovered depth map. This can improve the resolution, accuracy, and precision of the ToF depth map.

[0025] The method may further comprise combining the input time of flight map with the at least one correlation image to form a correlation enriched time of flight depth map.Enriching the input ToF depth map with encoded features of the low resolution RAW correlation signal may help reduce depth errors.

[0026] The method may further comprise hierarchically upsampling the correlation-enriched time-of-flight depth map according to the color representation of the corresponding image. This may help improve and ameliorate depth discontinuities and may increase the resolution of the improved time-of-flight depth map. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The present invention will now be described by way of examples with reference to the accompanying drawings. In the accompanying drawings:

[0028] Figure 1 (a) and (b) show the acquisition of ToF depth data.

[0029] Figure 2 (a) in FIG. 1 shows a photographic image. Figure 2 (b) shows the Figure 2 ToF depth map corresponding to the image in (a).

[0030] Figure 3 An overview of a pipeline example for processing ToF depth maps is shown.

[0031] Figure 4 An exemplary overview of the pipeline for processing ToF depth maps is shown. The shallow encoder takes a RAW correlation image as input. In the decoding stage, noisy ToF depth data is injected and upsampled to four times the original resolution under RGB guidance.

[0032] Figure 5 Results of the presented pipeline using multimodal guided upsampling (GU) for ToF upsampling are shown.

[0033] Figure 6 (a) to (c) in FIG. 5 show exemplary results for different scenarios.

[0034] Figure 7 (a) to (j) in Figure 3 show the results obtained using the provided pipeline and the results obtained by classical upsampling using U-Net without multimodal guidance for comparison.

[0035] Figure 8 (a) to (c) in FIG. 3 show an ablation study comparing images processed using only guided upsampling, images processed using only depth injection, and images processed using the multimodal method of the present invention.

[0036] Figure 9 An example of a camera for using the pipeline of the present invention to process images captured by the camera is shown. DETAILED DESCRIPTION

[0037] Figure 3 An overview of an exemplary pipeline for generating enhanced ToF depth maps is shown. Figure 3The pipeline includes an end-to-end trainable neural network. The pipeline takes as input a ToF depth map 301 of relatively low resolution and quality or density (compared to the output ToF depth map 305). The input ToF depth map 301 represents the distance between an object in an image and the camera at multiple pixel locations in the corresponding image.

[0038] The input time of flight depth map 301 is generated from at least one RAW correlation image that represents the overlap between the transmitted light signal and the reflected light signal at the plurality of pixel locations at a given phase shift. As is well known in the art, the RAW correlation image data is processed using the speed of light to generate the input ToF depth map. This processing of the RAW correlation data to form the input ToF depth map can be performed separately from the pipeline or in an initialization step of the pipeline. The noisy ToF input depth 301 is fed into a learning framework (labeled ToF upsampling, (ToFupsampling, ToFU)), indicated at 302.

[0039] The pipeline also takes as input a color representation 303 of the corresponding image for which the input depth map 301 has been generated. In this example, the color representation is a color-separated representation, specifically an RGB image. However, the color representation may include one or more channels.

[0040] The pipeline also takes as input at least one RAW related image, as shown at 304. Thus, multimodal input data is used.

[0041] The system is configured to generate an improved time-of-flight depth map 305 from an input time-of-flight depth map 301 based on a color representation of a corresponding image 303 and at least one dependent image 304 .

[0042] Now refer to Figure 4 The systems and methods are described in more detail.

[0043] In this example, the end-to-end neural network includes an encoder-decoder convolutional neural network with guided upsampling and depth injection, including a shallow encoder 401 and a decoder 402. The shallow encoder 401 takes as input a RAW image 403. The network encodes the RAW information 403 at the original resolution 1 / 1 from the ToF sensor to extract depth features for depth prediction.

[0044] In the decoding stage, the input ToF depth data (which may be noisy and corrupted) shown at 404 is injected at the original resolution 1 / 1 (i.e., combined with the RAW related data), and then hierarchically upsampled to four times the original resolution under the RGB guidance. The input ToF depth information is injected into the decoder at the ToF input resolution stage, thereby enabling the network to predict depth information at a metric scale.

[0045] During guided upsampling (GU), RGB images at 2x and 4x the original resolution of the ToF depth map, shown at 405 and 406 respectively, are used to support residual correction of the directly upsampled depth map and enhance boundary accuracy at depth discontinuities.

[0046] Therefore, noisy ToF depth data 404 is injected and upsampled to four times the original resolution using RGB guidance to generate an enhanced ToF depth map, as shown in 407.

[0047] The co-injection of RGB and RAW correlation image modalities helps to super-resolve the input ToF depth map by filling holes (black areas in the input ToF depth map) with additional information, predict farther areas, and resolve ambiguities caused by multipath reflections.

[0048] Even though the depth injection from the input ToF depth map is noisy and corrupted, and distant pixel values are invalid, the above method can reliably recover the depth of the entire scene. Guided upsampling helps improve and ameliorate depth discontinuities. In this example, the resolution of the final output is four times that of the original input ToF depth map. However, the depth map can also be upsampled to a higher resolution.

[0049] In summary, the modalities used are as follows:

[0050] Input: RAW correlation image (low resolution), input ToF depth map (low resolution) and RGB image (high resolution).

[0051] Output: Upsampled depth map (high resolution).

[0052] These modalities complement each other, and ToFU extracts useful information from each modality in order to generate a final super-resolved output ToF depth map.

[0053] An exemplary network architecture is described below. Other configurations are possible.

[0054] Encoder layer: 1x 2D convolution on RAW correlated input (-> 1 / 2 input resolution)

[0055] Layer before injection: 1x 2D upconvolution (from 1 / 2 input resolution to 1 / 1 input resolution)

[0056] Decoder and Guided Upsampler

[0057] Deep injection:

[0058] For each input:

[0059] 2D Conv->BatchNorm->LeakyReLu->ResNetBlock->ResNetBlock

[0060] cascade

[0061] 4x ResNetBlock

[0062] Residual = 2D convolution

[0063] Output = depth + residual

[0064] Cascade + injection output of upconvolution (before injection)

[0065] Cascaded convolution + upsampling using bilinear upsampling (depth prediction at 1x input resolution)

[0066] Layers before GU 1: 1x 2D upconvolution of cascaded convolutions (from 1 / 1 input resolution to 2x input resolution)

[0067] Guided upsampling stage 1:

[0068] For each input:

[0069] 2D Conv->BatchNorm->LeakyReLu->ResNetBlock->ResNetBlock

[0070] cascade

[0071] 4x ResNetBlock

[0072] Residual = 2D convolution

[0073] Output = depth + residual

[0074] Cascade of upconvolution + guided upsampling output

[0075] Cascaded convolutions and upsampling using bilinear upsampling (depth prediction at 2x input resolution)

[0076] Layers before GU 2: 1x 2D upconvolution of cascaded convolutions (from 2x input resolution to 4x input resolution)

[0077] Guided upsampling stage 2:

[0078] For each input:

[0079] 2D Conv->BatchNorm->LeakyReLu->ResNetBlock->ResNetBlock

[0080] cascade

[0081] 4x ResNetBlock

[0082] Residual = 2D convolution

[0083] Output = depth + residual

[0084] Cascade of upconvolution + guided upsampling output

[0085] Cascaded convolution and depth prediction (depth prediction at 4x input resolution)

[0086] The following equations (1) to (4) describe exemplary loss functions. For training the proposed network, the pixel difference between the predicted inverse depth of the simulated disparity and the ground truth is minimized by exploiting the fast-converging robust norm and smoothing term:

[0087] L Total =ω s L Smooth +ω D L Depth (1)

[0088] in,

[0089]

[0090] and:

[0091] L Depth =∑ω Scale |D(p)-D Pred (p)| Barron (3)

[0092] Among them, |*| Barron It is the Barron loss proposed by Barron in "A General and Adaptive Robust Loss Function" (CVPR 2019), which is a special form of the smoothed L1 norm:

[0093]

[0094] ω Scale Indicates that at each scale level L Depth where D is the inverse depth and I is the RGB image. Since the disparity values at lower scale levels should be scaled accordingly (e.g., half the resolution results in half the disparity value), the value of the loss term should be inversely scaled by the same scale parameter. Furthermore, the number of pixels decreases quadratically with each scale level, resulting in an equal proportional weighting of the contribution of each scale level: ω Scale =Scale*Scale 2=Scale 3 .

[0095] In one implementation, to generate training data, as well as accurate depth ground truth, a physics-based rendering pipeline (PBRT) can be used with a blender, as proposed in Su, Shuochen et al., “Deep end-to-end time-of-flight imaging” (Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018). A low-resolution version of the depth is clipped and corrupted with noise to simulate a ToF depth input signal.

[0096] Figure 5 The results of the proposed pipeline using multimodal guided upsampling for ToF upsampling are shown in FIG. The input ToF depth map is as follows: Figure 5 The predicted depth map after upsampling to 2x resolution is shown in (a). Figure 5 As shown in (b) in the figure, the resulting error is Figure 5 The predicted depth map after upsampling to 4x resolution is shown in (c). Figure 5 As shown in (d) in the figure, the resulting error is Figure 5 The corresponding RGB image and ground truth ToF depth map are shown in (e). Figure 5 (f) and (g) in Figure 3 for comparison. The proposed method helps to recover depth without losing information related to fine structures while improving edges along depth discontinuities.

[0097] Figure 6 (a) to (c) in FIG. 5 show other exemplary results of different scenarios. Figure 6 (a) in the figure shows the input ToF depth map of the scene, depicted as a small RGB image. Figure 6 (b) in the figure shows the corresponding upsampled output, Figure 6 (c) in Figure 5 shows the corresponding ground truth depth map of the scene.

[0098] Figure 7 (a) to (j) in FIG. 5 show the comparison between the results obtained using U-Net classical upsampling without multimodal guidance and the results obtained using the method of the present invention. Figure 7 (a) in FIG. 5 shows the input ToF depth map of the scene. Figure 7 (b) and (c) show the ToF depth map and the corresponding residual obtained using classical upsampling after 32k iterations, respectively. Figure 7(d) and (e) show the ToF depth map and the corresponding residual obtained using the method described in this paper after 32k iterations. Figure 7 (f) in FIG. 5 shows the input ToF depth map of the scene area at a higher magnification. Figure 7 (g) and (h) in Fig. 3 show the ToF depth map and the corresponding residual obtained using classical upsampling after convergence. Figure 7 (i) and (j) in FIG. 5 show the ToF depth map and the corresponding residual obtained using the method described in this paper after convergence. Figure 7 The ToF depth maps generated using the method described in this paper in (d) and (i) recover depth without losing information related to fine structures while improving edges along depth discontinuities.

[0099] exist Figure 8 The results obtained using the method of the present invention are as follows Figure 8 As shown in (a) of Figure 8 (b) uses only GU (no deep injection) and Figure 8 The corresponding ground truth image and RGB image are respectively Figure 8 During classical upsampling, edges and fine structures are not refined for “injection only” ( Figure 8 (c) in Figure 3), which can be seen by comparing the residuals along edges and fine structures in the dashed and solid circles. Low-resolution depth injection helps the network start with good depth estimation at low resolution, thus helping GU to achieve higher resolution. Figure 8 In (a), the residual error of depth prediction compared to the ground truth is reduced compared to other methods. Specifically, GU improves the residual along depth discontinuities, where image gradients are usually strong, thereby also recovering fine structures.

[0100] Therefore, depth injection can guide the network to predict well-defined depth at a lower resolution, which is refined with RGB guidance during layer-wise guided upsampling to recover the depth to four times the original resolution.

[0101] Figure 9 An example of a camera 901 is shown, which uses a pipeline to process images captured by an image sensor 902 in the camera. The camera also includes a depth sensor 903 for collecting Time-of-Flight depth data. Such a camera 901 typically includes some onboard processing capability, which may be provided by a processor 904. The processor 904 may also be used to perform basic functions of the device.

[0102] The transceiver 905 is capable of communicating with other entities 910 and 911 via a network. These entities may be physically remote from the camera 901. The network may be a publicly accessible network, such as the Internet. The entities 910 and 911 may be cloud-based. The entity 910 is a computing entity. The entity 911 is a command entity and a control entity. These entities are logical entities. In practice, these entities may be provided by one or more physical devices (such as a server and a data storage device), and the functions of two or more entities may be provided by a single physical device. Each physical device that implements an entity includes a processor and a memory. These devices may also include a transceiver for sending data to the transceiver 905 of the camera 901 and receiving data from the transceiver 905 of the camera 901. The memory stores code in a non-transient manner, which can be executed by the processor to implement the corresponding entity in the manner described herein.

[0103] The command and control entity 911 can train the artificial intelligence models used in the pipeline. This is typically a computationally intensive task, even if the resulting model can be efficiently described, so it may be efficient to develop and execute the algorithm in the cloud, where significant energy and computing resources are expected to be available. This is expected to be more efficient than developing such models in a typical camera.

[0104] In one implementation, once the algorithm is developed in the cloud, the command and control entity can automatically form a corresponding model and transmit it to the relevant camera device. In this example, the pipeline is implemented in the camera 901 by the processor 904.

[0105] In another possible implementation, an image may be captured by a camera sensor 902 and the image data may be sent to the cloud by a transceiver 905 for processing in a pipeline. The resulting image may then be sent back to the camera 901, as shown in FIG. Figure 9 As shown in 912.

[0106] Therefore, the method can be deployed in a variety of ways, such as in the cloud, on-device, or in dedicated hardware. As mentioned above, cloud facilities can perform training to develop new algorithms or refine existing ones. Depending on the computing power near the data corpus, training can be performed near the source data or in the cloud, for example using an inference engine. The method can also be implemented in-camera, in dedicated hardware, or in the cloud.

[0107] Therefore, the present invention uses an end-to-end trainable deep learning pipeline to achieve ToF depth super-resolution. This end-to-end trainable deep learning pipeline enriches the input ToF depth map with encoded features of low-resolution RAW correlation signals. The synthesized feature map is hierarchically upsampled by common modality guidance from aligned high-resolution RGB images. By injecting the encoded RAW correlation signal, the ToF depth is enriched with the RAW correlation signal for domain stabilization and modality guidance.

[0108] This method takes advantage of cross-modality. For example, ToF works well in low-light or textureless areas, while RGB works well in bright scenes or scenes with dark textured objects.

[0109] Since the pipeline is trainable end-to-end, all three different modalities (RGB, depth, and RAW correlation) are accessed simultaneously, and they can mutually improve the overall recovered depth map. This can improve the resolution, accuracy, and precision of the ToF depth map.

[0110] ToF depth errors can also be corrected. Missing data can be recovered because the method measures farther areas, and multipath ambiguity can be resolved through RGB guidance.

[0111] The network can utilize supervised or unsupervised training. The network can utilize multimodal training with synthetic correlation, RGB, ToF depth, and ground truth rendering for direct supervision.

[0112] Based on the ground truth image, additional adjustments can be made to the output ToF depth map.

[0113] Applicants hereby disclose individually each individual feature described herein and any combination of two or more such features. It would be within the ordinary skill of those skilled in the art to implement such features or combinations as a whole based on this specification, without regard to whether such features or combinations of features solve any of the problems disclosed herein, and without limiting the scope of the claims. Applicants indicate that aspects of the present invention may consist of any such individual features or combinations of features. In view of the foregoing description, it would be apparent to those skilled in the art that various modifications may be made within the scope of the present invention.

Claims

1. An image processing system, characterized in that: Used to receive an input time-of-flight depth map, the input time-of-flight depth map represents the distance between an object in an image and a camera at multiple pixel positions in the corresponding image, and to generate an improved time-of-flight depth map of the image based on the input time-of-flight depth map, the input time-of-flight depth map is generated from at least one related image, the related image represents the overlap between the emitted light signal and the reflected light signal at the multiple pixel positions under a given phase shift, and the system is used to generate the improved time-of-flight depth map from the input time-of-flight depth map based on the color representation of the corresponding image and at least one related image.

2. The image processing system according to claim 1, wherein The color representation of the corresponding image has a higher resolution than the input time-of-flight depth map and / or the at least one correlation image.

3. The image processing system according to claim 1, wherein The system is configured to generate the improved time-of-flight depth map through a trained artificial intelligence model.

4. The image processing system according to claim 3, wherein: The trained artificial intelligence model is an end-to-end trainable neural network.

5. The image processing system according to claim 3 or 4, characterized in that The model is trained using at least one of: an input time-of-flight depth map, a correlation image, and a color representation of the image.

6. The image processing system according to claim 1, wherein: The system is configured to combine the input time-of-flight depth map with the at least one correlation image to form a correlation-enriched time-of-flight depth map.

7. The image processing system according to claim 6, wherein: The system is configured to generate the improved time-of-flight depth map by hierarchically upsampling the correlation-enriched time-of-flight depth map according to the color representation of the corresponding image.

8. The image processing system according to claim 1, wherein: The improved time-of-flight depth map has a higher resolution than the input time-of-flight depth map.

9. The image processing system according to claim 1, wherein: The color representation of the corresponding image is a color separated representation.

10. A method for generating an improved time-of-flight depth map for an image from an input time-of-flight depth map, characterized in that The input time-of-flight depth map represents the distance between an object in the image and the camera at multiple pixel positions in the corresponding image, the input time-of-flight depth map is generated from at least one related image, the related image represents the overlap between the emitted light signal and the reflected light signal at the multiple pixel positions under a given phase shift, and the method includes generating the improved time-of-flight depth map from the input time-of-flight depth map based on the color representation of the corresponding image and at least one related image.

11. The method according to claim 10, characterized in that The color representation of the corresponding image has a higher resolution than the input time-of-flight depth map and / or the at least one correlation image.

12. The method according to claim 10 or 11, characterized in that The method includes generating the improved time-of-flight depth map by a trained artificial intelligence model.

13. The method according to claim 12, characterized in that The trained artificial intelligence model is an end-to-end trainable neural network.

14. The method according to claim 10, characterized in that The method further includes combining the input time-of-flight depth map with the at least one correlation image to form a correlation-enriched time-of-flight depth map.

15. The method according to claim 14, characterized in that The method further includes hierarchically upsampling the correlation-enriched time-of-flight depth map according to the color representation of the corresponding image.

Citation Information

Patent Citations

  • Method and system to increase X-Y resolution in a depth (Z) camera using red, blue, green (RGB) sensing

    US8134637B2

  • Depth from time-of-flight using machine learning

    US9760837B1