De-aliasing indirect time-of-flight measurements

By comparing iToF depth information with image depth information, multi-period aliasing error is identified and corrected, solving the accuracy problem of iToF depth camera under multi-period aliasing and generating a more accurate depth map.

CN120936902APending Publication Date: 2025-11-11QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480018640.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-21
Filing Date
2024-03-08
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Indirect time-of-flight (iToF) depth cameras cannot accurately distinguish depth under multi-period aliasing conditions, resulting in decreased measurement accuracy.

Method used

By comparing iToF depth information with image-based depth information, inconsistencies are identified and the iToF depth information is adjusted. Correction is performed using an integer number of half wavelengths, and dealiasing is achieved by combining image depth information.

Benefits of technology

It improves the accuracy and consistency of iToF depth measurement, generates more accurate depth maps, and combines the absolute distance advantage of iToF depth measurement with the low cost of image depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120936902A_ABST
    Figure CN120936902A_ABST
Patent Text Reader

Abstract

Systems and techniques for determining depth information are described herein. For example, a method for determining depth information is provided. The method may include transmitting electromagnetic (EM) radiation toward a plurality of points in an environment; comparing the phase of the transmitted EM radiation with the phase of the received EM radiation to determine, for each of the plurality of points in the environment, a respective time-of-flight estimate of the EM radiation between transmission and reception; determining first depth information based on the respective time-of-flight estimates determined for each of the plurality of points in the environment; obtaining second depth information based on the image of the environment; comparing the first depth information with the second depth information to determine an inconsistency between the first depth information and the second depth information; and adjusting the depth of the first depth information based on the inconsistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates in general to processing time-of-flight measurements. For example, aspects of this disclosure include systems and techniques for dealiasing indirect time-of-flight measurements (e.g., to remove the effects of multi-period aliasing from indirect time-of-flight measurements). Background Technology

[0002] An indirect time-of-flight (iToF) depth camera measures the phase difference between an emitted light pulse and a light pulse received by the iToF depth camera after the light pulse has been reflected by an object in the environment. The iToF depth camera can correlate the phase difference with the time of flight of the light pulse between emission and reception based on the speed of light and the frequency of the light pulse. The iToF depth camera can calculate the distance between the iToF depth camera and an object in the environment based on the time of flight and the speed of light. In this disclosure, the terms "light," "light pulse," and similar terms can refer to electromagnetic radiation of any frequency, whether or not it is in the visible light spectrum. In this disclosure, the term "object," when referring to the environment, can include discrete objects, points on objects, and points in the environment including the ground and / or walls, etc. In this disclosure, the term "depth" can refer to the distance between the sensor (e.g., of an iToF depth camera) and the object.

[0003] iToF phase difference measurements can be cyclic. For example, phase measurements can be repeated for every integer number of half-wavelengths of a light pulse from an object to an iToF depth camera. For example, an iToF depth camera can emit a light pulse with a frequency of 20 MHz and a wavelength of approximately 15 meters. The iToF depth camera can measure the phase difference based on a first reflection from a first object three meters away from the iToF depth camera. The iToF depth camera can measure the same phase difference based on a second reflection from a second object 10.5 meters away from the iToF depth camera (3 meters plus half a wavelength). Therefore, iToF depth measurements may suffer from multi-cycle aliasing. In this disclosure, the term "multi-cycle aliasing" can refer to the iToF depth camera's inability to distinguish two or more depths based on phase difference measurements. Multi-cycle aliasing limits the accuracy of iToF depth cameras. Summary of the Invention

[0004] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Accordingly, the following summary presents certain concepts in a simplified form relating to one or more aspects of the mechanisms disclosed herein, preceding the detailed description presented below.

[0005] Systems and techniques for determining depth information are described. According to at least one example, a method for determining depth information is provided. The method includes: transmitting electromagnetic (EM) radiation toward a plurality of points in an environment; comparing the phase of the transmitted EM radiation with the phase of a received EM radiation to determine a corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment; determining first depth information based on the corresponding time-of-flight estimates determined for each of the plurality of points in the environment; obtaining second depth information based on an image of the environment; comparing the first depth information with the second depth information to determine an inconsistency between the first depth information and the second depth information; and adjusting the depth of the first depth information based on the inconsistency.

[0006] In another example, an apparatus for determining depth information is provided, the apparatus comprising: at least one memory; and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: cause at least one transmitter to transmit electromagnetic (EM) radiation toward a plurality of points in an environment; compare the phase of the transmitted EM radiation with the phase of received EM radiation to determine a corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment; determine first depth information based on the corresponding time-of-flight estimate determined for each of the plurality of points in the environment; obtain second depth information based on an image of the environment; compare the first depth information with the second depth information to determine an inconsistency between the first depth information and the second depth information; and adjust the depth of the first depth information based on the inconsistency. In some cases, the apparatus includes at least one transmitter configured to transmit EM radiation toward a plurality of points in the environment.

[0007] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: instruct at least one transmitter to transmit electromagnetic (EM) radiation toward a plurality of points in an environment; compare the phase of the transmitted EM radiation with the phase of a received EM radiation to determine a corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment; determine first depth information based on the corresponding time-of-flight estimate determined for each of the plurality of points in the environment; obtain second depth information based on an image of the environment; compare the first depth information with the second depth information to determine an inconsistency between the first depth information and the second depth information; and adjust the depth of the first depth information based on the inconsistency.

[0008] In another example, an apparatus for determining depth information is provided. The apparatus includes: components for transmitting electromagnetic (EM) radiation toward a plurality of points in an environment; components for comparing the phase of the transmitted EM radiation with the phase of a received EM radiation to determine a corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment; components for determining first depth information based on the corresponding time-of-flight estimates determined for each of the plurality of points in the environment; components for obtaining second depth information based on an image of the environment; components for comparing the first depth information with the second depth information to determine inconsistencies between the first depth information and the second depth information; and components for adjusting the depth of the first depth information based on the inconsistencies.

[0009] In some aspects, one or more of the devices described herein are, may be part of, or may include: mobile devices (e.g., mobile phones or so-called "smartphones," tablet computers, or other types of mobile devices); extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices); vehicles (or computing devices or systems of vehicles); smart or connected devices (e.g., Internet of Things (IoT) devices); wearable devices; personal computers; laptop computers; video servers; television sets (e.g., network-connected television sets); robotic devices or systems; or other devices. In some aspects, each device may include one image sensor (e.g., a camera) or multiple image sensors (e.g., multiple cameras) for capturing one or more images. In some aspects, each device may include one or more displays for displaying one or more images, notifications, and / or other displayable data. In some aspects, each device may include one or more speakers, one or more light-emitting devices, and / or one or more microphones. In some aspects, each device may include one or more sensors. In some cases, the one or more sensors may be used to determine the location of the device, the state of the device (e.g., tracking state, operating state, temperature, humidity level and / or another state) and / or for other purposes.

[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0011] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0012] This application contains at least one drawing in color. A copy of this patent application disclosure with color drawings will be provided by the Patent Office upon request and payment of the necessary fees.

[0013] The following description, with reference to the accompanying drawings, details exemplary examples of this application:

[0014] Figure 1 This is a block diagram illustrating a system for dealiasing in indirect time-of-flight (iToF) depth measurements according to various aspects of this disclosure;

[0015] Figure 2 This is a block diagram illustrating another system for dealiasing iToF depth measurements according to various aspects of this disclosure;

[0016] Figure 3 Examples of visual representations of image-based depth information according to various aspects of this disclosure include example images including depth partitions and example visual representations of iToF-based depth information partitioned by depth partition 310;

[0017] Figure 4 This includes three representations of three corresponding example deep partitions according to various aspects of this disclosure and a representation of the merged deep partitions;

[0018] Figure 5 This is a block diagram illustrating a system for dealiasing iToF depth measurements according to various aspects of this disclosure;

[0019] Figure 6 This is a flowchart illustrating a process for dealiasing iToF depth measurements according to various aspects of this disclosure;

[0020] Figure 7 Examples of deep learning neural networks that can be used to implement a perception module and / or one or more verification modules, based on some aspects of the disclosed technology, are illustrated.

[0021] Figure 8 These are illustrative examples of convolutional neural networks (CNNs) according to various aspects of this disclosure; and

[0022] Figure 9 Example computing device architectures are illustrated, showing example computing devices that can implement the various technologies described herein. Detailed Implementation

[0023] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0024] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0025] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0026] Indirect Time-of-Flight (iToF) depth cameras emit one or more light pulses into the environment and determine depth information relative to the environment (e.g., iToF-based depth information). For example, an iToF depth camera can emit one or more light pulses and receive reflected light pulses, focusing them onto an array of sensors. Using the array of sensors, an iToF depth camera can determine the depth of each of a plurality of points within its field of view. The number of depths can be depth information representing the depth of objects in the environment.

[0027] As described above, multi-period aliasing limits the accuracy of indirect time-of-flight (iToF) depth measurements because objects located far from the iToF depth camera can reflect light pulses with the same phase difference when compared to the emitted light pulse. Therefore, an iToF depth camera may not be able to distinguish depth based solely on phase difference. In other words, iToF depth cameras may suffer from multi-period aliasing.

[0028] One existing technique for mitigating multi-period aliasing in iToF depth measurement is to use the intensity of reflected signals as a factor when determining depth. This technique relies on the assumption that reflections from objects closer to the sensor will be stronger than those from objects farther away. However, many factors influence the intensity of reflected signals. For example, any or all of the following can affect the intensity of reflected signals: the material reflectivity of the object (e.g., metallic contrast with fabric), the object's color (e.g., dark contrast with light), and / or the object's orientation (e.g., the correlation between the direction from which light pulses are received at the object and the object's surface normal).

[0029] This document describes systems, apparatuses, methods (also referred to as processes), and computer-readable media (collectively, "systems and techniques") for dealiasing iToF depth measurements. In this disclosure, the term "dealiasing," as used in relation to depth measurement, can refer to determining a depth that may have multiple depths based on multi-period aliasing of phase difference measurements. The systems and techniques described herein can obtain first depth information about the environment (e.g., iToF-based depth information). The systems and techniques can also obtain second depth information (e.g., image-based depth information determined using image-based depth estimation techniques, such as monocular depth estimation techniques or stereo depth estimation techniques). The systems and techniques can use the second depth information to dealias the first depth information.

[0030] iToF depth cameras can be more expensive and / or larger than image sensors, at least in part because iToF depth cameras are active sensors (emitting light pulses), while image sensors are passive (receiving light but not necessarily emitting it). iToF-based depth information can be more accurate than image-based depth information; for example, monocular image-based depth information may suffer from scale blur, and stereo image-based depth information may contain holes, spikes, or other defects. Furthermore, iToF-based depth information can include absolute distances (e.g., based on the time-of-flight of light), while image-based depth information may not include absolute distances but may include relative distances or approximations based on image blur, etc. However, image-based depth information can provide valuable insights into sequences of objects (e.g., from near to far from the camera).

[0031] The systems and techniques described herein compare sorted image-based depth information with sorted iToF-based depth information. Based on this comparison, these systems and techniques can determine which depths (or groups of depths) in the iToF-based depth information are inconsistent with their corresponding depths in the image-based depth information. Inconsistencies can indicate errors in the iToF-based depth information caused by multi-period aliasing. For example, according to image-based depth information, a wall may appear far from the camera. iToF-based depth information may indicate that the same wall is close to the iToF depth camera (juxtaposed with the camera), because multi-period aliasing can affect iToF depth measurements. By comparing the sorting of depths in the iToF-based depth information with the sorting of depths in the image-based depth information, these systems and techniques can determine errors in the iToF-based depth information.

[0032] These systems and techniques can utilize the ordering of object depths in image-based depth information, rather than depths in the image-based depth information itself. Given the limitations of many image-based depth estimation techniques, such as temporal stability and scale ambiguity, the use of image-based depth information systems and techniques may be more general and robust.

[0033] When these systems and techniques determine an inconsistency between the ordering of depths based on iToF-based depth information and the ordering of depths based on image-based depth information, they can adjust the iToF-based depth information by adding or subtracting an integer number of half-wavelengths of the light pulse from the depth. These systems and techniques can add an integer number of half-wavelengths because iToF depth measurements may deviate from an integer number of half-wavelengths due to multi-period aliasing. After adjusting the iToF-based depth information, these systems and techniques can reorder the depths based on the iToF-based depth information and compare the reordered depths with the ordered depths based on image-based depth information, making additional adjustments based on any additional inconsistencies.

[0034] These systems and techniques provide a robust and efficient solution to the multi-cycle aliasing problem in iToF depth measurement. They combine the benefits of iToF depth measurement (e.g., measured absolute depth, low power, etc.) with the benefits of image-based depth estimation techniques (low cost) to generate more accurate and coherent depth maps than might be produced by using one or the other system and technique alone.

[0035] Various aspects of this application will be described with reference to the accompanying drawings.

[0036] Figure 1This is a block diagram illustrating a system 100 for dealiasing indirect time-of-flight (iToF) depth measurements according to various aspects of the present disclosure. System 100 may include an iToF dealiasing unit 110. System 100 may provide iToF depth information 108 and image-based depth information 106 to the iToF dealiasing unit 110. The iToF dealiasing unit 110 may use the image-based depth information 106 to adjust the iToF depth information 108 to dealias the iToF depth information 108, thereby generating depth information 112.

[0037] In some cases, system 100 may acquire an image 102 of the environment captured using an image sensor. Image 102 may be formatted according to any suitable format, such as red, green, blue (RGB), luminance or lightness, blue projection and red projection (YUV), or grayscale.

[0038] In some cases, system 100 may include a depth estimator 104 to generate image-based depth information 106 based on image 102. The depth estimator 104 may use one or more image-based depth estimation techniques (such as, for example, monocular depth estimation or stereo depth estimation) to generate the image-based depth information 106. Where the depth estimator 104 determines the image-based depth information 106 based on a stereo depth estimation technique, image 102 may include two or more stereo-paired images.

[0039] In cases where system 100 includes a depth estimator 104, the depth estimator 104 can provide image-based depth information 106 to the iToF dealiasing unit 110. In other cases, system 100 may not include a depth estimator 104. In these other cases, system 100 can obtain image-based depth information 106 and provide image-based depth information 106 to the iToF dealiasing unit 110.

[0040] Image-based depth information 106 may be or may include depth information of the environment of image 102. Image-based depth information 106 may include the depth of each of a plurality of points in the environment. According to some aspects, image-based depth information 106 may include the depth of each pixel of image 102.

[0041] The iToF depth information 108 may be or may include depth information of the environment of image 102 and / or image-based depth information 106. The iToF depth information 108 may include the depth of each of a plurality of points in the environment. The iToF depth information 108 may be determined by an iToF depth camera. The number of points in the iToF depth information 108 may be the same as or different from the number of points in the image-based depth information 106 (e.g., based on the differences in resolution and / or location between the camera that captured image 102 and the iToF depth camera).

[0042] The iToF dealiasing unit 110 can compare iToF depth information 108 with image-based depth information 106 to determine inconsistencies between the iToF depth information 108 and the image-based depth information 106. For example, the iToF dealiasing unit 110 can correlate points in the iToF depth information 108 with points in the image-based depth information 106 and compare the depth order of the correlated points between the iToF depth information 108 and the image-based depth information 106. Based on this comparison, the iToF depth information 108 can determine inconsistencies between the depth order of the iToF depth information 108 and the depth order of the image-based depth information 106. The iToF dealiasing unit 110 can adjust the depth of the iToF depth information 108 based on the determined inconsistencies. For example, the iToF dealiasing unit 110 can adjust the depth of points in the iToF depth information 108 that have inconsistent depth ordering between the iToF depth information 108 and the image-based depth information 106. The iToF depth information 108 can be adjusted by adding or subtracting an integer number of half wavelengths from the depth, since inconsistencies may be caused by multi-period aliasing in iToF depth measurements.

[0043] Depth information 112 may be or may include iToF depth information 108, which includes the adjusted depth.

[0044] Figure 2 This is a block diagram illustrating a system 200 for dealiasing iToF depth measurements according to various aspects of the present disclosure. The system 200 can obtain iToF-based depth information 216 and image-based depth information 206, and use the image-based depth information 206 to adjust the iToF-based depth information 216 to dealias the iToF-based depth information 216 to generate depth information 228.

[0045] Image 202 can be used with Figure 1 The image 102 is the same as or substantially similar to the image 102. The depth estimator 204 can be used with... Figure 1The depth estimator 104 is the same as, substantially similar to, or performs the same or substantially the same operation as it. Image-based depth information 206 can be used with... Figure 1 The image-based depth information 106 is the same as or substantially similar to that in system 200. Similar to depth estimator 104 in system 100, depth estimator 204 may be optional in system 200. The iToF-based depth information 216 may be... Figure 1 The iToF depth information 108 is the same as or substantially similar to that of the other two. Depth information 228 can be compared with... Figure 1 The depth information 112 is the same or substantially similar.

[0046] Partitioner 208 can partition image 202 and / or image-based depth information 206 into depth partitions (e.g., using object detection techniques, saliency map techniques, or superpixel techniques). In this disclosure, the term "depth partition" can refer to a set of points in the depth information. A depth partition can include adjacent points with similar depths in an image plane (e.g., in image 202). For example, partitioner 208 can determine a depth partition of all adjacent pixels of image 202 with similar depths based on image-based depth information 206. Partitioner 208 can generate a partition map 210 that defines the depth partitions within image 202 and / or image-based depth information 206.

[0047] Figure 3 This includes a visual representation of example image-based depth information 302 and example image 304 including depth partitions 306. Image 304 can be an example of image 202. Image-based depth information 302 can be an example of image-based depth information 206. Depth partitions 306 can be an example of depth partitions in partitioned diagram 210. Depth in image-based depth information 302 is represented by color; for example, blue pixels represent the shortest depth, and red pixels represent the farthest depth.

[0048] Return to Figure 2 As described above, as an example, partitioner 208 may generate partitioned images 202 and / or depth partitions (e.g., depth partition 306) based on image-based depth information 206. Each of the various depth partitions may include points that are adjacent in the image plane (e.g., in image 202) and have similar depths based on image-based depth estimation techniques.

[0049] Sequencer 212 can sort the depth partitions of image-based depth information 206 by depth to generate sorted depth partitions 214 (e.g., sorting the depth partitions of image-based depth information 206 from shortest depth to farthest depth). The depth of a depth partition can be defined based on a statistical measure of the depths of all points included in the partition. For example, the depth of a partition can be defined by the average or median of the depths of all points included in the partition.

[0050] Sequencer 218 can apply partition map 210 to iToF-based depth information 216. For example, sequencer 218 can correlate one or more pixels of image 202 and / or one or more points of image-based depth information 206 with points of iToF-based depth information 216. Additionally, based on the correlation between iToF-based depth information 216 and image-based depth information 206, sequencer 218 can correlate depth partitions of partition map 210 with iToF-based depth information 216 to partition iToF-based depth information 216 into depth partitions. Sequencer 218 can sort the depth partitions of iToF-based depth information 216 by depth to generate sorted depth partitions 220.

[0051] Figure 3 This includes a visual representation of iToF-based depth information 308, partitioned by depth partition 310. The iToF-based depth information 308 can be an example of iToF-based depth information 216. The depth partition 310 applied to the iToF-based depth information 308 can be an example of a depth partition of a partition map 210 that partitions the iToF-based depth information 216.

[0052] Return to Figure 2 As described, comparator 222 can compare the sorted depth partitions 214 and 220 to determine inconsistencies between image-based depth information 206 and iToF-based depth information 216. For example, comparator 222 can determine corresponding depth partitions of image-based depth information 206 and iToF-based depth information 216 that are sorted differently in the sorted depth partitions 214 and 220 (e.g., outside a sorting threshold). The depth partitions of image-based depth information 206 can be correlated with the depth partitions of iToF-based depth information 216 via partition map 210 (e.g., spatially).

[0053] As an example of identifying inconsistencies, partitioning map 210 may define rectangular depth partitions near the center of image 202. Image-based depth information 206 may include the corresponding rectangular depth partitions and may include the depths corresponding to the rectangular depth partitions. Similarly, iToF-based depth information 216 may include corresponding rectangular depth partitions and may include the depths corresponding to the rectangular depth partitions. The depths of the rectangular depth partitions in image-based depth information 206 may differ from the depths of the rectangular depth partitions in iToF-based depth information 216. The depths may differ because image-based depth information 206 may include scale blur and / or because iToF-based depth information 216 may include errors caused by multi-period aliasing. Therefore, a direct comparison between the depths of image-based depth information 206 and the depths of iToF-based depth information 216 may be particularly useful or may not be particularly useful.

[0054] Continuing with examples of identifying inconsistencies, the sorted depth partition 214 may include sorting information indicating the depth order of rectangular-shaped depth partitions in the image-based depth information 206. For example, according to the sorted depth partition 214, the rectangular-shaped depth partitions of the sorted depth partition 214 may be among the farthest depth partitions in the image-based depth information 206. Additionally, the sorted depth partition 220 may include sorting information indicating the depth order of rectangular-shaped depth partitions in the iToF-based depth information 216. For example, according to the sorted depth partition 220, the rectangular-shaped depth partitions of the sorted depth partition 220 may be among the nearest depth partitions in the sorted depth partition 220. Comparator 222 may determine an inconsistency between the image-based depth information 206 and the iToF-based depth information 216 based on the difference between the depth order of the rectangular-shaped depth partitions in the sorted depth partition 214 and the depth order of the rectangular-shaped depth partitions in the sorted depth partition 220. Determining inconsistencies based on sorted depth partitions 214 and 220 can be a useful method for comparing image-based depth information 206 and iToF-based depth information 216. This comparison can be useful because it highlights errors in the iToF-based depth information 216 caused by multi-period aliasing. Furthermore, this comparison can overcome the effects of scale blurring in image-based depth information 206 by not directly relying on its depth.

[0055] Based on the identified inconsistencies, depth adjuster 224 can adjust the iToF-based depth information 216 (e.g., by adding or subtracting from the depth of the iToF-based depth information 216) to generate the adjusted depth partition 226. Because some errors in the iToF-based depth information 216 are caused by multi-period aliasing, these errors may cause the depth of the iToF-based depth information 216 to deviate from an integer number of half-wavelengths. Therefore, depth adjuster 224 can adjust the iToF-based depth information 216 by adding or subtracting from the depth of the iToF-based depth information 216 using an integer number of half-wavelengths.

[0056] Depth information 228 may be more accurate than image-based depth information 206 because depth information 228 may include absolute depth, such as iToF-based depth information 216. Depth information 228 may be more accurate than iToF-based depth information 216 because system 200 may have already de-aliased the iToF-based depth information 216 to correct for errors based on multi-period aliasing.

[0057] System 200 can iteratively improve depth information 228 by adjusting iToF-based depth information 216 to generate adjusted depth partitions 226, ordering the adjusted depth partitions 226, and comparing the adjusted depth partitions 226 with the sorted depth partitions 214 at comparator 222 to determine whether there are additional inconsistencies between the sorted adjusted depth partitions 226 and the sorted depth partitions 214.

[0058] Additionally or alternatively, depending on some aspects, system 200 may merge two or more depth partitions. For example, system 200 may merge depth partitions where two or more depth partitions are adjacent in the image plane and include similar depths (e.g., based on image-based depth information 206 and / or iToF-based depth information 216 (or adjusted depth partition 226)).

[0059] For example, Figure 4 This includes three representations (representation 402, representation 404, and representation 406) of three corresponding example depth partitions (including depth partition 408, depth partition 410, and depth partition 412). Depth partitions 408, 410, and 412 are adjacent to each other. Furthermore, depth partitions 408, 410, and 412 have similar depths. Therefore, according to some aspects, depth partitions 408, 410, and 412 can be merged into depth partition 416 of representation 414.

[0060] Figure 5This is a block diagram illustrating a system 500 for dealiasing iToF depth measurements according to various aspects of this disclosure.

[0061] System 500 may include camera 502 to generate image 504 of the environment. Image 504 may be used with... Figure 1 Image 102 and / or Figure 2 The image 202 is the same as or substantially similar to the image 202.

[0062] System 500 may include an iToF depth camera 506 to generate iToF depth information 508 of the environment. The iToF depth camera 506 may be positioned relative to camera 502 such that image 504 and iToF depth information 508 represent substantially the same view of the environment. The iToF depth information 508 may be... Figure 1 iToF depth information 108 and / or Figure 2 The depth information based on iToF is the same as or substantially similar to that of 216.

[0063] System 500 may include an iToF dealiasing unit 510. The iToF dealiasing unit 510 can be used with... Figure 1 Image 102 and / or Figure 2 The system is the same as, substantially similar to, or performs the same or substantially the same operation as the system 200.

[0064] The iToF dealiasing unit 510 can generate depth information 512 based on the iToF depth information 508 and the image 504. The depth information 512 can be used with... Figure 1 Depth information 112 and / or Figure 2 The depth information of 228 is the same or substantially similar.

[0065] Figure 6 This is a flowchart illustrating a process 600 for dealiasing iToF depth measurements according to various aspects of this disclosure. One or more operations of process 600 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., chipset, codec, etc.). The computing device may be a vehicle or a component or system of a vehicle, a mobile device (e.g., a mobile phone), a network-connected wearable device such as a watch, an extended reality (XR) device such as a virtual reality (VR) device or an augmented reality (AR) device, a desktop computing device, a tablet computing device, a server computer, a robotic device, a television, and / or any other computing device having the resource capability to perform process 600. One or more operations of process 600 may be implemented as software components that execute and run on one or more processors.

[0066] At box 602, a computing device (or one or more components thereof) may send (or instruct or cause a transmitter to send) electromagnetic (EM) radiation toward multiple points in the environment. For example, Figure 5 The iToF depth camera 506 can emit EM radiation toward the environment.

[0067] At box 604, a computing device (or one or more components thereof) may compare the phase of the transmitted EM radiation with the phase of the received EM radiation to determine a corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of a plurality of points in the environment. For example, an iToF depth camera 506 may receive reflected EM radiation and compare the reflected EM radiation with the transmitted EM radiation to estimate the time of flight of the EM radiation.

[0068] At box 606, a computing device (or one or more components thereof) may determine first depth information based on a corresponding time-of-flight estimate determined for each of a plurality of points in the environment. For example, an iToF depth camera 506 may determine iToF depth information 508 based on the estimated time-of-flight information. Figure 2 The iToF-based depth information 216 can be an example of the first depth information determined at box 606.

[0069] At box 608, the computing device (or one or more components thereof) can obtain second depth information based on an image of the environment. For example, it can obtain... Figure 2 Image-based depth information 206. Image-based depth information 206 can be based on... Figure 2 Image 202.

[0070] In some respects, a computing device (or one or more components thereof) can acquire an image of the environment and use monocular depth estimation techniques to determine second depth information. For example, Figure 2 The depth estimator 204 can obtain image 202 and can generate image-based depth information 206 based on image 202.

[0071] At box 610, the computing device (or one or more components thereof) may compare the first depth information with the second depth information to determine inconsistencies between the two depth information. For example, Figure 2 The comparator 222 can Figure 2 The sorted depth partition 220 (which is depth information based on iToF depth information 216) and Figure 2 The sorted depth partitions 214 (which are based on the depth information of image 202) are compared. Figure 2 The depth information 228 may be based on or include information about inconsistencies between the sorted depth partitions 214 and 220.

[0072] In some aspects, the computing device (or one or more components thereof) may generate a partitioning map based on an image; partition first depth information according to the partitioning map to generate a first depth partition; and partition second depth information according to the partitioning map to generate a second depth partition. Comparing the first depth information with the second depth information (e.g., in box 610) may include comparing the first depth partition with the second depth partition to determine inconsistencies. For example, Figure 2 The partitioner 208 can generate based on image 202. Figure 2 Partition diagram 210. Figure 2 The sequencer 212 can partition the image-based depth information 206 according to the partition map 210 to generate sorted depth partitions 214. The sequencer 218 can partition the iToF-based depth information 216 according to the partition map 210 to generate sorted depth partitions 220. The comparator 222 can compare the sorted depth partitions 214 with the sorted depth partitions 220. In some aspects, generating the partition map may include generating the partition map based on the image using at least one of object detection techniques, saliency map techniques, or superpixel techniques. For example, the partitioner 208 may use object detection techniques, saliency map techniques, or superpixel techniques to generate the partition map 210 based on the image 202.

[0073] In some respects, the partition map can be further generated based on a second depth information. For example, partitioner 208 can generate partition map 210 based on image 202 and iToF-based depth information 216.

[0074] In some aspects, a computing device (or one or more components thereof) may sort a first depth partition by depth and a second depth partition by depth. Comparing the first depth partition with the second depth partition may include comparing the sorted first depth partition with the sorted second depth partition to determine inconsistencies. For example, sequencer 212 may sort depth partitions based on image-based depth information 206 (such as those partitioned according to partition map 210) to generate sorted depth partitions 214. Sequencer 218 may sort depth partitions based on iToF-based depth information 216 (such as those partitioned according to partition map 210) to generate sorted depth partitions 220. Comparator 222 may compare the sorted depth partitions 214 with the sorted depth partitions 220. In some aspects, the first depth partition may be sorted according to a statistical measure of the depth of each first depth partition in the first depth partition indicated by the first depth information. Additionally, the second depth partition may be sorted according to a statistical measure of the depth of each second depth partition in the second depth partition indicated by the second depth information.

[0075] In some aspects, a computing device (or one or more components thereof) may determine multiple inconsistencies based on comparing a sorted first depth partition with a sorted second depth partition; merge multiple partitions of the first depth partition associated with the multiple inconsistencies; and adjust the depth of the merged multiple partitions based on the multiple inconsistencies. For example, comparator 222 may determine multiple inconsistencies when comparing sorted depth partition 214 with sorted depth partition 220. Additionally, comparator 222 may merge multiple partitions (e.g., Figure 4 (Deep partitions 408, 410, and 412). Additionally, the depth adjuster 224 can adjust the distance of the depth information 228 (e.g., including merged depth partitions).

[0076] At box 612, the computing device (or one or more components thereof) may adjust the depth of the first depth information based on inconsistencies. For example, Figure 2 The depth adjuster 224 can adjust the sorted depth partitions 220 based on the inconsistencies determined at box 610.

[0077] In some aspects, adjusting the depth of the first depth information based on inconsistency (e.g., at box 12) may include adjusting the depth of depth partitions within the first depth partition. For example, comparator 222 may adjust the depth of partitions within the sorted depth partition 220 based on inconsistencies between the sorted depth partition 220 and the sorted depth partition 214. In some aspects, a computing device (or one or more components thereof) may generate third depth information based on the first depth information and the adjusted depths of the depth partitions. For example, comparator 222 may generate depth information 228, which may include all depths of the sorted depth partitions 220 and one or more adjusted depths.

[0078] In some aspects, adjusting the depth of the first depth information may include adding or subtracting the distance based on the wavelength of the EM radiation from the depth of the first depth information. For example, comparator 222 may add or subtract an integer number of half wavelengths of the EM radiation transmitted at block 602 from the depth of the sorted depth partition 220 to adjust the sorted depth partition 220.

[0079] In some examples, as previously noted, the methods described herein (e.g., process 600 and / or other methods described herein) may be performed wholly or partially by a computing device or apparatus. In one example, one or more methods (e.g., process 600 and / or other methods described herein) may be performed by... Figure 1 System 100 Figure 1 iToF dealiasing device 110 Figure 2 System 200 Figure 2 Partitioner 208 Figure 2 Sequencer 212, Figure 2 Sequencer 218 Figure 2 Comparator 222, Figure 2 Depth adjuster 224 Figure 5 System 500 and / or Figure 5 The iToF dealiasing unit 510 is used to perform this. In another example, one or more methods can be performed by... Figure 9 The computing device architecture 900 shown is implemented wholly or partially. For example, it has Figure 9 The computing device of the computing device architecture 900 shown may include Figure 1 System 100 Figure 1 iToF dealiasing device 110 Figure 2 System 200 Figure 2 Partitioner 208 Figure 2 Sequencer 212, Figure 2 Sequencer 218 Figure 2 Comparator 222, Figure 2 Depth adjuster 224 Figure 5 System 500 and / or Figure 5 The components of the iToF dealiasing unit 510, and can achieve Figure 6 The operation of process 600 and / or other processes described herein. In some cases, a computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other types of data.

[0080] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0081] Process 600 and / or other processes described herein are illustrated as logic flowcharts, whose operations represent sequences of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0082] Furthermore, process 600 and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors, implemented in hardware, or implemented in a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0083] As noted above, various aspects of this disclosure may utilize machine learning models or systems.

[0084] Figure 7 This is an exemplary example of a neural network 700 (e.g., a deep learning neural network) that can be used to implement machine learning-based partitioning, image-based depth estimation, feature segmentation, implicit neural representation generation, rendering, and / or classification as described above.

[0085] Input layer 702 includes input data. In one exemplary example, input layer 702 may include data representing an image. Neural network 700 includes multiple hidden layers: hidden layers 706a, 706b through 706n. Hidden layers 706a, 706b through 706n comprise "n" hidden layers, where "n" is an integer greater than or equal to one. Multiple hidden layers can be made to include as many layers as needed for a given application. Neural network 700 also includes an output layer 704, which provides the output produced by the processing performed by hidden layers 706a, 706b through 706n. In one exemplary example, output layer 704 may provide depth information.

[0086] Neural network 700 may be or may include a multi-layer neural network with interconnected nodes. Each node may represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains information while processing it. In some cases, neural network 700 may include a feedforward network, in which case there are no feedback connections in which the network's output is fed back into itself. In some cases, neural network 700 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.

[0087] Information can be exchanged between nodes via node-to-node interconnects between layers. Nodes in input layer 702 can activate a set of nodes in the first hidden layer 706a. For example, as shown, each input node in input layer 702 is connected to each node in the first hidden layer 706a. Nodes in the first hidden layer 706a can transform the information of each input node by applying an activation function to the input node information. The information derived from this transformation can then be passed to nodes in the next hidden layer 706b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 706b can then activate nodes in the next hidden layer, and so on. Finally, the output of hidden layer 706n can activate one or more nodes in output layer 704, providing the output at those nodes. In some cases, although a node in neural network 700 (e.g., node 708) is shown as having multiple output lines, the node has a single output and is shown as all lines output from the node representing the same output value.

[0088] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 700. Once the neural network 700 is trained, it can be called a trained neural network, which can be used to perform one or more operations. For example, the interconnection between nodes may represent a piece of information about what the interconnected nodes have learned. The interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 700 to adapt to the input and learn as more and more data is processed.

[0089] The neural network 700 can be pre-trained to process features from the data in the input layer 702 using different hidden layers 706a, 706b to 706n, so as to provide an output through the output layer 704. In an example where the neural network 700 is used to identify features in an image, the neural network 700 can be trained using training data that includes both images and labels, as described above. For example, training images can be input into the network, where each training image has a label indicating a feature in the image (for a feature segmentation machine learning system) or a label indicating the category of activity in each image. In an example where object classification is used for illustrative purposes, the training images may include images of the number 2, in which case the image label could be [0 0 1 0 0 0 0 0 0 0].

[0090] In some cases, the neural network 700 can use a training process called backpropagation to adjust the weights of its nodes. As noted above, the backpropagation process can include forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. This process can be repeated a certain number of times for each set of training images until the neural network 700 is trained well enough that the layer weights are accurately tuned.

[0091] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 700. The weights are initially randomized before training the neural network 700. As an illustrative example, the image may include a numerical array representing the pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).

[0092] As noted above, for the first training iteration of the neural network 700, the output will likely include values ​​due to the weights being randomly chosen during initialization, without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each class may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). Using the initial weights, the neural network 700 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined, such as cross-entropy loss. Another example of a loss function is mean squared error (MSE), defined as... The loss can be set to equal E. 总计 The value of .

[0093] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training label. The Neural Network 700 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce the loss until it is minimized. The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... Where w represents the weight, w i Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.

[0094] Neural Network 700 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN include a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural Network 700 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.

[0095] Figure 8 This is an exemplary example of a Convolutional Neural Network (CNN) 800. The input layer 802 of the CNN 800 includes data representing an image or frame. For example, the data may include a numerical array representing pixels of an image, where each number in the array includes a value from 0 to 255 describing the intensity of the pixel at that location in the array. Using the previous example from above, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or lightness and two chroma components, etc.). The image can be passed through a convolutional hidden layer 804, an optional non-linear activation layer, a pooling hidden layer 806, and a fully connected layer 808 (which may be hidden) to obtain an output at the output layer 810. Although Figure 8 Only one hidden layer is shown in the diagram, but those skilled in the art will understand that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers may be included in a CNN 800. As previously described, the output may indicate a single category of an object, or may include probabilities that best describe the category of an object in an image.

[0096] The first layer of a CNN 800 can be a convolutional hidden layer 804. The convolutional hidden layer 804 analyzes the image data from the input layer 802. Each node in the convolutional hidden layer 804 is connected to a region of the input image called a receptive field (pixel). The convolutional hidden layer 804 can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in the convolutional hidden layer 804. For example, the region of the input image covered by the filter at each convolutional iteration will be the receptive field of the filter. In an exemplary example, if the input image comprises a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in the convolutional hidden layer 804. Each connection between a node and its receptive field learns weights and, in some cases, learns an overall bias, such that each node learns to analyze its specific local receptive field in the input image. Each node in the convolutional hidden layer 804 will have the same weights and biases (called shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the image frame example, the filter would have a depth of 3 (based on the three color components of the input image). An exemplary example of the filter array size is 5×5×3, corresponding to the size of the receptive field of a node.

[0097] The convolutional nature of the convolutional hidden layer 804 is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 804 may begin at the top left corner of the input image array and may convolve around the input image. As noted above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 804. In each convolutional iteration, the filter value is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain the sum of that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 804. For example, the filter may move a step size (called stride) to the next receptive field. The stride may be set to 1 or another suitable amount. For example, if the stride is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location in the input volume produces a number representing the filter result at that location, thus determining a sum value for each node of the convolutional hidden layer 804.

[0098] The map construction from the input layer to the convolutional hidden layer 804 is called an activation map (or feature map). An activation map includes values ​​for each node representing the filter results at each location in the input volume. Activation maps can include arrays containing various sums of values ​​produced by each iteration of the filter over the input volume. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 804 can include several activation maps to identify multiple features in the image. Figure 8 The example shown includes three activation maps. Using these three activation maps, the convolutional hidden layer 804 can detect three different types of features, each of which is detectable across the entire image.

[0099] In some examples, nonlinear hidden layers can be applied after convolutional hidden layers 804. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An exemplary example of a nonlinear layer is the Corrected Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0,x) to all values ​​in the input volume, which changes all negative activations to 0. Therefore, ReLU can increase the nonlinearity of CNN 800 without affecting the receptive field of convolutional hidden layers 804.

[0100] A pooling hidden layer 806 can be applied after the convolutional hidden layer 804 (and, in use, after a non-linear hidden layer). The pooling hidden layer 806 is used to simplify the information in the output of the convolutional hidden layer 804. For example, the pooling hidden layer 806 takes each activation map output from the convolutional hidden layer 804 and uses a pooling function to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by the pooling hidden layer. The pooling hidden layer 806 uses other forms of pooling functions, such as average pooling, L2 norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 804. Figure 8 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 804.

[0101] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a stride (e.g., equal to the dimension of the filter, such as stride 2) to the activation map output from convolutional hidden layer 804. The output from the max pooling filter includes the maximum number in each sub-region of the filter convolution. Using a 2×2 filter as an example, each unit in the pooling layer summarizes a region of 2×2 nodes from the previous layer (each node is a value in the activation map). For example, four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum of the four values ​​being output as the "maximum" value. If such a max pooling filter is applied to an activation filter of 24×24 nodes from convolutional hidden layer 804, the output from pooling hidden layer 806 will be an array of 12×12 nodes.

[0102] In some examples, L2 norm pooling filters may also be used. L2 norm pooling filters involve calculating the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as done in max pooling), and using the calculated value as the output.

[0103] Pooling functions (e.g., max pooling, L2 norm pooling, or other pooling functions) determine whether a given feature is found anywhere in a region of an image and discard the exact localization information. This can be done without affecting the results of feature detection because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of having far fewer pooling features, thus reducing the number of parameters required in subsequent layers of a CNN 800.

[0104] The final connection in the network is a fully connected layer, which connects each node from the pooling hidden layer 806 to each output node in the output layer 810. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 804 comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling hidden layer 806 comprises 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region in each of the three feature maps. Extending this example, the output layer 810 may include ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 806 is connected to each node of the output layer 810.

[0105] The fully connected layer 808 takes the output of the previous pooling hidden layer 806 (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 808 can determine the high-level features most relevant to a particular class and may include weights (nodes) for those high-level features. The product between the weights of the fully connected layer 808 and the pooling hidden layer 806 can be computed to obtain the probabilities for different classes. For example, if the CNN 800 is used to predict that the object in an image is a person, there will be high values ​​in the activation map representing the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the top left and top right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features common to people).

[0106] In some examples, the output from output layer 810 may include an M-dimensional vector (M = 10 in the previous example). M indicates the number of classes from which the CNN 800 must choose when classifying objects in an image. Other example outputs may also be provided. Each number in the M-dimensional vector represents the probability that an object belongs to a certain class. In an exemplary example, if the 10-dimensional output vector represents objects of ten different classes as [0 0 0.05 0.8 0 0.15 0 0 0 0], then the vector indicates a 5% probability that the image is an object of the third class (e.g., a dog), an 80% probability that the image is an object of the fourth class (e.g., a person), and a 15% probability that the image is an object of the sixth class (e.g., a kangaroo). The probability of a class can be considered as the confidence level that an object is part of that class.

[0107] Figure 9 An example computing device architecture 900 is illustrated, illustrating example computing devices that can implement the various technologies described herein. In some examples, the computing device may include a mobile device, a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a video server, a vehicle (or a computing device within a vehicle), or other devices. For example, computing device architecture 900 may include or implement... Figure 1 System 100 Figure 1 iToF dealiasing device 110 Figure 2 System 200 Figure 2 Partitioner 208 Figure 2 Sequencer 212, Figure 2 Sequencer 218 Figure 2 Comparator 222, Figure 2 Depth adjuster 224 Figure 5 System 500 and / or Figure 5 Either of the iToF dealiasers 510.

[0108] The components of the computing device architecture 900 are shown to communicate electrically with each other using a connection 912, such as a bus. The example computing device architecture 900 includes a processing unit (CPU or processor) 902 and a computing device connection 912 that couples various computing device components, including computing device memories 910 (such as read-only memory (ROM) 908 and random access memory (RAM) 906), to the processor 902.

[0109] The computing device architecture 900 may include a cache of high-speed memory that is directly connected to, very close to, or integrated as part of the processor 902. The computing device architecture 900 may copy data from memory 910 and / or storage device 914 to cache 904 for fast access by the processor 902. In this way, the cache provides a performance improvement by avoiding latency for the processor 902 while waiting for data. These and other modules may control or be configured to control the processor 902 to perform various actions. Other computing device memory 910 may also be used. Memory 910 may include various different types of memory with different performance characteristics. The processor 902 may include any general-purpose processor and hardware or software services configured to control the processor 902 (such as services 1 916, service 2 918, and service 3 920 stored in storage device 914), as well as dedicated processors in which software instructions are incorporated into the processor design. The processor 902 may be a self-contained system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0110] To enable user interaction with the computing device architecture 900, input device 922 can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. Output device 924 can also be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, projector, television, speaker device, etc. In some instances, a multi-mode computing device allows the user to provide multiple types of input to communicate with the computing device architecture 900. Communication interface 926 typically controls and manages user input and computing device output. There are no limitations on operation on any particular hardware arrangement, and therefore the underlying features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0111] Storage device 914 is a non-volatile memory and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as a magnetic tape cassette, flash memory card, solid-state memory device, digital multifunction disk, magnetic tape cartridge, random access memory (RAM) 906, read-only memory (ROM) 908, and hybrid forms thereof. Storage device 914 may include services 916, 918, and 920 for controlling processor 902. Other hardware or software modules are envisioned. Storage device 914 may be connected to computing device connection 912. In one aspect, a hardware module performing a specific function may include software components stored in a computer-readable medium connected to necessary hardware components, such as processor 902, connection 912, output device 924, etc., to perform the function.

[0112] For example, the term "substantially" in relation to a given parameter, characteristic, or condition may mean the degree to which a given parameter, characteristic, or condition is satisfied with a reasonable range of variation, as would be understood by one of ordinary skill in the art, for example, within acceptable manufacturing tolerances. For instance, depending on the specific parameter, characteristic, or condition that is substantially satisfied, the parameter, characteristic, or condition may be satisfied at least 90%, at least 95%, or even at least 99%.

[0113] Various aspects of this disclosure are applicable to any suitable electronic device (such as a security system, smartphone, tablet computer, laptop computer, vehicle, drone, or other device) that includes or is coupled to one or more active depth sensing systems. Although devices having or coupled to a light projector are described below, various aspects of this disclosure are applicable to devices having any number of light projectors and are therefore not limited to any particular device.

[0114] The term "device" is not limited to one or a specific number of physical objects (such as a smartphone, a controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that implement at least some parts of this disclosure. Although the following description and examples use the term "device" to describe various aspects of this disclosure, the term "device" is not limited to a particular configuration, type, or number of objects. Additionally, the term "system" is not limited to multiple components or a particular aspect. For example, a system may be implemented on one or more printed circuit boards or other substrates and may have movable or static components. Although the following description and examples use the term "system" to describe various aspects of this disclosure, the term "system" is not limited to a particular configuration, type, or number of objects.

[0115] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks comprising devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these aspects in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring aspects.

[0116] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but it may have additional steps not included in the diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.

[0117] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, cause or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion of the computer resources used may be accessible via a network. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc.

[0118] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media can include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as flash memory), memory or memory devices, magnetic disks or optical discs, flash memory, USB devices provided with non-volatile memory, network storage devices, compressed optical discs (CDs) or digital versatile optical discs (DVDs), any suitable combinations thereof, etc. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments can be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0119] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as power consumption, carrier signals, electromagnetic waves, and the signals themselves.

[0120] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor performs the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or intercalation cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed within a single device.

[0121] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0122] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the various inventive concepts can be implemented and employed in a variety of other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above can be used individually or in combination. Furthermore, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0123] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols without departing from the scope of this specification.

[0124] When a component is described as being “configured” to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0125] The phrase “coupled to” means any component is physically connected directly or indirectly to another component, and / or any component communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0126] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, the claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the claim language stating "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0127] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the aspects disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0128] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0129] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0130] The exemplary aspects of this disclosure include:

[0131] Aspect 1. An apparatus for determining depth information, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: transmit electromagnetic (EM) radiation toward a plurality of points in an environment by at least one transmitter; compare the phase of the transmitted EM radiation with the phase of received EM radiation to determine a corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment; determine first depth information based on the corresponding time-of-flight estimate determined for each of the plurality of points in the environment; obtain second depth information based on an image of the environment; compare the first depth information with the second depth information to determine an inconsistency between the first depth information and the second depth information; and adjust the depth of the first depth information based on the inconsistency. In some cases, the apparatus includes the at least one transmitter configured to transmit the EM radiation toward a plurality of points in the environment.

[0132] Aspect 2. The apparatus according to aspect 1, wherein the at least one processor is further configured to: acquire the image of the environment; and determine the second depth information using a monocular depth estimation technique.

[0133] Aspect 3. The apparatus according to any one of Aspect 1 or 2, wherein the at least one processor is further configured to: generate a partition map based on the image; partition the first depth information according to the partition map to generate a first depth partition; and partition the second depth information according to the partition map to generate a second depth partition; wherein, in order to compare the first depth information with the second depth information, the at least one processor is further configured to compare the first depth partition with the second depth partition to determine the inconsistency.

[0134] Aspect 4. The apparatus according to aspect 3, wherein, in order to generate the partition map, the at least one processor is further configured to generate the partition map based on the image using at least one of object detection technology, saliency map technology, or superpixel technology.

[0135] Aspect 5. The apparatus according to any one of Aspects 3 or 4, wherein the at least one processor is further configured to: sort the first depth partition by depth; and sort the second depth partition by depth; wherein, in order to compare the first depth partition with the second depth partition, the at least one processor is further configured to compare the sorted first depth partition with the sorted second depth partition to determine the inconsistency.

[0136] Aspect 6. The apparatus according to aspect 5, wherein the first depth partitions are sorted according to a statistical measure of the depth of each first depth partition in the first depth partitions as indicated by the first depth information.

[0137] Aspect 7. The apparatus according to any one of Aspects 5 or 6, wherein the at least one processor is further configured to: determine a plurality of inconsistencies based on comparing a sorted first depth partition with a sorted second depth partition; merge a plurality of partitions of the first depth partition associated with the plurality of inconsistencies; and adjust the depth of the merged plurality of partitions based on the plurality of inconsistencies.

[0138] Aspect 8. The apparatus according to any one of Aspects 3 to 7, wherein, in order to adjust the depth of the first depth information based on the inconsistency, the at least one processor is further configured to adjust the depth of the depth partition in the first depth partition.

[0139] Aspect 9. The apparatus according to aspect 8, wherein the at least one processor is further configured to generate third depth information based on the first depth information and the adjusted depth of the depth partition.

[0140] Aspect 10. The apparatus according to any one of Aspects 3 to 9, wherein the partition map is further generated based on the second depth information.

[0141] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein, in order to adjust the depth of the first depth information, the at least one processor is further configured to add or subtract a distance based on the wavelength of the EM radiation from the depth of the first depth information.

[0142] Aspect 12. The apparatus according to any one of aspects 1 to 11, wherein the at least one processor is further configured to generate third depth information based on the first depth information including the first depth information of an adjusted depth.

[0143] Aspect 13. A method for determining depth information, the method comprising: transmitting electromagnetic (EM) radiation toward a plurality of points in an environment; comparing the phase of the transmitted EM radiation with the phase of a received EM radiation to determine a corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment; determining first depth information based on the corresponding time-of-flight estimate determined for each of the plurality of points in the environment; obtaining second depth information based on an image of the environment; comparing the first depth information with the second depth information to determine an inconsistency between the first depth information and the second depth information; and adjusting the depth of the first depth information based on the inconsistency.

[0144] Aspect 14. The method according to aspect 13, the method further comprising: obtaining the image of the environment; and using a monocular depth estimation technique to determine the second depth information.

[0145] Aspect 15. The method according to any one of Aspects 13 or 14, the method further comprising: generating a partition map based on the image; partitioning the first depth information according to the partition map to generate a first depth partition; and partitioning the second depth information according to the partition map to generate a second depth partition; wherein comparing the first depth information with the second depth information comprises: comparing the first depth partition with the second depth partition to determine the inconsistency.

[0146] Aspect 16. The method according to aspect 15, wherein generating the partition map comprises: generating the partition map based on the image using at least one of object detection technology, saliency map technology, or superpixel technology.

[0147] Aspect 17. The method according to any one of Aspects 15 or 16, the method further comprising: sorting the first depth partition by depth; and sorting the second depth partition by depth; wherein comparing the first depth partition with the second depth partition comprises: comparing the sorted first depth partition with the sorted second depth partition to determine the inconsistency.

[0148] Aspect 18. The method according to aspect 17, wherein the first depth partition is sorted according to a statistical measure of the depth of each first depth partition in the first depth partition as indicated by the first depth information.

[0149] Aspect 19. The method according to any one of Aspects 17 or 18, the method further comprising: determining a plurality of inconsistencies based on comparing a sorted first depth partition with a sorted second depth partition; merging a plurality of partitions of the first depth partition associated with the plurality of inconsistencies; and adjusting the depth of the merged plurality of partitions based on the plurality of inconsistencies.

[0150] Aspect 20. The method according to any one of Aspects 15 to 19, wherein adjusting the depth of the first depth information based on the inconsistency comprises: adjusting the depth of the depth partition in the first depth partition.

[0151] Aspect 21. The method according to aspect 20, the method further comprising generating third depth information based on the first depth information and the adjusted depth of the depth partition.

[0152] Aspect 22. The method according to any one of Aspects 15 to 21, wherein the partition map is further generated based on the second depth information.

[0153] Aspect 23. The method according to any one of Aspects 13 to 22, wherein adjusting the depth of the first depth information comprises: adding or subtracting a distance based on the wavelength of the EM radiation from the depth of the first depth information.

[0154] Aspect 24. The method according to any one of aspects 13 to 23, the method further comprising: generating third depth information based on the first depth information including the first depth information of an adjusted depth.

[0155] Aspect 25. The method according to any one of aspects 15 to 24, wherein determining the inconsistency comprises: identifying a first depth partition in the sorted first depth partition that is outside the sorting threshold from a corresponding second depth partition in the sorted second depth partition.

[0156] Aspect 26. The method according to any one of Aspects 15 to 25, the method further comprising: reordering the first depth partitions by depth based on adjusting the depth of the depth partitions; comparing the reordered first depth partitions with the sorted second depth partitions to determine a second inconsistency; and adjusting the depth of the second depth partitions in the first depth partitions based on the second inconsistency.

[0157] Aspect 27. The apparatus according to any one of aspects 3 to 12, wherein, in order to determine the inconsistency, the at least one processor is further configured to: identify a first depth partition in the sorted first depth partition that is outside the sorting threshold from a corresponding second depth partition in the sorted second depth partition.

[0158] Aspect 28. The apparatus according to any one of Aspects 2 to 12 or 27, wherein the at least one processor is further configured to: reorder the first depth partition by depth based on the adjusted depth of the depth partition; compare the reordered first depth partition with the ordered second depth partition to determine a second inconsistency; and adjust the depth of the second depth partition in the first depth partition based on the second inconsistency.

[0159] Aspect 29. A method for determining depth information, the method comprising: obtaining first depth information of an environment using an indirect time-of-flight depth estimation technique; obtaining an image of the environment; generating second depth information based on the image; generating a partition map based on the image; partitioning the first depth information according to the partition map to generate first depth partitions; partitioning the second depth information according to the partition map to generate second depth partitions; sorting the first depth partitions by depth; sorting the second depth partitions by depth; comparing the sorted first depth partitions with the sorted second depth partitions to determine inconsistencies between the first depth information and the second depth information; and adjusting the depth of the depth partitions in the first depth partitions based on the inconsistencies.

[0160] Aspect 30. An apparatus for determining depth information, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: obtain first depth information of an environment using an indirect time-of-flight depth estimation technique; obtain an image of the environment; generate second depth information based on the image; generate a partition map based on the image; partition the first depth information according to the partition map to generate first depth partitions; partition the second depth information according to the partition map to generate second depth partitions; sort the first depth partitions by depth; sort the second depth partitions by depth; compare the sorted first depth partitions with the sorted second depth partitions to determine inconsistencies between the first depth information and the second depth information; and adjust the depth of the depth partitions in the first depth partitions based on the inconsistencies.

[0161] Aspect 31. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform any one of aspects 13 to 26.

[0162] Aspect 32. An apparatus for providing virtual content for display, the apparatus comprising one or more components for performing operations according to any one of aspects 13 to 26.

Claims

1. An apparatus for determining depth information, the apparatus comprising: At least one memory; and At least one processor, the at least one processor being coupled to the at least one memory and being configured to: At least one transmitter sends electromagnetic (EM) radiation toward multiple points in the environment; The phase of the transmitted EM radiation is compared with the phase of the received EM radiation to determine the corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment. The first depth information is determined based on the corresponding time-of-flight estimate determined for each of the plurality of points in the environment; Second depth information is obtained based on images of the environment. The first depth information is compared with the second depth information to determine the inconsistency between the first depth information and the second depth information; as well as The depth of the first depth information is adjusted based on the inconsistency.

2. The apparatus of claim 1, wherein the at least one processor is further configured to: Obtain the image of the environment; and The second depth information is determined using monocular depth estimation techniques.

3. The apparatus of claim 1, wherein the at least one processor is further configured to: Generate a partition map based on the image; The first depth information is partitioned according to the partition map to generate a first depth partition; as well as The second depth information is partitioned according to the partition map to generate a second depth partition; In order to compare the first depth information with the second depth information, the at least one processor is further configured to compare the first depth partition with the second depth partition to determine the inconsistency.

4. The apparatus of claim 3, wherein, in order to generate the partition map, the at least one processor is further configured to generate the partition map based on the image using at least one of object detection technology, saliency mapping technology, or superpixel technology.

5. The apparatus of claim 3, wherein the at least one processor is further configured to: Sort the first depth partition by depth; and Sort the second depth partitions by depth; In order to compare the first depth partition with the second depth partition, the at least one processor is further configured to compare the sorted first depth partition with the sorted second depth partition to determine the inconsistency.

6. The apparatus of claim 5, wherein the first depth partitions are ordered according to a statistical measure of the depth of each first depth partition in the first depth partitions as indicated by the first depth information.

7. The apparatus of claim 5, wherein the at least one processor is further configured to: Multiple inconsistencies are identified by comparing the sorted first-depth partition with the sorted second-depth partition; Merge multiple partitions of the first depth partition associated with the multiple inconsistencies; and The depth of the merged partitions is adjusted based on the aforementioned inconsistencies.

8. The apparatus of claim 3, wherein, in order to adjust the depth of the first depth information based on the inconsistency, the at least one processor is further configured to adjust the depth of the depth partition in the first depth partition.

9. The apparatus of claim 8, wherein the at least one processor is further configured to generate third depth information based on the first depth information and the adjusted depth of the depth partition.

10. The apparatus of claim 3, wherein the partition map is further generated based on the second depth information.

11. The apparatus of claim 1, wherein, in order to adjust the depth of the first depth information, the at least one processor is further configured to add or subtract a distance based on the wavelength of the EM radiation from the depth of the first depth information.

12. The apparatus of claim 1, wherein the at least one processor is further configured to generate third depth information based on the first depth information, which includes the first depth information and an adjusted depth.

13. The apparatus of claim 1, further comprising the at least one transmitter configured to transmit electromagnetic (EM) radiation toward a plurality of points in the environment.

14. A method for determining depth information, the method comprising: It emits electromagnetic (EM) radiation toward multiple points in the environment; The phase of the transmitted EM radiation is compared with the phase of the received EM radiation to determine the corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment. The first depth information is determined based on the corresponding time-of-flight estimate determined for each of the plurality of points in the environment; Second depth information is obtained based on images of the environment. The first depth information is compared with the second depth information to determine the inconsistency between the first depth information and the second depth information; as well as The depth of the first depth information is adjusted based on the inconsistency.

15. The method according to claim 14, further comprising: Obtain the image of the environment; as well as The second depth information is determined using monocular depth estimation techniques.

16. The method of claim 14, further comprising: Generate a partition map based on the image; The first depth information is partitioned according to the partition map to generate a first depth partition; as well as The second depth information is partitioned according to the partition map to generate a second depth partition; The comparison of the first depth information with the second depth information includes: comparing the first depth partition with the second depth partition to determine the inconsistency.

17. The method of claim 16, wherein generating the partition map comprises: The partition map is generated based on the image using at least one of object detection techniques, saliency mapping techniques, or superpixel techniques.

18. The method according to claim 16, further comprising: Sort the first depth partitions by depth; as well as Sort the second depth partitions by depth; The comparison of the first depth partition with the second depth partition includes: comparing the sorted first depth partition with the sorted second depth partition to determine the inconsistency.

19. The method of claim 18, wherein the first depth partitions are sorted according to a statistical measure of the depth of each first depth partition in the first depth partitions as indicated by the first depth information.

20. The method according to claim 18, further comprising: Multiple inconsistencies are identified by comparing the sorted first-depth partition with the sorted second-depth partition; Merge multiple partitions of the first depth partitions associated with the multiple inconsistencies; as well as The depth of the merged partitions is adjusted based on the aforementioned inconsistencies.

21. The method of claim 16, wherein adjusting the depth of the first depth information based on the inconsistency comprises: Adjust the depth of the depth partition in the first depth partition.

22. The method of claim 21, further comprising generating third depth information based on the first depth information and the adjusted depth of the depth partition.

23. The method of claim 16, wherein the partition map is further generated based on the second depth information.

24. The method of claim 14, wherein adjusting the depth of the first depth information comprises: The distance based on the wavelength of the EM radiation is added to or subtracted from the depth of the first depth information.

25. The method according to claim 14, further comprising: The third depth information is generated based on the first depth information, which includes the first depth information and is adjusted to the first depth information.

26. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the at least one processor, when executed, to: Instruct at least one transmitter to send electromagnetic (EM) radiation toward multiple points in the environment; The phase of the transmitted EM radiation is compared with the phase of the received EM radiation to determine the corresponding time-of-flight estimate of the EM radiation between transmission and reception for each of the plurality of points in the environment. The first depth information is determined based on the corresponding time-of-flight estimate determined for each of the plurality of points in the environment; Second depth information is obtained based on images of the environment. The first depth information is compared with the second depth information to determine the inconsistency between the first depth information and the second depth information; as well as The depth of the first depth information is adjusted based on the inconsistency.

27. The non-transitory computer-readable storage medium of claim 26, wherein the instructions, when executed by the at least one processor, cause the at least one processor to: Obtain the image of the environment; and The second depth information is determined using monocular depth estimation techniques.

28. The non-transitory computer-readable storage medium of claim 26, wherein the instructions, when executed by the at least one processor, cause the at least one processor to: Generate a partition map based on the image; The first depth information is partitioned according to the partition map to generate a first depth partition; as well as The second depth information is partitioned according to the partition map to generate a second depth partition; In order to compare the first depth information with the second depth information, the instruction, when executed by the at least one processor, causes the at least one processor to: compare the first depth partition with the second depth partition to determine the inconsistency.

29. The non-transitory computer-readable storage medium of claim 28, wherein, in order to generate the partition map, the instructions, when executed by the at least one processor, cause the at least one processor to generate the partition map based on the image using at least one of object detection technology, saliency map technology, or superpixel technology.

30. The non-transitory computer-readable storage medium of claim 28, wherein the instructions, when executed by the at least one processor, cause the at least one processor to: Sort the first depth partitions by depth; as well as Sort the second depth partitions by depth; In order to compare the first depth partition with the second depth partition, the instruction, when executed by the at least one processor, causes the at least one processor to: compare the sorted first depth partition with the sorted second depth partition to determine the inconsistency.