Device for measuring and method for determining the distance between two points in an environment
By combining an RGB camera and an i-ToF sensor to form a distance estimation system, and using a machine learning model to fuse distance data, the uncertainty and error problems in the sensor system are solved, resulting in more accurate distance measurement and improved depth resolution.
Patent Information
- Application Number
- CN202110924037.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-06-30
- Filing Date
- 2021-08-12
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2041-08-12
AI Technical Summary
Existing sensor systems suffer from uncertainties and errors when measuring distances. In particular, i-ToF sensors have limited discernible distances at high modulation frequencies, leading to unknown errors. Furthermore, multi-frequency modulation sensing may cause dealiasing errors, affecting depth resolution and accuracy.
The distance estimation system combines an RGB camera and an i-ToF sensor. It integrates the distance data from both through a machine learning model, using the RGB camera to provide an initial estimate and the i-ToF sensor to provide a refined estimate. The final distance is then calculated by a computing component, reducing errors and improving accuracy.
It achieves more accurate distance measurement, improves depth resolution and field of view, reduces uncertainty and error in the sensor system, and enhances the system's functionality and performance.
Smart Images

Figure CN114076951B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application generally relates to a sensor system. More specifically, the present application relates to a system that can detect an optical signal, estimate a distance of an object, and refine the estimated distance to determine an accurate distance measurement. The system can initiate one or more operations based on the accurate distance measurement. BACKGROUND
[0002] Sensors are widely used in applications such as smartphones, robotics, and autonomous vehicles to determine characteristics of objects in an environment (e.g., object recognition, object classification, depth information, edge detection, motion information, etc.). However, the determined results can have uncertainty, i.e., contain estimation errors. For example, an RGB camera can be used to detect edges of an object, but if the object partially overlaps with another object, a software (e.g., a machine learning model) that uses RGB images to detect edges of an object can produce an output with uncertainty.
[0003] In another example, a three-dimensional image sensor, or depth sensor, can use principles such as stereo, direct time-of-flight (d-ToF), and indirect time-of-flight (i-ToF) techniques. However, i-ToF sensors can have problems with ambiguous errors. For example, in an i-ToF system, a high modulation frequency (e.g., 300 MHz) can help improve depth resolution and accuracy, but the unambiguous range is limited (e.g., 50 cm). Thus, an i-ToF sensor can not be able to distinguish between a distance of 10 cm and a distance of 60 cm, resulting in ambiguous errors, or aliasing.
[0004] One solution to the ambiguous error problem of i-ToF systems is to use an additional low modulation frequency (e.g., 75 MHz) in the time domain to extend the range of visibility. However, if the multiple modulation frequencies used in this solution are separated in time, the system frame rate can be reduced, resulting in motion artifacts and other hardware side effects. Alternatively, if the system frame rate is maintained constant when adding multiple modulation frequencies, the integration time for each frequency can be shortened, resulting in reduced depth resolution and accuracy. Furthermore, multiple frequency modulation can still have de-aliasing errors, especially when the system signal-to-noise ratio (SNR) is reduced, such as at a long distance.
[0005] To overcome the sensor uncertainty or error problem, multiple types of sensors can be used in conjunction to obtain more accurate results. For example, to resolve the unknown error and / or de-aliasing error while maintaining the system frame rate and depth resolution, a low depth resolution depth map generated by a distance estimation system (e.g., an RGB camera connected with a machine learning model) can be combined with a high depth resolution depth map generated by an i-ToF sensor to obtain a final depth map with high depth resolution / accuracy and wide visibility range. SUMMARY
[0006] Various aspects and advantages of the embodiments will become apparent from the following description, or can be learned by practice of the embodiments.
[0007] One example aspect of the present disclosure is a device for measurement, including a distance estimation system having a sensor system. The distance estimation system is configured to receive a first optical signal; and determine a first distance between two points in an environment based on the first optical signal. The device includes a distance refinement system having one or more indirect time-of-flight sensors. The distance refinement system is configured to receive a second optical signal; and determine a second distance between the two points in the environment based on the second optical signal. The device includes a processing system having one or more computing components. The processing system is configured to receive information representative of the first distance and the second distance; and determine a third distance between the two points in the environment based on the first distance and the second distance. A difference between (i) a true distance between the two points in the environment and (ii) the first distance is greater than a difference between (i) the true distance between the two points in the environment and (ii) the third distance.
[0008] In some example aspects, the sensor system includes: a sensor array configured to receive the first optical signal and generate one or more first electrical signals; and one or more storage media configured to store one or more machine learning models trained to: receive, as input, a representation of the one or more first electrical signals; and provide an output representative of the first distance.
[0009] In some example aspects, the sensor array is a sensor array of an RGB camera, and the representation of the one or more first electrical signals includes at least a portion of an optical image output by the RGB camera.
[0010] In some example aspects, the one or more machine learning models include a convolutional neural network model.
[0011] In some example aspects, the sensor system includes a stereo camera having a plurality of lenses, and the first optical signal includes at least a portion of a three-dimensional image output by the stereo camera.
[0012] In some embodiments, the sensor system includes a structured light detection system.
[0013] In some embodiments, the sensor system includes one or more direct time-of-flight ranging sensors, and determining the first distance includes determining the first distance based on a round-trip time of a light ray between the two points in the environment.
[0014] In some embodiments, the one or more indirect time-of-flight ranging sensors are configured to operate at a first frequency, the sensor system includes one or more second indirect time-of-flight ranging sensors operating at a second frequency, and the second frequency is less than the first frequency.
[0015] In some embodiments, the sensor system includes an image sensor system and a time-of-flight ranging sensor system, and determining the first distance includes obtaining an output of the image sensor system, obtaining an output of the time-of-flight ranging sensor system, and determining the first distance based on the output of the image sensor system and the output of the time-of-flight ranging sensor system.
[0016] In some embodiments, determining the third distance between the two points in the environment includes determining a multiplier of a resolvable distance related to the one or more indirect time-of-flight ranging sensors based on the first distance, and adding the second distance to a product of the resolvable distance and the multiplier to determine the third distance.
[0017] In some embodiments, a first one of the two points in the environment represents a location of the device, and a second one of the two points in the environment represents a location of an object in the environment.
[0018] In some embodiments, the device is a mobile device, the sensor system includes one or more cameras embedded in the mobile device, and the one or more indirect time-of-flight ranging sensors include a three-dimensional sensor array embedded in the mobile device.
[0019] In some embodiments, the one or more cameras and the three-dimensional sensor array are embedded on a back side of the mobile device, opposite a screen of the mobile device.
[0020] In some embodiments, the one or more cameras and the three-dimensional sensor array are embedded on a front side of the mobile device, on a same side of a screen of the mobile device.
[0021] In some embodiments, the one or more operational components include one or more processors, one or more hardwired circuitry, one or more field programmable gate arrays, or a combination thereof.
[0022] In certain embodiments, a difference between (i) a true distance between two points in an environment and (ii) the first distance represents a first error, the first error being less than a resolvable distance associated with the one or more indirect time-of-flight ranging sensors; a value of the second distance is within the resolvable distance; and a difference between (i) the true distance between the two points in the environment and (ii) the third distance represents a second error, the second error being less than the first error.
[0023] Another example aspect of the application is a method for determining a distance between two points in an environment. The method includes receiving, by a sensor system, a first optical signal. The method includes determining, by the sensor system, a first distance between two points in an environment based on the first optical signal. The method includes receiving, by one or more indirect time-of-flight ranging sensors separate from the sensor system, a second optical signal. The method includes determining, by the one or more indirect time-of-flight ranging sensors, a second distance between the two points in the environment based on the second optical signal. The method includes receiving, by one or more processing components, information representative of the first distance and the second distance. The method includes determining, by the one or more processing components, a third distance between the two points in the environment based on the first distance and the second distance. A difference between (i) a true distance between the two points in the environment and (ii) the first distance is greater than a difference between (i) the true distance between the two points in the environment and (ii) the third distance.
[0024] In certain embodiments, an autonomous vehicle or a user device includes the sensor system, the one or more indirect time-of-flight ranging sensors, and the one or more processing components.
[0025] Yet another example aspect of the application is a device for measuring. The device includes a distance estimation system having a sensor system. The distance estimation system is configured to receive a first optical signal; and generate a first electrical signal based on the first optical signal to determine a first distance between two points in an environment. The device includes a distance refinement system having one or more indirect time-of-flight ranging sensors. The distance refinement system is configured to receive a second optical signal; and generate a second electrical signal based on the second optical signal to determine a second distance between the two points in the environment. The device includes a processing system having one or more processing components. The processing system is configured to receive the first electrical signal and the second electrical signal; provide input information representative of the first electrical signal and the second electrical signal to a machine learning model; receive output information representative of a third distance between the two points in the environment; and determine the third distance between the two points in the environment. A maximum value of the second distance is less than a resolvable distance associated with the one or more indirect time-of-flight ranging sensors. A maximum value of the third distance is greater than the resolvable distance associated with the one or more indirect time-of-flight ranging sensors.
[0026] Yet another example aspect of the application is a device for measurement. The device includes an estimation system having a sensor system. The estimation system is configured to receive a first signal; and determine, based on the first signal, a first value of a characteristic associated with a target object in an environment. The device includes a refinement system having one or more indirect time-of-flight sensors. The refinement system is configured to receive a second signal; and determine, based on the second signal, a second value of the characteristic associated with the target object in the environment. The device includes a processing system having one or more computational components. The processing system is configured to receive information representative of the first value and the second value; and determine, based on the first value and the second value, a third value of the characteristic associated with the target object in the environment. A difference between (i) a true value of the characteristic associated with the target object in the environment and (ii) the first value is greater than a difference between (i) the true value of the characteristic and (ii) the third value.
[0027] Other example aspects of the application include systems, methods, devices, sensors, computational components, tangible, non-transitory computer-readable media, and memory components relating to the technology.
[0028] The above and other features, aspects, and advantages of various embodiments are described in more detail with reference to the claims. Various embodiments are described in more detail with reference to the drawings, which are intended to illustrate but not to limit the application. BRIEF DESCRIPTION OF DRAWINGS
[0029] The above and other aspects and advantages of the application are more fully understood in conjunction with the following detailed description, when considered in connection with the accompanying drawings, in which:
[0030] Figure 1 A block diagram of an example depth sensing system is depicted in accordance with an example aspect of the application.
[0031] Figure 2 A block diagram of an example distance estimation system is depicted in accordance with an example aspect of the application.
[0032] Figure 3 A block diagram of an example sensor system is depicted in accordance with an example aspect of the application.
[0033] Figure 4 An example system is depicted in accordance with an example aspect of the application.
[0034] Figure 5 An example operation of a depth sensing system is depicted in accordance with an example aspect of the application.
[0035] Figure 6 A flow diagram of an example procedure is depicted in accordance with an example aspect of the application.
[0036] Figure 7FIG. 1 illustrates a block diagram of an example depth sensing system, in accordance with an example aspect of the present disclosure.
[0037] Figure 8 FIG. 2 illustrates a flow diagram of an example process, in accordance with an example aspect of the present disclosure.
[0038] Figure 9 FIG. 3 illustrates an example system, in accordance with an example aspect of the present disclosure.
[0039] Figure 10 FIG. 4 illustrates an example system, in accordance with an example aspect of the present disclosure.
[0040] Figure 11 FIG. 5 illustrates a cross-sectional view of an example portion of an example light detector, in accordance with an example aspect of the present disclosure.
[0041] Figure 12 FIG. 6 illustrates example computing system components and components, in accordance with an example aspect of the present disclosure.
[0042] Legend of Figures
[0043] 100: depth sensing system 102: component
[0044] 110: distance estimation system 112: sensor system
[0045] 114: depth mapping system 120: distance refinement system
[0046] 122: i-ToF system 124: depth mapping system
[0047] 130: processing system 140: object
[0048] 200: distance estimation system 202: RGB camera
[0049] 204: machine learning model 212: stereo camera
[0050] 222: structured light system 232: d-ToF sensor
[0051] 242: i-ToF sensor 252: motion estimator
[0052] 254: machine learning model 302: transmitter unit
[0053] 304: receiver unit 306: controller
[0054] 400: system 602, 604, 606, 608, 610: step
[0055] 700: depth sensing system 702: component
[0056] 710: distance estimation system 712: sensor system
[0057] 720: distance refinement system 730: processing system
[0058] 734: machine learning model 736: depth mapping system
[0059] 802, 804, 806, 808, 810: steps 900: system
[0060] 910: estimation system 914: processing system
[0061] 920: refinement system 924: processing system
[0062] 930: processing system 940: object
[0063] 1000: system 1010: estimation system
[0064] 1020: refinement system 1030: processing system
[0065] 1040: object 1100: germanium-on-silicon platform
[0066] 1101: first wafer 1102: second wafer
[0067] 1103: wafer bonding interface 1104: differential demodulation clock
[0068] 1105: first node 1106: second node
[0069] 1200: system 1202: computing system
[0070] 1204: computing component 1206: processor
[0071] 1208: memory 1210: data
[0072] 1212: instructions 1214: communication interface
[0073] 1216: machine learning model 1232: machine learning computing system
[0074] 1234: processor 1236: memory
[0075] 1238: data 1240: instructions
[0076] 1242: machine learning model 1244: model trainer
[0077] 1246: training data 1248: communication interface
[0078] 1250: network D1 : first distance
[0079] D2: second distance DETAILED DESCRIPTION
[0080] This application claims priority to U.S. Provisional Patent Application No. 63 / 065,482, filed August 13, 2020, the entire contents of which are incorporated herein by reference.
[0081] Example aspects of the present application are improved systems and methods for determining a distance between two points in an environment. For example, a depth sensing system can utilize distance estimates from a first type of sensor (e.g., an RGB camera) and refined distance estimates from a second type of sensor (e.g., an indirect time-of-flight sensor) to determine more accurate estimates of a distance between a component (e.g., a user device, an automated platform, etc.) and an object in an environment in which the component is located. In some implementations, the depth sensing system can utilize one or more trained machine learning models. For example, a machine learning model can receive distance estimates as its input and output, in response thereto, more accurate estimates of a distance between the component and the object, as will be described in more detail below.
[0082] The systems and methods of the present application provide a variety of technical benefits. For example, the techniques of the present application provide improved depth and distance estimates that assist systems that employ such techniques to improve functionality, including improved prediction and motion control for automated platforms, application operation for user devices, etc. In addition, the systems and methods of the present application can detect and correct for aliasing errors by comparing different distance estimates from different types of sensors. As such, the techniques of the present application can improve three-dimensional sensor accuracy and thus assist systems that employ such techniques to improve functionality.
[0083] Example embodiments of the present application will be described in more detail below with reference to the accompanying drawings. It should be noted that embodiments, features, hardware, software, and / or other components described with reference to one figure can also be used in systems and procedures shown in another figure.
[0084] Figure 1 A depth sensing system 100 is shown that is used to determine a distance between a component 102 and an object 140 in an environment. The component 102 can be a mobile device (e.g., a smartphone, a tablet computer, a wearable device, a vehicle, a drone, etc.) or a stationary device (e.g., a surveillance system, a robotic arm, etc.). The component 102 includes a distance estimation system 110, a distance refinement system 120, and a processing system 130.
[0085] The distance estimation system 110 can be configured to receive a first optical signal and determine a first distance between two points in an environment based on the first optical signal. For example, the distance estimation system 110 can be configured to receive an optical signal from the object 140 and determine a distance between the component 102 and the object 140 based on the optical signal. The distance estimation system 110 can include a sensor system 112 and a depth mapping system 114. Generally, the sensor system 112 receives an optical signal (e.g., reflected light) from the object 140 and generates one or more electrical signals. The sensor system 112 can include a sensor array configured to receive the first optical signal and generate one or more first electrical signals. The sensor array can be a sensor array of an RGB camera. In some implementations, the sensor system 112 can include a stereo camera with multiple lenses, a structured light detection system, and / or one or more direct time-of-flight (d-ToF) sensors.
[0086] For example, referring to the example of the distance estimation system 200 of Figure 2 , implementations of the sensor system 112 can employ an RGB camera 202 (or an IR camera) coupled with a machine learning model 204, or a stereo camera 212, or a structured light system 222, or a d-ToF sensor 232, or an i-ToF sensor 242, or a motion estimator 252 coupled with a machine learning model 254 (e.g., to perform depth estimation based on motion estimates determined from multiple images captured by one or more image sensors over time), or a sonar sensor, or a radar sensor, or a position sensor (e.g., an accelerometer, a gyroscope, NFC, GPS, etc.), or a combination of any of the above. The depth mapping system 114 receives the one or more electrical signals from the sensor system 112 and determines the distance between the component 102 and the object 140. For example, when one or more direct time-of-flight sensors are used, a first distance can be determined based on a round-trip time of light between two points in the environment (e.g., the distance between the component 102 and the object 140). Implementations of the depth mapping system 114 can employ hardware circuitry (e.g., ASICs, FPGAs), lookup tables (LUTs), one or more processors with software (e.g., MCUs, CPUs, GPUs, TPUs), or any other suitable approach.
[0087] Referring again to Figure 1 , the distance refinement system 120 can be configured to receive a second optical signal and determine a second distance between two points in the environment based on the second optical signal. For example, in the example of the distance estimation system 200 of Figure 1In some embodiments, the distance refinement system 120 is configured to receive an optical signal and determine a distance between the component 102 and the object 140 based on the optical signal. The distance refinement system 120 can include one or more i-ToF systems 122 and a depth mapping system 124. Implementations of the depth mapping system 124 can also employ hardware circuitry (e.g., ASICs, FPGAs), lookup tables (LUTs), one or more processors with software (e.g., MCUs, CPUs, GPUs, TPUs), or any other suitable means.
[0088] Referring to Figure 3 In some embodiments, the one or more i-ToF systems 122 can include a transmitter unit 302, a receiver unit 304, and a controller 306. In operation, the transmitter unit 302 can emit a transmitted light toward a target object (e.g., the object 140). The receiver unit 304 can receive a reflected light reflected from the target object. The controller 306 can drive at least the transmitter unit 302 and the receiver unit 304. In some embodiments, the one or more i-ToF systems 122 can be configured to operate at a first frequency. The sensor system 112 can include one or more second indirect time-of-flight sensors operating at a second frequency, which is different from the first frequency.
[0089] The transmitter unit 302 can include one or more light sources with a peak wavelength of the emitted light in the non-visible light wavelength range above 800 nm, such as 850 nm, 940 nm, 1050 nm, 1064 nm, 1310 nm, 1350 nm, or 1550 nm.
[0090] The receiver unit 304 can include one or more i-ToF sensors, such as a pixel array. The material of the i-ToF sensor can be a III-V semiconductor material (e.g., GaAs / AlAs, InP / InGaAs / InAlAs, GaSb / InAs, or InSb), a semiconductor material containing a group IV element (e.g., Ge, Si, or Sn), or a compound such as SixGeySn1-x-y (where 0≤x≤1, 0≤y≤1, x+y≤1), Ge1-aSna (where 0≤a≤1), or Ge1-xSix (where 0≤x≤1).
[0091] For example, a pixel array can be implemented on a Ge-on-Si platform. Figure 11 An example pixel array cross-section on a Ge-on-Si platform 1100 (and clock signals suitable for the pixel array) is shown. Figure 11An example pixel array employs a silicon-on-germanium architecture that can absorb wavelengths in the near-infrared (NIR, e.g., wavelength range from 780 nm to 1400 nm, or any similar wavelength range defined by a particular application) and short-wave infrared (SWIR, e.g., wavelength range from 1400 nm to 3000 nm, or any similar wavelength range defined by a particular application) spectrum. This enables better signal-to-noise ratio (SNR) without exceeding the maximum permissible exposure (MPE) limit.
[0092] The silicon-on-germanium platform 1100 can be associated with an i-ToF image sensor. For example, the i-ToF image sensor can employ a back-side illumination (BSI) configuration in which a germanium region is formed on a first wafer 1101 (e.g., top wafer) and one or more circuits are located on a second wafer 1102 (e.g., bottom wafer). The first wafer 1101 and the second wafer 1102 can be bonded to each other via a wafer bonding interface 1103. In some embodiments, the pixels can be implemented in a double-switched pinned pixel architecture. One or more differential demodulation clocks 1104 (e.g., CLKP, CLKN) can be distributed across the first wafer 1101 to create a continuous switching lateral electric field between first nodes 1105 (e.g., Demodl, Demod2) in each pixel, for example, on the germanium surface (e.g., the side closer to the via, VIA). Photocharges can be collected via second nodes 1106 (e.g., FD1, FD2). In some embodiments, because most photocharges are generated within the germanium layer, and the germanium layer is thin, the lateral electric field on the germanium surface can efficiently sweep the photocharges to the second nodes 1106. Moreover, because the germanium layer is not thick, the transition time for the photocharges to drift to the one or more second nodes 1106 (e.g., FD1 and / or FD2) is not long, thus enabling a large improvement in demodulation speed. In some embodiments, to avoid coupling to any sensitive high-impedance nodes and to relax design rule requirements, the second nodes 1106 (e.g., FD1 and / or FD2) can be made to interact with the wafer bonding interface that partially covers the pixel area. The one or more differential demodulation clocks 1104 (e.g., CLKP and / or CLKN) can be routed to the second wafer 1102 clock drivers outside the pixel area. The pixel demodulation drivers can be implemented in a tapered inverter chain, and maximum performance can be achieved by adjusting the supply of the inverter chain. In some embodiments, the pixel circuit implementation can employ a differential four-transistor architecture. Figure 11The bottom also shows a simplified timing diagram. Before each exposure, all pixels can be reset via Msh1 / Msh2 and Mrt1 / Mrt2 controlled by signal RST. After exposure, integration, and demodulation, the collected photocharges can be stored on C1 and C2 controlled by signal SH. Finally, via source follower Msf1 / Msf2 and column select switch Mbt1 / Mbt2 controlled by signal BS, the data can be read out to an ADC. In some embodiments, four-phase step measurement can be used to recover depth information without worrying about analog non-idealities.
[0093] Referring again to Figure 3 In some embodiments, the controller 306 includes a timing generator and a processing unit. The timing generator receives a reference clock signal and provides timing signals to the transmitter unit 302 for modulating the emitted light. The timing signals are also sent to the receiver unit 304 for controlling the collection of the optical carriers. The processing unit processes the optical carriers generated and collected by the receiver unit 304 and then determines the raw data of the target object. The processing unit can include control circuitry, one or more signal processors for processing the information output by the receiver unit 304, and / or a computer storage medium that can store instructions for determining or storing the raw data of the target object. In general, the controller 306 determines a distance between two points using the phase difference between the light emitted by the transmitter unit 302 and the light received by the receiver unit 304.
[0094] In some cases, the receiver unit 304 and the controller 306 are implemented on the same semiconductor chip, such as a system-on-chip (SoC). In some cases, the transmitter unit 302 is implemented by two different semiconductor chips, such as a laser emitter chip on a III-V substrate and a silicon laser driver chip on a silicon substrate.
[0095] Figure 4An example system 400 is shown, in which the sensor system 112 includes an RGB camera 202 and a machine learning model 204. The RGB camera 202 can be any suitable camera that can optically capture an image of an environment and convert it into an electrical signal (e.g., a CMOS digital camera). One or more storage media, which can be configured to store one or more machine learning models. One or more machine learning models, which can be trained to receive a representation of the one or more first electrical signals as its input, and provide an output representative of the first distance. For example, the representation of the one or more first electrical signals can include at least a portion of an optical image output by the RGB camera 202. The machine learning model 204 can be, for example, trained to receive as its input a portion of an image captured by the RGB camera 202 (e.g., a portion of the image that includes the object 140 in the environment), and produce an output representative of the distance between the component 102 and the object 140. Meanwhile, or as an alternative, the sensor system 112 can include a stereo camera. The one or more first electrical signals can include at least a portion of a three-dimensional image output by the stereo camera. In some embodiments, meanwhile, or as an alternative, the machine learning model 204 is trained to receive as its input at least a portion of a three-dimensional image output by the stereo camera, and produce an output representative of the distance between the component 102 and the object 140. In some implementations, the machine learning model 204 can be a convolutional neural network model, a deep neural network model, or any other suitable machine learning model. Generally, the distance, or depth, determined by the distance estimation system 110 can provide a good estimate. However, due to the training of the machine learning model 204 based on empirical data (e.g., a set of training images with distance information), and not necessarily accurately fitting all possible operating scenarios, there can be a higher error and / or lower depth resolution.
[0096] Conversely, the depth resolution of an i-ToF sensor depends on the demodulation frequency of the i-ToF sensor. When the demodulation frequency is high (e.g., 300 MHz), the i-ToF sensor can generally produce a result with high depth resolution and low error. However, the i-ToF sensor is limited by its unambiguous range, which is defined as:
[0097]
[0098] where c is the speed of light, and fdemod is the demodulation frequency. For example, when the demodulation frequency is 300 MHz, the unambiguous range of the i-ToF sensor is 50 cm. Thus, when operating at 300 MHz, the i-ToF sensor can not be able to distinguish between a 10 cm distance and a 60 cm distance, and thus can have an ambiguous error.
[0099] Because the distance estimation system 110 can provide a distance estimate with a large error but a wide range of distances between the component 102 and the object 140, and the distance refinement system 120 can provide a distance estimate with a low error within a resolvable distance, the processing system 130 can combine distance data from both the distance estimation system 110 and the distance refinement system 120 to provide a distance estimate between the component 102 and the object 140 with a high accuracy and a wide range. Implementations of the processing system 130 can employ hardware circuitry (e.g., ASICs, FPGAs), lookup tables, one or more processors with software (e.g., MCUs, CPUs, GPUs, TPUs), or any other suitable means.
[0100] Figure 5 An example operation of a depth sensing system is illustrated. Here, the distance estimation system 110 (e.g., implemented with an RGB camera connected with a machine learning model) determines a distance between the component 102 and the object 140 as Dl. Further, the distance refinement system 120 determines a distance between the component 102 and the object 140 as D2. The processing system 130 can include one or more computing components (e.g., one or more processors, one or more hardwired circuitry, one or more field programmable gate arrays, other components, or a combination thereof). The processing system 130 can be configured to receive information representative of the first distance Dl and the second distance D2, and determine a third distance between two points in an environment based on the first distance Dl and the second distance D2. A first point between the two points in the environment can represent, for example, a location of a device (e.g., the component 102). A second point between the two points in the environment can represent a location of the object 140 in the environment. Based on the first distance Dl, a multiplier of a resolvable distance associated with one or more indirect time-of-flight sensors can be determined. For example, a third distance can be determined by adding the second distance to a product of the resolvable distance and the multiplier. A difference (e.g., an absolute difference) between (i) a true distance (e.g., an actual distance) between the two points in the environment and (ii) the first distance can be greater than (e.g., statistically greater than) a difference between (i) the true distance between the two points in the environment and (ii) the third distance. Further, in some implementations, the difference between (i) the true distance between the two points in the environment and (ii) the first distance Dl can represent a first error, which is less than a resolvable distance associated with the one or more indirect time-of-flight sensors. The second distance D2 has a value within the resolvable distance. The difference between (i) the true distance between the two points in the environment and (ii) the third distance can represent a second error, which is less than (e.g., statistically less than) the first error.
[0101] For example, based on Dl and a known resolvable distance of the distance refinement system 120 (based on an operating frequency of the i-ToF sensor shown in equation (1)), the processing system 130 can determine a long-range and high-depth resolution distance Dout between the component 102 and the object 140:
[0102] (2)
[0103] (3)D out = D2+ N x (Unambiguous Range)
[0104] where N is the nearest multiple of the unambiguous range between the component 102 and the object 140, and floor() is a function that calculates the largest integer less than or equal to the division quotient between D1 and the unambiguous range. As long as the distance error from the range estimation system 110 is less than the unambiguous range associated with the range refinement system 120, the distance Dout can be considered accurate.
[0105] In some embodiments, the aliasing problem of i-ToF sensors can be solved by alternating the use of two or more different frequencies (e.g., 75 MHz and 50 MHz) to operate the i-ToF sensor over time. However, even with multi-frequency operation, aliasing errors can still occur, especially when the system SNR is low. For example, assume that the unambiguous ranges are 2 meters and 3 meters at 75 MHz and 50 MHz, respectively. The corresponding unambiguous range for aliasing of the i-ToF sensor operating at these two frequencies is 6 meters. If the SNR is high, the difference in the measured distances at 75 MHz and 50 MHz is small, e.g., if the true value of the reference is 1 meter for the first measurement, the distances measured by the i-ToF system 122 at 75 MHz and 50 MHz can be 0.9 meters and 1.2 meters, respectively, and the distance calculated by the depth mapping system 124 can be 0.9 meters or 1.2 meters (depending on the selected frequency data). However, if the SNR is low, the difference in the measured distances at 75 MHz and 50 MHz can be large, e.g., if the second measurement of the 1 meter reference is done in a noisy environment, the distances measured by the i-ToF system 122 at 75 MHz and 50 MHz can be 0.7 meters and 1.4 meters, respectively, and the distance calculated by the depth mapping system 124 in this case can be 4.7 meters or 4.4 meters (depending on the selected frequency data), resulting in an aliasing error.
[0106] To resolve the aliasing error, the processing system 130 can use the distance determined by the distance estimation system 110 as a reference distance. Using the example above, the processing system 130 can receive a distance measurement of 1 meter (e.g., Dl) from the distance estimation system 110. If the processing system 130 then receives distance measurements of 4.7 meters and / or 4.4 meters (e.g., D2) from the distance refinement system 120, the processing system 130 can determine that an aliasing error has occurred by comparing Dl to D2. If the processing system 130 determines that an aliasing error has occurred, the processing system 130 can adjust the distance measurements from the distance refinement system 120 using an integer multiple of the resolvable distance (e.g., 2 x 2 meters or 1 x 3 meters, depending on whether 75 MHz or 50 MHz is selected). Accordingly, the depth sensing system can detect and correct for aliasing errors.
[0107] Figure 6 An example procedure for determining a long-range and high-depth resolution distance using a depth sensing system (e.g., the depth sensing system 100) is shown.
[0108] The depth sensing system receives a first optical signal (step 602). In particular, the first optical signal is received by a sensor system (e.g., of the depth sensing system). For example, the RGB camera 202 of the distance estimation system 110 can capture an image of an object in an environment. The depth sensing system determines a first distance between two points in the environment (step 604). In particular, the first distance between the two points in the environment is determined by the sensor system based on the first optical signal. For example, the machine learning model 204 of the distance estimation system 110 can use a portion of the image captured by the RGB camera 202 to generate an output representing a distance between the component 102 and the object 140. The depth mapping system 114 can use the output of the machine learning model 204 to determine a distance between the component 102 and the object 140, which is a long-range distance with a large error.
[0109] The depth sensing system receives a second optical signal (step 606). In particular, the second optical signal is received by one or more indirect time-of-flight sensors separate from the sensor system. For example, the one or more i-ToF sensors of the distance refinement system 120 can detect light reflected from an object in an environment. The depth sensing system determines a second distance between two points in the environment (step 608). In particular, the second distance between the two points in the environment is determined by the one or more indirect time-of-flight sensors based on the second optical signal. For example, the depth mapping system 124 of the distance refinement system 120 can use a phase difference between an emitted light and a reflected light to determine a distance within a resolvable distance.
[0110] The depth sensing system determines a third distance between two points in the environment (step 610). Specifically, one or more computing components of the depth sensing system receive information representative of the first distance and the second distance, and determine, based on the first distance and the second distance, a third distance between two points in the environment. A difference (e.g., an absolute difference) between (i) a true distance between two points in the environment and (ii) the third distance can be greater than (e.g., statistically greater than) a difference between (i) the true distance between two points in the environment and (ii) the first distance. For example, the processing system 130 can determine a long-range and high-depth resolution distance between the component 102 and the object 140 using the first distance determined by the distance estimation system 110, the resolvable distance associated with the operating frequency of the i-ToF sensor, and the second distance determined by the distance refinement system 120 (e.g., refer to Figure 5 the described Dout).
[0111] Figure 7 A depth sensing system 700 is shown that utilizes a machine learning model trained with sensor fusion to determine a distance between a component 702 and an object 140 in an environment. The component 702 can be a mobile device (e.g., a smartphone, a tablet, a wearable device, a vehicle, a drone, etc.) or a stationary device (e.g., a surveillance system, a robotic arm, etc.). The component 702 includes a distance estimation system 710, a distance refinement system 720, and a processing system 730.
[0112] The distance estimation system 710 is configured to receive an optical signal and provide an electrical signal (e.g., a digital image) based on the optical signal that can be used to estimate a distance between the component 702 and the object 140. The distance estimation system 710 can include a sensor system 712. An implementation of the sensor system 712 can employ an RGB camera 202. An implementation of the sensor system 712 can also employ a stereo camera 212, or a structured light system 222, or a d-ToF sensor 232, or an i-ToF sensor 242, or a motion estimator 252, or a combination of any of the above sensors. At this time, the distance estimation system 710 can be configured to receive a first optical signal and generate a first electrical signal based on the first optical signal to determine a first distance between two points in the environment.
[0113] The distance refinement system 720 is configured to receive an optical signal and provide an electrical signal (e.g., quadrature amplitude) based on the optical signal that can be used to determine a distance between the component 702 and the object 140. The distance refinement system 720 can include one or more i-ToF systems 122, at which time the distance refinement system 720 can be configured to receive a second optical signal and generate a second electrical signal based on the second optical signal to determine a second distance between two points in the environment.
[0114] The processing system 730 can combine electrical signals from both the distance estimation system 710 and the distance refinement system 720 to determine a distance between the component 702 and the object 140 with a wider range and higher accuracy. The processing system 730 can include a machine learning model 734 and a depth mapping system 736. Implementations of the processing system 730 can employ hardware circuitry (e.g., ASICs, FPGAs), lookup tables, one or more processors with software (e.g., MCUs, CPUs, GPUs, TPUs), or any other suitable means.
[0115] The processing system 730 (e.g., including one or more operational components) can be configured to receive a first electrical signal and a second electrical signal, provide input information representing the first electrical signal and the second electrical signal to a machine learning model, receive output information representing a third distance between two points in an environment, and determine the third distance between the two points in the environment. A maximum value of the second distance can be less than a resolvable distance associated with the one or more indirect time-of-flight ranging sensors, and a maximum value of the third distance can be greater than the resolvable distance associated with the one or more indirect time-of-flight ranging sensors.
[0116] For example, the machine learning model 734 can be trained to receive as a combined input a signal from the distance estimation system 710 (e.g., a portion of a digital image including the object 140 in the environment captured by the RGB camera 202) and a signal from the distance refinement system 720 (e.g., a portion of quadrature amplitude including the object 140 in the environment captured by the i-ToF system 122) and generate an output representing a distance between the component 702 and the object 140. In some implementations, the machine learning model 734 can be a convolutional neural network model, a deep neural network model, or any other suitable machine learning model. For example, the machine learning model 734 can be trained to output distance data using RGB image data of an object and corresponding phase data generated by an i-ToF sensor at a particular operating frequency (e.g., 300 MHz). Because the sensors of the distance estimation system 710 (e.g., the RGB camera) can provide a higher error but a wider range of estimates of a distance between the component 702 and the object 140, and the sensors of the distance refinement system 720 (e.g., the i-ToF sensor) can provide lower error distance estimates within a resolvable distance, the machine learning model 734 can be trained to combine information from both the distance estimation system 710 and the distance refinement system 720 to determine a distance between the component 702 and the object 140 with a higher accuracy and a wider range.
[0117] The depth mapping system 736 receives the output of the machine learning model 734 and determines the distance between the component 702 and the object 140. The depth mapping system 736 can be implemented using hardware circuitry (e.g., ASICs, FPGAs), lookup tables, one or more processors with software (e.g., MCUs, CPUs, GPUs, TPUs), or any other suitable manner. In some implementations, the depth mapping system 736 can be incorporated into the machine learning model 734 (e.g., as one or more output layers of the machine learning model 734).
[0118] Figure 8 An example procedure for displaying a remote and high depth resolution distance using a depth sensing system (e.g., the depth sensing system 700) is described.
[0119] The depth sensing system receives a first optical signal (step 802). For example, an RGB camera of the distance estimation system 710 can capture an image of an object in an environment. The depth sensing system provides a first electrical signal representative of the environment (step 804). For example, the distance estimation system 710 can provide information representative of at least a portion of the image and / or any associated metadata to the processing system 730.
[0120] The depth sensing system receives a second optical signal (step 806). For example, one or more i-ToF sensors of the distance refinement system 720 can detect light reflected from an object in an environment. The depth sensing system provides a second electrical signal representative of the environment (step 808). For example, the i-ToF system 122 of the distance refinement system 720 can output information representative of a phase difference between a transmitted light of the component 702 and a reflected light of the object 140 to the processing system 730.
[0121] The depth sensing system determines a distance between two points (step 810). For example, the processing system 730 can receive (i) information from the distance estimation system 710 representative of at least a portion of the image and / or any associated metadata, and (ii) information from the distance refinement system 720 representative of a phase difference between a transmitted light of the component 702 and a reflected light of the object 140. The processing system 730 can fuse the received information to generate an input (e.g., one or more multi-dimensional vectors) to the machine learning model 734. The machine learning model 734 can utilize the input to generate an output representative of a distance between the component 702 and the object 140. The depth mapping system 736 can utilize the output from the machine learning model 734 to determine the distance between the component 702 and the object 140 (e.g., refer to Doutdescribed above). Figure 5
[0122] Figure 9 A system 900 is shown for determining one or more properties of an object 140 in an environment. In general terms, the system 900 includes an estimation system 910, a refinement system 920, and a processing system 930. The estimation system 910 is configured to utilize the sensor system 112 to determine an output representative of an estimate of one or more properties of an object 140 in an environment and determined by the processing system 914. The refinement system 920 is configured to utilize the i-ToF system 122 to determine an output representative of a refinement of one or more properties of the object 140 and determined by the processing system 924. The processing system 930 is configured to receive outputs from the estimation system 910 and the refinement system 920 to determine a more accurate value of the one or more properties (e.g., a value with higher certainty, a value with a lower margin of error, etc.). Example properties associated with the object 140 can include a distance between the object 140 and the assembly 102, an edge of the object 140, a recognition of the object 140, a classification of the object 140, a motion of the object 140, and any other application-specific property.
[0123] For example, Figure 1 The depth sensing system 100 of FIG. 1 can be an example of the system 900, where the distance estimation system 110 can be an example of the estimation system 910, the distance refinement system 120 can be an example of the refinement system 920, and the processing system 130 can be an example of the processing system 930.
[0124] In another example, the system 900 can be used to distinguish between the object 140 and another object 940 that partially overlaps with the object 140. In this example, the estimation system 910 can utilize an RGB camera coupled to a machine learning model to determine an output representative of an edge of the object 140. Since the object 940 partially overlaps with the object 140 in the RGB image, the estimation system 910 can assign an uncertainty value to each pixel in a group of pixels adjacent to the object 140, where the uncertainty value can represent a likelihood that the corresponding pixel belongs to the object 140. The refinement system 920 can utilize the i-ToF system 122 to determine a relative depth difference between the object 140 and the object 940. The processing system 930 can utilize outputs from both the estimation system 910 and the refinement system 920 to determine whether the adjacent pixels are part of the object 140.
[0125] Figure 10Another system 1000 is shown for determining one or more properties of an object 140 in an environment. Generally, the system 1000 includes an estimation system 1010, a refinement system 1020, and a processing system 1030. The estimation system 1010 is configured to utilize a sensor system 112 to provide an output that can be used to determine an estimate of one or more properties of an object 140 in an environment. The refinement system 1020 is configured to utilize an i-ToF system 122 to provide an output that can be used to determine a refined result of one or more properties of the object 140. The processing system 1030 is configured to receive outputs from both the estimation system 1010 and the refinement system 1020 to determine a value of the one or more properties. For example, Figure 7 The depth sensing system 700 of FIG. 7 can be an example of the system 1000, where the distance estimation system 710 can be an example of the estimation system 1010, the distance refinement system 720 can be an example of the refinement system 1020, and the processing system 730 can be an example of the processing system 1030. In some embodiments, the object 140 partially overlaps with another object 1040 thereof.
[0126] In this manner, Figure 9 and Figure 10 The system 900 and / or 1000 can include an estimation system 910, 1010 (e.g., including a sensor system 112) configured to receive a first signal and determine a first value of a property associated with a target object in an environment based on the first signal. The system 910, 1010 can include a refinement system 920, 1020 (e.g., including one or more indirect time-of-flight sensors) configured to receive a second signal and determine a second value of the property associated with the target object in the environment based on the second signal. The system 910, 1010 can include a processing system 930, 1030 (e.g., including one or more arithmetic components) configured to receive information representative of the first value and the second value and determine a third value of the property associated with the target object in the environment based on the first value and the second value. A difference (e.g., an absolute difference) between (i) a true value of the property associated with the target object in the environment and (ii) the first value can be greater than (e.g., statistically greater than) a difference between (i) the true value of the property and (ii) the third value.
[0127] Example implementations of the systems and methods of the present application are provided below. These implementations are for illustrative purposes only and are not limiting.
[0128] In some implementations, an automated platform can utilize the systems and methods of the present application to improve its functionality and operation. The automated platform can include any system or perform any of the methods / procedures described herein and shown in the figures.
[0129] In some embodiments, the automated platform can include an autonomous vehicle. The autonomous vehicle can include an on-board computing system that functions to autonomously perceive the environment in which the vehicle is located and control the actions of the vehicle in the environment. The autonomous vehicle can include a depth sensing system described herein that aids in improving the perception capabilities of the autonomous vehicle. For example, the depth sensing system can aid the autonomous vehicle in more accurately classifying objects and predicting their behavior, e.g., by the vehicle computing system. This functionality can be used to distinguish between a first object (e.g., a first vehicle) and a second object (e.g., a second vehicle) that at least partially overlaps with the first object in the environment surrounding the autonomous vehicle. An estimation system of the depth sensing system can utilize a first sensor (e.g., an RGB camera in some implementations that is coupled to a machine learning model) to determine an output representing an edge of the first object. Because the second object at least partially overlaps with the first object (e.g., in the RGB image), the estimation system can assign an uncertainty value (e.g., a likelihood that the pixel belongs to the first object) to each pixel in a group of pixels that are adjacent to the object. A refinement system can utilize a second sensor system (e.g., an i-ToF system) to determine a relative depth difference between the first object and the second object. A processing system can utilize outputs from both the estimation system and the refinement system to determine whether the adjacent pixels belong to the first object. This can aid the autonomous vehicle in determining a boundary of the first object and / or the second object.
[0130] The autonomous vehicle can classify the first object (e.g., as a vehicle, etc.) based on the depth difference, a determined distance of the first object, and / or a determination that the pixel belongs to the first object. For example, this functionality can assist the autonomous vehicle in better understanding the shape of the object and, thus, semantically labeling the object as a particular type. Further, by determining the identity and / or type of the object, the autonomous vehicle can be better able to predict the actions of the object. For example, the autonomous vehicle can improve the accuracy of a prediction of a trajectory of actions of the first object based on the type of the object. In some implementations, the autonomous vehicle can determine that the object is a dynamic object (e.g., a vehicle) that is likely to move in the surrounding environment and / or a static object (e.g., a light pole) that is likely to remain stationary in the surrounding environment. In some implementations, the autonomous vehicle can determine a particular manner of movement of the object based on the type of the object. For example, a car is more likely to travel within a lane boundary of a road than a bicycle.
[0131] The automated platform (e.g., autonomous vehicle) can, for example, determine a depth difference between the first object and the second object based on, for example, as described herein. Figure 6 and Figure 8The programs described can determine more accurate characteristics (e.g., distance between an object and an autonomous vehicle). The determined distances can be used to more accurately determine the location of the automated platform in its surrounding environment. For example, an autonomous vehicle can determine a third distance between two points based on a first distance determined by an estimation system and a second distance determined by a refinement system. The two points can include the automated platform and an object that can be used to confirm the location. For example, an autonomous vehicle can determine its location in an environment based on a landmark. The autonomous vehicle can determine a more accurate distance between itself and the landmark (e.g., a building, a statue, an intersection, etc.). This can assist the autonomous vehicle (or a remote system) in more accurately determining the location of the autonomous vehicle and planning its actions in the environment.
[0132] In some embodiments, the automated platform can include an aerial automated platform, such as a drone. The drone can include a depth sensing system as described herein and as shown in the figures. The depth sensing system can improve the operation and functionality of the drone, including, for example, the precision and accuracy of the drone landing in a landing area or identifying an imaging area. For example, a drone can receive a first optical signal via a distance estimation system and determine a first distance between the drone and a landing / imaging area in an environment based on the first optical signal. The drone can receive a second optical signal via a distance refinement system and determine a second distance between the drone and the landing / imaging area based on the second optical signal. The drone can determine a third distance between the drone and the landing / imaging area based on the first distance and the second distance via a processing system. The third distance more accurately determines the distance between the drone and the landing / imaging area than the first distance. This can allow the drone to more accurately land and / or focus its imaging sensors (e.g., on-board cameras) on the area to be imaged.
[0133] In another example, the automation platform can relate to manufacturing applications and / or medical applications. For example, the automation platform can include a robotic arm. The robotic arm can be an electromechanical arm configured to assist in assembling at least a portion of a manufacturing item (e.g., a hardware computing architecture including a plurality of electronic components). The robotic arm can include a depth sensing system as described herein and as shown in the figures. With the depth sensing system, the robotic arm can more accurately position a component (e.g., a microprocessor) to be installed in an item (e.g., a circuit board). For example, the robotic arm (e.g., in association with a computing system) can receive a first optical signal via a distance estimation system and determine a first distance between the robotic arm (or a portion thereof) and a placement location based on the first optical signal. The robotic arm can receive a second optical signal via a distance refinement system and determine a second distance between the robotic arm (or a portion thereof) and the placement location based on the second optical signal. The robotic arm can determine a third distance between the robotic arm and the placement location based on the first distance and the second distance via a processing system. The third distance more accurately determines the distance between the robotic arm and the placement location than the first distance. This can enable the robotic arm to assemble the manufacturing item with better efficiency and accuracy.
[0134] In some embodiments, a user device can improve its functionality and operation with the systems and methods of the present application. The user device can be, for example, a mobile device, a tablet computer, a wearable headset, etc. The user device can include a depth sensing system and perform the methods / procedures described herein and as shown in the figures.
[0135] For example, a user device (e.g., a mobile device) can include one or more sensor systems. The sensor systems can include a first sensor (e.g., an RGB camera) and / or a second sensor (e.g., an i-ToF sensor). For example, the sensor systems can include one or more cameras embedded in the mobile device. At the same time, or as an alternative, the sensor systems can include one or more indirect time-of-flight sensors including a three-dimensional sensor array embedded in the mobile device. In some embodiments, the one or more cameras and the three-dimensional sensor array can be embedded on a back side of the mobile device, i.e., on a side opposite to a screen of the mobile device. At the same time, or as an alternative, the one or more cameras and the three-dimensional sensor array can be embedded on a front side of the mobile device, i.e., on a side same as the screen of the mobile device.
[0136] Depth sensing systems can utilize sensor systems to improve the operation and functionality of user devices, including imaging of objects, launching / operation of software applications, access to user devices, etc. For example, a user device can receive a first optical signal via a distance estimation system and determine a first distance between the user device and an object based on the first optical signal. The user device can receive a second optical signal via a distance refinement system and determine a second distance between the user device and an object based on the second optical signal. The user can determine a third distance between the user device and the object based on the first distance and the second distance via a processing system, as described herein. The user device can perform one or more functions or operations based on the third distance. For example, the user device can perform an access function (e.g., unlock a mobile phone) based on the distance of an authorized user's face or a gesture made by an authorized user. In another example, the user device can cause a camera of the user device to focus based on a determined distance between the user device and / or an object to be imaged (e.g., a user). At the same time, or as an alternative, the user device can launch an application and / or initiate operation of a particular function based on the determined distance. For example, a user can make a gesture associated with launching an imaging application on the user device, causing the user device to launch the imaging application in response thereto. The user device can initiate a filter function based on the determined distance. The filter function can be applied to an object to be imaged in a more accurate manner by virtue of the distance determined by the user device.
[0137] In some embodiments, a user device can have augmented reality (AR) and virtual reality (VR) functionality. For example, the techniques of the present application can provide more immersive augmented reality (AR) and virtual reality (VR) experiences. For example, a user device can utilize the depth sensing systems and processes of the present application to more accurately determine a distance between the user device (and its user) and an object in an augmented reality. The user device can perform one or more movements based at least in part on the determined distance between the user device and the object. For example, the user device can render components of the augmented reality around the object and / or render a virtual environment based on the determined distance.
[0138] In some implementations, the systems and methods of the present application can be used in a surveillance system. For example, a surveillance system can be configured to perform one or more operations based on a distance of an object and / or detection of an object. The surveillance system can receive a first optical signal via a distance estimation system and determine a first distance between the surveillance system and an object (e.g., a person) based on the first optical signal. The surveillance system can receive a second optical signal via a distance refinement system and determine a second distance between the surveillance system and an object based on the second optical signal. The surveillance system can determine a third distance between the surveillance system and the object based on the first distance and the second distance via a processing system, as described herein. The surveillance system can perform one or more functions or operations based on the third distance. For example, the surveillance system can perform an access function (e.g., deny access) based on the distance of the object (e.g., an intruder) and / or its actions.
[0139] Figure 12 An example implementation in accordance with the present application depicts a block diagram of an example computing system 1200. The example system 1200 includes a computing system 1202 and a machine learning computing system 1232, which are communicatively coupled over a network 1250.
[0140] In some implementations, the computing system 1202 can perform the operations and functions of the various computing components described herein. For example, the computing system 1202 can represent a depth sensing system, an estimation system, a refinement system, a processing system, and / or other systems described herein, and perform the functions of such systems. The computing system 1202 can include one or more physically distinct entity computing components.
[0141] The computing system 1202 can include one or more computing components 1204. The one or more computing components 1204 can include one or more processors 1206 and a memory 1208. The one or more processors 1206 can be any suitable processing component (e.g., processor cores, microprocessors, ASICS, FPGAs, controllers, microcontrollers, etc.) and can be single or multiple processing units operating in concert. The memory 1208 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, one or more memory components, flash memory components, etc., and combinations thereof.
[0142] Memory 1208 can store information accessed by one or more processors 1206. For example, memory 1208 (e.g., one or more non-transitory computer-readable storage medium, memory components) can store data 1210 that is retrieved, received, accessed, written, manipulated, created, and / or stored by one or more processors 1206. Data 1210 can include, for example, data representing distances, objects, signals, errors, ranges, model inputs, model outputs, and / or one or more of any other data and / or information described herein. In some embodiments, computing system 1202 can retrieve the data from one or more memory components that are remote from the system 1202.
[0143] Memory 1208 can also store computer-readable instructions 1212 executed by one or more processors 1206. Instructions 1212 can be software written in any suitable programming language or can be implemented in hardware. Also, or instead, instructions 1212 can be executed in logically and / or virtually separate threads on processor 1206.
[0144] For example, when instructions 1212 stored by memory 1208 are executed by one or more processors 1206, one or more processors 1206 can be caused to perform any operations and / or functions described herein, including, for example, operations or functions of any systems described herein, one or more portions of methods / procedures described herein, and / or any other functions or operations.
[0145] According to an aspect of the present disclosure, computing system 1202 can store or include one or more machine learning models 1216. For example, machine learning models 1216 can be or otherwise include various machine learning models such as neural networks (e.g., deep neural networks), support vector machines, decision trees, ensemble models, k-nearest neighbor algorithms, Bayesian networks, or other types of models (including linear and / or non-linear models). Example neural networks include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks, etc.), convolutional neural networks, and / or other forms of neural networks.
[0146] In some implementations, the computing system 1202 can receive one or more machine learning models 1216 from the machine learning computing system 1232 via the network 1250 and can store the one or more machine learning models 1216 in the memory 1208. The computing system 1202 can then use or otherwise employ the one or more machine learning models 1216 (e.g., via the processor 1206). In particular, the computing system 1202 can utilize the machine learning model 1216 to output distance data. The distance data can represent a depth estimate. For example, the machine learning model 1216 can determine a depth estimate as determined by a plurality of images taken over time by one or more image sensors. The machine learning model 1216 can receive at least a portion of an image that includes an object in an environment and output a depth estimate.
[0147] In some implementations, the machine learning model 1216 can receive a fused input. For example, the fused input can be based on a signal from a distance estimation system (e.g., a portion of a digital image of an object in an environment taken by an RGB camera) and a signal from a distance refinement system (e.g., a portion of quadrature amplitude of an object in an environment taken by an i-ToF system). The machine learning model 1216 can be configured to receive the fused input and generate an output representative of a distance between the computing system 1202 and the object.
[0148] The machine learning computing system 1232 includes one or more processors 1234 and a memory 1236. The one or more processors 1234 can be any suitable processing component (e.g., processor cores, microprocessors, ASICs, FPGAs, controllers, microcontrollers, etc.) and can be single or multiple operating processors. The memory 1236 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, one or more memory components, flash memories, etc., and combinations thereof.
[0149] The memory 1236 can store information accessible by the one or more processors 1234. For example, the memory 1236 (e.g., one or more non-transitory computer-readable storage media, memory components) can store data 1238 that can be retrieved, received, accessed, written, manipulated, created, and / or stored. The data 1238 can include, for example, any data described herein and / or information related thereto. In some implementations, the machine learning computing system 1232 can retrieve the data from one or more memory components that are remote from the machine learning computing system 1232.
[0150] Memory 1236 can also store computer-readable instructions 1240 executed by the one or more processors 1234. Instructions 1240 can be software written in any suitable programming language or can be implemented in hardware. Meanwhile, or instead, instructions 1240 can execute on logically and / or virtually separate threads on processor 1234.
[0151] For example, when instructions 1240 stored by memory 1236 are executed by the one or more processors 1234, the one or more processors 1234 can be caused to perform any of the operations and / or functions described herein, including, for example, operations or functions of any of the systems described herein, one or more portions of the methods / procedures described herein, and / or any other functions or procedures.
[0152] In some implementations, machine learning computing system 1232 includes one or more server computing components. If machine learning computing system 1232 includes multiple server computing components, these server computing components can operate according to various computing architectures, including, for example, a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0153] Meanwhile or instead of machine learning models 1216 of computing system 1202, machine learning computing system 1232 can include one or more machine learning models 1242. For example, machine learning models 1242 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks), support vector machines, decision trees, ensemble models, k-nearest neighbor computing models, Bayesian networks, or other types of models (including linear and / or non-linear models). Example neural networks include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks, etc.), convolutional neural networks, and / or other forms of neural networks.
[0154] In an example, machine learning computing system 1232 can communicate with computing system 1202 according to a client / server architecture. For example, machine learning computing system 1232 can utilize machine learning models 1242 to provide a web service to computing system 1202. For example, the web service can provide functionality and operations of a depth sensing system and / or other systems described herein (e.g., determining a distance between two points, a distance between a component / system and an object, etc.).
[0155] Accordingly, machine learning models 1216 can be located at and used by computing system 1202, and / or machine learning models 1242 can be located at and used by machine learning computing system 1232.
[0156] In some implementations, the machine learning computing system 1232 and / or the computing system 1202 can train the machine learning models 1216 and / or 1242 using a model trainer 1244. The model trainer 1244 can train the machine learning models 1216 and / or 1242 using one or more training or learning algorithms. One example training technique is backwards propagation of error. In some implementations, the model trainer 1244 can perform a supervised training technique using a set of labeled training data. In other implementations, the model trainer 1244 can perform an unsupervised training technique using a set of unlabeled training data. The model trainer 1244 can perform a number of generalization techniques to improve the generalization ability of the trained models. Generalization techniques include weight decaying, dropout, or other techniques.
[0157] In particular, the model trainer 1244 can train the machine learning models 1216 and / or 1242 based on a set of training data 1246. The training data 1246 can include, for example, labeled input data (e.g., from RGB and / or i-ToF sensors) and / or fused sensor data representing distance information. The model trainer 1244 can be implemented as hardware, firmware, and / or software to control one or more processors.
[0158] The computing system 1202 can also include a communication interface 1214 for communicating with one or more systems or components, including systems or components that are remote from the computing system 1202. The communication interface 1214 can include any circuitry, components, software, etc. for communicating over the one or more networks 1250. In some implementations, the communication interface 1214 can include, for example, one or more of a communication controller, a receiver, a transceiver, a transmitter, a port, a conductor, software and / or hardware for communicating data. The machine learning computing system 1232 can likewise include a communication interface 1248.
[0159] The network 1250 can be any type of network or combination of networks that can enable communication between components. In some embodiments, the network can include one or more of a local area network, a wide area network, the Internet, a secure network, a mobile communication network, a mesh network, a peer-to-peer communication link, and / or some combination thereof, and can include any number of wired or wireless links. Communication over the network 1250 can be enabled, for example, using a network interface using any type of protocol, security scheme, encoding, format, encapsulation, etc.
[0160] Figure 12An example computing system 1200 that can be used to implement the present application is shown, although other computing systems can also be used. For example, in some embodiments, the computing system 1202 can include a model trainer 1244 and training data 1246. In such embodiments, the machine learning model 1216 can be trained and used locally on the computing system 1202. In another example, according to some embodiments, the computing system 1202 is not connected to other computing systems.
[0161] Further, components shown as being included in one of the computing systems 1202 or 1232 can instead be included in the other. Implementations of such configurations are also within the scope of the present application. Various possible configurations, combinations, and task and function divisions can be implemented using computer-based systems. Computer-implemented operations can be performed on a single component or across multiple components. Computer-implemented tasks and / or operations can be performed sequentially, in parallel, or in a different order.
[0162] Various devices can be configured to perform the methods, operations, and procedures described herein. For example, any system (e.g., an estimation system, a refinement system, a processing system, a depth sensing system) can include means for performing its operations and functions described herein. In some embodiments, one or more of the above-described means can be implemented separately. In some embodiments, one or more means can be part of or included in one or more other means. These devices can include processors, microprocessors, graphics processing units, logic circuitry, specialized circuitry, application-specific integrated circuits, programmable array logic, field-programmable gate arrays, controllers, microcontrollers, and / or other suitable hardware. The devices can also, or instead, include software-controlled devices, e.g., implemented using a processor or logic circuitry. The devices can include memory, or be able to otherwise access memory, e.g., one or more non-transitory computer-readable storage media, such as random access memory, read only memory, electrically erasable programmable read only memory, erasable programmable read only memory, flash / other memory components, data buffers, databases, and / or other suitable hardware.
[0163] While the above focuses on preferred embodiments of the present application, it is to be understood that the application is not limited to these precise embodiments. Rather, various modifications and similar arrangements and procedures can be implemented, and thus the scope of the claims should be interpreted as broadly as possible so as to encompass all such modifications and similar arrangements and procedures.
Claims
1. An apparatus for measurement, characterized by, The apparatus for measuring comprises: a distance estimation system comprising a sensor system, the distance estimation system configured to: receive a first optical signal; and based on the first optical signal, determine a first distance between two points in an environment; a distance refinement system comprising one or more indirect time-of-flight ranging sensors, the distance refinement system configured to: receive a second optical signal; and based on the second optical signal, determine a second distance between the two points in the environment; and a processing system comprising one or more computing components, the processing system configured to:
2. The device for measuring of claim 1, wherein, receive information representative of the first distance and the second distance; and based on the first distance and the second distance, determine a third distance between the two points in the environment; wherein a difference between (i) a true distance between the two points in the environment and (ii) the first distance is greater than a difference between (i) the true distance between the two points in the environment and (ii) the third distance, and wherein determining the third distance between the two points in the environment comprises:
3. The device for measuring of claim 2, wherein, based on the first distance, determining a multiplier of a resolvable distance associated with the one or more indirect time-of-flight ranging sensors; and 4. The device for measuring of claim 2, wherein, adding the second distance to a product of the resolvable distance and the multiplier to determine the third distance.
5. The device for measuring of claim 1, wherein, The sensor system comprises:
6. The device for measuring of claim 1, wherein, a sensor array configured to receive the first optical signal and generate one or more first electrical signals; and 7. The device for measuring of claim 1, wherein, one or more storage media configured to store one or more machine learning models trained to:
8. The device for measuring of claim 1, wherein, receive, as input, a representation of the one or more first electrical signals; and 9. The device for measuring of claim 1, wherein, provide an output representative of the first distance. The sensor array is a sensor array of an RGB camera, and the representation of the one or more first electrical signals comprises at least a portion of an optical image output by the RGB camera. The one or more machine learning models comprise a convolutional neural network model. The sensor system comprises a stereo camera having multiple lenses, and the first optical signal comprises at least a portion of a three-dimensional image output by the stereo camera. The sensor system comprises a structured light detection system. The sensor system comprises one or more direct time-of-flight ranging sensors, and determining the first distance comprises: based on a round-trip time of a light ray between the two points in the environment, determining the first distance. The one or more indirect time-of-flight ranging sensors are configured to operate at a first frequency, the sensor system comprises one or more second indirect time-of-flight ranging sensors operating at a second frequency, and the second frequency is less than the first frequency. The sensor system comprises an image sensor system and a time-of-flight ranging sensor system, and determining the first distance comprises: obtaining an output of the image sensor system; obtaining an output of the time-of-flight ranging sensor system; and based on the output of the image sensor system and the output of the time-of-flight ranging sensor system, determining the first distance.
10. The device for measuring of claim 1, wherein, A first point between the two points in the environment represents a location of the device, and a second point between the two points in the environment represents a location of an object in the environment.
11. The device for measuring of claim 1, wherein, The device is a mobile device, the sensor system includes one or more cameras embedded in the mobile device, and the one or more indirect time-of-flight sensors include a three-dimensional sensor array embedded in the mobile device.
12. The device for measuring of claim 11, wherein, The one or more cameras and the three-dimensional sensor array are embedded on a back side of the mobile device, i.e., on an opposite side of a screen of the mobile device.
13. The device for measuring of claim 11, wherein, The one or more cameras and the three-dimensional sensor array are embedded on a front side of the mobile device, i.e., on a same side of a screen of the mobile device.
14. The device for measuring of claim 1, wherein, The one or more computing components include one or more processors, one or more hardwired circuits, one or more field-programmable gate arrays, or a combination thereof.
15. The device for measuring of claim 1, wherein, A difference between (i) the real distance between the two points in the environment and (ii) the first distance represents a first error, the first error being smaller than a resolvable distance associated with the one or more indirect time-of-flight sensors; a value of the second distance is within the resolvable distance; and a difference between (i) the real distance between the two points in the environment and (ii) the third distance represents a second error, the second error being smaller than the first error.
16. A method for determining a distance between two points in an environment, the method comprising: The method includes: receiving, by a sensor system, a first optical signal; determining, by the sensor system based on the first optical signal, a first distance between two points in an environment; receiving, by one or more indirect time-of-flight sensors separate from the sensor system, a second optical signal; determining, by the one or more indirect time-of-flight sensors based on the second optical signal, a second distance between the two points in the environment; receiving, by one or more computing components, information representing the first distance and the second distance; and determining, by the one or more computing components based on the first distance and the second distance, a third distance between the two points in the environment; wherein a difference between (i) a real distance between the two points in the environment and (ii) the first distance is greater than a difference between (i) the real distance between the two points in the environment and (ii) the third distance, and wherein determining the third distance between the two points in the environment includes: determining, based on the first distance, a multiplier of a resolvable distance associated with the one or more indirect time-of-flight sensors; and adding the second distance to a product of the resolvable distance and the multiplier to determine the third distance. An autonomous vehicle or a user device includes the sensor system, the one or more indirect time-of-flight sensors, and the one or more computing components.
17. The method of claim 16, wherein, The device for measuring includes:
18. An apparatus for measurement, comprising: a distance estimation system including a sensor system, the distance estimation system configured to: receive a first optical signal; and generate a first electrical signal based on the first optical signal to determine a first distance between two points in an environment; A range refinement system including one or more indirect time-of-flight ranging sensors, the range refinement system configured to: receive a second optical signal; and generate a second electrical signal based on the second optical signal to determine a second range between the two points in the environment; and A processing system including one or more computational components, the processing system configured to: receive the first electrical signal and the second electrical signal; provide input information representative of the first electrical signal and the second electrical signal to a machine learning model; receive output information representative of a third range between the two points in the environment; and determine the third range between the two points in the environment. wherein a maximum value of the second range is less than a resolvable range associated with the one or more indirect time-of-flight ranging sensors, and a maximum value of the third range is greater than the resolvable range.
19. A measuring device, characterized in that, The apparatus for measuring includes: An estimation system including a sensor system, the estimation system configured to: receive a first signal; and determine, based on the first signal, a first value of a property associated with a target object in an environment; A refinement system including one or more indirect time-of-flight ranging sensors, the refinement system configured to: receive a second signal; and determine, based on the second signal, a second value of the property associated with the target object in the environment; and A processing system including one or more computational components, the processing system configured to: receive information representative of the first value and the second value; and determine, based on the first value and the second value, a third value of the property associated with the target object in the environment; wherein a difference between (i) a true value of the property associated with the target object in the environment and (ii) the first value is greater than a difference between (i) the true value of the property and (ii) the third value, and wherein determining the third value of the property associated with the target object in the environment includes: determining, based on the first value, a multiplier of a resolvable range associated with the one or more indirect time-of-flight ranging sensors; and adding the second value to a product of the resolvable range and the multiplier to determine the third value.
Citation Information
Patent Citations
Intensity and Depth Measurements in Time-of-Flight Sensors
US20200158876A1