Hybrid mode depth imaging
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2021-09-24
- Publication Date
- 2026-08-07
AI Technical Summary
然而,以足够的分辨率和/或准确度估计距离信息可能是非常有力且计算密集的
Smart Images

Figure CN117043547B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to depth imaging. For example, aspects of this disclosure relate to combined techniques for structured light and time-of-flight (ToF) depth imaging. Background Technology
[0002] Image sensors are typically integrated into a wide range of electronic devices, such as cameras, mobile phones, autonomous systems (e.g., autonomous drones, cars, robots, etc.), smart wearables, extended reality (e.g., augmented reality, virtual reality, mixed reality) devices, and many others. Image sensors allow users to capture video and images from any electronic device equipped with an image sensor. Video and images can be captured for entertainment, professional photography, surveillance, automation, and other applications. Video and images captured by image sensors can be manipulated in various ways to improve video or image quality and create artistic effects.
[0003] In some cases, light signals and image data captured by an image sensor can be analyzed to identify certain characteristics about the image data and / or the scene captured by the image data. These characteristics can then be used to modify the captured image data or perform various tasks. For example, light signals and / or image data can be analyzed to estimate the distance to the scene captured by the image data. Estimated distance information can be used in a variety of applications, such as 3D photography, extended reality experiences, object scanning, autonomous vehicle operation, earth topography, computer vision systems, facial recognition systems, robotics, games, and creating various artistic effects, such as blurring and bokeh effects (e.g., out-of-focus effects). However, estimating distance information with sufficient resolution and / or accuracy can be very powerful and computationally intensive. Summary of the Invention
[0004] This document describes systems and techniques for performing hybrid-mode depth imaging, at least in part, by combining techniques for structured light and time-of-flight (ToF) depth imaging. According to an illustrative example, a method for generating one or more depth maps is provided. The method includes: obtaining a frame comprising a reflection pattern of light generated based on a light pattern emitted by a structured light source, the light pattern being based on primitives comprising a set of uniquely identifiable features; determining a first distance measurement associated with pixels of the frame using a ToF sensor; determining a search space within the primitives, the search space comprising a subset of features from the set of uniquely identifiable features of the primitives, at least in part based on the first distance measurement; determining features of the primitives corresponding to a region surrounding the pixels of the frame based on searching the search space within the primitives; determining a second distance measurement associated with the pixels of the frame, at least in part based on the features of the primitives determined from the search space within the primitives; and generating a depth map at least in part based on the second distance measurement.
[0005] In another example, an apparatus for generating one or more depth maps is provided. The apparatus includes a structured light source configured to emit a light pattern based on primitives. Each primitive includes a uniquely identifiable set of features. The apparatus also includes a time-of-flight (ToF) sensor, at least one memory, and one or more processors coupled to the at least one memory. The one or more processors are configured to: obtain a frame including a reflection pattern of light generated based on the light pattern emitted by the structured light source; determine a first distance measurement associated with a pixel of the frame using the ToF sensor; determine a search space within the primitives based at least in part on the first distance measurement, the search space including a subset of features from the uniquely identifiable set of features of the primitives; determine features of the primitives corresponding to a region surrounding the pixel of the frame based on searching the search space within the primitives; determine a second distance measurement associated with the pixel of the frame based at least in part on the features of the primitives determined from the search space within the primitives; and generate a depth map based at least in part on the second distance measurement.
[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: obtain a frame containing a light reflection pattern generated based on a light pattern emitted by a structured light source, the light pattern being based on primitives containing a uniquely identifiable set of features; determine a first distance measurement associated with a pixel of the frame using a ToF sensor; determine a search space within the primitives based at least in part on the first distance measurement, the search space including a subset of features from the uniquely identifiable set of features of the primitives; determine features of the primitives corresponding to a region surrounding the pixel of the frame based on searching the search space within the primitives; determine a second distance measurement associated with the pixel of the frame based at least in part on the features of the primitives determined from the search space within the primitives; and generate a depth map based at least in part on the second distance measurement.
[0007] In another example, an apparatus for performing temporal blending on one or more frames is provided. The apparatus includes: means for obtaining a frame comprising a reflection pattern of light generated based on a light pattern emitted by a structured light source, the light pattern being based on primitives comprising a set of uniquely identifiable features; means for determining a first distance measurement associated with a pixel of the frame using a ToF sensor; means for determining a search space within the primitives based at least partially on the first distance measurement, the search space comprising a subset of features from the set of uniquely identifiable features of the primitives; means for determining features of the primitives corresponding to a region surrounding the pixel of the frame based on searching the search space within the primitives; means for determining a second distance measurement associated with the pixel of the frame based at least partially on the features of the primitives determined from the search space within the primitives; and means for generating a depth map based at least partially on the second distance measurement.
[0008] In some aspects, the method, apparatus, and computer-readable medium may include (or be configured to): obtain a first exposure of a frame associated with a first illumination level; obtain a second exposure of the frame associated with a second illumination level different from the first illumination level; and determine a first distance measurement associated with the pixels of the frame based at least in part on a comparison between a first light amplitude associated with the pixels in the first exposure and a second light amplitude associated with the pixels in the second exposure.
[0009] In some aspects, the first distance measurement includes a distance measurement range. In some aspects, the method, apparatus, and computer-readable medium may include determining (or being configured to determine) the size of a search space within a primitive based at least in part on the range of the distance measurement. For example, a large range of distance measurements is associated with a large search space. In some cases, the method, apparatus, and computer-readable medium may include determining (or being configured to determine) the range of the distance measurement based at least in part on the ambiguity level associated with the ToF sensor. For example, a high ambiguity level is associated with a large range of distance measurements.
[0010] In some aspects, the method, apparatus, and computer-readable medium may include (or be configured to): determining, at least in part, an offset between a first position of the pixel of the frame and a second position of the feature of the primitive, based on a first distance measurement, wherein the offset is inversely proportional to the first distance measurement; and determining, at least in part, the search space within the primitive based on the offset. In some cases, the method, apparatus, and computer-readable medium may include setting (or configuring to set) the central axis of the search space within the primitive to the second position of a feature of the primitive.
[0011] In some aspects, the region surrounding a pixel of a frame has a predetermined size. In such aspects, the method, apparatus, and computer-readable medium may include (or be configured to): determine a first region of a search space having a predetermined size; and determine whether image data within the region surrounding the pixel of the frame corresponds to image data within the first region of the search space. In some cases, the method, apparatus, and computer-readable medium may include (or be configured to): determine that image data within the region surrounding the pixel of the frame corresponds to image data within the first region of the search space; and determine a second distance measurement based at least in part on determining the distance between corresponding features of a pixel of the frame and the first region of the search space. In some cases, the method, apparatus, and computer-readable medium may include (or be configured to): determine that image data within the region surrounding the pixel of the frame does not correspond to image data within the first region of the search space within the primitive; determine a second region of the search space having the predetermined size; and determine whether image data within the region surrounding the pixel of the frame corresponds to image data within the second region of the search space.
[0012] In some aspects, the light pattern emitted by a structured light source comprises multiple light spots. In some aspects, the features within the set of uniquely identifiable features of a primitive include two or more light spots. In some cases, the light spots of a feature correspond to two or more pixels of a frame.
[0013] In some respects, the structured light source is configured to emit light patterns using diffractive optical elements that simultaneously project multiple light patterns corresponding to primitives.
[0014] In some aspects, the method, apparatus, and computer-readable medium may include (or be configured to): obtain an additional frame based on the primitives when the structured light source does not emit the light pattern; determine an ambient light signal based at least in part on the additional frame; and subtract the ambient light signal from the frame before determining a first distance measurement associated with the pixel of the frame. In some cases, the method, apparatus, and computer-readable medium may include (or be configured to): after subtracting the ambient light signal from the frame, use the frame to determine a light signal corresponding to multipath interference; and subtract the light signal corresponding to multipath interference from the frame before determining the first distance measurement associated with the pixel of the frame.
[0015] In some aspects, the methods, apparatus, and computer-readable media may include fitting (or configuring to fit) a function to an optical signal corresponding to the pixel of the frame prior to determining a first distance measurement associated with the pixel of the frame.
[0016] In some aspects, the device is, is part of, and / or includes: a camera, a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, a vehicle or a computing device or component of a vehicle, or other devices. In some aspects, the device includes a camera or multiple cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the aforementioned device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more accelerometers, any combination thereof, and / or other sensors).
[0017] This invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to define the scope of the claimed subject matter. The subject matter should be understood by referring to the appropriate portions of the entire specification, any or all of the drawings, and each claim.
[0018] The foregoing and other features and embodiments will become more apparent from the following description, claims and drawings. Attached Figure Description
[0019] The illustrative embodiments of this application are described in detail below with reference to the accompanying drawings:
[0020] Figure 1 This is a block diagram illustrating an example architecture of a time-of-flight (ToF) depth imaging system based on some examples;
[0021] Figure 2A This is a simplified block diagram illustrating an example of a direct ToF sensing process according to some examples of this disclosure;
[0022] Figure 2B This is a simplified block diagram illustrating an example of an indirect ToF sensing process according to some examples of this disclosure;
[0023] Figure 3A This is a block diagram illustrating the example architecture of a structured optical depth imaging system based on some examples;
[0024] Figure 3B This is a diagram illustrating examples of parallax caused by the parallax between an image sensor receiver and a pattern projector, based on several instances.
[0025] Figure 3C This is a diagram illustrating an example of a projected vertical-cavity surface-emitting laser (VCSEL) primitive reproduced by a diffractive optical element (DOE) according to some examples;
[0026] Figure 3D This is a diagram illustrating a depth imaging system, according to some examples, including a DOE and a lens placed in front of a VCSEL array;
[0027] Figure 4 This is a block diagram illustrating an example architecture of a hybrid mode depth imaging system based on some examples;
[0028] Figure 5A The following are example frame exposures captured by a ToF depth imaging system, based on some examples;
[0029] Figure 5B The following are example frame exposures captured by a hybrid-mode depth imaging system, based on some examples;
[0030] Figure 5C Example ToF depth maps generated by a hybrid-mode depth imaging system are shown, based on some examples.
[0031] Figure 6A and Figure 6B This is a diagram illustrating an example process for structured light decoding guided by ToF distance measurement, based on some examples;
[0032] Figure 7A and Figure 7BThe example frame exposure with reduced signal noise is shown according to some examples;
[0033] Figure 8 These are examples of graphs showing the captured amplitude and ideal amplitude of light received by a sensor, based on some examples.
[0034] Figure 9A and Figure 9B These are examples of depth maps generated by various depth imaging systems, based on a number of examples.
[0035] Figure 10 This is a flowchart illustrating an example of a process for hybrid-mode depth imaging, based on some examples;
[0036] Figure 11 This is a diagram illustrating an example of a system used to implement some of the aspects described herein. Detailed Implementation
[0037] Certain aspects and embodiments of this disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some can be applied in combination, as will be apparent to those skilled in the art. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The drawings and description are not intended to be limiting.
[0038] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of exemplary embodiments will provide those skilled in the art with enabling descriptions for implementing the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0039] Various systems and / or applications utilize three-dimensional (3D) information representing scenes, such as systems and / or applications performing facial recognition, authentication systems using facial identifiers (IDs) of objects, object scanning, object detection, object grasping, object tracking, autonomous driving, robotics, aerial navigation (e.g., for unmanned aerial vehicles, aircraft, etc.), indoor navigation, extended reality (e.g., augmented reality (AR), virtual reality (VR), mixed reality (MR), etc.), 3D scene understanding, and other tasks. The recent need to capture 3D information from scenes has generated a high demand for active depth sensing technologies.
[0040] Structured light systems are an example of technologies that provide reliable and highly accurate depth capture systems. Generally, a structured light system may include one or more structured light projectors and sensors for scanning and / or determining the size and / or movement of a scene and / or one or more objects (e.g., people, equipment, animals, vehicles, etc.) within the scene. The structured light projector projects light of a known shape or pattern onto a scene including one or more objects, and the sensors can receive light reflected from one or more objects in the scene. The structured light system can determine the size of the scene and / or movement within the scene (e.g., the size and / or movement of one or more objects within the scene) based on measured or detected deformations of the shape or pattern.
[0041] Time-of-flight (ToF) technology is another example of providing efficient and high-resolution depth capture systems. Typically, a ToF system may include one or more light emitters and one or more sensors. For example, a light emitter emits a light signal toward a target (e.g., one or more objects in a scene), which may hit the target and return to one or more sensors based on the light signal reflected from the target. One or more sensors can detect and / or measure the reflected light, which can then be used to determine the target's depth and / or distance information. A direct ToF system can determine depth and / or distance information based on the travel time of the emitted light signal (e.g., the time from the time the light signal is emitted to the time it is received with the corresponding return / reflected light signal). An indirect ToF system can determine depth and / or distance information based on two frames captured using two exposures of pulsed light spaced a certain time apart. The depth and / or distance information of a point in a frame may correspond to the ratio of the light amplitude of a point in one frame to the light amplitude of a point in another frame. An indirect ToF system can also determine depth and / or distance information based on the phase shift between the emitted light signal and the corresponding return / reflected light signal.
[0042] Structured light and Time-of-Flight (ToF) depth acquisition systems offer a variety of advantages and disadvantages. For example, structured light systems can determine highly accurate depth information but may have limited resolution and / or high computational complexity. ToF systems can generate high-resolution depth maps with low computational complexity, but the accuracy of the depth maps may be degraded by noise and / or light scattering.
[0043] This disclosure describes systems, apparatuses, methods, and computer-readable media (collectively, the “Systems and Technologies”) that provide improved depth imaging. These systems and technologies provide depth imaging systems with the ability to generate depth maps based on a combination of structured light and Time-of-Flight (ToF) technologies. Such depth imaging systems may be referred to as “hybrid-mode” depth imaging systems capable of performing hybrid-mode depth imaging. In some cases, the hybrid-mode depth imaging system can project a light pattern based on primitives from a structured light emitter. In some cases, the projected light pattern may include a pattern having primitives that repeat or tessellate in an overlapping or non-overlapping manner (e.g., using diffractive optical elements or DOEs). The depth imaging system can then use a ToF sensor to determine distance measurements associated with the returned (or reflected) light. ToF distance measurements can be used to accelerate (and more efficiently) the use of structured light technologies to determine additional (e.g., more accurate) distance measurements. For example, the depth imaging system can use ToF distance measurements to reduce the search space of a structured light decoder.
[0044] As described above, in some examples, a depth imaging system may include a structured light emitter configured to emit a light pattern comprising a checkerboard (e.g., or repeating) primitive pattern (also referred to as a primitive), such as using a DOE positioned relative to the structured light emitter. A primitive may contain multiple uniquely identifiable features (also referred to as "codewords"). For example, a feature or codeword in a primitive may comprise a 4×4 arrangement of light dots (also referred to as "dots"). As described herein, the feature (or codeword) can be used to perform a matching between a captured frame and the primitive pattern, as described herein. In some cases, the depth imaging system may use a vertical-cavity surface-emitting laser (VCSEL) to generate the primitive pattern, and a DOE may be used to tessellate the primitive pattern. In some examples, each dot in a primitive may correspond to a single VCSEL in a VCSEL array. The depth imaging system may project the checkerboard primitive pattern onto objects within the scene.
[0045] One or more sensors in a depth imaging system can capture frames based on light patterns (including repeating primitive patterns) reflected from objects within a scene and returned to the depth imaging system. Each point of the primitive pattern can occupy multiple pixels in the captured frame. In an illustrative example, the system (e.g., lenses for the receiver and transmitter) can be configured such that each point of the primitive corresponds to a 4×4 pixel arrangement in the captured frame. As described above, a feature (or codeword) can include a 4×4 point arrangement, which, when each point occupies 4×4 pixels, can result in the feature occupying 16×16 pixels in the captured frame.
[0046] Pixels in a frame can be offset (e.g., shifted) relative to corresponding points in the original primitive pattern. The values of these offsets correspond to and / or indicate the depth of the object associated with the pixel. In some cases, conventional structured light systems determine the depth associated with a pixel in a frame by obtaining a pixel region (or block) (e.g., a 16×16 pixel block) around the pixel and searching within the primitive for uniquely identifiable features corresponding to (e.g., matching or most similar to) the “features” in the pixel region surrounding the pixel. This technique can involve searching the entire primitive to identify corresponding (e.g., most similar) pixels, which can require significant time and / or processing power. For example, structured light decoding typically involves identifying the region (e.g., a 16×16 block) around each pixel in a frame from a pattern of possibly similar sizes from a primitive (e.g., which may have a size of 124×64). Using block-matching type decoding, a depth imaging system can compare each 4×4 region of a point in the primitive with a 16×16 neighborhood around the current pixel from the frame.
[0047] To avoid searching entire primitives, hybrid-mode depth imaging systems can determine a search space that includes the region (e.g., block, slice, or portion) of the primitive to be searched, based on Time-of-Flight (ToF) distance measurements associated with light spots in the frame. For example, a depth imaging system can determine ToF distance measurements of all or part of the light spots in a reflected primitive pattern (e.g., indirect ToF distance measurements). Based on a ToF distance measurement associated with a pixel in the frame, a depth imaging system can determine a search space within the primitive that may and / or is expected to include unique identifiable features corresponding to features in the pixel region surrounding the pixel in the frame. For example, the ToF distance measurement may correspond to an estimated (e.g., unrefined) offset between a feature of the frame and a corresponding feature of the primitive. The search space within the primitive to be searched can include points of the primitive at or near the estimated offset. In some cases, the size (e.g., width) of the search space can be defined at least in part based on the level of ambiguity associated with the ToF measurement. The ambiguity level may be a result of the configuration of the ToF sensor and / or inherent inaccuracies in the ToF system (which may generally be less accurate than structured light systems). In the illustrative example, the search space within a primitive can be defined as being centered on the offset and having a width corresponding to the ambiguity level measured by the Time-of-Flight (ToF) method. Higher ambiguity levels can correspond to a larger width. Furthermore, the search space within a primitive can span all or part of the primitive's height.
[0048] After defining a search space within the primitives, the hybrid-mode depth imaging system can search within the search space to identify features of the primitives corresponding to “features” formed by the pixel regions surrounding the pixels of a frame. In some cases, the hybrid-mode depth imaging system can search the search space of the primitives by comparing a block of frame pixels surrounding a pixel of a particular frame with a dot patch of the primitive having the corresponding size. For example, a 16×16 pixel block surrounding a pixel in a frame (e.g., where the pixel is in the middle of the 16×16 block) can be compared with various 16×16 dot patches within the primitive's search space. The blocks or regions can have any suitable and / or predetermined size (e.g., 16×16, 32×32, 64×64, etc.). In an illustrative example, the hybrid-mode depth imaging system can use a dot product similarity measure or any other suitable similarity measure to compare the blocks.
[0049] Once the hybrid-mode depth imaging system identifies the corresponding features within the search space of primitives, it can determine more precise (e.g., refined) distance measurements associated with a specific frame pixel. For example, the system can determine the precise offset between the location of a feature in a pixel region surrounding a frame pixel and the location of the corresponding primitive feature. In some examples, the system can repeat the hybrid-mode depth imaging process for all or part of additional pixels in the frame.
[0050] Hybrid-mode depth imaging systems can generate depth maps of a scene based on determined distance measurements (e.g., refined distance measurements). In some cases, the depth map can be compared with the depth map generated using a conventional structured light system. Figure 1 The results are accurate and / or precise. Furthermore, by determining a relatively small region (within the search space) of the primitives to be searched based on ToF distance measurements, hybrid-mode depth imaging systems can generate depth maps in less time and / or with lower computational complexity than conventional structured light systems.
[0051] In some cases, the systems and techniques described herein (e.g., hybrid-mode depth imaging systems) can perform one or more operations to improve the accuracy of Time-of-Flight (ToF) distance measurements. Improving the accuracy of ToF distance measurements can reduce the level of ambiguity associated with the measurement, which in turn can reduce the size of the region to be searched within the primitives. In one example, the intensity of light associated with pixels (or pixel patterns) within a frame can have a desired distribution (e.g., a Gaussian bell distribution). The systems and techniques can reduce noise within the depth map by fitting the light signal corresponding to the captured frame to the desired distribution before determining the ToF distance measurement. In another example, the systems and techniques can reduce noise associated with ambient light and / or multipath interference by determining the ToF distance measurement based on multiple frames (e.g., multiple exposures of a primitive pattern). For example, the systems and techniques can capture frames corresponding to ambient light signals (e.g., frames captured when the structured light emitter is off) and subtract the ambient light signal from one or more frames used to determine the ToF distance measurement.
[0052] Furthermore, systems and techniques can determine (and then remove) light signals corresponding to multipath interference based on captured light from one or more pixels of a defined frame that do not correspond to light points in the primitive pattern. For example, any light that is not ambient light and is not directly from the pattern will be due to multipath interference (e.g., reflections of the projected pattern from objects in the scene). The projected pattern can have bright and dark areas. Light caused by multipath interference is highly diffused light, at least in part because multipath interference-based light comes from glow reflected from the surroundings. For example, if a spotlight is projected onto a wall in a room, the entire room will be flooded with light, which includes light reflected multiple times from various objects. Because systems and techniques can rely on relative brightness to perform Time-of-Flight (ToF) measurements, multipath interference-based light can affect ToF measurements. For example, multipath interference-based light can make sharp corners appear curved in the resulting depth map (e.g., on the depth axis in a point cloud).
[0053] When using diffused light (e.g., floodlight illuminators), multipath interference mixes with direct light, making it difficult in some cases to separate multipath interference-based light from direct light. The system and techniques described herein utilize structured light sources (which have dark regions between light spots). By using structured light sources, multipath interference can be measured in the dark regions. For example, the system can measure multipath interference on the expected dark regions of a frame's pattern (e.g., obtained using a ToF sensor), resulting in sparse measurements of multipath interference-based light. To generate a complete map of the scene, the system can perform interpolation using the sparse multipath interference measurements. For example, the system can interpolate across the sparse multipath interference measurements, thereby providing the system with an accurate representation of multipath interference across the entire frame received by the ToF sensor, including the contribution of multipath interference to the bright regions of the projected pattern. The system can subtract multipath interference from the frame's pattern. Subtracting multipath interference improves the ToF accuracy of the hybrid-mode depth imaging system and thus reduces the ambiguity level of structured light computation in the hybrid-mode depth imaging system (and therefore reduces the search space). As described in this paper, reducing the level of ambiguity (and search space) can reduce the computational load of a hybrid-mode depth imaging system by reducing the area that needs to be searched when performing structured light calculations.
[0054] As described above, structured light systems use points to construct features. The depth resolution of such systems (e.g., the width and height of the resulting depth map) is determined by the number of points in the projected pattern, rather than the resolution of the captured frame. Using the system and techniques described herein, because the light points (dots) occupy a certain number of pixels in the captured frame (e.g., a 4×4 pixel arrangement), the structured light decoded depth map is a fraction of the frame resolution (e.g., when each point occupies 4×4 pixels, the depth map has one-quarter of the frame resolution). In an illustrative example where each point occupies 4×4 pixels, if the frame resolution is 640×480, the depth map will be 160×120. On the other hand, Time-of-Flight (ToF) measures the depth value for each pixel of the returned frame, in which case the depth map resolution is equal to the frame resolution.
[0055] The reduced depth map resolution in structured light systems is due to practical reasons rather than fundamental limitations. For example, structured light matching algorithms can be used, which return depth values for each frame pixel, or depth values can be returned at the sub-pixel level by performing interpolation between pixel locations. In a more general sense, when matching features or codewords, there is no need for "alignment" to a point; in this case, the system can match any arbitrary offset. In most applications, the high computational cost makes using full-resolution depth maps impractical. The hybrid-mode system and techniques described in this paper provide a way to recover full-frame resolution using a more complex structured light decoding process but with a reduced search space, making it practical from both a computational and efficiency standpoint.
[0056] Furthermore, many Time-of-Flight (ToF) systems employ floodlight emitters (e.g., uniform light sources). Some systems use two separate emitters or add a configurable diffuser and capture two frames, one for structured light and one for ToF. The systems and techniques described herein can be used for ToF measurements using a structured light source (instead of using two different emitters). In this case, using a structured light source for ToF measurements can mean that ToF does not return depth values for every pixel in a frame because there are unilluminated areas (e.g., areas between points). Therefore, ToF measurements are sparse. With the systems and techniques described herein, sparsity is not an issue because, for example, ToF measurements are used as guidance in the structured light decoding (e.g., matching) process. The structured light system essentially fills the gaps in sparse ToF measurements.
[0057] As described above, this system and technique can improve the accuracy of Time-of-Flight (ToF) distance measurements. For example, the system and technique can utilize regions between points (e.g., so-called "dark" regions of a frame) to measure multipath interference that affects ToF accuracy, which can be subtracted from the measurement. This solution might be impossible for floodlight-based ToF systems.
[0058] The system and technology also inherently offer a higher signal-to-noise ratio (SNR). For example, because the system and technology use structured light emitters and perform Time-of-Flight (ToF) at a point, the contrast is higher than that of a typical ToF system (e.g., using a floodlight) for the same emitter power. As a result, the system performs better under interference (e.g., outdoors in direct sunlight) and in high-absorption areas compared to a similarly powered floodlight.
[0059] This document provides further details regarding hybrid-mode depth imaging systems with reference to the accompanying figures. As will be explained in more detail with reference to the figures, the disclosed hybrid-mode depth imaging systems may include all or part of structured light depth imaging systems and / or Time-of-Flight (ToF) depth imaging systems.
[0060] Figure 1 This is a figure illustrating an example of a hybrid-mode depth imaging system 100 that can implement the hybrid-mode depth imaging techniques described herein. Additionally or alternatively, the hybrid-mode depth imaging system may include components for... Figure 4 The example depth imaging system 400 with structured optical signal processing shown is in whole or in part, and will be described in more detail below.
[0061] like Figure 1As shown, the hybrid-mode depth imaging system 100 may include a time-of-flight (ToF) sensor system 102, an image sensor 104, a storage device 106, and an application processor 110. In some examples, the depth imaging system 100 may optionally include other computing components 108, such as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), and / or an image signal processor (ISP), which the depth imaging system 100 may use to perform one or more of the operations / functions described herein with respect to the application processor 110. In some cases, the application processor 110 and / or other computing components 108 may implement a ToF engine 130, an image processing engine 134, and / or a rendering engine 136.
[0062] It should be noted that in some examples, the application processor 110 and / or other computing components 108 can also implement Figure 1 One or more computing engines are not shown in the document. The ToF engine 130, image processing engine 134, and rendering engine 136 are provided herein for illustrative and explanatory purposes, and other possible computing engines are not shown for simplicity. Furthermore, for illustrative and explanatory purposes, the operations disclosed herein by the ToF engine 130, image processing engine 134, rendering engine 136, and their various operations will be described as being implemented by the application processor 110. However, those skilled in the art will recognize that in other examples, the operations disclosed by the ToF engine 130, image processing engine 134, rendering engine 136, and / or their various operations may be implemented by other computing components 108.
[0063] The depth imaging system 100 may be part of or implemented by a computing device or a plurality of computing devices. In some examples, the depth imaging system 100 may be part of an electronic device (or device) such as a camera system (e.g., a digital camera, IP camera, video camera, security camera, etc.), a telephone system (e.g., a smartphone, cellular phone, conferencing system, etc.), a laptop or notebook computer, a tablet computer, a set-top box, a television, a display device, a digital media player, a game console, a video streaming device, a head-mounted display (HMD), an extended reality (XR) device, a drone, a computer in a car, an IoT (Internet of Things) device, a smart wearable device, or any other suitable electronic device. In some implementations, the ToF sensor system 102, the image sensor 104, the storage 106, other computing components 108, the application processor 110, the ToF engine 130, the image processing engine 134, and the rendering engine 136 may be part of the same computing device.
[0064] For example, in some cases, the ToF sensor system 102, image sensor 104, storage device 106, other computing components 108, application processor 110, ToF engine 130, image processing engine 134, and rendering engine 136 can be integrated into a camera, smartphone, laptop computer, tablet computer, smart wearable device, HMD, XR device, IoT device, gaming system, and / or any other computing device. However, in some implementations, one or more of the ToF sensor system 102, image sensor 104, storage device 106, other computing components 108, application processor 110, ToF engine 130, image processing engine 134, and / or rendering engine 136 may be part of or implemented by two or more separate computing devices.
[0065] The ToF sensor system 102 can use light such as near-infrared (NIR) light to determine depth and / or distance information about a target (e.g., surrounding / nearby scene, one or more surrounding / nearby objects, etc.). In some examples, the ToF sensor system 102 can measure both the distance and intensity of each pixel in the target (such as a scene). The ToF sensor system 102 may include a light emitter to emit a light signal toward the target (e.g., scene, object, etc.), which may hit the target and return / reflect back to the ToF sensor system 102. The ToF sensor system 102 may include sensors for detecting and / or measuring the returned / reflected light, which can then be used to determine the target's depth and / or distance information. The distance of the target relative to the ToF sensor system 102 can be used to perform depth mapping. The distance to the target can be calculated using either direct ToF or indirect ToF.
[0066] In direct Time-of-Flight (ToF), distance can be calculated based on the travel time of the emitted and reflected light pulses (e.g., the time from the emission of the light pulse to the receipt of the reflected light pulse). For example, the round-trip distance of the emitted and reflected light pulses can be calculated by multiplying the travel time of the emitted and reflected light pulses by the speed of light (typically denoted as c). The calculated round-trip distance can then be divided by 2 to determine the distance from the ToF sensor system 102 to the target.
[0067] In indirect Time-of-Flight (ToF), distance can be calculated by sending modulated light to the target and measuring the phase of the returned / reflected light. Knowing the frequency (f) of the emitted light, the phase shift of the returned / reflected light, and the speed of light allows for the calculation of the distance to the target. For example, the difference in travel time between the paths of the emitted and returned / reflected light causes a phase shift in the returned / reflected light. The phase difference between the emitted and returned / reflected light, along with the modulation frequency (f) of the light, can be used to calculate the distance between the ToF sensor system 102 and the target. For example, the formula for the distance between the ToF sensor system 102 and the target could be c / 2f × phase shift / 2π. As shown, higher frequency light can provide higher measurement accuracy but will result in a shorter maximum measurable distance.
[0068] Therefore, in some examples, dual frequencies can be used to improve measurement accuracy and / or distance, as further explained herein. For example, a 60MHz optical signal can be used to measure a target 2.5 meters away, and a 100MHz optical signal can be used to measure a target 1.5 meters away. In a dual-frequency scenario, both the 60MHz and 100MHz optical signals can be used to calculate a target 7.5 meters away.
[0069] Image sensor 104 may include any image and / or video sensor or capture device, such as a digital camera sensor, video camera sensor, smartphone camera sensor, image / video capture device on an electronic device (such as a television or computer), camera, etc. In some cases, image sensor 104 may be part of a camera or computing device (such as a digital camera, camcorder, IP camera, smartphone, smart TV, gaming system, etc.). In some examples, image sensor 104 may include multiple image sensors, such as a rear sensor device and a front sensor device, and may be part of a dual-camera or other multi-camera assembly (e.g., including two cameras, three cameras, four cameras, or other numbers of cameras). Image sensor 104 may capture image and / or video frames (e.g., raw image and / or video data), which may then be processed by application processor 110, ToF engine 130, image processing engine 134, and / or rendering engine 136, as further described herein.
[0070] Storage device 106 can be any storage device used for storing data. Furthermore, storage device 106 can store data from any component of the depth imaging system 100. For example, storage device 106 may store data from ToF sensor system 102 (e.g., ToF sensor data or measurements), data from image sensor 104 (e.g., frames, videos, etc.), data from other computing components 108 and / or application processor 110 and / or data used by other computing components 108 and / or application processor 110 (e.g., processing parameters, image data, ToF measurements, depth maps, tuning parameters, processing outputs, software, files, settings, etc.), data from ToF engine 130 and / or data used by ToF engine 130 (e.g., one or more neural networks, image data, tuning parameters, auxiliary metadata, ToF sensor data, ToF measurements, depth maps, training datasets, etc.), data from image processing engine 134 (e.g., image processing data and / or parameters, etc.), data from rendering engine 136 and / or data used by rendering engine 136 (e.g., output frames), operating system 100 of the depth imaging system, and / or any other type of data.
[0071] Application processor 110 may include, for example, but not limited to, CPU 112, GPU 114, DSP 116 and / or ISP 118, which application processor 110 may use to perform various computational operations, such as image / video processing, ToF signal processing, graphics rendering, machine learning, data processing, computation and / or any other operations. Figure 1 In the example shown, application processor 110 implements ToF engine 130, image processing engine 134, and rendering engine 136. In other examples, application processor 110 may also implement one or more other processing engines. Furthermore, in some cases, ToF engine 130 may implement one or more machine learning algorithms (e.g., one or more neural networks) configured to perform ToF signal processing and / or generate depth maps.
[0072] In some cases, application processor 110 may also include memory 122 (e.g., random access memory (RAM), dynamic RAM, etc.) and cache 120. Memory 122 may include one or more memory devices and may include any type of memory, such as volatile memory (e.g., RAM, DRAM, SDRAM, DDR, static RAM, etc.), flash memory, flash-based memory (e.g., solid-state drive), etc. In some examples, memory 122 may include one or more DDR (e.g., DDR, DDR2, DDR3, DDR4, etc.) memory modules. In other examples, memory 122 may include other types of memory modules. Memory 122 may be used to store data, such as image data, ToF data, processing parameters (e.g., ToF parameters, tuning parameters, etc.), metadata, and / or any type of data. In some examples, memory 122 may be used to store data from the ToF sensor system 102, image sensor 104, storage device 106, other computing components 108, application processor 110, ToF engine 130, image processing engine 134 and / or rendering engine 136 and / or data used by the ToF sensor system 102, image sensor 104, storage device 106, other computing components 108, application processor 110, ToF engine 130, image processing engine 134 and / or rendering engine 136.
[0073] Cache 120 may include one or more hardware and / or software components for storing data, such that future requests for that data can be served faster than if it were stored on memory 122 or storage device 106. For example, cache 120 may include any type of cache or buffer, such as, for example, a system cache or L2 cache. Cache 120 may be faster and / or more cost-effective than memory 122 and storage device 106. Furthermore, cache 120 may have lower power and / or operating requirements or footprint than memory 122 and storage device 106. Therefore, in some cases, cache 120 may be used to store / buffer and quickly serve certain types of data, such as image data or ToF data, expected to be processed and / or requested in the future by one or more components of depth imaging system 100 (e.g., application processor 110).
[0074] In some examples, the operation of the ToF engine 130, image processing engine 134, and rendering engine 136 (and any other processing engines) can be implemented by any computing component of the application processor 110. In one illustrative example, the operation of the rendering engine 136 can be implemented by the GPU 114, and the operation of the ToF engine 130, image processing engine 134, and / or one or more other processing engines can be implemented by the CPU 112, DSP 116, and / or ISP 118. In some examples, the operation of the ToF engine 130 and image processing engine 134 can be implemented by the ISP 118. In other examples, the operation of the ToF engine 130 and / or image processing engine 134 can be implemented by the ISP 118, CPU 112, DSP 116, and / or a combination of the ISP 118, CPU 112, and DSP 116.
[0075] In some cases, application processor 110 may include other electronic circuitry or hardware, computer software, firmware, or any combination thereof to perform any of the various operations described herein. In some examples, ISP 118 may receive data (e.g., image data, ToF data, etc.) captured or generated by ToF sensor system 102 and / or image sensor 104, and process the data to generate an output depth map and / or frames. Frames may include video frames from a video sequence or still images. Frames may include an array of pixels representing a scene. For example, a frame may be a red-green-blue (RGB) frame with red, green, and blue color components per pixel; a luminance, chrominance-red, chrominance-blue (YCbCr) frame with a luminance component and two chrominance (color) components (chrominance-red and chrominance-blue) per pixel; or any other suitable type of color or monochrome image.
[0076] In some examples, the ISP 118 can implement one or more processing engines (e.g., ToF engine 130, image processing engine 134, etc.) and can perform ToF signal processing and / or image processing operations, such as depth calculation, depth mapping, filtering, de-mosaicing, scaling, color correction, color conversion, noise reduction filtering, spatial filtering, artifact correction, etc. The ISP 118 can process data from other components in the ToF sensor system 102, image sensor 104, storage device 106, memory 122, cache 120, application processor 110, and / or data received from remote sources (such as remote cameras, servers, or content providers).
[0077] Although the depth imaging system 100 is shown as including certain components, those skilled in the art will understand that the depth imaging system 100 may include more than Figure 1The components shown may include more or fewer components. For example, in some cases, the depth imaging system 100 may also include one or more other memory devices (e.g., RAM, ROM, cache, etc.), one or more network interfaces (e.g., wired and / or wireless communication interfaces, etc.), one or more display devices, and / or Figure 1 Other hardware or processing devices not shown in the diagram. The following section discusses… Figure 11 Illustrative examples of computing devices and hardware components that can be implemented using the depth imaging system 100.
[0078] Figure 2A This is a simplified block diagram illustrating an example of a direct ToF sensing process 200. Figure 2A In the example, the ToF sensor system 102 first emits a light pulse 202 toward a target 210. The target 210 may include, for example, a scene, one or more objects, one or more animals, one or more people, etc. The light pulse 202 may travel to the target 210 until it hits the target 210. When the light pulse 202 hits the target 210, at least some portion of the light pulse 202 may be reflected back to the ToF sensor system 102.
[0079] The Time-of-Flight (ToF) sensor system 102 can receive reflected light pulses 204, which include at least some portions of the light pulses 202 reflected back from the target 210. The ToF sensor system 102 can sense the reflected light pulses 204 and calculate a distance 206 to the target 210 based on them. To calculate the distance 206, the ToF sensor system 102 can calculate the total travel time of the emitted light pulses 202 and 204 (e.g., the time from the emission of the light pulse 202 to the receipt of the reflected light pulse 204). The ToF sensor system 102 can multiply the total travel time of the emitted light pulses 202 and 204 by the speed of light (c) to determine the total distance traveled by the light pulses 202 and 204 (e.g., round-trip time). The ToF sensor system 102 can then divide the total travel time by 2 to obtain the distance 206 from the ToF sensor system 102 to the target 210.
[0080] Figure 2B This is a simplified block diagram illustrating an example of an indirect ToF sensing process 220. In this example, the phase shift of the reflected light can be calculated to determine the depth and distance of the target 210. Here, the ToF sensor system 102 first emits modulated light 222 toward the target 210. The modulated light 222 may have a known or predetermined frequency. The modulated light 222 can travel to the target 210 until it hits the target 210. When the modulated light 222 hits the target 210, at least some portion of the modulated light 222 can be reflected back to the ToF sensor system 102.
[0081] The ToF sensor system 102 can receive reflected light 224, and the phase shift 226 of the reflected light 224 and the distance 206 to the target 210 can be determined using the following formula:
[0082] Distance (206) = c / 2f × phase shift / 2π,
[0083] Where f is the frequency of the modulated light and c is the speed of light.
[0084] In some cases, when calculating depth and distance (e.g., distance 206), one or more factors affecting how light is reflected can be considered or used to adjust the calculation. For example, objects and surfaces can have specific properties that can cause light to be reflected differently. To illustrate, different surfaces can have different refractive indices, which can affect how light travels or interacts with the surface and / or the materials within it. Furthermore, inhomogeneities (such as material irregularities or scattering centers) can cause light to be reflected, refracted, transmitted, or absorbed, and sometimes can lead to energy loss. Thus, when light strikes a surface, it can be absorbed, reflected, transmitted, etc. The proportion of light reflected by a surface is called its reflectivity. However, reflectivity depends not only on the surface (e.g., refractive index, material properties, homogeneity or inhomogeneity, etc.) but also on the type of light reflected and the surrounding environment (e.g., temperature, ambient light, water vapor, etc.). Therefore, as further explained below, in some cases, when calculating distance 206 and / or depth information of target 210, information about the surrounding environment, the type of light, and / or the characteristics of target 210 can be considered.
[0085] Figure 3A This is a depiction of an example depth imaging system 300 configured to use light distribution to determine the depth of objects 306A and 306B in scene 306. The depth imaging system 300 can be used to generate a depth map (not depicted) of scene 306. For example, scene 306 may include objects (e.g., faces), and the depth imaging system 300 can be used to generate a depth map including multiple depth values that indicate the depth of portions of an object used for object identification or authentication (e.g., for facial authentication). The depth imaging system 300 includes a projector 302 and a receiver 308. The projector 302 may be referred to as a “structured light source,” “transmitter,” “emitter,” “light source,” or other similar terms, and should not be limited to a particular transmission component. Throughout the following disclosure, the terms projector, transmitter, and light source are used interchangeably. The receiver 308 may be referred to as a “detector,” “sensor,” “sensing element,” “photodetector,” etc., and should not be limited to a particular receiving component.
[0086] Projector 302 can be configured to project or transmit a distribution 304 of light points onto scene 306. White circles in distribution 304 indicate locations where no projected light is received for possible point locations, and black circles in distribution 304 indicate locations where projected light is received for possible point locations. Distribution 304 may be referred to as a codeword distribution or pattern, wherein defined portions of distribution 304 are codewords (also referred to as codes or features). As used herein, a codeword is a rectangular (e.g., square) portion of light distribution 304. For example, a 5×5 codeword 340 is illustrated in distribution 304. As shown, codeword 340 includes five rows of possible light points and five columns of possible light points. Distribution 304 can be configured to contain an array of codewords. For active depth sensing, codewords can be unique to each other in distribution 304. For example, codeword 340 is different from all other codewords in distribution 304. Furthermore, the positions of unique codewords relative to each other are known. In this way, one or more codewords in the distribution can be identified in the reflection, and the position of the identified codewords relative to each other, the shape or distortion of the identified codewords relative to the shape of the transmitted codewords, and the position of the identified codewords on the receiver sensor are used to determine the depth of the object reflecting the codewords in the scene.
[0087] Projector 302 includes one or more light sources 324 (such as one or more lasers). In some embodiments, the one or more light sources 324 include laser arrays. In one illustrative example, each laser may be a vertical-cavity surface-emitting laser (VCSEL). In another illustrative example, each laser may include a distributed feedback (DFB) laser. In another illustrative example, the one or more light sources 324 may include an array of resonant cavity light-emitting diodes (RC-LEDs). In some embodiments, the projector may also include a lens 326 and a light modulator 328. Projector 302 may also include an aperture 322 through which transmitted light escapes. In some embodiments, projector 302 may also include a diffractive optical element (DOE) to diffract the emission from the one or more light sources 324 into additional emission. In some aspects, the light modulator 328 (for adjusting the intensity of the emission) may include a DOE. When projecting the distribution of light spots 304 onto scene 306, projector 302 can emit one or more lasers from light source 324 through lens 326 (and / or through DOE or light modulator 328) and project them onto objects 306A and 306B in scene 306. Projector 302 can be located on the same reference plane as receiver 308, and projector 302 and receiver 308 can be separated by a distance referred to as baseline 312.
[0088] In some example implementations, the light projected by projector 302 may be infrared (IR) light. IR light may include portions of the visible spectrum and / or portions of the spectrum invisible to the naked eye. In one example, IR light may include near-infrared (NIR) light, which may or may not include light within the visible spectrum, and / or IR light outside the visible spectrum (such as far-infrared (FIR) light). The term IR light should not be limited to light having a specific wavelength within or near the wavelength range of IR light. Furthermore, IR light is provided as an example emission from the projector. In the following description, light of other suitable wavelengths may be used. For example, portions of the visible light spectrum outside the IR light wavelength range or ultraviolet light may be used.
[0089] Scene 306 may contain objects at different depths from the structured light system (e.g., from projector 302 and receiver 308). For example, objects 306A and 306B in scene 306 may be at different depths. Receiver 308 may be configured to receive reflections 310 of the transmitted light spot distribution 304 from scene 306. To receive reflections 310, receiver 308 may capture frames. When capturing frames, receiver 308 may receive reflections 310, as well as (i) other reflections from the light spot distribution 304 at other depths of scene 306, and (ii) ambient light. Noise may also be present in the capture.
[0090] In some example implementations, receiver 308 may include lens 330 to focus or direct received light (including reflections 310 from objects 306A and 306B) onto sensor 332 of receiver 308. Receiver 308 may also include aperture 320. Assuming an example where only reflection 310 is received, the depths of objects 306A and 306B can be determined based on baseline 312, displacement and distortion (such as in codewords) of light distribution 304 in reflection 310, and the intensity of reflection 310. For example, distance 334 along sensor 332 from position 316 to center 314 can be used to determine the depth of object 306B in scene 306. Similarly, distance 336 along sensor 332 from position 318 to center 314 can be used to determine the depth of object 306A in scene 306. The distance along sensor 332 can be measured based on the number of pixels of sensor 332 or distance units (such as millimeters).
[0091] In some example implementations, sensor 332 may include an array of photodiodes (e.g., avalanche photodiodes) for capturing frames. To capture a frame, each photodiode in the array can capture light striking it and can provide a value indicating the light intensity (the capture value). Therefore, a frame can be an array of capture values provided by the photodiode array.
[0092] As a complement to or alternative to the sensor 332, which includes a photodiode array, the sensor 332 may include a complementary metal-oxide-semiconductor (CMOS) sensor. To capture an image via a photosensitive CMOS sensor, each pixel of the sensor can capture light impacting the pixel and can provide a value indicating the light intensity. In some example embodiments, the photodiode array may be coupled to the CMOS sensor. In this way, electrical pulses generated by the photodiode array can trigger the corresponding pixel of the CMOS sensor to provide the captured value.
[0093] Sensor 332 may include at least a plurality of pixels equal to the number of possible light spots in distribution 304. For example, a photodiode array or a CMOS sensor may each include at least a plurality of photodiodes or a plurality of pixels corresponding to the number of possible light spots in distribution 304. Sensor 332 may be logically divided into groups of pixels or photodiodes corresponding to the size of bits of a codeword (e.g., a 4×4 group of 4×4 codewords). A group of pixels or photodiodes may also be referred to as a bit, and a portion of the data captured from a bit of sensor 332 may also be referred to as a bit. In some example embodiments, sensor 332 may include at least the same number of bits as distribution 304. If light source 324 transmits IR light (such as NIR light with a wavelength of, for example, 940 nanometers (nm), then sensor 332 may be an IR sensor to receive reflections of NIR light.
[0094] As shown in the figure, distance 334 (corresponding to reflection 310 from object 306B) is less than distance 336 (corresponding to reflection 310 from object 306A). Using triangulation based on baseline 312 and distances 334 and 336, the different depths of objects 306A and 306B in scene 306 can be determined when generating a depth map of scene 306. Depth determination can also be based on displacement or distortion of distribution 304 in reflection 310.
[0095] In some implementations, the projector 302 is configured to project a fixed light distribution, in which case the same light distribution is used for active depth sensing in each instance. In some implementations, the projector 302 is configured to project different light distributions at different times. For example, the projector 302 may be configured to project a first light distribution at a first time and a second light distribution at a second time. Thus, the resulting depth map of one or more objects in the scene is based on one or more reflections from the first light distribution and one or more reflections from the second light distribution. The codewords between the light distributions may be different, and the depth imaging system 300 may be able to identify codewords in the second light distribution that correspond to positions in the first light distribution where codewords cannot be identified. In this way, more efficient depth values can be generated when generating the depth map without reducing the resolution of the depth map (e.g., by increasing the size of the codewords).
[0096] Despite Figure 3A Several individual components are shown, but one or more of these components may be implemented together or include additional functionality. The depth imaging system 300 may not require all the described components, or the functionality of the components may be separated into individual components. Additional components, not shown, may also be present. For example, receiver 308 may include a bandpass filter to allow signals with a defined wavelength range to pass onto sensor 332 (thus filtering out signals with wavelengths outside that range). In this way, some incidental signals (such as ambient light) can be prevented from being received as interference during capture by sensor 332. The range of the bandpass filter may be centered on the transmission wavelength of projector 302. For example, if projector 302 is configured to transmit NIR light with a wavelength of 940 nm, receiver 308 may include a bandpass filter configured to allow NIR light with wavelengths in, for example, the range of 920 nm to 960 nm. Therefore, regarding... Figure 3A The examples described are for illustrative purposes.
[0097] Structured light depth imaging systems can rely on measuring "parallax" (pixel displacement along an axis) caused by the parallax between an image sensor receiver (e.g., receiver 308) and a pattern projected into the scene (e.g., via a projector, such as projector 302). Figure 3B This diagram illustrates an example of determining parallax 356 caused by the parallax between image sensor receiver 358 and pattern projector 352. Generally, the closer an object is to image sensor receiver 358, the greater the pixel shift (and therefore parallax 356). Depth is the distance of a point from receiver 358. Depth is inversely proportional to the offset represented by parallax 356. This phenomenon is similar to stereoscopic vision, where two views are compared to each other (e.g., left and right eyes) to infer depth. One difference is that, in the case of structured light, one of the "views" is a known reference pattern (projection).
[0098] Measuring parallax 356 may include identifying each unique portion of the projected pattern at the image sensor receiver 358 (referred to as decoding). Then, taking into account that the parallax is 0 at infinity, the structured optical depth imaging system can measure the parallax along baseline 362 (the horizontal axis, which can be similar to...) due to the parallax. Figure 3AThe perceived offset of the baseline 312 in the structured optical depth imaging system. In some cases, the projected pattern may not be unique in the field of view (FoV) of the structured optical depth imaging system, in which case the parallax measurement will be around. This determines the nearest distance (e.g., distance 354) that the structured optical depth imaging system can infer. In some cases, the uniqueness associated with the structured optical depth imaging system is uniqueness along the baseline, in which case the projected pattern can be repeated with a small margin in the orthogonal (vertical) direction without interfering with the above measurements.
[0099] Structured light projectors can take many forms and have various types of projection patterns. As mentioned above, an example of a structured light projector is a vertical-cavity surface-emitting laser (VCSEL) array, which emits coded patterns of laser points that are either on or off. As described above, by using diffractive optical elements (DOEs), the primitive pattern of the VCSEL array can be optically copied (or tessellated) into the projection scene to form M×N tiles. Figure 3C An example of a projected VCSEL primitive copied by the DOE is shown. VCSEL array primitive 370 is shown in the center (order 0) highlighted by a box. Each blue dot represents a VCSEL laser point. Figure 3C In the example, the primitive 370 array is replicated as diffraction order +8 / -8x+3 / -3 tiles or 17x7 tiles, shown as various tiles (e.g., tile 372). Replication is performed by a DOE placed in front of the VCSEL array. For example, as... Figure 3D As shown, the DOE 386 is placed in front of the VCSEL array 382 and the lens 384.
[0100] Figure 4 This is a block diagram illustrating an example of a depth imaging system 400 for hybrid-mode depth imaging. In some examples, the depth imaging system 400 may be composed of... Figure 1 Depth imaging system 100 and / or Figure 3A The depth imaging system 300 implements and / or includes all or part of the depth imaging system. Figure 1 Depth imaging system 100 and / or Figure 3AThe depth imaging system 300 may include all or part of it. For example, the depth imaging system 400 may include a structured light system 402, which includes a structured light source 404. In one example, the structured light source 404 may correspond to the projector 302 of the depth imaging system 300. In an illustrative example, the structured light source 404 may include a VCSEL array and a DOE configured to diffract and project light emitted by the VCSEL array. As shown, the structured light source 404 may be configured to emit a light pattern 412 (e.g., corresponding to a distribution 304 emitted by the projector 302). For example, the structured light source 404 may project the pattern 412 to illuminate a scene. In some cases, the pattern 412 may correspond to a primitive pattern tessellated (e.g., repeated) by the structured light source 404. In one instance, the primitive pattern may contain multiple uniquely identifiable features (e.g., codewords), each corresponding to two or more light spots. Each light spot may correspond to two or more pixels (e.g., when captured by the ToF sensor 410 of the depth imaging system 400). In one example, the ToF sensor 410 may correspond to the image sensor 104 of the depth imaging system 100. In some cases, the ToF sensor 410 may be configured to receive, capture, and / or process incident light. For example, the ToF sensor 410 may capture pattern reflections 414, which correspond to light reflected and / or returned to the ToF sensor 410 by a pattern 412. In one example, the ToF sensor 410 may generate one or more frames 418 corresponding to the pattern reflections 414. As will be explained in more detail below, the processor 420 of the depth imaging system 400 may process and / or analyze the frames 418 to generate a depth map 416 of the scene illuminated by the pattern 412.
[0101] In some cases, depth imaging system 400 may include the full functionality of a ToF depth imaging system (such as depth imaging system 100) and / or a structured light depth imaging system (such as depth imaging system 300). For example, although not in Figure 4 As shown, however, the depth imaging system 400 may include a projector configured for flood illumination (e.g., to facilitate conventional ToF distance measurements). However, in some cases, the depth imaging system 400 may perform mixed-mode depth imaging using a single light source (e.g., structured light source 404) and a single image sensor (e.g., ToF sensor 410), in which case a projector for flood illumination is not used or is not included in the system 400. For example, the depth imaging system 400 may determine the ToF distance measurement of the pixel corresponding to the light spot in pattern 412 (instead of determining the ToF distance measurement of each pixel in frame 418).
[0102] In one example, the depth imaging system 400 may perform a hybrid-mode depth imaging process involving two stages. This process may be performed on all or a portion of the pixels corresponding to the light point in frame 418. In some cases, the first stage of the process may include determining a Time-of-Flight (ToF) distance measurement associated with the pixel. The ToF distance measurement may represent an estimated and / or unrefined distance measurement. The second stage of the process may include using the ToF distance measurement to determine a search space for primitives and performing a search within that search space. For example, the search space may correspond to a subset of primitives (e.g., a subset of features from a set of uniquely identifiable features in the primitives), in which case the “features” corresponding to the pixel in frame 418 are located (e.g., expected and / or likely to be located). As used herein, “searching” the search space for primitives may include comparing image data within a region surrounding the pixels of the frame with image data within a region of the same or reduced size within the search space of the primitives. For example, in some cases as described above, a point in a primitive may occupy a certain number of pixels in the capture frame (e.g., a point may occupy 4×4 pixels in the frame). In an example where a point occupies 4×4 pixels, a 4×4 region of primitives can be searched from the captured frame to obtain a 16×16 pixel arrangement. Once the depth imaging system 400 identifies the region of primitives corresponding to the region of the frame, the depth imaging system 400 can determine a structured optical distance measurement associated with the frame pixels (e.g., displacement between the region of the frame and the region of the primitives). In some cases, the structured optical distance measurement can represent a precise and / or refined distance measurement. The depth imaging system 400 can incorporate the structured optical distance measurement into the depth map 416.
[0103] The depth imaging system 400 can use various types of Time-of-Flight (ToF) sensing processes and / or techniques to determine ToF distance measurements. In an illustrative example, the depth imaging system 400 can implement an amplitude-based ToF process, which involves determining the ratio between the light amplitudes corresponding to a pixel within two or more exposures of a frame (referred to herein as "frame exposures"). The distance is based on the delay of the second exposure of the second image relative to the first exposure of the first image (e.g., the ratio between the brightness of the two exposures is proportional to the distance). In some cases, different levels of illumination (e.g., different levels of light intensity) are used to generate frame exposures. For example, the structured light system 402 can generate a first frame F1 with a first exposure by projecting a pattern 412 (e.g., using a structured light source 404) at a first illumination level with light illumination having a specific duration t. The first frame F1 corresponds to the light measured by the ToF sensor 410 during that duration t. The structured light system 402 can then generate a second frame F2 with a second exposure by measuring the light arriving at the sensor during a duration between t and 2t. The ToF distance measurement associated with a pixel can correspond to the ratio F1 / (F1+F2) of the light amplitude associated with the pixel in two frame exposures. In some embodiments, a third frame F3 can be measured without any illumination to capture background illumination caused by an external light source, such as light or the sun. The ToF distance can ultimately be measured as (F1-F3) / (F1+F2-2F3)*c*t, where c is the speed of light.
[0104] Figure 5A A first frame 502 and a second frame 504 are shown corresponding to an example of frame exposures generated by a conventional amplitude-based Time-of-Flight (ToF) system (e.g., a ToF system using floodlight illumination). In this example, the first frame 502 has a first frame exposure corresponding to a high brightness level, and the second frame 504 has a second frame exposure corresponding to a low brightness level. As shown, each pixel of the first frame 502 and the second frame 504 corresponds to reflected light (e.g., due to the frame exposure being generated using floodlight illumination). Based on the same light pulse propagating through the scene, the two exposures of the first frame 502 and the second frame 504 have the same duration. The two exposures are separated one after the other in time. The reason why frames 502 and 504 appear to have different brightness is how much reflected light the sensor captures for each of frames 502 and 504. The first exposure of the first frame 502 (e.g., corresponding to the duration t mentioned above) begins almost immediately after the light pulse is emitted. If there is no nearby object reflecting the light back, the exposure will appear dark because the duration is short enough to avoid sensing light reflected back from distant objects. The second exposure of the second frame 504 is delayed (e.g., corresponding to the duration between t and 2t mentioned above), and the sensor will be able to capture the light returning from a distant object, which is why the ratio between the brightness of the two exposures is proportional to the distance.
[0105] Figure 5B A first frame 506 with a first exposure and a second frame 508 with a second exposure are shown, corresponding to examples of frame exposures generated by the depth imaging system 400. For example, the depth imaging system 400 can generate the first frame 506 with a first exposure and the second frame 508 with a second exposure by using illumination projection patterns 412 with different levels. In this example, the first exposure of the first frame 506 corresponds to a high brightness level, and the second exposure of the second frame 508 corresponds to a low brightness level. In some cases, the two exposures of the first frame 506 and the second frame 508 have the same duration, similar to the above description regarding... Figure 5A The duration described. For example, the structured light system 402 may have an acquisition event for both ToF and structured light (e.g., ToF sensor 410 captures a frame based on a projected pattern). For ToF, the structured light system 402 may apply the aforementioned ratio formula (e.g., F1 / (F1+F2) or (F1-F3) / (F1+F2-2F3)*c*t). To calculate the structured light depth, the structured light system 402 may add together the two exposures of frames 506 and 508, thus recovering most or all of the light from the projected pattern. The structured light system 402 may perform structured light pattern matching on frames with combined exposures, guided by a search space defined by the ToF depth.
[0106] like Figure 5B As shown, the light spots shown in the first frame 506 and the second frame 508 correspond to pattern 412. For example, because pattern 412 contains multiple codewords (e.g., a uniquely identifiable feature composed of multiple light spots), a portion of the pixels in the first frame 506 and the second frame 508 are not associated with reflected light (e.g., the light corresponding to pattern reflection 414). Therefore, in some cases, ToF distance measurements may not be applicable to pixels not associated with reflected light. Figure 5C An example depth map 510 is shown that can be generated based on a first frame 506 with a first exposure and a second frame 508 with a second exposure. In this example, objects with darker shadows are associated with a shorter depth compared to objects with brighter shadows.
[0107] Figure 6A and Figure 6B An example is shown that uses Time-of-Flight (ToF) distance measurements to determine and search the search space for primitives. For example, Figure 6A Display frame 618 (for example, corresponding to Figure 4Pixel 602 within frame 418. In one example, the depth imaging system 400 may generate frame 618 based on tessellizing and projecting primitive 606 into space and receiving reflected light at a sensor (e.g., at ToF sensor 410). In such an example, features defined by pixel 602 and one or more other pixels in frame 618 are associated with the light spot of primitive 606. The depth imaging system 400 may determine a ToF distance measurement corresponding to pixel 602. The depth imaging system 400 may use the ToF distance measurement to determine features (e.g., feature 616) within primitive 606 that correspond to features defined by pixel 602 and one or more other pixels in frame 618. For example, the depth imaging system 400 may use the ToF distance measurement to determine a search space 608 for primitive 606, in which feature 616 corresponding to pixel 602 may and / or is expected to be located.
[0108] In some cases, a structured optical depth imaging system can determine whether a primitive feature corresponds to a frame pixel by defining a region (e.g., a pixel block) surrounding the frame pixel and searching for primitives in the corresponding region. In one example, the corresponding region may be a region containing a feature (or codeword) corresponding to the pixel region surrounding the frame pixel. Therefore, the region can have any suitable size such that the region includes one or more features (or codewords). In an illustrative example, the region of a primitive may contain a 4×4 dot arrangement (constituting a feature or codeword) corresponding to a 16×16 pixel block covering the frame pixel. In other examples, the region may correspond to a 4×4 pixel block, an 8×8 pixel block, an 8×16 pixel block, etc. In some cases, a structured optical depth imaging system can determine whether a region of a primitive corresponds to a region of a frame by comparing the primitive data (e.g., features or codewords) in the region of the primitive with image data in the region of the frame. For example, a structured optical depth imaging system can determine that the region of a primitive (e.g., a feature or codeword containing a 4×4 dot arrangement) corresponds to a region of the frame (e.g., a 16×16 pixel region) based on the fact that the data in the region exceeds a similarity threshold. In another example, a structured optical depth imaging system can determine that a region of a primitive corresponds to a region of a frame based on determining that the region of the primitive is most similar to a region of the frame (e.g., throughout the primitive). In an illustrative example, the structured optical depth imaging system can determine the similarity between the primitive's region's metadata and the frame's region's image data based on convolution operations, as described below. Other techniques for determining the similarity between data from two regions include block matching, normalized cross-correlation, matched filter techniques, and others.
[0109] In some instances, conventional structured light depth imaging systems (e.g., structured light depth imaging systems that do not incorporate Time-of-Flight (ToF) technology) can determine features (or codewords) within a primitive corresponding to features defined by frame pixels based on the entire primitive being searched. For example, given a region of a specific size in a frame, a conventional structured light depth imaging system can compare the region of the frame with each distinct region (of the same size) within the primitive. In some cases, it may be necessary to search the entire primitive to obtain an accurate distance measurement (e.g., to identify the correct corresponding region). In contrast, the disclosed hybrid-mode depth imaging system and techniques can use an initial ToF distance measurement to determine a search space that is a subset of the primitive. For example, depth imaging system 400 can identify corresponding points of primitives within the search space (without searching regions of primitives outside the search space). In this way, depth imaging system 400 can implement a hybrid-mode depth imaging process that combines both ToF and structured light technologies. Because determining a ToF distance measurement can involve significantly less time and / or processing power than searching the entire primitive, the hybrid-mode depth imaging process can improve the efficiency of conventional structured light systems (while maintaining the same accuracy and / or precision).
[0110] Return to Figure 6B The depth imaging system 400 can determine a search space 608 based on a ToF distance measurement associated with pixel 602. In some cases, the search space 608 can be defined at least partially by an offset 620. As shown, the offset 620 extends between the central vertical axis of the search space 608 and point 614. In some cases, such as... Figure 6B As shown, based on the fact that the emitter (e.g., structured light source 404) and sensor (e.g., ToF sensor 410) are placed on the same horizontal plane (in this case, parallax should only cause horizontal displacement), the centers of point 614 and feature 616 are on the same row (aligned in the horizontal plane). Point 614 is located in primitive 606 at a position corresponding to the position of pixel 602 within frame 618. Therefore, offset 620 can be used to define the horizontal position of search space 608 within primitive 606. In some cases, offset 620 can represent an estimate of the displacement (also known as parallax) between pixel 602 and feature 616 (see [reference]). Figure 3A , Figure 3B (and a corresponding description of the displacement used in structured light systems). In one example, depth imaging system 400 can determine the amplitude of offset 620 based on a ToF distance measurement associated with pixel 602. For example, the size of offset 620 can be inversely proportional to the ToF distance measurement. In structured light systems, nearby objects within a scene are associated with higher displacements than distant objects. Therefore, depth imaging system 400 can determine the amplitude of offset 620 based on a lower ToF distance measurement. The relatively high offset value of the pixels is determined, and the relatively low offset value of the pixels associated with the high ToF distance measurement is determined. For example, structured light distance (SL distance) can be determined as: displacement, also known as parallax, can be expressed in pixels. As mentioned above, offset 620 is an example of displacement (or parallax). An example of the baseline is shown as Figure 3A Baseline 312 and Figure 3B The baseline is 362. The above text is about... Figure 3B Describe an instance of determining depth based on displacement (or parallax or offset).
[0111] like Figure 6B As shown, the search space 608 can also be defined at least in part by the ambiguity level 610. In some cases, the ambiguity level 610 may correspond to the ambiguity level of a ToF distance measurement. For example, as mentioned above, ToF sensors may have inherent limitations that reduce the accuracy of ToF distance measurements. While ToF depth imaging systems can generate higher resolution depth maps faster and / or than structured light depth imaging systems, ToF distance measurements may generally be less accurate than structured light distance measurements. Therefore, the depth imaging system 400 can determine the ambiguity level 610 based on the expected amount of error associated with the ToF distance measurement determined using the ToF system 408. For example, the depth imaging system 400 can determine that the ToF system 408 is expected to accurately calculate the ToF distance measurement within a range of distance measurements and / or a range of errors (which may be referred to as an error tolerance). In some cases, the error tolerance can be kept constant (e.g., based on the ToF acquisition system (such as...). Figure 4 (The inherent characteristics of the ToF sensor 410). In an illustrative example, the depth imaging system 400 may determine that the distance measurement determined by the ToF system 408 is expected to have an error tolerance or range of 0.78% (e.g., a correct distance measurement result is expected to be within ±0.78% of the determined distance measurement). In this example, the depth imaging system 400 may determine the amplitude of the blur level 610 based on the 0.78% error tolerance. Typically, the depth imaging system 400 may determine a high blur level for a ToF distance measurement associated with a high error tolerance, and a low blur level for a ToF distance measurement associated with a low error tolerance. In some examples, the blur level 610 may be determined based on automatic exposure. For example, photon noise (also known as shot noise) is scene-dependent, and outdoor conditions worsen it; in this case, the blur level 610 may be associated with automatic exposure.
[0112] As shown in the figure, the ambiguity level 610 can be used to define the width of the search space 608 (e.g., Figure 6B(As shown). In some cases, the height of search space 608 may correspond to the height of primitive 606. In other cases, the height of search space 608 may be less than the height of primitive 606 (e.g., half the height of primitive 606, one-third the height of primitive 606, etc.). In the illustrative example, primitive 606 may have a height of 64 points and a width of 124 points, region 612 may have a height of 16 points and a width of 16 points, and search space 608 may have a height of 64 points and a width of 20 points (e.g., as defined by blur level 610). Depth imaging system 400 may determine a search space of any suitable size based on the size of primitive 606, the size of offset 620, and / or the size of blur level 610.
[0113] After defining the search space 608, the depth imaging system 400 can search the search space 608 to identify features including the corresponding feature 616. For example, Figure 6B The description covers region 612 at a first location within search space 608. Depth imaging system 400 can determine (e.g., based on convolution operations) whether metadata associated with a point in region 612 at the first location corresponds to image data in region 604. In one example, if depth imaging system 400 determines that metadata in region 612 corresponds to image data in region 604, then depth imaging system 400 can determine that a feature in region 612 corresponds to pixel 602. For example, depth imaging system 400 can determine the location of the corresponding feature (e.g., feature 616) within region 612, which corresponds to the location of pixel 602 in region 604. Depth imaging system 400 can then determine a structured optical distance measurement associated with pixel 602 based on the displacement between point 614 and feature 616. In other examples, if depth imaging system 400 determines that metadata in region 612 does not correspond to image data in region 604, then depth imaging system 400 can move region 612 to a second location within search space 608. For example, the depth imaging system 400 can be horizontal or vertical (e.g., as shown in the image). Figure 6BThe depth imaging system 400 then "slides" a point or feature up or down in region 612. The depth imaging system 400 can then compare the image data within region 612 at the second location with the image data of region 604. In some examples, the depth imaging system 400 may continue this process of analyzing image data at various locations within search space 608 until a feature in primitive 606 corresponding to pixel 602 is identified (or until every possible location in region 612 is analyzed). Using such a technique, the depth imaging system 400 can identify features in primitive 606 corresponding to pixel 602 without analyzing data of primitive 606 outside search space 608. For example, when performing the mixed-mode technique described herein, analyzing image data outside search space 608 may be unnecessary.
[0114] In some cases, the depth imaging system 400 can incorporate the structured light distance measurement associated with pixel 602 into the depth map 416. Furthermore, the depth imaging system 400 can repeat the process of determining the structured light distance measurement based on all or part of the initial ToF distance measurements of additional pixels in frame 618 associated with the light spot of pattern 412. By using ToF distance measurements as a “guide” for structured light decoding, the disclosed depth imaging system can improve the efficiency of structured light decoding without sacrificing accuracy.
[0115] The disclosed hybrid-mode depth imaging system can implement one or more techniques and / or processes to further improve the quality of the generated depth map. In one example, using a structured light source to obtain ToF distance measurements can improve the accuracy of ToF distance measurements. For example, the light spot projected by a structured light source (such as structured light source 404 of depth imaging system 400) can typically have a greater intensity than floodlight illumination used by a conventional ToF system. Because pattern 412 includes dark areas (e.g., areas not associated with the light spot), the light spot can consist of a more focused and / or concentrated light signal. Therefore, the light signal in frame 418 can have a greater signal-to-noise ratio (SNR) than light captured by a conventional ToF system. A greater SNR enables more accurate ToF distance measurements.
[0116] Furthermore, in some examples, the depth imaging system 400 can reduce the effects of multipath interference (MPI) and / or ambient light in frame 418. For example, MPI can occur when emitted light signals return to the sensor via two or more different paths, which can blur the emitted light signals returning to the sensor after being reflected by a single object. Ambient light signals are another source of blurring. In some cases, the depth imaging system 400 can determine (and then eliminate or reduce) the MPI signals and / or ambient light signals within frame 418. To determine the ambient light signals, the depth imaging system 400 can capture the frame exposure without projecting light from the structured light source 404. This frame exposure can represent a third frame exposure (in addition to frame exposures corresponding to low and high illumination levels, such as...). Figure 5A and Figure 5B (As shown). In some cases, the depth imaging system 400 can subtract the light signal from the third frame exposure from the light signals in the other two exposures before determining the ToF distance measurement based on the other two exposures. In this way, the depth imaging system 400 can eliminate or reduce noise from the ambient light signal within the ToF distance measurement.
[0117] Furthermore, in some examples, the amplitudes of the MPI signals within frame exposures corresponding to low and high illumination levels can differ. For example, high illumination levels can disproportionately introduce MPI signals into a frame exposure (e.g., relative to MPI signals introduced by low illumination levels). Because the depth imaging system 400 can be configured to determine ToF distance measurements based on the ratio of light amplitudes within the two frame exposures, eliminating or reducing the MPI signals within the two frame exposures can be beneficial. In one example, the depth imaging system 400 can eliminate or reduce the MPI signals within a frame exposure by determining the light signals within a region of the frame exposure that is not associated with the bright spots of pattern 412 (referred to as the dark regions of pattern 412). The depth imaging system 400 can then subtract the determined light signals from other regions of the frame exposures, resulting in a subtraction of the estimated MPI. Figure 7A An example frame exposure 702 including an MPI signal is illustrated. In this example, the depth imaging system 400 has subtracted the ambient light signal from the frame exposure. Therefore, the light signal within the pattern region 704 (e.g., a dark area of frame exposure 702 not associated with the light spot of pattern 412) corresponds to the MPI signal (rather than other light signals). The depth imaging system 400 can determine the amplitude (e.g., average amplitude) of the light signal within the pattern region 704 and can subtract the amplitude of light signals from other regions of frame exposure 702. For example, the depth imaging system 400 can subtract the light signal from all or a portion of the light spot of frame exposure 702. Figure 7BThe illustration shows a frame exposure 706 corresponding to the frame exposure 702 after which the MPI signal from frame exposure 702 is eliminated or reduced by subtracting the MPI from frame exposure 702 in the depth imaging system 400. As shown, the spot in frame exposure 706 is more defined (e.g., less noisy) than the spot in frame exposure 702. In some cases, reducing or eliminating the MPI signal and / or ambient light signal from the frame exposure can improve the accuracy of ToF distance measurement and / or structured light distance measurement.
[0118] In some cases, the depth imaging system 400 may perform one or more additional or alternative operations to reduce noise within frame 418. For example, the optical amplitude of a spot within frame 418 may be described and / or approximated by a Gaussian function (e.g., a Gaussian bell curve) or a similar function. In some cases, noise within the captured optical signal may cause the optical amplitude to deviate from the ideal Gaussian function. Figure 8 The illustration shows an example graph 802 of the captured amplitude 806 corresponding to the amplitude of the captured optical signal. Graph 802 also shows an ideal amplitude 804 generated based on fitting the captured amplitude 806 to a Gaussian function. In some cases, the depth imaging system 400 may fit all or a portion of the captured optical signal of frame 418 to an ideal Gaussian function before determining the ToF distance measurement associated with the signal. Furthermore, in one example, the depth imaging system 400 may determine functions corresponding to various codewords associated with frame 418. For example, the depth imaging system 400 may use the pattern of the codewords as a signal for denoising. In some cases, fitting the captured optical signal to a function corresponding to the codewords (which may be an ideal function) can further reduce noise within frame 418. These processes of fitting the optical signal to a function (e.g., an ideal function) can improve the noise characteristics of the ToF distance measurement, which can result in a smaller level of ambiguity associated with the ToF distance measurement (and therefore a smaller search space for structured light decoding).
[0119] Figure 9A This illustrates an instance of depth map 902, and Figure 9B An example of depth map 904 is illustrated. Depth maps 902 and 904 demonstrate the advantages of the disclosed hybrid-mode depth imaging system. For example, depth maps 902 and 904 represent hybrid-mode depth maps, ToF depth maps, and structured light depth maps generated under various ambient light conditions. Ambient light conditions include ideal conditions (e.g., no ambient light), 10,000 lumens ambient light, 50,000 lumens ambient light, and 100,000 lumens ambient light. As shown, the hybrid-mode depth map generated under each ambient light condition generally has higher quality (e.g., more accurate) than the ToF depth map and the structured light depth map.
[0120] Figure 10 This is a flowchart illustrating an example of a process 1000 for generating one or more depth maps using the techniques described herein. At block 1002, process 1000 includes obtaining a frame comprising a reflection pattern of light generated based on a light pattern emitted by a structured light source (e.g., light pattern 412 emitted by structured light source 404). The light pattern is based on primitives containing a set of uniquely identifiable features (or codewords), as described herein. In some aspects, the light pattern emitted by the structured light source comprises multiple light spots (e.g., corresponding to a primitive pattern). In some aspects, features within the set of uniquely identifiable features of the primitives include two or more light spots. In some cases, the light spots of the primitives correspond to two or more pixels of the frame. For example, each light spot (or point) may correspond to a 4×4, 8×8, or other arrangement of pixels in the frame. In some examples, the structured light source is configured to emit light patterns using diffractive optical elements that simultaneously project multiple light patterns corresponding to primitives (effectively repeating primitives, such as...). Figure 3C (The primitives shown).
[0121] At box 1004, process 1000 includes using a time-of-flight (ToF) sensor (e.g., ToF sensor 410) to determine a first distance measurement associated with a pixel of the frame. In an illustrative example, the first distance measurement may include the above-mentioned... Figure 6A The described ToF distance measurement. In some aspects, process 1000 includes obtaining a first exposure of a frame associated with a first illumination level (e.g., a first brightness), and obtaining a second exposure of a frame associated with a second illumination level (e.g., a second brightness) different from the first illumination level. As described above, the first illumination level (or brightness) may differ from the second illumination level (or brightness) based on how much reflected light the sensor captures for each of the frames (e.g., based on how far the light is reflected from an object). Process 1000 may include determining a first distance measurement associated with a pixel of the frame based at least in part on a comparison between a first light amplitude associated with a pixel in the first exposure and a second light amplitude associated with a pixel in the second exposure. In some aspects, process 1000 includes fitting a function to the light signal corresponding to the pixel of the frame before determining the first distance measurement associated with the pixel of the frame. Fitting a function to the light signal can improve the noise characteristics of the ToF distance measurement, which can result in a smaller level of ambiguity associated with the ToF distance measurement. A smaller level of ambiguity can result in a smaller search space for structured light decoding, which can allow for more efficient feature recognition within primitives.
[0122] At box 1006, process 1000 includes determining a search space within a primitive based at least in part on a first distance measurement. The search space includes a subset of features from a set of uniquely identifiable features of the primitive. For example, the search space may include... Figure 6B The search space 608 is shown. In some cases, the first distance measurement includes a distance measurement range. Process 1000 may include determining the size of the search space within the primitive based at least in part on the range of the distance measurement. For example, a large range of distance measurements may be associated with a large search space. In some cases, process 1000 may include at least in part based on the ambiguity level associated with the ToF sensor (e.g., Figure 6B The ambiguity level 610 shown determines the range of distance measurements. For example, a high ambiguity level is associated with a large distance measurement range.
[0123] In some aspects, process 1000 includes determining, at least in part, an offset between a first position of a pixel of a frame and a second position of a feature of a primitive, based on a first distance measurement. In one illustrative example, the offset may include... Figure 6B The offset 620 is shown. In some cases, the offset is inversely proportional to the first distance measurement. Process 1000 may include determining a search space within a primitive based at least in part on the offset. In some cases, process 1000 includes setting the central axis of the search space within the primitive to a second location of a feature of the primitive.
[0124] Process 1000 may include searching a search space within a primitive. At box 1008, process 1000 includes determining features of primitives corresponding to the region surrounding a pixel of a frame based on the search space within the primitive. In an illustrative example, each feature includes points arranged in a 4×4 pattern, and each point corresponds to a pixel arranged in a 4×4 pattern, resulting in the feature occupying 16×16 pixels in the captured frame when each point occupies 4×4 pixels. In such an example, the region surrounding a pixel of a frame may include a 16×16 pixel region surrounding a pixel of a frame, and primitives (and thus each feature within the primitive) may be searched on a 4×4 point basis.
[0125] At box 1010, process 1000 includes determining a second distance measurement associated with pixels of a frame, based at least in part on features of primitives determined from a search space within the primitives. In some cases, the second distance measurement may include a disparity value (e.g., Figure 3B The disparity is shown as 356. At box 1012, process 1000 includes generating a depth map at least in part based on a second distance measurement. For example, as described above, the depth is inversely proportional to the offset represented by the disparity (or the second distance measurement).
[0126] In some examples, the region surrounding the pixels of a frame has a predetermined size. In some cases, process 1000 includes determining the region surrounding the pixels of a frame with a predetermined size. Process 1000 may include determining a first region of a search space within a primitive with a predetermined size. Process 1000 can search the first region by determining whether image data within the region surrounding the pixels of the frame corresponds to image data (e.g., points of the primitive) within the first region of the search space. Process 1000 may include determining that image data within the region surrounding the pixels of the frame corresponds to image data within the first region of the search space. In this case, process 1000 may include determining a second distance measurement based at least in part on determining the distance between corresponding features of the pixels of the frame and the first region of the search space. In some cases, process 1000 may include determining that image data within the region surrounding the pixels of the frame does not correspond to image data within the first region of the search space within a primitive. In this case, process 1000 may include determining a second region of the search space. For example, as described above regarding... Figure 6B As described, the depth imaging system 400 can be horizontal or vertical (e.g., as...). Figure 6B The upward or downward "sliding" region 612 shown is a point or feature. The second region of the search space also has a predetermined size. Process 1000 may include determining whether image data within a region around a pixel of a frame corresponds to image data within the second region of the search space.
[0127] In some aspects, process 1000 may include obtaining additional frames when the structured light source does not emit primitive-based light patterns. Process 1000 may include determining an ambient light signal at least in part based on the additional frames. Process 1000 may include subtracting the ambient light signal from the frames before determining a first distance measurement associated with a pixel of the frame. In some cases, process 1000 may include using the frames after subtracting the ambient light signal from the frames to determine a light signal corresponding to multipath interference. Process 1000 may include subtracting the light signal corresponding to multipath interference from the frames before determining a first distance measurement associated with a pixel of the frame. In some cases, these aspects may be used to reduce the effects of multipath interference (MPI) and / or ambient light in one or more frames, as described above.
[0128] In some examples, the processes described herein (e.g., process 1000 and / or other processes described herein) may be performed by a computing device or apparatus. In some examples, process 1000 may be performed by... Figure 1 Depth imaging system 100 Figure 4 Depth imaging system 400 and / or Figure 11 The computational system 1100 executes the process. In one example, process 1000 can be performed by a computing system 1100. Figure 11The computing device or system with the architecture of the computing system 1100 shown executes the computation. For example, it has... Figure 11 The computing device of the computing system 1100 architecture may include Figure 1 Components and / or of the depth imaging system 100 Figure 4 The components of the depth imaging system 400, and can achieve Figure 10 The operation.
[0129] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, wearable devices (e.g., VR headsets, AR headsets, AR glasses, network-connected watches or smartwatches or other wearable devices), server computers, vehicles or computing components or devices of vehicles, robotic devices, televisions, and / or any other computing device with the resource capability to perform the processes described herein (including process 1000). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP) based data or other types of data.
[0130] Components of a computing device can be implemented in circuitry. For example, components may include electronic circuitry or other electronic hardware and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.
[0131] Process 1000 is shown as a logic flowchart, whose operations represent a series of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the operation. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0132] Furthermore, the process 1000 and / or other processes described herein can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors, via hardware, or a combination thereof. As described above, the code can be stored, for example, on a computer-readable or machine-readable storage medium in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0133] Figure 11 This is a diagram illustrating an example of a system used to implement certain aspects of this technology. Specifically, Figure 11 An example of computing system 1100 is shown, which can be any computing device, such as constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 1105. Connection 1105 can be a physical connection using a bus, or a direct connection to processor 1110, such as in a chipset architecture. Connection 1105 can also be a virtual connection, a networking connection, or a logical connection.
[0134] In some embodiments, the computing system 1100 is a distributed system, wherein the functions described herein may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent a plurality of such components, each performing some or all of the functions described for that component. In some embodiments, a component may be a physical or virtual device.
[0135] Example system 1100 includes at least one processing unit (CPU or processor) 1110 and a connection 1105 that couples various system components, including system memory 1115 (such as read-only memory (ROM) 1120 and random access memory (RAM) 1125), to processor 1110. Computing system 1100 may include a cache 1112 of high-speed memory that is directly connected to, adjacent to, or integrated into processor 1110.
[0136] Processor 1110 may include any general-purpose processor and hardware or software services, such as services 1132, 1134, and 1136 stored in storage device 1130, which are configured to control processor 1110 and dedicated processors, wherein software instructions are incorporated into the actual processor design. Processor 1110 may be a substantially independent computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0137] To enable user interaction, the computing system 1100 includes an input device 1145, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. The computing system 1100 may also include an output device 1135, which can be one or more of multiple output mechanisms. In some instances, a multimodal system allows the user to provide multiple types of input / output to communicate with the computing system 1100. The computing system 1100 may include a communication interface 1140, which typically controls and manages user input and system output. The communication interface can use wired and / or wireless transceivers to perform or facilitate the reception and / or transmission of wired or wireless communications, including using audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, proprietary wired ports / plugs, wireless signal transmission, low-power (BLE) wireless signal transmission, wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), global microwave access interoperability (WiMAX), and infrared (IR) wireless signal transmission. In some embodiments, the public switched telephone network (PSTN) signal transmission may be, for example, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 1140 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1100 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based Global Positioning System (GPS), the Russian-based Global Navigation Satellite System (GLONASS), the Chinese-based BeiDou Navigation Satellite System (BDS), and the European-based Galileo GNSS. There are no limitations on operation on any particular hardware arrangement, and therefore the basic features herein can be readily replaced with improved hardware or firmware arrangements as they are developed.
[0138] Storage device 1130 may be a non-volatile and / or non-transitory and / or computer-readable storage medium, and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as magnetic tape cassettes, flash memory cards, solid-state storage devices, digital universal disks, cassette tapes, floppy disks, flexible disks, hard disks, magnetic tapes, magnetic stripes / stripes, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, optical disc read-only memory (CD-ROM), rewritable optical disc (CD), digital video disc (DVD), Blu-ray disc (BDD), holographic disc, another optical medium, secure digital (SD) cards, microsecure digital (microSD) cards, memory stick cards, smart card chips, EMV chips, etc. Subscriber Identity Module (SIM) card, mini / micro / nano / micro SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cassette, and / or combinations thereof.
[0139] Storage device 1130 may include software services, servers, etc., which enable the system to perform functions when the code defining such software is executed by processor 1110. In some embodiments, hardware services that perform a particular function may include software components stored in a computer-readable medium and associated with necessary hardware components (such as processor 1110, connection 1105, output device 1135, etc.) to perform that function.
[0140] As used herein, the term "computer-readable medium" includes (but is not limited to) portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. A computer-readable medium may include non-transitory media in which data can be stored and which do not contain carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include (but are not limited to) magnetic disks or magnetic tapes, optical storage media (e.g., optical discs (CDs) or digital versatile optical discs (DVDs)), flash memory, memory, or memory devices. A computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent programs, functions, subroutines, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, variables, parameters, data, etc., can be transmitted, forwarded, or transferred using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.
[0141] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.
[0142] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that embodiments can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the accompanying drawings and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.
[0143] The various embodiments described above can be illustrated as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structural diagrams, or block diagrams. Although flowcharts can describe operations as sequential processes, many operations can be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0144] The processes and methods described in the examples above can be implemented using computer-executable instructions stored in or otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a particular function or group of functions. Parts of the computer resources used may be accessible via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include hard disks or optical disks, flash memory, USB devices equipped with non-volatile memory, networked storage devices, etc.
[0145] Devices implementing the processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. One or more processors may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mount devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. As another example, such functionality may also be implemented on a circuit board between different chips or different processes that execute in a single device.
[0146] Instructions, media for transporting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means of providing the functionality described in this disclosure.
[0147] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, while illustrative embodiments of this application have been described in detail herein, it should be understood that the inventive concept can be implemented and employed in other different ways, and the appended claims are intended to be construed as including such variations, in addition to being limited by the prior art. Various features and aspects of the above-described applications can be used individually or in combination. Furthermore, embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in a different order than that described.
[0148] Those skilled in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.
[0149] When a component is described as being “configured” to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform those operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform those operations, or any combination thereof.
[0150] The phrase “coupled to” means that any component is physically connected directly or indirectly to another component, and / or that any component communicates directly or indirectly with another component (e.g., via a wired or wireless connection, and / or other suitable communication interface).
[0151] The use of claim language or other languages to state "at least one" and / or "one or more" in a set indicates that one or more members of the set (in any combination) satisfy the claim. For example, claim language stating "at least one of A and B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The use of "at least one" and / or "one or more" in the language set does not limit the set to items listed in the set. For example, claim language stating "at least one of A and B" can mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0152] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application; however, such implementation decisions should not be construed as departing from the scope of this application.
[0153] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication handheld devices, or integrated circuit devices with multiple uses (including applications in wireless communication handheld devices and other devices). Any feature described as a module or component can be implemented together in an integrated logic device or separately as a discrete but interoperable logic device. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code containing instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may contain packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Alternatively or concurrently, the technology may be implemented at least in part by a computer-readable communication medium carrying or conveying program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagating signals or waves.
[0154] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. This processor can be configured to perform any of the techniques described herein. A general-purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or implementation suitable for the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0155] The illustrative aspects of this disclosure include:
[0156] Aspect 1: An apparatus for generating one or more depth maps, comprising: a structured light source configured to emit a light pattern based on primitives, the primitives containing a uniquely identifiable set of features; a time-of-flight (ToF) sensor; at least one memory; and one or more processors (e.g., implemented in a circuit) coupled to the at least one memory. The one or more processors are configured to: obtain a frame comprising a reflection pattern of light generated based on the light pattern emitted by the structured light source; determine a first distance measurement associated with a pixel of the frame using the ToF sensor; determine a search space within the primitives based at least in part on the first distance measurement, the search space comprising a subset of features from the uniquely identifiable set of features of the primitives; determine features of the primitives corresponding to a region surrounding the pixel of the frame based on searching the search space within the primitives; determine a second distance measurement associated with the pixel of the frame based at least in part on the features of the primitives determined from the search space within the primitives; and generate a depth map based at least in part on the second distance measurement.
[0157] Aspect 2: The apparatus according to aspect 1, wherein one or more processors are configured to: obtain a first exposure of a frame associated with a first illumination level; obtain a second exposure of the frame associated with a second illumination level different from the first illumination level; and determine a first distance measurement associated with the pixel of the frame based at least in part on a comparison between a first light amplitude associated with the pixel in the first exposure and a second light amplitude associated with the pixel in the second exposure.
[0158] Aspect 3: The apparatus according to any one of Aspect 1 or 2, wherein: the first distance measurement includes a distance measurement range; and the one or more processors are configured to determine the size of the search space within the primitive at least in part based on the distance measurement range, wherein a large distance measurement range is associated with a large size of the search space.
[0159] Aspect 4: The apparatus according to aspect 3, wherein the one or more processors are configured to determine the distance measurement range at least in part based on the ambiguity level associated with the ToF sensor, wherein a high ambiguity level is associated with a large distance measurement range.
[0160] Aspect 5: An apparatus according to any one of Aspects 1 to 4, wherein the one or more processors are configured to: determine, at least in part, an offset between a first position of the pixel of the frame and a second position of the feature of the primitive based on the first distance measurement, wherein the offset is inversely proportional to the first distance measurement; and determine, at least in part, the search space within the primitive based on the offset.
[0161] Aspect 6: The apparatus according to aspect 5, wherein the one or more processors are configured to set the central axis of the search space within the primitive as the second position of the feature of the primitive.
[0162] Aspect 7: An apparatus according to any one of Aspects 1 to 6, wherein the region surrounding a pixel of a frame has a predetermined size, and wherein one or more processors are configured to: determine a first region of a search space having a predetermined size; and determine whether image data within the region surrounding the pixel of the frame corresponds to image data within the first region of the search space.
[0163] Aspect 8: The apparatus according to aspect 7, wherein one or more processors are configured to: determine that image data in a region surrounding pixels of a frame corresponds to image data in a first region of a search space; and determine a second distance measurement based at least in part on the distance between pixels of a frame and corresponding features of the first region of the search space.
[0164] Aspect 9: The apparatus according to aspect 7, wherein the one or more processors are configured to: determine that the image data in the region surrounding the pixel of the frame does not correspond to the image data in the first region of the search space within the primitive; determine a second region of the search space having the predetermined size; and determine whether the image data in the region surrounding the pixel of the frame corresponds to the image data in the second region of the search space.
[0165] Aspect 10: The apparatus according to any one of Aspects 1 to 9, wherein: the light pattern emitted by the structured light source comprises a plurality of light spots; and the features within the uniquely identifiable feature set of the primitives comprise two or more of the plurality of light spots.
[0166] Aspect 11: The apparatus according to aspect 10, wherein the featured light spots correspond to two or more pixels of a frame.
[0167] Aspect 12: The apparatus according to any one of aspects 1 to 11, wherein the structured light source is configured to emit the light pattern using a diffractive optical element, the diffractive optical element simultaneously projecting a plurality of light patterns corresponding to the primitive.
[0168] Aspect 13: An apparatus according to any one of aspects 1 to 12, wherein the one or more processors are configured to: obtain an additional frame when the structured light source does not emit the light pattern based on the primitives; determine an ambient light signal based at least in part on the additional frame; and subtract the ambient light signal from the frame before determining the first distance measurement associated with the pixel of the frame.
[0169] Aspect 14: The apparatus according to aspect 13, wherein one or more processors are configured to: use the frame to determine an optical signal corresponding to multipath interference after subtracting an ambient light signal from the frame; and subtract the optical signal corresponding to multipath interference from the frame before determining the first distance measurement associated with the pixel of the frame.
[0170] Aspect 15: The apparatus according to any one of aspects 1 to 14, wherein one or more processors are configured to fit a function to an optical signal corresponding to a pixel of a frame before determining a first distance measurement associated with a pixel of a frame.
[0171] Aspect 16: The apparatus according to any one of aspects 1 to 15, wherein the apparatus includes a mobile device.
[0172] Aspect 17: The apparatus according to any one of aspects 1 to 16 further includes a display.
[0173] Aspect 18: A method for generating one or more depth maps, the method comprising: obtaining a frame containing a reflection pattern of light generated based on a light pattern emitted by a structured light source, the light pattern being based on primitives containing a set of uniquely identifiable features; determining a first distance measurement associated with a pixel of the frame using a time-of-flight (ToF) sensor; determining a search space within the primitives based at least in part on the first distance measurement, the search space including a subset of features from the set of uniquely identifiable features of the primitives; determining features of the primitives corresponding to a region surrounding the pixel of the frame based on searching the search space within the primitives; determining a second distance measurement associated with the pixel of the frame based at least in part on the features of the primitives determined from the search space within the primitives; and generating a depth map based at least in part on the second distance measurement.
[0174] Aspect 19: The method according to aspect 18 further includes: obtaining a first exposure of a frame associated with a first illumination level; obtaining a second exposure of the frame associated with a second illumination level different from the first illumination level; and determining a first distance measurement associated with the pixel of the frame based at least in part on a comparison between a first light amplitude associated with the pixel in the first exposure and a second light amplitude associated with the pixel in the second exposure.
[0175] Aspect 20: The method according to any one of Aspects 18 or 19, wherein the first distance measurement includes a distance measurement range, and further includes determining the size of the search space within the primitive at least in part based on the distance measurement range, wherein a large range of distance measurements is associated with a large size of the search space.
[0176] Aspect 21: The method according to aspect 20 further includes determining the range of distance measurements based at least in part on the ambiguity level associated with the ToF sensor, wherein a high ambiguity level is associated with a large distance measurement range.
[0177] Aspect 22: The method according to any one of aspects 18 to 21 further includes: determining, at least in part, an offset between a first position of the pixel of the frame and a second position of the feature of the primitive based on the first distance measurement, wherein the offset is inversely proportional to the first distance measurement; and determining, at least in part, the search space within the primitive based on the offset.
[0178] Aspect 23: The method according to aspect 22 further includes setting the central axis of the search space within the primitive as a second position of a feature of the primitive.
[0179] Aspect 24: The method according to any one of aspects 18 to 23, wherein the region surrounding a pixel of a frame has a predetermined size, the method further comprising: determining a first region of a search space having a predetermined size; and determining whether image data within the region surrounding the pixel of the frame corresponds to image data within the first region of the search space.
[0180] Aspect 25: The method according to aspect 24 further includes: determining that image data in a region surrounding a pixel of a frame corresponds to image data in a first region of a search space; and determining a second distance measurement based at least in part on determining the distance between a pixel of a frame and a corresponding feature of the first region of the search space.
[0181] Aspect 26: The method according to aspect 24 further includes: determining that the image data in the region surrounding the pixel of the frame does not correspond to the image data in the first region of the search space within the primitive; determining a second region of the search space having the predetermined size; and determining whether the image data in the region surrounding the pixel of the frame corresponds to the image data in the second region of the search space.
[0182] Aspect 27: The method according to any one of Aspects 18 to 26, wherein: the light pattern emitted by the structured light source comprises a plurality of light spots; and the features within the uniquely identifiable feature set of the primitives comprise two or more of the plurality of light spots.
[0183] Aspect 28: According to the method of aspect 27, the light spot of the feature corresponds to two or more pixels of the frame.
[0184] Aspect 29: The method according to any one of aspects 18 to 28 further includes: using the structured light source, emitting the light pattern using a diffractive optical element, the diffractive optical element simultaneously projecting a plurality of light patterns corresponding to the primitive.
[0185] Aspect 30: The method according to any one of aspects 18 to 29 further includes: obtaining an additional frame when the structured light source does not emit a primitive-based light pattern; determining an ambient light signal based at least in part on the additional frame; and subtracting the ambient light signal from the frame before determining the first distance measurement associated with the pixel of the frame.
[0186] Aspect 31: The method according to aspect 30 further includes: using the frame to determine an optical signal corresponding to multipath interference after subtracting an ambient light signal from the frame; and subtracting the optical signal corresponding to multipath interference from the frame before determining the first distance measurement associated with the pixel of the frame.
[0187] Aspect 32: The method according to aspect 30 further includes: fitting a function to an optical signal corresponding to the pixel of the frame before determining the first distance measurement associated with the pixel of the frame.
[0188] Aspect 33: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform any of the operations described in aspects 1 to 32.
[0189] Aspect 34: An apparatus comprising means for performing any of the operations of aspects 1 to 32.
Claims
1. An apparatus for generating one or more depth maps, comprising: A structured light source configured to emit light patterns based on primitives, the primitives comprising a unique set of identifiable features; Time-of-flight (ToF) sensor; At least one memory; as well as One or more processors coupled to the at least one memory, the one or more processors being configured to: Obtain a frame including a reflection pattern of light generated based on the light pattern emitted by the structured light source; The ToF sensor is used to determine a first distance measurement associated with the pixels of the frame; The width of the search space within the primitive is determined based on the ambiguity level associated with the first distance measurement determined using the ToF sensor. The search space includes a subset of features from the uniquely identifiable feature set of the primitive, wherein a larger ambiguity level is associated with a larger width of the search space. Based on searching the search space within the primitives, the features of the primitives corresponding to the region surrounding the pixel of the frame are determined; A second distance measurement associated with the pixel of the frame is determined at least in part based on the features of the primitive determined from the search space within the primitive; and A depth map is generated, at least in part, based on the second distance measurement.
2. The apparatus of claim 1, wherein the one or more processors are configured to: Obtain the first exposure of the frame associated with the first illumination level; Obtain a second exposure of the frame associated with a second illumination level different from the first illumination level; and The first distance measurement associated with the pixel of the frame is determined at least in part based on a comparison between a first light amplitude associated with the pixel in the first exposure and a second light amplitude associated with the pixel in the second exposure.
3. The apparatus according to claim 1, wherein: The first distance measurement is determined to be within the range of distance measurements with error tolerance; as well as The one or more processors are configured to determine the size of the search space within the primitive based at least in part on the range of the distance measurements, wherein the range of the distance measurements is positively correlated with the size of the search space.
4. The apparatus of claim 1, wherein the one or more processors are configured to: Based at least in part on the first distance measurement, determine the offset between a first position of the pixel of the frame and a second position of the feature of the primitive, wherein the offset is inversely proportional to the first distance measurement; and The search space within the primitive is determined at least in part based on the offset.
5. The apparatus of claim 4, wherein the one or more processors are configured to set the central axis of the search space within the primitive as the second position of the feature of the primitive.
6. The apparatus of claim 1, wherein the region surrounding the pixel of the frame has a predetermined size, and wherein the one or more processors are configured to: Determine a first region of the search space, the first region of the search space having the predetermined size; and Determine whether the image data in the region surrounding the pixel of the frame corresponds to the image data in the first region of the search space.
7. The apparatus of claim 6, wherein the one or more processors are configured to: Determine that the image data within the region surrounding the pixel of the frame corresponds to the image data within the first region of the search space; and The second distance measurement is determined at least in part based on determining the distance between the pixel of the frame and the corresponding feature of the first region of the search space.
8. The apparatus of claim 6, wherein the one or more processors are configured to: The image data in the region surrounding the pixel of the frame does not correspond to the image data in the first region of the search space within the primitive; A second region of the search space is determined, the second region of the search space having the predetermined size; as well as Determine whether the image data in the region surrounding the pixel of the frame corresponds to the image data in the second region of the search space.
9. The apparatus according to claim 1, wherein: The light pattern emitted by the structured light source includes multiple light spots; and The features within the uniquely identifiable feature set of the graphic element include two or more light spots from a plurality of light spots.
10. The apparatus of claim 9, wherein the light spot of the feature corresponds to two or more pixels of the frame.
11. The apparatus according to claim 1, wherein, The structured light source is configured to emit the light pattern using diffractive optical elements, which simultaneously project multiple light patterns corresponding to the primitive.
12. The apparatus of claim 1, wherein the one or more processors are configured to: An additional frame is obtained when the structured light source does not emit the light pattern based on the primitives; The ambient light signal is determined at least in part based on the additional frame; and Before determining the first distance measurement associated with the pixel of the frame, the ambient light signal is subtracted from the frame.
13. The apparatus of claim 12, wherein the one or more processors are configured to: After subtracting the ambient light signal from the frame, the frame is used to determine the optical signal corresponding to multipath interference; and Before determining the first distance measurement associated with the pixel of the frame, the optical signal corresponding to multipath interference is subtracted from the frame.
14. The apparatus of claim 1, wherein the one or more processors are configured to fit a function to the light signal corresponding to the pixel of the frame before determining the first distance measurement associated with the pixel of the frame.
15. The apparatus according to claim 1, wherein, The device includes a mobile device.
16. The apparatus of claim 1, further comprising a display.
17. A method for generating one or more depth maps, the method comprising: Obtain frames comprising reflection patterns of light generated based on light patterns emitted by structured light sources, the light patterns being based on primitives comprising a set of uniquely identifiable features; A first distance measurement associated with the pixels of the frame is determined using a time-of-flight (ToF) sensor; The width of the search space within the primitive is determined based on the ambiguity level associated with the first distance measurement determined using the ToF sensor. The search space includes a subset of features from the uniquely identifiable feature set of the primitive, wherein a larger ambiguity level is associated with a larger width of the search space. Based on searching the search space within the primitives, the features of the primitives corresponding to the region surrounding the pixel of the frame are determined; A second distance measurement associated with the pixel of the frame is determined at least in part based on the features of the primitive determined from the search space within the primitive; and A depth map is generated, at least in part, based on the second distance measurement.
18. The method of claim 17, further comprising: Obtain the first exposure of the frame associated with the first illumination level; Obtain a second exposure of the frame associated with a second illumination level different from the first illumination level; as well as The first distance measurement associated with the pixel of the frame is determined at least in part based on a comparison between a first light amplitude associated with the pixel in the first exposure and a second light amplitude associated with the pixel in the second exposure.
19. The method of claim 17, wherein, The first distance measurement is determined to be within a distance measurement range with an error tolerance, and the method further includes: The size of the search space within the primitive is determined at least in part based on the range of the distance measurement, wherein the range of the distance measurement is positively correlated with the size of the search space.
20. The method of claim 17, further comprising: Based at least in part on the first distance measurement, determine the offset between a first position of the pixel of the frame and a second position of the feature of the primitive, wherein the offset is inversely proportional to the first distance measurement; and The search space within the primitive is determined at least in part based on the offset.
21. The method of claim 20, further comprising setting the central axis of the search space within the primitive as the second position of the feature of the primitive.
22. The method of claim 17, wherein the region surrounding the pixel of the frame has a predetermined size, the method further comprising: A first region of the search space is determined, the first region of the search space having the predetermined size; as well as Determine whether the image data in the region surrounding the pixel of the frame corresponds to the image data in the first region of the search space.
23. The method of claim 22, further comprising: The image data in the region surrounding the pixel of the frame corresponds to the image data in the first region of the search space; as well as The second distance measurement is determined at least in part based on determining the distance between the pixel of the frame and the corresponding feature of the first region of the search space.
24. The method of claim 22, further comprising: The image data in the region surrounding the pixel of the frame does not correspond to the image data in the first region of the search space within the primitive; A second region of the search space is determined, the second region of the search space having the predetermined size; as well as Determine whether the image data in the region surrounding the pixel of the frame corresponds to the image data in the second region of the search space.
25. The method of claim 17, wherein: The light pattern emitted by the structured light source includes multiple light spots; and The features within the uniquely identifiable feature set of the graphic element include two or more light spots from a plurality of light spots.
26. The method according to claim 25, wherein, The light spot of the feature corresponds to two or more pixels of the frame.
27. The method of claim 17, further comprising: Using the structured light source, the light pattern is emitted using diffractive optical elements, which simultaneously project multiple light patterns corresponding to the primitive.
28. The method of claim 17, further comprising: An additional frame is obtained when the structured light source does not emit the light pattern based on the primitives; The ambient light signal is determined at least in part based on the additional frame; as well as Before determining the first distance measurement associated with the pixel of the frame, the ambient light signal is subtracted from the frame.
29. The method of claim 28, further comprising: After subtracting the ambient light signal from the frame, the frame is used to determine the optical signal corresponding to multipath interference; as well as Before determining the first distance measurement associated with the pixel of the frame, the optical signal corresponding to multipath interference is subtracted from the frame.
30. The method of claim 17, further comprising: Before determining the first distance measurement associated with the pixel of the frame, a function is fitted to the light signal corresponding to the pixel of the frame.
31. An apparatus for generating one or more depth maps, comprising: Components for obtaining frames comprising a reflection pattern of light generated based on a light pattern emitted by a structured light source, the light pattern being based on primitives comprising a set of uniquely identifiable features; Components for determining a first distance measurement associated with pixels of the frame using a time-of-flight (ToF) sensor; Components for determining the width of a search space within the primitive based on an ambiguity level associated with the first distance measurement determined using the ToF sensor, the search space comprising a subset of features from the uniquely identifiable feature set of the primitive, wherein a larger ambiguity level is associated with a larger width of the search space; A component for determining features of a primitive corresponding to a region surrounding a pixel in a frame based on searching the search space within the primitive; Components for determining a second distance measurement associated with the pixel of the frame based at least in part on the features of the primitive determined from the search space within the primitive; as well as Components for generating a depth map based at least in part on the second distance measurement.
32. A computer-readable medium storing program code for generating one or more depth maps at a device, wherein, The program code can be executed by one or more processors to enable the device: Obtain frames comprising reflection patterns of light generated based on light patterns emitted by structured light sources, the light patterns being based on primitives comprising a set of uniquely identifiable features; A first distance measurement associated with the pixels of the frame is determined using a time-of-flight (ToF) sensor; The width of the search space within the primitive is determined based on the ambiguity level associated with the first distance measurement determined using the ToF sensor. The search space includes a subset of features from the uniquely identifiable feature set of the primitive, wherein a larger ambiguity level is associated with a larger width of the search space. Based on searching the search space within the primitives, the features of the primitives corresponding to the region surrounding the pixel of the frame are determined; A second distance measurement associated with the pixel of the frame is determined at least in part based on the features of the primitive determined from the search space within the primitive; as well as A depth map is generated, at least in part, based on the second distance measurement.
Citation Information
Patent Citations
Time-of-flight augmented structured light range-sensor
US20190079192A1
Mixed active depth
US20200386540A1