Methods and systems for depth and memory foveation of SPAD cameras

Foveation techniques in SPAD-based LiDAR systems address computational bottlenecks by reducing bins around an estimated peak, enhancing efficiency and SNR, thus optimizing power and memory usage.

WO2026049863A2PCT designated stage Publication Date: 2026-03-05PORTLAND STATE UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Existing SPAD-based LiDAR systems face challenges with ambient light susceptibility and large amounts of raw photon data, leading to computational bottlenecks, increased power consumption, and reduced signal-to-noise ratio (SNR) due to the processing of high-resolution histograms.

Method used

Implementing foveation techniques to reduce the number of bins in histograms, centering them around an estimated peak, and discarding data outside the foveation window to improve computational efficiency, reduce memory usage, and maintain high SNR.

Benefits of technology

This approach reduces processing time, power consumption, and memory requirements while maintaining depth resolution and improving SNR, making it suitable for power-constrained systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025036692_05032026_PF_FP_ABST
    Figure US2025036692_05032026_PF_FP_ABST
Patent Text Reader

Abstract

The application relates to methods and systems for determining depth of objects in a scene using single-photon sensing detectors. The method may include receiving, at a histogrammer, a stream of photon return events from a pixel of an imaging detector, the stream of photon return events generated by photons transmitted from a pulsed light source and reflected off an object in a scene, binning, with the histogrammer, only a subset of the photon return events into a plurality of bins of a histogram, the plurality of bins centered around an estimated histogram peak, and outputting, from the histogrammer, the histogram, the histogram usable to determine a distance of the object in the scene.
Need to check novelty before this filing date? Find Prior Art

Description

Docket No. PSU24302PCTMETHODS AND SYSTEMS FOR DEPTH AND MEMORY FOVEATION OF SPAD CAMERASCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 668,731, entitled “METHODS AND SYSTEMS FOR DEPTH AND MEMORY FOVEATION OF SPAD CAMERAS,” and filed July 8, 2024, the entire contents of which are hereby incorporated by reference for all purposes.GOVERNMENT SUPPORT

[0002] This invention was made with government support under Grant No. 2138471 awarded by the National Science Foundation, Grant No. 1942444 awarded by the National Science Foundation, and Grant No. N00014-23-1-2429 awarded by the Office of Naval Research. The U.S. Government has certain rights in the invention.FIELD

[0003] This disclosure relates to image processing, more particularly to foveating histograms captured by single-photon sensing 3D cameras.BACKGROUND

[0004] The use of digital video, imaging sensors, and distance sensors has become widespread and is still expanding in many industries and fields of endeavor. Digital video and sensor data may be combined and processed through various means to measure and compute depth. Fast, efficient, and accurate depth-sensing is desired for navigation and contact prevention applications, such as for driving autonomous vehicles. Efficiency may be in this context is related to efficiency of computation, and may be related to both power consumed and / or time taken for a computer operation to complete. Accuracy is related to the accuracy of a depth for a pixel compared to a pixel of ground truth image or a ground truth measurement at a point on the pixel. Direct time-of-flight light detection and ranging (LiDAR) has the potential to fulfill these demands, thanks to its ability to provide high-precision depth measurement at long standoff distances. More specifically, LiDAR may use a laser as a light source for imaging and be referred to alternatively as laser imaging, detection, and ranging. While conventional LiDAR systems and sensors rely on avalanche photodiodes (APDs). single-photon avalanche diodes (SPADs) are an emerging image-sensing technology. Raw data captured by an array of SPAD pixels may be thought of as a spatio-temporal photon stream. Each photon detection is represented as an x, y, t coordinate, where the x and y coordinates denote the pixel location and the t denotes the photon detection timestamp. Each SPAD pixel captures a round trip time of a laser pulse to and from a given scene point. The data captured by each pixel may be used to construct a photon timing histogram. The photon timing histogram records the number of photons at various time delays with respect to the timeDocket No. PSU24302PCT die laser pulse was transmitted. Each pixel may have a histogram that comprises a thousand or more bins (e.g., data points). The depth of each scene point may be determined based on each histogram, such as based on the peak of each histogram.SUMMARY

[0005] As discussed further herein below, various systems and methods are provided that significantly improve the computation of depth maps (e.g., distance maps) of an imaged scene using a single-photon-sensing 3D camera. In one embodiment, a method comprises receiving, at a histogrammer. a stream of photon return events from a pixel of an imaging detector, the stream of photon return events generated by photons transmitted from a pulsed light source and reflected off an object in a scene; binning, with the histogrammer, only a subset of the photon return events into a plurality of bins of a histogram, the plurality of bins centered around an estimated histogram peak; and outputting, from the histogrammer. the histogram, wherein the histogram is usable to determine a distance of the object in the scene.

[0006] This approach leverages: the advantages in reducing processing time, computation power, and memory' for generating depth maps with a SPAD system using fewer bins, and the advantages of image depth resolution and maintaining a high signal-to-noise ratio (SNR). This foveated approach allows the method of the present disclosure to “zoom into" the signal of interest, reducing the amount of raw photon data to be stored and transferred from the SPAD sensor. Said in another way, the foveated approach allows the method to bin data only in a range that includes the peak to be stored and transferred from the SPAD sensor. Data outside of the range may be mostly noise and may be removed, therein removing noise data from the data to be stored and transferred from the SPAD sensor. The foveated approach may increase resilience of reflected light detected by a SPAD to ambient light and other sources of noise. Said in another way , reflected light detected by a SPAD may be more distinguishable, where peaks of reflected light may not be as dampened by ambient light and other sources of noise compared to a method using the same amount of bins in a histogram without the foveated approach.

[0007] The foveation algorithm may calculate a reduced number of bins to use for creating a histogram and may center a foveation window (e.g.. the reduced number of bins) around an estimated depth of the scene point (e.g.. around an estimated or predicted peak of the histogram). Data at and within a selected range of distances from the estimated depth are then binned, while data outside the range are assumed to be noise data and not binned. The noise data is removed from the histogram and the data binned data is used to create a histogram for the pixel. The histogram may be used to assign the pixel a depth value.

[0008] It should be understood that the brief description above is provided to introduce in simplified form a selection of concepts that are further described in the detailed description. It is not meant to identify key or essential features of the claimed subject matter, the scope of which is definedDocket No. PSU24302PCT uniquely by the claims that follow the detailed description. Furthermore, the claimed subject matter is not limited to implementations that solve any disadvantages noted above or in any part of this disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0010] The disclosure may be better understood from reading the following description of nonlimiting embodiments, with reference to the attached drawings, wherein below:

[0011] FIG. 1 shows a schematic of a system that may capture and process data to generate a depth images via a SPAD-based LiDAR;

[0012] FIG. 2A shows a top-level view of an architecture for processing images using foveation techniques, in accordance with one or more embodiments of the present disclosure;

[0013] FIG. 2B a table of non-foveated images compared to foveated images using the architecture of FIG. 2A.

[0014] FIG. 3 shows a table of sets non-foveated and foveated images for comparison, where foveated images are created from a monocular depth images using memory' foveation and depth foveation techniques.

[0015] FIG. 4 shows a table of sets of non-foveated and foveated images for comparison, where the foveated images are created via spatio-temporal foveation techniques, including memory foveation and depth foveation techniques.

[0016] FIG. 5 shows a table of sets of non-foveated and foveated images for comparison, where the foveated images are created using optical flow techniques.

[0017] FIG. 6A shows a set of non-foveated and foveated images for comparison, the images created via hardw are emulation.

[0018] FIG. 6B show s a set of graphs showing non-foveated and foveated of histograms for a pixel of an image.

[0019] FIG. 6C shows a set of red blue green (RGB), ground truth depth, and super pixel segmentation images.

[0020] FIG. 7 shows a graph of the relationship between signal-to-background ratio (SBR) and root mean square (RMSE) for non-foveated and foveated images.

[0021] FIG. 8 shows a lower level view of an architecture for processing images using foveation techniques.

[0022] FIG. 9 show s a method of collecting and foveating photon data into a histogram.

[0023] FIG. 10 shows a method to generate a depth image using memory' foveation and depth foveation techniques of the present disclosure.Docket No. PSU24302PCT

[0024] FIG. 11 shows a method for foveating and generating a histogram for pixels of the depth image using memory foveation or depth foveation.

[0025] FIG. 12 shows a method for generating a depth prior using optical flow.

[0026] FIG. 13 shows a method for generating a depth prior from super pixel sampling techniques.

[0027] FIG. 14 shows a method for generating a depth prior from a monocular image.

[0028] FIG. 15 shows schematically an example of a control and signal diagram of one or moreSPADs for hardware of the present disclosure.

[0029] FIG. 16 shows schematically an example of a SPAD detection array of the present disclosure.DETAILED DESCRIPTION

[0030] The following description relates to processing of photon data received from a light detection and ranging (LiDAR) system using single-photon avalanche diodes (SPADs), referred to herein as a SPAD-based LiDAR. The SPAD-based LiDAR is part of a larger system that may processes data using computational methods and processors from the SPAD-based LiDAR, referred to herein as a SPAD system. The SPAD system may produce SPAD-generated images, such as depth images. A depth image may include a depth map and / or a depth image from a monocular image and the depth map. During capture, the method uses external signals to foveate or guide how the SPAD system produces depth. The method may foveate using algorithms disclosed herein. This foveated approach of the method allows our method to “zoom into” the signal of interest, reducing the amount of raw photon data to be stored and transferred from the SPAD sensor, while also improving resilience to ambient light and other sources of noise. Results of foveation may be shown in both in simulation and also with real hardware emulation. The algorithms can be applied to newly available and upcoming SPAD designs.

[0031] The present disclosure takes inspiration from biology. Many animals have regions of high special acuity (e.g., the fovea) which may scan over a scene, such as a landscape or an enclosed space. The present disclosure allies with foveated imaging research in computer vision and computational photography.

[0032] There are challenges to widespread adoption of SPAD-based LiDARs. One such example includes the susceptibility of SPAD to ambient light and a large amount of raw photon data that is processed to obtain in-pixel depth estimates. Such a large amount of raw photon may generate a data bottleneck such as when communicating and / or processing the data. The histogram of each pixel and the bins of each histogram is calculated via computational means, such as via algorithms. At higher resolutions, the number of pixels for an image to be processed increases, increasing the computational power that may be demanded to determine the depth of each point in an imaged scene. Likewise, greater variations of depth and greater distances from the SPAD LiDAR being measured may increase the binsDocket No. PSU24302PCT per histogram and variations of bins per histogram, increasing die computation power that may be demanded for determining the depth of the points in the imaged scene. Further, at higher resolutions, die memory to store the photon data / histograms may be increased compared to an image from the APDs, due to the quantity of data assigned to each pixel with a full resolution histogram. An increase in computational power to determine depth may increase the power consumption of the SPAD LiDAR system and may increase the time to process an image via the SPAD LiDAR system.

[0033] To reduce the processing of the photon data to determine depth, the number of bins per histogram may be reduced, such as via reducing number of bins from the thousands to an eighth or to a sixteenth of the original number of bins. Reducing the bins may decrease computational power, time for processing, and memory used for each histogram compared to depth determinations processed using a higher amount of bins. However, reducing the number of bins may decrease the depth resolution, where the range of depths per bin may increase. Further, reducing the number of bins may decrease the signal to noise ratio (SNR). Noise from data received and processed via the SPAD LiDAR may be received from multiple sources, such as ambient light. As the SNR decreases, peaks on a histogram may become less distinct from noise. As the SNR decreases, the quality of the depth image may decrease, as peaks in photons on the histogram appear more similar to noise and noise may be detected as peaks. At an SNR below a threshold, the depth of the image may be indistinguishable from noise. Both the increased range of depth per bin and the increased SNR ratio may result in an image processed using less bins to have a poorer and flatter resolution of depth and more incorrect depths compared to a ground truth three dimensional image (e.g., using the full histogram resolution).

[0034] Thus, according to embodiments disclosed herein, foveation techniques may be applied to reduce the total number of bins of each histogram without sacrificing depth resolution by centering the reduced number of bins (also referred to as the foveation window) around an estimated or predicted peak of the histogram using a depth prior. In other examples, the full number of bins may be used but the bin width of each bin may be reduced and the bins centered around the estimated or predicted peak of the histogram. In either case, photon events that fall outside the foveation window may be discarded and not included in the histogram, which may reduce the amount of data generated per pixel of the SPAD array.

[0035] A theoretical model is presented for expected gains (in terms of increased signal-to-noise- ratio and depth resolution) from foveation with SPADs. A method is shown of how to foveate in space and time at a single time instant by leveraging monocular depth estimates, which can come either from the SPAD-generated image or an external color camera. The method may be referred to herein as FoveaSPAD. Different categories of FoveaSPAD techniques may be used where each is optimized for memory / bandwidth and depth resolution. For example, optimization for some categories may decrease memory / bandwidth consumption more than increasing depth resolution compared to a non-FoveaSPAD method. Optimization for other categories may increase depth resolution more than decreasingDocket No. PSU24302PCT memory / bandwidth consumption compared to a non-FoveaSPAD method. For images of moving scenes, the use optical flow cues to direct SPAD foveation is demonstrated. A worst-case analysis of die limits of foveation with ambient light in SPAD sensors are shown. Results both in simulation and using recently available real SPAD dataset are shown.

[0036] In contrast to other approaches, the work of the present disclosure may scale with a large number of SPAD pixels. Other efforts, including partial histogram methods such as methods using sliding windows for sub-range gating, have been investigated which has linear efficiency and two-stage coarse-to-fine resolution scaling which provide logarithmic efficiency. Methods of the present disclosure uses context from cues such as optical flow to provide approximately near-constant time efficiency. Finally, other work has used external sensors for guided upsampling or upscaling, but these are post-capture processes of photons via SPADs. In contrast, foveation using FoveaSPAD may be performed during capture of photons via SPADs.

[0037] A complementary approach to foveation is to use an adaptive histogramming approach for the signal peak. The approach of the present disclosure is also complementary to adaptive gating approaches for SPAD that often assume a uniform prior on the true scene depth, whereas memory and depth foveation may provide evaluations for less uniform priors.

[0038] The work of the present disclosure, including FoveaSPAD, may be used via post-capture methods for up-sampling and super-resolution shown on data from many modes, such as depth images, color photographs etc. These post-capture methods may include using deep learning algorithms into the process of deciding where to sample. In fact, some of these algorithms are mature enough that commercial depth and LIDAR sensors allow post-capture foveation of the 3D point cloud through, for example, LIDAR-RGB fusion. In contrast to methods of prior art, FoveaSPAD may adapt during capture, and the efficiencies can effect small autonomous systems with power constraints. Directionally controlled LIDAR systems foveate spatially.

[0039] Different variations of the FoveaSPAD method may be referred to herein as different flavors. Different flavors of practical FoveaSPAD methods may be used where each is optimized for mcmory / bandwidth and depth resolution, such optimization for certain flavors may decrease memory / bandwidth consumption and optimization for other flavors may increase depth resolution. Results both in simulation and using recently available real SPAD datasets are shown.

[0040] A summary the mathematical symbols the algorithms and other mathematical models the method may use to foveate SPAD data is shown in TABLE I.TABLE I. TABLE OF SYMBOLS AND EXPLANATIONSDocket No. PSU24302PCT\

[0041] Simulations used for testing the method of foveation may be created with the SPAD simulation framework of “Gutierrez-Barragan, H. Chen, M. Gupta. A. Velten, and J. G. “itof2dtof: A robust and flexible representation for data-driven time-of-flight imaging,’’ IEEE Transactions on Computational Imaging, vol. 7, pp. 1205-1214, 2021” and “Gutierrez-Barragan, A. Ingle, T. Sects, M. Gupta, and A. Velten, “Compressive single-photon 3d cameras,” in Proceedings of the IEEE ('Vh Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 854-17 864 each of which is hereby incorporated by reference for all purposes, using the GitHub code available. While the simulation is initialized with red blue green depth (RGBD) datasets, all the “ground truth” depth images in this disclosure are the result of SPAD simulation on full high-resolution histograms. The simplified imaging model described below assumes that all laser photons arrive in a single bin i. In practice, the laser pulse spans several bins “smearing” the signal photons over more than one bin. The laser peak is often modeled as a Gaussian shaped pulse; where the simulation results use a 1 nanosecond full width at half maximum (FWHM). It is also possible to obtain a pseudo intensity image by aggregating photon counts across histograms for each pixel which can be used in lieu of a collocated red blue green (RGB) or monochrome camera image for monocular depth cues.Docket No. PSU24302PCT

[0042] Each pixel in the SPAD sensor array may be co-located with a pulsed laser illumination source via a Gaussian pulse shape. Assuming no multi-path or subsurface scattering effects, the photon flux incident on each pixel comprises a superposition of laser photons (that arrive in a short time window corresponding to the round-trip time-of-flight to and from the scene point) and background photons due to ambient light (that arrive uniformly randomly distributed throughout the capture duration). The laser repetition period (T) determines the maximum depth range of the SPAD LiDAR. The period T is discretized into N bins (N is often on the order of 1000’s of bins in conventional SPAD cameras). The number of photons captured by the SPAD pixel in the nth bin (1 < n < N) is Poisson distributed with a mean of d’sig l (n = i) + d’bk. where i is the bin location corresponding to the true scene depth. Various sources of noise such as dark counts and afterpulsing are assumed to be absorbed in the d’bkg term. A full histogram (e.g., a histogram from a trace that has all photons and points of time that are recorded binned) captured by this SPAD pixel over C laser cycles is given by a Poisson random vector with mean C ’sig 1 (n = i) + C bkg for 1 < i < N.

[0043] The time taken for each foveation experiment disclosed herein is the same. The experiments were to show how foveation may save memory or improve depth resolution, and how the signal-to- noise ratio changes depending on ambient light, bin width, and other considerations. A SPAD pixel may image a scene point illuminated by a pulsed laser. The scene point may be a point or an area representing a point on a surface scanned by the LiDAR via the laser. Assuming there are no multi-bounce effects and no ambient light, photon detections from the SPAD pixel are used to generate a histogram of arrival times. A conventional approach would use all N bins across the full histogram, whereas the method of die present disclosure foveates attention onto a subset M < N of these bins. Therefore, via the SNR analysis of the SPAD system, the ratio M / N appears since this represents the advantage due to foveation. Assumptions are not made as to how the foveated bins M were obtained, instead the effect of foveated bins M are explored on imaging, such as resolution and accuracy compared to a ground truth depth. The method shows algorithms to drive the selection of the foveated bins M. A worst case analysis is shown for whether the foveated bins M capture the histogram peak or not.

[0044] In an imaging case, the SPAD sensor detects time-of-arrival of photons and accumulates into a photon timing histogram to find a time that corresponds to the true depth of the scene point. If the histogram has a full scale range of T seconds which is related to the maximum unambiguous depth range Z as T = Z / c where c is the speed of light. Consider N histogram bins that are uniformly distributed across the full scale range T. The width of each bin is T / N. Since narrower bins produce fewer photon reads, the SNR for each bin is proportional to the width of that time bin as represented by equation 1.

[0045] C is the number of laser cycles. Said in another way, C is the total exposure time used to capture the histogram. There are two types of foveation used via the method of the present disclosure: memory foveation and depth foveation.Docket No. PSU24302PCT

[0046] In memory' foveation a set number of bytes in the memory are dedicated to finding the histogram peak. Therefore, placing the foveated bins at the peak is most efficient. In memory foveation, M number of bins are identified where the true depth exists and M « N. The width of the bins remains the same at T / N, and therefore the SNR is identical to the conventional case represented by equation 2.<2>

[0047] In depth foveation. the memory is fixed. However, depth foveation may distribute the bins more closely together near the histogram peak, improving depth resolution by increasing the accuracy of depth a foveated image compared to a ground truth depth where the number of bins are N where true depth exists. In depth foveation, the N bins that may have been distributed across an entire depth range of received photons are concentrated across a smaller region of the depth range of received photons. The smaller region is the region used in memory foveation. The smaller region is determined via multiplying the number of memory’ foveation bins M with the original bin width giving M*T / N. The smaller region is divided into N bins, therein the new bin width is M*T / (N2). The SNR is proportional to the bin width, and may be represented for depth foveation by equation 3.

[0048] The depth resolution may be increased but with a decreased SNR. To increase the SNR of the foveated depth, C may be increased, increasing the number of cycles the laser pulses to create the histogram. The new cycle number Cnewis equal to or greater than Cnew / C >N2 / M2. The new SNR (e.g., SNRnew) is proportional to the bin width, and may be represented for depth foveation via equation 4.

[0049] Memory foveation may reduce memory usage with approximately no change in SNR. Depth foveation may increase depth resolution but with reduced SNR that can be compensated by more laser photons received by the SPAD (e.g.. longer exposure for the SPAD).

[0050] The signal-to background ratio (SBR) may be defined for SPADs as the ratio of the total number of signal photons to the total number of background photons received over each laser cycle. The SBR is proportional to the probability of receiving signal photons to the probability of receiving background photons. With ambient light, photons from both the laser source and the ambient illumination may be measured by the SPAD. Each time a photon is detected, the SPAD sensor resets creating a pause. It is this pause that creates a binomial model for image capture in SPADs. Therefore, the SBR analysis cannot simply compare the photon bin widths as in the prior section for the full resolution (N bins) and the foveated resolution (M bins). Instead, SBR calculations include the probability of photon from the source vs. the background.

[0051] In a conventional case, with no foveation, using the Poisson model for photon distribution, tire probability’ of a photon from the laser incident on the bin may be represented byDocket No. PSU24302PCT corresponds to the correct depth, where correct is relative to the accuracy of depth of a ground truth image, represented by pia3er= (1 ~ e ®slg). Correct depth detection will happen even if an ambient photon is detected at the correct depth, so the probability of correct depth detection is represented by pcorreot. Pconect may be represented by pCOnect = (1 -e-i'®sig+*bkg')^ T]lcsymbol i is the location of the bin corresponding to tire correct depth of tire scene point. A photon of pCorrect may be detected at i if no photon from the laser is detected at any prior bin. Since the laser photons only show up at bin i, constrained by depth, the probability of the photon showing up at any other bin is zero. However, in a conventional case, photons from ambient light may appear at any prior bin to i, pausing detection at bin i. Therefore, the probability that the photon from the laser is detected at the correct depth may be represented by the equation psig= (1 —However, ambient photons may arrive at any time instant before photons from the 1thbin arrive. The probability’ that an ambient photon may be detected at a location q may be represented via a symbol p^kqand via the following equation A SRB proportionality may therein be represented via an equation

[0052] Memory and depth foveation may be used with ambient light conditions, such as when N is given. With memory and depth foveation the arrival of photons may be modeled from ambient sources of light and the laser source.

[0053] In memory foveation, the index for foveated bins N is j. The SRB may increase, since the histogram is unaffected photons received before bin j. A SBR proportionality may be represented by equation 6.

[0054] In cases where there is approximately perfect foveation. where i is equal to j, then the terms for ambient light before bin i become 1. When i = j, the SBR proportionality of equation 6 may therein simplified to and be represented by an equation 7.SBR oc 1 -e-(®«s+® ) (7)

[0055] Said in another way, the effect of foveation may remove the dependence on prior photon arrival for detection, since these no longer delay the measurement of photons at the zth bin. This "‘perfect foveation” SBR term is dependent on the ratio of the strength of the laser and ambient signal directly and is not constrained by the binomial nature of SPAD photon capture.

[0056] In depth foveation. the N bins are concentrated into the foveation window, and thus the histogram may be susceptible to the binomial nature of SPAD photon capture. In addition, the bins are a smaller to fit within the window of data graphed for the histogram compared to the memory foveation.Docket No. PSU24302PCTAs described regarding using depth foveation, the bin width is reduced as M / N. The probability that an ambient photon is detected at location q may be represented by the equation

[0057] In summary, memory foveation may increase the SBR. While depth foveation may have the same SBR as conventional capture, it improves depth resolution, and processing techniques using SPAD.

[0058] Turning to FIG. 1, it shows a schematic of a system 100 that may be used for imaging, and more specifically estimating and assigning depth to objects imaged. For example, the system 100 may image, estimate, and assign depth to an object 106. The system 100 is a light detection and ranging (LiDAR) system using SPADs for imaging, e.g., a SPAD LiDAR. The system 100 may include three subsy stems: a control system 110, a LiDAR sy stem 112, and a user sy stem 114. The control system 110 and LiDAR system 112 may be mounted or housed by an assembly 104. The assembly 104 may be or may be part of a larger system, such as a vehicle or a robotic appendage, assisted by the system 100 to interpret space and image environments. A user 118 may interact with the user system 116 to interact with images and other data gathered via the LiDAR system 112. Data, such as photon data, may be processed via an image processor system 141 before being received by user system 114. The processor system 141 may use foveation techniques to perform foveation during capture of data by the LiDAR system 112. The user system 114 may be a system that an operator of the system 100, referred to herein as a user 118, may control to give commands to or alter the function and settings of the LiDAR system 112. A plurality of dotted lines 122 may represent communication couplings, via which components may be communicatively coupled and data may be transferred.

[0059] The LiDAR system 112 may generate a beam 120 of light to image the object 106. Light from the beam 120 may be reflected by the object 106 and received by the LiDAR system 112 during capture to measure the distance of a feature the object 106 from the LiDAR system, imaging the shape and depth of the object 106. Generation of a beam, such as beam 120, may be referred to herein as firing the beam.

[0060] The LiDAR system 112 includes a light source, referred to herein as a light emitter 130, and a detector 131. The light emitter 130 may generate the beam 120. The light emitter 130 may be a laser. The detector 131 may be or incorporate a single-photon avalanche diode SPAD array 124 for detection of light, such as a SPAD sensor with single-photon sampling or another suitable image sensor capable of capturing single photon.Docket No. PSU24302PCT

[0061] In the example shown, the detector 131 may house the light emitter 130. However, it is to be appreciated that there may be other configurations of detectors, and the light emitter 130 may be external to the detector 131. The SPAD array 124 may receive and detect photons during capture of photons. The number of photons captured by the SPAD array 124 may be recorded as photon data. The time that photons are captured by the SPAD array 124 are recorded as time data. The SPAD array 124 may include a plurality of pixels. Each pixel may comprise one or more SPADs. The light emitter 130 may fire the beam 120. The beam 120 may exit the detector 131 via an aperture 126. Alternatively, where the light emitter 130 is external to the detector 131, the beam 120 may not exit the aperture 126 and be positioned to be unimpeded by features of the detector 131.

[0062] The beam 120 may be directed toward a surface 128 of the object 106. Photons of the beam 120 may contact the object 106. such as at the surface 128, and be reflected. For an example, a plurality of photons 136 from beam 120 may contact and be reflected at an area 134 of the surface 128. The photons 136 may be reflected back toward the SPAD array 124. The photons 136 may be received and pass through the aperture 126. The photons 136 may be received by the SPAD array 124. The quantity of the photons 136 received via the SPAD array 124 may be used to determine the target distance 142. The target distance 142 may be the distance between the aperture 126 and the surface 128. The target distance 142 may also represent depth in a LiDAR image, such as a depth mask or a depth processed image. For example, the target distance 142 may be the depth of a pixel of a LiDAR image created by the image processor system 141. The image processor system 141 may calculate the target distance 142 based on the peak of the distribution of the arrival times of the received photons.

[0063] The light emitter 130 and SPAD array 124 may be communicatively coupled to the clock 160. The SPAD array 124 and detector / detector system including tire SPAD array 124 may receive pulse frequency information from the clock 160 (e.g., indicating the start and end of each pulse of the light source 102). The pulse frequency information may be part of a clock (CLK) signal generated by the clock with the counting of time. The light emitter 130 may transmit light in pulses at a frequency dictated by the clock 160, and more specifically the CLK signal of the clock.

[0064] The control system 110 may include a controller 139 and a database 137. The database 137 may store images and other data from the LiDAR system 112. Photon data captured by the SPAD array 124may be recorded in the database 137. The database may store data from other sources than the LiDAR system 112, such as images from a camera, such as monocular images. The database 137 may also include settings of the LiDAR system 112. such as settings for patterns of scanning and laser pulses, that may be used for scanning methods and other frmetions of the LiDAR system 112. The controller 139 and image processor system 141 may be or include a computer, such as microcomputer. The computers of the controller 139 and processor system 141 each includes elements such as a microprocessor unit, input / output ports, an electronic storage medium for executable programs and calibration values, e.g., a read-only memory chip, random access memory, keep alive memory, and aDocket No. PSU24302PCT data bus. The storage medium can be programmed with computer readable data representing instructions executable by a processor for performing the methods described below as well as other variants that are anticipated but not specifically listed.

[0065] The controller 139 receives signals from the various sensors including a single or plurality of sensors of the LiDAR system as well as input from the user system 114. The controller 139 may also receive suggestions from the image processor system 141. The signals from the LiDAR system 112, user system 114, and suggestions from the image processor system 141 may influence the controller 139 to execute instructions. Instructions executed by the controller 139 may be sent as command signals to adjust conditions in the LiDAR system 112. For one example, the controller 139 may send command signals to various actuators of the LiDAR system 112 to change the position of the light emitter 130. The controller 139 may also send command signals to change the color or power of the beam 120 generated by the light emitter 130 based on user input or the suggestion of an image processor system 141.

[0066] For example, the controller 139 may adjust, such as via increasing or decreasing, the length of time that light is emitted by the LiDAR system 112 via a first signal. For another example, the controller 139 may adjust, such as via increasing or decreasing, the length of time that photons are received and captured by the LiDAR system 112 via a second signal.

[0067] The image processor system 141 may include a processor, such as a computer, with non- transitory memory. The image processor system 141 may communicatively couple to the database 137. The signatures of photons received read by the photo-sensitive surface may be sent to the image processor system 141. The image processor system 141 may create depth images, such as depth masks from data of tire LiDAR system 112 or depth images processed from monocular images and one or more depth masks from the database 137. The image processor system 141 may perform histogram computations for each pixel of a depth mask or a depth image using photon data from the SPAD array 124. For example, a pixel of a depth mask or a depth image may be created using the data of photons received via the photo-sensitive surface 132. The image processor system 141 may generate and foveate a histogram to assign depth to the pixel.

[0068] The signatures of photons received at the SPAD array 124 may be sent to the image processor system 141. The LiDAR system 112 may be communicatively coupled to the user system 114 via a wired connection 146 or a wireless connection. The wired connection 146 may communicatively couple to a computer 152 of the user system 114 via a port or another wired interconnect. Likewise, the computer 152 may be paired with a wireless interconnect 148 of the LiDAR system 112, to form a wireless connection between the LiDAR system 112 and user system 114. The wired connection 146 and wireless interconnect 148 may be communicatively coupled to the LiDAR system 112 at a communication junction 150. Likewise, the control system 110 and components of the LiDAR system 112 may be communicatively coupled at the communication junction 150.Docket No. PSU24302PCT

[0069] The user system 114 may send data, such as commands from the user 118, from the computer 152 to the LiD AR system 112 via either the wired comiection 146 and / or wireless interconnect 148. For example, the settings of laser, such as power, intensity', and color may be adjusted by the user 118 via the user system 114. The computer 152 may output visual data, such as images from the database 137 and images processed via the image processor system 141. The display device 154 may also display a graphical user interface (GUI) 158. The GUI 158 may give the user 118 options to control the LiD AR system 112, such as via selecting options to adjust the settings of the LiD AR system 112. For example, the user 118 may select an option in the GUI 158 to start scanning using the LiD AR system 112 and fire the light emitter 130. Likewise, the user 118 may interact with images from the LiD AR system 112, stored by the computer 152, and / or processed via the image processor system 141 via the GUI 158. The user 118 may interact with the GUI 158 via a single or plurality of inputs 156, such as a mouse and keyboard, joysticks, game console or touch screen.

[0070] Thus, FIG. 1 shows an example imaging environment for imaging a scene with a singlephoton-sensing 3D imaging system. The environment includes a light source (e.g., light emitter 130), an object (e.g., object 106), and a detector (e.g., detector 131). The environment further includes a computing device (e.g., control system 110). The light source and detector may form a single-photonsensing 3D camera (SPC) that captures distance (e.g., depth) information using the time-of-flight principle — akin to echolocation, but with light instead of sound. Consider a single scene point where distance needs to be estimated as shown in FIG. 1. The light source (which may be a laser) illuminates a scene point (which may be part of the object) with a short light pulse (e.g., the beam shown in FIG. 1). The detector may include an array of detector elements, with each detector element configured to (separately) detect photons. In some examples, each detector element may be a diode, such as a singlephoton avalanche diode or an avalanche photodiode. As used herein, the detector elements may be referred to as pixels and thus the detector may include a plurality of pixels. A pixel of the detector may capture a stream of return events (e.g., the return photon events of FIG. 1) as photons arrive at different time delays with respect to the time the original light pulse was transmitted. (The detector in general captures more than one return event in response to each laser pulse sent into the scene.) Moreover, this return stream may also contain spurious photon events not due to the light pulse (signal), but rather to ambient background light and other sources of noise in the image sensing hardware. Traditionally, a histogram is constructed by accumulating photon counts at different delays over many light (e.g.. laser) cycles. The “arg max” peak location of this histogram gives an estimate of the true distance of the scene point (relying on the simple relationship that the speed of light multiplied by the time delay is equal to twice the distance to the scene point). As will be explained in more detail below, power consumption, memory' requirements, and processing power needed to determine object distances may be reduced if the true distance of the scene point is estimated using a foveated histogram instead.Docket No. PSU24302PCT

[0071] The computing device (which may incorporate the light source and detector in some examples) may be (or may be coupled to) a smartphone camera, a light detection and ranging (LiDAR) sensor (e.g.. for autonomous robotics), a camera for scientific imaging, a virtual reality device, an augmented reality device, a desktop computer, a laptop, a mobile device (e.g., smartphone or tablet), or another suitable device.

[0072] Due to their compatibility with CMOS fabrication technology, there is increasing availability of high (kilo-to-megapixel) resolution arrays of single-photon detecting (e.g.. SPAD) pixels with additional data processing embedded in the hardware chip that includes the single-photon detecting pixels. Unfortunately, the high sensitivity and high speed is a double-edged sword: the amount of raw data generated by these detectors is orders of magnitude higher than can be reasonably processed or transferred in real-time. This aspect limits their applicability in many real-world applications, especially those that are power and bandwidth constrained.

[0073] Accordingly, and as explained in more detail below, embodiments are provided herein that offer a different approach for direct time-of-flight imaging that is compatible with a variety of detector and illumination schemes. Capturing and transferring the entire received waveform (either through single-photon sampling with a SPAD pixel, or fast analog-to-digital conversion of an APD) is resource hungry’. Instead of attempting to capture the complete waveform in the digital domain (which often consumes a large fraction of the total power), the embodiments disclosed herein perform as much of die processing as possible in the analog domain. To this end, aspects in the field of race logic are applied, where information is encoded not in die voltage levels of signals but in the precise arrival times of the signals. This approach is naturally suited to single-photon time-of-flight 3D sensing because the arrival times of the photon-return events carry useful scene information (scene distances and reflectivity). Additionally, the embodiments disclosed herein utilize foveated histograms to represent the transient distribution of photon return events, rather than full resolution histograms that other methods employ. The power and bandwidth limitation of single-photon cameras severely limits wider applicability of high resolution singe-photon counting arrays. Creating the full histogram on-sensor is infeasible due to severe memory constraints, while moving photon timestamp data off-sensor is undesirable because it introduces latency and consumes power. The foveated histograms and use of photon-arrival times as described herein address these issues, as they reduce or eliminate the need for large storage on the sensor and utilize very little power by transferring very little data off-sensor.

[0074] As explained in more detail below, a foveated histogram as disclosed herein focuses the bins of the histogram around a predicted or estimated peak of the photon return events. Focusing the bins may include focusing a limited number of bins (e.g., fewer than the full number of bins used in a full-resolution histogram) around the predicted or estimated peak or focusing the full number of bins arormd the predicted or estimated peak (in which case the bins would have a smaller width than the bin width of the full-resolution histogram). In either case, the focused bins may be referred to as a foveationDocket No. PSU24302PCT window that represents a limited span of the temporal range of sensed photon events. The foveation window may be centered around the predicted or estimated peak by using a depth prior, which is an estimate, for each pixel of the SPC camera detector array (e.g., each pixel of the SPAD array), of the sensed distance / depth of the imaged scene. The depth prior may be generated from a visible light camera, such as visible light camera 111 of FIG. 1. The visible light camera 111 may be a color camera (e.g., RGB camera) or a monochrome camera. The visible light camera 111 may be included in assembly 104 or the visible light camera 111 may be an external camera. In either case, the pixels of the visible light camera may be co-located with pixels of the SPAD array 124, such that a given pixel of a specific x,y coordinate of an image obtained with the visible light camera is located at the same x-y coordinate of an image obtained with the SPAD array. The methods of calibration and pixel-matching between the pixels of the RGB camera and the pixels of the SPAD array may be completed via image processing methods executed via the image processor system 141 of FIG. 1. for example.

[0075] FIG. 2A schematically shows a process 200 for processing images using foveated binning techniques and non foveated binning techniques to estimate depths of an imaged scene using SPAD LiDAR data. FIG. 2A shows a simplified process for producing processed images using a reduced number of bins compared to a ground truth processed image, including binning a histogram with foveation and without foveation.

[0076] The depth and memory foveation techniques described herein rely on estimation or prediction of the true depth of each imaged point in the imaged scene, in order to center the foveation window . The estimation or prediction of the true depths may be referred to herein as a depth prior.

[0077] A first image 212 is a depth image generated using a SPAD-based LiDAR and SPAD-based processing system. Depth for the first image 212 is generated using processing that does not include foveation (e.g., using the full number of histogram bins, N, spread over the entire depth range T). A second image 214 is an RGB (e.g., color) image. The first image 212 and second image 214 are ground truth images for depth and RGB respectively. Depth in the first image 212 may be represented via an intensity of color (e.g.. a color intensity) and / or an intensity of brightness (e.g.. a brightness intensity) on a scale 218. The scale 218 shows a gradient of colors and / or brightness intensities that may be selected and assigned to pixels of depth images. The first image 212 and the second image 214 may be captured simultaneously using a depth camera (e.g., the SPAD-based LiDAR system of FIG. 1) and an RGB camera (e.g.. the RGB camera of FIG. 1).

[0078] For example, the SPAD array and RGB camera may be part of the same camera system. In another example, the RGB camera may be an external camera. The pixels of the monocular image may be co-located with pixels of the SPAD array through calibration and methods of matching the pixels of the monocular image with the pixels of the SPAD array, such that the pixels of monocular image and the pixels of the SPAD array are aligned to occupy the same coordinates (x, y) when overlaid.Docket No. PSU24302PCT

[0079] The first image 212 may have high depth resolution owing to the use of the full number of bins and depth range, but suffers from the data bottleneck issues described above. Thus, the photon data may be processed via a limited bin teclmique 232 to produce a limited bin image 242. The second image 214 may be processed via depth prior estimation technique 224 to produce a third image 226 (e.g., a depth prior). The third image 226 may be an image of depth estimated using monocular estimation (e.g., a monocular image), and therein the depth prior estimation technique 224 may be a monocular estimation technique. Additional details about generating a depth prior using monocular estimation are provided below.

[0080] Further, FIG. 2A shows that the photon data and the third image 226 may be processed using via a foveated bin technique 234 to create a fifth image referred to herein as a depth foveated image 244. The depth foveated image 244 appears to have a higher resolution with greater clarity of objects compared to the limited bin image 242. Additionally, the depth foveated image 244 appears to have greater differences in depth compared to the limited bin image 242. Said in another way, the limited bin image 242 appears flatter compared to the depth foveated image 244.

[0081] The limited bin technique may be illustrated via a first graph 252, and the foveated bin teclmique may be illustrated via second graph 254. The first graph 252 and second graph 254 are graphs of intensity of photons measured vs time, with the time being represented via first axis 256 and the intensity of photons being represented by second axis 258. The first graph 252 displays a first trace 262 and the second graph displays a second trace 264 of data from a SPAD LiDAR (e.g., from one pixeFelement of the SPAD array). The first trace 262 of the first graph 252 is bimied via a plurality of limited bins 266 distributed over the full trace. The second trace 264 of the second graph 254 is binned via a plurality of foveated bins 268. The limited bins 266 and the foveated bins 268 are bins based on time. The first trace 262 and the second trace 264 may be traces of the same data. There may be the same number of foveated bins 268 as limited bins 266. The foveated bins 268 are concentrated around a peak 272 of the second trace 264, such as within a range 274 of thresholds from the peak 272. The thresholds of the range 274 are thresholds of time on the first axis 256. The foveated bins 268 may selectively bin data from the second trace 264 within the range 274 and may exclude binning data outside of the range 274. The range 274 may therein be referred to alternatively and herein as a foveation window 274. Data outside of range 274 may be mostly noise compared to data within the range 274. The limited bins 266 may bin approximately all of the data from the first trace 262. The limited bins 266 may therein have wider bin widths, where each bin counts photons received over a greater length of time, compared to the foveated bins 268.

[0082] FIG. 2B shows a table 280 of processed depth images for comparison. Depth images of two regions of interest (ROIs) imaged in the first and second images are shown, each generated with different techniques. Two monocular images 282 are shown in a first row of table 280, showing depth images of the two ROIs created using monocular depth estimations. Two limited bin images 284 areDocket No. PSU24302PCT shown in a second row of table 280, showing depth images of the tw o ROIs created using the limited bin technique. Two ground truth images 286 are shown in a third row of table 280, showing depth images of the two ROIs created using the full histogram resolution. Two depth foveated images 288 are shown in a fourth row of table 280, showing depth images of the two ROIs created using the depth foveation technique. Each column of the table 280 are for images of the same ROI processed using different techniques.

[0083] Turning to FIG. 3. it shows a table 300. Table 300 is a table of figures showing different sets of images that include a plurality of RGB images, a plurality ground truth depth images, a plurality of monocular depth images, a plurality of simulated depth images obtained using the limited bin technique, a plurality of images obtained using memory foveation techniques, and a plurality of images obtained using depth foveation technique. The monocular depth images may be obtained via ZoeDepth technique. The simulated (e.g., limited bin) images may be obtained using limited bins as described above. The memory foveation technique minimizes memory usage during computing. Images obtained using the depth foveation technique may maximize the accuracy of the depth of an image, where accuracy is relative to the depth of a ground truth depth image of the same image set.

[0084] The first column 312 displays the relationship between M and N for each set of images and N’ values for the sets of images. M is the number of bins across tire foveated histogram, and N’ is number of bins used for limited depth foveation. M is also the number of bins used for images with limited set of bins that are un-foveated, such as the simulated images. M is shown in the first column related to N, w here N is the number of bins across a whole histogram w ithout foveation or limited bins. The second column 314 shows the RGB images. The third column 316 show s the ground truth depth images. The fourth column 318 shows the monocular depth images. The fifth column 320 shows the simulated images, using a M number of bins without foveation. The sixth column 322 displays images obtained using memory foveation to minimize memory consumption. The seventh column 324 displays image obtained using depth foveation to increase the accuracy of depth.

[0085] A first row 332 displays a first set of images. For the first set of images in the first row 332, M is l / 16th of N for the foveated and un-foveated limited bin images, and N’ is 16 for the foveated images. The second row 334 displays a second set of images. For the second set of images in the second row 334, M is l / 8th of N for the foveated and un-foveated limited bin images, and N’ is 16 for the foveated images. The third row 336 displays a third set of images. For the third set of images in the third row 336, M is l / 8th of N for the foveated and un-foveated images, and N’ is 64 for the foveated bin images. The fourth row 338 displays a fourth set of images. For the fourth set of images in the fourth row 338, M is 1 / 16th of N for the foveated and un-foveated limited bin images, and N’ is 64 for the foveated images. The images shown in FIG. 3 are grayscale images of images originally produced in color. Thus. FIG. 3 shows that the memory and depth foveation techniques produce quality depth reconstructions with a fraction of the niemon usage. Each row in FIG. 3 includes the NYUv2 groundDocket No. PSU24302PCT truth depth images, the monocular depth output from ZoeDepth, a simulated SPAD output with N' bins, and the foveation techniques of the disclosure. The rows show different combinations of M and N', where M is the number of bins in the foveated histograms, and N' is the limited number of bins used for depth foveation.

[0086] A Table II shown below corresponds with FIG. 3. Table II represents variables a local scale for the memory and depth foveation evaluation of pixels in a depth image.TABLE II. TABLE OF MEMORY AND DEPTH FOVEATION EVALUATION

[0087] Table II shows a quantitative comparison of root mean square error (RMSE) and depth inlier metrics for different depth and memory foveation strategies for a data set, for this example the NYUv2 dataset. For each memory foveation fraction, the number of histogram bins in the foveated subwindow is varied to achieve depth foveation.

[0088] Turning to FIG. 4, it shows a table 400. Table 400 is a table of figures showing different sets of images that include a plurality of RGB images, a plurality ground truth depth images, a plurality of quantized monocular images, a plurality of sparse sample images, a plurality of images obtained using memory fovea techniques, and a plurality of images obtained using depth fovea technique. The RGB images and the ground truth depth images are from the NYUv2 training set. The sparse samples images are images of selected pixels to be processed via the memory foveation and depth foveation techniques. Selected pixels of the sparse samples are selected from the quantized monocular images.

[0089] The first column 412 displays the relationship between N and M and the N’ values for each image set. The second column 414 shows the RGB images. The third column 416 shows the ground truth depth images. The fourth column 418 shows the quantized monocular depth images. The fifth column 420 shows the sparse sample images. The sixth column 422 displays images obtained using memory foveation to minimize memory consumption. The seventh column 424 displays images obtained using depth foveation to increase the accuracy of depth.

[0090] For the image sets of the first row 432 and the second row7434, M is 1 / 16 of N, and N’ is 16.

[0091] The sparse sample images of the fifth column 420 comprise a plurality7of sparse sample pixels 442, where each of the sparse sample pixels 442 has a depth value (e.g., as determined from the monocular depth estimate) and a specific depth color and / or brightness value. Other pixels of the sparseDocket No. PSU24302PCT sample images of the fifth column 420 are shown dark, where each of the other pixels lacks depth values and a specific depth color and / or brightness value.

[0092] A spatial temporal method may select, foveate, and histogram sparse sample pixels 442. The spatial temporal method may process spatially by exploiting depth coherencies and applying foveated windows to a smaller selection of pixels (e.g., the sparse sample pixels) compared to if the total quantity of pixels were histogrammed via full histogram for memory foveated images of the sixth column 422 and depth foveated image of the seventh column 424. The spatial temporal method may separate pixels of the quantized monocular images of the fourth column 418 into groups referred to herein as buckets. Each pixel in a given bucket may have the same depth value. At least a sparse sample pixel of the sparse sample pixels 442 is selected at random from each bucket. However, there may be a plurality of sparse sample pixels 442 selected at random for each bucket. The special temporal method may then histogram the photon data for the sparse sample pixels 442 using foveation windows. For example, the special temporal method may use memory foveation and memory foveation windows to bin a histogram and assign a depth value for each sparse sample pixel 442. Once a depth value is determined ria the foveation for a selected sparse sample pixel, all other pixels in the same bucket as the selected sparse sample pixel may be given the same depth value to create the memoir foveated images of the sixth column 422. As another example, the special temporal method may use depth foveation and depth foveation windows to bin a histogram and assign a depth value for each sparse sample pixels 442 to create the depth foveated images of the seventh column 424. For both examples, other pixels that are not selected as sparse sample pixels from each bucket are considered redundant by the special temporal method. Therein, the other pixels from the buckets use the depth values from the sparse sample pixels. The spatial temporal method described for FIG. 4 may be shown and described in greater detail as a spatial temporal optional step of 1060 of FIG. 10.

[0093] Turning to FIG. 5 it shows a third table 500. The third table 500 is a table of figures showing different sets of images that include a plurality of RGB images, a plurality ground truth depth images, a plurality of optical flow driven images, a plurality of optical flow error images, and a plurality of fovea images created via the optical flow driven images and optical flow error images. The optical flow driven images are a plurality of images obtained using Memory Fovea techniques, and a plurality of images are obtained via a Carla training simulator program. The RGB images and the ground truth depth images are from a Carla training simulator program.

[0094] The first column 512 displays the relationship between N and M values for each image set. The second column 514 shows the RGB images. The third column 516 shows the ground truth depth images. The fourth column 518 shows the optical flow driven images (e.g.. estimated depth based on an optical flow technique, which may be used as a depth prior for foveation). The fifth column 520 shows the optical flow error images. The optical flow error images show an error of depth where error in depth is relative to being within a depth threshold of the ground truth depth images. The ground truthDocket No. PSU24302PCT depth images may be made distinct from the depth gradient scale, such as being a different color. The sixth column 522 displays the fovea images after foveating using optical flow as a depth prior and correcting the optical flow error images.

[0095] A first row 532 displays a first set of images. For the first set of images in the first row 532, M is l / 4th of N for the foveated images. The second row 534 displays a second set of images. For the second set of images in the second row 534, M is 1 / 1 Oth of N for the foveated images.

[0096] Optical flow techniques may be used to generate a depth prior for a SPAD sensor that may be fixed to a moving platform, such as an autonomous vehicle, where high-frame rate and efficient depth capture are desired. Optical flow images are created using optical flow techniques that use temporal information by transferring foveation information from a previous frame to subsequent frames. Foveation information may include ranges of times to be binned (e.g.. foveation windows). In optical flow, depth from a previous frame is assigned to a subsequent frame pixel by pixel, taking into account movement of objects and / or the camera. The bins that includes the histogram peaks are different from frame to frame, but within a threshold of bins such that a window of pixels may enable the recovery of the histogram peak in the current frame. The approach transferring foveation information from previous frames to subsequent frames may reduce computation times even further compared to memory and depth foveation techniques discussed previously. More specifically, using a foveation windows for the same pixel between depth images at different temporal times (e.g., different frames), as opposed to recalculating the foveation window for each temporal time, may reduce the computation even further.

[0097] For example, consider multiple frames of both depth and reflectance from a scene. Assume the first frame in the sequence has high-quality depth reconstructed from full-resolution SPAD histograms, for example. Consider now a subsequent frame. Calculating optical flow between two color or grayscale frames (e.g., the first frame and the subsequent frame), a vector (u, v) can be created for every pixel at any given time t. such that the vectors uphold the brightness consistency principle I(x + u* 5t, y + v* 5t, t + St) = I(x, y, t). The depth from the previous frame is used to drive the location of the foveating window in the current time instant. For example, the depth of a given pixel from the previous frame may be assigned to a new pixel in the current frame, where the new pixel is identified from the vector (e.g., the movement of the object from the previous frame to the current frame). It is to be appreciated here that t refers to time since the first frame and not photon arrival time.

[0098] In optical flow errors may propagate. Errors may occur at the edges of a frame or edges of an object or surface while moving, as shown by the error images. To remove errors, such as the error captured in the error images, the distribution of the photons under a foveated region are compared to photons from a noise floor. If the foveated region and the noise floor match, the optical flow is ignored and the depth is recomputed from the full histogram. In practice this is done by thresholding the values in the foveated window. As the value of M decreases relative to N, the error of the error of the optical flow error images may increase.Docket No. PSU24302PCT

[0099] Turning to FIG. 6A, it shows a first depth image 612 and a second depth image 622. For the first depth image 612, depth is calculated and assigned to each pixel using histograms of the SPAD LiDAR data and with N number of bins. For the second depth image 622. depth is calculated and assigned to each pixel using histograms of the SPAD LiDAR data and with M number of bins, where the M number of bins are foveated using a super-pixel approach. The first depth image 612 and the second depth image 622 are created via the same photon data, the subject for an example being a reindeer head. The super-pixel approach may be applied when a color or grayscale camera is not available. In the super-pixel approach, pseudo-intensity maps are generated from the photon data collected with the SPAD array by summing the raw photon data cubes along tire temporal axis for each pixel, which may be realized with a counter in each SPAD pixel. A super-pixel algorithm is executed on the pseudo-intensity maps to obtain course segmentations (super-pixels) of the scene. A complete (e.g., non-foveated) histogram is captured for the centroid pixel of each super-pixel. For the remaining pixels in each super-pixel, foveation is performed in a one-quarter sub-window around the true peak of the corresponding centroid pixel.

[0100] Turning to FIG. 6B, it shows a first graph 614 and a second graph 624. The first graph 614 is a graph of photon data of a first pixel of a SPAD array bimred into a histogram of N bins, which may be used to create the first depth image 612 of FIG. 6A. The first pixel may be a centroid pixel of a superpixel. The second graph 624 is a graph of photon data of a second pixel binned into a foveated histogram of 1 / 4N bins centered at the peak of the histogram of the first graph 614. which is then used to create of the second depth image 622 of FIG. 6A. The second pixel may be in the same super-pixel as the first pixel.

[0101] The first graph and the second graph have a first axis 632 and a second axis 634, where the first axis 632 represents the bin number and the second axis 634 represents the number of photons received by the SPAD LiDAR. The first graph 614 shows a first histogram 652, and the second graph 624 shows a second histogram 654. The first histogram 652 is created using N number of bins. The second histogram 654 is created using M number of bins, where the bins are foveated. While having less bins, the second histogram 654 includes a second peak 664 that is clearly defined and has approximately the same number of photons at approximately the same numbered bin as a first peak 662 of the first histogram 652.

[0102] FIG. 6C shows a first image 672 that is an RGB image, shown for visualization purpose. Additionally, FIG. 6C shows a second image 674 that is a grayscale SPAD image (e.g., a pseudointensity map as described above), and a third image 676 that is a super-pixel segmented image. The third image 676 is segmented into a plurality of super pixels 682. The first image 672, the second image 674. and the third image 676 are of the same subject as the first depth image 612 and the second depth image 622, for this example being the reindeer head.Docket No. PSU24302PCT

[0103] A pseudo-intensity map may be created by summing up raw photon data cubes along the temporal axis for each pixel of the SPAD array. In real hardware, the pseudo-intensity map be realized by having a counter in each SPAD pixel; this capability is commonly available in existing commercial SPAD arrays. A super pixel algorithm is applied to the pseudo-intensity maps to obtain a plurality coarse segmentations, the course segmentations divide the second image 674. transforming the second image 674 into the third image 676. These coarse segmentations are the super pixels 682 segmenting the third image 676. Each of the super pixels 682 contains a plurality of pixels that comprise the third image 676, such as tens to hundreds of pixels. It is to be appreciated that the super-pixel segmentation divides the second image 674 into super-pixels based on the color / intensity of the pixels of the second image 674 in order to segment imaged features. For example, the pixels that correspond to the antlers of the reindeer are included in a first set of super-pixels, the pixels that are in the background are included in one or more second sets of super-pixels, etc. In this way, the pixels are not simply divided into larger, even groups of pixels (e.g.. as a grid) but instead the super-pixels are set based on the underlying image features. Next, a full (non-foveated) histogram is captured for a subset of pixels of the total number of pixels for the histogram. The subset of pixels may each be a centroid pixel (centennost pixel) 684 specific to each of the super pixels 682. The centroid pixel 684 is shown schematically and is not to scale. The true peak location of the centroid pixel is located in and found via binning a trace of the photon and time data via a full histogram. Each of the other pixels of each of the super pixels 682 are histogrammed using foveated histograms, where each of the peak locations of the other pixels may be histogrammed via binning photon and time data using a foveated histogram referred to as a super pixel adjusted foveated histogram. The other pixels neighbor the centroid pixel 684, and therein may alternatively be referred to as neighboring pixels in relation to the centroid pixel 684. The super pixel adjusted foveated histogram has a foveation window that is a l / 4th sub-window of a window for the full histogram centered around the true peak of the centroid pixel 684 specific to die super pixel that contains the other pixels. Using super pixel algorithm to generate super pixels, selectively using a full histogram of die centroid pixel 684 of the super pixels, and using foveated the remaining pixels of super pixels reduces the overall bandwidth requirement for each pixel of the third image 676 by a factor of 64. It is to be appreciated, that for an alternative example.

[0104] A method for generating images segmented via super pixels, such as the third image 676, from an SPAD grayscale image, such as the second image 674, may be shown in FIG. 13 (e.g., a method 1300). The method may be used to generate depth images from SPAD data using super pixel segmentation, where a full (non-foveated) histogram is generated on SPAD photon and time data of centroid pixel 684 from each of the super pixel to estimate a true peak, then use the true peak and estimate a foveation window to generate foveated histograms to estimate foveated peaks from the remaining pixels in the super pixel.Docket No. PSU24302PCT

[0105] Turning to FIG. 7, it shows a graph 700. The graph includes a first axis 712 and a second axis 714. The first axis 712 is an axis showing signal-to-background ratio (SBR), where the signal decreases from left to right. As the SBR increases, a signal, such as a peak shown via a histogram, may become more defined and higher compared to the noise of photon data. Likewise, as the SBR decreases, the signal becomes less defined and visible relative to the noise of photon data. A greater quantity of photons relative to photons reflected from a laser of a LiDAR, such as from ambient light, may increase the SBR. The second axis 714 is an axis showing values of root mean square error (RMSE). The second axis 714 is dependent on the first axis 712. The graph 700 includes a first trace 722 and a second trace 724, where each trace represents increasing RMSE with decreasing signal-to-background ratio. The first trace 722 shows RMSE as a function of SBR calculated for a plurality of depth maps / images (e.g.. of the reindeer head) generated using non-foveated histogram bins with increasing background illumination. The second trace 724 shows RMSE as a function of SBR calculated for a plurality of depth maps / images generated using foveated histogram bins (using the same photon data as the first trace). The first trace 722 shows that depth map quality degrades more rapidly as background illumination increases compared to the second depth map of the second trace 724. Using memory foveation. as shown by the second trace 724, allows reliable depth map recovery for a wider range of SBR levels.

[0106] In summary, tire graph 700 of FIG. 7 shows that foveation has an advantage of allowing an image processor, such as image processor system 141 of FIG. 1, to pick an approximately correct depth peak even in the presence of strong background illumination, expanding the operable SBR range in practice.

[0107] Turning to FIG. 8, it shows a process 800 for generating depth images using foveation techniques. The process 800 may be a lower level view of the process in FIG. 2A. The process 800 utilizes a monocular network 812 and a SPAD simulator 814.

[0108] An RGB image, such as the second image 214, may be input into the monocular network 812. The monocular network 812 may create a monocular image, such as the third image 226. via monocular estimation with an estimated depth from the RGB image.

[0109] Pixel photon data 822 may be taken from a ground truth depth image, such as the first image 212. Pixel photon data 822 may be input into the SPAD simulator 814. The SPAD simulator 814 may generate a simulated histogram 828 from pixel photon data 822. The SPAD simulator 814 may uses the simulated histogram 828 to create a simulated depth image 832 and / or a grayscale image 834. The simulated depth image 832 and the grayscale image 834 may comprise a plurality of pixels created and assigned depth values or brightness via a plurality of simulated histograms.

[0110] Pixel photon data 824 may be taken from the third image 226 or another estimated depth image. Using pixel photon data 824 and the simulated histogram 828, a foveated histogram 842 may be generated for a pixel w ith an estimated depth for the foveated bin image, such as a memory' foveated image 844. The memory foveated image 844 is comprised of a plurality of pixels with depth valuesDocket No. PSU24302PCT assigned via foveated histograms, including the foveated histogram 842, where each of the foveated histograms assign a depth value to a pixel. Alternatively, for another example, the foveated bin image may be a depth foveated image, such as the depth foveated image 244 of FIG. 2 A.

[0111] Turning to FIG. 9, it shows a method 900 shows a method of collecting and foveating time data and photon data from a SPAD into a histogram. The data may be used to generate and process a depth for a pixel of a foveated image using memory or depth foveation techniques. The SPAD may be one of a plurality of SPADS in a SPAD-based LIDAR or another SPAD imaging device.

[0112] Method 900 is described with regard to the systems and components of FIG. 1. though it should be appreciated that the method 900 may be implemented with other systems and components without departing from the scope of the present disclosure. Method 900 may be carried out via a singlephoton sensing camera or a computational system operably coupled to a single-photon sensing camera. For example, method 900 may be carried out by instructions stored in memory of the image processor 141 and / or control system 110 of FIG. 1.

[0113] At 908, a value is set for M to determine the amount of bins that may be used to create a foveated histogram. As explained above, M may be less than the full number of bins N and may be selected based on depth resolution demands and / or computational / memory constraints. M bins may be used when memory foveation is performed. At 910, a value of N’ is set for the number of depth foveated bins to be used. At 912, the method 900 sends a command for the LiDAR to fire a light emitter, such as a laser, at a target. The LiDAR may be part of the LiDAR system 112 and the light emitter may be the light emitter 130 of FIG. 3. The SPADs of the LiDAR may receive a plurality of photons from light, including light from the light emitter after being reflected from an object. At 914, the method 900 each SPAD receives a plurality of photons. The photons are captured over a period of time after firing light via the light emitter. Each photon detected and the time the photon was received (e.g., time data) are recorded during detection as photon and time data. Said in another way, the photon data for each SPAD includes a count of photons received by the SPAD, with each photon arrival time-stamped.

[0114] At 916, data of the quantity of photons received by the SPAD is assigned to a pixel of the SPAD array. At least one SPAD is associated to each pixel, where a photon received by the SPAD is assigned to the pixel. However, it is to be appreciated that a group of SPADs may be associated a pixel, and the photon detected by the SPADs of the group are assigned to the pixel. At 922, the method 900 sends the photon and time data to be processed computationally. The photon and time data may be processed by the image processor or another computational device. At 924. the method 900 generates a trace from the photon and time data. The trace is of photons received over time. Said in another way the trace includes a first axis recording times photons are received in rmits of time and a second axis records the number of photons received by the SPAD. with the resulting trace showing the change in photons received over time. The trace may be generated via the image processor or the otherDocket No. PSU24302PCT computation device. Examples of traces generated at 924 may include the first trace 262 and the second trace 264 of FIG. 2 A.

[0115] At 926, the method 900 foveates the trace by selecting a smaller region of time from a full set of time recorded. The smaller region of time includes a time with a first peak of the photons, where the first peak is the greatest number of photons received by the SP AD during the time span of measuring. At 928, the method 900 creates a histogram by binning the time of the smaller region via foveated bins. The number of foveated bins may vary depending on if method 900 is doing memory or depth foveation. For example, memory foveation may be selected, where the foveated number of bins is M and the smaller region of the trace is binned by M bins. For a second example, depth foveation may be selected, where the foveated number of bins is N’ and the smaller region of the trace is binned by N’ bins. The histogram includes a second peak that is a bin that includes the maximum number of photons in the histogram bounded by a bin of time. The second peak is flanked by a plurality of other histogram bins, that decrease in size and number of photons. At 930, method 900 repeats 914-928 to for each SPAD or SPAD group associated with a pixel that receives and detects photons to generate histogram map (e g., a histogram image). The histogram map is comprised of a plurality of pixels each with a histogram created via the techniques at 928, where each pixel is a pixel assigned to a SPAD or a group of SPADs that received photons during a pulse of the light from the light emitter. There may be thousands of pixels each associated with a SPAD or a SPAD group, and a portion or all of the pixels may receive photons, send photon and time data to be processed and histogram, to generate a histogram map.

[0116] At 932, a depth map, such as a depth image, is created from the histogram map generated via 930. At 932 the histogram map is transformed into a depth map via a process that may be described herein as a decoding the histogram. During decoding, the method 930 assigns a depth value to each pixel based on the histogram of the pixel. Each histogram generated for the histogram image corresponds with and is assigned a depth value by the method 930. The histogram includes a second peak that is a bin that includes the maximum number of photons in the histogram bounded by a bin of time. The second peak and the other bins are translated by the computational device to generate and assign a depth value to the pixel in a process that may be described herein as decoding. The pixels may be used to create a depth image.

[0117] At 934 a visual indicator, such as a color or brightness values, may be assigned to each pixel based on the depth value of each pixel. The visual indicator may be selected from a gradient of intensities, with each intensity corresponding with a different depth value. The method may read a depth value and select a visual indicator from the gradient to shade the pixel. The visual indicator may create a depth map perceivable by a person view an image. Additionally or alternatively, method 900 may have other steps in place of or after 934. For example, the method 900 may input the depth map into logic and decision making system of an autonomous system, such as an autonomous vehicle, to makeDocket No. PSU24302PCT decisions on a position to move and a path to move to the position based on the depth map of the environment generated via method 900.

[0118] FIG. 10 shows a method 1000 for memory and depth foveation to generate a histogram from photon and time data from an SPC camera, such as a SPAD array. Method 1000 is described with regard to the systems and components of FIGS. 1. 15, and 16, though it should be appreciated that the method 1000 may be implemented with other systems and components without departing from the scope of the present disclosure. Method 1000 may be carried out via a single-photon sensing camera (including a single-photon sensing detector array) or a computational system operably coupled to a single-photon sensing camera or a single-photon sensing detector array (e.g., a SPAD array). For example, method 1000 may be carried out by instructions stored in memory of the image processor 141 and / or control system 110 of FIG. 1. Further, some aspects of method 1000 may be carried out by histogrammers of the single-photon sensing detector array, such as histogrammer 1520 of FIG. 16, described in more detail below. At 1010, a bin width is calculated for bins of a foveated histogram. As described previously, a SPAD may produce a voltage signal each time a photon is incident on the SPAD. When a voltage signal is generated, the voltage signal may be counted and time-stamped via circuitry of the SPAD array (described in more detail below) and the time stamps may be used to determine the arrival time of each photon (e.g., the time from when the light pulse was emitted to when the photon was received at the SPAD). also referred to as a return time. Each SPAD of the SPAD array may sense a plurality’ of photons over one or more light cycles, referred to as a stream of photon return events. The photon arrival times / stream of photon return events may be binned into a foveated histogram to identify die peak (e.g., the time frame where the most photons were incident on a SPAD or group of SPADs comprising a pixel) that corresponds to the distance from the SPAD array to an object in the imaged scene. Foveation may include utilizing a smaller number of bins of larger width centered around an estimated / predicted peak or a larger number of bins of smaller width centered around the estimated / predicted peak. Thus, the bin width is a duration of photon arrival time. Said in another way, bin width may be a set of time (e.g., change in time) that photons are collected for a bin. There is a specific bin width for memory foveation and depth foveation. Bin width may be a different numerical value for memory foveation compared to depth foveation.

[0119] Calculating the bin width may include, at 1012, calculating the bin width for memory foveation. Bin width for memory foveation may be represented by the symbol At and may be calculated via equation 9.TWidth of Bin For Memory Foveation = At = — (9)T is Temporal volume calculated from Z (e.g., maximum depth range of the SPC camera) and the speed of light (c). N is the number of bins across a full histogram. For example, a full histogram may comprise 1000 or more bins.Docket No. PSU24302PCT

[0120] Calculating the bin width may include, at 1016, calculating the bin width for depth foveation. Bin width for depth foveation may be represented by tire symbol Atdep*. Atde th may be calculated via equation 10.TWidth of Bin For Memory Foveation = Atrfepth= — (10) where N’ is number of bins used for depth foveation. Here, T is the temporal volume but is not based on the maximum depth range of the camera but is instead based on the foveation window (which may be the number of bins M multiplied by the bin width calculated for memory foveation).

[0121] In some examples, the system may be pre-configured to perform depth foveation or memory foveation, and thus whether the bin width is calculated for depth or memory' foveation may be preselected prior to executing method 1000 by whether memory foveation and / or depth foveation is selected in the settings of the computational device processing the SPAD data. If memory foveation is desired and selected in the settings of the computational device, the method 1000 at 1010 may carry out 1012 but not 1016. Likewise, if depth foveation is desired and selected in the settings of the computational device, the method 1000 may carry out 1012 but not 1016.

[0122] At 1018, method 1000 gathers photon data for a plurality of pixels (S) of a SPAD array. In some examples, each pixel may include one SPAD element. In other examples, each pixel may include more than one SPAD element, such as four SPAD elements. S is equivalent to the amount of pixels of the camera / SPAD array. The data gathered includes photon and time data to be binned (e.g., time- stamped photon events), captured across one or more pulses of light.

[0123] At 1020. a depth prior is acquired to estimate a foveation window for emory foveation and / or depth foveation techniques used via method 1000. The depth prior is an initial estimate of a depth for each pixel (e.g., a depth of the scene point captured by that pixel). Because depth is determined based on photon arrival time, the depth prior may indicate a predicted / estimated peak of the photon data after binning to the histogram. A foveation window, such as a memory' foveation window or a depth foveation window , may be positioned around a time indicated by the depth prior of each pixel.

[0124] The depth prior may be obtained via a plurality of ways at 1020. In a first example, indicated at 1022, the depth prior may be obtained from a monocular image. The monocular image is a 2D image that may be obtained from a 2D camera, such as an RGB camera (e.g.. visible light camera 111), positioned to capture the imaged scene prior to and / or at the same time as the photon data is collected. A monocular depth image algorithm may be applied to each monocular pixel of the monocular image to generate a monocular depth for each monocular pixel. The monocular depth of each monocular pixel may be used as the depth prior for a respective pixel of the SPAD array, where the pixel is in the same location as the monocular pixel. Additional detail about generating the depth prior with a monocular image is presented below in a method 1400 of FIG. 14. It is to be appreciated that the SPAD array may include more pixels than are present in the monocular image. Thus, when depth is determined for each monocular pixel (e.g., the depth prior), the depth for a given monocularDocket No. PSU24302PCT pixel may be applied to more than one pixel of the SPAD array. Likewise, it is to be appreciated that die SPAD array may include less pixels than are present in the monocular image. Image processing methods may be applied to change the spatial resolution of the monocular image such that the monocular image aligns with the SPAD pixels, such as reducing the quantity of and combining the pixels of the monocular image into larger pixels, where the larger pixels of the monocular image may be matched to share the same coordinate location as the SPAD pixels of the SPAD image. With an external camera, such as the RGB camera, the pixels of the monocular image may be co-located to the SPAD pixels through calibration and pixel-matching, as explained previously.

[0125] In a second example, indicated at 1032, the depth prior may be obtained via an optical flow method. Optical flow techniques may be used for a SPAD sensor / SPAD array that may be fixed to moving platform, such as an autonomous vehicle, and / or depth video or other sequential depth images where high-frame rate and depth capture with efficient memory consumption are desired. The optical flow method generates the depth prior based on known depths of a previous frame and calculated movement of imaged object(s) between the previous frame and the current frame. Additional detail about generating the depth prior using the optical flow technique is presented below in method 1200 of FIG. 12.

[0126] In a third example, indicated at 1036, the depth prior may be obtained via a super pixel sampling method. The super pixel sampling method may be used for when low resolution sampling is desired and / or when a separate monochrome or RGB camera is not available. The super-pixel sampling method may include generating a pseudo-intensity map from the photon data. Using an algorithm, the pseudo intensity map is segmented into a plurality of super pixels each comprising a plurality of pixels. At least one of the segmented pixels of the super pixels (e.g., the centroid pixel) has a depth determined via binning the photon and time data for that pixel via a full histogram, which may constitute the depth prior. During foveation, the foveation window for the pixels in tire super-pixel is centered around the depth for the centroid pixel.

[0127] At 1042, a foveated histogram is generated for each pixel of S as additional photon data is captured. Additional details for generating the foveated histogram are provided below with respect to FIG. 11. Briefly, once the depth prior is obtained, the foveation window (e.g., M or N' bins of the set bin width) for a selected pixel is centered around the photon arrival time that corresponds to the depth of that pixel from the depth prior. The photon data for the selected pixel is binned into the foveation bins (e.g., via a histogrammer). and this process is performed for each pixel. At 1052, an image is returned from the histogrammed pixels S, referred to herein as a histogram image (H). The histogram image is comprised of the foveated histogram for each of the pixels S. In another example, the histogram image may comprise a value for each pixel, where each value is the histogram peak for that pixel (and thus each value is a photon arrival time). The SPAD sensor / array may output the histogram image. At 1054, a depth image is created from H, in a technique referred to as decoding. The depth image may beDocket No. PSU24302PCT referred to herein as decoded depth image (D). To decode D from H, each pixel of H is assigned a depth value based on the peak of the histogram, where the peaks of the histograms are calculated into depths (e.g., the photon arrival time is transformed into a depth value). In examples where a depth image is generated, each depth value is assigned a color and / or brightness value upon decoding, such that each pixel of D has a depth value represented by a color and / or brightness assigned based on the histogram peak of the pixel. Each pixel of D decoded from a pixel H shares the same coordinates (x, y) as the pixel of H.

[0128] Method 1000 may include an optional step at 1060 that may be referred to as a spatial temporal optional step. The spatial temporal step may include, at 1062. quantizing the depth prior into discrete buckets (B). For example, when the depth prior is the monocular depth estimate (e.g.. a monocular depth image), the monocular depth image may be quantized into buckets, such that each pixel of the monocular depth image with the same depth value (or in the same range of depth values) is placed into the same bucket. Method 1000 continues to 1064 where at least one pixel of S is randomly selected from each bucket of B. However, it is to be appreciated that more than one pixel of S may be selected from each bucket of B. The random pixels of S specific to each bucket may be referred and represent via S. 1064 may alternatively be described as selecting S from S or reducing S to S. Method 1000 continues to 1066, where 1010, 1020, 1040, 1042, 1052, and 1054 are repeated for the pixels of S. Said in another way, S is used in place of S for generating the foveated histograms and histogram image. After 1066, method 1000 continues to 1068 where a depth image such as a depth map is created using the quantized sparse data determined at 1066. The depth image D(B) is D, but calculated using S in place of S for 1040, 1042, 1052 and 1054. For example, a foveated histogram is generated for each pixel S and not the remaining pixels. The depth value determined for each pixel S is applied to the other pixels in the same bucket. The relationship may be illustrated using equation 11.D(B) = min (D (S ) £ S) (11)

[0129] Turning to FIG. 11, it shows a method 1100 for foveating and generating a histogram for the pixels S. The method 1100 may be performed as part of method 1000, such as at 1042 of FIG. 10. At 1110, a pixel (Si) is selected from the total quantity of pixels S of the SPAD array. Initially, Si may be a first pixel (e.g., where Si is set to a starting pixel such as So or Si).

[0130] Each pixel of S has a different coordinate on a grid in 2D space. The coordinates may be represented by a value of x and a value of y, (x. y). Each S, of S has a specific value of (x, y) that is different from the other pixels of S, where at least the value of x or the value of y is different between S; and another pixel of S.

[0131] At 1112. an estimated peak (d ) is obtained from the depth prior for the selected pixel Si. The d specific to Si may be additionally be referred to as d (x, y) for being the estimated peak at the coordinates (x, y) of the pixel. The depth prior may be a monocular depth image, determined based on a prior frame, or determined from a pseudo-intensity map. For example, the depth prior may be from aDocket No. PSU24302PCT monocular image with depth estimated via a monocular depth algorithm. As another example, the depth prior may be from a prior depth image adjusted based on an optical flow method. As another example, the depth prior may be from a super pixel image with depth prior estimated via a super pixel sampling method. The depth prior is retrieved from a depth prior pixel sharing the same coordinates (x, y) with the pixel Si. The estimated peak d may be the photon arrival time back-calculated from the depth prior.

[0132] At 1114. the foveation window is positioned around the d (x, y). where d (x, y) is the center of the foveation window.

[0133] At 1122, it is determined whether memory foveation is being performed. If memory foveation is being performed (1122 is YES), method 1100 continues to 1124. At 1124, the bin width and number of bins are set to be At and M, respectively. Method 1100 continues to 1126. where a memory foveated histogram is captured by binning the photon and time data (e.g., the stream of photon return events) within the memory foveation window (e.g., with the M bins centered around the estimated peak) for the pixel S;. The memory foveation window is centered around the d (x. y) determined from the depth prior. The photon and time data that are outside of the memory foveation window are not bimred. The time span within the memory foveation window is segmented into M number of bins that are each At amount of time in width. Each photon event recorded for different times within the bin are summed in the bin. In this way, each bin may include a total photon count that represents the number of photons that impinged on the SPAD element(s) of that pixel for the time span of the bin. The bin w ith the greatest amount of photons is a peak of the memory foveated histogram. The histogramming may occur on the SPAD sensor (e.g., on the hardware chip that includes the SPAD array / SPAD pixels, via a histogrammer coupled to the selected pixel). Thus, method 1100 may include sending the histogram parameters (e.g., bin width, bin number, and depth prior / center of the foveation window) to the SPAD sensor. After 1126, method 1100 proceeds to 1142.

[0134] Returning to 1122, if memory foveation is not being performed (1122 is NO), method 1100 thus performs depth foveation and continues to 1134. At 1134, the bin width and number of bins are set to Atdep* and N’. respectively. The foveation window is a depth foveation window that is N’* Atdepth in width of time. The foveation window is centered around the peak at d (x, y). Method 1100 continues to 1136. where a depth foveated histogram is captured by binning the photon and time data (e.g., the stream of photon return events) within the foveation window for the pixel Si. Photon and time data that are outside of the foveation window are not binned. The time span within the depth foveation window is segmented into N’ number of bins that are each Atdepth amount of time in width. Each photon event recorded for different times within a bin are summed in the bin. In this way. each bin may include a total photon count that represents the number of photons that impinged on the SPAD element(s) of that pixel for the time span of the bin. The bin with the greatest amount of photons is a peak of the depth foveated histogram. As explained above, the histogramming may occur on the SPAD sensor (e g., via a histogrammer coupled to the selected pixel). Thus, method 1100 may include sending the histogramDocket No. PSU24302PCT parameters (e.g., bin width, bin number, and depth prior / center of the foveation window) to the SPAD sensor. After 1126, method 1100 proceeds to 1142.

[0135] At 1140, method 1100 determines whether there are more pixels of S to be binned / liistogrammed. If there are still pixels of S to be histogrammed (1142 is YES), method 1100 may return to 1110 to select the next pixel. Returning to 1142, if there are no pixels remaining to be histogrammed (1142 is NO), method 1100 may end.

[0136] For example, to determine if there are pixels of S to be histogrammed, S, may be increased by a value of Si (e.g., Si = Si + Si) and then the count for the new pixel of Si may be compared to S. If S, is less than or equal to S (1142 is YES), method 1100 may return to 1110. If Si is greater than S (1106 is NO), then method 1100 ends.

[0137] However, it is to be appreciated that the pixels of the SPAD array may undergo foveation histogramming simultaneously, and method 1100 depicts the process for histogramming a selected pixel that may be repeated serially for each pixel or performed simultaneously for each pixel.

[0138] FIG. 12 shows a method 1200 for obtaining a depth prior using optical flow techniques. Method 1200 may be perfonned as part of method 1000, such as at 1032 of FIG. 10. The method 1200 may be used for a sequence of depth images taken in succession, such as for a video comprising depth images for each frame. In some examples, the method 1200 may be used when the SPADs that capture data for the depth image are fixed to a moving platform. Each frame may have a different optical time (each frame may be captured at a different time). Optical time (to) is different from the time (t) values that are binned during photon capture. Optical time represents the time at which each frame (e g., optical flow image) is created. Each frame (e.g., optical flow image) may therein be paired with a complementary optical time and vice versa.

[0139] At 1212, method 1200 captures photon data for a first frame via the SPAD sensor. At 1212, photons may be captured and the times which the photons are captured are recorded for each pixel of the SPAD array / sensor. At 1214, a first two dimensional (2D) image of the first frame may be generated. The first 2D image may be captured with a monochrome or RGB camera (such as visible light camera 111), for example, and may be captured at the same time the photon data for the first frame is captured or captured prior to the photon data for the first frame being captured. At 1216. a depth is determined for each pixel of the first frame. The depth is determined via a full histogram, where the photon and time data for each pixel are binned into a respective histogram without foveation. As explained above, the histogramming may occur on the SPAD sensor (e.g., via a histogrammer coupled to the selected pixel). Thus, method 1200 may include sending the histogram parameters (e.g., bin width, bin number, and depth prior / center of the foveation window, which in this example would be the full temporal volume of the camera since foveation is not occurring) to the SPAD sensor. The depth of each pixel is determined from the peak of the respective histogram.Docket No. PSU24302PCT

[0140] At 1218, photon data are captured for a next frame via the SPAD sensor. The next frame may be a second frame after the first frame. However, the next frame may be another frame, such as a frame after the second frame. At 1222, a 2D image is generated for the next frame, the 2D image referred to herein as the next 2D image. The next 2D image may be captured by the monochrome or RGB camera prior to when the photon data is captured.

[0141] At 1224, a vector (u.v) is created for each pixel of the next frame based on the 2D image of the prior frame (e.g., the 2D image of the first frame), referred to as a prior 2D image herein, and the next 2D image. Each vector (u.v) is generated such that is upholds the brightness consistency principle represented by equation 14.I(x, y, t) = l(x + u * 8t, y + v * St t, t + St) (14)Where St is the change in time between the 2D frames (e.g., a previous and a current frame). In this way, the brightness (I) of a pixel (x.y) at a current time t (in a current 2D image) is a function of the brightness of that pixel in the previous frame (in a previous 2D image), adjusted by the factors u and v, which account for movement of the object.

[0142] At 1226, a depth estimate is generated for each pixel of the next frame using the vector (u,v) for that pixel and the depth of that pixel from the prior frame. The depth estimate for each pixel collectively forms the depth prior. When the next frame is the second frame that follows the first frame, die depth of that pixel from the prior frame may be the depth determined from the full histogram. For subsequent frames, the depth may be the depth determined using foveation of the prior frame. The depth for a given pixel from a prior frame may be shifted to a new pixel or otherwise adjusted based on the vector to form the depth estimate. As explained previously, the depth prior may be used to center the foveation window of each pixel.

[0143] At 1228, a foveated region of each pixel of the next frame (after the photon data has been collected and binned via a foveated histogram) is compared to noise floor (e.g., noise threshold). For example, the comparison to noise may be done by comparing the distribution of photons in the foveated w indow to a photon distribution of a noise floor, where the noise floor is a threshold of photons below which a photon reading may be considered noise. The noise floor may be calculated from each prior pixel of the prior frame for each next pixel of the next frame, where prior pixel and the next pixel sharing the same noise floor share the same coordinates (x,y) on the prior frame and the next frame, respectively. For example, the noise floor may be determined via photon counting. More specifically, in 1212 a full histogram may be captured for each pixel which give the amount of photons captured per bin over the entire time axis of the SPAD array. From these photon counts, the bins with the highest photon cormts are identified and defined as a signal. The threshold value may be defined as a fraction of the maximum photon count in the previously -defined signals, where anything below this threshold is now defined as the noise floor. As another example, profiles may be generated for regions of the histogram defined as signal and noise. For example, a profile may be created based on the full resolution histograms (e.g.,Docket No. PSU24302PCT determined from the photon counts captured at 1212) to signal -match the regions of the histogram defined as the desired signal, based on photon count and the shape of the signal. Once a desired signal has been found, any region in the histogram not defined as the signal can be considered part of the noise floor. The noise floor may also be predetermined value, such as a value calculated from an average noise of the first pixels from the first image. Alternatively, the threshold may be a minimum SNR ratio for the foveated region of the histogram.

[0144] At 1232, method 1200 determines if the photon data of any pixel of the next frame meets a condition relative to the threshold of the noise floor. For example, if the photon distribution of a pixel (after foveation using the depth prior calculated with optical flow as described above) matches or is within a threshold range of the photon distribution of the noise floor, the condition may be not met. Essentially, in optical flow methods, noise may propagate / accumulate and each foveated histogram may be evaluated to determine if the depth prior used for foveation was accurate. If the photon data of each pixel meets the condition (1232 is YES), the depth generated at 1226 is determined to be accurate and usable for foveation, and method 1200 continues to 1236. Returning briefly to 1232, if the photon data of any pixel does not meet the condition (1232 is NO), method 1200 proceeds to 1242. At 1242, method 1242 discards the depth prior for the identified pixels and estimates depth of the identified pixel(s) via a full histogram when the next frame is captured. The calculation of depth using the full histogram may be performed for each pixel of the next frame that does not meet the condition relative to the noise threshold. It is to be appreciated that the depth error is determined once foveation has occurred and the depth is determined for each pixel of the current frame. The pixels that exhibit depth error may be flagged and for the subsequent frame, rather than use a depth prior and foveate the photon data for the flagged pixels, the depth for the flagged pixels is determined using a full resolution histogram.

[0145] At 1236. method 1200 determines if the next frame is the last frame of a sequence of frames. If the next frame is the last frame (1236 is YES), method 1200 may end. If the next frame is not the last frame (1236 is NO), method 1200 proceeds to 1252. At 1252, method 1200 store the data of the next frame and then sets the next frame to be a new prior frame. From 1252, method 1200 returns to 1218, where photon and time data are captured for another next frame, and the new prior frame is used as the prior frame therein.

[0146] In this way, optical flow methods may apply the depth of each pixel of a prior frame to a current frame, taking into account movement of objects from the prior frame to the current frame (determined from optical flow, which may generate movement vectors for each pixel based on 2D images corresponding to the prior frame and the current frame). The depth images generated using foveation based on optical flow methods may be prone to error and thus each depth image is evaluated for error. Pixels displaying error / noise above a threshold may have their depths calculated using a full resolution histogram on the next pass (e.g., the next frame).Docket No. PSU24302PCT

[0147] FIG. 13 shows the method 1300 for generating a depth prior from super pixel sampling techniques. In method 1300, an image is generated that is segmented into super pixels, such as the third image 676, from a pseudo intensity map, such as tire second image 674. The method 1300 also selectively creates non-foveated histograms and estimates a true photon peak from data of at least a pixel of the super pixel to assign depth to the true photon peak pixel, while using the non-foveated histogram and true photon peak from true photon peak pixel to estimate a foveation window to foveate the other pixels of the super pixel.

[0148] Method 1300 may be carried out as part of method 1000. such as at 1036 of method 1000. At 1312, method 1300 captures photon data via the SPAD sensor, as described previously.

[0149] At 1314. a pseudo-intensity map is generated from a first portion of the photon data. The pseudo-intensity’ map may be created by summing up raw photon data cubes along the temporal axis for each of the pixels or based on output from a counter in each SPAD pixel. In this way, an intensity of each pixel may be generated from the total number of photons that impinge on the SPAD(s) of the pixel. The pseudo-intensity map may have of the same spatial resolution as the SPAD sensor. However, in some examples, the pseudo-intensity map may be captured by the SPAD array using a lower temporal resolution, meaning the bin width is larger and that there are a lower number of bins in the captured photon-cube. For example, at 1312, the photon data may be captured at a lower temporal resolution, with 10 bins across the temporal axis each bin with a width of 10ns. Then, at 1318 (explained in more detail below), the full resolution histogram would be captured with 1000 bins and bin widths of 0. Ins.

[0150] At 1316, a segmentation algorithm is applied to the pseudo-intensity map to obtain a plurality of coarse segmentations. The course segmentations divide and group the pixels of the pseudointensity map into a plurality of super-pixels. The pixels grouped by the super-pixels may be referred to herein as segmented pixels. Each of the super-pixels comprises a plurality of segmented pixels, such as tens to hundreds of pixels.

[0151] At 1318, a full resolution histogram is generated for each centroid pixel (the centermost pixels of the super pixels), where each super pixel has one centroid pixel, using a second portion of the photon data to determine the depth of each centroid pixel. Each centroid pixel may be mapped back to a SPAD pixel, and the full resolution histogram is generated from the photon data captured by the SPAD pixel. The full resolution histograms are generated by respective histogrammers of the SPAD sensor. At least a centroid pixel (center pixel) is selected from each of the super pixels. The true peak location of the centroid pixel is evaluated using the full resolution histogram (non-foveated histogram). The full resolution histogram is decoded and has depth assigned to the centroid pixel of each pixel. The depth of each centroid pixel may collectively form the depth prior.

[0152] At 1320, the foveation w indow s for each of the other segmented pixels of the super-pixels are centered based on a respective depth prior. For example, the depth of a selected centroid pixel is a first depth prior for a specific super-pixel and thus the other segmented pixels of the selected superDocket No. PSU24302PCT pixel. The foveation window for the other segmented pixels of die selected super-pixel is a specific super-pixel foveation window, where the super-pixel foveation window is a l / 4th the width along time axis of the full window, centered based on die first depth prior (e.g., the depth of the centroid pixel). Method 1300 continues to 1322 where a third portion of the photon data is bimied into foveated histograms bins using the foveation windows. Thus, once the depth of each centroid pixel is determined, as further photon data is captured, the foveated histograms are created with the foveation windows centered around the depth of the respective centroid pixel.

[0153] FIG. 14 shows the method 1400 for generating a depth prior from a monocular image sampling technique. In method 1400, a monocular image and a frame generated from photon data recorded by a SPAD array are compared. More specifically, depth estimates of a set of the monocular image pixels of the monocular image are compared to depths of a set of a plurality pixels sampled from the frame taken using the SPAD array. A relationship may be created between the depths of the pixels of the monocular image and the depths of the sampled pixels of the SPAD frame to generate a depth prior to be used to foveate subsequent photon data.

[0154] Method 1400 may be performed as part of method 1000 of FIG. 10, such as at 1022 of method 1000. At 1410, photon data are captured via the SPAD array, as explained previously. In some examples, a frame of photon data may be captured.

[0155] At 1412, a monocular image is captured and a depth of each pixel of the monocular image is estimated (e g., to form a monocular depth image). The monocular image may be a 2D image captured via a 2D camera (e.g., visible light camera 111), such as image 212 of FIG. 2-8 or the images of column 318 of FIG. 3. The monocular image may have the same amount of pixels as the SPAD pixels. In other examples, image processing techniques may be applied to increase or decrease the resolution of the monocular image to match the spatial resolution of the SPAD array. The pixels of the monocular image may be referred to herein as monocular pixels. The monocular image may be captured prior to or at the same time that the photon data are captured. The depth of each pixel of the monocular image may be determined using a suitable monocular depth estimate algorithm, such as ZoeDepth.

[0156] At 1414, a first set of SPAD pixels are sampled from the frame of photon data, referred to as sampled SPAD pixels. The sampled SPAD pixels may be selected randomly across the entirety of the SPAD array. At 1416, each of the sampled SPAD pixels are binned using a full histogram. The histogram of each sampled SPAD pixel is decoded, and a depth decoded is assigned to the sampled pixel. Thus, for the photon data captured at 1410, a subset of the pixels is sampled and captured photon data thereof binned into a full resolution histogram to determine depth for the sampled pixels.

[0157] At 1418. corresponding pixels of the monocular image are sampled, referred to herein as sampled monocular pixels. The sampled monocular pixels share coordinates (x, y) with the sampled SPAD pixels. At 1420, is the depth estimate for each of tire sampled monocular pixels is obtained from die monocular depth image.Docket No. PSU24302PCT

[0158] At 1422, a relationship is derived betw een the depths of the sampled monocular pixels and the depths of the sampled SPAD pixels. For example, tire relationship may be modeled as a polynomial fit. At 1424, the relationship is applied to the depth estimates of the monocular pixels of the monocular image (e.g., to the pixels of the monocular depth image) to generate the depth prior. Each monocular pixel may have a first depth estimated using the monocular depth assigning algorithm. Each of the first depths may be transformed via the relationship into a respective second depth. Each of the second depths may collectively form the depth prior for the pixels of S when foveation through method 1100 of FIG. 11 is performed, as pixels of S are the SPAD pixels.

[0159] Thus, monocular depth estimates may be used to form the depth prior. Monocular depth estimation is brittle due to training dataset biases. In contrast. SPADs provide high-accuracy sensor measurements. The inaccurate monocular depth is leveraged to reduce the number of SPAD bins to capture, saving memory and improving depth resolution. To do this, the inaccurate monocular depth estimations are calibrated to the scene. This can be done either locally, fitting data to a particular scene, or generally across the dataset. In either case, a small set of pixels is sampled at full histogram resolution and the relationship between the monocular estimate and the SPAD estimate at these pixels is modeled by a polynomial fit.

[0160] FIG. 15 shows a schematic 1500 of a control and signal diagram for a SPAD array used in hardware of the present disclosure, such as SPAD array 124 of FIG. 1. The schematic 1500 comprises a detection circuit 1508 and a SPAD 1510. The detection circuit 1508 may be an enable, quench and reset circuit. In some examples, the detection circuit 1508 may be an active quenching circuit with readout circuitry.

[0161] The schematic 1500 comprises a plurality of pathways for electrical signal settings, including a SPAD bias 1512 and a global ramp voltage signal 1518. The SPAD bias 1512 is configured to supply the SPAD 1510 with reverse-bias voltage. The reverse-bias voltage allows avalanche current to develop when a photon impinges on the SPAD. and the avalanche current is quenched via the active quenching circuit of the detection circuit 1508. The detection circuit 1508 may be enabled and disabled based on the global ramp voltage signal 1518 relative to an enable threshold 1516 and a disable threshold 1514. The enable threshold 1516 may be selected to control the detection circuit 1508 to be enabled and therein detecting signals from the SPAD 1512. The disable threshold 1514 may be selected to control the detection circuit 1508 to be disabled and therein not detecting signals from the SPAD 1512. The global ramp signal 1518 may gradually increase, such as with a linear slope, and when increased above the disable threshold may disable the detection circuit. Likewise, the global ramp signal 1518 may rapidly decrease, such as in a stepwise manner, to decrease below the enable threshold 1516. In this way, the global ramp signal may be pulsed to enable photon detection and then signal quenching.

[0162] The SPAD 1510 may detect incident photons 1524 that contact the SPAD 1510. Incident photons 1524 and the detection time of the incident photon 1524 may be recorded by the detectionDocket No. PSU24302PCT circuit 1508. Recorded incident photons 1524 and their paired detection times may be sent by the detection circuit to a time-to-digital converter (TDC) and histogrammer 1520.

[0163] FIG. 16 shows a schematic diagram of a SPAD sensor 1600 shown schematically and electronic inputs and outputs shown schematically. In some examples, SPAD sensor 1600 may be included as part of an SPC camera, such as assembly 104 of FIG. 1. For example, the SPAD array 124 may be included on a SPAD sensor similar to SPAD sensor 1600, and thus SPAD array 124 is a nonlimiting example of SPAD array 1602 of FIG. 16.

[0164] The SPAD sensor 1600 includes a SPAD array 1602 that may be arranged on a backing 1604. The SPAD array 1602 includes a plurality of SPADs 1610. The SPADs 1610 may be non-limiting examples of the SPAD 1510 of FIG. 15. Each of the SPADs may be electrically and communicatively coupled to a detection circuit, such as the detection circuit 1508. The SPAD sensor 1600 may also include the TDC and histogrammer 1520. where the TDC and histogrammer 1520 is housed via the backing 1 04.

[0165] The SPADs of the SPAD array 1602 are organized into macropixel groups, where each macropixel group comprises a subset of SPADs. Each SPAD of a SPAD macropixel group is electrically and communicatively coupled with its own a detection circuit or a common detection circuit for the macropixel. Each SPAD macropixel group may be coupled to a respective TDC and histogrammer. For example, the SPAD array may include a SPAD macropixel group 1640. The SPAD macropixel group 1640 may comprise four SPAD pixels, including a first SPAD 1642, a second SPAD 1644, a third SPAD 1646, and a fourth SPAD 1648. The first SPAD 1642, the second SPAD 1644, the third SPAD 1646, and the fourth SPAD 1648 may be electrically and communicatively coupled to the TDC and histogrammer 1520. A SPAD of the SPAD pixel group 1640 may be the SPAD 1510. Photon and time data from incident photons received by the first SPAD 1642, the second SPAD 1644. the third SPAD 1646, and / or the fourth SPAD 1648 may be histogrammed as disclosed herein via the TDC and histogrammer 1520. In some examples, each macropixel may generate data (e.g., an M bin histogram) that may be used to determine depth for a depth map or image, with a 1 : 1 correspondence of macropixels to image pixels.

[0166] The SPAD sensor 1600 includes a plurality of row drivers 1612 and column drivers 1614 that may be electrically and communicatively coupled with the electronic components of the SPAD array 1602, including the SPADs 1610 to control the SPADs.

[0167] The SPAD sensor 1600 may receive and generate input / output signals 1632 from various devices. For example, the SPAD sensor 1600 may receive the global ramp signal 1518 from a global ramp generator 1616. Additionally, a VBD signal 1622, a VEX signal 1624, a laser synchronization signal 1626, and a clock (CLK) signal 1628 may be input into the SPAD sensor 1600 via electrical and communicative couplings. In some examples, the inputs may include bin width, bin number, and / or foveation windows for controlling the histogramming performed by the histograrmners. An M binDocket No. PSU24302PCT histogram (or N’ bin histogram in the example of depth foveation) may be generated and output as a signal by each histogrammer of the SPAD sensor 1600. More specifically, the M bin histogram may be generated by the TDC and histogrammer 1520 and output as a signal via the SPAD sensor 1600. The histogrammer may generate a histogram for each pixel (e.g., each SPAD) in the macropixel sequentially, or the histogrammer may generate a histogram for the macropixel as a whole.

[0168] Each SPAD of a macropixel may be electrically coupled to a selective earth leakage (SEL) component 1652. The SEL component 1652 may isolate and remove electrical abnormalities, such as currents and / or voltages above a threshold. For example, excess current and voltage may escape the detection circuit 1508 via the SEL component 1652 to prevent degradation to the detection circuit 1508.

[0169] Thus, SPAD sensor 1600. when incorporated in a SPC camera and / or otherw ise coupled to additional processing circuitry, such as the image processor system 141 of FIG. 1. may be configured to receive, at a histogrammer of the SPAD sensor, a stream of photon return events from a pixel of the SPAD sensor, the stream of photon return events generated by photons transmitted from a pulsed light source and reflected off an object in a scene. The SPAD sensor may further be configmed to bin. with the histogrammer, only a subset of the photon return events into a plurality7of bins of a histogram, the plurality7of bins centered around an estimated histogram peak. The SPAD sensor may be configmed to output, from the histogrammer, the histogram, the histogram usable (e.g., by the image processor system 141) to determine a distance of the object in the scene. The stream of photon return events may include time-stamped signals generated from the pixel (e.g., SPAD element) of the SPAD array in response to die photons impinging on the pixel. In this way, each photon return event may comprise a time delay of a voltage pulse generated by the pixel relative to a pulse time of the pulsed light somce. The histogrammer may be configured to, when burning only the subset of the photon return events into the plurality of bins of the histogram, bin each photon return event of the subset of the photon return events into a respective bin of the plurality of bins based on the time delay of that photon return event, and wherein any photon return events that fall outside the plurality of bins are not binned.

[0170] The estimated histogram peak may be determined based on a depth prior. In some examples, the depth prior may be generated based on a monocular depth estimate from a single image of the scene. The single image of the scene may be captured with a visible light camera (e.g., monochrome or RGB) calibrated such that pixels of the single image scene are co-located with or map to pixels of the SPAD sensor. The monocular depth estimate from the single image may be an estimated depth of each pixel of the single image using a monocular depth estimation algorithm. The monocular depth estimate from the single image may be corrected based on a set of depths determined from the SPAD sensor. A set of pixels of the SPAD sensor may be sampled in a prior frame and the depth of the sampled pixels determined using full-resolution histograms. A relationship (e.g., using a polynomial fit) may7be determined based on the depth of the sampled pixels relative to the depth estimates of the corresponding pixels of the single image. The monocular depth estimate from the single image may beDocket No. PSU24302PCT corrected by applying the relationship to the depth estimates of each pixel of the single image. The corrected depth of the pixel of the single image that corresponds to the pixel of the SPAD sensor may be used to center the plurality of bins for histogramming. In other examples, the stream of photon return events define a current frame, and tire depth prior is generated based a depth map of a prior frame and movement of the object between the prior frame and the current frame. The depth map of the prior frame may be determined from full-resolution histograms or by applying foveation. When movement of the object is detected, the depth of the pixels imaging the object from the depth map may be adjusted to account for the movement. The movement of the object may be determined based on a first visible light image captured at or nearly at the same time as the prior frame (e.g., of photon data) and a second visible light image captured after the prior frame and immediately prior to the current frame. The depth of the pixel from the prior frame, adjusted to account for movement of the object, may be used to center the plurality of bins for histogramming. In still further examples, the stream of photon return events define a current frame, and the depth prior is generated based a pseudo-intensity map generated from photon return events of a first prior frame and calculated depths of a subset of pixels of the first prior frame or a second prior frame. For example, the current frame may be a third frame following a first frame (the first prior frame) and a second frame (the second prior frame). The photon return events of the first frame (generated by each pixel of the SPAD sensor) may be used to calculate the pseudointensity map. The pseudo-intensity map may be generated by counting the photons received at each pixel of the SPAD sensor. The pseudo-intensity map may be segmented into a plurality of super-pixels. Each super-pixel may include a respective subset of pixels that are located in the same area of the pseudo-intensity map and have the same or nearly the same intensity. A respective centroid pixel may be identified for each super-pixel. The centroid pixel may be the center-most pixel or a selected pixel near the center of the super-pixel. The photon return events for the second frame (generated by each pixel of the SPAD array or only the centroid pixels) may be used to determine a depth of each centroid pixel. For example, the photon return events of the second frame for each centroid pixel may be binned into a full-resolution histogram to identity' the depth of each centroid pixel (e.g., based on the peak of each full-resolution histogram) in the second frame. Binning the photon return events of the current frame (the third frame) may then include binning the photon return events received from the pixel into a respective bin of the plurality of bins, wherein the plurality of bins is centered based on the depth of a corresponding centroid pixel. For example, the pixel may be assigned to a super-pixel as explained above. The centroid pixel of the assigned super-pixel may be identified and the depth of the centroid pixel determined as explained above. The depth of the centroid pixel may correspond to the estimated peak of the histogram.

[0171] The SPAD sensor 1600 may be a single-photon counting detector comprising a plurality of pixels, and the plurality of pixels may include a first pixel. The SPAD sensor 1600 may further include a histogrammer coupled to the first pixel and configured to generate a foveated histogram from a stream of photon return events from the first pixel generated by photons transmitted from a pulsed light sourceDocket No. PSU24302PCT and reflected off an object in a scene. The foveated histogram may include a plurality of bins spanning a temporal range that is smaller than a full temporal range of the single-photon counting detector, and wherein the plurality of bins is centered around an estimated peak of the foveated histogram. For example, the plurality of bins may be a first plurality of bins and tire histogrammer may be further configured to generate a full resolution histogram from a second stream of photon return events. The full resolution histogram may include a second plurality of bins spanning the full temporal range. In some examples, the first plurality of bins includes fewer bins than the second plurality of bins and each bin of the first plurality of bins has a same bin width as each bin of the second plurality of bins. In other examples, the first plurality of bins includes the same number of bins as the second plurality' of bins and each bin of the first plurality of bins has a smaller bin width than a bin width of each bin of the second plurality' of bins. The estimated peak of the foveated histogram may be determined using a depth prior, which may be generated as explained above.

[0172] This disclosure provides support for a method to enable efficient SPAD use on low-power platforms, where efficiency includes: data efficiency defined by reducing consumption of memory' for storage and bandwidth, and computational efficiency reducing computational power requested for SPAD data processing and creation of depth images.

[0173] The method disclosed herein may reduce memory consumption for the same depth accuracy or increase the accuracy of depth relative to a ground truth image compared to other methods of the art that uses SPAD LiDARs and other SPAD based imaging for imaging depth.EXAMPLES

[0174] In certain examples, a method includes: receiving, at a histogrammer, a stream of photon return events from a pixel of an imaging detector, the stream of photon return events generated by photons transmitted from a pulsed light source and reflected off an object in a scene; binning, with the histogrammer. only a subset of the photon return events into a plurality of bins of a histogram, the plurality' of bins centered around an estimated histogram peak; and outputting, from the histogrammer, the histogram, the histogram usable to determine a distance of the object in the scene.

[0175] In certain examples, each photon return event comprises a time delay of a voltage pulse generated by the pixel relative to a pulse time of the pulsed light source.

[0176] In certain examples, binning, with the histogrammer. only the subset of the photon return events into the plurality of bins of the histogram comprises binning each photon return event of the subset of the photon return events into a respective bin of the plurality of bins based on the time delay of that photon return event.

[0177] In certain examples, any photon return events that fall outside the plurality of bins arc not binned.Docket No. PSU24302PCT

[0178] Certain examples further include determining the estimated histogram peak based on a depth prior, wherein the stream of photon return events define a current frame.

[0179] In certain examples, the depth prior includes an estimate of a depth of the object for the pixel determined based on image data and / or photon data of one or more prior frames.

[0180] In certain examples, the depth prior is generated based on the image data of one or more prior frames including a monocular depth estimate from a single image of the scene, the monocular depth estimate corrected based on a set of depths calculated from the photon data of the one or more prior frames using full-resolution histograms.

[0181] In certain examples, the depth prior is generated based a depth map determined from the photon data of the one or more prior frames and movement of the object between the prior frame and the current frame, determined based on the image data of the one or more prior frames.

[0182] In certain examples, the depth prior is generated based a pseudo-intensity' map generated from photon return events of a first prior frame and calculated depths of a subset of pixels of the first prior frame or a second prior frame.

[0183] In certain examples, a system includes: a single-photon counting detector comprising a plurality of pixels, the plurality of pixels including a first pixel; and a histogrammer coupled to the first pixel and configured to generate a foveated histogram from a stream of photon return events from the first pixel generated by photons transmitted from a pulsed light source and reflected off an object in a scene.

[0184] In certain examples, the foveated histogram includes a plurality of bins spanning a temporal range that is smaller than a full temporal range of the single-photon counting detector.

[0185] In certain examples, the plurality of bins is centered around an estimated peak of the foveated histogram.

[0186] In certain examples, the plurality of bins is a first plurality of bins and wherein the histogrammer is further configured to generate a full resolution histogram from a second stream of photon return events, the full resolution histogram including a second plurality of bins spanning the full temporal range.

[0187] In certain examples, the first plurality of bins includes fewer bins than the second plurality of bins and each bin of the first plurality of bins has a same bin width as each bin of the second plurality of bins.

[0188] In certain examples, the first plurality of bins includes the same number of bins as the second plurality of bins and each bin of the first plurality of bins has a smaller bin width than a bin width of each bin of the second plurality of bins.

[0189] In certain examples, the estimated peak of the foveated histogram is determined using a depth prior.Docket No. PSU24302PCT

[0190] The control and estimation routines disclosed herein may be stored as executable instructions in non-transitoiy memory and may be carried out by a control sy stem. The specific routines described herein may represent one or more of any number of processing strategies such as event- driven, interrupt-driven, multi-tasking, multi-threading, and the like. As such, various actions, operations, and / or functions illustrated may be performed in the sequence illustrated, in parallel, or in some cases omitted. Likewise, the order of processing is not necessarily required to achieve the features and advantages of the example embodiments described herein, but is provided for ease of illustration and description. One or more of the illustrated actions, operations, and / or functions may be repeatedly performed depending on the particular strategy being used. Further, the described actions, operations, and / or functions may graphically represent code to be programmed into non-transitory memory of the computer readable storage medium.

[0191] The description of embodiments has been presented for purposes of illustration and description. Suitable modifications and variations to the embodiments may be performed in light of the above description or may be acquired from practicing the methods. The methods may be performed by executing stored instructions with one or more logic devices (e.g.. processors) in combination with one or more hardware elements, such as storage devices, memory, hardware network interface s / antennas, switches, actuators, clock circuits, and so on. The described methods and associated actions may also be performed in various orders in addition to the order described in this application, in parallel, and / or simultaneously. The described systems are exemplary in nature, and may include additional elements and / or omit elements. The subject matter of the present disclosure includes all novel and non-obvious combinations and sub-combinations of the various systems and configurations, and other features, functions, and / or properties disclosed.

[0192] As used herein, the terns “system” or “module” or “modulator” may include a hardware and / or software system that operates to perform one or more functions. For example, a module or system may include a computer processor, controller, or other logic-based device that performs operations based on instructions stored on a tangible and non-transitory computer readable storage medium, such as a computer memory. Alternatively, a module or system may include a hard-wired device that performs operations based on hard-wired logic of the device. Various modules or units shown in the attached figures may represent the hardware that operates based on software or hardwired instructions, the software that directs hardware to perform the operations, or a combination thereof.

[0193] The foregoing described aspects depict different components contained within, or connected with different other components. It is to be understood that such depicted architectures are merely exemplary, and that in fact many other architectures can be implemented which achieve the same functionality. In a conceptual sense, any arrangement of components to achieve the same functionality' is effectively "associated" such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as "associated with" eachDocket No. PSU24302PCT other such that the desired functionality is achieved, irrespective of architectures or intennedial components. Likewise, any two components so associated can also be viewed as being "operably connected," or "operably coupled," to each other to achieve the desired functionality.

[0194] As used in this application, an element or step recited in the singular and proceeded with the word “a” or “an” should be understood as not excluding plural of said elements or steps, unless such exclusion is stated. Furthermore, references to “one embodiment” or “one example” of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features. The terms “first,” "second,” “third,” and so on are used merely as labels, and are not intended to impose numerical requirements or a particular positional order on their objects.

[0195] The following claims particularly point out certain combinations and sub-combinations regarded as novel and non-obvious. These claims may refer to "an" element or "a first" element or the equivalent thereof. Such claims should be understood to include incorporation of one or more such elements, neither requiring nor excluding two or more such elements. Other combinations and subcombinations of the disclosed features, functions, elements, and / or properties may be claimed through amendment of the present claims or through presentation of new claims in this or a related application. Such claims, whether broader, narrower, equal, or different in scope to the original claims, also are regarded as included within the subject matter of the present disclosure.

Claims

Docket No. PSU24302PCTCLAIMS:

1. A method, comprising: receiving, at a histogrammer, a stream of photon return events from a pixel of an imaging detector, the stream of photon return events generated by photons transmitted from a pulsed light source and reflected off an object in a scene; binning, with the histogrammer, only a subset of the photon return events into a plurality of bins of a histogram, the plurality of bins centered around an estimated histogram peak; and outputting, from the histogrammer, the histogram, the histogram usable to determine a distance of the object in the scene.

2. The method of claim 1, wherein each photon return event comprises a time delay of a voltage pulse generated by the pixel relative to a pulse time of the pulsed light source.

3. The method of claim 2, wherein binning, with the histogrammer, only the subset of the photon return events into the plurality of bins of the histogram comprises binning each photon return event of the subset of the photon return events into a respective bin of the plurality of bins based on the time delay of that photon return event.

4. The method of claim 3. wherein any photon return events that fall outside the plurality of bins are not binned.

5. The method of claim 4, further comprising determining the estimated histogram peak based on a depth prior, wherein the stream of photon return events define a current frame.

6. The method of claim 5, wherein the depth prior includes an estimate of a depth of the object for die pixel determined based on image data and / or photon data of one or more prior frames.

7. The method of claim 6, wherein the depth prior is generated based on the image data of one or more prior frames including a monocular depth estimate from a single image of the scene, the monocular depth estimate corrected based on a set of depths calculated from the photon data of the one or more prior frames using full-resolution histograms.

8. The method of claim 6, wherein the depth prior is generated based a depth map determined from the photon data of the one or more prior frames and movement of the object between the prior frame and the current frame, determined based on the image data of the one or more prior frames.Docket No. PSU24302PCT9. The method of claim 6, wherein the depth prior is generated based a pseudo-intensity map generated from photon return events of a first prior frame and calculated depths of a subset of pixels of the first prior frame or a second prior frame.

10. A system, comprising: a single-photon counting detector comprising a plurality of pixels, the plurality of pixels including a first pixel; and a histogrammer coupled to the first pixel and configured to generate a foveated histogram from a stream of photon return events from the first pixel generated by photons transmitted from a pulsed light source and reflected off an object in a scene, , and.

11. The system of claim 10, wherein the foveated histogram includes a plurality of bins spanning a temporal range that is smaller than a full temporal range of the single-photon counting detector.

12. The system of claim 11, wherein the plurality of bins is centered around an estimated peak of the foveated histogram.

13. The system of claim 12, wherein the plurality of bins is a first plurality of bins and wherein the histogrammer is further configured to generate a full resolution histogram from a second stream of photon return events, the full resolution histogram including a second plurality of bins spanning the full temporal range.

14. The system of claim 13, wherein the first plurality of bins includes fewer bins than the second plurality' of bins and each bin of the first plurality' of bins has a same bin width as each bin of the second plurality' of bins.

15. The system of claim 13, wherein the first plurality of bins includes the same number of bins as the second plurality of bins and each bin of the first plurality of bins has a smaller bin width than a bin width of each bin of the second plurality of bins.

16. The system of claim 13, wherein the estimated peak of the foveated histogram is determined using a depth prior.