Systems, methods, and media for Euler-type single-photon computer vision
Euler-type single-photon computer vision techniques process raw data using phase-based methods to address the limitations of SPADs in passive imaging, achieving fast and efficient edge detection and motion estimation.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- WISCONSIN ALUMNI RES FOUND
- Filing Date
- 2024-05-01
- Publication Date
- 2026-05-19
AI Technical Summary
Conventional single-photon avalanche diodes (SPADs) are less suitable for passive imaging tasks under uncontrolled lighting conditions due to lower quality data generation, limiting their application in machine vision and other conventional imaging tasks.
Implementing Euler-type single-photon computer vision techniques that process raw single-photon data using computationally lightweight phase-based methods, such as 3D convolution kernels and velocity tuning filters, to directly extract scene information without explicit image reconstruction.
Achieves significant computational speedups of over two orders of magnitude for tasks like edge detection and motion estimation, making real-time processing feasible and reducing data transfer costs by analyzing data on-chip.
Smart Images

Figure 2026516024000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This application claims priority to U.S. Application No. 18 / 310,399, filed May 1, 2023, the content of which is hereby incorporated by reference in its entirety.
[0002] (Description of research funded by the federal government) This invention was made with government support under Grant No. 1943149 awarded by the National Science Foundation. The government has certain rights in this invention.
Background Art
[0003] Image sensors in conventional digital cameras usually capture hundreds to thousands of photons per pixel to generate an image. In recent years, single - photon avalanche diodes (SPADs) that can detect individual photons and accurately measure their arrival times have become popular. Due to the achievable sensitivity and picosecond time resolution of SPADs, such as imaging at extremely high frame rates (e.g., exceeding 1 billion frames per second), non - line - of - sight (NLOS) imaging, and microscopic imaging of nanoscale biological phenomena, it has promoted the development of new devices with new functions.
[0004] However, these new SPAD-based imaging techniques are typically active, with the SPAD being used in precise time synchronization with an active light source (e.g., a pulsed laser). This includes applications such as NLOS imaging, LiDAR, and microscopy. Due to the SPAD's output (e.g., the detection of a single photon in precise time), SPADs are less suitable for more conventional imaging tasks, such as capturing images of a scene under passive, uncontrolled lighting (e.g., sunlight, moonlight). While passive SPAD-based imaging systems have the potential to expand the SPAD's range to a wider range of applications, including machine vision, data generated from passive SPAD-based systems has so far been of relatively lower quality compared to images captured using conventional image sensors.
[0005] Therefore, new systems, methods, and media for Euler-type single-photon computer vision are desired. [Overview of the project]
[0006] According to some embodiments of the disclosed subject matter, systems, methods, and media for Euler-type single-photon computer vision are provided.
[0007] According to some embodiments of the disclosed subject matter, a system is provided for facilitating a single-photon computer vision task, comprising a plurality of detectors configured to detect the arrival of individual photons, the plurality of detectors comprising an image sensor arranged in an array, and at least one processor, the at least one processor being programmed to cause the image sensor to generate a sequence of images representing a scene, each of which includes a plurality of pixels; for each of the plurality of three-dimensional filters, to perform a convolution between the three-dimensional filter and a plurality of frames, each of which is based on one or more images of the sequence of images; for each of the plurality of frames, to generate a plurality of filter bank responses corresponding to the three-dimensional filters among the plurality of three-dimensional filters; and to perform a computer vision process based on the plurality of filter responses.
[0008] According to some embodiments, each of the multiple detectors includes a single-photon avalanche diode (SPAD).
[0009] According to some embodiments, each image in the sequence of images has an exposure time TIFF2026516024000002.tif6150 contains a binary image representing photons detected by the image sensor.
[0010] According to some embodiments, each of the plurality of three-dimensional filters includes a speed tuning filter, and a first subset of the plurality of three-dimensional filters is a three-dimensional frequency It is tuned to TIFF2026516024000003.tif6150. TIFF2026516024000004.tif6150 and TIFF2026516024000005.tif6150 represents a spatial pattern, TIFF2026516024000006.tif6150 represents a pattern in time, and each of the three-dimensional filters in the first subset has a different scale.
[0011] According to some embodiments, at least one processor further, Determine the z-score for each of the multiple filter bank responses; Each z-score is mapped to a weight associated with the respective filter bank for which the z-score was determined, thereby generating multiple weighted filter bank responses; Using multiple weighted filter bank responses, a computer vision process is executed based on these multiple filter bank responses. It is programmed that way.
[0012] According to some embodiments, at least one processor, Relationship Using TIFF2026516024000007.tif12150, the variance value Estimated TIFF2026516024000008.tif6150, here, TIFF2026516024000009.tif6150 contains multiple frames Filter TIFF2026516024000010.tif6150 This is a filter bank response generated by applying TIFF2026516024000011.tif6150. TIFF2026516024000012.tif6150 is, This is the variance of the estimated local mean flux over TIFF2026516024000013.tif6150. and TIFF2026516024000014.tif12150 is, Filter across TIFF2026516024000015.tif6150 This is the sum of TIFF2026516024000016.tif6150. and relational expressions The Z-score is determined using TIFF2026516024000017.tif16150. It is further programmed to do so.
[0013] According to some embodiments, at least one processor further: maps each z-score to the weight associated with each filter bank determined using the relational expression TIFF2026516024000018.tif8150, and is programmed such that where TIFF2026516024000019.tif6150 includes a threshold z-score.
[0014] According to some embodiments, the computer vision process is an edge detection process, and at least one processor further: executes the computer vision process based on the phase consistency associated with each of a plurality of filter bank responses, and is programmed as such.
[0015] According to some embodiments, at least one processor is further programmed to detect one or more corners based on the phase consistency associated with each of a plurality of filter bank responses.
[0016] According to some embodiments, the computer vision process is a motion estimation process, and at least one processor further: Based on the phase invariant relationship TIFF2026516024000020.tif6150, execute the computer vision process for each of a plurality of pixels, and is programmed such that, where TIFF2026516024000021.tif6150 is the velocity at each pixel, TIFF2026516024000022.tif6150 is each three-dimensional filter TIFF2026516024000023.tif6150's spatial direction The component velocity at TIFF2026516024000024.tif6150, and TIFF2026516024000025.tif6150, where TIFF2026516024000026.tif6150 is the spatio - temporal direction of TIFF2026516024000027.tif6150.
[0017] According to some embodiments of the disclosed subject matter, a method for facilitating single - photon computer vision tasks is provided. The method includes causing an image sensor to generate a sequence of images representing a scene, where each image includes a plurality of pixels and the image sensor includes a plurality of detectors configured to detect the arrival of individual photons, and the plurality of detectors are arranged in an array; for each of a plurality of three - dimensional filters, performing a convolution between the three - dimensional filter and a plurality of frames, where each of the plurality of frames is based on one or more of the images of the image sequence; for each of the plurality of frames, generating a plurality of filter bank responses corresponding to each of the three - dimensional filters among the plurality of three - dimensional filters; and performing a computer vision process based on the plurality of filter bank responses.
[0018] According to some embodiments of the disclosed subject matter, a non-temporary computer-readable medium is provided which, when executed by a processor, causes a processor to perform a method for facilitating a single-photon computer vision task, the method comprising: causing an image sensor to generate a sequence of images representing a scene, each image comprising a plurality of pixels, the image sensor comprising a plurality of detectors configured to detect the arrival of individual photons, the plurality of detectors arranged in an array; performing a convolution between a plurality of three-dimensional filters and a plurality of frames for each of a plurality of three-dimensional filters, each of the plurality of frames being based on one or more images of the sequence of images; generating a plurality of filter bank responses for each of the plurality of frames, each corresponding to a three-dimensional filter among the plurality of three-dimensional filters; and performing a computer vision process based on the plurality of filter bank responses.
[0019] This patent or application file includes at least one drawing performed in color. A copy of this patent or patent application publication containing the color drawing will be provided by the Japan Patent Office upon request and payment of the prescribed fee.
[0020] The various purposes, features, and advantages of the disclosed subject matter can be more fully understood by referring to the following detailed description of the disclosed subject matter, when considered in conjunction with the following drawings, where the same reference numerals indicate the same elements. [Brief explanation of the drawing]
[0021] [Figure 1A] This shows an example of a computer vision result generated from a scene image obtained by averaging a series of binary frames.
[0022] [Figure 1B] This shows an example of computer vision results generated from images of the same scene obtained using quantum burst imaging technology.
[0023] [Figure 1C] An example of the flow of Euler-type single-photon computer vision is shown, based on several embodiments of the disclosed subject matter.
[0024] [Figure 2] Examples of systems for Euler-type single-photon computer vision are shown, based on several embodiments of the disclosed subject matter.
[0025] [Figure 3] An example of a two-dimensional velocity tuning filter corresponding to a two-dimensional signal and velocity is shown.
[0026] [Figure 4] Examples of weight functions that can be used to determine the z-score for a response are shown, according to several embodiments of the disclosed subject matter.
[0027] [Figure 5] Examples of processes for Euler-type single-photon vision are shown, based on several embodiments of the disclosed subject matter.
[0028] [Figure 6] Examples of binary frames from two scenes and edge detection results generated using various techniques, including those described herein for Euler-type single-photon computer vision, are shown.
[0029] [Figure 7] Examples of edge detection results generated using the techniques described herein for Euler single-photon computer vision, for various levels of frame averaging and various levels of scene flux, are shown according to several embodiments of the disclosed subject matter.
[0030] [Figure 8]Examples of edge detection results generated using the techniques described herein for Euler-type single-photon computer vision are shown, using binary frames of scenes at two flux levels and various numbers and coarseness of scaling in the filter, according to several embodiments of the disclosed subject matter.
[0031] [Figure 9] Examples of motion estimates generated using various techniques, including those described herein for binary frames of scenes and Euler-type single-photon computer vision, are shown.
[0032] [Figure 10] Examples of motion estimates generated using the techniques described herein for Euler-type single-photon computer vision are shown, using binary frames of two scenes and various coarsenesses of scaling in the filter, according to several embodiments of the disclosed subject matter. [Modes for carrying out the invention]
[0033] According to various embodiments, mechanisms (which may include, for example, systems, methods, and media) for Euler-type single-photon vision are provided.
[0034] Single-photon sensors such as SPADs can measure optical signals with the finest possible resolution, at the level of individual photons. However, such sensors present two major challenges: strong Poisson noise and extremely high data acquisition rates. Previous research has primarily focused on image reconstruction problems, followed by the use of off-the-shelf techniques for downstream tasks. However, the most common solutions that account for motion are typically computationally expensive and not scalable to the large volumes of data generated by single-photon sensors.
[0035] According to several embodiments, the mechanisms described herein can facilitate the execution of computer vision tasks using data from a single-photon imaging device without performing explicit image reconstruction from the data. For example, as described herein, computationally lightweight phase-based techniques for computer vision tasks (e.g., edge detection and motion estimation) can be used to execute computer vision tasks that directly process raw single-photon data as a 3D volume (e.g., using velocity tuning filtering) and apply a 3D convolution kernel to the input photon stream. Experiments have been conducted demonstrating the results of using the techniques described herein for both edge detection and motion estimation tasks, as will be discussed later in relation to Figures 6-10, achieving speedups of more than two orders of magnitude compared to explicit reconstruction-based techniques.
[0036] Generally, digital image sensors record light for discrete sensing elements (often called pixels). The spatiotemporal density of these measurements continues to increase over time, and recent developments have led to single-photon quantum sensors such as single-photon avalanche diodes (SPADs) and JOTs (e.g., Fossum et al., "The Quanta Image Sensor: Every Photon Counts," Sensors, 16 1260 (2016)). Such single-photon sensors can be configured to record measurements at the granularity of individual photons, which can facilitate an array of exciting applications such as photography under challenging conditions (e.g., low light, fast motion, and / or high dynamic range), high-speed tracking, and 3D imaging.
[0037] Such single-photon sensors open up new possibilities by providing access to the arrival time of individual photons. The challenges presented by such sensors include the amount of raw data captured by these sensors (e.g., leading to difficulties in moving the data from image sensors), the high quantization of such raw data (e.g., down to 1 bit per pixel), and the generally high noise content of such raw data (e.g., due to the Poisson statistics of photons). Furthermore, the computational (and therefore power) cost of analyzing such raw data is generally high, and the amount of storage, computation, and communication costs associated with capturing and using such data increases because individual photons are processed independently instead of being aggregated (as in conventional sensors). These challenges hinder the large-scale practical application of this exciting technology.
[0038] According to several embodiments, the mechanisms described herein can be used to implement relatively light-weight (using relatively few computational resources such as memory, processing resources, and communication resources) computer vision techniques for single-photon imaging devices that capture binary single-photon frames (and / or multi-bit single-photon frames) at relatively high speeds. The most widely studied problem in single-photon imaging has been image reconstruction, under the assumption that recovering high-quality images from single-photon data is important for downstream inference. However, due to strong noise and heavy quantization, image reconstruction from binary frames (especially from single-binary frames) is a challenging problem and often requires powerful pre-configuration and computationally intensive techniques for image reconstruction.
[0039] Mechanisms described herein in several embodiments Figure 1A shows an example of computer vision results generated from a scene image obtained by averaging a series of binary frames.
[0040] As shown in Figure 1A, when the radiance of a scene changes over time, a SPAD array (or other suitable single-photon image sensor) can capture a high-speed sequence of binary frames. As shown in Figure 1A, a single frame is extremely noisy and quantized, and naively averaging the frames over time increases the signal but loses motion information.
[0041] An intuitive method to reduce data noise and quantization is to aggregate information across many frames. However, this method is prone to potentially serious motion blur. For example, as shown in Figure 1A, simply averaging binary frames results in a falling ball being completely blurred.
[0042] Figure 1B shows an example of computer vision results generated from images of the same scene obtained using quantum burst imaging techniques. The technique shown in Figure 1B can be described as a Lagrangian vision pipeline based on frame-by-frame reconstruction as an intermediate step. Quantum burst imaging techniques are described in U.S. Patent No. 11,170,549 by Ma et al., which is incorporated herein by reference in its entirety. As shown in Figure 1B, quantum burst imaging may include an alignment step in which motion is estimated at the patch level, followed by a robust merging (summing) of frames after compensating for the estimated motion (robustness can mitigate inaccurate motion estimation under shot noise). While high-quality frames can be obtained using quantum burst imaging techniques, these techniques generally incur significant computational and memory / bandwidth costs.
[0043] Figure 1C shows an example of the flow of Euler-type single-photon computer vision according to several embodiments of the disclosed subject matter.
[0044] As shown in Figure 1C, in some embodiments, an Euler-type vision pipeline can be implemented using the mechanism described herein. For example, a recorded photon stream can be processed in a single shot using a bank of velocity-tuned three-dimensional filters, and then local pixel-level calculations can be performed to extract low-level information such as edges and motion vectors. Such a single-pass technique can result in less computation and less data movement overall than the pipeline shown in Figure 1B for a fixed output frame rate.
[0045] Many computer vision tasks do not require a complete image of the scene being analyzed, and such computer vision tasks are therefore not necessarily bound by the same cost-to-quality trade-off as image reconstruction. In some embodiments, the mechanisms described herein can be used to perform signal phase reconstruction as a substitute to reconstruct information that can be reconstructed from image reconstruction and can be addressed without reconstructing the entire signal (image). Phase is a key feature in both visual perception and visual tasks. In the context of video, local phase from a directional 3D filter can directly encode information about the motion of the scene (as described later, for example, in relation to Figure 5). According to some embodiments, phase information can be extracted from single-photon sensor data using a family of 3D velocity tuning filter banks. Several computer vision techniques that can be used to extract scene information (e.g., edges, motion) from the reconstructed phase information are described herein. Such phase-based techniques may involve only linear filtering and pixel-level calculations, resulting in extremely fast execution compared to image reconstruction-based methods. As will be discussed later in relation to Figures 6 to 10, the results using the implementation of the mechanism described herein show a computational speedup of more than two orders of magnitude compared to explicit burst vision methods of comparable quality (see, for example, Figures 6 to 10).
[0046] In some embodiments, the significant difference in speed between explicit burst vision techniques (and other image reconstruction-based techniques) and the mechanisms described herein can be attributed to the different perspectives taken by the techniques. For example, burst reconstruction can be considered a form of search: given a patch, the task is to find similar patches across other frames of a video. Searching long sequences becomes degraded and costly if the search is repeated patch by patch. The general idea of tracking the trajectory of patches through an exposure volume can be likened to the Lagrangian specification in fluid dynamics, which describes the motion of individual particles in a flow field. In contrast, the mechanisms described herein can be likened to the Eulerian method in fluid dynamics, where the flow properties (such as velocity) are described at each point in space and time without the concept of particles.
[0047] As single-photon sensors become more widely used and specialized processor architectures for such sensors are developed, the simplicity of Euler-type methods can make them attractive candidates for on-chip implementation, which can be advantageous in practical single-photon imaging by reducing the cost of data transfer (for example, transmission of raw photon data from the image sensor chip can be omitted if the data can be analyzed on the image sensor chip). In some embodiments, the Euler-type single-photon vision techniques described herein can be used to provide a general strategy for designing lightweight algorithms for extremely fast vision tasks directly from raw single-photon data.
[0048] Figure 1C shows a SPAD array observing a scene capturing a sequence of frames (e.g., binary frames) over time. The average incident flux at each pixel is: It can be represented as TIFF2026516024000028.tif6150 (photons / second), where, TIFF2026516024000029.tif6150 represents the spatial position of pixels. TIFF2026516024000030.tif6150 and time frame index This represents TIFF2026516024000031.tif6150. The number of incident photons is average. It can be modeled as a Poisson random variable having TIFF2026516024000032.tif6150. In some embodiments, the number of photons detected by a pixel during each frame exposure can be limited to a maximum of one. Note that some SPAD arrays can capture a single frame in which multiple photon detections are permitted, and this can be used to generate a multi-bit frame (e.g., the number of detections within the frame capture time is recorded up to the top). In addition or alternatively, data from multiple frames (e.g., multiple binary frames, multiple multi-bit frames) can be combined (e.g., by the sum of the number of photon detections). Thus, pixel measurements TIFF2026516024000033.tif6150 can be binarized and follow a Bernoulli distribution, with the following relationship: TIFF2026516024000034.tif16150 TIFF2026516024000035.tif6150(1) It can be expressed using, where the exposure time for each frame is TIFF2026516024000036.tif6150 seconds, TIFF2026516024000037.tif6150 is a single-photon detector quantum efficiency, and TIFF2026516024000038.tif6150 is a dark count rate (DCR) representing spurious detection unrelated to the incident photon. Separate quantum sample TIFF2026516024000039.tif6150 and The files TIFF2026516024000040.tif6150 can be assumed to be statistically independent of each other.
[0049] In some embodiments, the mechanisms described herein can be used to extract information from photon cube data captured by an array of single-photon detectors (e.g., a SPAD array). Because individual frames are extremely noisy and quantized (binarized), information generally must be aggregated across a sequence of multiple single-photon frames. However, simply summing the frames in time can result in significant motion blur (depending on the scene composition, etc.), making it difficult to extract meaningful scene information from the photon cube. As mentioned above, it is possible to explicitly compensate for motion through techniques such as search-based burst photography to reconstruct a high-quality image from the photon cube, but such techniques are computationally and bandwidth-intensive and are not suitable for real-time processing with current technology.
[0050] In some embodiments, the mechanism described herein can directly extract scene information from a photon cube without an intermediate step of image reconstruction. In some embodiments, the mechanism described herein can be based on the analysis of motion as the spatiotemporal orientation of intensity or phase isosurfaces when viewing video as a 3D volume. Motion information can be extracted through a 3D orientation filter used. Such a filter may be called a velocity tuning filter, because a filter in a given direction in the 3D frequency domain responds only to motion at a specific range of velocities. In Figure 1C, the hues shown in the depiction of the velocity tuning filter and filter bank response represent the phase with complex values (e.g., two values per pixel). As illustrated, the phase of the filter and response can be varied based on the wavelet shape of the impulse response function.
[0051] In some embodiments, the advantage of using a velocity tuning filter in single-photon video is that, compared to frame-by-frame processing, the 3D filter can aggregate scene information (including fine details and motion) across a large spatiotemporal support, leading to a significant reduction in noise. Although reconstructing the entire flux signal from the filter response is still difficult, the phase information is well preserved and can therefore be used directly in downstream algorithms despite the strong noise and quantization of the raw photon cube. Note that, for a single-tone sine wave, the phase of the Fourier coefficients can be proven to be unbiased for almost all frequencies under the imaging model described above in relation to EQ(1). Furthermore, the simulation described in Appendix A shows that the variance of the Fourier coefficient phase is close to the Cramer-Rao lower bound on the unbiased estimator. Both the proof and the simulation results are included in Appendix A, which is incorporated herein by reference in its entirety. Although a velocity tuning filter is not a pure sine wave, because the filter resembles a sine wave, directly extracting the phase of such a filter can be expected to be close to the optimal (maximum-likelihood) estimate. For specific cases of low-level vision tasks, such as edge detection and motion estimation, which are described below in relation to Figures 5-10, velocity tuning filter banks (applied to single-photon data) were designed and analyzed.
[0052] Figure 2 shows an example of a system for Euler-type single-photon computer vision in several embodiments of the disclosed subject matter.
[0053] As shown in the figure, system 200 may include an image sensor 204 (e.g., an area sensor including a sensor array of single-photon detectors); optical elements 206 (e.g., one or more lenses, one or more attenuation elements such as filters, an aperture, and / or any other suitable optical elements such as a beam splitter); any suitable hardware processor (which may be a central processing unit (CPU), digital signal processor (DSP), microcontroller (MCU), graphics processing unit (GPU), accelerator processing unit (APU), etc.) or a combination of hardware processors; a processor 208 that can be configured to control the operation of system 200; input devices 210 (shutter button, menu button, microphone, touchscreen, motion); The system may include: a sensor, liquid crystal display, light-emitting diode display, etc., or any suitable combination thereof; memory 212; a signal generator 214 which can be configured to generate one or more signals for controlling the operation of the image sensor 204; one or more communication systems 216 which can be configured to facilitate communication between the system 200 and other devices such as smartphones, wearable computers, tablet computers, laptop computers, personal computers, servers, embedded computers (e.g., for controlling autonomous vehicles, robots, etc.) via a communication link; and / or a display 218 which can be configured to present information for user consumption (e.g., images, user interfaces, etc.). In some embodiments, the memory 212 may store image data and / or any other suitable data. The memory 212 may include a storage device (e.g., a hard disk, Blu-ray disc, digital video disc, random access memory (RAM), read-only memory (ROM), electronically erasable read-only memory (EEPROM), etc.) for storing computer programs for controlling the processor 208.In some embodiments, the memory 212 may contain instructions for causing the processor 208 to perform operations related to the mechanism described herein, such as operations described later in relation to Figure 5.
[0054] In some embodiments, the image sensor 204 may be an image sensor implemented using at least partially a sensor array of SPAD detectors (sometimes called Geiger-mode avalanche diodes) configured to detect the arrival time of individual photons and / or one or more other detectors. In some embodiments, one or more elements of the image sensor 204 may be configured to generate data indicating the arrival time of photons from the scene via an optical element 206. For example, in some embodiments, the image sensor 204 may be an array of multiple SPAD detectors. In yet another example, the image sensor 204 may be a hybrid array including SPAD detectors and one or more conventional photodetectors (e.g., CMOS-based pixels). In yet another example, the image sensor 204 may be multiple image sensors, such as a first image sensor including a sensor array of SPAD detectors that can be used to generate information about the brightness of the scene, and a second image sensor including one or more conventional pixels that can be used to generate information about the color of the scene. In such examples, the optical element 206 (e.g., multiple lenses, beam splitter, etc.) may include an optical system for directing an incident light orientation sensor towards a SPAD-based image sensor and another orientation sensor towards a conventional image sensor. In some embodiments, the image sensor 204 may have an imaging plane from which the optical element 206 can focus light from the scene.
[0055] In some embodiments, system 200 may include additional optical elements. For example, although optical lens 206 is shown as a single lens, it can be implemented as a composite lens or a combination of lenses. While the mechanisms described herein are generally described as using SPAD-based detectors, it should be noted that this is merely one example of a single-photon detector. As mentioned above, other single-photon detectors, such as jet-based image sensors, can also be used.
[0056] In some embodiments, the signal generator 214 may be one or more signal generators capable of generating signals to the control image sensor 204. For example, in some embodiments, the signal generator 214 may supply signals to enable and / or disable one or more pixels of the image sensor 204 (e.g., by controlling the gate signals of the SPADs used to implement the pixels). As another example, the signal generator 214 may supply signals to control the reading of image signals from the image sensor 204 (e.g., to memory 212, to processor 208, to cache memory associated with the image sensor 204, etc.).
[0057] In some embodiments, system 200 can communicate with remote devices over a network using a communication system 216 and a communication link. In addition or alternatively, system 200 can be included as part of another device such as a smartphone, tablet computer, laptop computer, autonomous vehicle, or robot. Parts of system 200 can be shared with the device into which system 200 is integrated. For example, if system 200 is integrated with an autonomous vehicle, processor 208 can be the processor of the autonomous vehicle and can be used to control the operation of system 200.
[0058] In some embodiments, the system 200 can communicate with any other suitable device, which may be either a general-purpose device such as a computer or a special-purpose device such as a client or server. Either of these general-purpose or special-purpose devices may include any suitable component such as a hardware processor (which may be a microprocessor, digital signal processor, controller, etc.), memory, communication interface, display controller, or input device. For example, the other device may be implemented as a digital camera, security camera, outdoor surveillance system, smartphone, wearable computer, tablet computer, personal data assistant (PDA), personal computer, laptop computer, multimedia terminal, game console, peripheral for a game counselor (or any of the above devices), or special-purpose device.
[0059] Communication by the communication system 216 over a communication link can be carried out using any suitable computer network, or any suitable combination of networks, including the Internet, intranet, wide area network (WAN), local area network (LAN), wireless network, digital subscriber line (DSL) network, frame relay network, asynchronous transfer mode (ATM) network, and virtual private network (VPN). The communication link may include any communication link suitable for communicating data between system 200 and another device, such as a network link, dial-up link, wireless link, hardwired link, any other suitable communication link, or any suitable combination of such links.
[0060] In some embodiments, the display 218 can be used to present images and / or videos generated by the system 200, to present a user interface, etc. In some embodiments, the display 218 can be implemented using any suitable device or combination of devices, and may include one or more inputs such as a touchscreen.
[0061] It should also be noted that data received via a communication link or other communication links can be received from any suitable source. In some embodiments, the processor 208, for example, a transmitter, receiver, transceiver, or any other suitable communication device can be used to send and receive data via a communication link or any other communication link.
[0062] Figure 3 shows an example of a two-dimensional velocity tuning filter corresponding to a two-dimensional signal and velocity.
[0063] As shown in Figure 3, the principle of speed tuning is two speeds: panel (a) TIFF2026516024000041.tif6150 and panel (b) This can be shown using a 1D box-shaped signal captured over time (vertically) at TIFF2026516024000042.tif6150 pixels / frame. In both cases, the 2D(xt) spectrum is: It follows the line given in TIFF2026516024000043.tif6150. EQ(3) described below extends this to a moving two-dimensional signal (video) in the three-dimensional frequency domain.
[0064] Figure 4 shows examples of weight functions that can be used to determine the z-score for a response according to several embodiments of the disclosed subject matter.
[0065] Quantum sample Three-dimensional complex linear filter applied to the input video stream of TIFF2026516024000044.tif6150 Considering TIFF2026516024000045.tif6150: The filename becomes TIFF2026516024000046.tif12150, and here TIFF2026516024000047.tif6150 is a bandpass filter and has three-dimensional frequency response. The tuning is centered around TIFF2026516024000048.tif6150. TIFF2026516024000049.tif6150 represents the spatiotemporal support of h, TIFF2026516024000050.tif6150 represents the location and time (e.g., frame) of a possible data point instant. TIFF2026516024000051.tif6150 is a quantum sample This is the filtered bank response generated by applying a filter to samples from the stream of TIFF2026516024000052.tif6150, where In TIFF2026516024000053.tif6150, k represents one of the filters in the filter bank, tuned to a specific frequency (which can also be expressed using k). Note that h can represent the size of the filter and can be any appropriate range. For example, h can range from 2 pixels to 100 pixels in space and from 2 frames to 5000 frames in time. Another example is h can range from 3 pixels to 75 pixels in space and from 3 frames to 2500 frames in time. Yet another example is h can range from 3 pixels to 50 pixels in space and from 3 frames to 2000 frames in time. Yet another example is h can range from approximately 3 to 5 pixels in space and from approximately 3 to 5 frames in time. As yet another example, h can range from approximately 30 to 50 pixels in space and from approximately 1000 to 3000 frames in time. In relation to the video, the 3D filter can be interpreted as being speed-tuned, as mentioned above. TIFF2026516024000054.tif6150 is a unit vector Speed according to TIFF2026516024000055.tif16150 Frequencies moving in TIFF2026516024000056.tif17150 It can respond to the spatial pattern of TIFF2026516024000057.tif6150 to the fullest extent.
[0066] Figure 3 shows a simplified 2D example that illustrates this relationship through an example of a moving one-dimensional signal (2D overall).
[0067] For computational efficiency, in some embodiments, spatiotemporal separation is possible, such as with logarithmic Gabor filters (e.g., Field, "Rationsations between the statistics of natural images and the response properties of cortical cells," International Journal of Computer Vision, 5(1):77-104 (1990), and Kovesi, "Image Features from Phase Congruency," Videre: Journal of Computer Vision Research, 1(3) (1999)).
[0068] For example, in some embodiments, the spatial filters used in connection with the mechanisms described herein can be polarly separable in the frequency domain (e.g., similar in shape to a steerable filter bank). Such filters can be tuned with equally spaced orientations (e.g., in the range of 2 to 12 orientations; in more specific examples, the filters can be equally spaced with 6 different orientations, 2 orientations, 3 orientations, 4 orientations, 5 orientations, 7 orientations, 8 orientations, 9 orientations, 10 orientations, 11 orientations, 12 orientations, etc.) and constructed with multiple scales (e.g., in the range of 2 to 6 scales; in more specific examples, 2 scales, 3 scales, 4 scales, 5 scales, 6 scales, etc.). Note that the scales and / or the number of scales can differ along the spatial dimension (e.g., along x,y) and along the temporal dimension (e.g., t). These relationships can be treated separately / independently, as they can be changed by the video content through a speed tuning formula. For example, filter design can be an appropriate combination of either spatial or temporal scales (e.g., a coarse spatial scale and a coarse temporal scale, a coarse spatial scale and a fine / fast temporal scale, a fine spatial scale and a fine / fast temporal scale, a fine spatial scale and a coarse temporal scale, etc.).
[0069] In some embodiments, the radial bandwidth of the filter can be any suitable range, such as approximately 1 to 3 octaves. In a particular example, the radial bandwidth of the filter can be 2 octaves. Note that the choice of bandwidth is unlikely to make a substantial difference to the system operation and can be a relatively minor implementation detail. In some embodiments, the time filter can be tuned separately for each scale and through EQ(3) to a predetermined set of speeds Center frequency for TIFF2026516024000058.tif7150 The file TIFF2026516024000059.tif6150 can be obtained. In some embodiments, any number of appropriate speeds can be used. For example, in the experiment described below, three speeds were generally used.
[0070] In some embodiments, the filter response can be used without subsampling the filter response at a coarse scale (unlike, for example, the pyramidal representation). This simplifies the implementation of the algorithm because there is no interpolation required to return to the native sensor resolution, but at the cost of higher memory usage. For example, the overcompleteness coefficient is It can be calculated as TIFF2026516024000060.tif6150. By using a more memory-efficient representation, further cost reductions can be expected in the future.
[0071] Since individual SPAD samples are binary and noisy, a filter bank is desirable that can extract relevant details while rejecting spurious responses (which tend to dominate the data) as much as possible. In some embodiments, robust estimation of noise and / or uncertainty in the filter response facilitates more accurate removal of spurious responses.
[0072] From the central limit theorem, the response of EQ.(2) TIFF2026516024000061.tif6150 follows an approximately (complex) normal distribution, with variance: It can be expected that TIFF2026516024000062.tif13150 exists. This dispersion is due to quantum efficiency TIFF2026516024000063.tif6150 and dark count This can be approximated by assuming that TIFF2026516024000064.tif6150 is an ideal sensor. Then, from EQ.1, The filename becomes TIFF2026516024000065.tif6150, and here, The filename is TIFF2026516024000066.tif6150.
[0073] Estimate of local average flux From TIFF2026516024000067.tif6150 ( (Through the blur kernel on TIFF2026516024000068.tif6150), EQ.(4) is expressed as follows: It can be further approximated using TIFF2026516024000069.tif12150. TIFF2026516024000070.tif11150 is known. In some embodiments, at runtime, TIFF2026516024000071.tif6150 is a z-score. It can be converted to TIFF2026516024000072.tif16150. Additionally or alternatively, in some embodiments, the z-score is, for example, as shown in Figure 4. Weights are assigned to TIFF2026516024000073.tif8150. It can be mapped to TIFF2026516024000074.tif6150. In some embodiments, the parameter The z-score can be set before time (e.g., within a range including 2 to 6). This can prevent weak responses from contributing. In some embodiments, the z-score can be used as a secondary input to a computer vision task as an indicator of response reliability (e.g., it can be provided to the input without modification). It should be noted that downstream computer vision tasks do not necessarily need to directly scale responses based on the z-score (although this is one possibility). For example, motion estimation techniques can ignore the magnitude of the raw response and rely on the z-score and response phase.
[0074] Due to Gabor's uncertainty principle, smaller bandwidths accommodate greater spatiotemporal support (e.g., for coarser scales or elongated filters with low angular sensitivity). Typically, such filters exhibit a classic trade-off: smaller EQ.4 variance but worse localization. Ultimately, the filter's performance depends on the true range of the signal structure (e.g., edges).
[0075] A common fact in single-photon vision is that noise levels vary with light intensity, meaning that a filter that is reliable in strong light may become unreliable in weak light. In some embodiments, filter designs and downstream algorithms can be configured to adapt to this variation, for example, through the use of multiscale filter banks. The use of z-scores can further facilitate techniques for adapting to noise variations (e.g., due to ambient light intensity).
[0076] Figure 5 shows an example of a process for Euler-type single-photon vision in several embodiments of the disclosed subject matter.
[0077] In step 502, process 500 can capture a sequence of binary frames of the scene using any suitable image sensor. For example, as described above in relation to Figures 1C and 2, the image sensor can be a SPAD-based image sensor or a jet-based image sensor. However, these are merely examples, and the mechanisms described herein can be used to facilitate computer vision tasks using any sensor including a single-photon detector.
[0078] In some embodiments, process 500 can be made to capture a sequence of frames at any appropriate frame rate and / or within any appropriate time budget. For example, process 500 can be made to capture a sequence of frames at a high frame rate in situations where the motion and / or intensity of the scene are likely to be high. In a more specific example, the frame rate can be set between approximately 300 frames per second (fps) and approximately 100,000 fps for current SPAD-based image sensors. In another more specific example, the frame rate can be set between approximately 10 fps and approximately 1000 fps for current jet-based image sensors.
[0079] In some embodiments, the total time budget can range from about 1 millisecond to about 1 second. In certain examples, the total time budget can range from about 10 milliseconds (ms) to about 100 ms for scenes with a relatively high dynamic range. In some embodiments, the total time budget can be constrained based on the amount of movement within the scene, especially if objects move outside the scene during the time budget, as it becomes more difficult to produce high-quality images for scenes with more movement with longer time budgets and / or more binary frames. Furthermore, in some embodiments, the total time budget can be constrained based on the amount of available memory, as longer time budgets and / or more binary frames require the availability of additional memory that can be written at a rate comparable to the image sensor's frame rate.
[0080] In some embodiments, the total time budget can be omitted, a stream of binary frames can be captured, and after frames have already been captured, a sequence of binary frames corresponding to a specific period can be selected. For example, process 500 can continuously capture binary frames of a scene, and at any appropriate time, a sequence of frames can be selected from the continuously captured sequence for use in a computer vision task. In another example, process 500 can continuously capture binary frames of a scene, and once frames are captured, process 500 can continuously analyze the most recent frames (for example, as described later in relation to 506 and / or 508), and information from the oldest frames can be omitted from use in a computer vision task (e.g., deleted, replaced, no longer considered, flagged for overwriting, etc.).
[0081] In some embodiments, in process 502, process 500 can capture a series of multi-bit frames of the scene using any suitable image sensor. For example, the image sensor can be configured to record until any suitable number of photons arrive in a frame (for example, recording until one photon arrives can be used to generate a 2-bit frame, recording until two or three photons arrives can be used to generate a 2-bit frame, recording until seven photons arrives can be used to generate a 3-bit frame, etc.).
[0082] In 504, process 500 can create one or more multibit frames from a series of binary frames. In addition or alternatively, in some embodiments, process 500 can create one or more long multibit frames from a series of short multibit frames.
[0083] In some embodiments, process 500 can determine the number of binary (or multibit) frames to use to create a multibit frame using any suitable criterion or combination of criteria. For example, process 500 can combine frames targeting that the amount of motion in each multibit frame is no more than 1 pixel per frame, thereby reducing blurring in the combined frame. In some embodiments, process 500 can determine the number of frames to combine to correspond to the amount of motion in the scene and / or approximately 1 pixel of motion per frame, using any suitable technique or combination of techniques. For example, process 500 can evaluate the data in the frequency domain (e.g., based on the Fourier transform), and blurring can be revealed by the absence of high-frequency information in the Fourier domain. In some embodiments, 504 can be omitted (e.g., if the motion in the scene is less than 1 pixel per frame, or if each binary frame is processed as a binary frame).
[0084] In 506, process 500 can perform convolution on a binary frame (or multi-bit frame) with each of several filters in the filter bank. For example, as shown in Figure 1C, a convolution can be performed between each binary frame in a sequence of binary frames and each speed tuning filter. For example, this can be a 3D convolution similar to the convolution performed in a convolutional neural network. In some embodiments, any appropriate stride can be used for the convolution. For example, a stride of 1 can be used for the convolution. As another example, the stride of the convolution can be set with respect to the tuning frequency of the filter. In a more specific example, a coarser scale (e.g., low frequency) response can be obtained with a larger stride (e.g., a lower sampling rate) without loss of information. Note that the filters can be applied in the frequency domain (e.g., implicitly assuming periodic boundary conditions).
[0085] In 508, process 500 can generate a filter response for each of the filters based on the result of the convolution between the filters and the binary frame (or multi-bit frame) information. For example, each filter can generate a response corresponding to each pixel. In a more specific example, if there are N filters and the video contains X pixels per frame, the responses would be: The TIFF2026516024000076.tif6150 value may be included. As another example, each filter may generate less than one response corresponding to each pixel (e.g., stride is greater than 1 pixel and / or frame). In some embodiments, process 500 may generate filter responses corresponding to each frame in the sequence of frames captured in 502 and / or each multi-bit frame created in 504. In some embodiments, the filter responses may be feature maps based on the results of convolution between the filter and one or more binary frames.
[0086] In 510, process 500 may execute one or more suitable computer vision processes to analyze a series of binary frames based on the filter response generated in 508.
[0087] In some embodiments, in 510, process 500 can perform an edge detection computer vision process using a phase-based technique. In addition or alternatively, in some embodiments, in 510, process 500 can perform a motion estimation computer vision process using a phase-based technique. These algorithms can be described as Euler-type because no search is performed, only local information is used, and the majority of the computation is pixel-based, and therefore they are easily parallelizable. These techniques can process sequences of single-photon frames directly without expensive image or video reconstruction, and can improve the speed at which computer vision tasks can be performed on single-photon data.
[0088] In some embodiments, process 500 can perform an edge detection computer vision process, which can be based on temporal phase matching. Phase matching is the insightful observation that an edge-like feature is a discontinuity in which the phases of all frequency components in the signal coincide. Phase matching also applies to video, where a moving edge traces a 3D plane over time. In this case, a multiscale bank of speed-tuned filters plays the role of frequency components, and temporal phase matching (TPC) can be detected. The tuning frequency of one filter in the multiscale bank TIFF2026516024000077.tif6150 is, It can be represented as TIFF2026516024000078.tif6150, and here TIFF2026516024000079.tif6150 represents a scale, TIFF2026516024000080.tif6150 is a unit vector along the spacetime direction. The filename is TIFF2026516024000081.tif6150. TIFF2026516024000082.tif6150 and TIFF2026516024000083.tif6150 can be expressed in spherical coordinates. TIFF2026516024000084.tif6150 corresponds to the angle of the x,y plane deflected from x=0, TIFF2026516024000085.tif6150 corresponds to an angle in the time dimension deflected from x,y=0. A phase-aligned PC along this direction can be expressed as follows: TIFF2026516024000086.tif13150 This is 1 when the response has the same phase at all scales. In some embodiments, EQ.(6) can localize the features well when the response has the same phase at all scales, It can be adjusted to be TIFF2026516024000087.tif8150. For example, adjustments to EQ.(6) can be made based on the discussion in Kovesi's "Image Features from Phase Congruency" to better handle blurred features. Note that phase congruency is a normalized quantity and is invariant to amplitude scaling (such as that caused by light intensity). In EQ.(6) above, only phase information is implicitly used to avoid the phase wrapping problem.
[0089] In some embodiments, Once TIFF2026516024000088.tif6150 has been calculated for all directions, process 500 can estimate edge strengths using (three-dimensional) principal component analysis, which can be efficiently implemented using a closed-form representation of the eigenvalues of a 3x3 matrix (based, e.g., Kopp, "Efficient Numerical Diagonalization of Hermitian 3×3 matrices", International Journal of Modern Physics C, 19(03):523-548 (2008), and Smith, "Eigenvalues of a Symmetric 3×3 matrix", Communications of the ACM, 4(4):168 (1961)). In some embodiments, a second eigenvalue (if significant) can indicate a spatiotemporal "corner" in the 3D volume.
[0090] In some embodiments, under noisy conditions for single-photon sensing, the right-hand side of EQ.(6) can be multiplied by the weighting term described above in relation to Figure 4, thereby excluding orientations with weak responses (for example, TIFF2026516024000089.tif6150 was set to 2 in all of the edge detection experiments described below.
[0091] In some embodiments, process 500 can perform a motion estimation computer vision process to estimate edge normal velocities, which can be based on temporal phase consistency. Such information can be used to estimate normal velocities from 3D edge direction estimates. Since the principal direction obtained by temporal phase consistency is 3D (which is the normal to the plane traced by the edge moving over time), process 500 can also directly receive normal velocity estimates at the edge location, similar to optical flow obtained from an event camera. Further explanation of these estimates is included below in relation to Figure 9.
[0092] In some embodiments, process 500 can perform a motion estimation computer vision process that can be based on local frequency information. In some embodiments, a filter can be defined at spatiotemporal frequency k, and process 500 can perform a motion estimation computer vision process that can perform a motion estimation computer vision process that can perform a motion estimation computer vision process that can perform a motion estimation computer vision process that can perform a motion estimation computer vision process that can be based on local frequency information. The velocity in the k-direction given by TIFF2026516024000090.tif6150 can be estimated. In some embodiments, the instantaneous frequency is It can be represented as TIFF2026516024000091.tif6150, and the component velocity TIFF2026516024000092.tif6150 is spatial direction TIFF2026516024000093.tif6150 and The file is given as TIFF2026516024000094.tif6150. It represents the 2D velocity (optical flow) at the pixel. To obtain TIFF2026516024000095.tif6150, process 500 uses the phase invariant equation from the component estimates: TIFF2026516024000096.tif6150 can be formed. In some embodiments, process 500 can obtain one equation from each reliable filter response, which can then be put together and solved as a weighted least squares problem with the weights described above in relation to Figure 4 (threshold). TIFF2026516024000097.tif6150 was set to 6 for the optical flow experiment described later). Appendix A, incorporated herein by reference, contains additional implementation details. In some embodiments, such motion estimation can be applied independently at each scale. The role of the scales is described later in relation to Figure 8.
[0093] In some embodiments, process 500 can perform steps 502-510 on any suitable block of frames. For example, process 500 can divide the sequence of binary frames captured in step 502 into any suitable number of blocks. In some embodiments, the sequence of binary images can be divided into blocks of a specific size (e.g., blocks of 50-10,000 frames for frame rates up to 100,000 fps, blocks corresponding to approximately 10 milliseconds of total exposure time, etc.). In some embodiments, a block may contain at least a minimum number of binary frames to ensure that sufficient information is included in the filtered response when convolved with the filter in step 506. For example, in some embodiments, each block may contain at least 20 binary frames. As another example, in a range of specific light intensities (e.g., approximately 1 photon / pixel), each block may contain a single binary frame.
[0094] In some embodiments, the process 500 may introduce some delay, for example, between the time when the first binary frame in the block of binary frames being analyzed is captured and the time when the last binary frame in the block of binary frames being analyzed is captured.
[0095] In some embodiments, process 500 can return to 502 and begin capturing additional binary frames of the scene after executing the computer vision process in 510, or in parallel with the execution of 504-510. In such embodiments, process 500 can analyze discrete blocks of frames (for example, a first block of frames can be analyzed starting from a first time, and a second block of frames that does not contain any of the frames included in the first block of frames can be analyzed starting from a later second time).
[0096] In addition or alternatively, in some embodiments, process 500 may move to 512 after executing (or starting to execute) the computer vision process in 510, and capture additional binary frames (or multiple frames) of the scene.
[0097] In 514, process 500 can create one or more additional multibit frames from a series of binary frames. In addition or alternatively, in some embodiments, process 500 can create one or more long multibit frames from a series of short multibit frames. In some embodiments, 514 can be omitted (for example, when the motion in the scene is less than one pixel per frame, or when all binary frames are processed as binary frames).
[0098] In 516, process 500 can perform a convolution of the additional binary frame (or multibit frame) captured in 512 to determine the contribution of that frame to the filter response generated in 508 (or in the iteration prior to 518) for each of the multiple filters in the filter bank.
[0099] In 518, process 500 can generate one or more new filter responses corresponding to the additional frame, based on the convolution performed in 516, and possibly in part on the convolution performed in 506 (or a previous iteration of 516). For example, the result of the filter convolution with previous frames in a series of frames can be added to the result of the convolution performed in 516 to generate a filter response for the current frame.
[0100] In 520, process 500 may remove and / or ignore the oldest filter response (e.g., from memory) and / or the oldest unremoved contribution to the oldest filter response (e.g., by removing the contribution to the oldest filter response from the oldest frame).
[0101] At 522, process 500 may execute any suitable computer vision process or process to analyze a series of binary frames based on the filter responses generated / updated at 518 and / or 520. In some embodiments, process 500 may return to 512 and capture additional binary frames of the scene after executing the computer vision process(es) at 522, or in parallel with the execution of 514-520. In such embodiments, process 500 may analyze the update stream of frames (for example, a first block of frames may be analyzed starting at a first time, and a second block of frames containing many of the same frames included in the first block of frames may be analyzed starting at a later second time).
[0102] Figures 6 and 8-10 demonstrate examples of results generated using the techniques described herein for actual binary frame sequences captured with a SwissSPAD sensor having a resolution of 256x512 and a maximum frame rate of 97,700 fps. Flux levels are reported as photons / pixel (abbreviated as ppp).
[0103] Sequences with slow motion (where the flow is << 1 pixel per frame, which is common at high frame rates) were subsampled temporally with a low-pass filter, which was roughly equivalent to sampling with a multi-bit sensor. A set of tuned speeds was adapted accordingly.
[0104] The SPAD prototype is a research device and contains several "hot pixels" with high dark count rates. These pixels were detected offline in the dark frame and interpolated.
[0105] Due to the extensive filter support, filtering was performed in the frequency domain. For fair comparison, the algorithms described herein and the techniques compared (e.g., BM3D published by Tampere Institute of Technology at https: / / webpages.tuni.fi / foi / GCF-BM3D / index.html, and burst reconstruction described in Ma et al., U.S. Patent No. 11,170,549) were implemented in MATLAB® and executed on a CPU.
[0106] Figure 6 shows binary frames of two scenes and examples of edge detection results generated using various techniques, including those described herein for Euler-type single-photon computer vision.
[0107] Figure 6 shows edge detection results for actual SPAD video. Edges from binary frames from SwissSPAD (top) and from the Euler-type time-phase matching algorithm (TPC, bottom) described in relation to Figure 5 were compared with various reconstruction-based methods (middle). Single-image denoising used BM3D, and Lagrangian vision used burst acquisition. Richer convolutional feature detectors were used for all reconstruction-based results. TPC restored sharp edges even under noise and motion at a fraction of the cost of burst reconstruction (and faster than BM3D). Direct detection and naive averaging from single frames are fast but each suffers from a noise-versus-motion blur trade-off. Single-image denoising avoids motion blur but has oversmoothing artifacts that cause loss of many edges.
[0108] The results shown in Figure 6 illustrate edges obtained by TPC from a low-pass filtered sequence of 120 frames of video captured by a SwissSPAD sensor. For comparison, results from reconstruction-based methods, including the Lagrangian reconstruction-based method shown in Figure 1B, are also presented. For comparison, Richer Convolutional Features (RCF) were used as a representative frame-based edge detector.
[0109] The Euler-type method implemented according to the mechanism described herein achieved results of similar quality to the Lagrangian technique, but at a speed more than two orders of magnitude faster (for example, the MATLAB implementation performed the analysis in 0.145 seconds, compared to 153 seconds for the Lagrangian MATLAB implementation. Furthermore, in the same hardware and software environment, the Euler method was an order of magnitude faster than the tested BM3D implementation).
[0110] Figure 7 shows examples of edge detection results generated using the techniques described herein for Euler-type single-photon computer vision for various levels of frame averaging and various levels of scene flux, according to several embodiments of the disclosed subject matter.
[0111] Figure 7 shows the effect of flux on edge detection in a simulated scene. Figure 7 shows edge detection using temporal phase matching, simulating a SPAD video with 51 frames of size 128x128, with 1 pixel of motion per frame, and at various flux levels (photons per pixel, in ppp units) and accuracies. The edge detection quality depends on the total incident flux (its contour lines are roughly the diagonal of the dotted lines), and edges can be detected even from binary video if the flux is at least 1 ppp.
[0112] The performance of the edge detector depends on the amount of light and motion in the scene. In fact, in slow-moving scenes, simple low-pass filtering can yield more accurate data. Figure 7 shows the changes when simulating a synthesized scene for a TPC detector with a fixed filter bank. An ideal sensor was assumed to have no dark count and perfect quantum efficiency. Even under extremely harsh conditions (1-bit sample and motion), the TPC successfully reconstructed edges. Ultimately, reconstruction depends on the total number of incident photons, and reasonable quality reconstruction is possible even at flux levels comparable to ~1 photon-pixel.
[0113] Figure 8 shows examples of edge detection results generated by the Euler-type single-photon computer vision technique described herein, using binary frames of a scene at two flux levels and filters having varying scaling numbers and coarseness, according to several embodiments of the disclosed subject matter.
[0114] Figure 8 shows the results of edge detection under various filter configurations. A scene without rigid motion (a person juggling two footballs) was captured. Faces have been blurred for privacy. The top row shows the input frame after time-lowpass filtering, where edges were detected with time-phase consistency across two scale ranges. The inset shows a tone-mapped burst reconstruction for reference. The bottom row shows input frames from the same scene re-captured under weaker light. Fine-scale edge maps are significantly degraded, while coarse-scale maps retain quality. In the last column, the filter's angular bandwidth is reduced, resulting in more reliable reconstruction of long edges, but the detector overshoots for curved edges such as footballs, heads, and elbows.
[0115] As mentioned above in relation to EQS.(4) and (5), filter design is important and can affect the results. Figure 8 is an example somewhat similar to Figure 7, but using real data. In this example, the same controlled scene (a person juggling two footballs) was captured twice under different lighting conditions (medium light and low light). TPC was run on filter banks with two different scale ranges, one covering spatial wavelengths of 3 to 13 pixels ("fine scale") and the other covering 6 to 28 pixels ("coarse scale"). The filters had six spatial directional tunings. TIFF2026516024000098.tif6150) and three speed tunings ( The image was created using TIFF2026516024000099.tif (6150 pixels / frame). A fine-scale filter bank yielded sharper edges under more light but deteriorated in low light. In contrast, a coarser-scale filter gave thicker (lower resolution) edges but was more reliable in low light. The SNR can be further improved by narrowing the angular bandwidth of the filter, which is effective for long edges but prone to overshoot around curved edges, which is also a form of delocalization or resolution reduction. These results are consistent with the detection versus localization trade-off that has been studied for many years in the literature on edge detection.
[0116] Figure 9 shows an example of a binary frame of the scene and motion estimates generated using various techniques, including those described herein for Euler-type single-photon computer vision.
[0117] Figure 9 shows motion estimation from actual SPAD video. Figure 9 shows a low-light scene with a moving object (a toy train on a track) and optical flow estimation. The reference flow was estimated using RAFT-it. In the top row, the left side shows the binary frame after time low-pass filtering (LPF) (original inserted), and the right side shows the optical flow based on the two binary frames. The second row shows the reconstructed image and optical flow results generated from frame-by-frame denoising by BM3D. The third row shows a sample of the reconstructed image generated using Lagrangian / explicit burst vision and the optical flow results generated from the reconstructed image. The bottom row shows the normal velocity from the edges extracted by temporal phase matching (magnified to the upper right), and the 2D velocity from the implementation of the motion estimation computer vision process described above in relation to Figure 5. The train's motion is clearly separated, unlike direct or 2-frame estimation by BM3D. While the Lagrangian technique offers good quality and reliability, it is quite costly (in terms of time).
[0118] Figure 9 shows the results for an actual SwissSPAD sequence using both the edge normal velocity and 2D velocity estimation techniques described in relation to Figure 5. For comparison, the state-of-the-art two-frame technique RAFT-it was applied with and without image reconstruction. The results produced using the implementation of the mechanism described herein are significantly better than directly applying RAFT-it to noisy frames in that they reliably separate moving objects from a static background. Furthermore, the results produced using the implementation of the mechanism described herein performed considerably better than single-image denoising devices due to the temporal incoherence of denoising artifacts. While the Lagrangian technique yielded the best quality results, it was also significantly more expensive.
[0119] Phase-based 2D velocity estimation reveals that the movement of the train's projected headlights is also detected as motion, but this is ignored in RAFT-it, and only the train is segmented. This can be attributed to more advanced knowledge in the learning-based method, which motivates the development of similar multi-frame or 3D flow estimators for single-photon sensors.
[0120] Figure 10 shows an example of motion estimates generated using the techniques described herein for Euler-type single-photon computer vision, using binary frames of two scenes and varying coarseness of scaling in the filter, according to several embodiments of the disclosed subject matter.
[0121] Figure 10 shows multiple velocity cues from multiscale motion estimation and edge detection. From left to right, the SPAD sequence after pre-filtering is shown, followed by velocity estimation from 3D edge orientation based on normal temporal phase consistency, and then 2D velocity estimation at two scales. A burst reconstruction is inserted for reference. The coarser-scale filtered response yields less noise and therefore more reliable estimates compared to the finer scale. As a result, a denser velocity map is obtained, but localization is worse. The basic algorithm (TPC) also uses coarser-scale filtering, but edge velocities are naturally sparser. Quality varies spatially with local light intensity and image contrast, as expected from Poisson noise.
[0122] Fine-scale and normal velocity estimates provide better localization but are not always reliable due to noise. They can also suffer from aperture problems. Coarse-scale estimates are more robust (responses have higher z-scores) and less affected by aperture problems, but they have poor localization and can blur across object boundaries.
[0123] Edge detection and edge normal velocity detection based on Euler-type temporal compatibility, as well as the implementation of phase-based 2D motion, were also quantitatively evaluated on simulation data and compared with single-image denoising methods that are faster than burst reconstruction. These results are presented in Appendix A, incorporated by reference.
[0124] While single-photon sensors offer the prospect of recording visual detail with the resolution of individual photons, they also present challenges such as extremely noisy and quantized imaging models and prohibitively large computational and bandwidth requirements due to the enormous amount of data generated. In some embodiments, the mechanisms described herein can be used to implement relatively lightweight visual algorithms based on linear filtering and local phase-based processing of raw single-photon data, avoiding the expensive intermediate step of image reconstruction.
[0125] In some embodiments, the mechanisms described herein can be used to implement at least a portion of a computer vision pipeline on a single-photon detector-based image sensor. As new hardware architectures for single-photon sensors capable of performing complex computations at the photon level are developed, the mechanisms described herein can facilitate full on-chip real-time photon processing once a photon is captured. This is made possible by the computational simplicity of the mechanisms described herein. In some embodiments, the mechanisms described herein can be implemented with better memory efficiency by performing filtering entirely online (e.g., by exponential smoothing), eliminating the need for memory of past frames. Such on-chip Euler-type vision systems can facilitate the widespread deployment of single-photon imaging in real-world computer vision applications, such as in scientific fields like SLAM and biomechanics, and in consumer areas like sports videography.
[0126] As described above, a specific family of speed-tuned log-Gabor filters is described in relation to the mechanisms described herein. In some embodiments, better and more efficient filters can be obtained by formulating a suitable loss function for downstream visual tasks, including filters learned end-to-end from data. In addition or alternatively, in some embodiments, 3D gradient and monogenic filters can be used as well as phase filters, and the Canny edge detector (and its 3D counterpart), and the Lucas-Kanade optical flow estimator are classic algorithms that can be adapted for use with the mechanisms described herein. Such techniques are expected to operate faster than the phase-based techniques described herein because they require fewer filters. Apart from SNR and localization, which are standard optimization criteria in this setting, other relevant constraints such as causality and resource cost may influence which type of filter is suitable for a particular computer vision application.
[0127] Generally, velocity tuning filters operate under the assumption of local linear motion, which can be violated by the sudden appearance or disappearance of objects. In some embodiments, explicit occlusion inference, such as that used in more modern optical flow techniques, can be used in specific practical implementations to mitigate errors that may result from violations of the linear motion assumption.
[0128] As described herein, reconstructing high SNR inputs (e.g., high SNR reconstructed images) is not always necessary in visual tasks. A relevant concept is the prospect of results over any given time, which improves as the algorithm is run longer. In some embodiments, the scale of the filter bank can be considered to correspond to time, as information is aggregated over a wider volume, but the algorithm attempts to detect features at a finer scale. In some embodiments, the mechanisms described herein can be used in conjunction with diffusion-based algorithms, which can further reduce computational and bandwidth costs on specific architectures.
[0129] Further examples with various features Implementation examples are described in the following numbered clauses.
[0130] 1. A method for facilitating single-photon computer vision tasks, wherein this method is The method of causing an image sensor to generate a sequence of images representing a scene, wherein each image contains multiple pixels, and the image sensor includes multiple detectors configured to detect the arrival of individual photons, and the multiple detectors are arranged in an array. Performing a convolution between a three-dimensional filter and multiple frames for each of multiple three-dimensional filters, wherein each of the multiple frames is based on one or more images of a sequence of images. For each of the multiple frames, generate multiple filter bank responses, each corresponding to one of the three-dimensional filters among the multiple three-dimensional filters, Executing a computer vision process based on multiple filter bank responses, Includes.
[0131] 2. The method according to Clause 1, wherein each of the multiple detectors includes a single-photon avalanche diode (SPAD).
[0132] 3. Each image in the sequence of images has an exposure time. The method according to either one of clauses 1 or 2, comprising a binary image representing photons detected by an image sensor during TIFF2026516024000100.tif6150.
[0133] 4. Each of the multiple three-dimensional filters includes a speed tuning filter, and the first subset of the multiple three-dimensional filters is a three-dimensional frequency filter. It is tuned to TIFF2026516024000101.tif6150. Here, TIFF2026516024000102.tif6150 and TIFF2026516024000103.tif6150 represents a spatial pattern, TIFF2026516024000104.tif6150 represents a pattern in time, and each of the three-dimensional filters in the first subset has a different scale. The method described in any one of the clauses 1 to 3.
[0134] 5. Determining the z-score for each of the multiple filter bank responses, This involves mapping each z-score to the weight associated with each filter bank from which the z-score was determined, Using weighted filter bank responses, a computer vision process is executed based on multiple filter responses, The method described in any one of the clauses 1 to 4, further including the method described in any one of the clauses 1 to 4.
[0135] 6. Relationship Using TIFF2026516024000105.tif12150, the variance value This involves estimating TIFF2026516024000106.tif6150, TIFF2026516024000107.tif6150 contains multiple frames Filter TIFF2026516024000108.tif6150 This is a filter bank response generated by applying TIFF2026516024000109.tif6150. TIFF2026516024000110.tif6150 is, This is the variance of the estimated local mean flux over TIFF2026516024000111.tif6150. and TIFF2026516024000112.tif12150 is, Filter across TIFF2026516024000113.tif6150 The total of TIFF2026516024000114.tif6150 is estimated, and Relationship The Z-score will be determined using TIFF2026516024000115.tif16150, The method described in Clause 5, further including the following:
[0136] 7. Relationship This further involves mapping the z-scores to the weights associated with each filter bank, using TIFF2026516024000116.tif8150, where TIFF2026516024000117.tif11150 is a method of implementation as described in Clause 6, including a threshold z-score.
[0137] 8. The method according to any one of clauses 1 to 7, wherein the computer vision process is an edge detection process based on phase consistency associated with each of a plurality of filter responses.
[0138] 9. The method according to clause 8, further comprising detecting one or more corners based on the phase matching associated with each of the multiple filter responses.
[0139] 10. The computer vision process is a motion estimation process based on phase invariance, as described in Clause 1.
[0140] 11. The computer vision process is a motion estimation process, and this method further involves phase invariance. This involves performing a computer vision process for each of multiple pixels based on TIFF2026516024000118.tif11150, where TIFF2026516024000119.tif6150 shows the speed at each pixel. TIFF2026516024000120.tif10150 contains each of the three-dimensional filters. Spatial direction of TIFF2026516024000121.tif11150 The component velocities in TIFF2026516024000122.tif11150, and The filename is TIFF2026516024000123.tif11150, and here TIFF2026516024000124.tif11150 is, The method described in Clause 1, which is the spatiotemporal direction of TIFF2026516024000125.tif11150.
[0141] 12. A non-temporary, computer-readable medium containing computer-executable code, which includes code that causes a computer to cause a processor to perform any of the methods described in clauses 1 to 11.
[0142] 13. A system for simulating interaction with an infant, comprising at least one processor configured to perform any of the methods described in clauses 1 to 11.
[0143] 14. The system according to Clause 13, further comprising an image sensor including a plurality of detectors configured to detect the arrival of individual photons, wherein the plurality of detectors are arranged in an array.
[0144] In some embodiments, any suitable computer-readable medium can be used to store instructions for performing the functions and / or processes described herein. For example, in some embodiments, the computer-readable medium can be temporary or non-temporary. For example, non-temporary computer-readable mediums can include magnetic media (hard disks, floppy disks, etc.), optical media (compact disks, digital video disks, Blu-ray disks, etc.), semiconductor media (RAM, flash memory, electrically programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.), any suitable medium that is not temporary or lacks any aspect of persistence during transmission, and / or any suitable tangible medium. In another embodiment, a temporary computer-readable medium can include a network, wires, conductors, optical fibers, computing circuits, or any suitable medium that is not temporary or lacks any aspect of persistence during transmission, and / or signals on any suitable intangible medium.
[0145] It should be noted that, as used herein, the term "mechanism" may encompass hardware, software, firmware, or any appropriate combination thereof.
[0146] It should be understood that the steps described above for the process in Figure 5 can be performed or implemented in any appropriate order or sequence, not limited to the order and sequence illustrated and described. Furthermore, some of the steps described above for the process in Figure 5 can be performed or implemented substantially simultaneously or in parallel, where appropriate, to reduce waiting and processing times.
[0147] Although the present invention has been described and illustrated in the above-mentioned exemplary embodiments, this disclosure is for illustrative purposes only, and numerous modifications to the details of the embodiment of the present invention can be made without departing from the spirit and scope of the invention, and it should be understood that the spirit and scope of the invention are limited by the appended claims. The features of the disclosed embodiments can be combined and reconfigured in various ways. [Explanation of symbols]
[0148] 200 Systems 204 Lens 206 optical elements 208 processors 210 inputs 212 memory 214 Signal Generator 216 Communication Systems 218 displays
Claims
1. A system for facilitating single-photon computer vision tasks, An image sensor comprising a plurality of detectors configured to detect the arrival of individual photons, wherein the plurality of detectors are arranged in an array, At least one processor, Equipped with, The aforementioned at least one processor, The image sensor generates a sequence of images representing a scene, each of which contains multiple pixels; For each of the multiple three-dimensional filters, a convolution is performed between the three-dimensional filter and the multiple frames, wherein each of the multiple frames is based on one or more of the images in the sequence of images; For each of the plurality of frames, generate a plurality of filter bank responses, each corresponding to one of the plurality of three-dimensional filters; and Based on the aforementioned multiple filter bank responses, a computer vision process is executed. A system that is programmed to do so.
2. The system according to claim 1, wherein each of the plurality of detectors includes a single-photon avalanche diode (SPAD).
3. Each image in the sequence of the aforementioned images is, The system according to claim 1, further comprising a binary image representing a photon detected by the image sensor during the interval.
4. Each of the plurality of three-dimensional filters includes a speed tuning filter, and a first subset of the plurality of three-dimensional filters is a three-dimensional frequency It is tuned to this, Here, and This represents a spatial pattern, represents a pattern in time, and each of the three-dimensional filters in the first subset has a different scale. The system according to claim 1.
5. The aforementioned at least one processor further: The z-score is determined for each of the above-mentioned filter bank responses; Each z-score is mapped to a weight associated with each filter bank for which the z-score was determined, thereby generating a plurality of weighted filter bank responses; Using the aforementioned weighted filter bank responses, a computer vision process is executed based on the aforementioned weighted filter bank responses. The system according to claim 1, which is programmed to do so.
6. The aforementioned at least one processor further: Relationship Using the variance value We estimate that, and here, is multiple frames Filter This is the filter bank response generated by applying the filter, teeth, This is the variance of the estimated local mean flux over the period, and teeth, Filters across It is the sum; and Relationship The z-score is determined using the above method. The system according to claim 5, further programmed to do so.
7. The aforementioned at least one processor further: Each z-score is related to the following equation Map to the weights associated with each of the filter banks determined using It is programmed to do so, Here This includes the threshold z-score. The system according to claim 6.
8. The aforementioned computer vision process is an edge detection process, The aforementioned at least one processor further: The computer vision process is executed based on the phase consistency associated with each of the plurality of filter bank responses. The system according to claim 1, which is programmed to do so.
9. The aforementioned at least one processor further: Based on the phase matching associated with each of the plurality of filter bank responses, one or more corners are detected. The system according to claim 8, which is programmed to do so.
10. The aforementioned computer vision process is a motion estimation process, The aforementioned at least one processor further: Phase invariance relation Based on this, the computer vision process is executed for each of the plurality of pixels. It is programmed to do so, here The velocity at each pixel, These are the respective three-dimensional filters spatial direction The component rates in, and And here teeth, The spatiotemporal direction of The system according to claim 1.
11. A method for facilitating single-photon computer vision tasks, The method of causing an image sensor to generate a sequence of images representing a scene, wherein each of the images includes a plurality of pixels, the image sensor includes a plurality of detectors configured to detect the arrival of individual photons, and the plurality of detectors are arranged in an array. Performing a convolution between each of a plurality of three-dimensional filters and a plurality of frames, wherein each of the plurality of frames is based on one or more of the images in the sequence of images, For each of the aforementioned multiple frames, a plurality of filter bank responses are generated, each corresponding to one of the three-dimensional filters among the plurality of three-dimensional filters. The computer vision process is executed based on the aforementioned multiple filter bank responses. Methods that include...
12. The method according to claim 11, wherein each of the plurality of detectors includes a single-photon avalanche diode (SPAD).
13. Each image in the sequence of the aforementioned images is, The method according to claim 11, further comprising a binary image representing a photon detected by the image sensor during the interval.
14. Each of the plurality of three-dimensional filters includes a speed tuning filter, and a first subset of the plurality of three-dimensional filters is a three-dimensional frequency It is tuned to this, Here, and This represents a spatial pattern, represents a pattern in time, and each of the three-dimensional filters in the first subset has a different scale. The method according to claim 11.
15. The z-score is determined for each of the aforementioned plurality of filter bank responses, Each z-score is mapped to a weight associated with each filter bank for which the z-score was determined, thereby generating multiple weighted filter bank responses. Using the aforementioned multiple weighted filter bank responses, a computer vision process is executed based on the aforementioned multiple filter bank responses. The method according to claim 11, further comprising:
16. Relationship Using the variance value This is to estimate, is multiple frames Filter This is the filter bank response generated by applying the filter, teeth, This is the variance of the estimated local mean flux over the period, and teeth, Filters across The sum of, estimating, Relationship The z-score is determined using the above method, The method according to claim 15, further comprising:
17. The computer vision process is an edge detection process based on phase consistency associated with each of the plurality of filter bank responses. The method according to claim 11.
18. The method according to claim 11, wherein the computer vision process is a motion estimation process based on a phase invariance relationship.
19. A non-temporary computer-readable medium containing computer-executable instructions that, when executed by a processor, cause the processor to perform a method for facilitating a single-photon computer vision task, The aforementioned method, The method of causing an image sensor to generate a sequence of images representing a scene, wherein each of the images includes a plurality of pixels, the image sensor includes a plurality of detectors configured to detect the arrival of individual photons, and the plurality of detectors are arranged in an array. Performing a convolution between each of a plurality of three-dimensional filters and a plurality of frames, wherein each of the plurality of frames is based on one or more of the images in the sequence of images, For each of the aforementioned multiple frames, a plurality of filter bank responses are generated, each corresponding to one of the three-dimensional filters among the plurality of three-dimensional filters. The computer vision process is executed based on the aforementioned multiple filter bank responses. Non-temporary computer-readable media, including [specific examples of such media].
20. Each of the aforementioned plurality of detectors includes a single-photon avalanche diode (SPAD), The non-temporary computer-readable medium according to claim 19.
21. Each image in the sequence of the aforementioned images is, The interval includes a binary image representing the photons detected by the image sensor, The non-temporary computer-readable medium according to claim 19.